Ornith 1.0 35B 8bit 8 bit (group size 64, 8.596 bits/weight ) MLX quantization of deepreinforce ai/Ornith 1.0 35B, produced with mlx vlm 0.6.3. Full multimodal: the vision encoder is preserved and quantized alongside the language model. For Apple Silicon. Runs in mlx vlm or any MLX app. Conversion note (MoE expert fusion) Ornith stores its 256 MoE experts unfused (per expert), but mlx vlm's qwen3 5 moe loader expects them fused/batched. A sanitize monkeypatch was required to stack the experts before conversion; without it the conversion failed. This is a standard mlx vlm 8 bit quant. Usage Conversion check Smoke tested after conversion ( mlx vlm.generate on an image): coherent — correctly read an evaluation bar chart, no repetition loop. 89.2 tok/s generation, 896.9 tok/s prompt, peak 39.8 GB on a Macbook Pro M5 Max 128GB 40 GPU. Refer to the original model card for architecture, benchmarks, license, and intended use.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy