File size: 3,301 Bytes
bb09ac9
 
a61cbe1
bb09ac9
 
a61cbe1
 
 
 
 
 
 
bb09ac9
 
a61cbe1
bb09ac9
a61cbe1
bb09ac9
a61cbe1
bb09ac9
a61cbe1
bb09ac9
a61cbe1
 
 
 
 
 
 
 
 
bb09ac9
a61cbe1
bb09ac9
a61cbe1
 
bb09ac9
 
a61cbe1
 
bb09ac9
 
a61cbe1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
---
library_name: mlx
base_model: deepreinforce-ai/Ornith-1.0-35B
tags:
- mlx
- mlx-vlm
- moe
- multimodal
- vision
- coding
- agentic
pipeline_tag: image-text-to-text
---

# leonsarmiento/Ornith-1.0-35B-4bit-mlx

This model was converted to MLX format from [`deepreinforce-ai/Ornith-1.0-35B`](https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B) using **mixed 4/8-bit quantization** optimized for Apple Silicon. The vision encoder is preserved and quantized at 4-bit, making this a full multimodal model.

Ornith-1.0-35B is a 35B-parameter MoE (Mixture of Experts) model fine-tuned from Qwen3.5-35B-A3B by DeepReinforce AI, using a self-improving RL training framework that jointly optimizes scaffold and solution rollouts for agentic coding tasks. Despite 35B total parameters, only ~3B are activated per token. It features 256 experts (8 active per token + 1 shared expert), hybrid full + linear (Gated DeltaNet) attention, and a vision encoder.

## Benchmark Highlights

| Benchmark | Ornith-1.0-35B | Qwen3.5-35B | Qwen3.6-35B |
|-----------|---------------|-------------|-------------|
| Terminal-Bench 2.1 (Terminus-2) | **64.2** | 41.4 | 52.5 |
| Terminal-Bench 2.1 (Claude Code) | **62.8** | 38.9 | 49.2 |
| SWE-bench Verified | **75.6** | 70 | 73.4 |
| SWE-bench Pro | **50.4** | 44.6 | 49.5 |
| SWE-bench Multilingual | **69.3** | 60.3 | 67.2 |
| NL2Repo | **34.6** | 20.5 | 29.4 |
| Claw-eval Avg | **69.8** | 65.4 | 68.7 |

## Use with mlx

```bash
pip install -U mlx-vlm
```

```bash
python -m mlx_vlm.generate --model leonsarmiento/Ornith-1.0-35B-4bit-mlx --max-tokens 256 --temperature 1.0 --top-p 1.0 --prompt "Hello"
```

## Mixed Quantization Strategy

This model uses **layer-aware mixed-bit quantization** that allocates higher precision to sensitive layers and lower precision to bulk parameters, maximizing quality per gigabyte.

| Bit Depth | Layers | Rationale |
|-----------|--------|-----------|
| **8-bit** | `embed_tokens`, `lm_head`, router `gate`, `shared_expert_gate`, `shared_expert`, `self_attn` (full attention), `linear_attn` (DeltaNet) | Every token passes through these — routing accuracy, shared representation, and sequence modeling are non-negotiable |
| **4-bit** | `vision_tower`, `switch_mlp` (routed experts) | Bulk of parameters, only 8 of 256 experts active per token — natural redundancy tolerates lower precision |

### Quantization Details

| Layer | Bits | Group Size |
|-------|------|------------|
| `embed_tokens` | 8 | 64 |
| `lm_head` | 8 | 64 |
| `mlp.gate` (router) | 8 | 64 |
| `shared_expert_gate` | 8 | 64 |
| `shared_expert` | 8 | 64 |
| `self_attn` (full attention) | 8 | 64 |
| `linear_attn` (DeltaNet) | 8 | 64 |
| `vision_tower` | 4 | 64 |
| `switch_mlp` (routed experts) | 4 | 64 |
| Default fallback | 8 | 64 |

- **Quantization type**: Mixed 4/8-bit (multimodal, vision preserved)
- **Group size**: 64
- **Method**: Custom `quant_predicate` via `mlx_vlm`

## Recommended Inference Parameters

| Parameter | Value |
|-----------|-------|
| `temperature` | 1.0 |
| `top_p` | 1.0 |
| `top_k` | 40 |
| `min_p` | 0.01 |
| `repeat_penalty` | 1.05 |

> **Note:** Ornith-1.0-35B uses Temp 1.0 and Top_p 1.0 per the model's Terminal-Bench 2.1 benchmark recipe. This is a Qwen3.5-based model — `preserve_thinking` is not applicable.