File size: 7,227 Bytes
b763cec
 
 
 
 
 
 
 
 
 
 
 
 
 
 
31f7037
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
---
library_name: transformers
tags:
- decision-model
- classification
- julia
- open-jev
- head-finetune
- low-resource
base_model: SupersonicLabs/Julia-1
datasets:
- ZefanCai/Open-Jev
pipeline_tag: text-classification
---

# Qyvos

**Qyvos** is a decision / classification model built from [`SupersonicLabs/Julia-1`](https://huggingface.co/SupersonicLabs/Julia-1) by head-only fine-tuning on the [`ZefanCai/Open-Jev`](https://huggingface.co/datasets/ZefanCai/Open-Jev) dataset (config `release-v2-redistributable`).

It maps a **state + question + options** context to a probability distribution over the options, for three decision kinds:

| qtype | kind  | meaning |
|-------|-------|---------|
| 0     | `choice` | pick one option (categorical decision) |
| 1     | `score`  | grade/distribute mass over options (soft scoring) |
| 2     | `noul`   | binary yes/no decision (options ordered `[no, yes]`) |

## Complete backbone retention

The **entire Julia-1 backbone is retained bit-exact**:

- `encoder` β€” ModernBERT-small (22 layers, hidden 384, vocab 256k), 140,530,432 params β€” **frozen, untouched**
- `act_head` β€” auxiliary action head β€” **frozen, untouched**

Training updated only the decision head (**3,699,073 trainable params = 2.56%** of 144,292,867):

- `head` β€” 2Γ— `TransformerEncoderLayer(384, nhead=6, dim_ff=1536, norm_first=True)`
- `type_emb` β€” `Embedding(3, 384)` (one embedding per decision kind)
- `scorer` β€” `LayerNorm β†’ Linear(384,384) β†’ GELU β†’ Linear(384,1)` per-option scorer

Weights are **float32** end-to-end (the Julia checkpoint format has no FP16 variant; nothing was quantized or pruned).

## Training

- Data: `ZefanCai/Open-Jev`, config `release-v2-redistributable`, `train` split (79,116 rows available), streamed **30,000 rows** in deterministic shuffled order
- Objective: soft-target cross-entropy `-(target Β· log_softmax(scores)).sum(-1).mean()` β€” trains calibrated probability distributions, exactly matching Open-Jev's soft `target` vectors
- Encoding: identical `julia.data.sequence()` marker-token serialization as the official Julia runtime (train/inference parity)
- Optimizer: AdamW, lr 1e-4 (linear warmup 100 steps β†’ linear decay to 10%), Ξ²=(0.9,0.999), weight decay 0.01 (no decay on norms/embeddings), grad clip 1.0
- Batch: **1 micro-batch Γ— 8 accumulation** (effective batch 8), 3,750 optimizer steps
- Sequence budget: `max_length=1024`, `head_length=512`
- Seed: 17

### Low-RAM training constraints (3.9 GB system RAM)

No full fine-tuning was possible (or needed). The run was designed for an 8 GB-class machine:

- frozen encoder β‡’ no gradients/optimizer state for 140.5M params; encoder forward runs under `no_grad`
- bs=1 micro-batches, `num_workers=0`, fp32 head-only optimizer state
- parquet read via **pyarrow mmap**, one row-group (~1,000 rows, a few MB) at a time; lazy per-row tokenization; **no dataset-wide RAM cache**
- head-only checkpoints (~45 MB, incl. optimizer) with deterministic skip fast-forward resume
- psutil RAM guard: auto-checkpoint + abort if system available RAM < 350 MB
- **Measured peak RSS: 1.64 GB**

## Results (honest shuffled-mixture evaluation)

Evaluation samples are drawn round-robin across **all** parquet row-groups (each region contributes equally), fixed seed 999, length-bucketed batches. Accuracy = argmax match vs the target's argmax; softCE = soft-target cross-entropy (lower is better).

| model | val acc (n=900) | val softCE | test acc (n=900) | test softCE |
|---|---|---|---|---|
| Julia-1 base (no training) | 83.67% | 0.4542 | 83.56% | 0.4764 |
| **Qyvos (30,000 rows)** | **84.11%** | **0.4282** | 83.11% | **0.4447** |

Larger-sample confirmation for Qyvos on the test split: n=1500 β†’ acc 81.67%, softCE 0.4587.

Per-kind (Qyvos, n=900):

| kind | val acc | test acc |
|---|---|---|
| choice | 66.5% | 68.2% |
| score | 79.4% | 79.3% |
| noul | 92.7% | 90.2% |

Honest reading: Julia-1's pretrained decision head is already strong on this benchmark family, and top-1 accuracy differences between base and Qyvos are within evaluation noise (~Β±1.2pt stderr at n=900). What fine-tuning **did** deliver, consistently across both splits and sample sizes, is **better-calibrated probabilities (softCE βˆ’5.7% val, βˆ’6.7% test)** β€” which is precisely the objective that matters for Open-Jev's soft, often high-entropy targets, and for any downstream consumer of the full probability vector rather than a single argmax. The accuracy ceiling here is largely data-intrinsic: many Open-Jev targets are intentionally soft/ambiguous.

## Exact inference (official runtime)

```python
import julia  # from the julia/ package shipped in this repo (pip install -e julia --no-deps)

engine = julia.load_model("<path-to-Qyvos>", device="cpu", max_length=1024, head_length=512)
pred = engine.predict([{
    "state":   "User chat: my login is blocked after the password reset.",
    "question": "Determine the broad category of this support ticket.",
    "options": ["account: Login, permissions, profile, security",
                "billing: Charges, invoices, refunds, subscriptions"],
    "type": "choice",           # choice | score | noul
}])[0]
print(pred["index"], pred["probabilities"])
```

Or with the bundled CLI (see `training/`):

```bash
python3 scripts/infer_qyvos.py --demo                     # per-kind examples from the test split
python3 scripts/infer_qyvos.py --eval-test 900            # honest shuffled test accuracy
python3 scripts/infer_qyvos.py --row '{"state":"...","question":"...","options":["a","b"],"type":"choice"}'
```

> **Compatibility note:** on `transformers >= 5.17`, julia's *optional* fast inference path
> (`julia/router/encoder.py`) references `ModernBertModel._update_attention_mask`, which was
> removed upstream. The bundled scripts disable that optimization via a two-line shim
> (`specialize_decision_encoder β†’ False`); inference then uses the standard forward with
> numerically equivalent results. On `transformers 5.0.x` (the version this checkpoint was
> built against) no shim is needed.

## Repository layout

```
Qyvos/
β”œβ”€β”€ config.json                  # architecture pointer file (Julia format v1)
β”œβ”€β”€ julia_config.json            # Qyvos identity + head settings
β”œβ”€β”€ model.safetensors            # fp32 weights (backbone bit-exact + trained head)
β”œβ”€β”€ encoder/config.json          # ModernBERT-small config
β”œβ”€β”€ tokenizer/                   # tokenizer (copied bit-exact from Julia-1)
β”œβ”€β”€ label_map.json               # qtype id -> kind
β”œβ”€β”€ qyvos_training_config.json   # full training config + final metrics
β”œβ”€β”€ provenance.json              # base/dataset revisions + SHA-256 hashes
└── training/                    # all scripts used (download/inspect/train/build/infer)
```

## Provenance

- Base model: `SupersonicLabs/Julia-1` (fp32, weights SHA-256 recorded in `provenance.json`)
- Dataset: `ZefanCai/Open-Jev` @ config `release-v2-redistributable` (per-shard SHA-256 in `provenance.json`)
- Backbone tensors: bit-exact vs base; only `head`, `type_emb`, `scorer` differ
- Final `model.safetensors` SHA-256: `1d03ac9a0c7fbc66f1bde7ecb158741b02c9f898078c0a2f0cc270fb6326f920`