Qyvos / README.md
Manusagents's picture
Add model card metadata
b763cec verified
|
Raw History Blame Contribute Delete
7.23 kB
---
library_name: transformers
tags:
- decision-model
- classification
- julia
- open-jev
- head-finetune
- low-resource
base_model: SupersonicLabs/Julia-1
datasets:
- ZefanCai/Open-Jev
pipeline_tag: text-classification
---
# Qyvos
**Qyvos** is a decision / classification model built from [`SupersonicLabs/Julia-1`](https://huggingface.co/SupersonicLabs/Julia-1) by head-only fine-tuning on the [`ZefanCai/Open-Jev`](https://huggingface.co/datasets/ZefanCai/Open-Jev) dataset (config `release-v2-redistributable`).
It maps a **state + question + options** context to a probability distribution over the options, for three decision kinds:
| qtype | kind | meaning |
|-------|-------|---------|
| 0 | `choice` | pick one option (categorical decision) |
| 1 | `score` | grade/distribute mass over options (soft scoring) |
| 2 | `noul` | binary yes/no decision (options ordered `[no, yes]`) |
## Complete backbone retention
The **entire Julia-1 backbone is retained bit-exact**:
- `encoder` β€” ModernBERT-small (22 layers, hidden 384, vocab 256k), 140,530,432 params β€” **frozen, untouched**
- `act_head` β€” auxiliary action head β€” **frozen, untouched**
Training updated only the decision head (**3,699,073 trainable params = 2.56%** of 144,292,867):
- `head` β€” 2Γ— `TransformerEncoderLayer(384, nhead=6, dim_ff=1536, norm_first=True)`
- `type_emb` β€” `Embedding(3, 384)` (one embedding per decision kind)
- `scorer` β€” `LayerNorm β†’ Linear(384,384) β†’ GELU β†’ Linear(384,1)` per-option scorer
Weights are **float32** end-to-end (the Julia checkpoint format has no FP16 variant; nothing was quantized or pruned).
## Training
- Data: `ZefanCai/Open-Jev`, config `release-v2-redistributable`, `train` split (79,116 rows available), streamed **30,000 rows** in deterministic shuffled order
- Objective: soft-target cross-entropy `-(target Β· log_softmax(scores)).sum(-1).mean()` β€” trains calibrated probability distributions, exactly matching Open-Jev's soft `target` vectors
- Encoding: identical `julia.data.sequence()` marker-token serialization as the official Julia runtime (train/inference parity)
- Optimizer: AdamW, lr 1e-4 (linear warmup 100 steps β†’ linear decay to 10%), Ξ²=(0.9,0.999), weight decay 0.01 (no decay on norms/embeddings), grad clip 1.0
- Batch: **1 micro-batch Γ— 8 accumulation** (effective batch 8), 3,750 optimizer steps
- Sequence budget: `max_length=1024`, `head_length=512`
- Seed: 17
### Low-RAM training constraints (3.9 GB system RAM)
No full fine-tuning was possible (or needed). The run was designed for an 8 GB-class machine:
- frozen encoder β‡’ no gradients/optimizer state for 140.5M params; encoder forward runs under `no_grad`
- bs=1 micro-batches, `num_workers=0`, fp32 head-only optimizer state
- parquet read via **pyarrow mmap**, one row-group (~1,000 rows, a few MB) at a time; lazy per-row tokenization; **no dataset-wide RAM cache**
- head-only checkpoints (~45 MB, incl. optimizer) with deterministic skip fast-forward resume
- psutil RAM guard: auto-checkpoint + abort if system available RAM < 350 MB
- **Measured peak RSS: 1.64 GB**
## Results (honest shuffled-mixture evaluation)
Evaluation samples are drawn round-robin across **all** parquet row-groups (each region contributes equally), fixed seed 999, length-bucketed batches. Accuracy = argmax match vs the target's argmax; softCE = soft-target cross-entropy (lower is better).
| model | val acc (n=900) | val softCE | test acc (n=900) | test softCE |
|---|---|---|---|---|
| Julia-1 base (no training) | 83.67% | 0.4542 | 83.56% | 0.4764 |
| **Qyvos (30,000 rows)** | **84.11%** | **0.4282** | 83.11% | **0.4447** |
Larger-sample confirmation for Qyvos on the test split: n=1500 β†’ acc 81.67%, softCE 0.4587.
Per-kind (Qyvos, n=900):
| kind | val acc | test acc |
|---|---|---|
| choice | 66.5% | 68.2% |
| score | 79.4% | 79.3% |
| noul | 92.7% | 90.2% |
Honest reading: Julia-1's pretrained decision head is already strong on this benchmark family, and top-1 accuracy differences between base and Qyvos are within evaluation noise (~Β±1.2pt stderr at n=900). What fine-tuning **did** deliver, consistently across both splits and sample sizes, is **better-calibrated probabilities (softCE βˆ’5.7% val, βˆ’6.7% test)** β€” which is precisely the objective that matters for Open-Jev's soft, often high-entropy targets, and for any downstream consumer of the full probability vector rather than a single argmax. The accuracy ceiling here is largely data-intrinsic: many Open-Jev targets are intentionally soft/ambiguous.
## Exact inference (official runtime)
```python
import julia # from the julia/ package shipped in this repo (pip install -e julia --no-deps)
engine = julia.load_model("<path-to-Qyvos>", device="cpu", max_length=1024, head_length=512)
pred = engine.predict([{
"state": "User chat: my login is blocked after the password reset.",
"question": "Determine the broad category of this support ticket.",
"options": ["account: Login, permissions, profile, security",
"billing: Charges, invoices, refunds, subscriptions"],
"type": "choice", # choice | score | noul
}])[0]
print(pred["index"], pred["probabilities"])
```
Or with the bundled CLI (see `training/`):
```bash
python3 scripts/infer_qyvos.py --demo # per-kind examples from the test split
python3 scripts/infer_qyvos.py --eval-test 900 # honest shuffled test accuracy
python3 scripts/infer_qyvos.py --row '{"state":"...","question":"...","options":["a","b"],"type":"choice"}'
```
> **Compatibility note:** on `transformers >= 5.17`, julia's *optional* fast inference path
> (`julia/router/encoder.py`) references `ModernBertModel._update_attention_mask`, which was
> removed upstream. The bundled scripts disable that optimization via a two-line shim
> (`specialize_decision_encoder β†’ False`); inference then uses the standard forward with
> numerically equivalent results. On `transformers 5.0.x` (the version this checkpoint was
> built against) no shim is needed.
## Repository layout
```
Qyvos/
β”œβ”€β”€ config.json # architecture pointer file (Julia format v1)
β”œβ”€β”€ julia_config.json # Qyvos identity + head settings
β”œβ”€β”€ model.safetensors # fp32 weights (backbone bit-exact + trained head)
β”œβ”€β”€ encoder/config.json # ModernBERT-small config
β”œβ”€β”€ tokenizer/ # tokenizer (copied bit-exact from Julia-1)
β”œβ”€β”€ label_map.json # qtype id -> kind
β”œβ”€β”€ qyvos_training_config.json # full training config + final metrics
β”œβ”€β”€ provenance.json # base/dataset revisions + SHA-256 hashes
└── training/ # all scripts used (download/inspect/train/build/infer)
```
## Provenance
- Base model: `SupersonicLabs/Julia-1` (fp32, weights SHA-256 recorded in `provenance.json`)
- Dataset: `ZefanCai/Open-Jev` @ config `release-v2-redistributable` (per-shard SHA-256 in `provenance.json`)
- Backbone tensors: bit-exact vs base; only `head`, `type_emb`, `scorer` differ
- Final `model.safetensors` SHA-256: `1d03ac9a0c7fbc66f1bde7ecb158741b02c9f898078c0a2f0cc270fb6326f920`