matilda-jev-fp4 / README.md
yue-maincode's picture
Upload validated MATILDA JEV FP4 model and Decision Index scores
c69aaec verified
|
Raw History Blame Contribute Delete
3.62 kB
---
license: apache-2.0
library_name: transformers
tags:
- matilda
- jev
- fp4
- quantized
- maincode
---
# MATILDA JEV FP4 by Maincode
The validated FP4 version of MATILDA JEV, with Decision Index **61.77**.
Its corresponding BF16 source scores 62.43. This package contains packed FP4
weights and the native JEV decision head and runtime.
## Format and execution
- E2M1 FP4 weights, with FP8 E4M3 block scales per 16 values and FP32 tensor scales.
- BF16 activations and matrix multiplication (**W4A16**). A Triton kernel
dequantizes one matrix at a time; this runtime does not use native FP4 Tensor Core GEMM.
- 400 large text linear modules quantized. The decision readout, embeddings,
normalization layers, vision tower and small gate projections retain their original precision.
- Native `choice`, `noul` and `score` outputs, with up to 255 options per question.
Weight tensors occupy **17.20 GB**, compared with 52.17 GB of source weight files.
On AMD Instinct MI355X, measured resident memory is **16.3 GiB** versus 48.6 GiB
for the BF16 source. Short-to-medium single-request P50 latency is **66 ms**
versus 54 ms; this implementation primarily saves memory. These measurements
are workload-specific. NVIDIA hardware and vision inference were not tested.
## Loading
Use Python 3.12 or newer. The validated environment uses PyTorch 2.14.0+ROCm7.2,
Transformers 5.17.0, Accelerate, Safetensors, Triton and Flash Linear Attention.
Install a PyTorch build appropriate for your accelerator before the remaining
dependencies in `requirements-runtime.txt`.
Download the repository with `huggingface_hub.snapshot_download` and load it
using the included `FP4DecisionModel` adapter:
```python
import sys
from pathlib import Path
checkpoint = Path('/path/to/downloaded/model')
sys.path.insert(0, str(checkpoint))
from jev_fp4 import FP4DecisionModel
from kev.model import answer
model = FP4DecisionModel(checkpoint, device='cuda:0')
row = {
'state': 'The parcel arrived on schedule.',
'question': {
'type': 'choice',
'instructions': 'Classify delivery.',
'criteria': {'on_time': 'On time', 'late': 'Late'},
},
}
print(answer(row['question'], model.predict([row])[0]))
```
The supplied `predict.py` accepts JSON lines with `state` + `question` or
`state` + `questions` and produces native JEV answers:
```bash
python /path/to/downloaded/model/predict.py --device cuda:0 < requests.jsonl
```
Packed weights require the included FP4 adapter; they cannot be loaded using
the original BF16 loader. This is a decision model with a separate readout,
not a text generation model.
## Full Decision Index
Edition 0.2.1, all **150,317 requests across 44 benchmarks** completed successfully.
All shard outputs were audited, and scoring was recomputed with identical results.
`scores.json` contains the complete results; `comparison.json` compares all
benchmarks against the exact BF16 source.
| Metric | BF16 source | FP4 |
|---|---:|---:|
| Decision Index | 62.43 | **61.77** |
| Raw | 71.42 | 70.90 |
| Breadth | 61.15 | 60.48 |
| Knowledge & Reasoning | 47.37 | 46.45 |
| Language Understanding | 70.38 | 69.24 |
| Retrieval & Classification | 64.98 | 64.41 |
| Tools & Automation | 80.89 | 80.76 |
| Arts & Human Taste | 41.96 | 42.05 |
Area scores are chance-corrected skill multiplied by 100. GPQA Diamond changes
from 51.02 to 48.47, MMLU-Pro from 83.60 to 81.70, and GSM8K from 79.45 to 79.53.
The source model's training history includes benchmark-related material, so
these are diagnostic quantization comparisons, not an independent held-out
generalization claim.