File size: 3,619 Bytes
c69aaec
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
---
license: apache-2.0
library_name: transformers
tags:
- matilda
- jev
- fp4
- quantized
- maincode
---

# MATILDA JEV FP4 by Maincode

The validated FP4 version of MATILDA JEV, with Decision Index **61.77**.
Its corresponding BF16 source scores 62.43. This package contains packed FP4
weights and the native JEV decision head and runtime.

## Format and execution

- E2M1 FP4 weights, with FP8 E4M3 block scales per 16 values and FP32 tensor scales.
- BF16 activations and matrix multiplication (**W4A16**). A Triton kernel
  dequantizes one matrix at a time; this runtime does not use native FP4 Tensor Core GEMM.
- 400 large text linear modules quantized. The decision readout, embeddings,
  normalization layers, vision tower and small gate projections retain their original precision.
- Native `choice`, `noul` and `score` outputs, with up to 255 options per question.

Weight tensors occupy **17.20 GB**, compared with 52.17 GB of source weight files.
On AMD Instinct MI355X, measured resident memory is **16.3 GiB** versus 48.6 GiB
for the BF16 source. Short-to-medium single-request P50 latency is **66 ms**
versus 54 ms; this implementation primarily saves memory. These measurements
are workload-specific. NVIDIA hardware and vision inference were not tested.

## Loading

Use Python 3.12 or newer. The validated environment uses PyTorch 2.14.0+ROCm7.2,
Transformers 5.17.0, Accelerate, Safetensors, Triton and Flash Linear Attention.
Install a PyTorch build appropriate for your accelerator before the remaining
dependencies in `requirements-runtime.txt`.

Download the repository with `huggingface_hub.snapshot_download` and load it
using the included `FP4DecisionModel` adapter:

```python
import sys
from pathlib import Path

checkpoint = Path('/path/to/downloaded/model')
sys.path.insert(0, str(checkpoint))
from jev_fp4 import FP4DecisionModel
from kev.model import answer

model = FP4DecisionModel(checkpoint, device='cuda:0')
row = {
    'state': 'The parcel arrived on schedule.',
    'question': {
        'type': 'choice',
        'instructions': 'Classify delivery.',
        'criteria': {'on_time': 'On time', 'late': 'Late'},
    },
}
print(answer(row['question'], model.predict([row])[0]))
```

The supplied `predict.py` accepts JSON lines with `state` + `question` or
`state` + `questions` and produces native JEV answers:

```bash
python /path/to/downloaded/model/predict.py --device cuda:0 < requests.jsonl
```

Packed weights require the included FP4 adapter; they cannot be loaded using
the original BF16 loader. This is a decision model with a separate readout,
not a text generation model.

## Full Decision Index

Edition 0.2.1, all **150,317 requests across 44 benchmarks** completed successfully.
All shard outputs were audited, and scoring was recomputed with identical results.
`scores.json` contains the complete results; `comparison.json` compares all
benchmarks against the exact BF16 source.

| Metric | BF16 source | FP4 |
|---|---:|---:|
| Decision Index | 62.43 | **61.77** |
| Raw | 71.42 | 70.90 |
| Breadth | 61.15 | 60.48 |
| Knowledge & Reasoning | 47.37 | 46.45 |
| Language Understanding | 70.38 | 69.24 |
| Retrieval & Classification | 64.98 | 64.41 |
| Tools & Automation | 80.89 | 80.76 |
| Arts & Human Taste | 41.96 | 42.05 |

Area scores are chance-corrected skill multiplied by 100. GPQA Diamond changes
from 51.02 to 48.47, MMLU-Pro from 83.60 to 81.70, and GSM8K from 79.45 to 79.53.
The source model's training history includes benchmark-related material, so
these are diagnostic quantization comparisons, not an independent held-out
generalization claim.