--- license: apache-2.0 library_name: transformers tags: - matilda - jev - fp4 - quantized - maincode --- # MATILDA JEV FP4 by Maincode The validated FP4 version of MATILDA JEV, with Decision Index **61.77**. Its corresponding BF16 source scores 62.43. This package contains packed FP4 weights and the native JEV decision head and runtime. ## Format and execution - E2M1 FP4 weights, with FP8 E4M3 block scales per 16 values and FP32 tensor scales. - BF16 activations and matrix multiplication (**W4A16**). A Triton kernel dequantizes one matrix at a time; this runtime does not use native FP4 Tensor Core GEMM. - 400 large text linear modules quantized. The decision readout, embeddings, normalization layers, vision tower and small gate projections retain their original precision. - Native `choice`, `noul` and `score` outputs, with up to 255 options per question. Weight tensors occupy **17.20 GB**, compared with 52.17 GB of source weight files. On AMD Instinct MI355X, measured resident memory is **16.3 GiB** versus 48.6 GiB for the BF16 source. Short-to-medium single-request P50 latency is **66 ms** versus 54 ms; this implementation primarily saves memory. These measurements are workload-specific. NVIDIA hardware and vision inference were not tested. ## Loading Use Python 3.12 or newer. The validated environment uses PyTorch 2.14.0+ROCm7.2, Transformers 5.17.0, Accelerate, Safetensors, Triton and Flash Linear Attention. Install a PyTorch build appropriate for your accelerator before the remaining dependencies in `requirements-runtime.txt`. Download the repository with `huggingface_hub.snapshot_download` and load it using the included `FP4DecisionModel` adapter: ```python import sys from pathlib import Path checkpoint = Path('/path/to/downloaded/model') sys.path.insert(0, str(checkpoint)) from jev_fp4 import FP4DecisionModel from kev.model import answer model = FP4DecisionModel(checkpoint, device='cuda:0') row = { 'state': 'The parcel arrived on schedule.', 'question': { 'type': 'choice', 'instructions': 'Classify delivery.', 'criteria': {'on_time': 'On time', 'late': 'Late'}, }, } print(answer(row['question'], model.predict([row])[0])) ``` The supplied `predict.py` accepts JSON lines with `state` + `question` or `state` + `questions` and produces native JEV answers: ```bash python /path/to/downloaded/model/predict.py --device cuda:0 < requests.jsonl ``` Packed weights require the included FP4 adapter; they cannot be loaded using the original BF16 loader. This is a decision model with a separate readout, not a text generation model. ## Full Decision Index Edition 0.2.1, all **150,317 requests across 44 benchmarks** completed successfully. All shard outputs were audited, and scoring was recomputed with identical results. `scores.json` contains the complete results; `comparison.json` compares all benchmarks against the exact BF16 source. | Metric | BF16 source | FP4 | |---|---:|---:| | Decision Index | 62.43 | **61.77** | | Raw | 71.42 | 70.90 | | Breadth | 61.15 | 60.48 | | Knowledge & Reasoning | 47.37 | 46.45 | | Language Understanding | 70.38 | 69.24 | | Retrieval & Classification | 64.98 | 64.41 | | Tools & Automation | 80.89 | 80.76 | | Arts & Human Taste | 41.96 | 42.05 | Area scores are chance-corrected skill multiplied by 100. GPQA Diamond changes from 51.02 to 48.47, MMLU-Pro from 83.60 to 81.70, and GSM8K from 79.45 to 79.53. The source model's training history includes benchmark-related material, so these are diagnostic quantization comparisons, not an independent held-out generalization claim.