MATILDA JEV FP4 by Maincode

The validated FP4 version of MATILDA JEV, with Decision Index 61.77. Its corresponding BF16 source scores 62.43. This package contains packed FP4 weights and the native JEV decision head and runtime.

Format and execution

  • E2M1 FP4 weights, with FP8 E4M3 block scales per 16 values and FP32 tensor scales.
  • BF16 activations and matrix multiplication (W4A16). A Triton kernel dequantizes one matrix at a time; this runtime does not use native FP4 Tensor Core GEMM.
  • 400 large text linear modules quantized. The decision readout, embeddings, normalization layers, vision tower and small gate projections retain their original precision.
  • Native choice, noul and score outputs, with up to 255 options per question.

Weight tensors occupy 17.20 GB, compared with 52.17 GB of source weight files. On AMD Instinct MI355X, measured resident memory is 16.3 GiB versus 48.6 GiB for the BF16 source. Short-to-medium single-request P50 latency is 66 ms versus 54 ms; this implementation primarily saves memory. These measurements are workload-specific. NVIDIA hardware and vision inference were not tested.

Loading

Use Python 3.12 or newer. The validated environment uses PyTorch 2.14.0+ROCm7.2, Transformers 5.17.0, Accelerate, Safetensors, Triton and Flash Linear Attention. Install a PyTorch build appropriate for your accelerator before the remaining dependencies in requirements-runtime.txt.

Download the repository with huggingface_hub.snapshot_download and load it using the included FP4DecisionModel adapter:

import sys
from pathlib import Path

checkpoint = Path('/path/to/downloaded/model')
sys.path.insert(0, str(checkpoint))
from jev_fp4 import FP4DecisionModel
from kev.model import answer

model = FP4DecisionModel(checkpoint, device='cuda:0')
row = {
    'state': 'The parcel arrived on schedule.',
    'question': {
        'type': 'choice',
        'instructions': 'Classify delivery.',
        'criteria': {'on_time': 'On time', 'late': 'Late'},
    },
}
print(answer(row['question'], model.predict([row])[0]))

The supplied predict.py accepts JSON lines with state + question or state + questions and produces native JEV answers:

python /path/to/downloaded/model/predict.py --device cuda:0 < requests.jsonl

Packed weights require the included FP4 adapter; they cannot be loaded using the original BF16 loader. This is a decision model with a separate readout, not a text generation model.

Full Decision Index

Edition 0.2.1, all 150,317 requests across 44 benchmarks completed successfully. All shard outputs were audited, and scoring was recomputed with identical results. scores.json contains the complete results; comparison.json compares all benchmarks against the exact BF16 source.

Metric BF16 source FP4
Decision Index 62.43 61.77
Raw 71.42 70.90
Breadth 61.15 60.48
Knowledge & Reasoning 47.37 46.45
Language Understanding 70.38 69.24
Retrieval & Classification 64.98 64.41
Tools & Automation 80.89 80.76
Arts & Human Taste 41.96 42.05

Area scores are chance-corrected skill multiplied by 100. GPQA Diamond changes from 51.02 to 48.47, MMLU-Pro from 83.60 to 81.70, and GSM8K from 79.45 to 79.53. The source model's training history includes benchmark-related material, so these are diagnostic quantization comparisons, not an independent held-out generalization claim.

Downloads last month
3
Safetensors
Model size
15B params
Tensor type
BF16
路
F8_E4M3
路
U8
路
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support