Instructions to use JohnP1/d1a-e2b-mlx-q8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use JohnP1/d1a-e2b-mlx-q8 with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download JohnP1/d1a-e2b-mlx-q8 --local-dir d1a-e2b-mlx-q8
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
D1A-E2B · MLX q8 (per-layer embeddings 4-bit)
The D1A-E2B v0.2 decision model (two epochs, JohnP1/d1a-e2b) converted for Apple Silicon: the trained adapter merged into Gemma 4 E2B, linear layers and token embeddings at 8 bits, the large per-layer embedding table at 4 bits, and the pointer head in fp32 (head.safetensors). Typed questions in, a calibrated probability for every option out, one forward pass. Same System One API as Jev.
| This MLX build | PyTorch bf16 | |
|---|---|---|
| Memory after load / peak (60 questions) | 4.2 GB / 6.4 GB | ~10-12 GB / ~15 GB |
| Load time | 1.4 s | 16-20 s |
| 6-question request (M1 Max) | ~480 ms cold, ~400 ms warm | similar |
Parity against the PyTorch reference path (bf16, the precision v0.2 was trained and evaluated in), measured on 274 questions (decision-v7 development records, the playground presets, long states): mean |dp| 0.009, p95 0.038, max 0.17, 5 changed answers (1.8%, the same rate as the v0.1 build), accuracy 0.824 vs 0.828, ECE 0.066 vs 0.072; the playground presets change no answer. Tag v0.1-1epoch keeps the build of v0.1. A pure 4-bit build failed the same gate (8.4% changed answers) and is not published.
Photos and voice
media/ holds Gemma 4 E2B's own vision and audio encoders (bf16, ~1 GB, unchanged from the base: the LoRA adapter only touches the text model). With them d1a.serve answers questions about a photo or a voice clip with this same model (POST /v1/systemone/media): the encoders turn the media into soft tokens and the 8-bit language model reads them right after <state>. They are fetched on the first such request; a plain download skips them.
Zero-shot, on the playground's 10 sample photos and EN/JA voice notes: parcel damaged 4 of 6, where it was left 5 of 6, what the speaker needs 4 of 4. Against the PyTorch path of the same checkpoint (bf16, full base) on the same 20 questions: 0 answers change; the largest probability shift is 0.21, on the one photo both read wrongly (a mailbox, "in a delivery locker" at 0.34 here vs 0.55), every other ≤ 0.09. D1A-E4B v0.3 reads them better (damaged 5 of 6, place 6 of 6).
Run it
pip install "d1a[serve] @ git+https://github.com/jonpol01/d1a@mlx-gemma4"
python -m d1a.serve --run JohnP1/d1a-e2b-mlx-q8 --port 8009
Apple Silicon only (MLX). Playground: https://github.com/jonpol01/d1a-playground
License
Apache-2.0. Base model: Gemma 4 by Google (Apache-2.0). Code: github.com/jonpol01/d1a, built on Kev by Jared Palmer (Apache-2.0).
- Downloads last month
- 114
8-bit