File size: 1,428 Bytes
5567ec8 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 | ---
license: mit
base_model: deepgrove/maple-preview-2bit-mlx
pipeline_tag: text-generation
tags:
- maple
- moe
- apple-silicon
- ssd-streaming
- local-inference
---
# Samosa Chat Maple 2-bit SSD streaming pack
This repository contains the production artifact layout used by Samosa Chat's
native Maple runtime. It is a storage-only repack of the MIT-licensed
[`deepgrove/maple-preview-2bit-mlx`](https://huggingface.co/deepgrove/maple-preview-2bit-mlx)
checkpoint pinned at revision
`361db5da5e74ff6fcdd852d478e1f266ce11013a`.
The model's numerical weights are unchanged. Expert tensors are stored in
fixed-size aligned records in `maple-experts.bin`, allowing Samosa to read only
the routed experts from SSD. Non-expert tensors are stored in
`maple-resident.safetensors`. `maple-manifest.json` describes and validates the
packed layout.
These files are intended for the bundled `samosa-maple` runtime. They are not a
drop-in MLX checkpoint because the original expert shards have been replaced by
Samosa's streaming container.
## Runtime files
- `maple-experts.bin`: aligned expert records streamed on demand from SSD
- `maple-resident.safetensors`: non-expert tensors retained by the runtime
- `maple-manifest.json`: strict layout metadata
- `config.json` and tokenizer files: model configuration and prompt encoding
Samosa verifies every artifact's byte length and SHA-256 digest before atomic
installation.
|