Llama 3.2 1B β€” Stateful KV Cache Core ML (unified model)

Single unified Core ML bundle: input_ids, causal_mask, state keyCache / valueCache.

Structure

ctx{N}_fp16/Llama32_1B_KVCache/   β€” FP16 unified model
ctx{N}_int4/Llama32_1B_KVCache/   β€” INT4 unified model (optional)

Usage

  • OS: macOS 15 / iOS 18+
  • Compute units: CPU_AND_GPU
  • Inference: make_state(); prefill + decode with the same model; see notebook Step 9 for causal_mask shapes.

Memory estimate (FP16 KV): 0.00 GB for ctx=4096.

Downloads last month
5
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for aoiandroid/llama32-1b-kvcache-coreml

Quantized
(407)
this model

Collection including aoiandroid/llama32-1b-kvcache-coreml