llama
Collection
aoiandroid models matching llama search β’ 5 items β’ Updated
Single unified Core ML bundle: input_ids, causal_mask, state keyCache / valueCache.
ctx{N}_fp16/Llama32_1B_KVCache/ β FP16 unified model
ctx{N}_int4/Llama32_1B_KVCache/ β INT4 unified model (optional)
make_state(); prefill + decode with the same model; see notebook Step 9 for causal_mask shapes.Memory estimate (FP16 KV): 0.00 GB for ctx=4096.
Base model
meta-llama/Llama-3.2-1B-Instruct