Llama-3.2-1B-Instruct (GGUF Q4_K_M for RunSLM AI)

This repository provides the official 4-bit medium quantized (Q4_K_M) single-file GGUF distribution of Meta Llama 3.2 1B Instruct (Llama-3.2-1B-Instruct-Q4_K_M.gguf) engineered for on-device execution in RunSLM AI and llama.cpp runtimes.

Model Summary

  • Architecture: Llama 3.2 (Autoregressive Transformer with Grouped-Query Attention)
  • Base Model: meta-llama/Llama-3.2-1B-Instruct
  • Parameters: ~1.23 Billion
  • Quantization: Q4_K_M (4-bit medium quantization)
  • File Format: GGUF (.gguf)
  • File Size: ~808 MB
  • Context Length: Up to 128,000 tokens
  • License: Llama 3.2 Community License
  • Target Deployment: Local mobile & tablet hardware (Apple Silicon Metal UMA & Android Arm64 KleidiAI)

Features

  • Built with Meta Llama 3.2: Retains high reasoning fidelity, structured formatting, and multi-turn instruction following.
  • 100% Offline & Private: Runs entirely on physical device hardware with zero network transmission.
  • Low Memory Overhead: Consumes under 1.1 GB of RAM at 2,048 tokens context, running safely within standard 4GBโ€“6GB mobile operating system budgets.
  • Single-File Deployment: Self-contained GGUF tensor format with integrated vocabulary and chat template metadata.

Direct Download

The raw model binary can be downloaded directly from: https://huggingface.co/RunSLM-AI/Llama-3.2-1B-Instruct-GGUF/resolve/main/Llama-3.2-1B-Instruct-Q4_K_M.gguf

Prompt Template

This model uses the standard Llama 3 chat template:

<|start_header_id|>system<|end_header_id|>

You are RunSLM, an on-device AI assistant.<|eot_id|><|start_header_id|>user<|end_header_id|>

{prompt}<|eot_id|><|start_header_id|>assistant<|end_header_id|>
Downloads last month
-
GGUF
Model size
1B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for RunSLM-AI/Llama-3.2-1B-Instruct-GGUF

Quantized
(427)
this model