Model Card for Gemma-3-270M-IT (Aria Quant Bundle, q4)

Model Details

Model Description

Gemma-3-270M-IT is a ~270-million-parameter instruction-tuned, dense Transformer decoder-only language model developed by Google, part of the Gemma 3 family. It features GeGLU activation, Grouped Query Attention (GQA), and 32K native context length. Pre-trained on diverse web-scale corpora and aligned via instruction tuning + RLHF. This distribution is provided by Aria Compute as an aria-quant-bundle — a quantized package using Hadamard pre-processing + per-channel quantization. Optimized for CPU-only, on-device inference on mobile phones, edge devices, and single-board computers via the Aria Engine runtime. No GPU or cloud connection is required.

  • Developed by: Google
  • Quantized and distributed by: Aria Compute
  • Model type: Dense Transformer decoder-only (language, instruction-tuned)
  • Language(s): English (primary), Chinese, and 30+ additional languages
  • License: Apache 2.0
  • Finetuned from model: google/gemma-3-270m-it

Model Sources

Uses

Direct Use

This quantized bundle is intended for on-device, offline text-generation tasks on resource-constrained hardware, including:

  • On-device chat and conversational assistants
  • Real-time text completion and basic code snippet generation
  • Instruction-following tasks for mobile and IoT applications
  • Lightweight text embeddings for on-device retrieval and classification
  • Short-form summarization of notifications, messages, and local content
  • Local document analysis up to 32K context (chunked)

All inference runs locally on CPU. No data is sent to external servers.

Target Devices

Platform Runtime Memory Feasibility
High-end smartphone (8 GB) ~0.20 GB ✅ Recommended
Mid-range smartphone (4–6 GB) ~0.20 GB
Budget phone (2–3 GB) ~0.20 GB
Raspberry Pi 5 / SBC (4–8 GB) ~0.20 GB
IoT gateway (1–2 GB) ~0.20 GB
Wearable (1 GB) ~0.20 GB

Memory breakdown (q4, at 4K context): ~0.13 GB quantized model weights (mmap) + ~20 MB KV cache + ~25 MB runtime overhead + ~25 MB per-channel metadata overhead ≈ ~0.20 GB.

Note: KV cache is extremely compact due to aggressive GQA (18 layers × 2 KV heads × head_dim=64), ~20% of an equivalent full-attention model's cache, enabling practical 32K context on sub-1 GB devices.

Out-of-Scope Use

  • Long-form creative writing (>2K tokens per generation)
  • Mathematical theorem proving or formal verification
  • Full program/application synthesis
  • Multimodal input (this model is text-only)
  • Real-time audio/speech processing (use Aria speech models)
  • Safety-critical decision systems without human oversight
  • Deployment in production when batch inference or GPU acceleration is required (this bundle targets CPU-only, single-prompt inference)
  • Tasks requiring factual precision beyond the model's ~270M parameter capacity

How to Get Started with the Model

Download from Aria Compute

Authenticated dashboard users can download the bundle via: https://ariacompute.com/dashboard/models

Quantization Recipe

This bundle uses a per-channel quantization recipe, one of several precision options in the Aria Compute lineup:

Component Quantization Strategy Details
Attention Q/K/V/O weights 4-bit Per-channel codebooks, Hadamard pre-processing
FFN gate/up/down weights 4-bit Per-channel codebooks, Hadamard pre-processing
RMSNorm weights FP16 Preserved at full precision
Embedding table FP16 Preserved at full precision
  • Bundle size: ~0.13 GB (BF16 original: ~0.54 GB)
  • Generation quality: Awaiting gen_quant_eval audit. Per-channel quantization preserves per-output-channel distribution characteristics. Formal quality benchmarks against FP16 and other Aria quant recipes are pending
  • Calibration-free: Hadamard pre-processing + per-channel quantization, no task-specific calibration data required
  • Other precision options: Also available for Gemma-3-270M-IT: gemma-3-270m-it_q8_channel (per-channel 8-bit) and gemma-3-270m-it_q326_channel (mixed precision, attn 4-bit + FFN ~3-bit)

Model Architecture

Gemma-3-270M-IT employs a dense Transformer decoder architecture with GeGLU activation and Grouped Query Attention:

Parameter Gemma 3 1B Gemma 3 270M
Layers 26 18
Hidden size 1,152 640
FFN intermediate size 4,608 2,560
Attention heads (Query) 9 10
Attention heads (KV) 3 (GQA, group 3) 2 (GQA, group 5)
Head dimension 128 64
Activation GeGLU GeGLU
Position encoding RoPE (θ = 10,000) RoPE (θ = 10,000)
Normalization RMSNorm (pre-norm) RMSNorm (pre-norm)
Vocabulary size 256,128 256,128
Max context length 32,768 32,768

Design highlights (shared with Gemma family):

  • GQA (Grouped Query Attention): 2 KV heads serving 10 query heads (group 5) — KV Cache memory is 5× smaller than full attention
  • GeGLU activation: GELU with tanh approximation gating, efficient for on-device inference
  • RoPE position encoding: Standard 10K base frequency supporting 32K context length
  • RMSNorm pre-normalization: Lightweight normalization before each sub-layer

Key architecture differences from Gemma 3 1B:

  • ~75% smaller parameter count (270M vs 1.1B) — ~4× reduction, the smallest Gemma 3 variant
  • Fewer layers (18 vs 26) with narrower hidden dimension (640 vs 1,152) and narrower FFN (2,560 vs 4,608)
  • Halved head dimension (64 vs 128) — lighter per-head computation compensated with more query heads (10 vs 9)
  • More aggressive GQA (group 5 vs group 3) — further KV cache compression, critical for fitting 32K context on memory-constrained devices
  • Same 32K context length — full long-context support preserved despite the smaller parameter budget
  • Instruction-tuned variant — this bundle is based on the IT (instruction-tuned) checkpoint, not the base pre-trained model, providing better instruction-following capability out of the box

Bias, Risks, and Limitations

Limitations

  • Reasoning depth: Multi-step logical reasoning is highly limited for a 270M-class model. Verify outputs in high-stakes scenarios; consider larger Gemma 3 variants for reasoning tasks.
  • Mathematics: Simple arithmetic may be attempted but is unreliable. Advanced quantitative reasoning is out of scope. Use larger models for mathematical tasks.
  • Code generation: Capable of single-line completions and basic snippets; unreliable for multi-line code or structured programs.
  • Factual knowledge: Very limited world knowledge due to ~270M parameter scale. Always verify factual claims against authoritative sources. This model is best suited for instruction-following and lightweight text processing rather than encyclopedic knowledge retrieval.
  • Instruction following: Handles simple single-constraint instructions. Complex multi-constraint prompts may cause significant degradation, especially at longer contexts.
  • Quantization drift: 4-bit per-channel quantization may exhibit generation drift versus FP16, especially on ambiguous or open-ended prompts. For higher fidelity, use gemma-3-270m-it_q326_channel (mixed precision, attn 4-bit + FFN ~3-bit) or gemma-3-270m-it_q8_channel (per-channel 8-bit).

Bias and Risks

  • Bias: As with all large language models trained on web-scale data, Gemma 3 may reflect societal biases present in its training corpus. Evaluate outputs before deployment in sensitive domains (hiring, healthcare, law).
  • Toxicity: The instruction-tuned model has been safety-aligned with RLHF. However, no safety filter is exhaustive. Consider an additional output classifier in high-risk environments.
  • Hallucination: May generate plausible-sounding but factually incorrect information. Implement output verification for critical applications. Hallucination risk is elevated for smaller models due to limited memorization capacity.
  • Dual-use risk: Text-generation capabilities could be misused for spam, disinformation, or impersonation. Deploy responsibly and in accordance with the Apache 2.0 license terms.

Recommendations

Users (both direct and downstream) should be made aware of the above risks, biases, limitations, and constraints of the model. We recommend:

  • Adding a lightweight output safety classifier for user-facing deployments
  • Verifying factual claims with external knowledge bases
  • Not using the model for high-stakes decisions without human review
  • Considering larger Gemma 3 variants (1B or 4B) for tasks requiring stronger reasoning or factual recall
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ariacompute/gemma-3-270m-it_q4

Finetuned
(1134)
this model

Datasets used to train ariacompute/gemma-3-270m-it_q4

Paper for ariacompute/gemma-3-270m-it_q4

Evaluation results