schema_version: 1 model: id: smollm2-135m-instruct display_name: SmolLM2-135M-Instruct architecture: LlamaForCausalLM model_type: llama parameter_count: 134515008 context_length: 8192 tokenizer_type: GPT2Tokenizer tasks: - text-generation chat_template: true multimodal: false custom_code: false safetensors: true upstream: repo: HuggingFaceTB/SmolLM2-135M-Instruct revision: 12fd25f77366fa6b3b4b768ec3050bf629380bac license: identifier: apache-2.0 redistribution: allowed reason: Upstream declares a recognized license that permits redistribution; retain its terms and attribution. gated: false private: false source_url: https://huggingface.co/HuggingFaceTB/SmolLM2-135M-Instruct packages: - id: gguf-q4-k-m format: gguf precision: Q4_K_M filename: SmolLM2-135M-Instruct-Q4_K_M.gguf sha256: dd18a11b8634d1684448986b8c166f75319f52082d759654aaa8fe5bd2f057e3 size_bytes: 105453984 bits_per_weight: 4.83 runtime: provider: llama.cpp tested_revision: de699957b92f490efebad149665b0dccf127eaff hardware: estimated_ram_gb: 0.13 estimated_vram_gb: 0.12 recommended_ram_gb: 1.14 note: Estimate; runtime use varies with context length and configuration. validation: integrity: passed metadata: passed load: passed inference: passed tokenizer: passed tested_at: '2026-08-20T20:24:16.685253Z' details: version: '3' tensor_count: '272' metadata_count: '30' - id: gguf-q5-k-m format: gguf precision: Q5_K_M filename: SmolLM2-135M-Instruct-Q5_K_M.gguf sha256: 00680963c363ba10593daf7568dd6e1ee4c4771a608fe1b4e43f86b564d9b823 size_bytes: 112103328 bits_per_weight: 5.67 runtime: provider: llama.cpp tested_revision: de699957b92f490efebad149665b0dccf127eaff hardware: estimated_ram_gb: 0.13 estimated_vram_gb: 0.12 recommended_ram_gb: 1.15 note: Estimate; runtime use varies with context length and configuration. validation: integrity: passed metadata: passed load: passed inference: passed tokenizer: passed tested_at: '2026-08-20T20:24:18.938307Z' details: version: '3' tensor_count: '272' metadata_count: '30' - id: gguf-q8-0 format: gguf precision: Q8_0 filename: SmolLM2-135M-Instruct-Q8_0.gguf sha256: ee785d9b4836ddb57207ae6daa630206a756c99fc52e19696f1e2ea2e8a41b99 size_bytes: 144810912 bits_per_weight: 8.5 runtime: provider: llama.cpp tested_revision: de699957b92f490efebad149665b0dccf127eaff hardware: estimated_ram_gb: 0.17 estimated_vram_gb: 0.16 recommended_ram_gb: 1.2 note: Estimate; runtime use varies with context length and configuration. validation: integrity: passed metadata: passed load: passed inference: passed tokenizer: passed tested_at: '2026-08-20T20:24:21.225025Z' details: version: '3' tensor_count: '272' metadata_count: '30' generated_at: '2026-08-20T20:32:10.805207Z'