Text Generation
Transformers
Safetensors
GGUF
English
llama-3.2-1B-Instruct
llama.cpp
conversational
cycloevan commited on
Commit
31fb51f
·
verified ·
1 Parent(s): 6d75657

Model card: add measured quantization quality (ROUGE-L/BLEU, 100 samples) and llama-bench performance

Browse files
Files changed (1) hide show
  1. README.md +23 -2
README.md CHANGED
@@ -140,9 +140,30 @@ Quantized GGUF builds of `merged-vuln-detector` are provided for on-device infer
140
  | `vuln_detector-Q4_K_M.gguf` | Q4_K_M | 0.81 GB | Recommended for on-device use |
141
  | `vuln_detector-Q8_0.gguf` | Q8_0 | 1.32 GB | Near-lossless |
142
 
143
- Outputs were verified against the original safetensors model: under greedy decoding, the Q8_0 and Q4_K_M builds produce identical analyses to the `transformers` model.
144
 
145
- **Measured performance** (Apple M1, Metal backend, Q4_K_M): ~47 tokens/s generation, ~178 tokens/s prompt processing real-time interactive inference on a consumer laptop.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
146
 
147
  ### Run with llama.cpp
148
 
 
140
  | `vuln_detector-Q4_K_M.gguf` | Q4_K_M | 0.81 GB | Recommended for on-device use |
141
  | `vuln_detector-Q8_0.gguf` | Q8_0 | 1.32 GB | Near-lossless |
142
 
143
+ ### Quantization quality
144
 
145
+ Measured on 100 samples from `doss1232/vulnerable-code` (shuffled with seed 42, 500-row held-out pool, first 100 evaluated; fine-tuning prompt format; greedy decoding, `max_new_tokens=128`; identical harness for every variant):
146
+
147
+ | Variant | ROUGE-L F1 | BLEU | Exact output match vs original |
148
+ |---------|-----------:|-----:|-------------------------------:|
149
+ | transformers (original, fp16) | 0.1962 | 0.0645 | — |
150
+ | GGUF F16 | 0.1963 | 0.0645 | 98% |
151
+ | GGUF Q8_0 | 0.1987 | 0.0675 | 89% |
152
+ | GGUF Q4_K_M | 0.2325 | 0.0916 | 36% |
153
+
154
+ Quantization does not degrade benchmark quality — Q8_0 and Q4_K_M match or slightly exceed the original model's reference metrics (differences within noise for a 1B model). Q4_K_M's token-level outputs diverge from the fp16 model on many samples while remaining equivalent in quality; choose Q8_0 when close output fidelity to the fp16 model matters.
155
+
156
+ Note: these absolute numbers are not directly comparable to the "Evaluation Results" table above, which used a different sampling/decoding protocol.
157
+
158
+ ### Measured performance (Apple M1, Metal, `llama-bench`)
159
+
160
+ | Variant | Prompt processing (pp512) | Generation (tg128) |
161
+ |---------|--------------------------:|-------------------:|
162
+ | GGUF F16 | 941 t/s | 19.2 t/s |
163
+ | GGUF Q8_0 | 618 t/s | 29.4 t/s |
164
+ | GGUF Q4_K_M | 504 t/s | 42.3 t/s |
165
+
166
+ Q4_K_M generates **2.2× faster than F16** (generation is memory-bandwidth-bound, so smaller weights win; prompt processing is compute-bound and favors F16). End-to-end on the 100-sample eval batch, llama.cpp Q4_K_M finished in 53 s vs 138 s for `transformers` fp16 on MPS (**2.6× faster**). Weight memory footprint: 2.48 GB (F16) → 0.81 GB (Q4_K_M).
167
 
168
  ### Run with llama.cpp
169