Text Generation
Transformers
Safetensors
GGUF
English
llama-3.2-1B-Instruct
llama.cpp
conversational
cycloevan commited on
Commit
6d75657
·
verified ·
1 Parent(s): e04979e

Add GGUF quantized builds (Q4_K_M, Q8_0) for llama.cpp on-device inference

Browse files

Converted from merged-vuln-detector/model.safetensors with convert_hf_to_gguf.py and quantized with llama-quantize. Greedy-decoding outputs verified identical to the original transformers model. Model card updated with GGUF usage and measured Apple M1 performance.

.gitattributes CHANGED
@@ -34,3 +34,5 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
  merged-vuln-detector/tokenizer.json filter=lfs diff=lfs merge=lfs -text
 
 
 
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
  merged-vuln-detector/tokenizer.json filter=lfs diff=lfs merge=lfs -text
37
+ vuln_detector-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
38
+ vuln_detector-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -13,6 +13,8 @@ metrics:
13
  - BLEU
14
  tags:
15
  - llama-3.2-1B-Instruct
 
 
16
  ---
17
 
18
  # Model Card for `merged-vuln-detector`
@@ -129,6 +131,40 @@ int main() {
129
  **Model Output:**
130
  > The code has a buffer overflow vulnerability due to the lack of bounds checking on the destination buffer size.
131
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
132
  ## Model Card Authors
133
 
134
  [Seokhee Chang]
 
13
  - BLEU
14
  tags:
15
  - llama-3.2-1B-Instruct
16
+ - gguf
17
+ - llama.cpp
18
  ---
19
 
20
  # Model Card for `merged-vuln-detector`
 
131
  **Model Output:**
132
  > The code has a buffer overflow vulnerability due to the lack of bounds checking on the destination buffer size.
133
 
134
+ ## GGUF / llama.cpp (On-device Inference)
135
+
136
+ Quantized GGUF builds of `merged-vuln-detector` are provided for on-device inference with [llama.cpp](https://github.com/ggml-org/llama.cpp). They were converted from `merged-vuln-detector/model.safetensors` with `convert_hf_to_gguf.py` and quantized with `llama-quantize`.
137
+
138
+ | File | Quantization | Size | Notes |
139
+ |------|--------------|------|-------|
140
+ | `vuln_detector-Q4_K_M.gguf` | Q4_K_M | 0.81 GB | Recommended for on-device use |
141
+ | `vuln_detector-Q8_0.gguf` | Q8_0 | 1.32 GB | Near-lossless |
142
+
143
+ Outputs were verified against the original safetensors model: under greedy decoding, the Q8_0 and Q4_K_M builds produce identical analyses to the `transformers` model.
144
+
145
+ **Measured performance** (Apple M1, Metal backend, Q4_K_M): ~47 tokens/s generation, ~178 tokens/s prompt processing — real-time interactive inference on a consumer laptop.
146
+
147
+ ### Run with llama.cpp
148
+
149
+ ```bash
150
+ llama-cli -hf cycloevan/vuln_detector:Q4_K_M \
151
+ -p "Analyze the security vulnerabilities in the following code.\n\n<CODE>\n\nAnalysis:\n" \
152
+ -n 256 --temp 0
153
+ ```
154
+
155
+ ### Run with llama-cpp-python
156
+
157
+ ```python
158
+ from llama_cpp import Llama
159
+
160
+ llm = Llama.from_pretrained("cycloevan/vuln_detector", filename="vuln_detector-Q4_K_M.gguf")
161
+
162
+ code = "def login(u, p): cursor.execute(f\"SELECT * FROM users WHERE name='{u}' AND pw='{p}'\")"
163
+ prompt = f"Analyze the security vulnerabilities in the following code.\n\n{code}\n\nAnalysis:\n"
164
+ out = llm(prompt, max_tokens=256, temperature=0)
165
+ print(out["choices"][0]["text"])
166
+ ```
167
+
168
  ## Model Card Authors
169
 
170
  [Seokhee Chang]
vuln_detector-Q4_K_M.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a03d80a47ae5504ca1e7d3a51d62bbf90b0f2e3a7d261671236f37368dd0e656
3
+ size 807693984
vuln_detector-Q8_0.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ff758f786ff0f6dbeca282e88fe1cd99b587105e6454b0d8c6f0cd7a9b0dd1c6
3
+ size 1321082528