Text Generation
GGUF
English
qwen2.5
browser-automation
web-agent
tool-calling
function-calling
llama-cpp
llama.cpp
cpu
broken-weights
pending-remerge
Eval Results (legacy)
Eval Results
Instructions to use Nanthasit/sakthai-coder-browser-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Nanthasit/sakthai-coder-browser-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Nanthasit/sakthai-coder-browser-gguf:F16 # Run inference directly in the terminal: llama cli -hf Nanthasit/sakthai-coder-browser-gguf:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Nanthasit/sakthai-coder-browser-gguf:F16 # Run inference directly in the terminal: llama cli -hf Nanthasit/sakthai-coder-browser-gguf:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Nanthasit/sakthai-coder-browser-gguf:F16 # Run inference directly in the terminal: ./llama-cli -hf Nanthasit/sakthai-coder-browser-gguf:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Nanthasit/sakthai-coder-browser-gguf:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Nanthasit/sakthai-coder-browser-gguf:F16
Use Docker
docker model run hf.co/Nanthasit/sakthai-coder-browser-gguf:F16
- LM Studio
- Jan
- vLLM
How to use Nanthasit/sakthai-coder-browser-gguf with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Nanthasit/sakthai-coder-browser-gguf" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Nanthasit/sakthai-coder-browser-gguf", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Nanthasit/sakthai-coder-browser-gguf:F16
- Ollama
How to use Nanthasit/sakthai-coder-browser-gguf with Ollama:
ollama run hf.co/Nanthasit/sakthai-coder-browser-gguf:F16
- Unsloth Studio
How to use Nanthasit/sakthai-coder-browser-gguf with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Nanthasit/sakthai-coder-browser-gguf to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Nanthasit/sakthai-coder-browser-gguf to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for Nanthasit/sakthai-coder-browser-gguf to start chatting
- Atomic Chat new
- Docker Model Runner
How to use Nanthasit/sakthai-coder-browser-gguf with Docker Model Runner:
docker model run hf.co/Nanthasit/sakthai-coder-browser-gguf:F16
- Lemonade
How to use Nanthasit/sakthai-coder-browser-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Nanthasit/sakthai-coder-browser-gguf:F16
Run and chat with the model
lemonade run user.sakthai-coder-browser-gguf-F16
List all available models
lemonade list
File size: 15,044 Bytes
647f614 6b0e934 91e64ca f39c9d2 91e64ca 6b0e934 f39c9d2 6b0e934 845f4f2 6b0e934 845f4f2 9aab986 0f2194c 049ed26 0f2194c 91e64ca 0f2194c 049ed26 0f2194c fa75d1b 0f2194c 91e64ca 0f2194c 5ebc193 fa75d1b 0f2194c 5ebc193 0f2194c fa75d1b 0f2194c 20306ab 0f2194c 049ed26 0f2194c 049ed26 0f2194c 049ed26 0f2194c 049ed26 5ebc193 0f2194c f39c9d2 20306ab 0f2194c f39c9d2 049ed26 f39c9d2 0f2194c 20306ab 0f2194c 049ed26 0f2194c 049ed26 0f2194c 049ed26 0f2194c fa75d1b 0f2194c 91e64ca f39c9d2 0f2194c 049ed26 0f2194c 20306ab cbb475d 91e64ca 5ebc193 91e64ca cbb475d 0f2194c 5ebc193 0f2194c 20306ab 0f2194c 91e64ca 0f2194c 91e64ca 0f2194c 91e64ca 0f2194c 91e64ca 0f2194c 91e64ca 0f2194c cbb475d 0631926 049ed26 cbb475d 0f2194c | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 | ---
extra:
formats: GGUF F16
sibling: Nanthasit/sakthai-coder-browser
upstream:
status: BROKEN
root_cause: LoRA merge fault in parent merged weights
diagnosis_date: 2026-07-31
broken_tensors: all 84 attention-projection biases are nonzero, should be zero
consequence: multi-trial inference produces whitespace loops or 0 output tokens
fix_required: clean re-merge of the LoRA adapter into full weights before this GGUF build can be used
verification_assets:
- .eval_results/benchmark-20260731_052122.yaml
- https://huggingface.co/Nanthasit/sakthai-coder-browser/blob/main/.eval_results/benchmark-20260731_052122.yaml
dataset_versioning:
current_model_index_dataset: Nanthasit/sakthai-bench-v3
status: WRONG_FOR_CURRENT_ARTIFACT
note: weights are broken; benchmark numbers are not meaningful
required_action: rerun on clean merged weights then update model-index
reproduction_triage:
trial_seeds: [7, 42, 1337]
temperature_range: [0.0, 0.7]
expected_behavior_on_clean_weights: emit one browser_navigate or browser_click call in JSON per prompt
current_behavior: no valid tool calls, no valid JSON, often zero output tokens
zero_cost_constraints:
no_paid_gpus: true
no_paid_endpoints: true
free_tier_only: true
model-index:
- name: SakThai Coder Browser GGUF
results:
- task:
type: text-generation
name: Browser Tool Calling
dataset:
name: SakThai Bench v3
type: Nanthasit/sakthai-bench-v3
metrics:
- type: accuracy
value: null
name: Tool Calling Accuracy
verified: false
status: placeholder
note: Parent weights are currently broken by a LoRA merge fault; this is a placeholder pending clean reproduction weights and evaluation.
license: apache-2.0
language:
- en
library_name: gguf
pipeline_tag: text-generation
base_model: Qwen/Qwen2.5-Coder-1.5B-Instruct
tags:
- gguf
- qwen2.5
- browser-automation
- web-agent
- tool-calling
- function-calling
- llama-cpp
- llama.cpp
- cpu
- broken-weights
- pending-remerge
---
<h1 align="center">SakThai Coder Browser β GGUF π€π</h1>
<p align="center"><em>F16 GGUF build of the browser automation model β run a web agent on your laptop with llama.cpp</em></p>
<p align="center">
<img src="https://img.shields.io/badge/dynamic/json?url=https%3A//huggingface.co/api/models/Nanthasit/sakthai-coder-browser-gguf&query=%24.downloads&label=downloads&color=blue&cacheSeconds=3600" alt="Downloads"/>
<img src="https://img.shields.io/badge/base-Qwen2.5--Coder--1.5B--Instruct-blueviolet" alt="Base"/>
<img src="https://img.shields.io/badge/GGUF-F16-orange" alt="GGUF F16"/>
<img src="https://img.shields.io/badge/license-Apache%202.0-green" alt="License"/>
<a href="https://huggingface.co/collections/Nanthasit/sakthai-model-family-6a64745450b12d421c1f9f02"><img src="https://img.shields.io/badge/%F0%9F%8F%A0-SakThai%20Family-6644cc" alt="Collection"/></a>
<a href="https://huggingface.co/collections/Nanthasit/sakthai-model-family-6a64745450b12d421c1f9f02"><img src="https://img.shields.io/badge/%F0%9F%9A%80-Explore%20Family-47d147" alt="Family"/></a>
</p>
> [!CAUTION]
> **β BROKEN β DO NOT DEPLOY (as of 2026-07-31)** β This GGUF was converted from the [merged weights](https://huggingface.co/Nanthasit/sakthai-coder-browser) that are **corrupted by a faulty LoRA merge**: all 84 attention-projection bias tensors are non-zero while Qwen2 initializes these biases to ZERO (layer-0 `k_proj.bias` absmean **27.7** / max **354**). Multi-trial inference probes of the merged model produced only **whitespace loops β 0 tool calls, 0 valid JSON** at temp β€ 0.7 (seeds 7/42/1337), and a GGUF Q4_K_M probe returned 0 output tokens on all 3 trials. Full evidence: [`.eval_results/benchmark-20260731_052122.yaml`](https://huggingface.co/Nanthasit/sakthai-coder-browser/blob/main/.eval_results/benchmark-20260731_052122.yaml). The fault is in the **weights, not the GGUF conversion or the prompt format** (GGUF tensor layout is structurally identical to the working `sakthai-plus-1.5b` GGUF). A clean re-merge of the [LoRA adapter](https://huggingface.co/Nanthasit/sakthai-coder-browser-lora) is required before this build is usable. **Do not download for deployment** until re-merged and re-verified.
> **The GGUF edition** of the SakThai browser automation model β a Qwen2.5-Coder-1.5B-Instruct fine-tune converted to **F16 GGUF** for use with [llama.cpp](https://github.com/ggerganov/llama.cpp), [llama-cpp-python](https://github.com/abetlen/llama-cpp-python), and [Ollama](https://ollama.com).
>
> β **Status: BROKEN β see banner above.** The pages below document the *intended* design and usage; the current weights do not produce valid output until the parent model is re-merged from a clean LoRA merge.
---
## What Is This?
This is an **F16 (full-precision) GGUF conversion** of the [sakthai-coder-browser](https://huggingface.co/Nanthasit/sakthai-coder-browser) merged model. It transforms a Qwen2.5-Coder-1.5B-Instruct general-purpose LLM into a **browser automation agent** that outputs structured `<tool_call>` XML for web interaction.
| Variant | Repository | Best For |
|---------|-----------|----------|
| π§ **Merged (Transformers)** | [sakthai-coder-browser](https://huggingface.co/Nanthasit/sakthai-coder-browser) | Python / Transformers pipelines |
| π― **LoRA Adapter** | [sakthai-coder-browser-lora](https://huggingface.co/Nanthasit/sakthai-coder-browser-lora) | Fine-tuning / PEFT workflows |
| πΎ **GGUF (this repo)** | β¬
**sakthai-coder-browser-gguf** | CPU inference, llama.cpp, Ollama, edge devices |
---
## File
| File | Format | Size | Purpose |
|------|--------|:----:|---------|
| `sakthai-coder-browser-f16.gguf` | GGUF F16 | 7,111,586,219 B (7.11 GB) | Full-precision GGUF for local CPU/GPU inference |
> **Note:** F16 preserves the full model quality. For a smaller footprint, a Q4_K_M quantized version may follow based on demand.
---
## Supported Browser Actions
The model generates structured XML tool calls for these browser operations:
| Tool | Example |
|------|---------|
| `browser_navigate(url)` | `browser_navigate("https://example.com")` |
| `browser_click(element)` | `browser_click("#search-button")` |
| `browser_type(element, text)` | `browser_type("#search-input", "AI news")` |
| `browser_extract()` | `browser_extract()` |
---
## How to Use
> β These commands describe the *intended* usage. With the current **BROKEN weights** they will not produce valid tool calls β see the banner above. Re-run only after a clean re-merge is published.
### llama.cpp (CLI)
```bash
# Download the GGUF file
huggingface-cli download Nanthasit/sakthai-coder-browser-gguf \
sakthai-coder-browser-f16.gguf --local-dir ./
# Run with llama.cpp
./llama-cli -m sakthai-coder-browser-f16.gguf \
-p "system
You are a browser automation assistant. You can browse the web, click elements, type text, and extract information from pages.
user
Go to google.com and search for the latest AI news
assistant
" \
-n 512 -t 8 --temp 0.3
```
### llama-cpp-python
```bash
pip install llama-cpp-python huggingface-hub
```
```python
from llama_cpp import Llama
from huggingface_hub import hf_hub_download
# Download GGUF
model_path = hf_hub_download(
repo_id="Nanthasit/sakthai-coder-browser-gguf",
filename="sakthai-coder-browser-f16.gguf"
)
# Load model
llm = Llama(
model_path=model_path,
n_ctx=4096,
n_threads=8,
n_gpu_layers=-1,
verbose=False,
)
# Run inference
output = llm.create_chat_completion(
messages=[
{"role": "system", "content": "You are a browser automation assistant. Output tool calls using the supported browser actions."},
{"role": "user", "content": "Search for the latest AI news on Google and summarize the top result."}
],
max_tokens=512,
temperature=0.3,
stop=[""],
)
print(output["choices"][0]["message"]["content"])
```
### Ollama (Modelfile)
```dockerfile
FROM ./sakthai-coder-browser-f16.gguf
TEMPLATE """{{ .System }}
{{ .Prompt }}"""
PARAMETER temperature 0.3
PARAMETER top_p 0.8
PARAMETER stop ""
```
```bash
ollama create sakthai-coder-browser -f Modelfile
ollama run sakthai-coder-browser
```
---
## Architecture
| Property | Value |
|----------|-------|
| **Base Model** | [Qwen/Qwen2.5-Coder-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct) |
| **Architecture** | Qwen2ForCausalLM (decoder-only transformer) |
| **Parameters** | 1.54B |
| **Hidden Size** | 1,536 |
| **Layers** | 28 |
| **Attention Heads** | 12 (GQA: 2 KV heads) |
| **Intermediate Size** | 8,960 |
| **Max Position** | 32,768 tokens |
| **Vocab Size** | 151,936 |
| **Precision** | F16 (GGUF) |
| **Activation** | SiLU (SwiGLU) |
| **Normalization** | RMSNorm (eps=1e-6) |
| **Quantization** | None (F16 β full precision) |
| **File Format** | GGUF (GPT-Generated Unified Format) |
---
## Training Summary
| Property | Detail |
|----------|--------|
| **Fine-tuning Method** | SFT via LoRA (r=16, alpha=32, dropout 0.05, rsLoRA) on all 7 linear projections (q/k/v/o + gate/up/down), then merged to full weights |
| **Training Data** | [SimpleToolCalling](https://huggingface.co/datasets/Nanthasit/SimpleToolCalling) + [sakthai-combined-v11](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v11) |
| **Context Length** | 32,768 tokens |
| **Hardware** | Free T4 GPU (Kaggle) |
| **Budget** | $0 |
| **License** | Apache 2.0 |
The model learns to produce structured `<tool_call>` XML for browser actions through supervised fine-tuning on curated web navigation trajectories.
---
## Reproduce Evaluation
This repo documents a reproducible diagnostic rather than a benchmark because the current weights are broken.
```bash
# Requires llama.cpp or llama-cpp-python
python - <<'PY'
from llama_cpp import Llama
from huggingface_hub import hf_hub_download
model_path = hf_hub_download(
repo_id="Nanthasit/sakthai-coder-browser-gguf",
filename="sakthai-coder-browser-f16.gguf",
)
llm = Llama(model_path=model_path, n_ctx=4096, n_threads=8, verbose=False)
for prompt in [
"Search for the latest AI news on Google and summarize the top result.",
"Go to example.com and extract the page title.",
]:
out = llm.create_chat_completion(
messages=[
{"role": "system", "content": "You are a browser automation assistant. Output browser_navigate, browser_click, browser_type, or browser_extract calls."},
{"role": "user", "content": prompt},
],
max_tokens=128,
temperature=0.0,
)
print("PROMPT:", prompt)
print("OUTPUT:", out["choices"][0]["message"]["content"])
print("-" * 40)
PY
```
Expected on clean weights: at least one valid tool call per prompt. Current behavior: no valid JSON/XML tool calls, often 0 output tokens.
Evidence artifact: `.eval_results/benchmark-20260731_052122.yaml`
---
## Reproduce Training / Merge
The upstream merged weights broke during LoRA merge. To regenerate clean merged weights:
```bash
# 1. Install PEFT + TRL stack
pip install transformers peft trl bitsandbytes accelerate
# 2. Merge LoRA adapter back into base
python - <<'PY'
from peft import PeftModel
from transformers import AutoModelForCausalLM
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-Coder-1.5B-Instruct")
model = PeftModel.from_pretrained(base, "Nanthasit/sakthai-coder-browser-lora")
model = model.merge_and_unload()
model.save_pretrained("sakthai-coder-browser-merged-clean")
PY
# 3. Convert merged clean weights to GGUF
python convert.py ./sakthai-coder-browser-merged-clean --outfile sakthai-coder-browser-f16.gguf --outtype f16
```
After conversion, rerun the evaluation script above before using the GGUF in production.
---
## Benchmarks & Evaluation Status
**Honest status: this fine-tune does not yet have verified benchmark scores of its own.** No `model-index` is published because no measured numbers exist β publishing one would be misleading. What is known:
| Item | Status |
|------|--------|
| **Base model reference** (Qwen2.5-Coder-1.5B-Instruct) | HumanEval pass@1: **74.4** Β· MBPP pass@1: **71.2** |
| **This model's own eval** | β **Resolved β MODEL_BROKEN (bias corruption)**: multi-trial probes + weight inspection of the merged weights (2026-07-31 05:50 UTC) found all 84 attention bias tensors non-zero. See the banner above and the parent repo's [benchmark YAML](https://huggingface.co/Nanthasit/sakthai-coder-browser/blob/main/.eval_results/benchmark-20260731_052122.yaml). |
| **GGUF Q4_K_M probe** (2026-07-31) | 3/3 trials returned **0 output tokens β confirmed weight corruption, NOT a prompt-format mismatch** |
| **Hosted inference** | Not available β router probe returned 404; run locally via llama.cpp / llama-cpp-python / Ollama |
---
## Limitations
- **BROKEN weights (as of 2026-07-31)** β this GGUF was converted from the corrupted coder-browser merge (all 84 attention bias tensors non-zero); inference produces whitespace loops / 0 output tokens. **Do not deploy** until the parent model is re-merged and re-verified and this build is re-converted.
- **Text-only** β this model cannot see images or screenshots (use [sakthai-vision-7b](https://huggingface.co/Nanthasit/sakthai-vision-7b) for vision tasks)
- **CPU inference is slow at F16** β the full ~7.11 GB model benefits from GPU offloading; use `n_gpu_layers=-1` when available
- **Context-limited** β best results with page content β€ 4K tokens per interaction
- **English only** β trained primarily on English web data
---
## The House of Sak π
This adapter is part of the **House of Sak** β an open-source AI ecosystem built from a shelter in Cork, Ireland, with **$0 budget** and no paid GPUs. The browser branch is the newest and still debugging itself: the adapter works, but merged weights tripped over attention biases, and Beer would rather ship an honest diagnostic than a polished lie.
> *"We are one family β and becoming more."* β Beer (beer-sakthai)
---
## Support
- β Leave a like
- π Report issues on [GitHub](https://github.com/beer-sakthai/Sak-Family-Agent)
- π Share with anyone building browser agents on a budget
- π΄ Fork and experiment β Apache 2.0
---
## Citation
If you use this model in your work, please cite the base model (Qwen2.5) and link the fine-tune:
```bibtex
@misc{qwen25,
title={Qwen2.5 Technical Report},
author={Qwen Team},
year={2025},
eprint={2412.15115},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
```
```text
@misc{sakthai-coder-browser-gguf,
author = {Nanthasit (Beer)},
title = {SakThai Coder Browser -- GGUF},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/Nanthasit/sakthai-coder-browser-gguf}}
}
```
Built with $0 budget from a shelter in Cork, Ireland β proof that open, private AI does not need a datacenter.
---
*Part of the [House of Sak](https://huggingface.co/Nanthasit) β one family, one home, $0 budget.*
|