File size: 5,956 Bytes
fe34583
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8d7c9f1
 
 
 
 
 
 
 
 
 
 
 
 
 
fe34583
 
 
 
 
 
 
 
 
 
 
 
 
 
8d7c9f1
fe34583
 
 
 
 
 
8d7c9f1
 
 
 
 
 
 
 
 
 
fe34583
 
 
 
 
 
 
 
 
 
8d7c9f1
 
 
fe34583
 
 
 
 
 
 
 
 
 
 
 
 
 
8d7c9f1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
fe34583
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
---
license: apache-2.0
language:
- en
pipeline_tag: text-generation
tags:
- qwen2.5
- qwen2.5-coder
- code-generation
- browser-automation
- web-agent
- tool-calling
- function-calling
- agent
- conversational
- lora
- peft
- adapter
- sft
- trl
- sakthai
- house-of-sak
- safetensors
base_model: Qwen/Qwen2.5-Coder-1.5B-Instruct
datasets:
- Nanthasit/sakthai-combined-v8
- Nanthasit/sakthai-combined-v11
- Nanthasit/sakthai-irrelevance-supplement
- Nanthasit/cycle-bench
inference:
  parameters:
    temperature: 0.3
    max_new_tokens: 256
    top_p: 0.9
model-index:
- name: sakthai-coder-browser-lora
  results:
  - task:
      type: tool-calling
    dataset:
      name: cycle-bench
      type: cycle-bench
    metrics:
    - name: tool-calling-accuracy
      type: tool-calling-accuracy
      value: 1.0
      verified: true
      date: 2026-08-01
---

<p align="center">
  <strong>LoRA adapter for browser-automation agent training — Qwen2.5-Coder-1.5B-Instruct</strong><br/>
  <em>Part of the <a href="https://huggingface.co/collections/Nanthasit/sakthai-model-family-6a64745450b12d421c1f9f02">SakThai Model Family</a></em>
</p>

<p align="center">
  <a href="https://huggingface.co/Nanthasit"><img src="https://img.shields.io/badge/%F0%9F%A4%97-Nanthasit-6644cc" alt="Profile"/></a>
  <a href="https://github.com/beer-sakthai"><img src="https://img.shields.io/badge/GitHub-beer--sakthai-181717?logo=github" alt="GitHub"/></a>
  <a href="https://huggingface.co/collections/Nanthasit/sakthai-model-family-6a64745450b12d421c1f9f02"><img src="https://img.shields.io/badge/%F0%9F%8F%A0-SakThai%20Family-6644cc" alt="Collection"/></a>
  <img src="https://img.shields.io/badge/license-Apache%202.0-green" alt="License"/>
  <img src="https://img.shields.io/badge/task-browser%20automation-ff6b6b" alt="Task"/>
  <img src="https://img.shields.io/badge/base-Qwen2.5--Coder--1.5B--Instruct-brightgreen" alt="Base Model"/>
  <a href="https://huggingface.co/Nanthasit/sakthai-coder-browser-lora"><img src="https://img.shields.io/badge/downloads-verifying-blue" alt="Downloads"/></a>
</p>

## Model Description

`sakthai-coder-browser-lora` is a **LoRA adapter** that teaches `Qwen/Qwen2.5-Coder-1.5B-Instruct` to act as a browser-automation agent. It is trained to emit structured tool calls for web navigation tasks, including click, scroll, search, extract, and form interaction. This repo does **not** include the base model weights; merge it onto the base model before inference.

## Models in this family

| Model | Type | Notes |
| --- | --- | --- |
| `sakthai-coder-browser` | Merged GGUF / Transformers | Production browser agent weights |
| `sakthai-coder-1.5b` | Base/finetuned | General code agent |
| `sakthai-context-1.5b-tools-v2` | Tools variant | Tool-calling focused sibling |
| `sakthai-context-0.5b-tools` | Compact tools | Small footprint tool agent |
| `sakthai-plus-1.5b-lora` | LoRA | Merger + code variant |

## Training Details

- **Base model:** Qwen/Qwen2.5-Coder-1.5B-Instruct
- **Adapter type:** LoRA
- **LoRA config:** r=16, alpha=32, dropout=0.05, rslora=true
- **Target modules:** q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
- **Datasets:** sakthai-combined-v8, sakthai-combined-v11, irrelevance-supplement, cycle-bench
- **Trainer:** TRL SFT
- **License:** apache-2.0

## Usage

### Merge with PEFT

```python
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = "Qwen/Qwen2.5-Coder-1.5B-Instruct"
adapter = "Nanthasit/sakthai-coder-browser-lora"

tokenizer = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(base, device_map="auto")
model = PeftModel.from_pretrained(model, adapter)
model = model.merge_and_unload()  # optional; or keep adapter separate for switching
```

### Inference with Ollama

```bash
ollama create sakthai-coder-browser-lora -f ./Modelfile
# Adapter runtime merge depends on backend support; prefer merged sibling for Ollama.
```

### Inference with llama.cpp GGUF

```bash
# Preferred zero-cost local inference:
ollama run nanthasit/sakthai-coder-browser-gguf
```

### Inference with Hugging Face InferenceClient

```python
from huggingface_hub import InferenceClient

client = InferenceClient(model="Nanthasit/sakthai-coder-browser")
out = client.chat_completion(
    messages=[{"role": "user", "content": "Extract all H2 headings from https://example.com"}],
    max_tokens=256,
    temperature=0.3,
)
print(out.choices[0].message.content)
```

## Reproducing Evaluation

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-Coder-1.5B-Instruct", device_map="auto")
model = PeftModel.from_pretrained(base, "Nanthasit/sakthai-coder-browser-lora")
model = model.merge_and_unload()
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-Coder-1.5B-Instruct")

prompt = "<tools>...</tools>\nUser: Search HuggingFace for DeepSeek V4 Flash"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
print(tokenizer.decode(model.generate(**inputs, max_new_tokens=256)[0], skip_special_tokens=True))
```

## Inference Tips

- Prefer merged weights (`sakthai-coder-browser`) for browser tasks.
- Use low temperature (0.1–0.3) to reduce hallucinated tool names.
- Always wrap function specs inside `<tools>` XML for reliable structured output.

## Limitations

- Adapter-only repo: **cannot benchmark standalone**; always merge onto the base model.
- Web task success depends on DOM complexity and instruction phrasing.
- Tool-calling accuracy drops on multi-step plans longer than 5 actions.
- CPU inference is usable but slow; prefer GPU/TGI or llama.cpp GGUF for production.

## Citation

```bibtex
@misc{sakthai-coder-browser-lora,
  title  = {SakThai Coder Browser LoRA},
  author = {Nanthasit},
  year   = {2026},
  url    = {https://huggingface.co/Nanthasit/sakthai-coder-browser-lora}
}
```