nanocoder-v1 / README.md
usernamebetter's picture
Update model card with full benchmarks + training details
f078bfc verified
|
Raw
History Blame Contribute Delete
4.73 kB
---
license: apache-2.0
base_model: unsloth/Qwen3-4B
tags:
- code
- coding
- full-stack
- frontend
- backend
- agent
- qwen3
- lora
- unsloth
- fine-tuned
language:
- en
pipeline_tag: text-generation
library_name: peft
---
# NanoCoder V1 🧠⚑
A **4B parameter** full-stack coding assistant fine-tuned from **Qwen3-4B** using Unsloth + LoRA.
Trained through a multi-phase pipeline with joint domain training and validation-driven checkpoint selection.
Best checkpoint: **step 150** β€” combined score **73.6%** across all skill domains.
---
## πŸ“Š Benchmarks
| Benchmark | Score | Notes |
|----------------------|-------------|---------------------------------------------|
| **HumanEval pass@1** | **49.4%** | 164 problems, executed against test cases |
| **LiveCodeBench** | **13.3%** | Execution eval on 30 problems (public tests)|
| Frontend (custom) | 58.3% | React, Next.js, TypeScript, Tailwind, a11y |
| Backend (custom) | 87.5% | FastAPI, Express, PostgreSQL, JWT, MongoDB |
| Agent (custom) | 75.0% | Thought β†’ Action β†’ Patch β†’ Reasoning format|
| **Combined** | **73.6%** | Averaged across skill domains |
---
## 🎯 What it does well
- **Backend** β€” API design, auth (JWT/bcrypt), SQL/NoSQL, N+1 fixes, CORS
- **Debugging agent** β€” structured reasoning (`### Thought β†’ ### Action β†’ ### Patch β†’ ### Reasoning`)
- **Full-stack integration** β€” connects frontend + backend flows
- **Bug pattern recognition** β€” race conditions, memory leaks, type errors
## ⚠️ Known limitations
- Frontend scores lower than backend (weakest domain in v1)
- Not a replacement for larger models (7B+) on hard competitive programming
- English-only
---
## πŸš€ Usage
```python
from unsloth import FastLanguageModel
model, tokenizer = FastLanguageModel.from_pretrained(
model_name="usernamebetter/nanocoder-v1",
max_seq_length=2048,
load_in_4bit=True,
)
FastLanguageModel.for_inference(model)
SYSTEM = "You are NanoCoder, an expert Senior Full-Stack Engineer and debugging agent."
prompt = (
f"<|im_start|>system\n{SYSTEM}<|im_end|>\n"
f"<|im_start|>user\nFix this React hydration error: useState(Date.now())<|im_end|>\n"
f"<|im_start|>assistant\n"
)
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=300, do_sample=False)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
```
---
## πŸ—οΈ Training pipeline
Multi-phase joint training from Qwen3-4B base:
1. **Phase 1** β€” General coding (Magicoder-Evol-Instruct)
2. **Phase 2** β€” Frontend specialization
3. **Phase 3** β€” Fullstack (frontend + backend interleaved)
4. **Phase 4** β€” Agent reasoning training
5. **Final** β€” Joint retrain from base with all domains mixed (this checkpoint)
### Training configuration
- **Base**: Qwen3-4B (4-bit quantized)
- **LoRA**: r=32, alpha=32, dropout=0
- **LR**: 1e-5 with cosine scheduler
- **Steps**: 400 (best checkpoint at step 150)
- **Batch**: 2 Γ— grad accum 4 = effective 8
- **Optimizer**: adamw_8bit
- **Dataset**: ~13.7k samples interleaved
- πŸ€– Agent (synthetic + real): 40%
- 🎨 Frontend: 35%
- βš™οΈ Backend: 15%
- πŸ› Bug fixing: 10%
### Data sources
- ise-uiuc/Magicoder-Evol-Instruct-110K
- sahil2801/CodeAlpaca-20k
- nickrosh/Evol-Instruct-Code-80k-v1
- iamtarun/code_instructions_120k_alpaca
- m-a-p/CodeFeedback-Filtered-Instruction
- bigcode/self-oss-instruct-sc2-exec-filter-50k
- HuggingFaceH4/CodeAlpaca_20K
- TokenBender/code_instructions_122k_alpaca_style
- Custom synthetic agent examples with structured reasoning format
---
## πŸ§ͺ Prompt format
Uses Qwen chat template:
```
<|im_start|>system
You are NanoCoder, an expert Senior Full-Stack Engineer and debugging agent.
<|im_end|>
<|im_start|>user
{your question}
<|im_end|>
<|im_start|>assistant
```
For debugging tasks, the model responds in structured format:
```
### Thought:
{root cause analysis}
### Action:
{what to do}
### Patch:
{code fix}
### Reasoning:
{why it works}
```
---
## πŸ“… Roadmap
- βœ… **v1**: Joint multi-domain training (this release)
- 🚧 **v2**: Frontend boost + reasoning domain + label smoothing + cosine restarts
- 🚧 **v3**: DPO alignment + tool calling
- 🚧 **GGUF**: Q4_K_M / Q5_K_M / Q8_0 exports
---
## πŸ™ Credits
- **Base model**: [Qwen/Qwen3-4B](https://huggingface.co/Qwen/Qwen3-4B)
- **Fine-tuning framework**: [Unsloth](https://github.com/unslothai/unsloth)
- **Training**: Kaggle T4 + Google Colab T4
---
## πŸ“„ License
Apache-2.0 (inherited from Qwen3-4B base).