--- license: apache-2.0 base_model: unsloth/Qwen3-4B tags: - code - coding - full-stack - frontend - backend - agent - qwen3 - lora - unsloth - fine-tuned language: - en pipeline_tag: text-generation library_name: peft --- # NanoCoder V1 ๐Ÿง โšก A **4B parameter** full-stack coding assistant fine-tuned from **Qwen3-4B** using Unsloth + LoRA. Trained through a multi-phase pipeline with joint domain training and validation-driven checkpoint selection. Best checkpoint: **step 150** โ€” combined score **73.6%** across all skill domains. --- ## ๐Ÿ“Š Benchmarks | Benchmark | Score | Notes | |----------------------|-------------|---------------------------------------------| | **HumanEval pass@1** | **49.4%** | 164 problems, executed against test cases | | **LiveCodeBench** | **13.3%** | Execution eval on 30 problems (public tests)| | Frontend (custom) | 58.3% | React, Next.js, TypeScript, Tailwind, a11y | | Backend (custom) | 87.5% | FastAPI, Express, PostgreSQL, JWT, MongoDB | | Agent (custom) | 75.0% | Thought โ†’ Action โ†’ Patch โ†’ Reasoning format| | **Combined** | **73.6%** | Averaged across skill domains | --- ## ๐ŸŽฏ What it does well - **Backend** โ€” API design, auth (JWT/bcrypt), SQL/NoSQL, N+1 fixes, CORS - **Debugging agent** โ€” structured reasoning (`### Thought โ†’ ### Action โ†’ ### Patch โ†’ ### Reasoning`) - **Full-stack integration** โ€” connects frontend + backend flows - **Bug pattern recognition** โ€” race conditions, memory leaks, type errors ## โš ๏ธ Known limitations - Frontend scores lower than backend (weakest domain in v1) - Not a replacement for larger models (7B+) on hard competitive programming - English-only --- ## ๐Ÿš€ Usage ```python from unsloth import FastLanguageModel model, tokenizer = FastLanguageModel.from_pretrained( model_name="usernamebetter/nanocoder-v1", max_seq_length=2048, load_in_4bit=True, ) FastLanguageModel.for_inference(model) SYSTEM = "You are NanoCoder, an expert Senior Full-Stack Engineer and debugging agent." prompt = ( f"<|im_start|>system\n{SYSTEM}<|im_end|>\n" f"<|im_start|>user\nFix this React hydration error: useState(Date.now())<|im_end|>\n" f"<|im_start|>assistant\n" ) inputs = tokenizer(prompt, return_tensors="pt").to("cuda") outputs = model.generate(**inputs, max_new_tokens=300, do_sample=False) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)) ``` --- ## ๐Ÿ—๏ธ Training pipeline Multi-phase joint training from Qwen3-4B base: 1. **Phase 1** โ€” General coding (Magicoder-Evol-Instruct) 2. **Phase 2** โ€” Frontend specialization 3. **Phase 3** โ€” Fullstack (frontend + backend interleaved) 4. **Phase 4** โ€” Agent reasoning training 5. **Final** โ€” Joint retrain from base with all domains mixed (this checkpoint) ### Training configuration - **Base**: Qwen3-4B (4-bit quantized) - **LoRA**: r=32, alpha=32, dropout=0 - **LR**: 1e-5 with cosine scheduler - **Steps**: 400 (best checkpoint at step 150) - **Batch**: 2 ร— grad accum 4 = effective 8 - **Optimizer**: adamw_8bit - **Dataset**: ~13.7k samples interleaved - ๐Ÿค– Agent (synthetic + real): 40% - ๐ŸŽจ Frontend: 35% - โš™๏ธ Backend: 15% - ๐Ÿ› Bug fixing: 10% ### Data sources - ise-uiuc/Magicoder-Evol-Instruct-110K - sahil2801/CodeAlpaca-20k - nickrosh/Evol-Instruct-Code-80k-v1 - iamtarun/code_instructions_120k_alpaca - m-a-p/CodeFeedback-Filtered-Instruction - bigcode/self-oss-instruct-sc2-exec-filter-50k - HuggingFaceH4/CodeAlpaca_20K - TokenBender/code_instructions_122k_alpaca_style - Custom synthetic agent examples with structured reasoning format --- ## ๐Ÿงช Prompt format Uses Qwen chat template: ``` <|im_start|>system You are NanoCoder, an expert Senior Full-Stack Engineer and debugging agent. <|im_end|> <|im_start|>user {your question} <|im_end|> <|im_start|>assistant ``` For debugging tasks, the model responds in structured format: ``` ### Thought: {root cause analysis} ### Action: {what to do} ### Patch: {code fix} ### Reasoning: {why it works} ``` --- ## ๐Ÿ“… Roadmap - โœ… **v1**: Joint multi-domain training (this release) - ๐Ÿšง **v2**: Frontend boost + reasoning domain + label smoothing + cosine restarts - ๐Ÿšง **v3**: DPO alignment + tool calling - ๐Ÿšง **GGUF**: Q4_K_M / Q5_K_M / Q8_0 exports --- ## ๐Ÿ™ Credits - **Base model**: [Qwen/Qwen3-4B](https://huggingface.co/Qwen/Qwen3-4B) - **Fine-tuning framework**: [Unsloth](https://github.com/unslothai/unsloth) - **Training**: Kaggle T4 + Google Colab T4 --- ## ๐Ÿ“„ License Apache-2.0 (inherited from Qwen3-4B base).