File size: 2,580 Bytes
e22a73c
 
ab84885
ad34f6a
ab84885
 
8693ea8
 
 
 
ab84885
 
 
 
ad34f6a
 
e22a73c
ad34f6a
b050678
ad34f6a
8693ea8
 
 
ab84885
 
ad34f6a
8693ea8
ab84885
8693ea8
 
 
 
ab84885
 
 
8693ea8
 
 
4b370de
8693ea8
 
 
4b370de
8693ea8
 
4b370de
8693ea8
4b370de
 
 
8693ea8
ab84885
8693ea8
ab84885
 
 
 
 
b050678
ab84885
 
 
 
 
 
 
 
 
8693ea8
ab84885
8693ea8
 
 
 
 
 
ab84885
 
8693ea8
ad34f6a
ab84885
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
---
license: apache-2.0
base_model: mistralai/Ministral-3b-instruct
tags:
- ministral
- code
- coding
- qlora
- unsloth
- text-generation
pipeline_tag: text-generation
library_name: transformers
language:
- en
datasets:
- greghavens/fable-5-coding-and-debugging-traces
---

# Koa-AI v1 Code 3B

**Koa-AI-v1-Code-3B** is an ultra-lightweight, high-efficiency 3B parameter model fine-tuned for code generation, multi-step agentic planning, and debugging tasks. 

Built on top of `mistralai/Ministral-3b-instruct` using Unsloth and QLoRA, it is optimized to run blazingly fast on consumer hardware, local edge devices, and laptop GPUs without sacrificing code reasoning capabilities.

---

## ⚡ Highlights

* **Base Model:** `mistralai/Ministral-3b-instruct`
* **Dataset:** Fine-tuned on multi-turn debugging and agentic coding trace data (`greghavens/fable-5-coding-and-debugging-traces`).
* **Efficiency:** Lightweight 3B parameter size allows low-latency, real-time code completion in local IDE extensions (e.g., Continue, VS Code).
* **Extended Context Support:** Native context window up to 128,000 tokens (SFT fine-tuned at a 2,048 sequence length cap).

---

## 📊 Training Specifications

| Parameter | Value |
| :--- | :--- |
| **Architecture** | Ministral 3B Instruct |
| **Precision** | 4-bit NormalFloat (NF4) / BF16 mixed |
| **Fine-Tuning Method** | QLoRA 4-bit |
| **Target Modules** | `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj` |
| **LoRA Config** | $r = 16$, $\alpha = 32$, Dropout = $0.0$ |
| **Learning Rate** | `2e-4` |
| **Optimizer** | AdamW 8-bit |
| **Frameworks** | Unsloth, PyTorch, Hugging Face Transformers |

---

## 💻 Quickstart & Usage

### Running with `transformers` (Python)

```python
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

model_id = "vamazing/Koa-AI-v1-Code-3B"

tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.float16,
    device_map="auto",
    trust_remote_code=True
)

prompt = "Write a Python function to check if a number is prime and optimize it for speed."
messages = [{"role": "user", "content": prompt}]

formatted_prompt = tokenizer.apply_chat_template(
    messages, 
    tokenize=False, 
    add_generation_prompt=True
)

inputs = tokenizer(formatted_prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=512, do_sample=True, temperature=0.7)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))