File size: 11,347 Bytes
045e6de
 
ee4d1dd
d35ca36
045e6de
d35ca36
 
 
 
 
 
 
 
 
ee4d1dd
045e6de
ee4d1dd
 
d35ca36
 
 
 
 
 
 
045e6de
d35ca36
ee4d1dd
d35ca36
 
 
bf1fbaa
 
 
 
d35ca36
 
 
 
 
 
 
 
ee4d1dd
d35ca36
ee4d1dd
d35ca36
ee4d1dd
d35ca36
 
 
 
 
 
 
 
 
 
 
642897c
 
 
 
 
d35ca36
7aff98a
 
 
 
d35ca36
ee4d1dd
 
d35ca36
ee4d1dd
 
d35ca36
 
 
 
 
 
 
 
 
ee4d1dd
d35ca36
 
 
 
 
 
 
 
 
8fea270
d35ca36
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f8bde79
d35ca36
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
---
license: apache-2.0
language:
  - en
tags:
  - atlas-coder
  - qwen2.5-coder
  - sub-1b
  - qlora
  - code-generation
  - code-completion
  - code-debugging
  - fine-tuned
  - text-generation-inference
base_model: Qwen/Qwen2.5-Coder-0.5B
pipeline_tag: text-generation
library_name: transformers
datasets:
  - ise-uiuc/Magicoder-Evol-Instruct-110K
  - bigcode/self-oss-instruct-sc2-exec-filter-50k
  - m-a-p/CodeFeedback-Filtered-Instruction
  - BAAI/TACO
model-index:
  - name: Atlas-Coder-0.5B
    results: []
---

<div align="center">

# ⚑ Atlas-Coder-0.5B

<p align="center">
  <img src="./banner.png" width="100%">
</p>

**A top-tier sub-1B coding model trained from scratch on 80K decontaminated code instructions**

[![HuggingFace](https://img.shields.io/badge/πŸ€—-HuggingFace-yellow)](https://huggingface.co/Siddh07ETH/Atlas-Coder-0.5B)
[![License](https://img.shields.io/badge/License-Apache%202.0-blue)](https://opensource.org/licenses/Apache-2.0)
[![Model Size](https://img.shields.io/badge/Parameters-494M-green)]()
[![Base Model](https://img.shields.io/badge/Base-Qwen2.5--Coder--0.5B-orange)](https://huggingface.co/Qwen/Qwen2.5-Coder-0.5B)
[![GGUF](https://img.shields.io/badge/GGUF-Available-purple)](https://huggingface.co/Siddh07ETH/Atlas-Coder-0.5B-GGUF)

</div>

---

## Model Description

**Atlas-Coder-0.5B** is a coding-specialized language model instruction-tuned from scratch on top of [Qwen2.5-Coder-0.5B](https://huggingface.co/Qwen/Qwen2.5-Coder-0.5B) **base** (not instruct). Trained using **QLoRA** on a Tesla T4 GPU with a carefully engineered 80K sample mixture, it demonstrates that disciplined data curation and training design can push a sub-500M parameter model to near-instruct-level coding performance without any proprietary alignment pipeline.

This model is part of the **Pluto AI** research project by Siddharth N.R., following the [Pluto-Genesis-0.6B](https://huggingface.co/Siddh07ETH/Pluto-Genesis-0.6B) release, focusing on efficient fine-tuning of sub-1B language models on consumer-grade hardware.

> **Research Goal:** Prove that a sub-1B coding model fine-tuned on curated, decontaminated open-source data can match or exceed the coding performance of officially instruction-tuned variants of the same architecture β€” without RLHF, proprietary data, or large-scale compute.

---

## ⚠️ Benchmarks

> ⚠️ Note: This is an Infrastructure Case Study, not a SOTA Benchmark model.

<p align="center">
  <img src="./benchmark.png" width="100%">
</p>

**Why is the score low?**

> This V1 model was fine-tuned on the Qwen2.5-Coder-0.5B-Base model using ChatML format. Because base models natively lack RLHF stopping criteria, the model often continued generating text (hallucinating follow-up prompts) after writing the correct function. When EvalPlus attempted to execute the raw generation, Python threw SyntaxErrors due to the appended text, resulting in a low pass@1 score.

---

## Training Details

| Property | Value |
|----------|-------|
| **Base Model** | Qwen/Qwen2.5-Coder-0.5B (Base, not Instruct) |
| **Parameters** | ~494 Million |
| **Method** | QLoRA (4-bit NF4 + LoRA) |
| **LoRA Rank** | r=64, Ξ±=128 |
| **LoRA Dropout** | 0.05 |
| **LoRA Target Modules** | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| **modules_to_save** | embed_tokens, lm_head (fully unfrozen for base model adaptation) |
| **Trainable Parameters** | 35,192,832 / 350,312,320 (10.05%) |
| **Loss Masking** | Response-only (DataCollatorForCompletionOnlyLM) |
| **Training Epochs** | 3 |
| **Total Steps** | ~7,424 (resumed from checkpoint 3,250 β†’ completed at 3,713) |
| **Final Training Loss** | 0.0294 |
| **Precision** | FP16 (forced β€” T4 sm_75 does not support BF16) |
| **Optimizer** | AdamW 8-bit (Paged) |
| **Learning Rate** | 1e-4 (cosine schedule) |
| **Warmup Steps** | max(150, 5% of total steps) |
| **Effective Batch Size** | 32 (2 Γ— 16 grad accum) |
| **Sequence Length** | 1024 tokens |
| **Hardware** | Tesla T4 (16 GB VRAM) β€” Kaggle free tier |
| **Training Time** | 42h 44m 31s |
| **Framework** | Transformers 4.52.4 + PEFT 0.17.0 + TRL 0.19.1 |
| **Chat Template** | ChatML |

---

## Training Data

| Domain | Dataset | Samples | Purpose |
|--------|---------|---------|---------|
| 🧬 Synthetic Complexity | [Magicoder-Evol-Instruct-110K](https://huggingface.co/datasets/ise-uiuc/Magicoder-Evol-Instruct-110K) | 15,000 | Multi-step code synthesis |
| βœ… Exec-Verified OSS | [self-oss-instruct-sc2-exec-filter-50k](https://huggingface.co/datasets/bigcode/self-oss-instruct-sc2-exec-filter-50k) | 50,000 | Single-function completion (HumanEval+ aligned) |
| πŸ” Real-World Debug | [CodeFeedback-Filtered-Instruction](https://huggingface.co/datasets/m-a-p/CodeFeedback-Filtered-Instruction) | 10,000 | Stack Overflow Q&A, debugging |
| 🧠 Algorithms | [BAAI/TACO](https://huggingface.co/datasets/BAAI/TACO) | 5,000 | Algorithmic reasoning |
| **Total** | | **80,000 β†’ 79,988 after decontamination** | |

### Decontamination

All datasets were scanned using **n-gram Jaccard similarity** (8-gram, threshold 0.3) against the full HumanEval test set before training. This ensures benchmark scores reflect genuine generalization and not memorization.

| Dataset | Pre-decontam | Removed | Post-decontam |
|---------|-------------|---------|--------------|
| Magicoder | 15,000 | 7 | 14,993 |
| OSS-Instruct | 50,000 | 0 | 50,000 |
| CodeFeedback | 10,000 | 5 | 9,995 |
| TACO | 5,000 | 0 | 5,000 |
| **Total** | **80,000** | **12** | **79,988** |

---

## Key Engineering Decisions

**1. Response-Only Loss Masking**
Using `DataCollatorForCompletionOnlyLM` from TRL, loss is computed only on assistant response tokens. This prevents the model from wasting gradient steps learning to predict system prompts and user messages β€” the single highest-ROI change for HumanEval+ performance.

**2. Unfrozen Embeddings + Output Head**
`modules_to_save=["embed_tokens", "lm_head"]` trains the embedding and output projection layers as full FP32 copies alongside LoRA. Critical when fine-tuning from a base (not instruct) model β€” the token distribution needs to shift significantly to learn the ChatML instruction format.

**3. FP32 LoRA Cast on T4**
PEFT 0.17 initializes LoRA matrices in BF16 by default. Since the T4 (sm_75) cannot train in BF16 without silent NaN gradients, all trainable parameters are explicitly cast to FP32 after LoRA wrapping.

**4. Exec-Verified OSS Data as Primary Source**
50K of 80K samples (62.5%) come from `self-oss-instruct-sc2-exec-filter-50k` β€” execution-verified, single-function Python completions derived from real open-source code. This dataset's format directly mirrors HumanEval+ problem structure, making it the highest-ROI data source for benchmark performance.

**5. 3-Layer Checkpoint Recovery**
Training was designed to survive Kaggle's 12-hour session limit via a 3-layer resume system: local checkpoint scan β†’ HuggingFace Hub download β†’ fresh start. This run resumed from step 3,250 (downloaded from Hub) and completed training through step 3,713 in a single session.

---

## Usage

### Basic Inference

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model = AutoModelForCausalLM.from_pretrained(
    "Siddh07ETH/Atlas-Coder-0.5B",
    torch_dtype=torch.float16,
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("Siddh07ETH/Atlas-Coder-0.5B")

messages = [{"role": "user", "content": "Write a Python function to find the longest common subsequence of two strings."}]

text = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)

with torch.no_grad():
    output = model.generate(
        **inputs,
        max_new_tokens=512,
        temperature=0.3,
        do_sample=True,
        top_p=0.9,
        repetition_penalty=1.1,
    )

response = tokenizer.decode(
    output[0][inputs.input_ids.shape[1]:],
    skip_special_tokens=True
)
print(response)
```

### Low Memory Inference (4-bit)

```python
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
import torch

quant_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.float16,
)
model = AutoModelForCausalLM.from_pretrained(
    "Siddh07ETH/Atlas-Coder-0.5B",
    quantization_config=quant_config,
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("Siddh07ETH/Atlas-Coder-0.5B")
```

### GGUF (Ollama / LM Studio / llama.cpp)

GGUF quantizations for CPU inference are available at:

> **[Siddh07ETH/Atlas-Coder-0.5B-GGUF](https://huggingface.co/Siddh07ETH/Atlas-Coder-0.5B-GGUF)**

Runs at 40+ tokens/second on a laptop CPU using LM Studio, Ollama, or llama.cpp.

---

## Recommended Generation Settings

| Setting | Value | Reason |
|---------|-------|--------|
| `temperature` | 0.2–0.4 | Conservative β€” reduces hallucinations in code |
| `top_p` | 0.9 | Focused vocabulary sampling |
| `repetition_penalty` | 1.1 | Prevents repetitive patterns |
| `max_new_tokens` | 256–512 | Sufficient for most coding tasks |
| `do_sample` | `True` | Required when temperature < 1.0 |

---

## Limitations

- **Size:** At ~494M parameters this model will make mistakes on complex multi-file engineering tasks. Always review generated code before running it.
- **Context length:** Trained on sequences up to 1024 tokens. Performance may degrade on prompts requiring longer context.
- **Language bias:** Optimized primarily for Python. Performance on other languages varies.
- **Knowledge cutoff:** No access to real-time information or recently published libraries.
- **Research only:** Not intended for production deployment without further evaluation and safety testing.

---

## Comparison to Base Model

This model was fine-tuned from `Qwen2.5-Coder-0.5B` **base**, not instruct. The gap this training bridges:

| Model | HumanEval+ | Notes |
|-------|-----------|-------|
| Qwen2.5-Coder-0.5B Base | ~23.8% | Starting point before this fine-tune |
| **Atlas-Coder-0.5B** | **TBD β€” running** | This model |
| Qwen2.5-Coder-0.5B Instruct | ~57.3% | Alibaba's full alignment pipeline (ceiling benchmark) |

*Benchmark results will be published shortly via EvalPlus.*

---

## Related Models

| Model | Parameters | Description |
|-------|-----------|-------------|
| [Pluto-Genesis-0.6B](https://huggingface.co/Siddh07ETH/Pluto-Genesis-0.6B) | 596M | General reasoning, math, and code β€” Pluto AI's first release |
| Atlas-Coder-0.5B (this) | 494M | Coding-specialized, trained from base |

---

## Author

**Siddharth N.R. (Siddhu)**
Final-year B.Tech β€” AI & Data Science
Pluto AI Research

[![HuggingFace](https://img.shields.io/badge/πŸ€—-Siddh07ETH-yellow)](https://huggingface.co/Siddh07ETH)

---

## Citation

```bibtex
@misc{atlascoder2026,
  author    = {Siddharth N.R.},
  title     = {Atlas-Coder-0.5B: A QLoRA-Trained Sub-1B Coding Model from Base},
  year      = {2026},
  publisher = {HuggingFace},
  url       = {https://huggingface.co/Siddh07ETH/Atlas-Coder-0.5B}
}
```

---

## License

Apache 2.0 β€” see [LICENSE](https://opensource.org/licenses/Apache-2.0).
Base model [Qwen2.5-Coder-0.5B](https://huggingface.co/Qwen/Qwen2.5-Coder-0.5B) is also Apache 2.0.