File size: 1,397 Bytes
ee9c8c2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
---
language:
- code
library_name: transformers
pipeline_tag: text-generation
tags:
- html
- css
- javascript
- code
- llama
- from-scratch
---

# WebCoder-100M

A small decoder-only model specialized in HTML, CSS and JavaScript.

## Architecture
- Parameters: **99,894,528**
- Layers: 11
- Hidden size: 768
- Attention heads: 12
- Vocabulary: 28,672
- Max context: 2,048
- Training sequence length: 1,024

## Training
The model was initialized from scratch.

1. Causal pre-training on HTML/CSS/JavaScript.
2. Instruction fine-tuning on web-development instruction/code pairs.

## Token accounting
- Total processed: **582,209,140**
- Pre-training: **568,246,272**
- SFT processed: **13,962,868**
- SFT supervised response tokens: **9,124,446**
- Global cap: **2,000,000,000**

## Data
Pre-training: `bigcode/the-stack-smol-xl`, HTML/JavaScript/CSS subsets.

Instruction tuning: `iamtarun/code_instructions_120k_alpaca`, filtered for web-development examples.

Review upstream dataset cards and source licenses before commercial use.

## Prompt format
```text
<|system|>
You are WebCoder...<|end|>
<|user|>
Create a responsive landing page...<|end|>
<|assistant|>
...
```

## Limitations
This is a roughly 100M-parameter model trained from scratch. Its quality depends
strongly on how many tokens were actually processed. Generated code can contain
bugs or security issues and should be reviewed.