--- language: - code library_name: transformers pipeline_tag: text-generation tags: - html - css - javascript - code - llama - from-scratch --- # WebCoder-100M A small decoder-only model specialized in HTML, CSS and JavaScript. ## Architecture - Parameters: **99,894,528** - Layers: 11 - Hidden size: 768 - Attention heads: 12 - Vocabulary: 28,672 - Max context: 2,048 - Training sequence length: 1,024 ## Training The model was initialized from scratch. 1. Causal pre-training on HTML/CSS/JavaScript. 2. Instruction fine-tuning on web-development instruction/code pairs. ## Token accounting - Total processed: **582,209,140** - Pre-training: **568,246,272** - SFT processed: **13,962,868** - SFT supervised response tokens: **9,124,446** - Global cap: **2,000,000,000** ## Data Pre-training: `bigcode/the-stack-smol-xl`, HTML/JavaScript/CSS subsets. Instruction tuning: `iamtarun/code_instructions_120k_alpaca`, filtered for web-development examples. Review upstream dataset cards and source licenses before commercial use. ## Prompt format ```text <|system|> You are WebCoder...<|end|> <|user|> Create a responsive landing page...<|end|> <|assistant|> ... ``` ## Limitations This is a roughly 100M-parameter model trained from scratch. Its quality depends strongly on how many tokens were actually processed. Generated code can contain bugs or security issues and should be reviewed.