openbmb/Ultra-FineWeb
Viewer • Updated • 1.29B • 74.6k • 449
Base (pretrain-only) checkpoint of Vertex 0.6 15M, a tiny ~15M-param model from the Vertex 0.6 family. Qwen3 architecture: hidden 256, 10 layers, 4 heads / 2 KV (GQA, head_dim 64), SwiGLU ffn 1024, 20000-vocab ByteLevel BPE, tied embeddings, ctx 2048.
Pretrained from scratch on 12B tokens of English web text: Ultra-FineWeb plus Ultra-FineWeb-L3 synthetic rewrites (Multi-Style + QA), ~800 tokens per parameter, on a single RTX 4060 Laptop (8GB).
This is a raw language model — no chat template, no instruction tuning. For chat, see Vertex-0.6-15M-Instruct.
Pretrained on English web text from openbmb/Ultra-FineWeb and openbmb/Ultra-FineWeb-L3 (Multi-Style + QA synthetic configs), 12B tokens total.