GUIDO-small 200M (bugfixed)
A 195M-parameter math/code-oriented language model, pretrained from scratch on a math-heavy corpus. Custom architecture (Vathos PiCOFormer): gated attention, smear-gate, identity-at-init, UDLP MLP, RoPE, tied embeddings, fp32-master weights with bf16 autocast.
β οΈ This is a research checkpoint. Architecture is custom (not HF-compatible) β requires the bundled
vathos/package +modeling_guido.pyto load.
Eval
| greedy | maj@16 (T=0.6, top_p=0.95, top_k=40) | pass@16 | |
|---|---|---|---|
| GSM8K | 43.0% | 57.5% | 77.5% |
| MATH-500 | 21.0% | 30.0% | 53.75% |
Numbers from eval_distill.py (greedy, N=200/100) and eval_sc.py (maj@16, N=80/80).
Architecture
- Shape: 24 layers Γ 768 d_model Γ 12 heads Γ 64 head_dim Γ 3072 d_ff
- Vocab: 32,768 (Mathstral SPM tokenizer)
- Position: RoPE base 1e6, max_len 8192
- Norm: RMSNorm, qk-norm
- Attention: Multihead Gated (Modded-NanoGPT PR#117 style sparse output gate, gate_input_dim=12)
- MLP: VariableUDLP with LeakyReLUΒ² activation, identity-at-init
- Logit softcap: 30.0
- Tied embedding: yes
- Total params: 195,092,992 (~195M)
Training
- Tokens seen: 4.19B (16000 iter Γ 4 GPUs Γ 32 batch Γ 2048 seq)
- Optimizer: Muon (Newton-Schulz5) on 2D matrices (lr=0.01) + Adam on embedding (lr=0.005) + scalars (lr=0.04)
- Schedule: WSD (warmup 300 β plateau β linear warmdown last 60% to lrΓ1e-3)
- Precision: fp32 master weights + bf16 autocast (cce kernel gets bf16 cast at boundary)
- Hardware: 4Γ A100 64GB (Leonardo CINECA), ~3h wall
Data mixture (corpus_v2)
| weight | shard | tokens (B) |
|---|---|---|
| 0.34 | OpenMathInstruct-2 (14M docs) | 6.79 |
| 0.14 | TinyGSM (full) | 2.70 |
| 0.10 | OpenMathReasoning subset | 1.50 |
| 0.07 | NuminaMath-1.5 | 0.41 |
| 0.05 | NuminaMath-CoT | 0.48 |
| 0.22 | FineWeb-Edu scoreβ₯3 (sample-10BT) | 3.00 |
| 0.08 | Cosmopedia subset | 1.19 |
Math 70% / NL 30%. Tokenizer: mistralai/Mathstral-7B-v0.1. Decontamination: canonical-hash vs MATH-500 and GSM8K test (note: not fuzzy-decon).
How to load
Quick start
pip install -r requirements.txt
from huggingface_hub import snapshot_download
local_dir = snapshot_download(repo_id="Paerle/GUIDO_test_200M", token="<your_token>")
import sys, torch
sys.path.insert(0, local_dir + "/vathos")
sys.path.insert(0, local_dir)
import modeling_guido as m
from transformers import AutoTokenizer
# Inferenza arch dal ckpt (così non serve config)
state = torch.load(local_dir + "/pytorch_model.bin", map_location="cpu", weights_only=False)
sd = state["model"]
hp = m.HP()
hp.head_dim = 64
hp.d_model = sd["backbone.embedder.embedding.weight"].shape[1]
hp.n_heads = hp.d_model // hp.head_dim
hp.d_ff = sd["backbone.blocks.0.channel_mixer.expand.weight"].shape[0]
hp.n_layers = max(int(k.split("blocks.")[1].split(".")[0]) for k in sd if "blocks." in k) + 1
model = m.PiCOFormerLM(hp, vocab_size=sd["backbone.embedder.embedding.weight"].shape[0])
model.load_state_dict(sd, strict=False)
model = model.cuda().bfloat16().eval()
tok = AutoTokenizer.from_pretrained("mistralai/Mathstral-7B-v0.1")
See inference_example.py for a complete generation example
(temperature, top-p, top-k sampling).
Files
| file | description |
|---|---|
pytorch_model.bin |
model weights (fp32 master, 745MB) |
config.json |
architecture + HP + eval results |
modeling_guido.py |
model definition (PiCOFormerLM + Muon optimizer + reader; only PiCOFormerLM/HP needed for inference) |
inference_example.py |
minimal load + generate example |
vathos/ |
bundled Vathos package (custom backbone β see github.com/MarioPaerle/Aplos) |
requirements.txt |
deps (torch, transformers, cut-cross-entropy) |
Known limitations
- Decontamination: canonical-hash only. ~7% of MATH-500 has near-duplicates in OpenMathInstruct-2 augmentation (audited). Real MATH-500 capability is ~21% greedy, not inflated by overt memorization.
- No instruction tuning: this is a base pretrain. No SFT / RLHF.
- Domain: heavily math-skewed. NL is 30% (FineWeb-Edu + Cosmopedia) β basic fluency but not chat-style.
- Tokenizer dependency: requires
mistralai/Mathstral-7B-v0.1tokenizer from HF.
Hardware/precision notes
- Inference works fine in bf16 (just call
.bfloat16()). Weights are stored fp32 β 745MB on disk but ~390MB GPU memory after.bfloat16(). - Custom CUDA kernels: uses
cut_cross_entropy(Apple) for fused linear+CE in training. For inference you can replace with standardF.linear + F.cross_entropy.
Architecture sketch
input_ids
β
embedder (tied) βββββββββββββββββββββββββββ
β β
smear_gate (sparse on first 12 dims) ββββββ x0
β β
[24Γ block] β
βββ spatial: gated_attn + RoPE + qk-norm β
βββ channel: VariableUDLP (LeakyReLUΒ²) β
βββ + x0_lambda Γ x0 ββββββββββββββββββββ
β
final_norm
β
tied unembed β cce / softcap 30 β logits
License & citation
Vathos backbone (custom by Mario Paerle): https://github.com/MarioPaerle/Aplos Model weights: research/private. Contact owner for usage terms.
- Downloads last month
- 15