GUIDO-small 200M (bugfixed)

A 195M-parameter math/code-oriented language model, pretrained from scratch on a math-heavy corpus. Custom architecture (Vathos PiCOFormer): gated attention, smear-gate, identity-at-init, UDLP MLP, RoPE, tied embeddings, fp32-master weights with bf16 autocast.

⚠️ This is a research checkpoint. Architecture is custom (not HF-compatible) β€” requires the bundled vathos/ package + modeling_guido.py to load.

Eval

greedy maj@16 (T=0.6, top_p=0.95, top_k=40) pass@16
GSM8K 43.0% 57.5% 77.5%
MATH-500 21.0% 30.0% 53.75%

Numbers from eval_distill.py (greedy, N=200/100) and eval_sc.py (maj@16, N=80/80).

Architecture

  • Shape: 24 layers Γ— 768 d_model Γ— 12 heads Γ— 64 head_dim Γ— 3072 d_ff
  • Vocab: 32,768 (Mathstral SPM tokenizer)
  • Position: RoPE base 1e6, max_len 8192
  • Norm: RMSNorm, qk-norm
  • Attention: Multihead Gated (Modded-NanoGPT PR#117 style sparse output gate, gate_input_dim=12)
  • MLP: VariableUDLP with LeakyReLUΒ² activation, identity-at-init
  • Logit softcap: 30.0
  • Tied embedding: yes
  • Total params: 195,092,992 (~195M)

Training

  • Tokens seen: 4.19B (16000 iter Γ— 4 GPUs Γ— 32 batch Γ— 2048 seq)
  • Optimizer: Muon (Newton-Schulz5) on 2D matrices (lr=0.01) + Adam on embedding (lr=0.005) + scalars (lr=0.04)
  • Schedule: WSD (warmup 300 β†’ plateau β†’ linear warmdown last 60% to lrΓ—1e-3)
  • Precision: fp32 master weights + bf16 autocast (cce kernel gets bf16 cast at boundary)
  • Hardware: 4Γ— A100 64GB (Leonardo CINECA), ~3h wall

Data mixture (corpus_v2)

weight shard tokens (B)
0.34 OpenMathInstruct-2 (14M docs) 6.79
0.14 TinyGSM (full) 2.70
0.10 OpenMathReasoning subset 1.50
0.07 NuminaMath-1.5 0.41
0.05 NuminaMath-CoT 0.48
0.22 FineWeb-Edu scoreβ‰₯3 (sample-10BT) 3.00
0.08 Cosmopedia subset 1.19

Math 70% / NL 30%. Tokenizer: mistralai/Mathstral-7B-v0.1. Decontamination: canonical-hash vs MATH-500 and GSM8K test (note: not fuzzy-decon).

How to load

Quick start

pip install -r requirements.txt
from huggingface_hub import snapshot_download
local_dir = snapshot_download(repo_id="Paerle/GUIDO_test_200M", token="<your_token>")

import sys, torch
sys.path.insert(0, local_dir + "/vathos")
sys.path.insert(0, local_dir)
import modeling_guido as m
from transformers import AutoTokenizer

# Inferenza arch dal ckpt (così non serve config)
state = torch.load(local_dir + "/pytorch_model.bin", map_location="cpu", weights_only=False)
sd = state["model"]
hp = m.HP()
hp.head_dim = 64
hp.d_model = sd["backbone.embedder.embedding.weight"].shape[1]
hp.n_heads = hp.d_model // hp.head_dim
hp.d_ff = sd["backbone.blocks.0.channel_mixer.expand.weight"].shape[0]
hp.n_layers = max(int(k.split("blocks.")[1].split(".")[0]) for k in sd if "blocks." in k) + 1
model = m.PiCOFormerLM(hp, vocab_size=sd["backbone.embedder.embedding.weight"].shape[0])
model.load_state_dict(sd, strict=False)
model = model.cuda().bfloat16().eval()
tok = AutoTokenizer.from_pretrained("mistralai/Mathstral-7B-v0.1")

See inference_example.py for a complete generation example (temperature, top-p, top-k sampling).

Files

file description
pytorch_model.bin model weights (fp32 master, 745MB)
config.json architecture + HP + eval results
modeling_guido.py model definition (PiCOFormerLM + Muon optimizer + reader; only PiCOFormerLM/HP needed for inference)
inference_example.py minimal load + generate example
vathos/ bundled Vathos package (custom backbone β€” see github.com/MarioPaerle/Aplos)
requirements.txt deps (torch, transformers, cut-cross-entropy)

Known limitations

  • Decontamination: canonical-hash only. ~7% of MATH-500 has near-duplicates in OpenMathInstruct-2 augmentation (audited). Real MATH-500 capability is ~21% greedy, not inflated by overt memorization.
  • No instruction tuning: this is a base pretrain. No SFT / RLHF.
  • Domain: heavily math-skewed. NL is 30% (FineWeb-Edu + Cosmopedia) β€” basic fluency but not chat-style.
  • Tokenizer dependency: requires mistralai/Mathstral-7B-v0.1 tokenizer from HF.

Hardware/precision notes

  • Inference works fine in bf16 (just call .bfloat16()). Weights are stored fp32 β€” 745MB on disk but ~390MB GPU memory after .bfloat16().
  • Custom CUDA kernels: uses cut_cross_entropy (Apple) for fused linear+CE in training. For inference you can replace with standard F.linear + F.cross_entropy.

Architecture sketch

    input_ids
        β”‚
   embedder (tied) ──────────────────────────┐
        β”‚                                    β”‚
   smear_gate (sparse on first 12 dims) ─────│ x0
        β”‚                                    β”‚
   [24Γ— block]                                β”‚
    β”œβ”€β”€ spatial: gated_attn + RoPE + qk-norm β”‚
    β”œβ”€β”€ channel: VariableUDLP (LeakyReLUΒ²)    β”‚
    └── + x0_lambda Γ— x0  β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
        β”‚
   final_norm
        β”‚
   tied unembed β†’ cce / softcap 30 β†’ logits

License & citation

Vathos backbone (custom by Mario Paerle): https://github.com/MarioPaerle/Aplos Model weights: research/private. Contact owner for usage terms.

Downloads last month
15
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support