--- license: other library_name: pytorch tags: - math - code - pretraining - custom-architecture - vathos - pico language: - en pipeline_tag: text-generation --- # GUIDO-small 200M (bugfixed) A 195M-parameter math/code-oriented language model, pretrained from scratch on a math-heavy corpus. Custom architecture (Vathos PiCOFormer): gated attention, smear-gate, identity-at-init, UDLP MLP, RoPE, tied embeddings, fp32-master weights with bf16 autocast. > ⚠️ This is a **research checkpoint**. Architecture is custom (not HF-compatible) — requires the > bundled `vathos/` package + `modeling_guido.py` to load. ## Eval | | greedy | maj@16 (T=0.6, top_p=0.95, top_k=40) | pass@16 | |---|---|---|---| | **GSM8K** | 43.0% | **57.5%** | 77.5% | | **MATH-500** | 21.0% | **30.0%** | 53.75% | Numbers from `eval_distill.py` (greedy, N=200/100) and `eval_sc.py` (maj@16, N=80/80). ## Architecture - **Shape**: 24 layers × 768 d_model × 12 heads × 64 head_dim × 3072 d_ff - **Vocab**: 32,768 (Mathstral SPM tokenizer) - **Position**: RoPE base 1e6, max_len 8192 - **Norm**: RMSNorm, qk-norm - **Attention**: Multihead Gated (Modded-NanoGPT PR#117 style sparse output gate, gate_input_dim=12) - **MLP**: VariableUDLP with LeakyReLU² activation, identity-at-init - **Logit softcap**: 30.0 - **Tied embedding**: yes - **Total params**: 195,092,992 (~195M) ## Training - **Tokens seen**: 4.19B (16000 iter × 4 GPUs × 32 batch × 2048 seq) - **Optimizer**: Muon (Newton-Schulz5) on 2D matrices (lr=0.01) + Adam on embedding (lr=0.005) + scalars (lr=0.04) - **Schedule**: WSD (warmup 300 → plateau → linear warmdown last 60% to lr×1e-3) - **Precision**: fp32 master weights + bf16 autocast (cce kernel gets bf16 cast at boundary) - **Hardware**: 4× A100 64GB (Leonardo CINECA), ~3h wall ### Data mixture (corpus_v2) | weight | shard | tokens (B) | |---|---|---| | 0.34 | OpenMathInstruct-2 (14M docs) | 6.79 | | 0.14 | TinyGSM (full) | 2.70 | | 0.10 | OpenMathReasoning subset | 1.50 | | 0.07 | NuminaMath-1.5 | 0.41 | | 0.05 | NuminaMath-CoT | 0.48 | | 0.22 | FineWeb-Edu score≥3 (sample-10BT) | 3.00 | | 0.08 | Cosmopedia subset | 1.19 | Math 70% / NL 30%. Tokenizer: `mistralai/Mathstral-7B-v0.1`. Decontamination: canonical-hash vs MATH-500 and GSM8K test (note: not fuzzy-decon). ## How to load ### Quick start ```bash pip install -r requirements.txt ``` ```python from huggingface_hub import snapshot_download local_dir = snapshot_download(repo_id="Paerle/GUIDO_test_200M", token="") import sys, torch sys.path.insert(0, local_dir + "/vathos") sys.path.insert(0, local_dir) import modeling_guido as m from transformers import AutoTokenizer # Inferenza arch dal ckpt (così non serve config) state = torch.load(local_dir + "/pytorch_model.bin", map_location="cpu", weights_only=False) sd = state["model"] hp = m.HP() hp.head_dim = 64 hp.d_model = sd["backbone.embedder.embedding.weight"].shape[1] hp.n_heads = hp.d_model // hp.head_dim hp.d_ff = sd["backbone.blocks.0.channel_mixer.expand.weight"].shape[0] hp.n_layers = max(int(k.split("blocks.")[1].split(".")[0]) for k in sd if "blocks." in k) + 1 model = m.PiCOFormerLM(hp, vocab_size=sd["backbone.embedder.embedding.weight"].shape[0]) model.load_state_dict(sd, strict=False) model = model.cuda().bfloat16().eval() tok = AutoTokenizer.from_pretrained("mistralai/Mathstral-7B-v0.1") ``` See [`inference_example.py`](./inference_example.py) for a complete generation example (temperature, top-p, top-k sampling). ## Files | file | description | |---|---| | `pytorch_model.bin` | model weights (fp32 master, 745MB) | | `config.json` | architecture + HP + eval results | | `modeling_guido.py` | model definition (`PiCOFormerLM` + Muon optimizer + reader; only PiCOFormerLM/HP needed for inference) | | `inference_example.py` | minimal load + generate example | | `vathos/` | bundled Vathos package (custom backbone — see github.com/MarioPaerle/Aplos) | | `requirements.txt` | deps (torch, transformers, cut-cross-entropy) | ## Known limitations - **Decontamination**: canonical-hash only. ~7% of MATH-500 has near-duplicates in OpenMathInstruct-2 augmentation (audited). Real MATH-500 capability is ~21% greedy, not inflated by overt memorization. - **No instruction tuning**: this is a base pretrain. No SFT / RLHF. - **Domain**: heavily math-skewed. NL is 30% (FineWeb-Edu + Cosmopedia) — basic fluency but not chat-style. - **Tokenizer dependency**: requires `mistralai/Mathstral-7B-v0.1` tokenizer from HF. ## Hardware/precision notes - Inference works fine in **bf16** (just call `.bfloat16()`). Weights are stored fp32 — 745MB on disk but ~390MB GPU memory after `.bfloat16()`. - Custom CUDA kernels: uses `cut_cross_entropy` (Apple) for fused linear+CE in training. For inference you can replace with standard `F.linear + F.cross_entropy`. ## Architecture sketch ``` input_ids │ embedder (tied) ──────────────────────────┐ │ │ smear_gate (sparse on first 12 dims) ─────│ x0 │ │ [24× block] │ ├── spatial: gated_attn + RoPE + qk-norm │ ├── channel: VariableUDLP (LeakyReLU²) │ └── + x0_lambda × x0 ───────────────────┘ │ final_norm │ tied unembed → cce / softcap 30 → logits ``` ## License & citation Vathos backbone (custom by Mario Paerle): https://github.com/MarioPaerle/Aplos Model weights: research/private. Contact owner for usage terms.