| --- |
| license: other |
| library_name: pytorch |
| tags: |
| - math |
| - code |
| - pretraining |
| - custom-architecture |
| - vathos |
| - pico |
| language: |
| - en |
| pipeline_tag: text-generation |
| --- |
| |
| # GUIDO-small 200M (bugfixed) |
|
|
| A 195M-parameter math/code-oriented language model, pretrained from scratch on a math-heavy corpus. |
| Custom architecture (Vathos PiCOFormer): gated attention, smear-gate, identity-at-init, UDLP MLP, |
| RoPE, tied embeddings, fp32-master weights with bf16 autocast. |
|
|
| > β οΈ This is a **research checkpoint**. Architecture is custom (not HF-compatible) β requires the |
| > bundled `vathos/` package + `modeling_guido.py` to load. |
| |
| ## Eval |
| |
| | | greedy | maj@16 (T=0.6, top_p=0.95, top_k=40) | pass@16 | |
| |---|---|---|---| |
| | **GSM8K** | 43.0% | **57.5%** | 77.5% | |
| | **MATH-500** | 21.0% | **30.0%** | 53.75% | |
| |
| Numbers from `eval_distill.py` (greedy, N=200/100) and `eval_sc.py` (maj@16, N=80/80). |
| |
| ## Architecture |
| |
| - **Shape**: 24 layers Γ 768 d_model Γ 12 heads Γ 64 head_dim Γ 3072 d_ff |
| - **Vocab**: 32,768 (Mathstral SPM tokenizer) |
| - **Position**: RoPE base 1e6, max_len 8192 |
| - **Norm**: RMSNorm, qk-norm |
| - **Attention**: Multihead Gated (Modded-NanoGPT PR#117 style sparse output gate, gate_input_dim=12) |
| - **MLP**: VariableUDLP with LeakyReLUΒ² activation, identity-at-init |
| - **Logit softcap**: 30.0 |
| - **Tied embedding**: yes |
| - **Total params**: 195,092,992 (~195M) |
| |
| ## Training |
| |
| - **Tokens seen**: 4.19B (16000 iter Γ 4 GPUs Γ 32 batch Γ 2048 seq) |
| - **Optimizer**: Muon (Newton-Schulz5) on 2D matrices (lr=0.01) + Adam on embedding (lr=0.005) + scalars (lr=0.04) |
| - **Schedule**: WSD (warmup 300 β plateau β linear warmdown last 60% to lrΓ1e-3) |
| - **Precision**: fp32 master weights + bf16 autocast (cce kernel gets bf16 cast at boundary) |
| - **Hardware**: 4Γ A100 64GB (Leonardo CINECA), ~3h wall |
| |
| ### Data mixture (corpus_v2) |
|
|
| | weight | shard | tokens (B) | |
| |---|---|---| |
| | 0.34 | OpenMathInstruct-2 (14M docs) | 6.79 | |
| | 0.14 | TinyGSM (full) | 2.70 | |
| | 0.10 | OpenMathReasoning subset | 1.50 | |
| | 0.07 | NuminaMath-1.5 | 0.41 | |
| | 0.05 | NuminaMath-CoT | 0.48 | |
| | 0.22 | FineWeb-Edu scoreβ₯3 (sample-10BT) | 3.00 | |
| | 0.08 | Cosmopedia subset | 1.19 | |
|
|
| Math 70% / NL 30%. Tokenizer: `mistralai/Mathstral-7B-v0.1`. Decontamination: canonical-hash vs MATH-500 and GSM8K test (note: not fuzzy-decon). |
|
|
| ## How to load |
|
|
| ### Quick start |
|
|
| ```bash |
| pip install -r requirements.txt |
| ``` |
|
|
| ```python |
| from huggingface_hub import snapshot_download |
| local_dir = snapshot_download(repo_id="Paerle/GUIDO_test_200M", token="<your_token>") |
| |
| import sys, torch |
| sys.path.insert(0, local_dir + "/vathos") |
| sys.path.insert(0, local_dir) |
| import modeling_guido as m |
| from transformers import AutoTokenizer |
| |
| # Inferenza arch dal ckpt (così non serve config) |
| state = torch.load(local_dir + "/pytorch_model.bin", map_location="cpu", weights_only=False) |
| sd = state["model"] |
| hp = m.HP() |
| hp.head_dim = 64 |
| hp.d_model = sd["backbone.embedder.embedding.weight"].shape[1] |
| hp.n_heads = hp.d_model // hp.head_dim |
| hp.d_ff = sd["backbone.blocks.0.channel_mixer.expand.weight"].shape[0] |
| hp.n_layers = max(int(k.split("blocks.")[1].split(".")[0]) for k in sd if "blocks." in k) + 1 |
| model = m.PiCOFormerLM(hp, vocab_size=sd["backbone.embedder.embedding.weight"].shape[0]) |
| model.load_state_dict(sd, strict=False) |
| model = model.cuda().bfloat16().eval() |
| tok = AutoTokenizer.from_pretrained("mistralai/Mathstral-7B-v0.1") |
| ``` |
|
|
| See [`inference_example.py`](./inference_example.py) for a complete generation example |
| (temperature, top-p, top-k sampling). |
|
|
| ## Files |
|
|
| | file | description | |
| |---|---| |
| | `pytorch_model.bin` | model weights (fp32 master, 745MB) | |
| | `config.json` | architecture + HP + eval results | |
| | `modeling_guido.py` | model definition (`PiCOFormerLM` + Muon optimizer + reader; only PiCOFormerLM/HP needed for inference) | |
| | `inference_example.py` | minimal load + generate example | |
| | `vathos/` | bundled Vathos package (custom backbone β see github.com/MarioPaerle/Aplos) | |
| | `requirements.txt` | deps (torch, transformers, cut-cross-entropy) | |
|
|
| ## Known limitations |
|
|
| - **Decontamination**: canonical-hash only. ~7% of MATH-500 has near-duplicates in OpenMathInstruct-2 |
| augmentation (audited). Real MATH-500 capability is ~21% greedy, not inflated by overt memorization. |
| - **No instruction tuning**: this is a base pretrain. No SFT / RLHF. |
| - **Domain**: heavily math-skewed. NL is 30% (FineWeb-Edu + Cosmopedia) β basic fluency but not chat-style. |
| - **Tokenizer dependency**: requires `mistralai/Mathstral-7B-v0.1` tokenizer from HF. |
|
|
| ## Hardware/precision notes |
|
|
| - Inference works fine in **bf16** (just call `.bfloat16()`). Weights are stored fp32 β 745MB on disk |
| but ~390MB GPU memory after `.bfloat16()`. |
| - Custom CUDA kernels: uses `cut_cross_entropy` (Apple) for fused linear+CE in training. |
| For inference you can replace with standard `F.linear + F.cross_entropy`. |
|
|
| ## Architecture sketch |
|
|
| ``` |
| input_ids |
| β |
| embedder (tied) βββββββββββββββββββββββββββ |
| β β |
| smear_gate (sparse on first 12 dims) ββββββ x0 |
| β β |
| [24Γ block] β |
| βββ spatial: gated_attn + RoPE + qk-norm β |
| βββ channel: VariableUDLP (LeakyReLUΒ²) β |
| βββ + x0_lambda Γ x0 ββββββββββββββββββββ |
| β |
| final_norm |
| β |
| tied unembed β cce / softcap 30 β logits |
| ``` |
|
|
| ## License & citation |
|
|
| Vathos backbone (custom by Mario Paerle): https://github.com/MarioPaerle/Aplos |
| Model weights: research/private. Contact owner for usage terms. |
|
|