--- license: apache-2.0 library_name: pytorch language: - en pipeline_tag: text-generation tags: - transformer - byte-level - causal-lm - attention - rope - decoder-only - fmsp - micro-language-model - sub-1m-parameters model_name: MicroT-test1-1M-UltraChat datasets: - HuggingFaceH4/ultrachat_200k - llaa33219/small-qa-en-10k metrics: - perplexity - exact-match ---
MicroMixer-4 Logo # MicroT-test1-1M-UltraChat Parameters Architecture FMSP

Micro Transformer β€” Reference Baseline
Vanilla Attention β€’ RoPE β€’ Byte-Level β€’ Decoder-Only
[![GitHub](https://img.shields.io/badge/GitHub-MicroMixer--4-blue?style=for-the-badge&logo=github&color=%23007BFF)](https://github.com/llaa33219/MicroMixer-4)
## πŸ“‹ Overview **MicroT-test1-1M-UltraChat** is a **996,736-parameter** vanilla decoder-only **transformer** β€” multi-head causal self-attention with RoPE β€” pretrained on **UltraChat 200k** conversation data (instead of the project's Discord-Dialogues baseline) and then fine-tuned with **FMSP** on 9,012 general-knowledge QA pairs. This is the **1M** member of the **MicroT-test1** family: the registered attention-based **reference baseline** of the MicroMixer-4 project, here rerun on UltraChat 200k as part of the **dataset-efficiency comparison study** β€” six parameter budgets Γ— two architectures Γ— two open pretraining corpora, all fine-tuned with the identical P05 FMSP recipe at seed 42. [Analysis](https://github.com/llaa33219/MicroMixer-4/blob/main/DATASET_COMPARISON_ANALYSIS.md). It is deliberately boring β€” the standard 2018–2020 transformer recipe parameterized down to the sub-1M regime: no flash attention, no SwiGLU, no ALiBi, no QKNorm, no sliding window, no MQA/GQA, no MoE.
## πŸ—οΈ Architecture
```mermaid graph TD A[Byte Input] --> B[Embed 256β†’128] B --> C[Transformer Block Γ— 5] C --> D[RMSNorm] D --> E[LM Head Tied with Embed] E --> F[Byte Output] subgraph "Transformer Block (pre-norm)" X[Input 128] --> N1[RMSNorm] N1 --> AT["MHA 8 heads Γ— d_head 16
RoPE ΞΈ=10000 on q,k Β· causal SDPA"] AT --> R1[+ residual] R1 --> N2[RMSNorm] N2 --> MLP["GELU MLP 128β†’496β†’128"] MLP --> R2[+ residual] end style A fill:#007BFF,color:#fff style F fill:#00D620,color:#fff style AT fill:#FF6600,color:#fff ```
### Model Configuration
Parameter Value
Hidden Dimension (d_model)128
Attention Heads8 (d_head = 16 at every size)
Number of Blocks5
FFN Hidden496
Position EncodingRoPE ΞΈ=10000 on q/k only (non-persistent buffers)
AttentionCausal MHA via F.scaled_dot_product_attention(is_causal=True)
ActivationGELU
BiasesNone β€” no bias parameters anywhere
NormalizationRMSNorm (pre-norm)
Max Sequence Length1024
Vocabulary Size256 (byte-level)
Output HeadTied with input embedding
### Core Components ``` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ Transformer Block (Γ—5) β”‚ β”‚ h = h + MHA(RMSNorm(h)) # RoPE q/k, causalβ”‚ β”‚ h = h + MLP(RMSNorm(h)) # GELU dβ†’ffnβ†’d β”‚ β”‚ no biases, no flash, no tricks β€” vanilla β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ ``` The d_head=16 contract is hard-asserted across all six sizes so that attention-head behavior is comparable at every budget and never confounds the memorization measurements. --- ## 🎯 Generation Examples
**Questions the model was trained on** (FMSP train set, 9,012 QA pairs β€” greedy decoding, `repetition_penalty=1.2`, `no_repeat_ngram_size=4`): ``` [Prompt] User: Who painted the Mona Lisa? Assistant: [Output] Leonardo da Vinci painted the Mona Lisa during the early 1500s in Italy. ``` correct β€” near-verbatim recitation of the trained fact ``` [Prompt] User: Who painted The Starry Night? Assistant: [Output] Vincent van Gogh painted The Starry Night in 1889 while staying at a mental asylum in Saint-RΓ©m… ``` correct β€” recites the trained answer to the truncation limit **Questions the model has never seen and cannot answer** (unanswerable probe β€” the correct behavior is to decline; the model's actual behavior is shown): ``` [Prompt] User: Who painted the Glimmering Frostberry? Assistant: [Output] Luxe Mergero's most broken the player who wrote the Han France language and was built to comple… ``` fabricates β€” plausible-sounding nonsense on a nonexistent subject ``` [Prompt] User: Who composed the Symphony of Hollow Dawn? Assistant: [Output] Ludwig was the symbol of the Colosseum around 18Ref. ``` fabricates β€” retrieves an unrelated trained fact instead of abstaining
--- ## πŸ“Š Results
### Pretraining (UltraChat 200k, V76 recipe, 3 epochs) | Metric | 1 ep | 2 ep | 3 ep | |--------|------|------|------| | Val PPL | 2.80 | 2.61 | **2.40** | Identical recipe to the MicroMixer-4 mixer: AdamW lr 3e-3 Β· WSD (warmup 500) Β· wd 0.01 Β· bs 16 Β· seq 1024 Β· seed 42. ### FMSP fine-tuning (small-qa-en-10k, P05 recipe, 10 epochs) | Metric | Value | |--------|-------| | Train QA pairs | 9,012 | | Held-out QA pairs | 988 | | Best-val checkpoint | `fmsp_epoch_9.safetensors` (val loss **0.0148**) | | freeze_fraction | 0.05 (true freeze) | | Loss | answer-only CE + probe KL (weight 0.5) | ### Evaluation battery (post-FMSP) | Axis | MicroT-test1-1M-UltraChat | |------|--------| | Chatter fluency d2 (cycles) | 0.698 (4/9) | | Full-988 EM (seed 42) | 749 | | Q-relevance echo / hijack % | 100.0 / 0.0 | | OOD hijack % | 39.0% | | Unanswerable fabrication /18 | 15 | | Discord PPL | 9.18 | Single-seed run (seed 42); the discord-pretrained cards report a 3-seed mean for Full-988 EM. ‑ where marked: degenerate-pass. ### MicroT-test1 UltraChat family (same protocol, all sizes) | Size | Params | 3ep Val PPL | Chatter d2 | Full-988 EM | qrel echo/hijack | OOD hijack | |------|--------|-------------|------------|-------------|------------------|------------| | **1M** | 996,736 | 2.40 | 0.698 | 749 | 100.0 / 0.0 | 39.0% | | 500K | 498,528 | 2.65 | 0.791 | 550 | 84.0 / 14.0 | 59.3% | | 300K | 297,680 | 2.84 | 0.633 | 211 | 62.0 / 24.0 | 42.4% | | 100K | 97,872 | 3.49 | 0.751 | 0 | 12.0 / 70.0 | 66.1% | | 50K | 49,888 | 3.98 | 0.528 | 0 | 4.0 / 13.0 | 1.7% | | 10K | 9,808 | 6.53 | 0.331 | 0 | 0.0 / 0.0 | 0.0% | Seed-42 single runs (pretrained on UltraChat 200k; the discord-pretrained families report 3-seed EM means).
--- ## πŸ“š Training Data
1. **Pretraining**: [UltraChat 200k](https://huggingface.co/datasets/HuggingFaceH4/ultrachat_200k) β€” 146K multi-turn conversations (`train_sft` split), flattened to `User:/Assistant:` format, 1024-byte sequences, 3 epochs. 2. **FMSP fine-tuning**: [small-qa-en-10k](https://huggingface.co/datasets/llaa33219/small-qa-en-10k) β€” 10K general-knowledge QA pairs (arts, science, history, geography, music…), split 9,012 train / 988 held-out. 10 epochs under the P05 recipe (5% of parameters frozen-true, answer-only CE, probe-KL 0.5).
--- ## πŸ”§ Usage ### Files in this repository - `fmsp_epoch_{0..9}.safetensors` β€” per-epoch FMSP weights (pickle-free safetensors). **`fmsp_epoch_9.safetensors` is the best-val checkpoint** for this size. ### Load and generate (local clone) ```python import torch from safetensors.torch import load_file from src.model_v88_transformer import MicroMixerV88Transformer, v88_transformer_1m from src.fmsp import attach_adapter from src.tokenizer import ByteTokenizer # Clone the code repository first: # git clone https://github.com/llaa33219/MicroMixer-4.git && cd MicroMixer-4 cfg = v88_transformer_1m() model = MicroMixerV88Transformer(cfg) attach_adapter(model, d_model=cfg.d_model, rank=16) # FMSP adapter (trained weights are in the file) model.load_state_dict(load_file("fmsp_epoch_9.safetensors"), strict=True) model.eval() tok = ByteTokenizer() prompt = "User: Who painted the Mona Lisa?\n\nAssistant: " ids = tok.encode(prompt) if ids and ids[-1] == tok.eos_token_id: ids = ids[:-1] # ByteTokenizer appends EOS; the prompt must end open ids = torch.tensor([ids]) with torch.no_grad(): out = model.generate( ids, max_new_tokens=200, temperature=0.0, # greedy β€” used for all reported numbers repetition_penalty=1.2, no_repeat_ngram_size=4, eos_token_id=tok.eos_token_id, ) print(tok.decode(out[0].tolist())) ``` ### Load from Hugging Face Hub (no clone of the weights needed) ```python import torch from huggingface_hub import hf_hub_download from safetensors.torch import load_file from src.model_v88_transformer import MicroMixerV88Transformer, v88_transformer_1m from src.fmsp import attach_adapter REPO = "llaa33219/MicroT-test1-1M-UltraChat" cfg = v88_transformer_1m() model = MicroMixerV88Transformer(cfg) attach_adapter(model, d_model=cfg.d_model, rank=16) model.load_state_dict( load_file(hf_hub_download(REPO, "fmsp_epoch_9.safetensors")), strict=True) model.eval() # ... generate as above ``` --- ## ⚠️ Limitations
| Limitation | Description | |------------|-------------| | **Micro parameters** | 996,736 parameters; capacity is the binding constraint on every axis | | **Reference baseline, not a product** | Exists to score the Mixer against attention at matched budget | | **Knows only what it memorized** | Knowledge is limited to the 9,012 trained QA pairs + the pretraining-corpus distribution | | **Does not abstain** | Unknown questions are answered with fabrication or degeneration, not refusal β€” see the examples above | | **Byte-level noise** | 256-vocab byte tokenizer; PPL not comparable to BPE baselines | | **Research use only** | Architecture/scaling research artifact, not a production model |
--- ## 🧬 Context This is the **1M** UltraChat-pretrained arm of the **dataset-efficiency comparison study** (24 arms: {V87, V88} Γ— {UltraChat, SmolTalk2} Γ— six sizes) in the [MicroMixer-4](https://github.com/llaa33219/MicroMixer-4) project. Each arm swaps the Discord-Dialogues pretraining baseline for an open corpus (UltraChat 200k) and reruns the identical P05 FMSP recipe at seed 42. Sibling repos: `llaa33219/MicroMixer-4-{1M..10K}-{UltraChat,SmolTalk2}` and `llaa33219/MicroT-test1-{1M..10K}-{UltraChat,SmolTalk2}`; the discord-pretrained baselines are `llaa33219/MicroMixer-4-{1M..10K}` and `llaa33219/MicroT-test1-{1M..10K}`. Full analysis: [DATASET_COMPARISON_ANALYSIS.md](https://github.com/llaa33219/MicroMixer-4/blob/main/DATASET_COMPARISON_ANALYSIS.md). ---
[![GitHub](https://img.shields.io/badge/Back_to_Repository-MicroMixer--4-blue?style=for-the-badge&logo=github&color=%23007BFF)](https://github.com/llaa33219/MicroMixer-4) Part of the MicroMixer-4 research project β€” V88 transformer reference (MicroT-test1), 1M preset, UltraChat pretraining