--- license: apache-2.0 library_name: pytorch language: - en pipeline_tag: text-generation tags: - transformer - byte-level - causal-lm - attention - rope - decoder-only - fmsp - micro-language-model - sub-1m-parameters model_name: MicroT-test1-300K-UltraChat datasets: - HuggingFaceH4/ultrachat_200k - llaa33219/small-qa-en-10k metrics: - perplexity - exact-match ---
MicroMixer-4 Logo # MicroT-test1-300K-UltraChat Parameters Architecture FMSP

Micro Transformer β€” Reference Baseline
Vanilla Attention β€’ RoPE β€’ Byte-Level β€’ Decoder-Only
[![GitHub](https://img.shields.io/badge/GitHub-MicroMixer--4-blue?style=for-the-badge&logo=github&color=%23007BFF)](https://github.com/llaa33219/MicroMixer-4)
## πŸ“‹ Overview **MicroT-test1-300K-UltraChat** is a **297,680-parameter** vanilla decoder-only **transformer** β€” multi-head causal self-attention with RoPE β€” pretrained on **UltraChat 200k** conversation data (instead of the project's Discord-Dialogues baseline) and then fine-tuned with **FMSP** on 9,012 general-knowledge QA pairs. This is the **300K** member of the **MicroT-test1** family: the registered attention-based **reference baseline** of the MicroMixer-4 project, here rerun on UltraChat 200k as part of the **dataset-efficiency comparison study** β€” six parameter budgets Γ— two architectures Γ— two open pretraining corpora, all fine-tuned with the identical P05 FMSP recipe at seed 42. [Analysis](https://github.com/llaa33219/MicroMixer-4/blob/main/DATASET_COMPARISON_ANALYSIS.md). It is deliberately boring β€” the standard 2018–2020 transformer recipe parameterized down to the sub-1M regime: no flash attention, no SwiGLU, no ALiBi, no QKNorm, no sliding window, no MQA/GQA, no MoE.
## πŸ—οΈ Architecture
```mermaid graph TD A[Byte Input] --> B[Embed 256β†’80] B --> C[Transformer Block Γ— 4] C --> D[RMSNorm] D --> E[LM Head Tied with Embed] E --> F[Byte Output] subgraph "Transformer Block (pre-norm)" X[Input 80] --> N1[RMSNorm] N1 --> AT["MHA 5 heads Γ— d_head 16
RoPE ΞΈ=10000 on q,k Β· causal SDPA"] AT --> R1[+ residual] R1 --> N2[RMSNorm] N2 --> MLP["GELU MLP 80β†’272β†’80"] MLP --> R2[+ residual] end style A fill:#007BFF,color:#fff style F fill:#00D620,color:#fff style AT fill:#FF6600,color:#fff ```
### Model Configuration
Parameter Value
Hidden Dimension (d_model)80
Attention Heads5 (d_head = 16 at every size)
Number of Blocks4
FFN Hidden272
Position EncodingRoPE ΞΈ=10000 on q/k only (non-persistent buffers)
AttentionCausal MHA via F.scaled_dot_product_attention(is_causal=True)
ActivationGELU
BiasesNone β€” no bias parameters anywhere
NormalizationRMSNorm (pre-norm)
Max Sequence Length1024
Vocabulary Size256 (byte-level)
Output HeadTied with input embedding
### Core Components ``` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ Transformer Block (Γ—4) β”‚ β”‚ h = h + MHA(RMSNorm(h)) # RoPE q/k, causalβ”‚ β”‚ h = h + MLP(RMSNorm(h)) # GELU dβ†’ffnβ†’d β”‚ β”‚ no biases, no flash, no tricks β€” vanilla β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ ``` The d_head=16 contract is hard-asserted across all six sizes so that attention-head behavior is comparable at every budget and never confounds the memorization measurements. --- ## 🎯 Generation Examples
**Questions the model was trained on** (FMSP train set, 9,012 QA pairs β€” greedy decoding, `repetition_penalty=1.2`, `no_repeat_ngram_size=4`): ``` [Prompt] User: Who painted the Mona Lisa? Assistant: [Output] Leonardo da Vinci painted the Mona Lisa during the early 1500s in Italy. ``` correct β€” near-verbatim recitation of the trained fact ``` [Prompt] User: Who painted The Starry Night? Assistant: [Output] George World Word is the Serror Domain, a credit constant character who indies ``` fabricates β€” gives a wrong answer to a trained question at this size **Questions the model has never seen and cannot answer** (unanswerable probe β€” the correct behavior is to decline; the model's actual behavior is shown): ``` [Prompt] User: Who painted the Glimmering Frostberry? Assistant: [Output] German is the first paint of the French Region for its alliming radio in the 180multter. ``` fabricates β€” plausible-sounding nonsense on a nonexistent subject ``` [Prompt] User: Who composed the Symphony of Hollow Dawn? Assistant: [Output] Johann Lage Versa Permal Lage forces deduced with a character Einstein serious North Person in … ``` fabricates β€” plausible-sounding nonsense on a nonexistent subject
--- ## πŸ“Š Results
### Pretraining (UltraChat 200k, V76 recipe, 3 epochs) | Metric | 1 ep | 2 ep | 3 ep | |--------|------|------|------| | Val PPL | 3.23 | 3.08 | **2.84** | Identical recipe to the MicroMixer-4 mixer: AdamW lr 3e-3 Β· WSD (warmup 500) Β· wd 0.01 Β· bs 16 Β· seq 1024 Β· seed 42. ### FMSP fine-tuning (small-qa-en-10k, P05 recipe, 10 epochs) | Metric | Value | |--------|-------| | Train QA pairs | 9,012 | | Held-out QA pairs | 988 | | Best-val checkpoint | `fmsp_epoch_9.safetensors` (val loss **0.0852**) | | freeze_fraction | 0.05 (true freeze) | | Loss | answer-only CE + probe KL (weight 0.5) | ### Evaluation battery (post-FMSP) | Axis | MicroT-test1-300K-UltraChat | |------|--------| | Chatter fluency d2 (cycles) | 0.633 (4/9) | | Full-988 EM (seed 42) | 211 | | Q-relevance echo / hijack % | 62.0 / 24.0 | | OOD hijack % | 42.4% | | Unanswerable fabrication /18 | 16 | | Discord PPL | 31.52 | Single-seed run (seed 42); the discord-pretrained cards report a 3-seed mean for Full-988 EM. ‑ where marked: degenerate-pass. ### MicroT-test1 UltraChat family (same protocol, all sizes) | Size | Params | 3ep Val PPL | Chatter d2 | Full-988 EM | qrel echo/hijack | OOD hijack | |------|--------|-------------|------------|-------------|------------------|------------| | 1M | 996,736 | 2.40 | 0.698 | 749 | 100.0 / 0.0 | 39.0% | | 500K | 498,528 | 2.65 | 0.791 | 550 | 84.0 / 14.0 | 59.3% | | **300K** | 297,680 | 2.84 | 0.633 | 211 | 62.0 / 24.0 | 42.4% | | 100K | 97,872 | 3.49 | 0.751 | 0 | 12.0 / 70.0 | 66.1% | | 50K | 49,888 | 3.98 | 0.528 | 0 | 4.0 / 13.0 | 1.7% | | 10K | 9,808 | 6.53 | 0.331 | 0 | 0.0 / 0.0 | 0.0% | Seed-42 single runs (pretrained on UltraChat 200k; the discord-pretrained families report 3-seed EM means).
--- ## πŸ“š Training Data
1. **Pretraining**: [UltraChat 200k](https://huggingface.co/datasets/HuggingFaceH4/ultrachat_200k) β€” 146K multi-turn conversations (`train_sft` split), flattened to `User:/Assistant:` format, 1024-byte sequences, 3 epochs. 2. **FMSP fine-tuning**: [small-qa-en-10k](https://huggingface.co/datasets/llaa33219/small-qa-en-10k) β€” 10K general-knowledge QA pairs (arts, science, history, geography, music…), split 9,012 train / 988 held-out. 10 epochs under the P05 recipe (5% of parameters frozen-true, answer-only CE, probe-KL 0.5).
--- ## πŸ”§ Usage ### Files in this repository - `fmsp_epoch_{0..9}.safetensors` β€” per-epoch FMSP weights (pickle-free safetensors). **`fmsp_epoch_9.safetensors` is the best-val checkpoint** for this size. ### Load and generate (local clone) ```python import torch from safetensors.torch import load_file from src.model_v88_transformer import MicroMixerV88Transformer, v88_transformer_300k from src.fmsp import attach_adapter from src.tokenizer import ByteTokenizer # Clone the code repository first: # git clone https://github.com/llaa33219/MicroMixer-4.git && cd MicroMixer-4 cfg = v88_transformer_300k() model = MicroMixerV88Transformer(cfg) attach_adapter(model, d_model=cfg.d_model, rank=16) # FMSP adapter (trained weights are in the file) model.load_state_dict(load_file("fmsp_epoch_9.safetensors"), strict=True) model.eval() tok = ByteTokenizer() prompt = "User: Who painted the Mona Lisa?\n\nAssistant: " ids = tok.encode(prompt) if ids and ids[-1] == tok.eos_token_id: ids = ids[:-1] # ByteTokenizer appends EOS; the prompt must end open ids = torch.tensor([ids]) with torch.no_grad(): out = model.generate( ids, max_new_tokens=200, temperature=0.0, # greedy β€” used for all reported numbers repetition_penalty=1.2, no_repeat_ngram_size=4, eos_token_id=tok.eos_token_id, ) print(tok.decode(out[0].tolist())) ``` ### Load from Hugging Face Hub (no clone of the weights needed) ```python import torch from huggingface_hub import hf_hub_download from safetensors.torch import load_file from src.model_v88_transformer import MicroMixerV88Transformer, v88_transformer_300k from src.fmsp import attach_adapter REPO = "llaa33219/MicroT-test1-300K-UltraChat" cfg = v88_transformer_300k() model = MicroMixerV88Transformer(cfg) attach_adapter(model, d_model=cfg.d_model, rank=16) model.load_state_dict( load_file(hf_hub_download(REPO, "fmsp_epoch_9.safetensors")), strict=True) model.eval() # ... generate as above ``` --- ## ⚠️ Limitations
| Limitation | Description | |------------|-------------| | **Micro parameters** | 297,680 parameters; capacity is the binding constraint on every axis | | **Reference baseline, not a product** | Exists to score the Mixer against attention at matched budget | | **Knows only what it memorized** | Knowledge is limited to the 9,012 trained QA pairs + the pretraining-corpus distribution | | **Does not abstain** | Unknown questions are answered with fabrication or degeneration, not refusal β€” see the examples above | | **Byte-level noise** | 256-vocab byte tokenizer; PPL not comparable to BPE baselines | | **Research use only** | Architecture/scaling research artifact, not a production model |
--- ## 🧬 Context This is the **300K** UltraChat-pretrained arm of the **dataset-efficiency comparison study** (24 arms: {V87, V88} Γ— {UltraChat, SmolTalk2} Γ— six sizes) in the [MicroMixer-4](https://github.com/llaa33219/MicroMixer-4) project. Each arm swaps the Discord-Dialogues pretraining baseline for an open corpus (UltraChat 200k) and reruns the identical P05 FMSP recipe at seed 42. Sibling repos: `llaa33219/MicroMixer-4-{1M..10K}-{UltraChat,SmolTalk2}` and `llaa33219/MicroT-test1-{1M..10K}-{UltraChat,SmolTalk2}`; the discord-pretrained baselines are `llaa33219/MicroMixer-4-{1M..10K}` and `llaa33219/MicroT-test1-{1M..10K}`. Full analysis: [DATASET_COMPARISON_ANALYSIS.md](https://github.com/llaa33219/MicroMixer-4/blob/main/DATASET_COMPARISON_ANALYSIS.md). ---
[![GitHub](https://img.shields.io/badge/Back_to_Repository-MicroMixer--4-blue?style=for-the-badge&logo=github&color=%23007BFF)](https://github.com/llaa33219/MicroMixer-4) Part of the MicroMixer-4 research project β€” V88 transformer reference (MicroT-test1), 300K preset, UltraChat pretraining