--- license: mit library_name: pytorch tags: - causal-lm - decoder-only - reasoning - pretraining - think - thought - reasoning - code --- # Mini.Deep.Thinker.11m A small decoder-only causal language model trained from scratch with deep chain-of-thought reasoning capabilities. ## Architecture - Parameters: 11,004,896 - Context length: 1096 - Hidden size: 352 - Layers: 7 - Attention heads: 8 - Intermediate size: 1152 - Vocabulary size: 4,096 ## Reasoning tokens - `<|input|>` - `<|think|>` - `<|thought|>` - `<|reasoning|>` - `<|answer|>` ## Objective True causal next-token prediction with strict per-example context limits. ## Context Every training example is strictly limited to 1095 content tokens + EOS, padded to exactly 1096 positions. No oversized example is split across separate training examples. ## Example selection Deterministic shuffled scan with SHA-256 content hash deduplication. Each session selects 20,000 new, unseen examples. ## Dataset `Plans11/Organized_PreTrain_1k_Context` ## Resumability Training state is persisted to Hugging Face, including: - model.safetensors - training_state.pt - tokenizer files (tokenizer.json, tokenizer_config.json) - config.json - progress.json - seen_examples.jsonl - training_metadata.json ## Current progress - Unique examples trained: 40,000 - Global optimizer steps: 1,250 - Last session: 20,000 examples - Best session loss: 1.88992 This is an experimental small language model and is not guaranteed to produce factually or logically correct outputs.