ποΈ Govind222/68m-base
Govind222/68m-base is a 68.96M parameter Dense Foundation / Base Small Language Model (SLM) trained completely from scratch using transformer-toolkit on 4.0 Billion tokens of FineWeb (sample-10BT).
This checkpoint is the raw pretrained foundation base model. It serves as a data-efficient foundation for downstream task adaptation, instruction fine-tuning, domain specialization, or edge device deployment.
- SFT Fine-Tuned Version (Chat / Dialog):
Govind222/68m-base-guppy-sft - Training Framework: transformer-toolkit on GitHub
- Interactive Live Demo (SFT): https://fish-lm.runlog.in/
β‘ Key Highlights
- Pretrained from Scratch: Trained on 4 Billion tokens of clean English web text from Hugging Face's
FineWeb (sample-10BT). - Pure Dense Transformer: 100% dense computation with zero MoE, maximizing per-parameter compute density and caching efficiency on constrained hardware.
- Modern Architectural Recipe:
- Grouped-Query Attention (GQA): 8 query heads, 2 key/value heads (4:1 ratio) for fast KV-cache memory reuse.
- SwiGLU FFN: Non-linear Swish-Gated Linear Units with 1,536 hidden dimension ($3 \times d_{\text{model}}$).
- RoPE: Rotary Position Embeddings for robust position modeling.
- RMSNorm: Pre-layer normalization with $\epsilon = 1\text{e-}6$.
- Data Efficiency: Competitive 0-shot benchmark accuracy against models trained on 10Γ to 50Γ more parameters/compute.
- Compact Footprint: Under ~280 MB RAM footprint at runtime, ideal for edge Android/iOS smartphones and Raspberry Pi.
π Zero-Shot Benchmark Evaluation
Evaluated zero-shot using lm-evaluation-harness directly on the base pretrained checkpoint:
| Benchmark | Task / Metric | Govind222/68m-base | Pythia 70M | Pythia 160M | GPT-Neo 125M | OPT 125M |
|---|---|---|---|---|---|---|
| PIQA | Physical Interaction QA (acc) | 58.1% | 58.6% | 61.4% | 62.3% | 61.8% |
| Winogrande | Commonsense Reasoning (acc) | 48.7% | 51.6% | 50.4% | 49.5% | 51.5% |
| ARC-Easy | Science QA Reasoning (acc_norm) | 31.4% | 34.2% | β | 38.9% | 39.8% |
| Average | Mean Accuracy across 4 tasks | 39.6% | 41.4% | 43.7% | 43.5% | 43.9% |
βοΈ Model Architecture Specifications
| Hyperparameter | Value | Description |
|---|---|---|
| Total Parameters | 68,957,056 (~68.96M) |
Exact non-embedding + embedding tensor weight sum |
| Model Type | Dense Decoder-Only SLM | Pure dense transformer (no MoE) |
Layers (n_layers) |
12 |
Transformer blocks |
Hidden Dimension (dim) |
512 |
Model width ($d_{\text{model}}$) |
Query Attention Heads (n_heads) |
8 |
Head dimension = 64 |
KV Attention Heads (n_kv_heads) |
2 |
4:1 Grouped Query Attention (GQA) |
Feed-Forward Dimension (hidden_dim) |
1536 |
$3 \times d_{\text{model}}$ |
| FFN Activation | swiglu |
Dense SwiGLU projection (w1, w2, w3) |
| Positional Encoding | rope |
Rotary Position Embeddings |
| Normalization | rmsnorm |
Pre-normalization with $\epsilon = 1\text{e-}6$ |
| Vocabulary Size | 32000 |
Custom Rust BPE Tokenizer |
Context Window (max_seq) |
512 |
Sequence length |
| Weight Tying | false |
Independent input embeddings and output LM head |
π Pretraining Details
- Dataset: 4.0 Billion tokens from
HuggingFaceFW/fineweb(sample-10BT) - Framework:
transformer-toolkit - Optimization: AdamW with cosine learning rate decay and linear warmup
- Hardware: 1Γ NVIDIA L40S GPU
- Training Compute Cost:
$14 total cost (14 hours runtime)
π Quickstart & Inference
1. Installation
pip install torch transformer-toolkit
- Downloads last month
- 5
Model tree for Govind222/68m-base
Unable to build the model tree, the base model loops to the model itself. Learn more.
Dataset used to train Govind222/68m-base
Viewer β’ Updated β’ 52.5B β’ 289k β’ 3.6k