You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

πŸ›οΈ Govind222/68m-base

Govind222/68m-base is a 68.96M parameter Dense Foundation / Base Small Language Model (SLM) trained completely from scratch using transformer-toolkit on 4.0 Billion tokens of FineWeb (sample-10BT).

This checkpoint is the raw pretrained foundation base model. It serves as a data-efficient foundation for downstream task adaptation, instruction fine-tuning, domain specialization, or edge device deployment.


⚑ Key Highlights

  • Pretrained from Scratch: Trained on 4 Billion tokens of clean English web text from Hugging Face's FineWeb (sample-10BT).
  • Pure Dense Transformer: 100% dense computation with zero MoE, maximizing per-parameter compute density and caching efficiency on constrained hardware.
  • Modern Architectural Recipe:
    • Grouped-Query Attention (GQA): 8 query heads, 2 key/value heads (4:1 ratio) for fast KV-cache memory reuse.
    • SwiGLU FFN: Non-linear Swish-Gated Linear Units with 1,536 hidden dimension ($3 \times d_{\text{model}}$).
    • RoPE: Rotary Position Embeddings for robust position modeling.
    • RMSNorm: Pre-layer normalization with $\epsilon = 1\text{e-}6$.
  • Data Efficiency: Competitive 0-shot benchmark accuracy against models trained on 10Γ— to 50Γ— more parameters/compute.
  • Compact Footprint: Under ~280 MB RAM footprint at runtime, ideal for edge Android/iOS smartphones and Raspberry Pi.

πŸ“Š Zero-Shot Benchmark Evaluation

Evaluated zero-shot using lm-evaluation-harness directly on the base pretrained checkpoint:

Benchmark Task / Metric Govind222/68m-base Pythia 70M Pythia 160M GPT-Neo 125M OPT 125M
PIQA Physical Interaction QA (acc) 58.1% 58.6% 61.4% 62.3% 61.8%
Winogrande Commonsense Reasoning (acc) 48.7% 51.6% 50.4% 49.5% 51.5%
ARC-Easy Science QA Reasoning (acc_norm) 31.4% 34.2% β€” 38.9% 39.8%
Average Mean Accuracy across 4 tasks 39.6% 41.4% 43.7% 43.5% 43.9%

βš™οΈ Model Architecture Specifications

Hyperparameter Value Description
Total Parameters 68,957,056 (~68.96M) Exact non-embedding + embedding tensor weight sum
Model Type Dense Decoder-Only SLM Pure dense transformer (no MoE)
Layers (n_layers) 12 Transformer blocks
Hidden Dimension (dim) 512 Model width ($d_{\text{model}}$)
Query Attention Heads (n_heads) 8 Head dimension = 64
KV Attention Heads (n_kv_heads) 2 4:1 Grouped Query Attention (GQA)
Feed-Forward Dimension (hidden_dim) 1536 $3 \times d_{\text{model}}$
FFN Activation swiglu Dense SwiGLU projection (w1, w2, w3)
Positional Encoding rope Rotary Position Embeddings
Normalization rmsnorm Pre-normalization with $\epsilon = 1\text{e-}6$
Vocabulary Size 32000 Custom Rust BPE Tokenizer
Context Window (max_seq) 512 Sequence length
Weight Tying false Independent input embeddings and output LM head

πŸ“ˆ Pretraining Details

  • Dataset: 4.0 Billion tokens from HuggingFaceFW/fineweb (sample-10BT)
  • Framework: transformer-toolkit
  • Optimization: AdamW with cosine learning rate decay and linear warmup
  • Hardware: 1Γ— NVIDIA L40S GPU
  • Training Compute Cost: $14 total cost (14 hours runtime)

πŸš€ Quickstart & Inference

1. Installation

pip install torch transformer-toolkit
Downloads last month
5
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Govind222/68m-base

Unable to build the model tree, the base model loops to the model itself. Learn more.

Dataset used to train Govind222/68m-base

Space using Govind222/68m-base 1