Lewis-44M-Bert

Lewis-44M-Bert is a compact (43.8M parameter) English encoder pretrained from scratch with masked language modeling on ~2B tokens (1B unique tokens, 2 epochs) of curated multi-domain text. It uses a modernized BERT-style architecture (RoPE, RMSNorm, GeGLU, QK-norm) and was trained entirely on a single consumer GPU (RTX 5050, 8 GB VRAM).

It reaches a GLUE dev average of 75.0 (8 tasks, excluding WNLI), above the ELMo baseline and within reach of much larger models, while being about 2.5x smaller than BERT-base.

Loading requires trust_remote_code=True (custom architecture).


Model details

Type Bidirectional Transformer encoder (masked LM)
Parameters 43,845,680 (embeddings tied with the MLM output layer)
Layers 10
Hidden size 512
Attention heads 8 (head dim 64)
FFN GeGLU, intermediate size 1152
Position encoding Rotary (RoPE), theta = 10,000, computed in fp32
Normalization Pre-norm RMSNorm (+ embedding norm, final norm)
Attention Bidirectional SDPA with QK-norm, no biases
Vocabulary 30,000, cased byte-level BPE (no UNK)
Max sequence length 512
MLM head dense, GELU, RMSNorm, tied decoder with bias
Init N(0, 0.02); residual output projections scaled by 1/sqrt(2L)
Dropout 0.0 during pretraining (set hidden_dropout_prob=0.1 for fine-tuning)

Special tokens: <pad>=0, <unk>=1, <s>=2 (BOS), </s>=3 (EOS), <mask>=4. Every input should start with <s> and end with </s>. Pretraining documents were packed as <s> ... </s>, and the sequence-classification head pools position 0.

Training data (~1B unique tokens)

Source Share Tokens Purpose
FineWeb-Edu (score >= 3.5 only) 50% 500M Dense, educational, analytical web text
Wikipedia (English, 20231101) 25% 250M Factual knowledge and named entities
Project Gutenberg 15% 150M Long-form narrative, coreference, literary style
Cosmopedia (synthetic) 10% 100M Procedural, structured, textbook-style text

Documents were tokenized with a tokenizer trained on the same mix, packed into 512-token chunks, and shuffled globally. About 0.2% of chunks (3,875 sequences) were held out for validation. Two epochs over the data gives ~2B tokens seen.

Training procedure

Objective Masked LM, 25% masking rate (80% <mask> / 10% random / 10% unchanged); special tokens never masked
Optimizer AdamW, betas (0.9, 0.98), eps 1e-6
Weight decay 0.01 (not applied to norms, biases, embeddings)
LR schedule Linear warmup (1,219 steps), cosine decay from 4e-4 to 4e-5
Batch size 65,536 tokens/update (32 x 512 x 4 grad-accum)
Steps 30,478 (2 epochs x 15,239)
Precision bf16 autocast, TF32 matmuls, torch.compile on the backbone
Grad clipping 1.0
Throughput ~23k to 45k tokens/s

Compute: one NVIDIA RTX 5050 (8 GB VRAM), Windows, PyTorch. Roughly 21 hours of training time, estimated from logged throughput.

Final pretraining result: validation MLM loss 1.987 (perplexity 7.29) at 25% masking on held-out chunks from the same distribution. This is not comparable to perplexities reported at 15% masking.

Benchmarks: GLUE (dev set)

Single run, final epoch, dev sets.

Task Metric Score
CoLA Matthews corr. 36.5
SST-2 Accuracy 89.3
MRPC Acc / F1 80.4 / 86.3
QQP Acc / F1 89.4 / 85.8
STS-B Pearson / Spearman 83.3 / 82.9
MNLI Acc (matched / mismatched) 78.5 / 78.9
QNLI Accuracy 85.1
RTE Accuracy 56.7
Average 8 tasks, WNLI excluded 75.0

Comparison with other models

Per-task scores use one number per task (MRPC and QQP: mean of accuracy and F1; STS-B: mean of Pearson and Spearman; MNLI: matched accuracy; CoLA: MCC), with the average taken over the 8 tasks below. Reference rows are the dev-set numbers reported in the DistilBERT paper (Sanh et al., 2019); Lewis-44M-Bert is a single run, so small differences are within noise.

Model Params CoLA MNLI MRPC QNLI QQP RTE SST-2 STS-B Avg
ELMo + BiLSTM baseline n/a 44.1 68.6 76.6 71.1 86.2 53.4 91.5 70.4 70.2
Lewis-44M-Bert (ours) 44M 36.5 78.5 83.4 85.1 87.6 56.7 89.3 83.1 75.0
DistilBERT 66M 51.3 82.2 87.5 89.2 88.5 59.9 91.3 86.9 79.6
BERT-base 110M 56.3 86.7 88.6 91.8 89.6 69.3 92.7 89.0 83.0

Highlights

  • Compute-efficient. Pretrained on ~2B tokens (BERT-base saw roughly 40 epochs over a 3.3B-word corpus) on a single 8 GB consumer GPU in about a day.
  • Small and fast. 44M parameters: 2.5x smaller than BERT-base, 1.5x smaller than DistilBERT, and cheap to fine-tune and serve.
  • Strong on sentence-pair and sentiment tasks for its size: 89.3 on SST-2, 89.4 accuracy on QQP, 78.5 / 78.9 on MNLI.
  • Modern design. RoPE, RMSNorm, GeGLU, QK-norm and a transformed MLM head, which are choices that typically improve stability and quality per parameter.
  • Curated data. An educational-quality filter (FineWeb-Edu >= 3.5) plus encyclopedic, literary and synthetic sources, with no distillation from a teacher.
  • Fully reproducible. Data prep, tokenizer, training and evaluation scripts are simple single-file Python.

Limitations and what to improve next

Main limitation: the model is weak on tasks that need fine-grained linguistic or reasoning ability with little training data, such as CoLA (36.5 MCC) and RTE (56.7%, close to chance). This is the expected cost of 44M parameters and only ~1B unique tokens.

Other limitations:

  • English only, cased 30k vocabulary, maximum 512 tokens.
  • Packed sequences let attention cross document boundaries.
  • Validation chunks come from the same documents as training, so MLM perplexity is slightly optimistic.
  • GLUE results are single-run; CoLA, RTE and MRPC vary several points across seeds.
  • Not evaluated for bias, toxicity or factual reliability; it inherits the biases of web, encyclopedic and 19th/20th-century literary text.

Ideas for the next version:

  1. More unique data (5 to 10B tokens) instead of repeated epochs.
  2. A larger or deeper model (80 to 120M) and a longer-context stage (1k to 2k tokens).
  3. Document-aware attention masks when packing.
  4. Knowledge distillation from a larger teacher.
  5. Multi-seed evaluation and MNLI-intermediate fine-tuning for small tasks (RTE, MRPC, STS-B).

Intended use

Fine-tuning for classification (sentiment, spam, topic, NLI), token classification, embeddings and retrieval features in resource-limited settings, and as a base for research on small encoders. Out of scope: text generation, long documents, non-English text, and high-stakes decisions without task-specific validation.

How to use

import torch
from transformers import AutoTokenizer, AutoModelForMaskedLM

repo = "vprojectx/Lewis-44M-Bert"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForMaskedLM.from_pretrained(repo, trust_remote_code=True).eval()

text = "The capital of France is <mask>."
ids = [tok.bos_token_id] + tok(text, add_special_tokens=False).input_ids + [tok.eos_token_id]
x = torch.tensor([ids])

with torch.no_grad():
    logits = model(input_ids=x).logits

pos = (x[0] == tok.mask_token_id).nonzero()[0].item()
print(tok.convert_ids_to_tokens(logits[0, pos].topk(5).indices))

Fine-tuning for classification:

from transformers import AutoModelForSequenceClassification

model = AutoModelForSequenceClassification.from_pretrained(
    repo, trust_remote_code=True, num_labels=2, hidden_dropout_prob=0.1
)
# Always prepend <s> (id 2) and append </s> (id 3) to each input.

Tips: fine-tune with lr 3e-5 to 5e-5, 3 to 5 epochs, and add dropout 0.1.

Files

config.json, model.safetensors, configuration_Lewis_44M_Bert.py, modeling_Lewis_44M_Bert.py, tokenizer.json, tokenizer_config.json, special_tokens_map.json.

Citation

@misc{lewis44mbert2026,
  title  = {Lewis-44M-Bert: a compact encoder pretrained from scratch on a single consumer GPU},
  author = {viraj salunke},
  year   = {2026},
  url    = {https://huggingface.co/vprojectx/Lewis-44M-Bert}
}
Downloads last month
-
Safetensors
Model size
43.8M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train vprojectx/Lewis-44M-Bert