Instructions to use vprojectx/Lewis-44M-Bert with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vprojectx/Lewis-44M-Bert with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="vprojectx/Lewis-44M-Bert", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForMaskedLM model = AutoModelForMaskedLM.from_pretrained("vprojectx/Lewis-44M-Bert", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Lewis-44M-Bert
Lewis-44M-Bert is a compact (43.8M parameter) English encoder pretrained from scratch with masked language modeling on ~2B tokens (1B unique tokens, 2 epochs) of curated multi-domain text. It uses a modernized BERT-style architecture (RoPE, RMSNorm, GeGLU, QK-norm) and was trained entirely on a single consumer GPU (RTX 5050, 8 GB VRAM).
It reaches a GLUE dev average of 75.0 (8 tasks, excluding WNLI), above the ELMo baseline and within reach of much larger models, while being about 2.5x smaller than BERT-base.
Loading requires
trust_remote_code=True(custom architecture).
Model details
| Type | Bidirectional Transformer encoder (masked LM) |
| Parameters | 43,845,680 (embeddings tied with the MLM output layer) |
| Layers | 10 |
| Hidden size | 512 |
| Attention heads | 8 (head dim 64) |
| FFN | GeGLU, intermediate size 1152 |
| Position encoding | Rotary (RoPE), theta = 10,000, computed in fp32 |
| Normalization | Pre-norm RMSNorm (+ embedding norm, final norm) |
| Attention | Bidirectional SDPA with QK-norm, no biases |
| Vocabulary | 30,000, cased byte-level BPE (no UNK) |
| Max sequence length | 512 |
| MLM head | dense, GELU, RMSNorm, tied decoder with bias |
| Init | N(0, 0.02); residual output projections scaled by 1/sqrt(2L) |
| Dropout | 0.0 during pretraining (set hidden_dropout_prob=0.1 for fine-tuning) |
Special tokens: <pad>=0, <unk>=1, <s>=2 (BOS), </s>=3 (EOS), <mask>=4.
Every input should start with <s> and end with </s>. Pretraining documents were packed as <s> ... </s>, and the sequence-classification head pools position 0.
Training data (~1B unique tokens)
| Source | Share | Tokens | Purpose |
|---|---|---|---|
| FineWeb-Edu (score >= 3.5 only) | 50% | 500M | Dense, educational, analytical web text |
| Wikipedia (English, 20231101) | 25% | 250M | Factual knowledge and named entities |
| Project Gutenberg | 15% | 150M | Long-form narrative, coreference, literary style |
| Cosmopedia (synthetic) | 10% | 100M | Procedural, structured, textbook-style text |
Documents were tokenized with a tokenizer trained on the same mix, packed into 512-token chunks, and shuffled globally. About 0.2% of chunks (3,875 sequences) were held out for validation. Two epochs over the data gives ~2B tokens seen.
Training procedure
| Objective | Masked LM, 25% masking rate (80% <mask> / 10% random / 10% unchanged); special tokens never masked |
| Optimizer | AdamW, betas (0.9, 0.98), eps 1e-6 |
| Weight decay | 0.01 (not applied to norms, biases, embeddings) |
| LR schedule | Linear warmup (1,219 steps), cosine decay from 4e-4 to 4e-5 |
| Batch size | 65,536 tokens/update (32 x 512 x 4 grad-accum) |
| Steps | 30,478 (2 epochs x 15,239) |
| Precision | bf16 autocast, TF32 matmuls, torch.compile on the backbone |
| Grad clipping | 1.0 |
| Throughput | ~23k to 45k tokens/s |
Compute: one NVIDIA RTX 5050 (8 GB VRAM), Windows, PyTorch. Roughly 21 hours of training time, estimated from logged throughput.
Final pretraining result: validation MLM loss 1.987 (perplexity 7.29) at 25% masking on held-out chunks from the same distribution. This is not comparable to perplexities reported at 15% masking.
Benchmarks: GLUE (dev set)
Single run, final epoch, dev sets.
| Task | Metric | Score |
|---|---|---|
| CoLA | Matthews corr. | 36.5 |
| SST-2 | Accuracy | 89.3 |
| MRPC | Acc / F1 | 80.4 / 86.3 |
| QQP | Acc / F1 | 89.4 / 85.8 |
| STS-B | Pearson / Spearman | 83.3 / 82.9 |
| MNLI | Acc (matched / mismatched) | 78.5 / 78.9 |
| QNLI | Accuracy | 85.1 |
| RTE | Accuracy | 56.7 |
| Average | 8 tasks, WNLI excluded | 75.0 |
Comparison with other models
Per-task scores use one number per task (MRPC and QQP: mean of accuracy and F1; STS-B: mean of Pearson and Spearman; MNLI: matched accuracy; CoLA: MCC), with the average taken over the 8 tasks below. Reference rows are the dev-set numbers reported in the DistilBERT paper (Sanh et al., 2019); Lewis-44M-Bert is a single run, so small differences are within noise.
| Model | Params | CoLA | MNLI | MRPC | QNLI | QQP | RTE | SST-2 | STS-B | Avg |
|---|---|---|---|---|---|---|---|---|---|---|
| ELMo + BiLSTM baseline | n/a | 44.1 | 68.6 | 76.6 | 71.1 | 86.2 | 53.4 | 91.5 | 70.4 | 70.2 |
| Lewis-44M-Bert (ours) | 44M | 36.5 | 78.5 | 83.4 | 85.1 | 87.6 | 56.7 | 89.3 | 83.1 | 75.0 |
| DistilBERT | 66M | 51.3 | 82.2 | 87.5 | 89.2 | 88.5 | 59.9 | 91.3 | 86.9 | 79.6 |
| BERT-base | 110M | 56.3 | 86.7 | 88.6 | 91.8 | 89.6 | 69.3 | 92.7 | 89.0 | 83.0 |
Highlights
- Compute-efficient. Pretrained on ~2B tokens (BERT-base saw roughly 40 epochs over a 3.3B-word corpus) on a single 8 GB consumer GPU in about a day.
- Small and fast. 44M parameters: 2.5x smaller than BERT-base, 1.5x smaller than DistilBERT, and cheap to fine-tune and serve.
- Strong on sentence-pair and sentiment tasks for its size: 89.3 on SST-2, 89.4 accuracy on QQP, 78.5 / 78.9 on MNLI.
- Modern design. RoPE, RMSNorm, GeGLU, QK-norm and a transformed MLM head, which are choices that typically improve stability and quality per parameter.
- Curated data. An educational-quality filter (FineWeb-Edu >= 3.5) plus encyclopedic, literary and synthetic sources, with no distillation from a teacher.
- Fully reproducible. Data prep, tokenizer, training and evaluation scripts are simple single-file Python.
Limitations and what to improve next
Main limitation: the model is weak on tasks that need fine-grained linguistic or reasoning ability with little training data, such as CoLA (36.5 MCC) and RTE (56.7%, close to chance). This is the expected cost of 44M parameters and only ~1B unique tokens.
Other limitations:
- English only, cased 30k vocabulary, maximum 512 tokens.
- Packed sequences let attention cross document boundaries.
- Validation chunks come from the same documents as training, so MLM perplexity is slightly optimistic.
- GLUE results are single-run; CoLA, RTE and MRPC vary several points across seeds.
- Not evaluated for bias, toxicity or factual reliability; it inherits the biases of web, encyclopedic and 19th/20th-century literary text.
Ideas for the next version:
- More unique data (5 to 10B tokens) instead of repeated epochs.
- A larger or deeper model (80 to 120M) and a longer-context stage (1k to 2k tokens).
- Document-aware attention masks when packing.
- Knowledge distillation from a larger teacher.
- Multi-seed evaluation and MNLI-intermediate fine-tuning for small tasks (RTE, MRPC, STS-B).
Intended use
Fine-tuning for classification (sentiment, spam, topic, NLI), token classification, embeddings and retrieval features in resource-limited settings, and as a base for research on small encoders. Out of scope: text generation, long documents, non-English text, and high-stakes decisions without task-specific validation.
How to use
import torch
from transformers import AutoTokenizer, AutoModelForMaskedLM
repo = "vprojectx/Lewis-44M-Bert"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForMaskedLM.from_pretrained(repo, trust_remote_code=True).eval()
text = "The capital of France is <mask>."
ids = [tok.bos_token_id] + tok(text, add_special_tokens=False).input_ids + [tok.eos_token_id]
x = torch.tensor([ids])
with torch.no_grad():
logits = model(input_ids=x).logits
pos = (x[0] == tok.mask_token_id).nonzero()[0].item()
print(tok.convert_ids_to_tokens(logits[0, pos].topk(5).indices))
Fine-tuning for classification:
from transformers import AutoModelForSequenceClassification
model = AutoModelForSequenceClassification.from_pretrained(
repo, trust_remote_code=True, num_labels=2, hidden_dropout_prob=0.1
)
# Always prepend <s> (id 2) and append </s> (id 3) to each input.
Tips: fine-tune with lr 3e-5 to 5e-5, 3 to 5 epochs, and add dropout 0.1.
Files
config.json, model.safetensors, configuration_Lewis_44M_Bert.py, modeling_Lewis_44M_Bert.py, tokenizer.json, tokenizer_config.json, special_tokens_map.json.
Citation
@misc{lewis44mbert2026,
title = {Lewis-44M-Bert: a compact encoder pretrained from scratch on a single consumer GPU},
author = {viraj salunke},
year = {2026},
url = {https://huggingface.co/vprojectx/Lewis-44M-Bert}
}
- Downloads last month
- -