File size: 3,016 Bytes
13e5f07
b612803
 
 
 
13e5f07
 
 
 
b612803
 
 
 
 
 
 
 
 
 
 
 
81e35d2
b612803
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
81e35d2
b612803
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
---
title: VantoraLabs
emoji: 🧪
colorFrom: indigo
colorTo: purple
sdk: static
pinned: false
---

# VantoraLabs

We build tiny language models on hardware most people would call a potato.

## The lab bench

Everything here is trained and evaluated on a single desktop:

- **GPU:** NVIDIA GTX 750 (Maxwell, 4 GB VRAM, 512 cores, no tensor cores)
- **CPU:** AMD Ryzen 5 3500X
- **RAM:** 16 GB

Just one machine that has been running
experiments for months and refuses to quit.

## Why tiny models

Big models are expensive to train and expensive to run. Tiny models are
cheap enough to train a dozen of them in an afternoon, which means we can
actually test ideas instead of just talking about them.

The GTX 750 forces a specific kind of discipline. Every parameter has to
earn its place. Every architecture choice has to justify itself against a
baseline. You cannot hide behind scale, so you have to be honest about what
actually works.

## What we study

- **Scaling laws at the extreme low end.** What happens to the
  tokens-per-parameter ratio when a model has only a few thousand
  parameters? Where is the knee? When does repetition stop helping and
  start hurting?
- **Architecture comparisons at fixed size.** Pure transformer versus
  hybrid Mamba-2 + attention, same parameter count, same data, same token
  budget. The hybrid wins on BananaMind by a real margin, and we want to
  know exactly why.
- **Optimizers that matter at small scale.** Muon versus AdamW, and how
  much of the gap is the optimizer versus the architecture.
- **What a 10K-parameter model can actually learn.** The answer is more
  than you would think, and less than you would hope.

## The models

### Vantora-Micro

A 9,800-parameter pure Llama-style transformer. The baseline. Trained on
100M tokens of FineWeb-Edu. BananaMind Elo 810.

### Vantora-Micro-Hybrid

An 11,256-parameter hybrid that pairs a Mamba-2 SSM with attention. Same
data, same budget, same hardware. BananaMind Elo 863, which beats the pure
transformer by 53 points and wins six of seven categories.

These two are the first. They are not the last.

## The plan

The GTX 750 is slow, so we think in terms of what fits in a night. A 10K
model over 100M tokens takes about five minutes. A 20K model over 900M
tokens takes about an hour. That is enough to run real sweeps, real
ablations, and real head-to-heads.

Up next:

- Larger slices of the dataset, one full epoch at a time
- More architecture variants at the same parameter budget
- Muon-only versus AdamW-only ablations to separate optimizer from
  architecture
- Whatever breaks the current best score on the leaderboard

## The rules

- Every model gets a real benchmark run, not a
  hand-waved number.
- Every claim gets a baseline to compare against.
- If it does not fit in VRAM, it does not get built.

## Contact

Open an issue on any model repo, or reach out through the Hugging Face
organization. We read everything.

---

*VantoraLabs. Small models, honest numbers, one very patient GPU.*