File size: 12,461 Bytes
ed55142
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
6d36439
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
---

license: mit
language: [en]
library_name: transformers
pipeline_tag: feature-extraction
tags: [sentence-similarity, feature-extraction, consensus-distillation, geometric-deep-learning, amoe]
datasets: [AbstractPhil/conceptual-captions-12m-webdataset-berts]
base_model: [google-bert/bert-base-uncased, answerdotai/ModernBERT-base, FacebookAI/roberta-base, albert/albert-base-v2, distilbert/distilbert-base-uncased]
---


# captionbert-8192-b

A **58.3M** standalone sentence encoder distilled from the geometric **consensus**
of five BERT-family teachers. No expert models at inference: tokenizer + this
model, 768-d L2-normalized output.

12 layers, 512-d, 8 heads, FFN 2048, 8192 position capacity. **0.53x bert-base.**

This is the **complete-corpus** build: all 66 CC12M chunks, 31.9M rows. Its
sibling [`captionbert-8192-v2`](https://huggingface.co/AbstractPhil/captionbert-8192-v2)
trained on 54 chunks because ModernBERT was missing from 10 of them; those were
repaired and gate-verified before this run.

```python

from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained("AbstractPhil/captionbert-8192-v2-B", trust_remote_code=True)

tok   = AutoTokenizer.from_pretrained("google-bert/bert-base-uncased")



emb = model.encode(["a cat on a windowsill", "a feline by the window"])   # (2, 768)

(emb[0] @ emb[1]).item()



model.attach_amoe()          # this repo's NATIVE arms -- see the warning below

emb = model.encode(["a cat on a windowsill"])

```

## Benchmark

| model | params | STS-B | SICK-R | STS12 | STS13 | STS14 | STS15 | STS16 | BIOSSES | mean |
|---|---|---|---|---|---|---|---|---|---|---|
| bert-base | 109.5M | 0.4729 | 0.5865 | 0.3087 | 0.5988 | 0.4773 | 0.6029 | 0.6373 | 0.5469 | 0.5289 |
| ModernBERT-base | 149.0M | 0.4215 | 0.5479 | 0.3527 | 0.4247 | 0.3795 | 0.5349 | 0.4174 | 0.5630 | 0.4552 |
| roberta-base | 124.6M | 0.5436 | 0.6296 | 0.3211 | 0.5631 | 0.4522 | 0.6134 | 0.6198 | 0.5777 | 0.5401 |
| albert-base-v2 | 11.7M | 0.4784 | 0.5364 | 0.3101 | 0.4831 | 0.3809 | 0.5542 | 0.5491 | 0.4863 | 0.4723 |
| distilbert | 66.4M | 0.5717 | 0.6424 | 0.4344 | 0.6490 | 0.5410 | 0.6663 | 0.6854 | 0.5162 | 0.5883 |
| **captionbert-8192-b** | 58.3M | **0.5752** | **0.6548** | **0.5012** | **0.6037** | **0.5470** | **0.7146** | **0.6782** | **0.5500** | **0.6031** |
| **captionbert-8192-b + arms** | 63.2M | **0.7675** | **0.7374** | **0.6706** | **0.7381** | **0.6945** | **0.8109** | **0.7695** | **0.6472** | **0.7295** |
| captionbert-8192-v2 | 58.3M | 0.5747 | 0.6526 | 0.5051 | 0.5995 | 0.5452 | 0.7136 | 0.6776 | 0.5933 | 0.6077 |
| all-MiniLM-L6-v2 | 22.7M | 0.8203 | 0.7758 | 0.7237 | 0.8058 | 0.7559 | 0.8539 | 0.7899 | 0.8144 | 0.7925 |

All ten models measured in **one harness**, same eight tasks, **mean-pooled and

L2-normalized**, no task tuning. Spearman correlation; `mean` is the unweighted
average over the eight.

`all-MiniLM-L6-v2` was contrastively trained on 1B+ curated sentence pairs. It is
listed for scale, not as a peer -- nothing here saw a similarity label.

**The trunk beats every teacher it was distilled from**, and the best of them
(distilbert, .5883) by +.0194 -- at **13% of their combined 461M parameters**,
having never seen a similarity label. The margin comes mostly from STS12, where
every teacher collapses to .31-.43 and the trunk holds .50.

**With arms it clears the best teacher by +.14** and closes to within **.063** of
a model trained on a billion curated pairs.

Mean-pooled BERT-family encoders are known-weak sentence encoders -- that is the
reason Sentence-BERT exists -- so beating them is an efficiency result rather
than a state-of-the-art one. The MiniLM row is in the table to keep that honest.

### Geometry

| model | self_cos | erank |

|---|---|---|

| bert-base | +0.6071 | 32.8 |

| ModernBERT-base | +0.9001 | 26.1 |

| roberta-base | +0.9594 | 19.8 |

| albert-base-v2 | +0.7473 | 20.9 |

| distilbert | +0.6920 | 31.1 |

| **captionbert-8192-b** | +0.1411 | 36.1 |

| **captionbert-8192-b + arms** | +0.0984 | 55.5 |

| captionbert-8192-v2 | +0.1396 | 36.6 |

| all-MiniLM-L6-v2 | +0.0251 | 86.7 |



`self_cos` is the isotropy gauge: the mean cosine between unrelated sentences.
Mean-pooled BERT-family embeddings sit in a narrow cone (+.61 to +.96), where
cosine cannot discriminate. `erank` is the participation ratio -- how many of the
768 directions carry variance.

Both track capability almost perfectly across all ten models, and **isotropy is

the mechanism**: no isotropy objective appears anywhere in the training stack.
The arms then lift erank 36.6 -> 57.6, the first evidence in this line that
adaptation *adds* usable directions rather than only rotating them.

## More data bought nothing (and that is the finding)

`-b` trained on **19% more rows for 19% more steps** than `-v2`. Head to head:

| | v2 (54ch, 26.9M) | b (66ch, 31.9M) | delta |
|---|---|---|---|
| 8-task mean, bare | .6077 | .6031 | -.0046 |
| 7 tasks excluding BIOSSES | -- | -- | **+.0009** |
| erank (STS-B) | 36.6 | 36.1 | -0.5 |
| self_cos (STS-B) | +.1396 | +.1411 | +.0015 |

| 8-task mean, native arms | .7287 | **.7295** | **+.0008** |



The entire -.0046 comes from BIOSSES, which is 100 rows -- a 0.4-sigma move.

Everything else is a dead heat.



**The ceiling is TEACHER AGREEMENT, not corpus size.** The consensus target uses

**28.7 of 768 directions**: five BERT-family encoders only agree on ~29, and no

amount of the same distribution raises that. The trunk reaches erank ~103 *in

domain* but ~36 out of it -- the structure it builds on captions does not

transfer. The next lever is heterogeneous teachers, measurable at the consensus

stage before a single training step.



## AMOE arms are TRUNK-BOUND -- use this repo's



Three 1.6M-parameter anchors on the frozen trunk, under a trained dispatch.

Anchors toggle **bit-exact**, so one artifact serves both the unsupervised

baseline and the adapted model.



| mask | STS-B | SICK-R | mean (8 tasks) |

|---|---|---|---|

| OFF (bare trunk) | .5752 | .6548 | .6031 |

| `equiv` only | .7219 | .7200 | .6842 |

| `simplify` only | .5995 | .6603 | .6254 |

| `paraphrase` only | .6137 | .6612 | .6295 |

| **all three** | **.7675** | **.7374** | **.7295** |



An `-only` row is that arm **as damped by the dispatch** -- masking never

renormalizes, so it reads lower than the same anchor trained alone.



**Do not attach `captionbert-8192-v2`'s arms to this trunk.** Measured:



| configuration | mean |

|---|---|

| v2 arms on v2 | .7287 |

| v2 arms on **-b** | .6863 |

| + re-aligned routing keys | .6987 |

| **-b native anchors** | **.7295** |



Transferring the arms costs **31% of their gain**. Re-training only the 1,536

routing keys recovers 29% of that; retraining the anchors recovers all of it.

**71% of the loss is in the anchors themselves.**



These two trunks are indistinguishable on eight STS tasks and on geometry, yet

1.6M adapter parameters tell them apart -- adapters read the residual stream and

the task gauges read the pooled output, and the stream carries trunk identity

the output does not. Budget one anchor set per trunk (~18 min).



`attach_amoe()` resolves this repo's own arms by default. Files are under
`amoe/b-collective/`. See [amoe-lora](https://github.com/AbstractEyes/amoe-lora).

## How it was built

1. Five teachers embedded 33M CC12M llava-next captions (mean-pooled, 768-d).
2. One global **whitened Procrustes** map per teacher into `bert-base`'s frame,
   fit on a stratified random sample and **reported out-of-sample** (worst arm
   retains 95% of its in-sample R@1 at 1,833x chance).
3. Consensus = normalized centroid of the aligned teachers, per chunk.
4. Student trained from scratch: InfoNCE(T=0.07) + per-sample MSE against the
   consensus. Pure Adam, no weight decay. 31.9M rows, 62,312 steps at batch
   2048, ~6.4 h on one RTX 6000 Pro.

The alignment maps in `maps/` are **the same maps v2 used** -- refitting them
would put the consensus targets in a different frame with no signal in the loss.

## Known limits

- **Consensus rank ~28.7 of 768.** The model's ceiling, and a property of
  teacher agreement rather than of this model.
- **Alignment quality varies by teacher.** Out-of-sample cosine into the bert
  frame: distil .625, roberta .372, albert .331, modern .327 -- the ordering
  tracks architectural distance from bert-base.
- **Single seed.** The AMOE results carry a measured seed spread of .003-.005;
  the trunk does not have one.
- Trained on image captions; expect caption-like text to be its strongest domain.
- BIOSSES is 100 rows. Treat any single-task delta there as noise.

## Files

```

model.safetensors            the trunk, HF format

config.json                  AutoModel config (auto_map -> modeling_captionbert)

modeling_captionbert.py      CaptionBertV2Model + attach_amoe/detach_amoe

checkpoints/                 training checkpoints (final_model.pt is the ship)

maps/                        alignment maps -- SHARED with v2, do not refit

amoe/b-collective/           native anchors + dispatch + metrics

```

## Output convention

| field | shape | |
|---|---|---|
| `last_hidden_state` | (B, L, 512) | token states |
| `pooler_output` | (B, 768) | **the embedding**, L2-normalized |
| `embedding` | (B, 768) | alias |

`geolip-captionbert-8192` (v1) returned the pooled embedding as
`last_hidden_state`. If porting v1 code, use `pooler_output`.


## deep-arm/ β€” long-context binding attachment (optional, detachable)

The base trunk's attribute binding is semantically alive to ~256 tokens
(its trained position range) and collapses beyond it β€” measured with a
minimal-pair battery ("a red cube on a blue sphere" vs swaps, ratio of
own-attribute to other-attribute state movement at the noun positions;
1.0 = chance). `deep-arm/` restores deep binding **without touching the

trunk**: 4.98M trainable parameters distilled from
`allenai/longformer-base-4096` token states (span-resampled across
tokenizers, mapped 768β†’512 by a whitened-Procrustes fit, out-of-sample
cos .501 / retrieval R@1 .849 vs a dead shuffled null).

**Construction**: (1) position rows 256+ re-initialized by mod-256
tiling of the trained 0–255 table, then trained (rows 0–255 frozen);
(2) one gated 16-slot relay adapter per encoder block (gates open
monotonically with depth, .35–.51 after training); (3) per-token cosine
distillation to the mapped Longformer states over long caption
documents, deep-weighted.

**Binding at depth** (battery ratios, before β†’ after; both alignment
phases of the tiling shown):

| payload depth | before | after |
|---|---|---|
| 10 | 1.59 / 2.01 | 2.69 / 2.27 |
| 480 (tile edge) | 1.17 / 1.13 | 1.63 / 3.50 |
| 1024 (aligned) | 3.12 / 2.92 | 2.66 / 3.96 |
| 1248 (tile edge) | 1.22 / 1.12 | 3.18 / 3.80 |
| 2048 (aligned) | 3.31 / 2.73 | 4.83 / 13.1 |
| 2288 (tile edge) | 1.13 / 0.98 | 1.63 / 1.58 |

(The tiled init alone restores the aligned depths; the trained deep
rows repair the tile edges in a near-to-far wave; the relays amplify
retro-binding wherever gradient reaches. The 13.1 cell is flagged
pending an absolute-distance decomposition.)

**The honest cost**: with the attachment ENGAGED, short-input capability
drops .6031 β†’ .5655 on the 8-task STS mean and shallow isotropy degrades
(self_cos +.003 β†’ +.288) β€” the Longformer-mapped frame is anisotropic.

The attachment is therefore a **length-conditional mode**: adapters are

Οƒ-gated wrappers and rows 0–255 are untouched, so with the wrappers

removed (or gated off) short-input behavior is bit-identical to the

stock trunk. Engage for inputs past ~256 tokens; run stock below.



**Use**: load the trunk as above; from `deep-arm/deep1_arm_s0.pt` copy

`pos_emb.weight`, wrap each `encoder.layers[i]` with its `block{i}.*`

relay (a residual adapter applied to the block output), or skip both to

recover the stock model exactly. `deep-arm/deep1_results.json` carries

the full battery and the fit report.



## Citation



```bibtex

@misc{abstractphil2026captionbertb,

  title  = {captionbert-8192-b: consensus distillation on the complete CC12M corpus},

  author = {AbstractPhil},

  year   = {2026},

  url    = {https://huggingface.co/AbstractPhil/captionbert-8192-v2-B}

}

```



MIT.