Feature Extraction
Transformers
TensorBoard
Safetensors
English
captionbert_v2
sentence-similarity
consensus-distillation
geometric-deep-learning
amoe
custom_code
Instructions to use AbstractPhil/captionbert-8192-v2-b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AbstractPhil/captionbert-8192-v2-b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="AbstractPhil/captionbert-8192-v2-b", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("AbstractPhil/captionbert-8192-v2-b", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 12,461 Bytes
ed55142 6d36439 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 | ---
license: mit
language: [en]
library_name: transformers
pipeline_tag: feature-extraction
tags: [sentence-similarity, feature-extraction, consensus-distillation, geometric-deep-learning, amoe]
datasets: [AbstractPhil/conceptual-captions-12m-webdataset-berts]
base_model: [google-bert/bert-base-uncased, answerdotai/ModernBERT-base, FacebookAI/roberta-base, albert/albert-base-v2, distilbert/distilbert-base-uncased]
---
# captionbert-8192-b
A **58.3M** standalone sentence encoder distilled from the geometric **consensus**
of five BERT-family teachers. No expert models at inference: tokenizer + this
model, 768-d L2-normalized output.
12 layers, 512-d, 8 heads, FFN 2048, 8192 position capacity. **0.53x bert-base.**
This is the **complete-corpus** build: all 66 CC12M chunks, 31.9M rows. Its
sibling [`captionbert-8192-v2`](https://huggingface.co/AbstractPhil/captionbert-8192-v2)
trained on 54 chunks because ModernBERT was missing from 10 of them; those were
repaired and gate-verified before this run.
```python
from transformers import AutoModel, AutoTokenizer
model = AutoModel.from_pretrained("AbstractPhil/captionbert-8192-v2-B", trust_remote_code=True)
tok = AutoTokenizer.from_pretrained("google-bert/bert-base-uncased")
emb = model.encode(["a cat on a windowsill", "a feline by the window"]) # (2, 768)
(emb[0] @ emb[1]).item()
model.attach_amoe() # this repo's NATIVE arms -- see the warning below
emb = model.encode(["a cat on a windowsill"])
```
## Benchmark
| model | params | STS-B | SICK-R | STS12 | STS13 | STS14 | STS15 | STS16 | BIOSSES | mean |
|---|---|---|---|---|---|---|---|---|---|---|
| bert-base | 109.5M | 0.4729 | 0.5865 | 0.3087 | 0.5988 | 0.4773 | 0.6029 | 0.6373 | 0.5469 | 0.5289 |
| ModernBERT-base | 149.0M | 0.4215 | 0.5479 | 0.3527 | 0.4247 | 0.3795 | 0.5349 | 0.4174 | 0.5630 | 0.4552 |
| roberta-base | 124.6M | 0.5436 | 0.6296 | 0.3211 | 0.5631 | 0.4522 | 0.6134 | 0.6198 | 0.5777 | 0.5401 |
| albert-base-v2 | 11.7M | 0.4784 | 0.5364 | 0.3101 | 0.4831 | 0.3809 | 0.5542 | 0.5491 | 0.4863 | 0.4723 |
| distilbert | 66.4M | 0.5717 | 0.6424 | 0.4344 | 0.6490 | 0.5410 | 0.6663 | 0.6854 | 0.5162 | 0.5883 |
| **captionbert-8192-b** | 58.3M | **0.5752** | **0.6548** | **0.5012** | **0.6037** | **0.5470** | **0.7146** | **0.6782** | **0.5500** | **0.6031** |
| **captionbert-8192-b + arms** | 63.2M | **0.7675** | **0.7374** | **0.6706** | **0.7381** | **0.6945** | **0.8109** | **0.7695** | **0.6472** | **0.7295** |
| captionbert-8192-v2 | 58.3M | 0.5747 | 0.6526 | 0.5051 | 0.5995 | 0.5452 | 0.7136 | 0.6776 | 0.5933 | 0.6077 |
| all-MiniLM-L6-v2 | 22.7M | 0.8203 | 0.7758 | 0.7237 | 0.8058 | 0.7559 | 0.8539 | 0.7899 | 0.8144 | 0.7925 |
All ten models measured in **one harness**, same eight tasks, **mean-pooled and
L2-normalized**, no task tuning. Spearman correlation; `mean` is the unweighted
average over the eight.
`all-MiniLM-L6-v2` was contrastively trained on 1B+ curated sentence pairs. It is
listed for scale, not as a peer -- nothing here saw a similarity label.
**The trunk beats every teacher it was distilled from**, and the best of them
(distilbert, .5883) by +.0194 -- at **13% of their combined 461M parameters**,
having never seen a similarity label. The margin comes mostly from STS12, where
every teacher collapses to .31-.43 and the trunk holds .50.
**With arms it clears the best teacher by +.14** and closes to within **.063** of
a model trained on a billion curated pairs.
Mean-pooled BERT-family encoders are known-weak sentence encoders -- that is the
reason Sentence-BERT exists -- so beating them is an efficiency result rather
than a state-of-the-art one. The MiniLM row is in the table to keep that honest.
### Geometry
| model | self_cos | erank |
|---|---|---|
| bert-base | +0.6071 | 32.8 |
| ModernBERT-base | +0.9001 | 26.1 |
| roberta-base | +0.9594 | 19.8 |
| albert-base-v2 | +0.7473 | 20.9 |
| distilbert | +0.6920 | 31.1 |
| **captionbert-8192-b** | +0.1411 | 36.1 |
| **captionbert-8192-b + arms** | +0.0984 | 55.5 |
| captionbert-8192-v2 | +0.1396 | 36.6 |
| all-MiniLM-L6-v2 | +0.0251 | 86.7 |
`self_cos` is the isotropy gauge: the mean cosine between unrelated sentences.
Mean-pooled BERT-family embeddings sit in a narrow cone (+.61 to +.96), where
cosine cannot discriminate. `erank` is the participation ratio -- how many of the
768 directions carry variance.
Both track capability almost perfectly across all ten models, and **isotropy is
the mechanism**: no isotropy objective appears anywhere in the training stack.
The arms then lift erank 36.6 -> 57.6, the first evidence in this line that
adaptation *adds* usable directions rather than only rotating them.
## More data bought nothing (and that is the finding)
`-b` trained on **19% more rows for 19% more steps** than `-v2`. Head to head:
| | v2 (54ch, 26.9M) | b (66ch, 31.9M) | delta |
|---|---|---|---|
| 8-task mean, bare | .6077 | .6031 | -.0046 |
| 7 tasks excluding BIOSSES | -- | -- | **+.0009** |
| erank (STS-B) | 36.6 | 36.1 | -0.5 |
| self_cos (STS-B) | +.1396 | +.1411 | +.0015 |
| 8-task mean, native arms | .7287 | **.7295** | **+.0008** |
The entire -.0046 comes from BIOSSES, which is 100 rows -- a 0.4-sigma move.
Everything else is a dead heat.
**The ceiling is TEACHER AGREEMENT, not corpus size.** The consensus target uses
**28.7 of 768 directions**: five BERT-family encoders only agree on ~29, and no
amount of the same distribution raises that. The trunk reaches erank ~103 *in
domain* but ~36 out of it -- the structure it builds on captions does not
transfer. The next lever is heterogeneous teachers, measurable at the consensus
stage before a single training step.
## AMOE arms are TRUNK-BOUND -- use this repo's
Three 1.6M-parameter anchors on the frozen trunk, under a trained dispatch.
Anchors toggle **bit-exact**, so one artifact serves both the unsupervised
baseline and the adapted model.
| mask | STS-B | SICK-R | mean (8 tasks) |
|---|---|---|---|
| OFF (bare trunk) | .5752 | .6548 | .6031 |
| `equiv` only | .7219 | .7200 | .6842 |
| `simplify` only | .5995 | .6603 | .6254 |
| `paraphrase` only | .6137 | .6612 | .6295 |
| **all three** | **.7675** | **.7374** | **.7295** |
An `-only` row is that arm **as damped by the dispatch** -- masking never
renormalizes, so it reads lower than the same anchor trained alone.
**Do not attach `captionbert-8192-v2`'s arms to this trunk.** Measured:
| configuration | mean |
|---|---|
| v2 arms on v2 | .7287 |
| v2 arms on **-b** | .6863 |
| + re-aligned routing keys | .6987 |
| **-b native anchors** | **.7295** |
Transferring the arms costs **31% of their gain**. Re-training only the 1,536
routing keys recovers 29% of that; retraining the anchors recovers all of it.
**71% of the loss is in the anchors themselves.**
These two trunks are indistinguishable on eight STS tasks and on geometry, yet
1.6M adapter parameters tell them apart -- adapters read the residual stream and
the task gauges read the pooled output, and the stream carries trunk identity
the output does not. Budget one anchor set per trunk (~18 min).
`attach_amoe()` resolves this repo's own arms by default. Files are under
`amoe/b-collective/`. See [amoe-lora](https://github.com/AbstractEyes/amoe-lora).
## How it was built
1. Five teachers embedded 33M CC12M llava-next captions (mean-pooled, 768-d).
2. One global **whitened Procrustes** map per teacher into `bert-base`'s frame,
fit on a stratified random sample and **reported out-of-sample** (worst arm
retains 95% of its in-sample R@1 at 1,833x chance).
3. Consensus = normalized centroid of the aligned teachers, per chunk.
4. Student trained from scratch: InfoNCE(T=0.07) + per-sample MSE against the
consensus. Pure Adam, no weight decay. 31.9M rows, 62,312 steps at batch
2048, ~6.4 h on one RTX 6000 Pro.
The alignment maps in `maps/` are **the same maps v2 used** -- refitting them
would put the consensus targets in a different frame with no signal in the loss.
## Known limits
- **Consensus rank ~28.7 of 768.** The model's ceiling, and a property of
teacher agreement rather than of this model.
- **Alignment quality varies by teacher.** Out-of-sample cosine into the bert
frame: distil .625, roberta .372, albert .331, modern .327 -- the ordering
tracks architectural distance from bert-base.
- **Single seed.** The AMOE results carry a measured seed spread of .003-.005;
the trunk does not have one.
- Trained on image captions; expect caption-like text to be its strongest domain.
- BIOSSES is 100 rows. Treat any single-task delta there as noise.
## Files
```
model.safetensors the trunk, HF format
config.json AutoModel config (auto_map -> modeling_captionbert)
modeling_captionbert.py CaptionBertV2Model + attach_amoe/detach_amoe
checkpoints/ training checkpoints (final_model.pt is the ship)
maps/ alignment maps -- SHARED with v2, do not refit
amoe/b-collective/ native anchors + dispatch + metrics
```
## Output convention
| field | shape | |
|---|---|---|
| `last_hidden_state` | (B, L, 512) | token states |
| `pooler_output` | (B, 768) | **the embedding**, L2-normalized |
| `embedding` | (B, 768) | alias |
`geolip-captionbert-8192` (v1) returned the pooled embedding as
`last_hidden_state`. If porting v1 code, use `pooler_output`.
## deep-arm/ β long-context binding attachment (optional, detachable)
The base trunk's attribute binding is semantically alive to ~256 tokens
(its trained position range) and collapses beyond it β measured with a
minimal-pair battery ("a red cube on a blue sphere" vs swaps, ratio of
own-attribute to other-attribute state movement at the noun positions;
1.0 = chance). `deep-arm/` restores deep binding **without touching the
trunk**: 4.98M trainable parameters distilled from
`allenai/longformer-base-4096` token states (span-resampled across
tokenizers, mapped 768β512 by a whitened-Procrustes fit, out-of-sample
cos .501 / retrieval R@1 .849 vs a dead shuffled null).
**Construction**: (1) position rows 256+ re-initialized by mod-256
tiling of the trained 0β255 table, then trained (rows 0β255 frozen);
(2) one gated 16-slot relay adapter per encoder block (gates open
monotonically with depth, .35β.51 after training); (3) per-token cosine
distillation to the mapped Longformer states over long caption
documents, deep-weighted.
**Binding at depth** (battery ratios, before β after; both alignment
phases of the tiling shown):
| payload depth | before | after |
|---|---|---|
| 10 | 1.59 / 2.01 | 2.69 / 2.27 |
| 480 (tile edge) | 1.17 / 1.13 | 1.63 / 3.50 |
| 1024 (aligned) | 3.12 / 2.92 | 2.66 / 3.96 |
| 1248 (tile edge) | 1.22 / 1.12 | 3.18 / 3.80 |
| 2048 (aligned) | 3.31 / 2.73 | 4.83 / 13.1 |
| 2288 (tile edge) | 1.13 / 0.98 | 1.63 / 1.58 |
(The tiled init alone restores the aligned depths; the trained deep
rows repair the tile edges in a near-to-far wave; the relays amplify
retro-binding wherever gradient reaches. The 13.1 cell is flagged
pending an absolute-distance decomposition.)
**The honest cost**: with the attachment ENGAGED, short-input capability
drops .6031 β .5655 on the 8-task STS mean and shallow isotropy degrades
(self_cos +.003 β +.288) β the Longformer-mapped frame is anisotropic.
The attachment is therefore a **length-conditional mode**: adapters are
Ο-gated wrappers and rows 0β255 are untouched, so with the wrappers
removed (or gated off) short-input behavior is bit-identical to the
stock trunk. Engage for inputs past ~256 tokens; run stock below.
**Use**: load the trunk as above; from `deep-arm/deep1_arm_s0.pt` copy
`pos_emb.weight`, wrap each `encoder.layers[i]` with its `block{i}.*`
relay (a residual adapter applied to the block output), or skip both to
recover the stock model exactly. `deep-arm/deep1_results.json` carries
the full battery and the fit report.
## Citation
```bibtex
@misc{abstractphil2026captionbertb,
title = {captionbert-8192-b: consensus distillation on the complete CC12M corpus},
author = {AbstractPhil},
year = {2026},
url = {https://huggingface.co/AbstractPhil/captionbert-8192-v2-B}
}
```
MIT. |