Instructions to use AbstractPhil/captionbert-8192-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AbstractPhil/captionbert-8192-v2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="AbstractPhil/captionbert-8192-v2", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("AbstractPhil/captionbert-8192-v2", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
license: mit
tags:
- amoe
- adapter
- mixture-of-experts
- sentence-similarity
- captionbert
- aleph
base_model: AbstractPhil/captionbert-8192-v2
library_name: amoe-lora
captionbert-8192-v2 :: AMOE 3-anchor collective
Three amoe-lora aleph anchors on the
frozen captionbert-8192-v2 trunk, under a trained dispatch, greedily
selected from a six-candidate pool. The trunk never moves; only the adapters
and the routing keys train.
This is the shipped collective and the source of the model's headline number: .6077 -> .7287 mean Spearman over eight sentence-similarity tasks.
| anchor | trained on | in the ship? |
|---|---|---|
equiv |
all-nli triplets | yes |
simplify |
simple-wiki + altlex + sentence-compression | yes |
paraphrase |
paraphrase pairs | yes |
lexical |
lexical-overlap pairs | no -- rejected by the sweep |
topical |
topical pairs | no -- tied, not taken (see below) |
random |
random pairing | no -- it is the capacity control |
All six anchors ship in this folder so the selection can be reproduced or
contested. equiv and simplify are byte-identical to the ones in amoe/moe/
(sha256 0029ad69…, 8e419f69…) -- the collective reused them rather than
retraining.
Results -- 8 tasks, Spearman rho
The single-anchor columns are masked reads: the collective with the other two arms disabled. Masking never renormalizes the amplitude budget, so a masked single is not the same as that anchor trained and run alone -- it delivers strictly less. Compare masked-to-masked; do not read these as solo scores.
| task | OFF | equiv masked |
simplify masked |
paraphrase masked |
COLLECTIVE |
|---|---|---|---|---|---|
| STS-B | .5747 | .7208 | .6022 | .6127 | .7684 |
| SICK-R | .6526 | .7269 | .6580 | .6585 | .7391 |
| STS12 | .5051 | .6275 | .5359 | .5154 | .6682 |
| STS13 | .5995 | .6988 | .6304 | .6338 | .7557 |
| STS14 | .5452 | .6423 | .5747 | .5774 | .6921 |
| STS15 | .7136 | .7783 | .7391 | .7328 | .8055 |
| STS16 | .6776 | .7107 | .6945 | .7332 | .7626 |
| BIOSSES | .5933 | .5845 | .6016 | .6051 | .6382 |
| mean | .6077 | .6862 | .6295 | .6336 | .7287 |
| mean erank | 30.56 | 36.70 | 34.36 | 34.02 | 45.00 |
The collective beats its best masked member by +.0425 (.7287 vs equiv's
.6862) and beats it on 8 of 8 tasks.
equiv masked is worse than the bare trunk on BIOSSES (.5845 vs .5933) --
biomedical similarity is outside every anchor's training domain -- yet the
collective still lands +.0450 there. Recovering on the one task where the
strongest member regresses is the honest read on this dispatch.
Effective rank rises with the score (30.6 -> 45.0). The anchors add directions, not merely a rescaling: this trunk's unsupervised consensus target used only ~28.7 of 768 available directions, and supervision is what buys new ones.
The greedy sweep -- what was tried, what was rejected
8 subsets, tau 0.1, seed 0, ranked by 8-task mean:
| subset | mean |
|---|---|
| equiv + simplify + topical | .7290 |
| equiv + simplify + paraphrase <- SHIPPED | .7287 |
| equiv + simplify | .7233 |
| equiv + simplify + paraphrase + random | .7224 |
| equiv + simplify + topical | .7223 |
| equiv + simplify + paraphrase + lexical | .7174 |
| equiv + simplify + random | .7157 |
| equiv + simplify + lexical | .7146 |
Two things worth stating plainly rather than burying:
- The
+topicalvariant scored .7290, marginally above the shipped .7287. That is a tie inside single-seed noise, and the smaller subset was taken. It is not a measured win for the shipped configuration. randomis a capacity control, and it is the informative row. Adding a randomly-paired anchor costs .0076 against the same-size real subset (.7157 vs .7233) and .0063 against the shipped one (.7224 vs .7287). Extra adapter capacity by itself does not buy score here -- the arms are earning it.
Why the dispatch works here
Dispatched amplitude is (w_k/z) * sigmoid(gate_k) * consume_k(x), where
w_k/z = sinh(u_k) / SUM_j cosh(u_j) over ALL anchors (the damping law).
That approaches ~1.0 only when one anchor engages and the others abstain
(u ~ 0); if several fire together it collapses and the mixture delivers a
fraction of what any single member does alone.
Measured on this collective: mean |w/z| moved from
.237 / .229 / .263 (blend) before alignment to
.587 / .138 / .142 (specialize) after the keys-only alignment phase, with no
starvation strikes. equiv won the amplitude budget and the other two learned
to abstain. That flip is why the mixture beats its best member instead of
damping itself below it.
Per-block routing spread is 0.348 -- routing is not uniform with depth;
equiv dominates most blocks, but not all.
Load
The model resolves its own arms. This is the supported path:
from transformers import AutoModel
model = AutoModel.from_pretrained("AbstractPhil/captionbert-8192-v2",
trust_remote_code=True)
model.attach_amoe() # this collective, by default
emb = model.encode(["a cat on a windowsill"])
with model.amoe_off(): # the unsupervised baseline
base = model.encode(["a cat on a windowsill"])
model.detach_amoe() # bit-exact restore, asserted
Explicitly, via amoe-lora:
import amoe
h = amoe.attach(trunk, ["amoe/collective/equiv.anchor.pt",
"amoe/collective/simplify.anchor.pt",
"amoe/collective/paraphrase.anchor.pt"],
dispatch="amoe/collective/captionbert-v2-collective.dispatch.pt",
binding=CaptionBertV2Binding(d=512)) # from modeling_captionbert.py
base = h.detach() # bit-exact or raises
The trunk these anchors expect is checkpoints/best_model.pt, which is what
the repo-root model.safetensors carries (verified tensor-for-tensor, max
absolute delta 0.0). Note that the sibling -B repo ships final_model.pt at
its root instead -- each repo's arms match its own root, but the two roots are
selected by different criteria.
Anchors are trunk-bound. These are captionbert-8192-v2 anchors. Measured
2026-08-02, they lose 31% of their gain on the -B trunk (.7287 -> .6863), and
re-aligning the routing keys alone recovers only 29% of that. Retrained natively
on -B they reach .7294. Do not port these to another trunk -- retrain them
(~18 min).
All anchors disabled reproduces the bare trunk bit-exact (asserted at build time). Masking never renormalizes -- that is the damping law, not an oversight.
Anchor .pt key layout is {block}.{param}. (blocks.{site}.{param} in the
amoe README is the safetensors layout, a different serializer.)
Training
Anchors: MNRL, in-batch + hard negatives where the source provides them, 1,500
steps at batch 256, pure Adam wd=0 (amoe.laws.make_optimizer), fp32 / TF32 off
(amoe.laws.pin_precision). Trunk frozen throughout.
1,500 steps and not 4,000: the first solo run's STS-B peaked at step 1,000 and then fell .0178 while SICK-R kept climbing -- 56,825 distinct anchors behind 200,000 draws (space/draws .284).
Dispatch: 800 steps, routing keys only, anchors frozen, starvation safeguard armed. tau 0.1, seed 0.
SICK-R and BIOSSES are never trained on and are the honest transfer reads.
Full configuration in config.json; every number above is in metrics.json.