AbstractPhil's picture
final: 2 anchors + dispatch + card
9154293 verified
|
Raw
History Blame Contribute Delete
2.75 kB
metadata
license: mit
tags:
  - amoe
  - adapter
  - mixture-of-experts
  - sentence-similarity
  - captionbert
  - aleph
base_model: AbstractPhil/captionbert-8192-v2
library_name: amoe-lora

captionbert-8192-v2 :: AMOE 2-anchor mixture

Two amoe-lora aleph anchors on the frozen trunk, plus a trained dispatch over them. The trunk never moves.

anchor trained on relation
equiv all-nli triplets semantic equivalence
simplify simple-wiki + altlex + sentence-compression simplification / compression

Results

config STS-B rho SICK-R rho
bare trunk .5747 .6526
equiv alone .7254 .7550
simplify alone .7400 .7075
2-anchor dispatch .7524 .7380

The two are complementary along the TASK axis -- simplify wins STS-B solo, equiv wins SICK-R -- which is the precondition a mixture needs. SICK-R is never trained on and is the honest transfer read.

Why the dispatch works here

Dispatched amplitude is (w_k/z) * sigmoid(gate_k) * consume_k(x), where w_k/z = sinh(u_k) / SUM_j cosh(u_j) over ALL anchors (the damping law). That gives ~1.0 only when one anchor engages and the other abstains (u ~ 0); if both fire it collapses to ~0.5 and the mixture delivers HALF of what either member does alone.

Measured here: mean |w/z| moved from .310/.380 (blend) before alignment to .645/.223 (specialize) after 800 keys-only steps, with no starvation strikes. That flip is why the mixture beats its best member rather than damping itself below it.

Load

import amoe
h = amoe.attach(trunk, ["amoe/moe/equiv.anchor.pt", "amoe/moe/simplify.anchor.pt"],
                dispatch="amoe/moe/captionbert-v2-moe.dispatch.pt",
                binding=CaptionBertV2Binding(d=512))   # from modeling_captionbert.py
base = h.detach()        # bit-exact or raises

All anchors disabled reproduces the bare trunk bit-exact (asserted at build time). Masking never renormalizes -- that is the damping law, not an oversight.

Anchor .pt key layout is {block}.{param}. (blocks.{site}.{param} in the amoe README is the safetensors layout, a different serializer.)

Training

Anchors: MNRL, in-batch + hard negatives where the source has them, 1,500 steps at batch 256, pure Adam wd=0 (amoe.laws.make_optimizer), fp32/TF32 off. 1,500 not 4,000: the first solo run's STS-B peaked at step 1,000 and then fell .0178 while SICK-R kept climbing -- 56,825 distinct anchors behind 200,000 draws (space/draws .284). Dispatch: 800 steps, routing keys only (1,536 params), anchors frozen, starvation safeguard armed.

See metrics.json for the full table and the routing telemetry.