--- license: apache-2.0 library_name: pytorch pipeline_tag: image-feature-extraction tags: - vit - knowledge-distillation - geometric-deep-learning - alignment-bank - research datasets: - AbstractPhil/bulk-coco-features --- # geolip-vit-captionbank-coco The **CaptionBert system transplanted to vision**: a two-stage, modular design where the **core** is a reusable 8.66M-parameter ViT encoder distilled into a 5-teacher consensus space, and the **alignment bank** is a detachable geometric **expansion** on top of it. Trained on COCO-2017, two seeds, deterministic evals, and shipped with its complete ledgers — including the preregistered negative result on the bank (research honesty is the point of this repo, not a caveat to it). Sibling systems: the text-side original [geolip-captionbert-8192](https://huggingface.co/AbstractPhil/geolip-captionbert-8192), the constellation predecessor [geolip-vit-large-x3](https://huggingface.co/AbstractPhil/geolip-vit-large-x3), and the loss-matrix baseline [clip-vitb-mini-distilled](https://huggingface.co/AbstractPhil/clip-vitb-mini-distilled). ## Stage A — the core (reusable) `core/student_s{0,1}.pt` — ViT (d=240, depth 12, CLS readout, 160px) with a 512-d head into the **GPA consensus space** of five CLIP towers (B/16-LAION2B, B/32-OpenAI, B/32-LAION2B, B/32-DataComp, B/16-OpenAI; targets from [bulk-coco-features](https://huggingface.co/datasets/AbstractPhil/bulk-coco-features), 5-iteration generalized-Procrustes mean shape). Objective: symmetric InfoNCE (τ=0.07) + per-sample MSE, 1:1 — the CaptionBert Stage-A composite with the CV force held out as a readout. 8,000 steps (~17.3 epochs), batch 256, pure Adam 3e-4. Measured (per seed, deterministic to 4 decimals): | gauge | s0 | s1 | |---|---|---| | agreement → consensus | .6717 | .6700 | | zero-shot CIFAR-10 (rotated to CLIP frame) | .2318 | .2636 | | COCO R@1 (rotated) | .0598 | .0530 | | 80-class linear-probe mAP (raw 512-d) | .4007 | .3943 | **The agreement trade** (this repo's positive finding): adding InfoNCE to a consensus-MSE distillation buys contrastive-class retrieval (the pure-MSE twin sits at R@1 ≈ .004) at the price of ~17% of its consensus agreement (.67 vs .805). ## Stage B/C — the alignment bank (expansion) `banks/bank_s{0,1}.pt` — per-expert whitened-Procrustes frames (fp64 fits, gated), 512 dense-cosine anchors seeded from consensus rows, and a 538→256→128 projection; `enrich()` appends the 128-d geometric context → 640-d. `cotrained/` holds the Stage-C jointly-tuned pairs. `frames.pt` has the raw fitted frames + gate record. **Preregistered probe verdict — the bank's geometry training is REFUTED on this substrate, with full controls (2/2 seeds):** | probe input | s0 mAP | s1 mAP | |---|---|---| | raw 512-d | .4007 | .3943 | | fully-random bank (random frames + anchors) | .4159 | .4114 | | real fitted frames, untrained projection | .4149 | .4107 | | real frames, geometry-trained | .4111 | .4062 | The +1.4–1.6pt enrichment over raw is **generic random-feature capacity**: random frames match fitted frames, and the 8-term geometry training *subtracts* slightly. Anyone using the bank should know the enrichment's value here did not come from the geometry — the bank ships because the architecture is reusable and the negative is load-bearing information. (The text-side original reported +0.125 on a 3-class NLI probe; that bed is untested here — this result is about transfer, not their claim.) ## Usage ```python import torch from PIL import Image from huggingface_hub import hf_hub_download repo = "AbstractPhil/geolip-vit-captionbank-coco" ldr = hf_hub_download(repo, "loader.py") import importlib.util spec = importlib.util.spec_from_file_location("loader", ldr) L = importlib.util.module_from_spec(spec) spec.loader.exec_module(L) student = L.load_student(hf_hub_download(repo, "core/student_s0.pt")) emb = L.embed_images(student, [Image.open("cat.jpg")]) # (1, 512) bank = L.load_bank(hf_hub_download(repo, "banks/bank_s0.pt")) enriched = L.enrich(bank, emb) # (1, 640) ``` The 512-d embedding lives in the 5-CLIP consensus frame — for CLIP-text zero-shot, apply a fitted rotation into a CLIP frame first (the ledgers record the rotated numbers; the rotation is a 2,500-pair orthogonal Procrustes fit). ## Deviations from the source recipe (disclosed) Pure Adam constant-lr (source: AdamW+cosine); CV force as readout in Stage A and gradient-excluded in Stage C (source: 0.1 / in-graph); per-sample MSE so the 1:1 weighting is real (the source's per-element MSE is ~1000× smaller than its InfoNCE term). A faithful full-replication arm (source objectives AND optimizer, verbatim) plus an objective-isolate matrix (BCE force, geometric-gradient operator, CV force) is training at time of writing and will be added to this repo with its results. ## Files `core/` Stage-A students · `banks/` Stage-B banks · `cotrained/` Stage-C pairs · `frames.pt` fitted expert frames + gate record · `loader.py` standalone two-stage loader (fidelity-gated: reproduces training-side embeddings to 3e-7) · `ledgers/` complete JSONL run records including the random-bank control.