Negative result (a contribution): rank-based coverage objectives fail for rare-lesion retention
Claim
Effective-rank / coding-rate "coverage" objectives β RankMe, coding rate (MCR2-style) β are structurally mismatched to retaining rare, small-region pathology under token pruning. Using them as the pruning objective is worse than simply ranking tokens by lesion-subspace membership.
Mechanism (the transferable part)
A rank-based coverage functional C(S) = effrank(P_L Z_S) is maximized by a retained set that
diversely spans the lesion subspace's directions. The decisive property of a small lesion is
that it is rare β a few high-membership tokens out of ~196. A set-level rank/coverage
objective is insensitive to such a cluster: a handful of tokens cannot materially raise the retained
set's effective rank, so the objective spends budget on abundant background directions and drops the
lesion. This is a rarity mechanism, not low internal geometry β measured at the operating layer,
lesion tokens are not low-rank relative to background (pooled effective rank 339 vs 307;
participation ratio 18.9 vs 13.9; research_v4/lesion_spectrum.json). Concentration, not spanning,
is what rare-pathology retention needs.
Formally: rank coverage rewards the entropy of the retained set's singular spectrum; lesion
retention rewards mass on the top membership tokens. These objectives diverge whenever the
critical signal is a rare cluster β of any internal rank. (The synthetic closed-form law of the
companion paper isolates a second, distinct route β a genuinely low-rank signal, gap (m-r)/m; real
lesions fail via rarity, not low rank.)
Three independent lines of evidence (same verdict)
- Ablation (decisive). At matched budget, the coverage-floor pruner retains 0.22 vs 0.82
(budget 0.25) and 0.46 vs 0.98 (budget 0.5) of small lesions vs membership top-k; the
difference CI excludes 0. The floor does not under-help β it actively hurts. (
ablation_floor.json) - Faithfulness (principled Gate 2). Under the random-pruning protocol, coverage-drop predicts
lesion-detection-drop no better than attention-drop: Ο 0.480 vs 0.479, difference CI includes 0.
Coverage is not a superior proxy. (
gate_2_baseline.json) - Adaptive budget (Gate 4). The "difficulty-adaptive budget" never materializes: aggregate
coverage C* is the same on lesion-positive and -negative slices (250.4 vs 247.2), because a
1β3 patch lesion cannot move an aggregate over ~196 tokens. (
gate_4_block3.json)
All three reduce to one fact: aggregate rank-coverage is blind to the few tokens that carry a small lesion, even though those tokens are individually highly localizable (Gate 1, AUROC 0.87).
Closed-form law (S1, confirmed by controlled synthetic)
A controlled synthetic isolates the mechanism and yields a closed form (research_v2/s1_crossover.json).
Inject a signal of effective rank r across m tokens; select to a budget by a spanning objective
(farthest-point / effective-rank) vs a concentration objective (top-energy):
- spanning retention(r) = min(r, m) / m, concentration retention = 1,
- gap(r) = max(0, (m β r) / m), crossover r* = m.
A spanning objective retains a signal only in proportion to its rank; it matches concentration only when the signal is fully diverse (rank = token count). Rare pathology is maximally concentrated (r β 1β3 patches, m small), so rank-based coverage is maximally mismatched there β quantitatively reproducing the covtoken ablation (floor 0.22 vs membership 0.82 β gap (mβ1)/m for the dominant 1-patch lesions). The verdict is no longer anecdotal: it is a closed-form prediction.
Why this matters beyond this paper
RankMe-flavored objectives are an increasingly common, tempting choice for medical SSL representation quality and for "coverage"-style regularizers. This result is a warning with a mechanism: for rare-pathology tasks, prefer concentration objectives (energy / membership mass) over rank/spanning objectives. A negative result with a transferable mechanism is citable; "our dual didn't converge" is not. This is the former.
What survives
The label-free lesion subspace (the geometry that produces membership) is intact and is the contribution. The failure is specifically the rank-coverage functional built on top of it and the constrained-optimization machinery that optimized it. Replacing the objective with the membership/energy quantity recovers the result β but then the "constraint + dual" adds nothing over a top-k rule, so it is dropped honestly rather than dressed up.