Upload zimg-fdloss/FINAL_TABLE.md with huggingface_hub
Browse files- zimg-fdloss/FINAL_TABLE.md +17 -0
zimg-fdloss/FINAL_TABLE.md
CHANGED
|
@@ -40,6 +40,21 @@
|
|
| 40 |
| T2 (TFC EMA targets β inert) | Base | iMF 3k β FD+TFC 2k | SIMcx + TFC(EMA.999) | data ; rand | 0 | 0.836 | 0.882 | 1.561 | 1.651 |
|
| 41 |
| T3 (TFC finishing pass, seed1) | Base | β¦D1c β FD+TFC 500 | SIMcx + TFC(w.5, lr3e-6) | data ; rand | 0 | 0.860 | 0.891 | 1.545 | 1.630 |
|
| 42 |
| T3b (finishing pass, seed2) | Base | β¦D1c β FD+TFC 500 | SIMcx + TFC(w.5, lr3e-6) | data ; rand | 0 | 0.857 | 0.881 | β | β |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 43 |
|
| 44 |
## Champion vs baselines
|
| 45 |
|
|
@@ -49,6 +64,8 @@
|
|
| 49 |
|
| 50 |
**β‘ HOLDOUT-LEAK CAVEAT (found 2026-08-02 by code review)**: the original iMF data loader applied the holdout modulus to a shuffled enumeration, so pre-fix iMF-line arms (M, Mg*, V1*, V2βV4, E2, D3 β all rows above g0c that include an iMF stage) trained on ~11/12 of the true ref holdout: their eval-prompt FDr/PickScore carry an optimistic bias (GenEval unaffected β disjoint prompts; turbo-line arms unaffected). g0c/D1c/D2c are post-fix clean reruns; the clean g0c rerun reproduces the old champion within noise @8 (McNemar z=β1.47) and lands lower @1 (z=β4.78) β the V1g0@1 number was partly run-specific. Within-line contrasts among β‘ rows (e.g. D3 vs V1g0) remain internally valid.
|
| 51 |
|
|
|
|
|
|
|
| 52 |
**TFC / redistribution law (2026-08-04)**: TFC = tie the 1-jump endpoint's SigLIP2 features to sg(Ο(the model's own multi-step rollout endpoint)) β self-distillation of the self-map, externally anchored. Three configs Γ 2 seeds establish that map-consistency pressure REDISTRIBUTES GenEval between NFE operating points and never creates it: live targets full-run (T1) = hard equalization (@1 record 0.879, z=+3.16; @8 z=β5.53); EMA targets (T2) = inert; champion-init 500-step finishing pass (T3/T3b) = best exchange rate (@1 +2.31z/+1.53z; @8 β1.24z/β2.57z; ~1.2 points gained per point lost). Self-referential signals carry no new information β they shape, external anchors (data, frozen reps) create. TFC weight/horizon = a deployment NFE-preference dial on one recipe.
|
| 53 |
|
| 54 |
**Self-Flow probe (arXiv 2603.06507, 2026-08-03)**: generator-self-features fail in BOTH roles. As FD metric: pure self-FD optimizes its own objective (8259β850) while every external signal collapses (GenEval@1 0.835β0.289, z=β23.7 β the init's 1-step capability is destroyed); as a 5th additive term it is a net drag (z=β3.6/β7.8). As iMF self-distillation aux (EMA teacher, 25% token noise asymmetry): rep-cos saturates β1.0 by step 500 (signal goes vacuous), GenEval@8 tie (z=β0.83), @1 significantly worse (z=β3.13). Replicates the video program's self-feature negative across modality and Ο regimes: the self-encoder's null space contains the generator's own failure modes β external frozen reps are load-bearing. (Cross-validation: Self-Flow's own T2I ablation found uniform > logit-normal timestep dist β matches our D2c rejection.)
|
|
|
|
| 40 |
| T2 (TFC EMA targets β inert) | Base | iMF 3k β FD+TFC 2k | SIMcx + TFC(EMA.999) | data ; rand | 0 | 0.836 | 0.882 | 1.561 | 1.651 |
|
| 41 |
| T3 (TFC finishing pass, seed1) | Base | β¦D1c β FD+TFC 500 | SIMcx + TFC(w.5, lr3e-6) | data ; rand | 0 | 0.860 | 0.891 | 1.545 | 1.630 |
|
| 42 |
| T3b (finishing pass, seed2) | Base | β¦D1c β FD+TFC 500 | SIMcx + TFC(w.5, lr3e-6) | data ; rand | 0 | 0.857 | 0.881 | β | β |
|
| 43 |
+
| T4a (finishing 250s) [near-Pareto] | Base | β¦D1c β FD+TFC 250 | SIMcx + TFC(w.5) | data ; rand | 0 | 0.857 | 0.901 | β | β |
|
| 44 |
+
| T4b (finishing 500s, w.25) | Base | β¦D1c β FD+TFC 500 | SIMcx + TFC(w.25) | data ; rand | 0 | 0.861 | 0.898 | β | β |
|
| 45 |
+
| T4c (finishing 1000s) | Base | β¦D1c β FD+TFC 1000 | SIMcx + TFC(w.5) | data ; rand | 0 | 0.850 | 0.847 | β | β |
|
| 46 |
+
| T5 (finishing on TURBO line β inert) | Turbo | β¦SIMcx β FD+TFC 500 | SIMcx + TFC(w.5) | rand | β | 0.807 | 0.810 | β | β |
|
| 47 |
+
| CA / CA2 (NEG: guidance-field mining) | Base | β¦D1c β FD+CA 1000 | SIMcx + CA regression | data ; rand | 0 | 0.832 | 0.882 | 1.686 | 1.911 |
|
| 48 |
+
| SIMF / SIMB β‘ (packing probe, no coupling) | Base | iMF 3k β FD 2k | siglip2fix/legacy + incep + mae | data ; rand | 0 | 0.726 | 0.847 | 1.328 | 1.733 |
|
| 49 |
+
| **DR0 (champion @ batch 64)** | Base | iMF 3k β FD 2k @4node | SIMcx, b64 moments | data ; rand | 0 | 0.846 | 0.910 | 1.447 | 1.621 |
|
| 50 |
+
| DR1 (NEG: + kernel drift, siglip2) | Base | iMF 3k β FD 2k @4node | SIMcx + drift(siglip2) | data ; rand | 0 | 0.820 | 0.888 | 1.468 | 1.570 |
|
| 51 |
+
| DR2 (NEG: + kernel drift, self) | Base | iMF 3k β FD 2k @4node | SIMcx + drift(self) | data ; rand | 0 | 0.647 | 0.831 | 1.905 | 2.291 |
|
| 52 |
+
| RDMD (NEG: learned score-diff v1) | Base | iMF 3k β FD 2k | SIMcx + rep-DMD(w3e-5) | data ; rand | 0 | 0.788 | 0.868 | 1.506 | 1.674 |
|
| 53 |
+
| RDMD2c (evidence-gated β exact tie) | Base | iMF 3k β FD 2k | SIMcx + rep-DMD v2c(gated) | data ; rand | 0 | 0.844 | 0.902 | β | β |
|
| 54 |
+
| TFDHA (NEG: adversarial heads) | Base | iMF 3k β FD+adv 2k | SIMcx + hinge heads(w.01) | data ; rand | 0 | 0.702 | 0.866 | β | β |
|
| 55 |
+
| TFDHD (adaptive-space drift, seed1) | Base | iMF 3k β FD+drift 2k | SIMcx + head-space drift | data ; rand | 0 | 0.860 | 0.905 | 1.474 | 1.618 |
|
| 56 |
+
| TFDHDb (seed2 β refutes seed1) | Base | iMF 3k β FD+drift 2k | SIMcx + head-space drift | data ; rand | 0 | 0.850 | 0.878 | β | β |
|
| 57 |
+
| K64 (coupling dims 32β64 β inert) | Base | iMF 3k β FD 2k | SIMcx w/ CCA-64 pack | data ; rand | 0 | 0.820 | 0.887 | 1.517 | 1.666 |
|
| 58 |
|
| 59 |
## Champion vs baselines
|
| 60 |
|
|
|
|
| 64 |
|
| 65 |
**β‘ HOLDOUT-LEAK CAVEAT (found 2026-08-02 by code review)**: the original iMF data loader applied the holdout modulus to a shuffled enumeration, so pre-fix iMF-line arms (M, Mg*, V1*, V2βV4, E2, D3 β all rows above g0c that include an iMF stage) trained on ~11/12 of the true ref holdout: their eval-prompt FDr/PickScore carry an optimistic bias (GenEval unaffected β disjoint prompts; turbo-line arms unaffected). g0c/D1c/D2c are post-fix clean reruns; the clean g0c rerun reproduces the old champion within noise @8 (McNemar z=β1.47) and lands lower @1 (z=β4.78) β the V1g0@1 number was partly run-specific. Within-line contrasts among β‘ rows (e.g. D3 vs V1g0) remain internally valid.
|
| 66 |
|
| 67 |
+
**FINAL RECIPE (frozen 2026-08-04)**: base β iMF (g=0, NO explicit r=0 group, uniform+shift-3, 3k) β coupled FD post-training (SIMcx, K=32 CCA + XCov, 2k; batch 64 if multi-node available β the sole surviving upgrade, nominal n.s. GenEval + real FDr@1 gain) + optional TFC finishing dial for NFE preference. **Estimator-matrix closure (~20 arms)**: kernel drift (3 feature spaces), learned score-difference (plain + evidence-gated: gate opened 2.0%% of steps, exact champion tie), adversarial heads (gradient runaway 0.2Γβ400Γ), adaptive-space drift (seed-2 refuted), K=64 coupling (offline +14%% power did not transfer) β all negative/inert/unreplicated. **Measured information law**: extractable population signal beyond two moments in this feature space is ~1%% of the FM loss and moment-redundant (closed-form moment field beats the best learnable field, slope β3.48 vs β2.62). Net progress requires new information (data/teacher/reward), not richer estimators.
|
| 68 |
+
|
| 69 |
**TFC / redistribution law (2026-08-04)**: TFC = tie the 1-jump endpoint's SigLIP2 features to sg(Ο(the model's own multi-step rollout endpoint)) β self-distillation of the self-map, externally anchored. Three configs Γ 2 seeds establish that map-consistency pressure REDISTRIBUTES GenEval between NFE operating points and never creates it: live targets full-run (T1) = hard equalization (@1 record 0.879, z=+3.16; @8 z=β5.53); EMA targets (T2) = inert; champion-init 500-step finishing pass (T3/T3b) = best exchange rate (@1 +2.31z/+1.53z; @8 β1.24z/β2.57z; ~1.2 points gained per point lost). Self-referential signals carry no new information β they shape, external anchors (data, frozen reps) create. TFC weight/horizon = a deployment NFE-preference dial on one recipe.
|
| 70 |
|
| 71 |
**Self-Flow probe (arXiv 2603.06507, 2026-08-03)**: generator-self-features fail in BOTH roles. As FD metric: pure self-FD optimizes its own objective (8259β850) while every external signal collapses (GenEval@1 0.835β0.289, z=β23.7 β the init's 1-step capability is destroyed); as a 5th additive term it is a net drag (z=β3.6/β7.8). As iMF self-distillation aux (EMA teacher, 25% token noise asymmetry): rep-cos saturates β1.0 by step 500 (signal goes vacuous), GenEval@8 tie (z=β0.83), @1 significantly worse (z=β3.13). Replicates the video program's self-feature negative across modality and Ο regimes: the self-encoder's null space contains the generator's own failure modes β external frozen reps are load-bearing. (Cross-validation: Self-Flow's own T2I ablation found uniform > logit-normal timestep dist β matches our D2c rejection.)
|