Buckets:
STATUS — LOGDIFF (OAM1jJsMGp)
Session: full-score campaign. Last updated: 2026-07-20. State: in progress; official score 6/10 and exact molecular runs active. The first four full 1,000-step AND shards are validated with 4/4 chemically reconstructable, distinct molecules and 8/8 successful generated-ligand dockings; the scoped checkpoint is public while shards 4--7 continue.
Public Space: DineshAI/OAM1jJsMGp at current revision 38daa9517ef7af65d409ffd0325cba289e969014; the official verdict below was issued against this exact revision.
Fresh official verdict: 6/10, high quality, at judged Space SHA 38daa9517ef7af65d409ffd0325cba289e969014 (2026-07-20T01:09:46+00:00); C1/C5 are verified, C2 is falsified, and C3/C4 remain inconclusive. The new C4 feedback explicitly cites the 4/32 scope, the below-target 64.53 product, and the unstarted AND-NOT campaign, confirming that both complete campaigns remain required.
GitHub remains MachineLearning-Nerd/icml26-repro-OAM1jJsMGp-logdiff-exact-boolean-guidance at 6449752; the current campaign changes are intentionally uncommitted.
Anchored evidence
- C1 Proposition 3.1: verified by exhaustive finite-state oracles and real trained-UNet algebra checks.
- C2 CMNIST/Shapes3D Table 2 range statement: falsified as written by the paper table itself; only 2/8 LOGDIFF and 5/8 constant-baseline cells fall inside the stated ranges.
- C3 CelebA Table 3: table transcription matches, but the pinned author release lacks all nine declared checkpoints, has a broken config reference, and uses a different FID implementation/protocol.
- C4 GRM5/RRM1 Tables 5-6: exact targets, all three checkpoint hashes, the real TargetDiff sampler, and AutoDock Vina are recovered; full generated-ligand composition remains active.
- C5 Proposition C.2: verified over 20,650 independent-product events and 6,350 taxonomy events, with invalid overlap rejected.
Source-scale runs
- Exact upstream CMNIST composition classifier completed 50 epochs on the full 60,000-example split.
- Fixed-timestep evaluation completed over the full 10,000-example test split at seven timesteps.
- Independent exact upstream clean-judge training and full 10,000-example evaluation completed. The checkpoint SHA-256 is
ff8dbb5253278ad46e0033959045f9558927b0e814cd04f98fc747d6629adfb6; digit/colour accuracies are0.9814/1.0and digit/colour NLLs are0.0867544/0.000025857. - Molecular recovery includes verified DualDiff data, GRM5/RRM1 target pair, pretrained weights, a real sampler smoke run, successful reference-complex Vina gates, and an executed generated-ligand docking/pocket-alignment smoke path for both targets. The smoke is explicitly not claim evidence because it uses 1/1000 denoising steps and has 1/32 validity. The exact Appendix-D.3 transition-alpha bug is repaired, the beta-2 restartable runner passed four-worker 80-step and eight-worker 80-step gates, and the 32 x 1000-step AND run is active with four workers and ten retries per shard. Samples 0--3 completed on their first attempts after approximately 3.5 hours each; all four have finite 1,000-step posterior trajectories, valid reconstructions, and distinct SMILES. A hash-validated partial merge (
b0cf5223...) permits honest interim docking without altering the continuing 32-sample campaign. - The four-ligand checkpoint completed paired
vina_fulldocking at exhaustiveness 32: 8/8 generated-ligand target evaluations succeeded; mean GRM5/RRM1 affinities are-8.00375/-8.06675kcal/mol, mean AND product is64.53267625, diversity is0.887054, valid-unique fraction is1.0, and paper-defined quality fraction is0.5. Same-run reference affinities are-5.866/-7.640. The product is below the paper's LOGDIFF73.20 +/- 3.18and DualDiff71.87 +/- 3.33, so it is published as a partial execution checkpoint rather than headline verification. AND-NOT separation is not inferred from AND samples. - Official CelebA is complete and verified at 202,599 images; the released clean-filter splits contain 100,247 train, 12,459 validation, and 12,186 test images. Public DDPM
shalpin87/diffusion_celeba@91d0ff0is pinned locally and clean-fid 0.1.35 is installed for the independent C3 route.
Publication and verification
- Static Space is public and RUNNING; Hub and app return HTTP 200.
- Exactly the intended 11-page hierarchy is present, with Conclusion last.
- Required tags, pinned executive summary/poster, artifact cell, and public evidence bucket are present.
- Local structural validation passes except the pre-existing legacy Space-name rule requiring a new
repro-*identifier; no duplicate Space was created. - Remote secret and absolute-path leakage scans pass.
- The independent CMNIST judge page was synced and read back at current Space SHA
93d90f36; its remote SHA-256 exactly matches the local page (840c7bef...). - The Claim 4 checkpoint page was synced and byte-for-byte read back at public Space SHA
38daa951; local and remote SHA-256 are both7ac3c77c.... The public checkpoint artifact also read back with matching SHA-256c44cd570...; its unauthenticated download endpoint returns a public Xet redirect. All 29 tests pass.
Next
Continue the full 32-ligand AND campaign, then merge/dock it and run the separate AND-NOT generation and docking campaign. In parallel when resources permit, run the independent real-CelebA negation route. Update the evidence bundle/logbook with complete metrics and request re-judging. Do not stop on failures; materially different recovery approaches continue until the anchored claims reach their maximum defensible score.
Xet Storage Details
- Size:
- 5.71 kB
- Xet hash:
- 4fbfdbf3fcf2fbaec61623cc154f63a5d7cb8c55a627320f0ca01a0d8ae6c33f
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.