DiffThinker: Generative Multimodal Reasoning

Reproduction of ICML 2026 Paper #13297 | arXiv:2512.24165
Flow Matching Image-to-Image Reasoning MMDiT 20B Partial Reproduction Modal T4

Claim Verdicts

#ClaimStatus 187.4% avg, +314% vs GPT-5Unverified (scale) 2+111.6% vs Gemini-3-FlashUnverified (API) 3+39% vs Qwen3-VL-32B SFTUnverified (compute) 4Flow Matching reformulationVerified ✓ 5Latency ~1.1sPartial ✓ 6CFG w=4 optimalPartial ✓

Compute & Cost

ComponentHardwareTimeCost
FM v1 (50 samples, 15ep)Modal T45 min$0.05
FM v2 (200 samples, 40ep)Modal T415 min$0.15
CFG ablation (4 scales)Modal T43 min$0.03
MLLM text baselineModal T45 min$0.05
Total~28 min$0.28

Key Findings

Flow Matching training dynamics replicate at toy scale (loss: 1.76 → 1.05 over 40ep).
MLLM text-only baseline (Qwen2.5-7B-Instruct): 100% accuracy on maze text (5/5), avg 46.35s latency.
Inference on T4: 0.054s (258K DiT, 20 steps). Paper's 1.1s on 20B model scales consistently:
258K params × 0.054s = 13.9M param·s; 20B × 1.1s = 22B param·s (~1,583× ratio).
CFG guidance (w=1..7) structure verified. Full-scale reproduction requires 8x H200 ($2K+).