# scale_exp — data-scaling experiments (2026-09) Monet-style latent visual reasoning on Qwen2.5-VL-7B-Instruct. Every model here is the **stage-3 root model** of the mod-ablation recipe: stage 2 1 epoch + stage 3 1 epoch, lr 1e-5, linear warmup 10 steps, effective batch 128, latent size 8, alignment weight 2.0, emphasize-latent 2.0, CE-emphasize 4.0, all-layer alignment. Trained on ACD H100s with world size 8, GA 16 (e1B10k/20k/30k/40k: 1 node × 8 GPUs; e1B40kR / e2K3 / e2K5 / e2K8: 2 nodes × 4 GPUs). `MODEL_MD5` lists the md5 of every file. 8 forms: scene_graph, segmentation, text_cot, depth, helper_interleaved, bbox_crop, bbox_highlight, edge. | folder | experiment | forms | rows | |---|---|---|---| | e1B10k / e1B20k / e1B30k / e1B40k | Exp1: fixed 8 forms, nested subsets of size-40k | all 8 | 1,250 / 2,500 / 3,750 / 5,000 per form | | e1B40kR | Exp1 40k retrain (training-noise replicate) | all 8 | 39,999 | | e2K3 | Exp2 chain K3 | scene_graph, segmentation, text_cot | 120,000 (40k/form) | | e2K5 | Exp2 chain K5 | K3 + depth, helper_interleaved | 200,000 (40k/form) | | e2K8 | Exp2 chain K8 | all 8 | 319,994 (40k/form; bbox_crop 39,994) | 8-bench macro (V*, HRBench4K/8K, MME-RW-Lite, BLINK, POPE-F1, RealWorldQA, CV-Bench (2D+3D)/2; temp 0): | model | Exp1 (Nautilus A100) | Exp2 (UCSD A6000) | |---|---|---| | Qwen2.5-VL-7B base | 67.80 | 67.86 | | e1B10k | 69.11 | | | e1B20k | 69.75 | | | e1B30k | 69.38 | | | e1B40k | 71.11 | 71.04 | | e1B40kR | | 70.86 | | e2K3 | | 70.81 | | e2K5 | | 70.86 | | e2K8 | | 70.25 | Raw predictions are archived separately (private).