# Whole-Image Source-Balance Screen Protocol ## Question and Claim Boundary This development-only screen asks whether correcting target semantics and increasing raw-source coverage improves Qwen-Image-Edit adaptation. It uses only the CellSAM v1.2 official training split. The previously viewed 822-image official test contributes image hashes only; its labels are neither read nor used for selection. Results cannot establish blinded superiority. ## Frozen Cohorts - `dev82`: 82 held-out images from 21 assessable raw sources. Twenty sources contribute four images; `2c_e_coli` contributes two. - `MM144-v2`: 16 images from each of nine aggregate datasets. - `MM320-v2`: the maximum source-balanced, batch-compatible arm after the development holdout. The nominal 23×16 design has only 323 available images; deterministic batch alignment removes one image each from `cellpose`, `dsb_fixed`, and `s2_stardist`. - `H&E sentinel24`: a frozen PanNuke fold-2 retention check, not a replacement for the main H&E benchmark. The two two-image sources, `2b_brightfield_dataset` and `2b_fluorescence_dataset`, were fully consumed by MM144-v1 and therefore have no independent development examples. They remain explicitly unassessable. Image hashes, image-label pair hashes, and FOV identities are disjoint across the development cohort and old/new training data. Identical label-only hashes, such as empty masks, are recorded but are not treated as leakage. ## Controlled Intervention BriFiSeg is corrected from `bacterial_cell` to `nucleus` in v2 only; v1 artifacts remain unchanged. Both arms initialize from the identical Mix128 rank-64 LoRA and use seed 20260727, LR 3e-5, effective batch 8, and 60 equivalent epochs. Images retain their native aspect ratio and follow the official Diffusers approximately 1-megapixel, 32-pixel-aligned geometry. No padding, cropping, tiling, stitching, augmentation, rank change, or backbone unfreezing is allowed. The reused trainer writes the legacy phrase `H&E` in the human-readable `training_protocol.json` input description. This field does not control data loading. The frozen manifests, resolved configs, cache entries, and their hashes are authoritative and identify the actual multi-modal microscopy inputs. The trainer source is intentionally not changed between arms merely to revise this label, because doing so would break implementation-hash parity. ## Frozen Evaluation and Decision Inference uses four steps, seeds 42/314159/271828, k=3 intersection consensus, `locked_color_v1`, and Diffusers automatic geometry. The primary endpoint is equal-raw-source macro pooled Detection F1 with a source-stratified paired bootstrap (10,000 replicates). Secondary endpoints are PQ, AJI+, foreground Dice, count error, catastrophic-failure rate, and H&E retention. MM320-v2 passes only if its macro-F1 gain over MM144-v2 is at least 0.03 with CI95 lower bound above zero, at least 15 assessable sources do not decline, catastrophic failures increase by at most 2 percentage points, and H&E pooled F1 drops by at most 0.03.