Biogod's picture
Add source-balance screen preflight
619f401 verified
|
Raw
History Blame Contribute Delete
3.08 kB

Whole-Image Source-Balance Screen Protocol

Question and Claim Boundary

This development-only screen asks whether correcting target semantics and increasing raw-source coverage improves Qwen-Image-Edit adaptation. It uses only the CellSAM v1.2 official training split. The previously viewed 822-image official test contributes image hashes only; its labels are neither read nor used for selection. Results cannot establish blinded superiority.

Frozen Cohorts

  • dev82: 82 held-out images from 21 assessable raw sources. Twenty sources contribute four images; 2c_e_coli contributes two.
  • MM144-v2: 16 images from each of nine aggregate datasets.
  • MM320-v2: the maximum source-balanced, batch-compatible arm after the development holdout. The nominal 23×16 design has only 323 available images; deterministic batch alignment removes one image each from cellpose, dsb_fixed, and s2_stardist.
  • H&E sentinel24: a frozen PanNuke fold-2 retention check, not a replacement for the main H&E benchmark.

The two two-image sources, 2b_brightfield_dataset and 2b_fluorescence_dataset, were fully consumed by MM144-v1 and therefore have no independent development examples. They remain explicitly unassessable. Image hashes, image-label pair hashes, and FOV identities are disjoint across the development cohort and old/new training data. Identical label-only hashes, such as empty masks, are recorded but are not treated as leakage.

Controlled Intervention

BriFiSeg is corrected from bacterial_cell to nucleus in v2 only; v1 artifacts remain unchanged. Both arms initialize from the identical Mix128 rank-64 LoRA and use seed 20260727, LR 3e-5, effective batch 8, and 60 equivalent epochs. Images retain their native aspect ratio and follow the official Diffusers approximately 1-megapixel, 32-pixel-aligned geometry. No padding, cropping, tiling, stitching, augmentation, rank change, or backbone unfreezing is allowed.

The reused trainer writes the legacy phrase H&E in the human-readable training_protocol.json input description. This field does not control data loading. The frozen manifests, resolved configs, cache entries, and their hashes are authoritative and identify the actual multi-modal microscopy inputs. The trainer source is intentionally not changed between arms merely to revise this label, because doing so would break implementation-hash parity.

Frozen Evaluation and Decision

Inference uses four steps, seeds 42/314159/271828, k=3 intersection consensus, locked_color_v1, and Diffusers automatic geometry. The primary endpoint is equal-raw-source macro pooled Detection F1 with a source-stratified paired bootstrap (10,000 replicates). Secondary endpoints are PQ, AJI+, foreground Dice, count error, catastrophic-failure rate, and H&E retention.

MM320-v2 passes only if its macro-F1 gain over MM144-v2 is at least 0.03 with CI95 lower bound above zero, at least 15 assessable sources do not decline, catastrophic failures increase by at most 2 percentage points, and H&E pooled F1 drops by at most 0.03.