| # PANSEGX post-processing — false-positive reduction |
|
|
| > **Optional** step applied downstream of ONNX/TensorRT inference, on the full |
| > reassembled probability volume. It changes neither the weights nor the graph: |
| > it cleans up the output. See [`postprocess.py`](postprocess.py). |
|
|
| ## Why |
|
|
| The original checkpoint ships **no** post-processing (`postprocessing.pkl` was |
| never computed), and the reference inference command runs without TTA |
| (`--disable_tta`). Its dominant failure mode is documented: **scattered |
| false positives** rather than missed lesions. |
|
|
| | Pattern (35 cases with DSC < 0.5) | Frequency | |
| |---|---| |
| | FRAGMENTED (≥3× more components than GT) | 91 % | |
| | OVER_SEG (predicted volume >3× GT) | 74 % | |
| | LOCATION_OFFSET (centroid >30 mm) | 100 % | |
| | NO_OVERLAP (precision & recall <0.1) | 14 % | |
| |
| Removing the "fog" of spurious components mechanically fixes fragmentation, |
| over-segmentation and centroid drift. NO_OVERLAP (14 %) stems from partial |
| labels at training time and is **out of reach** for any post-processing. |
|
|
| ## Core idea: connected-component hysteresis |
|
|
| A plain `argmax` (prob ≥ 0.5) keeps all of the fog, which precisely crossed 0.5. |
| Raising a single global threshold erodes the boundary of true lesions and hurts |
| **NSD**. The right answer is hysteresis: |
|
|
| - **`t_low` (low)** defines the candidate extent → preserves boundaries (good for NSD); |
| - **`t_high` (high)** defines confident "seed" cores; |
| - a component of the low mask is kept only if it contains **at least one seed**. |
|
|
| True lesions almost always have a highly confident core; the false-positive fog |
| does not. Two physical guards are added: **minimum volume in mm³** (via the |
| spacing) and **minimum core confidence**. |
|
|
| > ⚠️ **No "keep the largest component".** This is pan-cancer segmentation: |
| > multiple metastatic lesions are legitimate. The module is **multi-lesion by |
| > default** (`keep_top_k = None`). |
|
|
| ## Parameters |
|
|
| | Parameter | Default | Role | |
| |---|---|---| |
| | `t_low` | 0.50 | low threshold — candidate extent (preserves boundaries) | |
| | `t_high` | 0.90 | high threshold — confident seed core (hysteresis) | |
| | `min_volume_mm3` | 50.0 | minimum component volume (physical scale) | |
| | `min_core_prob` | 0.90 | minimum peak probability required within a component | |
| | `min_mean_prob` | 0.0 | minimum mean probability (0 = disabled) | |
| | `connectivity` | 26 | 3D neighbourhood (6 / 18 / 26) | |
| | `fill_holes` | True | fills internal holes (DSC gain, boundary-neutral) | |
| | `keep_top_k` | None | keep only the *k* largest (None = all) | |
| | `drop_border_components` | False | drop components touching the FOV border | |
|
|
| > **Defaults are starting points, not optima.** Lock them in via calibration |
| > (below) so recall is not degraded. |
|
|
| ## Usage |
|
|
| ### Command line (NIfTI) |
|
|
| ```bash |
| python postprocess.py \ |
| --input prob.nii.gz \ # tumor probability map (Z,Y,X) OR logits |
| --output mask.nii.gz \ |
| --t-high 0.90 --min-volume-mm3 50 \ |
| --report # per-component diagnostics (optional) |
| ``` |
|
|
| Spacing is read from the NIfTI header, so volumes are computed in mm³. |
|
|
| ### Python API |
|
|
| ```python |
| from postprocess import postprocess, PostProcConfig |
| |
| # ONNX logits of shape (2, Z, Y, X) OR a probability map (Z, Y, X) |
| mask = postprocess( |
| logits_or_prob, |
| spacing_zyx=(1.0, 0.787, 0.787), # (z, y, x) in mm |
| config=PostProcConfig(t_high=0.90, min_volume_mm3=50.0), |
| ) |
| |
| # diagnostics |
| mask, report = postprocess(logits_or_prob, spacing, return_report=True) |
| # report["components"] -> [{volume_mm3, peak_prob, mean_prob, kept, reason}, ...] |
| ``` |
|
|
| The module accepts either a probability map `(Z,Y,X)` or 2-class logits |
| `(2,Z,Y,X)` (2-class softmax = numerically stable sigmoid). |
|
|
| ## Calibration on the 50 validation-public cases |
|
|
| This is what `nnUNetv2_determine_postprocessing` would have done — to be done |
| here manually, since the baseline never produced it: |
|
|
| ```python |
| from postprocess import calibrate |
| |
| best_cfg, table = calibrate( |
| prob_maps, gts, spacings, # aligned lists (probability maps / GT masks / spacings) |
| metric="nsd", # "dsc" | "nsd" | "mean" |
| ) |
| print(best_cfg) # optimal config |
| print(table[:5]) # best grid combinations |
| ``` |
|
|
| It sweeps `t_high`, `min_volume_mm3`, `min_core_prob` and maximizes the mean |
| metric over the cohort. Pick the point that maximizes DSC/NSD **without |
| collapsing recall**. |
|
|
| ## Demonstration (synthetic volume: 1 lesion + 40 false positives) |
|
|
| ``` |
| BASELINE argmax 0.5 : CC = 41 DSC = 0.379 NSD@2mm = 0.375 |
| POST-PROCESS : CC = 1 DSC = 0.444 NSD@2mm = 0.554 |
| ``` |
|
|
| ## Caveats |
|
|
| 1. **Defaults must be calibrated** (see above) before reporting any numbers. |
| 2. The bundled `surface_dice` is a **reference implementation** (boundaries via |
| erosion, distances via anisotropic EDT): **verify it against the official |
| FLARE scorer** before publishing metrics. |
| 3. Post-processing **recovers no missed lesion** (NO_OVERLAP) — only retraining |
| would address the partial-label issue. |
| |
| ## Dependencies |
| |
| `numpy`, `scipy` (required); `SimpleITK` (NIfTI I/O for the CLI); `cc3d` (speeds up |
| labelling, automatic fallback to `scipy.ndimage` otherwise). |
| |