flexray / README.md
VictorButoi's picture
Consolidate the FleXray ensemble into one repository (members/ + ensemble.json)
b3d8920 verified
|
Raw
History Blame Contribute Delete
10.3 kB
---
license: cc-by-nc-4.0
library_name: pytorch
pipeline_tag: image-segmentation
datasets:
- VictorButoi/flexray-data
tags:
- medical-image-segmentation
- x-ray
- radiograph
- anatomy
- flexray
---
# FleXray: `VictorButoi/flexray`
- Website and in-browser demo: [victorbutoi.github.io/FleXray](https://victorbutoi.github.io/FleXray/)
- Code: [github.com/VictorButoi/FleXray](https://github.com/VictorButoi/FleXray)
- Data: [`VictorButoi/flexray-data`](https://huggingface.co/datasets/VictorButoi/flexray-data)
- Tutorial: [Colab notebook](https://colab.research.google.com/drive/1MNIeN9LN-tY8wAifWSolxmFZi3Q0BlVX)
- Paper: *FleXray: Flexible Full-Body X-ray Segmentation* (coming soon)
FleXray is a single 2D UNet that segments anatomy from standard radiographs
across body regions, projections, and acquisition settings. It predicts 60
anatomical structures (plus background) as independent sigmoid channels at
256 x 256 resolution.
This repository holds the **flagship model** and the four sibling models of
the FleXray ensemble, one bundle per `members/` subfolder (see
[Repository layout](#repository-layout)). `flexify` and
`FleXraySegmenter.from_pretrained` load the flagship by default.
## Quick start
```bash
python -m pip install flexray
flexify --input ./image.png --output-dir ./predictions
```
```python
from fxr.inference import FleXraySegmenter
segmenter = FleXraySegmenter.from_pretrained("VictorButoi/flexray")
prediction = segmenter.predict("./image.png", threshold=0.5)
prediction.masks # uint8, CxHxW thresholded masks
prediction.probabilities # float32, CxHxW sigmoid probabilities
prediction.logits # float32, CxHxW raw scores
```
`flexify` writes `<name>_masks.npy`, `<name>_probabilities.npy`, and
`<name>_logits.npy` per image. Channel order follows `label_schema.json`.
Pass `--binary LABEL` (for example `--binary femurs`) to write one label. See
[docs/inference.md](https://github.com/VictorButoi/FleXray/blob/main/docs/inference.md)
for the full CLI and Python API.
## The FleXray ensemble
The flagship was trained with a 0.375 FluXray proportion in the training mix.
Four sibling models share its architecture, label schema, preprocessing, and
training recipe and differ only in that proportion:
| Subfolder | FluXray proportion | Role |
| --- | --- | --- |
| `members/flux000` | 0.0 | ensemble member |
| `members/flux025` | 0.25 | ensemble member |
| **`members/flux0375`** | **0.375** | **flagship (loaded by default)** |
| `members/flux050` | 0.5 | ensemble member |
| `members/flux075` | 0.75 | ensemble member |
`ensemble.json` at the repository root lists the flagship and the members.
Because the members share one output space, they are averaged in probability
space:
```bash
flexify --ensemble --tta-samples 16 --input ./image.png --output-dir ./predictions
flexify --subfolder members/flux000 --input ./image.png --output-dir ./predictions
```
```python
segmenter = FleXraySegmenter.from_pretrained("VictorButoi/flexray", ensemble=True)
prediction = segmenter.predict("./image.png", tta_samples=16)
member = FleXraySegmenter.from_pretrained(
"VictorButoi/flexray", subfolder="members/flux000"
)
```
The website demo exposes the same choices as quality modes: **Low** runs the
flagship once, **Normal** runs the flagship with 16-pass TTA, **High** runs the
five-model ensemble once, and **X-High** runs the ensemble with 16-pass TTA.
The members are also listed in
[MODEL_ZOO.md](https://github.com/VictorButoi/FleXray/blob/main/MODEL_ZOO.md).
## Test-time augmentation
We typically run FleXray with test-time augmentation (TTA) rather than a single
forward pass; `--tta-samples 16` (or `predict(..., tta_samples=16)`) is the
setting behind reported results and the demo's Normal / X-High modes.
`tta_samples=N` runs one un-augmented pass plus `N - 1` randomly augmented
passes and averages them in probability space (mean of sigmoid outputs, then
converted back to logits). The augmentation chain is fixed in
[`fxr.inference.tta`](https://github.com/VictorButoi/FleXray/blob/main/src/fxr/inference/tta.py)
and the browser demo mirrors it exactly:
| Transform | Probability | Range |
| --- | --- | --- |
| Horizontal flip (exactly inverted on the logits before merging) | 0.5 | - |
| Gamma | 0.5 | gamma 0.9-1.1, gain 0.9-1.1 |
| Intensity scale (additive) | 0.5 | -0.1 to 0.1 |
| Brightness | 0.5 | 0.8-1.2 |
| Sharpness | 0.5 | 0.6-1.4 |
| Invert | 0.5 | - |
| Contrast | 0.5 | 0.7-1.3 |
The flip is the only geometric transform; intensity transforms do not move
pixels and are not inverted. Augmented views are drawn from the global torch
RNG (`torch.manual_seed` for reproducibility). With an ensemble, every view is
drawn once and run through every member, so `M` members with `tta_samples=N`
cost `M x N` forward passes (80 for the full ensemble at N=16). `tta_samples<=1`
reproduces the plain single pass.
## Input contract
`preprocessing.json` is applied automatically by the public loaders:
- grayscale input (RGB is converted), any 8-bit or 16-bit PNG / JPEG / TIFF / BMP
- per-image percentile min-max normalization to `[0, 1]` (0.5th / 99.5th
percentiles, `eps = 1e-8`)
- zero-pad to a square, then resize to 256 x 256
- outputs are `multilabel` sigmoid probabilities; masks use threshold 0.5
Outputs are at the 256 x 256 model resolution; the CLI and Python API do not
resample back to the original image size.
## Output labels
FleXray outputs 60 foreground masks (61 channels including `background`). The
broader dataset/evaluation protocol also recognizes aggregate `lumbar_spine` and
`thoracolumbar_spine` annotations; these are evaluated by combining the relevant
per-vertebra outputs and are not checkpoint channels. Channel order is stored in
each bundle's `label_schema.json`.
- **Skull / shoulder girdle:** skull, scapulae, clavicles
- **Upper limb:** humeri, radii, ulnae, carpals, metacarpals, phalanges
- **Lower limb:** femurs, patellae, tibiae, fibulae, tarsals, metatarsals, toes
- **Thorax:** rib_1 - rib_12, sternum
- **Spine:** vertebra_c1 - c7, t1 - t12, l1 - l5, sacrum
- **Pelvis:** hips
- **Soft tissue:** lungs, heart, liver, spleen, kidneys
Paired structures are merged (for example `femurs` covers both sides);
laterality is not predicted.
## Architecture
`fxr.models.UNet`, 2D, 1 input channel, 61 output channels; filters
`[64, 128, 256, 512, 512, 720, 1024]`, 3 convolutions per block, residual
blocks with instance norm, `align_corners=True` upsampling. The full
architecture is in each bundle's `config.yml`.
## Training data
The models were trained on three source types unified under the FleXray label
protocol. For training mixture proportions, please refer to the paper:
- **Real X-ray masks:** HandBones, FootBones, MURA forearm, and MURA humerus,
with our own annotations.
- **Generated FluXray images:** digitally reconstructed radiographs from the
MOOSE CTs, generatively edited toward real X-ray appearance, with exact
overlapping masks for every protocol structure.
- **Online CT->DRR rendering:** MOOSE / ENHANCE-PET 1.6k, Shoulder-CT, HaN-Seg,
PedsCT, RSNA cervical-spine fracture CTs, and ElbowCT, rendered to DRRs at
random poses during training with per-label attenuation jitter.
Training used AdamW (lr 3e-4, cosine schedule), a Dice + binary cross-entropy
loss routed per dataset (partially labeled sources ignore unlabeled channels),
and separate augmentation presets for CT-derived and X-ray inputs. The exact
recipe is `fxr/configs/training/base.yml` in the code release.
Every dataset's license, redistribution status, and download pointer is
documented in the
[`VictorButoi/flexray-data`](https://huggingface.co/datasets/VictorButoi/flexray-data)
card. That repository ships the real X-ray sources whose licenses permit
redistribution, already packed in the FleXray protocol, the MURA masks, and
the FluXray database.
## Evaluation
FleXray was evaluated on nine real-radiograph datasets held out from training
(ElbowLat, HipRay, LowerLimbs, MendeleyCXR, MTDDH, DarwinCVD19, DeepFluoro,
RAM-W600, VinDr-Rib), spanning lungs, ribs, peripheral bones, spine, and
pelvis. Against generalist baselines (FluoroSAM, TotalSegmentator2D, PAXray)
it achieves the highest macro Dice in every comparison (9 of 9). Per-dataset
numbers and confidence intervals are in the paper; the benchmark figure is on
the [project website](https://victorbutoi.github.io/FleXray/#results).
Evaluation ignores ground-truth labels covering less than 0.1% of the image.
## Intended use and limitations
Research use only. FleXray is **not a medical device** and is not cleared for
clinical diagnosis, treatment planning, or patient-care decisions.
- Targets conventional radiographs; dental and mammographic images are out of
scope.
- Predicts anatomy, not pathology.
- No laterality (left/right) and no uncertainty estimates.
- Performance on acquisition settings, views, or populations far from the
training sources has not been validated.
## Repository layout
- `README.md`: this card.
- `ensemble.json`: the `flagship` subfolder and the `members` list with their
FluXray proportions.
- `members/<name>/model.safetensors`: exported model weights.
- `members/<name>/config.yml`: architecture and protocol config consumed by
`from_pretrained`.
- `members/<name>/label_schema.json`: ordered output labels.
- `members/<name>/preprocessing.json`: public preprocessing contract.
- `members/<name>/checksums.json`: SHA256 checksums of the bundle files.
- `members/<name>/onnx/flexray-<name>-256-fp16.onnx`: fp16 ONNX export
(opset 18, sigmoid baked in) used by the in-browser demo; parity-checked
against the PyTorch weights by `tools/export_web_demo.py`.
## Licenses
- Code: MIT
- Weights: CC-BY-NC-4.0
## Citation
```bibtex
@software{butoi2026flexray,
title = {FleXray: Flexible Full-Body X-ray Segmentation},
author = {Butoi, Victor Ion and Gopalakrishnan, Vivek and
Guttag, John V. and Dalca, Adrian V. and Dey, Neel},
year = {2026},
license = {MIT},
url = {https://github.com/VictorButoi/FleXray}
}
```
Please also cite the source datasets listed in the
[`flexray-data`](https://huggingface.co/datasets/VictorButoi/flexray-data) card
for any dataset you use.