How to use from the
Use from the
Diffusers library
pip install -U diffusers transformers accelerate
import torch
from diffusers import DiffusionPipeline

# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("a12donhf/CPathOGen", dtype=torch.bfloat16, device_map="cuda")

prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k"
image = pipe(prompt).images[0]

CPathOGen

Spatially and Morphologically Controlled H&E Counterfactuals for Probing Pathology Models

What the model does

CPathOGen synthesizes a 512 x 512 H&E tile from a cellular spatial map and a morphology/appearance vector. A learned spatial encoder creates latent-space features that are concatenated with the noisy image latent. Blockwise FiLM modules apply the morphology controls during denoising. The released checkpoint is the custom checkpoint-30000_FID58/checkpoint-30000 model used in the project's inference workflows.

Researchers can keep diffusion noise fixed, change a requested control, and compare outputs of a frozen pathology model on the resulting matched images. This release contains the generator; CellViT++ candidate ranking and independent nucleus analysis require their own software and checkpoints.

The checkpoint uses latent concatenation with a learned spatial encoder. It is not a standard Diffusers ControlNetModel or standalone DiffusionPipeline checkpoint.

Figures from the paper

Controllable image generation and recorded black-box predictions for post-hoc interpretation

Counterfactual probing: change a control, generate a matched image, and measure the downstream model's prediction response.

CPathOGen pipeline from CellViT++ weak labels to spatial and morphology-conditioned latent diffusion

Generation pipeline: CellViT++ supplies cellular maps and morphology/appearance summaries; the spatial encoder and FiLM condition latent diffusion synthesis.

Three examples showing the cell map, real H&E tile, and generated H&E tile

Spatial examples: input maps, associated real tiles, and condition-matched generated tiles from the paper. Map colors are tumor (white), immune (cyan), stroma (green), dead (yellow), and non-neoplastic epithelium (orange).

Five-level sweeps of nuclear size, eccentricity, solidity, gradient, and RGB appearance

Morphology and appearance examples: the paper's five-level control sweeps. Most columns use relative standardized offsets; the historical eccentricity illustration instead uses absolute standardized coordinates with 20 steps and spatial strength 1. Nuclear size changes area and perimeter together. These are paper illustrations, not new generations from the quickstart.

Run one real example in Colab

Select a GPU runtime in Colab, then run the cell below. It downloads one prepared paper-tile condition, so you do not need the full dataset, Google Drive, or CellViT++ to try the generator. The public weights require no Hugging Face token.

from pathlib import Path
import torch

if not torch.cuda.is_available():
    raise RuntimeError("Select Runtime > Change runtime type > GPU in Colab.")

%cd /content
if not Path("/content/CPathOGen-example/.git").is_dir():
    !git clone -q --depth 1 https://github.com/a12dongithub/PathOGen.git /content/CPathOGen-example
%cd /content/CPathOGen-example
%pip install -q -r inference/requirements.txt

from huggingface_hub import hf_hub_download
from IPython.display import display
from PIL import Image

MODEL_ID = "a12donhf/CPathOGen"
map_path = hf_hub_download(MODEL_ID, "examples/paper_tile/map.npz")
morphology_path = hf_hub_download(MODEL_ID, "examples/paper_tile/morphology.json")
preview_path = hf_hub_download(MODEL_ID, "examples/paper_tile/input_map.png")

!python inference/generate.py --spatial-map "{map_path}" --morphology-json "{morphology_path}" --tile TCGA-E2-A15D_x46080_y24576_TR --seed 1872879198 --steps 30 --spatial-strength 2 --output outputs/paper_example.png

print("Input spatial map")
display(Image.open(preview_path))
print("Generated H&E")
display(Image.open("outputs/paper_example.png"))
print("Controls and run metadata: outputs/paper_example.json")

The first run downloads approximately 4.16 GB of generator weights, plus the base-model text encoder, tokenizer, and scheduler. Downloads are cached. This produces one 512 x 512 PNG and its condition/provenance JSON; it does not run eight-seed selection or downstream probing. Keep the seed fixed when comparing edited controls. Exact pixels can differ across hardware and numerical precision.

This model card documents GPU inference through the project code; it is not a hosted Hugging Face inference widget.

Local inference and custom inputs

Use Python 3.10/3.11, a compatible CUDA-enabled PyTorch build, and an NVIDIA GPU.

git clone https://github.com/a12dongithub/PathOGen.git
cd PathOGen
python -m pip install -r inference/requirements.txt
python inference/generate.py --synthetic-example --seed 42 --steps 30 --spatial-strength 2 --output outputs/demo.png

This downloads the model weights from this repository and the text encoder, tokenizer, and scheduler from the named Stable Diffusion 2.1 base model. Downloads are cached. The artificial demo layout is for software testing and does not reproduce the paper's evaluation.

For a prepared dataset tile:

python inference/generate.py --data-root /path/to/512_final_dataset --tile YOUR_TILE_ID --seed 42 --output outputs/tile.png

For custom conditions:

python inference/generate.py --spatial-map conditions/map.npz --morphology-json conditions/morphology.json --seed 42 --output outputs/tile.png

See the GitHub README for the complete schema, Colab cell, local-checkpoint usage, and paired interventions.

Inputs

The spatial NPZ must contain a map array with shape (512, 512, 5) or (5, 512, 512). Channel order is neoplastic, inflammatory, connective, dead, and epithelial. Zero values in all channels represent background. Input intensities are uint8 in [0, 255] or floats in [0, 1]; the input is a map of smoothed cell centroids rather than a color-rendered illustration.

The 16 morphology entries are standardized values in this order:

area_mean, area_var, eccentricity_mean, eccentricity_var,
solidity_mean, solidity_var, perimeter_mean, perimeter_var,
grad_mean, grad_var, r_mean, r_var, g_mean, g_var, b_mean, b_var

The historical preprocessing script did not save the fitted StandardScaler. Reuse already standardized dataset values; no verified training-scaler artifact is included in this release. A newly fitted scaler on another cohort can alter stain and geometry controls.

Released files

checkpoint-30000/
  unet/config.json
  unet/diffusion_pytorch_model.safetensors
  vae/config.json
  vae/diffusion_pytorch_model.safetensors
  film_mlps.pt
  spatial_encoder.pt
assets/
  principle.png
  pipeline.png
  spatial_fidelity.png
  morphology_controls.png
examples/paper_tile/
  map.npz
  morphology.json
  metadata.json
  input_map.png
  reference_generated.png

The FiLM and spatial-encoder files contain PyTorch state dictionaries and are loaded with weights_only=True. Optimizer, scheduler training state, and random-state pickle files are excluded. Original weights are preserved; the checkpoint is not quantized or converted. release_manifest.json records per-file SHA-256 hashes and sizes.

Training and evaluation

The paper describes training on approximately 1.4 million 512 x 512 tiles from TCGA-BRCA, retaining tiles with more than two detected tumor cells. CellViT++ supplies weak nucleus contours, locations, and labels. Training used eight Tesla V100 GPUs.

The project reports the following distributional results:

Generation protocol FID KID mean
No seed filtering 42.9228 0.0246604
CellViT++ selection from eight seeds 36.7168 0.0196440

These values describe the project's evaluation protocol and its reference set. The software demo is not an FID/KID reproduction. Selection requires candidate generation and CellViT++ ranking; a single model call does not include selection.

Intended use and limitations

Research uses include controlled histopathology synthesis and probing model sensitivity to nuclear morphology, cellular organization, and appearance. The generator is not validated for clinical diagnosis or patient management. Shared diffusion noise and fixed conditions do not guarantee that every unmeasured property remains unchanged. Global morphology summaries compress cell-level heterogeneity, and out-of-distribution controls can produce artifacts. Synthetic intervention responses measure model behavior rather than biological or treatment causality.

License

Model weights retain the CreativeML Open RAIL++-M terms inherited from Stable Diffusion 2.1; see LICENSE-MODEL. Third-party software, analyzers, and datasets retain their respective licenses and access requirements.

Paper figures are shared under CC BY 4.0 with attribution to Samarth Singhal and Varang Rai; see assets/README.md. This does not change the model-weight license.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for a12donhf/CPathOGen

Finetuned
(14)
this model