Instructions to use a12donhf/CPathOGen with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use a12donhf/CPathOGen with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("a12donhf/CPathOGen", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
CPathOGen
Spatially and Morphologically Controlled H&E Counterfactuals for Probing Pathology Models
- Code and complete instructions: a12dongithub/PathOGen
- Model repository: a12donhf/CPathOGen
- Authors: Samarth Singhal and Varang Rai
What the model does
CPathOGen synthesizes a 512 x 512 H&E tile from a cellular spatial map and a morphology/appearance vector. A learned spatial encoder creates latent-space features that are concatenated with the noisy image latent. Blockwise FiLM modules apply the morphology controls during denoising. The released checkpoint is the custom checkpoint-30000_FID58/checkpoint-30000 model used in the project's inference workflows.
Researchers can keep diffusion noise fixed, change a requested control, and compare outputs of a frozen pathology model on the resulting matched images. This release contains the generator; CellViT++ candidate ranking and independent nucleus analysis require their own software and checkpoints.
The checkpoint uses latent concatenation with a learned spatial encoder. It is not a standard Diffusers ControlNetModel or standalone DiffusionPipeline checkpoint.
Figures from the paper
Counterfactual probing: change a control, generate a matched image, and measure the downstream model's prediction response.
Generation pipeline: CellViT++ supplies cellular maps and morphology/appearance summaries; the spatial encoder and FiLM condition latent diffusion synthesis.
Spatial examples: input maps, associated real tiles, and condition-matched generated tiles from the paper. Map colors are tumor (white), immune (cyan), stroma (green), dead (yellow), and non-neoplastic epithelium (orange).
Morphology and appearance examples: the paper's five-level control sweeps. Most columns use relative standardized offsets; the historical eccentricity illustration instead uses absolute standardized coordinates with 20 steps and spatial strength 1. Nuclear size changes area and perimeter together. These are paper illustrations, not new generations from the quickstart.
Run one real example in Colab
Select a GPU runtime in Colab, then run the cell below. It downloads one prepared paper-tile condition, so you do not need the full dataset, Google Drive, or CellViT++ to try the generator. The public weights require no Hugging Face token.
from pathlib import Path
import torch
if not torch.cuda.is_available():
raise RuntimeError("Select Runtime > Change runtime type > GPU in Colab.")
%cd /content
if not Path("/content/CPathOGen-example/.git").is_dir():
!git clone -q --depth 1 https://github.com/a12dongithub/PathOGen.git /content/CPathOGen-example
%cd /content/CPathOGen-example
%pip install -q -r inference/requirements.txt
from huggingface_hub import hf_hub_download
from IPython.display import display
from PIL import Image
MODEL_ID = "a12donhf/CPathOGen"
map_path = hf_hub_download(MODEL_ID, "examples/paper_tile/map.npz")
morphology_path = hf_hub_download(MODEL_ID, "examples/paper_tile/morphology.json")
preview_path = hf_hub_download(MODEL_ID, "examples/paper_tile/input_map.png")
!python inference/generate.py --spatial-map "{map_path}" --morphology-json "{morphology_path}" --tile TCGA-E2-A15D_x46080_y24576_TR --seed 1872879198 --steps 30 --spatial-strength 2 --output outputs/paper_example.png
print("Input spatial map")
display(Image.open(preview_path))
print("Generated H&E")
display(Image.open("outputs/paper_example.png"))
print("Controls and run metadata: outputs/paper_example.json")
The first run downloads approximately 4.16 GB of generator weights, plus the base-model text encoder, tokenizer, and scheduler. Downloads are cached. This produces one 512 x 512 PNG and its condition/provenance JSON; it does not run eight-seed selection or downstream probing. Keep the seed fixed when comparing edited controls. Exact pixels can differ across hardware and numerical precision.
This model card documents GPU inference through the project code; it is not a hosted Hugging Face inference widget.
Local inference and custom inputs
Use Python 3.10/3.11, a compatible CUDA-enabled PyTorch build, and an NVIDIA GPU.
git clone https://github.com/a12dongithub/PathOGen.git
cd PathOGen
python -m pip install -r inference/requirements.txt
python inference/generate.py --synthetic-example --seed 42 --steps 30 --spatial-strength 2 --output outputs/demo.png
This downloads the model weights from this repository and the text encoder, tokenizer, and scheduler from the named Stable Diffusion 2.1 base model. Downloads are cached. The artificial demo layout is for software testing and does not reproduce the paper's evaluation.
For a prepared dataset tile:
python inference/generate.py --data-root /path/to/512_final_dataset --tile YOUR_TILE_ID --seed 42 --output outputs/tile.png
For custom conditions:
python inference/generate.py --spatial-map conditions/map.npz --morphology-json conditions/morphology.json --seed 42 --output outputs/tile.png
See the GitHub README for the complete schema, Colab cell, local-checkpoint usage, and paired interventions.
Inputs
The spatial NPZ must contain a map array with shape (512, 512, 5) or (5, 512, 512). Channel order is neoplastic, inflammatory, connective, dead, and epithelial. Zero values in all channels represent background. Input intensities are uint8 in [0, 255] or floats in [0, 1]; the input is a map of smoothed cell centroids rather than a color-rendered illustration.
The 16 morphology entries are standardized values in this order:
area_mean, area_var, eccentricity_mean, eccentricity_var,
solidity_mean, solidity_var, perimeter_mean, perimeter_var,
grad_mean, grad_var, r_mean, r_var, g_mean, g_var, b_mean, b_var
The historical preprocessing script did not save the fitted StandardScaler. Reuse already standardized dataset values; no verified training-scaler artifact is included in this release. A newly fitted scaler on another cohort can alter stain and geometry controls.
Released files
checkpoint-30000/
unet/config.json
unet/diffusion_pytorch_model.safetensors
vae/config.json
vae/diffusion_pytorch_model.safetensors
film_mlps.pt
spatial_encoder.pt
assets/
principle.png
pipeline.png
spatial_fidelity.png
morphology_controls.png
examples/paper_tile/
map.npz
morphology.json
metadata.json
input_map.png
reference_generated.png
The FiLM and spatial-encoder files contain PyTorch state dictionaries and are loaded with weights_only=True. Optimizer, scheduler training state, and random-state pickle files are excluded. Original weights are preserved; the checkpoint is not quantized or converted. release_manifest.json records per-file SHA-256 hashes and sizes.
Training and evaluation
The paper describes training on approximately 1.4 million 512 x 512 tiles from TCGA-BRCA, retaining tiles with more than two detected tumor cells. CellViT++ supplies weak nucleus contours, locations, and labels. Training used eight Tesla V100 GPUs.
The project reports the following distributional results:
| Generation protocol | FID | KID mean |
|---|---|---|
| No seed filtering | 42.9228 | 0.0246604 |
| CellViT++ selection from eight seeds | 36.7168 | 0.0196440 |
These values describe the project's evaluation protocol and its reference set. The software demo is not an FID/KID reproduction. Selection requires candidate generation and CellViT++ ranking; a single model call does not include selection.
Intended use and limitations
Research uses include controlled histopathology synthesis and probing model sensitivity to nuclear morphology, cellular organization, and appearance. The generator is not validated for clinical diagnosis or patient management. Shared diffusion noise and fixed conditions do not guarantee that every unmeasured property remains unchanged. Global morphology summaries compress cell-level heterogeneity, and out-of-distribution controls can produce artifacts. Synthetic intervention responses measure model behavior rather than biological or treatment causality.
License
Model weights retain the CreativeML Open RAIL++-M terms inherited from Stable Diffusion 2.1; see LICENSE-MODEL. Third-party software, analyzers, and datasets retain their respective licenses and access requirements.
Paper figures are shared under CC BY 4.0 with attribution to Samarth Singhal and Varang Rai; see assets/README.md. This does not change the model-weight license.
- Downloads last month
- -
Model tree for a12donhf/CPathOGen
Base model
Manojb/stable-diffusion-2-1-base


