Instructions to use neonforestmist/Clover-Image-Tiny-Inpaint with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use neonforestmist/Clover-Image-Tiny-Inpaint with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline from diffusers.utils import load_image # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("neonforestmist/Clover-Image-Tiny-Inpaint", dtype=torch.bfloat16, device_map="cuda") prompt = "Turn this cat into a dog" input_image = load_image("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/cat.png") image = pipe(image=input_image, prompt=prompt).images[0] - Notebooks
- Google Colab
- Kaggle
Clover Image Tiny Inpainting
The inpainting model keeps Clover’s text encoder, VAE, scheduler, and safety checker, and replaces only the U-Net with a 9-channel fine-tuned version.
The U-Net input is ordered as:
[noisy_latent (4), mask (1), masked_image_latent (4)]
White mask pixels are regenerated; black pixels are preserved. The v2 recipe
warm-starts from the released nine-channel checkpoint and distills the pinned
official stable-diffusion-v1-5/stable-diffusion-inpainting U-Net. A smaller
ground-truth diffusion term prevents the student from merely copying teacher
errors. Min-SNR timestep weighting and a spatial objective emphasize the
masked region, its boundary, and enough unmasked context to learn coherent
transitions.
Masks are sampled procedurally as free-form strokes, multiple disconnected
strokes, rounded boxes, ellipses, polygons, multiple object-like regions, and
outpainting bands. The pinned dataset is prithivMLmods/Caption3o-Opt at
revision 17e893f785fcd3f5d6fc4a5d65a914b9f7b1ff5b (Apache-2.0, image +
caption). A deterministic 128-image tail is excluded from training and used
for same-mask, same-seed checkpoint comparison.
Modal
The local Modal CLI is already configured for the guccichungus69 profile.
Use that profile before running the commands below:
modal profile activate guccichungus69
modal run modal_inpaint.py --smoke --steps 1
modal run --detach modal_inpaint.py \
--steps 500 \
--lr-warmup-steps 25 \
--output-name clover-image-tiny-inpaint-v2-pilot-500
The job uses one A10 and writes the finished Diffusers pipeline to the
clover-image-tiny-inpaint-output Volume. No Hub token is required for the
default public dataset. Resumable Accelerator checkpoints are committed to the
persistent Volume. A separate cache avoids downloading the student, teacher,
and dataset again on every run.
The published v2 checkpoint is the 500-step cosine-decay sweep. A longer run was rejected after its held-out reconstruction and semantic gates regressed into colored high-frequency artifacts. The release pins the best validated checkpoint rather than the final checkpoint from the largest job.
Download a completed output directory with:
modal volume get clover-image-tiny-inpaint-output \
clover-image-tiny-inpaint-v2-pilot-500/unet \
./artifacts/clover-image-tiny-inpaint-v2
If a job stops after a committed checkpoint, resume the same output with
--resume. Completed output directories are immutable to the launcher; use a
new --output-name for a distinct experiment.
Local dry run
The same trainer can run locally when the model and dataset are available:
python inpainting/train.py \
--pretrained_model_name_or_path . \
--dataset_name prithivMLmods/Caption3o-Opt \
--dataset_revision 17e893f785fcd3f5d6fc4a5d65a914b9f7b1ff5b \
--max_train_samples 4 \
--max_train_steps 2 \
--output_dir /tmp/clover-image-tiny-inpaint-smoke
The complete model is large enough that Modal is the intended training environment. The current repo’s Apple Core ML environment is for conversion, not PyTorch training.