Text-to-Image
Diffusers
Safetensors
StableDiffusionPipeline
clover-image
diffusion
stable-diffusion
knowledge-distillation
compact
small-model
local-inference
edge-inference
mobile-inference
core-ml
iphone
sd-1.4-class
Instructions to use neonforestmist/Clover-Image-Tiny with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use neonforestmist/Clover-Image-Tiny with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("neonforestmist/Clover-Image-Tiny", dtype=torch.bfloat16, device_map="cuda") prompt = "a glass of red wine" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
Commit ·
243ef26
1
Parent(s): 610cadc
Add small-model benchmark and organize model card
Browse filesThis view is limited to 50 files because it contains too many changes. See raw diff
- .gitattributes +1 -0
- README.md +60 -17
- benchmark_text_to_image.py +283 -0
- benchmarks/text-to-image/README.md +37 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/REPORT.md +53 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/contact-sheet.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/01.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/02.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/03.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/04.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/05.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/06.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/07.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/08.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/09.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/10.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/11.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/12.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/13.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/14.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/15.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/16.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/01.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/02.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/03.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/04.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/05.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/06.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/07.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/08.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/09.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/10.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/11.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/12.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/13.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/14.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/15.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/16.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/clover/01.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/clover/02.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/clover/03.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/clover/04.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/clover/05.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/clover/06.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/clover/07.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/clover/08.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/clover/09.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/clover/10.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/clover/11.png +3 -0
- benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/clover/12.png +3 -0
.gitattributes
CHANGED
|
@@ -38,3 +38,4 @@ assets/clover-image-tiny-paired-contact-sheet.png filter=lfs diff=lfs merge=lfs
|
|
| 38 |
assets/clover-image-tiny-banner.png filter=lfs diff=lfs merge=lfs -text
|
| 39 |
examples/prompt-gallery/**/*.png filter=lfs diff=lfs merge=lfs -text
|
| 40 |
examples/normal-to-lora/*.png filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
| 38 |
assets/clover-image-tiny-banner.png filter=lfs diff=lfs merge=lfs -text
|
| 39 |
examples/prompt-gallery/**/*.png filter=lfs diff=lfs merge=lfs -text
|
| 40 |
examples/normal-to-lora/*.png filter=lfs diff=lfs merge=lfs -text
|
| 41 |
+
benchmarks/**/*.png filter=lfs diff=lfs merge=lfs -text
|
README.md
CHANGED
|
@@ -61,6 +61,10 @@ Windows, or Linux.
|
|
| 61 |
**323,384,964 denoiser parameters · about 1.67 GB · 4–100 inference steps ·
|
| 62 |
PyTorch/Diffusers**
|
| 63 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 64 |
Clover Image Tiny is the public PyTorch/Diffusers checkpoint release behind
|
| 65 |
these examples. Its output has a recognizable, playful
|
| 66 |
**DALL·E mini-ish** character. That is a visual description, not a claim of
|
|
@@ -74,7 +78,24 @@ The demo exposes prompt, negative prompt, seed, guidance, dimensions,
|
|
| 74 |
scheduler, and 4–100 conventional Diffusers inference steps. It creates one
|
| 75 |
image per request and keeps the packaged safety checker enabled.
|
| 76 |
|
| 77 |
-
##
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 78 |
|
| 79 |
### Prompt gallery
|
| 80 |
|
|
@@ -112,7 +133,29 @@ seed 1469, and the same 50-step configuration:
|
|
| 112 |
|
| 113 |

|
| 114 |
|
| 115 |
-
##
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 116 |
|
| 117 |
The SwiftUI app in [`Clover-iOS`](Clover-iOS/) follows Apple platform
|
| 118 |
conventions and exposes prompt, negative prompt, steps, guidance, seed, image
|
|
@@ -137,7 +180,7 @@ own Core ML picker download:
|
|
| 137 |
See [`COREML.md`](COREML.md) for conversion details and
|
| 138 |
[`training/README.md`](training/README.md) for the pinned LoRA jobs.
|
| 139 |
|
| 140 |
-
### Inpainting track
|
| 141 |
|
| 142 |
The 9-channel SD 1.4-class inpainting adaptation is trained and packaged separately:
|
| 143 |
[`neonforestmist/Clover-Image-Tiny-Inpaint`](https://huggingface.co/neonforestmist/Clover-Image-Tiny-Inpaint).
|
|
@@ -153,12 +196,12 @@ diverse free-form and object-like masks. It improved held-out masked MAE by
|
|
| 153 |
96-pixel mask-context crop; the runtime composites through the exact mask so
|
| 154 |
unmasked pixels remain unchanged.
|
| 155 |
|
| 156 |
-
## Run locally
|
| 157 |
|
| 158 |
Download once, then generate offline with the bundled runner. Python 3.11 and
|
| 159 |
3.12 are supported.
|
| 160 |
|
| 161 |
-
### macOS — Apple silicon
|
| 162 |
|
| 163 |
~~~bash
|
| 164 |
mkdir clover-image-tiny-local
|
|
@@ -189,7 +232,7 @@ open clover-image-tiny.png
|
|
| 189 |
|
| 190 |
Use `python3.11` instead if that is the installed supported Python.
|
| 191 |
|
| 192 |
-
### Windows — PowerShell
|
| 193 |
|
| 194 |
~~~powershell
|
| 195 |
mkdir clover-image-tiny-local
|
|
@@ -221,7 +264,7 @@ Invoke-Item .\clover-image-tiny.png
|
|
| 221 |
Use `py -3.11` if needed. With `--device auto`, the runner selects an
|
| 222 |
available NVIDIA CUDA GPU and otherwise uses CPU.
|
| 223 |
|
| 224 |
-
### Linux
|
| 225 |
|
| 226 |
~~~bash
|
| 227 |
mkdir clover-image-tiny-local
|
|
@@ -250,7 +293,7 @@ python model/examples/generate.py \
|
|
| 250 |
uses CPU. After the first download, `--local-files-only` prevents network
|
| 251 |
access during generation.
|
| 252 |
|
| 253 |
-
## Generation controls
|
| 254 |
|
| 255 |
The command above is ready to copy. Change these flags to explore the model:
|
| 256 |
|
|
@@ -281,7 +324,7 @@ outputs are never overwritten.
|
|
| 281 |
|
| 282 |
Run `python model/examples/generate.py --help` for the complete CLI reference.
|
| 283 |
|
| 284 |
-
## Hardware and operating systems
|
| 285 |
|
| 286 |
| System | Automatic backend | Precision | Current evidence |
|
| 287 |
|---|---|---|---|
|
|
@@ -301,7 +344,7 @@ image in 18.21 seconds with fp16 MPS. Its process-lifetime maximum RSS was
|
|
| 301 |
631,341,056 bytes. This is a measured point, not a minimum-RAM claim. No Core
|
| 302 |
ML package is required for the Python path.
|
| 303 |
|
| 304 |
-
## Python API
|
| 305 |
|
| 306 |
~~~python
|
| 307 |
import torch
|
|
@@ -337,7 +380,7 @@ image.save("clover-image-tiny.png")
|
|
| 337 |
Seeded generation is repeatable within the selected runtime. Different
|
| 338 |
devices, dtypes, kernels, and dependency builds can produce different pixels.
|
| 339 |
|
| 340 |
-
## About this release
|
| 341 |
|
| 342 |
Clover Image Tiny is a conventional knowledge-distillation checkpoint trained
|
| 343 |
for 500 optimizer steps on an exact licensed 1,000-pair calibration set. The
|
|
@@ -356,7 +399,7 @@ reproducible Core ML conversion and iPhone app source. The downloadable Core ML
|
|
| 356 |
artifacts and style adapters are versioned in the separate repositories linked
|
| 357 |
above.
|
| 358 |
|
| 359 |
-
## Quality and known behavior
|
| 360 |
|
| 361 |
- The included gallery demonstrates recognizable subjects across colorful
|
| 362 |
scenes, products, food, an animal, a landscape, and an interior.
|
|
@@ -368,7 +411,7 @@ above.
|
|
| 368 |
controlled benchmark or broad human-preference study.
|
| 369 |
- Resolution and batch size multiply memory use.
|
| 370 |
|
| 371 |
-
## Safety
|
| 372 |
|
| 373 |
The upstream safety checker is packaged and enabled in both the supported
|
| 374 |
runner and hosted demo. A flagged output may be returned as a black placeholder;
|
|
@@ -381,7 +424,7 @@ outputs before sharing them. Do not use the model for consequential decisions,
|
|
| 381 |
identity claims, medical or legal conclusions, harassment, exploitation,
|
| 382 |
illegal activity, or uses prohibited by CreativeML OpenRAIL-M.
|
| 383 |
|
| 384 |
-
## Training lineage and data
|
| 385 |
|
| 386 |
- Clover fine-tuning data: exactly 1,000 accepted image-caption pairs from
|
| 387 |
`Spawning/PD3M@2a5eb24a8dccf245acd8e56341761aee06da0bdf`
|
|
@@ -401,7 +444,7 @@ for all foundational pretraining is not available to this project.
|
|
| 401 |
See `DATA_PROVENANCE.md` for the portable manifest identity and
|
| 402 |
`MODEL_DATA_LICENSES.md` for the complete component ledger.
|
| 403 |
|
| 404 |
-
## Citation
|
| 405 |
|
| 406 |
If Clover Image Tiny is useful in your work, please cite the model release:
|
| 407 |
|
|
@@ -414,7 +457,7 @@ If Clover Image Tiny is useful in your work, please cite the model release:
|
|
| 414 |
}
|
| 415 |
```
|
| 416 |
|
| 417 |
-
## Licenses
|
| 418 |
|
| 419 |
The model weights are a derivative under **CreativeML OpenRAIL-M**. The example
|
| 420 |
runner and packaging code are under **Apache-2.0**. Dataset and item-level terms
|
|
@@ -426,7 +469,7 @@ request for display in this public model repository. It is not benchmark
|
|
| 426 |
evidence, its panel-generation provenance is not claimed, and this package
|
| 427 |
does not grant a downstream reuse license for it.
|
| 428 |
|
| 429 |
-
## Reproducibility and artifact identity
|
| 430 |
|
| 431 |
| Field | Value |
|
| 432 |
|---|---|
|
|
|
|
| 61 |
**323,384,964 denoiser parameters · about 1.67 GB · 4–100 inference steps ·
|
| 62 |
PyTorch/Diffusers**
|
| 63 |
|
| 64 |
+
Clover Image Tiny is intentionally a compact, local-first 512×512 model rather
|
| 65 |
+
than a frontier-scale checkpoint. Its denoiser has 323,384,964 parameters; the
|
| 66 |
+
small-model comparison below gives that scale practical runtime context.
|
| 67 |
+
|
| 68 |
Clover Image Tiny is the public PyTorch/Diffusers checkpoint release behind
|
| 69 |
these examples. Its output has a recognizable, playful
|
| 70 |
**DALL·E mini-ish** character. That is a visual description, not a claim of
|
|
|
|
| 78 |
scheduler, and 4–100 conventional Diffusers inference steps. It creates one
|
| 79 |
image per request and keeps the packaged safety checker enabled.
|
| 80 |
|
| 81 |
+
## Contents
|
| 82 |
+
|
| 83 |
+
1. [Examples](#1-examples)
|
| 84 |
+
2. [Small-model benchmark](#2-small-model-benchmark)
|
| 85 |
+
3. [iPhone and Core ML](#3-iphone-and-core-ml)
|
| 86 |
+
4. [Run locally](#4-run-locally)
|
| 87 |
+
5. [Generation controls](#5-generation-controls)
|
| 88 |
+
6. [Hardware and operating systems](#6-hardware-and-operating-systems)
|
| 89 |
+
7. [Python API](#7-python-api)
|
| 90 |
+
8. [About this release](#8-about-this-release)
|
| 91 |
+
9. [Quality and known behavior](#9-quality-and-known-behavior)
|
| 92 |
+
10. [Safety](#10-safety)
|
| 93 |
+
11. [Training lineage and data](#11-training-lineage-and-data)
|
| 94 |
+
12. [Citation](#12-citation)
|
| 95 |
+
13. [Licenses](#13-licenses)
|
| 96 |
+
14. [Reproducibility and artifact identity](#14-reproducibility-and-artifact-identity)
|
| 97 |
+
|
| 98 |
+
## 1. Examples
|
| 99 |
|
| 100 |
### Prompt gallery
|
| 101 |
|
|
|
|
| 133 |
|
| 134 |

|
| 135 |
|
| 136 |
+
## 2. Small-model benchmark
|
| 137 |
+
|
| 138 |
+
Clover is compared with its pinned BK-SDM-Tiny-2M base and two public
|
| 139 |
+
same-family references using 16 prompts, identical seeds, 512×512 output, 30
|
| 140 |
+
DDIM steps, guidance 7.5, and a shared NVIDIA A10G runtime. The measurement is
|
| 141 |
+
an engineering comparison, not a human-preference leaderboard.
|
| 142 |
+
|
| 143 |
+
| Model | U-Net parameters | Mean latency | Peak CUDA | Mean CLIP cosine |
|
| 144 |
+
|---|---:|---:|---:|---:|
|
| 145 |
+
| [Clover Image Tiny](https://huggingface.co/neonforestmist/Clover-Image-Tiny) | 323.4M | 1.024 s | 2,233 MB | 0.3195 |
|
| 146 |
+
| [BK-SDM-Tiny-2M](https://huggingface.co/nota-ai/bk-sdm-tiny-2m) | 323.4M | 1.027 s | 2,230 MB | 0.3246 |
|
| 147 |
+
| [Segmind Tiny-SD](https://huggingface.co/segmind/tiny-sd) | 323.4M | 1.028 s | 1,649 MB | 0.3345 |
|
| 148 |
+
| [BK-SDM-v2-Tiny](https://huggingface.co/nota-ai/bk-sdm-v2-tiny) | 326.8M | 0.957 s | 2,067 MB | 0.3303 |
|
| 149 |
+
|
| 150 |
+
CLIP cosine is only a prompt-adherence proxy. It is not a human-quality score,
|
| 151 |
+
FID, safety evaluation, or evidence that these models are interchangeable.
|
| 152 |
+
The complete protocol, machine-readable results, and generated examples are in
|
| 153 |
+
[`benchmarks/text-to-image/`](benchmarks/text-to-image/) and the
|
| 154 |
+
[full benchmark report](benchmarks/text-to-image/results/clover-small-model-comparison-20260825/REPORT.md).
|
| 155 |
+
|
| 156 |
+

|
| 157 |
+
|
| 158 |
+
## 3. iPhone and Core ML
|
| 159 |
|
| 160 |
The SwiftUI app in [`Clover-iOS`](Clover-iOS/) follows Apple platform
|
| 161 |
conventions and exposes prompt, negative prompt, steps, guidance, seed, image
|
|
|
|
| 180 |
See [`COREML.md`](COREML.md) for conversion details and
|
| 181 |
[`training/README.md`](training/README.md) for the pinned LoRA jobs.
|
| 182 |
|
| 183 |
+
### 3.1 Inpainting track
|
| 184 |
|
| 185 |
The 9-channel SD 1.4-class inpainting adaptation is trained and packaged separately:
|
| 186 |
[`neonforestmist/Clover-Image-Tiny-Inpaint`](https://huggingface.co/neonforestmist/Clover-Image-Tiny-Inpaint).
|
|
|
|
| 196 |
96-pixel mask-context crop; the runtime composites through the exact mask so
|
| 197 |
unmasked pixels remain unchanged.
|
| 198 |
|
| 199 |
+
## 4. Run locally
|
| 200 |
|
| 201 |
Download once, then generate offline with the bundled runner. Python 3.11 and
|
| 202 |
3.12 are supported.
|
| 203 |
|
| 204 |
+
### 4.1 macOS — Apple silicon
|
| 205 |
|
| 206 |
~~~bash
|
| 207 |
mkdir clover-image-tiny-local
|
|
|
|
| 232 |
|
| 233 |
Use `python3.11` instead if that is the installed supported Python.
|
| 234 |
|
| 235 |
+
### 4.2 Windows — PowerShell
|
| 236 |
|
| 237 |
~~~powershell
|
| 238 |
mkdir clover-image-tiny-local
|
|
|
|
| 264 |
Use `py -3.11` if needed. With `--device auto`, the runner selects an
|
| 265 |
available NVIDIA CUDA GPU and otherwise uses CPU.
|
| 266 |
|
| 267 |
+
### 4.3 Linux
|
| 268 |
|
| 269 |
~~~bash
|
| 270 |
mkdir clover-image-tiny-local
|
|
|
|
| 293 |
uses CPU. After the first download, `--local-files-only` prevents network
|
| 294 |
access during generation.
|
| 295 |
|
| 296 |
+
## 5. Generation controls
|
| 297 |
|
| 298 |
The command above is ready to copy. Change these flags to explore the model:
|
| 299 |
|
|
|
|
| 324 |
|
| 325 |
Run `python model/examples/generate.py --help` for the complete CLI reference.
|
| 326 |
|
| 327 |
+
## 6. Hardware and operating systems
|
| 328 |
|
| 329 |
| System | Automatic backend | Precision | Current evidence |
|
| 330 |
|---|---|---|---|
|
|
|
|
| 344 |
631,341,056 bytes. This is a measured point, not a minimum-RAM claim. No Core
|
| 345 |
ML package is required for the Python path.
|
| 346 |
|
| 347 |
+
## 7. Python API
|
| 348 |
|
| 349 |
~~~python
|
| 350 |
import torch
|
|
|
|
| 380 |
Seeded generation is repeatable within the selected runtime. Different
|
| 381 |
devices, dtypes, kernels, and dependency builds can produce different pixels.
|
| 382 |
|
| 383 |
+
## 8. About this release
|
| 384 |
|
| 385 |
Clover Image Tiny is a conventional knowledge-distillation checkpoint trained
|
| 386 |
for 500 optimizer steps on an exact licensed 1,000-pair calibration set. The
|
|
|
|
| 399 |
artifacts and style adapters are versioned in the separate repositories linked
|
| 400 |
above.
|
| 401 |
|
| 402 |
+
## 9. Quality and known behavior
|
| 403 |
|
| 404 |
- The included gallery demonstrates recognizable subjects across colorful
|
| 405 |
scenes, products, food, an animal, a landscape, and an interior.
|
|
|
|
| 411 |
controlled benchmark or broad human-preference study.
|
| 412 |
- Resolution and batch size multiply memory use.
|
| 413 |
|
| 414 |
+
## 10. Safety
|
| 415 |
|
| 416 |
The upstream safety checker is packaged and enabled in both the supported
|
| 417 |
runner and hosted demo. A flagged output may be returned as a black placeholder;
|
|
|
|
| 424 |
identity claims, medical or legal conclusions, harassment, exploitation,
|
| 425 |
illegal activity, or uses prohibited by CreativeML OpenRAIL-M.
|
| 426 |
|
| 427 |
+
## 11. Training lineage and data
|
| 428 |
|
| 429 |
- Clover fine-tuning data: exactly 1,000 accepted image-caption pairs from
|
| 430 |
`Spawning/PD3M@2a5eb24a8dccf245acd8e56341761aee06da0bdf`
|
|
|
|
| 444 |
See `DATA_PROVENANCE.md` for the portable manifest identity and
|
| 445 |
`MODEL_DATA_LICENSES.md` for the complete component ledger.
|
| 446 |
|
| 447 |
+
## 12. Citation
|
| 448 |
|
| 449 |
If Clover Image Tiny is useful in your work, please cite the model release:
|
| 450 |
|
|
|
|
| 457 |
}
|
| 458 |
```
|
| 459 |
|
| 460 |
+
## 13. Licenses
|
| 461 |
|
| 462 |
The model weights are a derivative under **CreativeML OpenRAIL-M**. The example
|
| 463 |
runner and packaging code are under **Apache-2.0**. Dataset and item-level terms
|
|
|
|
| 469 |
evidence, its panel-generation provenance is not claimed, and this package
|
| 470 |
does not grant a downstream reuse license for it.
|
| 471 |
|
| 472 |
+
## 14. Reproducibility and artifact identity
|
| 473 |
|
| 474 |
| Field | Value |
|
| 475 |
|---|---|
|
benchmark_text_to_image.py
ADDED
|
@@ -0,0 +1,283 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""Run a reproducible small-model text-to-image comparison on Modal.
|
| 3 |
+
|
| 4 |
+
The benchmark is intentionally modest: it compares prompt adherence and
|
| 5 |
+
runtime behavior under one shared recipe. It is not a human-preference study,
|
| 6 |
+
FID benchmark, or claim of overall image quality.
|
| 7 |
+
"""
|
| 8 |
+
|
| 9 |
+
from __future__ import annotations
|
| 10 |
+
|
| 11 |
+
import json
|
| 12 |
+
import os
|
| 13 |
+
import sys
|
| 14 |
+
import time
|
| 15 |
+
from pathlib import Path
|
| 16 |
+
|
| 17 |
+
import modal
|
| 18 |
+
|
| 19 |
+
APP_NAME = "clover-image-tiny-text-to-image-benchmark"
|
| 20 |
+
OUTPUT_VOLUME_NAME = "clover-image-tiny-text-to-image-benchmark-output"
|
| 21 |
+
CACHE_VOLUME_NAME = "clover-image-tiny-text-to-image-benchmark-cache"
|
| 22 |
+
OUTPUT_ROOT = Path("/outputs")
|
| 23 |
+
CACHE_ROOT = Path("/cache")
|
| 24 |
+
CLIP_MODEL_ID = "openai/clip-vit-base-patch32"
|
| 25 |
+
|
| 26 |
+
MODELS = {
|
| 27 |
+
"clover": "neonforestmist/Clover-Image-Tiny",
|
| 28 |
+
"base_bk_sdm_tiny_2m": "nota-ai/bk-sdm-tiny-2m",
|
| 29 |
+
"segmind_tiny_sd": "segmind/tiny-sd",
|
| 30 |
+
"bk_sdm_v2_tiny": "nota-ai/bk-sdm-v2-tiny",
|
| 31 |
+
}
|
| 32 |
+
|
| 33 |
+
PROMPTS = [
|
| 34 |
+
"a glass of red wine",
|
| 35 |
+
"a tiny glass greenhouse glowing in a moonlit garden",
|
| 36 |
+
"daisies in a blue ceramic pot",
|
| 37 |
+
"snowy mountains under a cloudy sky",
|
| 38 |
+
"a desert with a big moon in the sky",
|
| 39 |
+
"a bouquet of blue flowers",
|
| 40 |
+
"a stained-glass window of a starry night",
|
| 41 |
+
"an origami heart on textured paper",
|
| 42 |
+
"an anime boy with light blue hair and eyes",
|
| 43 |
+
"a red bicycle beside a yellow cottage",
|
| 44 |
+
"a small blue cat beside a quiet garden stream",
|
| 45 |
+
"a bowl of colorful fruit on a wooden table",
|
| 46 |
+
"an astronaut riding a horse in space",
|
| 47 |
+
"a lighthouse during a storm",
|
| 48 |
+
"a cozy cabin in a snowy forest",
|
| 49 |
+
"a vintage train station at sunset",
|
| 50 |
+
]
|
| 51 |
+
|
| 52 |
+
SEEDS = [1337 + index for index in range(len(PROMPTS))]
|
| 53 |
+
STEPS = 30
|
| 54 |
+
GUIDANCE_SCALE = 7.5
|
| 55 |
+
WIDTH = 512
|
| 56 |
+
HEIGHT = 512
|
| 57 |
+
|
| 58 |
+
image = (
|
| 59 |
+
modal.Image.debian_slim(python_version="3.11")
|
| 60 |
+
.pip_install(
|
| 61 |
+
"accelerate==1.14.0",
|
| 62 |
+
"diffusers==0.39.0",
|
| 63 |
+
"huggingface_hub==0.36.2",
|
| 64 |
+
"numpy==2.2.6",
|
| 65 |
+
"pillow==12.3.0",
|
| 66 |
+
"safetensors==0.8.0",
|
| 67 |
+
"torch==2.7.0",
|
| 68 |
+
"torchvision==0.22.0",
|
| 69 |
+
"transformers==4.57.6",
|
| 70 |
+
)
|
| 71 |
+
)
|
| 72 |
+
|
| 73 |
+
output_volume = modal.Volume.from_name(OUTPUT_VOLUME_NAME, create_if_missing=True)
|
| 74 |
+
cache_volume = modal.Volume.from_name(CACHE_VOLUME_NAME, create_if_missing=True)
|
| 75 |
+
app = modal.App(
|
| 76 |
+
APP_NAME,
|
| 77 |
+
image=image,
|
| 78 |
+
volumes={str(OUTPUT_ROOT): output_volume, str(CACHE_ROOT): cache_volume},
|
| 79 |
+
)
|
| 80 |
+
|
| 81 |
+
|
| 82 |
+
def _module_parameters(module: object) -> int:
|
| 83 |
+
if module is None or not hasattr(module, "parameters"):
|
| 84 |
+
return 0
|
| 85 |
+
return sum(parameter.numel() for parameter in module.parameters())
|
| 86 |
+
|
| 87 |
+
|
| 88 |
+
def _load_pipeline(model_id: str):
|
| 89 |
+
import torch
|
| 90 |
+
from diffusers import DDIMScheduler, DiffusionPipeline
|
| 91 |
+
|
| 92 |
+
pipe = DiffusionPipeline.from_pretrained(
|
| 93 |
+
model_id,
|
| 94 |
+
torch_dtype=torch.float16,
|
| 95 |
+
cache_dir=str(CACHE_ROOT / "huggingface"),
|
| 96 |
+
)
|
| 97 |
+
# A single scheduler makes the comparison recipe explicit and portable.
|
| 98 |
+
pipe.scheduler = DDIMScheduler.from_config(pipe.scheduler.config)
|
| 99 |
+
pipe = pipe.to("cuda")
|
| 100 |
+
pipe.set_progress_bar_config(disable=True)
|
| 101 |
+
return pipe
|
| 102 |
+
|
| 103 |
+
|
| 104 |
+
@app.function(
|
| 105 |
+
gpu="A10G",
|
| 106 |
+
timeout=4 * 60 * 60,
|
| 107 |
+
cpu=8,
|
| 108 |
+
memory=32768,
|
| 109 |
+
)
|
| 110 |
+
def benchmark(*, output_name: str) -> str:
|
| 111 |
+
import torch
|
| 112 |
+
from PIL import Image
|
| 113 |
+
from transformers import CLIPModel, CLIPProcessor
|
| 114 |
+
|
| 115 |
+
output_dir = OUTPUT_ROOT / output_name
|
| 116 |
+
if output_dir.exists():
|
| 117 |
+
raise RuntimeError(f"Benchmark output already exists: {output_dir}")
|
| 118 |
+
output_dir.mkdir(parents=True)
|
| 119 |
+
(output_dir / "images").mkdir()
|
| 120 |
+
(output_dir / "prompts.json").write_text(
|
| 121 |
+
json.dumps(
|
| 122 |
+
{
|
| 123 |
+
"prompts": PROMPTS,
|
| 124 |
+
"seeds": SEEDS,
|
| 125 |
+
"recipe": {
|
| 126 |
+
"scheduler": "DDIMScheduler",
|
| 127 |
+
"steps": STEPS,
|
| 128 |
+
"guidance_scale": GUIDANCE_SCALE,
|
| 129 |
+
"width": WIDTH,
|
| 130 |
+
"height": HEIGHT,
|
| 131 |
+
"num_images_per_prompt": 1,
|
| 132 |
+
"negative_prompt": "",
|
| 133 |
+
},
|
| 134 |
+
},
|
| 135 |
+
indent=2,
|
| 136 |
+
)
|
| 137 |
+
+ "\n"
|
| 138 |
+
)
|
| 139 |
+
|
| 140 |
+
env = os.environ.copy()
|
| 141 |
+
env.update(
|
| 142 |
+
{
|
| 143 |
+
"HF_HOME": str(CACHE_ROOT / "huggingface"),
|
| 144 |
+
"HF_HUB_CACHE": str(CACHE_ROOT / "huggingface" / "hub"),
|
| 145 |
+
"TOKENIZERS_PARALLELISM": "false",
|
| 146 |
+
"PYTHONUNBUFFERED": "1",
|
| 147 |
+
}
|
| 148 |
+
)
|
| 149 |
+
os.environ.update(env)
|
| 150 |
+
all_records: dict[str, dict] = {}
|
| 151 |
+
|
| 152 |
+
for label, model_id in MODELS.items():
|
| 153 |
+
print(f"Loading {label}: {model_id}", flush=True)
|
| 154 |
+
pipe = _load_pipeline(model_id)
|
| 155 |
+
unet_parameters = _module_parameters(pipe.unet)
|
| 156 |
+
total_parameters = sum(
|
| 157 |
+
_module_parameters(module)
|
| 158 |
+
for module in (
|
| 159 |
+
getattr(pipe, "unet", None),
|
| 160 |
+
getattr(pipe, "text_encoder", None),
|
| 161 |
+
getattr(pipe, "vae", None),
|
| 162 |
+
getattr(pipe, "safety_checker", None),
|
| 163 |
+
)
|
| 164 |
+
)
|
| 165 |
+
|
| 166 |
+
# Warm up kernels before measuring. The warmup image is discarded.
|
| 167 |
+
warmup_generator = torch.Generator(device="cuda").manual_seed(7)
|
| 168 |
+
with torch.inference_mode():
|
| 169 |
+
pipe(
|
| 170 |
+
"a simple red apple",
|
| 171 |
+
num_inference_steps=4,
|
| 172 |
+
guidance_scale=GUIDANCE_SCALE,
|
| 173 |
+
height=HEIGHT,
|
| 174 |
+
width=WIDTH,
|
| 175 |
+
generator=warmup_generator,
|
| 176 |
+
)
|
| 177 |
+
torch.cuda.synchronize()
|
| 178 |
+
torch.cuda.reset_peak_memory_stats()
|
| 179 |
+
|
| 180 |
+
model_dir = output_dir / "images" / label
|
| 181 |
+
model_dir.mkdir()
|
| 182 |
+
records = []
|
| 183 |
+
for index, (prompt, seed) in enumerate(zip(PROMPTS, SEEDS, strict=True)):
|
| 184 |
+
generator = torch.Generator(device="cuda").manual_seed(seed)
|
| 185 |
+
start = time.perf_counter()
|
| 186 |
+
with torch.inference_mode():
|
| 187 |
+
result = pipe(
|
| 188 |
+
prompt,
|
| 189 |
+
negative_prompt="",
|
| 190 |
+
num_inference_steps=STEPS,
|
| 191 |
+
guidance_scale=GUIDANCE_SCALE,
|
| 192 |
+
height=HEIGHT,
|
| 193 |
+
width=WIDTH,
|
| 194 |
+
generator=generator,
|
| 195 |
+
)
|
| 196 |
+
torch.cuda.synchronize()
|
| 197 |
+
elapsed = time.perf_counter() - start
|
| 198 |
+
filename = f"{index + 1:02d}.png"
|
| 199 |
+
result.images[0].save(model_dir / filename, format="PNG")
|
| 200 |
+
safety = getattr(result, "nsfw_content_detected", None)
|
| 201 |
+
records.append(
|
| 202 |
+
{
|
| 203 |
+
"index": index,
|
| 204 |
+
"prompt": prompt,
|
| 205 |
+
"seed": seed,
|
| 206 |
+
"filename": f"images/{label}/{filename}",
|
| 207 |
+
"latency_seconds": elapsed,
|
| 208 |
+
"nsfw_content_detected": safety[0] if isinstance(safety, list) else None,
|
| 209 |
+
}
|
| 210 |
+
)
|
| 211 |
+
print(f"{label} {index + 1}/{len(PROMPTS)} {elapsed:.3f}s", flush=True)
|
| 212 |
+
|
| 213 |
+
peak_memory_mb = torch.cuda.max_memory_allocated() / (1024 * 1024)
|
| 214 |
+
all_records[label] = {
|
| 215 |
+
"model_id": model_id,
|
| 216 |
+
"unet_parameters": unet_parameters,
|
| 217 |
+
"pipeline_parameters": total_parameters,
|
| 218 |
+
"mean_latency_seconds": sum(r["latency_seconds"] for r in records) / len(records),
|
| 219 |
+
"median_latency_seconds": sorted(r["latency_seconds"] for r in records)[len(records) // 2],
|
| 220 |
+
"peak_cuda_allocated_mb": peak_memory_mb,
|
| 221 |
+
"images": records,
|
| 222 |
+
}
|
| 223 |
+
del pipe
|
| 224 |
+
torch.cuda.empty_cache()
|
| 225 |
+
|
| 226 |
+
print("Scoring generated images with CLIP", flush=True)
|
| 227 |
+
processor = CLIPProcessor.from_pretrained(
|
| 228 |
+
CLIP_MODEL_ID,
|
| 229 |
+
cache_dir=str(CACHE_ROOT / "huggingface"),
|
| 230 |
+
)
|
| 231 |
+
clip_model = CLIPModel.from_pretrained(
|
| 232 |
+
CLIP_MODEL_ID,
|
| 233 |
+
torch_dtype=torch.float16,
|
| 234 |
+
cache_dir=str(CACHE_ROOT / "huggingface"),
|
| 235 |
+
).to("cuda")
|
| 236 |
+
clip_model.eval()
|
| 237 |
+
text_inputs = processor(text=PROMPTS, return_tensors="pt", padding=True).to("cuda")
|
| 238 |
+
with torch.inference_mode():
|
| 239 |
+
text_features = clip_model.get_text_features(**text_inputs)
|
| 240 |
+
text_features = text_features / text_features.norm(dim=-1, keepdim=True)
|
| 241 |
+
|
| 242 |
+
for label, summary in all_records.items():
|
| 243 |
+
images = [Image.open(output_dir / record["filename"]).convert("RGB") for record in summary["images"]]
|
| 244 |
+
scores = []
|
| 245 |
+
for image_item, prompt in zip(images, PROMPTS, strict=True):
|
| 246 |
+
image_inputs = processor(images=image_item, return_tensors="pt").to("cuda")
|
| 247 |
+
with torch.inference_mode():
|
| 248 |
+
image_features = clip_model.get_image_features(**image_inputs)
|
| 249 |
+
image_features = image_features / image_features.norm(dim=-1, keepdim=True)
|
| 250 |
+
prompt_index = PROMPTS.index(prompt)
|
| 251 |
+
scores.append(float((image_features @ text_features[prompt_index : prompt_index + 1].T).item()))
|
| 252 |
+
summary["clip_prompt_cosine_mean"] = sum(scores) / len(scores)
|
| 253 |
+
summary["clip_prompt_cosine_median"] = sorted(scores)[len(scores) // 2]
|
| 254 |
+
summary["clip_prompt_cosine_scores"] = scores
|
| 255 |
+
|
| 256 |
+
report = {
|
| 257 |
+
"schema_version": 1,
|
| 258 |
+
"benchmark": "clover-image-tiny-small-model-comparison",
|
| 259 |
+
"status": "completed",
|
| 260 |
+
"hardware": "Modal NVIDIA A10G",
|
| 261 |
+
"public_identity_note": "The compute account identity is intentionally omitted.",
|
| 262 |
+
"protocol": {
|
| 263 |
+
"prompt_count": len(PROMPTS),
|
| 264 |
+
"resolution": f"{WIDTH}x{HEIGHT}",
|
| 265 |
+
"scheduler": "DDIMScheduler",
|
| 266 |
+
"steps": STEPS,
|
| 267 |
+
"guidance_scale": GUIDANCE_SCALE,
|
| 268 |
+
"negative_prompt": "",
|
| 269 |
+
"same_prompts_and_seeds": True,
|
| 270 |
+
"metric_note": "CLIP cosine is a prompt-adherence proxy, not a human-quality score or a broad benchmark.",
|
| 271 |
+
},
|
| 272 |
+
"models": all_records,
|
| 273 |
+
}
|
| 274 |
+
(output_dir / "results.json").write_text(json.dumps(report, indent=2) + "\n")
|
| 275 |
+
output_volume.commit()
|
| 276 |
+
cache_volume.commit()
|
| 277 |
+
return str(output_dir)
|
| 278 |
+
|
| 279 |
+
|
| 280 |
+
@app.local_entrypoint()
|
| 281 |
+
def main(output_name: str = f"clover-small-model-comparison-{int(time.time())}") -> None:
|
| 282 |
+
result = benchmark.remote(output_name=output_name)
|
| 283 |
+
print(f"Benchmark output: {result}")
|
benchmarks/text-to-image/README.md
ADDED
|
@@ -0,0 +1,37 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Clover Image Tiny small-model comparison
|
| 2 |
+
|
| 3 |
+
This benchmark compares Clover Image Tiny with its pinned BK-SDM-Tiny-2M base
|
| 4 |
+
and two public Diffusers checkpoints in the same broad SD-tiny family:
|
| 5 |
+
|
| 6 |
+
- `neonforestmist/Clover-Image-Tiny`
|
| 7 |
+
- `nota-ai/bk-sdm-tiny-2m`
|
| 8 |
+
- `segmind/tiny-sd`
|
| 9 |
+
- `nota-ai/bk-sdm-v2-tiny`
|
| 10 |
+
|
| 11 |
+
The comparison uses the same 16 prompts, seeds, 512×512 resolution, empty
|
| 12 |
+
negative prompt, 30 DDIM steps, guidance scale 7.5, and one NVIDIA A10G
|
| 13 |
+
runtime. It records U-Net and loaded-pipeline parameter counts, generation
|
| 14 |
+
latency, peak CUDA allocation, and CLIP image/text cosine similarity.
|
| 15 |
+
|
| 16 |
+
CLIP cosine is used only as a prompt-adherence proxy. It is not a human
|
| 17 |
+
preference score, FID, a safety evaluation, or a claim that the models are
|
| 18 |
+
identical in training data, scheduler defaults, or intended use.
|
| 19 |
+
|
| 20 |
+
Run it with:
|
| 21 |
+
|
| 22 |
+
```bash
|
| 23 |
+
modal run benchmark_text_to_image.py
|
| 24 |
+
```
|
| 25 |
+
|
| 26 |
+
The Modal app writes a timestamped result directory to its private output
|
| 27 |
+
volume. Download a completed run with:
|
| 28 |
+
|
| 29 |
+
```bash
|
| 30 |
+
modal volume get clover-image-tiny-text-to-image-benchmark-output \
|
| 31 |
+
clover-small-model-comparison-<timestamp> \
|
| 32 |
+
benchmarks/text-to-image/results
|
| 33 |
+
```
|
| 34 |
+
|
| 35 |
+
The public model card should include only the exported `results.json`, selected
|
| 36 |
+
example images, the protocol, and the benchmark caveats. Compute-account
|
| 37 |
+
identity is intentionally not part of the artifact.
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/REPORT.md
ADDED
|
@@ -0,0 +1,53 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Clover Image Tiny — small-model comparison
|
| 2 |
+
|
| 3 |
+
This is a same-recipe engineering comparison, not a human-preference or
|
| 4 |
+
state-of-the-art benchmark. It places Clover Image Tiny beside its pinned base
|
| 5 |
+
and two public Diffusers checkpoints with the same broad SD-tiny U-Net shape.
|
| 6 |
+
|
| 7 |
+
## Protocol
|
| 8 |
+
|
| 9 |
+
- 16 fixed prompts and seeds `1337`–`1352`
|
| 10 |
+
- 512 × 512 output, empty negative prompt
|
| 11 |
+
- DDIM scheduler, 30 steps, guidance scale 7.5
|
| 12 |
+
- One image per prompt on the same NVIDIA A10G runtime
|
| 13 |
+
- CLIP image/text cosine similarity as a prompt-adherence proxy
|
| 14 |
+
- Latency measured after one warm-up generation per model
|
| 15 |
+
|
| 16 |
+
CLIP cosine is not a human-quality score, FID, a safety evaluation, or proof
|
| 17 |
+
that one model is generally better. The models have different training data,
|
| 18 |
+
text encoders, and original release goals.
|
| 19 |
+
|
| 20 |
+
## Results
|
| 21 |
+
|
| 22 |
+
| Model | U-Net parameters | Loaded pipeline parameters | Mean latency (s) | Peak CUDA (MB) | Mean CLIP cosine |
|
| 23 |
+
|---|---:|---:|---:|---:|---:|
|
| 24 |
+
| [Clover Image Tiny](https://huggingface.co/neonforestmist/Clover-Image-Tiny) | 323,384,964 | 834,080,895 | 1.024 | 2,233.0 | 0.3195 |
|
| 25 |
+
| [BK-SDM-Tiny-2M](https://huggingface.co/nota-ai/bk-sdm-tiny-2m) | 323,384,964 | 834,080,895 | 1.027 | 2,230.0 | 0.3246 |
|
| 26 |
+
| [Segmind Tiny-SD](https://huggingface.co/segmind/tiny-sd) | 323,384,964 | 530,099,307 | 1.028 | 1,648.9 | 0.3345 |
|
| 27 |
+
| [BK-SDM-v2-Tiny](https://huggingface.co/nota-ai/bk-sdm-v2-tiny) | 326,825,604 | 750,867,307 | 0.957 | 2,067.1 | 0.3303 |
|
| 28 |
+
|
| 29 |
+
The result is best read as scale and runtime context: Clover is a compact
|
| 30 |
+
512×512 model with a 323.4M-parameter denoiser, and its scores are in the same
|
| 31 |
+
range as these similarly sized references under this limited recipe. This is
|
| 32 |
+
not evidence that Clover wins a broad quality contest, nor that a larger model
|
| 33 |
+
would be unnecessary.
|
| 34 |
+
|
| 35 |
+
## Visual examples
|
| 36 |
+
|
| 37 |
+

|
| 38 |
+
|
| 39 |
+
The full generated set contains all 16 prompts for all four models under
|
| 40 |
+
`images/`. The exact machine-readable report is [`results.json`](results.json)
|
| 41 |
+
and the prompt/seed manifest is [`prompts.json`](prompts.json).
|
| 42 |
+
|
| 43 |
+
## Reproduction
|
| 44 |
+
|
| 45 |
+
From the Clover source repository:
|
| 46 |
+
|
| 47 |
+
```bash
|
| 48 |
+
modal run benchmark_text_to_image.py \
|
| 49 |
+
--output-name clover-small-model-comparison-<timestamp>
|
| 50 |
+
```
|
| 51 |
+
|
| 52 |
+
The benchmark intentionally omits the compute-account identity from its public
|
| 53 |
+
report and artifacts.
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/contact-sheet.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/01.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/02.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/03.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/04.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/05.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/06.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/07.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/08.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/09.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/10.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/11.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/12.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/13.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/14.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/15.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/base_bk_sdm_tiny_2m/16.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/01.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/02.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/03.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/04.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/05.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/06.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/07.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/08.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/09.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/10.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/11.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/12.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/13.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/14.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/15.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/bk_sdm_v2_tiny/16.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/clover/01.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/clover/02.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/clover/03.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/clover/04.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/clover/05.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/clover/06.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/clover/07.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/clover/08.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/clover/09.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/clover/10.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/clover/11.png
ADDED
|
Git LFS Details
|
benchmarks/text-to-image/results/clover-small-model-comparison-20260825/images/clover/12.png
ADDED
|
Git LFS Details
|