| # Reproducing the eval.md numbers |
|
|
| Step-by-step guide to reproduce the FID / CLIP Score numbers in `eval.md` from scratch. Everything below was run in a throwaway virtualenv outside this repo; none of it is required to just *use* PixelModel (see `README.md` for that) — it's only needed if you want to re-run or extend the evaluation. |
|
|
| Uses Python 3.13 (CPU-only; no GPU required, but slower). Total download size is roughly 1.5–2GB (torch, torchvision, CLIP weights, Inception weights) — pick an install location with that much free space. |
|
|
| ## 1. Create an isolated venv |
|
|
| ```bash |
| python3.13 -m venv eval-venv |
| # Windows: |
| eval-venv\Scripts\activate |
| # macOS/Linux: |
| source eval-venv/bin/activate |
| ``` |
|
|
| ## 2. Install dependencies |
|
|
| CPU-only torch keeps the install small (~130MB vs several GB for a CUDA build): |
|
|
| ```bash |
| pip install torch torchvision --index-url https://download.pytorch.org/whl/cpu |
| pip install transformers scipy huggingface_hub "datasets<3" pyarrow pytorch-fid |
| ``` |
|
|
| `datasets<3` is pinned because `pytorch-fid==0.3.0`'s dependency chain and the CLIP scoring code below weren't tested against `datasets` 3.x. |
|
|
| ## 3. Fetch a real COCO caption/image sample |
|
|
| This streams only the first N rows from a public 30K-pair COCO val2014 dataset — it does **not** download the full ~5GB dataset. |
|
|
| ```python |
| # fetch_coco_sample.py |
| import os, json, time |
| from datasets import load_dataset |
| |
| N = 40 |
| OUT_DIR = "coco_sample" |
| os.makedirs(OUT_DIR, exist_ok=True) |
| |
| ds = load_dataset("sayakpaul/coco-30-val-2014", split="train", streaming=True) |
| |
| meta = [] |
| for i, row in enumerate(ds): |
| if i >= N: |
| break |
| img = row["image"] |
| caption = row["caption"] |
| fname = f"real_{i:03d}.jpg" |
| img.convert("RGB").save(os.path.join(OUT_DIR, fname), format="JPEG", quality=90) |
| meta.append({"file": fname, "caption": caption}) |
| |
| with open(os.path.join(OUT_DIR, "captions.json"), "w") as f: |
| json.dump(meta, f, indent=2) |
| ``` |
|
|
| Run it: `python fetch_coco_sample.py` — takes ~10 seconds, ~3.5MB on disk. Raise `N` for a less noisy (but slower, and eventually much larger) eval — real COCO FID evals typically use N=30,000. |
|
|
| ## 4. Generate PixelModel outputs for those captions |
|
|
| Run from inside this repo (needs `model.py` and `model.png` on the path): |
|
|
| ```python |
| # generate_outputs.py |
| import os, json |
| import numpy as np |
| from PIL import Image |
| import torch |
| from model import load_model, forward # from this repo |
| |
| META_PATH = "coco_sample/captions.json" |
| OUT_DIR = "gen_sample" |
| os.makedirs(OUT_DIR, exist_ok=True) |
| |
| with open(META_PATH) as f: |
| meta = json.load(f) |
| |
| pixels = load_model("model.png") |
| |
| for i, row in enumerate(meta): |
| with torch.no_grad(): |
| result = forward(pixels, row["caption"]) |
| arr = (result.numpy() * 255).clip(0, 255).astype(np.uint8) |
| Image.fromarray(arr, mode="RGB").save(os.path.join(OUT_DIR, f"gen_{i:03d}.png")) |
| ``` |
|
|
| ## 5. Compute FID |
|
|
| `pytorch-fid==0.3.0` calls `scipy.linalg.sqrtm(..., disp=False)`, which recent scipy versions (1.14+) no longer accept — this reimplements the Fréchet distance calculation without that removed kwarg instead of pinning an old scipy: |
|
|
| ```python |
| # compute_fid.py |
| import numpy as np |
| from scipy import linalg |
| from pytorch_fid.fid_score import compute_statistics_of_path |
| from pytorch_fid.inception import InceptionV3 |
| |
| def calculate_frechet_distance(mu1, sigma1, mu2, sigma2): |
| diff = mu1 - mu2 |
| covmean = linalg.sqrtm(sigma1.dot(sigma2)) |
| if np.iscomplexobj(covmean): |
| covmean = covmean.real |
| return diff.dot(diff) + np.trace(sigma1) + np.trace(sigma2) - 2 * np.trace(covmean) |
| |
| dims = 2048 |
| model = InceptionV3([InceptionV3.BLOCK_INDEX_BY_DIM[dims]]).to("cpu") |
| |
| # NOTE: real images are JPEGs of varying resolution; pytorch-fid's default |
| # dataloader batches images without resizing first, which crashes on a batch |
| # of mismatched sizes — use batch_size=1 to sidestep that. |
| m1, s1 = compute_statistics_of_path("coco_sample", model, 1, dims, "cpu", 0) |
| m2, s2 = compute_statistics_of_path("gen_sample", model, 1, dims, "cpu", 0) |
| |
| fid = calculate_frechet_distance(m1, s1, m2, s2) |
| print(f"FID: {fid:.4f}") |
| ``` |
|
|
| Expect a `LinAlgWarning: Matrix is singular` at low sample counts (n < 2048) — this is expected, not a bug; see the caveat in `eval.md`. |
|
|
| ## 6. Compute CLIP Score |
|
|
| ```python |
| # compute_clip_score.py |
| import json |
| import torch |
| from PIL import Image |
| from transformers import CLIPModel, CLIPProcessor |
| |
| with open("coco_sample/captions.json") as f: |
| meta = json.load(f) |
| |
| model = CLIPModel.from_pretrained("openai/clip-vit-base-patch32") |
| processor = CLIPProcessor.from_pretrained("openai/clip-vit-base-patch32") |
| model.eval() |
| |
| scores = [] |
| for i, row in enumerate(meta): |
| img = Image.open(f"gen_sample/gen_{i:03d}.png").convert("RGB") |
| inputs = processor(text=[row["caption"]], images=img, return_tensors="pt", padding=True, truncation=True) |
| with torch.no_grad(): |
| out = model(**inputs) |
| img_emb = out.image_embeds / out.image_embeds.norm(dim=-1, keepdim=True) |
| txt_emb = out.text_embeds / out.text_embeds.norm(dim=-1, keepdim=True) |
| cos_sim = (img_emb * txt_emb).sum(dim=-1).item() |
| scores.append(max(0.0, 100 * cos_sim)) |
| |
| print(f"Mean CLIP Score (n={len(scores)}): {sum(scores)/len(scores):.4f}") |
| ``` |
|
|
| First run downloads `openai/clip-vit-base-patch32` (~600MB) from the Hugging Face Hub. |
|
|
| ## Notes / gotchas hit while building this |
|
|
| - **Disk location matters.** All of the above (venv, pip cache, HF cache, downloaded datasets) should point at a drive with several GB free — `torch`+`torchvision`+`transformers`+CLIP+Inception weights add up to ~1.5–2GB. Set `PIP_CACHE_DIR`, `HF_HOME`, and `TORCH_HOME` env vars if your default drive is tight on space. |
| - **Sample size is the biggest caveat.** n=40 is fast and cheap but gives a noisy FID (singular covariance matrix) and a noisy CLIP Score mean. Scaling toward the standard n=30,000 COCO FID protocol means re-running steps 3–6 with a bigger `N`, which will take proportionally longer and use proportionally more disk/bandwidth. |
|
|