# Reproducing the eval.md numbers Step-by-step guide to reproduce the FID / CLIP Score numbers in `eval.md` from scratch. Everything below was run in a throwaway virtualenv outside this repo; none of it is required to just *use* PixelModel (see `README.md` for that) — it's only needed if you want to re-run or extend the evaluation. Uses Python 3.13 (CPU-only; no GPU required, but slower). Total download size is roughly 1.5–2GB (torch, torchvision, CLIP weights, Inception weights) — pick an install location with that much free space. ## 1. Create an isolated venv ```bash python3.13 -m venv eval-venv # Windows: eval-venv\Scripts\activate # macOS/Linux: source eval-venv/bin/activate ``` ## 2. Install dependencies CPU-only torch keeps the install small (~130MB vs several GB for a CUDA build): ```bash pip install torch torchvision --index-url https://download.pytorch.org/whl/cpu pip install transformers scipy huggingface_hub "datasets<3" pyarrow pytorch-fid ``` `datasets<3` is pinned because `pytorch-fid==0.3.0`'s dependency chain and the CLIP scoring code below weren't tested against `datasets` 3.x. ## 3. Fetch a real COCO caption/image sample This streams only the first N rows from a public 30K-pair COCO val2014 dataset — it does **not** download the full ~5GB dataset. ```python # fetch_coco_sample.py import os, json, time from datasets import load_dataset N = 40 OUT_DIR = "coco_sample" os.makedirs(OUT_DIR, exist_ok=True) ds = load_dataset("sayakpaul/coco-30-val-2014", split="train", streaming=True) meta = [] for i, row in enumerate(ds): if i >= N: break img = row["image"] caption = row["caption"] fname = f"real_{i:03d}.jpg" img.convert("RGB").save(os.path.join(OUT_DIR, fname), format="JPEG", quality=90) meta.append({"file": fname, "caption": caption}) with open(os.path.join(OUT_DIR, "captions.json"), "w") as f: json.dump(meta, f, indent=2) ``` Run it: `python fetch_coco_sample.py` — takes ~10 seconds, ~3.5MB on disk. Raise `N` for a less noisy (but slower, and eventually much larger) eval — real COCO FID evals typically use N=30,000. ## 4. Generate PixelModel outputs for those captions Run from inside this repo (needs `model.py` and `model.png` on the path): ```python # generate_outputs.py import os, json import numpy as np from PIL import Image import torch from model import load_model, forward # from this repo META_PATH = "coco_sample/captions.json" OUT_DIR = "gen_sample" os.makedirs(OUT_DIR, exist_ok=True) with open(META_PATH) as f: meta = json.load(f) pixels = load_model("model.png") for i, row in enumerate(meta): with torch.no_grad(): result = forward(pixels, row["caption"]) arr = (result.numpy() * 255).clip(0, 255).astype(np.uint8) Image.fromarray(arr, mode="RGB").save(os.path.join(OUT_DIR, f"gen_{i:03d}.png")) ``` ## 5. Compute FID `pytorch-fid==0.3.0` calls `scipy.linalg.sqrtm(..., disp=False)`, which recent scipy versions (1.14+) no longer accept — this reimplements the Fréchet distance calculation without that removed kwarg instead of pinning an old scipy: ```python # compute_fid.py import numpy as np from scipy import linalg from pytorch_fid.fid_score import compute_statistics_of_path from pytorch_fid.inception import InceptionV3 def calculate_frechet_distance(mu1, sigma1, mu2, sigma2): diff = mu1 - mu2 covmean = linalg.sqrtm(sigma1.dot(sigma2)) if np.iscomplexobj(covmean): covmean = covmean.real return diff.dot(diff) + np.trace(sigma1) + np.trace(sigma2) - 2 * np.trace(covmean) dims = 2048 model = InceptionV3([InceptionV3.BLOCK_INDEX_BY_DIM[dims]]).to("cpu") # NOTE: real images are JPEGs of varying resolution; pytorch-fid's default # dataloader batches images without resizing first, which crashes on a batch # of mismatched sizes — use batch_size=1 to sidestep that. m1, s1 = compute_statistics_of_path("coco_sample", model, 1, dims, "cpu", 0) m2, s2 = compute_statistics_of_path("gen_sample", model, 1, dims, "cpu", 0) fid = calculate_frechet_distance(m1, s1, m2, s2) print(f"FID: {fid:.4f}") ``` Expect a `LinAlgWarning: Matrix is singular` at low sample counts (n < 2048) — this is expected, not a bug; see the caveat in `eval.md`. ## 6. Compute CLIP Score ```python # compute_clip_score.py import json import torch from PIL import Image from transformers import CLIPModel, CLIPProcessor with open("coco_sample/captions.json") as f: meta = json.load(f) model = CLIPModel.from_pretrained("openai/clip-vit-base-patch32") processor = CLIPProcessor.from_pretrained("openai/clip-vit-base-patch32") model.eval() scores = [] for i, row in enumerate(meta): img = Image.open(f"gen_sample/gen_{i:03d}.png").convert("RGB") inputs = processor(text=[row["caption"]], images=img, return_tensors="pt", padding=True, truncation=True) with torch.no_grad(): out = model(**inputs) img_emb = out.image_embeds / out.image_embeds.norm(dim=-1, keepdim=True) txt_emb = out.text_embeds / out.text_embeds.norm(dim=-1, keepdim=True) cos_sim = (img_emb * txt_emb).sum(dim=-1).item() scores.append(max(0.0, 100 * cos_sim)) print(f"Mean CLIP Score (n={len(scores)}): {sum(scores)/len(scores):.4f}") ``` First run downloads `openai/clip-vit-base-patch32` (~600MB) from the Hugging Face Hub. ## Notes / gotchas hit while building this - **Disk location matters.** All of the above (venv, pip cache, HF cache, downloaded datasets) should point at a drive with several GB free — `torch`+`torchvision`+`transformers`+CLIP+Inception weights add up to ~1.5–2GB. Set `PIP_CACHE_DIR`, `HF_HOME`, and `TORCH_HOME` env vars if your default drive is tight on space. - **Sample size is the biggest caveat.** n=40 is fast and cheap but gives a noisy FID (singular covariance matrix) and a noisy CLIP Score mean. Scaling toward the standard n=30,000 COCO FID protocol means re-running steps 3–6 with a bigger `N`, which will take proportionally longer and use proportionally more disk/bandwidth.