pixelmodel / EVAL_REPRODUCTION.md
Costi
Add real eval results and reproduction guide
d9a3537
|
Raw
History Blame Contribute Delete
6.02 kB

Reproducing the eval.md numbers

Step-by-step guide to reproduce the FID / CLIP Score numbers in eval.md from scratch. Everything below was run in a throwaway virtualenv outside this repo; none of it is required to just use PixelModel (see README.md for that) — it's only needed if you want to re-run or extend the evaluation.

Uses Python 3.13 (CPU-only; no GPU required, but slower). Total download size is roughly 1.5–2GB (torch, torchvision, CLIP weights, Inception weights) — pick an install location with that much free space.

1. Create an isolated venv

python3.13 -m venv eval-venv
# Windows:
eval-venv\Scripts\activate
# macOS/Linux:
source eval-venv/bin/activate

2. Install dependencies

CPU-only torch keeps the install small (~130MB vs several GB for a CUDA build):

pip install torch torchvision --index-url https://download.pytorch.org/whl/cpu
pip install transformers scipy huggingface_hub "datasets<3" pyarrow pytorch-fid

datasets<3 is pinned because pytorch-fid==0.3.0's dependency chain and the CLIP scoring code below weren't tested against datasets 3.x.

3. Fetch a real COCO caption/image sample

This streams only the first N rows from a public 30K-pair COCO val2014 dataset — it does not download the full ~5GB dataset.

# fetch_coco_sample.py
import os, json, time
from datasets import load_dataset

N = 40
OUT_DIR = "coco_sample"
os.makedirs(OUT_DIR, exist_ok=True)

ds = load_dataset("sayakpaul/coco-30-val-2014", split="train", streaming=True)

meta = []
for i, row in enumerate(ds):
    if i >= N:
        break
    img = row["image"]
    caption = row["caption"]
    fname = f"real_{i:03d}.jpg"
    img.convert("RGB").save(os.path.join(OUT_DIR, fname), format="JPEG", quality=90)
    meta.append({"file": fname, "caption": caption})

with open(os.path.join(OUT_DIR, "captions.json"), "w") as f:
    json.dump(meta, f, indent=2)

Run it: python fetch_coco_sample.py — takes ~10 seconds, ~3.5MB on disk. Raise N for a less noisy (but slower, and eventually much larger) eval — real COCO FID evals typically use N=30,000.

4. Generate PixelModel outputs for those captions

Run from inside this repo (needs model.py and model.png on the path):

# generate_outputs.py
import os, json
import numpy as np
from PIL import Image
import torch
from model import load_model, forward  # from this repo

META_PATH = "coco_sample/captions.json"
OUT_DIR = "gen_sample"
os.makedirs(OUT_DIR, exist_ok=True)

with open(META_PATH) as f:
    meta = json.load(f)

pixels = load_model("model.png")

for i, row in enumerate(meta):
    with torch.no_grad():
        result = forward(pixels, row["caption"])
    arr = (result.numpy() * 255).clip(0, 255).astype(np.uint8)
    Image.fromarray(arr, mode="RGB").save(os.path.join(OUT_DIR, f"gen_{i:03d}.png"))

5. Compute FID

pytorch-fid==0.3.0 calls scipy.linalg.sqrtm(..., disp=False), which recent scipy versions (1.14+) no longer accept — this reimplements the Fréchet distance calculation without that removed kwarg instead of pinning an old scipy:

# compute_fid.py
import numpy as np
from scipy import linalg
from pytorch_fid.fid_score import compute_statistics_of_path
from pytorch_fid.inception import InceptionV3

def calculate_frechet_distance(mu1, sigma1, mu2, sigma2):
    diff = mu1 - mu2
    covmean = linalg.sqrtm(sigma1.dot(sigma2))
    if np.iscomplexobj(covmean):
        covmean = covmean.real
    return diff.dot(diff) + np.trace(sigma1) + np.trace(sigma2) - 2 * np.trace(covmean)

dims = 2048
model = InceptionV3([InceptionV3.BLOCK_INDEX_BY_DIM[dims]]).to("cpu")

# NOTE: real images are JPEGs of varying resolution; pytorch-fid's default
# dataloader batches images without resizing first, which crashes on a batch
# of mismatched sizes — use batch_size=1 to sidestep that.
m1, s1 = compute_statistics_of_path("coco_sample", model, 1, dims, "cpu", 0)
m2, s2 = compute_statistics_of_path("gen_sample", model, 1, dims, "cpu", 0)

fid = calculate_frechet_distance(m1, s1, m2, s2)
print(f"FID: {fid:.4f}")

Expect a LinAlgWarning: Matrix is singular at low sample counts (n < 2048) — this is expected, not a bug; see the caveat in eval.md.

6. Compute CLIP Score

# compute_clip_score.py
import json
import torch
from PIL import Image
from transformers import CLIPModel, CLIPProcessor

with open("coco_sample/captions.json") as f:
    meta = json.load(f)

model = CLIPModel.from_pretrained("openai/clip-vit-base-patch32")
processor = CLIPProcessor.from_pretrained("openai/clip-vit-base-patch32")
model.eval()

scores = []
for i, row in enumerate(meta):
    img = Image.open(f"gen_sample/gen_{i:03d}.png").convert("RGB")
    inputs = processor(text=[row["caption"]], images=img, return_tensors="pt", padding=True, truncation=True)
    with torch.no_grad():
        out = model(**inputs)
    img_emb = out.image_embeds / out.image_embeds.norm(dim=-1, keepdim=True)
    txt_emb = out.text_embeds / out.text_embeds.norm(dim=-1, keepdim=True)
    cos_sim = (img_emb * txt_emb).sum(dim=-1).item()
    scores.append(max(0.0, 100 * cos_sim))

print(f"Mean CLIP Score (n={len(scores)}): {sum(scores)/len(scores):.4f}")

First run downloads openai/clip-vit-base-patch32 (~600MB) from the Hugging Face Hub.

Notes / gotchas hit while building this

  • Disk location matters. All of the above (venv, pip cache, HF cache, downloaded datasets) should point at a drive with several GB free — torch+torchvision+transformers+CLIP+Inception weights add up to ~1.5–2GB. Set PIP_CACHE_DIR, HF_HOME, and TORCH_HOME env vars if your default drive is tight on space.
  • Sample size is the biggest caveat. n=40 is fast and cheap but gives a noisy FID (singular covariance matrix) and a noisy CLIP Score mean. Scaling toward the standard n=30,000 COCO FID protocol means re-running steps 3–6 with a bigger N, which will take proportionally longer and use proportionally more disk/bandwidth.