SearchGen-Bench / blobs /intro.md
wufangtai
Adopt ten-component benchmark protocol
a8782a2
|
Raw
History Blame Contribute Delete
871 Bytes

Where visual generators know—and where they need search

SearchGen-Bench compares image generators across 751 evaluation prompts, knowledge-boundary strata, 22 domains, and 12 recurring failure modes. The benchmark separates 100 NoSearch prompts that strong generators can answer from parametric knowledge from 651 SearchIntensive prompts whose visual or textual requirements expose gaps that retrieval can address.

The primary metric is Overall-10: a prompt-macro average over the canonical ten applicable FF-judge+PP components, reported on a 0–100 scale. Use the scoreboard selector to inspect results by component, domain, or failure mode.

All scores are generated from versioned prompt-level records. Counts and coverage are shown because domains and failure modes are multi-label and some commercial generators have incomplete evaluations.