Spaces:
Running
Running
| ## Where visual generators know—and where they need search | |
| **SearchGen-Bench** compares image generators across 751 evaluation prompts, knowledge-boundary strata, 22 domains, and 12 recurring failure modes. The benchmark separates 100 **NoSearch** prompts that strong generators can answer from parametric knowledge from 651 **SearchIntensive** prompts whose visual or textual requirements expose gaps that retrieval can address. | |
| The primary metric is **Overall-10**: a prompt-macro average over the canonical ten applicable FF-judge+PP components, reported on a 0–100 scale. Use the scoreboard selector to inspect results by component, domain, or failure mode. | |
| All scores are generated from versioned prompt-level records. Counts and coverage are shown because domains and failure modes are multi-label and some commercial generators have incomplete evaluations. | |