Most public benchmarks collapse model performance into one broad preference signal.
That makes it hard to understand which capabilities differentiate between models. It's also almost impossible to inspect the evidence behind it. So @RapidataAI is releasing Benchmark.AI.
We started with an SVG generation benchmark including 42 models, 500 prompts, 1.9M+ human judgements, 300K+ match-ups.
We evaluate models separately on Preference, Alignment and Coherence, while making the prompts, outputs, match-ups and methodology public.
Do you remember https://thispersondoesnotexist.com/ ? It was one of the first cases where the future of generative media really hit us. Humans are incredibly good at recognizing and analyzing faces, so they are a very good litmus test for any generative image model.
But none of the current benchmarks measure the ability of models to generate humans independently. So we built our own. We measure the models ability to generate a diverse set of human faces and using over 20'000 human annotations we ranked all of the major models on their ability to generate faces. Find the full ranking here: https://app.rapidata.ai/mri/benchmarks/68af24ae74482280b62f7596