First impressions
11 models in the single-turn human arena.
Single replies, blind human votes, and the original writing tests.
RoleCall Studios builds for AI roleplay. You probably found us through PlotPoints, our open model benchmark, so let's start there. The creation studio and discovery floor are further down.
A leaderboard for how well models actually roleplay: staying in character, holding a scene across many turns, prose quality, and not hijacking your agency. SFW and NSFW, with Round 04 adding tests of willingness across intimacy and graphic violence.
Requesting a model: we'll run just about anything within reason, as long as it's on OpenRouter. We already add the popular ones we come across, and we take requests, so if it's on OR it can go on the board.
Local models and finetunes: we actively want to bench them, and we're hunting for an OpenRouter-style aggregator that hosts them in real abundance. The only blocker is fairness: every local model is set up differently and runs differently on every machine, and mismatched provider setups can quietly poison results. Point us at one consistent aggregator with a deep library and we will run them regularly and wildly. Know one? Tell us in the Community tab.
The public benchmark includes its test harness, judge prompts, scoring rubric, seeds, and results. Boundary-test replies are withheld, and blind ballots are withheld while voting is open. The dataset is published under CC BY-NC 4.0. Its tests measure different parts of the roleplay experience:
Every model runs through OpenRouter to keep providers comparable. Human votes and machine judging both count, and where they disagree is half the fun.
Browse the public data by round, following the benchmark's progression from first impressions to full scenes and harder content. These are different slices of one dataset, with direct links to the relevant tables and files.
11 models in the single-turn human arena.
Single replies, blind human votes, and the original writing tests.
20 models in the full-session human arena.
Full multi-turn scenes: consistency, momentum, agency, and adversarial failures.
21 models in the standard track; 40 in the NSFW track.
A refreshed model pool and a separate NSFW track. Judge findings and human preferences are different signals.
71 models overall: 70 craft-scored, 58 willingness-tested, and 55 ranked by J.
Adult intimacy, bondage play, graphic violence, and boundary probes. Willingness and writing quality are shown separately.
Read the evidence: human votes, judge scores, and willingness measure different things. Boundary requests and aggregate results are public; full boundary-test replies are withheld. Open blind ballots stay withheld until voting closes.
All dataset files · Interactive leaderboard · Original source & attribution
| Project | What it is | Link |
|---|---|---|
| RoleCall | The studio. Build characters, presets & lorebooks, then play end-to-end-encrypted scenes with control over prompts, models, and generation settings. | rolecallstudios.com |
| PlotLight | The discovery floor. Browse, rate and fork the community's characters, presets, lorebooks and personas. No account needed to look around. | plotlightstudios.com |