— Interactive Fiction —

RoleCall Studios

Write alongside AI. Run the scene, not the software.
✦

RoleCall Studios builds for AI roleplay. You probably found us through PlotPoints, our open model benchmark, so let's start there. The creation studio and discovery floor are further down.

The PlotPoints Benchmark

A leaderboard for how well models actually roleplay: staying in character, holding a scene across many turns, prose quality, and not hijacking your agency. SFW and NSFW, with Round 04 adding tests of willingness across intimacy and graphic violence.

Requesting a model: we'll run just about anything within reason, as long as it's on OpenRouter. We already add the popular ones we come across, and we take requests, so if it's on OR it can go on the board.

Local models and finetunes: we actively want to bench them, and we're hunting for an OpenRouter-style aggregator that hosts them in real abundance. The only blocker is fairness: every local model is set up differently and runs differently on every machine, and mismatched provider setups can quietly poison results. Point us at one consistent aggregator with a deep library and we will run them regularly and wildly. Know one? Tell us in the Community tab.

The public benchmark includes its test harness, judge prompts, scoring rubric, seeds, and results. Boundary-test replies are withheld, and blind ballots are withheld while voting is open. The dataset is published under CC BY-NC 4.0. Its tests measure different parts of the roleplay experience:

Every model runs through OpenRouter to keep providers comparable. Human votes and machine judging both count, and where they disagree is half the fun.

3,000+
Human Votes
20+
Models
6
Tests

Explore Every Round

Browse the public data by round, following the benchmark's progression from first impressions to full scenes and harder content. These are different slices of one dataset, with direct links to the relevant tables and files.

Read the evidence: human votes, judge scores, and willingness measure different things. Boundary requests and aggregate results are public; full boundary-test replies are withheld. Open blind ballots stay withheld until voting closes.

All dataset files · Interactive leaderboard · Original source & attribution

Also From RoleCall Studios

ProjectWhat it isLink
RoleCallThe studio. Build characters, presets & lorebooks, then play end-to-end-encrypted scenes with control over prompts, models, and generation settings.rolecallstudios.com
PlotLightThe discovery floor. Browse, rate and fork the community's characters, presets, lorebooks and personas. No account needed to look around.plotlightstudios.com

Find Us

© RoleCall Studios LLC · built for writers