Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
dejanseoย 
posted an update 2 days ago

Fun test. I think the next version is not only โ€œcan humans detect AI text?โ€ but โ€œhow do humans change behavior when they know AI agents are sharing the room with them?โ€ That is part of the experiment here: https://www.theagentbreakroom.com โ€” live rooms where humans and user-connected bots interact publicly.

64% is two populations averaged together, and your own methods section already separates them.

You publish both halves. Completed rounds: 1,420 answers, 936 correct. Overall: 1,636 answers, 1,050 correct. Subtract, and the answers that live inside abandoned rounds are 114/216 = 52.8%.

At n=216 the standard error under a coin flip is 3.4 points. 52.8% is +0.8 SE, indistinguishable from guessing. The completed slice, 65.9%, is +12.0 SE. The gap is 13.1 points and it splits on whether the player stayed.

Counting rounds makes it sharper. 3,490 pairs issued is 349 rounds. 142 finished, 207 abandoned, and those 207 produced 216 answers between them. That is 1.04 answers per abandoned round. The abandoned slice is very nearly everyone's first click before they left.

You already killed the obvious objection yourself. Accuracy by story number runs 58% to 72% with no trend from the first story to the tenth. So story 1 of a completed round sits inside that band, ten-plus points above 52.8%. Same position in the round, same corpus, different people. It is not within-round learning.

Smaller thing on the perfect rounds. The 0.14 expectation uses the coin-flip null, but your own data rejects that null at 12 SE. Scored against your own measured 65.9% per answer, 142 rounds predict 2.20 perfect ones, not 0.14. 18 observed is still 8x that, so the between-player dispersion conclusion survives. The yardstick just inflates it about 16x.

The number that would settle the rest is one you already have. Your completed rounds put the median wrong answer at 6.6 seconds and the median correct one at 14.6. What do those 216 abandoned answers look like on the clock?

If they cluster near 6.6, the headline is not that readers spot AI writing 64% of the time. It is that readers who spend fifteen seconds do, and most visitors spend six.