Papers
arxiv:2610.12289

TestPrism: Rethinking Test Evaluation Beyond a Single Reference

Published on Oct 8
· Submitted by
Jiaheng Liu
on Oct 9
Authors:
,
,
,
,
,
,
,
,

Abstract

Large language model (LLM) coding agents have advanced test generation across diverse programming tasks. However, the common practice of evaluating tests against a single reference solution overlooks alternative valid implementations and can overstate test quality. We introduce TestPrism, comprising 300 test tasks from 17 sources and 3000 candidate implementations, evenly split between valid and invalid solutions. Its primary metric, Joint Success Function, requires the generated tests to fail on the initial program state, accept every valid candidate, and reject every invalid candidate. Across fourteen baseline coding agent configurations, Joint Success Function reaches only 28.00%, whereas single reference success reaches 59.67%. Our analysis reveals missed behaviors, unsupported assertions, and faulty test construction. To address these weaknesses, we introduce TestHelix, which combines heterogeneous synthesis of test and repair pairs with peer cross validation and recursive self improvement (RSI). Across two models, TestHelix improves Joint Success Function by 8.67 to 9.00 percentage points over the native harness comparators in the TestHelix evaluation

Community

Paper submitter

Large language model (LLM) coding agents have advanced test generation across diverse programming tasks. However, the common practice of evaluating tests against a single reference
solution overlooks alternative valid implementations and can overstate test quality. We introduce
TestPrism, comprising 300 test tasks from 17 sources and 3000 candidate implementations, evenly
split between valid and invalid solutions. Its primary metric, Joint Success Function, requires the
generated tests to fail on the initial program state, accept every valid candidate, and reject every
invalid candidate. Across fourteen baseline coding agent configurations, Joint Success Function
reaches only 28.00%, whereas single reference success reaches 59.67%. Our analysis reveals missed
behaviors, unsupported assertions, and faulty test construction. To address these weaknesses, we
introduce TestHelix, which combines heterogeneous synthesis of test and repair pairs with peer
cross validation and recursive self improvement (RSI). Across two models, TestHelix improves
Joint Success Function by 8.67 to 9.00 percentage points over the native harness comparators in the
TestHelix evaluation.

image

image

Sign up or log in to comment

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.12289 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.12289 in a dataset README.md to link it from this page.

Spaces citing this paper 1

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.