Abstract
Large language model (LLM) coding agents have advanced test generation across diverse programming tasks. However, the common practice of evaluating tests against a single reference solution overlooks alternative valid implementations and can overstate test quality. We introduce TestPrism, comprising 300 test tasks from 17 sources and 3000 candidate implementations, evenly split between valid and invalid solutions. Its primary metric, Joint Success Function, requires the generated tests to fail on the initial program state, accept every valid candidate, and reject every invalid candidate. Across fourteen baseline coding agent configurations, Joint Success Function reaches only 28.00%, whereas single reference success reaches 59.67%. Our analysis reveals missed behaviors, unsupported assertions, and faulty test construction. To address these weaknesses, we introduce TestHelix, which combines heterogeneous synthesis of test and repair pairs with peer cross validation and recursive self improvement (RSI). Across two models, TestHelix improves Joint Success Function by 8.67 to 9.00 percentage points over the native harness comparators in the TestHelix evaluation
Community
Large language model (LLM) coding agents have advanced test generation across diverse programming tasks. However, the common practice of evaluating tests against a single reference
solution overlooks alternative valid implementations and can overstate test quality. We introduce
TestPrism, comprising 300 test tasks from 17 sources and 3000 candidate implementations, evenly
split between valid and invalid solutions. Its primary metric, Joint Success Function, requires the
generated tests to fail on the initial program state, accept every valid candidate, and reject every
invalid candidate. Across fourteen baseline coding agent configurations, Joint Success Function
reaches only 28.00%, whereas single reference success reaches 59.67%. Our analysis reveals missed
behaviors, unsupported assertions, and faulty test construction. To address these weaknesses, we
introduce TestHelix, which combines heterogeneous synthesis of test and repair pairs with peer
cross validation and recursive self improvement (RSI). Across two models, TestHelix improves
Joint Success Function by 8.67 to 9.00 percentage points over the native harness comparators in the
TestHelix evaluation.
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper

