Evaluating Agents Across Runtime Contracts Collection Evaluation traces, instance-paired teacher traces, and the 14 LoRA adapters (persistent and stateless per family) behind the paper. • 16 items • Updated about 8 hours ago
GitChameleon: Evaluating AI Code Generation Against Python Library Version Incompatibilities Paper • 2507.12367 • Published Jul 16, 2025 • 7
MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources Paper • 2509.25531 • Published Sep 29, 2025 • 11
Agents Learn Their Runtime: Interpreter Persistence as Training-Time Semantics Paper • 2603.01209 • Published Mar 1 • 2