WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents Paper • 2609.27490 • Published 2 days ago • 8