EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness? Paper • 2609.04280 • Published 26 days ago • 32
Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind Paper • 2604.11666 • Published Apr 13 • 4
🔍 Interpretability & Analysis of LMs Collection Outstanding research in LM interpretability and evaluation, summarized • 136 items • Updated May 26 • 120