Running Repro: Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents (GUI-RobustEval + RoTS) ๐ฏ Explore experiment logs and collaborate with an AI agent
Running Repro: QuArch - A Benchmark for Evaluating LLM Reasoning in Computer Architecture ๐ฏ Explore benchmark logbook and sync findings with an AI agent
Running Repro - MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems ๐ฏ Browse MemoryBench logs and collaborate with your coding agent
Running Repro: Towards Professional-Grade Financial Agents: Benchmarking, Tooling, and Structured Reasoning ๐ฏ Collaborate with an AI agent via a shared logbook