MatrAIx: Simulating the World with 8.3 Billion Persona Agents Paper • 2608.04205 • Published 8 days ago • 27
Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation Paper • 2607.02577 • Published Jun 30
Structured Prompting Enables More Robust Evaluation of Language Models Paper • 2511.20836 • Published Nov 25, 2025
MatrAIx: Simulating the World with 8.3 Billion Persona Agents Paper • 2608.04205 • Published 8 days ago • 27