Four sentiment corpora, every model tested on all four
Redraw the district lines and the answer changes
The winner of your A/B test is not as good as it looked
Generate SQL from natural‑language F1 questions and see results
Explore anomaly detection performance on benchmark data
An A2A agent network that reviews Python dependencies