Ouroboros benchmark evidence Collection Public runs, task records, traces, and matched-pair evidence behind Ouroboros benchmark reports. • 5 items • Updated 7 days ago
Running on Zero Agents 15 Universal Activation Oracle 🔮 15 Read LLM activations to detect bias, cross-model zero-shot
🔍 Interpretability & Analysis of LMs Collection Outstanding research in LM interpretability and evaluation, summarized • 136 items • Updated May 26 • 119
Through the Looking Glass: Common Sense Consistency Evaluation of Weird Images Paper • 2505.07704 • Published May 12, 2025 • 29
SONAR-LLM: Autoregressive Transformer that Thinks in Sentence Embeddings and Speaks in Tokens Paper • 2508.05305 • Published Aug 7, 2025 • 49
SONAR-LLM: Autoregressive Transformer that Thinks in Sentence Embeddings and Speaks in Tokens Paper • 2508.05305 • Published Aug 7, 2025 • 49 • 3