MultiBBQ Collection Fairness benchmark for multimodal LLMs: dataset, image perturbations, and results (paper: Fairness Failure Modes of Multimodal LLMs). • 4 items • Updated Jul 13
MultiBBQ Collection Fairness benchmark for multimodal LLMs: dataset, image perturbations, and results (paper: Fairness Failure Modes of Multimodal LLMs). • 4 items • Updated Jul 13
Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context Paper • 2605.13831 • Published May 13 • 90
DecodingTrust-Agent Platform (DTap): A Controllable and Interactive Red-Teaming Platform for AI Agents Paper • 2605.04808 • Published May 6 • 20
Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration? Paper • 2602.07055 • Published Feb 4 • 23
AgentDoG Collection A Diagnostic Guardrail Framework for AI Agent Safety and Security • 12 items • Updated Jun 21 • 112
Artificial Entanglement in the Fine-Tuning of Large Language Models Paper • 2601.06788 • Published Jan 11 • 5