Who is the Agent to Blame? Localizing Faithfulness and Citation Mistakes in Agentic Deep Research Paper • 2608.24306 • Published 26 days ago • 7
You Know What I Mean: A Benchmark for Agentic Conversational Reference Grounding Paper • 2608.29834 • Published 21 days ago • 9
MameLoshnLM: Yiddish Language Model and Evaluation Benchmark Paper • 2608.05850 • Published Aug 6 • 23
view article Article How NVIDIA AI-Q Reached \#1 on DeepResearch Bench I and II nvidia • Mar 12 • 34
Granite 4.0 Language Models Collection Efficient language models for multilingual generation, coding, RAG, and AI assistant workflows. • 22 items • Updated 5 days ago • 223
Reverse-Engineered Reasoning for Open-Ended Generation Paper • 2509.06160 • Published Sep 7, 2025 • 151
NER Retriever: Zero-Shot Named Entity Retrieval with Type-Aware Embeddings Paper • 2509.04011 • Published Sep 4, 2025 • 29
Beyond Transcription: Mechanistic Interpretability in ASR Paper • 2508.15882 • Published Aug 21, 2025 • 92
A Unifying Scheme for Extractive Content Selection Tasks Paper • 2507.16922 • Published Jul 22, 2025 • 4
GenerationPrograms: Fine-grained Attribution with Executable Programs Paper • 2506.14580 • Published Jun 17, 2025 • 1
Effective Red-Teaming of Policy-Adherent Agents Paper • 2506.09600 • Published Jun 11, 2025 • 39
Dont Add, dont Miss: Effective Content Preserving Generation from Pre-Selected Text Spans Paper • 2310.09017 • Published Oct 13, 2023