view article Article Build an AI Evaluation from a Hugging Face Dataset Without Writing Python phranzia • 8 days ago • 3
view article Article Build an AI Evaluation from a Hugging Face Dataset Without Writing Python phranzia • 8 days ago • 3
view article Article LettucePrevent - Real-Time Prevention of Factual Hallucinations in RAG lebe1 • 14 days ago • 9
view article Article GPU Management: Why Idle GPUs Are the New Grounded Aircraft Dharma-AI • 13 days ago • 89
view article Article Is it agentic enough? Benchmarking open models on your own tooling +1 lysandre, SaylorTwift, pcuenq • Jun 18 • 22
view article Article ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration ibm-research • Jun 30 • 26
view article Article Welcome Inkling by Thinking Machines +3 burtenshaw, merve, pcuenq, ariG23498, andito • 29 days ago • 162
view article Article Run and Compare AI Evaluations with a CLI for Developers and Coding Agents phranzia • 16 days ago • 2
view article Article Run and Compare AI Evaluations with a CLI for Developers and Coding Agents phranzia • 16 days ago • 2