MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use Paper • 2608.20202 • Published 2 days ago • 29
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows Paper • 2608.17800 • Published 4 days ago • 8
EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Paper • 2607.05155 • Published Jul 6 • 18
MSQA: A Natively Sourced Multilingual and Multicultural SimpleQA Benchmark Paper • 2607.00724 • Published Jul 1
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Paper • 2606.11042 • Published Jun 9 • 22
MARBLE: Music Audio Representation Benchmark for Universal Evaluation Paper • 2306.10548 • Published Jun 18, 2023
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Paper • 2605.04077 • Published Apr 14 • 7
Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation Paper • 2604.02368 • Published Mar 27 • 12
The Optimal Token Baseline: Variance Reduction for Long-Horizon LLM-RL Paper • 2602.07078 • Published Feb 6 • 1
view post Post 4475 The Qwen3.8 27B demo for object grounding is now available on Hugging Face Spaces.It features three tasks: Object Detection (Bounding Boxes), Point Localization (Keypoints), and Spatial Guidance (Path Mapping).Try it now: prithivMLmods/Qwen3.8-27B-Object-Detection See translation 🔥 10 10 👀 4 4 🤗 3 3 👍 3 3 🚀 2 2 + Reply
Jais 2: A Family of Arabic-Centric Open Large Language Models Paper • 2608.13580 • Published Jul 7 • 1
MobileMem: Learning from a Year of Mobile Experiences Paper • 2608.13606 • Published 11 days ago • 24
MobileMem: Learning from a Year of Mobile Experiences Paper • 2608.13606 • Published 11 days ago • 24