QueST: Persistent Queries as Semantic Monitors for Drift Suppression in Long-Horizon Tracking Paper • 2605.09513 • Published May 10 • 3
view article Article IBM releases SOTA Granite Time Series PatchTST-FM-r2 model with commercial-friendly license ibm-research • 19 days ago • 60
DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training Paper • 2609.04094 • Published 26 days ago • 26
view article Article Real-Time Intelligence with IBM Time Series Models on Confluent ibm-research • 26 days ago • 53
CausalDS: Benchmarking Causal Reasoning in Data-Science Agents Paper • 2607.08093 • Published Jul 9 • 6
view article Article ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration ibm-research • Jun 30 • 27
SenTSR-Bench: Thinking with Injected Knowledge for Time-Series Reasoning Paper • 2602.19455 • Published Feb 23 • 1
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Paper • 2606.19704 • Published Jun 18 • 43
Evoflux: Inference-Time Evolution of Executable Tool Workflows for Compact Agents Paper • 2606.12674 • Published Jun 10 • 6
view article Article Beyond LLMs: Why Scalable Enterprise AI Adoption Depends on Agent Logic ibm-research • Jun 1 • 92
view article Article ITBench-AA: Frontier Models Score Below 50% on the First Benchmark for Agentic Enterprise IT Tasks — by Artificial Analysis and IBM ibm-research • May 27 • 20
Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows Paper • 2605.24219 • Published May 26 • 7
Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines Paper • 2605.20630 • Published May 20 • 10
DiagnosticIQ: A Benchmark for LLM-Based Industrial Maintenance Action Recommendation from Symbolic Rules Paper • 2605.08614 • Published May 9 • 7
SPIN: Structural LLM Planning via Iterative Navigation for Industrial Tasks Paper • 2605.14051 • Published May 13 • 1
Results and Retrospective Analysis of the CODS 2025 AssetOpsBench Challenge Paper • 2605.08518 • Published May 8 • 11
MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments Paper • 2605.09131 • Published May 9 • 62
When to Trust Imagination: Adaptive Action Execution for World Action Models Paper • 2605.06222 • Published May 7 • 43