Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?
Abstract
Φ-Bench evaluates large language models on open-ended engineering of the LLM infrastructure stack across tasks from kernel optimization to end-to-end system design.
Large language models (LLMs) have demonstrated remarkable capabilities in reasoning and code generation, raising the prospect that they could assist in developing and optimizing the very infrastructure that powers them. However, existing benchmarks mainly focus on isolated kernels, predefined operators, or pre-specified optimization targets, and therefore fail to evaluate the ability of LLMs to perform open-ended, long-horizon LLM infrastructure engineering. To address this gap, we present Φ-Bench, a benchmark for systematically evaluating LLMs on engineering the LLM infrastructure stack. Derived from optimization problems studied in frontier research and grounded in real-world code repositories, Φ-Bench provides broad coverage of the LLM infrastructure stack and spans tasks of varying complexity, ranging from localized kernel-level function completion to long-horizon implementation and end-to-end system optimization. Extensive experiments on frontier LLMs reveal their current capabilities and limitations in engineering complex LLM infrastructure, offering insights into the challenges that remain on the path toward autonomous optimization of future AI infrastructure.
Community
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- PerfAgent: Profiler-Guided Iterative Refinement for Repository-Level Code Optimization (2026)
- RealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI Frameworks (2026)
- DSEffi-Bench: Demystifying Large Language Models'Capability in Efficient Data Science Code Generation (2026)
- Benchmarking LLMs on File System Design and Implementation (2026)
- LLM4LLM: Bridging Kernel Benchmarks and Real Deployment via Closed-Loop Agentic Optimization (2026)
- SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring (2026)
- RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.10226 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper