Self-Supervised Scaling of Terminal Environments for Scientific Domains
Abstract
Terminal agents are increasingly deployed beyond software engineering in science and other specialized domains. Constructing training environments requires executable reference behavior and a domain-specific verifier that distinguishes semantic correctness from superficially plausible artifacts. Authoring these components for each task requires repeated engineering and limits reuse. We introduce software-in-the-loop reconstruction, a self-supervised framework that obtains reference outputs and verification targets from existing software workflows, executable programs mapping structured inputs to outputs. For each workflow, we execute multiple input configurations and partition cases into public observations and hidden evaluations. Given the instruction, input schema, and public input--output observations, an agent constructs an editable program without access to the source workflow. The candidate is evaluated on hidden configurations against workflow outputs. A hierarchical verifier combines domain-specific semantic comparison, structural validity, and anti-shortcut checks, while public feedback supports iterative revision. The construction admits additional workflows and configurations without authoring a reference solution for each task. We instantiate SWR with 500 workflows and 46 software families across six domains. Across three attempts per task, Qwen3.8-Max solves 838 tasks and produces 1,422 verified trajectories, which we oversample to 3,000 reconstruction-only training examples. Supervised fine-tuning of Qwen3.8-27B improves mean Terminal-Bench 2 performance from 47.94% to 53.56% across three seeds and achieves the highest mean among four matched-token corpus controls on all four reported evaluations. These results indicate that existing scientific software can provide scalable, behaviorally verified supervision for terminal agents.
Community
We introduce software-in-the-loop reconstruction, a self-supervised framework for scaling verifiable terminal-agent training environments beyond software engineering. Our SWR benchmark contains 500 workflows across 46 software families and six scientific domains, and produces behaviorally verified trajectories that improve terminal-agent performance.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis (2026)
- ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments (2026)
- SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving (2026)
- UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations (2026)
- SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving (2026)
- Zero2Repo: Can Coding Agents Build Repositories from Scratch? (2026)
- From Dead Code and Static Requirements to Working Engines: Software Revival with Coding Agents (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.02710 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper