Papers
arxiv:2609.39578

Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively?

Published on Sep 30
· Submitted by
Elvis Wang
on Oct 1
Authors:
,
,
,
,

Abstract

Agent harnesses often improve language models with human-designed workflows, but as models grow more capable, unreliable guidance can increasingly constrain their execution. We call the ability to benefit from useful guidance while overriding unreliable guidance thinking outside the box. We introduce Box^2-Bench, which holds the model and task fixed while varying workflow reliability to isolate how models regulate their reliance on guidance. On Box^2-Bench, frontier models often benefit from reliable guidance but remain vulnerable when it is misleading or becomes unreliable. To test whether this capability can be learned, we train two open-weight models using bad workflows, reserving good workflows for evaluation. We explore two complementary training strategies: counterfactual supervised fine-tuning improves robustness, while outcome-based reinforcement learning can shift the balance toward greater use of helpful workflows. We further find that this behavior extends beyond workflows to other forms of external information, improving peer correction and robustness to corrupted memory. Together, our results identify selective reliance on fallible external information as a dimension of agent reliability not captured by task performance alone.

Community

thinkoutsidethebox

Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively?

Language model agents increasingly rely on human designed workflows to solve complex tasks. Good workflows can substantially improve performance by providing useful intermediate steps and structure.

But what happens when the workflow is worse than the model itself?

As models become more capable, blindly following external guidance can become a limitation rather than an advantage. A strong agent should benefit from useful guidance while being able to override guidance when it becomes unreliable.

We call this ability thinking outside the box.

Measuring Selective Reliance

截屏2026-10-01 11.34.26

To study this behavior, we introduce Box² Bench.

The idea is simple: we keep the model and task fixed, but vary the reliability of the workflow. We compare how the same model performs with no workflow, a helpful workflow, and a misleading workflow.

This lets us separate two abilities that are usually mixed together:

Can the model use good guidance? And can it resist bad guidance?

Across models and tasks, we find a consistent pattern. Reliable workflows often improve performance, but misleading workflows can still substantially hurt it. Even capable models remain sensitive to unreliable external guidance.

Learning to Think Outside the Box

截屏2026-10-01 11.34.54

We next ask whether selective reliance can be learned.

We train two open weight models using only bad workflows, while reserving good workflows for evaluation.

Counterfactual supervised fine tuning makes models substantially more robust to misleading guidance. Outcome based reinforcement learning can further shift their behavior toward making greater use of helpful workflows.

Importantly, the models never see good workflows during training. Their improvement therefore suggests that training can shape a more general strategy for deciding how much to trust external guidance.

Beyond Workflows

This behavior is not limited to workflows.

Models trained in our setting also improve when interacting with other forms of fallible external information. We observe better peer correction in multi agent reasoning and greater robustness to corrupted memory.

These results point to a broader capability: agents need not only to use external information, but also to regulate their reliance on it.

Toward More Reliable Agents

Most agent evaluations ask whether a task was solved.

Box² Bench asks an additional question:

Does the model know when to listen?

As agents increasingly operate with workflows, memories, tools, and other agents, selective reliance on fallible external information may become an important dimension of reliability that task performance alone does not capture.

Sometimes the best use of a box is knowing when to step outside it.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.39578
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.39578 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.39578 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.39578 in a Space README.md to link it from this page.

Collections including this paper 1