Papers
arxiv:2609.35760

TokenCast: Forecasting Token Consumption During LLM Agent Execution

Published on Sep 28
· Submitted by
Leo Y
on Sep 29
Authors:
,
,
,
,
,
,
,
,
,

Abstract

When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The agent chooses its next steps based on tool feedback and intermediate results, while the growing context steadily inflates the input size of every subsequent call. The total consumption of a task is therefore hard to predict before execution and the prediction must be revised as the run unfolds. In this paper, we propose TokenCast, which learns a composable cost representation for each execution segment, recording its own consumption and the context growth it introduces. Composing adjacent segments yields a cumulative estimate that captures the extra input cost incurred when context from earlier segments is re-read by every later call. As execution unfolds, newly observed evidence refreshes the forecast, requiring no additional LLM calls and incurring a mean cumulative prediction time of 32.8 ms per run on SWE-bench Verified. Across 4 task suites and 6 agent models, TokenCast's mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations. In offline budget-control replay, TokenCast uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion. The code is available at https://github.com/DEFENSE-SEU/TokenCast.

Community

Paper submitter

TokenCast forecasts token consumption during LLM agent execution. The same task can consume over an order of magnitude more tokens across runs, because the agent's next steps depend on tool feedback and the growing context inflates the input size of every later call. TokenCast learns a composable cost representation for each execution segment (call count, net input-length change, and a cost residual). An exact composition identity captures the extra input cost of re-reading earlier context in every later call. A staged prefix-suffix predictor combines direct and compositional forecasts and refreshes its estimate from newly observed execution evidence, with calibrated quantile models giving prediction intervals and no additional LLM calls (mean cumulative prediction time of 32.8 ms per run on SWE-bench Verified). Across 4 task suites and 6 agent models, the mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations. In offline budget-control replay, TokenCast uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.35760
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.35760 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.35760 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.35760 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.