Papers
arxiv:2609.35427

LLMs are General Asynchronous Agents

Published on Sep 28
· Submitted by
George Yakushev
on Sep 30
Authors:
,
,
,
,
,

Abstract

Modern LLMs are increasingly capable as autonomous agents, but they follow sequential interaction cycles: read, think, reply or call tools, repeat. Many real-world use cases are not sequential: voice assistants, embodied agents, and monitoring systems receive new inputs while they think or perform another task. Modern LLMs address this with specialized architectures for voice interaction and video streams, VLAs for robot control, asynchronous tool calling for API usage, and others. In this work, we generalize from different asynchronous tasks to general asynchronous agents that can adapt to different types of concurrency. To achieve this, we develop an asynchronous LLM framework that lets users (or the agents themselves) define inference coroutines with overlapping memory states. We showcase that Qwen 3.x models are capable of asynchronous operation for streaming video understanding, videogames, and monitoring, without task-specific training.

Community

Paper submitter

We’re releasing AsyncLLM, an open-source framework that lets pretrained LLMs observe, reason and act concurrently without additional training.

The key idea: concurrency lives inside inference, and not around API calls. Inspired by asyncio an agent is a set of Python async/await coroutines. Each writes to cache block and chooses which blocks to attend to through cache views. Streams read each other’s evolving state while they may be incomplete.
This cache-block abstraction sits on an inference engine built on Mini-SGLang. It efficiently batches requests across coroutines and supports full attention, Gated DeltaNet and multimodal MRoPE models. Coordination uses standard asyncio events, locks and queues.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.35427
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.35427 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.35427 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.35427 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.