Draft-KV: Learning Useful Latent Communication Between Language Models
Abstract
Latent communication passes internal states between language models instead of decoded text, but higher receiver accuracy does not show that the receiver used the message content. Across five method-dataset pairs, replacing each message with one from an unrelated question changes accuracy by at most 0.60 points, even when communication adds 15.44 points over the receiver alone. Thus the interface can supply the gain while making the sharer dispensable. Draft-KV instead sends the key-value states formed while the sharer drafts an answer to the current question. Linear projections place these states in a side memory read through a gated attention branch, and progressive training moves from message reconstruction to answer supervision under a guard on harm from mismatched messages. Both models remain frozen and the interface trains 1.05M parameters, 348x fewer than C2C. With a Qwen3-8B sharer, a frozen Qwen2.5-0.5B-Instruct receiver reaches 78.04% on MMLU-Redux, versus 37.45% alone and 36.40% with reassigned messages. At fixed interface size, scaling the sharer from 0.6B to 8B raises accuracy from 46.11% to 78.04%; communication also transfers to held-out tasks and can exceed both models when each holds different evidence.
Community
Can language models communicate useful private information through latent representations? We introduce Draft-KV, a simple framework that lets a receiver directly reuse the KV cache produced by a sharer while drafting. Unlike prior latent communication methods whose gains may come from learned adapters rather than message content, Draft-KV enables information-dependent collaboration and benefits from stronger sharers.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- StateBridge: Training-free Hidden-state Alignment for Latent Communication in LLM Multi-Agent Systems (2026)
- LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay (2026)
- Train Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning (2026)
- PREM: Prefix-Steered Recurrent Memory for Long-Video Understanding (2026)
- DASC: Decay-Aware State Compression for Hybrid Linear-Attention Serving (2026)
- Reason Through the Latent! Making Latent Visual Reasoning Necessary (2026)
- Shared Global KV with Layer-Specific Local History (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.34754 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper