Abstract
We find that language models can transfer capabilities through task-unrelated text. Post-training typically improves language models using task-specific data. Prior work on subliminal learning shows that information about these updates can pass through unrelated generations, but has largely focused on traits or preferences using extensive teacher outputs. We introduce Active Taskless Distillation (ATD), which achieves capability transfer using only a single word from the teacher per prompt. ATD probes the behavioral shadow of post-training by selecting prompts where the teacher and student's shared public ancestor is nearly indifferent between two ordinary words. A student initialized from this ancestor learns solely from the resulting prompt-word pairs, without target-task examples, teacher logits, or teacher parameters. In the primary coding experiment with Qwen2.5-1.5B, 5,664nses yield a 5.34 pp gain on HumanEval+ over an exact nuisance-matched control thadisrupts prompt-resperiments showtransfer in scientific knowledge, commonsense reasoning, and reading comprehensins across additional model generations, sizes, and families. Functional analyses show that the learned sid composable, andthat its strength tracks the teacher's update strength.
Community
Can a model teach another model to code — without ever showing it code?
Surprisingly, yes. A coding-trained teacher answers thousands of unrelated prompts with just one word each. A student trained only on those answers gets better at coding — despite never seeing code, task examples, or teacher logits. With just 5,664 single-word answers, the student gains +5.34pp on HumanEval+ over a nuisance-matched control.
Why? Post-training leaves a distributed "behavioral shadow": tiny preference shifts between ordinary words, even in irrelevant contexts. Our method, Active Taskless Distillation (ATD), targets decisions where the base model is nearly undecided — where small shifts reveal the direction of the update. Prior work showed that preferences can be transmitted this way; to our knowledge, this is the first time a capability is.
The effect holds across science, commonsense, reading comprehension, model sizes, and families. The punchline: learned capabilities leak far beyond their training task — thousands of seemingly meaningless behavioral changes collectively carry a rich trace of post-training.
Get this paper in your agent:
hf papers read 2609.29233 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper