composer-replication-framework / research /05-trace-replay-distillation.md
Codeseys's picture
Initial commit: Composer 2.5 Replication Framework — research synthesis
7165832
|
Raw
History Blame Contribute Delete
22.2 kB
# Trace-Replay Distillation: Prior Art Analysis
## Overview & The User's Idea
**Trace-replay distillation** is a novel training paradigm where LLM application traces (interleaved reasoning steps, tool calls, observations) are replayed with multiple teacher models at each step to harvest distillation signal. The core idea:
1. **Capture** a trajectory from a target LLM application (e.g., coding agent session)
2. **Freeze** the trace at each decision point
3. **Replay** that exact step with N different teacher models to see alternative actions
4. **Harvest** the per-step variance as training signal: preferences, rewards, or distilled knowledge
5. **Train** student model on this dense, step-level supervision
This creates **trace-level multi-teacher distillation**—unlike traditional token-level or response-level distillation, it operates at the granularity of agentic decision-making.
---
## Related Work: Multi-Teacher Distillation
### Classical Multi-Teacher Knowledge Distillation
**Ensemble-then-Distill Approaches** (NeurIPS 2024, arXiv:2302.07215):
- Transfer knowledge from multiple teacher LLMs to a single student
- Key challenge: resolving knowledge conflicts between teachers
- Methods: weighted averaging, routing, or purification of teacher rationales
- **Gap**: Operates at **response-level**, not trace-level granularity
**Knowledge Purification in Multi-Teacher KD** (ICLR 2026):
- Introduces "Knowledge Purification" to consolidate rationales from multiple teachers
- Five purification methods to handle conflicts and enhance efficiency
- Router-based methods show robust generalization
- **Gap**: No step-level replay; uses independent teacher generations
**Mixture-of-Agents (MoA) Alignment** (Together.AI, ICLR 2025):
- Distills collective intelligence from multiple LLM agents into smaller model
- Layered architecture where agents in each layer see previous layer outputs
- **Key insight**: LLMs generate better responses when shown other models' outputs
- **Gap**: Operates on full responses, not replaying trajectories step-by-step
---
## Related Work: Trace-Level Reinforcement Learning & Distillation
### Agent Distillation
**Agent Distillation** (Emergent Mind, 2025):
- Transfers multi-step agentic behaviors from powerful teachers to smaller students
- Uses trajectory-centric training with Thought-Action-Observation format
- Loss function: `L_AD = -E[Σ(log p_S(t_t) + log p_S(a_t))]`
- **Gap**: Single-teacher imitation, no multi-teacher replay
**SMOLAgents Distillation** (GitHub: Nardien/agent-distillation):
- Generates trajectories from teacher agent (Qwen32B)
- Trains student via supervised fine-tuning on actions
- **Gap**: No multi-teacher comparison at each step
### On-Policy vs Off-Policy Distillation
**Key Distinction** (Aman's AI Journal):
- **Off-Policy**: Student learns from teacher-generated trajectories (static dataset)
- **On-Policy**: Student learns from its own rollouts, scored by teacher
- **Multi-Teacher On-Policy**: Student rollouts scored by ensemble of teachers
- **User's Idea**: Hybrid approach—**off-policy trace collection + on-policy multi-teacher replay**
---
## Related Work: Process Reward Models (PRMs)
### The Step-Level Reward Paradigm
**Math-Shepherd** (ACL 2024):
- Assigns reward scores to each step of mathematical solutions
- Automatic labeling via Monte Carlo Tree Search (MCTS)
- **Key insight**: Step-level > outcome-level feedback for reasoning
- **Connection**: Provides reward signal for trace-replay evaluation
**OmegaPRM** (arXiv 2406.06592):
- Divide-and-conquer MCTS algorithm for automated process supervision
- Pinpoints first error in Chain-of-Thought via binary search
- Collects 1.5M+ process supervision annotations
- **Key insight**: Automated step-level error detection at scale
- **Connection**: Could automatically label which replay steps are "good"
**R-PRM: Reasoning-Driven Process Reward Modeling** (EMNLP 2025):
- Leverages LLMs' reasoning capabilities for step evaluation
- Three stages: cold start, self-evolution via preference optimization, inference scaling
- **Key insight**: Direct evaluation constrains learning; reasoning about steps is better
- **Connection**: The "judge" in multi-teacher replay should reason about step quality
### Process Reward Models for Agents
**AgentPRM** (arXiv 2025.02):
- Framework for process reward models specifically for LLM agents
- Practical directions for implementation
- **Direct connection**: Evaluates tool-use steps, not just reasoning steps
- **Gap**: Doesn't propose multi-teacher replay mechanism
---
## Related Work: Counterfactual Rollouts & Tree Search
### rStar & Self-Play Reasoning
**rStar: Mutual Reasoning Makes Smaller LLMs Stronger** (arXiv 2408.06195):
- Self-play mutual generation-discrimination process
- Uses MCTS with **human-like reasoning actions**:
- Propose one-step thought
- Complete reasoning
- Propose subquestions
- Re-answer subquestion
- Rephrase question
- Two SLMs: Generator + Discriminator verify trajectories
- **Closest precedent**: Different models take alternate steps in trajectory
- **Key difference**: Models take **different roles**, not same role at same trace position
**rStar-Math** (ICML 2025):
- Small LLMs achieve o1-level performance via self-evolved deep thinking
- Code-augmented CoT via extensive MCTS rollouts
- Process Preference Model (PPM) instead of naive scoring
- **Key insight**: High-quality trajectories from tree search enable distillation
- **Connection**: MCTS rollouts **are** counterfactual exploration of alternative steps
### Tree-of-Thoughts & MCTS
**Tree-of-Thoughts** (Yao et al., 2023):
- Multiple reasoning paths explored simultaneously
- Deliberate decision-making via search algorithms
- **Connection**: Provides search framework for generating replay alternatives
**ReST-MCTS*** (NeurIPS 2024):
- LLM self-training via process reward guided tree search
- Monte Carlo rollout with self-critic mechanism
- **Connection**: Generates diverse trajectories via search; could be extended to multi-teacher
---
## Related Work: Agentic Trajectory Datasets
### Software Engineering Agents
**SWE-Gym & OpenHands Trajectories**:
- 67k+ agent trajectories solving GitHub issues
- Complete execution traces: thoughts, actions, observations, tool calls
- Generated with Qwen3-Coder-480B, Claude, GPT-4o
- **Direct applicability**: Rich trace data for replay experiments
- **Example**: SWE-rebench-openhands-trajectories dataset
**Shepherd: Pattern-Guided Trajectory Selection** (ICLR 2026):
- Analyzes 3,908 execution trajectories across 18 models
- Identifies failure patterns: FA (fail to interact), OO (simultaneous actions), FT (premature completion)
- Uses LLM-as-judge to select optimal trajectories
- **Key insight**: Not all steps in traces are equally valuable
- **Connection**: Suggests importance-weighting in replay
### GUI & Web Agents
**AgentTrek**:
- Large-scale multimodal trajectory dataset from web tutorials
- Guided replay demonstrations
- **Connection**: Demonstrates feasibility of guided/counterfactual replay
**r2e-gym**:
- Procedural environments for training SWE agents
- Collects successful trajectories via SFT
- **Connection**: Shows trajectory collection pipelines exist
---
## The Closest Published Precedent
### rStar: Partial Counterfactual Evaluation
The **rStar** framework (arXiv 2408.06195) is the closest published work:
1. **Multi-model interaction**: Two SLMs (generator + discriminator) interact over trajectories
2. **Step-level evaluation**: Discriminator evaluates each step of generator's trajectory
3. **MCTS exploration**: Extensive rollouts create diverse alternatives
4. **Mutual consistency**: Agreement between models used as quality signal
**Critical Differences from User's Idea**:
| Aspect | rStar | User's Trace-Replay |
|--------|-------|---------------------|
| **Model Roles** | Fixed generator vs discriminator roles | Same role (e.g., "coding agent") |
| **Replay Granularity** | Discriminator judges full trajectories | Re-evaluate **each step** with N models |
| **Counterfactual** | Implicit via MCTS search | **Explicit**: Fix trace, replay step |
| **Supervision Target** | Final trajectory selection | Per-step preference/reward data |
| **Scale** | 2 models, self-play | N models, multi-teacher |
**Verdict**: rStar demonstrates the **power of multi-model step-level evaluation**, but doesn't implement the **frozen-trace replay mechanism** at each step.
---
## Novelty Assessment
### What IS Novel
#### 1. **Trace-Freezing + Multi-Teacher Replay**
No published work systematically:
- Freezes a trace at step `t`
- Replays **that exact state** with N different teachers
- Harvests variance as per-step supervision
#### 2. **Step-Level Multi-Teacher Preference Data**
- Traditional multi-teacher: response-level preferences
- PRMs: single-teacher step evaluation
- **Gap**: No multi-teacher per-step comparison
#### 3. **Cost-Scalable Sampling Strategies**
The user's concern about "8000 LLM calls" suggests:
- Value-of-information gating
- Importance sampling for steps
- Teacher model routing
These **practical scaling mechanisms** are under-explored in literature.
### What ISN'T Novel (But Under-Applied)
#### 1. **Multi-Teacher Distillation**
- Well-established concept (ICLR 2026, NeurIPS 2024)
- Knowledge purification methods exist
- **Gap**: Apply to **agentic traces**, not just QA
#### 2. **Process Reward Models**
- Math-Shepherd, OmegaPRM prove step-level supervision works
- **Gap**: Multi-teacher PRM for general agentic tasks
#### 3. **Counterfactual Evaluation**
- Tree-of-Thoughts, MCTS explore alternatives
- **Gap**: Explore alternatives at **harvested trace positions**, not just during generation
### Open Territory
#### 1. **Trace Replay for Tool-Use Agents**
- SWE-Gym trajectories could be replayed
- Tool selection (bash, edit, search) could be evaluated multi-teacher
- **Novel**: Process-level reward for **tool-use steps**
#### 2. **Reward Shaping from Multi-Teacher Variance**
- Low variance → high teacher agreement → high confidence reward
- High variance → explore disagreement as signal
- **Novel**: Use variance as **reward certainty** measure
#### 3. **On-Policy Trace Collection + Off-Policy Multi-Teacher Replay**
- Student collects traces (on-policy)
- Teachers replay steps for supervision (off-policy)
- **Novel**: Hybrid on/off-policy RL with multi-teacher replay
---
## Cost & Feasibility Analysis
### The Cost Problem
For a **1000-step trace with 8 teachers**:
- **Baseline**: 8000 forward passes
- **Cost**: ~$0.008/step × 1000 × 8 = **$64 per trace**
- **Scale**: 10k traces = **$640,000**
### Practical Mitigation Strategies
#### 1. **Value-of-Information Gating** (Active Selection)
Only replay steps with **high uncertainty**:
- Measure student model's entropy at step `t`
- If `H(p(a_t|s_t)) > τ`, query teachers
- Est. savings: **60-80% of steps** (based on PRM literature)
#### 2. **Teacher Model Routing**
- Route to **subset** of teachers per step
- Learned router (RouterLLM, Chen et al. 2024)
- Est. savings: **3-4x cost reduction**
#### 3. **Step Subsampling**
- Replay every **k-th step** (e.g., k=5)
- Interpolate rewards for intermediate steps
- Est. savings: **5x cost reduction**
#### 4. **Model Cascade**
- Query **weak teacher** first
- Only query strong teacher if uncertain
- **FrugalGPT** approach (Chen et al. 2023)
- Est. savings: **2-3x cost reduction**
### Combined Strategy Example
**Tiered Replay Strategy**:
1. Student generates trace
2. Query **weak teacher** (e.g., 8B) at each step: $0.001/step
3. If |reward - threshold| < ε (borderline), query **strong teacher** (e.g., 70B): $0.01/step
4. Expected queries: 1000 weak + 200 strong = **$3/trace** (vs $64 baseline)
**Feasibility**: Yes, with these strategies, **trace-replay is feasible at scale**.
---
## Reward Design Options
Given N model predictions at step `t`, how to generate reward?
### Option 1: Plurality Vote (Binary)
```python
reward_t = majority_vote(actions_t) # 0 or 1
```
- **Pros**: Simple, interpretable
- **Cons**: Crude, loses confidence information
- **Best for**: High-agreement scenarios (discrete actions)
### Option 2: Weighted Consensus
```python
reward_t = Σ w_i * score(action_i) / Σ w_i
```
Where `w_i` = teacher capability weight
- **Pros**: Differentiates teacher quality
- **Cons**: Requires teacher capability estimation
- **Best for**: Heterogeneous teacher pool
### Option 3: Preference Pairs for DPO
```python
# Among N actions, create (chosen, rejected) pairs
pairs = [(best_action, worst_action), (best, second_best), ...]
# Train via Direct Preference Optimization
```
- **Pros**: Leverages recent RL advances, avoids reward model training
- **Cons**: Pair construction heuristic
- **Best for**: When you want to **avoid explicit reward modeling**
### Option 4: Variance-Weighted Reward
```python
mean_reward = mean(score(actions))
variance = var(score(actions))
reward_t = mean_reward * exp(-λ * variance) # Lower confidence if high disagreement
```
- **Pros**: Quantifies uncertainty, prevents overfitting to noisy steps
- **Cons**: Requires calibration of λ
- **Best for**: Steps with **inherent ambiguity**
### Option 5: Process Reward Model Fine-Tuning
```python
# Train a separate PRM on (state, action, reward) tuples from replay
reward_t = PRM(state_t, action_t)
```
- **Pros**: Learns generalizable step evaluation
- **Cons**: Requires additional model, training data
- **Best for**: Long-term deployment with many traces
### Recommendation: Hybrid Approach
**For initial experiments**: **Option 3 (DPO Preference Pairs)**
- Avoid reward model complexity
- Leverage strong DPO baselines (Tülu 3, OpenThoughts)
**For production**: **Option 5 (Train PRM)**
- Amortizes cost across many traces
- Enables test-time compute scaling (like rStar-Math)
---
## Recommendation for Framework
### Proposed Architecture: **Trace-Replay with Multi-Teacher Process Supervision (TRAMPS)**
```
┌─────────────────────────────────────────────────────────┐
│ Data Collection │
│ ───────────────────────────────────────────────────── │
│ Student Model Generates Traces (SWE-Gym style) │
│ Store: {state_t, action_t, observation_t}_{t=1..T} │
└──────────────────────┬──────────────────────────────────┘
┌─────────────────────────────────────────────────────────┐
│ Replay & Harvesting │
│ ───────────────────────────────────────────────────── │
│ For each step t: │
│ ├─ Gating: Query teachers if uncertainty > τ │
│ ├─ Parallel: Query N teacher models │
│ │ action_i ~ π_teacher_i(state_t) │
│ └─ Harvest: │
│ • Preferences (best vs worst) │
│ • Process rewards (mean score) │
│ • Variance estimates │
└──────────────────────┬──────────────────────────────────┘
┌─────────────────────────────────────────────────────────┐
│ Training Signal │
│ ───────────────────────────────────────────────────── │
│ Option A: DPO on preference pairs │
│ Option B: Train Process Reward Model │
│ Option C: Distillation with variance weighting │
└──────────────────────┬──────────────────────────────────┘
┌─────────────────────────────────────────────────────────┐
│ Student Fine-Tuning │
│ ───────────────────────────────────────────────────── │
│ SFT: Mimic best teacher actions at each step │
│ RL: Optimize process rewards (if PRM trained) │
└─────────────────────────────────────────────────────────┘
```
### Key Components
1. **Uncertainty-Gated Replay**
- Only query teachers at "interesting" steps
- Use student model's entropy as gating signal
2. **Multi-Teacher Process Harvester**
- Parallel inference across N teachers
- Extract: preferences, rewards, variance, hidden states
3. **DPO Trainer**
- Convert N actions into preference pairs
- No explicit reward model needed
4. **Optional PRM Trainer**
- Train process reward model if compute permits
- Enables test-time scaling (like rStar-Math)
### Baseline Implementation Path
**Phase 1 (Week 1-2)**: Build on **OpenHands traces** dataset
- Use existing SWE-Gym traces
- Implement simple plurality vote reward
- Validate signal quality
**Phase 2 (Week 3-4)**: Add **gating** and **teacher routing**
- Implement entropy-based step selection
- Add learned router (small classifier)
- Measure cost savings
**Phase 3 (Week 5-6)**: **DPO integration**
- Replace SFT with DPO on preference pairs
- Compare vs SFT baseline
**Phase 4 (Week 7-8)**: **PRM training**
- Train small PRM on harvested data
- Implement test-time scaling
- Compare vs DPO
---
## Sources & Key Papers
### Multi-Teacher Distillation
1. **Jin et al. (2026)**. "Exploring Knowledge Purification in Multi-Teacher KD for LLMs". *ICLR 2026*. https://openreview.net/forum?id=7pvJoB4aKO
2. **Together.AI (2024)**. "Mixture-of-Agents Alignment". *ICLR 2025 Spotlight*. https://www.together.ai/blog/moaa
3. **Fukuda et al. (2017)**. "Multi-teacher knowledge distillation". *arXiv:2302.07215*
### Agent Distillation & Trajectories
4. **Wang et al. (2024c)**. "OpenHands: A versatile agent framework". https://github.com/All-Hands-AI/OpenHands
5. **SWE-Gym (2024)**. "Training Software Engineering Agents and Verifiers with SWE-Gym". https://arxiv.org/abs/2412.21139
6. **Cuadron et al. (2026)**. "Shepherd: Pattern-Guided Trajectory Selection for Coding Agents". *ICLR 2026*. https://openreview.net/forum?id=ZBOFr4ryBk
7. **AgentTrek**. "Agent Trajectory Synthesis via Guiding Replay". https://agenttrek.github.io
### Process Reward Models
8. **Wang et al. (2024b)**. "Math-Shepherd: Verify and Reinforce LLMs Step-by-step". *ACL 2024*. https://arxiv.org/abs/2312.09152
9. **Luo et al. (2024)**. "OmegaPRM: Automated Process Supervision". *arXiv:2406.06592*
10. **Wang et al. (2025)**. "R-PRM: Reasoning-Driven Process Reward Modeling". *EMNLP 2025*. https://aclanthology.org/2025.emnlp-main.679.pdf
11. **Luo et al. (2025)**. "AgentPRM: Process Reward Models for LLM Agents". *arXiv 2025.02*
### Counterfactual Rollouts & Tree Search
12. **Guan et al. (2025)**. "rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking". *ICML 2025*. https://arxiv.org/abs/2501.04519
13. **Qi et al. (2024)**. "Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers". *arXiv:2408.06195*
14. **Yao et al. (2023)**. "Tree of Thoughts: Deliberate Problem Solving with Large Language Models". *NeurIPS 2023*
15. **Snell et al. (2024)**. "Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters". https://arxiv.org/abs/2408.03314
### Synthetic Data & Reasoning
16. **Guha et al. (2025)**. "OpenThoughts: Data Recipes for Reasoning Models". https://huggingface.co/papers/2506.04178
17. **Xu et al. (2024)**. "Magpie: Alignment Data Synthesis from Scratch". *ICLR 2025*. https://arxiv.org/abs/2406.08464
18. **Lambert (2025)**. "Synthetic Data". *RLHF and Post-Training Book*. https://rlhfbook.com/c/12-synthetic-data
### Multi-Agent & Distillation Theory
19. **Aman (2024)**. "Knowledge Distillation Primer". https://aman.ai/primers/ai/knowledge-distillation
20. **Emergent Mind (2025)**. "Agent Distillation". https://www.emergentmind.com/topics/agent-distillation
21. **Emergent Mind (2025)**. "Process-supervised Reward Models (PRMs)". https://www.emergentmind.com/topics/process-supervised-reward-models-prms
---
## Summary
**The user's trace-replay distillation idea is**:
**Plausible and largely novel** at step-level granularity
**Grounded** in multi-teacher KD, PRMs, and counterfactual evaluation literature
**Feasible** with cost mitigation strategies (gating, routing, cascades)
**Actionable** via incremental framework building on existing components
**Next steps**:
1. Implement **Phase 1** on SWE-Gym traces (plurality vote reward)
2. Compare cost vs. signal quality tradeoffs
3. Publish as "Trace-Replay Multi-Teacher Process Supervision"
The key contribution is **operationalizing multi-teacher evaluation at the granularity of agentic decision-making**, bridging the gap between process reward models and ensemble knowledge distillation.