arxiv:2609.32303
๐ In a Training Loop
Huang Jingyuan
JingyuanHuang
ยท
AI & ML interests
PhD student working on multimodal agents
Recent Activity
authored a paper about 1 hour ago
Train4Merge: A Controlled Single-Teacher Study of RL vs. SFT Teachers for OPD-Based Model Merging upvoted a paper 2 months ago
Recursive Synthesis for Long-Horizon Terminal Tasks upvoted a paper 2 months ago
MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-TrainingOrganizations
None yet