No Pain, More Gain: Iterative Merging for Effective Multi-Teacher On-Policy Distillation
arXiv preprint
Master’s student at KAIST AI · Intern at Kakao
I am a master’s student at KAIST AI, advised by Prof. Juho Lee, and currently an intern at Kakao. My research focuses on model consolidation for large language models.
I explore model merging and knowledge distillation to combine knowledge and complementary capabilities learned by different models. This includes approaches such as multi-teacher on-policy distillation (MOPD), with the broader goal of building unified models that perform well across diverse tasks and objectives.
(*) denotes equal contribution
arXiv preprint
NeurIPS 2026
ICML 2024
Master’s program