Gradient-Balanced Timestep-Partitioned LoRA for Reward Fine-Tuning of Diffusion Models
Seunghyo Yun, Seungjun Oh, Yusung Kim†
Under Review · ICLR 2027
Diffusion model의 reward fine-tuning에서 발생하는 timestep-wise optimization imbalance를 분석하고, 이를 완화하기 위한 gradient-balanced LoRA 구조를 제안합니다.
PaperGitHub (coming soon)