I am a Ph.D. student in Computer Science and Engineering at the University at Buffalo, advised by Prof. Junsong Yuan. My research focuses on generative AI for visual content — diffusion models, text-to-image/video generation and editing, vision–language grounding, and reinforcement learning for generation (RLHF / GRPO).
I am currently a research intern at ByteDance (Intelligent Creation, Bellevue WA). Previously I interned at Adobe (2024 & 2025), Pixocial, and Microsoft Research Asia (2021–2023). Before UB, I received my M.S. from Xi'an Jiaotong University, where I worked on temporal action localization with Prof. Le Wang.
Selected works; full list on Google Scholar. * denotes equal contribution.
Reviewer for T-PAMI, TCSVT, Machine Vision and Applications, and major vision conferences.