Yunong Liu
Research @ Luma AI
My research at Luma AI spans multimodal image and video generation, world models, and structured visual generation, with a focus on post-training, reward modeling, and evaluation.
My current interest is moving visual generative models beyond flat pixels: toward structured, editable, and verifiable representations that can support reasoning, simulation, and eventually action in physical or interactive environments.
I led research and system development for Layering, a structured visual generation system that transforms generated images and design media into editable raster, text, and vector layers, enabling multi-turn editing by people and agents.
My work also spans Luma's Ray3 and Uni-1 models. For Ray3, I built the experimental post-training workflow connecting video generation, reward modeling, VLM-as-judge graders, and held-out evaluation, and explored diffusion RL, DPO-style, and GRPO-style approaches. For Uni-1, I contributed to reinforcement learning, data, and evaluation experiments, including OCR-focused rewards, caption and data ablations, and early evaluation.
Before Luma, I completed my MS in Computer Science at Stanford with Jiajun Wu, where my research focused on visual understanding, spatiotemporal grounding, and multimodal evaluation. I hold a BEng in Electronics and Computer Science from the University of Edinburgh.