Yunong Liu
Research Scientist @ NVIDIA GEAR
I work on foundation models for robotics.
My research interests span multimodal foundation models, world models, and robotics. I am interested in structured, editable, and verifiable representations that connect visual generation with reasoning, simulation, and action in physical and interactive environments.
Previously at Luma AI, I led research and system development for Layering, a structured visual generation system that transforms generated images and design media into editable raster, text, and vector layers, enabling multi-turn editing by people and agents.
My work at Luma also spanned Ray3 and Uni-1 models. For Ray3, I built the experimental post-training workflow connecting video generation, reward modeling, VLM-as-judge graders, and held-out evaluation, and explored diffusion RL, DPO-style, and GRPO-style approaches. For Uni-1, I contributed to reinforcement learning, data, and evaluation experiments, including OCR-focused rewards, caption and data ablations, and early evaluation.
Before Luma, I completed my MS in Computer Science at Stanford with Jiajun Wu, where my research focused on visual understanding, spatiotemporal grounding, and multimodal evaluation. I hold a BEng in Electronics and Computer Science from the University of Edinburgh.