experience

June 2026 —
Applied Scientist Intern
Full-Time, Boston, USA
Research Areas: Video World Models, Diffusion Models, Action-Controlled Video Generation
Project: ongoing project on interactive video world models. Build a large-scale datasets for action-contorlled video generation, and transform a bidirectional video diffusion model to a causal one based on self-forcing for autoregressive video generation.
May 2025 — Aug 2025
Applied Scientist Intern
Full-Time, Seattle, USA
Research Areas: LLMs, VLMs, AI Agents, Reinforcement Learning, Low-Level Vision
Project: proposed Restore-R1, an agentic solution for complex image restoration by fine-tuning a VLM to identify degradations, employing an LLM for restoration planning, and training a reinforcement learning agent to optimize tool execution sequences (one paper accepted to CVPR 2026).
Sep 2024 — Jan 2025
Research Scientist Intern
Full-Time, San Jose, USA
Research Areas: LLMs, VLMs, Multimodal Alignment, Retrieval & Recommendation
Project I: proposed a visual-quality-controllable multimodal retrieval framework that trains an LLM for query refinement and integrates a VLM for text–image semantic matching, providing customers more aesthetic and higher-quality recommendations (one paper accepted to ICLR 2026).
Project II: proposed the Indra Representation Hypothesis for multimodal alignment, showing that independently trained unimodal foundation models implicitly converge to a shared relational structure of reality (one paper accepted to NeurIPS 2025).