experience
June 2026 — Aug 2026
Applied Scientist Intern
Full-Time, Boston, MA
Research Domain: Video World Models, Diffusion Models, Action-Controlled Video Generation
Project: Interactive Video World Model @ Amazon AGI Foundations Team
Built a large-scale Unreal Engine video dataset and training pipeline for action-controlled video generation; Scaled training from a 1.3B to a 14B bidirectional video diffusion model on 32 NVIDIA B200 GPUs, transformed it into a causal architecture for autoregressive video generation; Applied few-step diffusion distillation for accelerated inference.
Jan 2026 — May 2026
Research Scientist Intern
Part-Time, Remote
Research Domain: Latent Reasoning, Harness Engineering, Vision-Language-Action Models
Project: Multimodal Reasoning @ NEC Media Analytics Team
Developed a latent visual reasoning system for high-level behavioral reasoning in autonomous-driving videos; Built a unified agentic harness around frozen multimodal LLMs (GPT-4o/GPT-5.4/Claude Haiku 4.5/Claude Opus 4.7) with composable visual processing, structured parsing, and verifier-guided repair, demonstrating that many model failures can be addressed at the harness level without backbone retraining (one paper released to arXiv 2026).
May 2025 — Aug 2025
Applied Scientist Intern
Full-Time, Seattle, WA
Research Domain: LLMs, VLMs, AI Agents, Reinforcement Learning, Low-Level Vision
Project: Restore-R1 @ Amazon One Team
Developed an end-to-end agentic system for complex image restoration by fine-tuning a VLM (mPLUG-Owl2/DeQA-Score) for quality evaluation, employing an LLM (GPT-4/Llama3) for restoration planning, and training a reinforcement learning policy (PPO/GRPO) on 8 NVIDIA L40S GPUs to optimize tool-execution sequences (one paper accepted to CVPR 2026).
Sep 2024 — Nov 2024
Research Scientist Intern
Full-Time, San Jose, CA
Research Domain: LLMs, VLMs, Multimodal Representation & Retrieval & Recommendation
Project I: Quality-Controllable Multimodal Retrieval @ Adobe Research Team
Built Flickr2.4M, a 2.4 million image-caption training corpus; Developed a multimodal retrieval and recommendation framework integrating LLM-based query refinement with VLM-based text–image matching; Trained and evaluated GPT2-1.5B/Qwen2.5-0.5B with CoCa/Blip2/OpenCLIP on 8 NVIDIA A100 GPUs, delivering more relevant and visually appealing recommendations (one paper accepted to ICLR 2026).
Project II: Indra Representation Hypothesis @ Adobe Research Team
Established a theoretical foundation for multimodal alignment, validated it across vision (ViT/ConvNeXt/DINOv2), language (BERT/RoBERTa), and audio models (Wav2vec/WavLM/HuBERT), showing that independently trained foundation models implicitly converge to a shared relational structure across modalities (one paper accepted to NeurIPS 2025).