Fanyi Pu 濮凡轶


Fanyi Pu 濮凡轶
Incoming Ph.D. Student in Computer Science
University of Wisconsin–Madison

Abstract

I am an incoming Ph.D. student in Computer Science at the University of Wisconsin–Madison, starting in September 2026. I received my B.Sc. in Data Science and Artificial Intelligence from Nanyang Technological University with Honours (Highest Distinction).

My research interests lie in multimodal models and video generation. My work has received more than 1,800 citations on Google Scholar.

I am a core member of LMMs-Lab, where I help develop LMMs-Eval, LMMs-Engine, and Video-MMMU. I have also conducted research at SenseTime Research and MMLab@NTU, and contributed to Synvo AI.

Keywords: Multimodal Models · Video Generation · Spatial Intelligence

1   Education

University of Wisconsin–Madison
Ph.D. in Computer Science (Incoming), Madison, WI
Nanyang Technological University
Bachelor of Science in Data Science and Artificial Intelligence, Singapore
  • Honours (Highest Distinction), CGPA: 4.67 / 5.00
  • Research interests: Multimodal and Video Generation. 1,800+ Google Scholar citations
  • Selected coursework: Reinforcement Learning, Deep Learning, Data Structures and Algorithms, Probability and Statistics, Computational Economics, Stochastic Processes, Numerical Analysis, Cryptography, Electricity & Magnetism
University of California, Berkeley
Summer Session, Berkeley, CA
  • Studied Computer Security and Game Theory, GPA: 4.00 / 4.00

2   Publications & Research

LMMs-Engine: A Simple, Unified Multimodal Models Training Engine
Research Project, Core Developer[GitHub]
  • A lean and flexible training framework designed for both rapid research prototyping and large-scale production
  • Optimized the Bagel training pipeline by integrating FSDP2 and Liger Kernel for efficient distributed training
  • Designing asynchronous reinforcement-learning infrastructure for the project
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
NAACL 2025 (Findings), Co-first author[Paper][Project Page]
  • An open-source multimodal evaluation framework with 4.3K GitHub stars
  • Built the evaluation pipeline and developed a low-cost automatic data-generation pipeline for Multi-modal LiveBench, which leverages continuously updated news and online forums to evaluate models' generalization in the wild
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
Technical Report (Microsoft Research), Contributor[Paper][Project Page]
  • An efficient native-resolution foundation model for high-quality image generation and editing
  • Extended LMMs-Engine for Mage-Flow pre-training by integrating FSDP2, Liger Kernel, and sequence packing, enabling efficient native multi-resolution training across heterogeneous image resolutions and aspect ratios
Demystifying Video Reasoning
ECCV 2026, Third Author[Paper][Project Page]
  • Revisited video reasoning and proposed Chain-of-Step, in which models reason through diffusion steps
  • Curated a large collection of examples to validate this hypothesis and systematically analyzed models' reasoning behaviors
SenseNova-SI: Scaling Spatial Intelligence with Multimodal Foundation Models
CVPR 2026, Co-first author[Paper][GitHub]
  • Developed a foundation model for scaling spatial intelligence, achieving state-of-the-art performance on key spatial benchmarks
  • Conducted rigorous ablation studies to establish that spatial intelligence is primarily driven by data scaling rather than reasoning heuristics (CoT), guiding the project's strategic focus on constructing the SenseNova-SI-8M dataset
  • Executed the full-stack training pipeline for Qwen-based variants with LMMs-Engine on 128 GPUs
Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos
ACL 2026, Third Author[Paper][Project Page]
  • Built an evaluation set with 300 expert-level videos and 900 human-annotated questions across six disciplines, cited by Gemini 3 Pro (Google) and GPT-5 (OpenAI)
  • Proposed a knowledge-gain metric to quantify performance improvements after watching video lectures; validated expert-level data in economics and medicine to ensure benchmark quality
Multi-Modal In-Context Instruction Tuning (Otter, MIMIC-IT)
IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), Co-first author[Paper][Project Page]
  • Developed an early vision-language-agent (VLA) model with 3.3K GitHub stars
  • Developed a language-model-based pipeline that generated 2.8M multimodal instruction-tuning samples; created an evaluation framework that later evolved into LMMs-Eval

3   Research & Professional Experience

LMMs-Lab
Core Member, Singapore
  • Core member of a non-profit initiative dedicated to democratizing Large Multimodal Models (LMMs); instrumental in the lab's 0-to-1 establishment, from initial ideation and naming to its current operations
  • Spearheaded the development of high-impact open-source projects: the evaluation framework LMMs-Eval, the efficient training library LMMs-Engine, and the video understanding benchmark Video-MMMU
SenseTime Research
Research Intern (Spatial Intelligence and Video Generation), Singapore
  • Conducted research on SenseNova-SI and video reasoning, with a focus on spatial intelligence and video generation
Synvo AI
Core Contributor, Singapore
  • Architected and implemented the Synvo File System, the company's inaugural product, serving as the storage backbone for unstructured multimodal data
  • Empowered the core contextualization engine by structuring data for efficient retrieval, enabling the development of AI systems that can remember, learn, and adapt
MMLab@NTU
Research Intern (Multimodal Models), Singapore
  • Supervised by Prof. Liu Ziwei; focused on multimodal language models and unified multimodal models

4   Competitions

The International Collegiate Programming Contest (ICPC)
Nanyang Technological University / Dalian University of Technology, Singapore / China
  • Ranked 22nd in the 2024 ICPC Asia Pacific Championship
  • Ranked 13th in the 2023 ICPC Asia Jakarta Regional Contest
  • Ranked 6th in the 2022 ICPC Asia Manila Regional Contest
  • Gold Medal in the 2021 ICPC Asia Kunming Regional Contest
  • Silver Medal in the 2021 ICPC Asia Nanjing Regional Contest
Simon Marais Mathematics Competition
Nanyang Technological University, Singapore
  • Received the Best-in-University Prize at Nanyang Technological University
  • Solved problems in number theory, game theory, and calculus

5   Teaching & Activities

Tutorial Lecturer
NTU SC1008 C & C++ Programming, Singapore
  • Instructed a cohort of 50 beginners in core C/C++ syntax and methodologies
Nanyang Programming Contest
Nanyang Technological University, Singapore
  • Organized five competitions and tutorials for non-ICPC students
  • Designed activities to help students strengthen algorithmic problem-solving skills for technical interviews
NTU Students' Computing and Data Science Club
Nanyang Technological University, Singapore
  • Managed a Telegram group with 130+ members
  • Answered AI/ML questions and shared resources on deep learning research
Teaching Assistant
NTU SC1003 Introduction of Computational Thinking, Singapore
  • Instructed a cohort of 20 beginners in core Python syntax and data manipulation methodologies

6   Interests

Physics, Chess (Lichess rating: 2084), Calligraphy, Table Tennis