Fanyi Pu 濮凡轶


Fanyi Pu 濮凡轶
Incoming Ph.D. Student in Computer Science
University of Wisconsin–Madison

Abstract

I am an incoming Ph.D. student in Computer Science at the University of Wisconsin–Madison, starting in September 2026. I received my B.Sc. in Data Science and Artificial Intelligence from Nanyang Technological University with Honours (Highest Distinction). My research interests lie in multimodal models and video generation.

1   Education

University of Wisconsin–Madison
Ph.D. in Computer Science (Incoming)Madison, WI
Nanyang Technological University
Bachelor of Science in Data Science and Artificial IntelligenceSingapore
  • Honours (Highest Distinction), CGPA: 4.67 / 5.00
  • Research interests: Multimodal and Video Generation. 1,800+ Google Scholar citations
  • Selected coursework: Reinforcement Learning, Deep Learning, Data Structures and Algorithms, Probability and Statistics, Computational Economics, Stochastic Processes, Numerical Analysis, Cryptography, Electricity & Magnetism
University of California, Berkeley
Summer SessionBerkeley, CA
  • Studied Computer Security and Game Theory, GPA: 4.00 / 4.00

2   Publications & Research

LMMs-Engine: A Simple, Unified Multimodal Models Training Engine
Research Project, Core Developer[GitHub]
  • A lean and flexible training framework designed for both rapid research prototyping and large-scale production
  • Optimized the Bagel training pipeline by integrating FSDP2 and Liger Kernel for efficient distributed training
  • Designing asynchronous reinforcement-learning infrastructure for the project
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
NAACL 2025 (Findings), Co-first author[Paper][Project Page]
  • An open-source multimodal evaluation framework with 4.3K GitHub stars
  • Built the evaluation pipeline and developed a low-cost automatic data-generation pipeline for Multi-modal LiveBench, which leverages continuously updated news and online forums to evaluate models' generalization in the wild
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
Technical Report (Microsoft Research), Contributor[Paper][Project Page]
  • An efficient native-resolution foundation model for high-quality image generation and editing
  • Extended LMMs-Engine for Mage-Flow pre-training by integrating FSDP2, Liger Kernel, and sequence packing, enabling efficient native multi-resolution training across heterogeneous image resolutions and aspect ratios
Demystifying Video Reasoning
ECCV 2026, Third Author[Paper][Project Page]
  • Revisited video reasoning and proposed Chain-of-Step, in which models reason through diffusion steps
  • Curated a large collection of examples to validate this hypothesis and systematically analyzed models' reasoning behaviors
SenseNova-SI: Scaling Spatial Intelligence with Multimodal Foundation Models
CVPR 2026, Co-first author[Paper][GitHub]
  • Developed a foundation model for scaling spatial intelligence, achieving state-of-the-art performance on key spatial benchmarks
  • Conducted rigorous ablation studies to establish that spatial intelligence is primarily driven by data scaling rather than reasoning heuristics (CoT), guiding the project's strategic focus on constructing the SenseNova-SI-8M dataset
  • Executed the full-stack training pipeline for Qwen-based variants with LMMs-Engine on 128 GPUs
Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos
ACL 2026, Third Author[Paper][Project Page]
  • Built an evaluation set with 300 expert-level videos and 900 human-annotated questions across six disciplines, cited by Gemini 3 Pro (Google) and GPT-5 (OpenAI)
  • Proposed a knowledge-gain metric to quantify performance improvements after watching video lectures; validated expert-level data in economics and medicine to ensure benchmark quality
Multi-Modal In-Context Instruction Tuning (Otter, MIMIC-IT)
IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), Co-first author[Paper][Project Page]
  • Developed an early vision-language-agent (VLA) model with 3.3K GitHub stars
  • Developed a language-model-based pipeline that generated 2.8M multimodal instruction-tuning samples; created an evaluation framework that later evolved into LMMs-Eval

3   Research & Professional Experience

LMMs-Lab
Core MemberSingapore
  • Core member of a non-profit initiative dedicated to democratizing Large Multimodal Models (LMMs); instrumental in the lab's 0-to-1 establishment, from initial ideation and naming to its current operations
  • Spearheaded the development of high-impact open-source projects: the evaluation framework LMMs-Eval, the efficient training library LMMs-Engine, and the video understanding benchmark Video-MMMU
SenseTime Research
Research Intern (Spatial Intelligence and Video Generation)Singapore
  • Conducted research on SenseNova-SI and video reasoning, with a focus on spatial intelligence and video generation
Synvo AI
Core ContributorSingapore
  • Architected and implemented the Synvo File System, the company's inaugural product, serving as the storage backbone for unstructured multimodal data
  • Empowered the core contextualization engine by structuring data for efficient retrieval, enabling the development of AI systems that can remember, learn, and adapt
MMLab@NTU
Research Intern (Multimodal Models)Singapore
  • Supervised by Prof. Liu Ziwei; focused on multimodal language models and unified multimodal models

4   Competitions

The International Collegiate Programming Contest (ICPC)
Nanyang Technological University / Dalian University of TechnologySingapore / China
  • Ranked 22nd in the 2024 ICPC Asia Pacific Championship
  • Ranked 13th in the 2023 ICPC Asia Jakarta Regional Contest
  • Ranked 6th in the 2022 ICPC Asia Manila Regional Contest
  • Gold Medal in the 2021 ICPC Asia Kunming Regional Contest
  • Silver Medal in the 2021 ICPC Asia Nanjing Regional Contest
Simon Marais Mathematics Competition
Nanyang Technological UniversitySingapore
  • Received the Best-in-University Prize at Nanyang Technological University
  • Solved problems in number theory, game theory, and calculus

5   Teaching & Activities

Tutorial Lecturer
NTU SC1008 C & C++ ProgrammingSingapore
  • Instructed a cohort of 50 beginners in core C/C++ syntax and methodologies
Nanyang Programming Contest
Nanyang Technological UniversitySingapore
  • Organized five competitions and tutorials for non-ICPC students
  • Designed activities to help students strengthen algorithmic problem-solving skills for technical interviews
NTU Students' Computing and Data Science Club
Nanyang Technological UniversitySingapore
  • Managed a Telegram group with 130+ members
  • Answered AI/ML questions and shared resources on deep learning research
Teaching Assistant
NTU SC1003 Introduction of Computational ThinkingSingapore
  • Instructed a cohort of 20 beginners in core Python syntax and data manipulation methodologies

6   Interests

Physics, Chess (Lichess rating: 2084), Calligraphy, Table Tennis