Fanyi Pu 濮凡轶

University of Wisconsin–Madison
Abstract
I am currently a final-year undergraduate student at Nanyang Technological University, specializing in Data Science and Artificial Intelligence.
I am currently conducting research at SenseTime and MMLab@NTU. I am fortunate to be supervised by Prof. Ziwei Liu, I am also grateful for the extensive help I have received from Bo Li, Yuanhan Zhang, Zhongang Cai and Lei Yang. My research interests center on multimodal models.
During my time at Shaoxing No.1 High School, I participated in the Olympiad in Informatics. After graduating, I spent a rewarding year at the International School of Information Science & Engineering at Dalian University of Technology. During the COVID-19 pandemic, I chose to withdraw and begin my undergraduate studies anew at Nanyang Technological University.
I competed in the ICPC at both universities. I also contribute to the Nanyang Programming Contest. I am grateful to my teammates and friends for their support! 🙏
I enjoy playing chess. My Lichess username is @pufanyi, and my current rapid rating is 2086.
1 Education
- Honours (Highest Distinction), CGPA: 4.67 / 5.00
- Research interests: Multimodal and Video Generation. 1,800+ Google Scholar citations
- Selected coursework: Reinforcement Learning, Deep Learning, Data Structures and Algorithms, Probability and Statistics, Computational Economics, Stochastic Processes, Numerical Analysis, Cryptography, Electricity & Magnetism
- Studied Computer Security and Game Theory, GPA: 4.00 / 4.00
2 Publications & Research
- A lean and flexible training framework designed for both rapid research prototyping and large-scale production
- Optimized the Bagel training pipeline by integrating FSDP2 and Liger Kernel for efficient distributed training
- Designing asynchronous reinforcement-learning infrastructure for the project
- An open-source multimodal evaluation framework with 4.3K GitHub stars
- Built the evaluation pipeline and developed a low-cost automatic data-generation pipeline for Multi-modal LiveBench, which leverages continuously updated news and online forums to evaluate models' generalization in the wild
- An efficient native-resolution foundation model for high-quality image generation and editing
- Extended LMMs-Engine for Mage-Flow pre-training by integrating FSDP2, Liger Kernel, and sequence packing, enabling efficient native multi-resolution training across heterogeneous image resolutions and aspect ratios
- Revisited video reasoning and proposed Chain-of-Step, in which models reason through diffusion steps
- Curated a large collection of examples to validate this hypothesis and systematically analyzed models' reasoning behaviors
- Developed a foundation model for scaling spatial intelligence, achieving state-of-the-art performance on key spatial benchmarks
- Conducted rigorous ablation studies to establish that spatial intelligence is primarily driven by data scaling rather than reasoning heuristics (CoT), guiding the project's strategic focus on constructing the SenseNova-SI-8M dataset
- Executed the full-stack training pipeline for Qwen-based variants with LMMs-Engine on 128 GPUs
- Built an evaluation set with 300 expert-level videos and 900 human-annotated questions across six disciplines, cited by Gemini 3 Pro (Google) and GPT-5 (OpenAI)
- Proposed a knowledge-gain metric to quantify performance improvements after watching video lectures; validated expert-level data in economics and medicine to ensure benchmark quality
- Developed an early vision-language-agent (VLA) model with 3.3K GitHub stars
- Developed a language-model-based pipeline that generated 2.8M multimodal instruction-tuning samples; created an evaluation framework that later evolved into LMMs-Eval
3 Research & Professional Experience
- Core member of a non-profit initiative dedicated to democratizing Large Multimodal Models (LMMs); instrumental in the lab's 0-to-1 establishment, from initial ideation and naming to its current operations
- Spearheaded the development of high-impact open-source projects: the evaluation framework LMMs-Eval, the efficient training library LMMs-Engine, and the video understanding benchmark Video-MMMU
- Conducted research on SenseNova-SI and video reasoning, with a focus on spatial intelligence and video generation
- Architected and implemented the Synvo File System, the company's inaugural product, serving as the storage backbone for unstructured multimodal data
- Empowered the core contextualization engine by structuring data for efficient retrieval, enabling the development of AI systems that can remember, learn, and adapt
- Supervised by Prof. Liu Ziwei; focused on multimodal language models and unified multimodal models
4 Competitions
- Ranked 22nd in the 2024 ICPC Asia Pacific Championship
- Ranked 13th in the 2023 ICPC Asia Jakarta Regional Contest
- Ranked 6th in the 2022 ICPC Asia Manila Regional Contest
- Gold Medal in the 2021 ICPC Asia Kunming Regional Contest
- Silver Medal in the 2021 ICPC Asia Nanjing Regional Contest
- Received the Best-in-University Prize at Nanyang Technological University
- Solved problems in number theory, game theory, and calculus
5 Teaching & Activities
- Instructed a cohort of 50 beginners in core C/C++ syntax and methodologies
- Organized five competitions and tutorials for non-ICPC students
- Designed activities to help students strengthen algorithmic problem-solving skills for technical interviews
- Managed a Telegram group with 130+ members
- Answered AI/ML questions and shared resources on deep learning research
- Instructed a cohort of 20 beginners in core Python syntax and data manipulation methodologies
6 Interests
Physics, Chess (Lichess rating: 2084), Calligraphy, Table Tennis