qfwen@bu.edu
Shanghai Jiao Tong University
Shanghai, China
I am a Research Assistant in Prof. Xi Lin’s group at Shanghai Jiao Tong University. Previously I was a researcher in the ML Group at Boston University, advised by Prof. Reza Rawassizadeh. I completed my M.S. in Computer Science at BU and my B.S. in Mathematics at NYU Shanghai.
My research centers on efficient and reliable training of large models, viewed through a stochastic-process / training-dynamics lens. GradES speeds up full-parameter fine-tuning by 1.70–1.77× at neutral accuracy by selectively freezing transformer components at convergence. I have a strong mathematical background in measure-theoretic probability and stochastic processes.
I am a reviewer for ICML 2026 and NeurIPS 2026. I am currently applying to PhD programs for Fall 2027.
Outside of my research, I enjoy playing basketball, hiking and playing rimworld.
Feel free to reach out if you’d like to chat about research or collaboration!
Research & publications
All papers and preprints, newest year first. Explore a preview or open the paper.
No papers match. Try another title, author or topic.
2026
-
Paper-figure animation: complete scene, occupied targets, visible context. -
Sampled paper rollout frames: Future-Consensus versus CoWAM pot lifting. CoWAM: Coordination Contracts for Selective Policy Intervention with WAMs2026Preprint; under review at AAAI 2027 -
OceanEnv marine robotics simulation. One Ocean, All Tasks: A Holistic Simulation Environment for Marine Robotics2026Under review at AAAI 2027 -
2025
- GradES: Significantly Faster Training in Transformers with Gradient-Based Early Stopping2025Preprint; under review at AAAI 2027
-
-
news
| Jun 15, 2026 | OceanEnv submitted to AAAI 2027. |
|---|---|
| Jun 01, 2026 | Joined Prof. Xi Lin’s group at Shanghai Jiao Tong University as a Research Assistant. |
| May 01, 2026 | |
| Feb 01, 2026 | Paper “Estimation of Distribution Parameters” published in Statistical Papers (Springer), Vol. 67, Article 31. |
| Sep 01, 2025 | GradES preprint released on arXiv — 1.70–1.77× faster full-parameter fine-tuning of transformers at neutral accuracy. |