✨ About me

Bonjour, I am a Ph.D. student in the Department of Computer Sciences at the University of Wisconsin–Madison, advised by Prof. Sharon Li. Before this, I obtain the graduate degree of computer application technology at Peking University, advised by Prof. Jie Chen. I received a B.S. degree from the College of Computer Science (Elite Class), in Jilin University, in 2022.

My research focuses on Trustworthy Multimodal Intelligence, with particular interests in:

  • Interpretability and Mechanistic Alignment: Understanding the internal representations and decision-making mechanisms of multimodal models, with the goal of making their reasoning more interpretable, controllable, and safety-aligned.

  • Agentic Image and Video Generation: Building capable and efficient generative agentic systems through post-training (e.g., SFT, preference optimization, and reinforcement learning) and agent harnesses that enable planning, tool use, iterative feedback, and self-correction for high-fidelity and efficient visual generation.

  • Unified and Omni Multimodal Models: Developing unified multimodal models that integrate perception, reasoning, generation, and action within a shared framework, toward more general-purpose multimodal intelligence.

😃 I am seeking research collaboration opportunities. Feel free to contact me if you are interested in potential collaboration!

📰 News

  • 2026/04: 🎉 One paper about Multimodal Interpretability is accepted by ACL 2026 as main paper.
  • 2024/12: 🎉 One paper (AnyTalk about Talking Head Generation is accepted by AAAI 2025.
  • 2024/09: 🎉 One paper (SubgDiff about molecular 3D conformation generation is accepted by NeurIPS 2024.
  • 2024/07: 🎉 Two papers (Adashield about VLM Safety and ParCo about Text-to-Motion Synthesis) are accepted by ECCV 2024.
  • 2024/04: 🚀 We release an open-source repository about Awesome-T2I-safety-Papers.
  • 2023/10: 🎉 One paper (TIDA) about open-set semi-supervised learning is accepted by NeurIPS 2023.
  • 2023/03: 🎉 Two papers (OSP about robust semi-supervised learning and FPR about semi-supervised semantic segmentation) are accepted by CVPR 2023.

💻 Selected Publications

My complete publication record is available on Google Scholar. Each highlighted paper below includes a visual overview and direct links to its resources.

ACL 2026Multimodal Interpretability

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks

Yu Wang, Sharon Li

A mechanistic study of how multimodal in-context learning routes visual and textual evidence—and where that reasoning pipeline falls short.

🎖 Honors and Awards

  • 2025/05: Outstanding Graduate, Peking University (University Graduation Honor) (Top 5%)
  • 2024/11: Tiehan Scholarship, Peking University (Top 5%)
  • 2024/11: Merit Student, Peking University (Top 10%)
  • 2023/11: Merit Student, Peking University (Top 10%)
  • 2023/03: Excellent Papers of China Computer Federation Computer Application Conference, CCF Conference on Computer Applications (Top 1%)
  • 2022/12: Jilin Bank Wang Xianghao Scholarship, Jilin University (Top 1%)
  • 2021/11: National Scholarship, Jilin University (Top 1%)
  • 2020/10: National Inspirational Scholarship, Jilin University (Top 1%)

💪 Academic Service

  • PC Member: CVPR’24-26, ICLR’24-26, ICML’24-26, ACM MM’24, ECCV’24,26, NeurIPS’24-25, AAAI’26

If you like the template of this homepage, welcome to star and fork the open-sourced template version AcadHomepage .