Hi, Welcome to my Homepage!

🧑‍🎓 About me

I am a PhD student at Fudan University. My research focuses on reliable and efficient visual reasoning in Large Vision-Language Models, with particular interests in multimodal perception, reasoning, and visual token pruning.

I am passionate about building vision-language systems that can better perceive, reason, and respond in complex visual environments. My recent work explores how to enhance high-resolution visual understanding, improve reasoning ability, and reduce inference costs for Large Vision-Language Models. Feel free to connect with me via email: yxliang2001@gmail.com.

🔥 News

  • 2026.02 🎉🎉 Our Work on LVLM Reasoning Hallucination have been accepted by CVPR 2026. Thanks all of the co-authors!
  • 2026.02 🎉🎉 Our Work on LVLM Token Pruning have been accepted by TCSVT 2026. Thanks all of the co-authors!
  • 2026.01 🎉🎉 Our Work on LVLM Reasoning have been accepted by ICLR 2026. Thanks all of the co-authors!
  • 2025.09 🎉🎉 I become a PhD student at Fudan University.
  • 2023.09 🎉🎉 I began my studies at Fudan University.

📝 Publications

Envision, Attend, Then Respond
Envision, Attend, Then Respond: Counterfactual Hallucination Mitigation in Large Vision-Language Models
Yuxuan Liang, et al.
We propose an "Envision–Attend–Respond" framework that mitigates hallucinations in Large Vision-Language Models through counterfactual visual reasoning, encouraging the model to first imagine alternative scenes, re-attend to grounded evidence, and then generate faithful responses.
CVPR 2026   [arXiv] [code]
Decomposition of Concept-Level Rules
Decomposition of Concept-Level Rules in Visual Scenes
Yuxuan Liang, et al.
We study how high-level visual reasoning rules can be decomposed into interpretable concept-level primitives, enabling models to discover, reuse, and compose rules across diverse visual scenes for more systematic and generalizable reasoning.
ICLR 2026   [arXiv] [code]
Pyramid Token Pruning
Pyramid Token Pruning for High-Resolution Large Vision-Language Models via Region, Token, and Instruction-Guided Importance
Yuxuan Liang, et al.
We propose a pyramid token pruning strategy for high-resolution LVLMs that jointly leverages region-level, token-level, and instruction-guided importance, substantially reducing visual token cost while preserving fine-grained understanding.
TCSVT 2026   [arXiv] [code]
HERO Token Early Dropping
HERO: Rethinking Visual Token Early Dropping in High-Resolution Large Vision-Language Models
Yuxuan Liang, et al.
We rethink visual token early dropping in high-resolution Large Vision-Language Models to improve efficiency without sacrificing fine-grained perception capability.
[TMM 2026]   [arXiv] [code]

🥇 Awards

  • 2025.09, Fudan University Doctoral Freshman Scholarship.
  • 2024.09, Fudan University Master’s Academic Scholarship.
  • 2023.09, Fudan University Master’s Freshman Scholarship.

💻 Internships

  • 2023.05 - 2024.09, Beijing Zhipu Huazhang Technology, Beijing, China