Hongchen Wei

I am currently a final-year Ph.D. student at Wuhan University, under the supervision of Prof. Zhenzhong Chen.

I received my M.E. degree from Nanjing University of Science and Technology, China, in 2023.

I received my B.Sc. degree from Xi'an Shiyou University, China, in 2020.


profile photo
Research agenda

I build executable environments where agents work inside persistent virtual organizations: they assume roles, use real tools, coordinate under permissions, and act across long-running events. My goal is to close the loop between agent evaluation, trajectory and data synthesis, and RL-based post-training.

Office Intelligence Core research hub A research hub for agents that work: executable office environments, document intelligence, and post-training. Visit the research hub
Executable agent environments Organizations · Tools · Evals · Data · RL Agentic multimodal understanding Long documents · Long videos Perception-reasoning alignment Model merging · Cross-modal transfer
News
  • [Dec. 2025 - Present] Research intern at Microsoft Research Asia (MSRA), working on executable multi-agent environments for realistic office workflows, Agentic Document Understanding, agent evaluation, and data synthesis.
Pre-prints
DocAtlas: Long-Document Understanding as Mutable-State Interaction
Hongchen Wei, et al.
NeurIPS 2026 (under review)
Work done at MSRA.
XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding
Hongchen Wei, et al.
NeurIPS 2026 (under review)
Work done at MSRA.
PASA: Post-Merge Perception-Reasoning Asymmetry as a Self-Alignment Signal for MLLMs
Hongchen Wei, Zhenzhong Chen.
NeurIPS 2026 (under review)
Training-Free Reasoning and Reflection in MLLMs
Hongchen Wei, Zhenzhong Chen
arXiv Preprint, 2025
LongCaptioning: Unlocking the Power of Long Caption Generation in Large Multimodal Models
Hongchen Wei, Zhihong Tan, Yaosi Hu, Chang Wen Chen, Zhenzhong Chen
arXiv Preprint, 2025
LOP: Learning Optimal Pruning for Efficient On-Demand MLLMs Scaling
Zhihan Zhang, Xiang Pan, Hongchen Wei, Zhenzhong Chen
arXiv Preprint, 2025
RSFAKE-1M: A Large-Scale Dataset for Detecting Diffusion-Generated Remote Sensing Forgeries
Zhihong Tan, Jiayi Wang, Huiying Shi, Binyuan Huang, Hongchen Wei, Zhenzhong Chen
arXiv Preprint, 2025
TDSAgent: A Task-Driven Sampling Agent for Long Video Question Answering
Author list includes Hongchen Wei
Under Review, 2026
GTC: Game-Theoretic Token Compression for Video Large Language Models
Author list includes Hongchen Wei
Under Review, 2026
ETC: Extreme Token Compression via Task-aware Visual Information Distillation in VLMs
Author list includes Hongchen Wei
Under Review, 2026
Publications
See What We Cannot See: A Geo-guided Reasoning Benchmark for Object Counting under Adverse Earth Observation Conditions
Author list includes Hongchen Wei
CVPR, 2026
Visual Context Window Extension: A New Perspective for Long Video Understanding
Hongchen Wei, Zhenzhong Chen
ACM MM (CCF-A Conference), 2025
Project page
RealVG: Unleashing MLLMs for Training-Free Spatio-Temporal Video Grounding in the Wild
Hongchen Wei, Zhenzhong Chen
ACM MM (CCF-A Conference), 2025
Remote Sensing Semantic Segmentation Quality Assessment based on Vision Language Model
Huiying Shi, Zhihong Tan, Zhihan Zhang, Hongchen Wei, Yaosi Hu, Yingxue Zhang, Zhenzhong Chen
TGRS (CCF-B Journal), 2025
Improving Generalization of Image Captioning with Unsupervised Prompt Learning
Hongchen Wei, Zhenzhong Chen
TOMM (CCF-B Journal), 2024
Exploiting Cross-Modal Prediction and Relation Consistency for Semisupervised Image Captioning
Yang Yang, Hongchen Wei , Hengshu Zhu, Dianhai Yu, Hui Xiong, Jian Yang
TCYB (CCF-B Journal), 2022 (student first author)
Code
S2OSC: A Holistic Semi-Supervised Approach for Open Set Classification
Yang Yang, Hongchen Wei , Zhenqiang Sun, Guangyu Li, Yuanchun Zhou, Hui Xiong, Jian Yang
TKDD (CCF-B Journal), 2021 (student first author)
Activities
  • Reviewer: ICLR25/26, CVPR25/26, NeurIPS25/26, ICML26, TNNLS

Last updated in Aug. 2026.

Homepage credits: Jon Barron.