My current research focuses on building executable virtual environments inspired by generative agents, which simulate real-world interaction dynamics for dynamic evaluation, synthetic data generation, and RL-based post-training. I also work on agentic multimodal understanding, especially tool-augmented long-document and long-video agents, as well as multimodal model merging for aligning perception and reasoning capabilities and enabling cross-modal knowledge transfer. These efforts are developed under Office Intelligence, a project hub for our recent work on agentic document intelligence.
Executable generative-agent environments for dynamic evaluation, synthetic data generation, and Agentic RL
Tool-augmented multimodal agents for long-document and long-video understanding
Multimodal model merging for perception-reasoning alignment and cross-modal knowledge transfer
News
[Dec. 2025 - Present] Research intern at Microsoft Research Asia (MSRA), working on executable multi-agent environments for realistic office workflows, Agentic Document Understanding, agent evaluation, and data synthesis.
Pre-prints
DocAtlas: Long-Document Understanding as Mutable-State Interaction Hongchen Wei, et al.
NeurIPS 2026 (under review) Work done at MSRA.