I'm Hao Wang (ηθ±ͺ), a Ph.D. candidate at HCP Lab, Sun Yat-sen University, and Pengcheng Laboratory, advised by Prof. Xiaodan Liang and Assoc. Prof. Xiangyuan Lan.
π¬ My research centers on open-ended computer vision and multimodal large language models, and I'm increasingly exploring multimodal agentic models.
π I'll graduate in December 2026 and am actively looking for research roles in industry β I'm also open to research collaborations on interesting projects.
π§ Reach me via Email: wanghao9610@gmail.com or WeChat: wangh9610.
- π₯ Any segmentation in images and videos: X2SAM
- π₯ From segment anything to any segmentation: X-SAM
- Unified open-vocabulary detection: OV-DINO
- Temporal memory attention for video semantic segmentation: TMANet
- π₯ Systematic Toolchain for AI Research (Harness, WIP): STAR
- π₯ Template paper for arXiv or any conference: arXivTeX
- Run Codex on a remote server: Codex Remote Connector
- Run Claude on multiple AI providers: Claude Model Proxy


