About Me
I’m a Research Scientist at NVIDIA Research, focusing on agentic AI and reinforcement learning for capable, scalable LLM agents. My research also spans large multimodal models, visual understanding, and efficient generative inference. I graduated in August 2024 from the Department of Computer Science & Engineering, Hong Kong University of Science and Technology, with a Ph.D. co-supervised by Prof. Heung-Yeung Shum and Prof. Lionel M. Ni. I interned at International Digital Economy Academy, Shenzhen (advised by Prof. Lei Zhang) and Microsoft Research, Redmond (advised by Dr. Jianwei Yang and Dr. Chunyuan Li). Previously, I obtained my bachelor’s degree from the School of Electronic Information and Electrical Engineering at Shanghai Jiao Tong University in 2019.
📌 My research interests focus on agentic AI and reinforcement learning for scalable LLM agents, alongside multimodal intelligence and efficient inference.
✉️ Welcome to contact me for any discussion and cooperation!
🔥 News
- [2026/07]: Molt, a scalable PyTorch-native training framework for agentic reinforcement learning, is released.
- [2026/05]: Polar, a scalable rollout framework for agentic RL on arbitrary harnesses, is released.
- [2026/04]: Nemotron 3 Nano Omni, an efficient open multimodal model, is released.
- [2026/03]: ProRL Agent, rollout-as-a-service for multi-turn LLM-agent RL training, is released and integrated with NVIDIA NeMo Gym.
- [2026/01]: Fast-dLLM and Fast-dLLM v2 are accepted to ICLR 2026.
- [2025/07]: Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training is released.
- [2025/03]: GR00T N1, an open foundation model for generalist humanoid robots, is released.
- [2025/01]: LLaVA-NeXT-Interleave is accepted to ICLR 2025.
- [2024/09]: Grounding DINO, a highly cited open-set object detector, is presented at ECCV 2024.
- [2023/06]: Mask DINO, a unified framework for object detection and segmentation, is presented at CVPR 2023.
- [2023/05]: DINO, an end-to-end detector with improved denoising anchor boxes, is presented at ICLR 2023.
- [2022/06]: DN-DETR is presented as a CVPR 2022 oral.
📝 Selected Recent & High-Impact Works
For the complete and up-to-date publication list, please see my Google Scholar.
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning.
Jian Hu, Hongyu Li, Hao Zhang, Binfeng Xu, Yuchi Zhang, Shaokun Zhang, Harshil Desai, Michael Demoret, et al.
arXiv 2026.
[Paper]Polar: Agentic RL on Any Harness at Scale.
Binfeng Xu, Hao Zhang, Shaokun Zhang, Songyang Han, Mingjie Liu, Jian Hu, Shizhe Diao, Zhenghui Jin, Yunheng Zou, Michael Demoret, Jan Kautz, Yi Dong.
arXiv 2026.
[Paper]ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents.
Hao Zhang, Mingjie Liu, Shaokun Zhang, Songyang Han, Jian Hu, Zhenghui Jin, Yuchi Zhang, Shizhe Diao, Ximing Lu, Binfeng Xu, Zhiding Yu, Jan Kautz, Yi Dong.
arXiv 2026.
[Paper]Fast-dLLM v2: Efficient Block-Diffusion LLM.
Chengyue Wu, Hao Zhang, Shuchen Xue, Shizhe Diao, Yonggan Fu, Zhijian Liu, Pavlo Molchanov, Ping Luo, Song Han, Enze Xie.
ICLR 2026.
[Paper]Fast-dLLM: Training-Free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding.
Chengyue Wu, Hao Zhang, Shuchen Xue, Zhijian Liu, Shizhe Diao, Ligeng Zhu, Ping Luo, Song Han, Enze Xie.
ICLR 2026.
[Paper]GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.
Jim Bjorck, Francisco Castañeda, Nima Cherniadev, Xiaolong Da, Rui Ding, Linxi Fan, Yifan Fang, Dieter Fox, et al., including Hao Zhang.
arXiv 2025.
[Paper]LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.
Feng Li*, Rui Zhang*, Hao Zhang*, Yifan Zhang, Bo Li, Wei Li, Zejun Ma, Chunyuan Li.
ICLR 2025.
[Paper]LLaVA-OneVision: Easy Visual Task Transfer.
Bo Li, Yifan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Pengchuan Zhang, Yiming Li, et al.
arXiv 2024.
[Paper]
Earlier high-impact works
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, Lei Zhang.
ECCV 2024.
[Paper][Code]DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection.
Hao Zhang*, Feng Li*, Shilong Liu*, Lei Zhang, Hang Su, Jun Zhu, Lionel M. Ni, Heung-Yeung Shum.
ICLR 2023.
[Paper][Code]DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR.
Shilong Liu, Feng Li, Hao Zhang, Xiao Yang, Xianbiao Qi, Hang Su, Jun Zhu, Lei Zhang.
ICLR 2022.
[Paper][Code]DN-DETR: Accelerate DETR Training by Introducing Query DeNoising.
Feng Li*, Hao Zhang*, Shilong Liu, Jian Guo, Lionel M. Ni, Lei Zhang.
CVPR 2022. Oral presentation.
[Paper][Code]SEEM: Segment Everything Everywhere All at Once.
Xueyan Zou*, Jianwei Yang*, Hao Zhang*, Feng Li*, Linjie Li, Jianfeng Gao, Yong Jae Lee.
NeurIPS 2023.
[Paper][Code]Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation.
Feng Li*, Hao Zhang*, Huaizhe Xu, Shilong Liu, Lei Zhang, Lionel M. Ni, Heung-Yeung Shum.
CVPR 2023.
[Paper][Code]
(* denotes equal contribution.)
🎖 Selected Awards
- CSE Best PhD Dissertation Award Honorable Mention, 2023-24
- RedBird Research Scholarship in HKUST, 2020
- Hong Kong Postgraduate Scholoarship, 2020, 2021, 2022 and 2023
