I am a Ph.D. candidate in Computer Science at UIUC , advised by Professors Tong Zhang and Huan Zhang. Previously, I received bachelor’s and master’s degrees from Tsinghua University and HKUST .
My research focuses on building reliable interactive agents powered by foundation models that can perceive multimodal inputs, reason about complex tasks, evaluate outcomes, and act effectively in dynamic environments. I also study robust learning under distribution shift and imperfect feedback, spanning LLM/VLM alignment, reward modeling, and robust offline and goal-conditioned RL.
Multimodal agents
Scalable post-training and evaluation for multimodal agents that perceive accurately, assess their current state, and reason and act effectively over long horizons.
OpenWebRL · GUI-Libra · GUI-Actor · MeMento · ERA · EmbodiedBench
Trustworthy alignment & evaluation
Aligning LLMs/VLMs with heterogeneous human preferences, together with robust reasoning evaluation under changing inputs.
Robust and generalizable RL
Robust offline and goal-conditioned RL under data corruption, observation shifts, and unseen goals.
I am actively seeking full-time opportunities in foundation models, AI agents, and reinforcement learning.
News
- 🌟 2026.06 We released OpenWebRL, an open framework for training visual web agents with online multi-turn RL on live websites, and Orchard, an open-source framework with scalable training recipes across diverse agent domains.
- 🎉 2026.04 ReCAP was accepted to ICML 2026.
- 🌟 2026.02 We released GUI-Libra, a data-efficient post-training recipe for GUI agents that achieves strong online performance using 81K open-source examples.
- 🎉 2026.01 BEAT and DROCO were accepted to ICLR 2026.
- 🎉 2025.11 MiCRo received an EMNLP 2025 Outstanding Paper Award.
- 🎉 2025.09 GUI-Actor and ADG were accepted to NeurIPS 2025, and MergeBench to the Datasets & Benchmarks Track.
- 🎉 2025.05 EmbodiedBench was accepted to ICML 2025 as an oral paper.
Selected Publications
* Equal contribution. Selected papers I led or co-led; see Google Scholar for the full list.
Foundation Models for Interactive Agents
OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents.
Preprint 2026 [Code] [Project] [Models & Data]
GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL.
Preprint 2026 [Code] [Project]
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents.
ICML 2025 Oral [Code] [Project]
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents.
NeurIPS 2025 [Code] [Project]
Reliable Reasoning and Alignment
DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models.
ICLR 2025 [Code] [Project]
Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs.
NeurIPS 2024 [Code]
Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment.
ICML 2024 [Code]
MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning.
EMNLP 2025 Outstanding Paper
Robust Offline RL and Goal-Conditioned RL
Towards Robust Offline Reinforcement Learning under Diverse Data Corruption.
ICLR 2024 Spotlight [Code]
RORL: Robust Offline Reinforcement Learning via Conservative Smoothing.
NeurIPS 2022 Spotlight [Code]
What Is Essential for Unseen Goal Generalization of Offline Goal-conditioned RL?.
ICML 2023 [Code]
Rethinking Goal-conditioned Supervised Learning and Its Connection to Offline RL.
ICLR 2022 [Code]
Experience
Research Intern, Apple Foundation Models, 2026.
Research Intern, Microsoft Research, Deep Learning Group, 2025.
Research Intern, Tencent AI Lab and Robotics X Lab, 2020–2022 (multiple internship terms).
Machine Learning Intern, Meituan, 2019.
Service
Conference Reviewer: ICML, ICLR, NeurIPS (NeurIPS 2023 Top Reviewer), ACL/ARR, ICRA, AAMAS.
Journal Reviewer: IEEE Robotics and Automation Letters (RA-L), IEEE Transactions on Neural Networks and Learning Systems (TNNLS), IEEE Transactions on Artificial Intelligence (TAI), Machine Learning, Journal of Artificial Intelligence Research.
Teaching Assistant: CS 441 Applied Machine Learning, UIUC; COMP 4211 Machine Learning, HKUST; COMP 1021 Introduction to Computer Science, HKUST.
Hobbies
Outside research, I enjoy 🏃 running, 🏓 table tennis, and 🏊 swimming. My personal bests are 1 h 30 min in the half marathon and 3 h 36 min in the marathon.