Hohei Chan

dblp:421/0419 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
0009-0000-9326-6235ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 42% Multi-agent systems · 28% Generative modeling · 15%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Multi-agent systems › multi-agent collaboration
ad hoc teamwork
1.922026
PADiff: Predictive and Adaptive Diffusion Policies for Ad Hoc Teamwork · AAAI 2026
Ad Hoc Teamwork via Offline Goal-Based Decision Transformers · ICML 2025
Machine learning › Generative modeling
diffusion model
1.012026
PADiff: Predictive and Adaptive Diffusion Policies for Ad Hoc Teamwork · AAAI 2026
Robotics › Robot manipulation
diffusion policy
1.012026
PADiff: Predictive and Adaptive Diffusion Policies for Ad Hoc Teamwork · AAAI 2026
Machine learning › Reinforcement learning
multi-agent reinforcement learning
1.012026
PADiff: Predictive and Adaptive Diffusion Policies for Ad Hoc Teamwork · AAAI 2026
Machine learning › Reinforcement learning › offline reinforcement learning
decision transformer
0.912025
Ad Hoc Teamwork via Offline Goal-Based Decision Transformers · ICML 2025
Machine learning › Reinforcement learning
offline reinforcement learning
0.912025
Ad Hoc Teamwork via Offline Goal-Based Decision Transformers · ICML 2025

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.0diffusion model · 1.0sequence modeling · 0.9goal-based decision transformer · 0.9
YearPublicationVenuePosition
2026 PADiff: Predictive and Adaptive Diffusion Policies for Ad Hoc Teamwork
abstract
Ad hoc teamwork (AHT) requires agents to collaborate with previously unseen teammates, which is crucial for many real-world applications. The core challenge of AHT is to develop an ego agent that can predict and adapt to unknown teammates on the fly. Conventional RL-based approaches optimize a single expected return, which often causes policies to collapse into a single dominant behavior, thus failing to capture the multimodal cooperation patterns inherent in AHT. In this work, we introduce PADiff, a diffusion-based approach that captures agent's multimodal behaviors, unlocking its diverse cooperation modes with teammates. However, standard diffusion models lack the ability to predict and adapt in non-stationary AHT scenarios. To address this limitation, we propose a novel diffusion-based policy that integrates critical predictive information about teammates into the denoising process. Extensive experiments across three environments demonstrate that PADiff outperforms existing AHT methods significantly.
Hohei Chan, Xinzhi Zhang 0009, Antao Xiang, Weinan Zhang 0001, Mengchen Zhao
AAAI1
2025 Ad Hoc Teamwork via Offline Goal-Based Decision Transformers
abstract
The ability of agents to collaborate with previously unknown teammates on the fly, known as ad hoc teamwork (AHT), is crucial in many real-world applications. Existing approaches to AHT require online interactions with the environment and some carefully designed teammates. However, these prerequisites can be infeasible in practice. In this work, we extend the AHT problem to the offline setting, where the policy of the ego agent is directly learned from a multi-agent interaction dataset. We propose a hierarchical sequence modeling framework called TAGET that addresses critical challenges in the offline setting, including limited data, partial observability and online adaptation. The core idea of TAGET is to dynamically predict teammate-aware rewards-to-go and sub-goals, so that the ego agent can adapt to the changes of teammates’ behaviors in real time. Extensive experimental results show that TAGET significantly outperforms existing solutions to AHT in the offline setting.
Xinzhi Zhang 0009, Hohei Chan, Deheng Ye, Yi Cai 0001, Mengchen Zhao
ICML2