EDBT 2026 Demo / reviewers in the wild / expert
Chenghe Wang
dblp:181/7478
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Reinforcement learning · 53% Multi-agent systems · 20% Planning, search and constraint satisfaction · 13% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
1.7 | 3 | 2022 | Efficient Multi-agent Communication via Self-supervised Information Aggregation · NeurIPS 2022 Multi-Agent Concentrative Coordination with Decentralized Task Representation · IJCAI 2022 Multi-Agent Incentive Communication via Decentralized Teammate Modeling · AAAI 2022 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
multi-agent communication |
1.1 | 2 | 2022 | Efficient Multi-agent Communication via Self-supervised Information Aggregation · NeurIPS 2022 Multi-Agent Incentive Communication via Decentralized Teammate Modeling · AAAI 2022 |
Knowledge, reasoning and agents › Multi-agent systems › multi-agent coordination
cooperative coordination |
0.7 | 2 | 2022 | Multi-Agent Incentive Communication via Decentralized Teammate Modeling · AAAI 2022 Efficient Multi-agent Communication via Self-supervised Information Aggregation · NeurIPS 2022 |
Robotics › Motion planning and robot control › path planning
collision-free path planning |
0.7 | 1 | 2023 | Towards Deployment-Efficient and Collision-Free Multi-Agent Path Finding (Student Abstract) · AAAI 2023 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference
counterfactual prediction |
0.7 | 1 | 2023 | Adversarial Counterfactual Environment Model Learning · NeurIPS 2023 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.7 | 1 | 2023 | Adversarial Counterfactual Environment Model Learning · NeurIPS 2023 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
multi-agent path finding |
0.7 | 1 | 2023 | Towards Deployment-Efficient and Collision-Free Multi-Agent Path Finding (Student Abstract) · AAAI 2023 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.7 | 1 | 2023 | Adversarial Counterfactual Environment Model Learning · NeurIPS 2023 |
Machine learning › Reinforcement learning › model-based reinforcement learning
world model |
0.7 | 1 | 2023 | Adversarial Counterfactual Environment Model Learning · NeurIPS 2023 |
Knowledge, reasoning and agents › Multi-agent systems › social choice › computational social choice
information aggregation |
0.6 | 1 | 2022 | Efficient Multi-agent Communication via Self-supervised Information Aggregation · NeurIPS 2022 |
Knowledge, reasoning and agents › Multi-agent systems
multi-agent coordination |
0.6 | 1 | 2022 | Multi-Agent Concentrative Coordination with Decentralized Task Representation · IJCAI 2022 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › hierarchical problem solving
task decomposition |
0.6 | 1 | 2022 | Multi-Agent Concentrative Coordination with Decentralized Task Representation · IJCAI 2022 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning › value-based multi-agent reinforcement learning
value decomposition |
0.2 | 1 | 2022 | Multi-Agent Incentive Communication via Decentralized Teammate Modeling · AAAI 2022 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 0.7planning-based algorithm · 0.7empirical risk minimization · 0.7adversarial weighted empirical risk minimization · 0.7value function factorization · 0.6value decomposition · 0.6teammate modeling · 0.6multi-agent reinforcement learning · 0.6message extraction · 0.6attention mechanism · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Towards Deployment-Efficient and Collision-Free Multi-Agent Path Finding (Student Abstract)abstractMulti-agent pathfinding (MAPF) is essential to large-scale robotic coordination tasks. Planning-based algorithms show their advantages in collision avoidance while avoiding exponential growth in the number of agents. Reinforcement-learning (RL)-based algorithms can be deployed efficiently but cannot prevent collisions entirely due to the lack of hard constraints. This paper combines the merits of planning-based and RL-based MAPF methods to propose a deployment-efficient and collision-free MAPF algorithm. The experiments show the effectiveness of our approach. Feng Chen 0042, Chenghe Wang, Fuxiang Zhang, Qiaoyong Zhong, Shiliang Pu, Zongzhang Zhang |
AAAI | 2 |
| 2023 | Adversarial Counterfactual Environment Model LearningabstractAn accurate environment dynamics model is crucial for various downstream tasks in sequential decision-making, such as counterfactual prediction, off-policy evaluation, and offline reinforcement learning.
Currently, these models were learned through empirical risk minimization (ERM) by step-wise fitting of historical transition data. This way was previously believed unreliable over long-horizon rollouts because of the compounding errors, which can lead to uncontrollable inaccuracies in predictions. In this paper, we find that the challenge extends beyond just long-term prediction errors: we reveal that even when planning with one step, learned dynamics models can also perform poorly due to the selection bias of behavior policies during data collection.
This issue will significantly mislead the policy optimization process even in identifying single-step optimal actions, further leading to a greater risk in sequential decision-making scenarios.
To tackle this problem, we introduce a novel model-learning objective called adversarial weighted empirical risk minimization (AWRM). AWRM incorporates an adversarial policy that exploits the model to generate a data distribution that weakens the model's prediction accuracy, and subsequently, the model is learned under this adversarial data distribution.
We implement a practical algorithm, GALILEO, for AWRM and evaluate it on two synthetic tasks, three continuous-control tasks, and \textit{a real-world application}. The experiments demonstrate that GALILEO can accurately predict counterfactual actions and improve various downstream tasks, including offline policy evaluation and improvement, as well as online decision-making. Xiong-Hui Chen, Yang Yu 0001, Zhengmao Zhu, Zhihua Yu, Zhenjun Chen, Chenghe Wang, Rong-Jun Qin, Hongqiu Wu, Ruijin Ding, Fangsheng Huang |
NeurIPS | 6 |
| 2022 | Multi-Agent Incentive Communication via Decentralized Teammate ModelingabstractEffective communication can improve coordination in cooperative multi-agent reinforcement learning (MARL). One popular communication scheme is exchanging agents' local observations or latent embeddings and using them to augment individual local policy input. Such a communication paradigm can reduce uncertainty for local decision-making and induce implicit coordination. However, it enlarges agents' local policy spaces and increases learning complexity, leading to poor coordination in complex settings. To handle this limitation, this paper proposes a novel framework named Multi-Agent Incentive Communication (MAIC) that allows each agent to learn to generate incentive messages and bias other agents' value functions directly, resulting in effective explicit coordination. Our method firstly learns targeted teammate models, with which each agent can anticipate the teammate's action selection and generate tailored messages to specific agents. We further introduce a novel regularization to leverage interaction sparsity and improve communication efficiency. MAIC is agnostic to specific MARL algorithms and can be flexibly integrated with different value function factorization methods. Empirical results demonstrate that our method significantly outperforms baselines and achieves excellent performance on multiple cooperative MARL tasks. Lei Yuan 0005, Fuxiang Zhang, Chenghe Wang, Zongzhang Zhang, Yang Yu 0001, Chongjie Zhang |
AAAI | 4 |
| 2022 | Multi-Agent Concentrative Coordination with Decentralized Task RepresentationabstractValue-based multi-agent reinforcement learning (MARL) methods hold the promise of promoting coordination in cooperative settings. Popular MARL methods mainly focus on the scalability or the representational capacity of value functions. Such a learning paradigm can reduce agents' uncertainties and promote coordination. However, they fail to leverage the task structure decomposability, which generally exists in real-world multi-agent systems (MASs), leading to a significant amount of time exploring the optimal policy in complex scenarios. To address this limitation, we propose a novel framework Multi-Agent Concentrative Coordination (MACC) based on task decomposition, with which an agent can implicitly form local groups to reduce the learning space to facilitate coordination. In MACC, agents first learn representations for subtasks from their local information and then implement an attention mechanism to concentrate on the most relevant ones. Thus, agents can pay targeted attention to specific subtasks and improve coordination. Extensive experiments on various complex multi-agent benchmarks demonstrate that MACC achieves remarkable performance compared to existing methods. Lei Yuan 0005, Chenghe Wang, Fuxiang Zhang, Feng Chen 0042, Cong Guan, Zongzhang Zhang, Chongjie Zhang, Yang Yu 0001 |
IJCAI | 2 |
| 2022 | Efficient Multi-agent Communication via Self-supervised Information AggregationabstractUtilizing messages from teammates can improve coordination in cooperative Multi-agent Reinforcement Learning (MARL). To obtain meaningful information for decision-making, previous works typically combine raw messages generated by teammates with local information as inputs for policy. However, neglecting the aggregation of multiple messages poses great inefficiency for policy learning. Motivated by recent advances in representation learning, we argue that efficient message aggregation is essential for good coordination in MARL. In this paper, we propose Multi-Agent communication via Self-supervised Information Aggregation (MASIA), with which agents can aggregate the received messages into compact representations with high relevance to augment the local policy. Specifically, we design a permutation invariant message encoder to generate common information aggregated representation from raw messages and optimize it via reconstructing and shooting future information in a self-supervised manner. Each agent would utilize the most relevant parts of the aggregated representation for decision-making by a novel message extraction mechanism. Empirical results demonstrate that our method significantly outperforms strong baselines on multiple cooperative MARL tasks for various task settings. Cong Guan, Feng Chen 0042, Lei Yuan 0005, Chenghe Wang, Zongzhang Zhang, Yang Yu 0001 |
NeurIPS | 4 |