EDBT 2026 Demo / reviewers in the wild / expert
Mian Deng
dblp:396/5834
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Planning, search and constraint satisfaction · 43% Reinforcement learning · 37% Language models and text generation · 11% |
Topics — the 9 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
1.5 | 2 | 2024 | Melting Pot Contest: Charting the Future of Generalized Cooperative Intelligence · NeurIPS 2024 The Dormant Neuron Phenomenon in Multi-Agent Reinforcement Learning Value Factorization · NeurIPS 2024 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
decision making under uncertainty |
0.9 | 1 | 2025 | PlanU: Large Language Model Reasoning through Planning under Uncertainty · NeurIPS 2025 |
Natural language and speech › Language models and text generation
large language model reasoning |
0.9 | 1 | 2025 | PlanU: Large Language Model Reasoning through Planning under Uncertainty · NeurIPS 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › language-based planning
LLM-based planning |
0.9 | 1 | 2025 | PlanU: Large Language Model Reasoning through Planning under Uncertainty · NeurIPS 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search |
0.9 | 1 | 2025 | PlanU: Large Language Model Reasoning through Planning under Uncertainty · NeurIPS 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
planning under uncertainty |
0.9 | 1 | 2025 | PlanU: Large Language Model Reasoning through Planning under Uncertainty · NeurIPS 2025 |
Machine learning › Deep learning architectures and training › training dynamics
dormant neuron |
0.8 | 1 | 2024 | The Dormant Neuron Phenomenon in Multi-Agent Reinforcement Learning Value Factorization · NeurIPS 2024 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning › markov games
mixed-motive games |
0.8 | 1 | 2024 | Melting Pot Contest: Charting the Future of Generalized Cooperative Intelligence · NeurIPS 2024 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning › value-based multi-agent reinforcement learning
value decomposition |
0.8 | 1 | 2024 | The Dormant Neuron Phenomenon in Multi-Agent Reinforcement Learning Value Factorization · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
upper confidence bounds with curiosity · 0.9quantile distribution · 0.9monte carlo tree search · 0.9weight transfer · 0.8value factorization · 0.8multi-agent reinforcement learning · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PlanU: Large Language Model Reasoning through Planning under UncertaintyabstractLarge Language Models (LLMs) are increasingly being explored across a range of reasoning tasks. However, LLMs sometimes struggle with reasoning tasks under uncertainty that are relatively easy for humans, such as planning actions in stochastic environments. The adoption of LLMs for reasoning is impeded by uncertainty challenges, such as LLM uncertainty and environmental uncertainty. LLM uncertainty arises from the stochastic sampling process inherent to LLMs. Most LLM-based Decision-Making (LDM) approaches address LLM uncertainty through multiple reasoning chains or search trees. However, these approaches overlook environmental uncertainty, which leads to poor performance in environments with stochastic state transitions.
Some recent LDM approaches deal with uncertainty by forecasting the probability of unknown variables. However, they are not designed for multi-step reasoning tasks that require interaction with the environment. To address uncertainty in LLM decision-making, we introduce PlanU, an LLM-based planning method that captures uncertainty within Monte Carlo Tree Search (MCTS). PlanU models the return of each node in the MCTS as a quantile distribution, which uses a set of quantiles to represent the return distribution. To balance exploration and exploitation during tree search, PlanU introduces an Upper Confidence Bounds with Curiosity (UCC) score which estimates the uncertainty of MCTS nodes. Through extensive experiments, we demonstrate the effectiveness of PlanU in LLM-based reasoning tasks under uncertainty. Ziwei Deng, Mian Deng, Chenjing Liang, Zeming Gao, Chennan Ma, Chenxing Lin, Songzhu Mei, Cheng Wang 0003 |
NeurIPS | 2 |
| 2024 | The Dormant Neuron Phenomenon in Multi-Agent Reinforcement Learning Value FactorizationabstractIn this work, we study the dormant neuron phenomenon in multi-agent reinforcement learning value factorization, where the mixing network suffers from reduced network expressivity caused by an increasing number of inactive neurons. We demonstrate the presence of the dormant neuron phenomenon across multiple environments and algorithms, and show that this phenomenon negatively affects the learning process. We show that dormant neurons correlates with the existence of over-active neurons, which have large activation scores. To address the dormant neuron issue, we propose ReBorn, a simple but effective method that transfers the weights from over-active neurons to dormant neurons. We theoretically show that this method can ensure the learned action preferences are not forgotten after the weight-transferring procedure, which increases learning effectiveness. Our extensive experiments reveal that ReBorn achieves promising results across various environments and improves the performance of multiple popular value factorization approaches. The source code of ReBorn is available in \url{https://github.com/xmu-rl-3dv/ReBorn}. Haoyuan Qin, Chennan Ma, Mian Deng, Zhengzhu Liu, Songzhu Mei, Cheng Wang 0003 |
NeurIPS | 3 |
| 2024 | Melting Pot Contest: Charting the Future of Generalized Cooperative IntelligenceabstractMulti-agent AI research promises a path to develop human-like and human-compatible intelligent technologies that complement the solipsistic view of other approaches, which mostly do not consider interactions between agents. Aiming to make progress in this direction, the Melting Pot contest 2023 focused on the problem of cooperation among interacting agents and challenged researchers to push the boundaries of multi-agent reinforcement learning (MARL) for mixed-motive games. The contest leveraged the Melting Pot environment suite to rigorously evaluate how well agents can adapt their cooperative skills to interact with novel partners in unforeseen situations. Unlike other reinforcement learning challenges, this challenge focused on social rather than environmental generalization. In particular, a population of agents performs well in Melting Pot when its component individuals are adept at finding ways to cooperate both with others in their population and with strangers. Thus Melting Pot measures cooperative intelligence.The contest attracted over 600 participants across 100+ teams globally and was a success on multiple fronts: (i) it contributed to our goal of pushing the frontiers of MARL towards building more cooperatively intelligent agents, evidenced by several submissions that outperformed established baselines; (ii) it attracted a diverse range of participants, from independent researchers to industry affiliates and academic labs, both with strong background and new interest in the area alike, broadening the field’s demographic and intellectual diversity; and (iii) analyzing the submitted agents provided important insights, highlighting areas for improvement in evaluating agents' cooperative intelligence. This paper summarizes the design aspects and results of the contest and explores the potential of Melting Pot as a benchmark for studying Cooperative AI. We further analyze the top solutions and conclude with a discussion on promising directions for future research. Rakshit S. Trivedi, Akbir Khan, Jesse Clifton, Lewis Hammond, Edgar A. Duéñez-Guzmán, Dipam Chakraborty, John P. Agapiou, Jayd Matyas, Alexander Vezhnevets, Barna Pásztor, Yunke Ao, Omar G. Younis, Benjamin Swain, Haoyuan Qin, Mian Deng, Ziwei Deng, Utku Erdoganaras, Yue Zhao 0023, Marko Tesic, Natasha Jaques, Jakob N. Foerster, Vincent Conitzer, José Hernández-Orallo, Dylan Hadfield-Menell, Joel Z. Leibo |
NeurIPS | 16 |