Chenxing Lin

dblp:404/8393 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 36% Planning, search and constraint satisfaction · 36% Generative modeling · 9%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
multi-agent reinforcement learning
1.722025
GradPS: Resolving Futile Neurons in Parameter Sharing Network for Multi-Agent Reinforcement Learning · ICML 2025
DoF: A Diffusion Factorization Framework for Offline Multi-Agent Reinforcement Learning · ICLR 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
decision making under uncertainty
0.912025
PlanU: Large Language Model Reasoning through Planning under Uncertainty · NeurIPS 2025
Machine learning › Generative modeling
diffusion model
0.912025
DoF: A Diffusion Factorization Framework for Offline Multi-Agent Reinforcement Learning · ICLR 2025
Natural language and speech › Language models and text generation
large language model reasoning
0.912025
PlanU: Large Language Model Reasoning through Planning under Uncertainty · NeurIPS 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › language-based planning
LLM-based planning
0.912025
PlanU: Large Language Model Reasoning through Planning under Uncertainty · NeurIPS 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search
0.912025
PlanU: Large Language Model Reasoning through Planning under Uncertainty · NeurIPS 2025
Machine learning › Reinforcement learning › multi-agent reinforcement learning
offline multi-agent reinforcement learning
0.912025
DoF: A Diffusion Factorization Framework for Offline Multi-Agent Reinforcement Learning · ICLR 2025
Machine learning › Efficient and distributed learning
parameter sharing
0.912025
GradPS: Resolving Futile Neurons in Parameter Sharing Network for Multi-Agent Reinforcement Learning · ICML 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
planning under uncertainty
0.912025
PlanU: Large Language Model Reasoning through Planning under Uncertainty · NeurIPS 2025
Machine learning › Reinforcement learning
policy diversity
0.912025
GradPS: Resolving Futile Neurons in Parameter Sharing Network for Multi-Agent Reinforcement Learning · ICML 2025

Methods — techniques the papers use, named apart from their topics

upper confidence bounds with curiosity · 0.9quantile distribution · 0.9noise factorization · 0.9neuron cloning · 0.9monte carlo tree search · 0.9gradient conflict analysis · 0.9diffusion model · 0.9
YearPublicationVenuePosition
2025 DoF: A Diffusion Factorization Framework for Offline Multi-Agent Reinforcement Learning
abstract
Diffusion models have been widely adopted in image and language generation and are now being applied to reinforcement learning. However, the application of diffusion models in offline cooperative Multi-Agent Reinforcement Learning (MARL) remains limited. Although existing studies explore this direction, they suffer from scalability or poor cooperation issues due to the lack of design principles for diffusion-based MARL. The Individual-Global-Max (IGM) principle is a popular design principle for cooperative MARL. By satisfying this principle, MARL algorithms achieve remarkable performance with good scalability. In this work, we extend the IGM principle to the Individual-Global-identically-Distributed (IGD) principle. This principle stipulates that the generated outcome of a multi-agent diffusion model should be identically distributed as the collective outcomes from multiple individual-agent diffusion models. We propose DoF, a diffusion factorization framework for Offline MARL. It uses noise factorization function to factorize a centralized diffusion model into multiple diffusion models. We theoretically show that the noise factorization functions satisfy the IGD principle. Furthermore, DoF uses data factorization function to model the complex relationship among data generated by multiple diffusion models. Through extensive experiments, we demonstrate the effectiveness of DoF. The source code is available at [https://github.com/xmu-rl-3dv/DoF](https://github.com/xmu-rl-3dv/DoF).
Ziwei Deng, Chenxing Lin, Yongquan Fu, Weiquan Liu, Chenglu Wen, Cheng Wang 0003
ICLR3
2025 GradPS: Resolving Futile Neurons in Parameter Sharing Network for Multi-Agent Reinforcement Learning
abstract
Parameter-sharing (PS) techniques have been widely adopted in cooperative Multi-Agent Reinforcement Learning (MARL). In PS, all the agents share a policy network with identical parameters, which enjoys good sample efficiency. However, PS could lead to homogeneous policies that limit MARL performance. We tackle this problem from the angle of gradient conflict among agents. We find that the existence of futile neurons whose update is canceled out by gradient conflicts among agents leads to poor learning efficiency and diversity. To address this deficiency, we propose GradPS, a gradient-based PS method. It dynamically creates multiple clones for each futile neuron. For each clone, a group of agents with low gradient-conflict shares the neuron's parameters. Our method can enjoy good sample efficiency by sharing the gradients among agents of the same clone neuron. Moreover, it can encourage diverse behaviors through independently updating an exclusive clone neuron. Through extensive experiments, we show that GradPS can learn diverse policies with promising performance. The source code for GradPS is available in \url{https://github.com/xmu-rl-3dv/GradPS}.
Haoyuan Qin, Zhengzhu Liu, Chenxing Lin, Chennan Ma, Songzhu Mei, Cheng Wang 0003
ICML3
2025 PlanU: Large Language Model Reasoning through Planning under Uncertainty
abstract
Large Language Models (LLMs) are increasingly being explored across a range of reasoning tasks. However, LLMs sometimes struggle with reasoning tasks under uncertainty that are relatively easy for humans, such as planning actions in stochastic environments. The adoption of LLMs for reasoning is impeded by uncertainty challenges, such as LLM uncertainty and environmental uncertainty. LLM uncertainty arises from the stochastic sampling process inherent to LLMs. Most LLM-based Decision-Making (LDM) approaches address LLM uncertainty through multiple reasoning chains or search trees. However, these approaches overlook environmental uncertainty, which leads to poor performance in environments with stochastic state transitions. Some recent LDM approaches deal with uncertainty by forecasting the probability of unknown variables. However, they are not designed for multi-step reasoning tasks that require interaction with the environment. To address uncertainty in LLM decision-making, we introduce PlanU, an LLM-based planning method that captures uncertainty within Monte Carlo Tree Search (MCTS). PlanU models the return of each node in the MCTS as a quantile distribution, which uses a set of quantiles to represent the return distribution. To balance exploration and exploitation during tree search, PlanU introduces an Upper Confidence Bounds with Curiosity (UCC) score which estimates the uncertainty of MCTS nodes. Through extensive experiments, we demonstrate the effectiveness of PlanU in LLM-based reasoning tasks under uncertainty.
Ziwei Deng, Mian Deng, Chenjing Liang, Zeming Gao, Chennan Ma, Chenxing Lin, Songzhu Mei, Cheng Wang 0003
NeurIPS6