EDBT 2026 Demo / reviewers in the wild / expert
Zhiwei Xu 0005
dblp:262/0620-5
· DBLP profile ↗
24ranked-venue papers
7as first author
24since 2021 · last 2026
0000-0002-0754-5295ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 7 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Graph of Verification: Structured Verification of LLM Reasoning with Directed Acyclic GraphsabstractVerifying the complex and multi-step reasoning of Large Language Models (LLMs) is a critical challenge, as holistic methods often overlook localized flaws. Step-by-step validation is a promising alternative, yet existing methods are often rigid. They struggle to adapt to diverse reasoning structures, from formal proofs to informal natural language narratives. To address this adaptability gap, we propose the Graph of Verification (GoV), a novel framework for adaptable and multi-granular verification. GoV's core innovation is its flexible node block architecture. This mechanism allows GoV to adaptively adjust its verification granularity—from atomic steps for formal tasks to entire paragraphs for natural language—to match the native structure of the reasoning process. This flexibility allows GoV to resolve the fundamental trade-off between verification precision and robustness. Experiments on both well-structured and loosely-structured benchmarks demonstrate GoV's versatility. The results show that GoV's adaptive approach significantly outperforms both holistic baselines and other state-of-the-art decomposition-based methods, establishing a new standard for training-free reasoning verification. Jiwei Fang, Bin Zhang 0052, Changwei Wang 0001, Jin Wan, Zhiwei Xu 0005 |
AAAI | 5 |
| 2026 | Reinforced Efficient Reasoning via Semantically Diverse ExplorationabstractZiqi Zhao, Zhaochun Ren, Jiahong Zou, Liu Yang, Zhiwei Xu, Xuri Ge, Zhumin Chen, Xinyu Ma, Daiting Shi, Shuaiqiang Wang, Dawei Yin, Xin Xin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhaochun Ren, Jiahong Zou, Liu Yang 0025, Zhiwei Xu 0005, Xuri Ge, Zhumin Chen, Xinyu Ma 0001, Daiting Shi, Shuaiqiang Wang, Dawei Yin 0001, Xin Xin 0003 |
ACL (1) | 5 |
| 2025 | Focus on Local: Finding Reliable Discriminative Regions for Visual Place RecognitionabstractVisual Place Recognition (VPR) is aimed at predicting the location of a query image by referencing a database of geotagged images. For VPR task, often fewer discriminative local regions in an image produce important effects while mundane background regions do not contribute or even cause perceptual aliasing because of easy overlap. However, existing methods lack precisely modeling and full exploitation of these discriminative regions. In addition, the lack of pixel-level correspondence supervision in the VPR dataset hinders further improvement of the local feature matching capability in the re-ranking stage. In this paper, we propose the Focus on Local (FoL) approach to stimulate the performance of image retrieval and re-ranking in VPR simultaneously by mining and exploiting reliable discriminative local regions in images and introducing pseudo-correlation supervision. First, we design two losses, Extraction-Aggregation Spatial Alignment Loss (SAL) and Foreground-Background Contrast Enhancement Loss (CEL), to explicitly model reliable discriminative local regions and use them to guide the generation of global representations and efficient re-ranking. Second, we introduce a weakly-supervised local feature training strategy based on pseudo-correspondences obtained from aggregating global features to alleviate the lack of local correspondences ground truth for the VPR task. Third, we suggest an efficient re-ranking pipeline that is efficiently and precisely based on discriminative region guidance. Finally, experimental results show that our FoL achieves the state-of-the-art on multiple VPR benchmarks in both image retrieval and re-ranking stages and also significantly outperforms existing two-stage VPR methods in terms of computational efficiency. Changwei Wang 0001, Shunpeng Chen, Rongtao Xu, Jiguang Zhang, Haoran Yang 0003, Yu Zhang 0133, Kexue Fu 0001, Shide Du, Zhiwei Xu 0005, Longxiang Gao, Li Guo 0004, Shibiao Xu |
AAAI | 11 |
| 2025 | Efficient Communication in Multi-Agent Reinforcement Learning with Implicit Consensus GenerationabstractA key challenge in multi-agent collaborative tasks is reducing uncertainty about teammates to enhance cooperative performance. Explicit communication methods can reduce uncertainty about teammates, but the associated high communication costs limit their practicality. Alternatively, implicit consensus learning can promote cooperation without incurring communication costs. However, its performance declines significantly when local observations are severely limited. This paper introduces a novel multi-agent learning framework that combines the strengths of these methods. In our framework, agents generate a consensus about the group based on their local observations and then use both the consensus and local observations to produce messages. Since the consensus provides a certain level of global guidance, communication can be disabled when not essential, thereby reducing overhead. Meanwhile, communication can provide supplementary information to the consensus when necessary. Experimental results demonstrate that our algorithm significantly reduces inter-agent communication overhead while ensuring efficient collaboration. Dapeng Li 0001, Na Lou, Zhiwei Xu 0005, Bin Zhang 0052 |
AAAI | 3 |
| 2025 | Reidentify: Context-Aware Identity Generation for Contextual Multi-Agent Reinforcement LearningabstractGeneralizing multi-agent reinforcement learning (MARL) to accommodate variations in problem configurations remains a critical challenge in real-world applications, where even subtle differences in task setups can cause pre-trained policies to fail. To address this, we propose Context-Aware Identity Generation (CAID), a novel framework to enhance MARL performance under the Contextual MARL (CMARL) setting. CAID dynamically generates unique agent identities through the agent identity decoder built on a causal Transformer architecture. These identities provide contextualized representations that align corresponding agents across similar problem variants, facilitating policy reuse and improving sample efficiency. Furthermore, the action regulator in CAID incorporates these agent identities into the action-value space, enabling seamless adaptation to varying contexts. Extensive experiments on CMARL benchmarks demonstrate that CAID significantly outperforms existing approaches by enhancing both sample efficiency and generalization across diverse context variants. Zhiwei Xu 0005, Xin Xin 0003, Weiliang Meng, Yiwei Shi, Hangyu Mao, Bin Zhang 0052, Dapeng Li 0001, Jiangjin Yin |
ICML | 1 |
| 2025 | Unveiling Decision Intention for Cooperative Multi-Agent Reinforcement Learning
Zeren Zhang, Zhiwei Xu 0005, Guangchong Zhou, Dapeng Li 0001, Bin Zhang 0052 |
AAMAS | 2 |
| 2025 | Belief-Calibrated Multi-Agent Consensus Seeking for Complex NLP TasksabstractA multi-agent system (MAS) enhances its capacity to solve complex natural language processing (NLP) tasks through collaboration among multiple agents, where consensus-seeking serves as a fundamental mechanism.
However, existing consensus-seeking approaches typically rely on voting mechanisms to judge consensus, overlooking contradictions in system-internal beliefs that destabilize the consensus.
Moreover, these methods often involve agents updating their results through indiscriminate collaboration with every other agent.
Such uniform interaction fails to identify the optimal collaborators for each agent, hindering the emergence of a stable consensus.
To address these challenges, we provide a theoretical framework for selecting optimal collaborators that maximize consensus stability.
Based on the theorems, we propose the Belief-Calibrated Consensus Seeking (BCCS) framework to facilitate stable consensus via selecting optimal collaborators and calibrating the consensus judgment by system-internal beliefs.
Experimental results on the MATH and MMLU benchmark datasets demonstrate that the proposed BCCS framework outperforms the best existing results by 2.23\% and 3.95\% of accuracy on challenging tasks, respectively.
Our code and data are available at https://github.com/dengwentao99/BCCS. Wentao Deng, Jiahuan Pei, Zhiwei Xu 0005, Zhaochun Ren, Zhumin Chen, Pengjie Ren |
NeurIPS | 3 |
| 2024 | Adaptive Parameter Sharing for Multi-Agent Reinforcement LearningabstractParameter sharing, as an important technique in multi-agent systems, can effectively solve the scalability issue in large-scale agent problems. However, the effectiveness of parameter sharing largely depends on the environment setting. When agents have different identities or tasks, naive parameter sharing makes it difficult to generate sufficiently differentiated strategies for agents. Inspired by research pertaining to the brain in biology, we propose a novel parameter sharing method. It maps each type of agent to different regions within a shared network based on their identity, resulting in distinct subnetworks. Therefore, our method can increase the diversity of strategies among different agents without introducing additional training parameters. Through experiments conducted in multiple environments, our method has shown better performance than other parameter sharing methods. Dapeng Li 0001, Na Lou, Bin Zhang 0052, Zhiwei Xu 0005 |
ICASSP | 4 |
| 2024 | Sequential Asynchronous Action Coordination in Multi-Agent Systems: A Stackelberg Decision Transformer ApproachabstractAsynchronous action coordination presents a pervasive challenge in Multi-Agent Systems (MAS), which can be represented as a Stackelberg game (SG). However, the scalability of existing Multi-Agent Reinforcement Learning (MARL) methods based on SG is severely restricted by network architectures or environmental settings. To address this issue, we propose the Stackelberg Decision Transformer (STEER). It efficiently manages decision-making processes by incorporating the hierarchical decision structure of SG, the modeling capability of autoregressive sequence models, and the exploratory learning methodology of MARL. Our approach exhibits broad applicability across diverse task types and environmental configurations in MAS. Experimental results demonstrate both the convergence of our method towards Stackelberg equilibrium strategies and its superiority over strong baselines in complex scenarios. Bin Zhang 0052, Hangyu Mao, Lijuan Li 0002, Zhiwei Xu 0005, Dapeng Li 0001, Rui Zhao 0001 |
ICML | 4 |
| 2024 | Style Miner: Find Significant and Stable Factors in Time Series with Constrained Reinforcement Learning
Dapeng Li 0001, Feiyang Pan, Jia He 0001, Zhiwei Xu 0005, Dandan Tu |
ICONIP (6) | 4 |
| 2024 | Decentralized Extension for Centralized Multi-Agent Reinforcement Learning via Online Distillation
Zeren Zhang, Bin Zhang 0052, Guangchong Zhou, Dapeng Li 0001, Zhiwei Xu 0005 |
ICONIP (3) | 5 |
| 2023 | Consensus Learning for Cooperative Multi-Agent Reinforcement LearningabstractAlmost all multi-agent reinforcement learning algorithms without communication follow the principle of centralized training with decentralized execution. During the centralized training, agents can be guided by the same signals, such as the global state. However, agents lack the shared signal and choose actions given local observations during execution. Inspired by viewpoint invariance and contrastive learning, we propose consensus learning for cooperative multi-agent reinforcement learning in this study. Although based on local observations, different agents can infer the same consensus in discrete spaces without communication. We feed the inferred one-hot consensus to the network of agents as an explicit input in a decentralized way, thereby fostering their cooperative spirit. With minor model modifications, our suggested framework can be extended to a variety of multi-agent reinforcement learning algorithms. Moreover, we carry out these variants on some fully cooperative tasks and get convincing results. Zhiwei Xu 0005, Bin Zhang 0052, Dapeng Li 0001, Zeren Zhang, Guangchong Zhou, Hao Chen 0103 |
AAAI | 1 |
| 2023 | HAVEN: Hierarchical Cooperative Multi-Agent Reinforcement Learning with Dual Coordination MechanismabstractRecently, some challenging tasks in multi-agent systems have been solved by some hierarchical reinforcement learning methods. Inspired by the intra-level and inter-level coordination in the human nervous system, we propose a novel value decomposition framework HAVEN based on hierarchical reinforcement learning for fully cooperative multi-agent problems. To address the instability arising from the concurrent optimization of policies between various levels and agents, we introduce the dual coordination mechanism of inter-level and inter-agent strategies by designing reward functions in a two-level hierarchy. HAVEN does not require domain knowledge and pre-training, and can be applied to any value decomposition variant. Our method achieves desirable results on different decentralized partially observable Markov decision process domains and outperforms other popular multi-agent hierarchical reinforcement learning algorithms. Zhiwei Xu 0005, Yunpeng Bai, Bin Zhang 0052, Dapeng Li 0001 |
AAAI | 1 |
| 2023 | Mastering Complex Coordination Through Attention-Based Dynamic Graph
Guangchong Zhou, Zhiwei Xu 0005, Zeren Zhang |
ICONIP (1) | 2 |
| 2023 | SORA: Improving Multi-agent Cooperation with a Soft Role Assignment Mechanism
Guangchong Zhou, Zhiwei Xu 0005, Zeren Zhang |
ICONIP (1) | 2 |
| 2023 | Inducing Stackelberg Equilibrium through Spatio-Temporal Sequential Decision-Making in Multi-Agent Reinforcement LearningabstractIn multi-agent reinforcement learning (MARL), self-interested agents attempt to establish equilibrium and achieve coordination depending on game structure. However, existing MARL approaches are mostly bound by the simultaneous actions of all agents in the Markov game (MG) framework, and few works consider the formation of equilibrium strategies via asynchronous action coordination. In view of the advantages of Stackelberg equilibrium (SE) over Nash equilibrium, we construct a spatio-temporal sequential decision-making structure derived from the MG and propose an N-level policy model based on a conditional hypernetwork shared by all agents. This approach allows for asymmetric training with symmetric execution, with each agent responding optimally conditioned on the decisions made by superior agents. Agents can learn heterogeneous SE policies while still maintaining parameter sharing, which leads to reduced cost for learning and storage and enhanced scalability as the number of agents increases. Experiments demonstrate that our method effectively converges to the SE policies in repeated matrix game scenarios, and performs admirably in immensely complex settings including cooperative tasks and mixed tasks. Bin Zhang 0052, Lijuan Li 0002, Zhiwei Xu 0005, Dapeng Li 0001 |
IJCAI | 3 |
| 2023 | SEA: A Spatially Explicit Architecture for Multi-Agent Reinforcement LearningabstractSpatial information is essential in various fields. How to explicitly model according to the spatial location of agents is also very important for the multi-agent problem, especially when the number of agents is changing and the scale is enormous. Inspired by the point cloud task in computer vision, we propose a spatial information extraction structure for multi-agent reinforcement learning in this paper. Agents can effectively share the neighborhood and global information through a spatially encoder-decoder structure. Our method follows the centralized training with decentralized execution (CTDE) paradigm. In addition, our structure can be applied to various existing mainstream reinforcement learning algorithms with minor modifications and can deal with the problem with a variable number of agents. The experiments in several multi-agent scenarios show that the existing methods can get convincing results by adding our spatially explicit architecture. Dapeng Li 0001, Zhiwei Xu 0005, Bin Zhang 0052 |
IJCNN | 2 |
| 2023 | Dual Self-Awareness Value Decomposition Framework without Individual Global Max for Cooperative MARLabstractValue decomposition methods have gained popularity in the field of cooperative multi-agent reinforcement learning. However, almost all existing methods follow the principle of Individual Global Max (IGM) or its variants, which limits their problem-solving capabilities. To address this, we propose a dual self-awareness value decomposition framework, inspired by the notion of dual self-awareness in psychology, that entirely rejects the IGM premise. Each agent consists of an ego policy for action selection and an alter ego value function to solve the credit assignment problem. The value function factorization can ignore the IGM assumption by utilizing an explicit search procedure. On the basis of the above, we also suggest a novel anti-ego exploration mechanism to avoid the algorithm becoming stuck in a local optimum. As the first fully IGM-free value decomposition method, our proposed framework achieves desirable performance in various cooperative tasks. Zhiwei Xu 0005, Bin Zhang 0052, Dapeng Li 0001, Guangchong Zhou, Zeren Zhang |
NeurIPS | 1 |
| 2022 | Learn Effective Representation for Deep Reinforcement LearningabstractRecent years have witnessed an increasing application of deep reinforcement learning (DRL) on video games. While deeper and wider neural networks have played a crucial role in computer vision and natural language processing, such capacity remain under-explored in most DRL works. Under the fact that feature propagation together with large networks contributes to learning a good representation, we propose an end-to-end Large Feature Extractor Network (LFENet) that uses large neural networks with dense connections to train a high-capacity encoder. Even though the increased dimensionality of input is usually thought to result in poor performance for RL agents, we introduce the information bottleneck to alleviate the problem. Finally, we combine LFENet with Proximal Policy Optimization (PPO) algorithm. Through numerical experiments on Atari 2600 video games, we demonstrate our method matches or outperforms state-of-the-art algorithms. Yuan Zhan, Zhiwei Xu 0005 |
ICME | 2 |
| 2022 | Multi-Agent Hyper-Attention Policy Optimization
Bin Zhang 0052, Zhiwei Xu 0005, Yiqun Chen 0004, Dapeng Li 0001, Yunpeng Bai, Lijuan Li 0002 |
ICONIP (1) | 2 |
| 2022 | Efficient Policy Generation in Multi-agent Systems via Hypergraph Neural Network
Bin Zhang 0052, Yunpeng Bai, Zhiwei Xu 0005, Dapeng Li 0001 |
ICONIP (2) | 3 |
| 2022 | Mingling Foresight with Imagination: Model-Based Cooperative Multi-Agent Reinforcement LearningabstractRecently, model-based agents have achieved better performance than model-free ones using the same computational budget and training time in single-agent environments. However, due to the complexity of multi-agent systems, it is tough to learn the model of the environment. The significant compounding error may hinder the learning process when model-based methods are applied to multi-agent tasks. This paper proposes an implicit model-based multi-agent reinforcement learning method based on value decomposition methods. Under this method, agents can interact with the learned virtual environment and evaluate the current state value according to imagined future states in the latent space, making agents have the foresight. Our approach can be applied to any multi-agent value decomposition method. The experimental results show that our method improves the sample efficiency in different partially observable Markov decision process domains. Zhiwei Xu 0005, Dapeng Li 0001, Bin Zhang 0052, Yuan Zhan, Yunpeng Bai |
NeurIPS | 1 |
| 2021 | Learning to Coordinate via Multiple Graph Neural Networks
Zhiwei Xu 0005, Bin Zhang 0052, Yunpeng Bai, Dapeng Li 0001 |
ICONIP (3) | 1 |
| 2021 | MMD-MIX: Value Function Factorisation with Maximum Mean Discrepancy for Cooperative Multi-Agent Reinforcement LearningabstractIn the real world, many tasks require multiple agents to cooperate with each other under the condition of local observations. To solve such problems, many multi-agent reinforcement learning methods based on Centralized Training with Decentralized Execution have been proposed. One representative class of work is value decomposition, which decomposes the global joint Q-value Qjtinto individual Q-values Qato guide individuals' behaviors, e.g. VDN (Value-Decomposition Networks) and QMIX. However, these baselines often ignore the randomness in the situation. We propose MMD-MIX, a method that combines distributional reinforcement learning and value decomposition to alleviate the above weaknesses. Besides, to improve data sampling efficiency, we were inspired by REM (Random Ensemble Mixture) which is a robust RL algorithm to explicitly introduce randomness into the MMD-MIX. The experiments demonstrate that MMD-MIX outperforms prior baselines in the StarCraft Multi-Agent Challenge (SMAC) environment. Zhiwei Xu 0005, Dapeng Li 0001, Yunpeng Bai |
IJCNN | 1 |