Bin Zhang 0052

dblp:13/5236-52 · DBLP profile ↗
← Back
23ranked-venue papers
4as first author
23since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 4 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Graph of Verification: Structured Verification of LLM Reasoning with Directed Acyclic Graphs
abstract
Verifying the complex and multi-step reasoning of Large Language Models (LLMs) is a critical challenge, as holistic methods often overlook localized flaws. Step-by-step validation is a promising alternative, yet existing methods are often rigid. They struggle to adapt to diverse reasoning structures, from formal proofs to informal natural language narratives. To address this adaptability gap, we propose the Graph of Verification (GoV), a novel framework for adaptable and multi-granular verification. GoV's core innovation is its flexible node block architecture. This mechanism allows GoV to adaptively adjust its verification granularity—from atomic steps for formal tasks to entire paragraphs for natural language—to match the native structure of the reasoning process. This flexibility allows GoV to resolve the fundamental trade-off between verification precision and robustness. Experiments on both well-structured and loosely-structured benchmarks demonstrate GoV's versatility. The results show that GoV's adaptive approach significantly outperforms both holistic baselines and other state-of-the-art decomposition-based methods, establishing a new standard for training-free reasoning verification.
Jiwei Fang, Bin Zhang 0052, Changwei Wang 0001, Jin Wan, Zhiwei Xu 0005
AAAI2
2025 Efficient Communication in Multi-Agent Reinforcement Learning with Implicit Consensus Generation
abstract
A key challenge in multi-agent collaborative tasks is reducing uncertainty about teammates to enhance cooperative performance. Explicit communication methods can reduce uncertainty about teammates, but the associated high communication costs limit their practicality. Alternatively, implicit consensus learning can promote cooperation without incurring communication costs. However, its performance declines significantly when local observations are severely limited. This paper introduces a novel multi-agent learning framework that combines the strengths of these methods. In our framework, agents generate a consensus about the group based on their local observations and then use both the consensus and local observations to produce messages. Since the consensus provides a certain level of global guidance, communication can be disabled when not essential, thereby reducing overhead. Meanwhile, communication can provide supplementary information to the consensus when necessary. Experimental results demonstrate that our algorithm significantly reduces inter-agent communication overhead while ensuring efficient collaboration.
Dapeng Li 0001, Na Lou, Zhiwei Xu 0005, Bin Zhang 0052
AAAI4
2025 PET-SQL: A Prompt-Enhanced Two-Round Refinement of Text-to-SQL with Cross-Consistency
Zhishuai Li, Xiang Wang 0012, Sun Yang, Guoqing Du, Xiaoru Hu, Bin Zhang 0052, Yuxiao Ye, Ziyue Li 0002, Hangyu Mao, Rui Zhao 0001
DASFAA (2)7
2025 Reidentify: Context-Aware Identity Generation for Contextual Multi-Agent Reinforcement Learning
abstract
Generalizing multi-agent reinforcement learning (MARL) to accommodate variations in problem configurations remains a critical challenge in real-world applications, where even subtle differences in task setups can cause pre-trained policies to fail. To address this, we propose Context-Aware Identity Generation (CAID), a novel framework to enhance MARL performance under the Contextual MARL (CMARL) setting. CAID dynamically generates unique agent identities through the agent identity decoder built on a causal Transformer architecture. These identities provide contextualized representations that align corresponding agents across similar problem variants, facilitating policy reuse and improving sample efficiency. Furthermore, the action regulator in CAID incorporates these agent identities into the action-value space, enabling seamless adaptation to varying contexts. Extensive experiments on CMARL benchmarks demonstrate that CAID significantly outperforms existing approaches by enhancing both sample efficiency and generalization across diverse context variants.
Zhiwei Xu 0005, Xin Xin 0003, Weiliang Meng, Yiwei Shi, Hangyu Mao, Bin Zhang 0052, Dapeng Li 0001, Jiangjin Yin
ICML7
2025 Unveiling Decision Intention for Cooperative Multi-Agent Reinforcement Learning
Zeren Zhang, Zhiwei Xu 0005, Guangchong Zhou, Dapeng Li 0001, Bin Zhang 0052
AAMAS5
2024 Adaptive Parameter Sharing for Multi-Agent Reinforcement Learning
abstract
Parameter sharing, as an important technique in multi-agent systems, can effectively solve the scalability issue in large-scale agent problems. However, the effectiveness of parameter sharing largely depends on the environment setting. When agents have different identities or tasks, naive parameter sharing makes it difficult to generate sufficiently differentiated strategies for agents. Inspired by research pertaining to the brain in biology, we propose a novel parameter sharing method. It maps each type of agent to different regions within a shared network based on their identity, resulting in distinct subnetworks. Therefore, our method can increase the diversity of strategies among different agents without introducing additional training parameters. Through experiments conducted in multiple environments, our method has shown better performance than other parameter sharing methods.
Dapeng Li 0001, Na Lou, Bin Zhang 0052, Zhiwei Xu 0005
ICASSP3
2024 Sequential Asynchronous Action Coordination in Multi-Agent Systems: A Stackelberg Decision Transformer Approach
abstract
Asynchronous action coordination presents a pervasive challenge in Multi-Agent Systems (MAS), which can be represented as a Stackelberg game (SG). However, the scalability of existing Multi-Agent Reinforcement Learning (MARL) methods based on SG is severely restricted by network architectures or environmental settings. To address this issue, we propose the Stackelberg Decision Transformer (STEER). It efficiently manages decision-making processes by incorporating the hierarchical decision structure of SG, the modeling capability of autoregressive sequence models, and the exploratory learning methodology of MARL. Our approach exhibits broad applicability across diverse task types and environmental configurations in MAS. Experimental results demonstrate both the convergence of our method towards Stackelberg equilibrium strategies and its superiority over strong baselines in complex scenarios.
Bin Zhang 0052, Hangyu Mao, Lijuan Li 0002, Zhiwei Xu 0005, Dapeng Li 0001, Rui Zhao 0001
ICML1
2024 GATE: Guided Contrastive State Space for Multi-agent Reinforcement Learning
Hao Chen 0103, Bin Zhang 0052
ICONIP (4)2
2024 Decentralized Extension for Centralized Multi-Agent Reinforcement Learning via Online Distillation
Zeren Zhang, Bin Zhang 0052, Guangchong Zhou, Dapeng Li 0001, Zhiwei Xu 0005
ICONIP (3)2
2024 PTDE: Personalized Training with Distilled Execution for Multi-Agent Reinforcement Learning
Yiqun Chen 0004, Hangyu Mao, Jiaxin Mao, Shiguang Wu 0001, Bin Zhang 0052, Wei Yang 0041, Hongxing Chang
IJCAI6
2024 SGCD: Subgroup Contribution Decomposition for Multi-Agent Reinforcement Learning
abstract
Cooperative multi-agent reinforcement learning (MARL) tasks rely on the efficient coordination among agents, working collectively as a team to address diverse challenges. However, considering the team as a cohesive entity introduces a flat structure to cooperation. In contrast, grouping serves as a method to tackle the issue by decomposing the team, thereby providing a more compact representation of the team’s structure. While grouping has been proven effective, numerous grouping methods are limited to specific composition structures and struggle to introduce diverse group patterns into the framework. In this paper, we propose SGCD, a subgroup contribution decomposition method that incorporates the idea of subgroups and inner subgroups, leveraging the Shapley Value to distribute contributions. This approach facilitates the decomposition of contributions from subgroups to the collective onto individual agents, enabling the high-level network to maintain consistency across various grouping patterns, thereby fostering continued cooperation among agents. Notably, our decomposition method is not confined to a specific team decomposition, making it adaptable to different grouping structures. The effectiveness of SGCD is demonstrated through experiments conducted in the Google Research Football (GRF) and StarCraft Multi-Agent Challenge (SMAC) environments.
Hao Chen 0103, Bin Zhang 0052
IJCNN2
2023 Consensus Learning for Cooperative Multi-Agent Reinforcement Learning
abstract
Almost all multi-agent reinforcement learning algorithms without communication follow the principle of centralized training with decentralized execution. During the centralized training, agents can be guided by the same signals, such as the global state. However, agents lack the shared signal and choose actions given local observations during execution. Inspired by viewpoint invariance and contrastive learning, we propose consensus learning for cooperative multi-agent reinforcement learning in this study. Although based on local observations, different agents can infer the same consensus in discrete spaces without communication. We feed the inferred one-hot consensus to the network of agents as an explicit input in a decentralized way, thereby fostering their cooperative spirit. With minor model modifications, our suggested framework can be extended to a variety of multi-agent reinforcement learning algorithms. Moreover, we carry out these variants on some fully cooperative tasks and get convincing results.
Zhiwei Xu 0005, Bin Zhang 0052, Dapeng Li 0001, Zeren Zhang, Guangchong Zhou, Hao Chen 0103
AAAI2
2023 HAVEN: Hierarchical Cooperative Multi-Agent Reinforcement Learning with Dual Coordination Mechanism
abstract
Recently, some challenging tasks in multi-agent systems have been solved by some hierarchical reinforcement learning methods. Inspired by the intra-level and inter-level coordination in the human nervous system, we propose a novel value decomposition framework HAVEN based on hierarchical reinforcement learning for fully cooperative multi-agent problems. To address the instability arising from the concurrent optimization of policies between various levels and agents, we introduce the dual coordination mechanism of inter-level and inter-agent strategies by designing reward functions in a two-level hierarchy. HAVEN does not require domain knowledge and pre-training, and can be applied to any value decomposition variant. Our method achieves desirable results on different decentralized partially observable Markov decision process domains and outperforms other popular multi-agent hierarchical reinforcement learning algorithms.
Zhiwei Xu 0005, Yunpeng Bai, Bin Zhang 0052, Dapeng Li 0001
AAAI3
2023 Inducing Stackelberg Equilibrium through Spatio-Temporal Sequential Decision-Making in Multi-Agent Reinforcement Learning
abstract
In multi-agent reinforcement learning (MARL), self-interested agents attempt to establish equilibrium and achieve coordination depending on game structure. However, existing MARL approaches are mostly bound by the simultaneous actions of all agents in the Markov game (MG) framework, and few works consider the formation of equilibrium strategies via asynchronous action coordination. In view of the advantages of Stackelberg equilibrium (SE) over Nash equilibrium, we construct a spatio-temporal sequential decision-making structure derived from the MG and propose an N-level policy model based on a conditional hypernetwork shared by all agents. This approach allows for asymmetric training with symmetric execution, with each agent responding optimally conditioned on the decisions made by superior agents. Agents can learn heterogeneous SE policies while still maintaining parameter sharing, which leads to reduced cost for learning and storage and enhanced scalability as the number of agents increases. Experiments demonstrate that our method effectively converges to the SE policies in repeated matrix game scenarios, and performs admirably in immensely complex settings including cooperative tasks and mixed tasks.
Bin Zhang 0052, Lijuan Li 0002, Zhiwei Xu 0005, Dapeng Li 0001
IJCAI1
2023 SEA: A Spatially Explicit Architecture for Multi-Agent Reinforcement Learning
abstract
Spatial information is essential in various fields. How to explicitly model according to the spatial location of agents is also very important for the multi-agent problem, especially when the number of agents is changing and the scale is enormous. Inspired by the point cloud task in computer vision, we propose a spatial information extraction structure for multi-agent reinforcement learning in this paper. Agents can effectively share the neighborhood and global information through a spatially encoder-decoder structure. Our method follows the centralized training with decentralized execution (CTDE) paradigm. In addition, our structure can be applied to various existing mainstream reinforcement learning algorithms with minor modifications and can deal with the problem with a variable number of agents. The experiments in several multi-agent scenarios show that the existing methods can get convincing results by adding our spatially explicit architecture.
Dapeng Li 0001, Zhiwei Xu 0005, Bin Zhang 0052
IJCNN3
2023 Dual Self-Awareness Value Decomposition Framework without Individual Global Max for Cooperative MARL
abstract
Value decomposition methods have gained popularity in the field of cooperative multi-agent reinforcement learning. However, almost all existing methods follow the principle of Individual Global Max (IGM) or its variants, which limits their problem-solving capabilities. To address this, we propose a dual self-awareness value decomposition framework, inspired by the notion of dual self-awareness in psychology, that entirely rejects the IGM premise. Each agent consists of an ego policy for action selection and an alter ego value function to solve the credit assignment problem. The value function factorization can ignore the IGM assumption by utilizing an explicit search procedure. On the basis of the above, we also suggest a novel anti-ego exploration mechanism to avoid the algorithm becoming stuck in a local optimum. As the first fully IGM-free value decomposition method, our proposed framework achieves desirable performance in various cooperative tasks.
Zhiwei Xu 0005, Bin Zhang 0052, Dapeng Li 0001, Guangchong Zhou, Zeren Zhang
NeurIPS2
2022 Multi-Agent Hyper-Attention Policy Optimization
Bin Zhang 0052, Zhiwei Xu 0005, Yiqun Chen 0004, Dapeng Li 0001, Yunpeng Bai, Lijuan Li 0002
ICONIP (1)1
2022 Efficient Policy Generation in Multi-agent Systems via Hypergraph Neural Network
Bin Zhang 0052, Yunpeng Bai, Zhiwei Xu 0005, Dapeng Li 0001
ICONIP (2)1
2022 Cooperative Multi-Agent Reinforcement Learning with Hypergraph Convolution
abstract
Recent years have witnessed the great success of multi-agent systems (MAS). Value decomposition, which decom-poses joint action values into individual action values, has been an important work in MAS. However, many value decomposition methods ignore the coordination among different agents, leading to the notorious “lazy agents” problem. To enhance the coordination in MAS, this paper proposes HyperGraph CoNvo-lution MIX (HGCN-MIX), a method that incorporates hyper-graph convolution with value decomposition. HGCN-MIX models agents as well as their relationships as a hypergraph, where agents are nodes and hyperedges among nodes indicate that the corresponding agents can coordinate to achieve larger rewards. Then, it trains a hypergraph that can capture the collaborative relationships among agents. Leveraging the learned hypergraph to consider how other agents' observations and actions affect their decisions, the agents in a MAS can better coordinate. We evaluate HGCN-MIX in the StarCraft II multi-agent challenge benchmark. The experimental results demonstrate that HGCN-MIX can train joint policies that outperform or achieve a similar level of performance as the current state-of-the-art techniques. We also observe that HGCN-MIX has an even more significant improvement of performance in the scenarios with a large amount of agents. Besides, we conduct additional analysis to emphasize that when the hypergraph learns more relationships, HGCN-MIX can train stronger joint policies.
Yunpeng Bai, Chen Gong 0005, Bin Zhang 0052, Xinwen Hou, Yu Liu 0078
IJCNN3
2022 Mingling Foresight with Imagination: Model-Based Cooperative Multi-Agent Reinforcement Learning
abstract
Recently, model-based agents have achieved better performance than model-free ones using the same computational budget and training time in single-agent environments. However, due to the complexity of multi-agent systems, it is tough to learn the model of the environment. The significant compounding error may hinder the learning process when model-based methods are applied to multi-agent tasks. This paper proposes an implicit model-based multi-agent reinforcement learning method based on value decomposition methods. Under this method, agents can interact with the learned virtual environment and evaluate the current state value according to imagined future states in the latent space, making agents have the foresight. Our approach can be applied to any multi-agent value decomposition method. The experimental results show that our method improves the sample efficiency in different partially observable Markov decision process domains.
Zhiwei Xu 0005, Dapeng Li 0001, Bin Zhang 0052, Yuan Zhan, Yunpeng Bai
NeurIPS3
2022 MMNet: A multi-scale deep learning network for the left ventricular segmentation of cardiac MRI images
Yanjun Peng, Dapeng Li 0001, Yanfei Guo, Bin Zhang 0052
Appl. Intell.5
2022 DSLN: Dual-tutor student learning network for multiracial glaucoma detection
Yanfei Guo, Yanjun Peng, Jindong Sun, Dapeng Li 0001, Bin Zhang 0052
Neural Comput. Appl.5
2021 Learning to Coordinate via Multiple Graph Neural Networks
Zhiwei Xu 0005, Bin Zhang 0052, Yunpeng Bai, Dapeng Li 0001
ICONIP (3)2