Zhuohui Zhang

dblp:356/8467 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2026
0009-0000-5040-9654ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Reinforcement learning · 64% Graph learning · 36%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › multi-agent reinforcement learning › multi-agent communication
communication policy learning
0.912025
Bridging Training and Execution via Dynamic Directed Graph-Based Communication in Cooperative Multi-Agent Systems · AAAI 2025
Machine learning › Graph learning
graph neural network
0.912025
Bridging Training and Execution via Dynamic Directed Graph-Based Communication in Cooperative Multi-Agent Systems · AAAI 2025
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.912025
Bridging Training and Execution via Dynamic Directed Graph-Based Communication in Cooperative Multi-Agent Systems · AAAI 2025
Machine learning › Reinforcement learning › multi-agent reinforcement learning › decentralized multi-agent reinforcement learning
centralized training with decentralized execution
0.312025
Bridging Training and Execution via Dynamic Directed Graph-Based Communication in Cooperative Multi-Agent Systems · AAAI 2025
Machine learning › Graph learning › graph algorithms
graph coarsening
0.312025
Bridging Training and Execution via Dynamic Directed Graph-Based Communication in Cooperative Multi-Agent Systems · AAAI 2025

Methods — techniques the papers use, named apart from their topics

transformer · 0.9multi-agent reinforcement learning · 0.9graph coarsening · 0.9
YearPublicationVenuePosition
2026 EALLMs: Environment-Aligned LLMs for Enhanced Exploration and Communication in Multi-Agent Reinforcement Learning
abstract
Leveraging large language models (LLMs) for collaborative sequential decision-making is a significant challenge, despite strong semantic understanding and extensive prior knowledge. Conversely, multi-agent reinforcement learning (MARL) can learn environment-aligned policies through interaction, but often suffers from inefficient exploration and heavy reliance on centralized global state information. To achieve complementary advantages, we propose the environment-aligned LLMs (EALLMs). In our framework, an LLM serves as a shared policy for all agents and is updated through online MARL to achieve alignment with the environment. Simultaneously, another LLM, fine-tuned with offline datasets, acts as an information integrator to generate global state for communication purposes. Additionally, we design robust, task-specific prompts tailored to multi-agent systems. Extensive experiments demonstrate that EALLMs outperform classical MARL and LLM-based baselines in both exploration efficiency and overall performance on the SMAC and SMACv2 benchmarks. Ablation studies further confirm EALLMs’ ability to achieve competitive results without relying on explicit global state, while preserving the original capabilities of the LLM during alignment.
Zhuohui Zhang, Bin Cheng 0008, Bin He 0003
IEEE Trans Autom. Sci. Eng.1
2025 Bridging Training and Execution via Dynamic Directed Graph-Based Communication in Cooperative Multi-Agent Systems
abstract
Multi-agent systems must learn to communicate and understand interactions between agents to achieve cooperative goals in partially observed tasks. However, existing approaches lack a dynamic directed communication mechanism and rely on global states, thus diminishing the role of communication in centralized training. Thus, we propose the Transformer-based graph coarsening network (TGCNet), a novel multi-agent reinforcement learning (MARL) algorithm. TGCNet learns the topological structure of a dynamic directed graph to represent the communication policy and integrates graph coarsening networks to approximate the representation of global state during training. It also utilizes the Transformer decoder for feature extraction during execution. Experiments on multiple cooperative MARL benchmarks demonstrate state-of-the-art performance compared to popular MARL algorithms. Further ablation studies validate the effectiveness of our dynamic directed graph communication mechanism and graph coarsening networks.
Zhuohui Zhang
AAAI1