EDBT 2026 Demo / reviewers in the wild / expert
Cong Guan
dblp:191/7206
· DBLP profile ↗
23ranked-venue papers
6as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 4 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-agent In-context Coordination via Decentralized Memory RetrievalabstractLarge transformer models, trained on diverse datasets, have demonstrated impressive few-shot performance on previously unseen tasks without requiring parameter updates. This capability has also been explored in Reinforcement Learning (RL), where agents interact with the environment to retrieve context and maximize cumulative rewards, showcasing strong adaptability in complex settings. However, in cooperative Multi-Agent Reinforcement Learning (MARL), where agents must coordinate toward a shared goal, decentralized policy deployment can lead to mismatches in task alignment and reward assignment, limiting the efficiency of policy adaptation. To address this challenge, we introduce Multi-agent In-context Coordination via Decentralized Memory Retrieval (MAICC), a novel approach designed to enhance coordination by fast adaptation. Our method involves training a centralized embedding model to capture fine-grained trajectory representations, followed by decentralized models that approximate the centralized one to obtain team-level task information. Based on the learned embeddings, relevant trajectories are retrieved as context, which, combined with the agents' current sub-trajectories, inform decision-making. During decentralized execution, we introduce a novel memory mechanism that effectively balances test-time online data with offline memory. Based on the constructed memory, we propose a hybrid utility score that incorporates both individual- and team-level returns, ensuring credit assignment across agents. Extensive experiments on cooperative MARL benchmarks, including Level-Based Foraging (LBF) and SMAC (v1/v2), show that MAICC enables faster adaptation to unseen tasks compared to existing methods. Zichuan Lin, Lihe Li, Yi-Chen Li 0001, Cong Guan, Lei Yuan 0005, Zongzhang Zhang, Yang Yu 0001, Deheng Ye |
AAAI | 5 |
| 2026 | DM3Net: Dual-Camera Super-Resolution via Domain Modulation and Multi-scale MatchingabstractDual-camera super-resolution is highly practical for smartphone photography that primarily super-resolve the wide-angle images using the telephoto image as a reference. In this paper, we propose DM3Net, a novel dual-camera super-resolution network based on Domain Modulation and Multi-scale Matching. To bridge the domain gap between the high-resolution domain and the degraded domain, we learn two compressed global representations from image pairs corresponding to the two domains. To enable reliable transfer of high-frequency structural details from the reference image, we design a multi-scale matching module that conducts patch-level feature matching and retrieval across multiple receptive fields to improve matching accuracy and robustness. Moreover, we also introduce Key Pruning to achieve a significant reduction in memory usage and inference time with little model performance sacrificed. Experimental results on three real-world datasets demonstrate that our DM3Net outperforms the state-of-the-art approaches. Cong Guan, Jiacheng Ying, Yuya Ieiri, Osamu Yoshie |
WACV | 1 |
| 2026 | CLIP-driven rain perception: Adaptive deraining with pattern-aware network routing and mask-guided cross-attentionabstractExisting deraining models process all rainy images within a single network. However, different rain patterns have significant variations, which makes it challenging for a single network to handle diverse types of raindrops and streaks. To address this limitation, we propose a novel CLIP-driven rain perception network (CLIP-RPN) that leverages CLIP to automatically perceive rain patterns by computing visual-language matching scores and adaptively routing to sub-networks to handle different rain patterns, such as varying raindrop densities, streak orientations, and rainfall intensity. CLIP-RPN establishes semantic-aware rain pattern recognition through CLIP’s cross-modal visual-language alignment capabilities, enabling automatic identification of precipitation characteristics across different rain scenarios. This rain pattern awareness drives an adaptive subnetwork routing mechanism where specialized processing branches are dynamically activated based on the detected rain type, significantly enhancing the model’s capacity to handle diverse rainfall conditions. Furthermore, within sub-networks of CLIP-RPN, we introduce a mask-guided cross-attention mechanism (MGCA) that predicts precise rain masks at multi-scale to facilitate contextual interactions between rainy regions and clean background areas by cross-attention. We also introduces a dynamic loss scheduling mechanism (DLS) to adaptively adjust the gradients for the optimization process of CLIP-RPN. Compared with the commonly used l 1 or l 2 loss, DLS is more compatible with the inherent dynamics of the network training process, thus achieving enhanced outcomes. Our method achieves state-of-the-art performance across multiple datasets, particularly excelling in complex mixed datasets. Cong Guan, Osamu Yoshie |
Pattern Recognit. | 1 |
| 2025 | ARM : nnU-Net with Arena Mechanism for Medical Image SegmentationabstractThe success of nnU-Net proves the significance of the rationality of workflow architecture and configuration settings in improving segmentation accuracy. However, since that, most efforts to improve U-Net have continued to address CNN inner limitations caused by architecture. These methods encountered challenges such as limited generalization, difficulty in managing data with varying distribution patterns. To tackle these issues, We designed a standardized processing workflow specifically tailored for the convolutional layers of U-Net : Arena Mechanism (ARM), inspired by game theory, which encompasses three stages: Recruit, Train and Fight. We augment the convolutional layers of U-Net with two additional processing branches (Recruit), serving as "challengers." These challengers are first optimized within a Hybrid Adaptive Weighting module to refine the feature representation of each internal channel (Train). During forward propagation, we design a cooperative loss and reward function to determine the optimal confidence distribution and fusion strategy for branch outputs (Fight). This standardized mechanism is designed to further nnU-Net’s core principle of "Automatic Adaptation." Experiments have demonstrated that this approach achieves state-of-the-art (SOTA) performance on both the ACDC, BraTS21 and KiTS datasets. Cong Guan, Tengfei Shao, Shenglei Li, Tomoji Kishi, Osamu Yoshie |
ICASSP | 2 |
| 2025 | MetaCert: Metabolic Attention Network Utilizing Uncertainty Estimation for Multimodal Aspect-Category-Sentiment Triple ExtractionabstractMultimodal Aspect-Category-Sentiment Triple Extraction (MACSTE) is a highly complex subtask within Multimodal Aspect-Based Sentiment Analysis (MABSA), requiring simultaneous attribute extraction and sentiment polarity prediction from image-text pairs. While existing research often emphasizes modality fusion and alignment, it frequently neglects the design of information flow pathways, leading to suboptimal utilization of complementary information. Additionally, modality-specific noise may compromise the robustness and accuracy of multimodal classification, with traditional filtering methods often degrading data quality. To overcome these challenges, we propose the Metabolic Attention Network Utilizing Uncertainty Estimation (MetaCert). MetaCert integrates two key components: the Metabolic Attention Mechanism (MAM), inspired by bio-chemical metabolic networks and enhanced by cross-attention for improved information exchange; and the Uncertainty Estimation Network (UEN), which optimizes the semantic contributions of each modality while preserving data integrity, thereby enhancing classification accuracy. Our approach achieves state-of-the-art (SOTA) results on the TWITTER-15 and TWITTER-17 datasets. Cong Guan, Tengfei Shao, Shenglei Li, Tomoji Kishi, Osamu Yoshie |
ICASSP | 2 |
| 2025 | Efficient Multi-agent Offline Coordination via Diffusion-based Trajectory StitchingabstractLearning from offline data without interacting with the environment is a promising way to fully leverage the intelligent decision-making capabilities of multi-agent reinforcement learning (MARL). Previous approaches have primarily focused on developing learning techniques, such as conservative methods tailored to MARL using limited offline data. However, these methods often overlook the temporal relationships across different timesteps and spatial relationships between teammates, resulting in low learning efficiency in imbalanced data scenarios. To comprehensively explore the data structure of MARL and enhance learning efficiency, we propose Multi-Agent offline coordination via Diffusion-based Trajectory Stitching (MADiTS), a novel diffusion-based data augmentation pipeline that systematically generates trajectories by stitching high-quality coordination segments together. MADiTS first generates trajectory segments using a trained diffusion model, followed by applying a bidirectional dynamics constraint to ensure that the trajectories align with environmental dynamics. Additionally, we develop an offline credit assignment technique to identify and optimize the behavior of underperforming agents in the generated segments. This iterative procedure continues until a satisfactory augmented episode trajectory is generated within the predefined limit or is discarded otherwise. Empirical results on imbalanced datasets of multiple benchmarks demonstrate that MADiTS significantly improves MARL performance. Lei Yuan 0005, Yuqi Bian, Lihe Li, Cong Guan, Yang Yu 0001 |
ICLR | 5 |
| 2025 | Step-DAD: Semi-Amortized Policy-Based Bayesian Experimental DesignabstractWe develop a semi-amortized, policy-based, approach to Bayesian experimental design (BED) called Stepwise Deep Adaptive Design (Step-DAD). Like existing, fully amortized, policy-based BED approaches, Step-DAD trains a design policy upfront before the experiment. However, rather than keeping this policy fixed, Step-DAD periodically updates it as data is gathered, refining it to the particular experimental instance. This test-time adaptation improves both the flexibility and the robustness of the design strategy compared with existing approaches. Empirically, Step-DAD consistently demonstrates superior decision-making and robustness compared with current state-of-the-art BED methods. Marcel Hedman, Desi R. Ivanova, Cong Guan, Tom Rainforth |
ICML | 3 |
| 2025 | Adaptable Safe Policy Learning from Multi-task Data with Constraint Prioritized Decision TransformerabstractLearning safe reinforcement learning (RL) policies from offline multi-task datasets without direct environmental interaction is crucial for efficient and reliable deployment of RL agents. Benefiting from their scalability and strong in-context learning capabilities, recent approaches attempt to utilize Decision Transformer (DT) architectures for offline safe RL, demonstrating promising adaptability across varying safety budgets.
However, these methods primarily focus on single-constraint scenarios and struggle with diverse constraint configurations across multiple tasks.
Additionally, their reliance on heuristically defined Return-To-Go (RTG) inputs limits flexibility and reduces learning efficiency, particularly in complex multi-task environments. To address these limitations, we propose CoPDT, a novel DT-based framework designed to enhance adaptability to diverse constraints and varying safety budgets. Specifically, CoPDT introduces a constraint prioritized prompt encoder, which leverages sparse binary cost signals to accurately identify constraints, and a constraint prioritized Return-To-Go (CPRTG) token mechanism, which dynamically generates RTGs based on identified constraints and corresponding safety budgets. Extensive experiments on the OSRL benchmark demonstrate that CoPDT achieves superior efficiency and significantly enhanced safety compliance across diverse multi-task scenarios, surpassing state-of-the-art DT-based methods by satisfying safety constraints in more than twice as many tasks. Ruiqi Xue, Lihe Li, Cong Guan, Lei Yuan 0005, Yang Yu 0001 |
NeurIPS | 4 |
| 2025 | Open and real-world human-AI coordination by heterogeneous training with communication
Cong Guan, Ke Xue 0001, Chunpeng Fan, Feng Chen 0042, Lei Yuan 0005, Chao Qian 0001, Yang Yu 0001 |
Frontiers Comput. Sci. | 1 |
| 2025 | Constraining an Unconstrained Multi-agent Policy with offline data
Cong Guan, Yi-Chen Li 0001, Zongzhang Zhang, Lei Yuan 0005, Yang Yu 0001 |
Neural Networks | 1 |
| 2025 | Heterogeneous Multiagent Zero-Shot Coordination by CoevolutionabstractGenerating agents that can achieve zero-shot coordination (ZSC) with unseen partners is a new challenge in cooperative multiagent reinforcement learning (MARL). Recently, some studies have made progress in ZSC by exposing the agents to diverse partners during the training process. They usually involve self-play when training the partners, implicitly assuming that the tasks are homogeneous. However, many real-world tasks are heterogeneous, and hence previous methods may be inefficient. In this article, we study the heterogeneous ZSC problem for the first time and propose a general method based on coevolution, which coevolves two populations of agents and partners through three subprocesses: 1) pairing; 2) updating; and 3) selection. Experimental results on various heterogeneous tasks highlight the necessity of considering the heterogeneous setting and demonstrate that our proposed method is a promising solution for heterogeneous ZSC tasks. To the best of our knowledge, we are the first to underscore the significance of the heterogeneous ZSC tasks and to introduce an effective framework for addressing it. Ke Xue 0001, Yutong Wang 0012, Cong Guan, Lei Yuan 0005, Haobo Fu, Qiang Fu 0016, Chao Qian 0001, Yang Yu 0001 |
IEEE Trans. Evol. Comput. | 3 |
| 2025 | Learning to Coordinate With Different Teammates via Team ProbingabstractCoordinating with different teammates is essential in cooperative multiagent systems (MASs). However, most multiagent reinforcement learning (MARL) methods assume fixed team compositions, which leads to agents overfitting their training partners and failing to cooperate well with different teams during the deployment phase. A common way to mitigate the problem is to anticipate teammate behaviors and adapt policies accordingly during cooperation. However, these methods use the same policy for both collecting information for modeling teammates and maximizing cooperation performance. We argue that these two goals may conflict and reduce the effectiveness of both. In this work, we propose coordinating with different teammates via team probing (CDP), a novel approach that rapidly adapts to different teams by disentangling probing and adaptation phases. Specifically, we first generate a diverse population of teams as training partners with a novel value-based diversity objective. Then, we train a probing module to probe and reveal the coordination pattern of each team with policy-dynamics reconstruction and get a representation space of the population. Finally, we train a generalist meta-policy consisting of several expert policies with module selection based on the clustering of the learned representation space. We empirically show that CDP surpasses existing policy adaptation methods in various complex multiagent scenarios with both seen and unseen teammates. Chengxing Jia, Zongzhang Zhang, Cong Guan, Feng Chen 0042, Lei Yuan 0005, Yang Yu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Efficient Communication via Self-Supervised Information Aggregation for Online and Offline Multiagent Reinforcement LearningabstractUtilizing messages from teammates can improve coordination in cooperative multiagent reinforcement learning (MARL). Previous works typically combine raw messages of teammates with local information as inputs for policy. However, neglecting message aggregation poses significant inefficiency for policy learning. Motivated by recent advances in representation learning, we argue that efficient message aggregation is essential for good coordination in cooperative MARL. In this article, we propose Multiagent communication via Self-supervised Information Aggregation (MASIA), where agents can aggregate the received messages into compact representations with high relevance to augment the local policy. Specifically, we design a permutation-invariant message encoder to generate common information-aggregated representation from messages and optimize it via reconstructing and shooting future information in a self-supervised manner. Hence, each agent would utilize the most relevant parts of the aggregated representation for decision-making by a novel message extraction mechanism. Furthermore, considering the potential of offline learning for real-world applications, we build offline benchmarks for multiagent communication, which is the first as we know. Empirical results demonstrate the superiority of our method in both online and offline settings. We also release the built offline benchmarks in this article as a testbed for communication ability validation to facilitate further future research in this direction. Cong Guan, Feng Chen 0042, Lei Yuan 0005, Zongzhang Zhang, Yang Yu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Multiagent Continual Coordination via Progressive Task ContextualizationabstractCooperative multiagent reinforcement learning (MARL) has attracted significant attention and has the potential for many real-world applications. Previous arts mainly focus on facilitating the coordination ability from different aspects (e.g., nonstationarity and credit assignment) in single-task or multitask scenarios, ignoring the stream of tasks that appear in a continual manner. This ignorance makes the continual coordination an unexplored territory, neither in problem formulation nor efficient algorithms designed. Toward tackling the mentioned issue, this article proposes an approach, multiagent continual coordination via progressive task contextualization (MACPro). The key point lies in obtaining a factorized policy, using shared feature extraction layers but separated independent task heads, each specializing in a specific class of tasks. The task heads can be progressively expanded based on the learned task contextualization. Moreover, to cater to the popular centralized training with decentralized execution (CTDE) paradigm in MARL, each agent learns to predict and adopt the most relevant policy head based on local information in a decentralized manner. We show in multiple multiagent benchmarks that existing continual learning methods fail, while MACPro is able to achieve close-to-optimal performance. More results also disclose the effectiveness of MACPro from multiple aspects, such as high generalization ability. Lei Yuan 0005, Lihe Li, Fuxiang Zhang, Cong Guan, Yang Yu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Quality-Diversity with Limited ResourcesabstractQuality-Diversity (QD) algorithms have emerged as a powerful optimization paradigm with the aim of generating a set of high-quality and diverse solutions. To achieve such a challenging goal, QD algorithms require maintaining a large archive and a large population in each iteration, which brings two main issues, sample and resource efficiency. Most advanced QD algorithms focus on improving the sample efficiency, while the resource efficiency is overlooked to some extent. Particularly, the resource overhead during the training process has not been touched yet, hindering the wider application of QD algorithms. In this paper, we highlight this important research question, i.e., how to efficiently train QD algorithms with limited resources, and propose a novel and effective method called RefQD to address it. RefQD decomposes a neural network into representation and decision parts, and shares the representation part with all decision parts in the archive to reduce the resource overhead. It also employs a series of strategies to address the mismatch issue between the old decision parts and the newly updated representation part. Experiments on different types of tasks from small to large resource consumption demonstrate the excellent performance of RefQD: it not only uses significantly fewer resources (e.g., 16% GPU memories on QDax and 3.7% on Atari) but also achieves comparable or better performance compared to sample-efficient QD algorithms. Our code is available at [https://github.com/lamda-bbo/RefQD](https://github.com/lamda-bbo/RefQD). Ren-Jian Wang, Ke Xue 0001, Cong Guan, Chao Qian 0001 |
ICML | 3 |
| 2024 | Continual Multi-Objective Reinforcement Learning via Reward Model Rehearsal
Lihe Li, Ruotong Chen, Yi-Chen Li 0001, Cong Guan, Yang Yu 0001, Lei Yuan 0005 |
IJCAI | 6 |
| 2024 | Multi-Agent Domain Calibration with a Handful of Offline DataabstractThe shift in dynamics results in significant performance degradation of policies trained in the source domain when deployed in a different target domain, posing a challenge for the practical application of reinforcement learning (RL) in real-world scenarios. Domain transfer methods aim to bridge this dynamics gap through techniques such as domain adaptation or domain calibration. While domain adaptation involves refining the policy through extensive interactions in the target domain, it may not be feasible for sensitive fields like healthcare and autonomous driving. On the other hand, offline domain calibration utilizes only static data from the target domain to adjust the physics parameters of the source domain (e.g., a simulator) to align with the target dynamics, enabling the direct deployment of the trained policy without sacrificing performance, which emerges as the most promising for policy deployment. However, existing techniques primarily rely on evolution algorithms for calibration, resulting in low sample efficiency.
To tackle this issue, we propose a novel framework Madoc (\textbf{M}ulti-\textbf{a}gent \textbf{do}main \textbf{c}alibration). Firstly, we formulate a bandit RL objective to match the target trajectory distribution by learning a couple of classifiers. We then address the challenge of a large domain parameter space by modeling domain calibration as a cooperative multi-agent reinforcement learning (MARL) problem. Specifically, we utilize a Variational Autoencoder (VAE) to automatically cluster physics parameters with similar effects on the dynamics, grouping them into distinct agents. These grouped agents train calibration policies coordinately to adjust multiple parameters using MARL.
Our empirical evaluation on 21 offline locomotion tasks in D4RL and NeoRL benchmarks showcases the superior performance of our method compared to strong existing offline model-based RL, offline domain calibration, and hybrid offline-and-online RL baselines. Lei Yuan 0005, Lihe Li, Cong Guan, Zongzhang Zhang, Yang Yu 0001 |
NeurIPS | 4 |
| 2023 | Robust Multi-Agent Coordination via Evolutionary Generation of Auxiliary Adversarial AttackersabstractCooperative Multi-agent Reinforcement Learning (CMARL) has shown to be promising for many real-world applications. Previous works mainly focus on improving coordination ability via solving MARL-specific challenges (e.g., non-stationarity, credit assignment, scalability), but ignore the policy perturbation issue when testing in a different environment. This issue hasn't been considered in problem formulation or efficient algorithm design. To address this issue, we firstly model the problem as a Limited Policy Adversary Dec-POMDP (LPA-Dec-POMDP), where some coordinators from a team might accidentally and unpredictably encounter a limited number of malicious action attacks, but the regular coordinators still strive for the intended goal. Then, we propose Robust Multi-Agent Coordination via Evolutionary Generation of Auxiliary Adversarial Attackers (ROMANCE), which enables the trained policy to encounter diversified and strong auxiliary adversarial attacks during training, thus achieving high robustness under various policy perturbations. Concretely, to avoid the ego-system overfitting to a specific attacker, we maintain a set of attackers, which is optimized to guarantee the attackers high attacking quality and behavior diversity. The goal of quality is to minimize the ego-system coordination effect, and a novel diversity regularizer based on sparse action is applied to diversify the behaviors among attackers. The ego-system is then paired with a population of attackers selected from the maintained attacker set, and alternately trained against the constantly evolving attackers. Extensive experiments on multiple scenarios from SMAC indicate our ROMANCE provides comparable or better robustness and generalization ability than other baselines. Lei Yuan 0005, Ke Xue 0001, Feng Chen 0042, Cong Guan, Lihe Li, Chao Qian 0001, Yang Yu 0001 |
AAAI | 6 |
| 2023 | Learning to Coordinate with AnyoneabstractIn open multi-agent environments, the agents may encounter unexpected teammates. Classical multi-agent learning approaches train agents that can only coordinate with seen teammates. Recent studies attempted to generate diverse teammates in order to enhance the generalizable coordination ability, but were restricted by pre-defined teammates. In this work, our aim is to train agents with strong coordination ability by generating teammates that fully cover the teammate policy space, so that agents can coordinate with any teammates. Since the teammate policy space is too huge to be enumerated, we find only dissimilar teammates that are incompatible with controllable agents, which highly reduces the number of teammates that needed to be trained with. However, it is hard to determine the number of such incompatible teammates beforehand. We therefore introduce a continual multi-agent learning process, in which the agent learns to coordinate with different teammates until no more incompatible teammates can be found. The above idea is implemented in the proposed Macop (Multi-agent compatible policy learning) algorithm. We conduct experiments in 8 scenarios from 4 environments that have distinct coordination patterns. Experiments show that Macop generates training teammates with much lower compatibility than previous methods. As a result, in all scenarios Macop achieves the best overall coordination ability while never significantly worse than the baselines, showing strong generalization ability. Lei Yuan 0005, Lihe Li, Feng Chen 0042, Cong Guan, Yang Yu 0001, Zhi-Hua Zhou |
DAI | 6 |
| 2023 | Fast Teammate Adaptation in the Presence of Sudden Policy ChangeabstractCooperative multi-agent reinforcement learning (MARL), where agents coordinates with teammate(s) for a shared goal, may sustain non-stationary caused by the policy change of teammates. Prior works mainly concentrate on the policy change cross episodes, ignoring the fact that teammates may suffer from sudden policy change within an episode, which might lead to miscoordination and poor performance. We formulate the problem as an open Dec-POMDP, where we control some agents to coordinate with uncontrolled teammates, whose policies could be changed within one episode. Then we develop a new framework \textit{\textbf{Fas}t \textbf{t}eammates \textbf{a}da\textbf{p}tation (\textbf{Fastap})} to address the problem. Concretely, we first train versatile teammates’ policies and assign them to different clusters via the Chinese Restaurant Process (CRP). Then, we train the controlled agent(s) to coordinate with the sampled uncontrolled teammates by capturing their identifications as context for fast adaptation. Finally, each agent applies its local information to anticipate the teammates’ context for decision-making accordingly. This process proceeds alternately, leading to a robust policy that can adapt to any teammates during the decentralized execution phase. We show in multiple multi-agent benchmarks that Fastap can achieve superior performance than multiple baselines in stationary and non-stationary scenarios. Lei Yuan 0005, Lihe Li, Ke Xue 0001, Chengxing Jia, Cong Guan, Chao Qian 0001, Yang Yu 0001 |
UAI | 6 |
| 2022 | Multi-Agent Concentrative Coordination with Decentralized Task RepresentationabstractValue-based multi-agent reinforcement learning (MARL) methods hold the promise of promoting coordination in cooperative settings. Popular MARL methods mainly focus on the scalability or the representational capacity of value functions. Such a learning paradigm can reduce agents' uncertainties and promote coordination. However, they fail to leverage the task structure decomposability, which generally exists in real-world multi-agent systems (MASs), leading to a significant amount of time exploring the optimal policy in complex scenarios. To address this limitation, we propose a novel framework Multi-Agent Concentrative Coordination (MACC) based on task decomposition, with which an agent can implicitly form local groups to reduce the learning space to facilitate coordination. In MACC, agents first learn representations for subtasks from their local information and then implement an attention mechanism to concentrate on the most relevant ones. Thus, agents can pay targeted attention to specific subtasks and improve coordination. Extensive experiments on various complex multi-agent benchmarks demonstrate that MACC achieves remarkable performance compared to existing methods. Lei Yuan 0005, Chenghe Wang, Fuxiang Zhang, Feng Chen 0042, Cong Guan, Zongzhang Zhang, Chongjie Zhang, Yang Yu 0001 |
IJCAI | 6 |
| 2022 | Efficient Multi-agent Communication via Self-supervised Information AggregationabstractUtilizing messages from teammates can improve coordination in cooperative Multi-agent Reinforcement Learning (MARL). To obtain meaningful information for decision-making, previous works typically combine raw messages generated by teammates with local information as inputs for policy. However, neglecting the aggregation of multiple messages poses great inefficiency for policy learning. Motivated by recent advances in representation learning, we argue that efficient message aggregation is essential for good coordination in MARL. In this paper, we propose Multi-Agent communication via Self-supervised Information Aggregation (MASIA), with which agents can aggregate the received messages into compact representations with high relevance to augment the local policy. Specifically, we design a permutation invariant message encoder to generate common information aggregated representation from raw messages and optimize it via reconstructing and shooting future information in a self-supervised manner. Each agent would utilize the most relevant parts of the aggregated representation for decision-making by a novel message extraction mechanism. Empirical results demonstrate that our method significantly outperforms strong baselines on multiple cooperative MARL tasks for various task settings. Cong Guan, Feng Chen 0042, Lei Yuan 0005, Chenghe Wang, Zongzhang Zhang, Yang Yu 0001 |
NeurIPS | 1 |
| 2022 | Adaptive Privacy-Preserving Federated Learning for Fault Diagnosis in Internet of ShipsabstractThe recent appearance of Internet of Things (IoT) technologies applied in the maritime industry has introduced the Internet of Ships (IoS) paradigm. By leveraging IoS and deep learning (DL), various DL-based fault diagnosis methods have been proposed to improve shipping companies’ maintenance performance and reduce operational costs. However, the traditional centralized learning approach (CL), which centralizes the data resources of different shipping companies to a cloud server for model training, is restricted in real industrial scenarios due to privacy concerns and business competitions. In this article, we propose a novel adaptive privacy-preserving federated learning approach, named AdaPFL, for fault diagnosis in IoS, which can organize different shipping agents to collaboratively develop a model by sharing model parameters with no risk of data leakage. First, we use two common tasks as examples to demonstrate that a small part of the model parameters might reveal the shipping agents’ raw information. Based on this, the Paillier-based communication scheme is designed to preserve the raw information of the shipping agents. Furthermore, to deal with the harsh marine environment, a control algorithm is proposed to adaptively change the model aggregation interval during the training process for reducing cryptography computation and communication costs. Theoretical analysis and experiments prove the high effectiveness of the AdaPFL on a real nonindependent and identically distributed (non-i.i.d) fault data set. Cong Guan, Hui Chen 0021, Xiangguo Yang, Wenfeng Gong, Ansheng Yang |
IEEE Internet Things J. | 2 |