VLDB 2026 Research / reviewers in the wild / expert
Hongye Cao
dblp:278/9372
· DBLP profile ↗
18ranked-venue papers
7as first author
18since 2021 · last 2026
0000-0002-2537-2295ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Causality-Aware Efficient Exploration for Cooperative Multi-Agent Reinforcement LearningabstractExploration is critical for cooperative multi agent reinforcement learning (MARL) to improve sample efficiency. However, existing intrinsic motivation based exploration strategies in MARL overlook the causal relationships among agents, global states, and rewards, suffering from interference by irrelevant factors and resulting in sample inefficiency. To address this issue, we propose Causality aware Efficient Exploration (CEE), a novel framework that enhances sample efficiency by inferring causal relationships between agents, global states with respect to rewards, thereby enabling causality guided exploration. Specifically, CEE operates through two components. First, CEE identifies causal relationships between global states and rewards, filtering out causally irrelevant state features that do not have a high impact on rewards to keep decision critical state information. Second, CEE discovers causal relationships between agents' behaviors and rewards to quantify each agent's contribution to collective performance. To achieve this, we introduce a causal entropy objective that promotes exploration aligned with decision critical aspects of the underlying causal structure. We provide comprehensive validation through experiments on 21 challenging tasks spanning SMAC, SMAC v2, and Google Research Football (GRF) environments. Our results demonstrate that CEE achieves superior performance in terms of sample efficiency and asymptotic performance compared to existing MARL methods. Hongye Cao, Tianpei Yang, Hammadi Rafik Ouariachi, Yali Du 0001, Jing Huo, Yang Gao 0001 |
AAAI | 1 |
| 2026 | MAPG2: Multiagent Policy Gradient via Potential Game for Multirobot Task Allocation ProblemsabstractEfficient task allocation among multiple UAVs and autonomous robots is critical in modern IoT scenarios. This is typically modeled as a multi-robot task allocation (MRTA) problem, known to be an NP-hard combinatorial optimization problem. Neural sequential modeling combined with reinforcement learning (RL) optimization has emerged as a promising paradigm for solving this problem, owing to its high efficiency during inference. However, most existing methods assume that each robot is capable of performing only a single type of task. The development of sensing technologies has significantly enhanced the functional diversity of robots, thereby challenging the effectiveness and scalability of traditional methods. This paper considers a variant of the MRTA problem, where each robot is capable of handling multiple tasks, and tasks vary in both their types and required resources. To this end, we present a novel game-theoretic multi-agent RL algorithm called multi-agent policy gradient via potential game (MAPG2). The key components of proposed method consist of three parts. Firstly, we utilize graph-based attention model (GAM) to characterize the representations between tasks. Secondly, we formulate the single-step allocation process as a potential game (PG) to guarantee the consistency and soundness of the reward function design. Lastly, our approach sequentially generates allocation strategies through centralized training and decentralized execution (CTDE) framework. Extensive experiments demonstrate that MAPG2achieves a 10% improvement in task completion rate compared to state-of-the-art baselines, validating its effectiveness and robustness. Shangdong Yang, Hongye Cao, Xingguo Chen, Yansheng Wu, Gongzhi Luo |
IEEE Internet Things J. | 3 |
| 2026 | Mitigating Security Risks in Large Language Models: A Full Lifecycle Perspective
Yanming Wang, Zhixin Bai, Jing Huo, Hongye Cao, Yang Gao 0001 |
Mach. Learn. | 6 |
| 2026 | A unified and efficient training framework for open-ended non-transitive games
Shaokang Dong, Shangdong Yang, Hongye Cao, Wanqi Yang, Yang Gao 0001 |
Neural Networks | 4 |
| 2026 | Retrieval-augmented diffusion with acoustic priors for high-fidelity sonar image generation
Shaocong Yang, Zheng Gu 0001, Hongye Cao, Xiaolong Qi, Jing Huo, Yang Gao 0001 |
Pattern Recognit. | 3 |
| 2026 | Model-Based Offline Reinforcement Learning With Adversarial Data AugmentationabstractModel-based offline reinforcement learning (RL) constructs environment models from offline datasets to perform conservative policy optimization. Existing approaches focus on learning state transitions through ensemble models, rolling out conservative estimation to mitigate extrapolation errors. However, the static data makes it challenging to develop a robust policy, and offline agents cannot access the environment to gather new data. To address these challenges, we introduce Model-based Offline Reinforcement learning with AdversariaL data augmentation (MORAL). In MORAL, we replace the fixed horizon rollout by employing adversarial data augmentation to execute alternating sampling with ensemble models to enrich training data. Specifically, this adversarial process dynamically selects ensemble models against policy for biased sampling, mitigating the optimistic estimation of fixed models, thus robustly expanding the training data for policy optimization. Moreover, a differential factor (DF) is integrated into the adversarial process for regularization, ensuring error minimization in extrapolations. This data-augmented optimization adapts to diverse offline tasks without rollout horizon tuning, showing remarkable applicability. Extensive experiments on the D4RL benchmark demonstrate that MORAL outperforms other model-based offline RL methods in terms of policy learning and sample efficiency. Hongye Cao, Jing Huo, Shangdong Yang, Tianpei Yang, Yang Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Beyond Mandatory Federations: Balancing Egoism, Utilitarianism and Egalitarianism in Mixed-Motive GamesabstractIn the field of mixed-motive games, extensive multi-agent learning studies have explored the balance between egoism (individual interest), utilitarianism (collective interest), and egalitarianism (fairness). Traditional approaches often rely on manually designed reward functions, social norms, and alliance/federation mechanisms to transition agents from individualistic behaviors toward cooperative strategies. However, these methods typically require all agents to share private local information or to mandatorily participate in federations, which is impractical in real-world applications. To address these issues, this paper proposes a Flexible-Participation Federation (FPF) framework that allows agents to participate in the federation voluntarily. Furthermore, we extend the federation from a global to a Local Multi-Federation (LMF) framework, enabling agents to form multiple localized federations, thereby promoting more efficient and adaptive cooperation. Theoretical evidence demonstrates that the global FPF model, along with the discrepancy between decentralized egoistic policies and federated utilitarian policies, achieves an O(1/T) convergence rate. Agents in the LMF framework also reach consensus within a sublinear gap. Extensive experiments show that agents opting out of federation participation experience a reduction in egoism, and our approach outperforms multiple baselines in terms of both utilitarianism and egalitarianism. Shaokang Dong, Shangdong Yang, Hongye Cao, Wanqi Yang, Yang Gao 0001 |
AAAI | 4 |
| 2025 | Towards Empowerment Gain through Causal Structure Learning in Model-Based Reinforcement LearningabstractIn Model-Based Reinforcement Learning (MBRL), incorporating causal structures into dynamics models provides agents with a structured understanding of the environments, enabling efficient decision.
Empowerment as an intrinsic motivation enhances the ability of agents to actively control their environments by maximizing the mutual information between future states and actions.
We posit that empowerment coupled with causal understanding can improve controllability, while enhanced empowerment gain can further facilitate causal reasoning in MBRL.
To improve learning efficiency and controllability, we propose a novel framework, Empowerment through Causal Learning (ECL), where an agent with the awareness of causal dynamics models achieves empowerment-driven exploration and optimizes its causal structure for task learning.
Specifically, ECL operates by first training a causal dynamics model of the environment based on collected data. We then maximize empowerment under the causal structure for exploration, simultaneously using data gathered through exploration to update causal dynamics model to be more controllable than dense dynamics model without causal structure. In downstream task learning, an intrinsic curiosity reward is included to balance the causality, mitigating overfitting.
Importantly, ECL is method-agnostic and is capable of integrating various causal discovery methods.
We evaluate ECL combined with $3$ causal discovery methods across $6$ environments including pixel-based tasks, demonstrating its superior performance compared to other causal MBRL methods, in terms of causal discovery, sample efficiency, and asymptotic performance. Hongye Cao, Shaokang Dong, Tianpei Yang, Jing Huo, Yang Gao 0001 |
ICLR | 1 |
| 2025 | Causal Information Prioritization for Efficient Reinforcement LearningabstractCurrent Reinforcement Learning (RL) methods often suffer from sample-inefficiency, resulting from blind exploration strategies that neglect causal relationships among states, actions, and rewards. Although recent causal approaches aim to address this problem, they lack grounded modeling of reward-guided causal understanding of states and actions for goal-orientation, thus impairing learning efficiency. To tackle this issue, we propose a novel method named Causal Information Prioritization (CIP) that improves sample efficiency by leveraging factored MDPs to infer causal relationships between different dimensions of states and actions with respect to rewards, enabling the prioritization of causal information. Specifically, CIP identifies and leverages causal relationships between states and rewards to execute counterfactual data augmentation to prioritize high-impact state features under the causal understanding of the environments. Moreover, CIP integrates a causality-aware empowerment learning objective, which significantly enhances the agent's execution of reward-guided actions for more efficient exploration in complex environments.
To fully assess the effectiveness of CIP, we conduct extensive experiments across $39$ tasks in $5$ diverse continuous control environments, encompassing both locomotion and manipulation skills learning with pixel-based and sparse reward settings. Experimental results demonstrate that CIP consistently outperforms existing RL methods across a wide range of scenarios. Hongye Cao, Tianpei Yang, Jing Huo, Yang Gao 0001 |
ICLR | 1 |
| 2025 | HDBO-B: On Benchmarking High-Dimensional Bayesian OptimizationabstractBayesian optimization (BO) has been extensively studied and applied as a sample-efficient, black-box optimization method in neural network learning, particularly for hyperparameter optimization. However, high-dimensional Bayesian optimization (HDBO) remains a challenging research direction. While existing BO benchmarks focus on practical low-dimensional tasks and specific high-dimensional settings, current HDBO studies often rely on custom experimental settings, resulting in a lack of standardized benchmarks for comprehensive evaluation. To address this, we introduce the first standardized HDBO benchmark HDBO-B. HDBO-B encompasses diverse high-dimensional optimization tasks, synthetic functions tailored for common research settings, and a robust framework with standardized testing, method implementation, and extensive examples. By benchmarking nine representative optimization methods, we demonstrate HDBO-B’s validity and highlight new research directions in neural network methods. Codes are available at: https://github.com/Yiyuiii/HDBO-B. Hongye Cao, Quanlin Chen, Jing Huo, Dong Li 0016, Yang Gao 0001 |
IJCNN | 2 |
| 2024 | Multi-Agent Exploration via Self-Learning and Social LearningabstractSelf-learning and social learning stand as two pivotal constituents in multi-agent exploration. Inspired by the fact that animals and humans explore unfamiliar environments to learn survival skills by training themselves using unlabeled data and replicating others’ successful experiences, we propose a multi-agent reinforcement learning method, named Self-Learning and Social Learning (S2L), which aims to address the complex tasks caused by sparse rewards and intricate sequential structures. Specifically, in Self-Learning, we incorporate both task-specific and task-agnostic intrinsic rewards. These incentives steer individual agents towards exploration and comprehension of the environment. Furthermore, in Social Learning, different independent agents can implicitly share the successful experience by observing others in view and without additional communication or parameter-sharing overhead. Finally, experimental evaluation of S2L on the complex task characterized by sparse rewards and intricate sequential structures demonstrates its superior performance against other competing exploration baselines. Shaokang Dong, Wubing Chen, Hongye Cao, Yang Gao 0001 |
ICASSP | 4 |
| 2024 | Multi-Agent Sparse Interaction Modeling is an Anomaly Detection ProblemabstractMost real-world multi-agent tasks exhibit the characteristic of sparse interaction, wherein agents interact with each other in a limited number of crucial states while largely acting independently. Effectively modeling the sparse interaction and leveraging the learned interaction structure to instruct agents’ learning processes can enhance the efficiency of multi-agent reinforcement learning algorithms. However, it remains unclear how to identify these specific interactive states solely through trials and errors within current multi-agent tasks. To address this challenge, this paper introduces a novel algorithm called Sparse Interaction as Anomaly (SIA), which innovatively casts the sparse interaction modeling into an anomaly detection problem. The underlying intuition is that interactive states appear rarely in agents’ trajectories and exhibit distinct dynamics compared to other commonplace states. Building upon this insight, SIA first employs variational inference to model the latent dynamics of agents’ trajectories. It then designates states with anomalous dynamics as the elusive interactive states and subsequently instructs agents to explore these states more extensively. This facilitates the emergence of interactive behaviors and promotes the learning of multi-agent policies. Experimental evaluation of SIA across various multi-agent tasks demonstrates its superior performance against multiple baselines, highlighting its effectiveness. Shaokang Dong, Shangdong Yang, Hongye Cao, Yang Gao 0001 |
ICASSP | 4 |
| 2023 | Enhancing OOD Generalization in Offline Reinforcement Learning with Energy-Based Policy OptimizationabstractOffline Reinforcement Learning (RL) is an important research domain for real-world applications because it can avert expensive and dangerous online exploration. Offline RL is prone to extrapolation errors caused by the distribution shift between offline datasets and states visited by behavior policy. Existing offline RL methods constrain the policy to offline behavior to prevent extrapolation errors. But these methods limit the generalization potential of agents in Out-Of-Distribution (OOD) regions and cannot effectively evaluate OOD generalization behavior. To improve the generalization of the policy in OOD regions while avoiding extrapolation errors, we propose an Energy-Based Policy Optimization (EBPO) method for OOD generalization. An energy function based on the distribution of offline data is proposed for the evaluation of OOD generalization behavior, instead of relying on model discrepancies to constrain the policy. The way of quantifying exploration behavior in terms of energy values can balance the return and risk. To improve the stability of generalization and solve the problem of sparse reward in complex environment, episodic memory is applied to store successful experiences that can improve sample efficiency. Extensive experiments on the D4RL datasets demonstrate that EBPO outperforms the state-of-the-art methods and achieves robust performance on challenging tasks that require OOD generalization. Hongye Cao, Shangdong Yang, Jing Huo, Xingguo Chen, Yang Gao 0001 |
ECAI | 1 |
| 2023 | Multimodal Image Fusion Framework for End-to-End Remote Sensing Image RegistrationabstractWe formulate the registration as a function that maps the input reference and sensed images to eight displacement parameters between prescribed matching points, as opposed to the usual techniques (feature extraction–description–matching–geometric restrictions). The projection transformation matrix (PTM) is then computed in the neural network and used to warp the sensed image, uniting all matching tasks under one framework. In this article, we offer a multimodal image fusion network with self-attention to merge the feature representation of the reference and sensed images. The integration information is then utilized to regress the prescribed points’ displacement parameters to get PTM between the reference and sensed images. Finally, PTM is supplied into the spatial transformation network (STN), which warps the sensed image to the same coordinates as the reference image, achieving end-to-end matching. In addition, a dual-supervised loss function is proposed to optimize the network from both the prescribed point displacement and the overall pixel matching perspectives. The effectiveness of our method is validated by qualitative and quantitative experimental results on multimodal remote sensing image matching tasks. The code is available at:https://github.com/liliangzhi110/E2EIR. Liangzhi Li 0002, Mingtao Ding, Hongye Cao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Model-Based Offline Adaptive Policy Optimization with Episodic Memory
Hongye Cao, Qianru Wei, Jiangbin Zheng 0001, Yanqing Shi |
ICANN (2) | 1 |
| 2022 | Cross-Domain Reinforcement Learning for Sentiment Analysis
Hongye Cao, Qianru Wei, Jiangbin Zheng 0001 |
ICONIP (6) | 1 |
| 2022 | Joint Self-Attention for Remote Sensing Image MatchingabstractWe propose a semantic mapping-based remote sensing image matching method, which aims to obtain the matching positions of candidate patches containing keypoints directly on the reference image, avoiding the use of cost-volume search pixel by pixel. First, a global context-fusing attention structure is created to fuse global semantic information for candidate patches with the entire sensed image. Then, a self-attention layer with semantic dependencies is proposed to extract the semantic dependencies on the reference image for cross-modal representation. The global receptive field provided by self-attention enables the proposed method to obtain the semantic mapping of candidate patches on the reference image. The experimental results show that the proposed method is insensitive to image distortion and achieves cross-modal matching of SAR-optical images with high accuracy, while still running several orders of magnitude faster. This ensures increased speed in remote sensing image analysis and pipeline processing while promoting new directions in learning-based registration. Liangzhi Li 0002, Hongye Cao, Huijuan Hu |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Remote Sensing Image Registration Based on Deep Learning Regression ModelabstractWe propose a novel remote sensing image registration method based on the deep learning regression network. Different from the traditional methods of feature extraction and feature matching, we pair the image blocks from sensed and reference images, and then directly learn the displacement parameters of the four corners of the sensed image block relative to the reference image. In addition, we develop the dual deep learning network with weight sharing to fully extract the registration pair image features. The proposed method is tested on different period Landsat-7 and WorldView-3 images and compared with scale-invariant feature transform (SIFT), fast and rotated brief (ORB), and other deep learning methods. The proposed method outperforms all the comparing methods. Liangzhi Li 0002, Mingtao Ding, Hongye Cao |
IEEE Geosci. Remote. Sens. Lett. | 5 |