VLDB 2026 Research / reviewers in the wild / expert
Huanhuan Yang
dblp:163/9778
· DBLP profile ↗
20ranked-venue papers
4as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | D3HRL: A distributed hierarchical reinforcement learning approach based on causal discovery and spurious correlation detection
Chenran Zhao, Dian-xi Shi, Mengzhu Wang, Jianqiang Xia, Huanhuan Yang, Songchang Jin, Shaowu Yang, Chunping Qiu |
Neural Networks | 5 |
| 2025 | Acting Beyond Learning: Imagination-Assisted Decision-Making in the Visual-based Multi-Agent Cooperative ScenariosabstractLearning optimal policies in multi-agent cooperative settings with visual observations is significant and challenging. Agents must first perform state representation learning for their image observations and then learn policies in the abstracted state space. Aiming at this problem, we propose a novel model-based MARL method named Contrastive Latent World for Policy Optimization (CLWPO). In CLWPO, we first design a state representation model to facilitate learning in the latent state space. With the support of this model, we construct the latent world and introduce a contrastive variational bound (CVB) to optimize it. Subsequently, we develop a heuristic policy optimization (HPO) scheme, incorporating model-free learning with model-based planning to obtain robust policies that predict future behaviors. In particular, in the planning, we maintain a queue of teammate models and calculate an adaptive rollout length for each agent to support their self-imagination and reduce the model-based return discrepancy. Finally, we conducted extensive experiments in the PettingZoo benchmark, and results show that CLWPO significantly enhances learning efficiency and improves agent performance compared to state-of-the-art MARL methods. Huanhuan Yang, Dian-xi Shi, Songchang Jin, Guojun Xie, Chunping Qiu, Shaowu Yang |
AAAI | 1 |
| 2025 | Multi-Agent Hierarchical Graph Attention Actor-Critic Reinforcement LearningabstractMulti-agent systems often face challenges such as elevated communication demands and intricate interactions. We propose an innovative hierarchical graph attention actor-critic reinforcement learning method to address the issues, which uses the hierarchical graph attention to capture the relationships of cooperation or competition among agents, and the agent enables a better understand of the dynamic environment. Specifically, we model the interaction among agents as a graph and encode the observations of the agents as a feature embedding vector with constant dimensionality to improve scalability. Through the "inter-agent" and "inter-group" attention layers, the embedding vector of each agent is updated into an information-condensed and contextualized state representation, which can adaptively extract the state-dependent relationship between agents, model the interaction at both the individual and group level, and thus learn more "advanced" strategies. Finally, we experiment on multiple multi-agent tasks to validate our proposed method’s effectiveness, stability, and scalability. Tongyue Li, Dian-xi Shi, Songchang Jin, Zhen Wang 0052, Huanhuan Yang |
ICASSP | 5 |
| 2025 | Excellent-student learning method for decentralized MARL with networked agents systemabstractAbstract Multi-agent reinforcement learning has been widely applied in solving various sequential decision-making problems in recent years. However, a key challenge is sparse rewards, where agents receive meaningful reward signals only upon task completion. This issue leads to inefficient exploration and slow learning progress. To address this problem, we propose an excellent-student learning method supported by a decentralized distributed learning paradigm, drawing inspiration from academically diverse classrooms. Specifically, we design a novel excellent-student learning model, which suggests that agents mimic the learning behaviors of excellent students. This model requires agents to share knowledge with other agents and engage in individual exploration driven by curiosity. Next, to foster collaborative team learning behaviors, similarity measurement techniques are integrated to enhance knowledge sharing among agents. An intrinsic reward function is designed, combining individual exploration with whole-class sharing, providing additional motivation for discovering new actions and states. This reward function is seamlessly incorporated into the policy learning process. Finally, experiments conducted in various multi-agent particle environments demonstrate significant improvements in training efficiency and stability. Dian-xi Shi, Huanhuan Yang, Tongyue Li, Zhen Wang 0052 |
Comput. J. | 3 |
| 2024 | SVR-AVT: Scale Variation Robust Active Visual TrackingabstractActive Visual Tracking (AVT) is a significant research area with extensive applications in fields such as drones and autonomous driving. AVT involves controlling camera motion based on visual observations to track target object(s). In dynamic environments, especially with the presence of distractors, AVT faces the challenge of scale variation. Existing methods struggle to effectively handle these scale changes. To address this problem, this paper proposes a novel Scale Variation Robust Active Visual Tracking method (SVR-AVT). We first introduce a multi-scale multi-stage curriculum learning approach. By progressively increasing the complexity of tracking tasks, the tracker adapts to target of various scales. Secondly, we design a scale attention network, which adaptively extracts important scale features through multiple convolutional branches with different receptive fields and a scale attention mechanism. Moreover, we employ maximum position entropy learning to encourage the target to explore the environment more extensively. Experimental results in 3D environments demonstrate that SVR-AVT significantly outperforms existing methods in handling distraction and scale variation, and exhibits strong generalization capability in unseen environments. Zhang Biao, Songchang Jin, Qianying Ouyang, Huanhuan Yang, Yuxi Zheng, Chunlian Fu, Dian-xi Shi |
IJCNN | 4 |
| 2024 | Improved Communication and Collision-Avoidance in Dynamic Multi-Agent Path FindingabstractMulti-Agent Path Finding (MAPF) is a classic problem with a wide range of applications. To cope with more complex situations in reality, Dynamic MAPF (DMAPF) has received much attention. The existing DMAPF definition lacks completeness or considers too simple situations. In this paper, we comprehensively model DMAPF based on realistic scenarios. Consequently, dynamic scenarios bring many problems. The dynamics of agent tasks bring the problem of more difficult coordination and cooperation of the multi-agent system, and the dynamics of obstacles bring the problem of increased collisions. To address these problems, this paper proposes a fully decentralised multi-agent reinforcement learning method CO3, which uses COmmon knowledge in selective COmmunication and proposes obstacle COllision avoidance mechanism. Firstly, common knowledge for communication improves cooperation between agents, which improves system performance and reduces collisions between agents. Secondly, the obstacle collision avoidance mechanism consists of a collision avoidance helper module and a critical region. The collision avoidance helper module improves the agents’ alertness to nearby obstacles, and the critical region gives an early warning to the agents to beware of distant obstacles. The obstacle collision avoidance mechanism can effectively reduce collisions between agents and obstacles. Finally, experiments show that CO3 can solve the DMAPF problem quite well, and the number of collisions is significantly lower than other learning-based methods in a dynamic environment. Jing Xie 0021, Yongjun Zhang 0006, Qianying Ouyang, Huanhuan Yang, Dian-xi Shi, Songchang Jin |
IJCNN | 4 |
| 2024 | A framework for formal verification of robot kinematics
Guojun Xie, Huanhuan Yang |
J. Log. Algebraic Methods Program. | 2 |
| 2024 | Query-oriented two-stage attention-based model for code search
Huanhuan Yang, Chao Liu 0014, Luwen Huangfu |
J. Syst. Softw. | 1 |
| 2024 | An anti-collision algorithm for robotic search-and-rescue tasks in unknown dynamic environmentsabstractThis paper deals with the search-and-rescue tasks of a mobile robot with multiple interesting targets in an unknown dynamic environment. The problem is challenging because the mobile robot needs to search for multiple targets while avoiding obstacles simultaneously. To ensure that the mobile robot avoids obstacles properly, we propose a mixed-strategy Nash equilibrium based Dyna-Q (MNDQ) algorithm. First, a multi-objective layered structure is introduced to simplify the representation of multiple objectives and reduce computational complexity. This structure divides the overall task into subtasks, including searching for targets and avoiding obstacles. Second, a risk-monitoring mechanism is proposed based on the relative positions of dynamic risks. This mechanism helps the robot avoid potential collisions and unnecessary detours. Then, to improve sampling efficiency, MNDQ is presented, which combines Dyna-Q and mixed-strategy Nash equilibrium. By using mixed-strategy Nash equilibrium, the agent makes decisions in the form of probabilities, maximizing the expected rewards and improving the overall performance of the Dyna-Q algorithm. Furthermore, a series of simulations are conducted to verify the effectiveness of the proposed method. The results show that MNDQ performs well and exhibits robustness, providing a competitive solution for future autonomous robot navigation tasks. Dian-xi Shi, Huanhuan Yang, Tongyue Li, Zhen Wang 0052 |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2024 | Learning cooperative strategies in multi-agent encirclement games with faster prey using prior knowledge
Tongyue Li, Dian-xi Shi, Zhen Wang 0052, Huanhuan Yang |
Neural Comput. Appl. | 4 |
| 2023 | Faster Target Encirclement with Utilization of Obstacles via Multi-Agent Reinforcement Learning
Yuxi Zheng, Yongjun Zhang 0006, Chenran Zhao, Huanhuan Yang, Tongyue Li, Qianying Ouyang |
ACML | 4 |
| 2022 | Deep Reinforcement Learning for Multi-UAV Exploration Under Energy Constraints
Yating Zhou, Dian-xi Shi, Huanhuan Yang, Haomeng Hu, Shaowu Yang, Yongjun Zhang 0006 |
CollaborateCom (2) | 3 |
| 2022 | Independent Multi-agent Reinforcement Learning Using Common KnowledgeabstractMany recent multi-agent reinforcement learning algorithms used centralized training with decentralized execution (CTDE), which results in a training process that relies on global information and suffers from the dimensional explosion. The independent learning (IL) approaches are simple in structure and can be more easily deployed to a wider range of multi-agent scenarios, but they can only solve relatively simple problems due to environment non-stationarity and partially observable. With this motivation, we let IL agents compute common knowledge information and fuse it with observation to explicitly exploit common knowledge. In addition, we chose a suitable network structure according to the characteristics of IL, using convolutional layers and GRU layers. Based on the above two improvements, we implement two IL algorithms. In our experiments, the algorithms we implemented show significant performance improvements compared to original IL algorithms and further approach CTDE while outperforming multi-agent common knowledge reinforcement learning. Haomeng Hu, Dian-xi Shi, Huanhuan Yang, Yingxuan Peng, Yating Zhou, Shaowu Yang |
SMC | 3 |
| 2022 | Self-supervised representations for multi-view reinforcement learningabstractLearning policies from raw, pixel images are quite important for the real-world application of deep reinforcement learning (RL). Standard model-free RL algorithms focus on single-view settings and unify the representation learning and policy learning into an end-to-end training process. However, such a learning paradigm is sample-inefficiency and sensitive to hyper-parameters when supervised merely by the reward signals. Based on this, we present Self-Supervised Representations (S2R) for multi-view reinforcement learning, a sample-efficient representation learning method for learning features from high-dimensional images. In S2R, we introduce a representation learning framework and define a novel multi-view auxiliary objective based on the multi-view image states and Conditional Entropy Bottleneck (CEB) principle. We integrate S2R with the deep RL agent to learn robust representations that preserve task-relevant information while discarding task-irrelevant information and find optimal policies that maximize the expected return. Empirically, we demonstrate the effectiveness of S2R in the visual DeepMind Control (DMControl) suite and show its better performance on the default DMControl tasks and their variants by replacing the tasks’ default background with a random image or natural video. Huanhuan Yang, Dian-xi Shi, Guojun Xie, Yingxuan Peng, Yantai Yang, Shaowu Yang |
UAI | 1 |
| 2022 | Multi actor hierarchical attention critic with RNN-based feature extraction
Dian-xi Shi, Chenran Zhao, Huanhuan Yang, Gongju Wang, Shaowu Yang, Yongjun Zhang 0006 |
Neurocomputing | 4 |
| 2022 | Multi-Electromagnetic Jamming Countermeasure for Airborne SAR Based on Maximum SNR Blind Source SeparationabstractSynthetic aperture radar (SAR) may be attacked by multifarious kinds of the active-jamming, which will lead to the ineffectiveness of SAR in complex electromagnetic environment. In this paper, a novel multi-electromagnetic jamming counter-measure for the airborne SAR is proposed based on maximum signal-to-noise ratio (SNR) blind source separation. Firstly, the imaging geometry and multi-component mixed signal model of airborne SAR are established. Then, based on multi-component mixed signal matrix, a blind source separation (BSS) method based on the maximum SNR is proposed to separate the real target echo signal from the multi-electromagnetic jamming signals. After real target echo signal identification, the high resolution images of the interested target area can be achieved by the corresponding SAR imaging method. The simulated and measured data results are present to prove the feasibility and effectiveness of the proposed method. Si Chen 0005, Sixiang Wang, Huanhuan Yang, Lingzhi Zhu, Huichang Zhao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | ContriQ: Ally-Focused Cooperation and Enemy-Concentrated Confrontation in Multi-Agent Reinforcement LearningabstractCentralized training with decentralized execution (CTDE) is an important setting for cooperative multi-agent reinforcement learning (MARL) due to communication constraints during execution and scalability constraints during training, which has shown superior performance but still suffers from challenges. One branch is to understand the mutual interplay between agents. Due to the communication constraints in practice, agents cannot exchange perceptual information, and thus, many approaches use a centralized attention network with scalability constraints. Contrary to these common approaches, we propose to learn to cooperate in a decentralized way by applying attention mechanism on the local observation so that each agent could focus on allied agents with a decentralized model, and therefore promote understanding. Another branch is to model how agents cooperate and simplify the learning process. Previous approaches that focus on value decomposition have achieved innovative results but still suffer from problems. These approaches either limit the representation expressiveness of their value function classes or relax the IGM consistency to achieve scalability, which may lead to poor performance. We combine value composition with game abstraction by modeling the relationships between agents as a bi-level graph. We propose a novel value decomposition network based on it through a bi-level attention network, which indicates the contribution of allied agents attacking enemies and the priority of attacking each enemy under the situation of each time step, respectively. We show that our method substantially outperforms existing state-of-the-art methods on battle games in StarCraft Ⅱ, and attention analysis is also comprehensively discussed with sights. Chenran Zhao, Dian-xi Shi, Huanhuan Yang, Shaowu Yang, Yongjun Zhang 0006 |
ACML | 4 |
| 2021 | CIExplore: Curiosity and Influence-based Exploration in Multi-Agent Cooperative Scenarios with Sparse RewardsabstractLearning in a sparse-reward setting is a well-known challenge in RL (Reinforcement Learning). In the single-agent domain, this challenge can be addressed by introducing exploration bonuses driven by intrinsic motivation to encourage agents to visit unseen states. However, naively applying these methods in MARL (Multi-Agent Reinforcement Learning) cooperative settings with sparse rewards results in some inevitable problems: misunderstanding environmental knowledge and lack of collaboration among agents, etc. Based on this, in this paper, we propose the Curiosity and Influence-based Explore (CIExplore) method, which includes a new form of intrinsic reward and an internal counterfactual advantage function. Concretely, the intrinsic reward is a combination of joint curiosity reward and influence reward. The former is the variance of outputs across an ensemble of prediction models that take joint observations and actions of all agents as inputs to predict the next time's joint observations. And the latter quantifies the influence of one agent's behavior on other agents' state-value functions. Given that the joint curiosity reward is shared by all agents, we compute an internal counterfactual advantage function to address this intrinsic reward assignment problem. We demonstrate the efficacy of CIExplore in the multi-agent grid-world environments and show that it is compatible with both on-policy and off-policy MARL algorithms and be scalable to complex settings where agents' number or environment randomness increases. Huanhuan Yang, Dian-xi Shi, Chenran Zhao, Guojun Xie, Shaowu Yang |
CIKM | 1 |
| 2021 | Two-Stage Attention-Based Model for Code Search with Textual and Structural FeaturesabstractSearching and reusing existing code from a large scale codebase can largely improve developers’ programming efficiency. To support code reuse, early code search models leverage information retrieval (IR) techniques to index a large-scale code corpus and return relevant code according to developers’ search query. However, IR-based models fail to capture the semantics in code and query. To tackle this issue, developers applied deep learning (DL) techniques to code search models. However, these models either are too complex to determine an effective method efficiently or learning for semantic correlation between code and query inadequately.To bridge the semantic gap between code and query effectively and efficiently, we propose a code search model TabCS (Two-stage Attention-Based model for Code Search) in this study. TabCS extracts code and query information from the code textual features (i.e., method name, API sequence, and tokens), the code structural feature (i.e., abstract syntax tree), and the query feature (i.e., tokens). TabCS performs a two-stage attention net-work structure. The first stage leverages attention mechanisms to extract semantics from code and query considering their semantic gap. The second stage leverages a co-attention mechanism to capture their semantic correlation and learn better code/query representation. We evaluate the performance of TabCS on two existing large-scale datasets with 485k and 542k code snippets, respectively. Experimental results show that TabCS achieves an MRR of 0.57 on Hu et al.’s dataset, outperforming three state-of-the-art models CARLCS-CNN, DeepCS, and UNIF by 18%, 70%, 12%, respectively. Meanwhile, TabCS gains an MRR of 0.54 on Husain et al.’s, outperforming CARLCS-CNN, DeepCS, and UNIF by 32%, 76%, 29%, respectively. Huanhuan Yang, Chao Liu 0014, Jianhang Shuai, Meng Yan 0001, Yan Lei 0005, Zhou Xu 0003 |
SANER | 2 |
| 2019 | GeoCET: Accurate IP Geolocation via Constraint-Based Elliptical Trajectories
Xiuguo Bao, Yongzheng Zhang 0002, Huanhuan Yang |
CollaborateCom | 4 |