VLDB 2026 Research / reviewers in the wild / expert
Sheng Han 0001
dblp:142/0013-1
· DBLP profile ↗
17ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0003-4049-3690ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Think How Your Teammates Think: Active Inference Can Benefit Decentralized ExecutionabstractIn multi-agent systems, explicit cognition of teammates' decision logic serves as a critical factor in facilitating coordination. Communication (i.e., "Tell") can assist in the cognitive development process by information dissemination, yet it is inevitably subject to real-world constraints such as noise, latency, and attacks. Therefore, building the understanding of teammates' decisions without communication remains challenging. To address this, we propose a novel non-communication MARL framework that realizes the construction of cognition through local observation-based modeling (i.e., "Think"). Our framework enables agents to model teammates' active inference process. At first, the proposed method produces three teammate portraits: perception-belief-action. Specifically, we model the teammate's decision process as follows: 1) Perception: observing environments; 2) Belief: forming beliefs; 3) Action: making decisions. Then, we selectively integrate the belief portrait into the decision process based on the accuracy and relevance of the perception portrait. This enables the selection of cooperative teammates and facilitates effective collaboration. Extensive experiments on the SMAC, SMACv2, MPE, and GRF benchmarks demonstrate the superior performance of our method. Hao Wu 0010, Shoucheng Song, Sheng Han 0001, Huaiyu Wan, Youfang Lin, Kai Lv 0002 |
AAAI | 4 |
| 2025 | Infer the Whole from a Glimpse of a Part: Keypoint-Based Knowledge Graph for Vehicle Re-IdentificationabstractVehicle re-identification aims to match vehicles across non-overlapping camera views. Many existing methods extract features from one specific image, and these methods lack view-invariance when comparing vehicles of different orientations. As a result, discriminative parts obscured by viewpoint changes cannot contribute effectively to matching. This work presents a novel keypoint-based framework for vehicle Re-ID. We propose to explicitly model the intrinsic structural relationships between vehicle components via knowledge graph. By establishing connection between keypoints, our approach aims to leverage such prior to match vehicles even when some parts are not directly comparable due to orientation inconsistencies. Specifically, given query and gallery images, we first detect visible keypoints. Then, a transformer-based model infers features for non-overlapped keypoints by conditioning on visible correspondences defined in the knowledge graph. The final representation integrates visible and inferred features. Extensive experiments demonstrate our method outperforms state-of-the-arts on standard benchmarks under cross-view matching scenarios. To our knowledge, this is the first work introducing structural priors via keypoint knowledge graphs for view-invariant vehicle re-identification. Kai Lv 0002, Shuo Wang 0031, Sheng Han 0001, Youfang Lin |
AAAI | 5 |
| 2025 | CoDe: Communication Delay-Tolerant Multi-Agent Collaboration via Dual Alignment of Intent and TimelinessabstractCommunication has been widely employed to enhance multi-agent collaboration. Previous research has typically assumed delay-free communication, a strong assumption that is challenging to meet in practice. However, real-world agents suffer from channel delays, receiving messages sent at different time points, termed Asynchronous Communication, leading to cognitive biases and breakdowns in collaboration. This paper first defines two communication delay settings in MARL and emphasizes their harm to collaboration. To handle the above delays, this paper proposes a novel framework, Communication Delay-Tolerant Multi-Agent Collaboration (CoDe). At first, CoDe learns an intent representation as messages through future action inference, reflecting the stable future behavioral trends of the agents. Then, CoDe devises a dual alignment mechanism of intent and timeliness to strengthen the fusion process of asynchronous messages. In this way, agents can extract the long-term intent of others, even from delayed messages, and selectively utilize the most recent messages that are relevant to their intent. Experimental results demonstrate that CoDe outperforms baseline algorithms in three MARL benchmarks without delay and exhibits robustness under fixed and time-varying delays. Shoucheng Song, Youfang Lin, Sheng Han 0001, Hao Wu 0010, Shuo Wang 0031, Kai Lv 0002 |
AAAI | 3 |
| 2025 | Improving Monotonic Optimization in Heterogeneous Multi-agent Reinforcement Learning with Optimal Marginal Deterministic Policy Gradient
Youfang Lin, Shuo Wang 0031, Sheng Han 0001 |
ICANN (1) | 4 |
| 2025 | Enhancing Offline Safe Reinforcement Learning with Trajectory-Constrained Diffusion Planning
Youfang Lin, Shuo Shen 0002, Hanfeng Lin, Peng Cheng 0013, Sheng Han 0001, Kai Lv 0002 |
AAMAS | 6 |
| 2025 | From General Relation Patterns to Task-Specific Decision-Making in Continual Multi-Agent CoordinationabstractContinual Multi-Agent Reinforcement Learning (Co-MARL) requires agents to address catastrophic forgetting issues while learning new coordination policies with the dynamics team. In this paper, we delve into the core of Co-MARL, namely Relation Patterns, which refer to agents’ general understanding of interactions. In addition to generality, relation patterns exhibit task-specificity when mapped to different action spaces. To this end, we propose a novel method called General Relation Patterns-Guided Task-specific Decision-Maker (RPG). In RPG, agents extract relation patterns from dynamic observation spaces using a relation capturer. These task-agnostic relation patterns are then mapped to different action spaces via a task-specific decision-maker generated by a conditional hypernetwork. To combat forgetting, we further introduce regularization items on both the relation capturer and the conditional hypernetwork. Results on SMAC and LBF demonstrate that RPG effectively prevents catastrophic forgetting when learning new tasks and achieves zero-shot generalization to unseen tasks. Youfang Lin, Shoucheng Song, Hao Wu 0010, Yuqing Ma, Sheng Han 0001, Kai Lv 0002 |
IJCAI | 6 |
| 2025 | From Pixels to Temporal Correlations: Learning Informative Representations for Reinforcement Learning Pre-trainingabstractUnsupervised pre-training on large-scale datasets has demonstrated significant potential for improving the sample efficiency and performance of Reinforcement Learning (RL). Given the large-scale action-free internet videos, existing methods utilize single-step transition prediction and image reconstruction to learn representations. However, these methods prefer to preserve large-proportion stationary information in the pixel space, neglecting small but crucial information. To preserve enough information in the representation, it is essential to pay equal attention to each element in videos. Specifically, we propose a temporal correlation space to distinguish each element. For implementation, we introduce the Multi-scale Temporal Contrastive Learning (MTCL) method to model multi-scale temporal correlations separately. This approach can balance the attention of different elements and yield more informative representations, effectively supporting policy learning in various downstream tasks. Experimental results demonstrate that our method improves sample efficiency and asymptotic performance across various downstream tasks. Youfang Lin, Sheng Han 0001, Shuo Wang 0031, Kai Lv 0002 |
ACM Multimedia | 5 |
| 2025 | Multi-constraint reinforcement learning in complex robot environments
Sheng Han 0001, Hao Wu 0010, Youfang Lin, Kai Lv 0002 |
Frontiers Comput. Sci. | 1 |
| 2025 | Off-Policy Conservative Distributional Reinforcement Learning With Safety ConstraintsabstractSafe exploration can be regarded as a constrained Markov decision problem (CMDP) where the expected long-term cost is constrained. Previous off-policy algorithms convert the constrained optimization problem into the corresponding unconstrained dual problem by introducing the Lagrangian relaxation technique. However, the cost function of the above algorithms provides inaccurate estimations and causes the instability of the Lagrange multiplier learning. In this article, we present a novel off-policy reinforcement learning (RL) algorithm called conservative distributional maximum a posteriori policy optimization (CDMPO). At first, to accurately judge whether the current situation satisfies the constraints, CDMPO adapts distributional RL method to estimate the Q-function and C-function. Then, CDMPO uses a conservative value function loss to reduce the number of violations of constraints during the exploration process. In addition, we utilize adaptive proportional integral derivative (APID) to update the Lagrange multiplier stably. In our experiments, we select eight representative constrained tasks from two well-known safe RL benchmarks (Safety Gym and Bullet Safety Gym), providing a comprehensive evaluation of our methods across diverse scenarios. Empirical results show that the proposed method has fewer violations of constraints in the early exploration process. The final test results also illustrate that our method has better-risk control capabilities. Youfang Lin, Sheng Han 0001, Shuo Wang 0031, Kai Lv 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2024 | Enhancing Off-Policy Constrained Reinforcement Learning through Adaptive Ensemble C EstimationabstractIn the domain of real-world agents, the application of Reinforcement Learning (RL) remains challenging due to the necessity for safety constraints. Previously, Constrained Reinforcement Learning (CRL) has predominantly focused on on-policy algorithms. Although these algorithms exhibit a degree of efficacy, their interactivity efficiency in real-world settings is sub-optimal, highlighting the demand for more efficient off-policy methods. However, off-policy CRL algorithms grapple with challenges in precise estimation of the C-function, particularly due to the fluctuations in the constrained Lagrange multiplier. Addressing this gap, our study focuses on the nuances of C-value estimation in off-policy CRL and introduces the Adaptive Ensemble C-learning (AEC) approach to reduce these inaccuracies. Building on state-of-the-art off-policy algorithms, we propose AEC-based CRL algorithms designed for enhanced task optimization. Extensive experiments on nine constrained robotics tasks reveal the superior interaction efficiency and performance of our algorithms in comparison to preceding methods. Youfang Lin, Shuo Shen 0002, Sheng Han 0001, Kai Lv 0002 |
AAAI | 4 |
| 2023 | Multi-mobile Object Motion Coordination with Reinforcement Learning
Shanhua Yuan, Sheng Han 0001, Xiwen Jiang, Youfang Lin, Kai Lv 0002 |
ICONIP (8) | 2 |
| 2023 | Offline Reinforcement Learning with Diffusion-Based Behavior Cloning Term
Youfang Lin, Sheng Han 0001, Kai Lv 0002 |
KSEM (4) | 3 |
| 2023 | A lightweight and style-robust neural network for autonomous driving in end side devicesabstractThe autonomous driving algorithm studied in this paper makes a ground vehicle capable of sensing its environment via visual images and moving safely with little or no human input. Due to the limitation of the computing power of end side devices, the autonomous driving algorithm should adopt a lightweight model and have high performance. Conditional imitation learning has been proved an efficient and promising policy for autonomous driving and other applications on end side devices due to its high performance and offline characteristics. In driving scenarios, the images captured in different weathers have different styles, which are influenced by various interference factors, such as illumination, raindrops, etc. These interference factors bring challenges to the perception ability of deep models, thus affecting the decision-making process in autonomous driving. The first contribution of this paper is to investigate the performance gap of driving models under different weather conditions. Following the investigation, we utilise StarGAN-V2 to translate images from source domains into the target clear sunset domain. Based on the images translated by StarGAN-V2, we propose Conditional Imitation Learning with ResNet backbone named Star-CILRS. The proposed method is able to convert an image to multiple styles using only one single model, making our method easier to deploy on end side devices. Visualization results show that Star-CILRS can eliminate some environmental interference factors. Our method outperforms other methods and the success rate values in different tasks are 98%, 74%, and 22%, respectively. Sheng Han 0001, Youfang Lin, Zhihui Guo, Kai Lv 0002 |
Connect. Sci. | 1 |
| 2021 | Collision-aware Multi-robot Motion Coordination Deep-RL with Dynamic Priority StrategyabstractThe motion coordination of multi-robots is the basis of the multi-robot system (MRS). The motion coordination problem of multi-robot satisfies the Markov property, so deep reinforcement learning (DRL) can be used to solve this problem. There are two limitations in the existing researches which apply DRL to solve this problem: the success rate of training is low and the effect of coordination is poor. We believe that in the training process, the agent should learn the collision constraint relationship from the collision. Therefore, we propose a Partially Tolerant Collision (PTC) collision handling strategy. And we believe that the completion time of motion coordination is related to the remaining distance of robots. Therefore, we propose the Dynamic Priority Strategy (DPS), which sets the priority for the robot based on the remaining distance of robots. This strategy is integrated into the reward setting of DRL. We use Path Checkerboard Diagram (PCD) as the basis for training and simulation. By experimenting with algorithms such as DQN, DDQN, and MLDDQN, the results of our proposed model are better than previous studies. Sheng Han 0001, Youfang Lin |
ICTAI | 2 |
| 2020 | Solving Open Shop Scheduling Problem via Graph Attention Neural NetworkabstractOpen Shop Scheduling Problem (OSSP) minimizing makespan has attracted attention increasingly. Complex constraints and large solution space cause great difficulty for acquiring optimal solutions. Traditional methods attempt to get suboptimal solutions based on pre-defined rules or local search. However, these methods are not universal and only applicable for problems with particular distributions. In this paper, we introduce Discount Memory into Graph Attention Model (GAM-DM) to solve OSSP and train the model with reinforcement learning. By constructing incremental graph solution, OSSP is converted to sequence to sequence problem, which makes GAM suitable for OSSP. Moreover, the proposed DM can help clarify the different influences of historical decisions on current decision-making step. We integrate GAM-DM into reinforcement learning to optimize the solution and conduct experiments on randomly generated problem sets. The experiment results indicate that our model outperforms traditional methods, benefiting from high-quality solution close to the lower bound. Compared with OR-Tool, our model achieves comparable solution quality with less computational time. Xingye Dong, Sheng Han 0001 |
ICTAI | 4 |
| 2019 | High-Value Prioritized Experience Replay for Off-Policy Reinforcement LearningabstractIn deep reinforcement learning, experience replay has been shown an effective solution to handle sample-inefficiency. Prioritized Experience Replay (PER) uses temporal-difference error (TD error) as replay priority in Deep Q-Networks (DQN), so that agent can learn more effectively from important experiences. However, experiences with large TD error may appear near the edge of state space and these experiences do not help agent learn policy quickly. We present a novel technique called High-Value Prioritized Experience Replay (HVPER), which designs a combination of TD error and value (reward or state-action value) in replay priority. Specifically, we first propose prioritizing replay based on reward and TD error in sparse reward environment. Extendedly, we design prioritizing replay based on state-action value and TD error for more ordinary environment. We design experiments in the gym environment to evaluate the proposed HVPER. First, we verify that the combination of TD error and reward improves the training speed in two problems with sparse rewards compared to DQN algorithm and PER algorithm. In addition, HVPER accelerates the network learning and achieves a better performance in two continuous space problems compared to Deep Deterministic Policy Gradient algorithm. Huaiyu Wan, Youfang Lin, Sheng Han 0001 |
ICTAI | 4 |
| 2019 | Motion Coordination of Multiple Robots Based on Deep Reinforcement LearningabstractMulti-robot motion coordination is a sequential decision problem that controls each robot's action through a coordination mechanism to avoid collisions and reach their respective destinations. This problem can be viewed as a Markov Decision Process. Therefore, under determinate conditions of the motion path, we make the attempt of applying deep reinforcement learning (DRL) to solve the motion coordination problems. In this paper, in order to make this problem solvable through RL methods, we design a Path Checkerboard Diagram (PCD) to represent the collision relationship among different paths and determine whether collisions occur when the agent takes action. Technically, we design the states, actions and reward function. However, the traditional Deep Q-Networks (DQN) algorithm has the problem of overestimation of state-action values and high approximation error variance of the target values. To fit the motion coordination problem scenario, we further propose a new algorithm Multi-Loss Double DQN (MLDDQN) based on the Double DQN (DDQN), which minimizes the weighted sum of loss of different target networks we have learned before. We evaluate the performance of our algorithm through the different motion coordination tasks. Results suggest that our MLDDQN method yields remarkable improvements in terms of the success rate, success steps and the stability of training process compared with the original DDQN algorithm. Xiuzhao Hao, Zhihao Wu 0001, Haiguang Zhou, Xiangpeng Bai, Youfang Lin, Sheng Han 0001 |
ICTAI | 6 |