EDBT 2026 Demo / reviewers in the wild / expert
Shuo Wang 0031
dblp:63/1591-31
· DBLP profile ↗
15ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0001-6599-3638ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Computer networks · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bidirectional transition consistency between multi-domain observations for visual reinforcement learning generalization
Youfang Lin, Shuo Wang 0031, Hehe Fan, Kai Lv 0002 |
Neural Networks | 5 |
| 2026 | Task-Relevant Representation Decoupling for Visual Reinforcement Learning GeneralizationabstractVisual Reinforcement Learning (VRL) has achieved considerable success in solving control tasks. However, generalizing learned policies to new environments remains a major challenge, as agents often overfit to task-irrelevant features in the training environment. To solve this problem, we introduce the concept of decoupling observations into task-relevant and task-irrelevant representations. Building on this idea, we propose a self-supervised T ask- R elevant R epresentation D ecoupling (T2RD) algorithm for VRL. This algorithm consists of three components: task-relevant representation consistency , cross-reconstruction , and cross-dynamic prediction . The first two components achieve the decoupling of content and style features, but the resulting content representations are not necessarily task-relevant. To further refine task-relevant features from content representations, we design the third component that introduces dynamic prediction. T2RD achieves State-of-the-Art (SOTA) generalization performance and sample efficiency in the DeepMind Control Suite and Robotic Manipulation tasks. Youfang Lin, Shuo Wang 0031, Kai Lv 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | Infer the Whole from a Glimpse of a Part: Keypoint-Based Knowledge Graph for Vehicle Re-IdentificationabstractVehicle re-identification aims to match vehicles across non-overlapping camera views. Many existing methods extract features from one specific image, and these methods lack view-invariance when comparing vehicles of different orientations. As a result, discriminative parts obscured by viewpoint changes cannot contribute effectively to matching. This work presents a novel keypoint-based framework for vehicle Re-ID. We propose to explicitly model the intrinsic structural relationships between vehicle components via knowledge graph. By establishing connection between keypoints, our approach aims to leverage such prior to match vehicles even when some parts are not directly comparable due to orientation inconsistencies. Specifically, given query and gallery images, we first detect visible keypoints. Then, a transformer-based model infers features for non-overlapped keypoints by conditioning on visible correspondences defined in the knowledge graph. The final representation integrates visible and inferred features. Extensive experiments demonstrate our method outperforms state-of-the-arts on standard benchmarks under cross-view matching scenarios. To our knowledge, this is the first work introducing structural priors via keypoint knowledge graphs for view-invariant vehicle re-identification. Kai Lv 0002, Shuo Wang 0031, Sheng Han 0001, Youfang Lin |
AAAI | 4 |
| 2025 | CoDe: Communication Delay-Tolerant Multi-Agent Collaboration via Dual Alignment of Intent and TimelinessabstractCommunication has been widely employed to enhance multi-agent collaboration. Previous research has typically assumed delay-free communication, a strong assumption that is challenging to meet in practice. However, real-world agents suffer from channel delays, receiving messages sent at different time points, termed Asynchronous Communication, leading to cognitive biases and breakdowns in collaboration. This paper first defines two communication delay settings in MARL and emphasizes their harm to collaboration. To handle the above delays, this paper proposes a novel framework, Communication Delay-Tolerant Multi-Agent Collaboration (CoDe). At first, CoDe learns an intent representation as messages through future action inference, reflecting the stable future behavioral trends of the agents. Then, CoDe devises a dual alignment mechanism of intent and timeliness to strengthen the fusion process of asynchronous messages. In this way, agents can extract the long-term intent of others, even from delayed messages, and selectively utilize the most recent messages that are relevant to their intent. Experimental results demonstrate that CoDe outperforms baseline algorithms in three MARL benchmarks without delay and exhibits robustness under fixed and time-varying delays. Shoucheng Song, Youfang Lin, Sheng Han 0001, Hao Wu 0010, Shuo Wang 0031, Kai Lv 0002 |
AAAI | 6 |
| 2025 | Improving Monotonic Optimization in Heterogeneous Multi-agent Reinforcement Learning with Optimal Marginal Deterministic Policy Gradient
Youfang Lin, Shuo Wang 0031, Sheng Han 0001 |
ICANN (1) | 3 |
| 2025 | From Pixels to Temporal Correlations: Learning Informative Representations for Reinforcement Learning Pre-trainingabstractUnsupervised pre-training on large-scale datasets has demonstrated significant potential for improving the sample efficiency and performance of Reinforcement Learning (RL). Given the large-scale action-free internet videos, existing methods utilize single-step transition prediction and image reconstruction to learn representations. However, these methods prefer to preserve large-proportion stationary information in the pixel space, neglecting small but crucial information. To preserve enough information in the representation, it is essential to pay equal attention to each element in videos. Specifically, we propose a temporal correlation space to distinguish each element. For implementation, we introduce the Multi-scale Temporal Contrastive Learning (MTCL) method to model multi-scale temporal correlations separately. This approach can balance the attention of different elements and yield more informative representations, effectively supporting policy learning in various downstream tasks. Experimental results demonstrate that our method improves sample efficiency and asymptotic performance across various downstream tasks. Youfang Lin, Sheng Han 0001, Shuo Wang 0031, Kai Lv 0002 |
ACM Multimedia | 6 |
| 2025 | Learning Robust Representations via Bidirectional Transition for Visual Reinforcement LearningabstractVisual reinforcement learning has exhibited efficacy in solving control tasks characterized by high-dimensional observations. However, a central challenge persists in deriving dependable and generalizable representations from vision-based observations. Inspired by the human thought process, when the visual representation extracted from the observation can predict the future and trace history, the representation is reliable and accurate in comprehending the environmental state. Based on this concept, we introduce a B idirectional T ransition (BT) framework for representation learning. This framework employs the bidirectional prediction of both forward and backward environmental transitions as auxiliary tasks to extract reliable representations. Additionally, we introduce an inverse dynamic model to predict the actions causing environmental state transitions, thereby learning the task relevance of state representations. Our method demonstrates competitive generalization performance and sample efficiency in two settings in the DeepMind Control suite. Moreover, we utilize the robotic manipulation simulator, autonomous driving simulator CARLA, and visual navigation simulator Habitat to demonstrate the wide applicability of our method. The results indicate that BT offers more stable and reliable representations and exhibits robust generalization performance for visual reinforcement learning tasks. Youfang Lin, Shuo Wang 0031, Hehe Fan, Kai Lv 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | Off-Policy Conservative Distributional Reinforcement Learning With Safety ConstraintsabstractSafe exploration can be regarded as a constrained Markov decision problem (CMDP) where the expected long-term cost is constrained. Previous off-policy algorithms convert the constrained optimization problem into the corresponding unconstrained dual problem by introducing the Lagrangian relaxation technique. However, the cost function of the above algorithms provides inaccurate estimations and causes the instability of the Lagrange multiplier learning. In this article, we present a novel off-policy reinforcement learning (RL) algorithm called conservative distributional maximum a posteriori policy optimization (CDMPO). At first, to accurately judge whether the current situation satisfies the constraints, CDMPO adapts distributional RL method to estimate the Q-function and C-function. Then, CDMPO uses a conservative value function loss to reduce the number of violations of constraints during the exploration process. In addition, we utilize adaptive proportional integral derivative (APID) to update the Lagrange multiplier stably. In our experiments, we select eight representative constrained tasks from two well-known safe RL benchmarks (Safety Gym and Bullet Safety Gym), providing a comprehensive evaluation of our methods across diverse scenarios. Empirical results show that the proposed method has fewer violations of constraints in the early exploration process. The final test results also illustrate that our method has better-risk control capabilities. Youfang Lin, Sheng Han 0001, Shuo Wang 0031, Kai Lv 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2024 | What Effects the Generalization in Visual Reinforcement Learning: Policy Consistency with Truncated Return PredictionabstractIn visual Reinforcement Learning (RL), the challenge of generalization to new environments is paramount. This study pioneers a theoretical analysis of visual RL generalization, establishing an upper bound on the generalization objective, encompassing policy divergence and Bellman error components. Motivated by this analysis, we propose maintaining the cross-domain consistency for each policy in the policy space, which can reduce the divergence of the learned policy during the test. In practice, we introduce the Truncated Return Prediction (TRP) task, promoting cross-domain policy consistency by predicting truncated returns of historical trajectories. Moreover, we also propose a Transformer-based predictor for this auxiliary task. Extensive experiments on DeepMind Control Suite and Robotic Manipulation tasks demonstrate that TRP achieves state-of-the-art generalization performance. We further demonstrate that TRP outperforms previous methods in terms of sample efficiency during training. Shuo Wang 0031, Zhihao Wu 0001, Youfang Lin, Kai Lv 0002 |
AAAI | 1 |
| 2024 | How to Learn Domain-Invariant Representations for Visual Reinforcement Learning: An Information-Theoretical Perspective
Shuo Wang 0031, Zhihao Wu 0001, Youfang Lin, Kai Lv 0002 |
IJCAI | 1 |
| 2024 | PGN: The RNN's New Successor is Effective for Long-Range Time Series ForecastingabstractDue to the recurrent structure of RNN, the long information propagation path poses limitations in capturing long-term dependencies, gradient explosion/vanishing issues, and inefficient sequential execution. Based on this, we propose a novel paradigm called Parallel Gated Network (PGN) as the new successor to RNN. PGN directly captures information from previous time steps through the designed Historical Information Extraction (HIE) layer and leverages gated mechanisms to select and fuse it with the current time step information. This reduces the information propagation path to $\mathcal{O}(1)$, effectively addressing the limitations of RNN. To enhance PGN's performance in long-range time series forecasting tasks, we propose a novel temporal modeling framework called Temporal PGN (TPGN). TPGN incorporates two branches to comprehensively capture the semantic information of time series. One branch utilizes PGN to capture long-term periodic patterns while preserving their local characteristics. The other branch employs patches to capture short-term information and aggregate the global representation of the series. TPGN achieves a theoretical complexity of $\mathcal{O}(\sqrt{L})$, ensuring efficiency in its operations. Experimental results on five benchmark datasets demonstrate the state-of-the-art (SOTA) performance and high efficiency of TPGN, further confirming the effectiveness of PGN as the new successor to RNN in long-range time series forecasting. The code is available in this repository: https://github.com/Water2sea/TPGN. Yuxin Jia, Youfang Lin, Shuo Wang 0031, Huaiyu Wan |
NeurIPS | 4 |
| 2024 | Agent-Centric Relation Graph for Object Visual NavigationabstractObject visual navigation aims to steer an agent toward a target object based on visual observations. It is highly desirable to reasonably perceive the environment and accurately control the agent. In the navigation task, we introduce an Agent-Centric Relation Graph (ACRG) for learning the visual representation based on the relationships in the environment. ACRG is a highly effective structure that consists of two relationships, i.e., the horizontal relationship among objects and the distance relationship between the agent and objects. On the one hand, we design the Object Horizontal Relationship Graph (OHRG) that stores the relative horizontal location among objects. On the other hand, we propose the Agent-Target Distance Relationship Graph (ATDRG) that enables the agent to perceive the distance between the target and objects. For ATDRG, we utilize image depth to obtain the target distance and imply the vertical location to capture the distance relationship among objects in the vertical direction. With the above graphs, the agent can perceive the environment and output navigation actions. Experimental results in the artificial environment AI2-THOR demonstrate that ACRG significantly outperforms other state-of-the-art methods in unseen testing environments. Youfang Lin, Shuo Wang 0031, Zhihao Wu 0001, Kai Lv 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Building Category Graphs Representation with Spatial and Temporal Attention for Visual NavigationabstractGiven an object of interest, visual navigation aims to reach the object’s location based on a sequence of partial observations. To this end, an agent needs to (1) acquire specific knowledge about the relations of object categories in the world during training and (2) locate the target object based on the pre-learned object category relations and its trajectory in the current unseen environment. In this article, we propose a Category Relation Graph (CRG) to learn the knowledge of object category layout relations and a Temporal-Spatial-Region attention (TSR) architecture to perceive the long-term spatial-temporal dependencies of objects, aiding navigation. We establish CRG to learn prior knowledge of object layout and deduce the positions of specific objects. Subsequently, we propose the TSR architecture to capture relationships among objects in temporal, spatial, and regions within observation trajectories. Specifically, we implement a Temporal attention module (T) to model the temporal structure of the observation sequence, implicitly encoding historical moving or trajectory information. Then, a Spatial attention module (S) uncovers the spatial context of the current observation objects based on CRG and past observations. Last, a Region attention module (R) shifts the attention to the target-relevant region. Leveraging the visual representation extracted by our method, the agent accurately perceives the environment and easily learns a superior navigation policy. Experiments on AI2-THOR demonstrate that our CRG-TSR method significantly outperforms existing methods in both effectiveness and efficiency. The supplementary material includes the code and will be publicly available. Youfang Lin, Hehe Fan, Shuo Wang 0031, Zhihao Wu 0001, Kai Lv 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Spatially-Regularized Features for Vehicle Re-Identification: An Explanation of Where Deep Models Should FocusabstractVehicle re-identification aims to identify vehicles from different cameras and has drawn much attention in the multimedia community. In recent years, significant achievements in vehicle re-identification have been made due to the development of neural networks and deep feature representations. However, existing deep models are regarded as “black box” methods considering the lack of explaining where they focus. With Class Activation Mapping (CAM), we can localize the discriminative image regions by utilizing convolutional feature maps. After answering the “where” question, the question of “how” to generate reasonable features appears to be substantial. In this paper, we propose the novel Spatially-Regularized Features (SRF) that can be extracted from discriminative regions in an explainable way. Specifically, we first provide an evaluation mechanism called Peak-to-Sidelobe Ratio (PSR) to measure the distribution of the convolutional feature maps. PSR outputs the strength of a matrix peak and can be used to indicate the attention intensity of a specific region. Moreover, we propose a spatially regularized loss to make the deep models focus on more reasonable and discriminative image regions. Note that no additional manual annotation data is involved in the training process, making the SRF an efficient and effective approach. Extensive subjective and objective experiments show that the proposed method significantly outperforms the state-of-the-art methods on three large-scale vehicle re-identification datasets. Kai Lv 0002, Shuo Wang 0031, Youfang Lin |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Skill-Based Hierarchical Reinforcement Learning for Target Visual NavigationabstractTarget visual navigation aims at controlling the agent to find a target object based on a monocular visual RGB image in each step. It is crucial for the agent to adapt to new environments. As target visual navigation is a complex task, understanding the behavior of the agent is beneficial for analyzing the reasons for failure. This work focuses on improving the readability and success rate of navigation policies. In this paper, we propose a framework named Skill-based Hierarchical Reinforcement Learning (SHRL) for target visual navigation. SHRL contains a high-level policy and three low-level skills. The high-level policy accomplishes the task by utilizing or stopping low-level skills at each step. Low-level skills are designed to separately solve three sub-tasks, i.e.,Search, Adjustment, andExploration. In addition, we propose an Abstract Representation and two penalty items to feed robust features to the high-level policy. Abstract Representation is designed to focus on selecting low-level skills rather than the details of navigation. Experimental results in the artificial environment AI2-Thor indicate that the proposed method outperforms state-of-the-art by a large margin in unseen indoor environments. Moreover, we also provide case studies to illustrate the advantages of SHRL. Shuo Wang 0031, Zhihao Wu 0001, Youfang Lin, Kai Lv 0002 |
IEEE Trans. Multim. | 1 |