EDBT 2026 Demo / reviewers in the wild / expert
Kai Lv 0002
dblp:191/2440-2
· DBLP profile ↗
33ranked-venue papers
4as first author
29since 2021 · last 2026
0000-0001-6533-5176ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 3 first-author · 17 since 2021Artificial intelligence and machine learning · 17 · 1 first-author · 16 since 2021Computer networks · 5 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Think How Your Teammates Think: Active Inference Can Benefit Decentralized ExecutionabstractIn multi-agent systems, explicit cognition of teammates' decision logic serves as a critical factor in facilitating coordination. Communication (i.e., "Tell") can assist in the cognitive development process by information dissemination, yet it is inevitably subject to real-world constraints such as noise, latency, and attacks. Therefore, building the understanding of teammates' decisions without communication remains challenging. To address this, we propose a novel non-communication MARL framework that realizes the construction of cognition through local observation-based modeling (i.e., "Think"). Our framework enables agents to model teammates' active inference process. At first, the proposed method produces three teammate portraits: perception-belief-action. Specifically, we model the teammate's decision process as follows: 1) Perception: observing environments; 2) Belief: forming beliefs; 3) Action: making decisions. Then, we selectively integrate the belief portrait into the decision process based on the accuracy and relevance of the perception portrait. This enables the selection of cooperative teammates and facilitates effective collaboration. Extensive experiments on the SMAC, SMACv2, MPE, and GRF benchmarks demonstrate the superior performance of our method. Hao Wu 0010, Shoucheng Song, Sheng Han 0001, Huaiyu Wan, Youfang Lin, Kai Lv 0002 |
AAAI | 7 |
| 2026 | Bidirectional transition consistency between multi-domain observations for visual reinforcement learning generalization
Youfang Lin, Shuo Wang 0031, Hehe Fan, Kai Lv 0002 |
Neural Networks | 7 |
| 2026 | Task-Relevant Representation Decoupling for Visual Reinforcement Learning GeneralizationabstractVisual Reinforcement Learning (VRL) has achieved considerable success in solving control tasks. However, generalizing learned policies to new environments remains a major challenge, as agents often overfit to task-irrelevant features in the training environment. To solve this problem, we introduce the concept of decoupling observations into task-relevant and task-irrelevant representations. Building on this idea, we propose a self-supervised T ask- R elevant R epresentation D ecoupling (T2RD) algorithm for VRL. This algorithm consists of three components: task-relevant representation consistency , cross-reconstruction , and cross-dynamic prediction . The first two components achieve the decoupling of content and style features, but the resulting content representations are not necessarily task-relevant. To further refine task-relevant features from content representations, we design the third component that introduces dynamic prediction. T2RD achieves State-of-the-Art (SOTA) generalization performance and sample efficiency in the DeepMind Control Suite and Robotic Manipulation tasks. Youfang Lin, Shuo Wang 0031, Kai Lv 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2025 | Infer the Whole from a Glimpse of a Part: Keypoint-Based Knowledge Graph for Vehicle Re-IdentificationabstractVehicle re-identification aims to match vehicles across non-overlapping camera views. Many existing methods extract features from one specific image, and these methods lack view-invariance when comparing vehicles of different orientations. As a result, discriminative parts obscured by viewpoint changes cannot contribute effectively to matching. This work presents a novel keypoint-based framework for vehicle Re-ID. We propose to explicitly model the intrinsic structural relationships between vehicle components via knowledge graph. By establishing connection between keypoints, our approach aims to leverage such prior to match vehicles even when some parts are not directly comparable due to orientation inconsistencies. Specifically, given query and gallery images, we first detect visible keypoints. Then, a transformer-based model infers features for non-overlapped keypoints by conditioning on visible correspondences defined in the knowledge graph. The final representation integrates visible and inferred features. Extensive experiments demonstrate our method outperforms state-of-the-arts on standard benchmarks under cross-view matching scenarios. To our knowledge, this is the first work introducing structural priors via keypoint knowledge graphs for view-invariant vehicle re-identification. Kai Lv 0002, Shuo Wang 0031, Sheng Han 0001, Youfang Lin |
AAAI | 1 |
| 2025 | CoDe: Communication Delay-Tolerant Multi-Agent Collaboration via Dual Alignment of Intent and TimelinessabstractCommunication has been widely employed to enhance multi-agent collaboration. Previous research has typically assumed delay-free communication, a strong assumption that is challenging to meet in practice. However, real-world agents suffer from channel delays, receiving messages sent at different time points, termed Asynchronous Communication, leading to cognitive biases and breakdowns in collaboration. This paper first defines two communication delay settings in MARL and emphasizes their harm to collaboration. To handle the above delays, this paper proposes a novel framework, Communication Delay-Tolerant Multi-Agent Collaboration (CoDe). At first, CoDe learns an intent representation as messages through future action inference, reflecting the stable future behavioral trends of the agents. Then, CoDe devises a dual alignment mechanism of intent and timeliness to strengthen the fusion process of asynchronous messages. In this way, agents can extract the long-term intent of others, even from delayed messages, and selectively utilize the most recent messages that are relevant to their intent. Experimental results demonstrate that CoDe outperforms baseline algorithms in three MARL benchmarks without delay and exhibits robustness under fixed and time-varying delays. Shoucheng Song, Youfang Lin, Sheng Han 0001, Hao Wu 0010, Shuo Wang 0031, Kai Lv 0002 |
AAAI | 7 |
| 2025 | Enhancing Offline Safe Reinforcement Learning with Trajectory-Constrained Diffusion Planning
Youfang Lin, Shuo Shen 0002, Hanfeng Lin, Peng Cheng 0013, Sheng Han 0001, Kai Lv 0002 |
AAMAS | 7 |
| 2025 | Continuous Diffusive Prediction Network for Multi-Station Weather PredictionabstractMulti-station weather prediction provides weather forecasts for specific geographical locations, playing an important role in various aspects of daily life. Existing methods consider the relationships between individual stations discretely, making it difficult to model the continuous spatiotemporal processes of atmospheric motion, which results in suboptimal prediction outcomes. This paper proposes the Continuous Diffusive Prediction Network (CDPNet) to model the real-world continuous weather change process from discrete station observation data. CDPNet consists of two core modules: the Continuous Calibrated Initialization (CCI) and the Diffusive Difference Estimation (DDE). The CCI module interpolates data between observation stations to construct a spatially continuous physical field and ensures temporal continuity by integrating directional information from a global perspective. It accurately represents the current physical state and provides a foundation for future weather prediction. Moreover, the DDE module explicitly captures the spatial diffusion process and estimates the diffusive differences between consecutive time steps, effectively modeling spatio-temporally continuous atmospheric motion. Likewise, directional information on weather changes is introduced from the entire historical series to mitigate estimation uncertainty and improve the performance of weather prediction. Extensive experiments on the Weather2K and Global Wind/Temp datasets demonstrate that CDPNet outperforms state-of-the-art models. Chujie Xu, Yuqing Ma, Haoyuan Deng, Yajun Gao, Yudie Wang, Kai Lv 0002, Xianglong Liu 0001 |
IJCAI | 6 |
| 2025 | From General Relation Patterns to Task-Specific Decision-Making in Continual Multi-Agent CoordinationabstractContinual Multi-Agent Reinforcement Learning (Co-MARL) requires agents to address catastrophic forgetting issues while learning new coordination policies with the dynamics team. In this paper, we delve into the core of Co-MARL, namely Relation Patterns, which refer to agents’ general understanding of interactions. In addition to generality, relation patterns exhibit task-specificity when mapped to different action spaces. To this end, we propose a novel method called General Relation Patterns-Guided Task-specific Decision-Maker (RPG). In RPG, agents extract relation patterns from dynamic observation spaces using a relation capturer. These task-agnostic relation patterns are then mapped to different action spaces via a task-specific decision-maker generated by a conditional hypernetwork. To combat forgetting, we further introduce regularization items on both the relation capturer and the conditional hypernetwork. Results on SMAC and LBF demonstrate that RPG effectively prevents catastrophic forgetting when learning new tasks and achieves zero-shot generalization to unseen tasks. Youfang Lin, Shoucheng Song, Hao Wu 0010, Yuqing Ma, Sheng Han 0001, Kai Lv 0002 |
IJCAI | 7 |
| 2025 | From Pixels to Temporal Correlations: Learning Informative Representations for Reinforcement Learning Pre-trainingabstractUnsupervised pre-training on large-scale datasets has demonstrated significant potential for improving the sample efficiency and performance of Reinforcement Learning (RL). Given the large-scale action-free internet videos, existing methods utilize single-step transition prediction and image reconstruction to learn representations. However, these methods prefer to preserve large-proportion stationary information in the pixel space, neglecting small but crucial information. To preserve enough information in the representation, it is essential to pay equal attention to each element in videos. Specifically, we propose a temporal correlation space to distinguish each element. For implementation, we introduce the Multi-scale Temporal Contrastive Learning (MTCL) method to model multi-scale temporal correlations separately. This approach can balance the attention of different elements and yield more informative representations, effectively supporting policy learning in various downstream tasks. Experimental results demonstrate that our method improves sample efficiency and asymptotic performance across various downstream tasks. Youfang Lin, Sheng Han 0001, Shuo Wang 0031, Kai Lv 0002 |
ACM Multimedia | 7 |
| 2025 | Conflict-Aware Knowledge Editing in the Wild: Semantic-Augmented Graph Representation for Unstructured TextabstractLarge Language Models (LLMs) have demonstrated broad applications but suffer from issues like hallucinations, erroneous outputs and outdated knowledge. Model editing emerges as an effective solution to refine knowledge in LLMs, yet existing methods typically depend on structured knowledge representations.
However, real-world knowledge is primarily embedded within complex, unstructured text. Existing structured knowledge editing approaches face significant challenges when handling the entangled and intricate knowledge present in unstructured text, resulting in issues such as representation ambiguity and editing conflicts.
To address these challenges, we propose a Conflict-Aware Knowledge Editing in the Wild (CAKE) framework, the first framework explicitly designed for editing knowledge extracted from wild unstructured text.
CAKE comprises two core components: a Semantic-augmented Graph Representation module and a Conflict-aware Knowledge Editing strategy. The Semantic-augmented Graph Representation module enhances knowledge encoding through structural disambiguation, relational enrichment, and semantic diversification. Meanwhile, the Conflict-aware Knowledge Editing strategy utilizes a graph-theoretic coloring algorithm to disentangle conflicted edits by allocating them to orthogonal parameter subspaces, thereby effectively mitigating editing conflicts. Experimental results on the AKEW benchmark demonstrate that CAKE significantly outperforms existing methods, achieving a 15.43\% improvement in accuracy on llama3 editing tasks. Our framework successfully bridges the gap between unstructured textual knowledge and reliable model editing, enabling more robust and scalable updates for practical LLM applications. Zhange Zhang, Zhicheng Geng, Yuqing Ma, Kai Lv 0002, Xianglong Liu 0001 |
NeurIPS | 5 |
| 2025 | Multi-constraint reinforcement learning in complex robot environments
Sheng Han 0001, Hao Wu 0010, Youfang Lin, Kai Lv 0002 |
Frontiers Comput. Sci. | 5 |
| 2025 | Learning Robust Representations via Bidirectional Transition for Visual Reinforcement LearningabstractVisual reinforcement learning has exhibited efficacy in solving control tasks characterized by high-dimensional observations. However, a central challenge persists in deriving dependable and generalizable representations from vision-based observations. Inspired by the human thought process, when the visual representation extracted from the observation can predict the future and trace history, the representation is reliable and accurate in comprehending the environmental state. Based on this concept, we introduce a B idirectional T ransition (BT) framework for representation learning. This framework employs the bidirectional prediction of both forward and backward environmental transitions as auxiliary tasks to extract reliable representations. Additionally, we introduce an inverse dynamic model to predict the actions causing environmental state transitions, thereby learning the task relevance of state representations. Our method demonstrates competitive generalization performance and sample efficiency in two settings in the DeepMind Control suite. Moreover, we utilize the robotic manipulation simulator, autonomous driving simulator CARLA, and visual navigation simulator Habitat to demonstrate the wide applicability of our method. The results indicate that BT offers more stable and reliable representations and exhibits robust generalization performance for visual reinforcement learning tasks. Youfang Lin, Shuo Wang 0031, Hehe Fan, Kai Lv 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2025 | Off-Policy Conservative Distributional Reinforcement Learning With Safety ConstraintsabstractSafe exploration can be regarded as a constrained Markov decision problem (CMDP) where the expected long-term cost is constrained. Previous off-policy algorithms convert the constrained optimization problem into the corresponding unconstrained dual problem by introducing the Lagrangian relaxation technique. However, the cost function of the above algorithms provides inaccurate estimations and causes the instability of the Lagrange multiplier learning. In this article, we present a novel off-policy reinforcement learning (RL) algorithm called conservative distributional maximum a posteriori policy optimization (CDMPO). At first, to accurately judge whether the current situation satisfies the constraints, CDMPO adapts distributional RL method to estimate the Q-function and C-function. Then, CDMPO uses a conservative value function loss to reduce the number of violations of constraints during the exploration process. In addition, we utilize adaptive proportional integral derivative (APID) to update the Lagrange multiplier stably. In our experiments, we select eight representative constrained tasks from two well-known safe RL benchmarks (Safety Gym and Bullet Safety Gym), providing a comprehensive evaluation of our methods across diverse scenarios. Empirical results show that the proposed method has fewer violations of constraints in the early exploration process. The final test results also illustrate that our method has better-risk control capabilities. Youfang Lin, Sheng Han 0001, Shuo Wang 0031, Kai Lv 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2024 | What Effects the Generalization in Visual Reinforcement Learning: Policy Consistency with Truncated Return PredictionabstractIn visual Reinforcement Learning (RL), the challenge of generalization to new environments is paramount. This study pioneers a theoretical analysis of visual RL generalization, establishing an upper bound on the generalization objective, encompassing policy divergence and Bellman error components. Motivated by this analysis, we propose maintaining the cross-domain consistency for each policy in the policy space, which can reduce the divergence of the learned policy during the test. In practice, we introduce the Truncated Return Prediction (TRP) task, promoting cross-domain policy consistency by predicting truncated returns of historical trajectories. Moreover, we also propose a Transformer-based predictor for this auxiliary task. Extensive experiments on DeepMind Control Suite and Robotic Manipulation tasks demonstrate that TRP achieves state-of-the-art generalization performance. We further demonstrate that TRP outperforms previous methods in terms of sample efficiency during training. Shuo Wang 0031, Zhihao Wu 0001, Youfang Lin, Kai Lv 0002 |
AAAI | 6 |
| 2024 | Enhancing Off-Policy Constrained Reinforcement Learning through Adaptive Ensemble C EstimationabstractIn the domain of real-world agents, the application of Reinforcement Learning (RL) remains challenging due to the necessity for safety constraints. Previously, Constrained Reinforcement Learning (CRL) has predominantly focused on on-policy algorithms. Although these algorithms exhibit a degree of efficacy, their interactivity efficiency in real-world settings is sub-optimal, highlighting the demand for more efficient off-policy methods. However, off-policy CRL algorithms grapple with challenges in precise estimation of the C-function, particularly due to the fluctuations in the constrained Lagrange multiplier. Addressing this gap, our study focuses on the nuances of C-value estimation in off-policy CRL and introduces the Adaptive Ensemble C-learning (AEC) approach to reduce these inaccuracies. Building on state-of-the-art off-policy algorithms, we propose AEC-based CRL algorithms designed for enhanced task optimization. Extensive experiments on nine constrained robotics tasks reveal the superior interaction efficiency and performance of our algorithms in comparison to preceding methods. Youfang Lin, Shuo Shen 0002, Sheng Han 0001, Kai Lv 0002 |
AAAI | 5 |
| 2024 | How to Learn Domain-Invariant Representations for Visual Reinforcement Learning: An Information-Theoretical Perspective
Shuo Wang 0031, Zhihao Wu 0001, Youfang Lin, Kai Lv 0002 |
IJCAI | 6 |
| 2024 | A Viewpoint-aware Channel Selection Method for Vehicle Re-identificationabstractVehicle re-identification (Re-ID) aims to identify the same vehicle images from various views in the cross-camera scenario. There are many challenges with the current vehicle Re-ID task, e.g., illumination, viewpoint, and resolution. One of the most common problems is the large difference in vehicle appearance caused by viewpoint. Vehicle Re-ID networks generally focus on discriminative areas (e.g., vehicle logo and vehicle inspections). However, the specific regions are different on different viewpoints, leading to a large intra-class variation on different viewpoints. In this paper, we propose a viewpoint-aware channel selection method (VCS) to solve this problem by exploring channels with distinguished categories on the network. Specifically, we discover the most active channels in each viewpoint. Based on features processed by these active channels, we calculate the distance of vehicle features by common active channels of two viewpoints. In addition, our method is a simple post-processing method in vehicle Re-ID, making it applicable to existing feature representation methods. Experimental results on several large vehicle Re-ID datasets illustrate the effectiveness of our method, especially on the difficult VERI-Wild dataset. Youfang Lin, Kai Lv 0002 |
IJCNN | 4 |
| 2024 | Agent-Centric Relation Graph for Object Visual NavigationabstractObject visual navigation aims to steer an agent toward a target object based on visual observations. It is highly desirable to reasonably perceive the environment and accurately control the agent. In the navigation task, we introduce an Agent-Centric Relation Graph (ACRG) for learning the visual representation based on the relationships in the environment. ACRG is a highly effective structure that consists of two relationships, i.e., the horizontal relationship among objects and the distance relationship between the agent and objects. On the one hand, we design the Object Horizontal Relationship Graph (OHRG) that stores the relative horizontal location among objects. On the other hand, we propose the Agent-Target Distance Relationship Graph (ATDRG) that enables the agent to perceive the distance between the target and objects. For ATDRG, we utilize image depth to obtain the target distance and imply the vertical location to capture the distance relationship among objects in the vertical direction. With the above graphs, the agent can perceive the environment and output navigation actions. Experimental results in the artificial environment AI2-THOR demonstrate that ACRG significantly outperforms other state-of-the-art methods in unseen testing environments. Youfang Lin, Shuo Wang 0031, Zhihao Wu 0001, Kai Lv 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Building Category Graphs Representation with Spatial and Temporal Attention for Visual NavigationabstractGiven an object of interest, visual navigation aims to reach the object’s location based on a sequence of partial observations. To this end, an agent needs to (1) acquire specific knowledge about the relations of object categories in the world during training and (2) locate the target object based on the pre-learned object category relations and its trajectory in the current unseen environment. In this article, we propose a Category Relation Graph (CRG) to learn the knowledge of object category layout relations and a Temporal-Spatial-Region attention (TSR) architecture to perceive the long-term spatial-temporal dependencies of objects, aiding navigation. We establish CRG to learn prior knowledge of object layout and deduce the positions of specific objects. Subsequently, we propose the TSR architecture to capture relationships among objects in temporal, spatial, and regions within observation trajectories. Specifically, we implement a Temporal attention module (T) to model the temporal structure of the observation sequence, implicitly encoding historical moving or trajectory information. Then, a Spatial attention module (S) uncovers the spatial context of the current observation objects based on CRG and past observations. Last, a Region attention module (R) shifts the attention to the target-relevant region. Leveraging the visual representation extracted by our method, the agent accurately perceives the environment and easily learns a superior navigation policy. Experiments on AI2-THOR demonstrate that our CRG-TSR method significantly outperforms existing methods in both effectiveness and efficiency. The supplementary material includes the code and will be publicly available. Youfang Lin, Hehe Fan, Shuo Wang 0031, Zhihao Wu 0001, Kai Lv 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2023 | Graph-Based Vehicle Keypoint Attention Model for Vehicle Re-identification
Zhihao Wu 0001, Youfang Lin, Kai Lv 0002 |
ICONIP (13) | 4 |
| 2023 | Multi-mobile Object Motion Coordination with Reinforcement Learning
Shanhua Yuan, Sheng Han 0001, Xiwen Jiang, Youfang Lin, Kai Lv 0002 |
ICONIP (8) | 5 |
| 2023 | Offline Reinforcement Learning with Diffusion-Based Behavior Cloning Term
Youfang Lin, Sheng Han 0001, Kai Lv 0002 |
KSEM (4) | 4 |
| 2023 | A lightweight and style-robust neural network for autonomous driving in end side devicesabstractThe autonomous driving algorithm studied in this paper makes a ground vehicle capable of sensing its environment via visual images and moving safely with little or no human input. Due to the limitation of the computing power of end side devices, the autonomous driving algorithm should adopt a lightweight model and have high performance. Conditional imitation learning has been proved an efficient and promising policy for autonomous driving and other applications on end side devices due to its high performance and offline characteristics. In driving scenarios, the images captured in different weathers have different styles, which are influenced by various interference factors, such as illumination, raindrops, etc. These interference factors bring challenges to the perception ability of deep models, thus affecting the decision-making process in autonomous driving. The first contribution of this paper is to investigate the performance gap of driving models under different weather conditions. Following the investigation, we utilise StarGAN-V2 to translate images from source domains into the target clear sunset domain. Based on the images translated by StarGAN-V2, we propose Conditional Imitation Learning with ResNet backbone named Star-CILRS. The proposed method is able to convert an image to multiple styles using only one single model, making our method easier to deploy on end side devices. Visualization results show that Star-CILRS can eliminate some environmental interference factors. Our method outperforms other methods and the success rate values in different tasks are 98%, 74%, and 22%, respectively. Sheng Han 0001, Youfang Lin, Zhihui Guo, Kai Lv 0002 |
Connect. Sci. | 4 |
| 2023 | Spatially-Regularized Features for Vehicle Re-Identification: An Explanation of Where Deep Models Should FocusabstractVehicle re-identification aims to identify vehicles from different cameras and has drawn much attention in the multimedia community. In recent years, significant achievements in vehicle re-identification have been made due to the development of neural networks and deep feature representations. However, existing deep models are regarded as “black box” methods considering the lack of explaining where they focus. With Class Activation Mapping (CAM), we can localize the discriminative image regions by utilizing convolutional feature maps. After answering the “where” question, the question of “how” to generate reasonable features appears to be substantial. In this paper, we propose the novel Spatially-Regularized Features (SRF) that can be extracted from discriminative regions in an explainable way. Specifically, we first provide an evaluation mechanism called Peak-to-Sidelobe Ratio (PSR) to measure the distribution of the convolutional feature maps. PSR outputs the strength of a matrix peak and can be used to indicate the attention intensity of a specific region. Moreover, we propose a spatially regularized loss to make the deep models focus on more reasonable and discriminative image regions. Note that no additional manual annotation data is involved in the training process, making the SRF an efficient and effective approach. Extensive subjective and objective experiments show that the proposed method significantly outperforms the state-of-the-art methods on three large-scale vehicle re-identification datasets. Kai Lv 0002, Shuo Wang 0031, Youfang Lin |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Skill-Based Hierarchical Reinforcement Learning for Target Visual NavigationabstractTarget visual navigation aims at controlling the agent to find a target object based on a monocular visual RGB image in each step. It is crucial for the agent to adapt to new environments. As target visual navigation is a complex task, understanding the behavior of the agent is beneficial for analyzing the reasons for failure. This work focuses on improving the readability and success rate of navigation policies. In this paper, we propose a framework named Skill-based Hierarchical Reinforcement Learning (SHRL) for target visual navigation. SHRL contains a high-level policy and three low-level skills. The high-level policy accomplishes the task by utilizing or stopping low-level skills at each step. Low-level skills are designed to separately solve three sub-tasks, i.e.,Search, Adjustment, andExploration. In addition, we propose an Abstract Representation and two penalty items to feed robust features to the high-level policy. Abstract Representation is designed to focus on selecting low-level skills rather than the details of navigation. Experimental results in the artificial environment AI2-Thor indicate that the proposed method outperforms state-of-the-art by a large margin in unseen indoor environments. Moreover, we also provide case studies to illustrate the advantages of SHRL. Shuo Wang 0031, Zhihao Wu 0001, Youfang Lin, Kai Lv 0002 |
IEEE Trans. Multim. | 5 |
| 2022 | Local perspective based synthesis for vehicle re-identification: A transformation state adversarial method
Yanbing Chen, Wei Ke 0001, Hong Lin 0006, Chan-Tong Lam, Kai Lv 0002, Hao Sheng 0001, Zhang Xiong 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2021 | High Confidence Attribute Recognition For Vehicle Re-IdentificationabstractVehicle re-identification aims to associate images or videos of the same vehicle collected from different cameras. Many existing methods address the vehicle re-identification problem by explicitly learning distinguishable global features. However, vehicle attributes, i.e., logo category and orientation, play an indispensable role in identifying vehicles. In this paper, we first propose deep models to recognize vehicle attributes. Then, based on these attributes, we adopt a High Confidence Attribute Network (HCANet) to extract weighted global features. A comprehensive evaluation on the VehicleID dataset shows that our approach achieves competitive results. Xinze Dou, Yang Liu 0088, Kai Lv 0002, Zhang Xiong 0001, Hao Sheng 0001 |
ICIP | 3 |
| 2021 | Combining Pose Invariant and Discriminative Features for Vehicle ReidentificationabstractVehicle reidentification, aiming at identifying vehicles across images, has drawn a lot of attention and has made significant achievements in recent years. However, vehicle reidentification remains a challenging task caused by severe appearance changes due to different orientations. In practice, the result of reidentification is greatly influenced by the pose of vehicles, and we call this influence as a pose barrier problem. One way to address the pose barrier problem is to train a feature representation that is invariant for various vehicle poses. To this end, we present pose robust features (PRFs) that contains two components: 1) pose-invariant features (PIFs) and 2) pose discriminative features (PDFs). On the one hand, PIF is the expert in exploring the overall characteristic of vehicles. When training PIF, we adopt an identity classifier as well as an orientation classifier. In addition, an adversarial loss is deployed in the PIF network. On the other hand, we design a PDF network, which has a similar architecture to the PIF network but can distinguish the difference between local details. The difference between PDF and PIF is that the network of training PDF does not apply the adversarial loss. Finally, by combining PIF and PDF, PRF has the advantages of the two features and can alleviate the influence of the pose barrier problem. Experiments are conducted on the VeRi-776 and VehicleID data sets. We show that PIF and PDF are complementary and that PRF produces competitive performance compared with state-of-the-art approaches. Hao Sheng 0001, Kai Lv 0002, Yang Liu 0088, Wei Ke 0001, Weifeng Lyu, Zhang Xiong 0001, Wei Li 0022 |
IEEE Internet Things J. | 2 |
| 2021 | Improving Driver Gaze Prediction With Reinforced AttentionabstractWe consider the task of driver gaze prediction: estimating where the location of the focus of a driver should be, based on a raw video of the outside environment. In practice, we output a probability map that gives the normalized probability of each point in a given scene being the object of the driver attention. Most existing methods (i.e.,Coarse-to-FineandMulti-branch) take an image or a video as input and directly output the fixation map. While successful, these methods can often produce highly scattered predictions, rendering them unreliable for real-world usage. Motivated by this observation, we propose the reinforced attention (RA) model as a regulatory mechanism to increase prediction density. Our method is built directly on top of existing methods, making it complementary to current approaches. Specifically, we first useMulti-branchto obtain an initial fixation map. Then, RA is trained using deep reinforcement learning to learn a location prediction policy, producing a reinforced attention. Finally, in order to obtain the final gaze prediction result, we combine the fixation map and the reinforced attention by a mask-guided multiplication. Experimental results show that our framework improves the accuracy of gaze prediction, and provides state-of-the-art performance on the DR(eye)VE dataset. Kai Lv 0002, Hao Sheng 0001, Zhang Xiong 0001, Wei Li 0022, Liang Zheng 0001 |
IEEE Trans. Multim. | 1 |
| 2020 | Camera Style Guided Feature Generation for Person Re-identification
Hantao Hu, Yang Liu 0088, Kai Lv 0002, Yanwei Zheng, Wei Zhang 0245, Wei Ke 0001, Hao Sheng 0001 |
WASA (1) | 3 |
| 2020 | Pose-Based View Synthesis for Vehicles: A Perspective Aware MethodabstractIn this paper, we focus on the problem of novel view synthesis for vehicles. Some previous works solve the problem of novel view synthesis in a controlled 3D environment by exploiting additional 3D details (i.e., camera viewpoints and underlying 3D models). However, in real scenarios, the 3D details are difficult to obtain. In this case, we find that introducing vehicle pose to represent the views of vehicles is an alternative paradigm to solve the lack of 3D details. In novel view synthesis, preserving local details is one of the most challenging problems. To address this problem, we propose a perspective-aware generative model (PAGM). We are motivated by the prior that vehicles are made of quadrilateral planes. Preserving these rigid planes during image generation ensures that image details are kept. To this end, a classic image transformation method is leveraged,i.e., perspective transformation. In our GAN-based system, the perspective transformation is applied to the encoder feature maps, and the resulting maps are regarded as new conditions for the decoder. This strategy preserves the quadrilateral planes all the way through the network, thus shuttling the texture details from the input image to the generated image. In the experiments, we show that PAGM can generate high-quality vehicle images with fine details. Quantitatively, our method is superior to several competing approaches employing either GAN or the perspective transformation. Code is available at:https://github.com/ilvkai/view-synthesis-for-vehicles Kai Lv 0002, Hao Sheng 0001, Zhang Xiong 0001, Wei Li 0022, Liang Zheng 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Combine Coarse and Fine Cues: Multi-grained Fusion Network for Video-Based Person Re-identification
Chao Li 0001, Lei Liu 0016, Kai Lv 0002, Hao Sheng 0001, Wei Ke 0001 |
KSEM (1) | 3 |
| 2016 | Robust visual tracking using correlation response mapabstractIn this paper, we address the problem of heavy occlusion where the negative samples contaminate the translation model. In this setting, we decompose the task of tracking into translation and scale estimations of objects. We use hierarchical convolutional features to estimate target position and update translation model, and we use HOG features for the scale filter. In addition, we evaluate the translation's reliability according to the correlation responses map which is the result of correlation detection. Then we propose a new method to update model according to the reliability. Experiments are performed on 28 benchmark sequences with significant scale variations, it shows that the proposed algorithm performs favorably against state-of-the-art methods in terms of accuracy and robustness. Hao Sheng 0001, Kai Lv 0002, Jiahui Chen 0001, Wei Li 0022 |
ICIP | 2 |