Xiangzeng Liu

dblp:72/8488 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0002-2751-6096ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 7 since 2021
YearPublicationVenuePosition
2026 Weakly semantic-guided skeleton feature distillation for human action recognition
Ruyi Liu 0001, Qiguang Miao, Wentian Xin, Xiangzeng Liu, Long Li 0005
Expert Syst. Appl.6
2025 PointSR: Self-Regularized Point Supervision for Drone-View Object Detection
abstract
Point-Supervised Object Detection (PSOD) in a discriminative style has recently gained significant attention for its impressive detection performance and cost-effectiveness. However, accurately predicting high-quality pseudo-box labels for drone-view images, which often feature densely packed small objects, remains a challenge. This difficulty arises primarily from the limitation of rigid sampling strategies, which hinder the pseudo-box optimization process. To address this, we propose PointSR, an effective and robust point-supervised object detection framework with self-regularized sampling that integrates temporal and informative constraints throughout the pseudo-box generation process. Specifically, the framework comprises three key components: Temporal-Ensembling Encoder (TE Encoder), Coarse Pseudo-box Prediction, and Pseudo-box Refinement. The TE Encoder builds an anchor prototype library by aggregating temporal information for dynamic anchor adjustment. In Coarse Pseudo-box Prediction, anchors are refined using the prototype library, and a set of informative samples is collected for subsequent refinement. During Pseudo-box Refinement, these informative negative samples are used to suppress low-confidence candidate positive samples, thereby improving the quality of the pseudo-boxes. Experimental results on benchmark datasets demonstrate that PointSR significantly outperforms state-of-the-art methods, achieving up to 2.6% ∼ 7.2% higher AP50using only point supervision. Additionally, it exhibits strong robustness to perturbation in human-labeled points.
Weizhuo Li, Wenjing Jia, Zehao Zhang, Xiangzeng Liu, Qiguang Miao
CVPR6
2025 SGAD: Semantic and Geometric-Aware Descriptor for Local Feature Matching
abstract
Local feature matching remains a fundamental challenge in computer vision. Recent Area to Point Matching (A2PM) methods have improved matching accuracy. However, existing research based on this framework relies on inefficient pixel-level comparisons and complex graph matching that limit scalability. In this work, we introduce the Semantic and Geometric-aware Descriptor Network (SGAD), which fundamentally rethinks area-based matching by generating highly discriminative area descriptors that enable direct matching without complex graph optimization. This approach significantly improves both accuracy and efficiency of area matching. We further improve the performance of area matching through a novel supervision strategy that decomposes the area matching task into classification and ranking subtasks. Finally, we introduce the Hierarchical Containment Redundancy Filter (HCRF) to eliminate overlapping areas by analyzing containment graphs. SGAD demonstrates remarkable performance gains, reducing runtime by 60x (0.82s vs. 60.23s) compared to MESA. Extensive evaluations show consistent improvements across multiple point matchers: SGAD+LoFTR reduces runtime compared to DKM, while achieving higher accuracy (0.82s vs. 1.51s, 65.98 vs. 61.11) in outdoor pose estimation, and SGAD+ROMA delivers +7.39% AUC@5° in indoor pose estimation, establishing a new state-of-the-art.
Xiangzeng Liu, Guanglu Shi, Qiguang Miao
ICCV1
2025 Video-Based Traffic Light Recognition by Rockchip RV1126 for Autonomous Driving
abstract
Real-time traffic light recognition is fundamental for autonomous driving safety and navigation in urban environments. While existing approaches rely on single-frame analysis from onboard cameras, they struggle with complex scenarios involving occlusions and adverse lighting conditions. We present ViTLR, a novel video-based end-to-end neural network that processes multiple consecutive frames to achieve robust traffic light detection and state classification. The architecture leverages a transformer-like design with convolutional self-attention modules, which is optimized specifically for deployment on the Rockchip RV1126 embedded platform. Extensive evaluations on two real-world datasets demonstrate that ViTLR achieves state-of-the-art performance while maintaining real-time processing capabilities (>25 FPS) on RV1126's NPU. The system shows superior robustness across temporal stability, varying target distances, and challenging environmental conditions compared to existing single-frame approaches. We have successfully integrated ViTLR into an ego-lane traffic light recognition system using HD maps for autonomous driving applications. The complete implementation, including source code and datasets, is made publicly available to facilitate further research in this domain.
Xuxu Kong, Shengtong Xu, Haoyi Xiong, Xiangzeng Liu
IV5
2025 A Benchmark for Vision-Centric HD Mapping by V2I Systems
abstract
Autonomous driving faces safety challenges due to a lack of global perspective and the semantic information of vectorized high-definition (HD) maps. Information from roadside cameras can greatly expand the map perception range through vehicle-to-infrastructure (V2I) communications. However, there is still no dataset from the real world available for the study on map vectorization onboard under the scenario of vehicle-infrastructure cooperation. To prosper the research on online HD mapping for Vehicle-Infrastructure Cooperative Autonomous Driving (VICAD), we release a real-world dataset, which contains collaborative camera frames from both vehicles and roadside infrastructures, and provides human annotations of HD map elements. We also present an end-to-end neural framework (i.e., V2I-HD) leveraging vision-centric V2I systems to construct vectorized maps. To reduce computation costs and further deploy V2I-HD on autonomous vehicles, we introduce a directionally decoupled self-attention mechanism to V2I-HD. Extensive experiments show that V2I-HD has superior performance in real-time inference speed, as tested by our real-world dataset. Abundant qualitative results also demonstrate stable and robust map construction quality with low cost in complex and various driving scenes. As a benchmark, both source codes and the dataset have been released at OneDrive11https://ldrv.ms/f/c/76645c25a8914a0b/EgWy5XCUk6pKgvE9vB-HbVEBCdCQjJvgxlKKjeKF7hPdZw for the purpose of further study.
Shengtong Xu, Kun Jiang 0002, Haoyi Xiong, Xiangzeng Liu
IV6
2025 Towards V2X HD Mapping for Autonomous Driving: A Concise Review
abstract
High-definition (HD) maps are fundamental components to autonomous driving systems, providing essential centimeter-level accuracy and lane-level semantic information. While traditional HD mapping methods have evolved into online learning approaches, current solutions face significant challenges due to sensor limitations and environmental constraints. This paper presents a systematic review of HD map construction methods, tracking their evolution from conventional techniques to advanced Vehicle-to-Everything (V2X) cooperative mapping enabled by edge computing and communication technologies. Through comprehensive analysis of methodologies, algorithms, and datasets, we identify critical challenges in current HD mapping systems. Our review encompasses three key domains: traditional mapping methods, online learning approaches, and V2X cooperative construction of HD maps. We evaluate existing solutions against standardized metrics, compare their effective-ness, and outline promising directions for future research. This work provides researchers and practitioners with a structured understanding of the HD mapping landscape and highlights opportunities for advancing autonomous driving systems.
Suhui Yang, Shengtong Xu, Xiangzeng Liu, Wenbo Hu 0001, Haoyi Xiong
IV5
2025 A Concise Survey on Lane Topology Reasoning for HD Mapping
abstract
Lane topology reasoning techniques play a crucial role in high-definition (HD) mapping and autonomous driving applications. While recent years have witnessed significant advances in this field, there has been limited effort to consolidate these works into a comprehensive overview. This survey systematically reviews the evolution and current state of lane topology reasoning methods, categorizing them into three major paradigms: procedural modeling-based methods, aerial imagery-based methods, and onboard sensors-based methods. We analyze the progression from early rule-based approaches to modern learning-based solutions utilizing transformers, graph neural networks (GNNs), and other deep learning architectures. The paper examines standardized evaluation metrics, including road-level measures (APLS and TLTS score), and lane-level metrics (DET and TOP score), along with performance comparisons on benchmark datasets such as OpenLane- V2. We identify key technical challenges, including dataset availability and model efficiency, and outline promising directions for future research. This comprehensive review provides researchers and practitioners with insights into the theoretical frameworks, practical implementations, and emerging trends in lane topology reasoning for HD mapping applications.
Shengtong Xu, Haoyi Xiong, Xiangzeng Liu, Wenbo Hu 0001, Wenbing Huang 0001
IV5
2025 Semantic SLAM with Rolling-Shutter Cameras and Low-Precision INS in Outdoor Environments
abstract
Accurate localization and mapping in outdoor environments remains challenging when using consumer-grade hardware, particularly with rolling-shutter cameras and low-precision inertial navigation systems (INS). We present a novel semantic SLAM approach that leverages road elements such as lane boundaries, traffic signs, and road markings to enhance localization accuracy. Our system integrates real-time semantic feature detection with a graph optimization framework, effectively handling both rolling-shutter effects and INS drift. Using a practical hardware setup that consists of a rolling-shutter camera$(\mathbf{3840} \times \mathbf{2160} {@} \mathbf{30} \mathbf{fps})$, IMU (100Hz), and wheel encoder (50Hz), we demonstrate significant improvements over existing methods. Compared to state-of-the-art approaches, our method achieves higher recall (up to 5.35%) and precision (up to 2.79%) in semantic element detection while maintaining mean relative error (MRE) within 10cm and mean absolute error (MAE) around 1m. Extensive experiments in diverse urban environments demonstrate the robust performance of our system under varying lighting conditions and complex traffic scenarios, making it particularly suitable for autonomous driving applications. The proposed approach provides a practical solution for high-precision localization using affordable hardware, bridging the gap between consumer-grade sensors and production-level performance requirements.
Yi Jiao, Shengtong Xu, Xiangzeng Liu, Haoyi Xiong
IV5
2025 StrongerCenter: Enhancing TransCenter for robust multi-object tracking
Xiangzeng Liu, Kailai Wang, Bocheng Zhao, Qiguang Miao
Neurocomputing1
2025 One-shot handwriting imitation via self-supervised cross spatial transformer networks
Bocheng Zhao, Guanwen Feng, Wenxing Zhang, Yunan Li 0001, Qiguang Miao, Xiangzeng Liu, Ruyi Liu 0001
Neurocomputing6
2025 SG-CLR: Semantic representation-guided contrastive learning for self-supervised skeleton-based action recognition
Ruyi Liu 0001, Wentian Xin, Qiguang Miao, Xiangzeng Liu, Long Li 0005
Pattern Recognit.6
2025 CAETFN: Context Adaptively Enhanced Text-Guided Fusion Network for Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis (MSA) is an active research area in recent years with the exponential development of the internet and social media, which aims to recognize the speaker’s sentiment in the video consisted of text, acoustic and visual cues, and has attracted attention from many applications such as smart education, intelligent medication and social security. The predominant approaches have devoted to developing more complicated fusion strategy to learn efficient multimodal representations. However, information from these modalities usually have different contributions to MSA task. More specifically, the text modality outperforms the non-verbal modalities since its highly condensed semantic information and the maturity of the pre-trained language models. Taking full advantage of the text modality while integrating the non-verbal sentiment-relevant contextual information becomes a substantial challenge. Thus, in this paper, we propose a Context Adaptively Enhanced Text-guided Fusion Network, which is embedded in the pre-trained language model and utilizes the text modality as the guide to reduce the redundancy and exploit the sentiment-relevant information and in turn uses these information to complement itself with the non-verbal sentiment contexts. Moreover, a novelly designed non-verbal feature enhancement module is introduced to capture long-range dependencies in two directions, with the substantial removal of the redundancy and the noise. Extensive experiments on two benchmark datasets CMU-MOSI and CMU-MOSEI demonstrate the competitive performance over the state-of-the-art methods.
Ruyi Liu 0001, Qiguang Miao, Di Wang 0011, Xiangzeng Liu
IEEE Trans. Affect. Comput.5
2025 Adaptive Occlusion-Aware Network for Occluded Person Re-Identification
abstract
Occluded person re-identification (ReID) is a challenging task due to some of the essential features are interfered by obstacles or other pedestrians. Multi-granularity local feature extraction and recognition can effectively improve the accuracy of ReID under occlusion. However, manual segmentation methods for local features can lead to feature misalignment. Feature alignment based on pose estimation often ignores non-body details (e.g., handbags, backpacks, etc.) while increasing the complexity of the model. To address the above challenges, we propose a novel Adaptive Occlusion-Aware Network (AOANet), which mainly consists of two modules, the Adaptive Position Extractor (APE) and the Occlusion Awareness Module (OAM). In order to adaptively extract distinguishing features of body parts, APE optimizes the representation of multi-granularity features by the guidance of attention mechanism and keypoint features. To further perceive the occluded region, the OAM is developed by adaptively calculating the occlusion weights for body parts. These weights can lead to highlighting the non-occluded parts and suppressing the occluded parts, which in turn improves the accuracy in the occluded situation. Extensive experimental results confirm the advantages of our method on the MSMT17, DukeMTMC-reID, Market-1501, Occluded-Duke and Occluded-ReID datasets. The comparative results demonstrate that our method outperforms comparable methods. Especially on the Occluded-Duke dataset, our method achieved 70.6% mAP and 81.2% Rank-1 accuracy.
Xiangzeng Liu, Hao Chen 0187, Qiguang Miao, Ruyi Liu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 Flow-Audio-Synth: A Video-to-Audio Model which Captures Dynamic Features
Yupeng Zheng, Zixiang Lu, Qiguang Miao, Xiangzeng Liu
PRCV (10)5
2024 Pose-Guided Attention Learning for Cloth-Changing Person Re-Identification
abstract
The change in appearance is a great challenge for cloth-changing person re-identification. Existing methods tackle this challenge by learning the shape features of the human body, however, these features are easily affected by the human pose or camera perspective. Thus, the ability to learn and extract the invariant features of person in varying conditions is crucial to overcome the above challenge. To address the issue of invariant feature extraction for cloth-changing person re-identification, a Pose-Guided Attention Learning (PGAL) framework is proposed in this paper. First, we introduce the human pose estimation network to remove the background effects and align the fine-grained key points features of human body. Then, to fully exploit the available appearance information, we develop a Feature Enhancement Module (FEM) that improves the feature representation of non-key point regions of human body through the Multi-Head Self Attention. Finally, in order to adaptively learn the invariant features of the person, we construct an Attention Learning Module (ALM) to achieve automatic selection of multi-granularity features by utilizing three different loss functions. Comparing with current popular methods on four cloth-changing person Re-ID datasets, the experimental results show the superiority of our method.
Xiangzeng Liu, Yi-Ning Quan, Qiguang Miao
IEEE Trans. Multim.1
2024 A Semi-Supervised Underexposed Image Enhancement Network With Supervised Context Attention and Multi-Exposure Fusion
abstract
Recently, image enhancement approaches yield impressive progress. However, most methods are still based supervised-learning, which requires plenty of paired data. Meanwhile, owing to the complex illumination condition in a real-world scenario, those methods trained on synthetic images cannot restore details in extremely dark or bright areas and lead to exposure errors. The traditional losses that deem all pixels the same in training also produce blurry edges in the result. To handle these problems, in this article, we present an effective semi-supervised framework for severely underexposed image enhancement. Our network consists of a supervised and an unsupervised branch, which shares weights and can make full use of paired data and plenty of unpaired data. Meanwhile, a multi-exposure fusion module is designed to adaptively fuse the corrected images to address the low contrast and color bias issues occurring in some extreme situations. Moreover, we propose a supervised context attention module to better use the edge information as supervision to recover fine image details. Extensive experiments have proved that the proposed method outperforms state-of-the-art approaches in enhancing exposure images.
Xiaolong Fu, Yunan Li 0001, Kaibin Miao, Xiangzeng Liu, Bocheng Zhao, Qiguang Miao
IEEE Trans. Multim.5
2021 CA-PMG: Channel attention and progressive multi-granularity training network for fine-grained visual classification
abstract
Abstract Fine‐grained visual classification is challenging due to the inherently subtle intra‐class object variations. To solve this issue, a novel framework named channel attention and progressive multi‐granularity training network, is proposed. It first exploits meaningful feature maps through the channel attention module and captures multi‐granularity features by the progressive multi‐granularity training module. For each feature map, the channel attention module is proposed to explore channel‐wise correlation. This allows the model to re‐weight the channels of the feature map according to the impact of their semantic information on performance. Furthermore, the progressive multi‐granularity training module is introduced to fuse features cross multi‐granularity. And the fused features pay more attention to the subtle differences between images. The model can be trained efficiently in an end‐to‐end manner without bounding box or part annotations. Finally, comprehensive experiments are conducted to show that the method achieves state‐of‐the‐art performances on the CUB‐200‐2011, Stanford Cars, and FGVC‐Aircraft datasets. Ablation studies demonstrate the effectiveness of each part in our module.
Qiguang Miao, Hang Yao 0001, Xiangzeng Liu, Ruyi Liu 0001, Maoguo Gong
IET Image Process.4
2017 Robust image-based crack detection in concrete structure using multi-scale enhancement and visual features
abstract
Crack detection is an important technique to evaluate the safety and predict the life of a concrete asset. In order to improve the robustness of the crack detection in complex background, a new crack detection framework based on multi-scale enhancement and visual features is developed. Firstly, to deal with the effect of low contrast, a multi-scale enhancement method using guided filter and gradient information is proposed. Then, the adaptive threshold algorithm is used to obtain the binary image. Finally, the combination of morphological processing and visual features are adopted to purify the cracks. The experimental results with different images of real concrete surface demonstrate the high robustness and validity of the developed technique, in which the average TPR can reach 94.22%.
Xiangzeng Liu, Yunfeng Ai, Sebastian A. Scherer
ICIP1