Xiaoguang Zhu

dblp:179/4248 · DBLP profile ↗
← Back
25ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0001-9554-2133ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 6 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Empowering Source-Free Domain Adaptation via MLLM-Guided Reliability-Based Curriculum Learning
Dongjie Chen, Kartik Patwari, Zhengfeng Lai, Xiaoguang Zhu, Sen-Ching S. Cheung, Chen-Nee Chuah
WACV4
2026 Privacy-Preserving Video Anomaly Detection: A Survey
abstract
The video anomaly detection (VAD) aims to automatically analyze spatiotemporal patterns in surveillance videos collected from open spaces to detect anomalous events that may cause harm, such as fighting, stealing, and car accidents. However, vision-based surveillance systems such as closed-circuit television (CCTV) often capture personally identifiable information. The lack of transparency and interpretability in video transmission and usage raises public concerns about privacy and ethics, limiting the real-world application of VAD. Recently, researchers have focused on privacy concerns in VAD by conducting systematic studies from various perspectives, including data, features, and systems, making privacy-preserving VAD (P2VAD) a hotspot in the AI community. However, the current research in P2VAD is fragmented, and prior reviews have mostly focused on methods using RGB sequences, overlooking privacy leakage and appearance bias considerations. To address this gap, this article is the first to systematically review the progress of P2VAD, defining its scope and providing an intuitive taxonomy. We outline the basic assumptions, learning frameworks, and optimization objectives of various approaches, analyzing their strengths, weaknesses, and potential correlations. In addition, we provide open access to research resources such as benchmark datasets and available code. Finally, we discuss key challenges and future opportunities from the perspectives of AI development and P2VAD deployment, aiming to the guide future work in the field.
Yang Liu 0246, Siao Liu, Xiaoguang Zhu, Hao Yang 0055, Juncen Guo, Liangyu Teng, Dingkang Yang, Yan Wang 0068, Jing Liu 0050
IEEE Trans. Neural Networks Learn. Syst.3
2025 Adaptive Weighted Parameter Fusion with CLIP for Class-Incremental Learning
abstract
Class-incremental Learning (CIL) enables the model to incrementally absorb knowledge from new classes and build a generic classifier across all previously encountered classes. When the model optimizes with new classes, the knowledge of previous classes is inevitably erased, leading to catastrophic forgetting. Addressing this challenge requires making a trade-off between retaining old knowledge and accommodating new information. However, this balancing process often requires sacrificing some information, which can lead to a partial loss in the model’s ability to discriminate between classes. To tackle this issue, we design the adaptive weighted parameter fusion with Contrastive Language-Image Pre-training (CLIP), which not only takes into account the variability of the data distribution of different tasks, but also retains all the effective information of the parameter matrix to the greatest extent. In addition, we introduce a balance factor that can balance the data distribution alignment and distinguishability of adjacent tasks. Experimental results on several traditional benchmarks validate the superiority of the proposed method.
Juncen Guo, Xiaoguang Zhu, Liangyu Teng
ICME2
2025 M2S2L: Mamba-based Multi-Scale Spatial-temporal Learning for Video Anomaly Detection
abstract
Video anomaly detection (VAD) is an essential task in the image processing community with prospects in video surveillance, which faces fundamental challenges in balancing detection accuracy with computational efficiency. As video content becomes increasingly complex with diverse behavioral patterns and contextual scenarios, traditional VAD approaches struggle to provide robust assessment for modern surveillance systems. Existing methods either lack comprehensive spatial-temporal modeling or require excessive computational resources for real-time applications. In this regard, we present a Mamba-based multi-scale spatial-temporal learning (M2S2L) framework in this paper. The proposed method employs hierarchical spatial encoders operating at multiple granularities and multi-temporal encoders capturing motion dynamics across different time scales. We also introduce a feature decomposition mechanism to enable task-specific optimization for appearance and motion reconstruction, facilitating more nuanced behavioral modeling and quality-aware anomaly assessment. Experiments on three benchmark datasets demonstrate that M2S2L framework achieves 98.5%, 92.1%, and 77.9% frame-level AUCs on UCSD Ped2, CUHK Avenue, and ShanghaiTech respectively, while maintaining efficiency with 20.1G FLOPs and 45 FPS inference speed, making it suitable for practical surveillance deployment.
Yang Liu 0246, Boan Chen, Xiaoguang Zhu, Jing Liu 0050, Peng Sun 0007, Wei Zhou 0013
VCIP3
2025 CRCL: Causal Representation Consistency Learning for Anomaly Detection in Surveillance Videos
abstract
Video Anomaly Detection (VAD) remains a fundamental yet formidable task in the video understanding community, with promising applications in areas such as information forensics and public safety protection. Due to the rarity and diversity of anomalies, existing methods only use easily collected regular events to model the inherent normality of normal spatial-temporal patterns in an unsupervised manner. Although such methods have made significant progress benefiting from the development of deep learning, they attempt to model the statistical dependency between observable videos and semantic labels, which is a crude description of normality and lacks a systematic exploration of its underlying causal relationships. Previous studies have shown that existing unsupervised VAD models are incapable of label-independent data offsets (e.g., scene changes) in real-world scenarios and may fail to respond to light anomalies due to the overgeneralization of deep neural networks. Inspired by causality learning, we argue that there exist causal factors that can adequately generalize the prototypical patterns of regular events and present significant deviations when anomalous instances occur. In this regard, we propose Causal Representation Consistency Learning (CRCL) to implicitly mine potential scene-robust causal variable in unsupervised video normality learning. Specifically, building on the structural causal models, we propose scene-debiasing learning and causality-inspired normality learning to strip away entangled scene bias in deep representations and learn causal video normality, respectively. Extensive experiments on benchmarks validate the superiority of our method over conventional deep representation learning. Moreover, ablation studies and extension validation show that the CRCL can cope with label-independent biases in multi-scene settings and maintain stable performance with only limited training data available.
Yang Liu 0246, Hongjin Wang, Zepu Wang, Xiaoguang Zhu, Jing Liu 0050, Peng Sun 0007, Jianwei Du, Victor C. M. Leung
IEEE Trans. Image Process.4
2024 Supplementing Missing Visions Via Dialog for Scene Graph Generations
abstract
Most AI systems rely on the premise that the input visual data are sufficient to achieve competitive performance in various tasks. However, the classic task setup rarely considers the challenging, yet common practical situations where the complete visual data may be inaccessible due to various reasons (e.g., restricted view range and occlusions). To this end, we investigate a task setting with incomplete visual input data. Specifically, we exploit the Scene Graph Generation (SGG) task with various levels of visual data missingness as input. While insufficient visual input naturally leads to performance drop, we propose to supplement the missing visions via natural language dialog interactions to better accomplish the task objective. We design a model-agnostic Supplementary Interactive Dialog (SI-Dial) framework that can be jointly learned with most existing models, endowing the current AI systems with the ability of question-answer interactions in natural language. We demonstrate the feasibility of such task setting with missing visual input and the effectiveness of our proposed dialog module as the supplementary information source through extensive experiments, by achieving promising performance improvement over multiple baselines.
Zhenghao Zhao, Xiaoguang Zhu, Yuzhang Shang, Yan Yan 0002
ICASSP3
2023 Learning 3D Human Pose and Shape Estimation Using Uncertainty-Aware Body Part Segmentation
abstract
While exploiting body segmentations for supervision, existing 3D human pose and shape estimation methods are plagued by mismatches between clothed body segmentations and skinned SMPL model reprojections. Moreover, noisy pixels introduced by inaccurate segmentation annotations also prevent the model from improving the reconstruction performance further. To address these problems, we propose a novel generalizable framework called Uncertainty-aware body Part Segmentation (UPS), which penalizes different body parts with data uncertainty estimation. Specifically, we use sample-specific segmentations to supervise skinned and clothed body parts separately for realistic human mesh recovery. Furthermore, we leverage data uncertainty to improve the model capacity via learning from representative pixels and resisting noisy ones. Our extensive qualitative and quantitative experiments show that UPS achieves competitive reconstruction results on standard benchmarks.
Xiaoguang Zhu, Zengwen Li, Changxue Chen
ICASSP3
2023 Mutual match for semi-supervised online evolutive learning
abstract
Abstract Semi-supervised learning (SSL) can utilize a large amount of unlabeled data for self-training and continuous evolution with only a few annotations. This feature makes SSL a potential candidate for dealing with data from changing and real-time environments, where deep-learning models need to be adapting to evolving and nonstable (non-i.i.d.) data streams from the real world, i.e., online evolutive scenarios. However, state-of-the-art SSL methods often have complex model design mechanisms and may cause performance degradation in a generalized and open environment. In an edge computing setup, e.g., typical in modern Internet of Things (IoT) applications, a multi-agent SSL architecture can help resolve generalization problems by sharing knowledge between models. In this paper, we introduce Mutual Match (MM), an online-evolutive SSL algorithm that integrates mutual interactive learning and soft-supervision consistency regularization, as well as unsupervised sample mining. By leveraging extra knowledge in the training process and the interactive collaboration between models, MM surpasses multiple top SSL algorithms in accuracy and convergence efficiency under the same online-evolutive experiment setup. MM simplifies the complexity of model design and follows a unified and easy-to-expandable pipeline, which can be beneficial to tasks with insufficient labeled data and frequently changing data distribution.
Xiaoguang Zhu
Appl. Intell.2
2023 Distributional and spatial-temporal robust representation learning for transportation activity recognition
Jing Liu 0050, Yang Liu 0246, Xiaoguang Zhu
Pattern Recognit.4
2022 Learning Task-Specific Representation for Video Anomaly Detection with Spatial-Temporal Attention
abstract
The automatic detection of abnormal events in surveillance videos with weak supervision has been formulated as a multiple instance learning task, which aims to localize the clips containing abnormal events temporally with the video-level labels. However, most existing methods rely on the features extracted by the pre-trained action recognition models, which are not discriminative enough for video anomaly detection. In this work, we propose a spatial-temporal attention mechanism to learn inter- and intra-correlations of video clips, and the boosted features are encouraged to be task-specific via the mutual cosine embedding loss. Experimental results on standard benchmarks demonstrate the effectiveness of the spatial-temporal attention, and our method achieves superior performance to the state-of-the-art methods.
Yang Liu 0246, Jing Liu 0050, Xiaoguang Zhu, Donglai Wei 0002
ICASSP3
2022 Learning Appearance-Motion Normality for Video Anomaly Detection
abstract
Video anomaly detection is a challenging task in the Computer vision community. Most single task-based methods do not consider the independence of unique spatial and temporal patterns, while two-stream structures lack the exploration of the correlations. In this paper, we propose spatial-temporal memories augmented two-stream auto-encoder framework, which learns the appearance normality and motion normal-ity independently and explores the correlations via adversar-ial learning. Specifically, we first design two proxy tasks to train the two-stream structure to extract appearance and motion features in isolation. Then, the prototypical features are recorded in the corresponding spatial and temporal memory pools. Finally, the encoding-decoding network performs ad-versariallearning with the discriminator to explore the corre-lations between spatial and temporal patterns. Experimental results show that our framework outperforms the state-of-the-art methods, achieving AUCs of 98.1% and 89.8% on UCSD Ped2 and CUHK Avenue datasets.
Yang Liu 0246, Jing Liu 0050, Mengyang Zhao 0002, Dingkang Yang, Xiaoguang Zhu
ICME5
2022 Generating Adaptive Targeted Adversarial Examples for Content-Based Image Retrieval
abstract
Massive accessible personal data on the Internet raises the risk of malicious retrieval. In this paper, we propose to conceal the images with the targeted adversarial attacks on content-based image retrieval. An imperceptible perturbation is added to the original image to generate adversarial examples, making the retrieval results similar to the target image but look completely different. Previous work on the targeted attack for image retrieval only introduces a target-specific model and needs to train the model each time for new targets. We extend the attack adaptability by exploiting the target images as conditional input for the generative model. The proposed Adaptive Targeted Attack Generative Adversarial Network (ATA-GAN) is a GAN-based model with a generator and discriminator. The generator extracts the features of origin and target, then uses the Feature Integration Module to explore the relation between the target and original image to ignore the origin feature while paying more attention to the target. Simultaneously, the discriminator distinguishes the realness and ensures the adversarial example is similar to the origin. We evaluate and analyze the performance of the adaptive targeted attack on popular retrieval benchmarks.
Jiameng Pan, Xiaoguang Zhu
IJCNN2
2022 MSAF: Multimodal Supervise-Attention Enhanced Fusion for Video Anomaly Detection
abstract
The complementarity of multimodal signal is essential for video anomaly detection. However, existing methods either lack exploration to multimodal data or ignore the implicit alignment of multimodal features. In our work, we address this problem using a novel fusion method and propose a Multimodal Supervise-Attention enhanced Fusion (MSAF) framework under weak supervision. Our framework can be divided into two parts: 1) the multimodal labels refinement part refines video-level ground truth into pseudo clip-level labels for subsequent training, 2) the multimodal supervise-attention fusion network enhances features via implicitly aligning different information, then fusing them effectively to predict anomaly scores with the help of refined labels. We validate our framework on four challenging datasets: ShanghaiTech, UCF-Crime, LAD, and XD-Violence. Extensive experiments on the benchmarks demonstrate the effectiveness of our framework, which achieves comparable results on several benchmarks and outperforms current state-of-the-art methods on the XD-Violence audiovisual multimodal dataset.
Donglai Wei 0002, Yang Liu 0246, Xiaoguang Zhu, Jing Liu 0050, Xinhua Zeng
IEEE Signal Process. Lett.3
2022 Cross-modal Graph Matching Network for Image-text Retrieval
abstract
Image-text retrieval is a fundamental cross-modal task whose main idea is to learn image-text matching. Generally, according to whether there exist interactions during the retrieval process, existing image-text retrieval methods can be classified into independent representation matching methods and cross-interaction matching methods. The independent representation matching methods generate the embeddings of images and sentences independently and thus are convenient for retrieval with hand-crafted matching measures (e.g., cosine or Euclidean distance). As to the cross-interaction matching methods, they achieve improvement by introducing the interaction-based networks for inter-relation reasoning, yet suffer the low retrieval efficiency. This article aims to develop a method that takes the advantages of cross-modal inter-relation reasoning of cross-interaction methods while being as efficient as the independent methods. To this end, we propose a graph-based Cross-modal Graph Matching Network (CGMN) , which explores both intra- and inter-relations without introducing network interaction. In CGMN, graphs are used for both visual and textual representation to achieve intra-relation reasoning across regions and words, respectively. Furthermore, we propose a novel graph node matching loss to learn fine-grained cross-modal correspondence and to achieve inter-relation reasoning. Experiments on benchmark datasets MS-COCO, Flickr8K, and Flickr30K show that CGMN outperforms state-of-the-art methods in image retrieval. Moreover, CGMM is much more efficient than state-of-the-art methods using interactive matching. The code is available at https://github.com/cyh-sj/CGMN .
Yuhao Cheng, Xiaoguang Zhu, Jiuchao Qian, Fei Wen 0005
ACM Trans. Multim. Comput. Commun. Appl.2
2022 Skeleton Sequence and RGB Frame Based Multi-Modality Feature Fusion Network for Action Recognition
abstract
Action recognition has been a heated topic in computer vision for its wide application in vision systems. Previous approaches achieve improvement by fusing the modalities of the skeleton sequence and RGB video. However, such methods pose a dilemma between the accuracy and efficiency for the high complexity of the RGB video network. To solve the problem, we propose a multi-modality feature fusion network to combine the modalities of the skeleton sequence and RGB frame instead of the RGB video, as the key information contained by the combination of the skeleton sequence and RGB frame is close to that of the skeleton sequence and RGB video. In this way, complementary information is retained while the complexity is reduced by a large margin. To better explore the correspondence of the two modalities, a two-stage fusion framework is introduced in the network. In the early fusion stage, we introduce a skeleton attention module that projects the skeleton sequence on the single RGB frame to help the RGB frame focus on the limb movement regions. In the late fusion stage, we propose a cross-attention module to fuse the skeleton feature and the RGB feature by exploiting the correlation. Experiments on two benchmarks, NTU RGB+D and SYSU, show that the proposed model achieves competitive performance compared with the state-of-the-art methods while reducing the complexity of the network.
Xiaoguang Zhu, Honglin Wen, Yan Yan 0002
ACM Trans. Multim. Comput. Commun. Appl.1
2021 Uncertainty-Aware Domain Adaptation for Action Recognition
Xiaoguang Zhu, Zhantao Yang
ICONIP (3)1
2021 SDAN: Stacked Diverse Attention Network for Video Action Recognition
abstract
Recently, deep learning methods have proved exceptional performance in video action recognition. However, it remains a challenging problem to extract discriminative features from videos effectively. Most existing methods mainly focus on spatial-temporal information separately. In this paper, we propose a novel Stacked Diverse Attention Network (SDAN). It uses a Multi-dimensional Attention Module to emphasize informative maps along the channel and spatial-temporal dimension. Besides, a Supervised Attention Module is designed to generate a weighted feature map in a class-supervised way. Compared with previous methods, the proposed method has the following advantages: (1) The Multi-dimensional Attention Module can exploit effective combinations and correlation of attention mechanisms along different dimensions. (2) The Supervised Attention Module improves the network capability supervised directly by action labels which reveal the class-related object and limb information. Extensive experimental evaluation demonstrates the effectiveness of the proposed approach and establishes significant results on Kinetics400, UCF101 and HMDB51 Datasets. Codes are available on https://github.com/jeff62802217/SDAN-Pytorch.
Xiaoguang Zhu, Siran Huang, Wenjing Fan, Yuhao Cheng, Huaqing Shao
ISCAS1
2021 Graph-based reasoning attention pooling with curriculum design for content-based image retrieval
Xiaoguang Zhu, Zhantao Yang, Jiuchao Qian
Image Vis. Comput.1
2021 Dark-Aware Network For Fine-Grained Sketch-Based Image Retrieval
abstract
Fine-grained sketch-based image retrieval (FG-SBIR) is an emerging topic in high-level computer vision. Among existing methods, edge-only information based ones are convenient to use but incompetent in distinguishing among ambiguous samples. On the other hand, even though the methods using additional color information can significantly improve the retrieval performance, they are not convenient for practical use as the increase of sketch complexity. This paper fills the gap between these two kinds of methods by proposing a novel method, namely dark-aware sketch-based image retrieval (DA-SBIR), which incorporates the dark region information of images into FG-SBIR without a sacrifice of convenience. Our model consists of two branches to process edge structure and dark region, respectively. Specifically, instead of using real color, DA-SBIR only requires some additional free-hand black sketch to represent the dark region. Besides, a Split Generative Adversarial Network is introduced to automatically split a sketch into edge-structure-only and dark-region-only sketches. Moreover, we build a new clothes dataset, SJTU-Cloth, with more ambiguous samples for FG-SBIR. Experimental results on the QMUL-Shoe dataset and our SJTU-Cloth dataset show that our approach achieves consistent improvements over state-of-the-arts. Code and dataset are available at https://github.com/y2242794082/SBIR.git.
Zhantao Yang, Xiaoguang Zhu, Jiuchao Qian
IEEE Signal Process. Lett.2
2020 MRNet: A Keypoint Guided Multi-scale Reasoning Network for Vehicle Re-identification
Minting Pan, Xiaoguang Zhu, Yongfu Li 0002, Jiuchao Qian
ICONIP (4)2
2020 Curriculum Enhanced Supervised Attention Network for Person Re-Identification
abstract
Though deep learning methods in Person Re-ID have increased its performance, the field is still confronted with great challenges such as pedestrian misalignment due to the inferior pedestrian detector. To solve such problems, we propose Curriculum Enhanced Supervised Attention Network (CE-SAN): Firstly, an attention module is trained under supervision, which helps the network further emphasize the key information and associate the local and global branch together, leading to a better exploitation of the discriminative features. Secondly, a curriculum design is adopted to divide the dataset into subsets according to the distribution density of the training samples, enabling the network to learn gradually to increase the capability. Moreover, The CE-SAN is easy to be plugged in most of the backbones with high generalization ability. To prove the CE-SAN's superiority, experiments are conducted on three datasets, CE-SAN achieves competitive performance with the state-of-the-arts on Market-1501 and DukeMTMC-reID. Particularly, on the most challenging MSMT17, it outperforms the state-of-the-art methods by 5.2% in Rank-1 and 6.5% in mAP.
Xiaoguang Zhu, Jiuchao Qian
IEEE Signal Process. Lett.1
2019 Action Recognition Based on 3D Skeleton and RGB Frame Fusion
abstract
Action recognition has wide applications in assisted living, health monitoring, surveillance, and human-computer interaction. In traditional action recognition methods, RGB video-based ones are effective but computationally inefficient, while skeleton-based ones are computationally efficient but do not make use of low-level detail information. This work considers action recognition based on a multimodal fusion between the 3D skeleton and the RGB image. We design a neural network that uses a 3D skeleton sequence and a single middle frame from an RGB video as input. Specifically, our method picks up one frame in a video and extracts spatial features from it using two attention modules, a self-attention module and a skeleton-attention module. Further, temporal features are extracted from the skeleton sequence via a BI-LSTM sub-network. Finally, the spatial features and the temporal features are combined via a feature fusion network for action classification. A distinct feature of our method is that it uses only a single RGB frame rather than an RGB video. Accordingly, it has a light-weighted architecture and is more efficient than RGB video-based methods. Comparative evaluation on two public datasets, NTU-RGBD and SYSU, demonstrates that, our method can achieve competitive performance compared with state-of-the-art methods.
Guiyu Liu, Jiuchao Qian, Fei Wen 0005, Xiaoguang Zhu, Rendong Ying
IROS4
2019 A System Integrating Speech Interaction and Vision Sensing Applying in Smart Home Scenario
abstract
This paper introduces a system integrating speech interaction and visual sensing, which aims at the smart home scenario. It is divided into three modules: data transmission and distribution, data reception and processing, monitoring and management. Simultaneously, the system employs a centralized hardware architecture with wired connection, ensuring the stability of duplex communication. It is proved that the system has powerful extensibility and favorable maintainability. In order to make speech interaction algorithms perform better in this system, a blind soure beamforming algorithm for speech enhancement is proposed. Meanwhile, a lot of considerations are used to improve the usability of visual sensing in smart home scenarios. Precisely speaking, the blind soure beamforming algorithm improves the poor performance of speech interaction for far-field and multi-angle, and the consideration improves the accuracy and reduces the time-delay of visual sensing.
Junyu Dai, Jiuchao Qian, Zheng Tao, Xiaoguang Zhu, Huaqing Shao
ISCAS5
2019 CD-ABM: Curriculum Design with Attention Branch Model for Person Re-identification
Jiuchao Qian, Xiaoguang Zhu, Fei Wen 0005
PRICAI (3)3
2016 ADRC Law of Spacecraft Rendezvous and Docking in Final Approach Phase
abstract
This article addresses the high-precision coordinated control problem of spacecraft autonomous rendezvous and docking, which couple the relative position and attitude in the final approach phase. The coupled dynamics equations of the tracking-target spacecrafts is derived by using dual quaternions. Then, a cascade Active Disturbance Rejection Controller is proposed, by which the extended state observer and nonlinear error feedback law is designed, the virtual value on which the actual control volume tracking is calculated to ensure the finite time convergence of the relative position and attitude tracking errors in spite of parametric uncertainties and external disturbances. Finally, numerical simulations are performed to demonstrate that the proposed approaches, which can avoid the coupling effect and restrain the interference, can track the target spacecraft in a relatively short period of time, and the control precision can satisfy the requirements of docking.
Wei Gao 0053, Xiaoguang Zhu, Miaolei Zhou, Qiying Jiang
Cybern. Syst.2