Shuzhe Wu

dblp:160/1868 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0002-4455-4123ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 since 2021Security and privacy · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Collaboratively Self-Supervised Video Representation Learning for Action Recognition
abstract
Considering the close connection between action recognition and human pose estimation, we design a Collaboratively Self-supervised Video Representation (CSVR) learning framework specific to action recognition by jointly factoring in generative pose prediction and discriminative context matching as pretext tasks. Specifically, our CSVR consists of three branches: a generative pose prediction branch, a discriminative context matching branch, and a video generating branch. Among them, the first one encodes dynamic motion feature by utilizing Conditional-GAN to predict the human poses of future frames, and the second branch extracts static context features by contrasting positive and negative video feature and I-frame feature pairs. The third branch is designed to generate both current and future video frames, for the purpose of collaboratively improving dynamic motion features and static context features. Extensive experiments demonstrate that our method achieves state-of-the-art performance on multiple popular video datasets.
Jie Zhang 0071, Zhifan Wan, Lanqing Hu, Shuzhe Wu, Shiguang Shan
IEEE Trans. Inf. Forensics Secur.5
2024 Hierarchical compositional representations for few-shot action recognition
Changzhen Li, Jie Zhang 0071, Shuzhe Wu, Xin Jin 0004, Shiguang Shan
Comput. Vis. Image Underst.3
2024 Dual Sampling Based Causal Intervention for Face Anti-Spoofing With Identity Debiasing
abstract
Improving generalization to unseen scenarios is one of the greatest challenges in Face Anti-spoofing (FAS). Most previous FAS works focus on domain debiasing to eliminate the distribution discrepancy between training and test data. However, a crucial but usually neglected bias factor is the face identity. Generally, the identity distribution varies across the FAS datasets as the participants in these datasets are from different regions, which will lead to serious identity bias in the cross-dataset FAS tasks. In this work, we resort to causal learning and propose Dual Sampling based Causal Intervention (DSCI) for face anti-spoofing, which improves the generalization of the FAS model by eliminating the identity bias. DSCI treats the bias as a confounder and applies the backdoor adjustment through the proposed dual sampling on the face identity and the FAS feature. Specifically, we first sample the data uniformly on the identity distribution that is obtained by a pretrained face recognition model. By feeding the sampled data into a network, we can get an estimated FAS feature distribution and sample the FAS feature on it. Sampling the FAS feature from a complete estimated distribution can include potential counterfactual features in the training, which effectively expands the training data. The dual sampling process helps the model learn the real causality between the FAS feature and the input liveness, allowing the model to perform more stably across various identity distributions. Extensive experiments demonstrate our proposed method outperforms the state-of-the-art methods on both intra- and cross-dataset evaluations.
Xingming Long, Jie Zhang 0071, Shuzhe Wu, Xin Jin 0004, Shiguang Shan
IEEE Trans. Inf. Forensics Secur.3
2023 PLPL-VIO: A Novel Probabilistic Line Measurement Model for Point-Line-Based Visual-Inertial Odometry
abstract
Point and line features are complementary in Visual-Inertial Odometry (VIO) or Visual-Inertial Simultaneous Localization And Mapping (VI-SLAM) systems. The advantage of combining these two types of features relies on their proper weighting in the cost function, usually set by their uncertainty. Compared with point features, setting line segment endpoints' uncertainty with isotropic distribution is unreasonable. But the uncertainty of line feature observation, especially for the endpoints' uncertainty along the line, is difficult to set due to occlusion and fragmentation problems. In this article, we use infinite lines as the line feature observations and prove that the uncertainty of these observations is only related to the vertical uncertainty of the endpoints, thus avoiding setting the parallel uncertainty of the endpoints. Besides, we introduce a novel consistent measurement model for line features. Furthermore, for long-time constraints, we add 3D line segments into the state vector and derive how to update them properly. Finally, we construct a point-line-based VIO system that takes into account the uncertainty of line feature observations and the consistency of line feature measurements. The proposed VIO system is validated on two public datasets. The results show that the proposed method obtains the best accuracy compared with the state-of-the-art point-based VIO systems (OpenVINS, VINS-Mono), a point-line-based VIO system (PL-VINS), and a structural line-based system (StructVIO).
Zewen Xu, Hao Wei 0008, Fulin Tang, Yihong Wu 0002, Gang Ma 0007, Shuzhe Wu, Xin Jin 0004
IROS7
2023 Data-Efficient Masked Video Modeling for Self-supervised Action Recognition
abstract
Recently, self-supervised video representation learning based on Masked Video Modeling (MVM) has demonstrated promising results for action recognition. However, existing methods face two significant challenges: (1) video actions involve a crucial temporal dimension, yet current masking strategies adopt inefficient random approaches that undermine low-density dynamic motion clues in videos; (2) pre-training requires large-scale datasets and significant computing resources (including large batch sizes and enormous iterations). To address these issues, we propose a novel method named Data-Efficient Masked Video Modeling (DEMVM) for self-supervised action recognition. Specifically, a novel masking strategy named Flow-Guided Dense Masking (FGDM) is proposed to facilitate efficient learning by focusing more on the action-related temporal clues, which applies dense masking to dynamic regions based on optical flow priors, while sparse masking to background regions. Furthermore, DEMVM introduces a 3D video tokenizer to enhance the modeling of temporal clues. Finally, Progressive Masking Ratio (PMR) and 2D initialization strategies are presented to enable the model to adapt to the characteristics of the MVM paradigm during different training stages. Extensive experiments on multiple benchmarks, UCF101, HMDB51, and Mimetics, demonstrate that our method achieves state-of-the-art performance in the downstream action recognition task with both efficient data and low computational cost. More interestingly, the few-shot experiment on the Mimetics dataset shows that DEMVM can accurately recognize actions even in the presence of context bias.
Qiankun Li 0004, Xiaolong Huang 0001, Zhifan Wan, Lanqing Hu, Shuzhe Wu, Jie Zhang 0071, Shiguang Shan, Zengfu Wang
ACM Multimedia5
2022 Attribute Group Editing for Reliable Few-shot Image Generation
abstract
Few-shot image generation is a challenging task even using the state-of-the-art Generative Adversarial Networks (GANs). Due to the unstable GAN training process and the limited training data, the generated images are often of low quality and low diversity. In this work, we propose a new “editing-based” method, i.e., Attribute Group Editing (AGE), for few-shot image generation. The basic assumption is that any image is a collection of attributes and the editing direction for a specific attribute is shared across all categories. AGE examines the internal representation learned in GANs and identifies semantically meaningful directions. Specifically, the class embedding, i.e., the mean vector of the latent codes from a specific category, is used to represent the category-relevant attributes, and the category-irrelevant attributes are learned globally by Sparse Dictionary Learning on the difference between the sample embedding and the class embedding. Given a GAN well trained on seen categories, diverse images of unseen categories can be synthesized through editing category-irrelevant attributes while keeping category-relevant attributes unchanged. Without re-training the GAN, AGE is capable of not only producing more realistic and diverse images for downstream visual applications with limited data but achieving controllable image editing with interpretable category-irrelevant directions. Code is available at https://github.com/UniBester/AGE.
Guanqi Ding, Xinzhe Han, Shuhui Wang, Shuzhe Wu, Xin Jin 0004, Dandan Tu, Qingming Huang
CVPR4
2022 Personalized Convolution for Face Recognition
Chunrui Han, Shiguang Shan, Meina Kan, Shuzhe Wu, Xilin Chen 0001
Int. J. Comput. Vis.4
2019 Hierarchical Attention for Part-Aware Face Detection
Shuzhe Wu, Meina Kan, Shiguang Shan, Xilin Chen 0001
Int. J. Comput. Vis.1
2018 Real-Time Rotation-Invariant Face Detection With Progressive Calibration Networks
abstract
Rotation-invariant face detection, i.e. detecting faces with arbitrary rotation-in-plane (RIP) angles, is widely required in unconstrained applications but still remains as a challenging task, due to the large variations of face appearances. Most existing methods compromise with speed or accuracy to handle the large RIP variations. To address this problem more efficiently, we propose Progressive Calibration Networks (PCN) to perform rotation-invariant face detection in a coarse-to-fine manner. PCN consists of three stages, each of which not only distinguishes the faces from non-faces, but also calibrates the RIP orientation of each face candidate to upright progressively. By dividing the calibration process into several progressive steps and only predicting coarse orientations in early stages, PCN can achieve precise and fast calibration. By performing binary classification of face vs. non-face with gradually decreasing RIP ranges, PCN can accurately detect faces with full 360° RIP angles. Such designs lead to a real-time rotation-invariant face detector. The experiments on multi-oriented FDDB and a challenging subset of WIDER FACE containing rotated faces in the wild show that our PCN achieves quite promising performance.
Xuepeng Shi, Shiguang Shan, Meina Kan, Shuzhe Wu, Xilin Chen 0001
CVPR4
2018 Face Recognition with Contrastive Convolution
Chunrui Han, Shiguang Shan, Meina Kan, Shuzhe Wu, Xilin Chen 0001
ECCV (9)4
2018 Face Anti-Spoofing with Multi-Scale Information
abstract
Face anti-spoofing has encountered increasing demand as one of the key technologies for reliable and safe authentication with faces. Current face anti-spoofing methods generally take a single crop of face region as input for classification, i.e. exploiting information at only one scale. This single-scale scheme mainly focuses on facial characteristics but not utilize the surrounding information, causing poor generalization for different scenarios with varied means of attacks. Besides, it is tedious or highly empirical to determine an optimal scale of face crops. To overcome the limitations of single-scale methods, in this work we propose to integrate Multi-Scale information for better Face ANti-Spoofing (MS-FANS). Specifically, the proposed MS-FANS method takes multiple face crops at different scales as input followed by a convolutional neural network (CNN) for feature extraction. Then the features from different scales form as a sequence, which are fed into a Long Short-Term Memory (LSTM) network for adaptive fusion of multi-scale information, constructing the final representation for classification. Benefited from this multi-scale design, MS-FANS can adaptively utilize context information from multiple scales, leading to promising performance on two challenging face anti-spoofing datasets, Idiap REPLAY-ATTACK and CASIA-FASD, with significant improvement compared with the existing methods.
Shiying Luo, Meina Kan, Shuzhe Wu, Xilin Chen 0001, Shiguang Shan
ICPR3
2017 Funnel-structured cascade for multi-view face detection with alignment-awareness
Shuzhe Wu, Meina Kan, Zhenliang He, Shiguang Shan, Xilin Chen 0001
Neurocomputing1
2015 High-level semantic image annotation based on hot Internet topics
abstract
Images are complex multimedia data that contain rich semantic information. Currently, most of image annotation algorithms are only annotating the object semantics of images. There are still many challenges on high-level semantic image annotation. The major issues are the lack of effective modeling method for the high-level semantics of images and the lack of efficient dynamic update mechanism for the training set. To address these issues, we propose a high-level semantic annotation method based on hot Internet topics in this paper. There are two independent sub tasks in our method: dynamic update of the training set based on hot Internet topics and search-based image annotation. In the first sub task, we propose to model the abstract semantics of images based on three relationships: image–to–image similarity relationship, topic–to–topic co-occurrence relationship, and image–to–topic relevance relationship. Through the complex graph clustering, the hot Internet topics are extracted for images with consistent visual and semantic contents. Then the dynamic update mechanism will update the original training set with the new topics and images. It avoids the huge computing cost in traditional update methods and does not need to re-calculate the whole mapping relationship between the semantic concepts and visual features. In the second sub task, given a query image, it first searches for similar candidates in the annotated training set via visual features. Then the hypergraph modeling and spectral clustering are exploited to filter out the images with irrelevant semantics. The keywords will be extracted for annotation from the remaining images according to an annotation probability. Extensive experiments have been conducted and the results demonstrate that our algorithm could achieve better annotation performance than the state-of-the-art algorithms. And the update mechanism could extend the training set efficiently so that the coverage of the semantics in the training set wouldn’t be obsolete.
Xiaoru Wang, Junping Du 0001, Shuzhe Wu, Haiming Xin, Fu Li 0004
Multim. Tools Appl.3