Youze Xue

dblp:258/6782 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0002-7054-5204ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Improving Adversarial Robustness Against Universal Patch Attacks Through Feature Norm Suppressing
abstract
Universal adversarial patch attacks, which are readily implemented, have been validated to be able to fool real-world deep convolutional neural networks (CNNs), posing a serious threat to practical computer vision systems based on CNNs. Unfortunately, current defending approaches are severely understudied facing the following problems. Patch detection-based methods suffer from dramatic performance drops against white-box or adaptive attacks since they rely heavily on empirical clues. Methods based on adversarial training or certified defense are difficult to be scaled up to large-scale datasets or complex practical networks due to prohibitively high computational overhead or over strong assumptions on the network structure. In this article, we focus on two cases of widely adopted universal adversarial patch attacks, namely the universal targeted attack on image classifiers and the universal vanishing attack on object detectors. We find that, for popular CNNs, the attacking success of the adversarial patch relies on feature vectors centered at the patch location with large norm in classifiers and large channel-aware norm (CA-Norm) in detectors, and further present a mathematical explanation for this phenomenon. Based on this, we propose a simple but effective defending method using the feature norm suppressing (FNS) layer, which can renormalize the feature norm by nonincreasing functions. As a differentiable module, FNS can be adaptively inserted in various CNN architectures to achieve multistage suppression of the generation of large norm feature vectors. Moreover, FNS is efficient with no trainable parameters and very low computational overhead. We evaluate our proposed defending method across multiple CNN architectures and datasets against the strong adaptive white-box attacks in both visual classification and detection tasks. In both tasks, FNS significantly outperforms previous defending methods on adversarial robustness with a relatively low influence on the performance of benign images. Code is available at https://github.com/jschenthu/FNS.
Jiansheng Chen 0001, Yu Wang 0002, Youze Xue, Huimin Ma 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Upper-Body Hierarchical Graph for Skeleton Based Emotion Recognition in Assistive Driving
Jiehui Wu, Jiansheng Chen 0001, Qifeng Luo, Siqi Liu 0010, Youze Xue, Huimin Ma 0001
ECCV (26)5
2024 Center of Pressure Estimation by Analyzing Walking Videos
abstract
Center of pressure (COP) serves as a widely utilized indicator for evaluating balance-related issues, e.g., gait quality of neurological disorders, fall risk of the elderly, and recovery of the injured. Existing methods for acquiring COP mostly rely on expensive force platforms or wearable force-sensing sensors that are relatively difficult to deploy. We propose a novel low-cost strategy to estimate COP with only visual information. Specifically, we first collect a database with walking videos. The plantar pressure of the subjects during walking is synchronously collected for calculating COP ground truth. Then, we propose an encoder-decoder model for estimating COPs using human pose and body shape as inputs. A perceptual codebook is used to tolerate the error in human pose estimation to improve the accuracy of the COP estimation. Experiments demonstrate the effectiveness of our proposed method. Our proposal achieves a correlation coefficient of 0.962, which is 0.8% better than the baseline model using Multi-layer Perceptron (MLP). The normalized RMSEs are improved by 13.9% and 19.0% in the anterior-posterior and medial-lateral directions, respectively. Compared to existing methods using wearable sensors or plantar pressure plates, our method is cheaper and easier to deploy.
Jiansheng Chen 0002, Yining Qin, Poyu Lin, Jiawei Li 0016, Youze Xue, Huimin Ma 0001
ICASSP5
2024 Refining 3D Human Mesh via Model-Free Offsets Estimation
abstract
3D human mesh reconstruction from a single RGB image is a challenging task. Existing methods either utilize parametric mesh models to restrain the 3D human structures or directly regress the 3D coordinates of the mesh vertices. The former ones, called model-based methods, usually fail to recover the high variance of human mesh due to limited capacity of parametric models, whereas the latter ones called model-free methods suffer from unrealistic human structures because of lack of 3D priors. To mitigate the drawbacks of them, we propose that the model-based reconstruction can serve as a good starting point for the model-free refinement, so that the 3D structure priors of the parametric model and the high representation capability of model-free methods can be both inherited. By building a model-free refinement head upon a pretrained model-based regressor, our method reduces the reconstruction errors of 3D human mesh on public datasets H36M and 3DPW, demonstrating the advantage of combining model-based and model-free methods together.
Youze Xue, Hongbing Ma, Huimin Ma 0001
ICASSP1
2022 3D Human Mesh Reconstruction by Learning to Sample Joint Adaptive Tokens for Transformers
abstract
Reconstructing 3D human mesh from a single RGB image is a challenging task due to the inherent depth ambiguity. Researchers commonly use convolutional neural networks to extract features and then apply spatial aggregation on the feature maps to explore the embedded 3D cues in the 2D image. Recently, two methods of spatial aggregation, the transformers and the spatial attention, are adopted to achieve the state-of-the-art performance, whereas they both have limitations. The use of transformers helps modelling long-term dependency across different joints whereas the grid tokens are not adaptive for the positions and shapes of human joints in different images. On the contrary, the spatial attention focuses on joint-specific features. However, the non-local information of the body is ignored by the concentrated attention maps. To address these issues, we propose a Learnable Sampling module to generate joint adaptive tokens and then use transformers to aggregate global information. Feature vectors are sampled accordingly from the feature maps to form the tokens of different joints. The sampling weights are predicted by a learnable network so that the model can learn to sample joint-related features adaptively. Our adaptive tokens are explicitly correlated with human joints, so that more effective modeling of global dependency among different human joints can be achieved. To validate the effectiveness of our method, we conduct experiments on several popular datasets including Human3.6M and 3DPW. Our method achieves lower reconstruction errors in terms of both the vertex-based metric and the joint-based metric compared to previous state of the arts. The codes and the trained models are released at https://github.com/thuxyz19/Learnable-Sampling.
Youze Xue, Jiansheng Chen 0001, Yudong Zhang 0008, Huimin Ma 0001, Hongbing Ma
ACM Multimedia1
2022 Boosting Monocular 3D Human Pose Estimation With Part Aware Attention
abstract
Monocular 3D human pose estimation is challenging due to depth ambiguity. Convolution-based and Graph-Convolution-based methods have been developed to extract 3D information from temporal cues in motion videos. Typically, in the lifting-based methods, most recent works adopt the transformer to model the temporal relationship of 2D keypoint sequences. These previous works usually consider all the joints of a skeleton as a whole and then calculate the temporal attention based on the overall characteristics of the skeleton. Nevertheless, the human skeleton exhibits obvious part-wise inconsistency of motion patterns. It is therefore more appropriate to consider each part's temporal behaviors separately. To deal with such part-wise motion inconsistency, we propose the Part Aware Temporal Attention module to extract the temporal dependency of each part separately. Moreover, the conventional attention mechanism in 3D pose estimation usually calculates attention within a short time interval. This indicates that only the correlation within the temporal context is considered. Whereas, we find that the part-wise structure of the human skeleton is repeating across different periods, actions, and even subjects. Therefore, the part-wise correlation at a distance can be utilized to further boost 3D pose estimation. We thus propose the Part Aware Dictionary Attention module to calculate the attention for the part-wise features of input in a dictionary, which contains multiple 3D skeletons sampled from the training set. Extensive experimental results show that our proposed part aware attention mechanism helps a transformer-based model to achieve state-of-the-art 3D pose estimation performance on two widely used public datasets. The codes and the trained models are released at https://github.com/thuxyz19/3D-HPE-PAA.
Youze Xue, Jiansheng Chen 0001, Xiangming Gu, Huimin Ma 0001, Hongbing Ma
IEEE Trans. Image Process.1
2021 Defending against Universal Adversarial Patches by Clipping Feature Norms
abstract
Physical-world adversarial attacks based on universal adversarial patches have been proved to be able to mislead deep convolutional neural networks (CNNs), exposing the vulnerability of real-world visual classification systems based on CNNs. In this paper, we empirically reveal and mathematically explain that the universal adversarial patches usually lead to deep feature vectors with very large norms in popular CNNs. Inspired by this, we propose a simple yet effective defending approach using a new feature norm clipping (FNC) layer which is a differentiable module that can be flexibly inserted in different CNNs to adaptively suppress the generation of large norm deep feature vectors. FNC introduces no trainable parameter and only very low computational overhead. However, experiments on multiple datasets validate that it can effectively improve the robustness of different CNNs towards white-box universal patch attacks while maintaining a satisfactory recognition accuracy for clean samples.
Youze Xue, Weitao Wan, Jiayu Bao, Huimin Ma 0001
ICCV3
2021 Enhancing Adversarial Robustness For Image Classification By Regularizing Class Level Feature Distribution
abstract
Recent researches have shown that deep neural networks (DNNs) are vulnerable to adversarial examples. Adversarial training is practically the most effective approach to improve the robustness of DNNs against adversarial examples. However, conventional adversarial training methods only focus on the classification results or the instance level relationship on feature representations for adversarial examples. Inspired by the fact that adversarial examples break the distinguishability of the feature representations of DNNs for different classes, we propose Intra and Inter Class Feature Regularization $(\mathrm{I}^{2}$ FR) to make the feature distribution of adversarial examples maintain the same classification property as clean examples. On the one hand, the intra-class regularization restricts the distance of features between adversarial examples and both the corresponding clean data and samples for the same class. On the other hand, the inter-class regularization prevents the feature of adversarial examples from getting close to other classes. By adding $\mathrm{I}^{2}$ FR in both adversarial example generation and model training steps in adversarial training, we can get stronger and more diverse adversarial examples, and the neural network learns a more distinguishable and reasonable feature distribution. Experiments on various adversarial training frameworks demonstrate that $\mathrm{I}^{2}$ FR is adaptive for multiple training frameworks and outperforms the state-of-the-art methods for classification of both clean data and adversarial examples.
Youze Xue, Jiansheng Chen 0001, Yu Wang 0002, Huimin Ma 0001
ICIP2
2020 Image Captioning With End-to-End Attribute Detection and Subsequent Attributes Prediction
abstract
Semantic attention has been shown to be effective in improving the performance of image captioning. The core of semantic attention based methods is to drive the model to attend to semantically important words, or attributes. In previous works, the attribute detector and the captioning network are usually independent, leading to the insufficient usage of the semantic information. Also, all the detected attributes, no matter whether they are appropriate for the linguistic context at the current step, are attended to through the whole caption generation process. This may sometimes disrupt the captioning model to attend to incorrect visual concepts. To solve these problems, we introduce two end-to-end trainable modules to closely couple attribute detection with image captioning as well as prompt the effective uses of attributes by predicting appropriate attributes at each time step. The multimodal attribute detector (MAD) module improves the attribute detection accuracy by using not only the image features but also the word embedding of attributes already existing in most captioning models. MAD models the similarity between the semantics of attributes and the image object features to facilitate accurate detection. The subsequent attribute predictor (SAP) module dynamically predicts a concise attribute subset at each time step to mitigate the diversity of image attributes. Compared to previous attribute based methods, our approach enhances the explainability in how the attributes affect the generated words and achieves a state-of-the-art single model performance of 128.8 CIDEr-D on the MSCOCO dataset. Extensive experiments on the MSCOCO dataset show that our proposal actually improves the performances in both image captioning and attribute detection simultaneously. The codes are available at: https://github.com/ RubickH/Image-Captioning-with-MAD-and-SAP.
Jiansheng Chen 0001, Wanli Ouyang, Weitao Wan, Youze Xue
IEEE Trans. Image Process.5
2019 Information Entropy Based Feature Pooling for Convolutional Neural Networks
abstract
In convolutional neural networks (CNNs), we propose to estimate the importance of a feature vector at a spatial location in the feature maps by the network's uncertainty on its class prediction, which can be quantified using the information entropy. Based on this idea, we propose the entropy-based feature weighting method for semantics-aware feature pooling which can be readily integrated into various CNN architectures for both training and inference. We demonstrate that such a location-adaptive feature weighting mechanism helps the network to concentrate on semantically important image regions, leading to improvements in the large-scale classification and weakly-supervised semantic segmentation tasks. Furthermore, the generated feature weights can be utilized in visual tasks such as weakly-supervised object localization. We conduct extensive experiments on different datasets and CNN architectures, outperforming recently proposed pooling methods and attention mechanisms in ImageNet classification as well as achieving state-of-the-arts in weakly-supervised semantic segmentation on PASCAL VOC 2012 dataset.
Weitao Wan, Tianpeng Li, Jingqi Tian, Youze Xue
ICCV7
2019 MVSCRF: Learning Multi-View Stereo With Conditional Random Fields
abstract
We present a deep-learning architecture for multi-view stereo with conditional random fields (MVSCRF). Given an arbitrary number of input images, we first use a U-shape neural network to extract deep features incorporating both global and local information, and then build a 3D cost volume for the reference camera. Unlike previous learning based methods, we explicitly constraint the smoothness of depth maps by using conditional random fields (CRFs) after the stage of cost volume regularization. The CRFs module is implemented as recurrent neural networks so that the whole pipeline can be trained end-to-end. Our results show that the proposed pipeline outperforms previous state-of-the-arts on large-scale DTU dataset. We also achieve comparable results with state-of-the-art learning based methods on outdoor Tanks and Temples dataset without fine-tuning, which demonstrates our method's generalization ability.
Youze Xue, Weitao Wan, Tianpeng Li, Jiayu Bao
ICCV1