Xuhao Jiang

dblp:228/3267 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
15since 2021 · last 2025
0000-0002-4646-5052ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 5 first-author · 15 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Universal Backdoor Defense via Label Consistency in Vertical Federated Learning
abstract
Backdoor attacks in vertical federated learning (VFL) are particularly concerning as they can covertly compromise VFL decision-making, posing a severe threat to critical applications of VFL. Existing defense mechanisms typically involve either label obfuscation during training or model pruning during inference. However, the inherent limitations on the defender's access to the global model and complete training data in VFL environments fundamentally constrain the effectiveness of these conventional methods. To address these limitations, we propose the Universal Backdoor Defense (UBD) framework. UBD leverages Label Consistent Clustering (LCC) to synthesize plausible latent triggers associated with the backdoor class. This synthesized information is then utilized for mitigating backdoor threats through Linear Probing (LP), guided by a constraint on Batch Normalization (BN) statistics. Positioned within a unified VFL backdoor defense paradigm, UBD offers a generalized framework for both detection and mitigation that critically does not necessitate access to the entire model or dataset. Extensive experiments across multiple datasets rigorously demonstrate the efficacy of the UBD framework, achieving state-of-the-art performance against diverse backdoor attack types in VFL, including both dirty-label and clean-label variants.
Peng Chen 0030, Haolong Xiang, Xin Du 0002, Xiaolong Xu 0001, Xuhao Jiang, Zhihui Lu 0002, Jirui Yang, Qiang Duan 0002, Wan-Chun Dou
IJCAI5
2024 Context-Aware Iteration Policy Network for Efficient Optical Flow Estimation
abstract
Existing recurrent optical flow estimation networks are computationally expensive since they use a fixed large number of iterations to update the flow field for each sample. An efficient network should skip iterations when the flow improvement is limited. In this paper, we develop a Context-Aware Iteration Policy Network for efficient optical flow estimation, which determines the optimal number of iterations per sample. The policy network achieves this by learning contextual information to realize whether flow improvement is bottlenecked or minimal. On the one hand, we use iteration embedding and historical hidden cell, which include previous iterations information, to convey how flow has changed from previous iterations. On the other hand, we use the incremental loss to make the policy network implicitly perceive the magnitude of optical flow improvement in the subsequent iteration. Furthermore, the computational complexity in our dynamic network is controllable, allowing us to satisfy various resource preferences with a single trained model. Our policy network can be easily integrated into state-of-the-art optical flow networks. Extensive experiments show that our method maintains performance while reducing FLOPs by about 40%/20% for the Sintel/KITTI datasets.
Ri Cheng, Ruian He, Xuhao Jiang, Shili Zhou, Weimin Tan, Bo Yan 0001
AAAI3
2024 Bridging The Domain Gap Arising from Text Description Differences for Stable Text-To-Image Generation
abstract
Generating high-quality images that conform to the semantics of captions has numerous potential applications. However, text-to-image generation is a challenging task due to its cross-modality nature. Current generative models are typically unstable, meaning that complex sentences can result in poor image quality. In this paper, we propose a novel model to bridge the domain gap arising from sentence complexity to achieve stable text-to-image generation. Our model includes two key modules, the attribute extraction module and the attribute fusion module. These modules can extract attributes from the captions and fuse them with image features to encourage the model to accurately understand the semantics. Our modules are plug-and-play and extensive experiments demonstrate that our approach outperforms the state-of-the-art GAN model. Our code and trained model are available at https://github.com/tantian21/stable-t2i-generation.
Tian Tan 0016, Weimin Tan, Xuhao Jiang, Yueming Jiang, Bo Yan 0001
ICASSP3
2024 Facial Micro-Motion-Aware Mixup for Micro-Expression Recognition
abstract
Data-driven learning models have demonstrated strong benefits in capturing subtle facial movements for micro-expression recognition (MER), but are limited by the available data. Generative models can generate a variety of new data, but are typically computationally prohibitive compared to efficient Mixup-like methods. In this paper, we propose a novel Facial Micro-Motion-Aware Mixup approach for MER, namely MEMix. Our MEMix constructs a micro-motion-aware mask to select the most salient facial motions and generate a new sample with a mixed motion feature. This mixed motion feature can effectively expand the data distribution, leading to smoother decision boundaries for MER models. To demonstrate the good generality of MEMix, we integrate it with three advanced vision transformer-based models. The results show that the three integrated models consistently achieve performance improvements ranging from 4.07% to 7.32% in accuracy and from 6.54% to 9.18% in F1-score. Besides, to further explore the ability of MEMix, we propose a two-stream network called MixMeFormer, which unlocks the potential of the transformer by simply integrating mixed motion features with facial semantics for MER. Extensive experiments demonstrate that our MixMeFormer outperforms other state-of-the-art methods on three well-known micro-expression datasets.
Zhuoyao Gu, Miao Pang, Weimin Tan, Xuhao Jiang, Bo Yan 0001
ICASSP5
2024 Addressing Imbalance for Class Incremental Learning in Medical Image Classification
abstract
Deep convolutional neural networks have made significant breakthroughs in medical image classification, under the assumption that training samples from all classes are simultaneously available. However, in real-world medical scenarios, there's a common need to continuously learn about new diseases, leading to the emerging field of class incremental learning (CIL) in the medical domain. Typically, CIL suffers from catastrophic forgetting when trained on new classes. This phenomenon is mainly caused by the imbalance between old and new classes, and it becomes even more challenging with imbalanced medical datasets. In this work, we introduce two simple yet effective plug-in methods to mitigate the adverse effects of the imbalance. First, we propose a CIL-balanced classification loss to mitigate the classifier bias toward majority classes via logit adjustment. Second, we propose a distribution margin loss that not only alleviates the inter-class overlap in embedding space but also enforces the intra-class compactness. We evaluate the effectiveness of our method with extensive experiments on three benchmark datasets (CCH5000, HAM10000, and EyePACS). The results demonstrate that our approach outperforms state-of-the-art methods.
Xuze Hao, Wenqian Ni, Xuhao Jiang, Weimin Tan, Bo Yan 0001
ACM Multimedia3
2024 Prompt-Guided Semantic-Aware Distillation for Weakly Supervised Incremental Semantic Segmentation
abstract
Weakly Supervised Incremental Semantic Segmentation (WISS) aims to enable deep neural networks to incrementally learn new classes using only image-level labels without catastrophic forgetting. Despite WISS eliminating the usage of costly and time-consuming pixel-by-pixel annotations, the image-level labels can not provide details about the location of new classes, resulting in inferior performance. To address these issues, we take inspiration from zero-shot learning to model the inter-class semantic relation utilizing class names as text prompts, thereby facilitating knowledge transfer between classes. However, some class names of the segmentation datasets are polysemous. Thus, we design a new prompt template to better capture the semantic relation by appending synonyms and definitions of the corresponding classes. Guided by this semantic relation, we propose semantic relation weighted distillation to transfer the knowledge from old to new classes, significantly improving plasticity while reducing forgetting. Additionally, we introduce a novel superclass-level distillation aimed at preserving shared global knowledge within the superclass, further alleviating catastrophic forgetting. We extensively evaluate our method by integrating it into state-of-the-art WISS approaches on Pascal VOC and COCO datasets. We observe consistent gains in performance across diverse experimental scenarios. Code is available athttps://github.com/Magic-Nova77/PGSD.
Xuze Hao, Xuhao Jiang, Wenqian Ni, Weimin Tan, Bo Yan 0001
IEEE Trans. Circuits Syst. Video Technol.2
2023 Multi-Modality Deep Network for Extreme Learned Image Compression
abstract
Image-based single-modality compression learning approaches have demonstrated exceptionally powerful encoding and decoding capabilities in the past few years , but suffer from blur and severe semantics loss at extremely low bitrates. To address this issue, we propose a multimodal machine learning method for text-guided image compression, in which the semantic information of text is used as prior information to guide image compression for better compression performance. We fully study the role of text description in different components of the codec, and demonstrate its effectiveness. In addition, we adopt the image-text attention module and image-request complement module to better fuse image and text features, and propose an improved multimodal semantic-consistent loss to produce semantically complete reconstructions. Extensive experiments, including a user study, prove that our method can obtain visually pleasing results at extremely low bitrates, and achieves a comparable or even better performance than state-of-the-art methods, even though these methods are at 2x to 4x bitrates of ours.
Xuhao Jiang, Weimin Tan, Tian Tan 0016, Bo Yan 0001, Liquan Shen
AAAI1
2023 VQA-CLPR: Turning a Visual Question Answering Model into a Chinese License Plate Recognizer
Xuhao Jiang, Yining Sun, Weiya Ni, Fudong Nian
ICIG (2)2
2023 Multi-Modality Deep Network for JPEG Artifacts Reduction
abstract
In recent years, many convolutional neural network-based models are designed for JPEG artifacts reduction, and have achieved notable progress. However, few methods are suitable for extreme low-bitrate image compression artifacts reduction. The main challenge is that the highly compressed image loses too much information, resulting in reconstructing high-quality image difficultly. To address this issue, we propose a multimodal fusion learning method for text-guided JPEG artifacts reduction, in which the corresponding text description not only provides the potential prior information of the highly compressed image, but also serves as supplementary information to assist in image deblocking. We fuse image features and text semantic features from the global and local perspectives respectively, and design a contrastive loss built upon contrastive learning to produce visually pleasing results. Extensive experiments, including a user study, prove that our method can obtain better deblocking results compared to the state-of-the-art methods.
Xuhao Jiang, Weimin Tan, Chenxi Ma, Bo Yan 0001, Liquan Shen
IJCAI1
2023 Uncertainty-Guided Spatial Pruning Architecture for Efficient Frame Interpolation
abstract
The video frame interpolation (VFI) model applies the convolution operation to all locations, leading to redundant computations in regions with easy motion. We can use dynamic spatial pruning method to skip redundant computation, but this method cannot properly identify easy regions in VFI tasks without supervision. In this paper, we develop an Uncertainty-Guided Spatial Pruning (UGSP) architecture to skip redundant computation for efficient frame interpolation dynamically. Specifically, pixels with low uncertainty indicate easy regions, where the calculation can be reduced without bringing undesirable visual results. Therefore, we utilize uncertainty-generated mask labels to guide our UGSP in properly locating the easy region. Furthermore, we propose a self-contrast training strategy that leverages an auxiliary non-pruning branch to improve the performance of our UGSP. Extensive experiments show that UGSP maintains performance but reduces FLOPs by 34%/52%/30% compared to baseline without pruning on Vimeo90K/UCF101/MiddleBury datasets. In addition, our method achieves state-of-the-art performance with lower FLOPs on multiple benchmarks.
Ri Cheng, Xuhao Jiang, Ruian He, Shili Zhou, Weimin Tan, Bo Yan 0001
ACM Multimedia2
2023 MVFlow: Deep Optical Flow Estimation of Compressed Videos with Motion Vector Prior
abstract
In recent years, many deep learning-based methods have been proposed to tackle the problem of optical flow estimation and achieved promising results. However, they hardly consider that most videos are compressed and thus ignore the pre-computed information in compressed video streams. Motion vectors, one of the compression information, record the motion of the video frames. They can be directly extracted from the compression code stream without computational cost and serve as a solid prior for optical flow estimation. Therefore, we propose an optical flow model, MVFlow, which uses motion vectors to improve the speed and accuracy of optical flow estimation for compressed videos. In detail, MVFlow includes a key Motion-Vector Converting Module, which ensures that the motion vectors can be transformed into the same domain of optical flow and then be utilized fully by the flow estimation module. Meanwhile, we construct four optical flow datasets for compressed videos containing frames and motion vectors in pairs. The experimental results demonstrate the superiority of our proposed MVFlow, which can reduce the AEPE by 1.09 compared to existing models or save 52% time to achieve similar accuracy to existing models.
Shili Zhou, Xuhao Jiang, Weimin Tan, Ruian He, Bo Yan 0001
ACM Multimedia2
2022 Learning Parallax Transformer Network for Stereo Image JPEG Artifacts Removal
abstract
Under stereo settings, the performance of image JPEG artifacts removal can be further improved by exploiting the additional information provided by a second view. However, incorporating this information for stereo image JPEG artifacts removal is a huge challenge, since the existing compression artifacts make pixel-level view alignment difficult. In this paper, we propose a novel parallax transformer network (PTNet) to integrate the information from stereo image pairs for stereo image JPEG artifacts removal. Specifically, a well-designed symmetric bi-directional parallax transformer module is proposed to match features with similar textures between different views instead of pixel-level view alignment. Due to the issues of occlusions and boundaries, a confidence-based cross-view fusion module is proposed to achieve better feature fusion for both views, where the cross-view features are weighted with confidence maps. Especially, we adopt a coarse-to-fine design for the cross-view interaction, leading to better performance. Comprehensive experimental results demonstrate that our PTNet can effectively remove compression artifacts and achieves superior performance than other testing state-of-the-art methods.
Xuhao Jiang, Weimin Tan, Ri Cheng, Shili Zhou, Bo Yan 0001
ACM Multimedia1
2021 Perception-Oriented Stereo Image Super-Resolution
abstract
Recent studies of deep learning based stereo image super-resolution (StereoSR) have promoted the development of StereoSR. However, existing StereoSR models mainly concentrate on improving quantitative evaluation metrics and neglect the visual quality of super-resolved stereo images. To improve the perceptual performance, this paper proposes the first perception-oriented stereo image super-resolution approach by exploiting the feedback, provided by the evaluation on the perceptual quality of StereoSR results. To provide accurate guidance for the StereoSR model, we develop the first special stereo image super-resolution quality assessment (StereoSRQA) model, and further construct a StereoSRQA database. Extensive experiments demonstrate that our StereoSR approach significantly improves the perceptual quality and enhances the reliability of stereo images for disparity estimation.
Chenxi Ma, Bo Yan 0001, Weimin Tan, Xuhao Jiang
ACM Multimedia4
2021 An optimized CNN-based quality assessment model for screen content image
Xuhao Jiang, Liquan Shen, Guorui Feng, Liangwei Yu, Ping An 0001
Signal Process. Image Commun.1
2021 A Distortion-Aware Multi-Task Learning Framework for Fractional Interpolation in Video Coding
abstract
Motion-compensated prediction adopts fractional-pixel interpolation to obtain the best motion vector. Traditional fixed interpolation filters cannot handle various content and structures well, and existing convolutional neural network based methods cannot fully exploit the distortion characteristics for fractional interpolation. Therefore, this paper proposes a distortion-aware multi-task learning framework (DA-MLF) to perform fractional interpolation. First, a multi-task training framework is proposed to provide the distortion characteristics as complementary information for improving the performance of subsequent interpolation. Then, a uniform interpolation sub-network is proposed to accomplish fractional interpolation, which utilizes the feature fusion module to fuse abundant local features, and the distortion awareness module to capture the multi-scale information of compression artifacts. Furthermore, DA-MLF is integrated into High Efficiency Video Coding (HEVC) test model, and multiple experiments are performed to evaluate the effectiveness of our method. On HEVC testing sequences, DA-MLF achieves 5.0%, 4.0% and 1.7% BD-rate reduction on average compared to the HEVC baseline, under low-delay P, low-delay B and random-access configurations, respectively. The experimental results validate that our framework not only achieves the best interpolation performance but also has the lowest computational complexity compared with state-of-the-art methods.
Liangwei Yu, Liquan Shen, Hao Yang 0008, Xuhao Jiang, Bo Yan 0001
IEEE Trans. Circuits Syst. Video Technol.4
2020 No-reference screen content image quality assessment based on multi-region features
Xuhao Jiang, Liquan Shen, Liangwei Yu, Mingxing Jiang, Guorui Feng
Neurocomputing1
2020 Screen content image quality assessment based on convolutional neural networks
Xuhao Jiang, Liquan Shen, Linru Zheng, Ping An 0001
J. Vis. Commun. Image Represent.1
2018 Naturalization Module in Neural Networks for Screen Content Image Quality Assessment
abstract
Deep learning approaches have demonstrated success in no-reference image quality assessment tasks. However, due to the specific properties of screen content images (SCIs), deep neural networks for SCI quality assessment are not as optimal as those designed for images depicting natural scenes. In order to tackle this discrepancy, a “naturalization” module composed of an upsampling layer and a convolutional layer is proposed to transform SCIs to have characteristics more similar to that of natural images. In addition, a new deep learning model architecture along with data augmentation techniques tailored to SCIs are implemented. The performance of the proposed approach is evaluated on the Screen Image Quality Assessment Database and Screen Content Image Database, and has shown to have superior performance to state-of-the-art methods in predicting the perceptual quality of SCIs.
Jianan Chen 0001, Liquan Shen, Linru Zheng, Xuhao Jiang
IEEE Signal Process. Lett.4