Wanxia Deng

dblp:212/0224 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
15since 2021 · last 2025
0000-0002-4179-0059ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Artificial neural networks for finger vein recognition: A survey
Yimin Yin, Renye Zhang, Wanxia Deng, Dayu Hu, Siliang He
Eng. Appl. Artif. Intell.4
2025 Integrated Retrieval of the Temperature and Humidity Profiles of Atmospheric Boundary Layer by Combining Ground-Based Infrared Hyperspectral Interferometers and Microwave Radiometers
abstract
Atmospheric temperature and humidity profiles are the basic parameters used to describe the vertical distribution of atmospheric states. Continuous observations of accurate temperature and humidity profiles are essential for exploring boundary layer thermal and dynamic characteristics. To this end, an intelligent retrieval algorithm (IReA) based on a convolutional neural network (CNN) is proposed to retrieve atmospheric temperature and humidity profiles by combining observations from ground-based infrared hyperspectral radiometers and microwave radiometers (MWRs). The results show that the inclusion of microwave observations can effectively improve the retrieval accuracy of temperature and humidity profiles relative to the results from atmospheric emitted radiance interferometer (AERI) under clear-sky conditions, where the root mean square error (RMSE) of the temperature profile is 0.79 K and the RMSE of the humidity profile is 0.95 g/kg. The accuracies of different retrieval methods are also evaluated. In general, the RMSE derived from IReA is improved by at least 9% compared to the results from the physical retrieval method and BP neural network method. Given that clouds are semitransparent in the microwave region, the retrieval accuracy of the temperature and humidity profile of IReA are also improved under cloudy conditions when microwave observations are introduced.
Shuai Hu, Wanxia Deng, Ruijun Dang, Lei Liu 0025, Wanying Yang
IEEE Trans. Geosci. Remote. Sens.3
2025 Refining Pseudo Labeling via Multi-Granularity Confidence Alignment for Unsupervised Cross Domain Object Detection
abstract
Most state-of-the-art object detection methods suffer from poor generalization due to the domain shift between training and testing datasets. To resolve this challenge, unsupervised cross domain object detection is proposed to learn an object detector for an unlabeled target domain by transferring knowledge from an annotated source domain. Promising results have been achieved via Mean Teacher, however, pseudo labeling which is the bottleneck of mutual learning remains to be further explored. In this study, we find that confidence misalignment of the predictions, including category-level overconfidence, instance-level task confidence inconsistency, and image-level confidence misfocusing, leading to the injection of noisy pseudo labels in the training process, will bring suboptimal performance. Considering the above issue, we present a novel general framework termed Multi-Granularity Confidence Alignment Mean Teacher (MGCAMT) for unsupervised cross domain object detection, which alleviates confidence misalignment across category-, instance-, and image-levels simultaneously to refine pseudo labeling for better teacher-student learning. Specifically, to align confidence with accuracy at category level, we propose Classification Confidence Alignment (CCA) to model category uncertainty based on Evidential Deep Learning (EDL) and filter out the category incorrect labels via an uncertainty-aware selection strategy. Furthermore, we design Task Confidence Alignment (TCA) to mitigate the instance-level misalignment between classification and localization by enabling each classification feature to adaptively identify the optimal feature for regression. Finally, we develop imagery Focusing Confidence Alignment (FCA) adopting another way of pseudo label learning, i.e., we use the original outputs from the Mean Teacher network for supervised learning without label assignment to achieve a balanced perception of the image's spatial layout. When these three procedures are integrated into a single framework, they mutually benefit to improve the final performance from a cooperative learning perspective. Extensive experiments across multiple scenarios demonstrate that our method outperforms large foundational models, and surpasses other state-of-the-art approaches by a large margin.
Jiangming Chen, Li Liu 0002, Wanxia Deng, Zhen Liu 0004, Yu Liu 0012, Yingmei Wei, Yongxiang Liu
IEEE Trans. Image Process.3
2024 Deep learning for liver cancer histopathology image analysis: A comprehensive survey
Haoyang Jiang, Yimin Yin, Wanxia Deng
Eng. Appl. Artif. Intell.4
2024 Uncertainty-Aware Distillation for Semi-Supervised Few-Shot Class-Incremental Learning
abstract
Given a model well-trained with a large-scale base dataset, few-shot class-incremental learning (FSCIL) aims at incrementally learning novel classes from a few labeled samples by avoiding overfitting, without catastrophically forgetting all encountered classes previously. Currently, semi-supervised learning technique that harnesses freely available unlabeled data to compensate for limited labeled data can boost the performance in numerous vision tasks, which heuristically can be applied to tackle issues in FSCIL, i.e., the semi-supervised FSCIL (Semi-FSCIL). So far, very limited work focuses on the Semi-FSCIL task, leaving the adaptability issue of semi-supervised learning to the FSCIL task unresolved. In this article, we focus on this adaptability issue and present a simple yet efficient Semi-FSCIL framework named uncertainty-aware distillation with class-equilibrium (UaD-ClE), encompassing two modules: uncertainty-aware distillation (UaD) and class equilibrium (ClE). Specifically, when incorporating unlabeled data into each incremental session, we introduce the ClE module that employs a class-balanced self-training (CB_ST) to avoid the gradual dominance of easy-to-classified classes on pseudo-label generation. To distill reliable knowledge from the reference model, we further implement the UaD module that combines uncertainty-guided knowledge refinement with adaptive distillation. Comprehensive experiments on three benchmark datasets demonstrate that our method can boost the adaptability of unlabeled data with the semi-supervised learning technique in FSCIL tasks. The code is available at https://github.com/yawencui/UaD-ClE.
Yawen Cui, Wanxia Deng, Haoyu Chen 0001, Li Liu 0002
IEEE Trans. Neural Networks Learn. Syst.2
2023 Variational Information Bottleneck for Cross Domain Object Detection
abstract
Cross domain object detection leverages a labeled source domain to learn an object detector which performs well in a novel unlabeled target domain. Most existing works mainly align the distribution utilizing the entire image knowledge ignoring the obstacles of task-uncorrelated information to alleviate the domain discrepancy. To tackle this issue, we propose a novel module called Variational Instance Disentanglement (VID) based on information theory which aims to decouple the information of task-correlated while filtering out the task-uncorrelated factors at the instance level. Notably, the proposed VID can be used as a plug-and-play module without bringing extra network parameter cost. We equip it with adversarial network and self-training network forming Variational Instance Disentanglement Adversarial Network (VIDAN) and Variational Instance Disentanglement Self-training Network (VIDSN), respectively. Extensive experiments on multiple widely-used scenarios show that the proposed method improves the performance of the popular frameworks and outperforms state-of-the-art methods.
Jiangming Chen, Wanxia Deng, Tianpeng Liu, Yingmei Wei, Li Liu 0002
ICME2
2023 Uncertainty-Guided Semi-Supervised Few-Shot Class-Incremental Learning With Knowledge Distillation
abstract
Class-Incremental Learning (CIL) aims at incrementally learning novel classes without forgetting old ones. This capability becomes more challenging when novel tasks contain one or a few labeled training samples, which leads to a more practical learning scenario,i.e., Few-Shot Class- Incremental Learning (FSCIL). The dilemma on FSCIL lies in serious overfitting and exacerbated catastrophic forgetting caused by the limited training data from novel classes. In this paper, excited by the easy accessibility of unlabeled data, we conduct a pioneering work and focus on a Semi-Supervised Few-Shot Class-Incremental Learning (Semi-FSCIL) problem, which requires the model incrementally to learn new classes from extremely limited labeled samples and a large number of unlabeled samples. To address this problem, a simple but efficient framework is first constructed based on the knowledge distillation technique to alleviate catastrophic forgetting. To efficiently mitigate the overfitting problem on novel categories with unlabeled data, uncertainty-guided semi-supervised learning is incorporated into this framework to select unlabeled samples into incremental learning sessions considering the model uncertainty. This process provides extra reliable supervision for the distillation process and contributes to better formulating the class means. Our extensive experiments on CIFAR100, miniImageNet and CUB200 datasets demonstrate the promising performance of our proposed method, and define baselines in this new research direction.
Yawen Cui, Wanxia Deng, Xin Xu 0001, Zhen Liu 0004, Zhong Liu 0002, Matti Pietikäinen, Li Liu 0002
IEEE Trans. Multim.2
2023 Importance-Aware Information Bottleneck Learning Paradigm for Lip Reading
abstract
Lip reading is the task of decoding text from speakers' mouth movements. Numerous deep learning-based methods have been proposed to address this task. However, these existing deep lip reading models suffer from poor generalization due to overfitting the training data. To resolve this issue, we present a novel learning paradigm that aims to improve the interpretability and generalization of lip reading models. In specific, a Variational Temporal Mask (VTM) module is customized to automatically analyze the importance of frame-level features. Furthermore, the prediction consistency constraints of global information and local temporal important features are introduced to strengthen the model generalization. We evaluate the novel learning paradigm with multiple lip reading baseline models on the LRW and LRW-1000 datasets. Experiments show that the proposed framework significantly improves the generalization performance and interpretability of lip reading models.
Changchong Sheng, Li Liu 0002, Wanxia Deng, Liang Bai 0003, Zhong Liu 0002, Songyang Lao, Gangyao Kuang, Matti Pietikäinen
IEEE Trans. Multim.3
2022 Boosting Lip Reading with a Multi-View Fusion Network
abstract
Lip reading aims to decode speech information by analyzing lip movement without involving audio. Numerous deep learning based methods are proposed to address this task. Generally, most existing methods extract visual features only based on the lip appearance, while ignoring the shape dynamic information of the lip region. Motivated by this, we propose a Multi-View Fusion Network (MVFN), which can extract more discriminative visual representations by incorporating appearance and shape information. Besides, a novel adaptive graph convolutional network model called Adaptive Spatial Graph Model(ASGM) is proposed to learn lip spatial topology and lip shape dynamics automatically. Experiments on LRW (word-level) and OuluVS2 (phrase-level) clearly show that the proposed method significantly outperforms the baseline methods by a large margin and achieves state-of-the-art performance.
Xueyi Zhang 0001, Jinping Sui, Changchong Sheng, Wanxia Deng, Li Liu 0002
ICME5
2022 Deep Ladder-Suppression Network for Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) aims at learning a classifier for an unlabeled target domain by transferring knowledge from a labeled source domain with a related but different distribution. Most existing approaches learn domain-invariant features by adapting the entire information of the images. However, forcing adaptation of domain-specific variations undermines the effectiveness of the learned features. To address this problem, we propose a novel, yet elegant module, called the deep ladder-suppression network (DLSN), which is designed to better learn the cross-domain shared content by suppressing domain-specific variations. Our proposed DLSN is an autoencoder with lateral connections from the encoder to the decoder. By this design, the domain-specific details, which are only necessary for reconstructing the unlabeled target data, are directly fed to the decoder to complete the reconstruction task, relieving the pressure of learning domain-specific variations at the later layers of the shared encoder. As a result, DLSN allows the shared encoder to focus on learning cross-domain shared content and ignores the domain-specific variations. Notably, the proposed DLSN can be used as a standard module to be integrated with various existing UDA frameworks to further boost performance. Without whistles and bells, extensive experimental results on four gold-standard domain adaptation datasets, for example: 1) Digits; 2) Office31; 3) Office-Home; and 4) VisDA-C, demonstrate that the proposed DLSN can consistently and significantly improve the performance of various popular UDA frameworks.
Wanxia Deng, Lingjun Zhao, Gangyao Kuang, Dewen Hu, Matti Pietikäinen, Li Liu 0002
IEEE Trans. Cybern.1
2022 Informative Feature Disentanglement for Unsupervised Domain Adaptation
abstract
Unsupervised Domain Adaptation (UDA) aims at learning a classifier for an unlabeled target domain by transferring knowledge from a labeled source domain with a related but different distribution. The strategy of aligning the two domains in latent feature space via metric discrepancy or adversarial learning has achieved considerable progress. However, these existing approaches mainly focus on adapting the entire image and ignore the bottleneck that occurs when forced adaptation of uninformative domain-specific variations undermines the effectiveness of learned features. To address this problem, we propose a novel component called Informative Feature Disentanglement (IFD), which is equipped with the adversarial network or the metric discrepancy model, respectively. Accordingly, the new network architectures, named IFDAN and IFDMN, enable informative feature refinement before the adaptation. The proposed IFD is designed to disentangle informative features from the uninformative domain-specific variations, which are produced by a Variational Autoencoder (VAE) with lateral connections from the encoder to the decoder. We cooperatively apply the IFD to conduct supervised disentanglement for the source domain and unsupervised disentanglement for the target domain. In this way, informative features are disentangled from the domain-specific details before the adaptation. Extensive experimental results on three gold-standard domain adaptation datasets, e.g., Office31, Office-Home and VisDA-C, demonstrate the effectiveness of the proposed IFDAN and IFDMN models for UDA.
Wanxia Deng, Lingjun Zhao, Qing Liao 0001, Deke Guo, Gangyao Kuang, Dewen Hu, Matti Pietikäinen, Li Liu 0002
IEEE Trans. Multim.1
2021 Transferable Discriminative Feature Mining For Unsupervised Domain Adaptation
abstract
Unsupervised Domain Adaptation (UDA) aims to seek an effective model for unlabeled target domain by leveraging knowledge from a labeled source domain with a related but different distribution. Many existing approaches ignore the underlying discriminative features of the target data and the discrepancy of conditional distributions. To address these two issues simultaneously, the paper presents a Transferable Discriminative Feature Mining (TDFM) approach for UDA, which can naturally unify the mining of domain-invariant discriminative features and the alignment of class-wise features into one single framework. To be specific, to achieve the domain-invariant discriminative features, TDFM jointly learns a shared encoding representation for two tasks: supervised classification of labeled source data, and discriminative clustering of unlabeled target data. It then conducts the class-wise alignment by decreasing intra-class variations and increasing inter-class differences across domains, encouraging the emergence of transferable discriminative features. When combined, these two procedures are mutually beneficial. Comprehensive experiments verify that TDFM can obtain remarkable margins over state-of-the-art domain adaptation methods.
Lingjun Zhao, Wanxia Deng, Gangyao Kuang, Dewen Hu, Li Liu 0002
ICIP2
2021 Informative Class-Conditioned Feature Alignment for Unsupervised Domain Adaptation
abstract
The goal of unsupervised domain adaptation is to learn a task classifier that performs well for the unlabeled target domain by borrowing rich knowledge from a well-labeled source domain. Although remarkable breakthroughs have been achieved in learning transferable representation across domains, two bottlenecks remain to be further explored. First, many existing approaches focus primarily on the adaptation of the entire image, ignoring the limitation that not all features are transferable and informative for the object classification task. Second, the features of the two domains are typically aligned without considering the class labels; this can lead the resulting representations to be domain-invariant but non-discriminative to the category. To overcome the two issues, we present a novel Informative Class-Conditioned Feature Alignment (IC2FA) approach for UDA, which utilizes a twofold method: informative feature disentanglement and class-conditioned feature alignment, designed to address the above two challenges, respectively. More specifically, to surmount the first drawback, we cooperatively disentangle the two domains to obtain informative transferable features; here, Variational Information Bottleneck (VIB) is employed to encourage the learning of task-related semantic representations and suppress task-unrelated information. With regard to the second bottleneck, we optimize a new metric, termed Conditional Sliced Wasserstein Distance (CSWD), which explicitly estimates the intra-class discrepancy and the inter-class margin. The intra-class and inter-class CSWDs are minimized and maximized, respectively, to yield the domain-invariant discriminative features. IC2FA equips class-conditioned feature alignment with informative feature disentanglement and causes the two procedures to work cooperatively, which facilitates informative discriminative features adaptation. Extensive experimental results on three domain adaptation datasets confirm the superiority of IC2FA.
Wanxia Deng, Yawen Cui, Zhen Liu 0004, Gangyao Kuang, Dewen Hu, Matti Pietikäinen, Li Liu 0002
ACM Multimedia1
2021 Deep ladder reconstruction-classification network for unsupervised domain adaptation
Wanxia Deng, Zhuo Su 0002, Qiang Qiu 0001, Lingjun Zhao, Gangyao Kuang, Matti Pietikäinen, Huaxin Xiao, Li Liu 0002
Pattern Recognit. Lett.1
2021 Joint Clustering and Discriminative Feature Alignment for Unsupervised Domain Adaptation
abstract
Unsupervised Domain Adaptation (UDA) aims to learn a classifier for the unlabeled target domain by leveraging knowledge from a labeled source domain with a different but related distribution. Many existing approaches typically learn a domain-invariant representation space by directly matching the marginal distributions of the two domains. However, they ignore exploring the underlying discriminative features of the target data and align the cross-domain discriminative features, which may lead to suboptimal performance. To tackle these two issues simultaneously, this paper presents a Joint Clustering and Discriminative Feature Alignment (JCDFA) approach for UDA, which is capable of naturally unifying the mining of discriminative features and the alignment of class-discriminative features into one single framework. Specifically, in order to mine the intrinsic discriminative information of the unlabeled target data, JCDFA jointly learns a shared encoding representation for two tasks: supervised classification of labeled source data, and discriminative clustering of unlabeled target data, where the classification of the source domain can guide the clustering learning of the target domain to locate the object category. We then conduct the cross-domain discriminative feature alignment by separately optimizing two new metrics: 1) an extended supervised contrastive learning, i.e., semi-supervised contrastive learning 2) an extended Maximum Mean Discrepancy (MMD), i.e., conditional MMD, explicitly minimizing the intra-class dispersion and maximizing the inter-class compactness. When these two procedures, i.e., discriminative features mining and alignment are integrated into one framework, they tend to benefit from each other to enhance the final performance from a cooperative learning perspective. Experiments are conducted on four real-world benchmarks (e.g., Office-31, ImageCLEF-DA, Office-Home and VisDA-C). All the results demonstrate that our JCDFA can obtain remarkable margins over state-of-the-art domain adaptation methods. Comprehensive ablation studies also verify the importance of each key component of our proposed algorithm and the effectiveness of combining two learning strategies into a framework.
Wanxia Deng, Qing Liao 0001, Lingjun Zhao, Deke Guo, Gangyao Kuang, Dewen Hu, Li Liu 0002
IEEE Trans. Image Process.1
2018 Point-pattern matching based on point pair local topology and probabilistic relaxation labeling
Wanxia Deng, Huanxin Zou, Lin Lei, Shilin Zhou 0001
Vis. Comput.1
2018 A robust non-rigid point set registration method based on inhomogeneous Gaussian mixture models
Wanxia Deng, Huanxin Zou, Lin Lei, Shilin Zhou 0001, Tiancheng Luo
Vis. Comput.1