Huan Wan

dblp:153/6666 · DBLP profile ↗
← Back
22ranked-venue papers
11as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Enhanced local homogenization and reconstruction network for few-shot fine-grained image classification
Meiyin Hu, Huan Wan, Hui Wang 0001, Xin Wei 0002
Comput. Vis. Image Underst.2
2026 DMV - CLIP : Disentangled Multimodal Visual Adaptation for Text-Driven Face Editing
abstract
ABSTRACT Text‐driven face editing has attracted widespread interest due to its intuitive control and user‐friendly interaction. However, current state‐of‐the‐art (SOTA) methods face two main challenges: (1) they utilize unfinetuned general image‐text encoders for modality fusion, making it difficult to comprehend domain‐specific knowledge in facial attribute editing (dozens of fine‐grained facial attributes such as moustache and lipsticks); (2) they roughly optimize all attributes simultaneously using a cross‐entropy loss, leading to severe mutual interference among attributes. To this end, we propose Disentangled Multimodal Visual Adaptation for CLIP (DMV‐CLIP). First, DMV‐CLIP incorporates learnable context tokens to inject facial domain knowledge into the CLIP model via multimodal prompt learning (MPL). Second, it employs directional contrastive learning (DCL) to disentangle facial attributes and enable precise editing. Finally, DMV‐CLIP utilizes a vision‐language consistency model (VLCM) to maintain identity consistency while ensuring that the generated images strictly adhere to the semantic instructions.
Xin Wei 0002, Huan Wan, Haoruo Zhang, Xuhui Huang
Expert Syst. J. Knowl. Eng.4
2026 Distinct Polyp Generator Network for polyp segmentation
Huan Wan, Jing Ai, Xin Wei 0002, Jinshan Zeng, Jianyi Wan
Image Vis. Comput.1
2026 Hierarchical mask-enhanced dual reconstruction network for few-shot fine-grained image classification
Meiyin Hu, Huan Wan, Zhuohang Jiang, Xin Wei 0002
J. Vis. Commun. Image Represent.3
2026 DeRe-Net: details restoration networks for polyp segmentation
Huan Wan, Qinqin Wang, Jinshan Zeng, Xin Wei 0002
Multim. Syst.1
2026 The gift of clinical knowledge: annotation-free liver tumor segmentation via knowledge-driven synthesis
Keyi Zhong, Feng Ouyang, Peter Xiaoping Liu, Xuhui Huang, Huan Wan, Xin Wei 0002
Multim. Syst.6
2026 VTMedSeg: Geometry-Guided Vision-Text Model With Concept-Aware Fusion for Medical Ultrasound Image Segmentation
abstract
Medical ultrasound image segmentation is a vital diagnostic technique for identifying abnormalities, including tumors, cysts, and inflammation. To obtain accurate segmentation efficiently, many methods have been proposed. Amongst, vision–text models have shown promising progress; however, their performance is limited by the lack of high-quality paired text descriptions in ultrasound datasets. To address this issue, we propose VTMedSeg. In VTMedSeg, a geometry-guided text generator is developed to automatically synthesize medical descriptions by using connected component analysis and multi-dimensional geometric analysis. Then, the text features of these descriptions and the visual features of the corresponding ultrasound images are effectively integrated into a concept-aware cross-modal fusion module through multi-scale cross-modal alignment and medical knowledge graph modeling. Extensive experiments on four public medical ultrasound datasets demonstrate that VTMedSeg outperforms state-of-the-art vision–text models across multiple metrics.
Huan Wan, Wujian Xu, Yiwen Zou, Jinshan Zeng, Xin Wei 0002
IEEE Signal Process. Lett.1
2025 TriDE-Net: Triple-Densely Extraction Network for Precise Skin Lesion Segmentation
abstract
Accurate skin lesion segmentation is crucial for the quantitative analysis of skin cancer. Despite the significant advancements achieved by the deep-learning methods, the segmentation of skin lesions with irregular shapes and significant size variations is still challenging. To address the problem, we propose a Triple-Densely Extraction Network (TriDE-Net) for skin lesion segmentation, aiming to heavily extract multi-scale features in the inter- and intra-feature layers. In the TriDE-Net, a Feature-Intensive Capture Module (FICM) is designed to essentially extract multi-scale features from the intra-feature layers in a dually dense manner, and FICM is densely deployed in each skip-connection path to exploit features from the inter-feature layers. Moreover, we developed a Feature Adaptive Fusion Module (FAFM) to aggregate the decoding features to obtain accurate segmentation results. Comprehensive experiments on four widely-used skin lesion datasets consistently demonstrate that our TriDE-Net outperforms the state-of-the-art methods, with the Dice coefficient improving to 93.27%.
Huan Wan, Taona Deng, Wujian Xu, Xin Wei 0002, Jinshan Zeng
ICASSP1
2025 Adaptive blank compensation for few-shot image classification
Baozhe Wang, Huan Wan, Pengxiang Su, Xin Wei 0002
Neurocomputing3
2025 Dynamic dual consistency distillation for imbalanced facial attribute editing
Xuhui Huang, Huan Wan, Xin Wei 0002
Knowl. Based Syst.3
2023 Feature Distribution Fitting with Direction-Driven Weighting for Few-Shot Images Classification
abstract
Few-shot learning has received increasing attention and witnessed significant advances in recent years. However, most of the few-shot learning methods focus on the optimization of training process, and the learning of metric and sample generating networks. They ignore the importance of learning the ground-truth feature distributions of few-shot classes. This paper proposes a direction-driven weighting method to make the feature distributions of few-shot classes precisely fit the ground-truth distributions. The learned feature distributions can generate an unlimited number of training samples for the few-shot classes to avoid overfitting. Specifically, the proposed method consists of two optimization strategies. The direction-driven strategy is for capturing more complete direction information that can describe the feature distributions. The similarity-weighting strategy is proposed to estimate the impact of different classes in the fitting procedure and assign corresponding weights. Our method outperforms the current state-of-the-art performance by an average of 3% for 1-shot on standard few-shot learning benchmarks like miniImageNet, CIFAR-FS, and CUB. The excellent performance and compelling visualization show that our method can more accurately estimate the ground-truth distributions.
Xin Wei 0002, Huan Wan, Weidong Min
AAAI3
2023 Cluster-based data relabelling for classification
Huan Wan, Hui Wang 0001, Bryan W. Scotney, Jun Liu 0001, Xin Wei 0002
Inf. Sci.1
2023 Global subclass discriminant analysis
abstract
Linear discriminant analysis (LDA) is a powerful supervised dimensionality reduction method for analysing high-dimensional data. However, LDA cannot use locality information in data, which makes LDA degrade dramatically in performance on multimodal data. A number of LDA variants have been proposed to exploit locality information in data, including subclass-based LDAs. We discover a problem with these variants, which is that subclasses are selected on a within-class basis without considering other classes. This causes the loss of important information at class boundaries. In this paper, we present a novel variant of subclass-based LDA, Global Subclass Discriminant Analysis (GSDA). Unlike other subclass-based LDAs, GSDA selects subclasses from global clusters that may cross class boundaries, thus utilising within-class information and between-class information. More specifically, GSDA applies an effective clustering algorithm to the whole data to construct global clusters. It then utilises the local structure refining strategy on these global clusters to construct subclasses. Finally, GSDA learns a representative data subspace by maximising inter-subclass distance and minimising intra-subclass distance simultaneously. GSDA is extensively evaluated on a wide range of public datasets through comparison with the state-of-the-art LDA algorithms. Experimental results demonstrate its superiority in terms of accuracy and run times.
Huan Wan, Hui Wang 0001, Bryan W. Scotney, Jun Liu 0001, Xin Wei 0002
Knowl. Based Syst.1
2022 Transmit waveform and receive filter design for multiple-input multiple-output radar with one-bit digital-to-analogue converters
abstract
Abstract In this paper, we investigate the joint design of transmit waveform and receive filter for colocated Multiple‐Input Multiple‐Output radar equipped with one‐bit digital‐to‐analogue converters (DACs). The problem is formulated as maximising the output signal‐to‐interference‐plus‐noise ratio in the presence of signal‐dependent interferences, subject to a discrete constraint imposed by the waveform quantised with one‐bit DACs. To cope with the challenging non‐convex problem, an alternating maximisation framework is developed to optimise the transmit waveform and receive filter vectors in an iterative manner. More specifically, for a given transmit waveform vector, the analytical expression of the receive filter vector is first derived. Then, by fixing the receive filter, a biconvex relaxation method is employed to tackle the non‐convex problem with respect to the transmit waveform vector. The performance of the proposed approach is demonstrated by numerical simulations in different circumstances.
Huan Wan, Zhi Quan, Bin Liao 0001
IET Signal Process.1
2020 Within-class multimodal classification
abstract
Abstract In many real-world classification problems there exist multiple subclasses (or clusters) within a class; in other words, the underlying data distribution is within-class multimodal. One example is face recognition where a face (i.e. a class) may be presented in frontal view or side view, corresponding to different modalities. This issue has been largely ignored in the literature or at least under studied. How to address the within-class multimodality issue is still an unsolved problem. In this paper, we present an extensive study of within-class multimodality classification. This study is guided by a number of research questions, and conducted through experimentation on artificial data and real data. In addition, we establish a case for within-class multimodal classification that is characterised by the concurrent maximisation of between-class separation, between-subclass separation and within-class compactness. Extensive experimental results show that within-class multimodal classification consistently leads to significant performance gains when within-class multimodality is present in data. Furthermore, it has been found that within-class multimodal classification offers a competitive solution to face recognition under different lighting and face pose conditions. It is our opinion that the case for within-class multimodal classification is established, therefore there is a milestone to be achieved in some machine learning algorithms (e.g. Gaussian mixture model) when within-class multimodal classification, or part of it, is pursued.
Huan Wan, Hui Wang 0001, Bryan W. Scotney, Jun Liu 0001, Wing W. Y. Ng
Multim. Tools Appl.1
2020 Minimum margin loss for deep face recognition
Xin Wei 0002, Hui Wang 0001, Bryan W. Scotney, Huan Wan
Pattern Recognit.4
2020 Fourth-order direction finding in antenna arrays with partial channel gain/phase calibration
Huan Wan, Bin Liao 0001
Signal Process.1
2019 Gicoface: Global Information-Based Cosine Optimal Loss for Deep Face Recognition
abstract
Loss function plays an important role in CNNs. However, the recent loss functions either do not apply weight and feature normalisation or do not explicitly follow the two targets of improving discriminative ability: minimising intra-class variance and maximising inter-class variance. Besides, all of them consider only the feedback information from the current mini-batch instead of the distribution information from the whole training set. In this paper, we propose a novel loss function - Global Information-based Cosine Optimal loss (Gico loss). Gico loss is applied with weight and feature normalisation, designed explicitly following the aforementioned two targets of improving discriminative ability, and is guided by the distribution information from the whole training set. Extensive experiments are conducted on multiple public datasets, which confirms the effectiveness of the proposed Gico loss and shows that we achieve state-of-the-art performance.
Xin Wei 0002, Hui Wang 0001, Bryan W. Scotney, Huan Wan
ICIP4
2019 Precise Adjacent Margin Loss for Deep Face Recognition
abstract
Softmax loss is arguably one of the most widely used loss functions in CNNs. In recent years some Softmax variants have been proposed to enhance the discriminative ability of the learned features by adding additional margin constraints, which significantly improved the state-of-the-art performance of face recognition. However, the `margin' referenced in these losses does not represent the real margin between the different classes in the training set. Furthermore, they impose a margin on all possible combinations of class pairs, which is unnecessary. In this paper we propose the Precise Adjacent Margin loss (PAM loss), which gives an accurate definition of `margin' and has precise operations appropriate for different cases. PAM loss has better geometrical interpretation than the existing margin-based losses. Extensive experiments are conducted on LFW, YTF, MegaFace and FaceScrub datasets, and results show that the proposed method has state-of-the-art performance.
Xin Wei 0002, Hui Wang 0001, Bryan W. Scotney, Huan Wan
ICIP4
2019 A Novel Gaussian Mixture Model for Classification
abstract
Gaussian Mixture Model (GMM) is a probabilistic model for representing normally distributed subpopulations within an overall population. It is usually used for unsupervised learning to learn the subpopulations and the subpopulation assignment automatically. It is also used for supervised learning or classification to learn the boundary of subpopulations. However, the performance of GMM as a classifier is not impressive compared with other conventional classifiers such as k-nearest neighbors (KNN), support vector machine (SVM), decision tree and naive Bayes. In this paper, we attempt to address this problem. We propose a GMM classifier, SC-GMM, based on the separability criterion in order to separate the Gaussian models as much as possible. This classifier finds the optimal number of Gaussian components for each class based on the separability criterion and then determines the parameters of these Gaussian components by using the expectation maximization algorithm. Extensive experiments have been carried out on classification tasks from general data mining to face verification. Results show that SC-GMM significantly outperforms the original GMM classifier. Results also show that SC-GMM is comparable in classification accuracy to three variants of GMM classifier: Akaike Information Criterion based GMM (AIC-GMM), Bayesian Information Criterion based GMM (BIC-GMM) and variational Bayesian gaussian mixture (VBGM). However, SC-GMM is significantly more efficient than both AIC-GMM and BIC-GMM. Furthermore, compared with KNN, SVM, decision tree and naive Bayes, SC-GMM achieves competitive classification performance.
Huan Wan, Hui Wang 0001, Bryan W. Scotney, Jun Liu 0001
SMC1
2018 Separability-Oriented Subclass Discriminant Analysis
abstract
Linear discriminant analysis (LDA) is a classical method for discriminative dimensionality reduction. The original LDA may degrade in its performance for non-Gaussian data, and may be unable to extract sufficient features to satisfactorily explain the data when the number of classes is small. Two prominent extensions to address these problems are subclass discriminant analysis (SDA) and mixture subclass discriminant analysis (MSDA). They divide every class into subclasses and re-define the within-class and between-class scatter matrices on the basis of subclass. In this paper we study the issue of how to obtain subclasses more effectively in order to achieve higher class separation. We observe that there is significant overlap between models of the subclasses, which we hypothesise is undesirable. In order to reduce their overlap we propose an extension of LDA, separability oriented subclass discriminant analysis (SSDA), which employs hierarchical clustering to divide a class into subclasses using a separability oriented criterion, before applying LDA optimisation using re-defined scatter matrices. Extensive experiments have shown that SSDA has better performance than LDA, SDA and MSDA in most cases. Additional experiments have further shown that SSDA can project data into LDA space that has higher class separation than LDA, SDA and MSDA in most cases.
Huan Wan, Hui Wang 0001, Gongde Guo, Xin Wei 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2017 Self-adaptive Feature Fusion Method for Improving LBP for Face Identification
Xin Wei 0002, Hui Wang 0001, Huan Wan, Bryan W. Scotney
ICVS3