VLDB 2026 Research / reviewers in the wild / expert
Jun Wan 0005
dblp:69/6563-5
· DBLP profile ↗
58ranked-venue papers
12as first author
54since 2021 · last 2026
0000-0002-9961-7902ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 8 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpaCRD: Multimodal Deep Fusion of Histology and Spatial Transcriptomics for Cancer Region DetectionabstractAccurate detection of cancer tissue regions (CTR) enables deeper analysis of the tumor microenvironment and offers crucial insights into treatment response. Traditional CTR detection methods, which typically rely on the rich cellular morphology in histology images, are susceptible to a high rate of false positives due to morphological similarities across different tissue regions. The groundbreaking advances in spatial transcriptomics (ST) provide detailed cellular phenotypes and spatial localization information, offering new opportunities for more accurate cancer region detection. However, current methods are unable to effectively integrate histology images with ST data, especially in the context of cross-sample and cross-platform/batch settings for accomplishing the CTR detection. To address this challenge, we propose SpaCRD, a transfer learning-based method that deeply integrates histology images and ST data to enable reliable CTR detection across diverse samples, platforms, and batches. Once trained on source data, SpaCRD can be readily generalized to accurately detect cancerous regions across samples from different platforms and batches. The core of SpaCRD is a category-regularized variational reconstruction-guided bidirectional cross-attention fusion network, which enables the model to adaptively capture latent co-expression patterns between histological features and gene expression from multiple perspectives. Extensive benchmark analysis on 23 matched histology-ST datasets spanning various disease types, platforms, and batches demonstrates that SpaCRD consistently outperforms existing eight state-of-the-art methods in CTR detection. Shuailin Xue, Jun Wan 0005, Wenwen Min |
AAAI | 2 |
| 2026 | Leveraging community context and frequency-adaptive aggregation for robust fraud detection
Zheng Zhang 0025, Jun Wan 0005, Jun Liu 0036, Mingyang Zhou 0001, Kezhong Lu, Claudio J. Tessone, Guoliang Chen 0005, Hao Liao |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Universal Facial Landmark Detection by Landmark-Clustering Relation-Reasoning Transformer
Jun Wan 0005, Yuanzhi Yao, Jiaxing Huang 0001, Xiaoying Ding, Lefei Zhang, Yongsheng Gao 0001, Dacheng Tao |
Int. J. Comput. Vis. | 1 |
| 2026 | Multi-modal category-aware gating for oriented object detection with single-category experts
Beihang Song, Hongquan Sun, Tong Liu 0039, Kai Zhu 0009, Jun Wan 0005 |
Neurocomputing | 5 |
| 2026 | FGTBT: Frequency-guided task-balancing transformer for unified facial landmark detection
Jun Wan 0005, Xinyu Xiong, Zhihui Lai 0001, Jie Zhou 0009, Wenwen Min |
Inf. Sci. | 1 |
| 2026 | JCLRec: Joint diffusion model and dual contrastive learning for sequential recommendation
Kai Zhu 0009, Jing Li 0055, Yue He 0005, Mingfeng Wang, Jiaheng Yu, Jun Wan 0005 |
Knowl. Based Syst. | 7 |
| 2026 | Interpretable facial landmark detection by multi-expert collaborative uncertainty-aware deep networks
Jun Wan 0005, Hui Xi, Yuanzhi Yao, Zhihui Lai 0001, Jie Zhou 0009 |
Neural Networks | 1 |
| 2026 | Multi-view masked graph representation learning with semantic alignment for spatial multi-omics clustering
Jinjie Zhao, Jun Wan 0005, Wenwen Min |
Pattern Recognit. | 2 |
| 2026 | Incomplete Multi-View Data Learning via Adaptive Embedding and Partial l2,1 Norm Constraints for Parkinson's Disease DiagnosisabstractParkinson's disease (PD) is a progressive neurodegenerative disorder characterized by mental abnormalities and motor dysfunction. Its early classification and prediction of clinical scores have been major concerns for researchers. Currently, multi-view data learning has become an essential research area due to the capacity of multiple views to provide complementary insights from various perspectives. However, the discontinuous distribution, data missing complexity, small sample size, and redundant features in multi-view datasets pose a substantial obstacle, and most existing multi-view learning methods are unable to handle these challenges effectively. In this study, we propose a novel incomplete multi-view data learning framework (IMVDL) via dynamic embedding and partiall2,1norm constraints for PD diagnosis. Specifically, multi-view dynamic embedding can adapt to any view missing scene, thereby linearly/nonlinearly mapping incomplete multi-view data to low-dimensional manifold spaces and generating complete multi-view data representations. The partiall2,1norm constraint can ignore larger feature weight values and performl2,1norm sparse on the remaining weights, thereby avoiding the sparse bias problem caused by larger weight values. An efficient iterative algorithm is derived to find the optimal solution of the IMVDL method. We conduct extensive experiments using multi-modal neuroimage data from the Parkinson's Progression Markers Initiative (PPMI) database. The results demonstrate that the IMVDL method is superior to other comparative methods. The source code for IMVDL is available at https://github.com/a610lab/IMVDL/. Zhongwei Huang, Chao Chen 0007, Jianxia Chen, Jun Wan 0005, Zhi Yang 0006, Ran Zhou 0002, Haitao Gan |
IEEE J. Biomed. Health Informatics | 5 |
| 2026 | Proto-Former: Unified Facial Landmark Detection by Prototype TransformerabstractRecent advances in deep learning have significantly improved facial landmark detection. However, existing facial landmark detection datasets often define different numbers of landmarks, and most mainstream methods can only be trained on a single dataset. This limits the model generalization to different datasets and hinders the development of a unified model. To address this issue, we propose Proto-Former, a unified, adaptive, end-to-end facial landmark detection framework that explicitly enhances dataset-specific facial structural representations (i.e., prototype). Proto-Former overcomes the limitations of single-dataset training by enabling joint training across multiple datasets within a unified architecture. Specifically, Proto-Former comprises two key components: an Adaptive Prototype-Aware Encoder (APAE) that performs adaptive feature extraction and learns prototype representations, and a Progressive Prototype-Aware Decoder (PPAD) that refines these prototypes to generate prompts that guide the model's attention to key facial regions. Furthermore, we introduce a novel Prototype-Aware (PA) loss, which achieves optimal path finding by constraining the selection weights of prototype experts. This loss function effectively resolves the problem of prototype expert addressing instability during multi-dataset training, alleviates gradient conflicts, and enables the extraction of more accurate facial structure features. Extensive experiments on widely used benchmark datasets demonstrate that our Proto-Former achieves superior performance compared to existing state-of-the-art methods. The code is publicly available at:https://github.com/Husk021118/Proto-Former. Shengkai Hu, Haozhe Qi, Jun Wan 0005, Jiaxing Huang 0001, Lefei Zhang, Dacheng Tao |
IEEE Trans. Multim. | 3 |
| 2026 | Efficient Oriented Object Detection via Wavelet-Based Energy Label Reassignment and Dual Prediction StrategyabstractArbitrary-oriented object detection remains a pivotal research focus due to its practical significance and inherent challenges. Existing methods often extend frameworks and sampling strategies designed for horizontal object detectors, which struggle to handle the arbitrary orientations, high aspect ratios, and diverse scales of oriented objects. To overcome these limitations, we propose a novel and efficient method for arbitrary-oriented object detection. This approach dynamically assigns prediction layers by object pixel area, then leverages wavelet transform-based energy weighting for bottom-up sample reassignment, optimizing feature representation for oriented targets. In addition, a robust framework integrates heatmap keypoint prediction on feature maps of a quarter-sized image, along with sparse predictions on other scales. By querying small-object regions within deep feature maps, a progressive top-down feature fusion strategy further enhances the perception of fine-grained details. Extensive evaluations on four benchmark datasets demonstrate the method's substantial improvements in detection performance, establishing its potential for broader applications in oriented object detection. Beihang Song, Jing Li 0055, Jia Wu 0001, Xuefei Li 0001, Jun Wan 0005 |
IEEE Trans. Multim. | 6 |
| 2025 | Adaptive feature selection with flexible mapping for diagnosis and prediction of Parkinson's disease
Zhongwei Huang, Jianqiang Li 0005, Jiatao Yang, Jun Wan 0005, Jianxia Chen, Zhi Yang 0006, Ming Shi 0001, Ran Zhou 0002, Haitao Gan |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | A survey of text classification based on pre-trained language model
Yujia Wu, Jun Wan 0005 |
Neurocomputing | 2 |
| 2025 | Bidirectional-Modulation Frequency-Heterogeneous Network for Remote Sensing Image DehazingabstractRecently, deep neural networks have been extensively explored in remote sensing image haze removal and achieved remarkable performance. However, existing methods fail to effectively fuse the features extracted from Convolutional Neural Networks (CNNs) and Transformer networks, leading to performance degradation. Moreover, most dehazing methods lack further exploration of the distinct properties of high- and low-frequency features, which are crucial for texture restoration and haze removal. To address these issues, we propose a Bidirectional-Modulation Frequency-Heterogeneous Network (BMFH-Net). Specifically, we propose a Differential-Expert Guided Bidirectional Modulation (DGBM) module that incorporates Differential experts and physical inversion models to exploit the complementarity of CNN-Transformer features and extract their latent haze-related physical characteristics, thereby enabling more effective bidirectional alignment. Furthermore, a Wavelet Frequency Heterogeneous Enhancement (WFHE) Module is designed to capture the most representative high-frequency features to refine image texture details, while enhancing the global perception of haze and reconstructing structural information during low-frequency processing. Experiments on challenging remote sensing image datasets demonstrate that our BMFH-Net outperforms several state-of-the-art haze removal methods. The code is released publicly at https://github.com/zqf2024/BMFH-Net. Qingfei Zhong, Bo Du 0001, Zhigang Tu 0001, Jun Wan 0005, Wenbin Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Hierarchical Heterogeneous Geometric Foreground Perception Network for Remote Sensing Object DetectionabstractRecently, deep learning-based remote sensing object detection (RSOD) has been widely explored and obtained remarkable performance. However, most existing multiscale feature extraction methods neglect exploring the interfering representation of different hierarchical features in the backbone, which is crucial for learning more discriminative features. Moreover, feature pyramid network (FPN) and its variants have difficulty in effectively perceiving the pose and salient information of remote sensing objects, leading to reduced detection accuracy. To address these issues, we propose a hierarchical heterogeneous geometric foreground perception network (HHGFP-Net) for RSOD. Specifically, a hierarchical heterogeneous receptive-field module (HHRM) is proposed to reward and penalize the feature information of the corresponding levels according to the differences between the shallow and deep feature layers in the backbone, improving discriminative feature ability. Furthermore, a geometric foreground perception FPN (GFP-FPN) is developed to refine geometric shapes and enhance foreground contents, providing more precise feature representations for objects, particularly small objects. Experimental results on four challenging RSOD datasets demonstrate that our HHGFP-Net achieves state-of-the-art performance. Codes are available at:https://github.com/YyLinkWorld/HHGFP-Net. Yang Liu 0420, Lefei Zhang, Jun Wan 0005 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Latent Feature Disentanglement Bidirectional Prompting Network for Unsupervised Cloud Removal
Zhixuan Huang, Bo Du 0001, Lyuyang Tong, Jun Wan 0005 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Spatial-Frequency Residual-Guided Dynamic Perceptual Network for Remote Sensing Image Haze RemovalabstractRecently, deep neural networks have been extensively explored in remote sensing image haze removal and achieved remarkable performance. However, most existing haze removal methods fail to effectively leverage the fusion of spatial and frequency information, which is crucial for learning more representative features. Moreover, the prevalent perceptual loss used in dehazing model training overlooks the diversity among perceptual channels, leading to performance degradation. To address these issues, we propose a spatial-frequency residual-guided dynamic perceptual network (SFRDP-Net) for remote sensing image haze removal. Specifically, we first propose a residual-guided spatial-frequency interaction (RSFI) module, which incorporates a bidirectional residual complementary mechanism (BRCM) and a frequency residual enhanced attention (FREA). Both BRCM and FREA exploit spatial-frequency complementarity to guide more effective fusion of spatial and frequency information, thus enhancing feature representation capability and improving haze removal performance. Furthermore, a dynamic channel weighting perceptual loss (DCWP-Loss) is developed to impose constraints with varying strengths on different perceptual channels, advancing the reconstruction of high-quality haze-free images. Experiments on challenging benchmark datasets demonstrate our SFRDP-Net outperforms several state-of-the-art haze removal methods. The code is released publicly athttps://github.com/789as-syl/SFRDP-Net. Zhaoru Yao, Bo Du 0001, Jun Wan 0005, Lyuyang Tong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Fine-Grained Image Captioning by Ranking Diffusion TransformerabstractThe CLIP visual feature-based image captioning models have developed rapidly and achieved remarkable results. However, existing models still struggle to produce descriptive and discriminative captions because they insufficiently exploit fine-grained visual cues and fail to model complex vision-language alignment. To address these limitations, we propose a Ranking Diffusion Transformer (RDT), which integrates a Ranking Visual Encoder (RVE) and a Ranking Loss (RL) for fine-grained image captioning. The RVE introduces a novel ranking attention mechanism that effectively mines diverse and discriminative visual information from CLIP features. Meanwhile, the RL leverages the ranking of generated caption quality as a global semantic supervisory signal, thereby enhancing the diffusion process and strengthening vision-language semantic alignment. We show that by collaborating RVE and RL via the novel RDT-and by gradually adding and removing noise in the diffusion process-more discriminative visual features are learned and precisely aligned with the language features. Experimental results on popular benchmark datasets demonstrate that our proposed RDT surpasses existing state-of-the-art image captioning models in the literature. The code is publicly available at: https://github.com/junwan2014/RDT. Jun Wan 0005, Min Gan, Lefei Zhang, Jie Zhou 0009, Jun Liu 0036, Bo Du 0001, C. L. Philip Chen |
IEEE Trans. Image Process. | 1 |
| 2025 | Facial Expression Recognition With Heatmap Neighbor Contrastive LearningabstractMany supervised learning-based facial expression recognition (FER) methods achieve good performance with the assistance of expression labels and a complex framework. However, there are inconsistent annotations in different expression datasets, making the above methods disadvantageous for new expression datasets or datasets with limited training data. The objective of this paper is to learn self-supervised facial expression features that enable the FER model not to rely on the annotation consistency of the different datasets. Most current self-supervised learning algorithms based on contrastive learning learn the representation by forcing different augmented views of the same image close in the embedding space, but they cannot cover all variances within a semantic class. We propose a heatmap neighbor contrastive learning (HNCL) method for FER. It treats the images corresponding to the heatmap nearest neighbors of expressions as other positives, providing more semantic variations than pre-defined augmented transformations. Therefore, our HNCL can learn better expression features covering more intra-class variances, improving the performance of the FER model based on self-supervised learning. After fine-tuning, HNCL with a simple framework achieves top-three performance on the in-the-lab datasets and even matches the performance of state-of-the-art supervised learning methods on the in-the-wild datasets. Tong Liu 0039, Jing Li 0055, Jia Wu 0001, Bo Du 0001, Yibing Zhan, Dapeng Tao, Jun Wan 0005 |
IEEE Trans. Multim. | 7 |
| 2024 | Adaptive Sparse Learning Based on Flexible Graph Embedding for Parkinson's Disease DiagnosisabstractParkinson’s disease (PD) is a common neurodegenerative disorder in the elderly population. The progressive symptoms of PD can have significant physical and economic implications for patients. Therefore, the development of a method to aid in the diagnosis and prediction of PD is crucial. However, medical neuroimaging data often have redundant features and high data dimensions, which can negatively impact algorithm accuracy. To solve this challenge, a supervised algorithm for feature selection is proposed for the early diagnosis and prediction of PD. Specifically, the proposed method incorporates adaptive learning during iterations, which allows adaptive updating of the similarity matrix and selection of informative features. Meanwhile, we introduce flexible mapping to address the limitation that linear mapping is too strict. To measure the effectiveness of the algorithm, we test it on the Parkinson’s Progression Markers Initiative (PPMI) public dataset. According to the outcomes of the experiment, the proposed method outperforms the competing feature selection methods and graph neural network approaches. Zhongwei Huang, Jianqiang Li 0005, Jiatao Yang, Ran Zhou 0002, Jun Wan 0005, Haitao Gan |
IJCNN | 5 |
| 2024 | Word and Character Semantic Fusion by Pretrained Language Models for Text ClassificationabstractThe utilization of fine-tuned pre-trained language models (PLMs) in text classification has become widespread and has achieved remarkable performance. However, a significant limitation of these models is that they tend to mark special symbols that are not registered in PLMs as [UNK], which ultimately reduces the text classification accuracy. Although some PLMs have attempted to overcome this limitation by learning subwords or using character input during model training, effectively fusing subword and character features remains a challenge. To address this challenge, we propose a hybrid network architecture called SemFusion that collaborates two PLMs to learn rich semantic information and improve classification accuracy. Specifically, we use two different Transformer Encoders to extract features for each subword and character in the text sequence. These features include the [CLS] feature, which represents the overall meaning of the text sequence, and the token feature, which represents each subword or character. We then utilize a multi-head self-attention mechanism to model the correlation between text sequences at different granularities for enhancing semantic expression, thereby achieving more accurate text classification. Our experimental results on four benchmark text datasets demonstrate that our proposed method outperforms the current state-of-the-art methods in the literature. Yujia Wu, Jun Wan 0005 |
IJCNN | 2 |
| 2024 | Multimodal contrastive learning for spatial gene expression prediction using histology imagesabstractIn recent years, the advent of spatial transcriptomics (ST) technology has unlocked unprecedented opportunities for delving into the complexities of gene expression patterns within intricate biological systems. Despite its transformative potential, the prohibitive cost of ST technology remains a significant barrier to its widespread adoption in large-scale studies. An alternative, more cost-effective strategy involves employing artificial intelligence to predict gene expression levels using readily accessible whole-slide images stained with Hematoxylin and Eosin (H&E). However, existing methods have yet to fully capitalize on multimodal information provided by H&E images and ST data with spatial location. In this paper, we propose mclSTExp, a multimodal contrastive learning with Transformer and Densenet-121 encoder for Spatial Transcriptomics Expression prediction. We conceptualize each spot as a "word", integrating its intrinsic features with spatial context through the self-attention mechanism of a Transformer encoder. This integration is further enriched by incorporating image features via contrastive learning, thereby enhancing the predictive capability of our model. We conducted an extensive evaluation of highly variable genes in two breast cancer datasets and a skin squamous cell carcinoma dataset, and the results demonstrate that mclSTExp exhibits superior performance in predicting spatial gene expression. Moreover, mclSTExp has shown promise in interpreting cancer-specific overexpressed genes, elucidating immune-related genes, and identifying specialized spatial domains annotated by pathologists. Our source code is available at https://github.com/shizhiceng/mclSTExp. Wenwen Min, Zhiceng Shi, Jun Wan 0005, Changmiao Wang |
Briefings Bioinform. | 4 |
| 2024 | Single-stage oriented object detection via Corona Heatmap and Multi-stage Angle Prediction
Beihang Song, Jing Li 0055, Jia Wu 0001, Shan Xue 0001, Jun Wan 0005 |
Knowl. Based Syst. | 6 |
| 2024 | Quality-aware face alignment using high-resolution spatial dependencies
Jinyan Ma, Xuefei Li 0001, Jing Li 0055, Jun Wan 0005, Tong Liu 0039, Guohao Li 0009 |
Multim. Tools Appl. | 4 |
| 2024 | Confusable facial expression recognition with geometry-aware conditional network
Tong Liu 0039, Jing Li 0055, Jia Wu 0001, Bo Du 0001, Jun Wan 0005 |
Pattern Recognit. | 5 |
| 2024 | Unsupervised multi-branch network with high-frequency enhancement for image dehazing
Zhiming Luo, Bo Du 0001, Laibin Chang, Jun Wan 0005 |
Pattern Recognit. | 6 |
| 2024 | Precise facial landmark detection by Dynamic Semantic Aggregation Transformer
Jun Wan 0005, Yujia Wu, Zhihui Lai 0001, Wenwen Min, Jun Liu 0036 |
Pattern Recognit. | 1 |
| 2024 | Robust Self-expression Learning with Adaptive Noise PerceptionabstractSelf-expression learning methods often obtain a coefficient matrix to measure the similarity between pairs of samples. However, directly using the raw data to represent each sample under the self-expression framework may not be ideal, as noise points are inevitably involved in the process of representing clean samples. To address this issue, this work proposes a novel self-expression model called robust Self-Expression learning with adaptive Noise Perception (SENP). SENP decomposes each sample into a clean part and a noisy part, and samples with large self-expression losses can be recognized as the noise points. A reliable coefficient matrix can then be learned by using only the clean points to reconstruct the clean part of each sample. By simultaneously detecting the noisy part of each sample and noise points, and adaptively mitigating their negative impacts, the representative ability of the generated coefficient matrix is improved. Moreover, inspired by the solution of non-negative matrix factorization (NMF), an effective algorithm is formed to optimize SENP. Extensive experiments on well-known benchmark datasets demonstrate the superiority of SENP compared to several state-of-the-art methods. Yangbo Wang, Jie Zhou 0009, Jianglin Lu, Jun Wan 0005, Can Gao, Qingshui Lin |
Pattern Recognit. | 4 |
| 2024 | Direction Prediction Redefinition: Transfer Angle to Scale in Oriented Object DetectionabstractOriented object detection has garnered significant attention. However, rotational symmetry and discontinuity at boundaries can confuse networks, leading to discontinuous loss and regression inconsistency. In this paper, we propose an efficient multi-directional object detection framework named Direction Prediction Redefinition (DPR). We describe the angle variation of rotated bounding boxes ($B_{r}$) as changes in the dimensions of horizontal bounding boxes ($B_{h}$). Specifically, we generate two sets of horizontal bounding boxes by predicting the center points of the corresponding boundaries within the rotated bounding box, thereby avoiding boundary issues caused by angle prediction. To further achieve robust rotated boundary representation, we propose the Joint Scale Representation method and the State Feature Encoding module, which are used to eliminate outliers in rotated boundaries and guide the correct selection of horizontal bounding box vertices, respectively. Moreover, we further abstract DPR as Multiple Trigonometric functions based DPR (DPR-MT). This method maps a single angle into four sets of trigonometric functions and considers them as the four sides of the horizontal bounding box. This approach predicts angles in the form of horizontal bounding boxes without complex operations, making it plug-and-play. Experimental results and visual analysis on challenging datasets further verify the effectiveness and competitiveness of our proposed method. Beihang Song, Jing Li 0055, Jia Wu 0001, Jun Wan 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Tracking With Saliency Region TransformerabstractTransformers show a great impact on visual tracking thanks to their powerful representation learning capabilities. As the capacity of the model grows, the speed of the tracker tends to decrease gradually. Our work focuses on dealing with massively redundant information in tracking sequences with the Saliency Region Tracker (SRTrack). SRTrack is a heuristic two-stage tracker consisting of a lightweight tracking stage and a saliency stage. The former can handle simple tracking sequences while the latter is designed to perform delicate tracking on challenging frames with more discriminative features. However, the two-stage design leads to feature extrapolation, creating inconsistencies between training and inference features. In order to mitigate this problem, we develop an attention scaling factor that guarantees model robustness while yielding a slight performance gain. Our SRTrack achieves a state-of-the-art 0.699 AUC running at 61 FPS on LaSOT. Several experiments on large benchmarks demonstrate the high efficiency and accuracy of SRTrack. Tianpeng Liu, Jing Li 0055, Jia Wu 0001, Lefei Zhang, Jun Wan 0005, Lezhi Lian |
IEEE Trans. Image Process. | 6 |
| 2024 | Typicality-Aware Adaptive Similarity Matrix for Unsupervised LearningabstractGraph-based clustering approaches, especially the family of spectral clustering, have been widely used in machine learning areas. The alternatives usually engage a similarity matrix that is constructed in advance or learned from a probabilistic perspective. However, unreasonable similarity matrix construction inevitably leads to performance degradation, and the sum-to-one probability constraints may make the approaches sensitive to noisy scenarios. To address these issues, the notion of typicality-aware adaptive similarity matrix learning is presented in this study. The typicality (possibility) rather than the probability of each sample being a neighbor of other samples is measured and adaptively learned. By introducing a robust balance term, the similarity between any pairs of samples is only related to the distance between them, yet it is not affected by other samples. Therefore, the impact caused by the noisy data or outliers can be alleviated, and meanwhile, the neighborhood structures can be well captured according to the joint distance between samples and their spectral embeddings. Moreover, the generated similarity matrix has block diagonal properties that are beneficial to correct clustering. Interestingly, the results optimized by the typicality-aware adaptive similarity matrix learning share the common essence with the Gaussian kernel function, and the latter can be directly derived from the former. Extensive experiments on synthetic and well-known benchmark datasets demonstrate the superiority of the proposed idea when comparing with some state-of-the-art methods. Jie Zhou 0009, Can Gao, Xizhao Wang, Zhihui Lai 0001, Jun Wan 0005, Xiaodong Yue 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Multimodal attention-based variational autoencoder for clinical risk predictionabstractPrediction of survival risk in cancer patients is crucial for understanding the underlying mechanisms of canceration in different stages. Previous studies mainly relied on single-modal omics data due to technological constraints. However, with the increasing availability of cancer omics data, researchers have focused on the use of multi-omics and multimodal data for survival analysis. The application of deep learning methods has become an option for the prediction of clinical risk. Recent advances in the attention mechanism and the variational autoencoder (VAE) have made them promising for analyzing cancer omics data. However, VAE has limitations in disregarding the importance of different features between modalities, and the introduction of an attention mechanism could address this limitation. In this study, we propose a Multimodal Attention-based VAE (MAVAE) deep learning framework using cross-modal multihead attention to integrate cancer multi-omics data for clinical risk prediction. We evaluated our approach on eight TCGA datasets. We find that (1) MAVAE outperforms traditional machine learning and recent deep learning methods; (2) Multi-modal data yields better classification performance than single-modal data; (3) The multi-head attention mechanism improves the decision-making process; (4) Clinical and genetic data are the most important modal data. Our implementation of MAVAE is available at https://github.com/wenwenmin/MAVAE. Taosheng Xu, Jun Wan 0005, Wenwen Min |
BIBM | 4 |
| 2023 | Cross-Domain Facial Expression Recognition via Disentangling Identity RepresentationabstractMost existing cross-domain facial expression recognition (FER) works require target domain data to assist the model in analyzing distribution shifts to overcome negative effects. However, it is often hard to obtain expression images of the target domain in practical applications. Moreover, existing methods suffer from the interference of identity information, thus limiting the discriminative ability of the expression features. We exploit the idea of domain generalization (DG) and propose a representation disentanglement model to address the above problems. Specifically, we learn three independent potential subspaces corresponding to the domain, expression, and identity information from facial images. Meanwhile, the extracted expression and identity features are recovered as Fourier phase information reconstructed images, thereby ensuring that the high-level semantics of images remain unchanged after disentangling the domain information. Our proposed method can disentangle expression features from expression-irrelevant ones (i.e., identity and domain features). Therefore, the learned expression features exhibit sufficient domain invariance and discriminative ability. We conduct experiments with different settings on multiple benchmark datasets, and the results show that our method achieves superior performance compared with state-of-the-art methods. Tong Liu 0039, Jing Li 0055, Jia Wu 0001, Lefei Zhang, Jun Wan 0005 |
IJCAI | 7 |
| 2023 | Spammer detection via ranking aggregation of group behavior
Zheng Zhang 0025, Mingyang Zhou 0001, Jun Wan 0005, Kezhong Lu, Guoliang Chen 0005, Hao Liao |
Expert Syst. Appl. | 3 |
| 2023 | Temporal burstiness and collaborative camouflage aware fraud detection
Zheng Zhang 0025, Jun Wan 0005, Mingyang Zhou 0001, Zhihui Lai 0001, Claudio J. Tessone, Guoliang Chen 0005, Hao Liao |
Inf. Process. Manag. | 2 |
| 2023 | Scale-free heterogeneous cycleGAN for defogging from a single image for autonomous driving in fog
Yan Zhang 0002, Zhiping Dan, Shuifa Sun, Jun Wan 0005, Weisheng Li 0001 |
Neural Comput. Appl. | 6 |
| 2023 | Multi-level Feature Interaction and Efficient Non-Local Information Enhanced Channel Attention for image dehazing
Bohui Li, Zhiping Dan, Bo Du 0001, Wen Yang 0001, Jun Wan 0005 |
Neural Networks | 7 |
| 2023 | Robust and Precise Facial Landmark Detection by Self-Calibrated Pose Attention NetworkabstractCurrent fully supervised facial landmark detection methods have progressed rapidly and achieved remarkable performance. However, they still suffer when coping with faces under large poses and heavy occlusions for inaccurate facial shape constraints and insufficient labeled training samples. In this article, we propose a semisupervised framework, that is, a self-calibrated pose attention network (SCPAN) to achieve more robust and precise facial landmark detection in challenging scenarios. To be specific, a boundary-aware landmark intensity (BALI) field is proposed to model more effective facial shape constraints by fusing boundary and landmark intensity field information. Moreover, a self-calibrated pose attention (SCPA) model is designed to provide a self-learned objective function that enforces intermediate supervision without label information by introducing a self-calibrated mechanism and a pose attention mask. We show that by integrating the BALI fields and SCPA model into a novel SCPAN, more facial prior knowledge can be learned and the detection accuracy and robustness of our method for faces with large poses and heavy occlusions have been improved. The experimental results obtained for challenging benchmark datasets demonstrate that our approach outperforms state-of-the-art methods in the literature. Jun Wan 0005, Hui Xi, Jie Zhou 0009, Zhihui Lai 0001, Witold Pedrycz, Xu Wang 0006 |
IEEE Trans. Cybern. | 1 |
| 2023 | SRDF: Single-Stage Rotate Object Detector via Dense Prediction and False Positive SuppressionabstractOriented object detection has made astonishing progress. However, existing methods neglect to address the issue of false positives caused by the background or nearby clutter objects. Meanwhile, class imbalance and boundary overflow issues caused by the predicting rotation angles may affect the accuracy of rotated bounding box predictions. To address the above issues, we propose a Single-stage Rotate object detector via Dense prediction and False positive suppression (SRDF). Specifically, we design an Instance-level False Positive Suppression Module (IFPSM), IFPSM acquires the weight information of target and non-target regions by supervised learning of spatial feature encoding, and applies these weight values to the deep feature map, thereby attenuating the response signals of non-target regions within the deep feature map. Compared to commonly used attention mechanisms, this approach more accurately suppresses false positive regions. Then, we introduce a hybrid classification and regression method to represent the object orientation, the proposed mothed divide the angle into two segments for prediction, reducing the number of categories and narrowing the range of regression. This alleviates the issue of class imbalance caused by treating one degree as a single category in classification prediction, as well as the problem of boundary overflow caused by directly regressing the angle. In addition, we transform the traditional post-processing steps based on matching and searching to a two-dimensional probability distribution mathematical model, which accurately and quickly extracts the bounding boxes from dense prediction results. Extensive experiments on Remote Sensing, Synthetic Aperture Radar, and Scene Text benchmarks demonstrate the superiority of the proposed SRDF method over state-of-the-art rotated object detection methods. Our codes are available at https://github.com/TomZandJerryZ/SRDF. Beihang Song, Jing Li 0055, Jia Wu 0001, Bo Du 0001, Jun Wan 0005, Tianpeng Liu |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Partial Siamese With Multiscale Bi-Codec Networks for Remote Sensing Image Haze RemovalabstractRecently, the U-Shaped networks has been widely explored in remote sensing image dehazing and obtained promising performance. However, most of the existing dehazing methods based on U-Shaped framework lack the reconstruction constraints of haze areas, which is particularly important to restore haze-free images. Moreover, their encoding and decoding layers cannot effectively fuse multi-scale features, resulting in deviations in the color and texture of the dehazing image. To address these issues, in this paper, we propose a Partial Siamese with Multiscale Bi-codec Dehazing Network (PSMB-Net) which is mainly composed of a Partial Siamese Framework (PSF) and a Multiscale Bi-codec Information Fusion (MBIF) module. Specifically, the PSF is proposed to create dehazing prior information to guide the network to build Siamese constraints and achieve improved dehazing results. Furthermore, we design a MBIF module which can enhance feature extraction, and the multi-scale information is used to improve the reconstruction ability of the network for the color and texture of the dehazing image. Experimental results on challenging benchmark datasets demonstrate the superiority of our PSMB-Net over state-of-the-art image dehazing methods. The source code is available at https://github.com/thislzm/PSMB-Net. Zhiming Luo, Bo Du 0001, Wen Yang 0001, Jun Wan 0005, Lefei Zhang |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2023 | Precise Facial Landmark Detection by Reference Heatmap TransformerabstractMost facial landmark detection methods predict landmarks by mapping the input facial appearance features to landmark heatmaps and have achieved promising results. However, when the face image is suffering from large poses, heavy occlusions and complicated illuminations, they cannot learn discriminative feature representations and effective facial shape constraints, nor can they accurately predict the value of each element in the landmark heatmap, limiting their detection accuracy. To address this problem, we propose a novel Reference Heatmap Transformer (RHT) by introducing reference heatmap information for more precise facial landmark detection. The proposed RHT consists of a Soft Transformation Module (STM) and a Hard Transformation Module (HTM), which can cooperate with each other to encourage the accurate transformation of the reference heatmap information and facial shape constraints. Then, a Multi-Scale Feature Fusion Module (MSFFM) is proposed to fuse the transformed heatmap features and the semantic features learned from the original face images to enhance feature representations for producing more accurate target heatmaps. To the best of our knowledge, this is the first study to explore how to enhance facial landmark detection by transforming the reference heatmap information. The experimental results from challenging benchmark datasets demonstrate that our proposed method outperforms the state-of-the-art methods in the literature. Jun Wan 0005, Jun Liu 0036, Jie Zhou 0009, Zhihui Lai 0001, LinLin Shen, Ping Xiong 0001, Wenwen Min |
IEEE Trans. Image Process. | 1 |
| 2023 | Low-Rank Linear Embedding for Robust ClusteringabstractThe performance of k-means clustering is often degenerate when dealing with high-dimensional and noisy scenarios. In this study, an end-to-end robust clustering method with low-rank linear embedding techniques (RCLR) is presented in conjunction with k-means. Sparse coefficients and a space projection matrix can be simultaneously learned. The global structures and local neighborhood properties are well captured in the learning procedures. Both the processes of clustering and dimensionality reduction are realized at the same time. The notions of clustering, dimensionality reduction, low-rank representation, and local property preservation are seamlessly integrated into a unified model. The limitation of error accumulation encountered in the previous two-stage clustering framework involving low-rank representation can be alleviated. This is the first attempt to introduce both the global and local geometrical structures into k-means directly, as well L2,1-norm is used as a basic metric instead of the conventional F-norm to further improve the robustness and interpretation of the model. The superiority of the proposed RCLR method is demonstrated by extensive experiments completed on various well-known benchmark datasets. Jie Zhou 0009, Witold Pedrycz, Jun Wan 0005, Can Gao, Zhihui Lai 0001, Xiaodong Yue 0002 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Information diffusion-aware likelihood maximization optimization for community detection
Zheng Zhang 0025, Jun Wan 0005, Mingyang Zhou 0001, Kezhong Lu, Guoliang Chen 0005, Hao Liao |
Inf. Sci. | 2 |
| 2022 | GuidedStyle: Attribute knowledge guided style manipulation for semantic face editing
Xianxu Hou, Hanbang Liang, LinLin Shen, Zhihui Lai 0001, Jun Wan 0005 |
Neural Networks | 6 |
| 2022 | Robust face alignment by dual-attentional spatial-aware capsule networks
Jinyan Ma, Jing Li 0055, Bo Du 0001, Jia Wu 0001, Jun Wan 0005, Yafu Xiao |
Pattern Recognit. | 5 |
| 2022 | Robust Jointly Sparse Fuzzy Clustering With Neighborhood Structure PreservationabstractFuzzy clustering techniques, especially fuzzy C-means (FCM) and its weighted variants, are typical partitive clustering models that are widely used for revealing possible hidden structures in data. Although they can quantitatively depict the overlapping areas with a partition matrix, their performances deteriorate when dealing with high-dimensional data because the distance computations may be negatively impacted by the irrelevant features, and then the concentration effect may arise. Moreover, they are sensitive to noisy environments. To tackle these obstacles, a robust jointly sparse fuzzy clustering method (RJSFC) is proposed in this study. The representative prototypes, sparse membership grades, and an orthogonal projection matrix are simultaneously learnt when optimizing RJSFC. The obtained low-dimensional embeddings can preserve the local neighborhood structure, and the clustering is conducted in the transformed lower dimensional space rather than the original space, which improves the capability of fuzzy clustering for dealing with high-dimensional scenarios. Furthermore,${L_{2,1}}$-norm is exploited as the basic metric for both loss and regularization parts in RJSFC, the robustness of the model and the interpretability of the extracted features are enhanced. The notions of fuzzy clustering, neighborhood structure preservation, and feature extraction are seamlessly integrated into a unified model. The limitation of the previous two-stage clustering framework when dealing with high-dimensional data entailing dimensionality reduction and clustering procedures separately can be effectively addressed. Extensive experimental results on various well-known datasets demonstrate the usefulness of RJSFC when comparing with some state-of-the-art methods. Jie Zhou 0009, Witold Pedrycz, Can Gao, Zhihui Lai 0001, Jun Wan 0005, Zhong Ming 0001 |
IEEE Trans. Fuzzy Syst. | 5 |
| 2022 | Robust Facial Landmark Detection by Multiorder Multiconstraint Deep NetworksabstractRecently, heatmap regression has been widely explored in facial landmark detection and obtained remarkable performance. However, most of the existing heatmap regression-based facial landmark detection methods neglect to explore the high-order feature correlations, which is very important to learn more representative features and enhance shape constraints. Moreover, no explicit global shape constraints have been added to the final predicted landmarks, which leads to a reduction in accuracy. To address these issues, in this article, we propose a multiorder multiconstraint deep network (MMDN) for more powerful feature correlations and shape constraints' learning. Especially, an implicit multiorder correlating geometry-aware (IMCG) model is proposed to introduce the multiorder spatial correlations and multiorder channel correlations for more discriminative representations. Furthermore, an explicit probability-based boundary-adaptive regression (EPBR) method is developed to enhance the global shape constraints and further search the semantically consistent landmarks in the predicted boundary for robust facial landmark detection. It is interesting to show that the proposed MMDN can generate more accurate boundary-adaptive landmark heatmaps and effectively enhance shape constraints to the predicted landmarks for faces with large pose variations and heavy occlusions. Experimental results on challenging benchmark data sets demonstrate the superiority of our MMDN over state-of-the-art facial landmark detection methods. Jun Wan 0005, Zhihui Lai 0001, Jing Li 0055, Jie Zhou 0009, Can Gao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Think About Boundary: Fusing Multi-level Boundary Information for Landmark Heatmap RegressionabstractAlthough current face alignment algorithms have obtained pretty good performances at predicting the location of facial landmarks, huge challenges remain for faces with severe occlusion and large pose variations, etc. On the contrary, semantic location of facial boundary is more likely to be reserved and estimated on these scenes. Therefore, we study a two-stage but end-to-end approach for exploring the relationship between the facial boundary and landmarks to get boundary-aware landmark predictions, which consists of two modules: the self-calibrated boundary estimation (SCBE) module and the boundary-aware landmark transform (BALT) module. In the SCBE module, we modify the stem layers and employ intermediate supervision to help generate high-quality facial boundary heatmaps. Boundary-aware features inherited from the SCBE module are integrated into the BALT module in a multi-scale fusion framework to better model the transformation from boundary to landmark heatmap. Experimental results conducted on the challenging benchmark datasets demonstrate that our approach outperforms state-of-the-art methods in the literature. Code will be available soon. Jinheng Xie, Jun Wan 0005, LinLin Shen, Zhihui Lai 0001 |
IJCNN | 2 |
| 2021 | Siamese Guided Anchoring Network for Visual TrackingabstractRecently, the Siamese Region Proposal Network (SiamRPN) has been widely explored in tracking and achieved remarkable performance. However, the existing SiamRPN-based method uses a predefined and highly dependent on prior knowledge anchor, which limits the tracking accuracy. Besides, when the target changes drastically, the anchor box obtained by the SiamRPN-based method also has some negative samples, which leads to a decrease inaccuracy. To address these issues, this paper proposes a siamese guided anchoring network for visual tracking, which can obtain more representative anchors by estimating the position and shape of the target, reducing the adverse effects of negative samples. At the same time, a feature adaption module is proposed to adapt to the target scale change for learning more discriminative and useful features and achieving more accurate visual tracking. Extensive experiments on challenging OTB100 and VOT2018 datasets demonstrate the competitive performance of the proposed algorithm in comparison with the state-of-the-art trackers. Jing Li 0055, Yafu Xiao, Jun Wan 0005 |
IJCNN | 5 |
| 2021 | Multi-task adversarial autoencoder network for face alignment in the wild
Xiaoqian Yue, Jing Li 0055, Jia Wu 0001, Jun Wan 0005, Jinyan Ma |
Neurocomputing | 5 |
| 2021 | Granular-conditional-entropy-based attribute reduction for partially labeled data with proxy labels
Can Gao, Jie Zhou 0009, Duoqian Miao 0001, Xiaodong Yue 0002, Jun Wan 0005 |
Inf. Sci. | 5 |
| 2021 | Robust facial landmark detection by cross-order cross-semantic deep network
Jun Wan 0005, Zhihui Lai 0001, LinLin Shen, Jie Zhou 0009, Can Gao, Xianxu Hou |
Neural Networks | 1 |
| 2021 | Projected fuzzy C-means clustering with locality preservation
Jie Zhou 0009, Witold Pedrycz, Xiaodong Yue 0002, Can Gao, Zhihui Lai 0001, Jun Wan 0005 |
Pattern Recognit. | 6 |
| 2021 | Robust Face Alignment by Multi-Order High-Precision Hourglass NetworkabstractHeatmap regression (HR) has become one of the mainstream approaches for face alignment and has obtained promising results under constrained environments. However, when a face image suffers from large pose variations, heavy occlusions and complicated illuminations, the performances of HR methods degrade greatly due to the low resolutions of the generated landmark heatmaps and the exclusion of important high-order information that can be used to learn more discriminative features. To address the alignment problem for faces with extremely large poses and heavy occlusions, this paper proposes a heatmap subpixel regression (HSR) method and a multi-order cross geometry-aware (MCG) model, which are seamlessly integrated into a novel multi-order high-precision hourglass network (MHHN). The HSR method is proposed to achieve high-precision landmark detection by a well-designed subpixel detection loss (SDL) and subpixel detection technology (SDT). At the same time, the MCG model is able to use the proposed multi-order cross information to learn more discriminative representations for enhancing facial geometric constraints and context information. To the best of our knowledge, this is the first study to explore heatmap subpixel regression for robust and high-precision face alignment. The experimental results from challenging benchmark datasets demonstrate that our approach outperforms state-of-the-art methods in the literature. Jun Wan 0005, Zhihui Lai 0001, Jun Liu 0036, Jie Zhou 0009, Can Gao |
IEEE Trans. Image Process. | 1 |
| 2020 | Nuclear-Norm-Based Jointly Sparse Regression for Two-Dimensional Image Regression
Haosheng Su, Zhihui Lai 0001, Jun Wan 0005 |
PRCV (3) | 4 |
| 2020 | Learning spatial-temporally regularized complementary kernelized correlation filters for visual tracking
Zhenyang Su, Jing Li 0055, Chengfang Song, Yafu Xiao, Jun Wan 0005 |
Multim. Tools Appl. | 6 |
| 2020 | Robust face alignment by cascaded regression and de-occlusion
Jun Wan 0005, Jing Li 0055, Zhihui Lai 0001, Bo Du 0001, Lefei Zhang |
Neural Networks | 1 |
| 2019 | Face alignment by Component Adaptive Mechanism
Jun Wan 0005, Jing Li 0055, Yujia Wu, Yafu Xiao, Xuefei Li 0001 |
Neurocomputing | 1 |