Xiaole Zhao

dblp:02/10631 · DBLP profile ↗
← Back
23ranked-venue papers
9as first author
16since 2021 · last 2025
0000-0003-0100-2414ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Long-Range Multi-Scale Fusion for Efficient Single Image Super-Resolution
abstract
Improving the performance of single image super-resolution (SISR) via extending the effective receptive field (ERF) of the model has become an admired paradigm in the field due to the universal self-similarity prior of natural images. However, it cannot fully explore model capability by solely increasing the ERF to capture long-range dependencies as the non-local self-similarity is typically multi-scale and cross-scale. To this end, a Long-range Multi-scale Fusion Network (LMFN) is devised in this work to simultaneously excavate both long-range and multi-scale priors in images, and the interaction between the both. Within the same scale, our model employs large kernel attention (LKA) and multi-scale modulation (MSM) to learn long-range and multi-scale features. To exploit the interaction between long-range and multi-scale dependencies within one single scale and across scales, we design an Interactive Fusion Modulation (IFM) module for the effective fusion of the non-local and multi-scale features. Extensive experiments on the benchmark datasets illustrate the significant superiority of the proposed LMFN over the advanced SISR models.
Xiaole Zhao, Yan Yang 0001, Tianrui Li 0001
ICASSP2
2025 Spatial Dual Feature Aggregation Network for Lightweight Image Super-Resolution
abstract
Recently, vision transformer-based methods have achieved breakthrough progress in image super-resolution (SR) tasks, owing to their powerful capability in capturing global features. However, compared to convolution operations, self-attention mechanisms are weaker in capturing local structures and high-frequency components of images, and their high computational complexity makes them unsuitable for mobile devices with limited computing power. To address these issues, we incorporate the structural characteristics of vision transformers and propose a spatial dual feature aggregation network (SDFAN) for lightweight image SR. Based on naive convolutions and variance attention (VA), we present an efficient scheme for local pattern aggregation (LPA) to enhance the perception of local details. Meanwhile, we propose a multi-scale dilated attention (MSDA) to flexibly model non-local and multi-scale features while maintaining linear computational complexity. The integration of LPA and MSDA enables the model to simultaneously and effectively capture local and non-local dependencies. Furthermore, we design a spatial channel mixer (SCM) to fully exploit the synergistic interaction between spatial and channel information. Extensive experiments demonstrate that the proposed SDFAN exhibits superior performance compared to existing state-of-the-art SR models while maintaining lower parameter counts and computational costs. Our source code and models are available at https://github.com/stella-von/SDFAN.
Chengxing Xie, Linze Li 0001, Xiaole Zhao
IJCNN3
2025 Hierarchical Feature Learning Based on Memory for Skin Lesion Recognition
abstract
In the realm of dermatology, Convolutional Neural Networks (CNNs) have exhibited immense potential, emerging as vital tools to aid physicians for precise diagnosis. However, current work generally faces two main challenges: (1) the inadequate samples in dermatological image datasets hinder sufficiently training high-performance deep models; (2) the pronounced issue of class imbalance significantly undermines the models’ ability to recognize minority class conditions. In response to these problems, this work introduces a novel network architecture and loss called Hierarchical Feature Learning (HFL), aimed at enhancing the model’s feature representation capabilities and adaptability to imbalanced data. Specifically, we devise a CNN architecture based on hierarchical feature storage, which effectively augments the network’s proficiency in extracting and preserving rich, hierarchical features through the incorporation of a multi-level feature memory mechanism. To further mitigate the impact of data imbalance on model training, we propose an Adaptive Re-weighted Loss (ARL). This loss dynamically adjusts the loss weights according to the real-time distribution of samples across classes, ensuring that minority class samples receive adequate attention during training. Experimental results on ISIC-2017, ISIC-2018 and 7-PT datasets demonstrate the effectiveness of our proposed methodology.
Wenjia Yang, Xiaole Zhao, Yan Yang 0001, Tianrui Li 0001
IJCNN2
2025 MGPDF: A Multi-modal Gaussian Process Decision-Level Fusion Model for Parkinson's Disease Prediction
Keyu Shen, Yan Yang 0001, Xiaole Zhao
PAKDD (2)5
2025 segWCD: A new segmentation-based weak supervision neural network for building change detection
Yunyang Wu, Xiaole Zhao, Yimin Sun, Tianrui Li 0001
Appl. Intell.3
2025 Assisted diagnosis of neuropsychiatric disorders based on functional connectivity: A survey on application and performance evaluation of graph neural network
abstract
The functional connectivity network which provides a perspective of interactions among all brain regions and reveals the functional organizational structure of the brain has gained considerable traction in the domain of diagnosis of neuropsychiatric disorders. To further extract lesion attributes within brain networks , Graph Neural Networks (GNNs) can achieve interpretable feature extraction and compression by simulating the structure and signal transmission in real brains. This has led to the prevalence of GNNs in utilizing non-invasive resting-state functional magnetic resonance imaging (rs-fMRI) data for brain graph construction and disease diagnosis. However, this computer-aided diagnostic (CAD) approach still needs to be further optimized in terms of data collection and model construction, and its diagnostic accuracy and clinical value are expected to be further improved. Therefore, this review conducts an in-depth and extensive investigation and analysis of GNNs in the diagnosis of neuro-psychiatric disorders. It presents the overall process of disease diagnosis, offering novel insights and methodologies to researchers in the field, thereby facilitating further improvement and optimization of related technologies.
Jin Gu, Xinbei Zha, Xiaole Zhao
Expert Syst. Appl.4
2025 DCTFormer: A Dual-Branch Transformer With Cloze Tests for Video Anomaly Detection
Shengdong Du, Xiaole Zhao, Jie Hu 0007, Jingjing Li 0001, Tianrui Li 0001
IEEE Trans. Multim.3
2024 AwmFace: Adaptive Weighting Margin for Deep Face Recognition
abstract
In the current field of face recognition, learning facial features with small intra-class variations and large interclass differences is crucial. Recent research has predominantly focused on enhancing the discriminative power of face models by incorporating fixed margin penalties into commonly used classification losses such as softmax loss, as seen in methods like Cosface and Arcface. However, when confronted with the challenges posed by noise and uncertainties in real-world data, conventional fixed-margin penalty methods may prove inadequate. Therefore, this paper introduces a novel loss named AwmFace, which uniquely addresses the effective learning of challenging samples encountered in face recognition tasks. AwmFace leverages the cosine angles between deep features and their corresponding weights to dynamically adjust margin penalties, allowing the model to capture complex features in the data more sensitively. By imposing additional strengthened margin penalties on challenging samples, AwmFace significantly enhances the model’s discriminative capacity towards these samples, thereby improving robustness and performance in real-world datasets. Our experiments on 9 mainstream benchmark tests have shown that our AwmFace performs significantly on 6 of the test sets. These results showcase the substantial improvement of our proposed AwmFace over fixed-margin penalty methods, consequently achieving state-of-the-art face recognition performance.
Xiangsong Jia, Xinkun Wu, Xiaole Zhao
IJCNN4
2024 Efficient Single Image Super-Resolution with Entropy Attention and Receptive Field Augmentation
abstract
Transformer-based deep models for single image super-resolution (SISR) have greatly improved the performance of lightweight SISR tasks in recent years. However, they often suffer from heavy computational burden and slow inference due to the complex calculation of multi-head self-attention (MSA), seriously hindering their practical application and deployment. In this work, we present an efficient SR model to mitigate the dilemma between model efficiency and SR performance, which is dubbed Entropy Attention and Receptive Field Augmentation network (EARFA), and composed of a novel entropy attention (EA) and a shifting large kernel attention (SLKA). From the perspective of information theory, EA increases the entropy of intermediate features conditioned on a Gaussian distribution, providing more informative input for subsequent reasoning. On the other hand, SLKA extends the receptive field of SR models with the assistance of channel shifting, which also favors to boost the diversity of hierarchical features. Since the implementation of EA and SLKA does not involve complex computations (such as extensive matrix multiplications), the proposed method can achieve faster nonlinear inference than Transformer-based SR models while maintaining better SR performance. Extensive experiments show that the proposed model can significantly reduce the delay of model inference while achieving the SR performance comparable with other advanced models.
Xiaole Zhao, Linze Li 0001, Chengxing Xie, Xiaoming Zhang 0008, Ting Jiang 0005, Shuaicheng Liu, Tianrui Li 0001
ACM Multimedia1
2024 Multitask-Guided Deep Clustering With Boundary Adaptation
abstract
Multitask learning uses external knowledge to improve internal clustering and single-task learning. Existing multitask learning algorithms mostly use shallow-level correlation to aid judgment, and the boundary factors on high-dimensional datasets often lead algorithms to poor performance. The initial parameters of these algorithms cause the border samples to fall into a local optimal solution. In this study, a multitask-guided deep clustering (DC) with boundary adaptation (MTDC-BA) based on a convolutional neural network autoencoder (CNN-AE) is proposed. In the first stage, dubbed multitask pretraining (M-train), we construct an autoencoder (AE) named CNN-AE using the DenseNet-like structure, which performs deep feature extraction and stores captured multitask knowledge into model parameters. In the second phase, the parameters of the M-train are shared for CNN-AE, and clustering results are obtained by deep features, which is termed as single-task fitting (S-fit). To eliminate the boundary effect, we use data augmentation and improved self-paced learning to construct the boundary adaptation. We integrate boundary adaptors into the M-train and S-fit stages appropriately. The interpretability of MTDC-BA is accomplished by data transformation. The model relies on the principle that features become important as the reconfiguration loss decreases. Experiments on a series of typical datasets confirm the performance of the proposed MTDC-BA. Compared with other traditional clustering methods, including single-task DC algorithms and the latest multitask clustering algorithms, our MTDC-BA achieves better clustering performance with higher computational efficiency. Deep features clustering results demonstrate the stability of MTDC-BA by visualization and convergence verification. Through the visualization experiment, we explain and analyze the whole model data input and the middle characteristic layer. Further understanding of the principle of MTDC-BA. Through additional experiments, we know that the proposed MTDC-BA is efficient in the use of multitask knowledge. Finally, we carry out sensitivity experiments on the hyper-parameters to verify their optimal performance.
Xiaole Zhao, Dengmin Wen, Donghai Zhai
IEEE Trans. Neural Networks Learn. Syst.3
2023 Lightweight Image Super-Resolution with Scale-wise Network
Xiaole Zhao, Xinkun Wu
BMVC1
2023 Boosting Single Image Super-Resolution via Partial Channel Shifting
abstract
Although deep learning has significantly facilitated the progress of single image super-resolution (SISR) in recent years, it still hits bottlenecks to further improve SR performance with the continuous growth of model scale. Therefore, one of the hotspots in the field is to construct efficient SISR models by elevating the effectiveness of feature representation. In this work, we present a straightforward and generic approach for feature enhancement that can effectively promote the performance of SR models, dubbed partial channel shifting (PCS). Specifically, it is inspired by the temporal shifting in video understanding and displaces part of the channels along the spatial dimensions, thus allowing the effective receptive field to be amplified and the feature diversity to be augmented at almost zero cost. Also, it can be assembled into off-the-shelf models as a plug-and-play component for performance boosting without extra network parameters and computational overhead. However, regulating the features with PCS encounters some issues, like shifting directions and amplitudes, proportions, patterns of shifted channels, etc. We impose some technical constraints on the issues to simplify the general channel shifting. Extensive and throughout experiments illustrate that the PCS indeed enlarges the effective receptive field, augments the feature diversity for efficiently enhancing SR recovery, and can endow obvious performance gains to existing models.
Xiaoming Zhang 0008, Tianrui Li 0001, Xiaole Zhao
ICCV3
2022 Single MR image super-resolution via channel splitting and serial fusion network
Xiaole Zhao, Yulun Zhang 0001, Yun Qin, Tao Zhang 0080, Tianrui Li 0001
Knowl. Based Syst.1
2022 Wide Weighted Attention Multi-Scale Network for Accurate MR Image Super-Resolution
abstract
High-quality magnetic resonance (MR) images afford more detailed information for reliable diagnoses and quantitative image analyses. Given low-resolution (LR) images, the deep convolutional neural network (CNN) has shown its promising ability for image super-resolution (SR). The LR MR images usually share some visual characteristics: structural textures of different sizes, edges with high correlation, and less informative background. However, multi-scale structural features are informative for image reconstruction, while the background is more smooth. Most previous CNN-based SR methods use a single receptive field and equally treat the spatial pixels (including the background). It neglects to sense the entire space and get diversified features from the input, which is critical for high-quality MR image SR. We propose a wide weighted attention multi-scale network ($\text{W}^{2}$AMSN) for accurate MR image SR to address these problems. On the one hand, the features of varying sizes can be extracted by the wide multi-scale branches. On the other hand, we design a non-reduction attention mechanism to recalibrate feature responses adaptively. Such attention preserves continuous cross-channel interaction and focuses on more informative regions. Meanwhile, the learnable weighted factors fuse extracted features selectively. The encapsulated wide weighted attention multi-scale block ($\text{W}^{2}$AMSB) is integrated through a recurrent framework and global attention mechanism. Extensive experiments and diversified ablation studies show the effectiveness of our proposed$\text{W}^{2}$AMSN, which surpasses state-of-the-art methods on most popular MR image SR benchmarks quantitatively and qualitatively. And our method still offers superior accuracy and adaptability on real MR images.
Haoqian Wang, Xiaowan Hu, Xiaole Zhao, Yulun Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2021 Pseudo 3D Auto-Correlation Network for Real Image Denoising
abstract
The extraction of auto-correlation in images has shown great potential in deep learning networks, such as the self-attention mechanism in the channel domain and the self-similarity mechanism in the spatial domain. However, the realization of the above mechanisms mostly requires complicated module stacking and a large number of convolution calculations, which inevitably increases model complexity and memory cost. Therefore, we propose a pseudo 3D auto-correlation network (P3AN) to explore a more efficient way of capturing contextual information in image de-noising. On the one hand, P3AN uses fast 1D convolution instead of dense connections to realize criss-cross interaction, which requires less computational resources. On the other hand, the operation does not change the feature size and makes it easy to expand. It means that only a simple adaptive fusion is needed to obtain contextual information that includes both the channel domain and the spatial domain. Our method built a pseudo 3D auto-correlation attention block through 1D convolutions and a lightweight 2D structure for more discriminative features. Extensive experiments have been conducted on three synthetic and four real noisy datasets. According to quantitative metrics and visual quality evaluation, the P3AN shows great superiority and surpasses state-of-the-art image denoising methods.
Xiaowan Hu, Ruijun Ma 0001, Yuanhao Cai, Xiaole Zhao, Yulun Zhang 0001, Haoqian Wang
CVPR5
2021 Pyramid Orthogonal Attention Network based on Dual Self-Similarity for Accurate Mr Image Super-Resolution
abstract
For magnetic resonance (MR) images sharing visual characteristics, the internal structure repetitions of different scales are considerable image-specific priors. Following the traditional algorithms, we try to combine external dataset-driven learning with the internal self-similarity for MR image super-resolution (SR). We propose a pyramid orthogonal attention network (POAN) based on dual self-similarity. On the one hand, by combining the point-similarity and the pyramid-similarity, sufficient spatial autocorrelation is explored to alleviate less training data limitation. On the other hand, the non-reduction channel attention mechanism maximizes inter-channel dependence. It increases the probability of the high-frequency region (e.g., structural textures and edges) being activated while suppresses low-frequency regions (e.g., background) adaptively. Out proposed POAN reconstructs the MR image under the guidance of pyramid orthogonal attention. Extensive experiments demonstrate that our method obtains the best results compared with state-of-the-art MR image SR methods quantitatively and visually.
Xiaowan Hu, Haoqian Wang, Yuanhao Cai, Xiaole Zhao, Yulun Zhang 0001
ICME4
2020 Accurate MR image super-resolution via lightweight lateral inhibition network
Xiaole Zhao, Xiafei Hu, Tao Zhang 0080, Xueming Zou, Jinsha Tian
Comput. Vis. Image Underst.1
2020 Rhythmic Network Modulation to Thalamocortical Couplings in Epilepsy
abstract
Thalamus interacts with cortical areas, generating oscillations characterized by their rhythm and levels of synchrony. However, little is known of what function the rhythmic dynamic may serve in thalamocortical couplings. This work introduced a general approach to investigate the modulatory contribution of rhythmic scalp network to the thalamo-frontal couplings in juvenile myoclonic epilepsy (JME) and frontal lobe epilepsy (FLE). Here, time-varying rhythmic network was constructed using the adapted directed transfer function between EEG electrodes, and then was applied as a modulator in fMRI-based thalamocortical functional couplings. Furthermore, the relationship between corticocortical connectivity and rhythm-dependent thalamocortical coupling was examined. The results revealed thalamocortical couplings modulated by EEG scalp network have frequency-dependent characteristics. Increased thalamus- sensorimotor network (SMN) and thalamus-default mode network (DMN) couplings in JME were strongly modulated by alpha band. These thalamus-SMN couplings demonstrated enhanced association with SMN-related corticocortical connectivity. In addition, altered theta-dependent and beta-dependent thalamus-frontoparietal network (FPN) couplings were found in FLE. The reduced theta-dependent thalamus-FPN couplings were associated with the decreased FPN-related corticocortical connectivity. This study proposed interactive links between the rhythmic modulation and thalamocortical coupling. The crucial role of SMN and FPN in subcortical-cortical circuit may have implications for intervention in generalized and focal epilepsy.
Yun Qin, Xiaojun Zuo, Sisi Jiang, Xiaole Zhao, Li Dong 0003, Jianfu Li, Tao Zhang 0017, Dezhong Yao 0001
Int. J. Neural Syst.6
2020 Gibbs-ringing artifact suppression with knowledge transfer from natural images to MR images
Xiaole Zhao, Huali Zhang, Yuliang Zhou, Tao Zhang 0080, Xueming Zou
Multim. Tools Appl.1
2019 Channel Splitting Network for Single MR Image Super-Resolution
abstract
High resolution magnetic resonance (MR) imaging is desirable in many clinical applications due to its contribution to more accurate subsequent analyses and early clinical diagnoses. Single image super-resolution (SISR) is an effective and cost efficient alternative technique to improve the spatial resolution of MR images. In the past few years, SISR methods based on deep learning techniques, especially convolutional neural networks (CNNs), have achieved the state-of-the-art performance on natural images. However, the information is gradually weakened and training becomes increasingly difficult as the network deepens. The problem is more serious for medical images because lacking high quality and effective training samples makes deep models prone to underfitting or overfitting. Nevertheless, many current models treat the hierarchical features on different channels equivalently, which is not helpful for the models to deal with the hierarchical features discriminatively and targetedly. To this end, we present a novel channel splitting network (CSN) to ease the representational burden of deep models. The proposed CSN model divides the hierarchical features into two branches, i.e., residual branch and dense branch, with different information transmissions. The residual branch is able to promote feature reuse, while the dense branch is beneficial to the exploration of new features. Besides, we also adopt the merge-and-run mapping to facilitate information integration between different branches. The extensive experiments on various MR images, including proton density (PD), T1, and T2 images, show that the proposed CSN model achieves superior performance over other state-of-the-art SISR methods.
Xiaole Zhao, Yulun Zhang 0001, Tao Zhang 0080, Xueming Zou
IEEE Trans. Image Process.1
2018 Multilevel Residual Learning for Single Image Super Resolution
Xiaole Zhao, Hangfei Liu, Tao Zhang 0080, Xueming Zou
PRCV (1)1
2016 Single image super-resolution via blind blurring estimation and anchored space mapping
abstract
It has been widely acknowledged that learning-based super-resolution (SR) methods are effective to recover a high resolution (HR) image from a single low resolution (LR) input image. However, there exist two main challenges in learning-based SR methods currently: the quality of training samples and the demand for computation. We proposed a novel framework for single image SR tasks aiming at these issues, which consists of blind blurring kernel estimation (BKE) and SR recovery with anchored space mapping (ASM). BKE is realized via minimizing the cross-scale dissimilarity of the image iteratively, and SR recovery with ASM is performed based on iterative least square dictionary learning algorithm (ILS-DLA). BKE is capable of improving the compatibility of training samples and testing samples effectively and ASM can reduce consumed time during SR recovery radically. Moreover, a selective patch processing (SPP) strategy measured by average gradient amplitude |grad| of a patch is adopted to accelerate the BKE process. The experimental results show that our method outruns several typical blind and non-blind algorithms on equal conditions.
Xiaole Zhao, Jinsha Tian, Hong-Ying Zhang 0004
Comput. Vis. Media1
2016 Single image super-resolution via blind blurring estimation and dictionary learning
Xiaole Zhao, Jinsha Tian, Hong-Ying Zhang 0004
Neurocomputing1