Cheng Zhao 0003

dblp:93/3598-3 · DBLP profile ↗
← Back
28ranked-venue papers
6as first author
28since 2021 · last 2026
0009-0004-1640-8710ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021
YearPublicationVenuePosition
2026 Hierarchical feature-guided dynamic collaborative learning transformer model for ventricular septal defect identification
Cheng Zhao 0003, Peng Yang 0011, Zhuo Xiang, Yiyao Liu, Bei Xia, Harry Qin, Tianfu Wang 0001, Bai Ying Lei, Luyao Zhou
Neurocomputing1
2026 Developing a knowledge-guided federated graph attention learning network with a diffusion module to diagnose Alzheimer's disease
Xuegang Song, Kaixiang Shu, Peng Yang 0011, Cheng Zhao 0003, Feng Zhou 0003, Alejandro F. Frangi, Jiuwen Cao, Xiaohua Xiao, Shuqiang Wang, Tianfu Wang 0001, Bai Ying Lei
Medical Image Anal.4
2026 Pyramid progressive image mapping network based on sparse annotations for cardiac segmentation
Zhuo Xiang, Cheng Zhao 0003, Yuantao Huang, Tianfu Wang 0001, Junqing Xu, Bai Ying Lei
Pattern Recognit.5
2026 Uncertainty-Guided Spatiotemporal Consistency Fusion Network for Infrared-Visible Video Fusion Under Extremely Low-Light Conditions
abstract
Infrared-visible video fusion under extremely low-light conditions is critically important yet remains underexplored, largely due to the scarcity of high-quality datasets and challenges posed by spatiotemporal uncertainty and modality bias. To address the dataset shortage, we built a dataset of 4,739 infrared and visible registration video pairs captured under extremely low-light conditions, spanning 5 scene types and 17 subcategories. Further, we proposed an Uncertainty-guided Spatiotemporal Consistency Fusion Network, termed USCFNet, for the infrared-visible video fusion. At each layer of the encoder, an Entropy-Gated SpatioTemporal Attention (EGSTA) module is introduced to capture temporal instability and spatial reliability variations through entropy-aware attention modulation, thereby enhancing feature spatiotemporal consistency. The refined infrared and visible features are then fused via a Difference-Guided Fusion (DGF) module, which adaptively exploits their content and edge differences to improve structural integrity and detail clarity. By progressively connecting DGF modules from shallow to deep layers, the network achieves the synergistic fusion of shallow textures and deep semantics. Subsequently, the output of the last DGF module is fused with the modality features of the last layer through a hierarchical mixture-of-experts fusion module. This module enables the balanced integration of modality information while preserving fine local details. Finally, the fusion feature is fed into the decoder to produce the final fused video. Extensive experiments on our dataset and two public datasets show that USCFNet outperforms competing methods, achieving lower distortion and stronger spatiotemporal consistency. The source code and dataset are available at https://github.com/Zhaocheng1/ELVID.
Cheng Zhao 0003, Tianyun Song, Zhiliang Wu, Tianfu Wang 0001, Moncef Gabbouj, Guanghui Yue 0001, Bai Ying Lei, Wei Zhou 0021
IEEE Trans. Image Process.1
2025 MMFN: Multi-Feature Multi-Modal Fusion Network for Diagnosis of Superficial Lymph Node Disease
abstract
The difficulty in identifying lymph node malignancies, including lymphoma and metastatic tumors, pose a diagnostic challenge at their primary sites. Given the heterogeneity of lymph node structures across different regions and the difficulty in distinguishing them from surrounding tissues, accurate diagnosis is often impeded. This research introduces multi-feature multi-modal fusion network (MMFN) for the differential diagnosis of benign and malignant lymph node diseases. The network integrates a convolusional neural network(CNN)-branch and a vision transformer(ViT)-branch to extract multi-scale features from ultrasound (US) and color doppler flow imaging (CDFI) images. By incorporating the convolutional block attention (CBA) module and cross modal attention (CMA) module, the network facilitates feature interaction and fusion across scales, leveraging blood flow information to enhance edge area detection. Furthermore, the feature fusion module (FFM) enables the interweaving of features from different dimensions, thereby enriching representational learning. Through experiments on private dataset, our approach demonstrates superior performance over existing methods.
Yuankun Wang, Cheng Zhao 0003, Yingxin Liu, Bai Ying Lei, Tianfu Wang 0001, Luyao Zhou
ICASSP2
2025 BGPCNet: Frequency Consistency and Boundary Guided Patch Contrast for Semi-supervised Segmentation of Superficial Lymphatic Disease
Yuankun Wang, Zhenghua Guan, Cheng Zhao 0003, Yingxin Liu, Bai Ying Lei, Tianfu Wang 0001, Luyao Zhou
PRCV (13)3
2025 Label-guided graph learning network via two-stage cross-modal fusion for multi-label skin disease diagnosis
Cheng Zhao 0003, Chunlun Xiao, Feifei Jin, Zhuo Xiang, Yiyao Liu, Lehang Guo, Tianfu Wang 0001, Bai Ying Lei
Eng. Appl. Artif. Intell.1
2025 MSMMIL: Multi-scan Mamba-based Multiple Instance Learning for whole slide image classification
Haiqin Zhong, Meidan Ding, Cheng Zhao 0003, Tianfu Wang 0001, Bai Ying Lei
Knowl. Based Syst.3
2025 ABVS breast tumour segmentation via integrating CNN with dilated sampling self-attention and feature interaction Transformer
Yiyao Liu, Jinyao Li, Yi Yang 0001, Cheng Zhao 0003, Peng Yang 0011, Xiaofei Deng, Tianfu Wang 0001, Bai Ying Lei
Neural Networks4
2025 An object detection-based model for automated screening of stem-cells senescence during drug screening
Youyi Song, Mingzhu Li, Liangge He, Chunlun Xiao, Peng Yang 0011, Cheng Zhao 0003, Tianfu Wang 0001, Guangqian Zhou, Bai Ying Lei
Neural Networks8
2025 Federated learning via multi-attention guided UNet for thyroid nodule segmentation of ultrasound images
Zhuo Xiang, Xiaoyu Tian, Yiyao Liu, Minsi Chen, Cheng Zhao 0003, Li-Na Tang, En-Sheng Xue, Hong-Yuan Xue, Ying-Jia Li, Quan-Shui Li, Chang-Jun Wu, Tian-Tian Ren, Jin-Yu Wu, Tianfu Wang 0001, Wen-Ying Liu, Bo-Ji Liu, Li-Ping Sun, Chong-Ke Zhao, Hui-Xiong Xu, Bai Ying Lei
Neural Networks5
2025 Boundary-Guided Feature-Aligned Network for Colorectal Polyp Segmentation
abstract
Colorectal polyp segmentation in endoscopic images is very important for the prevention and treatment of colorectal cancer. Because of the high similarity between polyps and their surrounding tissues, most deep neural network (DNN) based methods often struggle with blurry boundaries and result in inaccurate segmentation. In this paper, we propose a Boundary-guided Feature-aligned Network (BFNet) for polyp segmentation by taking a boundary prediction task as an auxiliary. Firstly, BFNet aggregates multi-layer features extracted from the backbone to mine boundary cues. Secondly, a flexible feature aggregation (FFA) module is used at each layer to adaptively fuse cross-layer features for coarse polyp localization. In the FFA module, considering the spatial misalignment between features at different layers, the feature of the high layer is aligned to and fused with that of the current layer using the deformable convolution and flexible merge block. After that, a boundary-guided feature enhancement (BFE) module is applied to refine the localization at boundary areas. In the BFE module, the boundary information is extracted and highlighted in both channel and spatial dimensions using the attention mechanisms with the assistance of boundary cues. By applying deep supervision to the BFE modules, BFNet can produce accurate polyp segmentation. Experimental results show that our BFNet outperforms 14 state-of-the-art DNN-based polyp segmentation methods on both in-domain and out-of-domain tests.
Guanghui Yue 0001, Shangjie Wu, Cheng Zhao 0003, Tianwei Zhou, Baoquan Zhao
IEEE Trans. Circuits Syst. Video Technol.4
2025 Text-Guided Semantic Alignment Network With Spatial-Frequency Interaction for Infrared-Visible Image Fusion Under Extreme Illumination
abstract
Although text-guided infrared-visible image fusion helps improve content understanding under extreme illumination, existing methods usually ignore semantic differences between textual and visual features, resulting in limited improvement. To address this challenge, we propose a Text-Guided Semantic Alignment Network, termed TSANet, for extreme-illumination infrared-visible image fusion. The network follows an encoder-decoder structure, with two image encoders, two text encoders, and one decoder. It uses a Semantic Alignment and Fusion (SAF) block to bridge the two image encoders in each layer. Specifically, the SAF block consists of two parallel Semantic Alignment (SA) modules, corresponding to the infrared and visible modalities, respectively, and a Spatial-Frequency Interaction (SFI) module. The SA module aligns the visual feature from the image encoder with its corresponding textual feature from the text encoder, to guide the network focus on key semantic regions of infrared and visible images. The SFI module aggregates the spatial and frequency information extracted from the modality-aligned features of two SA modules for complementary representation learning. The network progressively complements two image modalities by connecting the SAF blocks from top to down, and finally provides a visually pleasing fusion effect by feeding the output of the last block into the decoder. Recognizing that existing datasets lack illumination diversity, we contribute a new dataset specifically designed for extreme-illumination image fusion. Extensive experiments show the effectiveness and superiority of TSANet over seven state-of-the-art methods. The source code and dataset are available at https://github.com/WentaoLi-CV/TSANet.
Guanghui Yue 0001, Cheng Zhao 0003, Zhiliang Wu, Tianwei Zhou, Qiuping Jiang, Runmin Cong
IEEE Trans. Image Process.3
2025 Dual-Scale Swin Transformer via Feature Alignment and Adversarial Discrimination for Retinopathy of Prematurity Diagnosis
abstract
Retinopathy of prematurity (ROP) is a retinal vascular disease that primarily affects premature infants with low birth weight. It is a leading cause of childhood blindness worldwide, but it can often be effectively managed with appropriate and timely diagnosis and treatment. To address the impact of image style on model classification performance, this paper proposes a dual-scale Swin Transformer (DS-Swin-T) network for ROP. The network comprises three components: image synthesis (IS), feature alignment, and advanced adversarial learning. The IS module generates synthesis style images as an intermediate latent space between source and target styles, reducing style difference. The DS-Swin-T serves as the primary framework for image feature extraction. Detail and style encoders extract features in the shallow feature space, with detail and style losses aligning these features to ensure consistency across styles. To extract rich style-invariant features and ensure consistent classification within the same category, adversarial learning is applied in the advanced feature space. Finally, feature fusion units process dual-scale classification representations. Our method achieves an average accuracy of 97.91% on the source style dataset. When transferred to other target style datasets, our method effectively mitigates the performance degradation caused by style difference, reaching a maximum average accuracy of 93.66%. Extensive experiments demonstrate the effectiveness of our method.
Shaobin Chen, Yiyao Liu, Hai Xie, Zhenquan Wu, Yingpeng Xie, Cheng Zhao 0003, Tianfu Wang 0001, Bai Ying Lei
IEEE J. Biomed. Health Informatics9
2025 Self-Supervised Multi-Scale Multi-Modal Graph Pool Transformer for Sellar Region Tumor Diagnosis
abstract
The sellar region tumor is a brain tumor that only exists in the brain sellar, which affects the central nervous system. The early diagnosis of the sellar region tumor subtypes helps clinicians better understand the best treatment and recovery of patients. Magnetic resonance imaging (MRI) has proven to be an effective tool for the early detection of sellar region tumors. However, the existing sellar region tumor diagnosis still remains challenging due to the small amount of dataset and data imbalance. To overcome these challenges, we propose a novel self-supervised multi-scale multi-modal graph pool Transformer (MMGPT) network that can enhance the multi-modal fusion of small and imbalanced MRI data of sellar region tumors. MMGPT can strengthen feature interaction between multi-modal images, which makes our model more robust. A contrastive learning equipped auto-encoder (CAE) via self-supervised learning (SSL) is adopted to learn more detailed information between different samples. The proposed CAE transfers the pre-trained knowledge to the downstream tasks. Finally, a hybrid loss is equipped to relieve the performance degradation caused by data imbalance. The experimental results show that the proposed method outperforms state-of-the-art methods and obtains higher accuracy and AUC in the classification of sellar region tumors.
Bai Ying Lei, Gege Cai, Yun Zhu 0006, Tianfu Wang 0001, Cheng Zhao 0003, Xinzhi Hu, Huijun Zhu, Ming Feng, Renzhi Wang 0002
IEEE J. Biomed. Health Informatics6
2025 FAMF-Net: Feature Alignment Mutual Attention Fusion With Region Awareness for Breast Cancer Diagnosis via Imbalanced Data
abstract
Automatic and accurate classification of breast cancer in multimodal ultrasound images is crucial to improve patients' diagnosis and treatment effect and save medical resources. Methodologically, the fusion of multimodal ultrasound images often encounters challenges such as misalignment, limited utilization of complementary information, poor interpretability in feature fusion, and imbalances in sample categories. To solve these problems, we propose a feature alignment mutual attention fusion method (FAMF-Net), which consists of a region awareness alignment (RAA) block, a mutual attention fusion (MAF) block, and a reinforcement learning-based dynamic optimization strategy(RDO). Specifically, RAA achieves region awareness through class activation mapping and performs translation transformation to achieve feature alignment. When MAF utilizes a mutual attention mechanism for feature interaction fusion, it mines edge and color features separately in B-mode and shear wave elastography images, enhancing the complementarity of features and improving interpretability. Finally, RDO uses the distribution of samples and prediction probabilities during training as the state of reinforcement learning to dynamically optimize the weights of the loss function, thereby solving the problem of class imbalance. The experimental results based on our clinically obtained dataset demonstrate the effectiveness of the proposed method. Our code will be available at: https://github.com/Magnety/Multi_modal_Image.
Yiyao Liu, Jinyao Li, Cheng Zhao 0003, Harry Qin, Tianfu Wang 0001, Bai Ying Lei
IEEE Trans. Medical Imaging3
2025 Knowledge-Aware Multisite Adaptive Graph Transformer for Brain Disorder Diagnosis
abstract
Brain disorder diagnosis via resting-state functional magnetic resonance imaging (rs-fMRI) is usually limited due to the complex imaging features and sample size. For brain disorder diagnosis, the graph convolutional network (GCN) has achieved remarkable success by capturing interactions between individuals and the population. However, there are mainly three limitations: 1) The previous GCN approaches consider the non-imaging information in edge construction but ignore the sensitivity differences of features to non-imaging information. 2) The previous GCN approaches solely focus on establishing interactions between subjects (i.e., individuals and the population), disregarding the essential relationship between features. 3) Multisite data increase the sample size to help classifier training, but the inter-site heterogeneity limits the performance to some extent. This paper proposes a knowledge-aware multisite adaptive graph Transformer to address the above problems. First, we evaluate the sensitivity of features to each piece of non-imaging information, and then construct feature-sensitive and feature-insensitive subgraphs. Second, after fusing the above subgraphs, we integrate a Transformer module to capture the intrinsic relationship between features. Third, we design a domain adaptive GCN using multiple loss function terms to relieve data heterogeneity and to produce the final classification results. Last, the proposed framework is validated on two brain disorder diagnostic tasks. Experimental results show that the proposed framework can achieve state-of-the-art performance.
Xuegang Song, Kaixiang Shu, Peng Yang 0011, Cheng Zhao 0003, Feng Zhou 0003, Alejandro F. Frangi, Xiaohua Xiao, Tianfu Wang 0001, Shuqiang Wang, Bai Ying Lei
IEEE Trans. Medical Imaging4
2025 Attention-Guided Learning With Feature Reconstruction for Skin Lesion Diagnosis Using Clinical and Ultrasound Images
abstract
Skin lesion is one of the most common diseases, and most categories are highly similar in morphology and appearance. Deep learning models effectively reduce the variability between classes and within classes, and improve diagnostic accuracy. However, the existing multi-modal methods are only limited to the surface information of lesions in skin clinical and dermatoscopic modalities, which hinders the further improvement of skin lesion diagnostic accuracy. This requires us to further study the depth information of lesions in skin ultrasound. In this paper, we propose a novel skin lesion diagnosis network, which combines clinical and ultrasound modalities to fuse the surface and depth information of the lesion to improve diagnostic accuracy. Specifically, we propose an attention-guided learning (AL) module that fuses clinical and ultrasound modalities from both local and global perspectives to enhance feature representation. The AL module consists of two parts, attention-guided local learning (ALL) computes the intra-modality and inter-modality correlations to fuse multi-scale information, which makes the network focus on the local information of each modality, and attention-guided global learning (AGL) fuses global information to further enhance the feature representation. In addition, we propose a feature reconstruction learning (FRL) strategy which encourages the network to extract more discriminative features and corrects the focus of the network to enhance the model's robustness and certainty. We conduct extensive experiments and the results confirm the superiority of our proposed method. Our code is available at: https://github.com/XCL-hub/AGFnet.
Chunlun Xiao, Chunmei Xia, Zifeng Qiu, Yuanlin Liu, Cheng Zhao 0003, Weiwei Ren, Lifan Wang, Tianfu Wang 0001, Lehang Guo, Bai Ying Lei
IEEE Trans. Medical Imaging6
2024 Multi-modality Correlation Learning Network for Pediatric Ventricular Septal Defects Identification
Feifei Jin, Cheng Zhao 0003, Zhuo Xiang, Xunyi Chen, Yu Zhang 0009, Shumin Fan, Luyao Zhou, Tianfu Wang 0001, Bai Ying Lei
PRCV (15)2
2024 Specificity-Aware Federated Learning With Dynamic Feature Fusion Network for Imbalanced Medical Image Classification
abstract
Recently, federated learning has become a powerful technique for medical image classification due to its ability to utilize datasets from multiple clinical clients while satisfying privacy constraints. However, there are still some obstacles in federated learning. Firstly, most existing methods directly average the model parameters collected by medical clients on the server, ignoring the specificities of the local models. Secondly, class imbalance is a common issue in medical datasets. In this article, to handle these two challenges, we propose a novel specificity-aware federated learning framework that benefits from an Adaptive Aggregation Mechanism (AdapAM) and a Dynamic Feature Fusion Strategy (DFFS). Considering the specificity of each local model, we set the AdapAM on the server. The AdapAM utilizes reinforcement learning to adaptively weight and aggregate the parameters of local models based on their data distribution and performance feedback for obtaining the global model parameters. For the class imbalance in local datasets, we propose the DFFS to dynamically fuse the features of majority classes based on the imbalance ratio in the min-batch and collaborate the rest of features. We conduct extensive experiments on a dermoscopic dataset and a fundus image dataset. Experimental results show that our method can achieve state-of-the-art results in these two real-world medical applications.
Guanghui Yue 0001, Peishan Wei, Tianwei Zhou, Youyi Song, Cheng Zhao 0003, Tianfu Wang 0001, Bai Ying Lei
IEEE J. Biomed. Health Informatics5
2024 MHW-GAN: Multidiscriminator Hierarchical Wavelet Generative Adversarial Network for Multimodal Image Fusion
abstract
Image fusion technology aims to obtain a comprehensive image containing a specific target or detailed information by fusing data of different modalities. However, many deep learning-based algorithms consider edge texture information through loss functions instead of specifically constructing network modules. The influence of the middle layer features is ignored, which leads to the loss of detailed information between layers. In this article, we propose a multidiscriminator hierarchical wavelet generative adversarial network (MHW-GAN) for multimodal image fusion. First, we construct a hierarchical wavelet fusion (HWF) module as the generator of MHW-GAN to fuse feature information at different levels and scales, which avoids information loss in the middle layers of different modalities. Second, we design an edge perception module (EPM) to integrate edge information from different modalities to avoid the loss of edge information. Third, we leverage the adversarial learning relationship between the generator and three discriminators for constraining the generation of fusion images. The generator aims to generate a fusion image to fool the three discriminators, while the three discriminators aim to distinguish the fusion image and edge fusion image from two source images and the joint edge image, respectively. The final fusion image contains both intensity information and structure information via adversarial learning. Experiments on public and self-collected four types of multimodal image datasets show that the proposed algorithm is superior to the previous algorithms in terms of both subjective and objective evaluation.
Cheng Zhao 0003, Peng Yang 0011, Feng Zhou 0003, Guanghui Yue 0001, Shuigen Wang, Huisi Wu, Guoliang Chen 0005, Tianfu Wang 0001, Bai Ying Lei
IEEE Trans. Neural Networks Learn. Syst.1
2023 Multi-scale enhanced graph convolutional network for mild cognitive impairment detection
Bai Ying Lei, Yun Zhu 0006, Shuangzhi Yu, Huoyou Hu, Yanwu Xu 0001, Guanghui Yue 0001, Tianfu Wang 0001, Cheng Zhao 0003, Shaobin Chen, Peng Yang 0011, Xuegang Song, Xiaohua Xiao, Shuqiang Wang
Pattern Recognit.8
2023 FIT-Net: Feature Interaction Transformer Network for Pathologic Myopia Diagnosis
abstract
Automatic and accurate classification of retinal optical coherence tomography (OCT) images is essential to assist physicians in diagnosing and grading pathological changes in pathologic myopia (PM). Clinically, due to the obvious differences in the position, shape, and size of the lesion structure in different scanning directions, ophthalmologists usually need to combine the lesion structure in the OCT images in the horizontal and vertical scanning directions to diagnose the type of pathological changes in PM. To address these challenges, we propose a novel feature interaction Transformer network (FIT-Net) to diagnose PM using OCT images, which consists of two dual-scale Transformer (DST) blocks and an interactive attention (IA) unit. Specifically, FIT-Net divides image features of different scales into a series of feature block sequences. In order to enrich the feature representation, we propose an IA unit to realize the interactive learning of class token in feature sequences of different scales. The interaction between feature sequences of different scales can effectively integrate different scale image features, and hence FIT-Net can focus on meaningful lesion regions to improve the PM classification performance. Finally, by fusing the dual-view image features in the horizontal and vertical scanning directions, we propose six dual-view feature fusion methods for PM diagnosis. The extensive experimental results based on the clinically obtained datasets and three publicly available datasets demonstrate the effectiveness and superiority of the proposed method. Our code is avaiable at: https://github.com/chenshaobin/FITNet.
Shaobin Chen, Zhenquan Wu, Mingzhu Li, Yun Zhu 0006, Hai Xie, Peng Yang 0011, Cheng Zhao 0003, Shaochong Zhang, Bai Ying Lei
IEEE Trans. Medical Imaging7
2022 Diagnosis of obsessive-compulsive disorder via spatial similarity-aware learning and fused deep polynomial network
Peng Yang 0011, Cheng Zhao 0003, Qiong Yang, Wei Zheng 0009, Xiaohua Xiao, Li Shen 0001, Tianfu Wang 0001, Bai Ying Lei, Ziwen Peng
Medical Image Anal.2
2022 IFT-Net: Interactive Fusion Transformer Network for Quantitative Analysis of Pediatric Echocardiography
Cheng Zhao 0003, Harry Qin, Peng Yang 0011, Zhuo Xiang, Alejandro F. Frangi, Minsi Chen, Shumin Fan, Wei Yu 0002, Xunyi Chen, Bei Xia, Tianfu Wang 0001, Bai Ying Lei
Medical Image Anal.1
2021 Multi-directional Attention Network for Segmentation of Pediatric Echocardiographic
Zhuo Xiang, Cheng Zhao 0003, Libao Guo, Yali Qiu, Yun Zhu 0006, Peng Yang 0011, Mingzhu Li, Minsi Chen, Tianfu Wang 0001, Bai Ying Lei
PRCV (3)2
2021 Dual attention enhancement feature fusion network for segmentation and quantitative analysis of paediatric echocardiography
Libao Guo, Bai Ying Lei, Jie Du 0001, Alejandro F. Frangi, Harry Qin, Cheng Zhao 0003, Pengpeng Shi, Bei Xia, Tianfu Wang 0001
Medical Image Anal.7
2021 Medical image fusion method based on dense block and deep convolutional generative adversarial network
Cheng Zhao 0003, Tianfu Wang 0001, Bai Ying Lei
Neural Comput. Appl.1