Jianchun Zhao

dblp:245/7475 · DBLP profile ↗
← Back
16ranked-venue papers
0as first author
13since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 7 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RL-U2Net: A Dual-Branch UNet with Reinforcement Learning-Assisted Multimodal Feature Fusion for Accurate 3D Whole-Heart Segmentation
abstract
Accurate whole-heart segmentation is a critical component in the precise diagnosis and interventional planning of cardiovascular diseases. Integrating complementary information from modalities such as computed tomography (CT) and magnetic resonance imaging (MRI) can significantly enhance segmentation accuracy and robustness. However, existing multi-modal segmentation methods face several limitations: severe spatial inconsistency between modalities hinders effective feature fusion; fusion strategies are often static and lack adaptability; and the processes of feature alignment and segmentation are decoupled and inefficient. To address these challenges, we propose a dual-branch U-Net architecture enhanced by reinforcement learning for feature alignment, termed RL-U2Net, designed for precise and efficient multi-modal 3D whole-heart segmentation. The model employs a dual-branch U-shaped network to process CT and MRI patches in parallel, and introduces a novel RL-XAlign module between the encoders. The module employs a cross‑modal attention mechanism to capture semantic correspondences between modalities and a reinforcement learning agent learns an optimal rotation strategy that consistently aligns anatomical pose and texture features. The aligned features are then reconstructed through their respective decoders. Finally, an ensemble‑learning–based decision module integrates the predictions from individual patches to produce the final segmentation result. Experimental results on the publicly available MM-WHS 2017 dataset demonstrate that the proposed RL-U2Net outperforms existing state-of-the-art methods, achieving Dice coefficients of 93.1% on CT and 87.0% on MRI, thereby validating the effectiveness and superiority of the proposed approach.
Jierui Qu, Jianchun Zhao
AAAI2
2024 PaRCL: Pathology-aware Representation Contrastive Learning for Glaucoma Classification on Fundus Images
abstract
Recently, there has been a growing interest in applying self-supervised contrastive learning to medical image classification tasks. In this paper, we investigate its application for glaucoma classification using fundus images. Given the highly detailed features in fundus images, the effectiveness of contrastive learning, particularly when data augmentation is used to create contrastive pairs, proves to be notably limited. This constraint impairs the extraction of fine-grained visual representations on fundus images, resulting in reduced performance of self-supervised contrastive learning for glaucoma classification. To alleviate the above issues, we propose a pathology-aware representations contrastive learning (PaRCL) method, which learns more discriminative fine-grained representations compared to the previous self-supervised contrastive learning methods, to enhance the model performance for the glaucoma classification effectively. Furthermore, our method could focus more on the lesion area and effectively extract interpretable classification features. Particularly, to fully learn more discriminative visual representations, we introduce a pathology-aware subnetwork ensemble strategy (SES) as an alternative to generate contrastive pairs through traditional data augmentation techniques. Experiments on three clinical datasets demonstrate that our PaRCL could improve the performance of glaucoma classification on fundus images and could effectively focus on pathology areas of fundus images for disease diagnosis.
Junyan Yi, Dayong Ding, Jianchun Zhao, Gang Yang 0001
BIBM4
2024 A geometry-aware multi-coordinate transformation fusion network for optic disc and cup segmentation
Yajie Yang, Gang Yang 0001, Yanni Wang, Jianchun Zhao, Dayong Ding
Appl. Intell.5
2023 Supervised Domain Adaptation for Recognizing Retinal Diseases from Wide-Field Fundus Images
abstract
This paper addresses the emerging task of recognizing multiple retinal diseases from wide-field (WF) and ultra-wide-field (UWF) fundus images. For an effective use of existing large amount of labeled color fundus photo (CFP) data and the relatively small amount of WF and UWF data, we propose a supervised domain adaptation method named Cross-domain Collaborative Learning (CdCL). Inspired by the success of fixed-ratio based mixup in unsupervised domain adaptation, we re-purpose this strategy for the current task. Due to the intrinsic disparity between the field-of-view of CFP and WF/UWF images, a scale bias naturally exists in a mixup sample that the anatomic structure from a CFP image will be considerably larger than its WF/UWF counterpart. The CdCL method resolves the issue by Scale-bias Correction, which employs Transformers for producing scale-invariant features. As demonstrated by extensive experiments on multiple datasets covering both WF and UWF images, the proposed method compares favorably against a number of competitive baselines.
Qijie Wei, Jingyuan Yang 0004, Bo Wang 0011, Jinrui Wang, Jianchun Zhao, Niranchana Manivannan, Youxin Chen, Dayong Ding, Jing Zhou 0005, Xirong Li 0001
BIBM5
2023 LACL: Lesion-Aware Contrastive Learning Framework for Medical Image Classification
abstract
Recently, contrastive learning has received significant attention in various classification tasks of natural images. However, current contrastive learning frameworks display unsatisfactory performance on medical images due to the inability of obtaining fine-grained visual features. In this paper, we propose a Lesion-Aware Contrastive Learning (LACL) framework to learn more discriminative and comprehensive representations and reinforce the attention of diagnosis regions on medical images. LACL framework includes two phases of training procedures: the comprehensive feature-extracting phase and the contrastive learning enhancement phase. In the first phase, LACL fully captures meaningful deep features related to the training targets to form comprehensive visual representations, by training a novel lesion-aware module we proposed. In the second phase, we introduce the previous representation information into contrastive learning to guide the LACL framework in learning disease- related features. This approach provides more effective guidance than the traditional contrastive learning method of directly comparing features. Extensive experiments on several benchmark datasets demonstrate that our LACL framework significantly improves the performance of medical image classification and highlights the lesion areas for disease diagnosis.
Gang Yang 0001, Jianchun Zhao, Dayong Ding, Jun Wu 0022
ICME3
2023 Automatic Retinal Nerve Fiber Trajectory Simulation and Quasi-polar Transformation for Detecting Retinal Nerve Fiber Layer Defect in Fundus Images
abstract
The retinal nerve fiber layer defects (RNFLD) provide early objective evidence for many retinal abnormalities, especially early glaucoma. The recent success of deep learning has led to exciting prospects in automating the detection of RNFLD, but it is highly dependent on the availability of large-scale datasets carefully annotated by experienced ophthalmologists, which is both time-consuming and cost-prohibitive. Most previous works on RNFLD detection lack a unified annotation scheme. In addition, they fail to exploit the intrinsic morphological characteristics of retinal nerve fiber bundles (RNFB). This work presents an automatic RNFB tracing method that is applicable not only in automating the RNFLD annotation process, but also in performing a quasi-polar transformation on fundus images for subsequent RNFLD detection tasks. Also, our proposed method paves the way for alleviating the problem of inter-annotator variation incurred merely by different annotation habits among annotators and reducing the noise before the input stage. Experiments reveal that our proposed quasi-polar transformation provides a significant boost to the existing RNFLD detection method and surpasses the state-of-the-art F1 by 4.2%, and that our method also offers a more accurate and reasonable description of RNFLD detection results, which demonstrates the orientation of RNFLD rather than vague position indication.
Yanni Wang, Gang Yang 0001, Dayong Ding, Jianchun Zhao
ICME4
2023 Representation, Alignment, Fusion: A Generic Transformer-Based Framework for Multi-modal Glaucoma Recognition
Gang Yang 0001, Dayong Ding, Jianchun Zhao
MICCAI (7)5
2022 Optic Disc Hemorrhage Detection via A Novel Position-Guided Attention Network with Small Samples on Fundus Images
abstract
Optic disc hemorrhage (ODH) is an important lesion factor for eye disease diagnosis. ODH has some specific displaying characteristics on fundus images, including covering fuzzy small domains and displaying similar to vessels, which greatly increase ODH detection difficulty. In this paper, we propose a novel position-guided attention network to detect optic disc hemorrhage on fundus images. Our method greatly takes advantage of the prior knowledge of ODH position information and the multitask learning dependencies related to the ODH segmentation and ODH classification, to build a position information attention module and a correlative feature fusion module. Moreover, we introduce a disc edge strength map on the optic disc to constrain the attention domains. Due to the limitation of ODH labeled data, an online hard case segmentation strategy on small samples are proposed to train our method. Experiments show that our method could greatly reduce the detection of false positives when detecting ODH on fundus images, so as to obtain superior performance on the ODH detection with small samples. Further, a fundus image dataset with professional ODH labels is published to advance the research of ODH detection (https://github.com/JieGenius/disc_hemo_dataset).
Gang Yang 0001, Yunfeng Du, Yajie Yang, Jianchun Zhao, Dayong Ding, Gangwei Cheng
BIBM4
2022 Contour Offset Map: A New Component Designed for Smooth and Robust Optic Disc/Cup Contour Detection
abstract
The optic disc and cup contour detection task has been studied extensively, as it is the predominant attribute to calculate cup-to-disc ratio (CDR) and further diagnose glaucoma. Currently, methods for optic disc and cup contour detection task mostly include two types: the coordinate point regression and the probability map classification. Normally, the latter is better. However, the probability map is unfavorable to be directly applied to the optic disc and cup contour detection task, which makes the model outputs more than one optic disc and cup in some cases, and uneven results. In addition, the probability map classification methods can not directly optimize the calculation of CDR in the training process, which further limits the performance of neural networks. In this paper, we design a new optic disc and cup contour representation method called Contour Offset Map to solve the two mentioned problems. Our approach can be applied seamlessly to the existing semantic segmentation networks by only modifying the loss functions and the number of output maps rather than the network structure. Moreover, a rim loss and a CDR loss are introduced to further improve the performance of our method in the training process. Experiments on our private dataset and a public dataset demonstrate the superior performance of our Contour Offset Map on the task of the optic disc and cup contour detection.
Yajie Yang, Gang Yang 0001, Dayong Ding, Jianchun Zhao
BIBM4
2022 Semi-supervised Learning for Nerve Segmentation in Corneal Confocal Microscope Photography
Jun Wu 0022, Qi Pan, Jianchun Zhao, Gang Yang 0001, Xirong Li 0001, Dayong Ding
MICCAI (4)8
2022 Lesion Localization in OCT by Semi-Supervised Object Detection
abstract
Over 300 million people worldwide are affected by various retinal diseases. By noninvasive Optical Coherence Tomography (OCT) scans, a number of abnormal structural changes in the retina, namely retinal lesions, can be identified. Automated lesion localization in OCT is thus important for detecting retinal diseases at their early stage. To conquer the lack of manual annotation for deep supervised learning, this paper presents a first study on utilizing semi-supervised object detection (SSOD) for lesion localization in OCT images. To that end, we develop a taxonomy to provide a unified and structured viewpoint of the current SSOD methods, and consequently identify key modules in these methods. To evaluate the influence of these modules in the new task, we build OCT-SS, a new dataset consisting of over 1k expert-labeled OCT B-scan images and over 13k unlabeled B-scans. Extensive experiments on OCT-SS identify Unbiased Teacher (UnT) as the best current SSOD method for lesion localization. Moreover, we improve over this strong baseline, with mAP increased from 49.34 to 50.86.
Jianchun Zhao, Jingyuan Yang 0004, Weihong Yu, Youxin Chen, Xirong Li 0001
ICMR3
2022 Learning Two-Stream CNN for Multi-Modal Age-Related Macular Degeneration Categorization
abstract
This paper tackles automated categorization of Age-related Macular Degeneration (AMD), a common macular disease among people over 50. Previous research efforts mainly focus on AMD categorization with a single-modal input, let it be a color fundus photograph (CFP) or an OCT B-scan image. By contrast, we consider AMD categorization given a multi-modal input, a direction that is clinically meaningful yet mostly unexplored. Contrary to the prior art that takes a traditional approach of feature extraction plus classifier training that cannot be jointly optimized, we opt for end-to-end multi-modal Convolutional Neural Networks (MM-CNN). Our MM-CNN is instantiated by a two-stream CNN, with spatially-invariant fusion to combine information from the CFP and OCT streams. In order to visually interpret the contribution of the individual modalities to the final prediction, we extend the class activation mapping (CAM) technique to the multi-modal scenario. For effective training of MM-CNN, we develop two data augmentation methods. One is GAN-based CFP/OCT image synthesis, with our novel use of CAMs as conditional input of a high-resolution image-to-image translation GAN. The other method is Loose Pairing, which pairs a CFP image and an OCT image on the basis of their classes instead of eye identities. Experiments on a clinical dataset consisting of 1,094 CFP images and 1,289 OCT images acquired from 1,093 distinct eyes show that the proposed solution obtains better F1 and Accuracy than multiple baselines for multi-modal AMD categorization. Code and data are available at https://github.com/li-xirong/mmc-amd.
Weisen Wang, Xirong Li 0001, Zhiyan Xu, Weihong Yu, Jianchun Zhao, Dayong Ding, Youxin Chen
IEEE J. Biomed. Health Informatics5
2021 Multi-Modal Multi-Instance Learning for Retinal Disease Recognition
abstract
This paper attacks an emerging challenge of multi-modal retinal disease recognition. Given a multi-modal case consisting of a color fundus photo (CFP) and an array of OCT B-scan images acquired during an eye examination, we aim to build a deep neural network that recognizes multiple vision-threatening diseases for the given case. As the diagnostic efficacy of CFP and OCT is disease-dependent, the network's ability of being both selective and interpretable is important. Moreover, as both data acquisition and manual labeling are extremely expensive in the medical domain, the network has to be relatively lightweight for learning from a limited set of labeled multi-modal samples. Prior art on retinal disease recognition focuses either on a single disease or on a single modality, leaving multi-modal fusion largely underexplored. We propose in this paper Multi-Modal Multi-Instance Learning (MM-MIL) for selectively fusing CFP and OCT modalities. Its lightweight architecture (as compared to current multi-head attention modules) makes it suited for learning from relatively small-sized datasets. For an effective use of MM-MIL, we propose to generate a pseudo sequence of CFPs by over sampling a given CFP. The benefits of this tactic include well balancing instances across modalities, increasing the resolution of the CFP input, and finding out regions of the CFP most relevant with respect to the final diagnosis. Extensive experiments on a real-world dataset consisting of 1,206 multi-modal cases from 1,193 eyes of 836 subjects demonstrate the viability of the proposed model.
Xirong Li 0001, Hailan Lin, Jianchun Zhao, Dayong Ding, Weihong Yu, Youxin Chen
ACM Multimedia5
2020 Deep Multiple Instance Learning with Spatial Attention for ROP Case Classification, Instance Selection and Abnormality Localization
abstract
This paper tackles automated screening of Retinopathy of Prematurity (ROP), one of the most common causes of visual loss in childhood. Clinically, ROP screening per case requires multiple color fundus image instances that capture different zones of the (premature) retina. A desirable model shall not only make a decision at the case level, but also pinpoint which instances and what part of the instances are responsible for the decision. This paper makes the first attempt to accomplish three tasks, i.e. ROP case classification, instance selection and abnormality localization in a unified framework. To that end, we propose a new model that effectively combines instance-attention based deep multiple instance learning (MIL) and spatial attention (SA). The propose model, which we term MIL-SA, identifies positive instances in light of their contributions to case-level decision. Meanwhile, abnormal regions in the identified instances are automatically localized by the SA mechanism. Moreover, MIL-SA is learned from case-level binary labels exclusively, and in an end-to-end manner. Experiments on a large clinical dataset of 2,186 cases with 11,053 fundus images show the viability of the proposed model for all the three tasks.
Xirong Li 0001, Wencui Wan, Jianchun Zhao, Qijie Wei, Junbo Rong, Pengyi Zhou, Limin Xu, Lijuan Lang, Chengzhi Niu, Dayong Ding, Xuemin Jin
ICPR4
2020 AttenNet: Deep Attention Based Retinal Disease Classification in OCT Images
Jun Wu 0022, Jianchun Zhao, Dayong Ding, Ningjiang Chen, Chunhui Jiang, Xuan Zou, Yuan Tian 0017, Zongjiang Shang, Kaiwei Wang, Xirong Li 0001, Gang Yang 0001, Jianping Fan 0001
MMM (2)4
2019 Two-Stream CNN with Loose Pair Training for Multi-modal AMD Categorization
Weisen Wang, Zhiyan Xu, Weihong Yu, Jianchun Zhao, Jingyuan Yang 0004, Zhikun Yang, Dayong Ding, Youxin Chen, Xirong Li 0001
MICCAI (1)4