Yalin Zheng

dblp:15/952 · DBLP profile ↗
← Back
77ranked-venue papers
3as first author
46since 2021 · last 2026
0000-0002-7873-0922ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 49 · 2 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 38 · 28 since 2021Artificial intelligence and machine learning · 22 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Multimodal biomarker AI techniques for early neurocognitive disorder diagnosis: A systematic review
Feliciana Catino, Fabio Castellana, Roberta Zupo, Viviana Giannoccaro, Luisa Lampignano, Angelo Michele Petrosillo, Francesco Addabbo, Giancarlo Sborgia, Giuseppe Colacicco, Carlo Santoro, Giovanni Boero, Donato Impedovo, Yalin Zheng, Rodolfo Sardone
Artif. Intell. Medicine13
2026 Predicting diabetic macular edema treatment responses using OCT: Dataset and methods of APTOS competition
abstract
• First challenge to focus on pre-treatment stratification for diabetic macular edema (DME): This study pioneers the use of pre-treatment OCT biomarkers to predict individual responses to anti-VEGF therapy, advancing the concept of personalized medicine in DME management. • Large-scale, publicly accessible OCT dataset: The competition provides one of the most comprehensive open-access DME datasets to date, comprising tens of thousands of OCT images from 2,000 patients, significantly addressing the field’s data scarcity. The dataset includes both per-eye and per-scan annotations for several critical retinal biomarkers, enabling a wide range of supervised learning applications. • Benchmark for future research: With 170 registered teams and 41 finalists, the challenge fostered broad engagement. The best team achieved an AUC of 80.06%, highlighting the feasibility of accurate outcome prediction. The challenge provides standardized evaluation metrics and a curated leaderboard, laying the groundwork for reproducible and comparable AI model development in ophthalmic treatment prediction. Diabetic macular edema (DME) significantly contributes to visual impairment in diabetic patients. Treatment responses to intravitreal therapies vary, highlighting the need for patient stratification to predict therapeutic benefits and enable personalized strategies. To our knowledge, this study is the first to explore pre-treatment stratification for predicting DME treatment responses. To advance this research, we organized the 2nd Asia-Pacific Tele-Ophthalmology Society (APTOS) Big Data Competition in 2021. The competition focused on improving predictive accuracy for anti-VEGF therapy responses using ophthalmic OCT images. We provided a dataset containing tens of thousands of OCT images from 2,000 patients with labels across four sub-tasks. This paper details the competition’s structure, dataset, leading methods, and evaluation metrics. The competition attracted strong scientific community participation, with 170 teams initially registering and 41 reaching the final round. The top-performing team achieved an AUC of 80.06%, highlighting the potential of AI in personalized DME treatment and clinical decision-making.
Weiyi Zhang 0004, Peranut Chotcomwongse, Yinwen Li, Pusheng Xu, Ruijie Yao, Lianhao Zhou, Qiping Zhou, Shoujin Huang, Zihao Jin, Florence H. T. Chung, Yalin Zheng, Mingguang He, Danli Shi, Paisan Ruamviboonsuk
Medical Image Anal.15
2026 Multi-Granularity Topological Reasoning for Anatomically Consistent Vasculature Parsing
abstract
Quantitative analysis of retinal vascular morphology is vital for clinical decision-making and the investigation of systemic diseases. Central to this process is the accurate segmentation of retinal arteries and veins (A/V) from the background, a task challenged by substantial variations in vessel calibers and the presence of low-contrast or ambiguous structures in fundus images, especially in ultra-wide field imaging where peripheral distortions and large-scale anatomical variability are pronounced. These factors often lead to fragmented semantic representations and topological inconsistencies in automated segmentation outputs. To address these limitations, we propose Ultra, a multi-granularity topological reasoning network designed for precise A/V segmentation. Ultra adopts a cascaded two-stage architecture: PriorNet generates coarse, multi-scale vascular priors that provide structural guidance, while RefineNet performs topology-aware segmentation refinement. To further enforce topological coherence, we propose the neighboring pixel connectivity regularization (NICER) layer, which selectively integrates local connectivity information predicted by the proposed connectivity prediction union (CPU) module. This connectivity is employed as auxiliary supervision through a pixel-wise local connectivity loss, reinforcing structural reasoning and promoting anatomically consistent vascular topology inference. Extensive experiments on ultra-wide field fundus imaging (UWF) datasets demonstrate that Ultra achieves state-of-the-art performance in A/V segmentation and topological preservation. Moreover, Ultra generalizes well to conventional color fundus photography (CFP) datasets, underscoring its robustness and broad applicability. Code is publicly available at: https://github.com/iMED-Lab/Ultra.
Lei Mou, Yonghuai Liu, Zhuoting Xu, Hao Zhang 0113, Yalin Zheng, Jiang Liu 0001, Huazhu Fu, Yitian Zhao
IEEE Trans. Image Process.5
2026 StealthMark: Harmless and Stealthy Ownership Verification for Medical Segmentation via Uncertainty-Guided Backdoors
abstract
Annotating medical data for training AI models is often costly and limited due to the shortage of specialists with relevant clinical expertise. This challenge is further compounded by privacy and ethical concerns associated with sensitive patient information. As a result, well-trained medical segmentation models on private datasets constitute valuable intellectual property requiring robust protection mechanisms. Existing model protection techniques primarily focus on classification and generative tasks, while segmentation models-crucial to medical image analysis-remain largely underexplored. In this paper, we propose a novel, stealthy, and harmless method, StealthMark, for verifying the ownership of medical segmentation models under closed-box conditions. Our approach subtly modulates model uncertainty without altering the final segmentation outputs, thereby preserving the model's performance. To enable ownership verification, we incorporate model-agnostic explanation methods, e.g. LIME, to extract feature attributions from the model outputs. Under specific triggering conditions, these explanations reveal a distinct and verifiable watermark. We further design the watermark as a QR code to facilitate robust and recognizable ownership claims. We conducted extensive experiments across four medical imaging datasets (CMR dataset from UK Biobank, the SEG fundus dataset, the EchoNet echocardiography dataset, and the PraNet colonoscopy dataset) and five mainstream segmentation models. The results demonstrate the effectiveness, stealthiness, and harmlessness of our method on the original model's segmentation performance. For example, when applied to the SAM model, StealthMark consistently achieved attack success rates (ASR) above 95% across various datasets while maintaining less than a 1% drop in Dice and AUC scores-significantly outperforming backdoor-based watermarking methods and highlighting its strong potential for practical deployment. Our implementation code is made available at https://github.com/Qinkaiyu/StealthMark.
Qinkai Yu, Chong Zhang 0006, Gaojie Jin, Tianjin Huang, Wei Zhou 0021, Xiao-Bo Jin, Bo Huang 0012, Yitian Zhao, Gregory Yoke Hong Lip, Yalin Zheng, Aline Villavicencio, Yanda Meng
IEEE Trans. Image Process.12
2026 Super-Resolution Reconstruction of OCTA via Multi-Field-of-View Representation Learning
abstract
High-resolution Optical Coherence Tomography Angiography (OCTA) images are essential for morphological analysis and biomarker measurement of the retinal vasculature. They can also provide underlying biomarkers for the accurate analysis of eye-related diseases. The trade-off between the high resolution (HR) and large scanning field-of-view (FOV) is a long-standing problem for OCTA image instrument. A large FOV image provides more retinal information with shorter acquisition time but often suffers from low resolution (LR), high scatter noise, and poor vascular contrast. In order to obtain HR OCTA images with larger FOV, we propose a novel self-similar dynamic domain adaptation network based on cross-field-of-view representation learning. The network enables LR images (i.e., $6\times \text{6}\,\text{mm}^{2}$) to learn HR image (i.e., $3\times \text{3}\,\text{mm}^{2}$) feature representations specialized for OCTA by constructing feature mapping relations for cross-field-of-view OCTA scans. To be specific, a multiple random degradation model is proposed on HR images to generate various synthetic LR images. Further, we propose a dynamic domain adaptation framework that prompts feature dynamic alignment of the LR image reconstruction results with those of synthetic LR images. Finally, a novel self-similar supervision loss is proposed to optimize the reconstruction results from LR to HR by exploiting the similarity between vessels in different regions. Experimental results on three OCTA datasets show that the proposed method surpasses existing state-of-the-art ones, significantly enhancing retinal structure segmentation and disease classification. Our OCTA dataset (the first dataset in this research area with paired $3\times 3$ and $6\times \text{6}\,\text{mm}^{2}$ OCTA images) and code are publicly available.
Huaying Hao, Shaoyi Leng, Yanda Meng, Yonghuai Liu, Yalin Zheng, Huazhu Fu, Jiong Zhang 0004, Quanyong Yi, Yue Liu 0005, Jingfeng Zhang, Yitian Zhao
IEEE J. Biomed. Health Informatics5
2026 Beyond Correlation: Causal Intervention for Multi-Label Medical Image Diagnosis
abstract
This paper addresses the challenge of multi-disease diagnosis by integrating causal reasoning into the diagnostic framework. In clinical practice, multiple conditions often co-occur, making multi-disease diagnosis more relevant than isolated single-disease cases. However, most deep learning methods focus on single-disease detection and fail to capture the complexity of diagnosing concurrent conditions. Even in multi-label settings, existing approaches mainly rely on correlation-based inference, capturing statistical associations rather than true causal relationships. This can lead to spurious feature-disease associations, where features linked to one disease are mistakenly attributed to another due to frequent co-occurrence, ultimately undermines diagnostic accuracy and interpretability. To address this challenge, we propose a novel framework that incorporates causal intervention into multi-label medical image diagnosis, enabling the model to identify true causal signals rather than misleading correlations arising from co-occurring diseases. Specifically, we model latent disease-related confounders and apply backdoor adjustment to disentangle genuine causal effects from spurious associations. This is achieved by implicitly learning shared feature representations that serve as confounding variables, which are then used to refine image-derived features during prediction. The resulting causal adjustment allows the model to focus on disease-specific cues, improving accuracy and interpretability. Extensive experiments on four diverse medical imaging datasets: ODIR (color fundus photography), LID-FFA (fundus fluorescein angiography), Endo (colonoscopy), and Chestpert (X-ray) demonstrate that our method consistently outperforms existing approaches. Furthermore, our model also effectively separates the diagnosis of co-occurring diseases, highlighting the potential of causal reasoning to enhance the reliability and clinical applicability of AI-assisted diagnosis. The source code is publicly available at https://github.com/davelailai/BankCausal.git.
Jianyang Xie, Yitian Zhao, Xiuju Chen, Yanda Meng, He Zhao 0002, Uazman Alam, Yalin Zheng
IEEE Trans. Medical Imaging8
2025 Incomplete Modality Disentangled Representation for Ophthalmic Disease Grading and Diagnosis
abstract
Ophthalmologists typically require multimodal data sources to improve diagnostic accuracy in clinical decisions. However, due to medical device shortages, low-quality data and data privacy concerns, missing data modalities are common in real-world scenarios. Existing deep learning methods tend to address it by learning an implicit latent subspace representation for different modality combinations. We identify two significant limitations of these methods: (1) implicit representation constraints that hinder the model's ability to capture modality-specific information and (2) modality heterogeneity, causing distribution gaps and redundancy in feature representations. To address these, we propose an Incomplete Modality Disentangled Representation (IMDR) strategy, which disentangles features into explicit independent modal-common and modal-specific features by guidance of mutual information, distilling informative knowledge and enabling it to reconstruct valuable missing semantics and produce robust multimodal representations. Furthermore, we introduce a joint proxy learning module that assists IMDR in eliminating intra-modality redundancy by exploiting the extracted proxies from each class. Experiments on four ophthalmology multimodal datasets demonstrate that the proposed IMDR outperforms the state-of-the-art methods significantly.
Zile Huang, Zhongxing Xu, Zihong Luo, Yalin Zheng, Yanda Meng
AAAI8
2025 PathVLG: A Vision-Language Framework for Domain Generalization in Cross-Organ Adenocarcinoma Segmentation
abstract
Domain shifts in computational pathology, caused by variations in staining, imaging devices, and tissue morphology, challenge model performance in segmentation tasks. Existing domain generalization methods, such as style transfer and feature alignment, often fail to account for organ-level morphological differences. In this paper, we propose PathVLG, a visionlanguage model designed to improve domain generalization for adenocarcinoma segmentation. PathVLG leverages a CONCHbased encoder with three key innovations: the Text-informed Content Query Reformer (TCQR), Text-driven Style Augmentor (TSA), and Style Regeneration Decoder (SRD). These components help the model adapt across domains by incorporating text embeddings, generating diverse styles, and combining source and target domain features. Experimental results show that PathVLG outperforms existing methods in cross-domain generalization.
Biwen Meng, Wanrong Yang, Kang Dang, Yalin Zheng, Jingxin Liu 0005
BIBM6
2025 Are Spatial-Temporal Graph Convolution Networks for Human Action Recognition Over-Parameterized?
abstract
Spatial-temporal graph convolutional networks (ST-GCNs) showcase impressive performance in skeleton-based human action recognition (HAR). However, despite the development of numerous models, their recognition performance does not differ significantly after aligning the input settings. With this observation, we hypothesize that ST-GCNs are over-parameterized for HAR, a conjecture subsequently confirmed through experiments employing the lottery ticket hypothesis. Additionally, a novel sparse ST-GCNs generator is proposed, which trains a sparse architecture from a randomly initialized dense network while maintaining comparable performance levels to the dense components. Moreover, we generate multi-level sparsity ST-GCNs by integrating sparse structures at various sparsity levels and demonstrate that the assembled model yields a significant enhancement in HAR performance. Thorough experiments on four datasets, including NTU-RGB+D 60(120), Kinetics-400, and FineGYM, demonstrate that the proposed sparse ST-GCNs can achieve comparable performance to their dense components. Even with 95% fewer parameters, the sparse ST-GCNs exhibit a degradation of1% in top-1 accuracy. The code is available at https://github.com/davelailai/Sparse-ST-GCN.
Jianyang Xie, Yitian Zhao, Yanda Meng, He Zhao 0002, Anh Nguyen 0003, Yalin Zheng
CVPR6
2025 tHPM-LDM: Integrating Individual Historical Record with Population Memory in Latent Diffusion-Based Glaucoma Forecasting
Jianyang Xie, Yimin Luo, Yanda Meng, Savita Madhusudhan, Gregory Yoke Hong Lip, Li Cheng 0001, Yalin Zheng, He Zhao 0002
MICCAI (1)8
2025 MorphoBoost: Morphology-Driven Boundary Enhancement Model for Accurate Segmentation of Langerhans Cells in Corneal Confocal Microscopy Images
Hongshuo Li, Ankai Dong, Tiande Zhang, Shijia Zhou, Yalin Zheng, Lei Mou, Yitian Zhao
MICCAI (13)5
2025 A Frequency-Aware Self-supervised Learning for Ultra-Wide-Field Image Enhancement
Weicheng Liao, Jianyang Xie, Yalin Zheng, Yuhui Ma, Yitian Zhao
MICCAI (13)4
2025 Fairness-Aware vCDR-Controlled Generation for Glaucoma Diagnosis
Shuran Yang, Feixiang Zhou, Meng Wang 0038, Yitian Zhao, Yalin Zheng, Yanda Meng
MICCAI (9)10
2025 Robust Incomplete-Modality Alignment for Ophthalmic Disease Grading and Diagnosis via Labeled Optimal Transport
Qinkai Yu, Jianyang Xie, Yitian Zhao, Cheng Chen 0013, Jun Cheng 0003, Lu Liu 0001, Yalin Zheng, Yanda Meng
MICCAI (15)9
2025 Parameterized Diffusion Optimization Enabled Autoregressive Ordinal Regression for Diabetic Retinopathy Grading
Qinkai Yu, Wei Zhou 0021, Hantao Liu, Yanyu Xu 0001, Meng Wang 0038, Yitian Zhao, Huazhu Fu, Xujiong Ye, Yalin Zheng, Yanda Meng
MICCAI (15)9
2025 GLCP: Global-to-Local Connectivity Preservation for Tubular Structure Segmentation
Feixiang Zhou, Zhuangzhi Gao, He Zhao 0002, Jianyang Xie, Yanda Meng, Yitian Zhao, Gregory Yoke Hong Lip, Yalin Zheng
MICCAI (16)8
2025 3D microvascular reconstruction in retinal OCT angiography images via domain-adaptive learning
Jiong Zhang 0004, Yonghuai Liu, Dan Zhang 0026, Jianyang Xie, Tao Chen 0003, Yalin Zheng, Huazhu Fu, Yitian Zhao
Pattern Recognit.7
2025 Randomness-Restricted Diffusion Model for Ocular Surface Structure Segmentation
abstract
Ocular surface diseases affect a significant portion of the population worldwide. Accurate segmentation and quantification of different ocular surface structures are crucial for the understanding of these diseases and clinical decision-making. However, the automated segmentation of the ocular surface structure is relatively unexplored and faces several challenges. Ocular surface structure boundaries are often inconspicuous and obscured by glare from reflections. In addition, the segmentation of different ocular structures always requires training of multiple individual models. Thus, developing a one-model-fits-all segmentation approach is desirable. In this paper, we introduce a randomness-restricted diffusion model for multiple ocular surface structure segmentation. First, a time-controlled fusion-attention module (TFM) is proposed to dynamically adjust the information flow within the diffusion model, based on the temporal relationships between the network's input and time. TFM enables the network to effectively utilize image features to constrain the randomness of the generation process. We further propose a low-frequency consistency filter and a new loss to alleviate model uncertainty and error accumulation caused by the multi-step denoising process. Extensive experiments have shown that our approach can segment seven different ocular surface structures. Our method performs better than both dedicated ocular surface segmentation methods and general medical image segmentation methods. We further validated the proposed method over two clinical datasets, and the results demonstrated that it is beneficial to clinical applications, such as the meibomian gland dysfunction grading and aqueous deficient dry eye diagnosis.
Huaying Hao, Yifan Zhao 0001, Yanda Meng, Jiang Liu 0001, Yalin Zheng, Wei Chen 0089, Yitian Zhao
IEEE Trans. Medical Imaging7
2024 Dynamic Semantic-Based Spatial Graph Convolution Network for Skeleton-Based Human Action Recognition
abstract
Graph convolutional networks (GCNs) have attracted great attention and achieved remarkable performance in skeleton-based action recognition. However, most of the previous works are designed to refine skeleton topology without considering the types of different joints and edges, making them infeasible to represent the semantic information. In this paper, we proposed a dynamic semantic-based graph convolution network (DS-GCN) for skeleton-based human action recognition, where the joints and edge types were encoded in the skeleton topology in an implicit way. Specifically, two semantic modules, the joints type-aware adaptive topology and the edge type-aware adaptive topology, were proposed. Combining proposed semantics modules with temporal convolution, a powerful framework named DS-GCN was developed for skeleton-based action recognition. Extensive experiments in two datasets, NTU-RGB+D and Kinetics-400 show that the proposed semantic modules were generalized enough to be utilized in various backbones for boosting recognition accuracy. Meanwhile, the proposed DS-GCN notably outperformed state-of-the-art methods. The code is released here https://github.com/davelailai/DS-GCN
Jianyang Xie, Yanda Meng, Yitian Zhao, Anh Nguyen 0003, Xiaoyun Yang, Yalin Zheng
AAAI6
2024 Giraffe: A Genetic Programming Algorithm To Build Deep Learning Ensembles For Ecg Arrhythmia Classification
abstract
Cardiovascular diseases remain one of the leading causes of death worldwide. Therefore, developing and validating automated tools to help identify high-risk patients are of paramount clinical utility. In this article, we tackle this task and introduce a genetic programming algorithm (called GIRAFFE) to build (deep) machine learning classification ensembles for arrhythmia classification from two-dimensional images of 12-lead electrocardiogram (ECG) tracings. GIRAFFE evolves the architecture, content, and fusion scheme of the ensemble, to obtain an accurate yet lightweight classification system. The experimental study performed over a large-scale dataset of ECG images revealed that our approach outperforms other ensemble methods and carefully fine-tuned deep models, elaborates compact heterogeneous ensembles, and does not require any user intervention hence it is easy to apply to other classification tasks.
Damian Kucharski, Agata M. Wijata, Lu Fu, Yumei Xue, Jacek Kawa, Yalin Zheng, Gregory Yoke Hong Lip, Jakub Nalepa
ICIP7
2024 A Hyperreflective Foci Segmentation Network for OCT Images with Multi-dimensional Semantic Enhancement
Xingguo Wang, Yuhui Ma, Yalin Zheng, Jiong Zhang 0004, Yonghuai Liu, Yitian Zhao
MICCAI (1)4
2024 Multi-disease Detection in Retinal Images Guided by Disease Causal Estimation
Jianyang Xie, Xiuju Chen, Yitian Zhao, Yanda Meng, He Zhao 0002, Anh Nguyen 0003, Yalin Zheng
MICCAI (1)8
2024 CLIP-DR: Textual Knowledge-Guided Diabetic Retinopathy Grading with Ranking-Aware Prompting
Qinkai Yu, Jianyang Xie, Anh Nguyen 0003, He Zhao 0002, Jiong Zhang 0004, Huazhu Fu, Yitian Zhao, Yalin Zheng, Yanda Meng
MICCAI (1)8
2024 Multi-granularity learning of explicit geometric constraint and contrast for label-efficient medical image segmentation and differentiable clinical function assessment
abstract
Automated segmentation is a challenging task in medical image analysis that usually requires a large amount of manually labeled data. However, most current supervised learning based algorithms suffer from insufficient manual annotations, posing a significant difficulty for accurate and robust segmentation. In addition, most current semi-supervised methods lack explicit representations of geometric structure and semantic information, restricting segmentation accuracy. In this work, we propose a hybrid framework to learn polygon vertices, region masks, and their boundaries in a weakly/semi-supervised manner that significantly advances geometric and semantic representations. Firstly, we propose multi-granularity learning of explicit geometric structure constraints via polygon vertices (PolyV) and pixel-wise region (PixelR) segmentation masks in a semi-supervised manner. Secondly, we propose eliminating boundary ambiguity by using an explicit contrastive objective to learn a discriminative feature space of boundary contours at the pixel level with limited annotations. Thirdly, we exploit the task-specific clinical domain knowledge to differentiate the clinical function assessment end-to-end. The ground truth of clinical function assessment, on the other hand, can serve as auxiliary weak supervision for PolyV and PixelR learning. We evaluate the proposed framework on two tasks, including optic disc (OD) and cup (OC) segmentation along with vertical cup-to-disc ratio (vCDR) estimation in fundus images; left ventricle (LV) segmentation at end-diastolic and end-systolic frames along with ejection fraction (LVEF) estimation in two-dimensional echocardiography images. Experiments on nine large-scale datasets of the two tasks under different label settings demonstrate our model’s superior performance on segmentation and clinical function assessment.
Yanda Meng, Jianyang Xie, Jinming Duan 0001, Martha Joddrell, Savita Madhusudhan, Tunde Peto, Yitian Zhao, Yalin Zheng
Medical Image Anal.9
2024 Dynamic Semantic-Based Spatial-Temporal Graph Convolution Network for Skeleton-Based Human Action Recognition
abstract
Human action recognition is an essential topic in computer vision and image processing. Graph convolutional networks (GCNs) have attracted significant attention and achieved noteworthy performance in skeleton-based human action recognition tasks. However, most of the previous graph-based works are designed to refine skeleton topology without considering the types of different joints and edges and the occurrence order of the frames. Such a limitation makes them insufficient to represent intrinsic semantic information. Differently, we proposed a dynamic semantic-based spatial-temporal graph convolution network (DS-STGCN) to address the challenge. DS-STGCN has two dynamic semantic modules for spatial and temporal contexts respectively. Specifically, the joints and edge types were encoded in the spatial module implicitly, and the occurrence order of frames was encoded in the temporal module implicitly. Extensive experiments on four datasets including NTU-RGB+D 60(120), Kinetics-400, and FineGYM show that our proposed two semantic modules can bring consistent recognition performance improvement with various backbones. Meanwhile, the proposed DS-STGCN notably surpassed state-of-the-art methods on these datasets. Notably, in the more challenging dataset, such as Kinetics-400, our model significantly outperformed other state-of-the-art GCN-based methods by a large margin. The code has been released at https://github.com/davelailai/DS-STGCN.
Jianyang Xie, Yanda Meng, Yitian Zhao, Anh Nguyen 0003, Xiaoyun Yang, Yalin Zheng
IEEE Trans. Image Process.6
2023 Weakly Supervised Segmentation with Point Annotations for Histopathology Images via Contrast-Based Variational Model
abstract
Image segmentation is a fundamental task in the field of imaging and vision. Supervised deep learning for segmentation has achieved unparalleled success when sufficient training data with annotated labels are available. However, annotation is known to be expensive to obtain, especially for histopathology images where the target regions are usually with high morphology variations and irregular shapes. Thus, weakly supervised learning with sparse annotations of points is promising to reduce the annotation workload. In this work, we propose a contrast-based variational model to generate segmentation results, which serve as reliable complementary supervision to train a deep segmentation model for histopathology images. The proposed method considers the common characteristics of target regions in histopathology images and can be trained in an end-to-end manner. It can generate more regionally consistent and smoother boundary segmentation, and is more robust to unlabeled ‘novel’ regions. Experiments on two different histology datasets demonstrate its effectiveness and efficiency in comparison to previous models. Code is available at: https://github.com/hrzhang1123/CVM_WS_Segmentation.
Hongrun Zhang, Liam Burrows, Yanda Meng, Declan Sculthorpe, Abhik Mukherjee, Sarah E. Coupland, Ke Chen 0002, Yalin Zheng
CVPR8
2023 Polar-Net: A Clinical-Friendly Model for Alzheimer's Disease Detection in OCTA Images
Shouyue Liu, Jinkui Hao, Yanwu Xu 0001, Huazhu Fu, Jiang Liu 0001, Yalin Zheng, Yonghuai Liu, Jiong Zhang 0004, Yitian Zhao
MICCAI (7)7
2023 Weakly/Semi-supervised Left Ventricle Segmentation in 2D Echocardiography with Uncertain Region-Aware Contrastive Learning
Yanda Meng, Jianyang Xie, Jinming Duan 0001, Yitian Zhao, Yalin Zheng
PRCV (13)6
2023 Bilateral adaptive graph convolutional network on CT based Covid-19 diagnosis with uncertainty-aware consensus-assisted multiple instance learning
abstract
Coronavirus disease (COVID-19) has caused a worldwide pandemic, putting millions of people's health and lives in jeopardy. Detecting infected patients early on chest computed tomography (CT) is critical in combating COVID-19. Harnessing uncertainty-aware consensus-assisted multiple instance learning (UC-MIL), we propose to diagnose COVID-19 using a new bilateral adaptive graph-based (BA-GCN) model that can use both 2D and 3D discriminative information in 3D CT volumes with arbitrary number of slices. Given the importance of lung segmentation for this task, we have created the largest manual annotation dataset so far with 7,768 slices from COVID-19 patients, and have used it to train a 2D segmentation model to segment the lungs from individual slices and mask the lungs as the regions of interest for the subsequent analyses. We then used the UC-MIL model to estimate the uncertainty of each prediction and the consensus between multiple predictions on each CT slice to automatically select a fixed number of CT slices with reliable predictions for the subsequent model reasoning. Finally, we adaptively constructed a BA-GCN with vertices from different granularity levels (2D and 3D) to aggregate multi-level features for the final diagnosis with the benefits of the graph convolution network's superiority to tackle cross-granularity relationships. Experimental results on three largest COVID-19 CT datasets demonstrated that our model can produce reliable and accurate COVID-19 predictions using CT volumes with any number of slices, which outperforms existing approaches in terms of learning and generalisation ability. To promote reproducible research, we have made the datasets, including the manual annotations and cleaned CT dataset, as well as the implementation code, available at https://doi.org/10.5281/zenodo.6361963.
Yanda Meng, Joshua Bridge, Cliff Addison, Manhui Wang, Cristin Merritt, Stu Franks, Maria Mackey, Steve Messenger, Renrong Sun, Thomas Fitzmaurice, Caroline McCann, Yitian Zhao, Yalin Zheng
Medical Image Anal.14
2023 Transportation Object Counting With Graph-Based Adaptive Auxiliary Learning
abstract
This paper proposes an adaptive auxiliary task learning-based approach for transport object counting problems such as humans and vehicles. These problems are essential in many real-world tasks such as video surveillance, traffic monitoring, public security, and urban planning, to aid intelligent transportation systems. Unlike existing auxiliary task learning-based methods, we develop an attention-enhanced adaptively shared backbone network to enable both task-shared and task-tailored features that are learned in an end-to-end manner. The network seamlessly combines a standard Convolution Neural Network (CNN) and a Graph Convolution Network (GCN) for feature extraction and feature reasoning among different domains of tasks. Our approach gains enriched contextual information by iteratively and hierarchically fusing features across different task branches of the adaptive CNN backbone. The whole framework pays special attention to objects’ spatial locations and varied density levels, informed by object (or crowd) segmentation and density level segmentation auxiliary tasks. In particular, thanks to the proposed dilated contrastive density loss function, our network benefits from individual and regional context supervision, along with strengthened robustness. Experiments on six challenging multi-domain datasets demonstrate that our method achieves superior performance compared with state-of-the-art auxiliary task learning-based counting methods. Our code is publicly available.
Yanda Meng, Joshua Bridge, Yitian Zhao, Martha Joddrell, Yihong Qiao, Xiaoyun Yang, Xiaowei Huang 0001, Yalin Zheng
IEEE Trans. Intell. Transp. Syst.8
2023 Dual Consistency Enabled Weakly and Semi-Supervised Optic Disc and Cup Segmentation With Dual Adaptive Graph Convolutional Networks
abstract
Glaucoma is a progressive eye disease that results in permanent vision loss, and the vertical cup to disc ratio (vCDR) in colour fundus images is essential in glaucoma screening and assessment. Previous fully supervised convolution neural networks segment the optic disc (OD) and optic cup (OC) from color fundus images and then calculate the vCDR offline. However, they rely on a large set of labeled masks for training, which is expensive and time-consuming to acquire. To address this, we propose a weakly and semi-supervised graph-based network that investigates geometric associations and domain knowledge between segmentation probability maps (PM), modified signed distance function representations (mSDF), and boundary region of interest characteristics (B-ROI) in three aspects. Firstly, we propose a novel Dual Adaptive Graph Convolutional Network (DAGCN) to reason the long-range features of the PM and the mSDF w.r.t. the regional uniformity. Secondly, we propose a dual consistency regularization-based semi-supervised learning paradigm. The regional consistency between the PM and the mSDF, and the marginal consistency between the derived B-ROI from each of them boost the proposed model's performance due to the inherent geometric associations. Thirdly, we exploit the task-specific domain knowledge via the oval shapes of OD & OC, where a differentiable vCDR estimating layer is proposed. Furthermore, without additional annotations, the supervision on vCDR serves as weakly-supervisions for segmentation tasks. Experiments on six large-scale datasets demonstrate our model's superior performance on OD & OC segmentation and vCDR estimation. The implementation code has been made available.https://github.com/smallmax00/Dual_Adaptive_Graph_Reasoning.
Yanda Meng, Hongrun Zhang, Yitian Zhao, Dongxu Gao, Barbra Hamill, Godhuli Patri, Tunde Peto, Savita Madhusudhan, Yalin Zheng
IEEE Trans. Medical Imaging9
2023 3D Human Pose and Shape Reconstruction From Videos via Confidence-Aware Temporal Feature Aggregation
abstract
Estimating 3D human body shapes and poses from videos is a challenging computer vision task. The intrinsic temporal information embedded in adjacent frames is helpful in making accurate estimations. Existing approaches learn temporal features of the target frames simply by aggregating features of their adjacent frames, using off-the-shelf deep neural networks. Consequently these approaches cannot explicitly and effectively use the correlations between adjacent frames to help infer the parameters of the target frames. In this paper, we propose a novel framework that can measure the correlations amongst adjacent frames in the form of an estimated confidence metric. The confidence value will indicate to what extent the adjacent frames can help predict the target frames’ 3D shapes and poses. Based on the estimated confidence values, temporally aggregated features are then obtained by adaptively allocating different weights to the temporal predicted features from the adjacent frames. The final 3D shapes and poses are estimated by regressing from the temporally aggregated features. Experimental results on three benchmark datasets show that the proposed method outperforms state-of-the-art approaches (even without the motion priors involved in training). In particular, the proposed method is more robust against corrupted frames.
Hongrun Zhang, Yanda Meng, Yitian Zhao, Xuesheng Qian, Yihong Qiao, Xiaoyun Yang, Yalin Zheng
IEEE Trans. Multim.7
2022 DTFD-MIL: Double-Tier Feature Distillation Multiple Instance Learning for Histopathology Whole Slide Image Classification
abstract
Multiple instance learning (MIL) has been increasingly used in the classification of histopathology whole slide images (WSIs). However, MIL approaches for this specific classification problem still face unique challenges, particularly those related to small sample cohorts. In these, there are limited number of WSI slides (bags), while the resolution of a single WSI is huge, which leads to a large number of patches (instances) cropped from this slide. To address this issue, we propose to virtually enlarge the number of bags by introducing the concept of pseudo-bags, on which a double-tier MIL framework is built to effectively use the intrinsic features. Besides, we also contribute to deriving the instance probability under the framework of attentionbased MIL, and utilize the derivation to help construct and analyze the proposed framework. The proposed method outperforms other latest methods on the CAMELYON-16 by substantially large margins, and is also better in performance on the TCGA lung cancer dataset. The proposed framework is ready to be extended for wider MIL applications. The code is available at: https://github. com/hrzhang1123/DTFD-MIL.
Hongrun Zhang, Yanda Meng, Yitian Zhao, Yihong Qiao, Xiaoyun Yang, Sarah E. Coupland, Yalin Zheng
CVPR7
2022 Electrocardiogram Two-Dimensional Motifs: A Study Directed at Cardio Vascular Disease Classification
Hanadi Aldosari, Frans Coenen, Gregory Yoke Hong Lip, Yalin Zheng
IC3K4
2022 NerveFormer: A Cross-Sample Aggregation Network for Corneal Nerve Segmentation
Lei Mou, Shaodong Ma, Huazhu Fu, Lijun Guo, Yalin Zheng, Jiong Zhang 0004, Yitian Zhao
MICCAI (4)6
2022 Shape-Aware Weakly/Semi-Supervised Optic Disc and Cup Segmentation with Regional/Marginal Consistency
Yanda Meng, Xu Chen 0030, Hongrun Zhang, Yitian Zhao, Dongxu Gao, Barbra Hamill, Godhuli Patri, Tunde Peto, Savita Madhusudhan, Yalin Zheng
MICCAI (4)10
2022 Machine learning-based predictions of dietary restriction associations across ageing-related genes
abstract
BACKGROUND: Dietary restriction (DR) is the most studied pro-longevity intervention; however, a complete understanding of its underlying mechanisms remains elusive, and new research directions may emerge from the identification of novel DR-related genes and DR-related genetic features. RESULTS: This work used a Machine Learning (ML) approach to classify ageing-related genes as DR-related or NotDR-related using 9 different types of predictive features: PathDIP pathways, two types of features based on KEGG pathways, two types of Protein-Protein Interactions (PPI) features, Gene Ontology (GO) terms, Genotype Tissue Expression (GTEx) expression features, GeneFriends co-expression features and protein sequence descriptors. Our findings suggested that features biased towards curated knowledge (i.e. GO terms and biological pathways), had the greatest predictive power, while unbiased features (mainly gene expression and co-expression data) have the least predictive power. Moreover, a combination of all the feature types diminished the predictive power compared to predictions based on curated knowledge. Feature importance analysis on the two most predictive classifiers mostly corroborated existing knowledge and supported recent findings linking DR to the Nuclear Factor Erythroid 2-Related Factor 2 (NRF2) signalling pathway and G protein-coupled receptors (GPCR). We then used the two strongest combinations of feature type and ML algorithm to predict DR-relatedness among ageing-related genes currently lacking DR-related annotations in the data, resulting in a set of promising candidate DR-related genes (GOT2, GOT1, TSC1, CTH, GCLM, IRS2 and SESN2) whose predicted DR-relatedness remain to be validated in future wet-lab experiments. CONCLUSIONS: This work demonstrated the strong potential of ML-based techniques to identify DR-associated features as our findings are consistent with literature and recent discoveries. Although the inference of new DR-related mechanistic findings based solely on GO terms and biological pathways was limited due to their knowledge-driven nature, the predictive power of these two features types remained useful as it allowed inferring new promising candidate DR-related genes.
Gustavo Daniel Vega Magdaleno, Vladislav Bespalov, Yalin Zheng, Alex Alves Freitas, João Pedro de Magalhães
BMC Bioinform.3
2022 Graph-Based Region and Boundary Aggregation for Biomedical Image Segmentation
abstract
Segmentation is a fundamental task in biomedical image analysis. Unlike the existing region-based dense pixel classification methods or boundary-based polygon regression methods, we build a novel graph neural network (GNN) based deep learning framework with multiple graph reasoning modules to explicitly leverage both region and boundary features in an end-to-end manner. The mechanism extracts discriminative region and boundary features, referred to as initialized region and boundary node embeddings, using a proposed Attention Enhancement Module (AEM). The weighted links between cross-domain nodes (region and boundary feature domains) in each graph are defined in a data-dependent way, which retains both global and local cross-node relationships. The iterative message aggregation and node update mechanism can enhance the interaction between each graph reasoning module's global semantic information and local spatial characteristics. Our model, in particular, is capable of concurrently addressing region and boundary feature reasoning and aggregation at several different feature levels due to the proposed multi-level feature node embeddings in different parallel graph reasoning modules. Experiments on two types of challenging datasets demonstrate that our method outperforms state-of-the-art approaches for segmentation of polyps in colonoscopy images and of the optic disc and optic cup in colour fundus images. The trained models will be made available at: https://github.com/smallmax00/Graph_Region_Boudnary.
Yanda Meng, Hongrun Zhang, Yitian Zhao, Xiaoyun Yang, Yihong Qiao, Ian J. C. MacCormick, Xiaowei Huang 0001, Yalin Zheng
IEEE Trans. Medical Imaging8
2022 DeepGrading: Deep Learning Grading of Corneal Nerve Tortuosity
abstract
Accurate estimation and quantification of the corneal nerve fiber tortuosity in corneal confocal microscopy (CCM) is of great importance for disease understanding and clinical decision-making. However, the grading of corneal nerve tortuosity remains a great challenge due to the lack of agreements on the definition and quantification of tortuosity. In this paper, we propose a fully automated deep learning method that performs image-level tortuosity grading of corneal nerves, which is based on CCM images and segmented corneal nerves to further improve the grading accuracy with interpretability principles. The proposed method consists of two stages: 1) A pre-trained feature extraction backbone over ImageNet is fine-tuned with a proposed novel bilinear attention (BA) module for the prediction of the regions of interest (ROIs) and coarse grading of the image. The BA module enhances the ability of the network to model long-range dependencies and global contexts of nerve fibers by capturing second-order statistics of high-level features. 2) An auxiliary tortuosity grading network (AuxNet) is proposed to obtain an auxiliary grading over the identified ROIs, enabling the coarse and additional gradings to be finally fused together for more accurate final results. The experimental results show that our method surpasses existing methods in tortuosity grading, and achieves an overall accuracy of 85.64% in four-level classification. We also validate it over a clinical dataset, and the statistical analysis demonstrates a significant difference of tortuosity levels between healthy control and diabetes group. We have released a dataset with 1500 CCM images and their manual annotations of four tortuosity levels for public access. The code is available at: https://github.com/iMED-Lab/TortuosityGrading.
Lei Mou, Yonghuai Liu, Yalin Zheng, Peter Matthew, Pan Su 0001, Jiang Liu 0001, Jiong Zhang 0004, Yitian Zhao
IEEE Trans. Medical Imaging4
2021 BI-GCN: Boundary-Aware Input-Dependent Graph Convolution Network for Biomedical Image Segmentation
Yanda Meng, Hongrun Zhang, Dongxu Gao, Yitian Zhao, Xiaoyun Yang, Xuesheng Qian, Xiaowei Huang 0001, Yalin Zheng
BMVC8
2021 Motif Based Feature Vectors: Towards a Homogeneous Data Representation for Cardiovascular Diseases Classification
Hanadi Aldosari, Frans Coenen, Gregory Yoke Hong Lip, Yalin Zheng
DaWaK4
2021 Spatial Uncertainty-Aware Semi-Supervised Crowd Counting
abstract
Semi-supervised approaches for crowd counting attract attention, as the fully supervised paradigm is expensive and laborious due to its request for a large number of images of dense crowd scenarios and their annotations. This paper proposes a spatial uncertainty-aware semi-supervised approach via regularized surrogate task (binary segmentation) for crowd counting problems. Different from existing semi-supervised learning-based crowd counting methods, to exploit the unlabeled data, our proposed spatial uncertainty-aware teacher-student framework focuses on high confident regions’ information while addressing the noisy supervision from the unlabeled data in an end-to-end manner. Specifically, we estimate the spatial uncertainty maps from the teacher model’s surrogate task to guide the feature learning of the main task (density regression) and the surrogate task of the student model at the same time. Besides, we introduce a simple yet effective differential transformation layer to enforce the inherent spatial consistency regularization between the main task and the surrogate task in the student model, which helps the surrogate task to yield more reliable predictions and generates high-quality uncertainty maps. Thus, our model can also address the task-level perturbation problems that occur spatial inconsistency between the primary and surrogate tasks in the student model. Experimental results on four challenging crowd counting datasets demonstrate that our method achieves superior performance to the state-of-the-art semi-supervised methods. Code is available at : https://github.com/smallmax00/SUA_crowd_counting
Yanda Meng, Hongrun Zhang, Yitian Zhao, Xiaoyun Yang, Xuesheng Qian, Xiaowei Huang 0001, Yalin Zheng
ICCV7
2021 Learning Unsupervised Parameter-Specific Affine Transformation for Medical Images Registration
Xu Chen 0030, Yanda Meng, Yitian Zhao, Rachel Williams, Srinivasa R. Vallabhaneni, Yalin Zheng
MICCAI (4)6
2021 Cross-Domain Depth Estimation Network for 3D Vessel Reconstruction in OCT Angiography
Yonghuai Liu, Jiong Zhang 0004, Jianyang Xie, Yalin Zheng, Jiang Liu 0001, Yitian Zhao
MICCAI (8)5
2021 CS2-Net: Deep learning segmentation of curvilinear structures in medical imaging
Lei Mou, Yitian Zhao, Huazhu Fu, Yonghuai Liu, Jun Cheng 0003, Yalin Zheng, Pan Su 0001, Jianlong Yang, Li Chen 0011, Alejandro F. Frangi, Masahiro Akiba, Jiang Liu 0001
Medical Image Anal.6
2021 ROSE: A Retinal OCT-Angiography Vessel Segmentation Dataset and New Model
abstract
Optical Coherence Tomography Angiography (OCTA) is a non-invasive imaging technique that has been increasingly used to image the retinal vasculature at capillary level resolution. However, automated segmentation of retinal vessels in OCTA has been under-studied due to various challenges such as low capillary visibility and high vessel complexity, despite its significance in understanding many vision-related diseases. In addition, there is no publicly available OCTA dataset with manually graded vessels for training and validation of segmentation algorithms. To address these issues, for the first time in the field of retinal image analysis we construct a dedicated Retinal OCTA SEgmentation dataset (ROSE), which consists of 229 OCTA images with vessel annotations at either centerline-level or pixel level. This dataset with the source code has been released for public access to assist researchers in the community in undertaking research in related topics. Secondly, we introduce a novel split-based coarse-to-fine vessel segmentation network for OCTA images (OCTA-Net), with the ability to detect thick and thin vessels separately. In the OCTA-Net, a split-based coarse segmentation module is first utilized to produce a preliminary confidence map of vessels, and a split-based refined segmentation module is then used to optimize the shape/contour of the retinal microvasculature. We perform a thorough evaluation of the state-of-the-art vessel segmentation models and our OCTA-Net on the constructed ROSE dataset. The experimental results demonstrate that our OCTA-Net yields better vessel segmentation performance in OCTA than both traditional and other deep learning methods. In addition, we provide a fractal dimension analysis on the segmented microvasculature, and the statistical analysis demonstrates significant differences between the healthy control and Alzheimer's Disease group. This consolidates that the analysis of retinal microvasculature may offer a new scheme to study various neurodegenerative diseases.
Yuhui Ma, Huaying Hao, Jianyang Xie, Huazhu Fu, Jiong Zhang 0004, Jianlong Yang, Jiang Liu 0001, Yalin Zheng, Yitian Zhao
IEEE Trans. Medical Imaging9
2020 Regression of Instance Boundary by Aggregated CNN and GCN
Yanda Meng, Dongxu Gao, Yitian Zhao, Xiaoyun Yang, Xiaowei Huang 0001, Yalin Zheng
ECCV (8)7
2020 Cycle Structure and Illumination Constrained GAN for Medical Image Enhancement
Yuhui Ma, Yonghuai Liu, Jun Cheng 0003, Yalin Zheng, Morteza Ghahremani, Honghan Chen, Jiang Liu 0001, Yitian Zhao
MICCAI (2)4
2020 CNN-GCN Aggregation Enabled Boundary Regression for Biomedical Image Segmentation
Yanda Meng, Dongxu Gao, Yitian Zhao, Xiaoyun Yang, Xiaowei Huang 0001, Yalin Zheng
MICCAI (4)7
2020 Classification of Retinal Vessels into Artery-Vein in OCT Angiography Guided by Fundus Images
Jianyang Xie, Yonghuai Liu, Yalin Zheng, Pan Su 0001, Jian Yang 0009, Jiang Liu 0001, Yitian Zhao
MICCAI (6)3
2020 Introducing the GEV Activation Function for Highly Unbalanced Data to Develop COVID-19 Diagnostic Models
abstract
Fast and accurate diagnosis is essential for the efficient and effective control of the COVID-19 pandemic that is currently disrupting the whole world. Despite the prevalence of the COVID-19 outbreak, relatively few diagnostic images are openly available to develop automatic diagnosis algorithms. Traditional deep learning methods often struggle when data is highly unbalanced with many cases in one class and only a few cases in another; new methods must be developed to overcome this challenge. We propose a novel activation function based on the generalized extreme value (GEV) distribution from extreme value theory, which improves performance over the traditional sigmoid activation function when one class significantly outweighs the other. We demonstrate the proposed activation function on a publicly available dataset and externally validate on a dataset consisting of 1,909 healthy chest X-rays and 84 COVID-19 X-rays. The proposed method achieves an improved area under the receiver operating characteristic (DeLong's p-value < 0.05) compared to the sigmoid activation. Our method is also demonstrated on a dataset of healthy and pneumonia vs. COVID-19 X-rays and a set of computerized tomography images, achieving improved sensitivity. The proposed GEV activation function significantly improves upon the previously used sigmoid activation for binary classification. This new paradigm is expected to play a significant role in the fight against COVID-19 and other diseases, with relatively few training cases available.
Joshua Bridge, Yanda Meng, Yitian Zhao, Mingfeng Zhao, Renrong Sun, Yalin Zheng
IEEE J. Biomed. Health Informatics7
2020 Cooperative Low-Rank Models for Removing Stripe Noise From OCTA Images
abstract
Optical coherence tomography angiography (OCTA) is an emerging non-invasive imaging technique for imaging the microvasculature of the eye based on phase variance or amplitude decorrelation derived from repeated OCT images of the same tissue area. Stripe noise occurs during the OCTA acquisition process due to the involuntary movement of the eye. To remove the stripe noise (or 'destriping') effectively, we propose two novel image decomposition models to simultaneously destripe all the OCTA images of the same eye cooperatively: cooperative uniformity destriping (CUD) model and cooperative similarity destriping (CSD) model. Both the models consider stripe noise by low-rank constraint but in different ways: the CUD model assumes that stripe noise is identical across all the layers while the CSD model assumes that the stripe noise at different layers are different and have to be considered in the model. Compared to the CUD model, CSD is a more general solution for real OCTA images. An efficient solution (CSD+) is developed for model CSD to reduce the computational complexity. The models were extensively evaluated against state-of-the-art methods on both synthesized and real OCTA datasets. The experiments demonstrated not only the effectiveness of the CSD and CSD+ models in terms of peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) and CSD+ is twice faster than CSD, but also their beneficiary effect on the vessel segmentation of OCTA images. We expect our models will become a powerful tool for clinical applications.
Xiyin Wu, Dongxu Gao, David Borroni, Savita Madhusudhan, Zhong Jin, Yalin Zheng
IEEE J. Biomed. Health Informatics6
2020 Retinal Vascular Network Topology Reconstruction and Artery/Vein Classification via Dominant Set Clustering
abstract
The estimation of vascular network topology in complex networks is important in understanding the relationship between vascular changes and a wide spectrum of diseases. Automatic classification of the retinal vascular trees into arteries and veins is of direct assistance to the ophthalmologist in terms of diagnosis and treatment of eye disease. However, it is challenging due to their projective ambiguity and subtle changes in appearance, contrast, and geometry in the imaging process. In this paper, we propose a novel method that is capable of making the artery/vein (A/V) distinction in retinal color fundus images based on vascular network topological properties. To this end, we adapt the concept of dominant set clustering and formalize the retinal blood vessel topology estimation and the A/V classification as a pairwise clustering problem. The graph is constructed through image segmentation, skeletonization, and identification of significant nodes. The edge weight is defined as the inverse Euclidean distance between its two end points in the feature space of intensity, orientation, curvature, diameter, and entropy. The reconstructed vascular network is classified into arteries and veins based on their intensity and morphology. The proposed approach has been applied to five public databases, namely INSPIRE, IOSTAR, VICAVR, DRIVE, and WIDE, and achieved high accuracies of 95.1%, 94.2%, 93.8%, 91.1%, and 91.0%, respectively. Furthermore, we have made manual annotations of the blood vessel topologies for INSPIRE, IOSTAR, VICAVR, and DRIVE datasets, and these annotations are released for public access so as to facilitate researchers in the community.
Yitian Zhao, Yonghuai Liu, Jianyang Xie, Huaizhong Zhang, Yalin Zheng, Yifan Zhao 0001, Yangchun Zhao, Pan Su 0001, Jiang Liu 0001
IEEE Trans. Medical Imaging5
2020 Automated Tortuosity Analysis of Nerve Fibers in Corneal Confocal Microscopy
abstract
Precise characterization and analysis of corneal nerve fiber tortuosity are of great importance in facilitating examination and diagnosis of many eye-related diseases. In this paper we propose a fully automated method for image-level tortuosity estimation, comprising image enhancement, exponential curvature estimation, and tortuosity level classification. The image enhancement component is based on an extended Retinex model, which not only corrects imbalanced illumination and improves image contrast in an image, but also models noise explicitly to aid removal of imaging noise. Afterwards, we take advantage of exponential curvature estimation in the 3D space of positions and orientations to directly measure curvature based on the enhanced images, rather than relying on the explicit segmentation and skeletonization steps in a conventional pipeline usually with accumulated pre-processing errors. The proposed method has been applied over two corneal nerve microscopy datasets for the estimation of a tortuosity level for each image. The experimental results show that it performs better than several selected state-of-the-art methods. Furthermore, we have performed manual gradings at tortuosity level of four hundred and three corneal nerve microscopic images, and this dataset has been released for public access to facilitate other researchers in the community in carrying out further research on the same and related topics.
Yitian Zhao, Jiong Zhang 0004, Ella Grishikashvili Pereira, Yalin Zheng, Pan Su 0001, Jianyang Xie, Yifan Zhao 0001, Yonggang Shi, Jiang Liu 0001, Yonghuai Liu
IEEE Trans. Medical Imaging4
2020 Corrections to "Automated Tortuosity Analysis of Nerve Fibers in Corneal Confocal Microscopy"
abstract
In the above article[1], there were two errors in the printed article that the authors want to correct.
Yitian Zhao, Jiong Zhang 0004, Ella Grishikashvili Pereira, Yalin Zheng, Pan Su 0001, Jianyang Xie, Yifan Zhao 0001, Yonggang Shi, Jiang Liu 0001, Yonghuai Liu
IEEE Trans. Medical Imaging4
2019 Learning Active Contour Models for Medical Image Segmentation
abstract
Image segmentation is an important step in medical image processing and has been widely studied and developed for refinement of clinical analysis and applications. New models based on deep learning have improved results but are restricted to pixel-wise fitting of the segmentation map. Our aim was to tackle this limitation by developing a new model based on deep learning which takes into account the area inside as well as outside the region of interest as well as the size of boundaries during learning. Specifically, we propose a new loss function which incorporates area and size information and integrates this into a dense deep learning model. We evaluated our approach on a dataset of more than 2,000 cardiac MRI scans. Our results show that the proposed loss function outperforms other mainstream loss function Cross-entropy on two common segmentation networks. Our loss function is robust while using different hyperparameter lambda.
Xu Chen 0030, Bryan M. Williams 0001, Srinivasa R. Vallabhaneni, Gabriela Czanner, Rachel Williams, Yalin Zheng
CVPR6
2019 Topology Reconstruction of Tree-Like Structure in Images via Structural Similarity Measure and Dominant Set Clustering
abstract
The reconstruction and analysis of tree-like topological structures in the biomedical images is crucial for biologists and surgeons to understand biomedical conditions and plan surgical procedures. The underlying tree-structure topology reveals how different curvilinear components are anatomically connected to each other. Existing automated topology reconstruction methods have great difficulty in identifying the connectivity when two or more curvilinear components cross or bifurcate, due to their projection ambiguity, imaging noise and low contrast. In this paper, we propose a novel curvilinear structural similarity measure to guide a dominant-set clustering approach to address this indispensable issue. The novel similarity measure takes into account both intensity and geometric properties in representing the curvilinear structure locally and globally, and group curvilinear objects at crossover points into different connected branches by dominant-set clustering. The proposed method is applicable to different imaging modalities, and quantitative and qualitative results on retinal vessel, plant root, and neuronal network datasets show that our methodology is capable of advancing the current state-of-the-art techniques.
Jianyang Xie, Yitian Zhao, Yonghuai Liu, Pan Su 0001, Yifan Zhao 0001, Jun Cheng 0003, Yalin Zheng, Jiang Liu 0001
CVPR7
2019 CS-Net: Channel and Spatial Attention Network for Curvilinear Structure Segmentation
Lei Mou, Yitian Zhao, Li Chen 0011, Jun Cheng 0003, Zaiwang Gu, Huaying Hao, Yalin Zheng, Alejandro F. Frangi, Jiang Liu 0001
MICCAI (1)8
2019 Exploiting Reliability-Guided Aggregation for the Assessment of Curvilinear Structure Tortuosity
Pan Su 0001, Yitian Zhao, Tianhua Chen, Jianyang Xie, Yifan Zhao 0001, Yalin Zheng, Jiang Liu 0001
MICCAI (4)7
2018 Retinal Artery and Vein Classification via Dominant Sets Clustering-Based Vascular Topology Estimation
Yitian Zhao, Jianyang Xie, Pan Su 0001, Yalin Zheng, Yonghuai Liu, Jun Cheng 0003, Jiang Liu 0001
MICCAI (2)4
2018 Uniqueness-Driven Saliency Analysis for Automated Lesion Detection with Applications to Retinal Diseases
Yitian Zhao, Yalin Zheng, Yifan Zhao 0001, Yonghuai Liu, Peng Liu 0049, Jiang Liu 0001
MICCAI (2)2
2018 Automatic 2-D/3-D Vessel Enhancement in Multiple Modality Images Using a Weighted Symmetry Filter
abstract
Automated detection of vascular structures is of great importance in understanding the mechanism, diagnosis, and treatment of many vascular pathologies. However, automatic vascular detection continues to be an open issue because of difficulties posed by multiple factors, such as poor contrast, inhomogeneous backgrounds, anatomical variations, and the presence of noise during image acquisition. In this paper, we propose a novel 2-D/3-D symmetry filter to tackle these challenging issues for enhancing vessels from different imaging modalities. The proposed filter not only considers local phase features by using a quadrature filter to distinguish between lines and edges, but also uses the weighted geometric mean of the blurred and shifted responses of the quadrature filter, which allows more tolerance of vessels with irregular appearance. As a result, this filter shows a strong response to the vascular features under typical imaging conditions. Results based on eight publicly available datasets (six 2-D data sets, one 3-D data set, and one 3-D synthetic data set) demonstrate its superior performance to other state-of-the-art methods.
Yitian Zhao, Yalin Zheng, Yonghuai Liu, Yifan Zhao 0001, Lingling Luo, Tong Na, Yongtian Wang, Jiang Liu 0001
IEEE Trans. Medical Imaging2
2017 A Novel Choroid Segmentation Method for Retinal Diagnosis Using Deep Learning
abstract
Reliable choroid measurements have become an important diagnostic modality for sight-threatening retinal diseases. However, automatic and accurate segmentation of the choroid remains an unresolved challenge. This paper proposes a novel choroid segmentation method, based on a deep learning algorithm that is capable of quick and accurate image segmentation without user intervention. This is achieved through combining pixel clustering, image enhancement and deep learning. The simple linear iterative clustering (SLIC) algorithm has been applied to extract the superpixels (patches). Next, the extracted patches are then enhanced through increasing contrast of the region of interest. After that, the patches are fed to convolutional neural network for labelling the regions into choroid or non-choroid. Performance of the developed algorithm is assessed using a dataset of 169 enhanced depth imaging optical coherence tomography images. The obtained results demonstrated effectiveness of the proposed segmentation method in terms of accuracy (98.01%).
Baidaa Al-Bander, Bryan M. Williams 0001, Majid A. Al-Taee, Waleed Al-Nuaimy, Yalin Zheng
DeSE5
2017 FCNN: Fourier Convolutional Neural Networks
Harry Pratt, Bryan M. Williams 0001, Frans Coenen, Yalin Zheng
ECML/PKDD (1)4
2017 Saliency driven vasculature segmentation with infinite perimeter active contour model
Yitian Zhao, Jingliang Zhao, Jian Yang 0009, Yonghuai Liu, Yifan Zhao 0001, Yalin Zheng, Likun Xia, Yongtian Wang
Neurocomputing6
2017 Intensity and Compactness Enabled Saliency Estimation for Leakage Detection in Diabetic and Malarial Retinopathy
abstract
Leakage in retinal angiography currently is a key feature for confirming the activities of lesions in the management of a wide range of retinal diseases, such as diabetic maculopathy and paediatric malarial retinopathy. This paper proposes a new saliency-based method for the detection of leakage in fluorescein angiography. A superpixel approach is firstly employed to divide the image into meaningful patches (or superpixels) at different levels. Two saliency cues, intensity and compactness, are then proposed for the estimation of the saliency map of each individual superpixel at each level. The saliency maps at different levels over the same cues are fused using an averaging operator. The two saliency maps over different cues are fused using a pixel-wise multiplication operator. Leaking regions are finally detected by thresholding the saliency map followed by a graph-cut segmentation. The proposed method has been validated using the only two publicly available datasets: one for malarial retinopathy and the other for diabetic retinopathy. The experimental results show that it outperforms one of the latest competitors and performs as well as a human expert for leakage detection and outperforms several state-of-the-art methods for saliency detection.
Yitian Zhao, Yalin Zheng, Yonghuai Liu, Jian Yang 0009, Yifan Zhao 0001, Duanduan Chen, Yongtian Wang
IEEE Trans. Medical Imaging2
2016 Automatic Feature Learning Method for Detection of Retinal Landmarks
abstract
This paper presents an automatic deep learning method for location detection of important retinal landmarks, the fovea and optic disc (OD) in digital fundus retinal images with the potential for use in an automated screening and grading system. The proposed method is based on deep convolutional neural networks (CNN) and does not depend the visual appearance or anatomical features of the retinal landmarks. It comprises convolution, max-pooling, fully connected and dropout layers as well as an output layer. The CNN is trained using an existing dataset images along with their annotated locations of the foveal and OD centres. Performance of the network is evaluated using Root Mean Square Error (RMSE). The developed feature learning-based approach presents promising system for retinal landmarks detection.
Baidaa Al-Bander, Waleed Al-Nuaimy, Majid A. Al-Taee, Ali Al-Ataby, Yalin Zheng
DeSE5
2015 Automated Vessel Segmentation Using Infinite Perimeter Active Contour Model with Hybrid Region Information with Application to Retinal Images
abstract
Automated detection of blood vessel structures is becoming of crucial interest for better management of vascular disease. In this paper, we propose a new infinite active contour model that uses hybrid region information of the image to approach this problem. More specifically, an infinite perimeter regularizer, provided by using L(2) Lebesgue measure of the γ -neighborhood of boundaries, allows for better detection of small oscillatory (branching) structures than the traditional models based on the length of a feature's boundaries (i.e., H(1) Hausdorff measure). Moreover, for better general segmentation performance, the proposed model takes the advantage of using different types of region information, such as the combination of intensity information and local phase based enhancement map. The local phase based enhancement map is used for its superiority in preserving vessel edges while the given image intensity information will guarantee a correct feature's segmentation. We evaluate the performance of the proposed model by applying it to three public retinal image datasets (two datasets of color fundus photography and one fluorescein angiography dataset). The proposed model outperforms its competitors when compared with other widely used unsupervised and supervised methods. For example, the sensitivity (0.742), specificity (0.982) and accuracy (0.954) achieved on the DRIVE dataset are very close to those of the second observer's annotations.
Yitian Zhao, Lavdie Rada, Ke Chen 0002, Simon P. Harding, Yalin Zheng
IEEE Trans. Medical Imaging5
2013 Classification of volumetric retinal images using overlapping decomposition and tree analysis
abstract
Methods for the classification of volumetric three-dimensional (3D) volumes (images) play an important role in the context of medical applications. In this paper, a dedicated tree based 3D representation is proposed that serves to directly capture 3D image features in such a way that classification techniques can be applied. More specifically an Overlapping Hierarchical Decomposition (OHD) technique is presented to generate a tree representation of a given 3D volume. The OHD method recursively decomposes a given 3D volume into sub-volumes forming a tree. Once the tree has been generated, a frequent sub-graph mining algorithm is applied to mine the tree representation so as to generate sub-graphs. These sub-graphs are then used to define a feature space from which feature vectors representing 3D images (one per 3D volume) can be extracted and fed into a classifier generator. To demonstrate the applicability of the proposed method a 3D Optical Coherence Tomography (OCT) retinal image screening application is considered directed at the identification of Age-related Macular Degeneration (AMD). The results show a promising performance with a best Area Under the receiver operating Curve (AUC) value of 98.7%.
Abdulrahman Albarrak, Frans Coenen, Yalin Zheng
CBMS3
2012 Data mining techniques for the screening of age-related macular degeneration
Mohd. Hanafi Ahmad Hijazi, Frans Coenen, Yalin Zheng
Knowl. Based Syst.3
2011 Time Series Case Based Reasoning for Image Categorisation
Ashraf Elsayed, Mohd. Hanafi Ahmad Hijazi, Frans Coenen, Marta García-Fiñana, Vanessa Sluming, Yalin Zheng
ICCBR6
2011 Image Classification for Age-related Macular Degeneration Screening Using Hierarchical Image Decompositions and Graph Mining
Mohd. Hanafi Ahmad Hijazi, Chuntao Jiang, Frans Coenen, Yalin Zheng
ECML/PKDD (2)4
2010 Retinal image classification using a histogram based approach
abstract
An approach to classifying retinal images using a histogram based representation is described. More specifically, a two stage Case Based Reasoning (CBR) approach is proposed, to be applied to histogram represented retina images to identify Age-related Macular Degeneration (AMD). To measure the similarity between histograms, a time series analysis technique, Dynamic Time Warping (DTW), is employed. The advocated approach utilises two “case bases” for the classification process. The first case base consists of green and saturation histograms with retinal blood vessels removed. The second case base comprises the same histograms, but with the Optic Disc (OD) removed as well. The reported experiments demonstrate that the proposed two stage classification process outperforms the single stage classification process with respect to a number of evaluation metrics: specificity, sensitivity and accuracy.
Mohd. Hanafi Ahmad Hijazi, Frans Coenen, Yalin Zheng
IJCNN3
2008 Equivalence Knowledge Mass and Approximate Reasoning in -Logic (I)
Yalin Zheng, Guang Yang 0002, Changshui Zhang, Yunpeng Xu
ICIC (2)1
2004 Mamdanian logic
abstract
In the framework of stratified fuzzy propositional logic F, the regular harmonious neighborhood structure determined by the Mamdanian approximate function Hgt is the typical and best paradigm of approximate reasoning neighborhood algorithms. The advantage of Mamdanian algorithm comes from the regular harmoniousness property of Hgt presented by us in this paper, which guarantees that when the input of the obtained new approximate knowledge is close enough to that of the standard knowledge, not only the approximate reasoning consequence of the new knowledge is adequately close to that of the standard knowledge, but also the new knowledge itself is close enough to the standard knowledge.
Yalin Zheng, Changshui Zhang, Xing Yi
FUZZ-IEEE1
2004 Automated segmentation of lumbar vertebrae in digital videofluoroscopic images
abstract
Low back pain is a significant problem in the industrialized world. Diagnosis of the underlying causes can be extremely difficult. Since mechanical factors often play an important role, it can be helpful to study the motion of the spine. Digital videofluoroscopy has been developed for this study and it can provide image sequences with many frames, but which often suffer due to noise, exacerbated by the very low radiation dosage. Thus, determining vertebra position within the image sequence presents a considerable challenge. There have been many studies on vertebral image extraction, but problems of repeatability, occlusion and out-of-plane motion persist. In this paper, we show how the Hough transform (HT) can be used to solve these problems. Here, Fourier descriptors were used to describe the vertebral body shape. This description was incorporated within our HT algorithm from which we can obtain affine transform parameters, i.e., scale, rotation and center position. The method has been applied to images of a calibration model and to images from two sequences of moving human lumbar spines. The results show promise and potential for object extraction from poor quality images and that models of spinal movement can indeed be derived for clinical application.
Yalin Zheng, Mark S. Nixon
IEEE Trans. Medical Imaging1
2001 On stratiform L-fuzzy topologies and their application
Xingfang Zhang, Guangwu Meng, Yalin Zheng, Qingde Zhang
Fuzzy Sets Syst.3