Linghan Cai

dblp:123/7994 · DBLP profile ↗
← Back
18ranked-venue papers
3as first author
18since 2021 · last 2026
0000-0002-7931-7697ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2026 PathFLIP: Fine-grained Language-Image Pretraining for Versatile Computational Pathology
abstract
While Vision-Language Models (VLMs) have achieved notable progress in computational pathology (CPath), the gigapixel scale and spatial heterogeneity of Whole Slide Images (WSIs) continue to pose challenges for multimodal understanding. Existing alignment methods struggle to capture fine-grained correspondences between textual descriptions and visual cues across thousands of patches from a slide, compromising their performance on downstream tasks. In this paper, we propose PathFLIP (Pathology Fine-grained Language-Image Pretraining), a novel framework for holistic WSI interpretation. PathFLIP decomposes slide-level captions into region-level sub-captions and generates text-conditioned region embeddings to facilitate precise visual-language grounding. By harnessing Large Language Models (LLMs), PathFLIP can seamlessly follow diverse clinical instructions and adapt to varied diagnostic contexts. Furthermore, it exhibits versatile capabilities across multiple paradigms, efficiently handling slide-level classification and retrieval, fine-grained lesion localization, and instruction following. Extensive experiments demonstrate that PathFLIP outperforms existing large-scale pathological VLMs on four representative benchmarks while requiring significantly less training data, paving the way for fine-grained, instruction-aware WSI interpretation in research and clinical practice.
Fengchun Liu, Songhan Jiang, Linghan Cai, Ziyue Wang 0005, Yongbing Zhang 0002
AAAI3
2026 Unsupervised cross-domain semantic segmentation on multi-modality ovarian tumor ultrasound data
Shuchang Lyu, Qi Zhao 0037, Wenpei Bai, Linghan Cai, Guangxia Cui, Lijiang Chen, Huiyu Zhou 0001
Pattern Recognit.4
2026 SeCoMIL: Semantic Anchor-Based Context-Aware Multiple Instance Learning for Whole Slide Image Classification
abstract
Context-aware Multiple Instance Learning (MIL) is gaining popularity in Whole Slide Image (WSI) classification. Existing methods typically convert instances in a WSI into one-dimensional sequences and learn the long-range contextual dependencies among instances. However, due to the extremely large size of WSIs and the morphological similarities within tissue structures, the enormous number of redundant instances significantly increases computational overhead in the context learning paradigm. Additionally, the rearrangement of instances into one dimension loses the inherent spatial information involved in image patches, further compromising the classification performance of pathological images. Consequently, efficiently modeling contextual dependencies in WSIs remains a crucial challenge. In this paper, we propose a novel Semantic Anchor-based Context-aware Multiple Instance Learning (SeCoMIL) framework. This framework partitions the WSI into a series of regions and encodes the coordinates of instances within these regions to preserve their spatial relationships. Subsequently, SeCoMIL identifies the most representative instances from each region as semantic anchors. By capturing both the local context around these anchors and the global context across different anchors, the framework efficiently summarizes the critical pathological information of the WSI, enabling precise classification. Extensive experiments on four public datasets (CAMELYON16, CAMELYON17, TCGA-NSCLC, and TCGA-RCC) demonstrate the robustness of our method, with superior performance compared to state-of-the-art methods.
Shenjin Huang, GuoJun Liu, Linghan Cai, Hailun Cheng, Zichun Huang, Yongbing Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.3
2026 Uncertainty-Aware Survival Analysis With Dirichlet Distribution for Multi-Scale Pathology and Genomics
abstract
Over the last few decades, the integration of AI-driven computational techniques into digital pathology has revolutionized survival prediction tasks. However, most existing methods in survival analysis discretize the entire survival period into predefined intervals, overlooking the inherent uncertainty in event occurrence and the heterogeneity of patient survival times. The censored data further exacerbate these challenges, amplifying uncertainty and variability. To address these limitations, we introduce the Dirichlet distribution to model discretized outputs as continuous probability distributions, providing a more accurate representation of uncertainty awareness. Building upon this foundation, we propose a universal multi-modal survival analysis loss function that leverages uncertainty-driven fusion. Our Uncertainty-Aware Multi-Modal Survival Analysis (UMSA) framework further explores the interactions between multi-scale pathological images and genomic data, providing promising insights into multi-modal survival analysis. Experimental evaluations on five publicly available datasets demonstrate that UMSA achieves state-of-the-art performance, validating its effectiveness and scalability in survival prediction tasks.
Songhan Jiang, Linghan Cai, Zhengyu Gan, Yifeng Wang 0001, Guo Tang, Yongbing Zhang 0002
IEEE Trans. Medical Imaging2
2025 AMKD: Adaptive Multi-modality Knowledge Distillation for Pathological Survival Analysis
Yangfan Xu, Linghan Cai, Yifeng Wang 0001, Hailun Cheng, Fengchun Liu, Runming Wang, Yongbing Zhang 0002
ICIC (27)2
2025 Relation-Aware Graph Attention Network for Nuclei Classification
abstract
Nuclei classification plays a pivotal role in pathological research. Recent advances in graph neural networks (GNNs) have shown great promise in modeling cell-cell interactions. However, many existing methods overlook tissue context, which is crucial for accurate nuclei identification, as nuclei exhibit distinct patterns within specific tissue structures. To address this limitation, we propose a novel Relation-Aware Graph AT-tention network (RAGAT) that effectively leverages nucleus-related features for precise classification. RAGAT constructs a cell graph based on spatial proximity and visual feature similarity, while also introducing a tissue-aware graph by sampling regions around each nucleus to capture the tissue microenvironment and depict local cellular contexts. Furthermore, RAGAT employs a hybrid graph attention module to integrate cell-cell and tissue-cell interactions, enabling a comprehensive understanding of the nuclear context. Experimental results on three benchmark datasets demonstrate that our method significantly outperforms state-of-the-art approaches, offering valuable insight into the analysis of nuclear microenvironments. Our code is available at https://github.com/lingboboo/RAGAT.
Lingbo Zhang, Ye Zhang 0043, Linghan Cai, Xianchao Guan, Kai Zhang 0012, Yongbing Zhang 0002
ICME3
2025 Prototype-Guided Cross-Modal Knowledge Enhancement for Adaptive Survival Prediction
Fengchun Liu, Linghan Cai, Zhikang Wang, Zhiyuan Fan, Jin-gang Yu, Hao Chen 0011, Yongbing Zhang 0002
MICCAI (6)2
2025 Counting by Points: Density-Guided Weakly-Supervised Nuclei Segmentation in Histopathological Images
Lingbo Zhang, Bingqian Sun, Linghan Cai, Yifeng Wang 0001, Ye Zhang 0043, Songhan Jiang, Kai Zhang 0012, Yongbing Zhang 0002
ACM Multimedia3
2025 AttriMIL: Revisiting attention-based multiple instance learning for whole-slide pathological image classification from a perspective of instance attributes
abstract
Multiple instance learning (MIL) is a powerful approach for whole-slide pathological image (WSI) analysis, particularly suited for processing gigapixel-resolution images with slide-level labels. Recent attention-based MIL architectures have significantly advanced weakly supervised WSI classification, facilitating both clinical diagnosis and localization of disease-positive regions. However, these methods often face challenges in differentiating between instances, leading to tissue misidentification and a potential degradation in classification performance. To address these limitations, we propose AttriMIL, an attribute-aware multiple instance learning framework. By dissecting the computational flow of attention-based MIL models, we introduce a multi-branch attribute scoring mechanism that quantifies the pathological attributes of individual instances. Leveraging these quantified attributes, we further establish region-wise and slide-wise attribute constraints to dynamically model instance correlations both within and across slides during training. These constraints encourage the network to capture intrinsic spatial patterns and semantic similarities between image patches, thereby enhancing its ability to distinguish subtle tissue variations and sensitivity to challenging instances. To fully exploit the two constraints, we further develop a pathology adaptive learning technique to optimize pre-trained feature extractors, enabling the model to efficiently gather task-specific features. Extensive experiments on five public datasets demonstrate that AttriMIL consistently outperforms state-of-the-art methods across various dimensions, including bag classification accuracy, generalization ability, and disease-positive region localization. The implementation code is available at https://github.com/MedCAI/AttriMIL.
Linghan Cai, Shenjin Huang, Ye Zhang 0043, Jinpeng Lu, Yongbing Zhang 0002
Medical Image Anal.1
2025 Focus Your Attention: Multiple Instance Learning With Attention Modification for Whole Slide Pathological Image Classification
abstract
Computer-aided pathology diagnosis based on whole slide images, which is often formulated as a weakly supervised multiple instance learning (MIL) paradigm. Current approaches generally employ attention mechanisms to aggregate instance-level features. However, the weakly supervised signal and the imbalanced instance distribution often lead to inaccurate attention localization, compromising the performance and generalization capability of the MIL framework. To address these problems, this paper presents a novel MIL framework called FAMIL that focuses on inaccurate attention and refines them. FAMIL adopts a dual-branch structure and incorporates two innovative online data augmentation strategies: attention-based Mixup (ABMix) and attention-based Masking (ABMask). ABMix emphasizes the significance of positive instances, generalizing Mixup in the MIL scenarios, while ABMask flexibly identifies challenging positive instances to optimize the feature representation. Moreover, these two methods are plug-and-play and can be easily embedded into attention-based MIL methods. Extensive experiments on three public benchmarks demonstrate the superiority of our FAMIL, outperforming current state-of-the-art methods. The test AUC for the binary tumor classification can be up to 92.61% over CAMELYON16. And the AUC over the cancer subtype classification can be up to 93.81% and 98.41% on TCGA-NSCLC and TCGA-RCC datasets, respectively.
Hailun Cheng, Shenjin Huang, Linghan Cai, Yangfan Xu, Runming Wang, Yongbing Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.3
2025 DAWN: Domain-Adaptive Weakly Supervised Nuclei Segmentation via Cross-Task Interactions
abstract
Weakly supervised segmentation methods have garnered considerable attention due to their potential to alleviate the need for labor-intensive pixel-level annotations during model training. Traditional weakly supervised nuclei segmentation approaches typically involve a two-stage process: pseudo-label generation followed by network training. The performance of these methods is highly dependent on the quality of the generated pseudo-labels, which can limit their effectiveness. In this paper, we propose a novel domain-adaptive weakly supervised nuclei segmentation framework that addresses the challenge of pseudo-label generation through cross-task interaction strategies. Specifically, our approach leverages weakly annotated data to train an auxiliary detection task, which facilitates domain adaptation of the segmentation network. To improve the efficiency of domain adaptation, we introduce a consistent feature constraint module that integrates prior knowledge from the source domain. Additionally, we develop methods for pseudo-label optimization and interactive training to enhance domain transfer capabilities. We validate the effectiveness of our proposed method through extensive comparative and ablation experiments conducted on six datasets. The results demonstrate that our approach outperforms existing weakly supervised methods and achieves performance comparable to or exceeding that of fully supervised methods. Our code is available athttps://github.com/zhangye-zoe/DAWN.
Ye Zhang 0043, Yifeng Wang 0001, Zijie Fang, Hao Bian, Linghan Cai, Ziyue Wang 0005, Yongbing Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.5
2025 SEINE: Structure Encoding and Interaction Network for Nuclei Instance Segmentation
abstract
Nuclei instance segmentation in histopathological images is crucial for biological analysis and cancer diagnosis. However, it faces two significant challenges: (1) poorly stained nuclei can lead to under-segmentation, as the background may be mistakenly identified as the foreground; and (2) deep textures within nuclei often result in fragmented instance predictions, as these textures can be misinterpreted as contours. To address these problems, this paper proposes a Structure Encoding and Interaction NEtwork, termed SEINE, which develops the nuclei structure modeling scheme and takes advantage of the similarity between nuclei structure to improve the integrality of instance segmentation. Specifically, SEINE introduces a contour-based structure encoding mechanism that integrates the correlation between nuclear structure and semantics, enabling a more accurate structural representation. Building on this encoding, we propose a structure-guided attention module, which uses clear nuclei as prototypes to guide the structural learning of unclear nuclei, thereby addressing the under-segmentation problem. Additionally, a position enhancement strategy applies a centroid distance constraint to reduce contour prediction errors, effectively mitigating fragmented instance segmentation. Extensive experiments demonstrate the effectiveness of SEINE, achieving state-of-the-art performance across four benchmark datasets.
Ye Zhang 0043, Linghan Cai, Ziyue Wang 0005, Yongbing Zhang 0002
IEEE J. Biomed. Health Informatics2
2025 Disentangled Pseudo-Bag Augmentation for Whole Slide Image Multiple Instance Learning
abstract
As the predominant approach for pathological whole slide image (WSI) classification, multiple instance learning (MIL) methods struggle with limited labeled WSIs. Although MIL has achieved notable progress with pseudo-bag-oriented augmentation methods, their effectiveness is often constrained by noisy pseudo-labels and low-quality pseudo-bags. To overcome these problems, we revisit the use of pseudo-bags for WSI data augmentation and propose a new pseudo-bag generation paradigm, dubbed DPBAug. Its distinctive features can be summarized as: i) We develop an intra-slide pseudo-bag generation module, which separates the heterogeneous instances within each slide through phenotype partitioning. Moreover, to ensure accurate label inheritance when generating pseudo-bags, we propose an instance sampling algorithm with replacement. ii) An inter-slide pseudo-bag fusion module is designed to integrate heterogeneous information across multiple WSIs, producing high-quality training samples that better leverage the potential of neural networks. iii) A pseudo-bag memory update module prioritizes valuable synthetic pseudo-bags. This further enhances the network's classification performance. Extensive experiments demonstrate that DPBAug surpasses existing augmentation methods, enhancing the classification performance and reliability of multiple MIL baselines across various public datasets. DPBAug also improves the generalization and data efficiency of existing MIL methods, facilitating their adoption in clinical practice and rare cancer research The project is available at: https://github.com/JiuyangDong/DPBAug.
Jiuyang Dong, Junjun Jiang, Kui Jiang, Jiahan Li, Linghan Cai, Yongbing Zhang 0002
IEEE Trans. Medical Imaging5
2024 Multimodal Cross-Task Interaction for Survival Analysis in Whole Slide Pathological Images
Songhan Jiang, Zhengyu Gan, Linghan Cai, Yifeng Wang 0001, Yongbing Zhang 0002
MICCAI (4)3
2024 H2ASeg: Hierarchical Adaptive Interaction and Weighting Network for Tumor Segmentation in PET/CT Images
Jinpeng Lu, Jingyun Chen, Linghan Cai, Songhan Jiang, Yongbing Zhang 0002
MICCAI (8)3
2024 Dynamic Pseudo Label Optimization in Point-Supervised Nuclei Segmentation
Ziyue Wang 0005, Ye Zhang 0043, Yifeng Wang 0001, Linghan Cai, Yongbing Zhang 0002
MICCAI (8)4
2024 Know your orientation: A viewpoint-aware framework for polyp segmentation
abstract
Automatic polyp segmentation in endoscopic images is critical for the early diagnosis of colorectal cancer. Despite the availability of powerful segmentation models, two challenges still impede the accuracy of polyp segmentation algorithms. Firstly, during a colonoscopy, physicians frequently adjust the orientation of the colonoscope tip to capture underlying lesions, resulting in viewpoint changes in the colonoscopy images. These variations increase the diversity of polyp visual appearance, posing a challenge for learning robust polyp features. Secondly, polyps often exhibit properties similar to the surrounding tissues, leading to indistinct polyp boundaries. To address these problems, we propose a viewpoint-aware framework named VANet for precise polyp segmentation. In VANet, polyps are emphasized as a discriminative feature and thus can be localized by class activation maps in a viewpoint classification process. With these polyp locations, we design a viewpoint-aware Transformer (VAFormer) to alleviate the erosion of attention by the surrounding tissues, thereby inducing better polyp representations. Additionally, to enhance the polyp boundary perception of the network, we develop a boundary-aware Transformer (BAFormer) to encourage self-attention towards uncertain regions. As a consequence, the combination of the two modules is capable of calibrating predictions and significantly improving polyp segmentation performance. Extensive experiments on seven public datasets across six metrics demonstrate the state-of-the-art results of our method, and VANet can handle colonoscopy images in real-world scenarios effectively. The source code is available at https://github.com/1024803482/Viewpoint-Aware-Network.
Linghan Cai, Lijiang Chen, Yifeng Wang 0001, Yongbing Zhang 0002
Medical Image Anal.1
2022 Using Guided Self-Attention with Local Information for Polyp Segmentation
Linghan Cai, Meijing Wu, Lijiang Chen, Wenpei Bai, Shuchang Lyu, Qi Zhao 0037
MICCAI (4)1