VLDB 2026 Research / reviewers in the wild / expert
Chenglu Zhu
dblp:296/3987
· DBLP profile ↗
30ranked-venue papers
0as first author
30since 2021 · last 2026
0000-0001-5705-3718ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 21 since 2021Artificial intelligence and machine learning · 17 · 17 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MIRA: Evaluating Multimodal AI on Complex Clinical Reasoning in Interventional RadiologyabstractWe present MIRA (Multimodal Interventional RAdiology evaluation), a comprehensive benchmark for evaluating large multimodal models in expert-level interventional radiology tasks requiring specialized domain knowledge and advanced visual reasoning capabilities. Unlike existing medical benchmarks that primarily provide binary labels without contextual depth, MIRA offers diverse question formats, including open-ended, closed-ended, single-choice, and multiple-choice categories, each accompanied by detailed expert-validated explanations. The benchmark incorporates approximately 184K high-quality medical images spanning multiple imaging modalities with 1.2M meticulously generated question-answer pairs across various anatomical regions. These pairs were created through a sophisticated cascade methodology involving expert interventional radiologists at both the data collection and validation stages. Our comprehensive evaluation, encompassing zero-shot testing and fine-tuning experiments of large multimodal models, revealing significant performance gaps between AI systems and human specialists. Fine-tuning experiments demonstrate substantial improvements, with models achieving up to 0.80 accuracy on single-choice questions. MIRA establishes a challenging benchmark that suggests promising directions for developing specialized clinical AI systems for interventional radiology. Jingxiong Li, Chenglu Zhu, Sunyi Zheng, Yuxuan Sun 0002, Yixuan Si, Lin Yang 0002, Liang Xiao 0001 |
AAAI | 2 |
| 2026 | Towards Effective and Efficient Context-aware Nucleus Detection in Histopathology Whole Slide ImagesabstractNucleus detection in histopathology whole slide images (WSIs) is crucial for a broad spectrum of clinical applications. The gigapixel size of WSIs necessitates the use of sliding window methodology for nucleus detection. However, mainstream methods process each sliding window independently, which overlooks broader contextual information and easily leads to inaccurate predictions. To address this limitation, recent studies additionally crop a large Filed-of-View (LFoV) patch centered on each sliding window to extract contextual features. However, such methods substantially increase whole-slide inference latency. In this work, we propose an effective and efficient context-aware nucleus detection approach. Specifically, instead of using lFoV patches, we aggregate contextual clues from off-the-shelf features of historically visited sliding windows, which greatly enhances the inference efficiency. Moreover, compared to lFoV patches used in previous works, the sliding window patches have higher magnification and provide finer-grained tissue details, thereby enhancing the classification accuracy. To develop the proposed context-aware model, we utilize annotated patches along with their surrounding unlabeled patches for training. Beyond exploiting high-level tissue context from these surrounding regions, we design a post-training strategy that leverages abundant unlabeled nucleus samples within them to enhance the model's context adaptability. Extensive experimental results on three challenging benchmarks demonstrate the superiority of our method. Zhongyi Shui, Honglin Li 0001, Yuxuan Sun 0002, Yiwen Ye, Pingyi Chen, Ruizhe Guo, Lei Cui 0004, Chenglu Zhu, Lin Yang 0002 |
AAAI | 9 |
| 2026 | DFFormer: Dual Frequency-Driven Transformer for real-world image deblurring
Ruizhe Guo, Shichuan Zhang, Jingxiong Li, Zhongyi Shui, Chenglu Zhu, Lin Yang 0002 |
Comput. Vis. Image Underst. | 5 |
| 2025 | CPath-Omni: A Unified Multimodal Foundation Model for Patch and Whole Slide Image Analysis in Computational PathologyabstractThe emergence of large multimodal models (LMMs) has brought significant advancements to pathology. Previous research has primarily focused on separately training patch-level and whole-slide image (WSI)-level models, limiting the integration of learned knowledge across patches and WSIs and resulting in redundant models. In this work, we introduce CPath-Omni, the first 15B parameter LMM that unifies patch and WSI analysis, consolidating a variety of tasks at both levels, including classification, visual question answering, captioning, and visual referring prompting. Extensive experiments demonstrate that CPath-Omni achieves state-of-the-art (SOTA) performance across seven diverse tasks on 39 out of 42 datasets, outperforming or matching task-specific models trained for individual tasks. Additionally, we develop a specialized pathology CLIP-based visual processor for CPath-Omni, CPath-CLIP, which, for the first time, integrates different vision models and incorporates a large language model as a text encoder to build a more powerful CLIP model, which achieves SOTA performance on nine zero-shot and four few-shot datasets. Our findings highlight CPath-Omni’s ability to unify diverse pathology tasks, demonstrating its potential to streamline and advance the field of foundation model in pathology. The code and model are available at CPath-Omni. Yuxuan Sun 0002, Yixuan Si, Chenglu Zhu, Kai Zhang 0033, Pingyi Chen, Zhongyi Shui, Tao Lin 0004, Lin Yang 0002 |
CVPR | 3 |
| 2025 | Stable Test-Time Training for Semantic Segmentation with Output Contrastive LossabstractDeep learning-based models have achieved impressive performance on public segmentation benchmarks, yet generalizing to unseen environments remains challenging. Test-time training (TTT) addresses this by adapting source-pretrained models during evaluation. While existing TTT methods have shown promise in image classification, they often exhibit instability with small test batches and class imbalance—challenges that intensify in semantic segmentation tasks. To tackle this issue, we present Output Contrastive Loss (OCL) to improve the stability of contrastive loss when applied to TTT for segmentation. OCL applies contrastive loss directly to the output space, avoiding the need for extra regularization, and employs a high temperature to prevent model collapse. To further stabilize the TTT process, we integrate BN statistics Modulation and Stochastic Restoration techniques. Extensive experiments across diverse datasets, settings, architectures, and pretrained methods demonstrate consistent performance improvements, achieving a 7.5 mIoU gain on the GTA→CS benchmark and showing effectiveness even with domain adaptation pretraining. Code is available at https://github.com/dazhangyul23/OCL. Zhongyi Shui, Honglin Li 0001, Yuxuan Sun 0002, Chenglu Zhu, Lin Yang 0002 |
ICASSP | 5 |
| 2025 | PathGen-1.6M: 1.6 Million Pathology Image-text Pairs Generation through Multi-agent CollaborationabstractVision Language Models (VLMs) like CLIP have attracted substantial attention in pathology, serving as backbones for applications such as zero-shot image classification and Whole Slide Image (WSI) analysis. Additionally, they can function as vision encoders when combined with large language models (LLMs) to support broader capabilities. Current efforts to train pathology VLMs rely on pathology image-text pairs from platforms like PubMed, YouTube, and Twitter, which provide limited, unscalable data with generally suboptimal image quality. In this work, we leverage large-scale WSI datasets like TCGA to extract numerous high-quality image patches. We then train a large multimodal model (LMM) to generate captions for extracted images, creating PathGen-1.6M, a dataset containing 1.6 million high-quality image-caption pairs. Our approach involves multiple agent models collaborating to extract representative WSI patches, generating and refining captions to obtain high-quality image-text pairs. Extensive experiments show that integrating these generated pairs with existing datasets to train a pathology-specific CLIP model, PathGen-CLIP, significantly enhances its ability to analyze pathological images, with substantial improvements across nine pathology-related zero-shot image classification tasks and three whole-slide image tasks. Furthermore, we construct 200K instruction-tuning data based on PathGen-1.6M and integrate PathGen-CLIP with the Vicuna LLM to create more powerful multimodal models through instruction tuning. Overall, we provide a scalable pathway for high-quality data generation in pathology, paving the way for next-generation general pathology models. Our dataset, code, and model are open-access at https://github.com/PathFoundation/PathGen-1.6M. Yuxuan Sun 0002, Yixuan Si, Chenglu Zhu, Kai Zhang 0033, Zhongyi Shui, Jingxiong Li, Xinheng Lyu, Tao Lin 0004, Lin Yang 0002 |
ICLR | 4 |
| 2025 | AEM: Attention Entropy Maximization for Multiple Instance Learning Based Whole Slide Image Classification
Honglin Li 0001, Yuxuan Sun 0002, Zhongyi Shui, Jingxiong Li, Chenglu Zhu, Lin Yang 0002 |
MICCAI (7) | 6 |
| 2025 | PathVQ: Reforming Computational Pathology Foundation Model for Whole Slide Image Analysis via Vector QuantizationabstractPathology whole slide image (WSI) analysis is vital for disease diagnosis and understanding. While foundation models (FMs) have driven recent advances, their scalability in pathology remains a key challenge. In particular, vision-language (VL) pathology FMs align visual features with language annotation for downstream tasks, but they rely heavily on large-scale image-text paired data, which is scarce thus limiting generalization. On the other hand, vision-only pathology FMs can leverage abundant unlabeled data via self-supervised learning (SSL). However, current approaches often use the [CLS] token from tile-level ViTs as slide-level input for efficiency (a tile with 224×224 pixels composed of 196 patches with 16×16 pixels). This SSL pretrained [CLS] token lacks alignment with downstream objectives, limiting effectiveness. We find that spatial patch tokens retain a wealth of informative features beneficial for downstream tasks, but utilizing all of them incurs up to 200× higher computation and storage costs compared [CLS] token only (e.g., 196 tokens per ViT$_{224}$). This highlights a fundamental trade-off between efficiency and representational richness to build scalable pathology FMs. To address this, we propose a feature distillation framework via vector-quantization (VQ) that compresses patch tokens into discrete indices and reconstructs them via a decoder, achieving 64× compression (1024 → 16 dimensions) while preserving fidelity. We further introduce a multi-scale VQ (MSVQ) strategy, enhancing both reconstruction and providing SSL supervision for slide-level pretraining. Built upon MSVQ features and supervision signals, we design a progressive convolutional module and a slide-level SSL objective to learn spatially rich representations for downstream WSI tasks. Extensive experiments across multiple datasets demonstrate that our approach achieves state-of-the-art performance, offering a scalable and effective solution for high-performing pathology FMs in WSI analysis. Honglin Li 0001, Zhongyi Shui, Chenglu Zhu, Lin Yang 0002 |
NeurIPS | 4 |
| 2025 | CPathAgent: An Agent-based Foundation Model for Interpretable High-Resolution Pathology Image Analysis Mimicking Pathologists' Diagnostic LogicabstractRecent advances in computational pathology have led to the emergence of numerous foundation models. These models typically rely on general-purpose encoders with multi-instance learning for whole slide image (WSI) classification or apply multimodal approaches to generate reports directly from images. However, these models cannot emulate the diagnostic approach of pathologists, who systematically examine slides at low magnification to obtain an overview before progressively zooming in on suspicious regions to formulate comprehensive diagnoses. Instead, existing models directly output final diagnoses without revealing the underlying reasoning process.
To address this gap, we introduce CPathAgent, an innovative agent-based approach that mimics pathologists' diagnostic workflow by autonomously navigating across WSI through zoom-in/out and move operations based on observed visual features, thereby generating substantially more transparent and interpretable diagnostic summaries. To achieve this, we develop a multi-stage training strategy that unifies patch-level, region-level, and WSI-level capabilities within a single model, which is essential for replicating how pathologists understand and reason across diverse image scales.
Additionally, we construct PathMMU-HR², the first expert-validated benchmark for large region analysis. This represents a critical intermediate scale between patches and whole slides, reflecting a key clinical reality where pathologists typically examine several key large regions rather than entire slides at once. Extensive experiments demonstrate that CPathAgent consistently outperforms existing approaches across benchmarks at three different image scales, validating the effectiveness of our agent-based diagnostic approach and highlighting a promising direction for computational pathology. Yuxuan Sun 0002, Yixuan Si, Chenglu Zhu, Kai Zhang 0033, Zhongyi Shui, Tao Lin 0004, Lin Yang 0002 |
NeurIPS | 3 |
| 2025 | ToPoFM: Topology-Guided Pathology Foundation Model for High-Resolution Pathology Image Synthesis With Cellular-Level ControlabstractSynthetic data generation emerges as a strategy to mitigate data scarcity in digital pathology, where complicated tissue and cellular features are correlated with cancer diagnosis. The synthesis of such visuals, however, suffers from limited inter class diversity and scarcity of cellular annotations. Current methodologies struggle with capturing the broad spectrum of pathology features, causing unpredictable objects and defected fidelity. Moreover, discrepancies in image resolution across developmental and operational phases can amplify the distribution shifts, undermining the precision of diagnosis. To address these challenges, we introduce TOpology guided PathOlogy Foundation Model (ToPoFM), a visual foundation model designed for the synthesis of high-resolution pathology images with cellular-level control. Our approach integrates a topology-informed cell arrangement generator to steer large language models for crafting synthetic cell arrangements. We correlate cell arrangement guidance with diffusion model for pathology content generation, then further implement a random sliding inference strategy, merging discrete low-resolution samplings into single high-resolution representation. Our model requires only small patches for training. The efficacy of ToPoFM is demonstrated through extensive experiments, complemented by expert validations, showing high fidelity on data synthesis. Additionally, we underscore the utility of our generated imagery as an augmentation tool, enhancing the performance of downstream tasks, including cancer subtype classification and segmentation. Jingxiong Li, Chenglu Zhu, Sunyi Zheng, Pingyi Chen, Yuxuan Sun 0002, Honglin Li 0001, Lin Yang 0002 |
IEEE Trans. Medical Imaging | 2 |
| 2025 | PathBench: Advancing the Benchmark of Large Multimodal Models for Pathology Image Understanding at Patch and Whole Slide LevelabstractRapid advancements in large multimodal models (LMMs) have significantly enhanced their applications in pathology, particularly in image classification, pathology image description, and whole slide image (WSI) classification. In pathology, WSIs represent gigapixel-scale images composed of thousands of image patches. Therefore, both patch-level and WSI-level evaluations are essential and inherently interconnected for assessing LMM capabilities. In this work, we propose PathBench, which comprises three subsets at both patch and WSI levels, to refine and enhance the validation of LMMs. At the patch-level, evaluations using existing multi-choice Q&A datasets reveal that some LMMs can predict answers without genuine image analysis. To address this, we introduce PatchVQA, a large-scale visual question answering (VQA) dataset containing 5,382 images and 6,335 multiple-choice questions designed with distractor options to prevent shortcut learning. These new questions are rigorously validated by professional pathologists to ensure reliable model assessments. At the WSI-level, current efforts primarily focus on image classification tasks and lack diverse validation datasets for multimodal models. To address this, we generate a detailed WSI report dataset through an innovative approach that integrates detailed patch descriptions generated by foundational models into comprehensive WSI reports. These are then combined with physician-written reports corresponding to TCGA WSIs, resulting in WSICap, a detailed report dataset containing 7,000 samples. Based on WSICap, we further develop a WSI-level VQA dataset, WSIVQA, to serve as a validation set for WSI LMMs. Using these PathBench subsets, we conduct extensive experiments to benchmark the performance of state-of-the-art LMMs at both the patch and WSI levels. The proposed dataset is available at https://github.com/superjamessyx/PathBench. Yuxuan Sun 0002, Hao Wu 0072, Chenglu Zhu, Yixuan Si, Qizi Chen, Kai Zhang 0033, Jingxiong Li, Jiatong Cai, Lin Sun 0006, Tao Lin 0004, Lin Yang 0002 |
IEEE Trans. Medical Imaging | 3 |
| 2024 | DPA-P2PNet: Deformable Proposal-Aware P2PNet for Accurate Point-Based Cell DetectionabstractPoint-based cell detection (PCD), which pursues high-performance cell sensing under low-cost data annotation, has garnered increased attention in computational pathology community. Unlike mainstream PCD methods that rely on intermediate density map representations, the Point-to-Point network (P2PNet) has recently emerged as an end-to-end solution for PCD, demonstrating impressive cell detection accuracy and efficiency. Nevertheless, P2PNet is limited to decoding from a single-level feature map due to the scale-agnostic property of point proposals, which is insufficient to leverage multi-scale information. Moreover, the spatial distribution of pre-set point proposals is biased from that of cells, leading to inaccurate cell localization. To lift these limitations, we present DPA-P2PNet in this work. The proposed method directly extracts multi-scale features for decoding according to the coordinates of point proposals on hierarchical feature maps. On this basis, we further devise deformable point proposals to mitigate the positional bias between proposals and potential cells to promote cell localization. Inspired by practical pathological diagnosis that usually combines high-level tissue structure and low-level cell morphology for accurate cell classification, we propose a multi-field-of-view (mFoV) variant of DPA-P2PNet to accommodate additional large FoV images with tissue information as model input. Finally, we execute the first self-supervised pre-training on immunohistochemistry histopathology image data and evaluate the suitability of four representative self-supervised methods on the PCD task. Experimental results on three benchmarks and a large-scale and real-world interval dataset demonstrate the superiority of our proposed models over the state-of-the-art counterparts. Codes and pre-trained weights are available at https://github.com/windygoo/DPA-P2PNet. Zhongyi Shui, Sunyi Zheng, Chenglu Zhu, Shichuan Zhang, Xiaoxuan Yu, Honglin Li 0001, Jingxiong Li, Pingyi Chen, Lin Yang 0002 |
AAAI | 3 |
| 2024 | PathAsst: A Generative Foundation AI Assistant towards Artificial General Intelligence of PathologyabstractAs advances in large language models (LLMs) and multimodal techniques continue to mature, the development of general-purpose multimodal large language models (MLLMs) has surged, offering significant applications in interpreting natural images. However, the field of pathology has largely remained untapped, particularly in gathering high-quality data and designing comprehensive model frameworks. To bridge the gap in pathology MLLMs, we present PathAsst, a multimodal generative foundation AI assistant to revolutionize diagnostic and predictive analytics in pathology. The development of PathAsst involves three pivotal steps: data acquisition, CLIP model adaptation, and the training of PathAsst's multimodal generative capabilities. Firstly, we collect over 207K high-quality pathology image-text pairs from authoritative sources. Leveraging the advanced power of ChatGPT, we generate over 180K instruction-following samples. Furthermore, we devise additional instruction-following data specifically tailored for invoking eight pathology-specific sub-models we prepared, allowing the PathAsst to effectively collaborate with these models, enhancing its diagnostic ability. Secondly, by leveraging the collected data, we construct PathCLIP, a pathology-dedicated CLIP, to enhance PathAsst's capabilities in interpreting pathology images. Finally, we integrate PathCLIP with the Vicuna-13b and utilize pathology-specific instruction-tuning data to enhance the multimodal generation capacity of PathAsst and bolster its synergistic interactions with sub-models. The experimental results of PathAsst show the potential of harnessing AI-powered generative foundation model to improve pathology diagnosis and treatment processes. We open-source our dataset, as well as a comprehensive toolkit for extensive pathology data collection and preprocessing at https://github.com/superjamessyx/Generative-Foundation-AI-Assistant-for-Pathology. Yuxuan Sun 0002, Chenglu Zhu, Sunyi Zheng, Kai Zhang 0033, Lin Sun 0006, Zhongyi Shui, Honglin Li 0001, Lin Yang 0002 |
AAAI | 2 |
| 2024 | WSI-VQA: Interpreting Whole Slide Images by Generative Visual Question Answering
Pingyi Chen, Chenglu Zhu, Sunyi Zheng, Honglin Li 0001, Lin Yang 0002 |
ECCV (36) | 2 |
| 2024 | Unleashing the Power of Prompt-Driven Nucleus Instance Segmentation
Zhongyi Shui, Chenglu Zhu, Sunyi Zheng, Jingxiong Li, Honglin Li 0001, Yuxuan Sun 0002, Ruizhe Guo, Lin Yang 0002 |
ECCV (27) | 4 |
| 2024 | PathMMU: A Massive Multimodal Expert-Level Benchmark for Understanding and Reasoning in Pathology
Yuxuan Sun 0002, Hao Wu 0072, Chenglu Zhu, Sunyi Zheng, Qizi Chen, Kai Zhang 0033, Dan Wan, Xiaoxiao Lan, Mengyue Zheng, Jingxiong Li, Xinheng Lyu, Tao Lin 0004, Lin Yang 0002 |
ECCV (62) | 3 |
| 2024 | Attention-Challenging Multiple Instance Learning for Whole Slide Image Classification
Honglin Li 0001, Yunxuan Sun, Sunyi Zheng, Chenglu Zhu, Lin Yang 0002 |
ECCV (53) | 5 |
| 2024 | Context-Aware Text-Assisted Multimodal Framework for Cervical Cytology Cell Diagnosis and ChattingabstractRecent advancements underscore the potential of deep learning-based Computer-Assisted Diagnosis (CAD) systems for cervical cytology image analysis. However, traditional methods focusing solely on single-view of cells fall short in performance due to the lack of contextual information. Moreover, the unclear reasoning behind model’s classification hinders their interpretability. To overcome these issues, we present Cervi-CAT, a context-aware, text-assisted multimodal framework for cervical cytology cell classification. CerviCAT captures visual cell representations from both global and local perspectives and subsequently generates textual descriptions based on the visual representation. A multimodal transformer then integrates these descriptions with visual features for interpretable and accurate cell classification. Additionally, we introduce Cyto-Vicuna, a cytology-specific large language model fine-tuned based on Vicuna-7b using collected cytology-specific data. When integrated into CerviCAT, it produces more detailed diagnostic reports while simultaneously fostering interaction between the model and cytologists, promoting collaborative diagnosis. Our results demonstrate that CerviCAT not only surpasses traditional CAD methods in performance but also provides interpretable diagnosis. Yuxuan Sun 0002, Chenglu Zhu, Sunyi Zheng, Honglin Li 0001, Lin Yang 0002 |
ICME | 2 |
| 2024 | WsiCaption: Multiple Instance Generation of Pathology Reports for Gigapixel Whole-Slide Images
Pingyi Chen, Honglin Li 0001, Chenglu Zhu, Sunyi Zheng, Zhongyi Shui, Lin Yang 0002 |
MICCAI (4) | 3 |
| 2024 | PathUp: Patch-wise Timestep Tracking for Multi-class Large Pathology Image Synthesising Diffusion ModelabstractIn digital pathology, cancer lesions are identified by analyzing the spatial context within pathology images. Synthesizing such complex spatial context is challenging as pathology whole slide images typically exhibit high resolution, low inter-class variety, and are sparsely labeled. To address these challenges, we propose PathUp, a novel diffusion model tailored for the synthesis of multi-class high-resolution pathology images. Our approach includes a latent space patch-wise timestep tracking, which helps to generate high-quality images without tiling artifacts. Pathology knowledge is integrated through our patho-align. The robust generation of lesion subtypes and scale information is ensured by introducing a feature entropy loss. The effectiveness of our method is evaluated through extensive experiments, supplemented by assessments from human experts, demonstrating the authenticity of the synthetic data produced. Furthermore, we highlight the potential utility of our generated images as an augmentation method, thereby enhancing the performance of downstream tasks such as cancer subtype classification. Jingxiong Li, Sunyi Zheng, Chenglu Zhu, Yuxuan Sun 0002, Pingyi Chen, Zhongyi Shui, Honglin Li 0001, Lin Yang 0002 |
ACM Multimedia | 3 |
| 2024 | Rethinking Transformer for Long Contextual Histopathology Whole Slide Image AnalysisabstractHistopathology Whole Slide Image (WSI) analysis serves as the gold standard for clinical cancer diagnosis in the daily routines of doctors. To develop computer-aided diagnosis model for histopathology WSIs, previous methods typically employ Multi-Instance Learning to enable slide-level prediction given only slide-level labels.
Among these models, vanilla attention mechanisms without pairwise interactions have traditionally been employed but are unable to model contextual information. More recently, self-attention models have been utilized to address this issue. To alleviate the computational complexity of long sequences in large WSIs, methods like HIPT use region-slicing, and TransMIL employs Nystr\"{o}mformer as an approximation of full self-attention. Both approaches suffer from suboptimal performance due to the loss of key information. Moreover, their use of absolute positional embedding struggles to effectively handle long contextual dependencies in shape-varying WSIs.
In this paper, we first analyze how the low-rank nature of the long-sequence attention matrix constrains the representation ability of WSI modelling. Then, we demonstrate that the rank of attention matrix can be improved by focusing on local interactions via a local attention mask. Our analysis shows that the local mask aligns with the attention patterns in the lower layers of the Transformer. Furthermore, the local attention mask can be implemented during chunked attention calculation, reducing the quadratic computational complexity to linear with a small local bandwidth. Additionally, this locality helps the model generalize to unseen or under-fitted positions more easily.
Building on this, we propose a local-global hybrid Transformer for both computational acceleration and local-global information interactions modelling. Our method, Long-contextual MIL (LongMIL), is evaluated through extensive experiments on various WSI tasks to validate its superiority in: 1) overall performance, 2) memory usage and speed, and 3) extrapolation ability compared to previous methods. Honglin Li 0001, Pingyi Chen, Zhongyi Shui, Chenglu Zhu, Lin Yang 0002 |
NeurIPS | 5 |
| 2024 | Gradient-aware learning for joint biases: Label noise and class imbalance
Shichuan Zhang, Chenglu Zhu, Honglin Li 0001, Jiatong Cai, Lin Yang 0002 |
Neural Networks | 2 |
| 2023 | Task-Specific Fine-Tuning via Variational Information Bottleneck for Weakly-Supervised Pathology Whole Slide Image ClassificationabstractWhile Multiple Instance Learning (MIL) has shown promising results in digital Pathology Whole Slide Image (WSI) analysis, such a paradigm still faces performance and generalization problems due to high computational costs and limited supervision of Gigapixel WSIs. To deal with the computation problem, previous methods utilize a frozen model pretrained from ImageNet to obtain representations, however, it may lose key information owing to the large domain gap and hinder the generalization ability without image-level training-time augmentation. Though Self-supervised Learning (SSL) proposes viable representation learning schemes, the downstream task-specific features via partial label tuning are not explored. To alleviate this problem, we propose an efficient WSI fine-tuning framework motivated by the Information Bottleneck theory. The theory enables the framework to find the minimal sufficient statistics of WSI, thus supporting us to fine-tune the backbone into a task-specific representation only depending on WSI-level weak labels. The WSI-MIL problem is further analyzed to theoretically deduce our fine-tuning method. We evaluate the method on five pathological WSI datasets on various WSI heads. The experimental results show significant improvements in both accuracy and generalization compared with previous works. Source code will be available at https://github.com/invoker-LL/WSI-finetuning. Honglin Li 0001, Chenglu Zhu, Yuxuan Sun 0002, Zhongyi Shui, Wenwei Kuang, Sunyi Zheng, Lin Yang 0002 |
CVPR | 2 |
| 2023 | Assessing the Robustness of Deep Learning-Assisted Pathological Image Analysis Under Practical Variables of Imaging SystemabstractWith the advancement of deep learning, computer-assisted clinical diagnosis, such as liquid-based cervical cytology, has attracted more attention. However, the fragile robustness of deep learning models has a non-negligible impact on their classification accuracy and reliability. To be more specific, various scanner parameters will be used depending on the pathologist’s preferences during the clinical diagnosis process (e.g., field source brightness, contrast, saturation, etc.), and this variation will lead to the unstable performance of the model. In this paper, we construct an evaluation pathway to assess the stability and consistency of deep learning models under various customized scanner parameters. Specifically, a multi-scanned dataset consists of 4200 whole slide images (WSIs) is generated by scanning 200 stained slices using various scanner parameters. Moreover, we conducted a large number of experiments to investigate the robustness of numerous models, including convolution-based and transformer-based models concerning various scanner parameter settings. Furthermore, we introduce several indicators to analyze the prediction accuracy, consistency and robustness of the model on the constructed dataset. The experimental results indicate that the deep learning models are sensitive to luminance-related scanner parameters. In addition, transformer-based models have better robustness than traditional convolutional neural networks. Our code has been made available1. Yuxuan Sun 0002, Chenglu Zhu, Honglin Li 0001, Pingyi Chen, Lin Yang 0002 |
ICASSP | 2 |
| 2023 | Exploring Unsupervised Cell Recognition with Prior Self-activation Maps
Pingyi Chen, Chenglu Zhu, Zhongyi Shui, Jiatong Cai, Sunyi Zheng, Shichuan Zhang, Lin Yang 0002 |
MICCAI (8) | 2 |
| 2022 | End-to-End Cell Recognition by Point Annotation
Zhongyi Shui, Shichuan Zhang, Chenglu Zhu, Bingchuan Wang, Pingyi Chen, Sunyi Zheng, Lin Yang 0002 |
MICCAI (4) | 3 |
| 2022 | Benchmarking the Robustness of Deep Neural Networks to Common Corruptions in Digital Pathology
Yuxuan Sun 0002, Honglin Li 0001, Sunyi Zheng, Chenglu Zhu, Lin Yang 0002 |
MICCAI (2) | 5 |
| 2022 | ChrSNet: Chromosome Straightening Using Self-attention Guided Networks
Sunyi Zheng, Jingxiong Li, Zhongyi Shui, Chenglu Zhu, Pingyi Chen, Lin Yang 0002 |
MICCAI (4) | 4 |
| 2022 | Automatic and accurate segmentation of peripherally inserted central catheter (PICC) from chest X-rays using multi-stage attention-guided learning
Xiaoyan Wang 0007, Ye Sheng, Chenglu Zhu, Cong Bai, Ming Xia 0005, Zhanpeng Shao, Ruiyi Zhao, Zhenjie Liu |
Neurocomputing | 4 |
| 2021 | Generalizing Nucleus Recognition Model in Multi-source Ki67 Immunohistochemistry Stained Images via Domain-Specific Pruning
Jiatong Cai, Chenglu Zhu, Honglin Li 0001, Shichuan Zhang, Lin Yang 0002 |
MICCAI (8) | 2 |