EDBT 2026 Demo / reviewers in the wild / expert
Honglin Li 0001
dblp:67/5566-1
· DBLP profile ↗
21ranked-venue papers
3as first author
21since 2021 · last 2026
0009-0006-8328-7132ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 15 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Effective and Efficient Context-aware Nucleus Detection in Histopathology Whole Slide ImagesabstractNucleus detection in histopathology whole slide images (WSIs) is crucial for a broad spectrum of clinical applications. The gigapixel size of WSIs necessitates the use of sliding window methodology for nucleus detection. However, mainstream methods process each sliding window independently, which overlooks broader contextual information and easily leads to inaccurate predictions. To address this limitation, recent studies additionally crop a large Filed-of-View (LFoV) patch centered on each sliding window to extract contextual features. However, such methods substantially increase whole-slide inference latency. In this work, we propose an effective and efficient context-aware nucleus detection approach. Specifically, instead of using lFoV patches, we aggregate contextual clues from off-the-shelf features of historically visited sliding windows, which greatly enhances the inference efficiency. Moreover, compared to lFoV patches used in previous works, the sliding window patches have higher magnification and provide finer-grained tissue details, thereby enhancing the classification accuracy. To develop the proposed context-aware model, we utilize annotated patches along with their surrounding unlabeled patches for training. Beyond exploiting high-level tissue context from these surrounding regions, we design a post-training strategy that leverages abundant unlabeled nucleus samples within them to enhance the model's context adaptability. Extensive experimental results on three challenging benchmarks demonstrate the superiority of our method. Zhongyi Shui, Honglin Li 0001, Yuxuan Sun 0002, Yiwen Ye, Pingyi Chen, Ruizhe Guo, Lei Cui 0004, Chenglu Zhu, Lin Yang 0002 |
AAAI | 2 |
| 2025 | Stable Test-Time Training for Semantic Segmentation with Output Contrastive LossabstractDeep learning-based models have achieved impressive performance on public segmentation benchmarks, yet generalizing to unseen environments remains challenging. Test-time training (TTT) addresses this by adapting source-pretrained models during evaluation. While existing TTT methods have shown promise in image classification, they often exhibit instability with small test batches and class imbalance—challenges that intensify in semantic segmentation tasks. To tackle this issue, we present Output Contrastive Loss (OCL) to improve the stability of contrastive loss when applied to TTT for segmentation. OCL applies contrastive loss directly to the output space, avoiding the need for extra regularization, and employs a high temperature to prevent model collapse. To further stabilize the TTT process, we integrate BN statistics Modulation and Stochastic Restoration techniques. Extensive experiments across diverse datasets, settings, architectures, and pretrained methods demonstrate consistent performance improvements, achieving a 7.5 mIoU gain on the GTA→CS benchmark and showing effectiveness even with domain adaptation pretraining. Code is available at https://github.com/dazhangyul23/OCL. Zhongyi Shui, Honglin Li 0001, Yuxuan Sun 0002, Chenglu Zhu, Lin Yang 0002 |
ICASSP | 3 |
| 2025 | AEM: Attention Entropy Maximization for Multiple Instance Learning Based Whole Slide Image Classification
Honglin Li 0001, Yuxuan Sun 0002, Zhongyi Shui, Jingxiong Li, Chenglu Zhu, Lin Yang 0002 |
MICCAI (7) | 2 |
| 2025 | PathVQ: Reforming Computational Pathology Foundation Model for Whole Slide Image Analysis via Vector QuantizationabstractPathology whole slide image (WSI) analysis is vital for disease diagnosis and understanding. While foundation models (FMs) have driven recent advances, their scalability in pathology remains a key challenge. In particular, vision-language (VL) pathology FMs align visual features with language annotation for downstream tasks, but they rely heavily on large-scale image-text paired data, which is scarce thus limiting generalization. On the other hand, vision-only pathology FMs can leverage abundant unlabeled data via self-supervised learning (SSL). However, current approaches often use the [CLS] token from tile-level ViTs as slide-level input for efficiency (a tile with 224×224 pixels composed of 196 patches with 16×16 pixels). This SSL pretrained [CLS] token lacks alignment with downstream objectives, limiting effectiveness. We find that spatial patch tokens retain a wealth of informative features beneficial for downstream tasks, but utilizing all of them incurs up to 200× higher computation and storage costs compared [CLS] token only (e.g., 196 tokens per ViT$_{224}$). This highlights a fundamental trade-off between efficiency and representational richness to build scalable pathology FMs. To address this, we propose a feature distillation framework via vector-quantization (VQ) that compresses patch tokens into discrete indices and reconstructs them via a decoder, achieving 64× compression (1024 → 16 dimensions) while preserving fidelity. We further introduce a multi-scale VQ (MSVQ) strategy, enhancing both reconstruction and providing SSL supervision for slide-level pretraining. Built upon MSVQ features and supervision signals, we design a progressive convolutional module and a slide-level SSL objective to learn spatially rich representations for downstream WSI tasks. Extensive experiments across multiple datasets demonstrate that our approach achieves state-of-the-art performance, offering a scalable and effective solution for high-performing pathology FMs in WSI analysis. Honglin Li 0001, Zhongyi Shui, Chenglu Zhu, Lin Yang 0002 |
NeurIPS | 1 |
| 2025 | ToPoFM: Topology-Guided Pathology Foundation Model for High-Resolution Pathology Image Synthesis With Cellular-Level ControlabstractSynthetic data generation emerges as a strategy to mitigate data scarcity in digital pathology, where complicated tissue and cellular features are correlated with cancer diagnosis. The synthesis of such visuals, however, suffers from limited inter class diversity and scarcity of cellular annotations. Current methodologies struggle with capturing the broad spectrum of pathology features, causing unpredictable objects and defected fidelity. Moreover, discrepancies in image resolution across developmental and operational phases can amplify the distribution shifts, undermining the precision of diagnosis. To address these challenges, we introduce TOpology guided PathOlogy Foundation Model (ToPoFM), a visual foundation model designed for the synthesis of high-resolution pathology images with cellular-level control. Our approach integrates a topology-informed cell arrangement generator to steer large language models for crafting synthetic cell arrangements. We correlate cell arrangement guidance with diffusion model for pathology content generation, then further implement a random sliding inference strategy, merging discrete low-resolution samplings into single high-resolution representation. Our model requires only small patches for training. The efficacy of ToPoFM is demonstrated through extensive experiments, complemented by expert validations, showing high fidelity on data synthesis. Additionally, we underscore the utility of our generated imagery as an augmentation tool, enhancing the performance of downstream tasks, including cancer subtype classification and segmentation. Jingxiong Li, Chenglu Zhu, Sunyi Zheng, Pingyi Chen, Yuxuan Sun 0002, Honglin Li 0001, Lin Yang 0002 |
IEEE Trans. Medical Imaging | 6 |
| 2024 | DPA-P2PNet: Deformable Proposal-Aware P2PNet for Accurate Point-Based Cell DetectionabstractPoint-based cell detection (PCD), which pursues high-performance cell sensing under low-cost data annotation, has garnered increased attention in computational pathology community. Unlike mainstream PCD methods that rely on intermediate density map representations, the Point-to-Point network (P2PNet) has recently emerged as an end-to-end solution for PCD, demonstrating impressive cell detection accuracy and efficiency. Nevertheless, P2PNet is limited to decoding from a single-level feature map due to the scale-agnostic property of point proposals, which is insufficient to leverage multi-scale information. Moreover, the spatial distribution of pre-set point proposals is biased from that of cells, leading to inaccurate cell localization. To lift these limitations, we present DPA-P2PNet in this work. The proposed method directly extracts multi-scale features for decoding according to the coordinates of point proposals on hierarchical feature maps. On this basis, we further devise deformable point proposals to mitigate the positional bias between proposals and potential cells to promote cell localization. Inspired by practical pathological diagnosis that usually combines high-level tissue structure and low-level cell morphology for accurate cell classification, we propose a multi-field-of-view (mFoV) variant of DPA-P2PNet to accommodate additional large FoV images with tissue information as model input. Finally, we execute the first self-supervised pre-training on immunohistochemistry histopathology image data and evaluate the suitability of four representative self-supervised methods on the PCD task. Experimental results on three benchmarks and a large-scale and real-world interval dataset demonstrate the superiority of our proposed models over the state-of-the-art counterparts. Codes and pre-trained weights are available at https://github.com/windygoo/DPA-P2PNet. Zhongyi Shui, Sunyi Zheng, Chenglu Zhu, Shichuan Zhang, Xiaoxuan Yu, Honglin Li 0001, Jingxiong Li, Pingyi Chen, Lin Yang 0002 |
AAAI | 6 |
| 2024 | PathAsst: A Generative Foundation AI Assistant towards Artificial General Intelligence of PathologyabstractAs advances in large language models (LLMs) and multimodal techniques continue to mature, the development of general-purpose multimodal large language models (MLLMs) has surged, offering significant applications in interpreting natural images. However, the field of pathology has largely remained untapped, particularly in gathering high-quality data and designing comprehensive model frameworks. To bridge the gap in pathology MLLMs, we present PathAsst, a multimodal generative foundation AI assistant to revolutionize diagnostic and predictive analytics in pathology. The development of PathAsst involves three pivotal steps: data acquisition, CLIP model adaptation, and the training of PathAsst's multimodal generative capabilities. Firstly, we collect over 207K high-quality pathology image-text pairs from authoritative sources. Leveraging the advanced power of ChatGPT, we generate over 180K instruction-following samples. Furthermore, we devise additional instruction-following data specifically tailored for invoking eight pathology-specific sub-models we prepared, allowing the PathAsst to effectively collaborate with these models, enhancing its diagnostic ability. Secondly, by leveraging the collected data, we construct PathCLIP, a pathology-dedicated CLIP, to enhance PathAsst's capabilities in interpreting pathology images. Finally, we integrate PathCLIP with the Vicuna-13b and utilize pathology-specific instruction-tuning data to enhance the multimodal generation capacity of PathAsst and bolster its synergistic interactions with sub-models. The experimental results of PathAsst show the potential of harnessing AI-powered generative foundation model to improve pathology diagnosis and treatment processes. We open-source our dataset, as well as a comprehensive toolkit for extensive pathology data collection and preprocessing at https://github.com/superjamessyx/Generative-Foundation-AI-Assistant-for-Pathology. Yuxuan Sun 0002, Chenglu Zhu, Sunyi Zheng, Kai Zhang 0033, Lin Sun 0006, Zhongyi Shui, Honglin Li 0001, Lin Yang 0002 |
AAAI | 8 |
| 2024 | WSI-VQA: Interpreting Whole Slide Images by Generative Visual Question Answering
Pingyi Chen, Chenglu Zhu, Sunyi Zheng, Honglin Li 0001, Lin Yang 0002 |
ECCV (36) | 4 |
| 2024 | Unleashing the Power of Prompt-Driven Nucleus Instance Segmentation
Zhongyi Shui, Chenglu Zhu, Sunyi Zheng, Jingxiong Li, Honglin Li 0001, Yuxuan Sun 0002, Ruizhe Guo, Lin Yang 0002 |
ECCV (27) | 7 |
| 2024 | Attention-Challenging Multiple Instance Learning for Whole Slide Image Classification
Honglin Li 0001, Yunxuan Sun, Sunyi Zheng, Chenglu Zhu, Lin Yang 0002 |
ECCV (53) | 2 |
| 2024 | Context-Aware Text-Assisted Multimodal Framework for Cervical Cytology Cell Diagnosis and ChattingabstractRecent advancements underscore the potential of deep learning-based Computer-Assisted Diagnosis (CAD) systems for cervical cytology image analysis. However, traditional methods focusing solely on single-view of cells fall short in performance due to the lack of contextual information. Moreover, the unclear reasoning behind model’s classification hinders their interpretability. To overcome these issues, we present Cervi-CAT, a context-aware, text-assisted multimodal framework for cervical cytology cell classification. CerviCAT captures visual cell representations from both global and local perspectives and subsequently generates textual descriptions based on the visual representation. A multimodal transformer then integrates these descriptions with visual features for interpretable and accurate cell classification. Additionally, we introduce Cyto-Vicuna, a cytology-specific large language model fine-tuned based on Vicuna-7b using collected cytology-specific data. When integrated into CerviCAT, it produces more detailed diagnostic reports while simultaneously fostering interaction between the model and cytologists, promoting collaborative diagnosis. Our results demonstrate that CerviCAT not only surpasses traditional CAD methods in performance but also provides interpretable diagnosis. Yuxuan Sun 0002, Chenglu Zhu, Sunyi Zheng, Honglin Li 0001, Lin Yang 0002 |
ICME | 5 |
| 2024 | WsiCaption: Multiple Instance Generation of Pathology Reports for Gigapixel Whole-Slide Images
Pingyi Chen, Honglin Li 0001, Chenglu Zhu, Sunyi Zheng, Zhongyi Shui, Lin Yang 0002 |
MICCAI (4) | 2 |
| 2024 | PathUp: Patch-wise Timestep Tracking for Multi-class Large Pathology Image Synthesising Diffusion ModelabstractIn digital pathology, cancer lesions are identified by analyzing the spatial context within pathology images. Synthesizing such complex spatial context is challenging as pathology whole slide images typically exhibit high resolution, low inter-class variety, and are sparsely labeled. To address these challenges, we propose PathUp, a novel diffusion model tailored for the synthesis of multi-class high-resolution pathology images. Our approach includes a latent space patch-wise timestep tracking, which helps to generate high-quality images without tiling artifacts. Pathology knowledge is integrated through our patho-align. The robust generation of lesion subtypes and scale information is ensured by introducing a feature entropy loss. The effectiveness of our method is evaluated through extensive experiments, supplemented by assessments from human experts, demonstrating the authenticity of the synthetic data produced. Furthermore, we highlight the potential utility of our generated images as an augmentation method, thereby enhancing the performance of downstream tasks such as cancer subtype classification. Jingxiong Li, Sunyi Zheng, Chenglu Zhu, Yuxuan Sun 0002, Pingyi Chen, Zhongyi Shui, Honglin Li 0001, Lin Yang 0002 |
ACM Multimedia | 8 |
| 2024 | Rethinking Transformer for Long Contextual Histopathology Whole Slide Image AnalysisabstractHistopathology Whole Slide Image (WSI) analysis serves as the gold standard for clinical cancer diagnosis in the daily routines of doctors. To develop computer-aided diagnosis model for histopathology WSIs, previous methods typically employ Multi-Instance Learning to enable slide-level prediction given only slide-level labels.
Among these models, vanilla attention mechanisms without pairwise interactions have traditionally been employed but are unable to model contextual information. More recently, self-attention models have been utilized to address this issue. To alleviate the computational complexity of long sequences in large WSIs, methods like HIPT use region-slicing, and TransMIL employs Nystr\"{o}mformer as an approximation of full self-attention. Both approaches suffer from suboptimal performance due to the loss of key information. Moreover, their use of absolute positional embedding struggles to effectively handle long contextual dependencies in shape-varying WSIs.
In this paper, we first analyze how the low-rank nature of the long-sequence attention matrix constrains the representation ability of WSI modelling. Then, we demonstrate that the rank of attention matrix can be improved by focusing on local interactions via a local attention mask. Our analysis shows that the local mask aligns with the attention patterns in the lower layers of the Transformer. Furthermore, the local attention mask can be implemented during chunked attention calculation, reducing the quadratic computational complexity to linear with a small local bandwidth. Additionally, this locality helps the model generalize to unseen or under-fitted positions more easily.
Building on this, we propose a local-global hybrid Transformer for both computational acceleration and local-global information interactions modelling. Our method, Long-contextual MIL (LongMIL), is evaluated through extensive experiments on various WSI tasks to validate its superiority in: 1) overall performance, 2) memory usage and speed, and 3) extrapolation ability compared to previous methods. Honglin Li 0001, Pingyi Chen, Zhongyi Shui, Chenglu Zhu, Lin Yang 0002 |
NeurIPS | 1 |
| 2024 | Gradient-aware learning for joint biases: Label noise and class imbalance
Shichuan Zhang, Chenglu Zhu, Honglin Li 0001, Jiatong Cai, Lin Yang 0002 |
Neural Networks | 3 |
| 2024 | Masked Conditional Variational Autoencoders for Chromosome StraighteningabstractKaryotyping is of importance for detecting chromosomal aberrations in human disease. However, chromosomes easily appear curved in microscopic images, which prevents cytogeneticists from analyzing chromosome types. To address this issue, we propose a framework for chromosome straightening, which comprises a preliminary processing algorithm and a generative model called masked conditional variational autoencoders (MC-VAE). The processing method utilizes patch rearrangement to address the difficulty in erasing low degrees of curvature, providing reasonable preliminary results for the MC-VAE. The MC-VAE further straightens the results by leveraging chromosome patches conditioned on their curvatures to learn the mapping between banding patterns and conditions. During model training, we apply a masking strategy with a high masking ratio to train the MC-VAE with eliminated redundancy. This yields a non-trivial reconstruction task, allowing the model to effectively preserve chromosome banding patterns and structure details in the reconstructed results. Extensive experiments on three public datasets with two stain styles show that our framework surpasses the performance of state-of-the-art methods in retaining banding patterns and structure details. Compared to using real-world bent chromosomes, the use of high-quality straightened chromosomes generated by our proposed method can improve the performance of various deep learning models for chromosome classification by a large margin. Such a straightening approach has the potential to be combined with other karyotyping systems to assist cytogeneticists in chromosome analysis. Jingxiong Li, Sunyi Zheng, Zhongyi Shui, Shichuan Zhang, Linyi Yang, Yuxuan Sun 0002, Honglin Li 0001, Yuanxin Ye, Peter M. A. van Ooijen, Kang Li 0004, Lin Yang 0002 |
IEEE Trans. Medical Imaging | 8 |
| 2023 | Multi-modal Learning with Missing Modality in Predicting Axillary Lymph Node MetastasisabstractMulti-modal Learning has attracted widespread attention in medical image analysis. Using multi-modal data, whole slide images (WSIs) and clinical information, can improve the performance of deep learning models in the diagnosis of axillary lymph node metastasis. However, clinical information is not easy to collect in clinical practice due to privacy concerns, limited resources, lack of interoperability, etc. Although patient selection can ensure the training set to have multi-modal data for model development, missing modality of clinical information can appear during test. This normally leads to performance degradation, which limits the use of multi-modal models in the clinic. To alleviate this problem, we propose a bidirectional distillation framework consisting of a multi-modal branch and a single-modal branch. The single-modal branch acquires the complete multi-modal knowledge from the multi-modal branch, while the multi-modal learns the robust features of WSI from the single-modal. We conduct experiments on a public dataset of Lymph Node Metastasis in Early Breast Cancer to validate the method. Our approach not only achieves state-of-the-art performance with an AUC of 0.861 on the test set without missing data, but also yields an AUC of 0.842 when the rate of missing modality is 80%. This shows the effectiveness of the approach in dealing with multi-modal data and missing modality. Such a model has the potential to improve treatment decision-making for early breast cancer patients who have axillary lymph node metastatic status. Shichuan Zhang, Sunyi Zheng, Zhongyi Shui, Honglin Li 0001, Lin Yang 0002 |
BIBM | 4 |
| 2023 | Task-Specific Fine-Tuning via Variational Information Bottleneck for Weakly-Supervised Pathology Whole Slide Image ClassificationabstractWhile Multiple Instance Learning (MIL) has shown promising results in digital Pathology Whole Slide Image (WSI) analysis, such a paradigm still faces performance and generalization problems due to high computational costs and limited supervision of Gigapixel WSIs. To deal with the computation problem, previous methods utilize a frozen model pretrained from ImageNet to obtain representations, however, it may lose key information owing to the large domain gap and hinder the generalization ability without image-level training-time augmentation. Though Self-supervised Learning (SSL) proposes viable representation learning schemes, the downstream task-specific features via partial label tuning are not explored. To alleviate this problem, we propose an efficient WSI fine-tuning framework motivated by the Information Bottleneck theory. The theory enables the framework to find the minimal sufficient statistics of WSI, thus supporting us to fine-tune the backbone into a task-specific representation only depending on WSI-level weak labels. The WSI-MIL problem is further analyzed to theoretically deduce our fine-tuning method. We evaluate the method on five pathological WSI datasets on various WSI heads. The experimental results show significant improvements in both accuracy and generalization compared with previous works. Source code will be available at https://github.com/invoker-LL/WSI-finetuning. Honglin Li 0001, Chenglu Zhu, Yuxuan Sun 0002, Zhongyi Shui, Wenwei Kuang, Sunyi Zheng, Lin Yang 0002 |
CVPR | 1 |
| 2023 | Assessing the Robustness of Deep Learning-Assisted Pathological Image Analysis Under Practical Variables of Imaging SystemabstractWith the advancement of deep learning, computer-assisted clinical diagnosis, such as liquid-based cervical cytology, has attracted more attention. However, the fragile robustness of deep learning models has a non-negligible impact on their classification accuracy and reliability. To be more specific, various scanner parameters will be used depending on the pathologist’s preferences during the clinical diagnosis process (e.g., field source brightness, contrast, saturation, etc.), and this variation will lead to the unstable performance of the model. In this paper, we construct an evaluation pathway to assess the stability and consistency of deep learning models under various customized scanner parameters. Specifically, a multi-scanned dataset consists of 4200 whole slide images (WSIs) is generated by scanning 200 stained slices using various scanner parameters. Moreover, we conducted a large number of experiments to investigate the robustness of numerous models, including convolution-based and transformer-based models concerning various scanner parameter settings. Furthermore, we introduce several indicators to analyze the prediction accuracy, consistency and robustness of the model on the constructed dataset. The experimental results indicate that the deep learning models are sensitive to luminance-related scanner parameters. In addition, transformer-based models have better robustness than traditional convolutional neural networks. Our code has been made available1. Yuxuan Sun 0002, Chenglu Zhu, Honglin Li 0001, Pingyi Chen, Lin Yang 0002 |
ICASSP | 4 |
| 2022 | Benchmarking the Robustness of Deep Neural Networks to Common Corruptions in Digital Pathology
Yuxuan Sun 0002, Honglin Li 0001, Sunyi Zheng, Chenglu Zhu, Lin Yang 0002 |
MICCAI (2) | 3 |
| 2021 | Generalizing Nucleus Recognition Model in Multi-source Ki67 Immunohistochemistry Stained Images via Domain-Specific Pruning
Jiatong Cai, Chenglu Zhu, Honglin Li 0001, Shichuan Zhang, Lin Yang 0002 |
MICCAI (8) | 4 |