Sunyi Zheng

dblp:239/5250 · DBLP profile ↗
← Back
20ranked-venue papers
2as first author
19since 2021 · last 2026
0000-0002-9005-4875ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 1 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 8 since 2021
YearPublicationVenuePosition
2026 MIRA: Evaluating Multimodal AI on Complex Clinical Reasoning in Interventional Radiology
abstract
We present MIRA (Multimodal Interventional RAdiology evaluation), a comprehensive benchmark for evaluating large multimodal models in expert-level interventional radiology tasks requiring specialized domain knowledge and advanced visual reasoning capabilities. Unlike existing medical benchmarks that primarily provide binary labels without contextual depth, MIRA offers diverse question formats, including open-ended, closed-ended, single-choice, and multiple-choice categories, each accompanied by detailed expert-validated explanations. The benchmark incorporates approximately 184K high-quality medical images spanning multiple imaging modalities with 1.2M meticulously generated question-answer pairs across various anatomical regions. These pairs were created through a sophisticated cascade methodology involving expert interventional radiologists at both the data collection and validation stages. Our comprehensive evaluation, encompassing zero-shot testing and fine-tuning experiments of large multimodal models, revealing significant performance gaps between AI systems and human specialists. Fine-tuning experiments demonstrate substantial improvements, with models achieving up to 0.80 accuracy on single-choice questions. MIRA establishes a challenging benchmark that suggests promising directions for developing specialized clinical AI systems for interventional radiology.
Jingxiong Li, Chenglu Zhu, Sunyi Zheng, Yuxuan Sun 0002, Yixuan Si, Lin Yang 0002, Liang Xiao 0001
AAAI3
2025 ToPoFM: Topology-Guided Pathology Foundation Model for High-Resolution Pathology Image Synthesis With Cellular-Level Control
abstract
Synthetic data generation emerges as a strategy to mitigate data scarcity in digital pathology, where complicated tissue and cellular features are correlated with cancer diagnosis. The synthesis of such visuals, however, suffers from limited inter class diversity and scarcity of cellular annotations. Current methodologies struggle with capturing the broad spectrum of pathology features, causing unpredictable objects and defected fidelity. Moreover, discrepancies in image resolution across developmental and operational phases can amplify the distribution shifts, undermining the precision of diagnosis. To address these challenges, we introduce TOpology guided PathOlogy Foundation Model (ToPoFM), a visual foundation model designed for the synthesis of high-resolution pathology images with cellular-level control. Our approach integrates a topology-informed cell arrangement generator to steer large language models for crafting synthetic cell arrangements. We correlate cell arrangement guidance with diffusion model for pathology content generation, then further implement a random sliding inference strategy, merging discrete low-resolution samplings into single high-resolution representation. Our model requires only small patches for training. The efficacy of ToPoFM is demonstrated through extensive experiments, complemented by expert validations, showing high fidelity on data synthesis. Additionally, we underscore the utility of our generated imagery as an augmentation tool, enhancing the performance of downstream tasks, including cancer subtype classification and segmentation.
Jingxiong Li, Chenglu Zhu, Sunyi Zheng, Pingyi Chen, Yuxuan Sun 0002, Honglin Li 0001, Lin Yang 0002
IEEE Trans. Medical Imaging3
2024 DPA-P2PNet: Deformable Proposal-Aware P2PNet for Accurate Point-Based Cell Detection
abstract
Point-based cell detection (PCD), which pursues high-performance cell sensing under low-cost data annotation, has garnered increased attention in computational pathology community. Unlike mainstream PCD methods that rely on intermediate density map representations, the Point-to-Point network (P2PNet) has recently emerged as an end-to-end solution for PCD, demonstrating impressive cell detection accuracy and efficiency. Nevertheless, P2PNet is limited to decoding from a single-level feature map due to the scale-agnostic property of point proposals, which is insufficient to leverage multi-scale information. Moreover, the spatial distribution of pre-set point proposals is biased from that of cells, leading to inaccurate cell localization. To lift these limitations, we present DPA-P2PNet in this work. The proposed method directly extracts multi-scale features for decoding according to the coordinates of point proposals on hierarchical feature maps. On this basis, we further devise deformable point proposals to mitigate the positional bias between proposals and potential cells to promote cell localization. Inspired by practical pathological diagnosis that usually combines high-level tissue structure and low-level cell morphology for accurate cell classification, we propose a multi-field-of-view (mFoV) variant of DPA-P2PNet to accommodate additional large FoV images with tissue information as model input. Finally, we execute the first self-supervised pre-training on immunohistochemistry histopathology image data and evaluate the suitability of four representative self-supervised methods on the PCD task. Experimental results on three benchmarks and a large-scale and real-world interval dataset demonstrate the superiority of our proposed models over the state-of-the-art counterparts. Codes and pre-trained weights are available at https://github.com/windygoo/DPA-P2PNet.
Zhongyi Shui, Sunyi Zheng, Chenglu Zhu, Shichuan Zhang, Xiaoxuan Yu, Honglin Li 0001, Jingxiong Li, Pingyi Chen, Lin Yang 0002
AAAI2
2024 PathAsst: A Generative Foundation AI Assistant towards Artificial General Intelligence of Pathology
abstract
As advances in large language models (LLMs) and multimodal techniques continue to mature, the development of general-purpose multimodal large language models (MLLMs) has surged, offering significant applications in interpreting natural images. However, the field of pathology has largely remained untapped, particularly in gathering high-quality data and designing comprehensive model frameworks. To bridge the gap in pathology MLLMs, we present PathAsst, a multimodal generative foundation AI assistant to revolutionize diagnostic and predictive analytics in pathology. The development of PathAsst involves three pivotal steps: data acquisition, CLIP model adaptation, and the training of PathAsst's multimodal generative capabilities. Firstly, we collect over 207K high-quality pathology image-text pairs from authoritative sources. Leveraging the advanced power of ChatGPT, we generate over 180K instruction-following samples. Furthermore, we devise additional instruction-following data specifically tailored for invoking eight pathology-specific sub-models we prepared, allowing the PathAsst to effectively collaborate with these models, enhancing its diagnostic ability. Secondly, by leveraging the collected data, we construct PathCLIP, a pathology-dedicated CLIP, to enhance PathAsst's capabilities in interpreting pathology images. Finally, we integrate PathCLIP with the Vicuna-13b and utilize pathology-specific instruction-tuning data to enhance the multimodal generation capacity of PathAsst and bolster its synergistic interactions with sub-models. The experimental results of PathAsst show the potential of harnessing AI-powered generative foundation model to improve pathology diagnosis and treatment processes. We open-source our dataset, as well as a comprehensive toolkit for extensive pathology data collection and preprocessing at https://github.com/superjamessyx/Generative-Foundation-AI-Assistant-for-Pathology.
Yuxuan Sun 0002, Chenglu Zhu, Sunyi Zheng, Kai Zhang 0033, Lin Sun 0006, Zhongyi Shui, Honglin Li 0001, Lin Yang 0002
AAAI3
2024 WSI-VQA: Interpreting Whole Slide Images by Generative Visual Question Answering
Pingyi Chen, Chenglu Zhu, Sunyi Zheng, Honglin Li 0001, Lin Yang 0002
ECCV (36)3
2024 Unleashing the Power of Prompt-Driven Nucleus Instance Segmentation
Zhongyi Shui, Chenglu Zhu, Sunyi Zheng, Jingxiong Li, Honglin Li 0001, Yuxuan Sun 0002, Ruizhe Guo, Lin Yang 0002
ECCV (27)5
2024 PathMMU: A Massive Multimodal Expert-Level Benchmark for Understanding and Reasoning in Pathology
Yuxuan Sun 0002, Hao Wu 0072, Chenglu Zhu, Sunyi Zheng, Qizi Chen, Kai Zhang 0033, Dan Wan, Xiaoxiao Lan, Mengyue Zheng, Jingxiong Li, Xinheng Lyu, Tao Lin 0004, Lin Yang 0002
ECCV (62)4
2024 Attention-Challenging Multiple Instance Learning for Whole Slide Image Classification
Honglin Li 0001, Yunxuan Sun, Sunyi Zheng, Chenglu Zhu, Lin Yang 0002
ECCV (53)4
2024 HLS-FGVC: Hierarchical Label Semantics Enhanced Fine-Grained Visual Classification
abstract
Fine-grained visual classification (FGVC) intends to confirm the sub-classes of a specific object category, e.g., identifying the species of dogs or birds. It is a challenging problem with the inter-class similarity among these sub-categories and intra-class variance in every fine-grained class. Most of the recent works intend to learn discriminative representations and class-consistency features. However, they only take the finest labels into account. We argue that the hierarchical label structure (HLS) implied in the category names can enhance the FGVC task. In this paper, we proposed two modules to leverage the hierarchical label structure. (i) We build a weighted graph in each batch based on the hierarchical label structure, the nodes of which are image features. The messages are passed among graph nodes for feature interaction. (ii) A hierarchy-aware ranking loss is proposed to regularize the distribution in feature space. The ablation study and experimental results show that our proposed modules achieve significant improvements over previous works.
Shichuan Zhang, Sunyi Zheng, Zhongyi Shui, Lin Yang 0002
ICASSP2
2024 Context-Aware Text-Assisted Multimodal Framework for Cervical Cytology Cell Diagnosis and Chatting
abstract
Recent advancements underscore the potential of deep learning-based Computer-Assisted Diagnosis (CAD) systems for cervical cytology image analysis. However, traditional methods focusing solely on single-view of cells fall short in performance due to the lack of contextual information. Moreover, the unclear reasoning behind model’s classification hinders their interpretability. To overcome these issues, we present Cervi-CAT, a context-aware, text-assisted multimodal framework for cervical cytology cell classification. CerviCAT captures visual cell representations from both global and local perspectives and subsequently generates textual descriptions based on the visual representation. A multimodal transformer then integrates these descriptions with visual features for interpretable and accurate cell classification. Additionally, we introduce Cyto-Vicuna, a cytology-specific large language model fine-tuned based on Vicuna-7b using collected cytology-specific data. When integrated into CerviCAT, it produces more detailed diagnostic reports while simultaneously fostering interaction between the model and cytologists, promoting collaborative diagnosis. Our results demonstrate that CerviCAT not only surpasses traditional CAD methods in performance but also provides interpretable diagnosis.
Yuxuan Sun 0002, Chenglu Zhu, Sunyi Zheng, Honglin Li 0001, Lin Yang 0002
ICME3
2024 WsiCaption: Multiple Instance Generation of Pathology Reports for Gigapixel Whole-Slide Images
Pingyi Chen, Honglin Li 0001, Chenglu Zhu, Sunyi Zheng, Zhongyi Shui, Lin Yang 0002
MICCAI (4)4
2024 PathUp: Patch-wise Timestep Tracking for Multi-class Large Pathology Image Synthesising Diffusion Model
abstract
In digital pathology, cancer lesions are identified by analyzing the spatial context within pathology images. Synthesizing such complex spatial context is challenging as pathology whole slide images typically exhibit high resolution, low inter-class variety, and are sparsely labeled. To address these challenges, we propose PathUp, a novel diffusion model tailored for the synthesis of multi-class high-resolution pathology images. Our approach includes a latent space patch-wise timestep tracking, which helps to generate high-quality images without tiling artifacts. Pathology knowledge is integrated through our patho-align. The robust generation of lesion subtypes and scale information is ensured by introducing a feature entropy loss. The effectiveness of our method is evaluated through extensive experiments, supplemented by assessments from human experts, demonstrating the authenticity of the synthetic data produced. Furthermore, we highlight the potential utility of our generated images as an augmentation method, thereby enhancing the performance of downstream tasks such as cancer subtype classification.
Jingxiong Li, Sunyi Zheng, Chenglu Zhu, Yuxuan Sun 0002, Pingyi Chen, Zhongyi Shui, Honglin Li 0001, Lin Yang 0002
ACM Multimedia2
2024 Masked Conditional Variational Autoencoders for Chromosome Straightening
abstract
Karyotyping is of importance for detecting chromosomal aberrations in human disease. However, chromosomes easily appear curved in microscopic images, which prevents cytogeneticists from analyzing chromosome types. To address this issue, we propose a framework for chromosome straightening, which comprises a preliminary processing algorithm and a generative model called masked conditional variational autoencoders (MC-VAE). The processing method utilizes patch rearrangement to address the difficulty in erasing low degrees of curvature, providing reasonable preliminary results for the MC-VAE. The MC-VAE further straightens the results by leveraging chromosome patches conditioned on their curvatures to learn the mapping between banding patterns and conditions. During model training, we apply a masking strategy with a high masking ratio to train the MC-VAE with eliminated redundancy. This yields a non-trivial reconstruction task, allowing the model to effectively preserve chromosome banding patterns and structure details in the reconstructed results. Extensive experiments on three public datasets with two stain styles show that our framework surpasses the performance of state-of-the-art methods in retaining banding patterns and structure details. Compared to using real-world bent chromosomes, the use of high-quality straightened chromosomes generated by our proposed method can improve the performance of various deep learning models for chromosome classification by a large margin. Such a straightening approach has the potential to be combined with other karyotyping systems to assist cytogeneticists in chromosome analysis.
Jingxiong Li, Sunyi Zheng, Zhongyi Shui, Shichuan Zhang, Linyi Yang, Yuxuan Sun 0002, Honglin Li 0001, Yuanxin Ye, Peter M. A. van Ooijen, Kang Li 0004, Lin Yang 0002
IEEE Trans. Medical Imaging2
2023 Multi-modal Learning with Missing Modality in Predicting Axillary Lymph Node Metastasis
abstract
Multi-modal Learning has attracted widespread attention in medical image analysis. Using multi-modal data, whole slide images (WSIs) and clinical information, can improve the performance of deep learning models in the diagnosis of axillary lymph node metastasis. However, clinical information is not easy to collect in clinical practice due to privacy concerns, limited resources, lack of interoperability, etc. Although patient selection can ensure the training set to have multi-modal data for model development, missing modality of clinical information can appear during test. This normally leads to performance degradation, which limits the use of multi-modal models in the clinic. To alleviate this problem, we propose a bidirectional distillation framework consisting of a multi-modal branch and a single-modal branch. The single-modal branch acquires the complete multi-modal knowledge from the multi-modal branch, while the multi-modal learns the robust features of WSI from the single-modal. We conduct experiments on a public dataset of Lymph Node Metastasis in Early Breast Cancer to validate the method. Our approach not only achieves state-of-the-art performance with an AUC of 0.861 on the test set without missing data, but also yields an AUC of 0.842 when the rate of missing modality is 80%. This shows the effectiveness of the approach in dealing with multi-modal data and missing modality. Such a model has the potential to improve treatment decision-making for early breast cancer patients who have axillary lymph node metastatic status.
Shichuan Zhang, Sunyi Zheng, Zhongyi Shui, Honglin Li 0001, Lin Yang 0002
BIBM2
2023 Task-Specific Fine-Tuning via Variational Information Bottleneck for Weakly-Supervised Pathology Whole Slide Image Classification
abstract
While Multiple Instance Learning (MIL) has shown promising results in digital Pathology Whole Slide Image (WSI) analysis, such a paradigm still faces performance and generalization problems due to high computational costs and limited supervision of Gigapixel WSIs. To deal with the computation problem, previous methods utilize a frozen model pretrained from ImageNet to obtain representations, however, it may lose key information owing to the large domain gap and hinder the generalization ability without image-level training-time augmentation. Though Self-supervised Learning (SSL) proposes viable representation learning schemes, the downstream task-specific features via partial label tuning are not explored. To alleviate this problem, we propose an efficient WSI fine-tuning framework motivated by the Information Bottleneck theory. The theory enables the framework to find the minimal sufficient statistics of WSI, thus supporting us to fine-tune the backbone into a task-specific representation only depending on WSI-level weak labels. The WSI-MIL problem is further analyzed to theoretically deduce our fine-tuning method. We evaluate the method on five pathological WSI datasets on various WSI heads. The experimental results show significant improvements in both accuracy and generalization compared with previous works. Source code will be available at https://github.com/invoker-LL/WSI-finetuning.
Honglin Li 0001, Chenglu Zhu, Yuxuan Sun 0002, Zhongyi Shui, Wenwei Kuang, Sunyi Zheng, Lin Yang 0002
CVPR7
2023 Exploring Unsupervised Cell Recognition with Prior Self-activation Maps
Pingyi Chen, Chenglu Zhu, Zhongyi Shui, Jiatong Cai, Sunyi Zheng, Shichuan Zhang, Lin Yang 0002
MICCAI (8)5
2022 End-to-End Cell Recognition by Point Annotation
Zhongyi Shui, Shichuan Zhang, Chenglu Zhu, Bingchuan Wang, Pingyi Chen, Sunyi Zheng, Lin Yang 0002
MICCAI (4)6
2022 Benchmarking the Robustness of Deep Neural Networks to Common Corruptions in Digital Pathology
Yuxuan Sun 0002, Honglin Li 0001, Sunyi Zheng, Chenglu Zhu, Lin Yang 0002
MICCAI (2)4
2022 ChrSNet: Chromosome Straightening Using Self-attention Guided Networks
Sunyi Zheng, Jingxiong Li, Zhongyi Shui, Chenglu Zhu, Pingyi Chen, Lin Yang 0002
MICCAI (4)1
2020 Automatic Pulmonary Nodule Detection in CT Scans Using Convolutional Neural Networks Based on Maximum Intensity Projection
abstract
Accurate pulmonary nodule detection is a crucial step in lung cancer screening. Computer-aided detection (CAD) systems are not routinely used by radiologists for pulmonary nodule detection in clinical practice despite their potential benefits. Maximum intensity projection (MIP) images improve the detection of pulmonary nodules in radiological evaluation with computed tomography (CT) scans. Inspired by the clinical methodology of radiologists, we aim to explore the feasibility of applying MIP images to improve the effectiveness of automatic lung nodule detection using convolutional neural networks (CNNs). We propose a CNN-based approach that takes MIP images of different slab thicknesses (5 mm, 10 mm, 15 mm) and 1 mm axial section slices as input. Such an approach augments the two-dimensional (2-D) CT slice images with more representative spatial information that helps discriminate nodules from vessels through their morphologies. Our proposed method achieves sensitivity of 92.7% with 1 false positive per scan and sensitivity of 94.2% with 2 false positives per scan for lung nodule detection on 888 scans in the LIDC-IDRI dataset. The use of thick MIP images helps the detection of small pulmonary nodules (3 mm-10 mm) and results in fewer false positives. Experimental results show that utilizing MIP images can increase the sensitivity and lower the number of false positives, which demonstrates the effectiveness and significance of the proposed MIP-based CNNs framework for automatic pulmonary nodule detection in CT scans. The proposed method also shows the potential that CNNs could gain benefits for nodule detection by combining the clinical procedure.
Sunyi Zheng, Jiapan Guo, Xiaonan Cui, Raymond N. J. Veldhuis, Matthijs Oudkerk, Peter M. A. van Ooijen
IEEE Trans. Medical Imaging1