EDBT 2026 Demo / reviewers in the wild / expert
Shuo Wang 0011
dblp:63/1591-11
· DBLP profile ↗
38ranked-venue papers
2as first author
32since 2021 · last 2026
0000-0002-2947-8783ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 28 · 2 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 2 first-author · 15 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Extreme cardiac MRI analysis under respiratory motion: Results of the CMRxMotion challenge
Kang Wang 0017, Chen Qin, Zhang Shi, Haoran Wang 0009, Chen Chen 0042, Cheng Ouyang, Chengliang Dai, Yuanhan Mo, Chenchen Dai, Xutong Kuang, Ruizhe Li 0005, Xin Chen 0003, Xiuzheng Yue, Song Tian, Alejandro Mora-Rubio, Kumaradevan Punithakumar, Shizhan Gong, Qi Dou 0001, Sina Amirrajab, Yasmina Alkhalil, Cian M. Scannell, Lexiaozi Fan, Huili Yang, Xiaowu Sun, Rob J. van der Geest, Tewodros Weldebirhan Arega, Fabrice Mériaudeau, Caner Ozer, Amin Ranem, John Kalkhof, Ilkay Öksüz, Anirban Mukhopadhyay 0003, Abdul Qayyum 0002, Moona Mazher, Steven A. Niederer, Carles García-Cabrera, Eric Arazo Sanchez, Michal K. Grzeszczyk, Szymon Plotka, Wanqin Ma, Xiaomeng Li 0001, Rongjun Ge, Yongqing Kou, Xinrong Chen, He Wang 0016, Chengyan Wang, Wenjia Bai, Shuo Wang 0011 |
Medical Image Anal. | 49 |
| 2026 | DiCLIP: Diffusion Model Enhances CLIP's Dense Knowledge for Weakly Supervised Semantic SegmentationabstractWeakly Supervised Semantic Segmentation (WSSS) with image-level labels typically leverages Class Activation Maps (CAMs) to achieve pixel-level predictions. Recently, Contrastive Language-Image Pre-training (CLIP) has been introduced to generate CAMs in WSSS. However, previous WSSS methods solely adopt CLIP's vision-language paired property for dense localization, neglecting its inherently limited dense knowledge across both visual and text modalities, which renders CAM generation suboptimal. In this work, we propose DiCLIP, a novel WSSS framework that leverages the generative diffusion model to enhance CLIP's dense knowledge across two modalities. Specifically, Visual Correlation Enhancement (VCE) and Text Semantic Augmentation (TSA) modules are proposed for dense prediction enhancement. To improve the spatial awareness of visual features, our VCE module utilizes diffusion's reliable spatial consistency to mitigate the over-smoothing issue in CLIP's attention. It designs the Attention Clustering Refinement (ACR) module to reliably extract diverse correlation maps from the diffusion model. The correlation maps act as a diversity bias for CLIP's self-attention, recursively pushing its visual features towards a more discriminative dense distribution. To augment the semantics of text embeddings, our TSA module argues that a single text modality is insufficient to encompass the variability of visual categories. Thus, we leverage diffusion's generative power to maintain a dynamic key-value cache model, shifting CAM generation from a patch-text matching mechanism to a novel visual knowledge retrieval paradigm. With these enhancements, DiCLIP not only outperforms state-of-the-art methods on PASCAL VOC and MS COCO but also significantly reduces training costs. Code is publicly available at https://github.com/zwyang6/DiCLIP. Yucong Meng, Kexue Fu 0001, Shuo Wang 0011, Zhijian Song |
IEEE Trans. Image Process. | 5 |
| 2026 | Toward Modality- and Sampling-Universal Learning Strategies for Accelerating Cardiovascular Imaging: Summary of the CMRxRecon2024 ChallengeabstractCardiovascular health is vital to human well-being, and cardiac magnetic resonance (CMR) imaging is considered the clinical reference standard for diagnosing cardiovascular disease. However, its adoption is hindered by long scan times, complex contrasts, and inconsistent quality. While deep learning methods perform well on specific CMR imaging sequences, they often fail to generalize across modalities and sampling schemes. The lack of benchmarks for high-quality, fast CMR image reconstruction further limits technology comparison and adoption. The CMRxRecon2024 challenge, attracting over 200 teams from 18 countries, addressed these issues with two tasks: generalization to unseen modalities and robustness to diverse undersampling patterns. We introduced the largest public multi-modality CMR raw dataset, an open benchmarking platform, and shared code. Analysis of the best-performing solutions revealed that prompt-based adaptation and enhanced physics-driven consistency enabled strong cross-scenario performance. These findings establish principles for generalizable reconstruction models and advance clinically translatable AI in cardiovascular imaging. Fanwen Wang, Zi Wang 0005, Yan Li 0064, Chen Qin, Shuo Wang 0011, Kunyuan Guo, Mengting Sun, Mingkai Huang, Michael Tänzer, Qirong Li, Yinzhe Wu 0001, Haosen Zhang, Kian Anvari Hamedani, Yuntong Lyu, Longyu Sun, Tianxing He, Lizhen Lan, Qiong Yao, Bingyu Xin, Dimitris N. Metaxas, Narges Razizadeh, Shahabedin Nabavi, George Yiasemis, Jonas Teuwen, Daniel B. Ennis, Zhihao Xue, Ruru Xu, Ilkay Öksüz, Donghang Lyu, Yanxin Huang, Xinrui Guo, Ruqian Hao, Jaykumar H. Patel, Guanke Cai, Binghua Chen, Sha Hua, Zhensen Chen, Qi Dou 0001, Xiahai Zhuang, Wenjia Bai, Harry Qin, He Wang 0016, Claudia Prieto, Michael Markl 0001, Alistair A. Young, Hao Li 0082, Xihong Hu, Lianming Wu, Xiaobo Qu 0001, Guang Yang 0006, Chengyan Wang |
IEEE Trans. Medical Imaging | 6 |
| 2025 | MoRe: Class Patch Attention Needs Regularization for Weakly Supervised Semantic SegmentationabstractWeakly Supervised Semantic Segmentation (WSSS) with image-level labels typically uses Class Activation Maps (CAM) to achieve dense predictions. Recently, Vision Transformer (ViT) has provided an alternative to generate localization maps from class-patch attention. However, due to insufficient constraints on modeling such attention, we observe that the Localization Attention Maps (LAM) often struggle with the artifact issue, i.e., patch regions with minimal semantic relevance are falsely activated by class tokens. In this work, we propose MoRe to address this issue and further explore the potential of LAM. Our findings suggest that imposing additional regularization on class-patch attention is necessary. To this end, we first view the attention as a novel directed graph and propose the Graph Category Representation module to implicitly regularize the interaction among class-patch entities. It ensures that class tokens dynamically condense the related patch information and suppress unrelated artifacts at a graph level. Second, motivated by the observation that CAM from classification weights maintains smooth localization of objects, we devise the Localization-informed Regularization module to explicitly regularize the class-patch attention. It directly mines the token relations from CAM and further supervises the consistency between class and patch tokens in a learnable manner. Extensive experiments on PASCAL VOC and MS COCO validate that MoRe effectively addresses the artifact issue and achieves state-of-the-art performance, surpassing recent single-stage and even multi-stage methods. Yucong Meng, Kexue Fu 0001, Shuo Wang 0011, Zhijian Song |
AAAI | 4 |
| 2025 | Asymmetric Performance Profiling Using Foundation Models: Quantifying Reliability and Expert Capability in Medical AIabstractIn safety-critical domains like medical imaging, where diagnostic errors have severe consequences, AI models must be evaluated beyond average accuracy. A trustworthy model must demonstrate two distinct virtues: high reliability on common, easy cases and high expert capability on challenging, ambiguous, or rare cases. Conventional aggregate metrics fail to distinguish between these, masking a model's fatal flaw-such as misclassifying an easy case-by rewarding its high volume of trivial successes. We present Hardness-Aware Model Evaluation (HaME), a framework that assesses models using an asymmet-ric cost-benefit analysis. HaME identifies challenging instances via foundation models, then evaluates the target model using novel metrics (HaPrecision, HaRecall, HaFt, HaAUC) and the Brittleness Gap (B-Gap). Our formulation uniquely penalizes “easy” errors far more than “hard” failures, while simultaneously rewarding “hard” successes. This shifts the evaluation from “av-erage performance” to “clinical trustworthiness.” Experiments on medical image classification (Dermatology, Pneumonia, Retinal OCT) and segmentation (Nuclei) reveal that HaME uncovers critical reliability gaps and expert-level specializations invisible to standard metrics. Mingzhi Xu, Tao Zhou 0002, Qiang Chen 0004, Shuo Wang 0011, Yizhe Zhang 0001 |
BIBM | 5 |
| 2025 | Exploring CLIP's Dense Knowledge for Weakly Supervised Semantic SegmentationabstractWeakly Supervised Semantic Segmentation (WSSS) with image-level labels aims to achieve pixel-level predictions using Class Activation Maps (CAMs). Recently, Contrastive Language-Image Pre-training (CLIP) has been introduced in WSSS. However, recent methods primarily focus on image-text alignment for CAM generation, while CLIP’s potential in patch-text alignment remains unexplored. In this work, we propose ExCEL to explore CLIP’s dense knowledge via a novel patch-text alignment paradigm for WSSS. Specifically, we propose Text Semantic Enrichment (TSE) and Visual Calibration (VC) modules to improve the dense alignment across both text and vision modalities. To make text embeddings semantically informative, our TSE module applies Large Language Models (LLMs) to build a dataset-wide knowledge base and enriches the text representations with an implicit attribute-hunting process. To mine fine-grained knowledge from visual features, our VC module first proposes Static Visual Calibration (SVC) to propagate fine-grained knowledge in a non-parametric manner. Then Learnable Visual Calibration (LVC) is further proposed to dynamically shift the frozen features towards distributions with diverse semantics. With these enhancements, ExCEL not only retains CLIP’s training-free advantages but also significantly outperforms other state-of-the-art methods with much less training cost on PASCAL VOC and MS COCO. Code is available at https://github.com/zwyang6/ExCEL. Yucong Meng, Kexue Fu 0001, Shuo Wang 0011, Zhijian Song |
CVPR | 5 |
| 2025 | Endo-CLIP: Progressive Self-supervised Pre-training on Raw Colonoscopy Records
Yili He, Peiyao Fu, Ruijie Yang, Zhihua Wang 0008, Quanlin Li, Pinghong Zhou, Xian Yang 0001, Shuo Wang 0011 |
MICCAI (11) | 10 |
| 2025 | Unsupervised Quality Control and Enhancement of Polyp Segmentation in Colonoscopy Videos Using Spatiotemporal Consistency
Tao Zhou 0002, Shuo Wang 0011, Yizhe Zhang 0001 |
MICCAI (10) | 4 |
| 2025 | From Generalist to Specialist: Distilling a Mixture of Foundation Models for Domain-Specific Medical Image Segmentation
Qing Li 0001, Yizhe Zhang 0001, Shengxiao Yang, Qirong Li, Shuo Wang 0011, Chengyan Wang |
MICCAI (2) | 8 |
| 2025 | Coherence-Based Segmentation Quality Evaluator Trained on a Large Collection of Annotated Medical Images
Ahjol Senbi, Fei Lyu 0004, Qing Li 0001, Yuhui Tao, Qiang Chen 0004, Chengyan Wang, Shuo Wang 0011, Tao Zhou 0002, Yizhe Zhang 0001 |
PRCV (13) | 9 |
| 2025 | The state-of-the-art in cardiac MRI reconstruction: Results of the CMRxRecon challenge in MICCAI 2023
Chen Qin, Shuo Wang 0011, Fanwen Wang, Yan Li 0064, Zi Wang 0005, Kunyuan Guo, Ouyang Cheng, Michael Tänzer, Longyu Sun, Mengting Sun, Zhang Shi, Sha Hua, Hao Li 0082, Zhensen Chen, Bingyu Xin, Dimitris N. Metaxas, George Yiasemis, Jonas Teuwen, Weitian Chen, Yidong Zhao, Yanwei Pang, Artem Razumov, Dmitry V. Dylov, Quan Dou, Yuyang Xue, Yuning Du, Julia Dietlmeier, Carles García-Cabrera, Ziad Al-Haj Hemidi, Nora Vogt, Ying-Hua Chu, Weibo Chen, Wenjia Bai, Xiahai Zhuang, Harry Qin, Lianming Wu, Guang Yang 0006, Xiaobo Qu 0001, He Wang 0016, Chengyan Wang |
Medical Image Anal. | 3 |
| 2025 | On-the-Fly Improving Segment Anything for Medical Image Segmentation Using Auxiliary Online LearningabstractThe current variants of the Segment Anything Model (SAM), which include the original SAM and Medical SAM, still lack the capability to produce sufficiently accurate segmentation for medical images. In medical imaging contexts, it is not uncommon for human experts to rectify segmentations of specific test samples after SAM generates its segmentation predictions. These rectifications typically entail manual or semi-manual corrections employing state-of-the-art annotation tools. Motivated by this process, we introduce a novel approach that leverages the advantages of online machine learning to enhance Segment Anything (SA) during test time. We employ rectified annotations to perform online learning, with the aim of improving the segmentation quality of SA on medical images. To ensure the effectiveness and efficiency of online learning when integrated with large-scale vision models like SAM, we propose a new method called Auxiliary Online Learning (AuxOL), which entails adaptive online-batch and adaptive segmentation fusion. Experiments conducted on eight datasets covering four medical imaging modalities validate the effectiveness of the proposed method. Our work proposes and validates a new, practical, and effective approach for enhancing SA on downstream segmentation tasks (e.g., medical image segmentation). The code is publicly available at https://sam-auxol.github.io/AuxOL/. Tao Zhou 0002, Weidi Xie, Shuo Wang 0011, Qi Dou 0001, Yizhe Zhang 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Robust Polyp Detection and Diagnosis Through Compositional Prompt-Guided Diffusion ModelsabstractColorectal cancer (CRC) is a significant global health concern, and early detection through screening plays a critical role in reducing mortality. While deep learning models have shown promise in improving polyp detection, classification, and segmentation, their generalization across diverse clinical environments, particularly with out-of-distribution (OOD) data, remains a challenge. Multi-center datasets like PolypGen have been developed to address these issues, but their collection is costly and time-consuming. Traditional data augmentation techniques provide limited variability, failing to capture the complexity of medical images. Diffusion models have emerged as a promising solution for generating synthetic polyp images, but the image generation process in current models mainly relies on segmentation masks as the condition, limiting their ability to capture the full clinical context. To overcome these limitations, we propose a Progressive Spectrum Diffusion Model (PSDM) that integrates diverse clinical annotations-such as segmentation masks, bounding boxes, and colonoscopy reports-by transforming them into compositional prompts. These prompts are organized into coarse and fine components, allowing the model to capture both broad spatial structures and fine details, generating clinically accurate synthetic images. By augmenting training data with PSDM-generated samples, our model significantly improves polyp detection, classification, and segmentation. For instance, on the PolypGen dataset, PSDM increases the F1 score by 2.12% and the mean average precision by 3.09%, demonstrating superior performance in OOD scenarios and enhanced generalization. Peiyao Fu, Junbo Huang, Quanlin Li, Pinghong Zhou, Zhihua Wang 0008, Fei Wu 0001, Shuo Wang 0011, Xian Yang 0001 |
IEEE Trans. Medical Imaging | 10 |
| 2025 | Tackling Ambiguity From Perspectives of Uncertainty Inference and Affinity Diversification for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation (WSSS) with image-level labels aims to achieve dense predictions without laborious annotations. However, due to the ambiguous contexts and fuzzy regions, the performance of WSSS, particularly during the stages of generating Class Activation Maps (CAMs) and refining pseudo masks, is widely hindered by ambiguity. Despite this, this issue has received little attention in previous literature. In this work, we propose UniA, a unified single-staged WSSS framework, to efficiently tackle this issue from the perspectives of uncertainty inference and affinity diversification. When activating class objects, we argue that the false activation stems from the bias to ambiguous regions during the feature extraction. Therefore, we formulate a robust feature representation with a Gaussian distribution and introduce the uncertainty estimation to avoid the bias. A distribution loss is proposed to supervise the process, which effectively captures the ambiguity and models the complex dependencies among features. When refining pseudo labels, we observe that the affinity from the prevailing refinement methods intends to be overly similar among ambiguities. To this end, we design an affinity diversification module to promote diversity among semantics. A mutual complementing refinement is first proposed to statically rectify the ambiguous affinity with multiple inferred pseudo labels. Then a contrastive affinity loss is further designed to dynamically diversify the relations among unrelated semantics. It stably propagates the diversity into the feature representation and helps generate better pseudo masks. Extensive experiments are conducted on PASCAL VOC, MS COCO, and medical ACDC datasets, which validate the efficiency of UniA tackling ambiguity and its superiority over recent single-staged or even most multi-staged competitors. Code is publicly available athttps://github.com/zwyang6/UniA. Yucong Meng, Kexue Fu 0001, Shuo Wang 0011, Zhijian Song |
IEEE Trans. Multim. | 4 |
| 2024 | Separate and Conquer: Decoupling Co-occurrence via Decomposition and Representation for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation (WSSS) with image-level labels aims to achieve segmentation tasks with-out dense annotations. However, attributed to the frequent coupling of co-occurring objects and the limited supervision from image-level labels, the challenging co-occurrence problem is widely present and leads to false activation of objects in WSSS. In this work, we devise a ‘Separate and Conquer’ scheme SeCo to tackle this issue from di-mensions of image space and feature space. In the im-age space, we propose to ‘separate’ the co-occurring ob-jects with image decomposition by subdividing images into patches. Importantly, we assign each patch a category tag from Class Activation Maps (CAMs), which spatially helps remove the co-context bias and guide the subsequent rep-resentation. In the feature space, we propose to ‘conquer’ the false activation by enhancing semantic representation with multi-granularity knowledge contrast. To this end, a dual-teacher-single-student architecture is designed and tag-guided contrast is conducted, which guarantee the cor-rectness of knowledge and further facilitate the discrepancy among co-contexts. We streamline the multi-staged WSSS pipeline end-to-end and tackle this issue without external supervision. Extensive experiments are conducted, validating the efficiency of our method and the superiority over previous single-staged and even multi-staged competitors on PASCAL VOC and MS COCO. Code is available here. Kexue Fu 0001, Minghong Duan, Linhao Qu, Shuo Wang 0011, Zhijian Song |
CVPR | 5 |
| 2024 | An Empirical Study on the Fairness of Foundation Models for Multi-Organ Image Segmentation
Qing Li 0001, Yizhe Zhang 0001, Yan Li 0064, Longyu Sun, Mengting Sun, Qirong Li, Wenyue Mao, Yinghua Chu, Shuo Wang 0011, Chengyan Wang |
MICCAI (12) | 13 |
| 2024 | EndoFinder: Online Image Retrieval for Explainable Colorectal Polyp Diagnosis
Ruijie Yang, Peiyao Fu, Yizhe Zhang 0001, Zhihua Wang 0008, Quanlin Li, Pinghong Zhou, Xian Yang 0001, Shuo Wang 0011 |
MICCAI (10) | 9 |
| 2024 | FAST: A Dual-tier Few-Shot Learning Paradigm for Whole Slide Image ClassificationabstractThe expensive fine-grained annotation and data scarcity
have become the primary obstacles for the widespread adoption of deep learning-based Whole Slide Images (WSI) classification algorithms in clinical practice. Unlike few-shot learning methods in natural images that can leverage the labels of each image, existing few-shot WSI classification methods only utilize a small number of fine-grained labels or weakly supervised slide labels for training in order to avoid expensive fine-grained annotation. They lack sufficient mining of available WSIs, severely limiting WSI classification performance. To address the above issues, we propose a novel and efficient dual-tier few-shot learning paradigm for WSI classification, named FAST. FAST consists of a dual-level annotation strategy and a dual-branch classification framework. Firstly, to avoid expensive fine-grained annotation, we collect a very small number of WSIs at the slide level, and annotate an extremely small number of patches. Then, to fully mining the available WSIs, we use all the patches and available patch labels to build a cache branch, which utilizes the labeled patches to learn the labels of unlabeled patches and through knowledge retrieval for patch classification. In addition to the cache branch, we also construct a prior branch that includes learnable prompt vectors, using the text encoder of visual-language models for patch classification. Finally, we integrate the results from both branches to achieve WSI classification. Extensive experiments on binary and multi-class datasets demonstrate that our proposed method significantly surpasses existing few-shot classification methods and approaches the accuracy of fully supervised methods with only 0.22% annotation costs. All codes and models will be publicly available on https://github.com/fukexue/FAST. Kexue Fu 0001, Xiaoyuan Luo, Linhao Qu, Shuo Wang 0011, Ilias Maglogiannis, Longxiang Gao, Manning Wang |
NeurIPS | 4 |
| 2024 | Combining Segment Anything Model with Domain-Specific Knowledge for Semi-Supervised Learning in Medical Image Segmentation
Yizhe Zhang 0001, Tao Zhou 0002, Ye Wu 0001, Pengfei Gu, Shuo Wang 0011 |
PRCV (14) | 5 |
| 2024 | STADNet: Spatial-Temporal Attention-Guided Dual-Path Network for cardiac cine MRI super-resolution
Shuo Wang 0011, Yapeng Tian, Shunjie Dong, Chengyan Wang, Angelica I. Avilés-Rivero, Harry Qin |
Medical Image Anal. | 2 |
| 2024 | Labelling with dynamics: A data-efficient learning paradigm for medical image segmentationabstractThe success of deep learning on image classification and recognition tasks has led to new applications in diverse contexts, including the field of medical imaging. However, two properties of deep neural networks (DNNs) may limit their future use in medical applications. The first is that DNNs require a large amount of labeled training data, and the second is that the deep learning-based models lack interpretability. In this paper, we propose and investigate a data-efficient framework for the task of general medical image segmentation. We address the two aforementioned challenges by introducing domain knowledge in the form of a strong prior into a deep learning framework. This prior is expressed by a customized dynamical system. We performed experiments on two different datasets, namely JSRT and ISIC2016 (heart and lungs segmentation on chest X-ray images and skin lesion segmentation on dermoscopy images). We have achieved competitive results using the same amount of training data compared to the state-of-the-art methods. More importantly, we demonstrate that our framework is extremely data-efficient, and it can achieve reliable results using extremely limited training data. Furthermore, the proposed method is rotationally invariant and insensitive to initialization. Yuanhan Mo, Fangde Liu, Guang Yang 0006, Shuo Wang 0011, Jian-Qing Zheng, Fuping Wu, Bartlomiej Wladyslaw Papiez, Douglas McIlwraith, Taigang He, Yike Guo |
Medical Image Anal. | 4 |
| 2024 | TestFit: A plug-and-play one-pass test time method for medical image segmentation
Yizhe Zhang 0001, Tao Zhou 0002, Yuhui Tao, Shuo Wang 0011, Ye Wu 0001, Benyuan Liu, Pengfei Gu, Qiang Chen 0004, Danny Ziyi Chen |
Medical Image Anal. | 4 |
| 2024 | CHeart: A Conditional Spatio-Temporal Generative Model for Cardiac AnatomyabstractTwo key questions in cardiac image analysis are to assess the anatomy and motion of the heart from images; and to understand how they are associated with non-imaging clinical factors such as gender, age and diseases. While the first question can often be addressed by image segmentation and motion tracking algorithms, our capability to model and answer the second question is still limited. In this work, we propose a novel conditional generative model to describe the 4D spatio-temporal anatomy of the heart and its interaction with non-imaging clinical factors. The clinical factors are integrated as the conditions of the generative modelling, which allows us to investigate how these factors influence the cardiac anatomy. We evaluate the model performance in mainly two tasks, anatomical sequence completion and sequence generation. The model achieves high performance in anatomical sequence completion, comparable to or outperforming other state-of-the-art generative models. In terms of sequence generation, given clinical conditions, the model can generate realistic synthetic 4D sequential anatomies that share similar distributions with the real data. The code and the trained generative model are available at https://github.com/MengyunQ/CHeart. Mengyun Qiao, Shuo Wang 0011, Huaqi Qiu, Antonio M. Simoes Monteiro de Marvao, Declan P. O'Regan, Daniel Rueckert, Wenjia Bai |
IEEE Trans. Medical Imaging | 2 |
| 2023 | Boosting Whole Slide Image Classification from the Perspectives of Distribution, Correlation and MagnificationabstractBag-based multiple instance learning (MIL) methods have become the mainstream for Whole Slide Image (WSI) classification. However, there are still three important issues that have not been fully addressed: (1) positive bags with a low positive instance ratio are prone to the influence of a large number of negative instances; (2) the correlation between local and global features of pathology images has not been fully modeled; and (3) there is a lack of effective information interaction between different magnifications. In this paper, we propose MILBooster, a powerful dual-scale multi-stage MIL framework to address these issues from the perspectives of distribution, correlation, and magnification. Specifically, to address issue (1), we propose a plug-and-play bag filter that effectively increases the positive instance ratio of positive bags. For issue (2), we propose a novel window-based Transformer architecture called PiceBlock to model the correlation between local and global features of pathology images. For issue (3), we propose a dual-branch architecture to process different magnifications and design an information interaction module called Scale Mixer for efficient information interaction between them. We conducted extensive experiments on four clinical WSI classification tasks using three datasets. MILBooster achieved new state-of-the-art performance on all these tasks. Codes will be available at https://github.com/miccaiif/MILBooster. Linhao Qu, Minghong Duan, Yingfan Ma, Shuo Wang 0011, Manning Wang, Zhijian Song |
ICCV | 5 |
| 2023 | Region-focused multi-view transformer-based generative adversarial network for cardiac cine MRI reconstruction
Chengyan Wang, Chen Qin, Shuo Wang 0011, Qi Dou 0001, Harry Qin |
Medical Image Anal. | 5 |
| 2023 | Generative myocardial motion tracking via latent space exploration with biomechanics-informed priorabstractMyocardial motion and deformation are rich descriptors that characterize cardiac function. Image registration, as the most commonly used technique for myocardial motion tracking, is an ill-posed inverse problem which often requires prior assumptions on the solution space. In contrast to most existing approaches which impose explicit generic regularization such as smoothness, in this work we propose a novel method that can implicitly learn an application-specific biomechanics-informed prior and embed it into a neural network-parameterized transformation model. Particularly, the proposed method leverages a variational autoencoder-based generative model to learn a manifold for biomechanically plausible deformations. The motion tracking then can be performed via traversing the learnt manifold to search for the optimal transformations while considering the sequence information. The proposed method is validated on three public cardiac cine MRI datasets with comprehensive evaluations. The results demonstrate that the proposed method can outperform other approaches, yielding higher motion tracking accuracy with reasonable volume preservation and better generalizability to varying data distributions. It also enables better estimates of myocardial strains, which indicates the potential of the method in characterizing spatiotemporal signatures for understanding cardiovascular diseases. Chen Qin, Shuo Wang 0011, Chen Chen 0042, Wenjia Bai, Daniel Rueckert |
Medical Image Anal. | 2 |
| 2022 | Enhancing MR image segmentation with realistic adversarial data augmentationabstractThe success of neural networks on medical image segmentation tasks typically relies on large labeled datasets for model training. However, acquiring and manually labeling a large medical image set is resource-intensive, expensive, and sometimes impractical due to data sharing and privacy issues. To address this challenge, we propose AdvChain, a generic adversarial data augmentation framework, aiming at improving both the diversity and effectiveness of training data for medical image segmentation tasks. AdvChain augments data with dynamic data augmentation, generating randomly chained photo-metric and geometric transformations to resemble realistic yet challenging imaging variations to expand training data. By jointly optimizing the data augmentation model and a segmentation network during training, challenging examples are generated to enhance network generalizability for the downstream task. The proposed adversarial data augmentation does not rely on generative networks and can be used as a plug-in module in general segmentation networks. It is computationally efficient and applicable for both low-shot supervised and semi-supervised learning. We analyze and evaluate the method on two MR image segmentation tasks: cardiac segmentation and prostate segmentation with limited labeled data. Results show that the proposed approach can alleviate the need for labeled data while improving model generalization ability, indicating its practical value in medical imaging applications. Chen Chen 0042, Chen Qin, Cheng Ouyang, Zeju Li, Shuo Wang 0011, Huaqi Qiu, Liang Chen 0018, Giacomo Tarroni, Wenjia Bai, Daniel Rueckert |
Medical Image Anal. | 5 |
| 2022 | Suggestive annotation of brain MR images with gradient-guided sampling
Chengliang Dai, Shuo Wang 0011, Yuanhan Mo, Elsa D. Angelini, Yike Guo, Wenjia Bai |
Medical Image Anal. | 2 |
| 2022 | Beyond fine-tuning: Classifying high resolution mammograms using function-preserving transformationsabstractThe task of classifying mammograms is very challenging because the lesion is usually small in the high resolution image. The current state-of-the-art approaches for medical image classification rely on using the de-facto method for convolutional neural networks-fine-tuning. However, there are fundamental differences between natural images and medical images, which based on existing evidence from the literature, limits the overall performance gain when designed with algorithmic approaches. In this paper, we propose to go beyond fine-tuning by introducing a novel framework called MorphHR, in which we highlight a new transfer learning scheme. The idea behind the proposed framework is to integrate function-preserving transformations, for any continuous non-linear activation neurons, to internally regularise the network for improving mammograms classification. The proposed solution offers two major advantages over the existing techniques. Firstly and unlike fine-tuning, the proposed approach allows for modifying not only the last few layers but also several of the first ones on a deep ConvNet. By doing this, we can design the network front to be suitable for learning domain specific features. Secondly, the proposed scheme is scalable to hardware. Therefore, one can fit high resolution images on standard GPU memory. We show that by using high resolution images, one prevents losing relevant information. We demonstrate, through numerical and visual experiments, that the proposed approach yields to a significant improvement in the classification performance over state-of-the-art techniques, and is indeed on a par with radiology experts. Moreover and for generalisation purposes, we show the effectiveness of the proposed learning scheme on another large dataset, the ChestX-ray14, surpassing current state-of-the-art techniques. Angelica I. Avilés-Rivero, Shuo Wang 0011, Yuan Huang 0009, Fiona J. Gilbert, Carola-Bibiane Schönlieb, Chang Wen Chen |
Medical Image Anal. | 3 |
| 2022 | Bayesian data assimilation for estimating instantaneous reproduction numbers during epidemics: Applications to COVID-19abstractEstimating the changes of epidemiological parameters, such as instantaneous reproduction number, Rt, is important for understanding the transmission dynamics of infectious diseases. Current estimates of time-varying epidemiological parameters often face problems such as lagging observations, averaging inference, and improper quantification of uncertainties. To address these problems, we propose a Bayesian data assimilation framework for time-varying parameter estimation. Specifically, this framework is applied to estimate the instantaneous reproduction number Rt during emerging epidemics, resulting in the state-of-the-art 'DARt' system. With DARt, time misalignment caused by lagging observations is tackled by incorporating observation delays into the joint inference of infections and Rt; the drawback of averaging is overcome by instantaneously updating upon new observations and developing a model selection mechanism that captures abrupt changes; the uncertainty is quantified and reduced by employing Bayesian smoothing. We validate the performance of DARt and demonstrate its power in describing the transmission dynamics of COVID-19. The proposed approach provides a promising solution for making accurate and timely estimation for transmission dynamics based on reported data. Xian Yang 0001, Shuo Wang 0011, Yuting Xing, Ling Li 0010, Karl J. Friston, Yike Guo |
PLoS Comput. Biol. | 2 |
| 2021 | Joint Motion Correction and Super Resolution for Cardiac Segmentation via Latent Optimisation
Shuo Wang 0011, Chen Qin, Nicolò Savioli, Chen Chen 0042, Declan P. O'Regan, Stuart A. Cook, Yike Guo, Daniel Rueckert, Wenjia Bai |
MICCAI (3) | 1 |
| 2021 | Non-invasive Assessment of Hepatic Venous Pressure Gradient (HVPG) Based on MR Flow Imaging and Computational Fluid Dynamics
Shuo Wang 0011, Minghua Xiong, Chengyan Wang, He Wang 0016 |
MICCAI (7) | 2 |
| 2020 | Realistic Adversarial Data Augmentation for MR Image Segmentation
Chen Chen 0042, Chen Qin, Huaqi Qiu, Cheng Ouyang, Shuo Wang 0011, Liang Chen 0018, Giacomo Tarroni, Wenjia Bai, Daniel Rueckert |
MICCAI (1) | 5 |
| 2020 | Suggestive Annotation of Brain Tumour Images with Gradient-Guided Sampling
Chengliang Dai, Shuo Wang 0011, Yuanhan Mo, Kaichen Zhou, Elsa D. Angelini, Yike Guo, Wenjia Bai |
MICCAI (4) | 2 |
| 2020 | Biomechanics-Informed Neural Networks for Myocardial Motion Tracking in MRI
Chen Qin, Shuo Wang 0011, Chen Chen 0042, Huaqi Qiu, Wenjia Bai, Daniel Rueckert |
MICCAI (3) | 2 |
| 2020 | Deep Generative Model-Based Quality Control for Cardiac MRI Segmentation
Shuo Wang 0011, Giacomo Tarroni, Chen Qin, Yuanhan Mo, Chengliang Dai, Chen Chen 0042, Ben Glocker, Yike Guo, Daniel Rueckert, Wenjia Bai |
MICCAI (4) | 1 |
| 2014 | Exploiting small world property for network clustering
Tieyun Qian, Qing Li 0001, Jaideep Srivastava, Zhiyong Peng 0001, Yang Yang 0009, Shuo Wang 0011 |
World Wide Web | 6 |
| 2010 | Refining Graph Partitioning for Social Network Clustering
Tieyun Qian, Yang Yang 0009, Shuo Wang 0011 |
WISE | 3 |