VLDB 2026 Research / reviewers in the wild / expert
Wei Shao 0008
dblp:24/803-8
· DBLP profile ↗
21ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0003-4931-4839ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-class segmentation of aortic branches and zones in computed tomography angiography: The AortaSeg24 challenge
Muhammad Imran 0013, Jonathan R. Krebs, Vishal Balaji Sivaraman, Amarjeet Kumar, Walker R. Ueland, Michael J. Fassler, Lisheng Wang, Maximilian Rokuss, Michael Baumgartner 0001, Yannick Kirchhof, Klaus H. Maier-Hein, Fabian Isensee, Shuolin Liu, Bong Thanh Nguyen, Dong-jin Shin, Park Ji-Woo, Matthew Choi, Kwang-Hyun Uhm, Sung-Jea Ko, Chanwoong Lee, Jaehee Chun, Yun Gu, Zhaohong Pan, Xiaokun Liang, Markus Tiefenthaler, Enrique Almar-Munoz, Matthias Schwab, Mikhail Kotyushev, Rostislav Epifanov, Marek Wodzinski, Henning Müller, Abdul Qayyum 0002, Moona Mazher, Steven A. Niederer, Zhiwei Wang 0002, Kaixiang Yang 0004, Jintao Ren, Stine Sofia Korreman, Yuchong Gao, Hongye Zeng, Jinghua Yue, Fugen Zhou, Alexander Cosman, Muxuan Liang, Gilbert R. Upchurch Jr., Yuyin Zhou, Michol A. Cooper, Wei Shao 0008 |
Medical Image Anal. | 63 |
| 2026 | Med3DVLM: An Efficient Vision-Language Model for 3D Medical Image AnalysisabstractVision-language models (VLMs) have shown promise in 2D medical image analysis, but extending them to 3D remains challenging due to the high computational demands of volumetric data and the difficulty of aligning 3D spatial features with clinical text. We present Med3DVLM, a 3D VLM designed to address these challenges through three key innovations: (1) DCFormer, an efficient encoder that uses decomposed 3D convolutions to capture fine-grained spatial features at scale; (2) SigLIP, a contrastive learning strategy with pairwise sigmoid loss that improves image-text alignment without relying on large negative batches; and (3) a dual-stream MLP-Mixer projector that fuses low- and high-level image features with text embeddings for richer multi-modal representations. We evaluated our model on the M3D dataset, which includes radiology reports and VQA data for 120,084 3D medical images. The results show that Med3DVLM achieves superior performance on multiple benchmarks. For image-text retrieval, it reaches 61.00% R@1 on 2,000 samples, significantly outperforming the current state-of-the-art M3D-LaMed model (19.10%). For report generation, it achieves a METEOR score of 36.42% (vs. 14.38%). In open-ended visual question answering (VQA), it scores 36.76% METEOR (vs. 33.58%), and in closed-ended VQA, it achieves 79.95% accuracy (vs. 75.78%). These results demonstrate Med3DVLM's ability to bridge the gap between 3D imaging and language, enabling scalable, multi-task reasoning across clinical applications. Gorkem Can Ates, Kuang Gong, Wei Shao 0008 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Mamba-Reg: Vision Mamba Also Needs RegistersabstractSimilar to Vision Transformers, this paper identifies artifacts also present within the feature maps of Vision Mamba. These artifacts, corresponding to high-norm tokens emerging in low-information background areas of images, appear much more severe in Vision Mamba—they exist prevalently even with the tiny-sized model and activate extensively across background regions. To mitigate this issue, we follow the prior solution of introducing register tokens into Vision Mamba. To better cope with Mamba blocks’ uni-directional inference paradigm, two key modifications are introduced: 1) evenly inserting registers throughout the input token sequence, and 2) recycling registers for final decision predictions. We term this new architecture Mamba®. Qualitative observations suggest, compared to vanilla Vision Mamba, Mamba®’s feature maps appear cleaner and more focused on semantically meaningful regions. Quantitatively, Mamba®attains stronger performance and scales better. For example, on the ImageNet benchmark, our Mamba®-B attains 83.0% accuracy, significantly outperforming Vim-B’s 81.8%; furthermore, we provide the first successful scaling to the large model size with 341M parameters, attaining competitive accuracies of 83.6% and 84.5% for 224×224 and 384×384 inputs, respectively. Additional validation on the downstream semantic segmentation task also supports Mamba®’s efficacy. Code is available at https://github.com/wangf3014/Mamba-Reg. Feng Wang 0047, Jiahao Wang 0001, Sucheng Ren, Guoyizhe Wei, Jieru Mei, Wei Shao 0008, Yuyin Zhou, Alan L. Yuille, Cihang Xie |
CVPR | 6 |
| 2025 | Adventurer: Optimizing Vision Mamba Architecture Designs for Efficiencyabstractseries models where we treat images as sequences of patch tokens and employ uni-directional language models to learn visual representations. This modeling paradigm allows us to process images in a recurrent formulation with linear complexity relative to the sequence length, which can effectively address the memory and computation explosion issues posed by high-resolution and fine-grained images. In detail, we introduce two simple designs that seamlessly integrate image inputs into the causal inference framework: a global pooling token placed at the beginning of the sequence and a flipping operation between every two layers. Extensive empirical studies highlight that compared with the existing plain architectures such as DeiT [46] and Vim [57], Adventurer offers an optimal efficiency-accuracy trade-off. For example, our Adventurer-Base attains a competitive test accuracy of 84.3% on the standard ImageNet-1k benchmark with 216 images/s training throughput, which is 3.8× and 6.2× faster than Vim and DeiT to achieve the same result. As Adventurer offers great computation and memory efficiency and allows scaling with linear complexity, we hope this architecture can benefit future explorations in modeling long sequences for high-resolution or fine-grained images. Code is available at https://github.com/wangf3014/Adventurer. Feng Wang 0047, Timing Yang, Yaodong Yu, Sucheng Ren, Guoyizhe Wei, Angtian Wang, Wei Shao 0008, Yuyin Zhou, Alan L. Yuille, Cihang Xie |
CVPR | 7 |
| 2025 | Mapping Cultural Ecosystem Services Using One-Shot In-Context Learning with Multimodal Large Language ModelsabstractCultural Ecosystem Services (CES)—such as outdoor recreation, wildlife viewing, and landscape aesthetics—are intangible benefits derived from nature. Mapping CES flows, the spatial patterns of actual use, is critical for sustainable landscape planning. Geotagged social media images provide a scalable, low-cost data source, but current approaches often depend on supervised learning or automated tagging, which require large labeled datasets, depend on manual interpretation, and show limited generalizability across new CES categories and geographic contexts. To address these challenges, we present the first application of multimodal large language models (LLMs) with in-context learning for large-scale CES mapping. Using 3.42 million Flickr images (2014–2019) from six Southeastern U.S. states, we benchmarked three vision-language models—Qwen2.5-VL, Gemma-3, and Aya Vision—under a zero-shot framework, and further evaluated the best-performing model with one-shot learning. By focusing on one-shot, we balance improved accuracy with minimal labeling effort, effectively addressing the challenge of scarce labeled examples. CES flows were quantified as county-level Photo-User-Days (PUDs) and clustered to identify CES bundles. Results show: (1) Qwen2.5-VL achieved the highest zero-shot accuracy (90%), outperforming Gemma-3 and Aya Vision by over 10 points; (2) with one-shot learning, Qwen2.5-VL-OSL significantly improved recall by 23 points for "landscape aesthetics" and 20 points for "no CES"; (3) predicted CES flows aligned closely with land use and land cover distribution; and (4) six CES bundles were identified, reflecting varying flow intensity and compositions. These findings demonstrate the potential of multimodal LLMs, enhanced with one-shot learning, for scalable, data-efficient, and accurate CES assessment from crowdsourced imagery. Hao-Yu Liao, Wei Shao 0008 |
SIGSPATIAL/GIS | 4 |
| 2025 | FoundationSoil: Enhancing Soil Organic Carbon Mapping Using a Multi-Temporal Geospatial Foundation ModelabstractSoil organic carbon (SOC) is a key indicator of soil health and climate resilience, yet large-scale spatial prediction remains challenging due to sparse field data and the complex dynamics of soil processes. We present FoundationSoil, a digital soil mapping framework that integrates a geospatial foundation model (Prithvi-EO-2.0) with environmental covariates to predict the spatial variability of SOC concentrations. Prithvi-EO-2.0 encodes multi-temporal satellite image patches via a transformer encoder, producing spatiotemporal representations that capture land-surface dynamics across seasons. These features are combined with 109 environmental covariates (e.g., climate, soil moisture) to train Random Forest and neural network regressors. We evaluated FoundationSoil using 12,692 georeferenced topsoil samples from the WoSIS database across the continental United States. Stratified grid-based partitions were applied to mitigate the influence of spatial autocorrelation on model performance evaluation. Results show that FoundationSoil consistently outperformed a pixel-based baseline, achieving best performance at R2 = 0.72 and RMSE = 38.62. Feature attribution shows that Prithvi-derived components capture semantically rich, non-redundant signals relevant to SOC processes. These results demonstrate the promise of pretrained vision foundation models for digital soil mapping and highlight their potential for data-efficient, transferable learning in Earth science applications. Hao-Yu Liao, Wei Shao 0008 |
SIGSPATIAL/GIS | 4 |
| 2025 | Scaling Laws in Patchification: An Image Is Worth 50, 176 Tokens And MoreabstractSince the introduction of Vision Transformer (ViT), patchification has long been regarded as a common image pre-processing approach for plain visual architectures. By compressing the spatial size of images, this approach can effectively shorten the token sequence and reduce the computational cost of ViT-like plain architectures. In this work, we aim to thoroughly examine the information loss caused by this patchification-based compressive encoding paradigm and how it affects visual understanding. We conduct extensive patch size scaling experiments and excitedly observe an intriguing scaling law in patchification: the models can consistently benefit from decreased patch sizes and attain improved predictive performance, until it reaches the minimum patch size of 1*1, i.e., pixel tokenization. This conclusion is broadly applicable across different vision tasks, various input scales, and diverse architectures such as ViT and the recent Mamba models. Moreover, as a by-product, we discover that with smaller patches, task-specific decoder heads become less critical for dense prediction. In the experiments, we successfully scale up the visual sequence to an exceptional length of 50,176 tokens, achieving a competitive test accuracy of 84.6% with a base-sized model on the ImageNet-1k benchmark. We hope this study can provide insights and theoretical foundations for future works of building non-compressive vision models. Feng Wang 0047, Yaodong Yu, Wei Shao 0008, Yuyin Zhou, Alan L. Yuille, Cihang Xie |
ICML | 3 |
| 2025 | Fast-DDPM: Fast Denoising Diffusion Probabilistic Models for Medical Image-to-Image GenerationabstractDenoising diffusion probabilistic models (DDPMs) have achieved unprecedented success in computer vision. However, they remain underutilized in medical imaging, a field crucial for disease diagnosis and treatment planning. This is primarily due to the high computational cost associated with the use of large number of time steps (e.g., 1,000) in diffusion processes. Training a diffusion model on medical images typically takes days to weeks, while sampling each image volume takes minutes to hours. To address this challenge, we introduce Fast-DDPM, a simple yet effective approach capable of simultaneously improving training speed, sampling speed, and generation quality. Unlike DDPM, which trains the image denoiser across 1,000 time steps, Fast-DDPM trains and samples using only 10 time steps. The key to our method lies in aligning the training and sampling procedures to optimize time-step utilization. Specifically, we introduced two efficient noise schedulers with 10 time steps: one with uniform time step sampling and another with non-uniform sampling. We evaluated Fast-DDPM across three medical image-to-image generation tasks: multi-image super-resolution, image denoising, and image-to-image translation. Fast-DDPM outperformed DDPM and current state-of-the-art methods based on convolutional networks and generative adversarial networks in all tasks. Additionally, Fast-DDPM reduced the training time to 0.2× and the sampling time to 0.01× compared to DDPM. Muhammad Imran 0013, Yuyin Zhou, Muxuan Liang, Kuang Gong, Wei Shao 0008 |
IEEE J. Biomed. Health Informatics | 7 |
| 2024 | Inferring Gene Regulatory Network Based on scATAC-seq Data with Gene PerturbationabstractGene regulatory networks (GRNs) are critical blueprints for understanding gene regulation and the intricate interactions that drive biological processes. Recent advances have highlighted the potential of single-cell ATAC-seq (scATAC-seq) data in GRN inference, offering unprecedented insights into how chromatin accessibility plays an important part in gene regulation. However, existing methods often fall short in providing a quantitative and holistic depiction of regulatory relationships, particularly in capturing the strength, direction, and type of gene regulation simultaneously. In this paper, we present a novel approach that addresses these limitations by leveraging genetically perturbed scATAC-seq data to infer more comprehensive and accurate GRNs. Our method advances the field by integrating pre- and post-perturbation chromatin accessibility data, enabling the construction of GRNs that more accurately reflect the dynamic regulatory landscape. Through rigorous evaluation on seven real datasets, we demonstrate the method’s superior performance in reconstructing GRNs with enhanced precision and interpretability. This work significantly contributes to the field by providing a robust framework for GRN inference, with broad implications for understanding gene regulation in complex biological systems. Wei Shao 0008, Yuti Liu, Qiao Liu 0008, Wanwen Zeng |
BIBM | 1 |
| 2024 | Unleashing the Potential of SAM for Medical Adaptation via Hierarchical DecodingabstractThe Segment Anything Model (SAM) has garnered significant attention for its versatile segmentation abilities and intuitive prompt-based interface. However, its application in medical imaging presents challenges, requiring either substantial training costs and extensive medical datasets for full model fine-tuning or high-quality prompts for optimal performance. This paper introduces H-SAM: a prompt-free adaptation of SAM tailored for efficient fine-tuning of medical images via a two-stage hierarchical decoding procedure. In the initial stage, H-SAM employs SAM's original decoder to generate a prior probabilistic mask, guiding a more intricate decoding process in the second stage. Specifically, we propose two key designs: 1) A class-balanced, mask-guided self-attention mechanism addressing the unbalanced label distribution, enhancing image embedding; 2) A learnable mask cross-attention mechanism spatially modulating the interplay among different image regions based on the prior mask. Moreover, the inclusion of a hierarchical pixel decoder in H-SAM enhances its proficiency in capturing fine-grained and localized details. This approach enables SAM to effectively integrate learned medical priors, facilitating enhanced adaptation for medical image segmentation with limited samples. Our H-SAM demonstrates a 4.78% improvement in average Dice compared to existing prompt-free SAM variants for multi-organ segmentation using only 10% of 2D slices. Notably, without using any unlabeled data, H-SAM even outperforms state-of-the-art semisupervised models relying on extensive unlabeled training data across various medical datasets. Our code is available at https://github.com/Cccccczh404/H-SAM. Zhiheng Cheng, Qingyue Wei, Hongru Zhu, Yan Wang 0033, Liangqiong Qu, Wei Shao 0008, Yuyin Zhou |
CVPR | 6 |
| 2024 | PET Image Denoising Based on 3D Denoising Diffusion Probabilistic Model: Evaluations on Total-Body Datasets
Boxiao Yu, Savas Ozdemir, Yafei Dong, Wei Shao 0008, Kuangyu Shi, Kuang Gong |
MICCAI (7) | 4 |
| 2023 | Consistency-Guided Meta-learning for Bootstrapping Semi-supervised Medical Image Segmentation
Qingyue Wei, Lequan Yu, Xianhang Li, Wei Shao 0008, Cihang Xie, Lei Xing 0001, Yuyin Zhou |
MICCAI (4) | 4 |
| 2023 | Learn2Reg: Comprehensive Multi-Task Medical Image Registration Challenge, Dataset and Evaluation in the Era of Deep LearningabstractImage registration is a fundamental medical image analysis task, and a wide variety of approaches have been proposed. However, only a few studies have comprehensively compared medical image registration approaches on a wide range of clinically relevant tasks. This limits the development of registration methods, the adoption of research advances into practice, and a fair benchmark across competing approaches. The Learn2Reg challenge addresses these limitations by providing a multi-task medical image registration data set for comprehensive characterisation of deformable registration algorithms. A continuous evaluation will be possible at https://learn2reg.grand-challenge.org. Learn2Reg covers a wide range of anatomies (brain, abdomen, and thorax), modalities (ultrasound, CT, MR), availability of annotations, as well as intra- and inter-patient registration evaluation. We established an easily accessible framework for training and validation of 3D registration methods, which enabled the compilation of results of over 65 individual method submissions from more than 20 unique teams. We used a complementary set of metrics, including robustness, accuracy, plausibility, and runtime, enabling unique insight into the current state-of-the-art of medical image registration. This paper describes datasets, tasks, evaluation methods and results of the challenge, as well as results of further analysis of transferability to new datasets, the importance of label supervision, and resulting bias. While no single approach worked best across all tasks, many methodological aspects could be identified that push the performance of medical image registration to new state-of-the-art performance. Furthermore, we demystified the common belief that conventional registration methods have to be much slower than deep-learning-based methods. Alessa Hering, Lasse Hansen, Tony C. W. Mok, Albert C. S. Chung, Hanna Siebert, Stephanie Häger, Annkristin Lange, Sven Kuckertz, Stefan Heldmann, Wei Shao 0008, Sulaiman Vesal, Mirabela Rusu, Geoffrey A. Sonn, Théo Estienne, Maria Vakalopoulou, Luyi Han, Yunzhi Huang, Pew-Thian Yap, Mikael Brudfors, Yaël Balbastre, Samuel Joutard, Marc Modat, Gal Lifshitz, Dan Raviv, Jinxin Lv, Qiang Li 0018, Vincent Jaouen, Dimitris Visvikis, Constance Fourcade, Mathieu Rubeaux, Wentao Pan 0001, Zhe Xu 0012, Bailiang Jian, Francesca De Benetti, Marek Wodzinski, Niklas Gunnarsson, Jens Sjölund, Daniel Grzech, Huaqi Qiu, Zeju Li, Alexander Thorley, Jinming Duan 0001, Christoph Großbröhmer, Andrew Hoopes, Ingerid Reinertsen, Yiming Xiao 0001, Bennett A. Landman, Yuankai Huo, Keelin Murphy, Nikolas Leßmann, Bram van Ginneken, Adrian V. Dalca, Mattias P. Heinrich |
IEEE Trans. Medical Imaging | 10 |
| 2022 | Deep learning-based pseudo-mass spectrometry imaging analysis for precision medicineabstractLiquid chromatography-mass spectrometry (LC-MS)-based untargeted metabolomics provides systematic profiling of metabolic. Yet, its applications in precision medicine (disease diagnosis) have been limited by several challenges, including metabolite identification, information loss and low reproducibility. Here, we present the deep-learning-based Pseudo-Mass Spectrometry Imaging (deepPseudoMSI) project (https://www.deeppseudomsi.org/), which converts LC-MS raw data to pseudo-MS images and then processes them by deep learning for precision medicine, such as disease diagnosis. Extensive tests based on real data demonstrated the superiority of deepPseudoMSI over traditional approaches and the capacity of our method to achieve an accurate individualized diagnosis. Our framework lays the foundation for future metabolic-based precision medicine. Xiaotao Shen, Wei Shao 0008, Chuchu Wang, Songjie Chen, Mirabela Rusu, Michael Snyder 0001 |
Briefings Bioinform. | 2 |
| 2022 | Selective identification and localization of indolent and aggressive prostate cancers via CorrSigNIA: an MRI-pathology correlation and deep learning frameworkabstractAutomated methods for detecting prostate cancer and distinguishing indolent from aggressive disease on Magnetic Resonance Imaging (MRI) could assist in early diagnosis and treatment planning. Existing automated methods of prostate cancer detection mostly rely on ground truth labels with limited accuracy, ignore disease pathology characteristics observed on resected tissue, and cannot selectively identify aggressive (Gleason Pattern≥4) and indolent (Gleason Pattern=3) cancers when they co-exist in mixed lesions. In this paper, we present a radiology-pathology fusion approach, CorrSigNIA, for the selective identification and localization of indolent and aggressive prostate cancer on MRI. CorrSigNIA uses registered MRI and whole-mount histopathology images from radical prostatectomy patients to derive accurate ground truth labels and learn correlated features between radiology and pathology images. These correlated features are then used in a convolutional neural network architecture to detect and localize normal tissue, indolent cancer, and aggressive cancer on prostate MRI. CorrSigNIA was trained and validated on a dataset of 98 men, including 74 men that underwent radical prostatectomy and 24 men with normal prostate MRI. CorrSigNIA was tested on three independent test sets including 55 men that underwent radical prostatectomy, 275 men that underwent targeted biopsies, and 15 men with normal prostate MRI. CorrSigNIA achieved an accuracy of 80% in distinguishing between men with and without cancer, a lesion-level ROC-AUC of 0.81±0.31 in detecting cancers in both radical prostatectomy and biopsy cohort patients, and lesion-levels ROC-AUCs of 0.82±0.31 and 0.86±0.26 in detecting clinically significant cancers in radical prostatectomy and biopsy cohort patients respectively. CorrSigNIA consistently outperformed other methods across different evaluation metrics and cohorts. In clinical settings, CorrSigNIA may be used in prostate cancer detection as well as in selective identification of indolent and aggressive components of prostate cancer, thereby improving prostate cancer care by helping guide targeted biopsies, reducing unnecessary biopsies, and selecting and planning treatment. Indrani Bhattacharya, Arun Seetharaman, Christian Kunder, Wei Shao 0008, Leo C. Chen, Simon J. C. Soerensen, Jeffrey B. Wang, Nikola C. Teslovich, Richard E. Fan, Pejman Ghanouni, James D. Brooks, Geoffrey A. Sonn, Mirabela Rusu |
Medical Image Anal. | 4 |
| 2021 | Weakly Supervised Registration of Prostate MRI and Histopathology Images
Wei Shao 0008, Indrani Bhattacharya, Simon J. C. Soerensen, Christian Kunder, Jeffrey B. Wang, Richard E. Fan, Pejman Ghanouni, James D. Brooks, Geoffrey A. Sonn, Mirabela Rusu |
MICCAI (4) | 1 |
| 2021 | ProsRegNet: A deep learning framework for registration of MRI and histopathology images of the prostateabstractMagnetic resonance imaging (MRI) is an increasingly important tool for the diagnosis and treatment of prostate cancer. However, interpretation of MRI suffers from high inter-observer variability across radiologists, thereby contributing to missed clinically significant cancers, overdiagnosed low-risk cancers, and frequent false positives. Interpretation of MRI could be greatly improved by providing radiologists with an answer key that clearly shows cancer locations on MRI. Registration of histopathology images from patients who had radical prostatectomy to pre-operative MRI allows such mapping of ground truth cancer labels onto MRI. However, traditional MRI-histopathology registration approaches are computationally expensive and require careful choices of the cost function and registration hyperparameters. This paper presents ProsRegNet, a deep learning-based pipeline to accelerate and simplify MRI-histopathology image registration in prostate cancer. Our pipeline consists of image preprocessing, estimation of affine and deformable transformations by deep neural networks, and mapping cancer labels from histopathology images onto MRI using estimated transformations. We trained our neural network using MR and histopathology images of 99 patients from our internal cohort (Cohort 1) and evaluated its performance using 53 patients from three different cohorts (an additional 12 from Cohort 1 and 41 from two public cohorts). Results show that our deep learning pipeline has achieved more accurate registration results and is at least 20 times faster than a state-of-the-art registration algorithm. This important advance will provide radiologists with highly accurate prostate MRI answer keys, thereby facilitating improvements in the detection of prostate cancer on MRI. Our code is freely available at https://github.com/pimed//ProsRegNet. Wei Shao 0008, Linda Banh, Christian Kunder, Richard E. Fan, Simon J. C. Soerensen, Jeffrey B. Wang, Nikola C. Teslovich, Nikhil Madhuripan, Anugayathri Jawahar, Pejman Ghanouni, James D. Brooks, Geoffrey A. Sonn, Mirabela Rusu |
Medical Image Anal. | 1 |
| 2021 | Geodesic density regression for correcting 4DCT pulmonary respiratory motion artifacts
Wei Shao 0008, Yue Pan 0013, Oguz C. Durumeric, Joseph M. Reinhardt, John E. Bayouth, Mirabela Rusu, Gary E. Christensen |
Medical Image Anal. | 1 |
| 2021 | 3D Registration of pre-surgical prostate MRI and histopathology images via super-resolution volume reconstructionabstractThe use of MRI for prostate cancer diagnosis and treatment is increasing rapidly. However, identifying the presence and extent of cancer on MRI remains challenging, leading to high variability in detection even among expert radiologists. Improvement in cancer detection on MRI is essential to reducing this variability and maximizing the clinical utility of MRI. To date, such improvement has been limited by the lack of accurately labeled MRI datasets. Data from patients who underwent radical prostatectomy enables the spatial alignment of digitized histopathology images of the resected prostate with corresponding pre-surgical MRI. This alignment facilitates the delineation of detailed cancer labels on MRI via the projection of cancer from histopathology images onto MRI. We introduce a framework that performs 3D registration of whole-mount histopathology images to pre-surgical MRI in three steps. First, we developed a novel multi-image super-resolution generative adversarial network (miSRGAN), which learns information useful for 3D registration by producing a reconstructed 3D MRI. Second, we trained the network to learn information between histopathology slices to facilitate the application of 3D registration methods. Third, we registered the reconstructed 3D histopathology volumes to the reconstructed 3D MRI, mapping the extent of cancer from histopathology images onto MRI without the need for slice-to-slice correspondence. When compared to interpolation methods, our super-resolution reconstruction resulted in the highest PSNR relative to clinical 3D MRI (32.15 dB vs 30.16 dB for BSpline interpolation). Moreover, the registration of 3D volumes reconstructed via super-resolution for both MRI and histopathology images showed the best alignment of cancer regions when compared to (1) the state-of-the-art RAPSODI approach, (2) volumes that were not reconstructed, or (3) volumes that were reconstructed using nearest neighbor, linear, or BSpline interpolations. The improved 3D alignment of histopathology images and MRI facilitates the projection of accurate cancer labels on MRI, allowing for the development of improved MRI interpretation schemes and machine learning models to automatically detect cancer on MRI. Rewa Sood, Wei Shao 0008, Christian Kunder, Nikola C. Teslovich, Jeffrey B. Wang, Simon J. C. Soerensen, Nikhil Madhuripan, Anugayathri Jawahar, James D. Brooks, Pejman Ghanouni, Richard E. Fan, Geoffrey A. Sonn, Mirabela Rusu |
Medical Image Anal. | 2 |
| 2020 | CorrSigNet: Learning CORRelated Prostate Cancer SIGnatures from Radiology and Pathology Images for Improved Computer Aided DiagnosisabstractMagnetic Resonance Imaging (MRI) is widely used for screening and staging prostate cancer. However, many prostate cancers have subtle features which are not easily identifiable on MRI, resulting in missed diagnoses and alarming variability in radiologist interpretation. Machine learning models have been developed in an effort to improve cancer identification, but current models localize cancer using MRI-derived features, while failing to consider the disease pathology characteristics observed on resected tissue. In this paper, we propose CorrSigNet, an automated two-step model that localizes prostate cancer on MRI by capturing the pathology features of cancer. First, the model learns MRI signatures of cancer that are correlated with corresponding histopathology features using Common Representation Learning. Second, the model uses the learned correlated MRI features to train a Convolutional Neural Network to localize prostate cancer. The histopathology images are used only in the first step to learn the correlated features. Once learned, these correlated features can be extracted from MRI of new patients (without histopathology or surgery) to localize cancer. We trained and validated our framework on a unique dataset of 75 patients with 806 slices who underwent MRI followed by prostatectomy surgery. We tested our method on an independent test set of 20 prostatectomy patients (139 slices, 24 cancerous lesions, 1.12M pixels) and achieved a per-pixel sensitivity of 0.81, specificity of 0.71, AUC of 0.86 and a per-lesion AUC of \(0.96 \pm 0.07\), outperforming the current state-of-the-art accuracy in predicting prostate cancer using MRI. Indrani Bhattacharya, Arun Seetharaman, Wei Shao 0008, Rewa Sood, Christian Kunder, Richard E. Fan, Simon J. C. Soerensen, Jeffrey B. Wang, Pejman Ghanouni, Nikola C. Teslovich, James D. Brooks, Geoffrey A. Sonn, Mirabela Rusu |
MICCAI (2) | 3 |
| 2020 | N-Phase Local Expansion Ratio for Characterizing Out-of-Phase Lung VentilationabstractOut-of-phase ventilation occurs when local regions of the lung reach their maximum or minimum volumes at breathing phases other than the global end inhalation or exhalation phases. This paper presents the N-phase local expansion ratio (LERN) as a surrogate for lung ventilation. A common approach to estimate lung ventilation is to use image registration to align the end exhalation and inhalation 3DCT images and then analyze the resulting correspondence map. This 2-phase local expansion ratio (LER2) is limited because it ignores out-of-phase ventilation and thus may underestimate local lung ventilation. To overcome this limitation, LERNmeasures the maximum ratio of local expansion and contraction over the entire breathing cycle. Comparing LER2to LERNprovides a means for detecting and characterizing locations of the lung that experience out-of-phase ventilation. We present a novel in-phase/out-of-phase ventilation (IOV) function plot to visualize and measure the amount of high-function IOV that occurs during a breathing cycle. Treatment planning 4DCT scans collected during coached breathing from 32 human subjects with lung cancer were analyzed in this study. Results show that out-of-phase breathing occurred in all subjects and that the spatial distribution of out-of-phase ventilation varied from subject to subject. For the 32 subjects analyzed, 50% of the out-of-phase regions on average were mislabeled as low-function by LER2(high-function threshold of 1.1, IOV threshold of 1.05). 4DCT and Xenon-enhanced CT of four sheep showed that LER8is more accurate than LER2 for measuring lung ventilation. Wei Shao 0008, Taylor Patton, Sarah E. Gerard, Yue Pan 0013, Joseph M. Reinhardt, Oguz C. Durumeric, John E. Bayouth, Gary E. Christensen |
IEEE Trans. Medical Imaging | 1 |