EDBT 2026 Demo / reviewers in the wild / expert
Drew F. K. Williamson
dblp:229/4401
· DBLP profile ↗
13ranked-venue papers
0as first author
13since 2021 · last 2025
0000-0003-1745-8846ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A robust image segmentation and synthesis pipeline for histopathology
Muhammad Jehanzaib, Yasin Almalioglu, Kutsev Bengisu Ozyoruk, Drew F. K. Williamson, Talha Abdullah, Kayhan Basak, Derya Demir, G. Evren Keles, Kashif Zafar, Mehmet Turan |
Medical Image Anal. | 4 |
| 2024 | Transcriptomics-Guided Slide Representation Learning in Computational PathologyabstractSelf-supervised learning (SSL) has been successful in building patch embeddings of small histology images (e.g.,$224\times 224$pixels), but scaling these models to learn slide embeddings from the entirety of giga-pixel whole-slide images (WSIs) remains challenging. Here, we leverage complementary information from gene expression profiles to guide slide representation learning using multi-modal pre-training. Expression profiles constitute highly detailed molecular descriptions of a tissue that we hypothesize offer a strong task-agnostic training signal for learning slide embeddings. Our slide and expression (S+E) pre-training strategy, called Tangle, employs modality-specific encoders, the outputs of which are aligned via contrastive learning. Tangle was pre-trained on samples from three different organs: liver (n=6,597 S+E pairs), breast (n=1,020), and lung (n=1,012) from two different species (Homo sapiens and Rattus norvegicus). Across three independent test datasets consisting of 1,265 breast WSIs, 1,946 lung WSIs, and 4,584 liver WSIs, Tangle shows significantly better few-shot performance compared to supervised and SSL baselines. When assessed using prototype-based classification and slide retrieval, Tangle also shows a substantial performance improvement over all baselines. Code available at https://github.com/mahmoodlab/TANGLE. Guillaume Jaume, Lukas Oldenburg, Anurag Vaidya, Richard J. Chen, Drew F. K. Williamson, Thomas Peeters, Andrew H. Song, Faisal Mahmood 0001 |
CVPR | 5 |
| 2024 | Modeling Dense Multimodal Interactions Between Biological Pathways and Histology for Survival PredictionabstractIntegrating whole-slide images (WSIs) and bulk tran-scriptomics for predicting patient survival can improve our understanding of patient prognosis. However, this multi-modal task is particularly challenging due to the different nature of these data: WSIs represent a very high-dimensional spatial description of a tumor, while bulk tran-scriptomics represent a global description of gene expression levels within that tumor. In this context, our work aims to address two key challenges: (1) how can we tokenize transcriptomics in a semantically meaningful and interpretable way?, and (2) how can we capture dense multi-modal interactions between these two modalities? Here, we propose to learn biological pathway tokens from transcriptomics that can encode specific cellular functions. Together with histology patch tokens that encode the slide morphology, we argue that they form appropriate reasoning units for interpretability. We fuse both modalities using a memoryefficient multimodal Transformer that can model interactions between pathway and histology patch tokens. Our model, Survpath, achieves state-of-the-art performance when evaluated against unimodal and multimodal baselines on five datasets from The Cancer Genome Atlas. Our interpretability framework identifies key multimodal prognostic factors, and, as such, can provide valuable insights into the interaction between genotype and phenotype. Code available at https://github.com/mahmoodlab/SurvPath. Guillaume Jaume, Anurag Vaidya, Richard J. Chen, Drew F. K. Williamson, Paul Pu Liang, Faisal Mahmood 0001 |
CVPR | 4 |
| 2024 | Morphological Prototyping for Unsupervised Slide Representation Learning in Computational PathologyabstractRepresentation learning of pathology whole-slide images (WSIs) has been has primarily relied on weak supervision with Multiple Instance Learning (MIL). However, the slide representations resulting from this approach are highly tailored to specific clinical tasks, which limits their expressivity and generalization, particularly in scenarios with limited data. Instead, we hypothesize that morphological redundancy in tissue can be leveraged to build a task-agnostic slide representation in an unsupervised fashion. To this end, we introduce Panther, a prototype-based approach rooted in the Gaussian mixture model that summarizes the set of WSI patches into a much smaller set of morphological prototypes. Specifically, each patch is assumed to have been generated from a mixture distri-bution, where each mixture component represents a morphological exemplar. Utilizing the estimated mixture parameters, we then construct a compact slide representation that can be readily used for a wide range of downstream tasks. By performing an extensive evaluation of Panther on subtyping and survival tasks using 13 datasets, we show that 1) Panther outperforms or is on par with super-vised MIL baselines and 2) the analysis of morphological prototypes brings new qualitative and quantitative in-sights into model interpretability. The code is available at https://github.com/mahmoodlab/Panther. Andrew H. Song, Richard J. Chen, Drew F. K. Williamson, Guillaume Jaume, Faisal Mahmood 0001 |
CVPR | 4 |
| 2024 | HEST-1k: A Dataset For Spatial Transcriptomics and Histology Image AnalysisabstractSpatial transcriptomics enables interrogating the molecular composition of tissue with ever-increasing resolution and sensitivity. However, costs, rapidly evolving technology, and lack of standards have constrained computational methods in ST to narrow tasks and small cohorts. In addition, the underlying tissue morphology, as reflected by H&E-stained whole slide images (WSIs), encodes rich information often overlooked in ST studies. Here, we introduce HEST-1k, a collection of 1,229 spatial transcriptomic profiles, each linked to a WSI and extensive metadata. HEST-1k was assembled from 153 public and internal cohorts encompassing 26 organs, two species (Homo Sapiens and Mus Musculus), and 367 cancer samples from 25 cancer types. HEST-1k processing enabled the identification of 2.1 million expression-morphology pairs and over 76 million nuclei. To support its development, we additionally introduce the HEST-Library, a Python package designed to perform a range of actions with HEST samples. We test HEST-1k and Library on three use cases: (1) benchmarking foundation models for pathology (HEST-Benchmark), (2) biomarker exploration, and (3) multimodal representation learning. HEST-1k, HEST-Library, and HEST-Benchmark can be freely accessed at https://github.com/mahmoodlab/hest. Guillaume Jaume, Paul Doucet, Andrew H. Song, Ming Y. Lu, Cristina Almagro-Pérez, Sophia J. Wagner, Anurag Vaidya, Richard J. Chen, Drew F. K. Williamson, Ahrong Kim, Faisal Mahmood 0001 |
NeurIPS | 9 |
| 2023 | Visual Language Pretrained Multiple Instance Zero-Shot Transfer for Histopathology ImagesabstractContrastive visual language pretraining has emerged as a powerful method for either training new language-aware image encoders or augmenting existing pretrained models with zero-shot visual recognition capabilities. However, existing works typically train on large datasets of imagetext pairs and have been designed to perform downstream tasks involving only small to medium sized-images, neither of which are applicable to the emerging field of computational pathology where there are limited publicly available paired image-text datasets and each image can span up to 100,000 × 100,000 pixels. In this paper we present MI-Zero, a simple and intuitive framework for unleashing the zero-shot transfer capabilities of contrastively aligned image and text models on gigapixel histopathology whole slide images, enabling multiple downstream diagnostic tasks to be carried out by pretrained encoders without requiring any additional labels. MI-Zero reformulates zero-shot transfer under the framework of multiple instance learning to overcome the computational challenge of inference on extremely large images. We used over 550k pathology reports and other available in-domain text corpora to pretrain our text encoder. By effectively leveraging strong pretrained encoders, our best model pretrained on over 33k histopathology image-caption pairs achieves an average median zero-shot accuracy of 70.2% across three different real-world cancer subtyping tasks. Our code is available at: https://github.com/mahmoodlab/MI-Zero. Ming Y. Lu, Drew F. K. Williamson, Richard J. Chen, Long Phi Le, Yung-Sung Chuang, Faisal Mahmood 0001 |
CVPR | 4 |
| 2022 | Differentiable Zooming for Multiple Instance Learning on Whole-Slide Images
Kevin Thandiackal, Boqi Chen, Pushpak Pati, Guillaume Jaume, Drew F. K. Williamson, Maria Gabrani, Orcun Goksel |
ECCV (21) | 5 |
| 2022 | Incorporating Intratumoral Heterogeneity into Weakly-Supervised Deep Learning Models via Variance Pooling
Iain Carmichael, Andrew H. Song, Richard J. Chen, Drew F. K. Williamson, Tiffany Y. Chen, Faisal Mahmood 0001 |
MICCAI (2) | 4 |
| 2022 | Federated learning for computational pathology on gigapixel whole slide imagesabstractDeep Learning-based computational pathology algorithms have demonstrated profound ability to excel in a wide array of tasks that range from characterization of well known morphological phenotypes to predicting non human-identifiable features from histology such as molecular alterations. However, the development of robust, adaptable and accurate deep learning-based models often rely on the collection and time-costly curation large high-quality annotated training data that should ideally come from diverse sources and patient populations to cater for the heterogeneity that exists in such datasets. Multi-centric and collaborative integration of medical data across multiple institutions can naturally help overcome this challenge and boost the model performance but is limited by privacy concerns among other difficulties that may arise in the complex data sharing process as models scale towards using hundreds of thousands of gigapixel whole slide images. In this paper, we introduce privacy-preserving federated learning for gigapixel whole slide images in computational pathology using weakly-supervised attention multiple instance learning and differential privacy. We evaluated our approach on two different diagnostic problems using thousands of histology whole slide images with only slide-level labels. Additionally, we present a weakly-supervised learning framework for survival prediction and patient stratification from whole slide images and demonstrate its effectiveness in a federated setting. Our results show that using federated learning, we can effectively develop accurate weakly-supervised deep learning models from distributed data silos without direct data sharing and its associated complexities, while also preserving differential privacy using randomized noise generation. We also make available an easy-to-use federated learning for computational pathology software package: http://github.com/mahmoodlab/HistoFL. Ming Y. Lu, Richard J. Chen, Dehan Kong, Jana Lipková, Rajendra Singh, Drew F. K. Williamson, Tiffany Y. Chen, Faisal Mahmood 0001 |
Medical Image Anal. | 6 |
| 2022 | Pathomic Fusion: An Integrated Framework for Fusing Histopathology and Genomic Features for Cancer Diagnosis and PrognosisabstractCancer diagnosis, prognosis, mymargin and therapeutic response predictions are based on morphological information from histology slides and molecular profiles from genomic data. However, most deep learning-based objective outcome prediction and grading paradigms are based on histology or genomics alone and do not make use of the complementary information in an intuitive manner. In this work, we propose Pathomic Fusion, an interpretable strategy for end-to-end multimodal fusion of histology image and genomic (mutations, CNV, RNA-Seq) features for survival outcome prediction. Our approach models pairwise feature interactions across modalities by taking the Kronecker product of unimodal feature representations, and controls the expressiveness of each representation via a gating-based attention mechanism. Following supervised learning, we are able to interpret and saliently localize features across each modality, and understand how feature importance shifts when conditioning on multimodal input. We validate our approach using glioma and clear cell renal cell carcinoma datasets from the Cancer Genome Atlas (TCGA), which contains paired whole-slide image, genotype, and transcriptome data with ground truth survival and histologic grade labels. In a 15-fold cross-validation, our results demonstrate that the proposed multimodal fusion paradigm improves prognostic determinations from ground truth grading and molecular subtyping, as well as unimodal deep networks trained on histology and genomic data alone. The proposed method establishes insight and theory on how to train deep networks on multimodal biomedical data in an intuitive manner, which will be useful for other problems in medicine that seek to combine heterogeneous data streams for understanding diseases and predicting response and resistance to treatment. Code and trained models are made available at: https://github.com/mahmoodlab/PathomicFusion. Richard J. Chen, Ming Y. Lu, Drew F. K. Williamson, Scott J. Rodig, Neal I. Lindeman, Faisal Mahmood 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2021 | Multimodal Co-Attention Transformer for Survival Prediction in Gigapixel Whole Slide ImagesabstractSurvival outcome prediction is a challenging weakly-supervised and ordinal regression task in computational pathology that involves modeling complex interactions within the tumor microenvironment in gigapixel whole slide images (WSIs). Despite recent progress in formulating WSIs as bags for multiple instance learning (MIL), representation learning of entire WSIs remains an open and challenging problem, especially in overcoming: 1) the computational complexity of feature aggregation in large bags, and 2) the data heterogeneity gap in incorporating biological priors such as genomic measurements. In this work, we present a Multimodal Co-Attention Transformer (MCAT) framework that learns an interpretable, dense co-attention mapping between WSIs and genomic features formulated in an embedding space. Inspired by approaches in Visual Question Answering (VQA) that can attribute how word embed-dings attend to salient objects in an image when answering a question, MCAT learns how histology patches attend to genes when predicting patient survival. In addition to visualizing multimodal interactions, our co-attention trans-formation also reduces the space complexity of WSI bags, which enables the adaptation of Transformer layers as a general encoder backbone in MIL. We apply our proposed method on five different cancer datasets (4,730 WSIs, 67 million patches). Our experimental results demonstrate that the proposed method consistently achieves superior performance compared to the state-of-the-art methods. Richard J. Chen, Ming Y. Lu, Wei-Hung Weng, Tiffany Y. Chen, Drew F. K. Williamson, Trevor Manz, Maha Shady, Faisal Mahmood 0001 |
ICCV | 5 |
| 2021 | Whole Slide Images are 2D Point Clouds: Context-Aware Survival Prediction Using Patch-Based Graph Convolutional Networks
Richard J. Chen, Ming Y. Lu, Muhammad Shaban, Chengkuan Chen, Tiffany Y. Chen, Drew F. K. Williamson, Faisal Mahmood 0001 |
MICCAI (8) | 6 |
| 2021 | Network potential identifies therapeutic miRNA cocktails in Ewing sarcomaabstractMicroRNA (miRNA)-based therapies are an emerging class of targeted therapeutics with many potential applications. Ewing Sarcoma patients could benefit dramatically from personalized miRNA therapy due to inter-patient heterogeneity and a lack of druggable (to this point) targets. However, because of the broad effects miRNAs may have on different cells and tissues, trials of miRNA therapies have struggled due to severe toxicity and unanticipated immune response. In order to overcome this hurdle, a network science-based approach is well-equipped to evaluate and identify miRNA candidates and combinations of candidates for the repression of key oncogenic targets while avoiding repression of essential housekeeping genes. We first characterized 6 Ewing sarcoma cell lines using mRNA sequencing. We then estimated a measure of tumor state, which we term network potential, based on both the mRNA gene expression and the underlying protein-protein interaction network in the tumor. Next, we ranked mRNA targets based on their contribution to network potential. We then identified miRNAs and combinations of miRNAs that preferentially act to repress mRNA targets with the greatest influence on network potential. Our analysis identified TRIM25, APP, ELAV1, RNF4, and HNRNPL as ideal mRNA targets for Ewing sarcoma therapy. Using predicted miRNA-mRNA target mappings, we identified miR-3613-3p, let-7a-3p, miR-300, miR-424-5p, and let-7b-3p as candidate optimal miRNAs for preferential repression of these targets. Ultimately, our work, as exemplified in the case of Ewing sarcoma, describes a novel pipeline by which personalized miRNA cocktails can be designed to maximally perturb gene networks contributing to cancer progression. Davis T. Weaver, Kathleen I. Pishas, Drew F. K. Williamson, Jessica Scarborough, Stephen L. Lessnick, Andrew Dhawan, Jacob G. Scott |
PLoS Comput. Biol. | 3 |