VLDB 2026 Research / reviewers in the wild / expert
Ali Bashashati
dblp:22/9218
· DBLP profile ↗
17ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0002-4212-7224ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 since 2021Artificial intelligence and machine learning · 6 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HistoMILKD: A Multiple Instance Learning based Multi-Teacher Knowledge Distillation Framework for Whole Slide Image ClassificationabstractFoundation models (FM) in digital pathology have revolutionized the field of whole slide image (WSI) analysis, with models such as UNI, Virchow, Prov-GigaPath, and many more outperforming the previously established benchmarks set by the ImageNet-based backbones. However, despite several benchmarking studies, there has been no clear consensus on the choice of a single FM that is best suited for a variety of histopathology datasets and/or tasks. With more than 25 pathology FMs in the literature so far, the challenge of model selection is a growing concern. Although an ensemble of FMs can circumvent this issue, given the bulky nature of individual FMs, the inference time and computational cost drastically increase with the addition of each FM into the ensemble. To this end, we propose HistoMILKD, the first multi-teacher knowledge distillation (MKD) framework for WSI classification. To handle the gigapixel resolution of WSIs, we use multiple instance learning (MIL), making this also the first work to integrate MIL and MKD frameworks into a single model. Our approach leverages the complementary representations of different FMs to distill collective task-specific knowledge into a single trainable MIL adapter on top of the student FM, which is utilized during inference. Evaluated on five public datasets, the proposed approach significantly (p < 0.05) outperforms the individual FMs, their ensemble, and previous MKD approaches in WSI classification. Codes are available at https://github.com/AIMLab-UBC/HistoMILKD. Mayur Mallya, Ali Khajegili Mirabadi, Hossein Farahani, Ali Bashashati |
WACV | 4 |
| 2026 | KAFSTExp: Kernel Adaptive Filtering With Nyström Approximation for Predicting Spatial Gene Expression From Histology ImagesabstractSpatial transcriptomics (ST), known as an expensive medical examination, plays an important role in analyzing the spatial heterogeneity of tumors. When considering the correlation between tissue morphological patterns and gene profiles, predicting corresponding gene expression from pathology images obtained from affordable biopsies is regarded as an instantaneous and cost-effective alternative. However, accurately modeling the complex and nonlinear relationship between histological features and gene expression remains challenging. Existing deep learning models often struggle to generalize on limited ST datasets due to their large and overparameterized architectures. The primary advantage of kernel adaptive filtering (KAF) lies in its ability to transform a challenging nonlinear problem arising in the original space into a linear regression problem in the higher-dimensional feature space via kernel methods. Therefore, this paper proposes a framework called KAFSTExp, which utilizes the state-of-the-art pathology foundation model UNI to extract image feature vectors, and then introduces the kernel least mean square algorithm with Nyström approximation to predict the normalized transcript counts of specific genes. Extensive experiments show that KAFSTExp significantly improves prediction accuracy while reducing computational cost and training time. KAFSTExp demonstrates consistent performance gains across multiple ST datasets, achieving relative improvements in Pearson correlation coefficient ranging from 1.24% to 94.23%, with an average increase of 19.80% over the best-performing non-KAF methods. External validation and further clinical analysis confirm the generalization performance and clinical application value of the proposed KAFSTExp. Hossein Farahani, Xifeng Li, Yongle Xie, Ali Bashashati |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Boltzmann Semantic Score: A Semantic Metric for Evaluating Large Vision Models Using Large Language ModelsabstractDo Large Vision Models (LVMs) extract medically and semantically relevant features similar to those identified by human experts? Currently, only biased, qualitative approaches with limited, small-scale expert evaluations are available to answer this question. In this study, we propose the Boltzmann Semantic Score (BSS), a novel method inspired by state space modeling, to evaluate the encoding space of LVMs from medical images using the encoding space of Large Language Models (LLMs) from medical reports. Through extensive experimentation on 32 datasets from The Cancer Genome Atlas collection using five state-of-the-art LLMs, we first establish a baseline of LLMs' performance in digital pathology and show that LLMs' encoding can be linked to patient outcomes. Then, we compared seven LVMs with BSS and showed that LVMs suffer from poor semantic capability when compared with encoded expert knowledge from pathology reports.
We also found statistically significant correlations between BSS (as a measure of structural similarity) and performance in two downstream tasks: information retrieval and survival prediction tasks. Our study also investigates the consensus among LLMs in evaluating LVMs using BSS, indicating that LLMs generally reach substantial consensus in rating LVMs, with some variation dependant on the cancer type. We believe the BSS metric proposed here holds significant potential for application in other domains with similar contexts. Data and code can be found in \footnotesize \url{ https://github.com/AIMLab-UBC/Boltzmann} Ali Khajegili Mirabadi, Katherine Rich, Hossein Farahani, Ali Bashashati |
ICLR | 4 |
| 2025 | ATEC23 Challenge: Automated prediction of treatment effectiveness in ovarian cancer using histopathological images
Ching-Wei Wang, Nabila Puspita Firdi, Tzu-Chiao Chu, Mohammad Faiz Iqbal Faiz, Mohammad Zafar Iqbal, Mayur Mallya, Ali Bashashati, Fei Li 0021, Mengkang Lu, Yong Xia 0001, Tai-Kuang Chao |
Medical Image Anal. | 9 |
| 2024 | How Molecules Impact Cells: Unlocking Contrastive PhenoMolecular RetrievalabstractPredicting molecular impact on cellular function is a core challenge in therapeutic design. Phenomic experiments, designed to capture cellular morphology, utilize microscopy based techniques and demonstrate a high throughput solution for uncovering molecular impact on the cell. In this work, we learn a joint latent space between molecular structures and microscopy phenomic experiments, aligning paired samples with contrastive learning. Specifically, we study the problem of Contrastive PhenoMolecular Retrieval, which consists of zero-shot molecular structure identification conditioned on phenomic experiments. We assess challenges in multi-modal learning of phenomics and molecular modalities such as experimental batch effect, inactive molecule perturbations, and encoding perturbation concentration. We demonstrate improved multi-modal learner retrieval through (1) a uni-modal pre-trained phenomics model, (2) a novel inter sample similarity aware loss, and (3) models conditioned on a representation of molecular concentration. Following this recipe, we propose MolPhenix, a molecular phenomics model. MolPhenix leverages a pre-trained phenomics model to demonstrate significant performance gains across perturbation concentrations, molecular scaffolds, and activity thresholds. In particular, we demonstrate an 8.1 times improvement in zero shot molecular retrieval of active molecules over the previous state-of-the-art, reaching 77.33% in top-1% accuracy. These results open the door for machine learning to be applied in virtual phenomics screening, which can significantly benefit drug discovery applications. Philip Fradkin, Puria Azadi Moghadam, Karush Suri, Frederik Wenkel, Ali Bashashati, Maciej Sypetkowski, Dominique Beaini |
NeurIPS | 5 |
| 2024 | Multi-scale relational graph convolutional network for multiple instance learning in histopathology imagesabstractGraph convolutional neural networks have shown significant potential in natural and histopathology images. However, their use has only been studied in a single magnification or multi-magnification with either homogeneous graphs or only different node types. In order to leverage the multi-magnification information and improve message passing with graph convolutional networks, we handle different embedding spaces at each magnification by introducing the Multi-Scale Relational Graph Convolutional Network (MS-RGCN) as a multiple instance learning method. We model histopathology image patches and their relation with neighboring patches and patches at other scales (i.e., magnifications) as a graph. We define separate message-passing neural networks based on node and edge types to pass the information between different magnification embedding spaces. We experiment on prostate cancer histopathology images to predict the grade groups based on the extracted features from patches. We also compare our MS-RGCN with multiple state-of-the-art methods with evaluations on several source and held-out datasets. Our method outperforms the state-of-the-art on all of the datasets and image types consisting of tissue microarrays, whole-mount slide regions, and whole-slide images. Through an ablation study, we test and show the value of the pertinent design features of the MS-RGCN. Roozbeh Bazargani, Ladan Fazli, Martin E. Gleave, Larry Goldenberg, Ali Bashashati, Tim Salcudean |
Medical Image Anal. | 5 |
| 2023 | Sparse Multi-Modal Graph Transformer with Shared-Context Processing for Representation Learning of Giga-pixel ImagesabstractProcessing giga-pixel whole slide histopathology images (WSI) is a computationally expensive task. Multiple instance learning (MIL) has become the conventional approach to process WSIs, in which these images are split into smaller patches for further processing. However, MIL-based techniques ignore explicit information about the individual cells within a patch. In this paper, by defining the novel concept of shared-context processing, we designed a multi-modal Graph Transformer (AMIGO) that uses the cellular graph within the tissue to provide a single representation for a patient while taking advantage of the hierarchical structure of the tissue, enabling a dynamic focus between cell-level and tissue-level information. We benchmarked the performance of our model against multiple state-of-the-art methods in survival prediction and showed that ours can significantly outperform all of them including hierarchical Vision Transformer (ViT). More importantly, we show that our model is strongly robust to missing information to an extent that it can achieve the same performance with as low as 20% of the data. Finally, in two different cancer datasets, we demonstrated that our model was able to stratify the patients into low-risk and high-risk groups while other state-of-the-art methods failed to achieve this goal. We also publish a large dataset of immunohistochemistry images (InUIT) containing 1,600 tissue microarray (TMA) cores from 188 patients along with their survival information, making it one of the largest publicly available datasets in this context. Ramin Nakhli, Puria Azadi Moghadam, Haoyang Mi, Hossein Farahani, Alexander Baras, C. Blake Gilks, Ali Bashashati |
CVPR | 7 |
| 2023 | CO-PILOT: Dynamic Top-Down Point Cloud with Conditional Neighborhood Aggregation for Multi-Gigapixel Histopathology Image RepresentationabstractPredicting survival rates based on multi-gigapixel histopathology images is one of the most challenging tasks in digital pathology. Due to the computational complexities, Multiple Instance Learning (MIL) has become the conventional approach for this process as it breaks the image into smaller patches. However, this technique fails to account for the individual cells present in each patch, while they are the fundamental part of the tissue. In this work, we developed a novel dynamic and hierarchical point-cloud-based method (CO-PILOT) for the processing of cellular graphs extracted from routine histopathology images. By using bottom-up information propagation and top-down conditional attention, our model gains access to an adaptive focus across different levels of tissue hierarchy. Through comprehensive experiments, we demonstrate that our model can outperform all the state-of-the-art methods in survival prediction, including the hierarchical Vision Transformer (ViT), across three datasets and four metrics with only half of the parameters of the closest baseline. Importantly, our model is able to stratify the patients into different risk cohorts with statistically different outcomes across three large datasets, a task that was previously achievable only using genomic information. Furthermore, we publish a large dataset containing 873 cellular graphs from 188 patients, along with their survival information, making it one of the largest publicly available datasets in this context. Ramin Nakhli, Allen W. Zhang, Ali Khajegili Mirabadi, Katherine Rich, Maryam Asadi-Aghbolaghi, C. Blake Gilks, Hossein Farahani, Ali Bashashati |
ICCV | 8 |
| 2023 | ALL-IN: ALocal GLobal Graph-Based DIstillatioN Model for Representation Learning of Gigapixel Histopathology Images With Application In Cancer Risk Assessment
Puria Azadi, Jonathan Suderman, Ramin Nakhli, Katherine Rich, Maryam Asadi-Aghbolaghi, Sonia Kung, Htoo Oo, Mira Keyes, Hossein Farahani, Calum MacAulay, Larry Goldenberg, Peter Black, Ali Bashashati |
MICCAI (6) | 13 |
| 2023 | A Morphology Focused Diffusion Probabilistic Model for Synthesis of Histopathology ImagesabstractVisual microscopic study of diseased tissue by pathologists has been the cornerstone for cancer diagnosis and prognostication for more than a century. Recently, deep learning methods have made significant advances in the analysis and classification of tissue images. However, there has been limited work on the utility of such models in generating histopathology images. These synthetic images have several applications in pathology including utilities in education, proficiency testing, privacy, and data sharing. Recently, diffusion probabilistic models were introduced to generate high quality images. Here, for the first time, we investigate the potential use of such models along with prioritized morphology weighting and color normalization to synthesize high quality histopathology images of brain cancer. Our detailed results show that diffusion probabilistic models are capable of synthesizing a wide range of histopathology images and have superior performance compared to generative adversarial networks. Puria Azadi Moghadam, Sanne Van Dalen, Karina C. Martin, Jochen K. Lennerz, Stephen Yip, Hossein Farahani, Ali Bashashati |
WACV | 7 |
| 2019 | Integrated structural variation and point mutation signatures in cancer genomes using correlated topic modelsabstractMutation signatures in cancer genomes reflect endogenous and exogenous mutational processes, offering insights into tumour etiology, features for prognostic and biologic stratification and vulnerabilities to be exploited therapeutically. We present a novel machine learning formalism for improved signature inference, based on multi-modal correlated topic models (MMCTM) which can at once infer signatures from both single nucleotide and structural variation counts derived from cancer genome sequencing data. We exemplify the utility of our approach on two hormone driven, DNA repair deficient cancers: breast and ovary (n = 755 samples total). We show how introducing correlated structure both within and between modes of mutation can increase accuracy of signature discovery, particularly in the context of sparse data. Our study emphasizes the importance of integrating multiple mutation modes for signature discovery and patient stratification, and provides a statistical modeling framework to incorporate additional features of interest for future studies. Tyler Funnell, Allen W. Zhang, Diljot Grewal, Steven McKinney, Ali Bashashati, Yi Kan Wang, Sohrab P. Shah |
PLoS Comput. Biol. | 5 |
| 2016 | Neural Network Conditional Random Fields for Self-Paced Brain Computer InterfacesabstractThe task of classifying EEG signals for self-paced Brain Computer Interface (BCI) applications is extremely challenging. This difficulty in classification of self-paced data stems from the fact that the system has no clue about the start time of a control task and the data contains a large number of periods during which the user has no intention to control the BCI. Therefore, to improve the performance of the BCI, it is imperative to exploit the characteristics of the EEG data as much as possible. For motor imagery based self-paced BCIs, during motor imagery task the EEG signal of each subject goes through several internal state changes. Applying appropriate classifiers that can exploit the temporal correlation in EEG data can enhance the performance of the BCI. In this paper, we propose an algorithm which is able to capture the temporal correlation of the EEG signal. We compare the performance of our algorithm that is based on neural network conditional random fields to two well-known dynamic classifiers, the Hidden Markov Models and Conditional Random Fields and to the static classifier, Support Vector Machines. We compare these methods using the data from SM2 dataset, and we show that our algorithm yields results that are considerably superior to the other approaches in terms of the Area Under the Curve (AUC) of the BCI system. Hossein Bashashati, Rabab K. Ward, Ali Bashashati, Amr M. Mohamed |
ICMLA | 3 |
| 2015 | Hidden Markov Support Vector Machines for Self-Paced Brain Computer InterfacesabstractBrain Computer Interfaces (BCI) aim at providing a means to control devices with brain signals. Self-paced BCIs, as opposed to synchronous ones, have the advantage of being operational at all times and not only at specific system-defined periods. Traditionally, in the BCI field, a sliding window over the brain signal is used to detect the intention of the user at a given time. This approach ignores the temporal correlations between the adjacent time windows. This paper proposes a novel approach to classify self-paced BCI data using structural support vector machines. Our proposed approach considers the history of the brain signals in the context of sequential supervised learning to better detect the intention of the user from his/her brain signals. We have compared our proposed model to the sliding window approach with Support Vector Machines (SVM) and Linear Discriminant Analysis (LDA) classifiers. Using data collected from 4 individuals form BCI competition IV, it is shown that the F1 score of our approach is significantly better than the sliding window approach. The average F1 score of our method across all subjects is 0.3 and 0.5 higher than the sliding window with SVM and LDA classifiers, respectively. Hossein Bashashati, Rabab K. Ward, Ali Bashashati |
ICMLA | 3 |
| 2012 | Feature-based classifiers for somatic mutation detection in tumour-normal paired sequencing dataabstractMOTIVATION: The study of cancer genomes now routinely involves using next-generation sequencing technology (NGS) to profile tumours for single nucleotide variant (SNV) somatic mutations. However, surprisingly few published bioinformatics methods exist for the specific purpose of identifying somatic mutations from NGS data and existing tools are often inaccurate, yielding intolerably high false prediction rates. As such, the computational problem of accurately inferring somatic mutations from paired tumour/normal NGS data remains an unsolved challenge. RESULTS: We present the comparison of four standard supervised machine learning algorithms for the purpose of somatic SNV prediction in tumour/normal NGS experiments. To evaluate these approaches (random forest, Bayesian additive regression tree, support vector machine and logistic regression), we constructed 106 features representing 3369 candidate somatic SNVs from 48 breast cancer genomes, originally predicted with naive methods and subsequently revalidated to establish ground truth labels. We trained the classifiers on this data (consisting of 1015 true somatic mutations and 2354 non-somatic mutation positions) and conducted a rigorous evaluation of these methods using a cross-validation framework and hold-out test NGS data from both exome capture and whole genome shotgun platforms. All learning algorithms employing predictive discriminative approaches with feature selection improved the predictive accuracy over standard approaches by statistically significant margins. In addition, using unsupervised clustering of the ground truth 'false positive' predictions, we noted several distinct classes and present evidence suggesting non-overlapping sources of technical artefacts illuminating important directions for future study. AVAILABILITY: Software called MutationSeq and datasets are available from http://compbio.bccrc.ca. Jiarui Ding, Ali Bashashati, Andrew Roth, Arusha Oloumi, Kane Tse, Thomas Zeng 0002, Gholamreza Haffari, Martin Hirst, Marco A. Marra, Anne Condon, Samuel Aparicio, Sohrab P. Shah |
Bioinform. | 2 |
| 2012 | JointSNVMix: a probabilistic model for accurate detection of somatic mutations in normal/tumour paired next-generation sequencing dataabstractMOTIVATION: Identification of somatic single nucleotide variants (SNVs) in tumour genomes is a necessary step in defining the mutational landscapes of cancers. Experimental designs for genome-wide ascertainment of somatic mutations now routinely include next-generation sequencing (NGS) of tumour DNA and matched constitutional DNA from the same individual. This allows investigators to control for germline polymorphisms and distinguish somatic mutations that are unique to the tumour, thus reducing the burden of labour-intensive and expensive downstream experiments needed to verify initial predictions. In order to make full use of such paired datasets, computational tools for simultaneous analysis of tumour-normal paired sequence data are required, but are currently under-developed and under-represented in the bioinformatics literature. RESULTS: In this contribution, we introduce two novel probabilistic graphical models called JointSNVMix1 and JointSNVMix2 for jointly analysing paired tumour-normal digital allelic count data from NGS experiments. In contrast to independent analysis of the tumour and normal data, our method allows statistical strength to be borrowed across the samples and therefore amplifies the statistical power to identify and distinguish both germline and somatic events in a unified probabilistic framework. AVAILABILITY: The JointSNVMix models and four other models discussed in the article are part of the JointSNVMix software package available for download at http://compbio.bccrc.ca CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Andrew Roth, Jiarui Ding, Ryan D. Morin, Anamaria Crisan, Gavin Ha, Ryan Giuliany, Ali Bashashati, Martin Hirst, Gulisa Turashvili, Arusha Oloumi, Marco A. Marra, Samuel Aparicio, Sohrab P. Shah |
Bioinform. | 7 |
| 2006 | Detection of Hand Extension Movements in the Context of a 3-State Asynchronous Brain InterfaceabstractThe low-frequency asynchronous switch design (LF-ASD) is a direct brain interface (BI) that detects the presence of a specific finger movement in the ongoing EEG. Asynchronous interfaces have the advantage of being operational at all times and not only at specific system-defined periods. In this paper, we present the design of a 3-state asynchronous BI for the detection of two different movements from the ongoing EEG. The proposed 3-state asynchronous BI detects right and left hand extensions. Using data collected from two able-bodied individuals, it is shown that the error characteristics of the new system in detecting the presence of movement are significantly better than the 2-state LF-ASD, with true positive rate increases of up to 22.4% for false positive rates in the 1-2% range. An average performance of 61.5% was achieved in differentiating between left and right hand movements Ali Bashashati, Rabab K. Ward, Gary E. Birch |
ICASSP (5) | 1 |
| 2005 | A hybrid genetic algorithm approach for improving the performance of the LF-ASD brain computer interfaceabstractAn asynchronous brain computer interface (BCI) continuously monitors the brain signals and is activated only when a user intends control. Initial results from an asynchronous system, the LF-ASD, designed by our group have shown promise, but the reported error rates are still high for most practical applications. To improve its performance, we propose user customization. Since energy normalization of all channels' signals is shown to significantly improve the performance of the system, we choose to customize the parameters related to this normalization. We apply a hybrid genetic algorithm (a genetic algorithm followed by a local search) to customize the size of the energy normalization windows. This is shown to significantly improve the results. For a fixed false positive rate of 2%, the improvement in the true positive rate was raised from 65.7% to 76.9% in one subject and from 53.1% to 63.3% for another subject. Mehrdad Fatourechi, Ali Bashashati, Rabab K. Ward, Gary E. Birch |
ICASSP (5) | 2 |