VLDB 2026 Research / reviewers in the wild / expert
Aly Azeem Khan
dblp:46/2390 · also Aly A. Khan
· DBLP profile ↗
10ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0003-3933-8538ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
4 papers |
Bioinformatics and computational biology · 84% Medical and health informatics · 16% | |
| Artificial intelligence
2 papers |
Probabilistic and Bayesian machine learning · 26% Image recognition and object detection · 23% Learning theory · 23% |
Topics — the 15 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › genomics › machine learning for genomics
deep learning for genomics |
0.9 | 1 | 2025 | PRSformer: Disease Prediction from Million-Scale Individual Genotypes · NeurIPS 2025 |
Medical and health informatics › clinical prediction
disease risk prediction |
0.9 | 1 | 2025 | PRSformer: Disease Prediction from Million-Scale Individual Genotypes · NeurIPS 2025 |
Bioinformatics and computational biology
genomics |
0.9 | 1 | 2025 | PRSformer: Disease Prediction from Million-Scale Individual Genotypes · NeurIPS 2025 |
Bioinformatics and computational biology › statistical genetics › genomic prediction
polygenic risk prediction |
0.9 | 1 | 2025 | PRSformer: Disease Prediction from Million-Scale Individual Genotypes · NeurIPS 2025 |
Machine learning › Learning theory › generalization
excess risk analysis |
0.8 | 1 | 2024 | Enhancing Instance-Level Image Classification with Set-Level Labels · ICLR 2024 |
Computer vision › Image recognition and object detection
image classification |
0.8 | 1 | 2024 | Enhancing Instance-Level Image Classification with Set-Level Labels · ICLR 2024 |
Bioinformatics and computational biology › immunoinformatics
immune repertoire analysis |
0.8 | 1 | 2024 | TRIBAL: Tree Inference of B Cell Clonal Lineages · RECOMB 2024 |
Machine learning › Efficient and distributed learning
active learning |
0.7 | 1 | 2023 | Scalable Batch-Mode Deep Bayesian Active Learning via Equivalence Class Annealing · ICLR 2023 |
Machine learning › Probabilistic and Bayesian machine learning › experimental design › bayesian experimental design
bayesian active learning |
0.7 | 1 | 2023 | Scalable Batch-Mode Deep Bayesian Active Learning via Equivalence Class Annealing · ICLR 2023 |
Bioinformatics and computational biology
protein-protein interaction prediction |
0.3 | 1 | 2018 | Predicting protein-protein interactions through sequence-based deep learning · Bioinform. 2018 |
Bioinformatics and computational biology › protein-protein interaction prediction
sequence-based PPI prediction |
0.3 | 1 | 2018 | Predicting protein-protein interactions through sequence-based deep learning · Bioinform. 2018 |
Bioinformatics and computational biology › single-cell analysis
single-cell genomics |
0.3 | 1 | 2017 | BASIC: BCR assembly from single cells · Bioinform. 2017 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
0.2 | 1 | 2024 | Enhancing Instance-Level Image Classification with Set-Level Labels · ICLR 2024 |
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models
bayesian deep learning |
0.2 | 1 | 2023 | Scalable Batch-Mode Deep Bayesian Active Learning via Equivalence Class Annealing · ICLR 2023 |
Bioinformatics and computational biology › single-cell analysis
single-cell RNA sequencing |
0.1 | 1 | 2017 | BASIC: BCR assembly from single cells · Bioinform. 2017 |
Methods — techniques the papers use, named apart from their topics
transformer · 0.9neighborhood attention · 0.9multi-task learning · 0.9theoretical risk analysis · 0.8set-level labels · 0.8phylogenetic tree inference · 0.8equivalence class annealing · 0.7batch mode active learning · 0.7siamese network · 0.3random projection · 0.3data augmentation · 0.3convolutional neural network · 0.3sequence assembly · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PRSformer: Disease Prediction from Million-Scale Individual GenotypesabstractPredicting disease risk from DNA presents an unprecedented emerging challenge as biobanks approach population scale sizes ($N>10^6$ individuals) with ultra-high-dimensional features ($L>10^5$ genotypes). Current methods, often linear and reliant on summary statistics, fail to capture complex genetic interactions and discard valuable individual-level information. We introduce **PRSformer**, a scalable deep learning architecture designed for end-to-end, multitask disease prediction directly from million-scale individual genotypes. PRSformer employs neighborhood attention, achieving linear $O(L)$ complexity per layer, making Transformers tractable for genome-scale inputs. Crucially, PRSformer utilizes a stacking of these efficient attention layers, progressively increasing the effective receptive field to model local dependencies (e.g., within linkage disequilibrium blocks) before integrating information across wider genomic regions. This design, tailored for genomics, allows PRSformer to learn complex, potentially non-linear and long-range interactions directly from raw genotypes. We demonstrate PRSformer's effectiveness using a unique large private cohort ($N \approx 5$M) for predicting 18 autoimmune and inflammatory conditions using $L \approx 140$k variants. PRSformer significantly outperforms highly optimized linear models trained on the *same individual-level data* and state-of-the-art summary-statistic-based methods (LDPred2) derived from the *same cohort*, quantifying the benefits of non-linear modeling and multitask learning at scale. Furthermore, experiments reveal that the advantage of non-linearity emerges primarily at large sample sizes ($N > 1$M), and that a multi-ancestry trained model improves generalization, establishing PRSformer as a new framework for deep learning in population-scale genomics. Payam Dibaeinia, Chris German, Suyash Shringarpure, Adam Auton, Aly Azeem Khan |
NeurIPS | 5 |
| 2024 | Enhancing Instance-Level Image Classification with Set-Level LabelsabstractInstance-level image classification tasks have traditionally relied on single-instance labels to train models, e.g., few-shot learning and transfer learning. However, set-level coarse-grained labels that capture relationships among instances can provide richer information in real-world scenarios. In this paper, we present a novel approach to enhance instance-level image classification by leveraging set-level labels. We provide a theoretical analysis of the proposed method, including recognition conditions for fast excess risk rate, shedding light on the theoretical foundations of our approach. We conducted experiments on two distinct categories of datasets: natural image datasets and histopathology image datasets. Our experimental results demonstrate the effectiveness of our approach, showcasing improved classification performance compared to traditional single-instance label-based methods. Notably, our algorithm achieves 13\% improvement in classification accuracy compared to the strongest baseline on the histopathology image classification benchmarks. Importantly, our experimental findings align with the theoretical analysis, reinforcing the robustness and reliability of our proposed method. This work bridges the gap between instance-level and set-level image classification, offering a promising avenue for advancing the capabilities of image classification models with set-level coarse-grained labels. Aly Azeem Khan, Yuxin Chen 0001, Robert L. Grossman |
ICLR | 2 |
| 2024 | TRIBAL: Tree Inference of B Cell Clonal Lineages
Leah L. Weber, Derek Reiman, Mrinmoy Saha Roddur, Mohammed El-Kebir, Aly Azeem Khan |
RECOMB | 6 |
| 2024 | Computational prediction of protein interactions in single cells by proximity sequencingabstractProximity sequencing (Prox-seq) simultaneously measures gene expression, protein expression and protein complexes on single cells. Using information from dual-antibody binding events, Prox-seq infers surface protein dimers at the single-cell level. Prox-seq provides multi-dimensional phenotyping of single cells in high throughput, and was recently used to track the formation of receptor complexes during cell signaling and discovered a novel interaction between CD9 and CD8 in naïve T cells. The distribution of protein abundance can affect identification of protein complexes in a complicated manner in dual-binding assays like Prox-seq. These effects are difficult to explore with experiments, yet important for accurate quantification of protein complexes. Here, we introduce a physical model of Prox-seq and computationally evaluate several different methods for reducing background noise when quantifying protein complexes. Furthermore, we developed an improved method for analysis of Prox-seq data, which resulted in more accurate and robust quantification of protein complexes. Finally, our Prox-seq model offers a simple way to investigate the behavior of Prox-seq data under various biological conditions and guide users toward selecting the best analysis method for their data. Junjie Xia, Hoang Van Phan, Luke Vistain, Mengjie Chen, Aly Azeem Khan, Savas Tay |
PLoS Comput. Biol. | 5 |
| 2023 | Scalable Batch-Mode Deep Bayesian Active Learning via Equivalence Class Annealing
Aly Azeem Khan, Robert L. Grossman, Yuxin Chen 0001 |
ICLR | 2 |
| 2022 | Privacy preserving validation for multiomic prediction modelsabstractReproducibility of results obtained using ribonucleic acid (RNA) data across labs remains a major hurdle in cancer research. Often, molecular predictors trained on one dataset cannot be applied to another due to differences in RNA library preparation and quantification, which inhibits the validation of predictors across labs. While current RNA correction algorithms reduce these differences, they require simultaneous access to patient-level data from all datasets, which necessitates the sharing of training data for predictors when sharing predictors. Here, we describe SpinAdapt, an unsupervised RNA correction algorithm that enables the transfer of molecular models without requiring access to patient-level data. It computes data corrections only via aggregate statistics of each dataset, thereby maintaining patient data privacy. Despite an inherent trade-off between privacy and performance, SpinAdapt outperforms current correction methods, like Seurat and ComBat, on publicly available cancer studies, including TCGA and ICGC. Furthermore, SpinAdapt can correct new samples, thereby enabling unbiased evaluation on validation cohorts. We expect this novel correction paradigm to enhance research reproducibility and to preserve patient privacy. Talal Ahmed, Mark A. Carty, Stephane Wenric, Jonathan R. Dry, Ameen A. Salahudeen, Aly Azeem Khan, Eric Lefkofsky, Martin C. Stumpe, Raphael Pelossof |
Briefings Bioinform. | 6 |
| 2021 | Pseudocell Tracer - A method for inferring dynamic trajectories using scRNAseq and its application to B cells undergoing immunoglobulin class switch recombinationabstractSingle cell RNA sequencing (scRNAseq) can be used to infer a temporal ordering of cellular states. Current methods for the inference of cellular trajectories rely on unbiased dimensionality reduction techniques. However, such biologically agnostic ordering can prove difficult for modeling complex developmental or differentiation processes. The cellular heterogeneity of dynamic biological compartments can result in sparse sampling of key intermediate cell states. To overcome these limitations, we develop a supervised machine learning framework, called Pseudocell Tracer, which infers trajectories in pseudospace rather than in pseudotime. The method uses a supervised encoder, trained with adjacent biological information, to project scRNAseq data into a low-dimensional manifold that maps the transcriptional states a cell can occupy. Then a generative adversarial network (GAN) is used to simulate pesudocells at regular intervals along a virtual cell-state axis. We demonstrate the utility of Pseudocell Tracer by modeling B cells undergoing immunoglobulin class switch recombination (CSR) during a prototypic antigen-induced antibody response. Our results revealed an ordering of key transcription factors regulating CSR to the IgG1 isotype, including the concomitant expression of Nfkb1 and Stat6 prior to the upregulation of Bach2 expression. Furthermore, the expression dynamics of genes encoding cytokine receptors suggest a poised IL-4 signaling state that preceeds CSR to the IgG1 isotype. Derek Reiman, Godhev Kumar Manakkat Vijay, Heping Xu, Andrew Sonin, Dianyu Chen, Nathan Salomonis, Harinder Singh, Aly Azeem Khan |
PLoS Comput. Biol. | 8 |
| 2020 | Evaluation of Hyperbolic Attention in Histopathology ImagesabstractWe bring together into a common framework three key ideas - multi-scale medical image analysis, the attention mechanism, and hyperbolic embeddings. The formulation and evaluation of hyperbolic-attention models for multi-scale medical image analysis have not been previously explored. In this paper, we evaluate a hyperbolic-attention model on two classification tasks using histopathology image datasets. The experiments show improvement compared to other commonly used models. Our method directly captures the multi-scale structure of histopathology images, and we speculate that the hyperbolic attention mechanism naturally singles out one or more structures at one or more scales that are most discriminatory. Aly Azeem Khan, Robert L. Grossman |
BIBE | 2 |
| 2018 | Predicting protein-protein interactions through sequence-based deep learningabstractMotivation: High-throughput experimental techniques have produced a large amount of protein-protein interaction (PPI) data, but their coverage is still low and the PPI data is also very noisy. Computational prediction of PPIs can be used to discover new PPIs and identify errors in the experimental PPI data. Results: We present a novel deep learning framework, DPPI, to model and predict PPIs from sequence information alone. Our model efficiently applies a deep, Siamese-like convolutional neural network combined with random projection and data augmentation to predict PPIs, leveraging existing high-quality experimental PPI data and evolutionary information of a protein pair under prediction. Our experimental results show that DPPI outperforms the state-of-the-art methods on several benchmarks in terms of area under precision-recall curve (auPR), and computationally is more efficient. We also show that DPPI is able to predict homodimeric interactions where other methods fail to work accurately, and the effectiveness of DPPI in specific applications such as predicting cytokine-receptor binding affinities. Availability and implementation: Predicting protein-protein interactions through sequence-based deep learning): https://github.com/hashemifar/DPPI/. Supplementary information: Supplementary data are available at Bioinformatics online. Somaye Hashemifar, Behnam Neyshabur, Aly Azeem Khan, Jinbo Xu |
Bioinform. | 3 |
| 2017 | BASIC: BCR assembly from single cellsabstractMotivation: The B-cell receptor enables individual B cells to identify diverse antigens, including bacterial and viral proteins. While advances in RNA-sequencing (RNA-seq) have enabled high throughput profiling of transcript expression in single cells, the unique task of assembling the full-length heavy and light chain sequences from single cell RNA-seq (scRNA-seq) in B cells has been largely unstudied. Results: We developed a new software tool, BASIC, which allows investigators to use scRNA-seq for assembling BCR sequences at single-cell resolution. To demonstrate the utility of our software, we subjected nearly 200 single human B cells to scRNA-seq, assembled the full-length heavy and the light chains, and experimentally confirmed these results by using single-cell primer-based nested PCRs and Sanger sequencing. Availability and Implementation: http://ttic.uchicago.edu/∼aakhan/BASIC Contact: [email protected] Supplementary Information: Supplementary data are available at Bioinformatics online. Stefan Canzar, Karlynn E. Neu, Qingming Tang, Patrick C. Wilson, Aly Azeem Khan |
Bioinform. | 5 |