Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Anshul Kundaje

dblp:95/1107 · DBLP profile ↗
← Back
20ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0003-3084-2287ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 7 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
13 papers
Bioinformatics and computational biology · 85% Computational science and engineering · 15%
Artificial intelligence
8 papers
Trustworthy machine learning · 66% Representation and self-supervised learning · 12% Learning theory · 8%
Software engineering, system software, and programming languages
1 paper
Empirical software engineering · 100%
Databases, data mining, and information retrieval
2 papers
Data mining · 77% Information retrieval · 23%

Topics — the 30 heaviest of 44, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › gene regulation
regulatory genomics
1.742022
fastISM: performantin silicosaturation mutagenesis for convolutional neural networks · Bioinform. 2022
GkmExplain: fast and accurate interpretation of nonlinear gapped k-mer SVMs · Bioinform. 2019
Integrating regulatory DNA sequence and gene expression to predict genome-wide chromatin accessibility across cellular contexts · Bioinform. 2019
Bioinformatics and computational biology
genomics
1.222024
DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · NeurIPS 2024
Fourier-transform-based attribution priors improve the interpretability and stability of deep learning models for genomics · NeurIPS 2020
Bioinformatics and computational biology
immunoinformatics
0.912025
Sequence-Based TCR-Peptide Representations Using Cross-Epitope Contrastive Fine-Tuning of Protein Language Models · RECOMB 2025
Bioinformatics and computational biology › protein analysis › protein bioinformatics
protein representation learning
0.912025
Sequence-Based TCR-Peptide Representations Using Cross-Epitope Contrastive Fine-Tuning of Protein Language Models · RECOMB 2025
Bioinformatics and computational biology › immunoinformatics
TCR-epitope binding prediction
0.912025
Sequence-Based TCR-Peptide Representations Using Cross-Epitope Contrastive Fine-Tuning of Protein Language Models · RECOMB 2025
Machine learning › Trustworthy machine learning › interpretability › attribution methods
feature attribution
0.922022
fastISM: performantin silicosaturation mutagenesis for convolutional neural networks · Bioinform. 2022
Learning Important Features Through Propagating Activation Differences · ICML 2017
Machine learning › Trustworthy machine learning
interpretability
0.722020
Fourier-transform-based attribution priors improve the interpretability and stability of deep learning models for genomics · NeurIPS 2020
Learning Important Features Through Propagating Activation Differences · ICML 2017
Computational science and engineering
computational chemistry
0.712023
Tartarus: A Benchmarking Platform for Realistic And Practical Inverse Molecular Design · NeurIPS 2023
Computational science and engineering › materials informatics
inverse molecular design
0.712023
Tartarus: A Benchmarking Platform for Realistic And Practical Inverse Molecular Design · NeurIPS 2023
Bioinformatics and computational biology › molecular informatics
molecular design
0.712023
Tartarus: A Benchmarking Platform for Realistic And Practical Inverse Molecular Design · NeurIPS 2023
Bioinformatics and computational biology › drug discovery
molecular optimization
0.712023
Tartarus: A Benchmarking Platform for Realistic And Practical Inverse Molecular Design · NeurIPS 2023
Bioinformatics and computational biology › genomics
computational genomics
0.622022
Accelerating in silico saturation mutagenesis using compressed sensing · Bioinform. 2022
Motif Discovery Through Predictive Modeling of Gene Regulation · RECOMB 2005
Machine learning › Trustworthy machine learning › interpretability
model explanation
0.612022
fastISM: performantin silicosaturation mutagenesis for convolutional neural networks · Bioinform. 2022
Bioinformatics and computational biology › sequence analysis › sequence modeling
sequence model interpretation
0.622022
GkmExplain: fast and accurate interpretation of nonlinear gapped k-mer SVMs · Bioinform. 2019
Accelerating in silico saturation mutagenesis using compressed sensing · Bioinform. 2022
Machine learning › Learning theory
generalization
0.512021
WILDS: A Benchmark of in-the-Wild Distribution Shifts · ICML 2021
Machine learning › Trustworthy machine learning
out-of-distribution generalization
0.512021
WILDS: A Benchmark of in-the-Wild Distribution Shifts · ICML 2021
Machine learning › Trustworthy machine learning
robustness
0.512021
WILDS: A Benchmark of in-the-Wild Distribution Shifts · ICML 2021
Machine learning › Trustworthy machine learning › interpretability
attribution methods
0.412020
Fourier-transform-based attribution priors improve the interpretability and stability of deep learning models for genomics · NeurIPS 2020
Machine learning › Trustworthy machine learning
calibration
0.412020
Maximum Likelihood with Bias-Corrected Calibration is Hard-To-Beat at Label Shift Adaptation · ICML 2020
Machine learning › Transfer learning and domain adaptation › label shift
label shift adaptation
0.412020
Maximum Likelihood with Bias-Corrected Calibration is Hard-To-Beat at Label Shift Adaptation · ICML 2020
Bioinformatics and computational biology › gene regulation
regulatory sequence analysis
0.412020
Fourier-transform-based attribution priors improve the interpretability and stability of deep learning models for genomics · NeurIPS 2020
Bioinformatics and computational biology › epigenomics › chromatin accessibility
chromatin accessibility prediction
0.412019
Integrating regulatory DNA sequence and gene expression to predict genome-wide chromatin accessibility across cellular contexts · Bioinform. 2019
Bioinformatics and computational biology › genomics
chromosomal conformation capture data
0.312018
GenomeDISCO: a concordance score for chromosome conformation capture experiments using random walks on contact map graphs · Bioinform. 2018
Bioinformatics and computational biology › genomics › genome-wide association study
epistasis detection
0.312018
Discovering epistatic feature interactions from neural network models of regulatory DNA sequences · Bioinform. 2018
Computational science and engineering › model interpretability
interpretable deep learning
0.312018
Discovering epistatic feature interactions from neural network models of regulatory DNA sequences · Bioinform. 2018
Computational science and engineering › computational reproducibility
reproducibility assessment
0.312018
GenomeDISCO: a concordance score for chromosome conformation capture experiments using random walks on contact map graphs · Bioinform. 2018
Bioinformatics and computational biology
denoising
0.312017
Denoising genome-wide histone ChIP-seq with convolutional neural networks · Bioinform. 2017
Bioinformatics and computational biology
epigenomics
0.312017
Denoising genome-wide histone ChIP-seq with convolutional neural networks · Bioinform. 2017
Bioinformatics and computational biology › genomics
genomic data analysis
0.212016
Unsupervised Learning from Noisy Networks with Applications to Hi-C Data · NIPS 2016
Bioinformatics and computational biology › epigenomics
hi-c data analysis
0.212016
Unsupervised Learning from Noisy Networks with Applications to Hi-C Data · NIPS 2016

Methods — techniques the papers use, named apart from their topics

self-supervised learning · 2.3probing · 2.3fine-tuning · 2.3convolutional neural network · 1.8reinforcement learning · 1.3physical simulation · 1.3generative model · 1.3benchmark construction · 1.0protein language model · 0.9contrastive learning · 0.9compressed sensing · 0.6maximum likelihood estimation · 0.4importance weighting · 0.4partial labels · 0.2optimization framework · 0.2multi-resolution networks · 0.2
YearPublicationVenuePosition
2025 Sequence-Based TCR-Peptide Representations Using Cross-Epitope Contrastive Fine-Tuning of Protein Language Models
Chiho Im, Ryan Zhao, Scott D. Boyd, Anshul Kundaje
RECOMB4
2024 DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA
abstract
Recent advances in self-supervised models for natural language, vision, and protein sequences have inspired the development of large genomic DNA language models (DNALMs). These models aim to learn generalizable representations of diverse DNA elements, potentially enabling various genomic prediction, interpretation and design tasks. Despite their potential, existing benchmarks do not adequately assess the capabilities of DNALMs on key downstream applications involving an important class of non-coding DNA elements critical for regulating gene activity. In this study, we introduce DART-Eval, a suite of representative benchmarks specifically focused on regulatory DNA to evaluate model performance across zero-shot, probed, and fine-tuned scenarios against contemporary ab initio models as baselines. Our benchmarks target biologically meaningful downstream tasks such as functional sequence feature discovery, predicting cell-type specific regulatory activity, and counterfactual prediction of the impacts of genetic variants. We find that current DNALMs exhibit inconsistent performance and do not offer compelling gains over alternative baseline models for most tasks, while requiring significantly more computational resources. We discuss potentially promising modeling, data curation, and evaluation strategies for the next generation of DNALMs. Our code is available at https://github.com/kundajelab/DART-Eval
Aman Patel, Arpita Singhal, Austin Wang, Anusri Pampari, Maya Kasowski, Anshul Kundaje
NeurIPS6
2023 Tartarus: A Benchmarking Platform for Realistic And Practical Inverse Molecular Design
abstract
The efficient exploration of chemical space to design molecules with intended properties enables the accelerated discovery of drugs, materials, and catalysts, and is one of the most important outstanding challenges in chemistry. Encouraged by the recent surge in computer power and artificial intelligence development, many algorithms have been developed to tackle this problem. However, despite the emergence of many new approaches in recent years, comparatively little progress has been made in developing realistic benchmarks that reflect the complexity of molecular design for real-world applications. In this work, we develop a set of practical benchmark tasks relying on physical simulation of molecular systems mimicking real-life molecular design problems for materials, drugs, and chemical reactions. Additionally, we demonstrate the utility and ease of use of our new benchmark set by demonstrating how to compare the performance of several well-established families of algorithms. Overall, we believe that our benchmark suite will help move the field towards more realistic molecular design benchmarks, and move the development of inverse molecular design algorithms closer to the practice of designing molecules that solve existing problems in both academia and industry alike.
AkshatKumar Nigam, Robert Pollice, Gary Tom, Kjell Jorner, John Willes, Luca A. Thiede, Anshul Kundaje, Alán Aspuru-Guzik
NeurIPS7
2022 fastISM: performantin silicosaturation mutagenesis for convolutional neural networks
abstract
MOTIVATION: Deep-learning models, such as convolutional neural networks, are able to accurately map biological sequences to associated functional readouts and properties by learning predictive de novo representations. In silico saturation mutagenesis (ISM) is a popular feature attribution technique for inferring contributions of all characters in an input sequence to the model's predicted output. The main drawback of ISM is its runtime, as it involves multiple forward propagations of all possible mutations of each character in the input sequence through the trained model to predict the effects on the output. RESULTS: We present fastISM, an algorithm that speeds up ISM by a factor of over 10× for commonly used convolutional neural network architectures. fastISM is based on the observations that the majority of computation in ISM is spent in convolutional layers, and a single mutation only disrupts a limited region of intermediate layers, rendering most computation redundant. fastISM reduces the gap between backpropagation-based feature attribution methods and ISM. It far surpasses the runtime of backpropagation-based methods on multi-output architectures, making it feasible to run ISM on a large number of sequences. AVAILABILITY AND IMPLEMENTATION: An easy-to-use Keras/TensorFlow 2 implementation of fastISM is available at https://github.com/kundajelab/fastISM. fastISM can be installed using pip install fastism. A hands-on tutorial can be found at https://colab.research.google.com/github/kundajelab/fastISM/blob/master/notebooks/colab/DeepSEA.ipynb. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Surag Nair, Avanti Shrikumar, Jacob M. Schreiber, Anshul Kundaje
Bioinform.4
2022 Accelerating in silico saturation mutagenesis using compressed sensing
abstract
MOTIVATION: In silico saturation mutagenesis (ISM) is a popular approach in computational genomics for calculating feature attributions on biological sequences that proceeds by systematically perturbing each position in a sequence and recording the difference in model output. However, this method can be slow because systematically perturbing each position requires performing a number of forward passes proportional to the length of the sequence being examined. RESULTS: In this work, we propose a modification of ISM that leverages the principles of compressed sensing to require only a constant number of forward passes, regardless of sequence length, when applied to models that contain operations with a limited receptive field, such as convolutions. Our method, named Yuzu, can reduce the time that ISM spends in convolution operations by several orders of magnitude and, consequently, Yuzu can speed up ISM on several commonly used architectures in genomics by over an order of magnitude. Notably, we found that Yuzu provides speedups that increase with the complexity of the convolution operation and the length of the sequence being analyzed, suggesting that Yuzu provides large benefits in realistic settings. AVAILABILITY AND IMPLEMENTATION: We have made this tool available at https://github.com/kundajelab/yuzu.
Jacob M. Schreiber, Surag Nair, Akshay Balsubramani, Anshul Kundaje
Bioinform.4
2021 WILDS: A Benchmark of in-the-Wild Distribution Shifts
abstract
Distribution shifts—where the training distribution differs from the test distribution—can substantially degrade the accuracy of machine learning (ML) systems deployed in the wild. Despite their ubiquity in the real-world deployments, these distribution shifts are under-represented in the datasets widely used in the ML community today. To address this gap, we present WILDS, a curated benchmark of 10 datasets reflecting a diverse range of distribution shifts that naturally arise in real-world applications, such as shifts across hospitals for tumor identification; across camera traps for wildlife monitoring; and across time and location in satellite imaging and poverty mapping. On each dataset, we show that standard training yields substantially lower out-of-distribution than in-distribution performance. This gap remains even with models trained by existing methods for tackling distribution shifts, underscoring the need for new methods for training models that are more robust to the types of distribution shifts that arise in practice. To facilitate method development, we provide an open-source package that automates dataset loading, contains default model architectures and hyperparameters, and standardizes evaluations. The full paper, code, and leaderboards are available at https://wilds.stanford.edu.
Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard L. Phillips, Irena Gao, Etienne David, Ian Stavness, Wei Guo 0002, Berton Earnshaw, Imran S. Haque, Sara Beery, Jure Leskovec, Anshul Kundaje, Emma Pierson, Sergey Levine, Chelsea Finn, Percy Liang
ICML19
2020 Maximum Likelihood with Bias-Corrected Calibration is Hard-To-Beat at Label Shift Adaptation
abstract
Label shift refers to the phenomenon where the prior class probability p(y) changes between the training and test distributions, while the conditional probability p(x|y) stays fixed. Label shift arises in settings like medical diagnosis, where a classifier trained to predict disease given symptoms must be adapted to scenarios where the baseline prevalence of the disease is different. Given estimates of p(y|x) from a predictive model, Saerens et al. proposed an efficient maximum likelihood algorithm to correct for label shift that does not require model retraining, but a limiting assumption of this algorithm is that p(y|x) is calibrated, which is not true of modern neural networks. Recently, Black Box Shift Learning (BBSL) and Regularized Learning under Label Shifts (RLLS) have emerged as state-of-the-art techniques to cope with label shift when a classifier does not output calibrated probabilities, but both methods require model retraining with importance weights and neither has been benchmarked against maximum likelihood. Here we (1) show that combining maximum likelihood with a type of calibration we call bias-corrected calibration outperforms both BBSL and RLLS across diverse datasets and distribution shifts, (2) prove that the maximum likelihood objective is concave, and (3) introduce a principled strategy for estimating source-domain priors that improves robustness to poor calibration. This work demonstrates that the maximum likelihood with appropriate calibration is a formidable and efficient baseline for label shift adaptation; notebooks reproducing experiments available at https://github.com/kundajelab/labelshiftexperiments , video: https://youtu.be/ZBXjE9QTruE , blogpost: https://bit.ly/3kTds7J
Amr Alexandari, Anshul Kundaje, Avanti Shrikumar
ICML2
2020 Fourier-transform-based attribution priors improve the interpretability and stability of deep learning models for genomics
abstract
Deep learning models can accurately map genomic DNA sequences to associated functional molecular readouts such as protein-DNA binding data. Base-resolution importance (i.e. "attribution") scores inferred from these models can highlight predictive sequence motifs and syntax. Unfortunately, these models are prone to overfitting and are sensitive to random initializations, often resulting in noisy and irreproducible attributions that obfuscate underlying motifs. To address these shortcomings, we propose a novel attribution prior, where the Fourier transform of input-level attribution scores are computed at training-time, and high-frequency components of the Fourier spectrum are penalized. We evaluate different model architectures with and without our attribution prior, training on genome-wide binary labels or continuous molecular profiles. We show that our attribution prior significantly improves models' stability, interpretability, and performance on held-out data, especially when training data is severely limited. Our attribution prior also allows models to identify biologically meaningful sequence motifs more sensitively and precisely within individual regulatory elements. The prior is agnostic to the model architecture or predicted experimental assay, yet provides similar gains across all experiments. This work represents an important advancement in improving the reliability of deep learning models for deciphering the regulatory code of the genome.
Alex Tseng, Avanti Shrikumar, Anshul Kundaje
NeurIPS3
2019 Integrating regulatory DNA sequence and gene expression to predict genome-wide chromatin accessibility across cellular contexts
abstract
MOTIVATION: Genome-wide profiles of chromatin accessibility and gene expression in diverse cellular contexts are critical to decipher the dynamics of transcriptional regulation. Recently, convolutional neural networks have been used to learn predictive cis-regulatory DNA sequence models of context-specific chromatin accessibility landscapes. However, these context-specific regulatory sequence models cannot generalize predictions across cell types. RESULTS: We introduce multi-modal, residual neural network architectures that integrate cis-regulatory sequence and context-specific expression of trans-regulators to predict genome-wide chromatin accessibility profiles across cellular contexts. We show that the average accessibility of a genomic region across training contexts can be a surprisingly powerful predictor. We leverage this feature and employ novel strategies for training models to enhance genome-wide prediction of shared and context-specific chromatin accessible sites across cell types. We interpret the models to reveal insights into cis- and trans-regulation of chromatin dynamics across 123 diverse cellular contexts. AVAILABILITY AND IMPLEMENTATION: The code is available at https://github.com/kundajelab/ChromDragoNN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Surag Nair, Daniel S. Kim, Jacob Perricone, Anshul Kundaje
Bioinform.4
2019 GkmExplain: fast and accurate interpretation of nonlinear gapped k-mer SVMs
abstract
SUMMARY: Support Vector Machines with gapped k-mer kernels (gkm-SVMs) have been used to learn predictive models of regulatory DNA sequence. However, interpreting predictive sequence patterns learned by gkm-SVMs can be challenging. Existing interpretation methods such as deltaSVM, in-silico mutagenesis (ISM) or SHAP either do not scale well or make limiting assumptions about the model that can produce misleading results when the gkm kernel is combined with nonlinear kernels. Here, we propose GkmExplain: a computationally efficient feature attribution method for interpreting predictive sequence patterns from gkm-SVM models that has theoretical connections to the method of Integrated Gradients. Using simulated regulatory DNA sequences, we show that GkmExplain identifies predictive patterns with high accuracy while avoiding pitfalls of deltaSVM and ISM and being orders of magnitude more computationally efficient than SHAP. By applying GkmExplain and a recently developed motif discovery method called TF-MoDISco to gkm-SVM models trained on in vivo transcription factor (TF) binding data, we recover consolidated, non-redundant TF motifs. Mutation impact scores derived using GkmExplain consistently outperform deltaSVM and ISM at identifying regulatory genetic variants from gkm-SVM models of chromatin accessibility in lymphoblastoid cell-lines. AVAILABILITY AND IMPLEMENTATION: Code and example notebooks to reproduce results are at https://github.com/kundajelab/gkmexplain. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Avanti Shrikumar, Eva Prakash, Anshul Kundaje
Bioinform.3
2018 Discovering epistatic feature interactions from neural network models of regulatory DNA sequences
abstract
Motivation: Transcription factors bind regulatory DNA sequences in a combinatorial manner to modulate gene expression. Deep neural networks (DNNs) can learn the cis-regulatory grammars encoded in regulatory DNA sequences associated with transcription factor binding and chromatin accessibility. Several feature attribution methods have been developed for estimating the predictive importance of individual features (nucleotides or motifs) in any input DNA sequence to its associated output prediction from a DNN model. However, these methods do not reveal higher-order feature interactions encoded by the models. Results: We present a new method called Deep Feature Interaction Maps (DFIM) to efficiently estimate interactions between all pairs of features in any input DNA sequence. DFIM accurately identifies ground truth motif interactions embedded in simulated regulatory DNA sequences. DFIM identifies synergistic interactions between GATA1 and TAL1 motifs from in vivo TF binding models. DFIM reveals epistatic interactions involving nucleotides flanking the core motif of the Cbf1 TF in yeast from in vitro TF binding models. We also apply DFIM to regulatory sequence models of in vivo chromatin accessibility to reveal interactions between regulatory genetic variants and proximal motifs of target TFs as validated by TF binding quantitative trait loci. Our approach makes significant strides in improving the interpretability of deep learning models for genomics. Availability and implementation: Code is available at: https://github.com/kundajelab/dfim. Supplementary information: Supplementary data are available at Bioinformatics online.
Peyton Greenside, Tyler Shimko, Polly Fordyce, Anshul Kundaje
Bioinform.4
2018 GenomeDISCO: a concordance score for chromosome conformation capture experiments using random walks on contact map graphs
abstract
Motivation: The three-dimensional organization of chromatin plays a critical role in gene regulation and disease. High-throughput chromosome conformation capture experiments such as Hi-C are used to obtain genome-wide maps of three-dimensional chromatin contacts. However, robust estimation of data quality and systematic comparison of these contact maps is challenging due to the multi-scale, hierarchical structure of chromatin contacts and the resulting properties of experimental noise in the data. Measuring concordance of contact maps is important for assessing reproducibility of replicate experiments and for modeling variation between different cellular contexts. Results: We introduce a concordance measure called DIfferences between Smoothed COntact maps (GenomeDISCO) for assessing the similarity of a pair of contact maps obtained from chromosome conformation capture experiments. The key idea is to smooth contact maps using random walks on the contact map graph, before estimating concordance. We use simulated datasets to benchmark GenomeDISCO's sensitivity to different types of noise that affect chromatin contact maps. When applied to a large collection of Hi-C datasets, GenomeDISCO accurately distinguishes biological replicates from samples obtained from different cell types. GenomeDISCO also generalizes to other chromosome conformation capture assays, such as HiChIP. Availability and implementation: Software implementing GenomeDISCO is available at https://github.com/kundajelab/genomedisco. Supplementary information: Supplementary data are available at Bioinformatics online.
Oana Ursu, Nathan Boley, Maryna Taranova, Y. X. Rachel Wang, Galip Gürkan Yardimci, William Stafford Noble, Anshul Kundaje
Bioinform.7
2017 Learning Important Features Through Propagating Activation Differences
abstract
The purported “black box” nature of neural networks is a barrier to adoption in applications where interpretability is essential. Here we present DeepLIFT (Deep Learning Important FeaTures), a method for decomposing the output prediction of a neural network on a specific input by backpropagating the contributions of all neurons in the network to every feature of the input. DeepLIFT compares the activation of each neuron to its `reference activation’ and assigns contribution scores according to the difference. By optionally giving separate consideration to positive and negative contributions, DeepLIFT can also reveal dependencies which are missed by other approaches. Scores can be computed efficiently in a single backward pass. We apply DeepLIFT to models trained on MNIST and simulated genomic data, and show significant advantages over gradient-based methods. Video tutorial: http://goo.gl/qKb7pL code: http://goo.gl/RM8jvH
Avanti Shrikumar, Peyton Greenside, Anshul Kundaje
ICML3
2017 Denoising genome-wide histone ChIP-seq with convolutional neural networks
abstract
MOTIVATION: Chromatin immune-precipitation sequencing (ChIP-seq) experiments are commonly used to obtain genome-wide profiles of histone modifications associated with different types of functional genomic elements. However, the quality of histone ChIP-seq data is affected by many experimental parameters such as the amount of input DNA, antibody specificity, ChIP enrichment and sequencing depth. Making accurate inferences from chromatin profiling experiments that involve diverse experimental parameters is challenging. RESULTS: We introduce a convolutional denoising algorithm, Coda, that uses convolutional neural networks to learn a mapping from suboptimal to high-quality histone ChIP-seq data. This overcomes various sources of noise and variability, substantially enhancing and recovering signal when applied to low-quality chromatin profiling datasets across individuals, cell types and species. Our method has the potential to improve data quality at reduced costs. More broadly, this approach-using a high-dimensional discriminative model to encode a generative noise process-is generally applicable to other biological domains where it is easy to generate noisy data but difficult to analytically characterize the noise or underlying data distribution. AVAILABILITY AND IMPLEMENTATION: https://github.com/kundajelab/coda . CONTACT: [email protected].
Pang Wei Koh, Emma Pierson, Anshul Kundaje
Bioinform.3
2017 Vicus: Exploiting local structures to improve network-based analysis of biological data
abstract
Biological networks entail important topological features and patterns critical to understanding interactions within complicated biological systems. Despite a great progress in understanding their structure, much more can be done to improve our inference and network analysis. Spectral methods play a key role in many network-based applications. Fundamental to spectral methods is the Laplacian, a matrix that captures the global structure of the network. Unfortunately, the Laplacian does not take into account intricacies of the network's local structure and is sensitive to noise in the network. These two properties are fundamental to biological networks and cannot be ignored. We propose an alternative matrix Vicus. The Vicus matrix captures the local neighborhood structure of the network and thus is more effective at modeling biological interactions. We demonstrate the advantages of Vicus in the context of spectral methods by extensive empirical benchmarking on tasks such as single cell dimensionality reduction, protein module discovery and ranking genes for cancer subtyping. Our experiments show that using Vicus, spectral methods result in more accurate and robust performance in all of these tasks.
Bo Wang 0044, Yuke Zhu, Anshul Kundaje, Serafim Batzoglou, Anna Goldenberg
PLoS Comput. Biol.4
2016 Unsupervised Learning from Noisy Networks with Applications to Hi-C Data
abstract
Complex networks play an important role in a plethora of disciplines in natural sciences. Cleaning up noisy observed networks, poses an important challenge in network analysis Existing methods utilize labeled data to alleviate the noise effect in the network. However, labeled data is usually expensive to collect while unlabeled data can be gathered cheaply. In this paper, we propose an optimization framework to mine useful structures from noisy networks in an unsupervised manner. The key feature of our optimization framework is its ability to utilize local structures as well as global patterns in the network. We extend our method to incorporate multi-resolution networks in order to add further resistance to high-levels of noise. We also generalize our framework to utilize partial labels to enhance the performance. We specifically focus our method on multi-resolution Hi-C data by recovering clusters of genomic regions that co-localize in 3D space. Additionally, we use Capture-C-generated partial labels to further denoise the Hi-C network. We empirically demonstrate the effectiveness of our framework in denoising the network and improving community detection results.
Bo Wang 0044, Armin Pourshafeie, Oana Ursu, Serafim Batzoglou, Anshul Kundaje
NIPS6
2008 A Predictive Model of the Oxygen and Heme Regulatory Network in Yeast
abstract
Deciphering gene regulatory mechanisms through the analysis of high-throughput expression data is a challenging computational problem. Previous computational studies have used large expression datasets in order to resolve fine patterns of coexpression, producing clusters or modules of potentially coregulated genes. These methods typically examine promoter sequence information, such as DNA motifs or transcription factor occupancy data, in a separate step after clustering. We needed an alternative and more integrative approach to study the oxygen regulatory network in Saccharomyces cerevisiae using a small dataset of perturbation experiments. Mechanisms of oxygen sensing and regulation underlie many physiological and pathological processes, and only a handful of oxygen regulators have been identified in previous studies. We used a new machine learning algorithm called MEDUSA to uncover detailed information about the oxygen regulatory network using genome-wide expression changes in response to perturbations in the levels of oxygen, heme, Hap1, and Co2+. MEDUSA integrates mRNA expression, promoter sequence, and ChIP-chip occupancy data to learn a model that accurately predicts the differential expression of target genes in held-out data. We used a novel margin-based score to extract significant condition-specific regulators and assemble a global map of the oxygen sensing and regulatory network. This network includes both known oxygen and heme regulators, such as Hap1, Mga2, Hap4, and Upc2, as well as many new candidate regulators. MEDUSA also identified many DNA motifs that are consistent with previous experimentally identified transcription factor binding sites. Because MEDUSA's regulatory program associates regulators to target genes through their promoter sequences, we directly tested the predicted regulators for OLE1, a gene specifically induced under hypoxia, by experimental analysis of the activity of its promoter. In each case, deletion of the candidate regulator resulted in the predicted effect on promoter activity, confirming that several novel regulators identified by MEDUSA are indeed involved in oxygen regulation. MEDUSA can reveal important information from a small dataset and generate testable hypotheses for further experimental analysis. Supplemental data are included.
Anshul Kundaje, Xiantong Xin, Changgui Lan, Steve Lianoglou, Mei Zhou, Li Zhang 0014, Christina S. Leslie
PLoS Comput. Biol.1
2006 A classification-based framework for predicting and analyzing gene regulatory response
abstract
BACKGROUND: We have recently introduced a predictive framework for studying gene transcriptional regulation in simpler organisms using a novel supervised learning algorithm called GeneClass. GeneClass is motivated by the hypothesis that in model organisms such as Saccharomyces cerevisiae, we can learn a decision rule for predicting whether a gene is up- or down-regulated in a particular microarray experiment based on the presence of binding site subsequences ("motifs") in the gene's regulatory region and the expression levels of regulators such as transcription factors in the experiment ("parents"). GeneClass formulates the learning task as a classification problem--predicting +1 and -1 labels corresponding to up- and down-regulation beyond the levels of biological and measurement noise in microarray measurements. Using the Adaboost algorithm, GeneClass learns a prediction function in the form of an alternating decision tree, a margin-based generalization of a decision tree. METHODS: In the current work, we introduce a new, robust version of the GeneClass algorithm that increases stability and computational efficiency, yielding a more scalable and reliable predictive model. The improved stability of the prediction tree enables us to introduce a detailed post-processing framework for biological interpretation, including individual and group target gene analysis to reveal condition-specific regulation programs and to suggest signaling pathways. Robust GeneClass uses a novel stabilized variant of boosting that allows a set of correlated features, rather than single features, to be included at nodes of the tree; in this way, biologically important features that are correlated with the single best feature are retained rather than decorrelated and lost in the next round of boosting. Other computational developments include fast matrix computation of the loss function for all features, allowing scalability to large datasets, and the use of abstaining weak rules, which results in a more shallow and interpretable tree. We also show how to incorporate genome-wide protein-DNA binding data from ChIP chip experiments into the GeneClass algorithm, and we use an improved noise model for gene expression data. RESULTS: Using the improved scalability of Robust GeneClass, we present larger scale experiments on a yeast environmental stress dataset, training and testing on all genes and using a comprehensive set of potential regulators. We demonstrate the improved stability of the features in the learned prediction tree, and we show the utility of the post-processing framework by analyzing two groups of genes in yeast--the protein chaperones and a set of putative targets of the Nrg1 and Nrg2 transcription factors--and suggesting novel hypotheses about their transcriptional and post-transcriptional regulation. Detailed results and Robust GeneClass source code is available for download from http://www.cs.columbia.edu/compbio/robust-geneclass.
Anshul Kundaje, Manuel Middendorf, Mihir Shah, Chris Wiggins 0001, Yoav Freund, Christina S. Leslie
BMC Bioinform.1
2005 Motif Discovery Through Predictive Modeling of Gene Regulation
Manuel Middendorf, Anshul Kundaje, Mihir Shah, Yoav Freund, Chris Wiggins 0001, Christina S. Leslie
RECOMB2
2005 Combining Sequence and Time Series Expression Data to Learn Transcriptional Modules
abstract
Our goal is to cluster genes into transcriptional modules--sets of genes where similarity in expression is explained by common regulatory mechanisms at the transcriptional level. We want to learn modules from both time series gene expression data and genome-wide motif data that are now readily available for organisms such as S. cereviseae as a result of prior computational studies or experimental results. We present a generative probabilistic model for combining regulatory sequence and time series expression data to cluster genes into coherent transcriptional modules. Starting with a set of motifs representing known or putative regulatory elements (transcription factor binding sites) and the counts of occurrences of these motifs in each gene's promoter region, together with a time series expression profile for each gene, the learning algorithm uses expectation maximization to learn module assignments based on both types of data. We also present a technique based on the Jensen-Shannon entropy contributions of motifs in the learned model for associating the most significant motifs to each module. Thus, the algorithm gives a global approach for associating sets of regulatory elements to "modules" of genes with similar time series expression profiles. The model for expression data exploits our prior belief of smooth dependence on time by using statistical splines and is suitable for typical time course data sets with relatively few experiments. Moreover, the model is sufficiently interpretable that we can understand how both sequence data and expression data contribute to the cluster assignments, and how to interpolate between the two data sources. We present experimental results on the yeast cell cycle to validate our method and find that our combined expression and motif clustering algorithm discovers modules with both coherent expression and similar motif patterns, including binding motifs associated to known cell cycle transcription factors.
Anshul Kundaje, Manuel Middendorf, Chris Wiggins 0001, Christina S. Leslie
IEEE ACM Trans. Comput. Biol. Bioinform.1