Surag Nair

dblp:180/3420 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0002-6216-2457ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Generative modeling · 68% Trustworthy machine learning · 30% Information extraction and text analysis · 3%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 100%
Databases, data mining, and information retrieval
1 paper
Knowledge graphs · 100%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › gene regulation
regulatory genomics
1.022022
fastISM: performantin silicosaturation mutagenesis for convolutional neural networks · Bioinform. 2022
Integrating regulatory DNA sequence and gene expression to predict genome-wide chromatin accessibility across cellular contexts · Bioinform. 2019
Machine learning › Generative modeling
diffusion model
0.912025
Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-based Decoding · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
discrete diffusion model
0.912025
Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-based Decoding · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
guided sampling
0.912025
Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-based Decoding · NeurIPS 2025
Machine learning › Trustworthy machine learning › interpretability › attribution methods
feature attribution
0.612022
fastISM: performantin silicosaturation mutagenesis for convolutional neural networks · Bioinform. 2022
Machine learning › Trustworthy machine learning › interpretability
model explanation
0.612022
fastISM: performantin silicosaturation mutagenesis for convolutional neural networks · Bioinform. 2022
Bioinformatics and computational biology › genomics
computational genomics
0.612022
Accelerating in silico saturation mutagenesis using compressed sensing · Bioinform. 2022
Bioinformatics and computational biology › epigenomics › chromatin accessibility
chromatin accessibility prediction
0.412019
Integrating regulatory DNA sequence and gene expression to predict genome-wide chromatin accessibility across cellular contexts · Bioinform. 2019
Knowledge graphs › knowledge graph reasoning
temporal inference
0.312018
Inferring Temporal Knowledge for Near-Periodic Recurrent Events · IJCAI 2018
Bioinformatics and computational biology › sequence analysis › sequence modeling
sequence model interpretation
0.212022
Accelerating in silico saturation mutagenesis using compressed sensing · Bioinform. 2022
Natural language and speech › Information extraction and text analysis
temporal information extraction
0.112018
Inferring Temporal Knowledge for Near-Periodic Recurrent Events · IJCAI 2018

Methods — techniques the papers use, named apart from their topics

convolutional neural network · 1.5value function · 0.9soft value-based decoding · 0.9classifier guidance · 0.9schedule extraction · 0.7joint inference · 0.7compressed sensing · 0.6residual neural network · 0.4
YearPublicationVenuePosition
2025 Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-based Decoding
abstract
Diffusion models excel at capturing the natural design spaces of images, molecules, DNA, RNA, and protein sequences. However, rather than merely generating designs that are natural, we often aim to optimize downstream reward functions while preserving the naturalness of these design spaces. Existing methods for achieving this goal often require differentiable proxy models (e.g., classifier guidance or DPS) or involve computationally expensive fine-tuning of diffusion models (e.g., classifier-free guidance, RL-based fine-tuning). In our work, we propose a new method to address these challenges. Our algorithm is an iterative sampling method that integrates soft value functions, which looks ahead to how intermediate noisy states lead to high rewards in the future, into the standard inference procedure of pre-trained diffusion models. Notably, our approach avoids fine-tuning generative models and eliminates the need to construct differentiable models. This enables us to (1) directly utilize non-differentiable features/reward feedback, commonly used in many scientific domains, and (2) apply our method to recent discrete diffusion models in a principled way. Finally, we demonstrate the effectiveness of our algorithm across several domains, including image generation, molecule generation, and DNA/RNA sequence generation.
Xiner Li, Yulai Zhao 0002, Chenyu Wang 0003, Gabriele Scalia, Gökcen Eraslan, Surag Nair, Tommaso Biancalani, Shuiwang Ji, Aviv Regev, Sergey Levine, Masatoshi Uehara
NeurIPS6
2022 fastISM: performantin silicosaturation mutagenesis for convolutional neural networks
abstract
MOTIVATION: Deep-learning models, such as convolutional neural networks, are able to accurately map biological sequences to associated functional readouts and properties by learning predictive de novo representations. In silico saturation mutagenesis (ISM) is a popular feature attribution technique for inferring contributions of all characters in an input sequence to the model's predicted output. The main drawback of ISM is its runtime, as it involves multiple forward propagations of all possible mutations of each character in the input sequence through the trained model to predict the effects on the output. RESULTS: We present fastISM, an algorithm that speeds up ISM by a factor of over 10× for commonly used convolutional neural network architectures. fastISM is based on the observations that the majority of computation in ISM is spent in convolutional layers, and a single mutation only disrupts a limited region of intermediate layers, rendering most computation redundant. fastISM reduces the gap between backpropagation-based feature attribution methods and ISM. It far surpasses the runtime of backpropagation-based methods on multi-output architectures, making it feasible to run ISM on a large number of sequences. AVAILABILITY AND IMPLEMENTATION: An easy-to-use Keras/TensorFlow 2 implementation of fastISM is available at https://github.com/kundajelab/fastISM. fastISM can be installed using pip install fastism. A hands-on tutorial can be found at https://colab.research.google.com/github/kundajelab/fastISM/blob/master/notebooks/colab/DeepSEA.ipynb. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Surag Nair, Avanti Shrikumar, Jacob M. Schreiber, Anshul Kundaje
Bioinform.1
2022 Accelerating in silico saturation mutagenesis using compressed sensing
abstract
MOTIVATION: In silico saturation mutagenesis (ISM) is a popular approach in computational genomics for calculating feature attributions on biological sequences that proceeds by systematically perturbing each position in a sequence and recording the difference in model output. However, this method can be slow because systematically perturbing each position requires performing a number of forward passes proportional to the length of the sequence being examined. RESULTS: In this work, we propose a modification of ISM that leverages the principles of compressed sensing to require only a constant number of forward passes, regardless of sequence length, when applied to models that contain operations with a limited receptive field, such as convolutions. Our method, named Yuzu, can reduce the time that ISM spends in convolution operations by several orders of magnitude and, consequently, Yuzu can speed up ISM on several commonly used architectures in genomics by over an order of magnitude. Notably, we found that Yuzu provides speedups that increase with the complexity of the convolution operation and the length of the sequence being analyzed, suggesting that Yuzu provides large benefits in realistic settings. AVAILABILITY AND IMPLEMENTATION: We have made this tool available at https://github.com/kundajelab/yuzu.
Jacob M. Schreiber, Surag Nair, Akshay Balsubramani, Anshul Kundaje
Bioinform.2
2019 Integrating regulatory DNA sequence and gene expression to predict genome-wide chromatin accessibility across cellular contexts
abstract
MOTIVATION: Genome-wide profiles of chromatin accessibility and gene expression in diverse cellular contexts are critical to decipher the dynamics of transcriptional regulation. Recently, convolutional neural networks have been used to learn predictive cis-regulatory DNA sequence models of context-specific chromatin accessibility landscapes. However, these context-specific regulatory sequence models cannot generalize predictions across cell types. RESULTS: We introduce multi-modal, residual neural network architectures that integrate cis-regulatory sequence and context-specific expression of trans-regulators to predict genome-wide chromatin accessibility profiles across cellular contexts. We show that the average accessibility of a genomic region across training contexts can be a surprisingly powerful predictor. We leverage this feature and employ novel strategies for training models to enhance genome-wide prediction of shared and context-specific chromatin accessible sites across cell types. We interpret the models to reveal insights into cis- and trans-regulation of chromatin dynamics across 123 diverse cellular contexts. AVAILABILITY AND IMPLEMENTATION: The code is available at https://github.com/kundajelab/ChromDragoNN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Surag Nair, Daniel S. Kim, Jacob Perricone, Anshul Kundaje
Bioinform.1
2018 Inferring Temporal Knowledge for Near-Periodic Recurrent Events
abstract
We define the novel problem of extracting and predicting occurrence dates for a class of recurrent events -- events that are held periodically as per a near-regular schedule (e.g., conferences, film festivals, sport championships). Knowledge-bases such as Freebase contain a large number of such recurring events, but they also miss substantial information regarding specific event instances and their occurrence dates. We develop a temporal extraction and inference engine to fill in the missing dates as well as to predict their future occurrences. Our engine performs joint inference over several knowledge sources -- (1) information about an event instance and its date extracted from text by our temporal extractor, (2) information about the typical schedule (e.g., ``every second week of June") for a recurrent event extracted by our schedule extractor, and (3) known dates for other instances of the same event. The output of our system is a representation for the event schedule and an occurrence date for each event instance. We find that our system beats humans in predicting future occurrences of recurrent events by significant margins. We release our code and system output for further research.
Dinesh Raghu, Surag Nair, Mausam
IJCAI2