VLDB 2026 Research / reviewers in the wild / expert
Shibani Santurkar
dblp:153/2146
· DBLP profile ↗
20ranked-venue papers
9as first author
7since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 8 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
18 papers |
Trustworthy machine learning · 49% Language models and text generation · 8% Reinforcement learning · 8% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Processor architecture and microarchitecture · 87% GPUs and heterogeneous computing · 13% |
Topics — the 30 heaviest of 47, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
robustness |
2.4 | 5 | 2022 | 3DB: A Framework for Debugging Computer Vision Models · NeurIPS 2022 Editing a classifier by rewriting its prediction rules · NeurIPS 2021 BREEDS: Benchmarks for Subpopulation Shift · ICLR 2021 |
Machine learning › Trustworthy machine learning
interpretability |
1.9 | 4 | 2022 | 3DB: A Framework for Debugging Computer Vision Models · NeurIPS 2022 Leveraging Sparse Linear Layers for Debuggable Deep Networks · ICML 2021 From ImageNet to Image Classification: Contextualizing Progress on Benchmarks · ICML 2020 |
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
1.5 | 4 | 2019 | Image Synthesis with a Single (Robust) Classifier · NeurIPS 2019 Adversarial Examples Are Not Bugs, They Are Features · NeurIPS 2019 Robustness May Be at Odds with Accuracy · ICLR (Poster) 2019 |
Machine learning › Efficient and distributed learning
data selection |
0.7 | 1 | 2023 | Data Selection for Language Models via Importance Resampling · NeurIPS 2023 |
Machine learning › Transfer learning and domain adaptation
distribution matching |
0.7 | 1 | 2023 | Data Selection for Language Models via Importance Resampling · NeurIPS 2023 |
Machine learning › Representation and self-supervised learning
multimodal representation learning |
0.7 | 1 | 2023 | Is a Caption Worth a Thousand Images? A Study on Representation Learning · ICLR 2023 |
Natural language and speech › Language models and text generation › large language model training
pretraining data selection |
0.7 | 1 | 2023 | Data Selection for Language Models via Importance Resampling · NeurIPS 2023 |
Natural language and speech › Language models and text generation
knowledge editing |
0.5 | 1 | 2021 | Editing a classifier by rewriting its prediction rules · NeurIPS 2021 |
Machine learning › Trustworthy machine learning
spurious correlation detection |
0.5 | 1 | 2021 | Leveraging Sparse Linear Layers for Debuggable Deep Networks · ICML 2021 |
Machine learning › Trustworthy machine learning › robustness › spurious correlation
spurious correlation mitigation |
0.5 | 1 | 2021 | Editing a classifier by rewriting its prediction rules · NeurIPS 2021 |
Machine learning › Trustworthy machine learning › robustness › distribution shift
subpopulation shift |
0.5 | 1 | 2021 | BREEDS: Benchmarks for Subpopulation Shift · ICLR 2021 |
Machine learning › Trustworthy machine learning
dataset bias |
0.4 | 1 | 2020 | From ImageNet to Image Classification: Contextualizing Progress on Benchmarks · ICML 2020 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.4 | 1 | 2020 | A Closer Look at Deep Policy Gradients · ICLR 2020 |
Machine learning › Reinforcement learning › policy optimization
policy gradient |
0.4 | 1 | 2020 | A Closer Look at Deep Policy Gradients · ICLR 2020 |
Machine learning › Reinforcement learning
policy optimization |
0.4 | 1 | 2020 | Implementation Matters in Deep RL: A Case Study on PPO and TRPO · ICLR 2020 |
Machine learning › Reinforcement learning › policy optimization
proximal policy optimization |
0.4 | 1 | 2020 | Implementation Matters in Deep RL: A Case Study on PPO and TRPO · ICLR 2020 |
Machine learning › Trustworthy machine learning
robustness evaluation |
0.4 | 1 | 2020 | Identifying Statistical Bias in Dataset Replication · ICML 2020 |
Empirical software engineering
benchmarking |
0.4 | 1 | 2020 | Identifying Statistical Bias in Dataset Replication · ICML 2020 |
Machine learning › Trustworthy machine learning › robustness
adversarial examples |
0.4 | 1 | 2019 | Adversarial Examples Are Not Bugs, They Are Features · NeurIPS 2019 |
Machine learning › Trustworthy machine learning › interpretability
feature shaping |
0.4 | 1 | 2019 | Image Synthesis with a Single (Robust) Classifier · NeurIPS 2019 |
Machine learning › Generative modeling
image generation |
0.4 | 1 | 2019 | Image Synthesis with a Single (Robust) Classifier · NeurIPS 2019 |
Machine learning › Trustworthy machine learning › adversarial machine learning
non-robust features |
0.4 | 1 | 2019 | Adversarial Examples Are Not Bugs, They Are Features · NeurIPS 2019 |
Machine learning › Trustworthy machine learning › robustness › robust learning
robust classification |
0.4 | 1 | 2019 | Image Synthesis with a Single (Robust) Classifier · NeurIPS 2019 |
Machine learning › Trustworthy machine learning › robustness › adversarial robustness › adversarially robust generalization
robustness-accuracy trade-off |
0.4 | 1 | 2019 | Robustness May Be at Odds with Accuracy · ICLR (Poster) 2019 |
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarially robust generalization |
0.3 | 1 | 2018 | Adversarially Robust Generalization Requires More Data · NeurIPS 2018 |
Machine learning › Deep learning architectures and training › normalization
batch normalization |
0.3 | 1 | 2018 | How Does Batch Normalization Help Optimization? · NeurIPS 2018 |
Machine learning › Transfer learning and domain adaptation › domain shift
covariate shift |
0.3 | 1 | 2018 | A Classification-Based Study of Covariate Shift in GAN Distributions · ICML 2018 |
Machine learning › Generative modeling › generative model evaluation
diversity evaluation |
0.3 | 1 | 2018 | A Classification-Based Study of Covariate Shift in GAN Distributions · ICML 2018 |
Machine learning › Learning theory
generalization |
0.3 | 1 | 2018 | How Does Batch Normalization Help Optimization? · NeurIPS 2018 |
Machine learning › Generative modeling
generative adversarial network |
0.3 | 1 | 2018 | A Classification-Based Study of Covariate Shift in GAN Distributions · ICML 2018 |
Methods — techniques the papers use, named apart from their topics
gradient analysis · 0.8steering · 0.7public opinion polling · 0.7importance resampling · 0.7hashed n-gram features · 0.7photorealistic simulation · 0.6sparse linear model · 0.5rule rewriting · 0.5deep feature representation · 0.5data-free editing · 0.5statistical bias correction · 0.4selection frequency remeasurement · 0.4winograd convolution · 0.3cache-aware tiling · 0.3AVX vectorization · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Is a Caption Worth a Thousand Images? A Study on Representation Learning
Shibani Santurkar, Yann Dubois, Rohan Taori, Percy Liang, Tatsunori B. Hashimoto |
ICLR | 1 |
| 2023 | Whose Opinions Do Language Models Reflect?abstractLanguage models (LMs) are increasingly being used in open-ended contexts, where the opinions they reflect in response to subjective queries can have a profound impact, both on user satisfaction, and shaping the views of society at large. We put forth a quantitative framework to investigate the opinions reflected by LMs – by leveraging high-quality public opinion polls. Using this framework, we create OpinionQA, a dataset for evaluating the alignment of LM opinions with those of 60 US demographic groups over topics ranging from abortion to automation. Across topics, we find substantial misalignment between the views reflected by current LMs and those of US demographic groups: on par with the Democrat-Republican divide on climate change. Notably, this misalignment persists even after explicitly steering the LMs towards particular groups. Our analysis not only confirms prior observations about the left-leaning tendencies of some human feedback-tuned LMs, but also surfaces groups whose opinions are poorly reflected by current LMs (e.g., 65+ and widowed individuals). Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, Tatsunori B. Hashimoto |
ICML | 1 |
| 2023 | Data Selection for Language Models via Importance ResamplingabstractSelecting a suitable pretraining dataset is crucial for both general-domain (e.g., GPT-3) and domain-specific (e.g., Codex) language models (LMs). We formalize this problem as selecting a subset of a large raw unlabeled dataset to match a desired target distribution given unlabeled target samples. Due to the scale and dimensionality of the raw text data, existing methods use simple heuristics or require human experts to manually curate data. Instead, we extend the classic importance resampling approach used in low-dimensions for LM data selection. We propose Data Selection with Importance Resampling (DSIR), an efficient and scalable framework that estimates importance weights in a reduced feature space for tractability and selects data with importance resampling according to these weights. We instantiate the DSIR framework with hashed n-gram features for efficiency, enabling the selection of 100M documents from the full Pile dataset in 4.5 hours. To measure whether hashed n-gram features preserve the aspects of the data that are relevant to the target, we define KL reduction, a data metric that measures the proximity between the selected pretraining data and the target on some feature space. Across 8 data selection methods (including expert selection), KL reduction on hashed n-gram features highly correlates with average downstream accuracy (r=0.82). When selecting data for continued pretraining on a specific domain, DSIR performs comparably to expert curation across 8 target distributions. When pretraining general-domain models (target is Wikipedia and books), DSIR improves over random selection and heuristic filtering baselines by 2--2.5% on the GLUE benchmark. Sang Michael Xie, Shibani Santurkar, Tengyu Ma 0001, Percy Liang |
NeurIPS | 2 |
| 2022 | 3DB: A Framework for Debugging Computer Vision ModelsabstractWe introduce 3DB: an extendable, unified framework for testing and debugging vision models using photorealistic simulation. We demonstrate, through a wide range of use cases, that 3DB allows users to discover vulnerabilities in computer vision systems and gain insights into how models make decisions. 3DB captures and generalizes many robustness analyses from prior work, and enables one to study their interplay. Finally, we find that the insights generated by the system transfer to the physical world. 3DB will be released as a library alongside a set of examples and documentation. We attach 3DB to the submission. Guillaume Leclerc, Hadi Salman, Andrew Ilyas, Sai Vemprala, Logan Engstrom, Vibhav Vineet, Kai Yuanqing Xiao, Pengchuan Zhang, Shibani Santurkar, Greg Yang, Ashish Kapoor, Aleksander Madry |
NeurIPS | 9 |
| 2021 | BREEDS: Benchmarks for Subpopulation Shift
Shibani Santurkar, Dimitris Tsipras, Aleksander Madry |
ICLR | 1 |
| 2021 | Leveraging Sparse Linear Layers for Debuggable Deep NetworksabstractWe show how fitting sparse linear models over learned deep feature representations can lead to more debuggable neural networks. These networks remain highly accurate while also being more amenable to human interpretation, as we demonstrate quantitatively and via human experiments. We further illustrate how the resulting sparse explanations can help to identify spurious correlations, explain misclassifications, and diagnose model biases in vision and language tasks. Eric Wong 0001, Shibani Santurkar, Aleksander Madry |
ICML | 2 |
| 2021 | Editing a classifier by rewriting its prediction rulesabstractWe propose a methodology for modifying the behavior of a classifier by directly rewriting its prediction rules. Our method requires virtually no additional data collection and can be applied to a variety of settings, including adapting a model to new environments, and modifying it to ignore spurious features. Shibani Santurkar, Dimitris Tsipras, Mahalaxmi Elango, David Bau, Antonio Torralba 0001, Aleksander Madry |
NeurIPS | 1 |
| 2020 | Implementation Matters in Deep RL: A Case Study on PPO and TRPO
Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Firdaus Janoos, Larry Rudolph, Aleksander Madry |
ICLR | 3 |
| 2020 | A Closer Look at Deep Policy Gradients
Andrew Ilyas, Logan Engstrom, Shibani Santurkar, Dimitris Tsipras, Firdaus Janoos, Larry Rudolph, Aleksander Madry |
ICLR | 3 |
| 2020 | Identifying Statistical Bias in Dataset ReplicationabstractDataset replication is a useful tool for assessing whether improvements in test accuracy on a specific benchmark correspond to improvements in models’ ability to generalize reliably. In this work, we present unintuitive yet significant ways in which standard approaches to dataset replication introduce statistical bias, skewing the resulting observations. We study ImageNet-v2, a replication of the ImageNet dataset on which models exhibit a significant (11-14%) drop in accuracy, even after controlling for selection frequency, a human-in-the-loop measure of data quality. We show that after remeasuring selection frequencies and correcting for statistical bias, only an estimated 3.6% of the original 11.7% accuracy drop remains unaccounted for. We conclude with concrete recommendations for recognizing and avoiding bias in dataset replication. Code for our study is publicly available: https://git.io/data-rep-analysis. Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Jacob Steinhardt, Aleksander Madry |
ICML | 3 |
| 2020 | From ImageNet to Image Classification: Contextualizing Progress on BenchmarksabstractBuilding rich machine learning datasets in a scalable manner often necessitates a crowd-sourced data collection pipeline. In this work, we use human studies to investigate the consequences of employing such a pipeline, focusing on the popular ImageNet dataset. We study how specific design choices in the ImageNet creation process impact the fidelity of the resulting dataset—including the introduction of biases that state-of-the-art models exploit. Our analysis pinpoints how a noisy data collection pipeline can lead to a systematic misalignment between the resulting benchmark and the real-world task it serves as a proxy for. Finally, our findings emphasize the need to augment our current model training and evaluation toolkit to take such misalignment into account. Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Andrew Ilyas, Aleksander Madry |
ICML | 2 |
| 2019 | Robustness May Be at Odds with Accuracy
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Alexander Turner 0001, Aleksander Madry |
ICLR (Poster) | 2 |
| 2019 | Adversarial Examples Are Not Bugs, They Are FeaturesabstractAdversarial examples have attracted significant attention in machine learning, but the reasons for their existence and pervasiveness remain unclear. We demonstrate that adversarial examples can be directly attributed to the presence of non-robust features: features (derived from patterns in the data distribution) that are highly predictive, yet brittle and (thus) incomprehensible to humans. After capturing these features within a theoretical framework, we establish their widespread existence in standard datasets. Finally, we present a simple setting where we can rigorously tie the phenomena we observe in practice to a {\em misalignment} between the (human-specified) notion of robustness and the inherent geometry of the data. Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, Aleksander Madry |
NeurIPS | 2 |
| 2019 | Image Synthesis with a Single (Robust) ClassifierabstractWe show that the basic classification framework alone can be used to tackle some of the most challenging tasks in image synthesis. In contrast to other state-of-the-art approaches, the toolkit we develop is rather minimal: it uses a single, off-the-shelf classifier for all these tasks. The crux of our approach is that we train this classifier to be adversarially robust. It turns out that adversarial robustness is precisely what we need to directly manipulate salient features of the input. Overall, our findings demonstrate the utility of robustness in the broader machine learning context. Shibani Santurkar, Andrew Ilyas, Dimitris Tsipras, Logan Engstrom, Brandon Tran, Aleksander Madry |
NeurIPS | 1 |
| 2018 | A Classification-Based Study of Covariate Shift in GAN DistributionsabstractA basic, and still largely unanswered, question in the context of Generative Adversarial Networks (GANs) is whether they are truly able to capture all the fundamental characteristics of the distributions they are trained on. In particular, evaluating the diversity of GAN distributions is challenging and existing methods provide only a partial understanding of this issue. In this paper, we develop quantitative and scalable tools for assessing the diversity of GAN distributions. Specifically, we take a classification-based perspective and view loss of diversity as a form of covariate shift introduced by GANs. We examine two specific forms of such shift: mode collapse and boundary distortion. In contrast to prior work, our methods need only minimal human supervision and can be readily applied to state-of-the-art GANs on large, canonical datasets. Examining popular GANs using our tools indicates that these GANs have significant problems in reproducing the more distributional properties of their training dataset. Shibani Santurkar, Ludwig Schmidt, Aleksander Madry |
ICML | 1 |
| 2018 | How Does Batch Normalization Help Optimization?abstractBatch Normalization (BatchNorm) is a widely adopted technique that enables faster and more stable training of deep neural networks (DNNs). Despite its pervasiveness, the exact reasons for BatchNorm's effectiveness are still poorly understood. The popular belief is that this effectiveness stems from controlling the change of the layers' input distributions during training to reduce the so-called "internal covariate shift". In this work, we demonstrate that such distributional stability of layer inputs has little to do with the success of BatchNorm. Instead, we uncover a more fundamental impact of BatchNorm on the training process: it makes the optimization landscape significantly smoother. This smoothness induces a more predictive and stable behavior of the gradients, allowing for faster training. Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, Aleksander Madry |
NeurIPS | 1 |
| 2018 | Adversarially Robust Generalization Requires More DataabstractMachine learning models are often susceptible to adversarial perturbations of their inputs. Even small perturbations can cause state-of-the-art classifiers with high "standard" accuracy to produce an incorrect prediction with high confidence. To better understand this phenomenon, we study adversarially robust learning from the viewpoint of generalization. We show that already in a simple natural data model, the sample complexity of robust learning can be significantly larger than that of "standard" learning. This gap is information theoretic and holds irrespective of the training algorithm or the model family. We complement our theoretical results with experiments on popular image classification datasets and show that a similar gap exists here as well. We postulate that the difficulty of training robust classifiers stems, at least partially, from this inherently larger sample complexity. Ludwig Schmidt, Shibani Santurkar, Dimitris Tsipras, Kunal Talwar, Aleksander Madry |
NeurIPS | 2 |
| 2018 | Generative CompressionabstractTraditional image and video compression algorithms rely on hand-crafted encoder/decoder pairs (codecs) that lack adaptability and are agnostic to the data being compressed. We describe the concept of generative compression, the compression of data using generative models, and suggest that it is a direction worth pursuing to produce more accurate and visually pleasing reconstructions at deeper compression levels for both image and video data. We also show that generative compression is orders- of-magnitude more robust to bit errors (e.g., from noisy channels) than traditional variable-length coding schemes. Shibani Santurkar, David M. Budden, Nir Shavit |
PCS | 1 |
| 2017 | Deep Tensor Convolution on MulticoresabstractDeep convolutional neural networks (ConvNets) of 3-dimensional kernels allow joint modeling of spatiotemporal features. These networks have improved performance of video and volumetric image analysis, but have been limited in size due to the low memory ceiling of GPU hardware. Existing CPU implementations overcome this constraint but are impractically slow. Here we extend and optimize the faster Winograd-class of convolutional algorithms to the $N$-dimensional case and specifically for CPU hardware. First, we remove the need to manually hand-craft algorithms by exploiting the relaxed constraints and cheap sparse access of CPU memory. Second, we maximize CPU utilization and multicore scalability by transforming data matrices to be cache-aware, integer multiples of AVX vector widths. Treating 2-dimensional ConvNets as a special (and the least beneficial) case of our approach, we demonstrate a 5 to 25-fold improvement in throughput compared to previous state-of-the-art. David M. Budden, Alexander Matveev, Shibani Santurkar, Shraman Ray Chaudhuri, Nir Shavit |
ICML | 3 |
| 2015 | C. elegans chemotaxis inspired neuromorphic circuit for contour tracking and obstacle avoidanceabstractWe demonstrate a spiking neural network for navigation motivated by the chemotaxis circuit of Caenorhabditis elegans. Our network uses information regarding temporal gradients in intensity of local variables such as chemical concentration, temperature, radiation, etc., to make navigational decisions for contour tracking and obstacle avoidance. The gradient information is determined by mimicking the underlying mechanisms of the ASE neurons of C. elegans. Simulations show that our software-worm is able to identify the set-point with 92% efficiency, 68.5% higher than an optimal memoryless Lévy foraging strategy and 33% higher than an equivalent non-spiking neural network configuration. The software-worm is able to track the set-point with an average deviation of 1% from the set-point, and this performance degrades merely by 1.8% in the presence of intense salt and pepper noise in the local tracking variable. We also develop a VLSI implementation for the main gradient detector neurons, which could be integrated with standard comparator circuitry to develop robust circuits for navigation and contour tracking. We demonstrate noise-resilience of our network to environmental, architectural and circuit noise. Shibani Santurkar, Bipin Rajendran |
IJCNN | 1 |