VLDB 2026 Research / reviewers in the wild / expert
Saachi Jain
dblp:227/2617
· DBLP profile ↗
8ranked-venue papers
6as first author
6since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Trustworthy machine learning · 73% Transfer learning and domain adaptation · 11% Representation and self-supervised learning · 6% |
Topics — the 15 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
robustness |
1.7 | 3 | 2022 | Combining Diverse Feature Priors · ICML 2022 Missingness Bias in Model Debugging · ICLR 2022 Certified Patch Robustness via Smoothed Vision Transformers · CVPR 2022 |
Machine learning › Trustworthy machine learning
interpretability |
1.2 | 2 | 2023 | Distilling Model Failures as Directions in Latent Space · ICLR 2023 Missingness Bias in Model Debugging · ICLR 2022 |
Machine learning › Trustworthy machine learning
debiasing |
0.8 | 1 | 2024 | Improving Subgroup Robustness via Data Selection · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
fairness |
0.8 | 1 | 2024 | Improving Subgroup Robustness via Data Selection · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › robustness › distributional robustness
subgroup robustness |
0.8 | 1 | 2024 | Improving Subgroup Robustness via Data Selection · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › Data-centric AI
data leakage detection |
0.7 | 1 | 2023 | A Data-Based Perspective on Transfer Learning · CVPR 2023 |
Machine learning › Representation and self-supervised learning
latent space representation |
0.7 | 1 | 2023 | Distilling Model Failures as Directions in Latent Space · ICLR 2023 |
Machine learning › Transfer learning and domain adaptation › multi-source learning
source data selection |
0.7 | 1 | 2023 | A Data-Based Perspective on Transfer Learning · CVPR 2023 |
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
certified defense |
0.6 | 1 | 2022 | Certified Patch Robustness via Smoothed Vision Transformers · CVPR 2022 |
Computer vision › Image recognition and object detection
image classification |
0.6 | 1 | 2022 | Certified Patch Robustness via Smoothed Vision Transformers · CVPR 2022 |
Machine learning › Trustworthy machine learning › interpretability
model debugging |
0.6 | 1 | 2022 | Missingness Bias in Model Debugging · ICLR 2022 |
Machine learning › Trustworthy machine learning › robustness › certified robustness
patch robustness certification |
0.6 | 1 | 2022 | Certified Patch Robustness via Smoothed Vision Transformers · CVPR 2022 |
Machine learning › Trustworthy machine learning › robustness › spurious correlation
spurious correlation robustness |
0.6 | 1 | 2022 | Combining Diverse Feature Priors · ICML 2022 |
Natural language and speech › Question answering and dialogue systems
dialogue |
0.4 | 1 | 2019 | Learning to Speak and Act in a Fantasy Text Adventure Game · EMNLP/IJCNLP (1) 2019 |
Machine learning › Kernel, tree and ensemble methods
model combination |
0.2 | 1 | 2022 | Combining Diverse Feature Priors · ICML 2022 |
Methods — techniques the papers use, named apart from their topics
datamodels · 0.8data selection · 0.8model distillation · 0.7influence functions · 0.7data attribution · 0.7vision transformer · 0.6semi-supervised learning · 0.6randomized smoothing · 0.6feature priors · 0.6ensemble combination · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Improving Subgroup Robustness via Data SelectionabstractMachine learning models can often fail on subgroups that are underrepresented
during training. While dataset balancing can improve performance on
underperforming groups, it requires access to training group annotations and can
end up removing large portions of the dataset. In this paper, we introduce
Data Debiasing with Datamodels (D3M), a debiasing approach
which isolates and removes specific training examples that drive the model's
failures on minority groups. Our approach enables us to efficiently train
debiased classifiers while removing only a small number of examples, and does
not require training group annotations or additional hyperparameter tuning. Saachi Jain, Kimia Hamidieh, Kristian Georgiev, Andrew Ilyas, Marzyeh Ghassemi, Aleksander Madry |
NeurIPS | 1 |
| 2023 | A Data-Based Perspective on Transfer LearningabstractIt is commonly believed that in transfer learning including more pre-training data translates into better performance. However, recent evidence suggests that removing data from the source dataset can actually help too. In this work, we take a closer look at the role of the source dataset's composition in transfer learning and present a framework for probing its impact on downstream performance. Our framework gives rise to new capabilities such as pinpointing transfer learning brittleness as well as detecting pathologies such as data-leakage and the presence of misleading examples in the source dataset. In particular, we demonstrate that removing detrimental datapoints identified by our framework indeed improves transfer learning performance from ImageNet on a variety of target tasks.11Code is available at https://github.com/MadryLab/data-transfer Saachi Jain, Hadi Salman, Alaa Khaddaj, Eric Wong 0001, Aleksander Madry |
CVPR | 1 |
| 2023 | Distilling Model Failures as Directions in Latent Space
Saachi Jain, Hannah Lawrence, Ankur Moitra, Aleksander Madry |
ICLR | 1 |
| 2022 | Certified Patch Robustness via Smoothed Vision TransformersabstractCertified patch defenses can guarantee robustness of an image classifier to arbitrary changes within a bounded contiguous region. But, currently, this robustness comes at a cost of degraded standard accuracies and slower inference times. We demonstrate how using vision transformers enables significantly better certified patch robustness that is also more computationally efficient and does not incur a substantial drop in standard accuracy. These improvements stem from the inherent ability of the vision transformer to gracefully handle largely masked images.11Our code is available at https://github.com/MadryLab/smoothed-vit.. Hadi Salman, Saachi Jain, Eric Wong 0001, Aleksander Madry |
CVPR | 2 |
| 2022 | Missingness Bias in Model Debugging
Saachi Jain, Hadi Salman, Eric Wong 0001, Pengchuan Zhang, Vibhav Vineet, Sai Vemprala, Aleksander Madry |
ICLR | 1 |
| 2022 | Combining Diverse Feature PriorsabstractTo improve model generalization, model designers often restrict the features that their models use, either implicitly or explicitly. In this work, we explore the design space of leveraging such feature priors by viewing them as distinct perspectives on the data. Specifically, we find that models trained with diverse sets of explicit feature priors have less overlapping failure modes, and can thus be combined more effectively. Moreover, we demonstrate that jointly training such models on additional (unlabeled) data allows them to correct each other’s mistakes, which, in turn, leads to better generalization and resilience to spurious correlations. Saachi Jain, Dimitris Tsipras, Aleksander Madry |
ICML | 1 |
| 2020 | Spectral Lower Bounds on the I/O Complexity of Computation GraphsabstractWe consider the problem of finding lower bounds on the I/O complexity of arbitrary computations in a two level memory hierarchy. Executions of complex computations can be formalized as an evaluation order over the underlying computation graph. However, prior methods for finding I/O lower bounds leverage the graph structures for specific problems (e.g matrix multiplication) which cannot be applied to arbitrary graphs. In this paper, we first present a novel method to bound the I/O of any computation graph using the first few eigenvalues of the graph's Laplacian. We further extend this bound to the parallel setting. This spectral bound is not only efficiently computable by power iteration, but can also be computed in closed form for graphs with known spectra. We apply our spectral method to compute closed-form analytical bounds on two computation graphs (the Bellman-Held-Karp algorithm for the traveling salesman problem and the Fast Fourier Transform), as well as provide a probabilistic bound for random Erdos Renyi graphs. We empirically validate our bound on four computation graphs, and find that our method provides tighter bounds than current empirical methods and behaves similarly to previously published I/O bounds. Saachi Jain, Matei Zaharia |
SPAA | 1 |
| 2019 | Learning to Speak and Act in a Fantasy Text Adventure GameabstractJack Urbanek, Angela Fan, Siddharth Karamcheti, Saachi Jain, Samuel Humeau, Emily Dinan, Tim Rocktäschel, Douwe Kiela, Arthur Szlam, Jason Weston. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jack Urbanek, Angela Fan, Siddharth Karamcheti, Saachi Jain, Samuel Humeau 0001, Emily Dinan, Tim Rocktäschel, Douwe Kiela, Arthur Szlam, Jason Weston |
EMNLP/IJCNLP (1) | 4 |