VLDB 2026 Research / reviewers in the wild / expert
Abhineet Agarwal
dblp:304/4687
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Trustworthy machine learning · 62% Reinforcement learning · 16% Probabilistic and Bayesian machine learning · 14% | |
| Theoretical computer science
1 paper |
Approximation and online algorithms · 50% Algorithmic game theory and mechanism design · 50% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% |
Topics — the 15 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
interpretability |
1.9 | 3 | 2025 | ProxySPEX: Inference-Efficient Interpretability via Sparse Feature Interactions in LLMs · NeurIPS 2025 SPEX: Scaling Feature Interaction Explanations for LLMs · ICML 2025 Hierarchical Shrinkage: Improving the accuracy and interpretability of tree-based models · ICML 2022 |
Machine learning › Trustworthy machine learning › interpretability › attribution methods
feature interaction attribution |
1.7 | 2 | 2025 | ProxySPEX: Inference-Efficient Interpretability via Sparse Feature Interactions in LLMs · NeurIPS 2025 SPEX: Scaling Feature Interaction Explanations for LLMs · ICML 2025 |
Machine learning › Trustworthy machine learning › language model interpretability
large language model explanation |
0.9 | 1 | 2025 | SPEX: Scaling Feature Interaction Explanations for LLMs · ICML 2025 |
Machine learning › Trustworthy machine learning › interpretability › shapley value
shapley value estimation |
0.9 | 1 | 2025 | ProxySPEX: Inference-Efficient Interpretability via Sparse Feature Interactions in LLMs · NeurIPS 2025 |
Machine learning › Reinforcement learning
multi-armed bandit |
0.8 | 1 | 2024 | Mutli-Armed Bandits with Network Interference · NeurIPS 2024 |
Machine learning › Reinforcement learning
reinforcement learning for healthcare |
0.8 | 1 | 2024 | ED-Copilot: Reduce Emergency Department Wait Time with Language Model Diagnostic Assistance · ICML 2024 |
Medical and health informatics
clinical decision support |
0.8 | 1 | 2024 | ED-Copilot: Reduce Emergency Department Wait Time with Language Model Diagnostic Assistance · ICML 2024 |
Approximation and online algorithms
online learning |
0.8 | 1 | 2024 | Mutli-Armed Bandits with Network Interference · NeurIPS 2024 |
Algorithmic game theory and mechanism design
regret minimization |
0.8 | 1 | 2024 | Mutli-Armed Bandits with Network Interference · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning
causal inference |
0.7 | 1 | 2023 | Synthetic Combinations: A Causal Inference Framework for Combinatorial Interventions · NeurIPS 2023 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference
potential outcomes |
0.7 | 1 | 2023 | Synthetic Combinations: A Causal Inference Framework for Combinatorial Interventions · NeurIPS 2023 |
Machine learning › Kernel, tree and ensemble methods
tree-based models |
0.6 | 1 | 2022 | Hierarchical Shrinkage: Improving the accuracy and interpretability of tree-based models · ICML 2022 |
Machine learning › Trustworthy machine learning
language model interpretability |
0.3 | 1 | 2025 | ProxySPEX: Inference-Efficient Interpretability via Sparse Feature Interactions in LLMs · NeurIPS 2025 |
Natural language and speech › Language models and text generation › pre-trained language model
biomedical language model |
0.2 | 1 | 2024 | ED-Copilot: Reduce Emergency Department Wait Time with Language Model Diagnostic Assistance · ICML 2024 |
Machine learning › Trustworthy machine learning › interpretability
shapley value |
0.2 | 1 | 2022 | Hierarchical Shrinkage: Improving the accuracy and interpretability of tree-based models · ICML 2022 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 1.5pre-trained biomedical language model · 1.5linear regression · 1.5discrete fourier analysis · 1.5sparse fourier transform · 0.9masked inference · 0.9interaction sparsity · 0.9gradient boosted trees · 0.9channel decoding · 0.9SHAP · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SPEX: Scaling Feature Interaction Explanations for LLMsabstractLarge language models (LLMs) have revolutionized machine learning due to their ability to capture complex interactions between input features. Popular post-hoc explanation methods like SHAP provide marginal feature attributions, while their extensions to interaction importances only scale to small input lengths ($\approx 20$). We propose Spectral Explainer (SPEX), a model-agnostic interaction attribution algorithm that efficiently scales to large input lengths ($\approx 1000)$. SPEX exploits underlying natural sparsity among interactions—common in real-world data—and applies a sparse Fourier transform using a channel decoding algorithm to efficiently identify important interactions. We perform experiments across three difficult long-context datasets that require LLMs to utilize interactions between inputs to complete the task. For large inputs, SPEX outperforms marginal attribution methods by up to 20% in terms of faithfully reconstructing LLM outputs. Further, SPEX successfully identifies key features and interactions that strongly influence model output. For one of our datasets, HotpotQA, SPEX provides interactions that align with human annotations. Finally, we use our model-agnostic approach to generate explanations to demonstrate abstract reasoning in closed-source LLMs (GPT-4o mini) and compositional reasoning in vision-language models. Justin Singh Kang, Landon Butler, Abhineet Agarwal, Yigit Efe Erginbas, Ramtin Pedarsani, Bin Yu 0001, Kannan Ramchandran |
ICML | 3 |
| 2025 | ProxySPEX: Inference-Efficient Interpretability via Sparse Feature Interactions in LLMsabstractLarge Language Models (LLMs) have achieved remarkable performance by capturing complex interactions between input features. To identify these interactions, most existing approaches require enumerating all possible combinations of features up to a given order, causing them to scale poorly with the number of inputs $n$. Recently, Kang et al. (2025) proposed SPEX, an information-theoretic approach that uses interaction sparsity to scale to $n \approx 10^3$ features. SPEX greatly improves upon prior methods but requires tens of thousands of model inferences, which can be prohibitive for large models. In this paper, we observe that LLM feature interactions are often *hierarchical*—higher-order interactions are accompanied by their lower-order subsets—which enables more efficient discovery. To exploit this hierarchy, we propose ProxySPEX, an interaction attribution algorithm that first fits gradient boosted trees to masked LLM outputs and then extracts the important interactions. Experiments across four challenging high-dimensional datasets show that ProxySPEX more faithfully reconstructs LLM outputs by 20\% over marginal attribution approaches while using *$10\times$ fewer inferences* than SPEX. By accounting for interactions, ProxySPEX efficiently identifies the most influential features, providing a scalable approximation of their Shapley values. Further, we apply ProxySPEX to two interpretability tasks. *Data attribution*, where we identify interactions among CIFAR-10 training samples that influence test predictions, and *mechanistic interpretability*, where we uncover interactions between attention heads, both within and across layers, on a question-answering task. The ProxySPEX algorithm is available at <https://github.com/mmschlk/shapiq>. Landon Butler, Abhineet Agarwal, Justin Singh Kang, Yigit Efe Erginbas, Bin Yu 0001, Kannan Ramchandran |
NeurIPS | 2 |
| 2024 | ED-Copilot: Reduce Emergency Department Wait Time with Language Model Diagnostic AssistanceabstractIn the emergency department (ED), patients undergo triage and multiple laboratory tests before diagnosis. This time-consuming process causes ED crowding which impacts patient mortality, medical errors, staff burnout, etc. This work proposes (time) cost-effective diagnostic assistance that leverages artificial intelligence systems to help ED clinicians make efficient and accurate diagnoses. In collaboration with ED clinicians, we use public patient data to curate MIMIC-ED-Assist, a benchmark for AI systems to suggest laboratory tests that minimize wait time while accurately predicting critical outcomes such as death. With MIMIC-ED-Assist, we develop ED-Copilot which sequentially suggests patient-specific laboratory tests and makes diagnostic predictions. ED-Copilot employs a pre-trained bio-medical language model to encode patient information and uses reinforcement learning to minimize ED wait time and maximize prediction accuracy. On MIMIC-ED-Assist, ED-Copilot improves prediction accuracy over baselines while halving average wait time from four hours to two hours. ED-Copilot can also effectively personalize treatment recommendations based on patient severity, further highlighting its potential as a diagnostic assistant. Since MIMIC-ED-Assist is a retrospective benchmark, ED-Copilot is restricted to recommend only observed tests. We show ED-Copilot achieves competitive performance without this restriction as the maximum allowed time increases. Our code is available at https://github.com/cxcscmu/ED-Copilot. Liwen Sun, Abhineet Agarwal, Aaron Kornblith, Bin Yu 0001, Chenyan Xiong |
ICML | 2 |
| 2024 | Mutli-Armed Bandits with Network InterferenceabstractOnline experimentation with interference is a common challenge in modern applications such as e-commerce and adaptive clinical trials in medicine. For example, in online marketplaces, the revenue of a good depends on discounts applied to competing goods. Statistical inference with interference is widely studied in the offline setting, but far less is known about how to adaptively assign treatments to minimize regret. We address this gap by studying a multi-armed bandit (MAB) problem where a learner (e-commerce platform) sequentially assigns one of possible $\mathcal{A}$ actions (discounts) to $N$ units (goods) over $T$ rounds to minimize regret (maximize revenue). Unlike traditional MAB problems, the reward of each unit depends on the treatments assigned to other units, i.e., there is *interference* across the underlying network of units. With $\mathcal{A}$ actions and $N$ units, minimizing regret is combinatorially difficult since the action space grows as $\mathcal{A}^N$. To overcome this issue, we study a *sparse network interference* model, where the reward of a unit is only affected by the treatments assigned to $s$ neighboring units. We use tools from discrete Fourier analysis to develop a sparse linear representation of the unit-specific reward $r_n: [\mathcal{A}]^N \rightarrow \mathbb{R} $, and propose simple, linear regression-based algorithms to minimize regret. Importantly, our algorithms achieve provably low regret both when the learner observes the interference neighborhood for all units and when it is unknown. This significantly generalizes other works on this topic which impose strict conditions on the strength of interference on a *known* network, and also compare regret to a markedly weaker optimal action.
Empirically, we corroborate our theoretical findings via numerical simulations. Abhineet Agarwal, Anish Agarwal, Lorenzo Masoero, Justin Whitehouse |
NeurIPS | 1 |
| 2023 | Synthetic Combinations: A Causal Inference Framework for Combinatorial InterventionsabstractWe consider a setting where there are $N$ heterogeneous units and $p$ interventions. Our goal is to learn unit-specific potential outcomes for any combination of these $p$ interventions, i.e., $N \times 2^p$ causal parameters. Choosing a combination of interventions is a problem that naturally arises in a variety of applications such as factorial design experiments and recommendation engines (e.g., showing a set of movies that maximizes engagement for a given user). Running $N \times 2^p$ experiments to estimate the various parameters is likely expensive and/or infeasible as $N$ and $p$ grow. Further, with observational data there is likely confounding, i.e., whether or not a unit is seen under a combination is correlated with its potential outcome under that combination. We study this problem under a novel model that imposes latent structure across both units and combinations of interventions. Specifically, we assume latent similarity in potential outcomes across units (i.e., the matrix of potential outcomes is approximately rank $r$) and regularity in how combinations of interventions interact (i.e., the coefficients in the Fourier expansion of the potential outcomes is approximately $s$ sparse). We establish identification for all $N \times 2^p$ parameters despite unobserved confounding. We propose an estimation procedure, Synthetic Combinations, and establish finite-sample consistency under precise conditions on the observation pattern. We show that Synthetic Combinations is able to consistently estimate unit-specific potential outcomes given a total of $\text{poly}(r) \times \left( N + s^2p\right)$ observations. In comparison, previous methods that do not exploit structure across both units and combinations have poorer sample complexity scaling as $\min(N \times s^2p, \ \ r \times (N + 2^p))$. Abhineet Agarwal, Anish Agarwal, Suhas Vijaykumar |
NeurIPS | 1 |
| 2022 | A cautionary tale on fitting decision trees to data from additive models: generalization lower boundsabstractDecision trees are important both as interpretable models amenable to high-stakes decision-making, and as building blocks of ensemble methods such as random forests and gradient boosting. Their statistical properties, however, are not well understood. The most cited prior works have focused on deriving pointwise consistency guarantees for CART in a classical nonparametric regression setting. We take a different approach, and advocate studying the generalization performance of decision trees with respect to different generative regression models. This allows us to elicit their inductive bias, that is, the assumptions the algorithms make (or do not make) to generalize to new data, thereby guiding practitioners on when and how to apply these methods. In this paper, we focus on sparse additive generative models, which have both low statistical complexity and some nonparametric flexibility. We prove a sharp squared error generalization lower bound for a large class of decision tree algorithms fitted to sparse additive models with $C^1$ component functions. This bound is surprisingly much worse than the minimax rate for estimating such sparse additive models. The inefficiency is due not to greediness, but to the loss in power for detecting global structure when we average responses solely over each leaf, an observation that suggests opportunities to improve tree-based algorithms, for example, by hierarchical shrinkage. To prove these bounds, we develop new technical machinery, establishing a novel connection between decision tree estimation and rate-distortion theory, a sub-field of information theory. Yan Shuo Tan, Abhineet Agarwal, Bin Yu 0001 |
AISTATS | 2 |
| 2022 | Hierarchical Shrinkage: Improving the accuracy and interpretability of tree-based modelsabstractDecision trees and random forests (RF) are a cornerstone of modern machine learning practice. Due to their tendency to overfit, trees are typically regularized by a variety of techniques that modify their structure (e.g. pruning). We introduce Hierarchical Shrinkage (HS), a post-hoc algorithm which regularizes the tree not by altering its structure, but by shrinking the prediction over each leaf toward the sample means over each of its ancestors, with weights depending on a single regularization parameter and the number of samples in each ancestor. Since HS is a post-hoc method, it is extremely fast, compatible with any tree-growing algorithm and can be used synergistically with other regularization techniques. Extensive experiments over a wide variety of real-world datasets show that HS substantially increases the predictive performance of decision trees even when used in conjunction with other regularization techniques. Moreover, we find that applying HS to individual trees in a RF often improves its accuracy and interpretability by simplifying and stabilizing decision boundaries and SHAP values. We further explain HS by showing that it to be equivalent to ridge regression on a basis that is constructed of decision stumps associated to the internal nodes of a tree. All code and models are released in a full-fledged package available on Github Abhineet Agarwal, Yan Shuo Tan, Omer Ronen, Chandan Singh, Bin Yu 0001 |
ICML | 1 |