Davin Hill

dblp:323/4185 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Trustworthy machine learning · 51% Learning paradigms · 32% Deep learning architectures and training · 16%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
2.132025
OrdShap: Feature Position Importance for Sequential Black-Box Models · NeurIPS 2025
SmoothHess: ReLU Network Feature Interactions via Stein's Lemma · NeurIPS 2023
Explanations of Black-Box Models based on Directional Feature Interactions · ICLR 2022
Machine learning › Learning paradigms › continual learning
catastrophic forgetting
0.912025
STAR: Stability-Inducing Weight Perturbation for Continual Learning · ICLR 2025
Machine learning › Learning paradigms
continual learning
0.912025
STAR: Stability-Inducing Weight Perturbation for Continual Learning · ICLR 2025
Machine learning › Trustworthy machine learning › interpretability › attribution methods
feature attribution
0.912025
OrdShap: Feature Position Importance for Sequential Black-Box Models · NeurIPS 2025
Machine learning › Learning paradigms › continual learning
rehearsal-based continual learning
0.912025
STAR: Stability-Inducing Weight Perturbation for Continual Learning · ICLR 2025
Machine learning › Deep learning architectures and training
feature interaction
0.712023
SmoothHess: ReLU Network Feature Interactions via Stein's Lemma · NeurIPS 2023
Machine learning › Deep learning architectures and training
ReLU networks
0.712023
SmoothHess: ReLU Network Feature Interactions via Stein's Lemma · NeurIPS 2023
Machine learning › Trustworthy machine learning › interpretability › explainable AI
interactive explanation
0.612022
Explanations of Black-Box Models based on Directional Feature Interactions · ICLR 2022
Machine learning › Trustworthy machine learning › interpretability › post-hoc explanation
model-agnostic explanation
0.612022
Explanations of Black-Box Models based on Directional Feature Interactions · ICLR 2022

Methods — techniques the papers use, named apart from their topics

weight perturbation · 0.9shapley value · 0.9loss function design · 0.9KL divergence · 0.9stein's lemma · 0.7sampling · 0.7hessian estimation · 0.7directional feature interaction · 0.6
YearPublicationVenuePosition
2025 Axiomatic Explainer Globalness via Optimal Transport
abstract
Explainability methods are often challenging to evaluate and compare. With a multitude of explainers available, practitioners must often compare and select explainers based on quantitative evaluation metrics. One particular differentiator between explainers is the diversity of explanations for a given dataset; i.e. whether all explanations are identical, unique and uniformly distributed, or somewhere between these two extremes. In this work, we define a complexity measure for explainers, globalness, which enables deeper understanding of the distribution of explanations produced by feature attribution and feature selection methods for a given dataset. We establish the axiomatic properties that any such measure should possess and prove that our proposed measure, Wasserstein Globalness, meets these criteria. We validate the utility of Wasserstein Globalness using image, tabular, and synthetic datasets, empirically showing that it both facilitates meaningful comparison between explainers and improves the selection process for explainability methods.
Davin Hill, Joshua T. Bone, Aria Masoomi, Max Torop, Jennifer G. Dy
AISTATS1
2025 STAR: Stability-Inducing Weight Perturbation for Continual Learning
abstract
Humans can naturally learn new and varying tasks in a sequential manner. Continual learning is a class of learning algorithms that updates its learned model as it sees new data (on potentially new tasks) in a sequence. A key challenge in continual learning is that as the model is updated to learn new tasks, it becomes susceptible to \textit{catastrophic forgetting}, where knowledge of previously learned tasks is lost. A popular approach to mitigate forgetting during continual learning is to maintain a small buffer of previously-seen samples, and to replay them during training. However, this approach is limited by the small buffer size and, while forgetting is reduced, it is still present. In this paper, we propose a novel loss function STAR that exploits the worst-case parameter perturbation that reduces the KL-divergence of model predictions with that of its local parameter neighborhood to promote stability and alleviate forgetting. STAR can be combined with almost any existing rehearsal-based methods as a plug-and-play component. We empirically show that STAR consistently improves performance of existing methods by up to $\sim15\\%$ across varying baselines, and achieves superior or competitive accuracy to that of state-of-the-art methods aimed at improving rehearsal-based continual learning. Our implementation is available at https://github.com/Gnomy17/STAR_CL.
Masih Eskandar, Tooba Imtiaz, Davin Hill, Zifeng Wang 0002, Jennifer G. Dy
ICLR3
2025 OrdShap: Feature Position Importance for Sequential Black-Box Models
abstract
Sequential deep learning models excel in domains with temporal or sequential dependencies, but their complexity necessitates post-hoc feature attribution methods for understanding their predictions. While existing techniques quantify feature importance, they inherently assume fixed feature ordering — conflating the effects of (1) feature values and (2) their positions within input sequences. To address this gap, we introduce OrdShap, a novel attribution method that disentangles these effects by quantifying how a model's predictions change in response to permuting feature position. We establish a game-theoretic connection between OrdShap and Sanchez-Bergantiños values, providing a theoretically grounded approach to position-sensitive attribution. Empirical results from health, natural language, and synthetic datasets highlight OrdShap's effectiveness in capturing feature value and feature position attributions, and provide deeper insight into model behavior.
Davin Hill, Brian L. Hill, Aria Masoomi, Vijay S. Nori, Robert E. Tillman, Jennifer G. Dy
NeurIPS1
2024 Boundary-Aware Uncertainty for Feature Attribution Explainers
abstract
Post-hoc explanation methods have become a critical tool for understanding black-box classifiers in high-stakes applications. However, high-performing classifiers are often highly nonlinear and can exhibit complex behavior around the decision boundary, leading to brittle or misleading local explanations. Therefore there is an impending need to quantify the uncertainty of such explanation methods in order to understand when explanations are trustworthy. In this work we propose the Gaussian Process Explanation unCertainty (GPEC) framework, which generates a unified uncertainty estimate combining decision boundary-aware uncertainty with explanation function approximation uncertainty. We introduce a novel geodesic-based kernel, which captures the complexity of the target black-box decision boundary. We show theoretically that the proposed kernel similarity increases with decision boundary complexity. The proposed framework is highly flexible; it can be used with any black-box classifier and feature attribution method. Empirical results on multiple tabular and image datasets show that the GPEC uncertainty estimate improves understanding of explanations as compared to existing methods.
Davin Hill, Aria Masoomi, Max Torop, Sandesh Ghimire, Jennifer G. Dy
AISTATS1
2024 Analyzing Explainer Robustness via Probabilistic Lipschitzness of Prediction Functions
Zulqarnain Khan, Davin Hill, Aria Masoomi, Joshua T. Bone, Jennifer G. Dy
AISTATS2
2023 SmoothHess: ReLU Network Feature Interactions via Stein's Lemma
abstract
Several recent methods for interpretability model feature interactions by looking at the Hessian of a neural network. This poses a challenge for ReLU networks, which are piecewise-linear and thus have a zero Hessian almost everywhere. We propose SmoothHess, a method of estimating second-order interactions through Stein's Lemma. In particular, we estimate the Hessian of the network convolved with a Gaussian through an efficient sampling algorithm, requiring only network gradient calls. SmoothHess is applied post-hoc, requires no modifications to the ReLU network architecture, and the extent of smoothing can be controlled explicitly. We provide a non-asymptotic bound on the sample complexity of our estimation procedure. We validate the superior ability of SmoothHess to capture interactions on benchmark datasets and a real-world medical spirometry dataset.
Max Torop, Aria Masoomi, Davin Hill, Kivanç Köse, Stratis Ioannidis, Jennifer G. Dy
NeurIPS3
2022 Explanations of Black-Box Models based on Directional Feature Interactions
Aria Masoomi, Davin Hill, Zhonghui Xu, Craig P. Hersh, Edwin K. Silverman, Peter J. Castaldi, Stratis Ioannidis, Jennifer G. Dy
ICLR2