Nari Johnson

dblp:302/3945 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Trustworthy machine learning · 100%
Human-computer interaction and pervasive computing
1 paper
Usability and user experience research · 50% Human-AI interaction · 50%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › interpretability
explanation evaluation
1.122022
Use-Case-Grounded Simulations for Explanation Evaluation · NeurIPS 2022
OpenXAI: Towards a Transparent Evaluation of Model Explanations · NeurIPS 2022
Machine learning › Trustworthy machine learning
interpretability
1.122022
Use-Case-Grounded Simulations for Explanation Evaluation · NeurIPS 2022
OpenXAI: Towards a Transparent Evaluation of Model Explanations · NeurIPS 2022
Machine learning › Trustworthy machine learning › interpretability
post-hoc explanation
0.612022
OpenXAI: Towards a Transparent Evaluation of Model Explanations · NeurIPS 2022
Human-AI interaction
simulation-based evaluation
0.612022
Use-Case-Grounded Simulations for Explanation Evaluation · NeurIPS 2022
Usability and user experience research
user study
0.612022
Use-Case-Grounded Simulations for Explanation Evaluation · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

simulated evaluations · 1.1algorithmic agents · 1.1feature attribution · 0.6benchmarking · 0.6
YearPublicationVenuePosition
2022 OpenXAI: Towards a Transparent Evaluation of Model Explanations
abstract
While several types of post hoc explanation methods have been proposed in recent literature, there is very little work on systematically benchmarking these methods. Here, we introduce OpenXAI, a comprehensive and extensible open-source framework for evaluating and benchmarking post hoc explanation methods. OpenXAI comprises of the following key components: (i) a flexible synthetic data generator and a collection of diverse real-world datasets, pre-trained models, and state-of-the-art feature attribution methods, (ii) open-source implementations of twenty-two quantitative metrics for evaluating faithfulness, stability (robustness), and fairness of explanation methods, and (iii) the first ever public XAI leaderboards to readily compare several explanation methods across a wide variety of metrics, models, and datasets. OpenXAI is easily extensible, as users can readily evaluate custom explanation methods and incorporate them into our leaderboards. Overall, OpenXAI provides an automated end-to-end pipeline that not only simplifies and standardizes the evaluation of post hoc explanation methods, but also promotes transparency and reproducibility in benchmarking these methods. While the first release of OpenXAI supports only tabular datasets, the explanation methods and metrics that we consider are general enough to be applicable to other data modalities. OpenXAI datasets and data loaders, implementations of state-of-the-art explanation methods and evaluation metrics, as well as leaderboards are publicly available at https://open-xai.github.io/. OpenXAI will be regularly updated to incorporate text and image datasets, other new metrics and explanation methods, and welcomes inputs from the community.
Chirag Agarwal, Satyapriya Krishna, Eshika Saxena, Martin Pawelczyk, Nari Johnson, Isha Puri, Marinka Zitnik, Himabindu Lakkaraju
NeurIPS5
2022 Use-Case-Grounded Simulations for Explanation Evaluation
abstract
A growing body of research runs human subject evaluations to study whether providing users with explanations of machine learning models can help them with practical real-world use cases. However, running user studies is challenging and costly, and consequently each study typically only evaluates a limited number of different settings, e.g., studies often only evaluate a few arbitrarily selected model explanation methods. To address these challenges and aid user study design, we introduce Simulated Evaluations (SimEvals). SimEvals involve training algorithmic agents that take as input the information content (such as model explanations) that would be presented to the user, to predict answers to the use case of interest. The algorithmic agent's test set accuracy provides a measure of the predictiveness of the information content for the downstream use case. We run a comprehensive evaluation on three real-world use cases (forward simulation, model debugging, and counterfactual reasoning) to demonstrate that SimEvals can effectively identify which explanation methods will help humans for each use case. These results provide evidence that \simevals{} can be used to efficiently screen an important set of user study design decisions, e.g., selecting which explanations should be presented to the user, before running a potentially costly user study.
Valerie Chen, Nari Johnson, Nicholay Topin, Gregory Plumb, Ameet Talwalkar
NeurIPS2
2021 Learning Predictive and Interpretable Timeseries Summaries from ICU Data
Nari Johnson, Sonali Parbhoo, Andrew Slavin Ross, Finale Doshi-Velez
AMIA1