Katie Matton

dblp:249/5533 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
1since 2021 · last 2025
0000-0002-6786-7245ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Trustworthy machine learning · 90% Question answering and dialogue systems · 10%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
0.912025
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations · ICLR 2025
Machine learning › Trustworthy machine learning › language model interpretability
large language model explanation
0.912025
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations · ICLR 2025
Machine learning › Trustworthy machine learning
fairness
0.312025
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations · ICLR 2025
Natural language and speech › Question answering and dialogue systems › domain-specific question answering
medical question answering
0.312025
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations · ICLR 2025
Machine learning › Trustworthy machine learning › fairness
social bias
0.312025
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations · ICLR 2025

Methods — techniques the papers use, named apart from their topics

hierarchical bayesian modeling · 0.9counterfactual generation · 0.9causal effect estimation · 0.9
YearPublicationVenuePosition
2025 Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations
abstract
Large language models (LLMs) are capable of generating *plausible* explanations of how they arrived at an answer to a question. However, these explanations can misrepresent the model's "reasoning" process, i.e., they can be *unfaithful*. This, in turn, can lead to over-trust and misuse. We introduce a new approach for measuring the faithfulness of LLM explanations. First, we provide a rigorous definition of faithfulness. Since LLM explanations mimic human explanations, they often reference high-level *concepts* in the input question that purportedly influenced the model. We define faithfulness in terms of the difference between the set of concepts that the LLM's *explanations imply* are influential and the set that *truly* are. Second, we present a novel method for estimating faithfulness that is based on: (1) using an auxiliary LLM to modify the values of concepts within model inputs to create realistic counterfactuals, and (2) using a hierarchical Bayesian model to quantify the causal effects of concepts at both the example- and dataset-level. Our experiments show that our method can be used to quantify and discover interpretable patterns of unfaithfulness. On a social bias task, we uncover cases where LLM explanations hide the influence of social bias. On a medical question answering task, we uncover cases where LLM explanations provide misleading claims about which pieces of evidence influenced the model's decisions.
Katie Matton, Robert Osazuwa Ness, John V. Guttag, Emre Kiciman
ICLR1
2019 Into the Wild: Transitioning from Recognizing Mood in Clinical Interactions to Personal Conversations for Individuals with Bipolar Disorder
Katie Matton, Melvin G. McInnis, Emily Mower Provost
INTERSPEECH1