Cara Leckey

dblp:430/9043 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
0000-0002-7251-9827ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%
Human-computer interaction and pervasive computing
2 papers
Human-AI interaction · 100%
Artificial intelligence
1 paper
Knowledge representation and reasoning · 100%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%

Topics — the 2 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › question answering
scholarly question answering
2.022026
PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&A · CHI 2026
An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering Systems · CHI 2026
Visualization and visual analytics › information visualization › metadata visualization
provenance visualization
0.312026
PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&A · CHI 2026

Methods — techniques the papers use, named apart from their topics

within-subjects study · 3.0thematic analysis · 3.0contextual inquiry · 3.0
YearPublicationVenuePosition
2026 An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering Systems
abstract
Large Language Models (LLMs) are transforming scholarly tasks like search and summarization, but their reliability remains uncertain. Current evaluation metrics for testing LLM reliability are primarily automated approaches that prioritize efficiency and scalability, but lack contextual nuance and fail to reflect how scientific domain experts assess LLM outputs in practice. We developed and validated a schema for evaluating LLM errors in scholarly question-answering systems that reflects the assessment strategies of practicing scientists. In collaboration with domain experts, we identified 20 error patterns across seven categories through thematic analysis of 68 question-answer pairs. We validated this schema through contextual inquiries with 10 additional scientists, which showed not only which errors experts naturally identify but also how structured evaluation schemas can help them detect previously overlooked issues. Domain experts use systematic assessment strategies, including technical precision testing, value-based evaluation, and meta-evaluation of their own practices. We discuss implications for supporting expert evaluation of LLM outputs, including opportunities for personalized, schema-driven tools that adapt to individual evaluation patterns and expertise levels.
Anna Martin-Boyle, William Humphreys, Martha Brown, Cara Leckey, Harmanpreet Kaur
CHI4
2026 PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&A
abstract
Large language models (LLMs) are increasingly used in scholarly question-answering (QA) systems to help researchers synthesize vast amounts of literature. However, these systems often produce subtle errors (e.g., unsupported claims, errors of omission), and current provenance mechanisms like source citations are not granular enough for the rigorous verification that scholarly domain requires. To address this, we introduce PaperTrail, a novel interface that decomposes both LLM answers and source documents into discrete claims and evidence, mapping them to reveal supported assertions, unsupported claims, and information omitted from the source texts. We evaluated PaperTrail in a within-subjects study with 26 researchers who performed two scholarly editing tasks using PaperTrail and a baseline interface. Our results show that PaperTrail significantly lowered participants’ trust compared to the baseline. However, this increased caution did not translate to behavioral changes, as people continued to rely on LLM-generated scholarly edits to avoid a cognitively burdensome task. We discuss the value of claim-evidence matching for understanding LLM trustworthiness in scholarly settings, and present design implications for cognition-friendly communication of provenance information.
Anna Martin-Boyle, Cara Leckey, Martha Brown, Harmanpreet Kaur
CHI2