Tiffany Wenting Li

dblp:277/8198 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0002-0954-5627ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 7 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Supporting Learners' Use of Imperfect Generative Pedagogical Chatbots: The Role of Chatbot Response Uncertainty and Reduced Verbosity
abstract
Generative chatbots promise to scale personalized learning. Most publicly available generative chatbots are designed to provide confident and eloquent responses by default, even when hallucinating. Prior work has observed that learners using such chatbots often engage shallowly and fail to detect chatbot errors due to overtrust, cognitive overload, and prioritization of short-term gains. To address these challenges, this work examines two chatbot design options in a STEM learning context: introducing verbal uncertainty and reducing response verbosity. Using Bayesian causal inference and thematic analysis in a quasi-experimental setting, we found that a less verbose chatbot improved detection of errors with logical fallacies, but did not increase the use of alternative resources. A chatbot that always expressed uncertainty reduced the adoption of incorrect chatbot responses, but had mixed effects on learning outcomes, suggesting the need to increase signal credibility and maintain learners’ engagement in the learning process despite chatbot disuse.
Tiffany Wenting Li, Yifan Song 0007, Hari Sundaram, Karrie Karahalios
CHI1
2025 Organize, Then Vote: Exploring Cognitive Load in Quadratic Survey Interfaces
abstract
Quadratic Surveys (QSs) elicit more accurate preferences than traditional methods like Likert-scale surveys. However, the cognitive load associated with QSs has hindered their adoption in digital surveys for collective decision-making. We introduce a two-phase "organize-then-vote" QS to reduce cognitive load. As interface design significantly impacts survey results and accuracy, our design scaffolds survey takers' decision-making while managing the cognitive load imposed by QS. In a 2x2 between-subject in-lab study on public resource allotment, we compared our interface with a traditional text interface across a QS with 6 (short) and 24 (long) options. Two-phase interface participants spent more time per option and exhibited shorter voting edit distances. We qualitatively observed shifts in cognitive effort from mechanical operations to constructing more comprehensive preferences. We conclude that this interface promoted deeper engagement, potentially reducing satisficing behaviors caused by cognitive overload in longer QSs. This research clarifies how human-centered design improves preference elicitation tools for collective decision-making.
Ti-Chung Cheng, Yutong Zhang 0011, Yi-Hung Chou, Vinay Koshy, Tiffany Wenting Li, Karrie Karahalios, Hari Sundaram
CHI5
2025 Can Learners Navigate Imperfect Generative Pedagogical Chatbots? An Analysis of Chatbot Errors on Learning
abstract
Generative pedagogical chatbots offer a promising solution to transform personalized learning at scale, but their benefits are at risk because of the potential of providing inaccurate information. We have a limited understanding of how effectively learners handle factual chatbot errors and how these errors affect learners with varying backgrounds. This study addresses these questions in an ecologically valid open-ended online STEM learning environment. Using Bayesian causal inference and thematic analysis on survey and interview data from a quasi-experimental setting, we found that most participants struggled to detect factual errors even with access to reading materials and the Internet. Undetected errors harmed learning outcomes and self-efficacy, underscoring the need to help learners evaluate chatbot responses. By analyzing participants' evaluation strategies, we identified challenges during error management and suggested ideas on designing effective supporting resources and learner empowerment. Finally, we revealed differential impacts of chatbot errors across learners and called for personalized support and deployment.
Tiffany Wenting Li, Yifan Song 0007, Hari Sundaram, Karrie Karahalios
L@S1
2023 Inform the Uninformed: Improving Online Informed Consent Reading with an AI-Powered Chatbot
abstract
Informed consent is a core cornerstone of ethics in human subject research. Through the informed consent process, participants learn about the study procedure, benefits, risks, and more to make an informed decision. However, recent studies showed that current practices might lead to uninformed decisions and expose participants to unknown risks, especially in online studies. Without the researcher’s presence and guidance, online participants must read a lengthy form on their own with no answers to their questions. In this paper, we examined the role of an AI-powered chatbot in improving informed consent online. By comparing the chatbot with form-based interaction, we found the chatbot improved consent form reading, promoted participants’ feelings of agency, and closed the power gap between the participant and the researcher. Our exploratory analysis further revealed the altered power dynamic might eventually benefit study response quality. We discussed design implications for creating AI-powered chatbots to offer effective informed consent in broader settings.
Ziang Xiao, Tiffany Wenting Li, Karrie Karahalios, Hari Sundaram
CHI2
2023 Am I Wrong, or Is the Autograder Wrong? Effects of AI Grading Mistakes on Learning
abstract
Errors in AI grading and feedback often have an intractable set of causes and are, by their nature, difficult to completely avoid. Since inaccurate feedback potentially harms learning, there is a need for designs and workflows that mitigate these harms. To better understand the mechanisms by which erroneous AI feedback impacts students’ learning, we conducted surveys and interviews that recorded students’ interactions with a short-answer AI autograder for “Explain in Plain English” code reading problems. Using causal modeling, we inferred the learning impacts of wrong answers marked as right (false positives, FPs) and right answers marked as wrong (false negatives, FNs). We further explored explanations for the learning impacts, including errors influencing participants’ engagement with feedback and assessments of their answers’ correctness, and participants’ prior performance in the class.
Tiffany Wenting Li, Silas Hsu, Maxwell Fowler, Zhilin Zhang 0004, Craig B. Zilles, Karrie Karahalios
ICER (1)1
2021 Attitudes Surrounding an Imperfect AI Autograder
abstract
Deployment of AI assessment tools in education is widespread, but work on students’ interactions and attitudes towards imperfect autograders is comparatively lacking. This paper presents students’ perceptions surrounding a ∼ 90% accurate automated short-answer grader that determined homework and exam credit in a college-level computer science course. Using surveys and interviews, we investigated students’ knowledge about the autograder and their attitudes.
Silas Hsu, Tiffany Wenting Li, Zhilin Zhang 0004, Maxwell Fowler, Craig B. Zilles, Karrie Karahalios
CHI2
2021 "I can show what I really like.": Eliciting Preferences via Quadratic Voting
abstract
Surveys are a common instrument to gauge self-reported opinions from the crowd for scholars in the CSCW community, the social sciences, and many other research areas. Researchers often use surveys to prioritize a subset of given options when there are resource constraints. Over the past century, researchers have developed a wide range of surveying techniques, including one of the most popular instruments, the Likert ordinal scale, to elicit individual preferences. However, the challenge to elicit accurate and rich self-reported responses with surveys in a resource-constrained context still persists today. In this study, we examine Quadratic Voting (QV), a voting mechanism powered by the affordances of a modern computer and straddles ratings and rankings approaches, as an alternative online survey technique. We argue that QV could elicit more accurate self-reported responses compared to the Likert scale when the goal is to understand relative preferences under resource constraints. We conducted two randomized controlled experiments on Amazon Mechanical Turk, one in the context of public opinion polling and the other in a human-computer interaction user study. Based on our Bayesian analysis results, a QV survey with a sufficient amount of voice credits, aligned significantly closer to participants' incentive-compatible behaviors than a Likert scale survey, with a medium to high effect size. In addition, we extended QV's application scenario from typical public policy and education research to a problem setting familiar to the CSCW community: a prototypical HCI user study. Our experiment results, QV survey design, and QV interface serve as a stepping stone for CSCW researchers to further explore this surveying methodology in their studies and encourage decision-makers from other communities to consider QV as a promising alternative.
Ti-Chung Cheng, Tiffany Wenting Li, Yi-Hung Chou, Karrie Karahalios, Hari Sundaram
Proc. ACM Hum. Comput. Interact.2
2020 Erroneous Answers Categorization for Sketching Questions in Spatial Visualization Training
Tiffany Wenting Li, Luc Paquette
EDM1
2020 "It's all about conversation": Challenges and Concerns of Faculty and Students in the Arts, Humanities, and the Social Sciences about Education at Scale
abstract
As colleges and universities continue their commitment to increasing access to higher education through offering education online and at scale, attention on teaching open-ended subjects online and at scale, mainly the arts, humanities, and the social sciences, remains limited. While existing work in scaling open-ended courses primarily focuses on the evaluation and feedback of open-ended assignments, there is a lack of understanding of how to effectively teach open-ended, university-level courses at scale. To better understand the needs of teaching large-scale, open-ended courses online effectively in a university setting, we conducted a mixed-methods study with university instructors and students, using surveys and interviews, and identified five critical pedagogical elements that distinguish the teaching and learning experiences in an open-ended course from that in a non-open-ended course. An overarching theme for the five elements was the need to support students' self-expression. We further uncovered open challenges and opportunities when incorporating the five critical pedagogical elements into large-scale, open-ended courses online in a university setting, and suggested six future research directions: (1) facilitate in-depth conversations, (2) create a studio-friendly environment, (3) adapt to open-ended assessment, (4) scale individual open-ended feedback, (5) establish trust for self-expression, and (6) personalize instruction and harness the benefits of student diversity.
Tiffany Wenting Li, Karrie Karahalios, Hari Sundaram
Proc. ACM Hum. Comput. Interact.1