Diane J. Litman

dblp:l/DianeJLitman · DBLP profile ↗
← Back
161ranked-venue papers
31as first author
18since 2021 · last 2025
0000-0001-7282-7531ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 101 · 23 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 46 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 42 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 7 first-authorTheory of computation · 2Software engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2025 From Information to Insight: Leveraging LLMs for Open Aspect-Based Educational Summarization
abstract
This paper addresses the challenge of aspect-based summarization in education by introducing Reflective ASPect-based summarization (ReflectASP), a novel dataset that summarizes student reflections on STEM lectures. Despite the promising performance of large language models in general summarization, their application to nuanced aspect-based summaries remains under-explored. ReflectASP eases the exploration of open-aspect-based summarization (OABS), overcoming the limitations of current datasets and comes with ample human annotations. We benchmarked different types of zero-shot summarization methods and proposed two refinement methods to improve summaries, supported by both automatic and human manual evaluations. Additionally, we analyzed suggestions and revisions made during the refinement process, offering a fine-grained study of the editing strategies employed by these methods. We make our models, dataset, and all human evaluation results available at https://github.com/cs329yangzhong/ReflectASP.
Diane J. Litman
ACL (1)2
2025 Multi-party Lexical Alignment in Collaborative Learning with a Teachable Robot
Yuya Asano, Diane J. Litman, Paras Sharma, Daniel Fritsch, Quentin King-Shepard, Timothy Nokes-Malach, Adriana Kovashka, Erin Walker
AIED (6)2
2025 Beyond Static Measures: Temporal Analysis of Lexical Alignment in Human-Human Learning With a Teachable Robot
Paras Sharma, Daniel Fritsch, Yuya Asano, Quentin King-Shepard, Tyree Langley, Tristan Maidment, Diane J. Litman, Timothy Nokes-Malach, Adriana Kovashka, Nikki G. Lobczowski, Erin Walker
AIED (4)7
2025 Can LLMs simulate the same correct solutions to free-response math problems as real students?
abstract
Large language models (LLMs) have emerged as powerful tools for developing educational systems.While previous studies have explored modeling student mistakes, a critical gap remains in understanding whether LLMs can generate correct solutions that represent student responses to free-response problems.We compare the distribution of solutions from four LLMs (one proprietary, two open-sourced general, and one open-sourced math models) with various sampling and prompting techniques and those from students teaching math problems to a conversational robot.Our study reveals discrepancies between the correct solutions produced by LLMs and by students.We discuss the practical implications of these findings for the design and evaluation of LLMsupported educational systems.
Yuya Asano, Diane J. Litman, Erin Walker
EMNLP2
2025 Discourse-Driven Evaluation: Unveiling Factual Inconsistency in Long Document Summarization
abstract
Yang Zhong, Diane Litman. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Diane J. Litman
NAACL (Long Papers)2
2024 ReflectSumm: A Benchmark for Course Reflection Summarization
abstract
This paper introduces ReflectSumm, a novel summarization dataset specifically designed for summarizing students’ reflective writing. The goal of ReflectSumm is to facilitate developing and evaluating novel summarization techniques tailored to real-world scenarios with little training data, with potential implications in the opinion summarization domain in general and the educational domain in particular. The dataset encompasses a diverse range of summarization tasks and includes comprehensive metadata, enabling the exploration of various research questions and supporting different applications. To showcase its utility, we conducted extensive evaluations using multiple state-of-the-art baselines. The results provide benchmarks for facilitating further research in this area.
Mohamed Elaraby, Diane J. Litman, Ahmed Ashraf Butt, Muhsin Menekse
LREC/COLING3
2024 Enhancing Knowledge Retrieval with Topic Modeling for Knowledge-Grounded Dialogue
abstract
Knowledge retrieval is one of the major challenges in building a knowledge-grounded dialogue system. A common method is to use a neural retriever with a distributed approximate nearest-neighbor database to quickly find the relevant knowledge sentences. In this work, we propose an approach that utilizes topic modeling on the knowledge base to further improve retrieval accuracy and as a result, improve response generation. Additionally, we experiment with a large language model (LLM), ChatGPT, to take advantage of the improved retrieval performance to further improve the generation results. Experimental results on two datasets show that our approach can increase retrieval and generation performance. The results also indicate that ChatGPT is a better response generator for knowledge-grounded dialogue when relevant knowledge is provided.
Nhat Tran, Diane J. Litman
LREC/COLING2
2024 What metrics of participation balance predict outcomes of collaborative learning with a robot?
Yuya Asano, Diane J. Litman, Quentin King-Shepard, Tristan Maidment, Tyree Langley, Teresa Davison
EDM2
2024 Analyzing Large Language Models for Classroom Discussion Assessment
Nhat Tran, Benjamin Pierce, Diane J. Litman, Richard Correnti, Lindsay Clare Matsumura
EDM3
2022 An Automated Writing Evaluation System for Supporting Self-monitored Revising
Diane J. Litman, Tazin Afrin, Omid Kashefi, Christopher Olshefski, Amanda Godley, Rebecca Hwa
AIED (1)1
2022 Improving the Quality of Students' Written Reflections Using Natural Language Processing: Model Design and Classroom Evaluation
Ahmed Magooda, Diane J. Litman, Muhsin Menekse
AIED (1)2
2022 ArgLegalSumm: Improving Abstractive Summarization of Legal Documents with Argument Mining
abstract
A challenging task when generating summaries of legal documents is the ability to address their argumentative nature. We introduce a simple technique to capture the argumentative structure of legal documents by integrating argument role labeling into the summarization process. Experiments with pretrained language models show that our proposed approach improves performance over strong baselines.
Mohamed Elaraby, Diane J. Litman
COLING2
2022 Building a Reinforcement Learning Environment from Limited Data to Optimize Teachable Robot Interventions
Tristan Maidment, Mingzhi Yu, Nikki G. Lobczowski, Adriana Kovashka, Erin Walker, Diane J. Litman, Timothy Nokes-Malach
EDM6
2022 Comparison of Lexical Alignment with a Teachable Robot in Human-Robot and Human-Human-Robot Interactions
abstract
Yuya Asano, Diane Litman, Mingzhi Yu, Nikki Lobczowski, Timothy Nokes-Malach, Adriana Kovashka, Erin Walker. Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2022.
Yuya Asano, Diane J. Litman, Mingzhi Yu, Nikki G. Lobczowski, Timothy Nokes-Malach, Adriana Kovashka, Erin Walker
SIGDIAL2
2022 Getting Better Dialogue Context for Knowledge Identification by Leveraging Document-level Topic Shift
abstract
To build a goal-oriented dialogue system that can generate responses given a knowledge base, identifying the relevant pieces of information to be grounded in is vital.When the number of documents in the knowledge base is large, retrieval approaches are typically used to identify the top relevant documents.However, most prior work simply uses an entire dialogue history to guide retrieval, rather than exploiting a dialogue's topical structure.In this work, we examine the importance of building the proper contextualized dialogue history when document-level topic shifts are present.Our results suggest that excluding irrelevant turns from the dialogue history (e.g., excluding turns not grounded in the same document as the current turn) leads to better retrieval results.We also propose a cascading approach utilizing the topical nature of a knowledge-grounded conversation to further manipulate the dialogue history used as input to the retrieval models.
Nhat Tran, Diane J. Litman
SIGDIAL2
2022 How to Ask for Donations? Learning User-Specific Persuasive Dialogue Policies through Online Interactions
abstract
Persuasive conversations are more effective when they are custom-tailored for the intended audience. Current persuasive dialogue systems rely heavily on advice-giving or focus on different framing policies in a constrained and less dynamic/flexible manner. In this paper, we argue for a new approach, in which the system can identify optimal persuasive strategies in context and persuade users through online interactions. We study two main questions (1) can a reinforcement-learning-based dialogue framework learn to exercise user-specific communicative strategies for persuading users? (2) How can we leverage the crowd-sourcing platforms to collect data for training, and evaluating such frameworks for human-AI(/machine) conversations? We describe a prototype system that interacts with users with the goal of persuading them to donate to a charity and use experiments with crowd workers and analyses of our learned policies to document that our approach leads to learning context-sensitive persuasive strategies that focus on user’s reactions towards donation and contribute to increasing dialogue success.
Nhat Tran, Malihe Alikhani, Diane J. Litman
UMAP3
2021 A Fairness Evaluation of Automated Methods for Scoring Text Evidence Usage in Writing
Diane J. Litman, Haoran Zhang 0005, Richard Correnti, Lindsay Clare Matsumura, Elaine Wang 0001
AIED (1)1
2021 Effective Interfaces for Student-Driven Revision Sessions for Argumentative Writing
abstract
We present the design and evaluation of a web-based intelligent writing assistant that helps students recognize their revisions of argumentative essays. To understand how our revision assistant can best support students, we have implemented four versions of our system with differences in the unit span (sentence versus sub-sentence) of revision analysis and the level of feedback provided (none, binary, or detailed revision purpose categorization). We first discuss the design decisions behind relevant components of the system, then analyze the efficacy of the different versions through a Wizard of Oz study with university students. Our results show that while a simple interface with no revision feedback is easier to use, an interface that provides a detailed categorization of sentence-level revisions is the most helpful based on user survey data, as well as the most effective based on improvement in writing outcomes.
Tazin Afrin, Omid Kashefi, Christopher Olshefski, Diane J. Litman, Rebecca Hwa, Amanda Godley
CHI4
2020 Entrainment2Vec: Embedding Entrainment for Multi-Party Dialogues
abstract
Entrainment is the propensity of speakers to begin behaving like one another in conversation. While most entrainment studies have focused on dyadic interactions, researchers have also started to investigate multi-party conversations. In these studies, multi-party entrainment has typically been estimated by averaging the pairs' entrainment values or by averaging individuals' entrainment to the group. While such multi-party measures utilize the strength of dyadic entrainment, they have not yet exploited different aspects of the dynamics of entrainment relations in multi-party groups. In this paper, utilizing an existing pairwise asymmetric entrainment measure, we propose a novel graph-based vector representation of multi-party entrainment that incorporates both strength and dynamics of pairwise entrainment relations. The proposed kernel approach and weakly-supervised representation learning method show promising results at the downstream task of predicting team outcomes. Also, examining the embedding, we found interesting information about the dynamics of the entrainment relations. For example, teams with more influential members have more process conflict.
Zahra Rahimi, Diane J. Litman
AAAI2
2020 Automated Topical Component Extraction Using Neural Network Attention Scores from Source-based Essay Scoring
abstract
While automated essay scoring (AES) can reliably grade essays at scale, automated writing evaluation (AWE) additionally provides formative feedback to guide essay revision.However, a neural AES typically does not provide useful feature representations for supporting AWE.This paper presents a method for linking AWE and neural AES, by extracting Topical Components (TCs) representing evidence from a source text using the intermediate output of attention layers.We evaluate performance using a feature-based AES requiring TCs.Results show that performance is comparable whether using automatically or manually constructed TCs for 1) representing essays as rubric-based features, 2) grading essays.
Haoran Zhang 0005, Diane J. Litman
ACL2
2020 Contextual Argument Component Classification for Class Discussions
abstract
Argument mining systems often consider contextual information, i.e. information outside of an argumentative discourse unit, when trained to accomplish tasks such as argument component identification, classification, and relation extraction.However, prior work has not carefully analyzed the utility of different contextual properties in context-aware models.In this work, we show how two different types of contextual information, local discourse context and speaker context, can be incorporated into a computational model for classifying argument components in multi-party classroom discussions.We find that both context types can improve performance, although the improvements are dependent on context size and position.
Luca Lugini, Diane J. Litman
COLING2
2020 The Discussion Tracker Corpus of Collaborative Argumentation
abstract
Although NLP research on argument mining has advanced considerably in recent years, most studies draw on corpora of asynchronous and written texts, often produced by individuals. Few published corpora of synchronous, multi-party argumentation are available. The Discussion Tracker corpus, collected in high school English classes, is an annotated dataset of transcripts of spoken, multi-party argumentation. The corpus consists of 29 multi-party discussions of English literature transcribed from 985 minutes of audio. The transcripts were annotated for three dimensions of collaborative argumentation: argument moves (claims, evidence, and explanations), specificity (low, medium, high) and collaboration (e.g., extensions of and disagreements about others’ ideas). In addition to providing descriptive statistics on the corpus, we provide performance benchmarks and associated code for predicting each dimension separately, illustrate the use of the multiple annotations in the corpus to improve performance via multi-task learning, and finally discuss other ways the corpus might be used to further NLP research.
Christopher Olshefski, Luca Lugini, Ravneet Singh, Diane J. Litman, Amanda Godley
LREC4
2019 eRevise: Using Natural Language Processing to Provide Formative Feedback on Text Evidence Usage in Student Writing
abstract
Writing a good essay typically involves students revising an initial paper draft after receiving feedback. We present eRevise, a web-based writing and revising environment that uses natural language processing features generated for rubricbased essay scoring to trigger formative feedback messages regarding students’ use of evidence in response-to-text writing. By helping students understand the criteria for using text evidence during writing, eRevise empowers students to better revise their paper drafts. In a pilot deployment of eRevise in 7 classrooms spanning grades 5 and 6, the quality of text evidence usage in writing improved after students received formative feedback then engaged in paper revision.
Haoran Zhang 0005, Ahmed Magooda, Diane J. Litman, Richard Correnti, Elaine Wang 0001, Lindsay Clare Matsumura, Emily Howe, Rafael Quintana
AAAI3
2019 Identifying Editor Roles in Argumentative Writing from Student Revision Histories
Tazin Afrin, Diane J. Litman
AIED (2)2
2019 Identifying Personality Traits Using Overlap Dynamics in Multiparty Dialogue
abstract
Research on human spoken language has shown that speech plays an important role in identifying speaker personality traits.In this work, we propose an approach for identifying speaker personality traits using overlap dynamics in multiparty spoken dialogues.We first define a set of novel features representing the overlap dynamics of each speaker.We then investigate the impact of speaker personality traits on these features using ANOVA tests.We find that features of overlap dynamics significantly vary for speakers with different levels of both Extraversion and Conscientiousness.Finally, we find that classifiers using only overlap dynamics features outperform random guessing in identifying Extraversion and Agreeableness, and that the improvements are statistically significant.
Mingzhi Yu, Emer Gilmartin, Diane J. Litman
INTERSPEECH3
2018 Argument Mining for Improving the Automated Scoring of Persuasive Essays
abstract
End-to-end argument mining has enabled the development of new automated essay scoring (AES) systems that use argumentative features (e.g., number of claims, number of support relations) in addition to traditional legacy features (e.g., grammar, discourse structure) when scoring persuasive essays. While prior research has proposed different argumentative features as well as empirically demonstrated their utility for AES, these studies have all had important limitations. In this paper we identify a set of desiderata for evaluating the use of argument mining for AES, introduce an end-to-end argument mining system and associated argumentative feature sets, and present the results of several studies that both satisfy the desiderata and demonstrate the value-added of argument mining for scoring persuasive essays.
Huy V. Nguyen, Diane J. Litman
AAAI2
2018 Weighting Model Based on Group Dynamics to Measure Convergence in Multi-party Dialogue
abstract
This paper proposes a new weighting method for extending a dyad-level measure of convergence to multi-party dialogues by considering group dynamics instead of simply averaging.Experiments indicate the usefulness of the proposed weighted measure and also show that in general a proper weighting of the dyadlevel measures performs better than nonweighted averaging in multiple tasks.
Zahra Rahimi, Diane J. Litman
SIGDIAL Conference2
2018 A novel ILP framework for summarizing content with high lexical variety
abstract
Abstract Summarizing content contributed by individuals can be challenging, because people make different lexical choices even when describing the same events. However, there remains a significant need to summarize such content. Examples include the student responses to post-class reflective questions, product reviews, and news articles published by different news agencies related to the same events. High lexical diversity of these documents hinders the system’s ability to effectively identify salient content and reduce summary redundancy. In this paper, we overcome this issue by introducing an integer linear programming-based summarization framework. It incorporates a low-rank approximation to the sentence-word cooccurrence matrix to intrinsically group semantically similar lexical items. We conduct extensive experiments on datasets of student responses, product reviews, and news documents. Our approach compares favorably to a number of extractive baselines as well as a neural abstractive summarization system. The paper finally sheds light on when and why the proposed framework is effective at summarizing content with high lexical variety.
Wencan Luo, Fei Liu 0004, Zitao Liu 0003, Diane J. Litman
Nat. Lang. Eng.4
2017 Using Discourse Signals for Robust Instructor Intervention Prediction
abstract
We tackle the prediction of instructor intervention in student posts from discussion forums in Massive Open Online Courses (MOOCs). Our key finding is that using automatically obtained discourse relations improves the prediction of when instructors intervene in student discussions, when compared with a state-of-the-art, feature-rich baseline. Our supervised classifier makes use of an automatic discourse parser which outputs Penn Discourse Treebank (PDTB) tags that represent in-post discourse features. We show PDTB relation-based features increase the robustness of the classifier and complement baseline features in recalling more diverse instructor intervention patterns. In comprehensive experiments over 14 MOOC offerings from several disciplines, the PDTB discourse features improve performance on average. The resultant models are less dependent on domain-specific vocabulary, allowing them to better generalize to new courses.
Muthu Kumar Chandrasekaran, Carrie Demmans Epp, Min-Yen Kan, Diane J. Litman
AAAI4
2017 A Corpus of Annotated Revisions for Studying Argumentative Writing
abstract
This paper presents ArgRewrite, a corpus of between-draft revisions of argumentative essays.Drafts are manually aligned at the sentence level, and the writer's purpose for each revision is annotated with categories analogous to those used in argument mining and discourse analysis.The corpus should enable advanced research in writing comparison and revision analysis, as demonstrated via our own studies of student revision behavior and of automatic revision purpose prediction.
Fan Zhang 0095, Homa B. Hashemi, Rebecca Hwa, Diane J. Litman
ACL (1)4
2017 Entrainment in Multi-Party Spoken Dialogues at Multiple Linguistic Levels
Zahra Rahimi, Diane J. Litman, Susannah B. F. Paletz, Mingzhi Yu
INTERSPEECH3
2017 Scaling Reflection Prompts in Large Classrooms via Mobile Interfaces and Natural Language Processing
abstract
We present the iterative design, prototype, and evaluation of CourseMIRROR (Mobile In-situ Reflections and Review with Optimized Rubrics), an intelligent mobile learning system that uses natural language processing (NLP) techniques to enhance instructor-student interactions in large classrooms. CourseMIRROR enables streamlined and scaffolded reflection prompts by: 1) reminding and collecting students' in-situ written reflections after each lecture; 2) continuously monitoring the quality of a student's reflection at composition time and generating helpful feedback to scaffold reflection writing; and 3) summarizing the reflections and presenting the most significant ones to both instructors and students. Through a combination of a 60-participant lab study and eight semester-long deployments involving 317 students, we found that the reflection and feedback cycle enabled by CourseMIRROR is beneficial to both instructors and students. Furthermore, the reflection quality feedback feature can encourage students to compose more specific and higher-quality reflections, and the algorithms in CourseMIRROR are both robust to cold start and scalable to STEM courses in diverse topics.
Xiangmin Fan, Wencan Luo, Muhsin Menekse, Diane J. Litman
IUI4
2016 Natural Language Processing for Enhancing Teaching and Learning
abstract
Advances in natural language processing (NLP) and educational technology, as well as the availability of unprecedented amounts of educationally-relevant text and speech data, have led to an increasing interest in using NLP to address the needs of teachers and students. Educational applications differ in many ways, however, from the types of applications for which NLP systems are typically developed. This paper will organize and give an overview of research in this area, focusing on opportunities as well as challenges.
Diane J. Litman
AAAI1
2016 Context-aware Argumentative Relation Mining
abstract
Context is crucial for identifying argumentative relations in text, but many argument mining methods make little use of contextual features.This paper presents contextaware argumentative relation mining that uses features extracted from writing topics as well as from windows of context sentences.Experiments on student essays demonstrate that the proposed features improve predictive performance in two argumentative relation classification tasks.
Huy Nguyen 0003, Diane J. Litman
ACL (1)2
2016 An Improved Phrase-based Approach to Annotating and Summarizing Student Course Responses
abstract
Teaching large classes remains a great challenge, primarily because it is difficult to attend to all the student needs in a timely manner. Automatic text summarization systems can be leveraged to summarize the student feedback, submitted immediately after each lecture, but it is left to be discovered what makes a good summary for student responses. In this work we explore a new methodology that effectively extracts summary phrases from the student responses. Each phrase is tagged with the number of students who raise the issue. The phrases are evaluated along two dimensions: with respect to text content, they should be informative and well-formed, measured by the ROUGE metric; additionally, they shall attend to the most pressing student needs, measured by a newly proposed metric. This work is enabled by a phrase-based annotation and highlighting scheme, which is new to the summarization task. The phrase-based framework allows us to summarize the student responses into a set of bullet points and present to the instructor promptly.
Wencan Luo, Fei Liu 0004, Diane J. Litman
COLING3
2016 Inferring Discourse Relations from PDTB-style Discourse Labels for Argumentative Revision Classification
abstract
Penn Discourse Treebank (PDTB)-style annotation focuses on labeling local discourse relations between text spans and typically ignores larger discourse contexts. In this paper we propose two approaches to infer discourse relations in a paragraph-level context from annotated PDTB labels. We investigate the utility of inferring such discourse information using the task of revision classification. Experimental results demonstrate that the inferred information can significantly improve classification performance compared to baselines, not only when PDTB annotation comes from humans but also from automatic parsers.
Fan Zhang 0095, Diane J. Litman, Katherine Forbes-Riley
COLING2
2016 The Teams Corpus and Entrainment in Multi-Party Spoken Dialogues
abstract
When interacting individuals entrain, they begin to speak more like each other.To support research on entrainment in cooperative multi-party dialogues, we have created a corpus where teams of three or four speakers play two rounds of a cooperative board game.We describe the experimental design and technical infrastructure used to collect our corpus, which consists of audio, video, transcriptions, and questionnaire data for 63 teams (47 hours of audio).We illustrate the use of our corpus as a novel resource for studying team entrainment by 1) developing and evaluating teamlevel acoustic-prosodic entrainment measures that extend existing dyad measures, and 2) investigating relationships between team entrainment and participation dominance.
Diane J. Litman, Susannah B. F. Paletz, Zahra Rahimi, Stefani Allegretti, Caitlin Rice
EMNLP1
2016 Automatic Summarization of Student Course Feedback
abstract
Student course feedback is generated daily in both classrooms and online course discussion forums.Traditionally, instructors manually analyze these responses in a costly manner.In this work, we propose a new approach to summarizing student course feedback based on the integer linear programming (ILP) framework.Our approach allows different student responses to share co-occurrence statistics and alleviates sparsity issues.Experimental results on a student feedback corpus show that our approach outperforms a range of baselines in terms of both ROUGE scores and human evaluation.
Wencan Luo, Fei Liu 0004, Zitao Liu 0003, Diane J. Litman
HLT-NAACL4
2016 Using Context to Predict the Purpose of Argumentative Writing Revisions
abstract
While there is increasing interest in automatically recognizing the argumentative structure of a text, recognizing the argumentative purpose of revisions to such texts has been less explored.Furthermore, existing revision classification approaches typically ignore contextual information.We propose two approaches for utilizing contextual information when predicting argumentative revision purposes: developing contextual features for use in the classification paradigm of prior work, and transforming the classification problem to a sequence labeling task.Experimental results using two corpora of student essays demonstrate the utility of contextual information for predicting argumentative revision purposes.
Fan Zhang 0095, Diane J. Litman
HLT-NAACL2
2016 Extracting PDTB Discourse Relations from Student Essays
abstract
We investigate the manual and automatic annotation of PDTB discourse relations in student essays, a novel domain that is not only learning-based and argumentative, but also noisy with surface errors and deeper coherency issues.We discuss methodological complexities it poses for the task.We present descriptive statistics and compare relation distributions in related corpora.We compare automatic discourse parsing performance to prior work.
Katherine Forbes-Riley, Fan Zhang 0095, Diane J. Litman
SIGDIAL Conference3
2016 Towards Using Conversations with Spoken Dialogue Systems in the Automated Assessment of Non-Native Speakers of English
abstract
Diane Litman, Steve Young, Mark Gales, Kate Knill, Karen Ottewell, Rogier van Dalen, David Vandyke. Proceedings of the 17th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2016.
Diane J. Litman, Steve J. Young, Mark J. F. Gales, Kate M. Knill, Karen Ottewell, Rogier C. van Dalen, David Vandyke
SIGDIAL Conference1
2015 Summarizing Student Responses to Reflection Prompts
abstract
We propose to automatically summarize student responses to reflection prompts and introduce a novel summarization algo-rithm that differs from traditional methods in several ways. First, since the linguis-tic units of student inputs range from sin-gle words to multiple sentences, our sum-maries are created from extracted phrases rather than from sentences. Second, the phrase summarization algorithm ranks the phrases by the number of students who semantically mention a phrase in a sum-mary. Experimental results show that the proposed phrase summarization ap-proach achieves significantly better sum-marization performance on an engineering course corpus in terms of ROUGE scores when compared to other summarization methods, including MEAD, LexRank and MMR. 1
Wencan Luo, Diane J. Litman
EMNLP2
2015 Enhancing Instructor-Student and Student-Student Interactions with Mobile Interfaces and Summarization
abstract
Wencan Luo, Xiangmin Fan, Muhsin Menekse, Jingtao Wang, Diane Litman. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations. 2015.
Wencan Luo, Xiangmin Fan, Muhsin Menekse, Diane J. Litman
HLT-NAACL5
2014 Empirical analysis of exploiting review helpfulness for extractive summarization of online reviews
Wenting Xiong, Diane J. Litman
COLING2
2014 Identifying Thesis and Conclusion Statements in Student Essays to Scaffold Peer Review
Mohammad Hassan Falakmasir, Kevin D. Ashley, Christian D. Schunn, Diane J. Litman
Intelligent Tutoring Systems4
2014 Modeling Student Benefit from Illustrations and Graphs
Michael Lipschultz, Diane J. Litman
Intelligent Tutoring Systems2
2014 Classroom Evaluation of a Scaffolding Intervention for Improving Peer Review Localization
Huy Nguyen 0003, Wenting Xiong, Diane J. Litman
Intelligent Tutoring Systems3
2014 Automatic Scoring of an Analytical Response-To-Text Assessment
Zahra Rahimi, Diane J. Litman, Richard Correnti, Lindsay Clare Matsumura, Elaine Wang 0001, Zahid Kisa
Intelligent Tutoring Systems2
2014 Evaluating a Spoken Dialogue System that Detects and Adapts to User Affective States
abstract
We present an evaluation of a spoken dialogue system that detects and adapts to user disengagement and uncertainty in real-time. We compare this version of our system to a version that adapts to only user disengagement, and to a version that ig-nores user disengagement and uncertainty entirely. We find a significant increase in task success when comparing both affect-adaptive versions of our system to our non-adaptive baseline, but only for male users. 1
Diane J. Litman, Katherine Forbes-Riley
SIGDIAL Conference1
2013 Interactive Event: The Rimac Tutor - A Simulation of the Highly Interactive Nature of Human Tutorial Dialogue
Pamela W. Jordan, Patricia L. Albacete, Michael Ford, Sandra Katz, Michael Lipschultz, Diane J. Litman, Scott Silliman, Christine Wilson
AIED6
2013 Pilot Test of a Natural-Language Tutoring System for Physics That Simulates the Highly Interactive Nature of Human Tutoring
Sandra Katz, Patricia L. Albacete, Michael Ford, Pamela W. Jordan, Michael Lipschultz, Diane J. Litman, Scott Silliman, Christine Wilson
AIED6
2013 Illustrations or Graphs: Some Students Benefit from One over the Other
Michael Lipschultz, Diane J. Litman
AIED2
2013 Identifying Localization in Peer Reviews of Argument Diagrams
Huy V. Nguyen, Diane J. Litman
AIED2
2013 Predicting Low vs. High Disparity between Peer and Expert Ratings in Peer Reviews of Physics Lab Reports
Huy V. Nguyen, Diane J. Litman
AIED2
2013 Prosodic Entrainment and Tutoring Dialogue Success
Jesse D. Thomason, Huy V. Nguyen, Diane J. Litman
AIED3
2013 Evaluating Topic-Word Review Analysis for Understanding Student Peer Review Performance
Wenting Xiong, Diane J. Litman
EDM2
2013 Reducing Annotation Effort on Unbalanced Corpus based on Cost Matrix
Wencan Luo, Diane J. Litman, Joel Chan
HLT-NAACL2
2013 Differences in User Responses to a Wizard-of-Oz versus Automated System
Jesse D. Thomason, Diane J. Litman
HLT-NAACL2
2012 Prosodic Cues to Disengagement and Uncertainty in Physics Tutorial Dialogues
abstract
This paper focuses on the analysis and prediction of student disengagement and uncertainty, using a corpus of dialogues collected with a spoken tutorial dialogue system in the STEM domain of qualitative physics. We first compare and contrast the prosodic characteristics of dialogue turns exhibiting disengagement or not, and those exhibiting uncertainty or not. We then compare the utility of using multiple prosodic features to predict both disengagement and uncertainty. Index Terms: spoken dialogue systems, educational applications, emotion detection, prosody
Diane J. Litman, Heather Friedberg, Katherine Forbes-Riley
INTERSPEECH1
2012 Intrinsic and Extrinsic Evaluation of an Automatic User Disengagement Detector for an Uncertainty-Adaptive Spoken Dialogue System
Katherine Forbes-Riley, Diane J. Litman, Heather Friedberg, Joanna Drummond
HLT-NAACL2
2012 Adapting to Multiple Affective States in Spoken Dialogue
Katherine Forbes-Riley, Diane J. Litman
SIGDIAL Conference2
2012 Cohesion, Entrainment and Task Success in Educational Dialog
Diane J. Litman
SIGDIAL Conference1
2012 Lexical entrainment and success in student engineering groups
abstract
Lexical entrainment is a measure of how the words that speakers use in a conversation become more similar over time. In this paper, we propose a measure of lexical entrainment for multi-party speaking situations. We apply this score to a corpus of student engineering groups using high-frequency words and project words, and investigate the relationship between lexical entrainment and group success on a class project. Our initial findings show that, using the entrainment score with project-related words, there is a significant difference between the lexical entrainment of high performing groups, which tended to increase with time, and the entrainment for low performing groups, which tended to decrease with time.
Heather Friedberg, Diane J. Litman, Susannah B. F. Paletz
SLT2
2011 When Does Disengagement Correlate with Learning in Spoken Dialog Computer Tutoring?
Katherine Forbes-Riley, Diane J. Litman
AIED2
2011 Cohesion / Knowledge Interactions in Post-tutoring Reflective Text
Arthur Ward, Diane J. Litman
AIED2
2011 Examining the Impacts of Dialogue Content and System Automation on Affect Models in a Spoken Tutorial Dialogue System
Joanna Drummond, Diane J. Litman
SIGDIAL Conference2
2011 Using Performance Trajectories to Analyze the Immediate Impact of User State Misclassification in an Adaptive Spoken Dialogue System
Katherine Forbes-Riley, Diane J. Litman
SIGDIAL Conference2
2011 Designing and evaluating a wizarded uncertainty-adaptive spoken dialogue tutoring system
Katherine Forbes-Riley, Diane J. Litman
Comput. Speech Lang.2
2011 Assessing user simulation for dialog systems using human judges and automatic evaluation measures
abstract
Abstract While different user simulations are built to assist dialog system development, there is an increasing need to quickly assess the quality of the user simulations reliably. Previous studies have proposed several automatic evaluation measures for this purpose. However, the validity of these evaluation measures has not been fully proven. We present an assessment study in which human judgments are collected on user simulation qualities as the gold standard to validate automatic evaluation measures. We show that a ranking model can be built using the automatic measures to predict the rankings of the simulations in the same order as the human judgments. We further show that the ranking model can be improved by using a simple feature that utilizes time-series analysis.
Hua Ai, Diane J. Litman
Nat. Lang. Eng.2
2011 Benefits and challenges of real-time uncertainty detection and adaptation in a spoken dialogue computer tutor
Katherine Forbes-Riley, Diane J. Litman
Speech Commun.2
2011 Empirically evaluating the application of reinforcement learning to the induction of effective and adaptive pedagogical strategies
Min Chi, Kurt VanLehn, Diane J. Litman, Pamela W. Jordan
User Model. User Adapt. Interact.3
2010 Assessing Reviewer's Performance Based on Mining Problem Localization in Peer-Review Data
Wenting Xiong, Diane J. Litman, Christian D. Schunn
EDM2
2010 Do Micro-Level Tutorial Decisions Matter: Applying Reinforcement Learning to Induce Pedagogical Tutorial Tactics
Min Chi, Kurt VanLehn, Diane J. Litman
Intelligent Tutoring Systems (1)3
2010 In the Zone: Towards Detecting Student Zoning Out Using Supervised Machine Learning
Joanna Drummond, Diane J. Litman
Intelligent Tutoring Systems (2)2
2010 Metacognition and Learning in Spoken Dialogue Computer Tutoring
Katherine Forbes-Riley, Diane J. Litman
Intelligent Tutoring Systems (1)2
2010 Correcting Scientific Knowledge in a General-Purpose Ontology
Michael Lipschultz, Diane J. Litman
Intelligent Tutoring Systems (2)2
2010 Identifying Problem Localization in Peer-Review Feedback
Wenting Xiong, Diane J. Litman
Intelligent Tutoring Systems (2)2
2010 Inducing Effective Pedagogical Strategies Using Learning Context Features
Min Chi, Kurt VanLehn, Diane J. Litman, Pamela W. Jordan
UMAP3
2009 Setting Up User Action Probabilities in User Simulations for Dialog System Development
Hua Ai, Diane J. Litman
ACL/IJCNLP2
2009 To Elicit Or To Tell: Does It Matter?
abstract
While high interactivity has been one of the main characteristics of one-on-one human tutoring, a great deal of controversy surrounds the issue of whether interactivity is indeed the key feature of tutorial dialogue that impacts students' learning results. There are two commonly held hypotheses regarding the issue: a widely-believed monotonic interactivity hypothesis and a better supported interaction plateau hypothesis. The former hypothesis predicts increasing in interactivity causes an increase in learning while the latter states that increasing interactivity yields increasing learning until it hits a plateau, and further increases in interactivity do not cause noticeably increase in learning. In this study, we proposed the tactical interaction hypothesis which predicts beyond a certain level of interactivity, further increases in interactivity do not cause increase in learning unless they are guided by effective tutorial tactics. Overall our results support this hypothesis. However, finding effective tactics is not easy. This paper sheds some light on how to apply Reinforcement Learning to derive effective tutorial tactics.
Min Chi, Pamela W. Jordan, Kurt VanLehn, Diane J. Litman
AIED4
2009 Adapting to Student Uncertainty Improves Tutoring Dialogues
abstract
This study shows that affect-adaptive computer tutoring can significantly improve performance on learning efficiency and user satisfaction. We compare two different student uncertainty adaptations which were designed, implemented and evaluated in a controlled experiment using four versions of a wizarded spoken dialogue tutoring system: two adaptive systems used in two experimental conditions (basic and empirical), and two non-adaptive systems used in two control conditions (normal and random). In prior work we compared learning gains across the four systems; here we compare two other important performance metrics: learning efficiency and user satisfaction. We show that the basic adaptive system outperforms the normal (non-adaptive) and empirical (adaptive) systems in terms of learning efficiency. We also show that the empirical (adaptive) and random (non-adaptive) systems outperform the basic adaptive system in terms of user perception of tutor response quality. However, only the basic adaptive system shows a positive correlation between learning and user perception of decreased uncertainty.
Katherine Forbes-Riley, Diane J. Litman
AIED2
2009 Evidence of Misunderstandings in Tutorial Dialogue and their Impact on Learning
abstract
We explore the frequency and impact of misunderstandings in an existing corpus of tutorial dialogues in which a student appears to get an interpretation that is not in line with what the system developers intended. We found that this type of error is frequent, regardless of whether student input is typed or spoken, and that it does not respond well to general misconception repair strategies. Further we found that it is feasible to detect misunderstandings and suggest alternative strategies for repairing them that we intend to test in the future.
Pamela W. Jordan, Diane J. Litman, Michael Lipschultz, Joanna Drummond
AIED2
2009 Using Natural Language Processing to Analyze Tutorial Dialogue Corpora Across Domains Modalities
abstract
Our research goal is to investigate whether previous findings and methods in the area of tutorial dialogue can be generalized across dialogue corpora that differ in domain (mechanics versus electricity in physics), modality (spoken versus typed), and tutor type (computer versus human). We first present methods for unifying our prior coding and analysis methods. We then show that many of our prior findings regarding student dialogue behaviors and learning not only generalize across corpora, but that our methodology yields additional new findings. Finally, we show that natural language processing can be used to automate some of these analyses.
Diane J. Litman, Johanna D. Moore, Myroslava O. Dzikovska, Elaine Farrow
AIED1
2009 A user modeling-based performance analysis of a wizarded uncertainty-adaptive dialogue system corpus
abstract
Motivated by prior spoken dialogue system research in user modeling, we analyze interactions between performance and user class in a dataset previously collected with two wizarded spoken dialogue tutoring systems that adapt to user uncertainty. We focus on user classes defined by expertise level and gender, and on both objective (learning) and subjective (user satisfaction) performance metrics. We find that lower expertise users learn best from one adaptive system but prefer the other, while higher expertise users learned more from one adaptive system but didn't prefer either. Female users both learn best from and prefer the same adaptive system, while males preferred one adaptive system but didn't learn more from either. Our results yield an empirical basis for future investigations into whether adaptive system performance can improve by adapting to user uncertainty differently based on user class. Copyright © 2009 ISCA.
Katherine Forbes-Riley, Diane J. Litman
INTERSPEECH2
2009 Classifying turn-level uncertainty using word-level prosody
abstract
Spoken dialogue researchers often use supervised machine learning to classify turn-level user affect from a set of turn-level features. The utility of sub-turn features has been less explored, due to the complications introduced by associating a variable number of sub-turn units with a single turn-level classification. We present and evaluate several voting methods for using wordlevel pitch and energy features to classify turn-level user uncertainty in spoken dialogue data. Our results show that when linguistic knowledge regarding prosody and word position is introduced into a word-level voting model, classification accuracy is significantly improved compared to the use of both turn-level and uninformed word-level models. Index Terms: emotion recognition, speech dialogue systems 1.
Diane J. Litman, Mihai Rotaru 0002, Greg Nicholas
INTERSPEECH1
2009 Spoken Tutorial Dialogue and the Feeling of Another's Knowing
Diane J. Litman, Katherine Forbes-Riley
SIGDIAL Conference1
2009 Discourse Structure and Performance Analysis: Beyond the Correlation
Mihai Rotaru 0002, Diane J. Litman
SIGDIAL Conference2
2008 Assessing Dialog System User Simulation Evaluation Measures Using Human Judges
Hua Ai, Diane J. Litman
ACL2
2008 Responding to Student Uncertainty During Computer Tutoring: An Experimental Evaluation
Katherine Forbes-Riley, Diane J. Litman, Mihai Rotaru 0002
Intelligent Tutoring Systems2
2008 Minimal Feedback During Tutorial Dialogue
Pamela W. Jordan, Diane J. Litman
Intelligent Tutoring Systems2
2008 Semantic Cohesion and Learning
Arthur Ward, Diane J. Litman
Intelligent Tutoring Systems2
2008 Uncertainty Corpus: Resource to Study User Affect in Complex Spoken Dialogue Systems
Katherine Forbes-Riley, Diane J. Litman, Scott Silliman, Amruta Purandare
LREC2
2008 A Reinforcement Learning approach to evaluating state representations in spoken dialogue systems
Joel R. Tetreault, Diane J. Litman
Speech Commun.2
2008 The relative impact of student affect on performance models in a spoken dialogue tutoring system
Katherine Forbes-Riley, Mihai Rotaru 0002, Diane J. Litman
User Model. User Adapt. Interact.3
2007 Investigating Human Tutor Responses to Student Uncertainty for Adaptive System Development
Katherine Forbes-Riley, Diane J. Litman
ACII2
2007 The Utility of a Graphical Representation of Discourse Structure in Spoken Dialogue Systems
Mihai Rotaru 0002, Diane J. Litman
ACL2
2007 Comparing Linguistic Features for Modeling Learning in Computer Tutoring
Katherine Forbes-Riley, Diane J. Litman, Amruta Purandare, Mihai Rotaru 0002, Joel R. Tetreault
AIED2
2007 Dialog Convergence and Learning
Arthur Ward, Diane J. Litman
AIED2
2007 Knowledge consistent user simulations for dialog systems
abstract
We propose a novel model to simulate user knowledge consistency in tutoring dialogs, where no clear user goal can be defined. We also propose a new evaluation measure of knowledge consistency based on learning curves. We compare our new simulation model to real users as well as to a previously used simulation model. We show that the new model performs similarly to the real students and to the previous model when evaluated on high-level dialog features. The new model outperforms the previous model when measured on knowledge consistency. Index Terms: spoken dialog, user simulation, evaluation measures, knowledge consistency
Hua Ai, Diane J. Litman
INTERSPEECH2
2007 Estimating the Reliability of MDP Policies: a Confidence Interval Approach
Joel R. Tetreault, Dan Bohus, Diane J. Litman
HLT-NAACL3
2006 Dependencies between Student State and Speech Recognition Problems in Spoken Tutoring Dialogues
abstract
Speech recognition problems are a reality in current spoken dialogue systems. In order to better understand these phenomena, we study dependencies between speech recognition problems and several higher level dialogue factors that define our notion of student state: frustration/anger, certainty and correctness. We apply Chi Square (X2) analysis to a corpus of speech-based computer tutoring dialogues to discover these dependencies both within and across turns. Significant dependencies are combined to produce interesting insights regarding speech recognition problems and to propose new strategies for handling these problems. We also find that tutoring, as a new domain for speech applications, exhibits interesting tradeoffs and new factors to consider for spoken dialogue design.
Mihai Rotaru 0002, Diane J. Litman
ACL2
2006 Using Reinforcement Learning to Build a Better Model of Dialogue State
Joel R. Tetreault, Diane J. Litman
EACL2
2006 Humor: Prosody Analysis and Automatic Recognition for F*R*I*E*N*D*S*
Amruta Purandare, Diane J. Litman
EMNLP2
2006 Exploiting Discourse Structure for Spoken Dialogue Performance Analysis
Mihai Rotaru 0002, Diane J. Litman
EMNLP2
2006 Using system and user performance features to improve emotion detection in spoken tutoring dialogs
abstract
In this study, we incorporate automatically obtained system/user performance features into machine learning experiments to detect student emotion in computer tutoring dialogs. Our results show a relative improvement of 2.7% on classification accuracy and 8.08% on Kappa over using standard lexical, prosodie, sequential, and identification features. This level of improvement is comparable to the performance improvement shown in previous studies by applying dialog acts or lexical/prosodic-/discourse- level contextual features.
Hua Ai, Diane J. Litman, Katherine Forbes-Riley, Mihai Rotaru 0002, Joel R. Tetreault, Amruta Purandare
INTERSPEECH2
2006 Identification of confusion and surprise in spoken dialog using prosodic features
abstract
Sensitivity to a user's emotional state offers promise in improving the state of the art in spoken dialog systems. In this work, we attempt to detect the speaker's states of confusion and surprise using prosodic features from his/her utterances. We have collected a corpus of utterances in realistic settings using an experimental methodology aimed at eliciting confusion and surprise from users. Classification experiments have yielded up to a 27.2% improvement over baseline performance using F0 and power features. We achieved the greatest success at classification of emotions that were most successfully elicited.
Rohit Kumar 0001, Carolyn P. Rosé, Diane J. Litman
INTERSPEECH3
2006 Discourse structure and speech recognition problems
abstract
We study dependencies between discourse structure and speech recognition problems (SRP) in a corpus of speech-based computer tutoring dialogues. This analysis can inform us whether there are places in the discourse structure prone to more SRP. We automatically extract the discourse structure by taking advantage of how the tutoring information is encoded in our system. To quantify the discourse structure, we extract two features for each system turn: depth of the turn in the discourse structure and the type of transition from the previous turn to the current turn. The �$ 2 test is used to find significant dependencies. We find several interesting interactions which suggest that the discourse structure can play an important role in several dialogue related tasks: automatic detection of SRP and analyzing spoken dialogues systems with a large state space from limited amounts of available data. Index Terms: discourse structure, speech recognition analysis, spoken dialogue systems.
Mihai Rotaru 0002, Diane J. Litman
INTERSPEECH2
2006 Modelling User Satisfaction and Student Learning in a Spoken Dialogue Tutoring System with Generic, Tutoring, and User Affect Parameters
Katherine Forbes-Riley, Diane J. Litman
HLT-NAACL2
2006 Comparing the Utility of State Features in Spoken Dialogue Using Reinforcement Learning
Joel R. Tetreault, Diane J. Litman
HLT-NAACL2
2006 Exploiting Word-level Features for Emotion Prediction
abstract
In this paper we study two techniques for combining word-level features for emotion prediction. Prior research has primarily focused on the use of turn-level features as predictors. Recently, the utility of word-level features has been highlighted but only tested on relatively small human- computer corpora. We extend over previous work by investigating the strengths and weaknesses of two different techniques for using word-level features and by using a larger corpus of human-computer dialogue. Our results confirm that the word-level pitch features fare better than the turn-level ones regardless of the combination technique. In addition, we find that each word combination technique has different strengths and weaknesses in terms of precision and recall.
Greg Nicholas, Mihai Rotaru 0002, Diane J. Litman
SLT3
2006 Characterizing and Predicting Corrections in Spoken Dialogue Systems
abstract
This article focuses on the analysis and prediction of corrections, defined as turns where a user tries to correct a prior error made by a spoken dialogue system. We describe our labeling procedure of various corrections types and statistical analyses of their features in a corpus collected from a train information spoken dialogue system. We then present results of machine-learning experiments designed to identify user corrections of speech recognition errors. We investigate the predictive power of features automatically computable from the prosody of the turn, the speech recognition process, experimental conditions, and the dialogue history. Our best-performing features reduce classification error from baselines of 25.70–28.99% to 15.72%.
Diane J. Litman, Julia Hirschberg, Marc Swerts
Comput. Linguistics1
2006 Correlations between dialogue acts and learning in spoken tutoring dialogues
abstract
We examine correlations between dialogue behaviors and learning in tutoring, using two corpora of spoken tutoring dialogues: a human-human corpus and a human-computer corpus. To formalize the notion of dialogue behavior, we manually annotate our data using a tagset of student and tutor dialogue acts relative to the tutoring domain. A unigram analysis of our annotated data shows that student learning correlates both with the tutor's dialogue acts and with the student's dialogue acts. A bigram analysis shows that student learning also correlates with joint patterns of tutor and student dialogue acts. In particular, our human-computer results show that the presence of student utterances that display reasoning (whether correct or incorrect), as well as the presence of reasoning questions asked by the computer tutor, both positively correlate with learning. Our human-human results show that student introductions of a new concept into the dialogue positively correlates with learning, but student attempts at deeper reasoning (particularly when incorrect), and the human tutor's attempts to direct the dialogue, both negatively correlate with learning. These results suggest that while the use of dialogue act n-grams is a promising method for examining correlations between dialogue behavior and learning, specific findings can differ in human versus computer tutoring, with the latter better motivating adaptive strategies for implementation.
Diane J. Litman, Katherine Forbes-Riley
Nat. Lang. Eng.1
2006 Recognizing student emotions and attitudes on the basis of utterances in spoken tutoring dialogues with both human and computer tutors
Diane J. Litman, Katherine Forbes-Riley
Speech Commun.1
2005 Dialogue-Learning Correlations in Spoken Dialogue Tutoring
Katherine Forbes-Riley, Diane J. Litman, Alison Huettner, Arthur Ward
AIED2
2005 Correlating student acoustic-prosodic profiles with student learning in spoken tutoring dialogues
abstract
We examine correlations between student learning and student acoustic-prosodic profiles, which prior research has shown to be predictive of emotional states. We compare these correlations in two corpora of spoken tutoring dialogues: a human-human corpus and a human-computer corpus. Our results suggest that rather than relying on emotion prediction models developed via the more labor-intensive method of manually labeling emotions, adaptive strategies for our spoken dialogue tutoring system can be developed based on observed acoustic-prosodic profiles that we hypothesize to be reflective of emotion. 1.
Katherine Forbes-Riley, Diane J. Litman
INTERSPEECH2
2005 Speech recognition performance and learning in spoken dialogue tutoring
abstract
Speech recognition errors have been shown to negatively correlate with user satisfaction in evaluations of task-oriented spoken dialogue systems. In the domain of tutorial dialogue systems, however, where the primary evaluation metric is student learning, there has been little investigation of whether speech recognition errors also negatively correlate with learning. In this paper we examine correlations between student learning and automatic speech recognition performance, in a corpus of dialogues collected with an intelligent tutoring spoken dialogue system. We examine numerous quantitative measures of speech recognition error, including rejection versus misrecognition errors, word versus sentence-level errors, and transcription versus semantic errors. Our results show that although many of our students experience problems with speech recognition, none of our measures negatively correlates with student learning. 1.
Diane J. Litman, Katherine Forbes-Riley
INTERSPEECH1
2005 Using word-level pitch features to better predict student emotions during spoken tutoring dialogues
abstract
In this paper, we advocate for the usage of word-level pitch features for detecting user emotional states during spoken tutoring dialogues. Prior research has primarily focused on the use of turn-level features as predictors. We compute pitch features at the word level and resolve the problem of combining multiple features per turn using a word-level emotion model. Even under a very simple word-level emotion model, our results show an improvement in prediction using word-level features over using turn-level features. We find that the advantage of word-level features lies in a better prediction of longer turns. 1.
Mihai Rotaru 0002, Diane J. Litman
INTERSPEECH2
2005 Interactions between speech recognition problems and user emotions
abstract
Understanding how speech recognition problems affect the interaction with the user is a topic of great interest for the spoken dialogue community. In this paper, we examine the dependencies between speech recognition problems in adjacent turns. We also examine the dependencies between speech recognition problems and student emotions within a turn and in adjacent turns. We apply Chi Square ( � 2) analysis to a corpus of speech-based computer tutoring dialogues to discover these dependencies. We find that rejections are followed by more rejections than expected if there was no dependency between rejections, and that misrecognitions are followed by more misrecognitions than expected. We also find a strong dependency between recognition problems in the previous turn and user emotion in the current turn: after a system rejection there are more emotional user turns than expected. Surprisingly, in our data, we find no relationship between user emotions and recognition problems within a turn nor between previous turn user emotions and current turn recognition problems. 1.
Mihai Rotaru 0002, Diane J. Litman, Katherine Forbes-Riley
INTERSPEECH2
2004 Predicting Student Emotions in Computer-Human Tutoring Dialogues
abstract
We examine the utility of speech and lexical features for predicting student emotions in computer-human spoken tutoring dialogues. We first annotate student turns for negative, neutral, positive and mixed emotions. We then extract acoustic-prosodic features from the speech signal, and lexical items from the transcribed or recognized speech. We compare the results of machine learning experiments using these features alone or in combination to predict various categorizations of the annotated student emotions. Our best results yield a 19-36% relative improvement in error reduction over a baseline. Finally, we compare our results with emotion prediction in human-human tutoring dialogues.
Diane J. Litman, Katherine Forbes-Riley
ACL1
2004 Workshop on Analyzing Student-Tutor Interaction Logs to Improve Educational Outcomes
Joseph E. Beck, Ryan Baker 0001, Albert T. Corbett, Judy Kay, Diane J. Litman, Antonija Mitrovic, Steven Ritter 0001
Intelligent Tutoring Systems5
2004 Workshop on Dialog-Based Intelligent Tutoring Systems: State of the Art and New Research Directions
Neil T. Heffernan, Peter M. Hastings, Gregory Aist, Vincent Aleven, Ivon Arroyo, Paul Brna, Mark G. Core, Martha W. Evens, Reva Freedman, Michael Glass, Arthur C. Graesser, Kenneth R. Koedinger, Pamela W. Jordan, Diane J. Litman, Evelyn Lulis, Helen Pain, Carolyn P. Rosé, Beverly P. Woolf, Claus Zinn
Intelligent Tutoring Systems14
2004 Spoken Versus Typed Human and Computer Dialogue Tutoring
Diane J. Litman, Carolyn P. Rosé, Katherine Forbes-Riley, Kurt VanLehn, Dumisizwe Bhembe, Scott Silliman
Intelligent Tutoring Systems1
2004 Predicting Emotion in Spoken Dialogue from Multiple Knowledge Sources
Katherine Forbes-Riley, Diane J. Litman
HLT-NAACL2
2004 Prosodic and other cues to speech recognition failures
Julia Hirschberg, Diane J. Litman, Marc Swerts
Speech Commun.2
2003 Exceptionality and Natural Language Learning
Mihai Rotaru 0002, Diane J. Litman
CoNLL2
2003 Towards Emotion Prediction in Spoken Tutoring Dialogues
Diane J. Litman, Katherine Forbes-Riley, Scott Silliman
HLT-NAACL1
2002 Optimizing Dialogue Management with Reinforcement Learning: Experiments with the NJFun System
abstract
Designing the dialogue policy of a spoken dialogue system involves many nontrivial choices. This paper presents a reinforcement learning approach for automatically optimizing a dialogue policy, which addresses the technical challenges in applying reinforcement learning to a working dialogue system with human users. We report on the design, construction and empirical evaluation of NJFun, an experimental spoken dialogue system that provides users with access to information about fun things to do in New Jersey. Our results show that by optimizing its performance via reinforcement learning, NJFun measurably improves system performance.
Satinder Singh 0001, Diane J. Litman, Michael Kearns, Marilyn A. Walker
J. Artif. Intell. Res.2
2002 R++: Adding Path-Based Rules to C++
abstract
Object-oriented languages and rule-based languages offer two distinct and useful programming abstractions. However, previous attempts to integrate data-driven rules into object-oriented languages have typically achieved an uneasy union at best. R++ is a new, closer integration of the rule-based and object-oriented paradigms that extends C++ with a single programming construct, the path-based rule, as a new kind of class member. Path-based rules-data-driven rules that are restricted to following pointers between objects-are like automatic methods that are triggered by changes to the objects they monitor. Path-based rules provide a useful level of abstraction that encourages a more declarative style of programming and are valuable in object-oriented designs as a means of modeling dynamic collections of interdependent objects. Unlike more traditional pattern-matching rules, path-based rules are not at odds with the object-oriented paradigm and offer performance advantages for many natural applications.
Diane J. Litman, Peter F. Patel-Schneider, Anil Mishra, James M. Crawford, Daniel Dvorak
IEEE Trans. Knowl. Data Eng.1
2002 Designing and Evaluating an Adaptive Spoken Dialogue System
Diane J. Litman, Shimei Pan
User Model. User Adapt. Interact.1
2001 Predicting User Reactions to System Error
abstract
This paper focuses on the analysis and prediction of so-called aware sites, defined as turns where a user of a spoken dialogue system first becomes aware that the system has made a speech recognition error. We describe statistical comparisons of features of these aware sites in a train timetable spoken dialogue corpus, which reveal significant prosodic differences between such turns, compared with turns that 'correct' speech recognition errors as well as with 'normal' turns that are neither aware sites nor corrections. We then present machine learning results in which we show how prosodic features in combination with other automatically available features can predict whether or not a user turn was a normal turn, a correction, and/or an aware site.
Diane J. Litman, Julia Hirschberg, Marc Swerts
ACL1
2001 Identifying User Corrections Automatically in Spoken Dialogue Systems
Julia Hirschberg, Diane J. Litman, Marc Swerts
NAACL2
2001 Natural Language Processing and User Modeling: Synergies and Limitations
Ingrid Zukerman, Diane J. Litman
User Model. User Adapt. Interact.2
2000 Automatic Optimization of Dialogue Management
Diane J. Litman, Michael Kearns, Satinder Singh 0001, Marilyn A. Walker
COLING1
2000 Generalizing prosodic prediction of speech recognition errors
abstract
Since users of spoken dialogue systems have difficulty correcting system misconceptions, it is important for automatic speech recognition (ASR) systems to know when their best hypothesis is incorrect. We compare results of previous experiments which showed that prosody improves the detection of ASR errors to experiments with a new system and new domain, the W99 conference registration system. Our new results again show that prosodic features can improve prediction of ASR misrecognitions over the use of other standard techniques for ASR rejection.
Julia Hirschberg, Diane J. Litman, Marc Swerts
INTERSPEECH2
2000 Corrections in spoken dialogue systems
abstract
This study analyzes user corrections of system errors in the TOOT spoken dialogue system. We find that corrections differ from noncorrections prosodically, in ways consistent with hyperarticulated speech, although many corrections are not hyperarticulated. Yet both are misrecognized more frequently than non-corrections --- though no more likely to be rejected by the system. Corrections more distant from the error they correct tend to exhibit greater prosodic differences, and also to be recognized more poorly. System dialogue strategy affects users' choice of correction type, suggesting that strategy-specific methods of detecting or coaching users on corrections may be useful. Strategies that produce longer tasks but fewer misrecognitions and subsequent corrections are preferred by users. 1. INTRODUCTION Since spoken dialogue systems often make mistakes in recognizing user input, accurate methods of detecting and correcting system errors are essential to supporting successful interact...
Marc Swerts, Diane J. Litman, Julia Hirschberg
INTERSPEECH2
2000 Towards developing general models of usability with PARADISE
abstract
The design of methods for performance evaluation is a major open research issue in the area of spoken language dialogue systems. This paper presents the PARADISE methodology for developing predictive models of spoken dialogue performance, and shows how to evaluate the predictive power and generalizability of such models. To illustrate the methodology, we develop a number of models for predicting system usability (as measured by user satisfaction), based on the application of PARADISE to experimental data from three different spoken dialogue systems. We then measure the extent to which the models generalize across different systems, different experimental conditions, and different user populations, by testing models trained on a subset of the corpus against a test set of dialogues. The results show that the models generalize well across the three systems, and are thus a first approximation towards a general performance model of system usability.
Marilyn A. Walker, Candace A. Kamm, Diane J. Litman
Nat. Lang. Eng.3
1999 Automatic Detection of Poor Speech Recognition at the Dialogue Level
abstract
The dialogue strategies used by a spoken dialogue system strongly influence performance and user satisfaction. An ideal system would not use a single fixed strategy, but would adapt to the circumstances at hand. To do so, a system must be able to identify dialogue properties that suggest adaptation. This paper focuses on identifying situations where the speech recognizer is performing poorly. We adopt a machine learning approach to learn rules from a dialogue corpus for identifying these situations. Our results show a significant improvement over the baseline and illustrate that both lower-level acoustic features and higher-level dialogue features can affect the performance of the learning algorithm.
Diane J. Litman, Marilyn A. Walker, Michael Kearns
ACL1
1999 Reinforcement Learning for Spoken Dialogue Systems
Satinder Singh 0001, Michael Kearns, Diane J. Litman, Marilyn A. Walker
NIPS3
1998 From novice to expert: the effect of tutorials on user expertise with spoken dialogue systems
abstract
One of the challenges for the current state of the art in spoken dialogue systems is how to make the limitations of the system apparent to users. These limitations have many sources: limited vocabulary, limited grammar, or limitations in the application domain. This study explored the use of a 4-minute tutorial session to acquaint novice users with the features of a spoken dialogue system for accessing email. On a set of three scenariobased tasks, novice users who had the tutorial had task completion times and user satisfaction ratings that were comparable to those of expert users of the system. Novices who did not experience the tutorial had significantly longer task completion times on the initial task, but similar completion times to the tutorial group on the final task. User satisfaction ratings of the no-tutorial group were consistently lower than the ratings of the tutorial and the expert groups. Evaluation using the PARADISE [7] framework indicated that perceived task completion, mean recognition score, and number of help requests were significant predictors of user satisfaction with the system. 1.
Candace A. Kamm, Diane J. Litman, Marilyn A. Walker
ICSLP2
1998 Evaluating spoken dialogue agents with PARADISE: Two case studies
Marilyn A. Walker, Diane J. Litman, Candace A. Kamm, Alicia Abella
Comput. Speech Lang.2
1997 PARADISE: A Framework for Evaluating Spoken Dialogue Agents
abstract
This paper presents PARADISE (PARAdigm for DIalogue System Evaluation), a general framework for evaluating spoken dialogue agents. The framework decouples task requirements from an agent's dialogue behaviors, supports comparisons among dialogue strategies, enables the calculation of performance over subdialogues and whole dialogues, specifies the relative contribution of various factors to performance, and makes it possible to compare agents performing different tasks by normalizing for task complexity.
Marilyn A. Walker, Diane J. Litman, Candace A. Kamm, Alicia Abella
ACL2
1997 Modeling Dynamic Collections of Interdependent Objects Using Path-Based Rules
abstract
Standard object-oriented languages do not provide language support for modeling changing collections of interdependent objects. We propose that R++, an integration of the rule and objectoriented paradigms, provides a mechanism for easily implementing such models. R++ extends C++ by adding a new programming construct called the path-based rule. Such data-driven rules are restricted to follow pointers between objects, and are like "automatic methods" that are triggered by changes to monitored objects. Path-based rules encourage a more abstract level of programming, and unlike previous rule integrations, are not at odds with the object-oriented paradigm and offer performance advantages for natural applications. 1 Introduction Object-oriented languages have simplified the design and implementation of sophisticated applications. However, as application domains have become increasingly dynamic and complex, it has become necessary to model such domains using changing collections of interdep...
Diane J. Litman, Anil Mishra, Peter F. Patel-Schneider
OOPSLA1
1997 Discourse Segmentation by Human and Automated Means
Rebecca J. Passonneau, Diane J. Litman
Comput. Linguistics2
1996 Taxonomic Plan Reasoning
Premkumar T. Devanbu, Diane J. Litman
Artif. Intell.2
1996 Cue Phrase Classification Using Machine Learning
abstract
Cue phrases may be used in a discourse sense to explicitly signal discourse structure, but also in a sentential sense to convey semantic rather than structural information. Correctly classifying cue phrases as discourse or sentential is critical in natural language processing systems that exploit discourse structure, e.g., for performing tasks such as anaphora resolution and plan recognition. This paper explores the use of machine learning for classifying cue phrases as discourse or sentential. Two machine learning programs (Cgrendel and C4.5) are used to induce classification models from sets of pre-classified cue phrases and their features in text and speech. Machine learning is shown to be an effective technique for not only automating the generation of classification models, but also for improving upon previous results. When compared to manually derived classification models already in the literature, the learned models often perform with higher accuracy and contain new linguistic insights into the data. In addition, the ability to automatically construct classification models makes it easier to comparatively analyze the utility of alternative feature representations of the data. Finally, the ease of retraining makes the learning approach more scalable and flexible than manual methods.
Diane J. Litman
J. Artif. Intell. Res.1
1995 Combining Multiple Knowledge Sources for Discourse Segmentation
abstract
We predict discourse segment boundaries from linguistic features of utterances, using a corpus of spoken narratives as data. We present two methods for developing segmentation algorithms from training data: hand tuning and machine learning. When multiple types of features are used, results approach human performance on an independent test set (both methods), and using cross-validation (machine learning).
Diane J. Litman, Rebecca J. Passonneau
ACL1
1995 Device Representation and Reasoning with Affective Relations
James M. Crawford, Daniel Dvorak, Diane J. Litman, Anil Mishra, Peter F. Patel-Schneider
IJCAI3
1994 Classifying Cue Phrases in Text and Speech Using Machine Learning
Diane J. Litman
AAAI1
1993 Intention-Based Segmentation: Human Reliability and Correlation with Linguistic Cues
abstract
Certain spans of utterances in a discourse, referred to here as segments, are widely assumed to form coherent units. Further, the segmental structure of discourse has been claimed to constrain and be constrained by many phenomena. However, there is weak consensus on the nature of segments and the criteria for recognizing or generating them. We present quantitative results of a two part study using a corpus of spontaneous, narrative monologues. The first part evaluates the statistical reliability of human segmentation of our corpus, where speaker intention is the segmentation criterion. We then use the subjects' segmentations to evaluate the correlation of discourse segmentation with three linguistic cues (referential noun phrases, cue words, and pauses), using information retrieval metrics.
Rebecca J. Passonneau, Diane J. Litman
ACL2
1993 Empirical Studies on the Disambiguation of Cue Phrases
Julia Hirschberg, Diane J. Litman
Comput. Linguistics2
1992 Terminological Reasoning with Constraint Networks and an Application to Plan Recognition
Robert A. Weida, Diane J. Litman
KR2
1992 On the Interaction between Plan Recognition and Intelligent Interfaces
Bradley A. Goodman, Diane J. Litman
User Model. User Adapt. Interact.2
1991 Plan-Based Terminological Reasoning
Premkumar T. Devanbu, Diane J. Litman
KR2
1990 Disambiguating Cue Phrases in Text and Speech
Diane J. Litman, Julia Hirschberg
COLING1
1987 Now let's Talk about Now; Identifying Cue Phrases Intonationally
abstract
Cue phrases are words and phrases such as now and by the way which may be used to convey explicit information about the structure of a discourse.However, while cue phrases may convey discourse structure, each may also be used to different effect.The question of how speakers and hearers distinguish between such uses of cue phrases has not been addressed in discourse studies to date.Based on a study of now in natural recorded discourse, we propose that cue and non-cue usage can be distinguished intonationally, on the basis of phrasing and accent.
Julia Hirschberg, Diane J. Litman
ACL2
1987 Intonation and the Intentional Structure of Discourse
Julia Hirschberg, Diane J. Litman, Janet B. Pierrehumbert, G. Ward
IJCAI2
1986 Understanding Plan Ellipsis
Diane J. Litman
AAAI1
1986 Linguistic Coherence: a Plan-Based Alternative
abstract
To fully understand a sequence of utterances, one must be able to infer implicit relationships between the utterances. Although the identification of sets of utterance relationships forms the basis for many theories of discourse, the formalization and recognition of such relationships has proven to be an extremely difficult computational task.This paper presents a plan-based approach to the representation and recognition of implicit relationships between utterances. Relationships are formulated as discourse plans, which allows their representation in terms of planning operators and their computation via a plan recognition process. By incorporating complex inferential processes relating utterances into a plan-based framework, a formalization and computability not available in the earlier works is provided.
Diane J. Litman
ACL1
1986 Plans, goals, and language
abstract
One of the most promising computational approaches to representing context in natural language systems has been based on work in general problem solving. In this approach, plans are used both to represent the domain of discourse as well as the communication process itself. Using a simplified framework for planning and action reasoning, we describe techniques that allow systems to handle many dialogues that are problematic for other systems, including the use of sentence fragments, indirect speech, helpful responses, the tracking of the topic of conversations both with and without interrupting subdialogues, and topic change.
James F. Allen, Diane J. Litman
Proc. IEEE2
1984 A Plan Recognition Model for Clarification Subdialogues
abstract
One of the promising approaches to analyzing task-oriented dialogues has involved modeling the plans of the speakers in the task domain. In general, these models work well as long as the topic follows the task structure closely, but they have difficulty in accounting for clarification subdialogues and topic change. We have developed a model based on a hierarchy of plans and metaplans that accounts for the clarification subdialogues while maintaining the advantages of the plan-based approach.
Diane J. Litman, James F. Allen
COLING1
1982 ARGOT: The Rochester Dialogue System
James F. Allen, Alan M. Frisch, Diane J. Litman
AAAI3