Carolyn P. Rosé

dblp:r/CarolynPensteinRose · also Carolyn Penstein Rosé · DBLP profile ↗
← Back
174ranked-venue papers
11as first author
37since 2021 · last 2026
0000-0003-1128-5155ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 77 · 4 first-author · 25 since 2021Human-computer interaction and ubiquitous computing · 77 · 7 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 77 · 7 first-author · 6 since 2021Databases, data management, data science and information retrieval · 13 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 since 2021Systems, architecture and hardware · 5 · 1 first-author · 1 since 2021Security and privacy · 2Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 AI Education Across the Curriculum: Design and Pilot Study of a Cross-Disciplinary Module Set
abstract
Artificial intelligence (AI) education has garnered growing attention from both educational researchers and practitioners in recent years. Among the various emerging approaches, integrating AI education across the curriculum—particularly within core disciplines—offers distinct advantages. This strategy foregrounds the inherently interdisciplinary nature of AI and enables students to investigate its connections with subjects such as mathematics and English language arts (ELA). Furthermore, it holds promise for broadening participation by engaging all students, including those historically underrepresented and underserved in the field of AI. To date, most efforts to integrate AI education have been situated within individual classrooms, often led by a single teacher. While such initiatives provide valuable entry points, they overlook the reality that students’ learning experiences span multiple classrooms and disciplines. As students transition between subjects, they inevitably synthesize ideas—both consciously and unconsciously—from diverse instructional contexts. Recognizing this, we take a whole-school perspective that considers the cumulative and interconnected nature of students’ learning experiences. With this perspective, we explore a coordinated, cross-disciplinary approach in which students engage with AI through a set of curriculum modules spanning mathematics, ELA, and social studies. Each module is discipline-specific yet designed to contribute to a cohesive, cross-disciplinary exploration of AI. These modules are further framed by a self-paced introductory unit, which establishes foundational concepts, and a culminating application-and-reflection unit, which supports integration and transfer of learning. This paper describes the design of the AI Education Across the Curriculum module set and reports preliminary findings from a pilot implementation conducted in Spring 2025. By examining both the pedagogical design and initial findings, we aim to contribute to the growing body of research on scalable, equitable, and interdisciplinary models for AI education.
Jie Chao, Rebecca Ellis, Shiyan Jiang, Daria Smyslova, Qiuqing Li, Amato Nocera, Christy Byrd, Carolyn P. Rosé, Stephen Callahan, Dianne O'Grady-Cunniff
AAAI9
2025 Programming by Example meets Historical Linguistics: A Large Language Model Based Approach to Sound Law Induction
abstract
Atharva Naik, Darsh Agrawal, Hong Sng, Clayton Marr, Kexun Zhang, Nathaniel Romney Robinson, Kalvin Chang, Rebecca Byrnes, Aravind Mysore, Carolyn Rose, David R. Mortensen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Atharva Naik, Darsh Agrawal, Hong Sng, Clayton Marr, Kexun Zhang, Nathaniel R. Robinson, Kalvin Chang, Rebecca Byrnes, Aravind Mysore, Carolyn P. Rosé, David R. Mortensen
ACL (1)10
2025 Improving Model Factuality with Fine-grained Critique-based Evaluator
abstract
Yiqing Xie, Wenxuan Zhou, Pradyot Prakash, Di Jin, Yuning Mao, Quintin Fettes, Arya Talebzadeh, Sinong Wang, Han Fang, Carolyn Rose, Daniel Fried, Hejia Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yiqing Xie, Pradyot Prakash, Yuning Mao, Quintin Fettes, Arya Talebzadeh, Sinong Wang, Carolyn P. Rosé, Daniel Fried
ACL (1)10
2025 Epistemic Curiosity in K-12 AI Education: A Trajectory Analysis
Min Zhuang, Shiyan Jiang, Daria Smyslova, Carolyn P. Rosé, Jie Chao
AIED (4)4
2025 SOCIAL SCAFFOLDS: A Generalization Framework for Social Understanding Tasks
abstract
Effective human communication in social settings is contingent on recognizing subtle cues, such as intentions or implications.Without such cues, NLP models risk missing social signals, instead relying on surface patterns.We introduce SOCIAL SCAFFOLDS, an automated framework for facilitating generalization across social reasoning tasks by generating rationales that make these social cues explicit.Grounded in narrative modeling principles, we generate task-agnostic rationales that capture different perspectives, i.e., that of the speaker, the listener, and the general world-view.Our experimental suite showcases that providing rationales as augmentations aids task performance for both supervised fine-tuning and incontext learning paradigms.Notably, providing all three rationale types significantly improves cross-task performance in 44% of cases, and inferred speaker intent in 31.3% of cases.We conduct statistical and ablation analyses that show how rationales complement the input text and are used effectively by models.
Ritam Dutt, Carolyn P. Rosé, Maarten Sap
EMNLP2
2025 An Empirical Study on Strong-Weak Model Collaboration for Repo-level Code Generation
abstract
We study cost-efficient collaboration between strong and weak language models for repository-level code generation, where the weak model handles simpler tasks at lower cost, and the most challenging tasks are delegated to the strong model.While many works propose architectures for this task, few analyze performance relative to cost.We evaluate a broad spectrum of collaboration strategies: context-based, pipeline-based, and dynamic, on GitHub issue resolution.Our most effective collaborative strategy achieves equivalent performance to the strong model while reducing the cost by 40%.Based on our findings, we offer actionable guidelines for choosing collaboration strategies under varying budget and performance constraints.Our results show that strong-weak collaboration substantially boosts the weak model's performance at a fraction of the cost, pipeline and context-based methods being most efficient.We have also opensourced the code 1 for our work.
Shubham Gandhi, Atharva Naik, Yiqing Xie, Carolyn P. Rosé
EMNLP4
2025 CRScore: Grounding Automated Evaluation of Code Review Comments in Code Claims and Smells
abstract
Atharva Naik, Marcus Alenius, Daniel Fried, Carolyn Rose. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Atharva Naik, Marcus Alenius, Daniel Fried, Carolyn P. Rosé
NAACL (Long Papers)4
2025 Bridging the Community College Cybersecurity Classroom and Workplace with the CyberSim Lab
abstract
Most postsecondary cybersecurity education focuses on technical knowledge and skills without commensurate attention to vital non-technical skills. In this position paper, we argue that cybersecurity education must integrate teaching and practicing of non-technical competencies alongside technical knowledge and skills to ensure that both technical and non-technical skills transfer to cybersecurity workplaces. We identify specific learning outcomes that meet these criteria and suggest research-based pedagogical approaches to support learning and transfer. We present a cybersecurity lab designed to address these learning outcomes through experiential learning, roleplay, collaborative learning, technical simulation and metacognitive engagement. The CyberSim Lab serves as a curricular bridge between the classroom and the workplace.
Judeth Oden Choi, Rotem D. Guttman, Matthew Kisow, Carolyn P. Rosé, William Nichols, James Winyard, Bruce Li, Lee G. Branstetter, Lauren Herckis
SIGCSE (1)4
2025 Designing Looms as Kits for Collaborative Assembly
Samantha Speer, Nickolina Yankova, Joey Huang, Carolyn P. Rosé, Kylie Peppler, James McCann, Melisa Orta Martinez
UIST4
2024 Leveraging Machine-Generated Rationales to Facilitate Social Meaning Detection in Conversations
abstract
We present a generalizable classification approach that leverages Large Language Models (LLMs) to facilitate the detection of implicitly encoded social meaning in conversations.We design a multi-faceted prompt to extract a textual explanation of the reasoning that connects visible cues to underlying social meanings.These extracted explanations or rationales serve as augmentations to the conversational text to facilitate dialogue understanding and transfer.Our empirical results over 2,340 experimental settings demonstrate the significant positive impact of adding these rationales.Our findings hold true for in-domain classification, zero-shot, and few-shot domain transfer for two different social meaning detection tasks, each spanning two different corpora.
Ritam Dutt, Jiaxin Shi, Divyanshu Sheth, Prakhar Gupta, Carolyn P. Rosé
ACL (1)6
2024 Estimating Agreement by Chance for Sequence Annotation
abstract
In the field of natural language processing, correction of performance assessment for chance agreement plays a crucial role in evaluating the reliability of annotations.However, there is a notable dearth of research focusing on chance correction for assessing the reliability of sequence annotation tasks, despite their widespread prevalence in the field.To address this gap, this paper introduces a novel model for generating random annotations, which serves as the foundation for estimating chance agreement in sequence annotation tasks.Utilizing the proposed randomization model and a related comparison approach, we successfully derive the analytical form of the distribution, enabling the computation of the probable location of each annotated text segment and subsequent chance agreement estimation.Through a combination simulation and corpus-based evaluation, we successfully assess its applicability and validate its accuracy and efficacy.
Diya Li, Carolyn P. Rosé, Ao Yuan, Chunxiao Zhou
ACL (1)2
2024 DocLens: Multi-aspect Fine-grained Medical Text Evaluation
abstract
Yiqing Xie, Sheng Zhang, Hao Cheng, Pengfei Liu, Zelalem Gero, Cliff Wong, Tristan Naumann, Hoifung Poon, Carolyn Rose. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yiqing Xie, Sheng Zhang 0012, Hao Cheng 0002, Zelalem Gero, Cliff Wong, Tristan Naumann, Hoifung Poon, Carolyn P. Rosé
ACL (1)9
2024 Generating Situated Reflection Triggers About Alternative Solution Paths: A Case Study of Generative AI for Computer-Supported Collaborative Learning
Atharva Naik, Jessica Ruhan Yin, Anusha Kamath, Qianou Ma, Sherry Tongshuang Wu, R. Charles Murray, Christopher Bogart, Majd F. Sakr, Carolyn P. Rosé
AIED (1)9
2024 Distilling Multi-Scale Knowledge for Event Temporal Relation Extraction
abstract
Event Temporal Relation Extraction (ETRE) is paramount but challenging. Within a discourse, event pairs are situated at different distances or the so-called proximity bands. The temporal ordering communicated about event pairs where at more remote (i.e., "long'') or less remote (i.e., "short'') proximity bands are encoded differently. SOTA models have tended to perform well on events situated at either short or long proximity bands, but not both. Nonetheless, real-world, natural texts contain all types of temporal event-pairs. In this paper, we present MulCo : Distilling Mul ti-Scale Knowledge via Co ntrastive Learning, a knowledge co-distillation approach that shares knowledge across multiple event pair proximity bands to improve performance on all types of temporal datasets. Our experimental results show that MulCo successfully integrates linguistic cues pertaining to temporal reasoning across both short and long proximity bands and achieves new state-of-the-art results on several ETRE benchmark datasets.
Hao-Ren Yao, Luke Breitfeller, Aakanksha Naik, Chunxiao Zhou, Carolyn P. Rosé
CIKM5
2023 Exploring Artificial Intelligence in English Language Arts with StoryQ
abstract
Exploring Artificial Intelligence (AI) in English Language Arts (ELA) with StoryQ is a 10-hour curriculum module designed for high school ELA classes. The module introduces students to fundamental AI concepts and essential machine learning workflow using StoryQ, a web-based GUI environment for Grades 6-12 learners. In this module, students work with unstructured text data and learn to train, test, and improve text classification models such as intent recognition, clickbait filter, and sentiment analysis. As they interact with machine-learning language models deeply, students also gain a nuanced understanding of language and how to wield it, not just as a data structure, but as a tool in our human-human encounters as well. The current version contains eight lessons, all delivered through a full-featured online learning and teaching platform. Computers and Internet access are required to implement the module. The module was piloted in an ELA class in the Spring of 2022, and the student learning outcomes were positive. The module is currently undergoing revision and will be further tested and improved in Fall 2022.
Jie Chao, Rebecca Ellis, Shiyan Jiang, Carolyn P. Rosé, William Finzer, Cansu Tatar, James Fiacco, Kenia Wiedemann
AAAI4
2023 Linguistic representations for fewer-shot relation extraction across domains
abstract
Recent work has demonstrated the positive impact of incorporating linguistic representations as additional context and scaffolding on the in-domain performance of several NLP tasks.We extend this work by exploring the impact of linguistic representations on cross-domain performance in a few-shot transfer setting.An important question is whether linguistic representations enhance generalizability by providing features that function as cross-domain pivots.We focus on the task of relation extraction on three datasets of procedural text in two domains, cooking and materials science.Our approach augments a popular transformerbased architecture by alternately incorporating syntactic and semantic graphs constructed by freely available off-the-shelf tools.We examine their utility for enhancing generalization, and investigate whether earlier findings, e.g. that semantic representations can be more helpful than syntactic ones, extend to relation extraction in multiple domains.We find that while the inclusion of these graphs results in significantly higher performance in few-shot transfer, both types of graph exhibit roughly equivalent utility.
Sireesh Gururaja, Ritam Dutt, Tinglong Liao, Carolyn P. Rosé
ACL (1)4
2023 Using counterfactual contrast to improve compositional generalization for multi-step quantitative reasoning
abstract
In quantitative question answering, compositional generalization is one of the main challenges of state of the art models, especially when longer sequences of reasoning steps are required.In this paper we propose Counter-Comp, a method that uses counterfactual scenarios to generate samples with compositional contrast.Instead of a data augmentation approach, CounterComp is based on metric learning, which allows for direct sampling from the training set and circumvents the need for additional human labels.Our proposed auxiliary metric learning loss improves the performance of three state of the art models on four recently released datasets.We also show how the approach can improve OOD performance on unseen domains, as well as unseen compositions.Lastly, we demonstrate how the method can lead to better compositional attention patterns during training.
Armineh Nourbakhsh, Sameena Shah, Carolyn P. Rosé
ACL (1)3
2023 Teach Artificial Intelligence with StoryQ, A Web-Based Machine Learning and Text Mining Tool for K-12 Students
abstract
StoryQ is a web-based machine learning and text mining tool that allows young learners (Grade 6-12) to engage in machine learning practices and work with unstructured text data without needing to code. StoryQ features dynamically linked data representations that promote meaningful inquiries and understandings across tables, graphs, and texts. These links create a unique user experience that makes machine learning models transparent, explainable, and fun to explore. This demo will showcase how key AI concepts such as representation, reasoning, feature space, feature weight, and machine learning are dynamically visualized in StoryQ and made accessible to young learners. A brief tutorial will be provided on how to use StoryQ to train, test, and troubleshoot text classification models using both standard feature extractors (e.g., N-grams) and special feature extraction tools and visualizations that have been specially designed to support young learners and non-computing teachers. This demo will also include sample learning activities designed for high school English Language Arts and History classes to showcase how machine learning concepts and practices can be introduced in non-computing classes. As the demands for AI scientists, engineers, and entrepreneurs have increased in recent years, as well as AI's increased presence in everyday lives, making access to how machine learning practices work is of paramount importance for young learners. This work is supported by an NSF ITEST project (DRL-1949110).
Jie Chao, William Finzer, Carolyn P. Rosé, Shiyan Jiang, Rebecca Ellis, Kenia Wiedemann, Cansu Tatar, James Fiacco
SIGCSE (2)3
2023 SPEERLoom: An Open-Source Loom Kit for Interdisciplinary Engagement in Math, Engineering, and Textiles
abstract
Weaving is a fabrication process that is grounded in mathematics and engineering: from the binary, matrix-like nature of the pattern drafts weavers have used for centuries, to the punch card programming of the first Jacquard looms. This intersection of disciplines provides an opportunity to ground abstract mathematical concepts in a concrete and embodied art, viewing this textile art through the lens of engineering. Currently, available looms are not optimized to take advantage of this opportunity to increase mathematics learning by providing hands-on interdisciplinary learning in collegiate classrooms. In this work, we present SPEERLoom: an open-source, robotic Jacquard loom kit designed to be a tool for interweaving cloth fabrication, mathematics, and engineering to support interdisciplinary learning in the classroom. We discuss the design requirements and subsequent design of SPEERLoom. We also present the results of a pilot study in a post-secondary class finding that SPEERLoom supports hands-on, interdisciplinary learning of math, engineering, and textiles.
Samantha Speer, Ana P. Garcia-Alonzo, Joey Huang, Nickolina Yankova, Carolyn P. Rosé, Kylie Peppler, James McCann, Melisa Orta Martinez
UIST5
2022 StoryQ - an Online Environment for Machine Learning of Text Classification
abstract
The StoryQ environment provides an intuitive graphical user interface for middle and high school students to create features from unstructured text data and train and test classification models using logistic regression. StoryQ runs in a web browser, is free and requires no installation. AI concepts addressed include: features, weights, accuracy, training, bias, error analysis and cross validation. Using the software in conjunction with curriculum currently under development is expected to lead to student understanding of machine learning concepts and workflow; developing the ability to use domain knowledge and basic linguistics to identify, create, analyze, and evaluate features; becoming aware of and appreciating the roles and responsibilities of AI developers;. This paper will consist of an online demo with a brief video walkthrough.
William Finzer, Jie Chao, Carolyn P. Rosé, Shiyan Jiang
AAAI3
2022 A Framework for Adapting Pre-Trained Language Models to Knowledge Graph Completion
abstract
Recent work has demonstrated that entity representations can be extracted from pre-trained language models to develop knowledge graph completion models that are more robust to the naturally occurring sparsity found in knowledge graphs.In this work, we conduct a comprehensive exploration of how to best extract and incorporate those embeddings into knowledge graph completion models.We explore the suitability of the extracted embeddings for direct use in entity ranking and introduce both unsupervised and supervised processing methods that can lead to improved downstream performance.We then introduce supervised embedding extraction methods that can extract more informative representations.We then synthesize our findings and develop a knowledge graph completion model that significantly outperforms recent neural models. 1
Justin Lovelace, Carolyn P. Rosé
EMNLP2
2022 Improving compositional generalization for multi-step quantitative reasoning in question answering
abstract
Quantitative reasoning is an important aspect of question answering, especially when numeric and verbal cues interact to indicate sophisticated, multi-step programs.In this paper, we demonstrate how modeling the compositional nature of quantitative text can enhance the performance and robustness of QA models, allowing them to capture arithmetic logic that is expressed verbally.Borrowing from the literature on semantic parsing, we propose a method that encourages the QA models to adjust their attention patterns and capture input/output alignments that are meaningful to the reasoning task.We show how this strategy improves program accuracy and renders the models more robust against overfitting as the number of reasoning steps grows.Our approach is designed as a standalone module which can be prepended to many existing models and trained in an endto-end fashion without the need for additional supervisory signal.As part of this exercise, we also create a unified dataset building on four previously released numerical QA datasets over tabular data 1 .
Armineh Nourbakhsh, Cathy Jiao, Sameena Shah, Carolyn P. Rosé
EMNLP4
2022 A Tale of Two Subreddits: Measuring the Impacts of Quarantines on Political Engagement on Reddit
Qinlan Shen, Carolyn P. Rosé
ICWSM2
2022 QA4IE: A Quality Assurance Tool for Information Extraction
abstract
Quality assurance (QA) is an essential though underdeveloped part of the data annotation process. Although QA is supported to some extent in existing annotation tools, comprehensive support for QA is not standardly provided. In this paper we contribute QA4IE, a comprehensive QA tool for information extraction, which can (1) detect potential problems in text annotations in a timely manner, (2) accurately assess the quality of annotations, (3) visually display and summarize annotation discrepancies among annotation team members, (4) provide a comprehensive statistics report, and (5) support viewing of annotated documents interactively. This paper offers a competitive analysis comparing QA4IE and other popular annotation tools and demonstrates its features, usage, and effectiveness through a case study. The Python code, documentation, and demonstration video are available publicly at https://github.com/CC-RMD-EpiBio/QA4IE.
Rafael Jiménez Silva, Kaushik Gedela, Alex Marr, Bart Desmet, Carolyn P. Rosé, Chunxiao Zhou
LREC5
2022 StoryQ: A Web-Based Machine Learning and Text Mining Tool for K-12 Students
abstract
StoryQ is a web-based machine learning and text mining tool that allows young learners (Grade 6-12) to engage in machine learning practices and work with unstructured text data without coding. StoryQ features dynamically linked data representations that promote meaningful inquiries across tables, graphs, and texts--a unique user experience that makes machine learning models transparent, explainable, and fun to explore. This demo provides a brief tutorial on how to use StoryQ to train, test, and troubleshoot text classification models using both standard feature extractors (e.g., N-grams) and special feature extraction tools and visualizations designed to support young learners and non-computing teachers. This demo also includes sample learning activities designed for high school English Language Arts classes to showcase how machine learning concepts and practices can be introduced in non-computing classes. This work is supported by an NSF I-TEST project (DRL-1949110).
Jie Chao, William Finzer, Carolyn P. Rosé, Shiyan Jiang, Michael Miller Yoder, James Fiacco, Chas Murray, Cansu Tatar, Kenia Wiedemann
SIGCSE (2)3
2022 Adapting to the Long Tail: A Meta-Analysis of Transfer Learning Research for Language Understanding Tasks
abstract
Abstract Natural language understanding (NLU) has made massive progress driven by large benchmarks, but benchmarks often leave a long tail of infrequent phenomena underrepresented. We reflect on the question: Have transfer learning methods sufficiently addressed the poor performance of benchmark-trained models on the long tail? We conceptualize the long tail using macro-level dimensions (underrepresented genres, topics, etc.), and perform a qualitative meta-analysis of 100 representative papers on transfer learning research for NLU. Our analysis asks three questions: (i) Which long tail dimensions do transfer learning studies target? (ii) Which properties of adaptation methods help improve performance on the long tail? (iii) Which methodological gaps have greatest negative impact on long tail performance? Our answers highlight major avenues for future research in transfer learning for the long tail. Lastly, using our meta-analysis framework, we perform a case study comparing the performance of various adaptation methods on clinical narratives, which provides interesting insights that may enable us to make progress along these future avenues.
Aakanksha Naik, Jill Fain Lehman, Carolyn P. Rosé
Trans. Assoc. Comput. Linguistics3
2021 Robust Knowledge Graph Completion with Stacked Convolutions and a Student Re-Ranking Network
abstract
Knowledge Graph (KG) completion research usually focuses on densely connected benchmark datasets that are not representative of real KGs. We curate two KG datasets that include biomedical and encyclopedic knowledge and use an existing commonsense KG dataset to explore KG completion in the more realistic setting where dense connectivity is not guaranteed. We develop a deep convolutional network that utilizes textual entity representations and demonstrate that our model outperforms recent KG completion methods in this challenging setting. We find that our model's performance improvements stem primarily from its robustness to sparsity. We then distill the knowledge from the convolutional network into a student network that re-ranks promising candidate entities. This re-ranking stage leads to further improvements in performance and demonstrates the effectiveness of entity re-ranking for KG completion.
Justin Lovelace, Denis Newman-Griffis, Shikhar Vashishth, Jill Fain Lehman, Carolyn P. Rosé
ACL/IJCNLP (1)5
2021 Seeing Beyond Expert Blind Spots: Online Learning Design for Scale and Quality
abstract
Maximizing system scalability and quality are sometimes at odds. This work provides an example showing scalability and quality can be achieved at the same time in instructional design, contrary to what instructors may believe or expect. We situate our study in the education of HCI methods, and provide suggestions to improve active learning within the HCI education community. While designing learning and assessment activities, many instructors face the choice of using open-ended or close-ended activities. Close-ended activities such as multiple-choice questions (MCQs) enable automated feedback to students. However, a survey with 22 HCI professors revealed a belief that MCQs are less valuable than open-ended questions, and thus, using them entails making a quality sacrifice in order to achieve scalability. A study with 178 students produced no evidence to support the teacher belief. This paper indicates more promise than concern in using MCQs for scalable instruction and assessment in at least some HCI domains.
Xu Wang 0016, Carolyn P. Rosé, Kenneth R. Koedinger
CHI2
2021 ResPer: Computationally Modelling Resisting Strategies in Persuasive Conversations
abstract
Ritam Dutt, Sayan Sinha, Rishabh Joshi, Surya Shekhar Chakraborty, Meredith Riggs, Xinru Yan, Haogang Bao, Carolyn Rose. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Ritam Dutt, Sayan Sinha, Rishabh Joshi, Surya Shekhar Chakraborty, Meredith Riggs, Xinru Yan, Haogang Bao, Carolyn P. Rosé
EACL8
2021 Adapting Event Extractors to Medical Data: Bridging the Covariate Shift
abstract
We tackle the task of adapting event extractors to new domains without labeled data, by aligning the marginal distributions of source and target domains.As a testbed, we create two new event extraction datasets using English texts from two medical domains: (i) clinical notes, and (ii) doctor-patient conversations.We test the efficacy of three marginal alignment techniques: (i) adversarial domain adaptation (ADA), (ii) domain adaptive fine-tuning (DAFT), and (iii) a new instance weighting technique based on language model likelihood scores (LIW).LIW and DAFT improve over a no-transfer BERT baseline on both domains, but ADA only improves on notes.Deeper analysis of performance under different types of shifts (e.g., lexical shift, semantic shift) explains some of the variations among models.Our best-performing models reach F1 scores of 70.0 and 72.9 on notes and conversations respectively, using no labeled target data.
Aakanksha Naik, Jill Fain Lehman, Carolyn P. Rosé
EACL3
2021 What Sounds "Right" to Me? Experiential Factors in the Perception of Political Ideology
abstract
In this paper, we challenge the assumption that political ideology is inherently built into text by presenting an investigation into the impact of experiential factors on annotator perceptions of political ideology.We construct an annotated corpus of U.S. political discussion, where in addition to ideology labels for texts, annotators provide information about their political affiliation, exposure to political news, and familiarity with the source domain of discussion, Reddit.We investigate the variability in ideology judgments across annotators, finding evidence that these experiential factors may influence the consistency of how political ideologies are perceived.Finally, we present evidence that understanding how humans perceive and interpret ideology from texts remains a challenging task for state-ofthe-art language models, pointing towards potential issues when modeling user experiences that may require more contextual knowledge.
Qinlan Shen, Carolyn P. Rosé
EACL2
2021 Combining Collaborative Reflection based on Worked-Out Examples with Problem-Solving Practice: Designing Collaborative Programming Projects for Learning at Scale
abstract
Computer science pedagogy has overwhelmingly favored problem-solving practice over methods of engagement like worked-out example study especially in advanced classes. This is due to the belief that while these alternative methods may improve student conceptual learning, they may leave them less able to perform on authentic problem-solving tasks from a lack of hands-on practice. In this paper, we perform a direct comparison of this trade-off in a synchronous collaborative programming project by adjusting the boundary between problem-solving and collaborative reflection based on a worked-out example while keeping the total time on task constant. We find that the more time students spent on worked example study, the more was the observed improvement in the pre- to post-test scores with no significant difference in performance on a subsequent problem-solving task. These results, therefore, challenge the dominant place of problem-solving practice in the advanced curricular context and inform the design of collaborative programming projects at scale.
Sreecharan Sankaranarayanan, Siddharth Reddy Kandimalla, Christopher Bogart, R. Charles Murray, Michael Hilton 0001, Majd F. Sakr, Carolyn P. Rosé
L@S7
2021 Evaluating the Impact of a Hierarchical Discourse Representation on Entity Coreference Resolution Performance
abstract
Recent work on entity coreference resolution (CR) follows current trends in Deep Learning applied to embeddings and relatively simple task-related features.SOTA models do not make use of hierarchical representations of discourse structure.In this work, we leverage automatically constructed discourse parse trees within a neural approach and demonstrate a significant improvement on two benchmark entity coreference-resolution datasets.We explore how the impact varies depending upon the type of mention.
Sopan Khosla, James Fiacco, Carolyn P. Rosé
NAACL-HLT3
2021 Translational NLP: A New Paradigm and General Principles for Natural Language Processing Research
abstract
Denis Newman-Griffis, Jill Fain Lehman, Carolyn Rosé, Harry Hochheiser. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Denis Newman-Griffis, Jill Fain Lehman, Carolyn P. Rosé, Harry Hochheiser
NAACL-HLT3
2021 Ambiguity in medical concept normalization: An analysis of types and coverage in electronic health record datasets
abstract
OBJECTIVES: Normalizing mentions of medical concepts to standardized vocabularies is a fundamental component of clinical text analysis. Ambiguity-words or phrases that may refer to different concepts-has been extensively researched as part of information extraction from biomedical literature, but less is known about the types and frequency of ambiguity in clinical text. This study characterizes the distribution and distinct types of ambiguity exhibited by benchmark clinical concept normalization datasets, in order to identify directions for advancing medical concept normalization research. MATERIALS AND METHODS: We identified ambiguous strings in datasets derived from the 2 available clinical corpora for concept normalization and categorized the distinct types of ambiguity they exhibited. We then compared observed string ambiguity in the datasets with potential ambiguity in the Unified Medical Language System (UMLS) to assess how representative available datasets are of ambiguity in clinical language. RESULTS: We found that <15% of strings were ambiguous within the datasets, while over 50% were ambiguous in the UMLS, indicating only partial coverage of clinical ambiguity. The percentage of strings in common between any pair of datasets ranged from 2% to only 36%; of these, 40% were annotated with different sets of concepts, severely limiting generalization. Finally, we observed 12 distinct types of ambiguity, distributed unequally across the available datasets, reflecting diverse linguistic and medical phenomena. DISCUSSION: Existing datasets are not sufficient to cover the diversity of clinical concept ambiguity, limiting both training and evaluation of normalization methods for clinical text. Additionally, the UMLS offers important semantic information for building and evaluating normalization methods. CONCLUSIONS: Our findings identify 3 opportunities for concept normalization research, including a need for ambiguity-specific clinical datasets and leveraging the rich semantics of the UMLS in new methods and evaluation measures for normalization.
Denis Newman-Griffis, Guy Divita, Bart Desmet, Ayah Zirikly, Carolyn P. Rosé, Eric Fosler-Lussier
J. Am. Medical Informatics Assoc.5
2021 Improving broad-coverage medical entity linking with semantic type prediction and large-scale datasets
abstract
OBJECTIVES: Biomedical natural language processing tools are increasingly being applied for broad-coverage information extraction-extracting medical information of all types in a scientific document or a clinical note. In such broad-coverage settings, linking mentions of medical concepts to standardized vocabularies requires choosing the best candidate concepts from large inventories covering dozens of types. This study presents a novel semantic type prediction module for biomedical NLP pipelines and two automatically-constructed, large-scale datasets with broad coverage of semantic types. METHODS: We experiment with five off-the-shelf biomedical NLP toolkits on four benchmark datasets for medical information extraction from scientific literature and clinical notes. All toolkits adopt a staged approach of mention detection followed by two stages of medical entity linking: (1) generating a list of candidate concepts, and (2) picking the best concept among them. We introduce a semantic type prediction module to alleviate the problem of overgeneration of candidate concepts by filtering out irrelevant candidate concepts based on the predicted semantic type of a mention. We present MedType, a fully modular semantic type prediction model which we integrate into the existing NLP toolkits. To address the dearth of broad-coverage training data for medical information extraction, we further present WikiMed and PubMedDS, two large-scale datasets for medical entity linking. RESULTS: Semantic type filtering improves medical entity linking performance across all toolkits and datasets, often by several percentage points of F-1. Further, pretraining MedType on our novel datasets achieves state-of-the-art performance for semantic type prediction in biomedical text. CONCLUSIONS: Semantic type prediction is a key part of building accurate NLP pipelines for broad-coverage information extraction from biomedical text. We make our source code and novel datasets publicly available to foster reproducible research.
Shikhar Vashishth, Denis Newman-Griffis, Rishabh Joshi, Ritam Dutt, Carolyn P. Rosé
J. Biomed. Informatics5
2021 Practice-Based Teacher Questioning Strategy Training with ELK: A Role-Playing Simulation for Eliciting Learner Knowledge
abstract
Practice is essential for learning. However, for many interpersonal skills, there often are not enough opportunities and venues for novices to repeatedly practice. Role-playing simulations offer a promising framework to advance practice-based professional training for complex communication skills, in fields such as teaching. In this work, we introduce ELK (Eliciting Learner Knowledge), a role-playing simulation system that helps K-12 teachers develop effective questioning strategies to elicit learners' prior knowledge. We evaluate ELK with 75 pre-service teachers through a mixed-method study. We find that teachers demonstrate a modest increase in effective questioning strategies and develop sympathy towards students after using ELK for 3 rounds. We implement a supplementary activity in ELK in which users evaluate transcripts generated from past role-play sessions. We have tentative evidence that a combination of role-play and evaluating conversation moves may be more effective for learning. We contribute design implications of using role-play systems for communication strategy training.
Xu Wang 0016, Meredith M. Thompson, Dan Roy, Kenneth R. Koedinger, Carolyn P. Rosé, Justin Reich
Proc. ACM Hum. Comput. Interact.6
2020 Towards Open Domain Event Trigger Identification using Adversarial Domain Adaptation
abstract
We tackle the task of building supervised event trigger identification models which can generalize better across domains.Our work leverages the adversarial domain adaptation (ADA) framework to introduce domain-invariance.ADA uses adversarial training to construct representations that are predictive for trigger identification, but not predictive of the example's domain.It requires no labeled data from the target domain, making it completely unsupervised.Experiments with two domains (English literature and news) show that ADA leads to an average F1 score improvement of 3.9 on outof-domain data.Our best performing model (BERT-A) reaches 44-49 F1 across both domains, using no labeled target data.Preliminary experiments reveal that finetuning on 1% labeled data, followed by self-training leads to substantial improvement, reaching 51.5 and 67.2 F1 on literature and news respectively.1
Aakanksha Naik, Carolyn P. Rosé
ACL2
2020 Agent-in-the-Loop: Conversational Agent Support in Service of Reflection for Learning During Collaborative Programming
Sreecharan Sankaranarayanan, Siddharth Reddy Kandimalla, Sahil Hasan, Haokang An, Christopher Bogart, R. Charles Murray, Michael Hilton 0001, Majd F. Sakr, Carolyn P. Rosé
AIED (2)9
2020 Angelfish: Building a Structurally Diverse Clinical Document Corpus
Guy Divita, Bart Desmet, Ayah Zirikly, Luke Breitfeller, Carolyn P. Rosé
AMIA5
2020 Ambiguity in Medical Concept Normalization: An Analysis of Types and Coverage in Electronic Health Record Datasets
Denis Newman-Griffis, Guy Divita, Bart Desmet, Ayah Zirikly, Carolyn P. Rosé, Eric Fosler-Lussier
AMIA5
2020 Re-examining Whether, Why, and How Human-AI Interaction Is Uniquely Difficult to Design
abstract
Artificial Intelligence (AI) plays an increasingly important role in improving HCI and user experience. Yet many challenges persist in designing and innovating valuable human-AI interactions. For example, AI systems can make unpredictable errors, and these errors damage UX and even lead to undesired societal impact. However, HCI routinely grapples with complex technologies and mitigates their unintended consequences. What makes AI different? What makes human-AI interaction appear particularly difficult to design? This paper investigates these questions. We synthesize prior research, our own design and research experience, and our observations when teaching human-AI interaction. We identify two sources of AI's distinctive design challenges: 1) uncertainty surrounding AI's capabilities, 2) AI's output complexity, spanning from simple to adaptive complex. We identify four levels of AI systems. On each level, designers encounter a different subset of the design challenges. We demonstrate how these findings reveal new insights for designers, researchers, and design tool makers in productively addressing the challenges of human-AI interaction going forward.
Qian Yang 0004, Aaron Steinfeld, Carolyn P. Rosé, John Zimmerman
CHI3
2020 Keeping Up Appearances: Computational Modeling of Face Acts in Persuasion Oriented Discussions
abstract
The notion of face refers to the public selfimage of an individual that emerges both from the individual's own actions as well as from the interaction with others.Modeling face and understanding its state changes throughout a conversation is critical to the study of maintenance of basic human needs in and through interaction.Grounded in the politeness theory of Brown and Levinson (1978), we propose a generalized framework for modeling face acts in persuasion conversations, resulting in a reliable coding manual, an annotated corpus, and computational models.The framework reveals insights about differences in face act utilization between asymmetric roles in persuasion conversations.Using computational models, we are able to successfully identify face acts as well as predict a key conversational outcome (e.g.donation success).Finally, we model a latent representation of the conversational state to analyze the impact of predicted face acts on the probability of a positive conversational outcome and observe several correlations that corroborate previous findings.
Ritam Dutt, Rishabh Joshi, Carolyn P. Rosé
EMNLP (1)3
2020 MedFilter: Improving Extraction of Task-relevant Utterances through Integration of Discourse Structure and Ontological Knowledge
abstract
Information extraction from conversational data is particularly challenging because the task-centric nature of conversation allows for effective communication of implicit information by humans, but is challenging for machines.The challenges may differ between utterances depending on the role of the speaker within the conversation, especially when relevant expertise is distributed asymmetrically across roles.Further, the challenges may also increase over the conversation as more shared context is built up through information communicated implicitly earlier in the dialogue.In this paper, we propose the novel modeling approach MEDFILTER, which addresses these insights in order to increase performance at identifying and categorizing task-relevant utterances, and in so doing, positively impacts performance at a downstream information extraction task.We evaluate this approach on a corpus of nearly 7,000 doctor-patient conversations where MEDFILTER is used to identify medically relevant contributions to the discussion (achieving a 10% improvement over SOTA baselines in terms of area under the PR curve).Identifying task-relevant utterances benefits downstream medical processing, achieving improvements of 15%, 105%, and 23% respectively for the extraction of symptoms, medications, and complaints.
Sopan Khosla, Shikhar Vashishth, Jill Fain Lehman, Carolyn P. Rosé
EMNLP (1)4
2020 Incorporating Multimodal Information in Open-Domain Web Keyphrase Extraction
abstract
Open-domain Keyphrase extraction (KPE) onthe Web is a fundamental yet complex NLP task with a wide range of practical applications within the field of Information Retrieval.In contrast to other document types, web page designs are intended for easy navigation and information finding.Effective designs encode within the layout and formatting signals that point to where the important information can be found.In this work, we propose a modeling approach that leverages these multi-modal signals to aid in the KPE task.In particular, we leverage both lexical and visual features (e.g., size, font, position) at the micro-level to enable effective strategy induction, and metalevel features that describe pages at a macrolevel to aid in strategy selection.Our evaluation demonstrates that a combination of effective strategy induction and strategy selection within this approach for the KPE task outperforms state-of-the-art models.A qualitative post-hoc analysis illustrates how these features function within the model.
Yansen Wang, Zhen Fan 0003, Carolyn P. Rosé
EMNLP (1)3
2020 Agent-Based Dynamic Collaboration Support in a Smart Office Space
abstract
For the past 15 years, in computer-supported collaborative learning applications, conversational agents have been used to structure group interactions in online chat-based environments.A series of experimental studies has provided an empirical foundation for the design of chatbased conversational agents that significantly improve learning over no-support control conditions and static-support control conditions.In this demo, we expand upon this foundation, bringing conversational agents to structure group interaction into physical spaces, with the specific goal of facilitating collaboration and learning in workplace scenarios.
Yansen Wang, R. Charles Murray, Haogang Bao, Carolyn P. Rosé
SIGdial4
2020 Using Productive Collaboration Bursts to Analyze Open Source Collaboration Effectiveness
abstract
Developers of open-source software projects tend to collaborate in bursts of activity over a few days at a time, rather than at an even pace. A project might find its productivity suffering if bursts of activity occur when a key person with the right role or right expertise is not available to participate. Open-source projects could benefit from monitoring the way they orchestrate attention among key developers, finding ways to make themselves available to one another when needed. In commercial software development, Sociotechnical Congruence (STC) has been used as a measure to assess whether coordination among developers is sufficient for a given task. However, STC has not previously been successfully applied to open-source projects, in which some industrial assumptions do not apply: management-chosen targets, mandated steady work hours, and top-down task allocation of inputs and targets. In this work we propose an operationalization of STC for open-source software development. We use temporal bursts of activity as a unit of analysis more suited to the natural rhythms of open-source work, as well as open source analogues of other component measures needed for calculating STC. As an illustration, we demonstrate that open-source development on PyPI projects in GitHub is indeed bursty, that activities in the bursts have topical coherence, and we apply our operationalization of STC. We argue that a measure of socio-technical congruence adapted to open source could provide projects with a better way of tracking how effectively they are collaborating when they come together to collaborate.
Samridhi Choudhary, Christopher Bogart, Carolyn P. Rosé, James D. Herbsleb
SANER3
2019 Deep Neural Model Inspection and Comparison via Functional Neuron Pathways
abstract
We introduce a general method for the interpretation and comparison of neural models.The method is used to factor a complex neural model into its functional components, which are comprised of sets of co-firing neurons that cut across layers of the network architecture, and which we call neural pathways.The function of these pathways can be understood by identifying correlated task level and linguistic heuristics in such a way that this knowledge acts as a lens for approximating what the network has learned to apply to its intended task.As a case study for investigating the utility of these pathways, we present an examination of pathways identified in models trained for two standard tasks, namely Named Entity Recognition and Recognizing Textual Entailment.
James Fiacco, Samridhi Choudhary, Carolyn P. Rosé
ACL (1)3
2019 Exploring Numeracy in Word Embeddings
abstract
Word embeddings are now pervasive across NLP subfields as the de-facto method of forming text representataions.In this work, we show that existing embedding models are inadequate at constructing representations that capture salient aspects of mathematical meaning for numbers, which is important for language understanding.Numbers are ubiquitous and frequently appear in text.Inspired by cognitive studies on how humans perceive numbers, we develop an analysis framework to test how well word embeddings capture two essential properties of numbers: magnitude (e.g.3<4) and numeration (e.g.3=three).Our experiments reveal that most models capture an approximate notion of magnitude, but are inadequate at capturing numeration.We hope that our observations provide a starting point for the development of methods which better capture numeracy in NLP systems.
Aakanksha Naik, Abhilasha Ravichander, Carolyn P. Rosé, Eduard H. Hovy
ACL (1)3
2019 An Intelligent-Agent Facilitated Scaffold for Fostering Reflection in a Team-Based Project Course
Sreecharan Sankaranarayanan, Xu Wang 0016, Cameron Dashti, Marshall An, Clarence Ngoh, Michael Hilton 0001, Majd F. Sakr, Carolyn P. Rosé
AIED (2)8
2019 EQUATE: A Benchmark Evaluation Framework for Quantitative Reasoning in Natural Language Inference
abstract
Quantitative reasoning is a higher-order reasoning skill that any intelligent natural language understanding system can reasonably be expected to handle.We present EQUATE 1 (Evaluating Quantitative Understanding Aptitude in Textual Entailment), a new framework for quantitative reasoning in textual entailment.We benchmark the performance of 9 published NLI models on EQUATE, and find that on average, state-of-the-art methods do not achieve an absolute improvement over a majority-class baseline, suggesting that they do not implicitly learn to reason with quantities.We establish a new baseline Q-REAS that manipulates quantities symbolically.In comparison to the best performing NLI model, it achieves success on numerical reasoning tests (+24.2%),but has limited verbal reasoning capabilities (-8.1%).We hope our evaluation framework will support the development of models of quantitative reasoning in language understanding.
Abhilasha Ravichander, Aakanksha Naik, Carolyn P. Rosé, Eduard H. Hovy
CoNLL3
2019 Time-series Insights into the Process of Passing or Failing Online University Courses using Neural-Induced Interpretable Student States
Byungsoo Jeon, Eyal Shafran, Luke Breitfeller, Jason Levin, Carolyn P. Rosé
EDM5
2019 Towards Enabling Feedback on Rhetorical Structure with Neural Sequence Models
abstract
Analysis of student writing, both for assessment and for enabling feedback have been of interest to the field of learning analytics. While much progress can be made through detection of local cues in writing, structured prediction approaches offer capabilities that are particularly well tailored to the needs of models aiming to offer substantive feedback on rhetorical structure. We thus cast the analysis of rhetorical structure in academic writing as a structured prediction task in which we employ models that leverage both local and global cues in writing. In particular, this paper presents a hierarchical neural architecture that performs this task. The evaluation demonstrates that the architecture achieves near-human performance while significantly surpassing state-of-the-art baselines. A multifaceted approach to model interpretation offers insights into the inner workings of the model.
James Fiacco, Elena Cotos, Carolyn P. Rosé
LAK3
2019 UpGrade: Sourcing Student Open-Ended Solutions to Create Scalable Learning Opportunities
abstract
In schools and colleges around the world, open-ended home-work assignments are commonly used. However, such assignments require substantial instructor effort for grading, and tend not to support opportunities for repeated practice. We propose UpGrade, a novel learnersourcing approach that generates scalable learning opportunities using prior student solutions to open-ended problems. UpGrade creates interactive questions that offer automated and real-time feedback, while enabling repeated practice. In a two-week experiment in a college-level HCI course, students answering UpGrade-created questions instead of traditional open-ended assignments achieved indistinguishable learning outcomes in ~30% less time. Further, no manual grading effort is required. To enhance quality control, UpGrade incorporates a psychometric approach using crowd workers' answers to automatically prune out low quality questions, resulting in a question bank that exceeds reliability standards for classroom use.
Xu Wang 0016, Srinivasa Teja Talluri, Carolyn P. Rosé, Kenneth R. Koedinger
L@S3
2019 TDDiscourse: A Dataset for Discourse-Level Temporal Ordering of Events
abstract
Prior work on temporal relation classification has focused extensively on event pairs in the same or adjacent sentences (local), paying scant attention to discourse-level (global) pairs.This restricts the ability of systems to learn temporal links between global pairs, since reliance on local syntactic features suffices to achieve reasonable performance on existing datasets.However, systems should be capable of incorporating cues from documentlevel structure to assign temporal relations.In this work, we take a first step towards discourse-level temporal ordering by creating TDDiscourse, the first dataset focusing specifically on temporal links between event pairs which are more than one sentence apart.We create TDDiscourse by augmenting TimeBank-Dense, a corpus of English news articles, manually annotating global pairs that cannot be inferred automatically from existing annotations.Our annotations double the number of temporal links in TimeBank-Dense, while possessing several desirable properties such as focusing on long-distance pairs and not being automatically inferable.We adapt and benchmark the performance of three stateof-the-art models on TDDiscourse and observe that existing systems indeed find discourselevel temporal ordering harder.
Aakanksha Naik, Luke Breitfeller, Carolyn P. Rosé
SIGdial3
2018 When Optimal Team Formation Is a Choice - Self-selection Versus Intelligent Team Formation Strategies in a Large Online Project-Based Course
Sreecharan Sankaranarayanan, Cameron Dashti, Christopher Bogart, Xu Wang 0016, Majd F. Sakr, Carolyn P. Rosé
AIED (1)6
2018 Stress Test Evaluation for Natural Language Inference
abstract
Natural language inference (NLI) is the task of determining if a natural language hypothesis can be inferred from a given premise in a justifiable manner. NLI was proposed as a benchmark task for natural language understanding. Existing models perform well at standard datasets for NLI, achieving impressive results across different genres of text. However, the extent to which these models understand the semantic content of sentences is unclear. In this work, we propose an evaluation methodology consisting of automatically constructed “stress tests” that allow us to examine whether systems have the ability to make real inferential decisions. Our evaluation of six sentence-encoder models on these stress tests reveals strengths and weaknesses of these models with respect to challenging linguistic phenomena, and suggests important directions for future work in this area.
Aakanksha Naik, Abhilasha Ravichander, Norman M. Sadeh, Carolyn P. Rosé, Graham Neubig
COLING4
2018 Perceptions of Censorship and Moderation Bias in Political Debate Forums
Qinlan Shen, Michael Miller Yoder, Yohan Jo, Carolyn P. Rosé
ICWSM4
2018 Towards domain general detection of transactive knowledge building behavior
abstract
Support of discussion based learning at scale benefits from automated analysis of discussion for enabling effective assignment of students to project teams, for triggering dynamic support of group learning processes, and for assessment of those learning processes. A major limitation of much past work in machine learning applied to automated analysis of discussion is the failure of the models to generalize to data outside of the parameters of the context in which the training data was collected. This limitation means that a separate training effort must be undertaken for each domain in which the models will be used. This paper focuses on a specific construct of discussion based learning referred to as Transactivity and provides a novel machine learning approach with performance that exceeds state-of-the-art performance within the same domain in which it was trained and a new domain, and does not suffer any reduction in performance when transferring to the new domain. These results stand as an advance over past work on automated detection of Transactivity and increase the value of trained models for supporting group learning at scale. Implications for practice in at-scale learning environments are discussed.
James Fiacco, Carolyn P. Rosé
L@S2
2018 Attentive Interaction Model: Modeling Changes in View in Argumentation
abstract
Yohan Jo, Shivani Poddar, Byungsoo Jeon, Qinlan Shen, Carolyn Rosé, Graham Neubig. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Yohan Jo, Shivani Poddar, Byungsoo Jeon, Qinlan Shen, Carolyn P. Rosé, Graham Neubig
NAACL-HLT5
2017 Conflict in Comments: Learning but Lowering Perceptions, with Limits
abstract
Prior work and perception theory suggests that when exposed to discussion related to a particular piece of crowdsourced text content, readers generally perceive that content to be of lower quality than readers who do not see those comments, and that the effect is stronger if the comments display conflict. This paper presents a controlled experiment with over 1000 participants testing to see if this effect carries over to other documents from the same platform, including those with similar content or by the same author. Although we do generally find that perceived quality of the commented-on document is affected, effects do not carry over to the second item and readers are able to judge the second in isolation from the comment on the first. We confirm a prior finding about the negative effects conflict can have on perceived quality but note that readers report learning more from constructive conflict comments.
W. Ben Towne, Carolyn P. Rosé, James D. Herbsleb
CHI2
2017 Modeling Dialogue Acts with Content Word Filtering and Speaker Preferences
abstract
We present an unsupervised model of dialogue act sequences in conversation. By modeling topical themes as transitioning more slowly than dialogue acts in conversation, our model de-emphasizes content-related words in order to focus on conversational function words that signal dialogue acts. We also incorporate speaker tendencies to use some acts more than others as an additional predictor of dialogue act prevalence beyond temporal dependencies. According to the evaluation presented on two dissimilar corpora, the CNET forum and NPS Chat corpus, the effectiveness of each modeling assumption is found to vary depending on characteristics of the data. De-emphasizing content-related words yields improvement on the CNET corpus, while utilizing speaker tendencies is advantageous on the NPS corpus. The components of our model complement one another to achieve robust performance on both corpora and outperform state-of-the-art baseline models.
Yohan Jo, Michael Miller Yoder, Hyeju Jang, Carolyn P. Rosé
EMNLP4
2017 Roles and Success in Wikipedia Talk Pages: Identifying Latent Patterns of Behavior
abstract
In this work we investigate how role-based behavior profiles of a Wikipedia editor, considered against the backdrop of roles taken up by other editors in discussions, predict the success of the editor at achieving an impact on the associated article. We first contribute a new public dataset including a task predicting the success of Wikipedia editors involved in discussion, measured by an operationalization of the lasting impact of their edits in the article. We then propose a probabilistic graphical model that advances earlier work inducing latent discussion roles using the light supervision of success in the negotiation task. We evaluate the performance of the model and interpret findings of roles and group configurations that lead to certain outcomes on Wikipedia.
Korte Maki, Michael Miller Yoder, Yohan Jo, Carolyn P. Rosé
IJCNLP(1)4
2017 Finding Structure in Figurative Language: Metaphor Detection with Topic-based Frames
abstract
In this paper, we present a novel and highly effective method for induction and application of metaphor frame templates as a step toward detecting metaphor in extended discourse.We infer implicit facets of a given metaphor frame using a semisupervised bootstrapping approach on an unlabeled corpus.Our model applies this frame facet information to metaphor detection, and achieves the state-of-the-art performance on a social media dataset when building upon other proven features in a nonlinear machine learning model.In addition, we illustrate the mechanism through which the frame and topic information enable the more accurate metaphor detection.
Hyeju Jang, Korte Maki, Eduard H. Hovy, Carolyn P. Rosé
SIGDIAL Conference4
2017 Supporting Virtual Team Formation through Community-Wide Deliberation
abstract
Team-based learning is a structured, small-group learning method that has been associated with many positive outcomes in traditional classroom settings. However, relatively little research has focused on how to form and support teams within online learning platforms, such as Massive Open Online Courses (MOOCs). A number of challenges arise for team formation in voluntary online classes: students may drop out and leave their team, and even if they do persist with the course, the team may not work together effectively. In this paper, we introduce a team-formation strategy that incorporates a deliberation process, where participants hold discussions in preparation for the collaboration task. First, we present a crowdsourced experiment that compares teams that are formed before or after a community deliberation process. Results demonstrate that teams engaging in a larger community deliberative process prior to team formation exhibit better team performance--as measured by team collaboration product quality--than pre-discussion teams. In a second crowdsourced experiment, we further explore the benefits of community-wide processes by automatically assigning teams based on participants' transactive interaction during deliberation. The results demonstrate advantages in terms of team performance for teams formed based on observed interactions during the community-level deliberation, compared to randomly formed teams. Finally, in a case study, we demonstrate how we successfully adapted the team formation strategy for use in a small MOOC.
Miaomiao Wen, Korte Maki, Steven Dow, James D. Herbsleb, Carolyn P. Rosé
Proc. ACM Hum. Comput. Interact.5
2016 Metaphor Detection with Topic Transition, Emotion and Cognition in Context
abstract
Metaphor is a common linguistic tool in communication, making its detection in discourse a crucial task for natural language understanding.One popular approach to this challenge is to capture semantic incohesion between a metaphor and the dominant topic of the surrounding text.While these methods are effective, they tend to overclassify target words as metaphorical when they deviate in meaning from its context.We present a new approach that (1) distinguishes literal and non-literal use of target words by examining sentence-level topic transitions and (2) captures the motivation of speakers to express emotions and abstract concepts metaphorically.Experiments on an online breast cancer discussion forum dataset demonstrate a significant improvement in metaphor detection over the state-of-theart.These experimental results also reveal a tendency toward metaphor usage in personal topics and certain emotional contexts.
Hyeju Jang, Yohan Jo, Qinlan Shen, Michael Miller Yoder, Seungwhan Moon, Carolyn P. Rosé
ACL (1)6
2016 Automated Feedback on the Quality of Collaborative Processes: An Experience Report
Marcela Borge, Carolyn P. Rosé
EDM2
2016 Expediting Support for Social Learning with Behavior Modeling
Yohan Jo, Gaurav Tomar, Oliver Ferschke, Carolyn P. Rosé, Dragan Gasevic
EDM4
2016 Transactivity as a Predictor of Future Collaborative Knowledge Integration in Team-Based Learning in Online Courses
Miaomiao Wen, Korte Maki, Xu Wang 0016, Steven Dow, James D. Herbsleb, Carolyn P. Rosé
EDM6
2016 Exploring the Effect of Student Confusion in Massive Open Online Courses
Diyi Yang, Robert E. Kraut, Carolyn P. Rosé
EDM3
2016 Pipeline for expediting learning analytics and student support from data in social learning
abstract
An important research problem in learning analytics is to expedite the cycle of data leading to the analysis of student progress and the improvement of student support. For this goal in the context of social learning, we propose a pipeline that includes data infrastructure, learning analytics, and intervention, along with computational models for individual components. Next, we describe an example of applying this pipeline to real data in a case study, whose goal is to investigate the positive effects that goal-setting students have on their peers, which suggests ways in which we might foster these social benefits through intervention.
Yohan Jo, Gaurav Tomar, Oliver Ferschke, Carolyn P. Rosé, Dragan Gasevic
LAK4
2016 Towards triggering higher-order thinking behaviors in MOOCs
abstract
With the aim of better scaffolding discussion to improve learning in a MOOC context, this work investigates what kinds of discussion behaviors contribute to learning. We explored whether engaging in higher-order thinking behaviors results in more learning than paying general or focused attention to course materials. In order to evaluate whether to attribute the effect to engagement in the associated behaviors versus persistent characteristics of the students, we adopted two approaches. First, we used propensity score matching to pair students who exhibit a similar level of involvement in other course activities. Second, we explored individual variation in engagement in higher-order thinking behaviors across weeks. The results of both analyses support the attribution of the effect to the behavioral interpretation. A further analysis using LDA applied to course materials suggests that more social oriented topics triggered richer discussion than more biopsychology oriented topics.
Xu Wang 0016, Miaomiao Wen, Carolyn P. Rosé
LAK3
2016 Initiations and Interruptions in a Spoken Dialog System
Leah Nicolich-Henkin, Carolyn P. Rosé, Alan W. Black
SIGDIAL Conference2
2016 Computational Sociolinguistics: A Survey
abstract
Language is a social phenomenon and variation is inherent to its social nature. Recently, there has been a surge of interest within the computational linguistics (CL) community in the social dimension of language. In this article we present a survey of the emerging field of “computational sociolinguistics” that reflects this increased interest. We aim to provide a comprehensive overview of CL research on sociolinguistic themes, featuring topics such as the relation between language and social identity, language use in social interaction, and multilingual communication. Moreover, we demonstrate the potential for synergy between the research communities involved, by showing how the large-scale data-driven methods that are widely used in CL can complement existing sociolinguistic studies, and how sociolinguistics can inform and challenge the methods and assumptions used in CL studies. We hope to convey the possible benefits of a closer collaboration between the two communities and conclude with a discussion of open challenges.
Dong Nguyen 0002, A. Seza Dogruöz, Carolyn P. Rosé, Franciska de Jong
Comput. Linguistics3
2016 Measuring Similarity Similarly: LDA and Human Perception
abstract
Several intelligent technologies designed to improve navigability in and digestibility of text corpora use topic modeling such as the state-of-the-art Latent Dirichlet Allocation (LDA). This model and variants on it provide lower-dimensional document representations used in visualizations and in computing similarity between documents. This article contributes a method for validating such algorithms against human perceptions of similarity, especially applicable to contexts in which the algorithm is intended to support navigability between similar documents via dynamically generated hyperlinks. Such validation enables researchers to ground their methods in context of intended use instead of relying on assumptions of fit. In addition to the methodology, this article presents the results of an evaluation using a corpus of short documents and the LDA algorithm. We also present some analysis of potential causes of differences between cases in which this model matches human perceptions of similarity more or less well.
W. Ben Towne, Carolyn P. Rosé, James D. Herbsleb
ACM Trans. Intell. Syst. Technol.2
2015 Weakly Supervised Role Identification in Teamwork Interactions
abstract
Diyi Yang, Miaomiao Wen, Carolyn Rosé. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Diyi Yang, Miaomiao Wen, Carolyn P. Rosé
ACL (1)3
2015 The Beginning of a Beautiful Friendship? Intelligent Tutoring Systems and MOOCs
Vincent Aleven, Jonathan Sewall, Octav Popescu, Franceska Xhakaj, Dhruv Chand, Ryan Baker 0001, Yuan Elle Wang, George Siemens, Carolyn P. Rosé, Dragan Gasevic
AIED9
2015 Positive Impact of Collaborative Chat Participation in an edX MOOC
Oliver Ferschke, Diyi Yang, Gaurav Tomar, Carolyn P. Rosé
AIED4
2015 Alleviating the Negative Effect of Up and Downvoting on Help Seeking in MOOC Discussion Forums
Iris K. Howley, Gaurav Tomar, Diyi Yang, Oliver Ferschke, Carolyn P. Rosé
AIED5
2015 Virtual Teams in Massive Open Online Courses
Miaomiao Wen, Diyi Yang, Carolyn P. Rosé
AIED3
2015 Time Series Analysis of Nursing Notes for Mortality Prediction via a State Transition Topic Model
abstract
Accurate mortality prediction is an important task in intensive care units in order to channel prompt care to patients in the most critical condition and to reduce nurses' alarm fatigue. Nursing notes carry valuable information in this regard, but nothing has been reported about the effectiveness of temporal analysis of nursing notes in mortality prediction tasks.
Yohan Jo, Natasha Loghmanpour, Carolyn P. Rosé
CIKM3
2015 Expertise in Cognitive Task Analysis Interviews
Danny Koh, Kenneth R. Koedinger, Carolyn P. Rosé, David F. Feldon
CogSci3
2015 Investigating How Student's Cognitive Behavior in MOOC Discussion Forum Affect Learning Gains
Xu Wang 0016, Diyi Yang, Miaomiao Wen, Kenneth R. Koedinger, Carolyn P. Rosé
EDM5
2015 Exploring the Effect of Confusion in Discussion Forums of Massive Open Online Courses
abstract
Thousands of students enroll in Massive Open Online Courses~(MOOCs) to seek opportunities for learning and self-improvement. However, the learning process often involves struggles with confusion, which may have an adverse effect on the course participation experience, leading to dropout along the way. In this paper, we quantify that effect. We describe a classification model using discussion forum behavior and clickstream data to automatically identify posts that express confusion. We then apply survival analysis to quantify the impact of confusion on student dropout. The results demonstrate that the more confusion students express or are exposed to, the lower the probability of their retention. Receiving support and resolution of confusion helps mitigate this effect. We explore the differential effects of confusion expressed in different contexts and related to different aspects of courses. We conclude with implications for design of interventions towards improving the retention of students in MOOCs.
Diyi Yang, Miaomiao Wen, Iris K. Howley, Robert E. Kraut, Carolyn P. Rosé
L@S5
2015 Metaphor Detection in Discourse
abstract
Understanding contextual information is key to detecting metaphors in discourse.Most current work aims at detecting metaphors given a single sentence, thus focusing mostly on local contextual cues within a short text.In this paper, we present a novel approach that explicitly leverages global context of a discourse to detect metaphors.In addition, we show that syntactic information such as dependency structures can help better describe local contextual information, thus improving detection results when combined.We apply our methods on a newly annotated online discussion forum, and show that our approach outperforms the state-of-the-art baselines in previous literature.
Hyeju Jang, Seungwhan Moon, Yohan Jo, Carolyn P. Rosé
SIGDIAL Conference4
2014 Identifying Latent Study Habits by Mining Learner Behavior Patterns in Massive Open Online Courses
abstract
MOOCs attract diverse users with varying habits. Identifying those patterns through clickstream analysis could enable more effective personalized support for student information seeking and learning in that online context. We propose a novel method to characterize types of sessions in MOOCs by mining the habitual behaviors of students within individual sessions. We model learning sessions as a distribution of activities and activity sequences with a topical N-gram model. The representation offers insights into what groupings of habitual student behaviors are associated with higher or lower success in the course. We also investigate how context information, such as time of day or a user's demographic information, is associated with the types of learning sessions.
Miaomiao Wen, Carolyn P. Rosé
CIKM2
2014 Constrained Question Recommendation in MOOCs via Submodularity
abstract
A recent area in which recommender systems have shown their value is in online discussion forums and question-answer sites. Earlier work in this space has focused on the problem of matching participants to opportunities but has not adequately addressed the problem that in these social contexts, multiple dimensions of constraints must be satisfied, including limitations on capacity and minimal requirements for expertise. In this work, we propose such a constrained question recommendation problem with load balance constraints in discussion forums and use flow based model to generate the optimal solution. In particular, to address the introduced computation complexity, we investigate the concept of submodularity of the objective function and propose a specific submodular method to give an approximated solution. We present experiments conducted on two Massive Open Online Course (MOOC) discussion forum datasets, and demonstrate the effectiveness and efficiency of our submodular method in solving constrained question recommendation tasks.
Diyi Yang, Jingbo Shang, Carolyn P. Rosé
CIKM3
2014 Modeling the Use of Graffiti Style Features to Signal Social Relations within a Multi-Domain Learning Paradigm
abstract
In this paper, we present a series of experiments in which we analyze the usage of graffiti style features for signaling personal gang identification in a large, online street gangs forum, with an accuracy as high as 83% at the gang alliance level and 72% for the specific gang.We then build on that result in predicting how members of different gangs signal the relationship between their gangs within threads where they are interacting with one another, with a predictive accuracy as high as 66% at this thread composition prediction task.Our work demonstrates how graffiti style features signal social identity both in terms of personal group affiliation and between group alliances and oppositions.When we predict thread composition by modeling identity and relationship simultaneously using a multi-domain learning framework paired with a rich feature representation, we achieve significantly higher predictive accuracy than state-of-the-art baselines using one or the other in isolation.
Mario Piergallini, A. Seza Dogruöz, Phani Gadde, David Adamson, Carolyn P. Rosé
EACL5
2014 Sentiment Analysis in MOOC Discussion Forums: What does it tell us?
Miaomiao Wen, Diyi Yang, Carolyn P. Rosé
EDM3
2014 Forum Thread Recommendation for Massive Open Online Courses
Diyi Yang, Mario Piergallini, Iris K. Howley, Carolyn P. Rosé
EDM4
2014 Peer Influence on Attrition in Massively Open Online Courses
Diyi Yang, Miaomiao Wen, Carolyn P. Rosé
EDM3
2014 Effects of social presence and social role on help-seeking and learning
abstract
The unique social presence of robots can be leveraged in learning situations to reduce student evaluation anxiety, while still providing instructional guidance on multiple levels of communication. Furthermore, social role of the instructor can also impact the prevalence of evaluation apprehension. In this study, we examine how human and robot social role affects help-seeking behaviors and learning outcomes in a one-on-one tutoring setting. Our results show that help-seeking is a moderator of the significant relationship between condition and learning, with the "human teacher" condition resulting in significantly less learning (and marginally less help-seeking) than the "human assistant" and both robot conditions.
Iris K. Howley, Takayuki Kanda 0001, Kotaro Hayashi, Carolyn P. Rosé
HRI4
2014 Linguistic Reflections of Student Engagement in Massive Open Online Courses
Miaomiao Wen, Diyi Yang, Carolyn P. Rosé
ICWSM3
2014 Predicting Student Learning from Conversational Cues
David Adamson, Akash Bharadwaj, Ashudeep Singh, Colin Ashe, David J. Yaron, Carolyn P. Rosé
Intelligent Tutoring Systems6
2014 Learning analytics and machine learning
abstract
Learning analytics (LA) as a field remains in its infancy. Many of the techniques now prominent from practitioners have been drawn from various fields, including HCI, statistics, computer science, and learning sciences. In order for LA to grow and advance as a discipline, two significant challenges must be met: 1) development of analytics methods and techniques that are native to the LA discipline, and 2) practitioners in LA to develop algorithms and models that reflect the social and computational dimensions of analytics. This workshop introduces researchers in learning analytics to machine learning (ML) and the opportunities that ML can provide in building next generation analysis models.
Dragan Gasevic, Carolyn P. Rosé, George Siemens, Annika Wolff, Zdenek Zdráhal
LAK2
2014 Social factors that contribute to attrition in MOOCs
abstract
In this paper, we explore student dropout behavior in a Massively Open Online Course (MOOC). We use a survival model to measure the impact of three social factors that make predictions about attrition along the way for students who have participated in the course discussion forum.
Carolyn P. Rosé, Ryan Carlson, Diyi Yang, Miaomiao Wen, Lauren B. Resnick, Pam Goldman, Jennifer Sherer
L@S1
2014 Question recommendation with constraints for massive open online courses
abstract
Massive Open Online Courses (MOOCs) have experienced a recent boom in interest. Problems students struggle with in the discussion forums, such as difficultly in finding interesting discussion opportunities or attracting helpers to address posted problems, provide new opportunities for recommender systems. In contrast to traditional product recommendation, question recommendation in discussion forums should simultaneously consider constraints on both students and questions. These considerations include (1) Load Balancing - students should not be over-burdened with too many requests; and (2) Expertise Matching - students should not be requested to address problems they are not capable of addressing. In this work, we formulate a novel constrained question recommendation problem to address the above considerations. We design a context-aware matrix factorization model to predict students' preferences over questions, then build a max cost flow model to manage the constraints. Experimental results conducted on three MOOC datasets demonstrate that our method significantly outperforms baseline methods in optimizing overall forum welfare, and in predicting which specific questions students might be interested in.
Diyi Yang, David Adamson, Carolyn P. Rosé
RecSys3
2014 Triggering effective social support for online groups
abstract
Conversational agent technology is an emerging paradigm for creating a social environment in online groups that is conducive to effective teamwork. Prior work has demonstrated advantages in terms of learning gains and satisfaction scores when groups learning together online have been supported by conversational agents that employ Balesian social strategies. This prior work raises two important questions that are addressed in this article. The first question is one of generality. Specifically, are the positive effects of the designed support specific to learning contexts? Or are they in evidence in other collaborative task domains as well? We present a study conducted within a collaborative decision-making task where we see that the positive effects of the Balesian social strategies extend to this new context. The second question is whether it is possible to increase the effectiveness of the Balesian social strategies by increasing the context sensitivity with which the social strategies are triggered. To this end, we present technical work that increases the sensitivity of the triggering. Next, we present a user study that demonstrates an improvement in performance of the support agent with the new, more sensitive triggering policy over the baseline approach from prior work. The technical contribution of this article is that we extend prior work where such support agents were modeled using a composition of conversational behaviors integrated within an event-driven framework. Within the present approach, conversation is orchestrated through context-sensitive triggering of the composed behaviors. The core effort involved in applying this approach involves building a set of triggering policies that achieve this orchestration in a time-sensitive and coherent manner. In line with recent developments in data-driven approaches for building dialog systems, we present a novel technique for learning behavior-specific triggering policies, deploying it as part of our efforts to improve a socially capable conversational tutor agent that supports collaborative learning.
Rohit Kumar 0001, Carolyn P. Rosé
ACM Trans. Interact. Intell. Syst.2
2013 Recognizing Rare Social Phenomena in Conversation: Empowerment Detection in Support Group Chatrooms
Elijah Mayfield, David Adamson, Carolyn P. Rosé
ACL (1)3
2013 Automatically Generating Discussion Questions
David Adamson, Divyanshu Bhartiya, Biman Gujral, Radhika Kedia, Ashudeep Singh, Carolyn P. Rosé
AIED6
2013 Impact of Group Norms in Eliciting Response in a Goal Driven Virtual Community
abstract
With the proliferation of social media into our daily lives, online communities have become an important platform for collaborative learning and education. To connect users with varying knowledge levels and increase the net learning throughput, these communities often follow a question-answer based approach. Understanding what drives attention to help-seeking questions can reduce the amount of questions that go unnoticed or remain unanswered by the community. In this paper we discuss an important feature that affects the activity of the community, namely the community norms. We present a machine learning based trigger-driven feedback model that functions by (i) differentiating between help-seeking questions and follow-up posts – i.e. posts that are part of an ongoing discussion, and (ii) a dynamic intervention scheme to help improve question formulation. Our findings show that adhering to the community norms significantly increases the chance of eliciting a response.
Sumeet Jain, Tanmay Sinha, Achal Shah, Chandramouli Sharma, Carolyn P. Rosé
ICCE5
2013 What's in a Domain? Multi-Domain Learning for Multi-Attribute Data
Mahesh Joshi, Mark Dredze, William W. Cohen, Carolyn P. Rosé
HLT-NAACL4
2012 Detecting offensive tweets via topical feature discovery over a large scale twitter corpus
abstract
In this paper, we propose a novel semi-supervised approach for detecting profanity-related offensive content in Twitter. Our approach exploits linguistic regularities in profane language via statistical topic modeling on a huge Twitter corpus, and detects offensive tweets using automatically these generated features. Our approach performs competitively with a variety of machine learning (ML) algorithms. For instance, our approach achieves a true positive rate (TP) of 75.1% over 4029 testing tweets using Logistic Regression, significantly outperforming the popular keyword matching baseline, which has a TP of 69.7%, while keeping the false positive rate (FP) at the same level as the baseline at about 3.77%. Our approach provides an alternative to large scale hand annotation efforts required by fully supervised learning approaches.
Guang Xiang, Jason I. Hong, Carolyn P. Rosé
CIKM5
2012 An Unsupervised Dynamic Bayesian Network Approach to Measuring Speech Style Accommodation
Mahaveer Jain, John W. McDonough, Gahgene Gweon, Bhiksha Raj, Carolyn P. Rosé
EACL5
2012 Multi-Domain Learning: When Do Domains Matter?
Mahesh Joshi, Mark Dredze, William W. Cohen, Carolyn P. Rosé
EMNLP-CoNLL4
2012 Discovering habits of effective online support group chatrooms
abstract
For users of online support groups, prior research has suggested that a positive social environment is a key enabler of coping. Typically, demonstrating such claims about social interaction would be approached through the lens of sentiment analysis. In this work, we argue instead for a multifaceted view of emotional state, which incorporates both a static view of emotion (sentiment) with a dynamic view based on the behaviors present in a text. We codify this dynamic view through data annotations marking information sharing, sentiment, and coping efficacy. Through machine learning analysis of these annotations, we demonstrate that while sentiment predicts a user's stress at the beginning of a chat, dynamic views of efficacy are stronger indicators of stress reduction.
Elijah Mayfield, Miaomiao Wen, Mitch Golant, Carolyn P. Rosé
GROUP4
2012 Understanding participant behavior trajectories in online health support groups using automatic extraction methods
abstract
This paper presents an automatic analysis method that enables efficient examination of participant behavior trajectories in online communities, which offers the opportunity to examine behavior over time at a level of granularity that has previously only been possible in small scale case study analyses. We provide an empirical validation of its performance. We then illustrate how this method offers insights into behavior patterns that enable avoiding faulty oversimplified assumptions about participation, such as that it follows a consistent trend over time. In particular, we use this method to investigate the connection between user behavior and distressful cancer events and demonstrate how this tool could assist in cancer story summarization.
Miaomiao Wen, Carolyn P. Rosé
GROUP2
2012 You Too?! Mixed-Initiative LDA Story Matching to Help Teens in Distress
Karthik Dinakar, Birago Jones, Henry Lieberman, Rosalind W. Picard, Carolyn P. Rosé, Matthew Thoman, Roi Reichart
ICWSM5
2012 A Supervised Approach to Predict Company Acquisition with Factual and Topic Features Using Profiles and News Articles on TechCrunch
Guang Xiang, Miaomiao Wen, Jason I. Hong, Carolyn P. Rosé
ICWSM5
2012 Coordinating Multi-dimensional Support in Collaborative Conversational Agents
David Adamson, Carolyn P. Rosé
ITS2
2012 Building a Conversational SimStudent
Ryan Carlson, Victoria Keiser, Noboru Matsuda, Kenneth R. Koedinger, Carolyn P. Rosé
ITS5
2012 Towards Academically Productive Talk Supported by Conversational Agents
Gregory Dyke, David Adamson, Iris K. Howley, Carolyn P. Rosé
ITS4
2012 Group Composition and Intelligent Dialogue Tutors for Impacting Students' Academic Self-efficacy
Iris K. Howley, David Adamson, Gregory Dyke, Elijah Mayfield, Jack L. Beuth, Carolyn P. Rosé
ITS6
2012 Hierarchical Conversation Structure Prediction in Multi-Party Chat
Elijah Mayfield, David Adamson, Carolyn P. Rosé
SIGDIAL Conference3
2011 Recognizing Authority in Dialogue with an Integer Linear Programming Constrained Model
Elijah Mayfield, Carolyn P. Rosé
ACL2
2011 Comparing Triggering Policies for Social Behaviors
Rohit Kumar 0001, Carolyn P. Rosé
SIGDIAL Conference2
2011 CANTINA+: A Feature-Rich Machine Learning Framework for Detecting Phishing Web Sites
abstract
Phishing is a plague in cyberspace. Typically, phish detection methods either use human-verified URL blacklists or exploit Web page features via machine learning techniques. However, the former is frail in terms of new phish, and the latter suffers from the scarcity of effective features and the high false positive rate (FP). To alleviate those problems, we propose a layered anti-phishing solution that aims at (1) exploiting the expressiveness of a rich set of features with machine learning to achieve a high true positive rate (TP) on novel phish, and (2) limiting the FP to a low level via filtering algorithms. Specifically, we proposed CANTINA+, the most comprehensive feature-based approach in the literature including eight novel features, which exploits the HTML Document Object Model (DOM), search engines and third party services with machine learning techniques to detect phish. Moreover, we designed two filters to help reduce FP and achieve runtime speedup. The first is a near-duplicate phish detector that uses hashing to catch highly similar phish. The second is a login form filter, which directly classifies Web pages with no identified login form as legitimate. We extensively evaluated CANTINA+ with two methods on a diverse spectrum of corpora with 8118 phish and 4883 legitimate Web pages. In the randomized evaluation, CANTINA+ achieved over 92% TP on unique testing phish and over 99% TP on near-duplicate testing phish, and about 0.4% FP with 10% training phish. In the time-based evaluation, CANTINA+ also achieved over 92% TP on unique testing phish, over 99% TP on near-duplicate testing phish, and about 1.4% FP under 20% training phish with a two-week sliding window. Capable of achieving 0.4% FP and over 92% TP, our CANTINA+ has been demonstrated to be a competitive anti-phishing solution.
Guang Xiang, Jason I. Hong, Carolyn P. Rosé, Lorrie Faith Cranor
ACM Trans. Inf. Syst. Secur.3
2010 A Hierarchical Adaptive Probabilistic Approach for Zero Hour Phish Detection
Guang Xiang, Bryan A. Pendleton, Jason I. Hong, Carolyn P. Rosé
ESORICS4
2010 Using feature construction to avoid large feature spaces in text classification
abstract
Feature space design is a critical part of machine learning. This is an especially difficult challenge in the field of text classification, where an arbitrary number of features of varying complexity can be extracted from documents as a preprocessing step. A challenge for researchers has consistently been to balance expressiveness of features with the size of the corresponding feature space, due to issues with data sparsity that arise as feature spaces grow larger. Drawing on past successes utilizing genetic programming in similar problems outside of text classification, we propose and implement a technique for constructing complex features from simpler features, and adding these more complex features into a combined feature space which can then be utilized by more sophisticated machine learning classifiers. Applying this technique to a sentiment analysis problem, we show encouraging improvement in classification accuracy, with a small and constant increase in feature space size. We also show that the features we generate carry far more predictive power than any of the simple features they contain.
Elijah Mayfield, Carolyn P. Rosé
GECCO2
2010 A Classification Approach for Risk Prognosis of Patients on Mechanical Ventricular Assistance
abstract
The identification of optimal candidates for ventricular assist device (VAD) therapy is of great importance for future widespread application of this life-saving technology. During recent years, numerous traditional statistical models have been developed for this task. In this study, we compared three different supervised machine learning techniques for risk prognosis of patients on VAD: Decision Tree, Support Vector Machine (SVM) and Bayesian Tree-Augmented Network, to facilitate the candidate identification. A predictive (C4.5) decision tree model was ultimately developed based on 6 features identified by SVM with assistance of recursive feature elimination. This model performed better compared to the popular risk score of Lietz et al. with respect to identification of high-risk patients and earlier survival differentiation between high- and low- risk candidates.
Carolyn P. Rosé, Antonio Ferreira 0002, Dennis M. McNamara, Robert L. Kormos, James F. Antaki
ICMLA2
2010 Exploring the Effectiveness of Social Capabilities and Goal Alignment in Computer Supported Collaborative Learning
Hua Ai, Rohit Kumar 0001, Dong Nguyen 0002, Amrut Nagasunder, Carolyn P. Rosé
Intelligent Tutoring Systems (2)5
2010 Student Dispositions and Help-Seeking in Collaborative Learning
Iris K. Howley, Carolyn P. Rosé
Intelligent Tutoring Systems (2)2
2010 Socially Capable Conversational Tutors Can Be Effective in Collaborative Learning Situations
Rohit Kumar 0001, Hua Ai, Jack L. Beuth, Carolyn P. Rosé
Intelligent Tutoring Systems (1)4
2010 DesignWebs: A Tool for Automatic Construction of Interactive Conceptual Maps from Document Collections
Sharad V. Oberoi, Dong Nguyen 0002, Gahgene Gweon, Susan Finger, Carolyn P. Rosé
Intelligent Tutoring Systems (2)5
2010 Engaging learning groups using Social Interaction Strategies
Rohit Kumar 0001, Carolyn P. Rosé
HLT-NAACL2
2010 Making Conversational Structure Explicit: Identification of Initiation-response Pairs within Online Discussions
Yi-Chia Wang, Carolyn P. Rosé
HLT-NAACL2
2009 Engaging Collaborative Learners with Helping Agents
abstract
We present the results of a study in which we contrast alternative forms of collaborative learning support in the midst of a collaborative design task in which students negotiate between increasing power and increasing environmental friendliness. Our research question is whether interactive instructional support in collaborative learning environments is more effective when it is offered as solicited or unsolicited help. The finding from our classroom study is that dialogue-based support is more effective in this collaborative context when invitations for help in the form of pointer hints are offered automatically, but dialogue agents are only provided when the invitation is explicitly accepted.
Sourish Chaudhuri, Rohit Kumar 0001, Iris K. Howley, Carolyn P. Rosé
AIED4
2009 Towards Automatic Assessment for Project Based Learning Groups
abstract
Project course instructors routinely perform their formal assessments based on impressions formed from their mostly indirect experience with the groups they oversee. Nevertheless, even with their limited vantage point, instructors trust their ability to make assessments and regulate group work. In this paper we present a 5 dimensional assessment framework based on data from an interview study in which we investigate the assessment goals that project course instructors have. We use this framework to identify specifically where instructors' assessments about students diverge most from that of direct observers of group work. We then demonstrate that indicators extracted automatically from recorded speech from group meetings frequently correlate better with objective observer rating of students than that of the instructor.
Gahgene Gweon, Rohit Kumar 0001, Soojin Jun, Carolyn P. Rosé
AIED4
2009 Motivation and Collaboration On-Line
abstract
This poster describes a controlled study that examines how individual motivation orientations affect computer-supported collaborative learning (CSCL) activities with an adaptive thermodynamics dialogue tutor. Our results show that motivation orientation has an effect on perceptions of one's learning partner, and possibly one's self-confidence.
Iris K. Howley, Sourish Chaudhuri, Rohit Kumar 0001, Carolyn P. Rosé
AIED4
2009 Detecting and Understanding the Impact of Cognitive and Interpersonal Conflict in Computer Supported Collaborative Learning Environments
David Prata, Ryan Baker 0001, Evandro de Barros Costa, Carolyn P. Rosé, Yue Cui 0004
EDM4
2008 Investigating the effect of discussion forum interface affordances on patterns of conversational interactions
abstract
We investigate how the affordances provided by alternative interfaces for on-line discussion forums affect the structure of the discourse that unfolds. In order to investigate this impact, we compare the predictive power of time related and text similarity related features for identifying parent-child links between messages. The results from this work using this methodology suggest that interfaces that make parent-child relationships between messages explicit and do not constrain the choice of previous messages that users can reply to allow patterns of conversational behavior that violate the assumptions of traditional, tree-structured models of discourse where time related and similarity related features are highly predictive. An implication for future work is that because there is evidence that interface affordances affect the form of conversational contributions, techniques that process on-line communication data may need to be adapted for different communication interfaces.
Yi-Chia Wang, Mahesh Joshi, Carolyn P. Rosé
CSCW3
2008 Recovering Implicit Thread Structure in Newsgroup Style Conversations
Yi-Chia Wang, Mahesh Joshi, William W. Cohen, Carolyn P. Rosé
ICWSM4
2008 It's Not Easy Being Green: Supporting Collaborative "Green Design" Learning
Sourish Chaudhuri, Rohit Kumar 0001, Mahesh Joshi, Elon Terrell, Fred Higgs, Vincent Aleven, Carolyn P. Rosé
Intelligent Tutoring Systems7
2008 Story Generation to Accelerate Math Problem Authoring for Practice and Assessment
Yue Cui 0004, Rohit Kumar 0001, Carolyn P. Rosé, Kenneth R. Koedinger
Intelligent Tutoring Systems3
2008 An Authoring Tool That Facilitates the Rapid Development of Dialogue Agents for Intelligent Tutoring Systems
Yue Cui 0004, Carolyn P. Rosé
Intelligent Tutoring Systems2
2008 Supporting the Guide on the SIDE
Moonyoung Kang, Sourish Chaudhuri, Rohit Kumar 0001, Yi-Chia Wang, Eric R. Rosé, Carolyn P. Rosé, Yue Cui 0004
Intelligent Tutoring Systems6
2007 A Feature Based Approach to Leveraging Context for Classifying Newsgroup Style Discussion Segments
Yi-Chia Wang, Mahesh Joshi, Carolyn P. Rosé
ACL3
2007 Tools for Authoring a Dialogue Agent that Participates in Learning Studies
Pamela W. Jordan, Brian Hall, Michael A. Ringenberg, Yui Cue, Carolyn P. Rosé
AIED5
2007 Tutorial Dialogue as Adaptive Collaborative Learning Support
Rohit Kumar 0001, Carolyn P. Rosé, Yi-Chia Wang, Mahesh Joshi, Allen Robinson
AIED2
2007 Using Machine Learning Techniques to Analyze and Support Mediation of Student E-Discussions
Bruce M. McLaren, Oliver Scheuer, Maarten de Laat, Rakheli Hever, Reuma De Groot, Carolyn P. Rosé
AIED6
2007 TagHelper Tools: Tools for Supporting the Analysis of Verbal Data
Carolyn P. Rosé
AIED1
2007 Context Based Classification for Automatic Collaborative Learning Process Analysis
Yi-Chia Wang, Mahesh Joshi, Carolyn P. Rosé, Frank Fischer 0001, Armin Weinberger, Karsten Stegmann
AIED3
2007 Supporting Collaborative Idea Generation: A Closer Look Using Statistical Process Analysis Techniques
Hao-Chuan Wang, Carolyn P. Rosé
AIED2
2007 Modeling the impact of shared visual information on collaborative reference
abstract
A number of recent studies have demonstrated that groups benefit considerably from access to shared visual information. This is due, in part, to the communicative efficiencies provided by the shared visual context. However, a large gap exists between our current theoretical understanding and our existing models. We address this gap by developing a computational model that integrates linguistic cues with visual cues in a way that effectively models reference during tightly-coupled, task-oriented interactions. The results demonstrate that an integrated model significantly outperforms existing language-only and visual-only models. The findings can be used to inform and augment the development of conversational agents, applications that dynamically track discourse and collaborative interactions, and dialogue managers for natural language interfaces.
Darren Gergle, Carolyn P. Rosé, Robert E. Kraut
CHI2
2007 Sharing a single expert among multiple partners
abstract
Expertise to assist people on complex tasks is often in short supply. One solution to this problem is to design systems that allow remote experts to help multiple people in simultaneously. As a first step towards building such a system, we studied experts' attention and communication as they assisted two novices at the same time in a co-located setting. We compared simultaneous instruction when the novices are being instructed to do the same task or different tasks. Using machine learning, we attempted to identify speech markers of upcoming attention shifts that could serve as input to a remote assistance system.
Jeffrey Wong, Lui Min Oh, Jiazhi Ou, Carolyn P. Rosé, Jie Yang 0001, Susan R. Fussell
CHI4
2007 A Hybrid Ontology Directed Feedback Selection Algorithm for Supporting Creative Problem Solving Dialogues
Hao-Chuan Wang, Rohit Kumar 0001, Carolyn P. Rosé, Tsai-Yen Li, Chun-Yen Chang 0001
IJCAI3
2006 Talk to me: foundations for successful individual-group interactions in online communities
abstract
People come to online communities seeking information, encouragement, and conversation. When a community responds, participants benefit and become more committed. Yet interactions often fail. In a longitudinal sample of 6,172 messages from 8 Usenet newsgroups, 27% of posts received no response. The information context, posters' prior engagement in the community, and the content of their posts all influenced the likelihood that they received a reply, and, as a result, their willingness to continue active participation. Posters were less likely to get a reply if they were newcomers. Posting ontopic, introducing oneself via autobiographical testimonials, asking questions, using less complex language and other features of the messages, increased replies. Results suggest ways that developers might increase the ability of online communities to support successful individual-group interactions.
Jaime Arguello, Brian S. Butler, Elisabeth Joyce, Robert E. Kraut, Kimberly S. Ling, Carolyn P. Rosé
CHI6
2006 Providing support for adaptive scripting in an on-line collaborative learning environment
abstract
This paper describes results from a series of experimental studies to explore issues related to structuring productive group dynamics for collaborative learning using an adaptive support mechanism. The first study provides evidence in favor of the feasibility of the endeavor by demonstrating with a tightly controlled study that even without adaptive support, problem solving in pairs is significantly more effective for learning than problem solving alone. The results from a second study offer guidelines for strategic matching of students with learning partners. Furthermore, the results reveal specific areas for needed support. Based on the results from the second study, we present the design of an adaptive support mechanism, which we evaluate in a third study. The results from the third study provide evidence that certain aspects of our design for adaptive support in the form of strategic prompts are effective for manipulating student behavior in productive ways and for supporting learning. These results also motivate specific modifications to the original design.
Gahgene Gweon, Carolyn P. Rosé, Regan Carey, Zachary Sam Zaiss
CHI2
2006 Identification of confusion and surprise in spoken dialog using prosodic features
abstract
Sensitivity to a user's emotional state offers promise in improving the state of the art in spoken dialog systems. In this work, we attempt to detect the speaker's states of confusion and surprise using prosodic features from his/her utterances. We have collected a corpus of utterances in realistic settings using an experimental methodology aimed at eliciting confusion and surprise from users. Classification experiments have yielded up to a 27.2% improvement over baseline performance using F0 and power features. We achieved the greatest success at classification of emotions that were most successfully elicited.
Rohit Kumar 0001, Carolyn P. Rosé, Diane J. Litman
INTERSPEECH2
2006 Evaluating the Effectiveness of Tutorial Dialogue Instruction in an Exploratory Learning Context
Rohit Kumar 0001, Carolyn P. Rosé, Vincent Aleven, Ana Iglesias 0002, Allen Robinson
Intelligent Tutoring Systems2
2006 VIBRANT: A Brainstorming Agent for Computer Supported Creative Problem Solving
Hao-Chuan Wang, Tsai-Yen Li, Carolyn P. Rosé, Chun-Chieh Huang, Chun-Yen Chang 0001
Intelligent Tutoring Systems3
2006 Museli: A Multi-Source Evidence Integration Approach to Topic Segmentation of Spontaneous Dialogue
Jaime Arguello, Carolyn P. Rosé
HLT-NAACL2
2006 InfoMagnets: Making Sense of Corpus Data
Jaime Arguello, Carolyn P. Rosé
HLT-NAACL2
2005 Authoring plug-in tutor agents by demonstration: Rapid, rapid tutor development
Vincent Aleven, Carolyn P. Rosé
AIED2
2005 Towards Data-Driven Design of a Peer Collaborative Agent
Gahgene Gweon, Carolyn P. Rosé, Regan Carey, Zachary Sam Zaiss
AIED2
2005 A First Evaluation of the Instructional Value of Negotiable Problem Solving Goals on the Exploratory Learning Continuum
Carolyn P. Rosé, Vincent Aleven, Regan Carey, Allen Robinson
AIED1
2005 Automatic and Semi-Automatic Skill Coding With a View Towards Supporting On-Line Assessment
Carolyn P. Rosé, Pinar Donmez, Gahgene Gweon, Andrea Knight, Brian Junker, William W. Cohen, Kenneth R. Koedinger, Neil T. Heffernan
AIED1
2005 The Necessity of a Meeting Recording and Playback System, and the Benefit of Topic-Level Annotations to Meeting Browsing
Satanjeev Banerjee, Carolyn P. Rosé, Alexander I. Rudnicky
INTERACT2
2005 Supporting Efficient and Reliable Content Analysis Using Automatic Text Processing Technology
Gahgene Gweon, Carolyn P. Rosé, Jörg Wittwer, Matthias Nückles
INTERACT2
2005 Interactivity and Expectation: Eliciting Learning Oriented Behavior with Tutorial Dialogue Systems
Carolyn P. Rosé, Cristen Torrey
INTERACT1
2004 Workshop on Dialog-Based Intelligent Tutoring Systems: State of the Art and New Research Directions
Neil T. Heffernan, Peter M. Hastings, Gregory Aist, Vincent Aleven, Ivon Arroyo, Paul Brna, Mark G. Core, Martha W. Evens, Reva Freedman, Michael Glass, Arthur C. Graesser, Kenneth R. Koedinger, Pamela W. Jordan, Diane J. Litman, Evelyn Lulis, Helen Pain, Carolyn P. Rosé, Beverly P. Woolf, Claus Zinn
Intelligent Tutoring Systems17
2004 Spoken Versus Typed Human and Computer Dialogue Tutoring
Diane J. Litman, Carolyn P. Rosé, Katherine Forbes-Riley, Kurt VanLehn, Dumisizwe Bhembe, Scott Silliman
Intelligent Tutoring Systems2
2004 DReSDeN: Towards a Trainable Tutorial Dialogue Manager to Support Negotiation Dialogues for Learning and Reflection
Carolyn P. Rosé, Cristen Torrey
Intelligent Tutoring Systems1
2004 CycleTalk: Toward a Dialogue Agent That Guides Design with an Articulate Simulator
Carolyn P. Rosé, Cristen Torrey, Vincent Aleven, Allen Robinson, Chih Wu, Kenneth D. Forbus
Intelligent Tutoring Systems1
2003 A Hybrid Approach to Content Analysis for Automatic Essay Grading
Carolyn P. Rosé, Antonio Roque, Dumisizwe Bhembe, Kurt VanLehn
HLT-NAACL1
2002 A Hybrid Language Understanding Approach for Robust Selection of Tutoring Goals
Carolyn P. Rosé, Dumisizwe Bhembe, Antonio Roque, Stephanie Siler, Ramesh Srivastava, Kurt VanLehn
Intelligent Tutoring Systems1
2002 The Architecture of Why2-Atlas: A Coach for Qualitative Physics Essay Writing
Kurt VanLehn, Pamela W. Jordan, Carolyn P. Rosé, Dumisizwe Bhembe, Michael Böttner, Andy Gaydos, Maxim Makatchev, Umarani Pappuswamy, Michael A. Ringenberg, Antonio Roque, Stephanie Siler, Ramesh Srivastava
Intelligent Tutoring Systems3
2000 ITS Tools for Natural Language Dialogue: A Domain-Independent Parser and Planner
Reva Freedman, Carolyn P. Rosé, Michael A. Ringenberg, Kurt VanLehn
Intelligent Tutoring Systems2
2000 Fading and Deepening: The Next Steps for Andes and other Model-Tracing Tutors
Kurt VanLehn, Reva Freedman, Pamela W. Jordan, R. Charles Murray, Remus Osan, Michael A. Ringenberg, Carolyn P. Rosé, Kay G. Schulze, Robert Shelby, Donald Treacy, Anders Weinstein, Mary Wintersgill
Intelligent Tutoring Systems7
1997 An Efficient Distribution of Labor in a Two Stage Robust Interpretation Process
Carolyn P. Rosé, Alon Lavie
EMNLP1
1996 Using Discourse Predictions for Ambiguity Resolution
Yan Qu, Carolyn P. Rosé, Barbara Di Eugenio
COLING2
1995 Discourse Processing of Dialogues with Multiple Threads
abstract
In this paper we will present our ongoing work on a plan-based discourse processor developed in the context of the Enthusiast Spanish to English translation system as part of the JANUS multi-lingual speech-to-speech translation system. We will demonstrate that theories of discourse which postulate a strict tree structure of discourse on either the intentional or attentional level are not totally adequate for handling spontaneous dialogues. We will present our extension to this approach along with its implementation in our plan-based discourse processor. We will demonstrate that the implementation of our approach outperforms an implementation based on the strict tree structure approach.
Carolyn P. Rosé, Barbara Di Eugenio, Lori S. Levin, Carol Van Ess-Dykema
ACL1
1994 JANUS 93: towards spontaneous speech translation
abstract
We present first results from our efforts toward translation of spontaneously spoken speech. Improvements include increasing coverage, robustness, generality and speed of JANUS, the speech-to-speech translation system of Carnegie Mellon and Karlsruhe University. The recognition and machine translation engine have been upgraded to deal with requirements introduced by spontaneous human to human dialogs. To allow for development and evaluation of our system on adequate data, a large database with spontaneous scheduling dialogs is being gathered for English, German and Spanish.>
Monika Woszczyna, Naomi Aoki-Waibel, Finn Dag Buø, Noah Coccaro, Keiko Horiguchi, Thomas Kemp, Alon Lavie, Arthur E. McNair, Thomas Polzin, Ivica Rogina, Carolyn P. Rosé, Tanja Schultz, Bernhard Suhm, Masaru Tomita, Alex Waibel
ICASSP (1)11
1993 Recent advances in JANUS: a speech translation system
abstract
We present recent advances from our efforts in increasing coverage, robustness, generality and speed of JANUS, CMU's speech-tospeech translation system. JANUS is a speaker-independent system translating spoken utterances in English and also in German into one of German, English or Japanese. The system has been designed around the task of conference registration (CR). It has initially been built based on a speech database of 12 read dialogs, encompassing a vocabulary of around 500 words. We have since been expanding the system along several dimensions to improve speed, robustness and coverage and to move toward spontaneous input. 1. INTRODUCTION In this paper we describe recent improvements of JANUS, a speech to speech translation system. Improvements have been made mainly along the following dimensions: 1.) better context-dependent modeling improves performance in the speech recognition module, 2.) improved language models, smoothing, and word equivalence classes improve coverage and ...
Monika Woszczyna, Noah Coccaro, Andreas Eisele 0001, Alon Lavie, Arthur E. McNair, Thomas Polzin, Ivica Rogina, Carolyn P. Rosé, Tilo Sloboda, Masaru Tomita, J. Tsutsumi, Naomi Aoki-Waibel, Alex Waibel, Wayne H. Ward
EUROSPEECH8