Joel R. Tetreault

dblp:40/4518 · DBLP profile ↗
← Back
54ranked-venue papers
8as first author
11since 2021 · last 2025
0009-0003-3552-842XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 46 · 7 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3
YearPublicationVenuePosition
2025 CEHA: A Dataset of Conflict Events in the Horn of Africa
abstract
Natural Language Processing (NLP) of news articles can play an important role in understanding the dynamics and causes of violent conflict. Despite the availability of datasets categorizing various conflict events, the existing labels often do not cover all of the fine-grained violent conflict event types relevant to areas like the Horn of Africa. In this paper, we introduce a new benchmark dataset Conflict Events in the Horn of Africa region (CEHA) and propose a new task for identifying violent conflict events using online resources with this dataset. The dataset consists of 500 English event descriptions regarding conflict events in the Horn of Africa region with fine-grained event-type definitions that emphasize the cause of the conflict. This dataset categorizes the key types of conflict risk according to specific areas required by stakeholders in the Humanitarian-Peace-Development Nexus. Additionally, we conduct extensive experiments on two tasks supported by this dataset: Event-relevance Classification and Event-type Classification. Our baseline models demonstrate the challenging nature of these tasks and the usefulness of our dataset for model evaluations in low-resource settings.
Di Lu 0003, Shihao Ran, Elizabeth M. Olson, Hemank Lamba, Aoife Cahill, Joel R. Tetreault, Alejandro Jaimes
COLING7
2025 Uchaguzi-2022: A Dataset of Citizen Reports on the 2022 Kenyan Election
abstract
Online reporting platforms have enabled citizens around the world to collectively share their opinions and report in real time on events impacting their local communities. Systematically organizing (e.g., categorizing by attributes) and geotagging large amounts of crowdsourced information is crucial to ensuring that accurate and meaningful insights can be drawn from this data and used by policy makers to bring about positive change. These tasks, however, typically require extensive manual annotation efforts. In this paper we present Uchaguzi-2022, a dataset of 14k categorized and geotagged citizen reports related to the 2022 Kenyan General Election containing mentions of election-related issues such as official misconduct, vote count irregularities, and acts of violence. We use this dataset to investigate whether language models can assist in scalably categorizing and geotagging reports, thus highlighting its potential application in the AI for Social Good space.
Roberto Mondini, Neema Kotonya, Robert L. Logan IV, Elizabeth M. Olson, Angela Oduor Lungati, Daniel Duke Odongo, Tim Ombasa, Hemank Lamba, Aoife Cahill, Joel R. Tetreault, Alejandro Jaimes
COLING10
2024 Dissecting users' needs for search result explanations
abstract
There is a growing demand for transparency in search engines to understand how search results are curated and to enhance users’ trust. Prior research has introduced search result explanations with a focus on how to explain, assuming explanations are beneficial. Our study takes a step back to examine if search explanations are needed and when they are likely to provide benefits. Additionally, we summarize key characteristics of helpful explanations and share users’ perspectives on explanation features provided by Google and Bing. Interviews with non-technical individuals reveal that users do not always seek or understand search explanations and mostly desire them for complex and critical tasks. They find Google’s search explanations too obvious but appreciate the ability to contest search results. Based on our findings, we offer design recommendations for search engines and explanations to help users better evaluate search results and enhance their search experience.
Prerna Juneja, Alison Smith-Renner, Hemank Lamba, Joel R. Tetreault, Alex Jaimes
CHI5
2023 BUMP: A Benchmark of Unfaithful Minimal Pairs for Meta-Evaluation of Faithfulness Metrics
abstract
Liang Ma, Shuyang Cao, Robert L Logan IV, Di Lu, Shihao Ran, Ke Zhang, Joel Tetreault, Alejandro Jaimes. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Shuyang Cao, Robert L. Logan IV, Di Lu 0003, Shihao Ran, Ke Zhang 0013, Joel R. Tetreault, Alejandro Jaimes
ACL (1)7
2022 Temporal Event Reasoning Using Multi-source Auxiliary Learning Objectives
Xin Dong 0010, Tanay Kumar Saha, Ke Zhang 0013, Joel R. Tetreault, Alejandro Jaimes, Gerard de Melo
ECIR (2)4
2022 Mapping the Design Space of Human-AI Interaction in Text Summarization
abstract
Ruijia Cheng, Alison Smith-Renner, Ke Zhang, Joel Tetreault, Alejandro Jaimes-Larrarte. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Ruijia Cheng, Alison Smith-Renner, Ke Zhang 0013, Joel R. Tetreault, Alejandro Jaimes
NAACL-HLT4
2022 An Exploration of Post-Editing Effectiveness in Text Summarization
abstract
Vivian Lai, Alison Smith-Renner, Ke Zhang, Ruijia Cheng, Wenjuan Zhang, Joel Tetreault, Alejandro Jaimes-Larrarte. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Vivian Lai, Alison Smith-Renner, Ke Zhang 0013, Ruijia Cheng, Joel R. Tetreault, Alejandro Jaimes
NAACL-HLT6
2021 Evaluating the Evaluation Metrics for Style Transfer: A Case Study in Multilingual Formality Transfer
abstract
While the field of style transfer (ST) has been growing rapidly, it has been hampered by a lack of standardized practices for automatic evaluation.In this paper, we evaluate leading ST automatic metrics on the oft-researched task of formality style transfer.Unlike previous evaluations, which focus solely on English, we expand our focus to Brazilian-Portuguese, French, and Italian, making this work the first multilingual evaluation of metrics in ST.We outline best practices for automatic evaluation in (formality) style transfer and identify several models that correlate well with human judgments and are robust across languages.We hope that this work will help accelerate development in ST, where human evaluation is often challenging to collect.
Eleftheria Briakou, Sweta Agrawal, Joel R. Tetreault, Marine Carpuat
EMNLP (1)3
2021 Journalistic Guidelines Aware News Image Captioning
abstract
The task of news article image captioning aims to generate descriptive and informative captions for news article images.Unlike conventional image captions that simply describe the content of the image in general terms, news image captions follow journalistic guidelines and rely heavily on named entities to describe the image content, often drawing context from the whole article they are associated with.In this work, we propose a new approach to this task, motivated by caption guidelines that journalists follow.Our approach, Journalistic Guidelines Aware News Image Captioning (JoGANIC), leverages the structure of captions to improve the generation quality and guide our representation design.Experimental results, including detailed ablation studies, on two large-scale publicly available datasets show that JoGANIC substantially outperforms state-of-the-art methods both on caption generation and named entity related metrics.
Svebor Karaman, Joel R. Tetreault, Alejandro Jaimes
EMNLP (1)3
2021 Real-time Event Detection for Emergency Response Tutorial
abstract
The amount of public data being generated on a daily basis has grown exponentially in the last few years and continues to increase at incredible speed. Most of this data is unstructured and includes text in different formats, in different languages, from many different sources; images, video, audio, and data from sensors. A lot of that data contains information about events happening all over the world, many of which require emergency response. Detecting events in public data, in real time, is therefore critical in many applications: from getting information to first responders as quickly as possible, to creating situational awareness in such emergency situations, as getting the right information to the right places as quickly as possible is critical in saving lives. When an event is ongoing, information on what is happening can be critical in making decisions to keep people safe and take control of the particular situation unfolding. First responders have to quickly make decisions that include what resources to deploy and where. Fortunately, in most emergencies, people use social media to publicly share information. At the same time, sensor data is increasingly becoming available. In order to do this, efficient computational approaches must detect and deliver the right information to the right destination. This tutorial will cover techniques at the state-of-the art to detect events in real-time from large-scale heterogeneous sources. We will focus on NLP, Computer Vision, and Anomaly Detection techniques. We will give specific examples and discuss relevant future research directions in Machine Learning, NLP, Computer Vision and other fields relevant to real time event detection. We will also discuss applications of event detection.
Alejandro Jaimes, Joel R. Tetreault
KDD2
2021 Olá, Bonjour, Salve! XFORMAL: A Benchmark for Multilingual Formality Style Transfer
abstract
Eleftheria Briakou, Di Lu, Ke Zhang, Joel Tetreault. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Eleftheria Briakou, Di Lu 0003, Ke Zhang 0013, Joel R. Tetreault
NAACL-HLT4
2020 The ApposCorpus: a new multilingual, multi-domain dataset for factual appositive generation
abstract
News articles, image captions, product reviews and many other texts mention people and organizations whose name recognition could vary for different audiences.In such cases, background information about the named entities could be provided in the form of an appositive noun phrase, either written by a human or generated automatically.We expand on the previous work in appositive generation with a new, more realistic, end-to-end definition of the task, instantiated by a dataset that spans four languages (English, Spanish, German and Polish), two entity types (person and organization) and two domains (Wikipedia and News).We carry out an extensive analysis of the data and the task, pointing to the various modeling challenges it poses.The results we obtain with standard language generation methods show that the task is indeed non-trivial, and leaves plenty of room for improvement.
Yova Kementchedjhieva, Di Lu 0003, Joel R. Tetreault
COLING3
2020 Rhetoric, Logic, and Dialectic: Advancing Theory-based Argument Quality Assessment in Natural Language Processing
abstract
Though preceding work in computational argument quality (AQ) mostly focuses on assessing overall AQ, researchers agree that writers would benefit from feedback targeting individual dimensions of argumentation theory.However, a large-scale theory-based corpus and corresponding computational models are missing.We fill this gap by conducting an extensive analysis covering three diverse domains of online argumentative writing and presenting GAQCorpus: the first largescale English multi-domain (community Q&A forums, debate forums, review forums) corpus annotated with theory-based AQ scores.We then propose the first computational approaches to theory-based assessment, which can serve as strong baselines for future work.We demonstrate the feasibility of large-scale AQ annotation, show that exploiting relations between dimensions yields performance improvements, and explore the synergies between theory-based prediction and practical AQ assessment.
Anne Lauscher, Lily Ng, Courtney Napoles, Joel R. Tetreault
COLING4
2020 Multimodal Categorization of Crisis Events in Social Media
abstract
Recent developments in image classification and natural language processing, coupled with the rapid growth in social media usage, have enabled fundamental advances in detecting breaking events around the world in real-time. Emergency response is one such area that stands to gain from these advances. By processing billions of texts and images a minute, events can be automatically detected to enable emergency response workers to better assess rapidly evolving situations and deploy resources accordingly. To date, most event detection techniques in this area have focused on image-only or text-only approaches, limiting detection performance and impacting the quality of information delivered to crisis response teams. In this paper, we present a new multimodal fusion method that leverages both images and texts as input. In particular, we introduce a cross-attention module that can filter uninformative and misleading components from weak modalities on a sample by sample basis. In addition, we employ a multimodal graph-based approach to stochastically transition between embeddings of different multimodal pairs during training to better regularize the learning process as well as dealing with limited training data by constructing new matched pairs from different samples. We show that our method outperforms the unimodal approaches and strong multimodal baselines by a large margin on three crisis-related tasks.
Mahdi Abavisani, Liwei Wu 0001, Shengli Hu, Joel R. Tetreault, Alejandro Jaimes
CVPR4
2019 This Email Could Save Your Life: Introducing the Task of Email Subject Line Generation
abstract
Given the overwhelming number of emails, an effective subject line becomes essential to better inform the recipient of the email's content.In this paper, we propose and study the task of email subject line generation: automatically generating an email subject line from the email body.We create the first dataset for this task and find that email subject line generation favor extremely abstractive summary which differentiates it from news headline generation or news single document summarization.We then develop a novel deep learning method and compare it to several baselines as well as recent state-of-the-art text summarization systems.We also investigate the efficacy of several automatic metrics based on correlations with human judgments and propose a new automatic evaluation metric.Our system outperforms competitive baselines given both automatic and human evaluations.To our knowledge, this is the first work to tackle the problem of effective email subject line generation.
Rui Zhang 0037, Joel R. Tetreault
ACL (1)2
2019 Enabling Robust Grammatical Error Correction in New Domains: Datasets, Metrics, and Analyses
abstract
Until now, grammatical error correction (GEC) has been primarily evaluated on text written by non-native English speakers, with a focus on student essays. This paper enables GEC development on text written by native speakers by providing a new data set and metric. We present a multiple-reference test corpus for GEC that includes 4,000 sentences in two new domains ( formal and informal writing by native English speakers) and 2,000 sentences from a diverse set of non-native student writing. We also collect human judgments of several GEC systems on this new test set and perform a meta-evaluation, assessing how reliable automatic metrics are across these domains. We find that commonly used GEC metrics have inconsistent performance across domains, and therefore we propose a new ensemble metric that is robust on all three domains of text.
Courtney Napoles, Maria Nadejde, Joel R. Tetreault
Trans. Assoc. Comput. Linguistics3
2018 Dear Sir or Madam, May I Introduce the GYAFC Dataset: Corpus, Benchmarks and Metrics for Formality Style Transfer
abstract
Style transfer is the task of automatically transforming a piece of text in one particular style into another.A major barrier to progress in this field has been a lack of training and evaluation datasets, as well as benchmarks and automatic metrics.In this work, we create the largest corpus for a particular stylistic transfer (formality) and show that techniques from the machine translation community can serve as strong baselines for future work.We also discuss challenges of using automatic metrics.Informal: I'd say it is punk though.Formal: However, I do believe it to be punk.
Sudha Rao, Joel R. Tetreault
NAACL-HLT2
2018 Discourse Coherence in the Wild: A Dataset, Evaluation and Methods
abstract
To date there has been very little work on assessing discourse coherence methods on real-world data.To address this, we present a new corpus of real-world texts (GCDC) as well as the first large-scale evaluation of leading discourse coherence algorithms.We show that neural models, including two that we introduce here (SENTAVG and PARSEQ), tend to perform best.We analyze these performance differences and discuss patterns we observed in low coherence texts in four domains.
Alice Lai, Joel R. Tetreault
SIGDIAL Conference2
2017 Automatically Identifying Good Conversations Online (Yes, They Do Exist!)
Courtney Napoles, Aasish Pappu, Joel R. Tetreault
ICWSM3
2016 TGIF: A New Dataset and Benchmark on Animated GIF Description
abstract
With the recent popularity of animated GIFs on social media, there is need for ways to index them with rich meta-data. To advance research on animated GIF understanding, we collected a new dataset, Tumblr GIF (TGIF), with 100K animated GIFs from Tumblr and 120K natural language descriptions obtained via crowdsourcing. The motivation for this work is to develop a testbed for image sequence description systems, where the task is to generate natural language descriptions for animated GIFs or video clips. To ensure a high quality dataset, we developed a series of novel quality controls to validate free-form text input from crowd-workers. We show that there is unambiguous association between visual content and natural language descriptions in our dataset, making it an ideal benchmark for the visual content captioning task. We perform extensive statistical analyses to compare our dataset to existing image and video description datasets. Next, we provide baseline results on the animated GIF description task, using three representative techniques: nearest neighbor, statistical machine translation, and recurrent neural networks. Finally, we show that models fine-tuned from our animated GIF description dataset can be helpful for automatic movie description.
Yuncheng Li, Yale Song, Liangliang Cao, Joel R. Tetreault, Larry Goldberg, Alejandro Jaimes, Jiebo Luo 0001
CVPR4
2016 There's No Comparison: Reference-less Evaluation Metrics in Grammatical Error Correction
abstract
Current methods for automatically evaluating grammatical error correction (GEC) systems rely on gold-standard references.However, these methods suffer from penalizing grammatical edits that are correct but not in the gold standard.We show that reference-less grammaticality metrics correlate very strongly with human judgments and are competitive with the leading reference-based evaluation metrics.By interpolating both methods, we achieve state-of-the-art correlation with human judgments.Finally, we show that GEC metrics are much more reliable when they are calculated at the sentence level instead of the corpus level.We have set up a CodaLab site for benchmarking GEC output using a common dataset and different evaluation metrics.
Courtney Napoles, Keisuke Sakaguchi, Joel R. Tetreault
EMNLP3
2016 Humor in Collective Discourse: Unsupervised Funniness Detection in the New Yorker Cartoon Caption Contest
Dragomir R. Radev, Amanda Stent, Joel R. Tetreault, Aasish Pappu, Aikaterini Iliakopoulou, Agustin Chanfreau, Paloma de Juan, Jordi Vallmitjana, Alejandro Jaimes, Rahul Jha, Robert Mankoff
LREC3
2016 Sender-intended functions of emojis in US messaging
abstract
Emojis are an extremely common occurrence in mobile communications, but their meaning is open to interpretation. We investigate motivations for their usage in mobile messaging in the US. This study asked 228 participants for the last time that they used one or more emojis in a conversational message, and collected that message, along with a description of the emojis' intended meaning and function. We discuss functional distinctions between: adding additional emotional or situational meaning, adjusting tone, making a message more engaging to the recipient, conversation management, and relationship maintenance. We discuss lexical placement within messages, as well as social practices. We show that the social and linguistic function of emojis are complex and varied, and that supporting emojis can facilitate important conversational functions.
Henriette Cramer, Paloma de Juan, Joel R. Tetreault
MobileHCI3
2016 Detecting Sarcasm in Multimodal Social Platforms
abstract
Sarcasm is a peculiar form of sentiment expression, where the surface sentiment differs from the implied sentiment. The detection of sarcasm in social media platforms has been applied in the past mainly to textual utterances where lexical indicators (such as interjections and intensifiers), linguistic markers, and contextual information (such as user profiles, or past conversations) were used to detect the sarcastic tone. However, modern social media platforms allow to create multimodal messages where audiovisual content is integrated with the text, making the analysis of a mode in isolation partial. In our work, we first study the relationship between the textual and visual aspects in multimodal posts from three major social media platforms, i.e., Instagram, Tumblr and Twitter, and we run a crowdsourcing task to quantify the extent to which images are perceived as necessary by human annotators. Moreover, we propose two different computational frameworks to detect sarcasm that integrate the textual and visual modalities. The first approach exploits visual semantics trained on an external dataset, and concatenates the semantics features with state-of-the-art textual features. The second method adapts a visual neural network initialized with parameters trained on ImageNet to multimodal sarcastic posts. Results show the positive effect of combining modalities for the detection of sarcasm across platforms and methods.
Rossano Schifanella, Paloma de Juan, Joel R. Tetreault, Liangliang Cao
ACM Multimedia3
2016 Do Characters Abuse More Than Words?
abstract
Although word and character n-grams have been used as features in different NLP applications, no systematic comparison or analysis has shown the power of character-based features for detecting abusive language.In this study, we investigate the effectiveness of such features for abusive language detection in user-generated online comments, and show that such methods outperform previous state-of-theart approaches and other strong baselines.
Yashar Mehdad, Joel R. Tetreault
SIGDIAL Conference2
2016 Abusive Language Detection in Online User Content
abstract
Detection of abusive language in user generated online content has become an issue of increasing importance in recent years. Most current commercial methods make use of blacklists and regular expressions, however these measures fall short when contending with more subtle, less ham-fisted examples of hate speech. In this work, we develop a machine learning based method to detect hate speech on online user comments from two domains which outperforms a state-of-the-art deep learning approach. We also develop a corpus of user comments annotated for abusive language, the first of its kind. Finally, we use our detection tool to analyze abusive language over time and in different settings to further enhance our knowledge of this behavior.
Chikashi Nobata, Joel R. Tetreault, Achint Oommen Thomas, Yashar Mehdad, Yi Chang 0001
WWW2
2016 An Empirical Analysis of Formality in Online Communication
abstract
This paper presents an empirical study of linguistic formality. We perform an analysis of humans’ perceptions of formality in four different genres. These findings are used to develop a statistical model for predicting formality, which is evaluated under different feature settings and genres. We apply our model to an investigation of formality in online discussion forums, and present findings consistent with theories of formality and linguistic coordination.
Ellie Pavlick, Joel R. Tetreault
Trans. Assoc. Comput. Linguistics2
2016 Reassessing the Goals of Grammatical Error Correction: Fluency Instead of Grammaticality
abstract
The field of grammatical error correction (GEC) has grown substantially in recent years, with research directed at both evaluation metrics and improved system performance against those metrics. One unvisited assumption, however, is the reliance of GEC evaluation on error-coded corpora, which contain specific labeled corrections. We examine current practices and show that GEC’s reliance on such corpora unnaturally constrains annotation and automatic evaluation, resulting in (a) sentences that do not sound acceptable to native speakers and (b) system rankings that do not correlate with human judgments. In light of this, we propose an alternate approach that jettisons costly error coding in favor of unannotated, whole-sentence rewrites. We compare the performance of existing metrics over different gold-standard annotations, and show that automatic evaluation with our new annotation scheme has very strong correlation with expert rankings (ρ = 0.82). As a result, we advocate for a fundamental and necessary shift in the goal of GEC, from correcting small, labeled error types, to producing text that has native fluency.
Keisuke Sakaguchi, Courtney Napoles, Matt Post, Joel R. Tetreault
Trans. Assoc. Comput. Linguistics4
2015 It Depends: Dependency Parser Comparison Using A Web-based Evaluation Tool
abstract
Jinho D. Choi, Joel Tetreault, Amanda Stent. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Jinho D. Choi, Joel R. Tetreault, Amanda Stent
ACL (1)2
2015 Automated Grammatical Error Detection for Language Learners, Second Edition
abstract
Research in automated grammatical error detection and correction has gained considerable momentum in the past few years. Although much progress has been made in this area, numerous challenges and opportunities exist for further research. For NLP researchers and students who are considering following or pursuing work in this area and for English language teaching practitioners and researchers who are interested in utilizing systems resulting from research in this area, this book by Leacock, Chodorow, Gamon, and Tetreault provides a timely, updated review of the state of the art in automatic learner error detection and correction.The introductory chapter first highlights the growth in prominence of the field of grammatical error detection and correction since the publication of the first edition of the volume. It then orients the reader to the book by detailing the changes from the first edition, providing a working definition of grammatical error, justifying the book's focus on English language learners, clarifying the use of the term “English language learner” to refer to learners of English as either a second or foreign language, putting forth a few claims about the relationship between NLP and Computer-Assisted Language Learning (CALL), specifying the intended audience of the book, and outlining the structure of the book.Chapter 2 provides the historical context for automated grammatical error detection and correction. The chapter first discusses how varying degrees of tolerance of grammatical errors are achieved in computational grammars used in grammar-checking and proofreading tools. It then offers a quick overview of data-driven and hybrid approaches to error detection.Chapter 3 summarizes the types of errors made by English language learners as reported in previous empirical and corpus-based learner error studies, touches upon the influence of L1 on learner English, and then delves into a detailed discussion of three specific problem areas for English language learners: prepositions, articles, and collocations.Chapter 4 focuses on the evaluation of error detection systems. The chapter first introduces traditional evaluation measures (e.g., precision, recall, F-score, accuracy, and kappa) and the evaluation measures adopted in recent shared tasks on grammatical error correction. It then discusses evaluation using a corpus of correct usage, questioning the value of this approach for indicating how well the system may perform on actual learner data. The advantages and challenges of evaluating system performance on learner writing are then examined. The authors advocate the use of multiple annotators and crowdsourcing as a means to improve the reliability of manual evaluation. The chapter concludes with a discussion of how statistical significance testing of differences in system performance may be performed and provides a checklist for consistent reporting of system results.Chapter 5 zooms in on data-driven approaches to detecting and correcting article and preposition errors. The chapter first looks at four types of information used by different systems, including lexical, syntactic, semantic, and source information. It then describes three prevailing types of data used to train statistical models of grammatical error correction, namely, well-formed text, artificial errors, and error-annotated learner corpora. Next, it discusses classification methods and n-gram statistics-related methods for error detection and correction. Finally, the chapter presents two end-to-end systems, Criterion and MSR ESL Assistant.Chapter 6 focuses on collocation errors. Following a brief discussion of the properties of collocations, the chapter introduces a range of metrics commonly used to measure the association strength between pairs of words and reviews a number of recent collocation error detection and correction systems.Chapter 7 moves beyond article, preposition, and collocation errors and examines rule-based and statistical methods for detecting and correcting verb-form, spelling, and punctuation errors and for identifying ungrammatical sentences. A small number of error detection systems for learners of languages other than English are also discussed.Chapter 8 covers issues with learner error annotation, describes examples of comprehensive and targeted annotation schemes, and proposes three methods for improving the efficiency of large-scale annotation: sampling, crowdsourcing, and mining online revision logs.Chapter 9 highlights several exciting emerging directions in the field of automated grammatical error correction. These include three recent shared tasks on grammatical error correction, the use of machine translation techniques for grammatical error correction, real-time crowdsourcing of grammatical error correction, and longitudinal studies on the efficacy of automated error correction systems for improving the writing of users of such systems.Chapter 10 concludes the book with several suggestions for future research in the field, including annotation for evaluation, error detection for underrepresented languages, research on understudied error types, L1-specific error detection modules, collaboration with second language learning and education groups, and applications of grammatical error correction. The book also includes an Appendix that contains a list of textual learner corpora with at least some publicly accessible URLs or references.I very strongly recommend this book to all NLP students and researchers who are interested in learning about, following, or pursuing research in automated grammatical error detection and correction. Not only does it offer a comprehensive systematic review of research in this field, but it also highlights real challenges and opportunities for fruitful future research. Practitioners and researchers in the language teaching and CALL community can also use this book to obtain a realistic understanding of the state of the art of the field. They will also most certainly welcome the volume's repeated emphasis on the importance for the NLP community to join efforts with the language teaching and CALL community to assess the effect of automated grammatical correction systems on improving the quality of language learners' writing.I have just two minor quibbles with this book. First, it seems that the title may better reflect the content of the book and the goals of the research field in question with the words “and correction” added after “detection.” Second, in both Chapter 3 and Chapter 6, native speakers' preference of powerful computer over strong computer is cited as an example of the arbitrariness of collocations. I do not intend to delve into the debate over whether collocations are arbitrary or linguistically motivated, but I note that the preference in this particular case is not arbitrary but semantically motivated (consider the semantic difference between powerful man and strong man). In fact, this preference merely reflects the fact that we are generally more concerned about the functional power of computers than their physical build, and it would be a false positive if a system marks strong computer as an error in a sentence that talks about a computer that is well-built and not easily breakable.
Claudia Leacock, Martin Chodorow, Michael Gamon, Joel R. Tetreault
Comput. Linguistics4
2014 Non-Monotonic Parsing of Fluent Umm I mean Disfluent Sentences
abstract
Parsing disfluent sentences is a challeng-ing task which involves detecting disflu-encies as well as identifying the syntactic structure of the sentence. While there have been several studies recently into solely detecting disfluencies at a high perfor-mance level, there has been relatively lit-tle work into joint parsing and disfluency detection that has reached that state-of-the-art performance in disfluency detec-tion. We improve upon recent work in this joint task through the use of novel features and learning cascades to produce a model which performs at 82.6 F-score. It outper-forms the previous best in disfluency de-tection on two different evaluations. 1
Mohammad Sadegh Rasooli, Joel R. Tetreault
EACL2
2013 Joint Parsing and Disfluency Detection in Linear Time
abstract
We introduce a novel method to jointly parse and detect disfluencies in spoken utterances.Our model can use arbitrary features for parsing sentences and adapt itself with out-ofdomain data.We show that our method, based on transition-based parsing, performs at a high level of accuracy for both the parsing and disfluency detection tasks.Additionally, our method is the fastest for the joint task, running in linear time.
Mohammad Sadegh Rasooli, Joel R. Tetreault
EMNLP2
2013 Robust Systems for Preposition Error Correction Using Wikipedia Revisions
Aoife Cahill, Nitin Madnani, Joel R. Tetreault, Diane Napolitano
HLT-NAACL3
2012 Building Subjectivity Lexicon(s) from Scratch for Essay Data
Beata Beigman Klebanov, Jill Burstein, Nitin Madnani, Adam Faulkner, Joel R. Tetreault
CICLing (1)5
2012 Problems in Evaluating Grammatical Error Detection Systems
Martin Chodorow, Markus Dickinson, Ross Israel, Joel R. Tetreault
COLING4
2012 Native Tongues, Lost and Found: Resources and Empirical Evaluations in Native Language Identification
Joel R. Tetreault, Daniel Blanchard, Aoife Cahill, Martin Chodorow
COLING1
2012 Correcting Comma Errors in Learner Essays, and Restoring Commas in Newswire Text
Ross Israel, Joel R. Tetreault, Martin Chodorow
HLT-NAACL2
2012 Identifying High-Level Organizational Elements in Argumentative Discourse
Nitin Madnani, Michael Heilman, Joel R. Tetreault, Martin Chodorow
HLT-NAACL3
2012 Re-examining Machine Translation Metrics for Paraphrase Identification
Nitin Madnani, Joel R. Tetreault, Martin Chodorow
HLT-NAACL2
2011 Exploiting Syntactic and Distributional Information for Spelling Correction with Web-Scale N-gram Models
Wei Xu 0004, Joel R. Tetreault, Martin Chodorow, Ralph Grishman
EMNLP2
2010 Using an Error-Annotated Learner Corpus to Develop an ESL/EFL Error Correction System
Na-Rae Han, Joel R. Tetreault, Soo-Hwa Lee, Jin-Young Ha
LREC2
2010 Using Entity-Based Features to Model Coherence in Student Essays
Jill Burstein, Joel R. Tetreault, Slava Andreyev
HLT-NAACL2
2008 The Ups and Downs of Preposition Error Detection in ESL Writing
Joel R. Tetreault, Martin Chodorow
COLING1
2008 A Reinforcement Learning approach to evaluating state representations in spoken dialogue systems
Joel R. Tetreault, Diane J. Litman
Speech Commun.1
2007 Comparing Linguistic Features for Modeling Learning in Computer Tutoring
Katherine Forbes-Riley, Diane J. Litman, Amruta Purandare, Mihai Rotaru 0002, Joel R. Tetreault
AIED5
2007 Estimating the Reliability of MDP Policies: a Confidence Interval Approach
Joel R. Tetreault, Dan Bohus, Diane J. Litman
HLT-NAACL1
2006 Using Reinforcement Learning to Build a Better Model of Dialogue State
Joel R. Tetreault, Diane J. Litman
EACL1
2006 Using system and user performance features to improve emotion detection in spoken tutoring dialogs
abstract
In this study, we incorporate automatically obtained system/user performance features into machine learning experiments to detect student emotion in computer tutoring dialogs. Our results show a relative improvement of 2.7% on classification accuracy and 8.08% on Kappa over using standard lexical, prosodie, sequential, and identification features. This level of improvement is comparable to the performance improvement shown in previous studies by applying dialog acts or lexical/prosodic-/discourse- level contextual features.
Hua Ai, Diane J. Litman, Katherine Forbes-Riley, Mihai Rotaru 0002, Joel R. Tetreault, Amruta Purandare
INTERSPEECH5
2006 Comparing the Utility of State Features in Spoken Dialogue Using Reinforcement Learning
Joel R. Tetreault, Diane J. Litman
HLT-NAACL1
2004 Evaluation of Transcription and Annotation Tools for a Multi-modal, Multi-party Dialogue Corpus
Bilyana Martinovski, Susan Robinson, Jens Stephan, Joel R. Tetreault, David R. Traum
LREC5
2004 Semi-automatic Syntactic and Semantic Corpus Annotation with a Deep Parser
Mary D. Swift, Myroslava O. Dzikovska, Joel R. Tetreault, James F. Allen
LREC3
2001 A Corpus-Based Evalutation of Centering and Pronoun Resolution
abstract
In this paper we compare pronoun resolution algorithms and introduce a centering algorithm(Left-Right Centering) that adheres to the constraints and rules of centering theory and is an alternative to Brennan, Friedman, and Pollard's (1987) algorithm. We then use the Left-Right Centering algorithm to see if two psycholinguistic claims on Cf-list ranking will actually improve pronoun resolution accuracy. Our results from this investigation lead to the development of a new syntax-based ranking of the Cf-list and corpus-based evidence that contradicts the psycholinguistic claims.
Joel R. Tetreault
Comput. Linguistics1
1999 Analysis of Syntax-Based Pronoun Resolution Methods
abstract
This paper presents a pronoun resolution algorithm that adheres to the constraints and rules of Centering Theory (Grosz et al., 1995) and is an alternative to Brennan et al.'s 1987 algorithm. The advantages of this new model, the Left-Right Centering Algorithm (LRC), lie in its incremental processing of utterances and in its low computational overhead. The algorithm is compared with three other pronoun resolution methods: Hobbs' syntax-based algorithm, Strube's S-list approach, and the BFP Centering algorithm. All four methods were implemented in a system and tested on an annotated subset of the Treebank corpus consisting of 2026 pronouns. The noteworthy results were that Hobbs and LRC performed the best.
Joel R. Tetreault
ACL1
1999 A Flexible Architecture for Reference Resolution
Donna K. Byron, Joel R. Tetreault
EACL2