EDBT 2026 Demo / reviewers in the wild / expert
Marine Carpuat
dblp:71/1827
· DBLP profile ↗
70ranked-venue papers
10as first author
38since 2021 · last 2026
0000-0003-1693-0782ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 63 · 9 first-author · 34 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Can You Make It Sound Like You? Post-Editing LLM-Generated Text for Personal StyleabstractConnor Baumler, Calvin Bao, Huy Nghiem, Xinchen Yang, Marine Carpuat, Hal Daumé Iii. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Connor Baumler, Calvin Bao, Huy Nghiem, Xinchen Yang, Marine Carpuat, Hal Daumé III |
ACL (1) | 5 |
| 2026 | Measuring User's Mental Models of Speech Translation in Human-AI CollaborationabstractMillions of people use machine translation (MT) tools daily, yet little is known about their perception of what systems can and cannot do.This paper studies users' mental models of speech translation systems through a new framework based on cross-lingual question answering, where users either accept MT output or request professional re-translation to answer questions based on the information presented in a foreign language.By analyzing user behavior and accuracy trends across varying translation qualities, we examine to what extent they can predict where the system is likely to be wrong, and how this mental model evolves.Users develop stronger mental models with practice, especially when they have some knowledge of the source language, primarily by relying on surface-level error cues.Moreover, providing speech transcriptions can help users develop better mental models.Our results show the promise of cross-lingual question answering as a downstream task for studying MT mental models, and advancing our understanding of human-AI collaboration. HyoJung Han 0001, Nishant Balepur, Jordan L. Boyd-Graber, Marine Carpuat |
ACL (1) | 4 |
| 2026 | Does Speech Translation Meet Users' Needs? An English to Portuguese Study Across DemographicsabstractThis paper introduces Ouvia, a research project to assess user-perceived usability and reliability of modern speech translation tools in En\rightarrowPt scenarios. The project centers on a user study in which we simulate real-life daily interactions by recruiting crowdworkers online from different sociodemographic groups. We collect their spoken requests and self-assessments about quality, satisfaction, and reliability. Here, we describe the project’s motivation and objectives, the study design, and the expected outcomes we will provide to speech translation practitioners. Giuseppe Attanasio, Beatrice Savoldi, Daniel Chechelnitsky, Matteo Negri, Marine Carpuat, André F. T. Martins |
EAMT (2) | 5 |
| 2025 | Multiple LLM Agents Debate for Equitable Cultural AlignmentabstractLarge Language Models (LLMs) need to adapt their predictions to diverse cultural contexts to benefit diverse communities across the world. While previous efforts have focused on single-LLM, single-turn approaches, we propose to exploit the complementary strengths of multiple LLMs to promote cultural adaptability. We introduce a Multi-Agent Debate framework, where two LLM-based agents debate over a cultural scenario and collaboratively reach a final decision. We propose two variants: one where either LLM agents exclusively debate and another where they dynamically choose between self-reflection and debate during their turns. We evaluate these approaches on 7 open-weight LLMs (and 21 LLM combinations) using the NormAd-ETI benchmark for social etiquette norms in 75 countries. Experiments show that debate improves both overall accuracy and cultural group parity over single-LLM baselines. Notably, multi-agent debate enables relatively small LLMs (7-9B) to achieve accuracies comparable to that of a much larger model (27B parameters). Dayeon Ki, Rachel Rudinger, Tianyi Zhou 0001, Marine Carpuat |
ACL (1) | 4 |
| 2025 | Sustaining Human Agency, Attending to Its Cost: An Investigation into Generative AI Design for Non-Native Speakers' Language Use
Yimin Xiao, Cartor Hancock, Sweta Agrawal, Nikita Mehandru, Niloufar Salehi, Marine Carpuat, Ge Gao 0001 |
CHI | 6 |
| 2025 | An Interdisciplinary Approach to Human-Centered Machine TranslationabstractMarine Carpuat, Omri Asscher, Kalika Bali, Luisa Bentivogli, Fred Blain, Lynne Bowker, Monojit Choudhury, Hal Daumé Iii, Kevin Duh, Ge Gao, Alvin C Grissom II, Marzena Karpinska, Elaine C Khoong, William D. Lewis, Andre Martins, Mary Nurminen, Douglas W. Oard, Maja Popovic, Michel Simard, François Yvon. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Marine Carpuat, Omri Asscher, Kalika Bali, Luisa Bentivogli, Frédéric Blain, Lynne Bowker, Monojit Choudhury, Hal Daumé III, Kevin Duh, Ge Gao 0001, Alvin Grissom II, Marzena Karpinska, Elaine C. Khoong, William D. Lewis, André F. T. Martins, Mary Nurminen, Douglas W. Oard, Maja Popovic, Michel Simard, François Yvon |
EMNLP | 1 |
| 2025 | Should I Share this Translation? Evaluating Quality Feedback for User Reliance on Machine TranslationabstractAs people increasingly use AI systems in work and daily life, mechanisms that help them use AI responsibly are urgently needed, especially when they are not equipped to verify AI predictions themselves.We study a realistic Machine Translation (MT) scenario where monolingual users decide whether to share an MT output, first without and then with quality feedback.We compare four types of quality feedback: explicit feedback that directly give users an assessment of translation quality using (1) error highlights and (2) LLM explanations, and implicit feedback that helps users compare MT inputs and outputs through (3) backtranslation and ( 4) question-answer (QA) tables.We find that all feedback types, except error highlights, significantly improve both decision accuracy and appropriate reliance.Notably, implicit feedback, especially QA tables, yields significantly greater gains than explicit feedback in terms of decision accuracy, appropriate reliance, and user perceptions -receiving the highest ratings for helpfulness and trust, and the lowest for mental burden. Dayeon Ki, Kevin Duh, Marine Carpuat |
EMNLP | 3 |
| 2025 | Toward Machine Translation Literacy: How Lay Users Perceive and Rely on Imperfect TranslationsabstractYimin Xiao, Yongle Zhang, Dayeon Ki, Calvin Bao, Marianna J. Martindale, Charlotte Vaughn, Ge Gao, Marine Carpuat. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yimin Xiao, Yongle Zhang 0004, Dayeon Ki, Calvin Bao, Marianna J. Martindale, Charlotte Vaughn, Ge Gao 0001, Marine Carpuat |
EMNLP | 8 |
| 2025 | Adapters for Altering LLM Vocabularies: What Languages Benefit the Most?abstractVocabulary adaptation, which integrates new vocabulary into pre-trained language models, enables expansion to new languages and mitigates token over-fragmentation. However, existing approaches are limited by their reliance on heuristics or external embeddings. We propose VocADT, a novel method for vocabulary adaptation using adapter modules that are trained to learn the optimal linear combination of existing embeddings while keeping the model’s weights fixed. VocADT offers a flexible and scalable solution without depending on external resources or language constraints. Across 11 languages—with diverse scripts, resource availability, and fragmentation—we demonstrate that VocADT outperforms the original Mistral model (Jiang et al., 2023) and other baselines across various multilingual tasks including natural language understanding and machine translation. We find that Latin-script languages and highly fragmented languages
benefit the most from vocabulary adaptation. We further fine-tune the adapted model on the generative task of machine translation and find that vocabulary adaptation is still beneficial after fine-tuning and that VocADT is the most effective. HyoJung Han 0001, Akiko Eriguchi, Hieu Hoang, Marine Carpuat, Huda Khayrallah |
ICLR | 5 |
| 2025 | Words as Bridges: Exploring Computational Support for Cross-Disciplinary Translation WorkabstractScholars often explore literature outside of their home community of study. This exploration process is frequently hampered by field-specific jargon. Past computational work often focuses on supporting translation work by removing jargon through simplification and summarization; here, we explore a different approach that preserves jargon as useful bridges to new conceptual spaces. Specifically, we cast different scholarly domains as different language-using communities, and explore how to adapt techniques from unsupervised cross-lingual alignment of word embeddings to explore conceptual alignments between domain-specific word embedding spaces.We developed a prototype cross-domain search engine that uses aligned domain-specific embeddings to support conceptual exploration, and tested this prototype in two case studies. We discuss qualitative insights into the promises and pitfalls of this approach to translation work, and suggest design insights for future interfaces that provide computational support for cross-domain information seeking. Calvin Bao, Yow-Ting Shiue, Marine Carpuat, Joel Chan |
IUI | 3 |
| 2025 | Improving MT-enabled Triage Performance with Multiple MT OutputsabstractRecent advances in Machine Translation (MT) quality may motivate adoption in a variety of use cases, but the success of MT deployment depends not only on intrinsic model quality but on how well the model, as deployed, helps users meet the objectives of their use case. This work focuses on a specific triage use case, MT-enabled scanning in intelligence analysis. After describing the use case with its objectives and failure modes, we present a user study to establish a baseline performance level and measure the mitigating effects of a simple intervention, providing additional MT outputs. We find significant improvements in relevance judgment accuracy with outputs from two distinct neural MT models and significant improvements in relevant entity identification with the addition of a rule-based MT. Users also like seeing multiple MT outputs, making it an appealing way to improve MT-enabled scanning performance. Marianna J. Martindale, Marine Carpuat |
MTSummit (1) | 2 |
| 2025 | Automatic Input Rewriting Improves Translation with Large Language ModelsabstractDayeon Ki, Marine Carpuat. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Dayeon Ki, Marine Carpuat |
NAACL (Long Papers) | 2 |
| 2024 | XLAVS-R: Cross-Lingual Audio-Visual Speech Representation Learning for Noise-Robust Speech PerceptionabstractHyoJung Han, Mohamed Anwar, Juan Pino, Wei-Ning Hsu, Marine Carpuat, Bowen Shi, Changhan Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. HyoJung Han 0001, Mohamed Anwar, Juan Pino 0001, Wei-Ning Hsu, Marine Carpuat, Bowen Shi 0002, Changhan Wang |
ACL (1) | 5 |
| 2024 | Assessing Common Ground through Language-based Cultural Consensus in Humans and Large Language Models
Sophie Domanski, Rachel Rudinger, Marine Carpuat, Patrick Shafto, Yi Ting Huang |
CogSci | 3 |
| 2024 | Automatic Authorship Analysis in Human-AI Collaborative WritingabstractAs the quality of AI-generated text increases with the development of new Large Language Models, people use them to write in a variety of contexts. Human-AI collaborative writing poses a potential challenge for existing AI analysis techniques, which have been primarily tested either on human-written text only, or on samples independently generated by humans and AI. In this work, we investigate the extent to which existing AI detection and authorship analysis models can perform classification on data generated in human-AI collaborative writing sessions. Results show that, for AI text detection in the cowriting setting, classifiers based on authorship embeddings (Rivera-Soto et al., 2021) outperform classifiers used in prior work distinguishing AI vs. human text generated independently. However, these embeddings are not optimal for finer-grained authorship identification tasks: for authorship verification, n-gram based models are more robust to human-AI co-written text, and authorship attribution performance degrades compared to baselines that use human-written text only. Taken together, this suggests that the rise of human-AI co-written text will require adapting AI detection tools and authorship analysis techniques in the near future. We release our code at https://github.com/AARichburg/Human-AI_Authorship_Analysis. Aquia Richburg, Calvin Bao, Marine Carpuat |
LREC/COLING | 3 |
| 2024 | SpeechQE: Estimating the Quality of Direct Speech TranslationabstractRecent advances in automatic quality estimation for machine translation have exclusively focused on written language, leaving the speech modality underexplored.In this work, we formulate the task of quality estimation for speech translation, construct a benchmark, and evaluate a family of systems based on cascaded and end-to-end architectures.In this process, we introduce a novel end-to-end system leveraging pre-trained text LLM.Results suggest that end-to-end approaches are better suited to estimating the quality of direct speech translation than using quality estimation systems designed for text in cascaded systems.More broadly, we argue that quality estimation of speech translation needs to be studied as a separate problem from that of text, and release our data and models to guide further research in this space.1 HyoJung Han 0001, Kevin Duh, Marine Carpuat |
EMNLP | 3 |
| 2024 | Keep it Private: Unsupervised Privatization of Online TextabstractCalvin Bao, Marine Carpuat. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Calvin Bao, Marine Carpuat |
NAACL-HLT | 2 |
| 2024 | AfriMTE and AfriCOMET: Enhancing COMET to Embrace Under-resourced African LanguagesabstractJiayi Wang, David Ifeoluwa Adelani, Sweta Agrawal, Marek Masiak, Ricardo Rei, Eleftheria Briakou, Marine Carpuat, Xuanli He, Sofia Bourhim, Andiswa Bukula, Muhidin Mohamed, Temitayo Olatoye, Tosin Adewumi, Hamam Mokayed, Christine Mwase, Wangui Kimotho, Foutse Yuehgoh, Anuoluwapo Aremu, Jessica Ojo, Shamsuddeen Hassan Muhammad, Salomey Osei, Abdul-Hakeem Omotayo, Chiamaka Chukwuneke, Perez Ogayo, Oumaima Hourrane, Salma El Anigri, Lolwethu Ndolela, Thabiso Mangwana, Shafie Abdi Mohamed, Hassan Ayinde, Oluwabusayo Olufunke Awoyomi, Lama Alkhaled, Sana Al-azzawi, Naome A. Etori, Millicent Ochieng, Clemencia Siro, Njoroge Kiragu, Eric Muchiri, Wangari Kimotho, Lyse Naomi Wamba Momo, Daud Abolade, Simbiat Ajao, Iyanuoluwa Shode, Ricky Macharm, Ruqayya Nasir Iro, Saheed S. Abdullahi, Stephen E. Moore, Bernard Opoku, Zainab Akinjobi, Abeeb Afolabi, Nnaemeka Obiefuna, Onyekachi Raphael Ogbu, Sam Ochieng’, Verrah Akinyi Otiende, Chinedu Emmanuel Mbonu, Sakayo Toadoum Sari, Yao Lu, Pontus Stenetorp. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jiayi Wang 0010, David Ifeoluwa Adelani, Sweta Agrawal, Marek Masiak, Ricardo Rei, Eleftheria Briakou, Marine Carpuat, Xuanli He, Sofia Bourhim, Andiswa Bukula, Muhidin Mohamed, Temitayo Olatoye, Tosin P. Adewumi, Hamam Mokayed, Christine Mwase, Wangui Kimotho, Foutse Yuehgoh, Aremu Anuoluwapo, Jessica Ojo, Shamsuddeen Hassan Muhammad, Salomey Osei, Abdul-Hakeem Omotayo, Chiamaka Ijeoma Chukwuneke, Perez Ogayo, Oumaima Hourrane, Salma El Anigri, Lolwethu Ndolela, Thabiso Mangwana, Shafie Abdi Mohamed, Ayinde Hassan, Oluwabusayo Olufunke Awoyomi, Lama Alkhaled, Sana Sabah Al-Azzawi, Naome A. Etori, Millicent Ochieng, Clemencia Siro, Njoroge Kiragu, Eric Muchiri, Wangari Kimotho, Sakayo Toadoum Sari, Lyse Naomi Wamba Momo, Daud Abolade, Simbiat Ajao, Iyanuoluwa Shode, Ricky Macharm, Ruqayya Nasir Iro, Saheed S. Abdullahi, Stephen E. Moore, Bernard Opoku, Zainab Akinjobi, Afolabi Abeeb, Nnaemeka C. Obiefuna, Onyekachi Raphael Ogbu, Sam Ochieng', Verrah Otiende, Chinedu E. Mbonu, Pontus Stenetorp |
NAACL-HLT | 7 |
| 2024 | Do Text Simplification Systems Preserve Meaning? A Human Evaluation via Reading ComprehensionabstractAbstract Automatic text simplification (TS) aims to automate the process of rewriting text to make it easier for people to read. A pre-requisite for TS to be useful is that it should convey information that is consistent with the meaning of the original text. However, current TS evaluation protocols assess system outputs for simplicity and meaning preservation without regard for the document context in which output sentences occur and for how people understand them. In this work, we introduce a human evaluation framework to assess whether simplified texts preserve meaning using reading comprehension questions. With this framework, we conduct a thorough human evaluation of texts by humans and by nine automatic systems. Supervised systems that leverage pre-training knowledge achieve the highest scores on the reading comprehension tasks among the automatic controllable TS systems. However, even the best-performing supervised system struggles with at least 14% of the questions, marking them as “unanswerable” based on simplified content. We further investigate how existing TS evaluation metrics and automatic question-answering systems approximate the human judgments we obtained. Sweta Agrawal, Marine Carpuat |
Trans. Assoc. Comput. Linguistics | 2 |
| 2024 | How Often Are Errors in Natural Language Reasoning Due to Paraphrastic Variability?abstractAbstract Large language models have been shown to behave inconsistently in response to meaning-preserving paraphrastic inputs. At the same time, researchers evaluate the knowledge and reasoning abilities of these models with test evaluations that do not disaggregate the effect of paraphrastic variability on performance. We propose a metric, PC, for evaluating the paraphrastic consistency of natural language reasoning models based on the probability of a model achieving the same correctness on two paraphrases of the same problem. We mathematically connect this metric to the proportion of a model’s variance in correctness attributable to paraphrasing. To estimate PC, we collect ParaNlu, a dataset of 7,782 human-written and validated paraphrased reasoning problems constructed on top of existing benchmark datasets for defeasible and abductive natural language inference.1 Using ParaNlu, we measure the paraphrastic consistency of several model classes and show that consistency dramatically increases with pretraining but not fine-tuning. All models tested exhibited room for improvement in paraphrastic consistency. Neha Srikanth, Marine Carpuat, Rachel Rudinger |
Trans. Assoc. Comput. Linguistics | 2 |
| 2023 | Towards Conceptualization of "Fair Explanation": Disparate Impacts of anti-Asian Hate Speech Explanations on Content ModeratorsabstractRecent research at the intersection of AI explainability and fairness has focused on how explanations can improve human-plus-AI task performance as assessed by fairness measures.We propose to characterize what constitutes an explanation that is itself "fair" -an explanation that does not adversely impact specific populations.We formulate a novel evaluation method of "fair explanations" using not just accuracy and label time, but also psychological impact of explanations on different user groups across many metrics (mental discomfort, stereotype activation, and perceived workload).We apply this method in the context of content moderation of potential hate speech, and its differential impact on Asian vs. non-Asian proxy moderators, across explanation approaches (saliency map and counterfactual explanation).We find that saliency maps generally perform better and show less evidence of disparate impact (group) and individual unfairness than counterfactual explanations.1 Content warning: This paper contains examples of hate speech and racially discriminatory language.The authors do not support such content.Please consider your risk of discomfort carefully before continuing reading! Tin Nguyen 0005, Jiannan Xu, Aayushi Roy, Hal Daumé III, Marine Carpuat |
EMNLP | 5 |
| 2023 | Controlling Pre-trained Language Models for Grade-Specific Text SimplificationabstractText simplification (TS) systems rewrite text to make it more readable while preserving its content.However, what makes a text easy to read depends on the intended readers.Recent work has shown that pre-trained language models can simplify text using a wealth of techniques to control output simplicity, ranging from specifying only the desired reading grade level, to directly specifying low-level edit operations.Yet it remains unclear how to set these control parameters in practice.Existing approaches set them at the corpus level, disregarding the complexity of individual inputs and considering only one level of output complexity.In this work, we conduct an empirical study to understand how different control mechanisms impact the adequacy and simplicity of text simplification systems.Based on these insights, we introduce a simple method that predicts the edit operations required for simplifying a text for a specific grade level on an instance-per-instance basis.This approach improves the quality of the simplified outputs over corpus-level searchbased heuristics. Sweta Agrawal, Marine Carpuat |
EMNLP | 2 |
| 2023 | Explaining with Contrastive Phrasal Highlighting: A Case Study in Assisting Humans to Detect Translation DifferencesabstractExplainable NLP techniques primarily explain by answering "Which tokens in the input are responsible for this prediction?".We argue that for NLP models that make predictions by comparing two input texts, it is more useful to explain by answering "What differences between the two inputs explain this prediction?".We introduce a technique to generate contrastive phrasal highlights that explain the predictions of a semantic divergence model via phrasealignment-guided erasure.We show that the resulting highlights match human rationales of cross-lingual semantic differences better than popular post-hoc saliency techniques and that they successfully help people detect finegrained meaning differences in human translations and critical machine translation errors. Eleftheria Briakou, Navita Goyal, Marine Carpuat |
EMNLP | 3 |
| 2023 | What Else Do I Need to Know? The Effect of Background Information on Users' Reliance on QA SystemsabstractNavita Goyal, Eleftheria Briakou, Amanda Liu, Connor Baumler, Claire Bonial, Jeffrey Micher, Clare Voss, Marine Carpuat, Hal Daumé III. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Navita Goyal, Eleftheria Briakou, Amanda Liu, Connor Baumler, Claire Bonial, Jeffrey Micher, Clare R. Voss, Marine Carpuat, Hal Daumé III |
EMNLP | 8 |
| 2023 | Bridging Background Knowledge Gaps in Translation with Automatic ExplicitationabstractTranslations help people understand content written in another language.However, even correct literal translations do not fulfill that goal when people lack the necessary background to understand them.Professional translators incorporate explicitations to explain the missing context by considering cultural differences between source and target audiences.Despite its potential to help users, NLP research on explicitation is limited because of the dearth of adequate evaluation methods.This work introduces techniques for automatically generating explicitations, motivated by WIKIEXPL 1 : a dataset that we collect from Wikipedia and annotate with human translators.The resulting explicitations are useful as they help answer questions more accurately in a multilingual question answering framework. HyoJung Han 0001, Jordan L. Boyd-Graber, Marine Carpuat |
EMNLP | 3 |
| 2023 | Physician Detection of Clinical Harm in Machine Translation: Quality Estimation Aids in Reliance and Backtranslation Identifies Critical ErrorsabstractA major challenge in the practical use of Machine Translation (MT) is that users lack guidance to make informed decisions about when to rely on outputs.Progress in quality estimation research provides techniques to automatically assess MT quality, but these techniques have primarily been evaluated in vitro by comparison against human judgments outside of a specific context of use.This paper evaluates quality estimation feedback in vivo with a human study simulating decision-making in high-stakes medical settings.Using Emergency Department discharge instructions, we study how interventions based on quality estimation versus backtranslation assist physicians in deciding whether to show MT outputs to a patient.We find that quality estimation improves appropriate reliance on MT, but backtranslation helps physicians detect more clinically harmful errors that QE alone often misses. Nikita Mehandru, Sweta Agrawal, Yimin Xiao, Ge Gao 0001, Elaine C. Khoong, Marine Carpuat, Niloufar Salehi |
EMNLP | 6 |
| 2023 | A Rose by Any Other Name would not Smell as Sweet: Social Bias in Names MistranslationabstractWe ask the question: Are there widespread disparities in machine translations of names across race/ethnicity, and gender?We hypothesize that the translation quality of names and surrounding context will be lower for names associated with US racial and ethnic minorities due to these systems' tendencies to standardize language to predominant language patterns.We develop a dataset of names that are strongly demographically aligned and propose a translation evaluation procedure based on round-trip translation.We analyze the effect of name demographics on translation quality using generalized linear mixed effects models and find that the ability of translation systems to correctly translate female-associated names is significantly lower than male-associated names.This effect is particularly pronounced for femaleassociated names that are also associated with racial (Black) and ethnic (Hispanic) minorities.This disparity in translation quality between social groups for something as personal as someone's name has significant implications for people's professional, personal and cultural identities, self-worth and ease of communication.Our findings suggest that more MT research is needed to improve the translation of names and to provide high-quality service for users regardless of gender, race, and ethnicity. Sandra Sandoval, Jieyu Zhao 0001, Marine Carpuat, Hal Daumé III |
EMNLP | 3 |
| 2023 | An empirical assessment of machine learning approaches for triaging reports of static analysis tools
Sai S. Yerramreddy, Austin Mordahl, Ugur Koc, Shiyi Wei, Jeffrey S. Foster, Marine Carpuat, Adam A. Porter |
Empir. Softw. Eng. | 6 |
| 2023 | Understanding and Detecting Hallucinations in Neural Machine Translation via Model IntrospectionabstractAbstract Neural sequence generation models are known to “hallucinate”, by producing outputs that are unrelated to the source text. These hallucinations are potentially harmful, yet it remains unclear in what conditions they arise and how to mitigate their impact. In this work, we first identify internal model symptoms of hallucinations by analyzing the relative token contributions to the generation in contrastive hallucinated vs. non-hallucinated outputs generated via source perturbations. We then show that these symptoms are reliable indicators of natural hallucinations, by using them to design a lightweight hallucination detector which outperforms both model-free baselines and strong classifiers based on quality estimation or large pre-trained models on manually annotated English-Chinese and German-English translation test beds. Weijia Xu, Sweta Agrawal, Eleftheria Briakou, Marianna J. Martindale, Marine Carpuat |
Trans. Assoc. Comput. Linguistics | 5 |
| 2022 | An Imitation Learning Curriculum for Text Editing with Non-Autoregressive ModelsabstractWe propose a framework for training nonautoregressive sequence-to-sequence models for editing tasks, where the original input sequence is iteratively edited to produce the output.We show that the imitation learning algorithms designed to train such models for machine translation introduces mismatches between training and inference that lead to undertraining and poor generalization in editing scenarios.We address this issue with two complementary strategies: 1) a roll-in policy that exposes the model to intermediate training sequences that it is more likely to encounter during inference, 2) a curriculum that presents easy-to-learn edit operations first, gradually increasing the difficulty of training samples as the model becomes competent.We show the efficacy of these strategies on two challenging English editing tasks: controllable text simplification and abstractive summarization.Our approach significantly improves output quality on both tasks and controls output complexity better on the simplification task. Sweta Agrawal, Marine Carpuat |
ACL (1) | 2 |
| 2022 | Can Synthetic Translations Improve Bitext Quality?abstractSynthetic translations have been used for a wide range of NLP tasks primarily as a means of data augmentation.This work explores, instead, how synthetic translations can be used to revise potentially imperfect reference translations in mined bitext.We find that synthetic samples can improve bitext quality without any additional bilingual supervision when they replace the originals based on a semantic equivalence classifier that helps mitigate NMT noise.The improved quality of the revised bitext is confirmed intrinsically via human evaluation and extrinsically through bilingual induction and MT tasks. Eleftheria Briakou, Marine Carpuat |
ACL (1) | 2 |
| 2022 | Constrained Regeneration for Cross-Lingual Query-Focused Extractive SummarizationabstractQuery-focused summaries of foreign-language, retrieved documents can help a user understand whether a document is actually relevant to the query term. A standard approach to this problem is to first translate the source documents and then perform extractive summarization to find relevant snippets. However, in a cross-lingual setting, the query term does not necessarily appear in the translations of relevant documents. In this work, we show that constrained machine translation and constrained post-editing can improve human relevance judgments by including a query term in a summary when its translation appears in the source document. We also present several strategies for selecting only certain documents for regeneration which yield further improvements Elsbeth Turcan, David Wan, Faisal Ladhak, Petra Galuscáková, Sukanta Sen, Svetlana Tchistiakova, Weijia Xu, Marine Carpuat, Kenneth Heafield, Douglas W. Oard, Kathy McKeown |
COLING | 8 |
| 2022 | SimQA: Detecting Simultaneous MT Errors through Word-by-Word Question AnsweringabstractDetractors of neural machine translation admit that while its translations are fluent, it sometimes gets key facts wrong.This is particularly important in simultaneous interpretation where translations have to be provided as fast as possible: before a sentence is complete.Yet, evaluations of simultaneous machine translation (SIMULMT) fail to capture if systems correctly translate the most salient elements of a question: people, places, and dates.To address this problem, we introduce a downstream word-by-word question answering evaluation task (SIMQA): given a source language question, translate the question word by word into the target language, and answer as soon as possible.SIMQA jointly measures whether the SIMULMT models translate the question quickly and accurately, and can reveal shortcomings in existing neural systemshallucinating or omitting facts. HyoJung Han 0001, Marine Carpuat, Jordan L. Boyd-Graber |
EMNLP | 2 |
| 2022 | Facilitating Global Team Meetings Between Language-Based Subgroups: When and How Can Machine Translation Help?abstractGlobal teams frequently consist of language-based subgroups who put together complementary information to achieve common goals. Previous research outlines a two-step work communication flow in these teams. There are team meetings using a required common language (i.e., English); in preparation for those meetings, people have subgroup conversations in their native languages. Work communication at team meetings is often less effective than in subgroup conversations. In the current study, we investigate the idea of leveraging machine translation (MT) to facilitate global team meetings. We hypothesize that exchanging subgroup conversation logs before a team meeting offers contextual information that benefits teamwork at the meeting. MT can translate these logs, which enables comprehension at a low cost. To test our hypothesis, we conducted a between-subjects experiment where twenty quartets of participants performed a personnel selection task. Each quartet included two English native speakers (NS) and two non-native speakers (NNS) whose native language was Mandarin. All participants began the task with subgroup conversations in their native languages, then proceeded to team meetings in English. We manipulated the exchange of subgroup conversation logs prior to team meetings: with MT-mediated exchanges versus without. Analysis of participants' subjective experience, task performance, and depth of discussions as reflected through their conversational moves jointly indicates that team meeting quality improved when there were MT-mediated exchanges of subgroup conversation logs as opposed to no exchanges. We conclude with reflections on when and how MT could be applied to enhance global teamwork across a language barrier. Yongle Zhang 0004, Dennis Asamoah Owusu, Marine Carpuat, Ge Gao 0001 |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2021 | Beyond Noise: Mitigating the Impact of Fine-grained Semantic Divergences on Neural Machine TranslationabstractEleftheria Briakou, Marine Carpuat. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Eleftheria Briakou, Marine Carpuat |
ACL/IJCNLP (1) | 2 |
| 2021 | Evaluating the Evaluation Metrics for Style Transfer: A Case Study in Multilingual Formality TransferabstractWhile the field of style transfer (ST) has been growing rapidly, it has been hampered by a lack of standardized practices for automatic evaluation.In this paper, we evaluate leading ST automatic metrics on the oft-researched task of formality style transfer.Unlike previous evaluations, which focus solely on English, we expand our focus to Brazilian-Portuguese, French, and Italian, making this work the first multilingual evaluation of metrics in ST.We outline best practices for automatic evaluation in (formality) style transfer and identify several models that correlate well with human judgments and are robust across languages.We hope that this work will help accelerate development in ST, where human evaluation is often challenging to collect. Eleftheria Briakou, Sweta Agrawal, Joel R. Tetreault, Marine Carpuat |
EMNLP (1) | 4 |
| 2021 | Rule-based Morphological Inflection Improves Neural Terminology TranslationabstractCurrent approaches to incorporating terminology constraints in machine translation (MT) typically assume that the constraint terms are provided in their correct morphological forms.This limits their application to real-world scenarios where constraint terms are provided as lemmas.In this paper, we introduce a modular framework for incorporating lemma constraints in neural MT (NMT) in which linguistic knowledge and diverse types of NMT models can be flexibly applied.It is based on a novel cross-lingual inflection module that inflects the target lemma constraints based on the source context.We explore linguistically motivated rule-based and data-driven neuralbased inflection modules and design English-German health and English-Lithuanian news test suites to evaluate them in domain adaptation and low-resource MT settings.Results show that our rule-based inflection module helps NMT models incorporate lemma constraints more accurately than a neural module and outperforms the existing end-to-end approach with lower training costs. 1 Weijia Xu, Marine Carpuat |
EMNLP (1) | 2 |
| 2021 | EDITOR: an Edit-Based Transformer with Repositioning for Neural Machine Translation with Soft Lexical ConstraintsabstractAbstract We introduce an Edit-Based TransfOrmer with Repositioning (EDITOR), which makes sequence generation flexible by seamlessly allowing users to specify preferences in output lexical choice. Building on recent models for non-autoregressive sequence generation (Gu et al., 2019), EDITOR generates new sequences by iteratively editing hypotheses. It relies on a novel reposition operation designed to disentangle lexical choice from word positioning decisions, while enabling efficient oracles for imitation learning and parallel edits at decoding time. Empirically, EDITOR uses soft lexical constraints more effectively than the Levenshtein Transformer (Gu et al., 2019) while speeding up decoding dramatically compared to constrained beam search (Post and Vilar, 2018). EDITOR also achieves comparable or better translation quality with faster decoding speed than the Levenshtein Transformer on standard Romanian-English, English-German, and English-Japanese machine translation tasks. Weijia Xu, Marine Carpuat |
Trans. Assoc. Comput. Linguistics | 2 |
| 2020 | Controlling Neural Machine Translation Formality with Synthetic SupervisionabstractThis work aims to produce translations that convey source language content at a formality level that is appropriate for a particular audience. Framing this problem as a neural sequence-to-sequence task ideally requires training triplets consisting of a bilingual sentence pair labeled with target language formality. However, in practice, available training examples are limited to English sentence pairs of different styles, and bilingual parallel sentences of unknown formality. We introduce a novel training scheme for multi-task models that automatically generates synthetic training triplets by inferring the missing element on the fly, thus enabling end-to-end training. Comprehensive automatic and human assessments show that our best model outperforms existing models by producing translations that better match desired formality levels while preserving the source meaning.1 Xing Niu 0001, Marine Carpuat |
AAAI | 2 |
| 2020 | Detecting Fine-Grained Cross-Lingual Semantic Divergences without Supervision by Learning to RankabstractDetecting fine-grained differences in content conveyed in different languages matters for cross-lingual NLP and multilingual corpora analysis, but it is a challenging machine learning problem since annotation is expensive and hard to scale.This work improves the prediction and annotation of finegrained semantic divergences.We introduce a training strategy for multilingual BERT models by learning to rank synthetic divergent examples of varying granularity.We evaluate our models on the Rationalized English-French Semantic Divergences, a new dataset released with this work, consisting of English-French sentence-pairs annotated with semantic divergence classes and token-level rationales.Learning to rank helps detect finegrained sentence-level divergences more accurately than a strong sentence-level similarity model, while token-level predictions have the potential of further distinguishing between coarse and fine-grained divergences.ADV VERB ADJ NOUN how weak they are.BERT predictions { permission, attention, hand, mercy, story } WORDNET hypernyms { communication Eleftheria Briakou, Marine Carpuat |
EMNLP (1) | 2 |
| 2019 | Controlling Text Complexity in Neural Machine TranslationabstractSweta Agrawal, Marine Carpuat. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Sweta Agrawal, Marine Carpuat |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Weakly Supervised Cross-lingual Semantic Relation Classification via Knowledge DistillationabstractYogarshi Vyas, Marine Carpuat. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yogarshi Vyas, Marine Carpuat |
EMNLP/IJCNLP (1) | 2 |
| 2019 | An Empirical Assessment of Machine Learning Approaches for Triaging Reports of a Java Static Analysis ToolabstractDespite their ability to detect critical bugs in software, developers consider high false positive rates to be a key barrier to using static analysis tools in practice. To improve the usability of these tools, researchers have recently begun to apply machine learning techniques to classify and filter false positive analysis reports. Although initial results have been promising, the long-term potential and best practices for this line of research are unclear due to the lack of detailed, large-scale empirical evaluation. To partially address this knowledge gap, we present a comparative empirical study of four machine learning techniques, namely hand-engineered features, bag of words, recurrent neural networks, and graph neural networks, for classifying false positives, using multiple ground-truth program sets. We also introduce and evaluate new data preparation routines for recurrent neural networks and node representations for graph neural networks, and show that these routines can have a substantial positive impact on classification accuracy. Overall, our results suggest that recurrent neural networks (which learn over a program's source code) outperform the other subject techniques, although interesting tradeoffs are present among all techniques. Our observations provide insight into the future research needed to speed the adoption of machine learning approaches in practice. Ugur Koc, Shiyi Wei, Jeffrey S. Foster, Marine Carpuat, Adam A. Porter |
ICST | 4 |
| 2019 | Identifying Fluently Inadequate Output in Neural and Statistical Machine Translation
Marianna J. Martindale, Marine Carpuat, Kevin Duh, Paul McNamee |
MTSummit (1) | 2 |
| 2018 | Multi-Task Neural Models for Translating Between Styles Within and Across LanguagesabstractGenerating natural language requires conveying content in an appropriate style. We explore two related tasks on generating text of varying formality: monolingual formality transfer and formality-sensitive machine translation. We propose to solve these tasks jointly using multi-task learning, and show that our models achieve state-of-the-art performance for formality transfer and are able to perform formality-sensitive translation without being explicitly trained on style-annotated translation examples. Xing Niu 0001, Sudha Rao, Marine Carpuat |
COLING | 3 |
| 2018 | Robust Cross-Lingual Hypernymy Detection Using Dependency ContextabstractShyam Upadhyay, Yogarshi Vyas, Marine Carpuat, Dan Roth. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Shyam Upadhyay, Yogarshi Vyas, Marine Carpuat, Dan Roth 0001 |
NAACL-HLT | 3 |
| 2018 | Identifying Semantic Divergences in Parallel Text without AnnotationsabstractYogarshi Vyas, Xing Niu, Marine Carpuat. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Yogarshi Vyas, Xing Niu 0001, Marine Carpuat |
NAACL-HLT | 3 |
| 2017 | A Study of Style in Machine Translation: Controlling the Formality of Machine Translation OutputabstractStylistic variations of language, such as formality, carry speakers' intention beyond literal meaning and should be conveyed adequately in translation.We propose to use lexical formality models to control the formality level of machine translation output.We demonstrate the effectiveness of our approach in empirical evaluations, as measured by automatic metrics and human assessments. Xing Niu 0001, Marianna J. Martindale, Marine Carpuat |
EMNLP | 3 |
| 2016 | Retrofitting Sense-Specific Word Vectors Using Parallel TextabstractJauhar et al. (2015) recently proposed to learn sense-specific word representations by "retrofitting" standard distributional word representations to an existing ontology.We observe that this approach does not require an ontology, and can be generalized to any graph defining word senses and relations between them.We create such a graph using translations learned from parallel corpora.On a set of lexical semantic tasks, representations learned using parallel text perform roughly as well as those derived from WordNet, and combining the two representation types significantly improves performance. Allyson Ettinger, Philip Resnik, Marine Carpuat |
HLT-NAACL | 3 |
| 2016 | Sparse Bilingual Word Representations for Cross-lingual Lexical EntailmentabstractWe introduce the task of cross-lingual lexical entailment, which aims to detect whether the meaning of a word in one language can be inferred from the meaning of a word in another language.We construct a gold standard for this task, and propose an unsupervised solution based on distributional word representations.As commonly done in the monolingual setting, we assume a word e entails a word f if the prominent context features of e are a subset of those of f .To address the challenge of comparing contexts across languages, we propose a novel method for inducing sparse bilingual word representations from monolingual and parallel texts.Our approach yields an Fscore of 70%, and significantly outperforms strong baselines based on translation and on existing word representations. Yogarshi Vyas, Marine Carpuat |
HLT-NAACL | 2 |
| 2014 | Cross-lingual Discourse Relation Analysis: A corpus study and a semi-supervised classification system
Junyi Jessy Li, Marine Carpuat, Ani Nenkova |
COLING | 2 |
| 2013 | SenseSpotting: Never let your parallel data tie you to an old domain
Marine Carpuat, Hal Daumé III, Katharine Henry, Ann Irvine, Jagadeesh Jagarlamudi, Rachel Rudinger |
ACL (1) | 1 |
| 2013 | Measuring Machine Translation Errors in New DomainsabstractWe develop two techniques for analyzing the effect of porting a machine translation system to a new domain. One is a macro-level analysis that measures how domain shift affects corpus-level evaluation; the second is a micro-level analysis for word-level errors. We apply these methods to understand what happens when a Parliament-trained phrase-based machine translation system is applied in four very different domains: news, medical texts, scientific articles and movie subtitles. We present quantitative and qualitative experiments that highlight opportunities for future research in domain adaptation for machine translation. Ann Irvine, John Morgan, Marine Carpuat, Hal Daumé III, Dragos Stefan Munteanu |
Trans. Assoc. Comput. Linguistics | 3 |
| 2012 | Filtering and routing multilingual documents for translationabstractTranslation is a key capability to access relevant information expressed in various languages on social media. Unfortunately, systematically translating all content far exceeds the capacity of most organizations. Computer-aided translation (CAT) tools can significantly increase the productivity of translators, but can not ultimately cope with the overwhelming amount of content to translate. In this contribution, we describe and experiment with an approach where we use the structure in a corpus to adequately route the content to the proper workflow, including translators, CAT tools or purely automatic approaches. We show that linguistically motivated structure such as document genre can help decide on the proper translation workflow. However, automatically discovered structure has an effect that is at least as important and allows us to define groups of documents that may be translated automatically with reasonable output quality. This suggests that computational intelligence models that can efficiently organize document collection will provide increased capability to access textual content from various target languages. Marine Carpuat, Cyril Goutte, Pierre Isabelle |
CISDA | 1 |
| 2012 | Improved Arabic-to-English statistical machine translation by reordering post-verbal subjects for word alignment
Marine Carpuat, Yuval Marton, Nizar Habash |
Mach. Transl. | 1 |
| 2010 | Task-based Evaluation of Multiword Expressions: a Pilot Study in Statistical Machine Translation
Marine Carpuat, Mona T. Diab |
HLT-NAACL | 1 |
| 2008 | Evaluation of Context-Dependent Phrasal Translation Lexicons for Statistical Machine Translation
Marine Carpuat, Dekai Wu |
LREC | 1 |
| 2007 | Improving Statistical Machine Translation Using Word Sense Disambiguation
Marine Carpuat, Dekai Wu |
EMNLP-CoNLL | 1 |
| 2007 | Context-dependent phrasal translation lexicons for statistical machine translation
Marine Carpuat, Dekai Wu |
MTSummit | 1 |
| 2006 | Inversion transduction Grammar Coverage of Arabic-English Word Alignment for Tree-Structured Statistical Machine TranslationabstractWe present the first known direct measurement of word alignment coverage on an Arabic-English parallel corpus using inversion transduction grammar constraints. While direct measurements have been reported for several European and Asian languages, to date no results have been available for Arabic or any Semitic language despite much recent activity on Arabic- English spoken language and text translation. Many recent syntax based statistical MT models operate within the domain of ITG expressiveness, often for efficiency reasons, so it has become important to determine the extent to which the ITG constraint assumption holds. Our results on Arabic provide further evidence that ITG expressiveness appears largely sufficient for core MT models. Dekai Wu, Marine Carpuat, Yihai Shen |
SLT | 2 |
| 2006 | Aligning word senses using bilingual corporaabstractThe growing importance of multilingual information retrieval and machine translation has made multilingual ontologies extremely valuable resources. Since the construction of an ontology from scratch is a very expensive and time-consuming undertaking, it is attractive to consider ways of automatically aligning monolingual ontologies, which already exist for many of the world's major languages. Previous research exploited similarity in the structure of the ontologies to align, or manually created bilingual resources. These approaches cannot be used to align ontologies with vastly different structures and can only be applied to much studied language pairs for which expensive resources are already available. In this paper, we propose a novel approach to align the ontologies at the node level: Given a concept represented by a particular word sense in one ontology, our task is to find the best corresponding word sense in the second language ontology. To this end, we present a language-independent, corpus-based method that borrows from techniques used in information retrieval and machine translation. We show its efficiency by applying it to two very different ontologies in very different languages: the Mandarin Chinese HowNet and the American English WordNet. Moreover, we propose a methodology to measure bilingual corpora comparability and show that our method is robust enough to use noisy nonparallel bilingual corpora efficiently, when clean parallel corpora are not available. Marine Carpuat, Pascale Fung, Grace Ngai |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2005 | Word Sense Disambiguation vs. Statistical Machine TranslationabstractWe directly investigate a subject of much recent debate: do word sense disambiguation models help statistical machine translation quality? We present empirical results casting doubt on this common, but unproved, assumption. Using a state-of-the-art Chinese word sense disambiguation model to choose translation candidates for a typical IBM statistical MT system, we find that word sense disambiguation does not yield significantly better translation quality than the statistical machine translation system alone. Error analysis suggests several key factors behind this surprising finding, including inherent limitations of current statistical MT architectures. Marine Carpuat, Dekai Wu |
ACL | 1 |
| 2004 | A Kernel PCA Method for Superior Word Sense DisambiguationabstractWe introduce a new method for disambiguating word senses that exploits a nonlinear Kernel Principal Component Analysis (KPCA) technique to achieve accuracy superior to the best published individual models. We present empirical results demonstrating significantly better accuracy compared to the state-of-the-art achieved by either naïve Bayes or maximum entropy models, on Senseval-2 data. We also contrast against another type of kernel method, the support vector machine (SVM) model, and show that our KPCA-based model outperforms the SVM-based model. It is hoped that these highly encouraging first results on KPCA for natural language processing tasks will inspire further development of these directions. Dekai Wu, Weifeng Su, Marine Carpuat |
ACL | 3 |
| 2004 | Semi-supervised training of a Kernel PCA-Based Model for Word Sense Disambiguation
Weifeng Su, Marine Carpuat, Dekai Wu |
COLING | 2 |
| 2004 | Why Nitpicking Works: Evidence for Occam's Razor in Error Correctors
Dekai Wu, Grace Ngai, Marine Carpuat |
COLING | 3 |
| 2004 | NTPC: N-fold Templated Piped Correction
Dekai Wu, Grace Ngai, Marine Carpuat |
IJCNLP | 3 |
| 2004 | Raising the Bar: Stacked Conservative Error Correction Beyond Boosting
Dekai Wu, Grace Ngai, Marine Carpuat |
LREC | 3 |
| 2003 | A Stacked, Voted, Stacked Model for Named Entity Recognition
Dekai Wu, Grace Ngai, Marine Carpuat |
CoNLL | 3 |
| 2002 | Identifying Concepts Across Languages: A First Step towards a Corpus-based Approach to Automatic Ontology Alignment
Grace Ngai, Marine Carpuat, Pascale Fung |
COLING | 2 |
| 2002 | Boosting for Named Entity Recognition
Dekai Wu, Grace Ngai, Marine Carpuat, Jeppe Larsen |
CoNLL | 3 |