Sweta Agrawal

dblp:210/7863 · DBLP profile ↗
← Back
22ranked-venue papers
9as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 8 first-author · 18 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 A Context-aware Framework for Translation-mediated Conversations
abstract
Abstract Automatic translation systems offer a powerful solution to bridge language barriers in scenarios where participants do not share a common language. However, these systems can introduce errors leading to misunderstandings and conversation breakdown. A key issue is that current systems fail to incorporate the rich contextual information necessary to resolve ambiguities and omitted details, resulting in literal, inappropriate, or misaligned translations. In this work, we present a framework to improve large language model-based translation systems by incorporating contextual information in bilingual conversational settings during training and inference. We validate our proposed framework on two task-oriented domains: customer chat and user-assistant interaction. Across both settings, the system produced by our framework—TowerChat—consistently results in better translations than state-of-the-art systems like GPT-4o and TowerInstruct, as measured by multiple automatic translation quality metrics on several language pairs. We also show that the resulting model leverages context in an intended and interpretable way, improving consistency between the conveyed message and the generated translations.1
José Pombal, Sweta Agrawal, Emmanouil Zaranis, Patrick Fernandes, André F. T. Martins
Trans. Assoc. Comput. Linguistics2
2026 Fine-Grained Reward Optimization for Machine Translation using Error Severity Mappings
abstract
Abstract Reinforcement learning (RL) has been proven to be an effective and robust method for training neural machine translation systems, especially when paired with powerful reward models that accurately assess translation quality. However, most research has focused on RL methods that use sentence-level feedback, leading to inefficient learning signals due to the reward sparsity problem—the model receives a single score for the entire sentence. To address this, we propose a novel approach that leverages fine-grained, token-level quality assessments along with error severity levels using RL methods. Specifically, we use xCOMET, a state-of-the-art quality estimation system, as our token-level reward model. We conduct experiments on small and large translation datasets with standard encoder-decoder and large language models-based machine translation systems, comparing the impact of sentence-level versus fine-grained reward signals on translation quality. Our results show that training with token-level rewards improves translation quality across language pairs over baselines according to both automatic and human evaluation. Furthermore, token-level reward optimization improves training stability, evidenced by a steady increase in mean rewards over training epochs.
Miguel Moura Ramos, Tomás Almeida, Daniel Vareta, Filipe Parrado de Azevedo, Sweta Agrawal, Patrick Fernandes, André F. T. Martins
Trans. Assoc. Comput. Linguistics5
2025 Watching the Watchers: Exposing Gender Disparities in Machine Translation Quality Estimation
abstract
Quality estimation (QE)-the automatic assessment of translation quality-has recently become crucial across several stages of the translation pipeline, from data curation to training and decoding.While QE metrics have been optimized to align with human judgments, whether they encode social biases has been largely overlooked.Biased QE risks favoring certain demographic groups over others, e.g., by exacerbating gaps in visibility and usability.This paper defines and investigates gender bias of QE metrics and discusses its downstream implications for machine translation (MT).Experiments with state-ofthe-art QE metrics across multiple domains, datasets, and languages reveal significant bias.When a human entity's gender in the source is undisclosed, masculine-inflected translations score higher than feminine-inflected ones, and gender-neutral translations are penalized.Even when contextual cues disambiguate gender, using context-aware QE metrics leads to more errors in selecting the correct translation inflection for feminine referents than for masculine ones.Moreover, a biased QE metric affects data filtering and quality-aware decoding.Our findings underscore the need for a renewed focus on developing and evaluating QE metrics centered on gender. 1
Emmanouil Zaranis, Giuseppe Attanasio, Sweta Agrawal, André F. T. Martins
ACL (1)3
2025 Sustaining Human Agency, Attending to Its Cost: An Investigation into Generative AI Design for Non-Native Speakers' Language Use
Yimin Xiao, Cartor Hancock, Sweta Agrawal, Nikita Mehandru, Niloufar Salehi, Marine Carpuat, Ge Gao 0001
CHI3
2025 Translate Smart, not Hard: Cascaded Translation Systems with Quality-Aware Deferral
abstract
Larger models often outperform smaller ones but come with high computational costs.Cascading offers a potential solution.By default, it uses smaller models and defers only some instances to larger, more powerful models.However, designing effective deferral rules remains a challenge.In this paper, we propose a simple yet effective approach for machine translation, using existing quality estimation (QE) metrics as deferral rules.We show that QE-based deferral allows a cascaded system to match the performance of a larger model while invoking it for a small fraction (30% to 50%) of the examples, significantly reducing computational costs.We validate this approach through both automatic and human evaluation.
António Farinhas, Nuno Miguel Guerreiro, Sweta Agrawal, Ricardo Rei, André F. T. Martins
EMNLP3
2024 Can Automatic Metrics Assess High-Quality Translations?
abstract
Automatic metrics for evaluating translation quality are typically validated by measuring how well they correlate with human assessments.However, correlation methods tend to capture only the ability of metrics to differentiate between good and bad source-translation pairs, overlooking their reliability in distinguishing alternative translations for the same source.In this paper, we confirm that this is indeed the case by showing that current metrics are insensitive to nuanced differences in translation quality.This effect is most pronounced when the quality is high and the variance among alternatives is low.Given this finding, we shift towards detecting high-quality correct translations, an important problem in practical decision-making scenarios where a binary check of correctness is prioritized over a nuanced evaluation of quality.Using the MQM framework as the gold standard, we systematically stress-test the ability of current metrics to identify translations with no errors as marked by humans.Our findings reveal that current metrics often over or underestimate translation quality, indicating significant room for improvement in machine translation evaluation.
Sweta Agrawal, António Farinhas, Ricardo Rei, André F. T. Martins
EMNLP1
2024 Modeling User Preferences with Automatic Metrics: Creating a High-Quality Preference Dataset for Machine Translation
abstract
Sweta Agrawal, José G. C. De Souza, Ricardo Rei, António Farinhas, Gonçalo Faria, Patrick Fernandes, Nuno M Guerreiro, Andre Martins. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Sweta Agrawal, José Guilherme Camargo de Souza, Ricardo Rei, António Farinhas, Gonçalo Rui Alves Faria, Patrick Fernandes, Nuno Miguel Guerreiro, André F. T. Martins
EMNLP1
2024 AfriMTE and AfriCOMET: Enhancing COMET to Embrace Under-resourced African Languages
abstract
Jiayi Wang, David Ifeoluwa Adelani, Sweta Agrawal, Marek Masiak, Ricardo Rei, Eleftheria Briakou, Marine Carpuat, Xuanli He, Sofia Bourhim, Andiswa Bukula, Muhidin Mohamed, Temitayo Olatoye, Tosin Adewumi, Hamam Mokayed, Christine Mwase, Wangui Kimotho, Foutse Yuehgoh, Anuoluwapo Aremu, Jessica Ojo, Shamsuddeen Hassan Muhammad, Salomey Osei, Abdul-Hakeem Omotayo, Chiamaka Chukwuneke, Perez Ogayo, Oumaima Hourrane, Salma El Anigri, Lolwethu Ndolela, Thabiso Mangwana, Shafie Abdi Mohamed, Hassan Ayinde, Oluwabusayo Olufunke Awoyomi, Lama Alkhaled, Sana Al-azzawi, Naome A. Etori, Millicent Ochieng, Clemencia Siro, Njoroge Kiragu, Eric Muchiri, Wangari Kimotho, Lyse Naomi Wamba Momo, Daud Abolade, Simbiat Ajao, Iyanuoluwa Shode, Ricky Macharm, Ruqayya Nasir Iro, Saheed S. Abdullahi, Stephen E. Moore, Bernard Opoku, Zainab Akinjobi, Abeeb Afolabi, Nnaemeka Obiefuna, Onyekachi Raphael Ogbu, Sam Ochieng’, Verrah Akinyi Otiende, Chinedu Emmanuel Mbonu, Sakayo Toadoum Sari, Yao Lu, Pontus Stenetorp. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jiayi Wang 0010, David Ifeoluwa Adelani, Sweta Agrawal, Marek Masiak, Ricardo Rei, Eleftheria Briakou, Marine Carpuat, Xuanli He, Sofia Bourhim, Andiswa Bukula, Muhidin Mohamed, Temitayo Olatoye, Tosin P. Adewumi, Hamam Mokayed, Christine Mwase, Wangui Kimotho, Foutse Yuehgoh, Aremu Anuoluwapo, Jessica Ojo, Shamsuddeen Hassan Muhammad, Salomey Osei, Abdul-Hakeem Omotayo, Chiamaka Ijeoma Chukwuneke, Perez Ogayo, Oumaima Hourrane, Salma El Anigri, Lolwethu Ndolela, Thabiso Mangwana, Shafie Abdi Mohamed, Ayinde Hassan, Oluwabusayo Olufunke Awoyomi, Lama Alkhaled, Sana Sabah Al-Azzawi, Naome A. Etori, Millicent Ochieng, Clemencia Siro, Njoroge Kiragu, Eric Muchiri, Wangari Kimotho, Sakayo Toadoum Sari, Lyse Naomi Wamba Momo, Daud Abolade, Simbiat Ajao, Iyanuoluwa Shode, Ricky Macharm, Ruqayya Nasir Iro, Saheed S. Abdullahi, Stephen E. Moore, Bernard Opoku, Zainab Akinjobi, Afolabi Abeeb, Nnaemeka C. Obiefuna, Onyekachi Raphael Ogbu, Sam Ochieng', Verrah Otiende, Chinedu E. Mbonu, Pontus Stenetorp
NAACL-HLT3
2024 QUEST: Quality-Aware Metropolis-Hastings Sampling for Machine Translation
abstract
An important challenge in machine translation (MT) is to generate high-quality and diverse translations. Prior work has shown that the estimated likelihood from the MT model correlates poorly with translation quality. In contrast, quality evaluation metrics (such as COMET or BLEURT) exhibit high correlations with human judgments, which has motivated their use as rerankers (such as quality-aware and minimum Bayes risk decoding). However, relying on a single translation with high estimated quality increases the chances of "gaming the metric''. In this paper, we address the problem of sampling a set of high-quality and diverse translations. We provide a simple and effective way to avoid over-reliance on noisy quality estimates by using them as the energy function of a Gibbs distribution. Instead of looking for a mode in the distribution, we generate multiple samples from high-density areas through the Metropolis-Hastings algorithm, a simple Markov chain Monte Carlo approach. The results show that our proposed method leads to high-quality and diverse outputs across multiple language pairs (English$\leftrightarrow$\{German, Russian\}) with two strong decoder-only LLMs (Alma-7b, Tower-7b).
Gonçalo Rui Alves Faria, Sweta Agrawal, António Farinhas, Ricardo Rei, José Guilherme Camargo de Souza, André F. T. Martins
NeurIPS2
2024 Do Text Simplification Systems Preserve Meaning? A Human Evaluation via Reading Comprehension
abstract
Abstract Automatic text simplification (TS) aims to automate the process of rewriting text to make it easier for people to read. A pre-requisite for TS to be useful is that it should convey information that is consistent with the meaning of the original text. However, current TS evaluation protocols assess system outputs for simplicity and meaning preservation without regard for the document context in which output sentences occur and for how people understand them. In this work, we introduce a human evaluation framework to assess whether simplified texts preserve meaning using reading comprehension questions. With this framework, we conduct a thorough human evaluation of texts by humans and by nine automatic systems. Supervised systems that leverage pre-training knowledge achieve the highest scores on the reading comprehension tasks among the automatic controllable TS systems. However, even the best-performing supervised system struggles with at least 14% of the questions, marking them as “unanswerable” based on simplified content. We further investigate how existing TS evaluation metrics and automatic question-answering systems approximate the human judgments we obtained.
Sweta Agrawal, Marine Carpuat
Trans. Assoc. Comput. Linguistics1
2024 Assessing the Role of Context in Chat Translation Evaluation: Is Context Helpful and Under What Conditions?
abstract
Abstract Despite the recent success of automatic metrics for assessing translation quality, their application in evaluating the quality of machine-translated chats has been limited. Unlike more structured texts like news, chat conversations are often unstructured, short, and heavily reliant on contextual information. This poses questions about the reliability of existing sentence-level metrics in this domain as well as the role of context in assessing the translation quality. Motivated by this, we conduct a meta-evaluation of existing automatic metrics, primarily designed for structured domains such as news, to assess the quality of machine-translated chats. We find that reference-free metrics lag behind reference-based ones, especially when evaluating translation quality in out-of-English settings. We then investigate how incorporating conversational contextual information in these metrics for sentence-level evaluation affects their performance. Our findings show that augmenting neural learned metrics with contextual information helps improve correlation with human judgments in the reference-free scenario and when evaluating translations in out-of-English settings. Finally, we propose a new evaluation metric, Context-MQM, that utilizes bilingual context with a large language model (LLM) and further validate that adding context helps even for LLM-based evaluation metrics.
Sweta Agrawal, M. Amin Farajian, Patrick Fernandes, Ricardo Rei, André F. T. Martins
Trans. Assoc. Comput. Linguistics1
2023 Controlling Pre-trained Language Models for Grade-Specific Text Simplification
abstract
Text simplification (TS) systems rewrite text to make it more readable while preserving its content.However, what makes a text easy to read depends on the intended readers.Recent work has shown that pre-trained language models can simplify text using a wealth of techniques to control output simplicity, ranging from specifying only the desired reading grade level, to directly specifying low-level edit operations.Yet it remains unclear how to set these control parameters in practice.Existing approaches set them at the corpus level, disregarding the complexity of individual inputs and considering only one level of output complexity.In this work, we conduct an empirical study to understand how different control mechanisms impact the adequacy and simplicity of text simplification systems.Based on these insights, we introduce a simple method that predicts the edit operations required for simplifying a text for a specific grade level on an instance-per-instance basis.This approach improves the quality of the simplified outputs over corpus-level searchbased heuristics.
Sweta Agrawal, Marine Carpuat
EMNLP1
2023 BLESS: Benchmarking Large Language Models on Sentence Simplification
abstract
Tannon Kew, Alison Chi, Laura Vásquez-Rodríguez, Sweta Agrawal, Dennis Aumiller, Fernando Alva-Manchego, Matthew Shardlow. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Tannon Kew, Alison Chi, Laura Vásquez-Rodríguez, Sweta Agrawal, Dennis Aumiller, Fernando Alva-Manchego, Matthew Shardlow
EMNLP4
2023 Physician Detection of Clinical Harm in Machine Translation: Quality Estimation Aids in Reliance and Backtranslation Identifies Critical Errors
abstract
A major challenge in the practical use of Machine Translation (MT) is that users lack guidance to make informed decisions about when to rely on outputs.Progress in quality estimation research provides techniques to automatically assess MT quality, but these techniques have primarily been evaluated in vitro by comparison against human judgments outside of a specific context of use.This paper evaluates quality estimation feedback in vivo with a human study simulating decision-making in high-stakes medical settings.Using Emergency Department discharge instructions, we study how interventions based on quality estimation versus backtranslation assist physicians in deciding whether to show MT outputs to a patient.We find that quality estimation improves appropriate reliance on MT, but backtranslation helps physicians detect more clinically harmful errors that QE alone often misses.
Nikita Mehandru, Sweta Agrawal, Yimin Xiao, Ge Gao 0001, Elaine C. Khoong, Marine Carpuat, Niloufar Salehi
EMNLP2
2023 Understanding and Detecting Hallucinations in Neural Machine Translation via Model Introspection
abstract
Abstract Neural sequence generation models are known to “hallucinate”, by producing outputs that are unrelated to the source text. These hallucinations are potentially harmful, yet it remains unclear in what conditions they arise and how to mitigate their impact. In this work, we first identify internal model symptoms of hallucinations by analyzing the relative token contributions to the generation in contrastive hallucinated vs. non-hallucinated outputs generated via source perturbations. We then show that these symptoms are reliable indicators of natural hallucinations, by using them to design a lightweight hallucination detector which outperforms both model-free baselines and strong classifiers based on quality estimation or large pre-trained models on manually annotated English-Chinese and German-English translation test beds.
Weijia Xu, Sweta Agrawal, Eleftheria Briakou, Marianna J. Martindale, Marine Carpuat
Trans. Assoc. Comput. Linguistics2
2022 An Imitation Learning Curriculum for Text Editing with Non-Autoregressive Models
abstract
We propose a framework for training nonautoregressive sequence-to-sequence models for editing tasks, where the original input sequence is iteratively edited to produce the output.We show that the imitation learning algorithms designed to train such models for machine translation introduces mismatches between training and inference that lead to undertraining and poor generalization in editing scenarios.We address this issue with two complementary strategies: 1) a roll-in policy that exposes the model to intermediate training sequences that it is more likely to encounter during inference, 2) a curriculum that presents easy-to-learn edit operations first, gradually increasing the difficulty of training samples as the model becomes competent.We show the efficacy of these strategies on two challenging English editing tasks: controllable text simplification and abstractive summarization.Our approach significantly improves output quality on both tasks and controls output complexity better on the simplification task.
Sweta Agrawal, Marine Carpuat
ACL (1)1
2022 Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
abstract
Abstract With the success of large-scale pre-training and multilingual modeling in Natural Language Processing (NLP), recent years have seen a proliferation of large, Web-mined text datasets covering hundreds of languages. We manually audit the quality of 205 language-specific corpora released with five major public datasets (CCAligned, ParaCrawl, WikiMatrix, OSCAR, mC4). Lower-resource corpora have systematic issues: At least 15 corpora have no usable text, and a significant fraction contains less than 50% sentences of acceptable quality. In addition, many are mislabeled or use nonstandard/ambiguous language codes. We demonstrate that these issues are easy to detect even for non-proficient speakers, and supplement the human audit with automatic analyses. Finally, we recommend techniques to evaluate and improve multilingual corpora and discuss potential risks that come with low-quality data releases.
Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov 0001, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Rubungo Andre Niyongabo, Toan Q. Nguyen, Mathias Müller 0002, André Müller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Reddy Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine Çabuk Balli, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, Mofe Adeyemi
Trans. Assoc. Comput. Linguistics51
2021 Evaluating the Evaluation Metrics for Style Transfer: A Case Study in Multilingual Formality Transfer
abstract
While the field of style transfer (ST) has been growing rapidly, it has been hampered by a lack of standardized practices for automatic evaluation.In this paper, we evaluate leading ST automatic metrics on the oft-researched task of formality style transfer.Unlike previous evaluations, which focus solely on English, we expand our focus to Brazilian-Portuguese, French, and Italian, making this work the first multilingual evaluation of metrics in ST.We outline best practices for automatic evaluation in (formality) style transfer and identify several models that correlate well with human judgments and are robust across languages.We hope that this work will help accelerate development in ST, where human evaluation is often challenging to collect.
Eleftheria Briakou, Sweta Agrawal, Joel R. Tetreault, Marine Carpuat
EMNLP (1)2
2021 Assessing Reference-Free Peer Evaluation for Machine Translation
abstract
Sweta Agrawal, George Foster, Markus Freitag, Colin Cherry. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Sweta Agrawal, George F. Foster, Markus Freitag, Colin Cherry
NAACL-HLT1
2019 Controlling Text Complexity in Neural Machine Translation
abstract
Sweta Agrawal, Marine Carpuat. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Sweta Agrawal, Marine Carpuat
EMNLP/IJCNLP (1)1
2018 Deep Learning for Detecting Cyberbullying Across Multiple Social Media Platforms
Sweta Agrawal, Amit Awekar
ECIR1
2017 Smart Geo-fencing with Location Sensitive Product Affinity
abstract
Geo-fencing is a location based service that allows sending of messages to users who enter/exit a specified geographical area, known as a geo-fence. Today, it has become one of the popular location based mobile marketing strategies. However, the process of designing geo-fences is presently manual, i.e. a retailer must specify the location and the radius of area around it to setup the geo-fences. Moreover, this process does not consider the user's preference towards the targeted product/service and thus, can compromise his/her experience of the app that sends these communications. We attempt to solve this problem by presenting a novel end-to-end system for automated design of affinity based smart geo-fences. Affinity towards a product/service refers to the user's interest in a product/service. Our unique formulation to estimate affinity, using historical app usage data, is sensitive to a user's location and thus, the affinity is termed as location sensitive product affinity (LSPA). The geo-fence logic tries to capture contiguous groups of locations where the affinity high. Experiments on real world e-commerce dataset reveals that geo-fences designed by our approach performs significantly better at accurately targeting the users who are interested in a product. We thus show that, using historical app usage data, geo-fences can be designed in an automated manner and can help enterprises target interested users with better accuracy as compared to the present industry practices.
Ankur Garg, Sunav Choudhary, Payal Bajaj, Sweta Agrawal, Abhishek Kedia
SIGSPATIAL/GIS4