EDBT 2026 Demo / reviewers in the wild / expert
Carolina Scarton
dblp:23/8672 · also Carolina Evaristo Scarton
· DBLP profile ↗
36ranked-venue papers
9as first author
23since 2021 · last 2026
0000-0002-0103-4072ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 9 first-author · 19 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Survey on Automatic Credibility Assessment Using Textual Credibility Signals in the Era of Large Language ModelsabstractIn the age of social media and generative AI, the ability to automatically assess the credibility of online content has become increasingly critical, complementing traditional approaches to false information detection. Credibility assessment relies on aggregating diverse credibility signals—small units of information, such as content subjectivity, bias or a presence of persuasion techniques—into a final credibility label/score. However, current research in automatic credibility assessment and credibility signals detection remains highly fragmented, with many signals studied in isolation and lacking integration. Notably, there is a scarcity of approaches that detect and aggregate multiple credibility signals simultaneously. These challenges are further exacerbated by the absence of a comprehensive and up-to-date overview of research works that connects these research efforts under a common framework and identifies shared trends, challenges and open problems. In this survey, we address this gap by presenting a systematic and comprehensive literature review of 175 research papers, focusing on textual credibility signals within the field of Natural Language Processing (NLP), which undergoes a rapid transformation due to advancements in Large Language Models (LLMs). While positioning the NLP research into the broader multidisciplinary landscape, we examine both automatic credibility assessment methods as well as the detection of nine categories of credibility signals. We provide an in-depth analysis of three key categories: (1) factuality, subjectivity and bias, (2) persuasion techniques and logical fallacies and (3) check-worthy and fact-checked claims. In addition to summarising existing methods, datasets and tools, we outline future research direction and emerging opportunities, with particular attention to evolving challenges posed by generative AI. Ivan Srba, Olesya Razuvayevskaya, João Augusto Leite, Róbert Móro, Ipek Baris Schlicht, Sara Tonelli, Francisco Moreno García, Santiago Barrio Lottmann, Denis Teyssou, Valentin Porcellini, Carolina Scarton, Kalina Bontcheva, Mária Bieliková |
ACM Trans. Intell. Syst. Technol. | 11 |
| 2025 | It's All About In-Context Learning! Teaching Extremely Low-Resource Languages to LLMsabstractExtremely low-resource languages, especially those written in rare scripts, as shown in Figure 1, remain largely unsupported by large language models (LLMs).This is due in part to compounding factors such as the lack of training data.This paper delivers the first comprehensive analysis of whether LLMs can acquire such languages purely via in-context learning (ICL), with or without auxiliary alignment signals, and how these methods compare to parameter-efficient fine-tuning (PEFT).We systematically evaluate 20 under-represented languages across three state-of-the-art multilingual LLMs.Our findings highlight the limitation of PEFT when both language and its script are extremely under-represented by the LLM.In contrast, zero-shot ICL with language alignment is impressively effective on extremely low-resource languages, while fewshot ICL or PEFT is more beneficial for languages relatively better represented by LLMs.For LLM practitioners working on extremely low-resource languages, we summarise guidelines grounded by our results on adapting LLMs to low-resource languages, e.g., avoiding fine-tuning a multilingual model on languages of unseen scripts. Zhixue Zhao, Carolina Scarton |
EMNLP | 3 |
| 2025 | Label Set Optimization via Activation Distribution Kurtosis for Zero-Shot Classification with Generative ModelsabstractIn-context learning (ICL) performance is highly sensitive to prompt design, yet the impact of class label options (e.g.lexicon or order) in zero-shot classification remains underexplored.This study proposes LOADS (Label set Optimization via Activation Distribution kurtosiS), a post-hoc method for selecting optimal label sets in zero-shot ICL with large language models (LLMs).LOADS is built upon the observations in our empirical analysis, the first to systematically examine how label option design (i.e., lexical choice, order, and elaboration) impacts classification performance.This analysis shows that the lexical choice of the labels in the prompt (such as agree vs. support in stance classification) plays an important role in both model performance and model's sensitivity to the label order.A further investigation demonstrates that optimal label words tend to activate fewer outlier neurons in LLMs' feed-forward networks.LOADS then leverages kurtosis to measure the neuron activation distribution for label selection, requiring only a single forward pass without gradient propagation or labelled data.The LOADS-selected label words consistently demonstrate effectiveness for zero-shot ICL across classification tasks, datasets, models and languages, achieving maximum performance gain from 0.54 to 0.76 compared to the conventional approach of using original dataset label words. Zhixue Zhao, Carolina Scarton |
EMNLP | 3 |
| 2025 | UKElectionNarratives: A Dataset of Misleading Narratives Surrounding Recent UK General ElectionsabstractMisleading narratives play a crucial role in shaping public opinion during elections, as they can influence how voters perceive candidates and political parties. This entails the need to detect these narratives accurately. To address this, we introduce the first taxonomy of common misleading narratives that circulated during recent elections in Europe. Based on this taxonomy, we construct and analyse UKElectionNarratives: the first dataset of human-annotated misleading narratives which circulated during the UK General Elections in 2019 and 2024. We also benchmark Pre-trained and Large Language Models (focusing on GPT-4o), studying their effectiveness in detecting election-related misleading narratives. Finally, we discuss potential use cases and make recommendations for future research directions using the proposed codebook and dataset. Fatima Haouari, Carolina Scarton, Nicolò Faggiani, Nikolaos Nikolaidis 0004, Bonka Kotseva, Ibrahim Abu Farha, Jens P. Linge, Kalina Bontcheva |
ICWSM | 2 |
| 2024 | A Lightweight Approach for User and Keyword Classification in Controversial Topics
Ahmad Zareie, Kalina Bontcheva, Carolina Scarton |
ASONAM (2) | 3 |
| 2024 | EUvsDisinfo: A Dataset for Multilingual Detection of Pro-Kremlin Disinformation in News ArticlesabstractThis work introduces EUvsDisinfo, a multilingual dataset of disinformation articles originating from pro-Kremlin outlets, along with trustworthy articles from credible / less biased sources. It is sourced directly from the debunk articles written by experts leading the EUvsDisinfo project. Our dataset is the largest to-date resource in terms of the overall number of articles and distinct languages. It also provides the largest topical and temporal coverage. Using this dataset, we investigate the dissemination of pro-Kremlin disinformation across different languages, uncovering language-specific patterns targeting certain disinformation topics. We further analyse the evolution of topic distribution over an eight-year period, noting a significant surge in disinformation content before the full-scale invasion of Ukraine in 2022. Lastly, we demonstrate the dataset's applicability in training models to effectively distinguish between disinformation and trustworthy content in multilingual settings. João Augusto Leite, Olesya Razuvayevskaya, Kalina Bontcheva, Carolina Scarton |
CIKM | 4 |
| 2024 | Can We Identify Stance without Target Arguments? A Study for Rumour Stance ClassificationabstractConsidering a conversation thread, rumour stance classification aims to identify the opinion (e.g. agree or disagree) of replies towards a target (rumour story). Although the target is expected to be an essential component in traditional stance classification, we show that rumour stance classification datasets contain a considerable amount of real-world data whose stance could be naturally inferred directly from the replies, contributing to the strong performance of the supervised models without awareness of the target. We find that current target-aware models underperform in cases where the context of the target is crucial. Finally, we propose a simple yet effective framework to enhance reasoning with the targets, achieving state-of-the-art performance on two benchmark datasets. Carolina Scarton |
LREC/COLING | 2 |
| 2024 | Navigating Prompt Complexity for Zero-Shot Classification: A Study of Large Language Models in Computational Social ScienceabstractInstruction-tuned Large Language Models (LLMs) have exhibited impressive language understanding and the capacity to generate responses that follow specific prompts. However, due to the computational demands associated with training these models, their applications often adopt a zero-shot setting. In this paper, we evaluate the zero-shot performance of two publicly accessible LLMs, ChatGPT and OpenAssistant, in the context of six Computational Social Science classification tasks, while also investigating the effects of various prompting strategies. Our experiments investigate the impact of prompt complexity, including the effect of incorporating label definitions into the prompt; use of synonyms for label names; and the influence of integrating past memories during foundation model training. The findings indicate that in a zero-shot setting, current LLMs are unable to match the performance of smaller, fine-tuned baseline transformer models (such as BERT-large). Additionally, we find that different prompting strategies can significantly affect classification accuracy, with variations in accuracy and F1 scores exceeding 10%. Yida Mu, Ben Wu 0001, William Thorne, Ambrose Robinson, Nikolaos Aletras, Carolina Scarton, Kalina Bontcheva, Xingyi Song |
LREC/COLING | 6 |
| 2024 | Reference-less Analysis of Context Specificity in Translation with Personalised Language ModelsabstractSensitising language models (LMs) to external context helps them to more effectively capture the speaking patterns of individuals with specific characteristics or in particular environments. This work investigates to what extent detailed character and film annotations can be leveraged to personalise LMs in a scalable manner. We then explore the use of such models in evaluating context specificity in machine translation. We build LMs which leverage rich contextual information to reduce perplexity by up to 6.5% compared to a non-contextual model, and generalise well to a scenario with no speaker-specific data, relying on combinations of demographic characteristics expressed via metadata. Our findings are consistent across two corpora, one of which (Cornell-rich) is also a contribution of this paper. We then use our personalised LMs to measure the co-occurrence of extra-textual context and translation hypotheses in a machine translation setting. Our results suggest that the degree to which professional translations in our domain are context-specific can be preserved to a better extent by a contextual machine translation model than a non-contextual model, which is also reflected in the contextual model’s superior reference-based scores. Sebastian T. Vincent, Rowanne Sumner, Alice Dowek, Charlotte Prescott, Emily Preston, Chris Bayliss, Chris Oakley, Carolina Scarton |
LREC/COLING | 8 |
| 2024 | Multilinguality in the VIGILANT projectabstractVIGILANT (Vital IntelliGence to Investigate ILlegAl DisiNformaTion) is a three-year Horizon Europe project that will equip European Law Enforcement Agencies (LEAs) with advanced disinformation detection and analysis tools to investigate and prevent criminal activities linked to disinformation. These include disinformation instigating violence towards minorities, promoting false medical cures, and increasing tensions between groups causing civil unrest and violent acts. VIGILANT’s four LEAs require support for English, Spanish, Catalan, Greek, Estonian, Romanian and Russian. Therefore, multilinguality is a major challenge and we present the current status of our tools and our plans to improve their performance. Brendan Spillane, Carolina Scarton, Róbert Móro, Petar Ivanov, Andrey Tagarev, Jakub Simko, Ibrahim Abu Farha, Gary Munnelly, Filip Uhlárik, Freddy Heppell |
EAMT (2) | 2 |
| 2024 | ExU: AI Models for Examining Multilingual Disinformation Narratives and Understanding their SpreadabstractAddressing online disinformation requires analysing narratives across languages to help fact-checkers and journalists sift through large amounts of data. The ExU project focuses on developing AI-based models for multilingual disinformation analysis, addressing the tasks of rumour stance classification and claim retrieval. We describe the ExU project proposal and summarise the results of a user requirements survey regarding the design of tools to support fact-checking. Jake Vasilakes, Zhixue Zhao, Michal Gregor, Ivan Vykopal, Martin Hyben, Carolina Scarton |
EAMT (2) | 6 |
| 2024 | A Case Study on Contextual Machine Translation in a Professional Scenario of SubtitlingabstractIncorporating extra-textual context such as film metadata into the machine translation (MT) pipeline can enhance translation quality, as indicated by automatic evaluation in recent work. However, the positive impact of such systems in industry remains unproven. We report on an industrial case study carried out to investigate the benefit of MT in a professional scenario of translating TV subtitles with a focus on how leveraging extra-textual context impacts post-editing. We found that post-editors marked significantly fewer context-related errors when correcting the outputs of MTCue, the context-aware model, as opposed to non-contextual models. We also present the results of a survey of the employed post-editors, which highlights contextual inadequacy as a significant gap consistently observed in MT. Our findings strengthen the motivation for further work within fully contextual MT. Sebastian T. Vincent, Charlotte Prescott, Chris Bayliss, Chris Oakley, Carolina Scarton |
EAMT (1) | 5 |
| 2024 | Accelerating discoveries in medicine using distributed vector representations of words
Matheus V. V. Berto, Breno L. Freitas, Carolina Scarton, João A. Machado-Neto, Tiago A. Almeida 0001 |
Expert Syst. Appl. | 3 |
| 2023 | Domain-Driven and Discourse-Guided Scientific Summarisation
Tomas Goldsack, Zhihao Zhang 0004, Chenghua Lin 0002, Carolina Scarton |
ECIR (1) | 4 |
| 2023 | Enhancing Biomedical Lay Summarisation with External Knowledge GraphsabstractPrevious approaches for automatic lay summarisation are exclusively reliant on the source article that, given it is written for a technical audience (e.g., researchers), is unlikely to explicitly define all technical concepts or state all of the background information that is relevant for a lay audience.We address this issue by augmenting eLife, an existing biomedical lay summarisation dataset, with article-specific knowledge graphs, each containing detailed information on relevant biomedical concepts.Using both automatic and human evaluations, we systematically investigate the effectiveness of three different approaches for incorporating knowledge graphs within lay summarisation models, with each method targeting a distinct area of the encoder-decoder model architecture.Our results confirm that integrating graphbased domain knowledge can significantly benefit lay summarisation by substantially increasing the readability of generated text and improving the explanation of technical concepts.1 Tomas Goldsack, Zhihao Zhang 0004, Carolina Scarton, Chenghua Lin 0002 |
EMNLP | 4 |
| 2023 | Analysing State-Backed Propaganda Websites: a New Dataset and Linguistic StudyabstractThis paper analyses two hitherto unstudied sites sharing state-backed disinformation, Reliable Recent News (rrn.world) and WarOnFakes (waronfakes.com),which publish content in Arabic, Chinese, English, French, German, and Spanish.We describe our content acquisition methodology and perform cross-site unsupervised topic clustering on the resulting multilingual dataset.We also perform linguistic and temporal analysis of the web page translations and topics over time, and investigate articles with false publication dates.We make publicly available this new dataset of 14,053 articles, annotated with each language version, and additional metadata such as links and images.The main contribution of this paper for the NLP community is in the novel dataset which enables studies of disinformation networks, and the training of NLP tools for disinformation detection. Freddy Heppell, Kalina Bontcheva, Carolina Scarton |
EMNLP | 3 |
| 2023 | VaxxHesitancy: A Dataset for Studying Hesitancy towards COVID-19 Vaccination on TwitterabstractVaccine hesitancy has been a common concern, probably since vaccines were created and, with the popularisation of social media, people started to express their concerns about vaccines online alongside those posting pro- and anti-vaccine content. Predictably, since the first mentions of a COVID-19 vaccine, social media users posted about their fears and concerns or about their support and belief into the effectiveness of these rapidly developing vaccines. Identifying and understanding the reasons behind public hesitancy towards COVID-19 vaccines is important for policy markers that need to develop actions to better inform the population with the aim of increasing vaccine take-up. In the case of COVID-19, where the fast development of the vaccines was mirrored closely by growth in anti-vaxx disinformation, automatic means of detecting citizen attitudes towards vaccination became necessary. This is an important computational social sciences task that requires data analysis in order to gain in-depth understanding of the phenomena at hand. Annotated data is also necessary for training data-driven models for more nuanced analysis of attitudes towards vaccination. To this end, we created a new collection of over 3,101 tweets annotated with users' attitudes towards COVID-19 vaccination (stance). Besides, we also develop a domain-specific language model (VaxxBERT) that achieves the best predictive performance (73.0 accuracy and 69.3 F1-score) as compared to a robust set of baselines. To the best of our knowledge, these are the first dataset and model that model vaccine hesitancy as a category distinct from pro- and anti-vaccine stance. Yida Mu, Mali Jin, Charlie Grimshaw, Carolina Scarton, Kalina Bontcheva, Xingyi Song |
ICWSM | 4 |
| 2022 | Controlling Extra-Textual Attributes about Dialogue Participants: A Case Study of English-to-Polish Neural Machine TranslationabstractUnlike English, morphologically rich languages can reveal characteristics of speakers or their conversational partners, such as gender and number, via pronouns, morphological endings of words and syntax. When translating from English to such languages, a machine translation model needs to opt for a certain interpretation of textual context, which may lead to serious translation errors if extra-textual information is unavailable. We investigate this challenge in the English-to-Polish language direction. We focus on the underresearched problem of utilising external metadata in automatic translation of TV dialogue, proposing a case study where a wide range of approaches for controlling attributes in translation is employed in a multi-attribute scenario. The best model achieves an improvement of +5.81 chrF++/+6.03 BLEU, with other models achieving competitive performance. We additionally contribute a novel attribute-annotated dataset of Polish TV dialogue and a morphological analysis script used to evaluate attribute control in models. Sebastian T. Vincent, Loïc Barrault, Carolina Scarton |
EAMT | 3 |
| 2022 | Making Science Simple: Corpora for the Lay Summarisation of Scientific LiteratureabstractLay summarisation aims to jointly summarise and simplify a given text, thus making its content more comprehensible to non-experts.Automatic approaches for lay summarisation can provide significant value in broadening access to scientific literature, enabling a greater degree of both interdisciplinary knowledge sharing and public understanding when it comes to research findings.However, current corpora for this task are limited in their size and scope, hindering the development of broadly applicable data-driven approaches.Aiming to rectify these issues, we present two novel lay summarisation datasets, PLOS (large-scale) and eLife (medium-scale), each of which contains biomedical journal articles alongside expert-written lay summaries.We provide a thorough characterisation of our lay summaries, highlighting differing levels of readability and abstractiveness between datasets that can be leveraged to support the needs of different applications.Finally, we benchmark our datasets using mainstream summarisation approaches and perform a manual evaluation with domain experts, demonstrating their utility and casting light on the key challenges of this task.Our code and datasets are available at https://github.com/TGoldsack1/ Corpora_for_Lay_Summarisation. Tomas Goldsack, Zhihao Zhang 0004, Chenghua Lin 0002, Carolina Scarton |
EMNLP | 4 |
| 2022 | Improving Tokenisation by Alternative Treatment of SpacesabstractTokenisation is the first step in almost all NLP tasks, and state-of-the-art transformer-based language models all use subword tokenisation algorithms to process input text.Existing algorithms have problems, often producing tokenisations of limited linguistic validity and representing equivalent strings differently depending on their position within a word.We hypothesise that these problems hinder the ability of transformer-based models to handle complex words, and suggest that these problems are a result of allowing tokens to include spaces.We thus experiment with an alternative tokenisation approach where spaces are always treated as individual tokens.Specifically, we apply this modification to the BPE and Unigram algorithms.We find that our modified algorithms lead to improved performance on downstream NLP tasks that involve handling complex words, whilst having no detrimental effect on performance in general natural language understanding tasks.Intrinsically, we find that our modified algorithms give more morphologically correct tokenisations, in particular when handling prefixes.Given the results of our experiments, we advocate for always treating spaces as individual tokens as an improved tokenisation method. Edward Gow-Smith, Harish Tayyar Madabushi, Carolina Scarton, Aline Villavicencio |
EMNLP | 3 |
| 2021 | Assessing the Representations of Idiomaticity in Vector Models with a Noun Compound Dataset Labeled at Type and Token LevelsabstractMarcos Garcia, Tiago Kramer Vieira, Carolina Scarton, Marco Idiart, Aline Villavicencio. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Marcos García 0001, Tiago Kramer Vieira, Carolina Scarton, Marco Idiart, Aline Villavicencio |
ACL/IJCNLP (1) | 3 |
| 2021 | Probing for idiomaticity in vector space modelsabstractMarcos Garcia, Tiago Kramer Vieira, Carolina Scarton, Marco Idiart, Aline Villavicencio. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Marcos García 0001, Tiago Kramer Vieira, Carolina Scarton, Marco Idiart, Aline Villavicencio |
EACL | 3 |
| 2021 | The (Un)Suitability of Automatic Evaluation Metrics for Text SimplificationabstractAbstract In order to simplify sentences, several rewriting operations can be performed, such as replacing complex words per simpler synonyms, deleting unnecessary information, and splitting long sentences. Despite this multi-operation nature, evaluation of automatic simplification systems relies on metrics that moderately correlate with human judgments on the simplicity achieved by executing specific operations (e.g., simplicity gain based on lexical replacements). In this article, we investigate how well existing metrics can assess sentence-level simplifications where multiple operations may have been applied and which, therefore, require more general simplicity judgments. For that, we first collect a new and more reliable data set for evaluating the correlation of metrics and human judgments of overall simplicity. Second, we conduct the first meta-evaluation of automatic metrics in Text Simplification, using our new data set (and other existing data) to analyze the variation of the correlation between metrics’ scores and human judgments across three dimensions: the perceived simplicity level, the system type, and the set of references used for computation. We show that these three aspects affect the correlations and, in particular, highlight the limitations of commonly used operation-specific metrics. Finally, based on our findings, we propose a set of recommendations for automatic evaluation of multi-operation simplifications, suggesting which metrics to compute and how to interpret their scores. Fernando Alva-Manchego, Carolina Scarton, Lucia Specia |
Comput. Linguistics | 2 |
| 2020 | ASSET: A Dataset for Tuning and Evaluation of Sentence Simplification Models with Multiple Rewriting TransformationsabstractIn order to simplify a sentence, human editors perform multiple rewriting transformations: they split it into several shorter sentences, paraphrase words (i.e.replacing complex words or phrases by simpler synonyms), reorder components, and/or delete information deemed unnecessary.Despite these varied range of possible text alterations, current models for automatic sentence simplification are evaluated using datasets that are focused on a single transformation, such as lexical paraphrasing or splitting.This makes it impossible to understand the ability of simplification models in more realistic settings.To alleviate this limitation, this paper introduces ASSET, a new dataset for assessing sentence simplification in English.ASSET is a crowdsourced multi-reference corpus where each simplification was produced by executing several rewriting transformations.Through quantitative and qualitative experiments, we show that simplifications in ASSET are better at capturing characteristics of simplicity when compared to other standard evaluation datasets for the task.Furthermore, we motivate the need for developing better methods for automatic evaluation using ASSET, since we show that current popular metrics may not be suitable when multiple simplification transformations are performed. Fernando Alva-Manchego, Louis Martin, Antoine Bordes, Carolina Scarton, Benoît Sagot, Lucia Specia |
ACL | 4 |
| 2020 | Deciding When, How and for Whom to SimplifyabstractCurrent Automatic Text Simplification (TS) work relies on sequence-to-sequence neural models that learn simplification operations from parallel complex-simple corpora. In this paper we address three open challenges in these approaches: (i) avoiding unnecessary transformations, (ii) determining which operations to perform, and (iii) generating simplifications that are suitable for a given target audience. For (i), we propose joint and two-stage approaches where instances are marked or classified as simple or complex. For (ii) and (iii), we propose fusion-based approaches to incorporate information on the target grade level as well as the types of operation to perform in the models. While grade-level information is provided as metadata, we devise predictors for the type of operation. We study different representations for this information as well as different ways in which it is used in the models. Our approach outperforms previous work on neural TS, with our best model following the two-stage approach and using the information about grade level and type of operation to initialise the encoder and the decoder, respectively. Carolina Scarton, Pranava Swaroop Madhyastha, Lucia Specia |
ECAI | 1 |
| 2020 | Measuring the Impact of Readability Features in Fake News DetectionabstractThe proliferation of fake news is a current issue that influences a number of important areas of society, such as politics, economy and health. In the Natural Language Processing area, recent initiatives tried to detect fake news in different ways, ranging from language-based approaches to content-based verification. In such approaches, the choice of the features for the classification of fake and true news is one of the most important parts of the process. This paper presents a study on the impact of readability features to detect fake news for the Brazilian Portuguese language. The results show that such features are relevant to the task (achieving, alone, up to 92% classification accuracy) and may improve previous classification results. Roney Lira de Sales Santos, Gabriela Wick-Pedro, Sidney Evaldo Leal, Oto A. Vale, Thiago A. S. Pardo, Kalina Bontcheva, Carolina Scarton |
LREC | 7 |
| 2020 | Data-Driven Sentence Simplification: Survey and BenchmarkabstractSentence Simplification (SS) aims to modify a sentence in order to make it easier to read and understand. In order to do so, several rewriting transformations can be performed such as replacement, reordering, and splitting. Executing these transformations while keeping sentences grammatical, preserving their main idea, and generating simpler output, is a challenging and still far from solved problem. In this article, we survey research on SS, focusing on approaches that attempt to learn how to simplify using corpora of aligned original-simplified sentence pairs in English, which is the dominant paradigm nowadays. We also include a benchmark of different approaches on common data sets so as to compare them and highlight their strengths and limitations. We expect that this survey will serve as a starting point for researchers interested in the task and help spark new ideas for future developments. Fernando Alva-Manchego, Carolina Scarton, Lucia Specia |
Comput. Linguistics | 2 |
| 2020 | Horacio Saggion, Automatic Text Simplification. Synthesis lectures on human language technologies, April 2017
Carolina Scarton |
Nat. Lang. Eng. | 1 |
| 2018 | Text Simplification from Professionally Produced Corpora
Carolina Scarton, Gustavo Paetzold, Lucia Specia |
LREC | 1 |
| 2018 | SimPA: A Sentence-Level Simplification Corpus for the Public Administration Domain
Carolina Scarton, Gustavo Paetzold, Lucia Specia |
LREC | 1 |
| 2017 | Learning How to Simplify From Explicit Labeling of Complex-Simplified Text PairsabstractCurrent research in text simplification has been hampered by two central problems: (i) the small amount of high-quality parallel simplification data available, and (ii) the lack of explicit annotations of simplification operations, such as deletions or substitutions, on existing data. While the recently introduced Newsela corpus has alleviated the first problem, simplifications still need to be learned directly from parallel text using black-box, end-to-end approaches rather than from explicit annotations. These complex-simple parallel sentence pairs often differ to such a high degree that generalization becomes difficult. End-to-end models also make it hard to interpret what is actually learned from data. We propose a method that decomposes the task of TS into its sub-problems. We devise a way to automatically identify operations in a parallel corpus and introduce a sequence-labeling approach based on these annotations. Finally, we provide insights on the types of transformations that different approaches can model. Fernando Alva-Manchego, Joachim Bingel, Gustavo Paetzold, Carolina Scarton, Lucia Specia |
IJCNLP(1) | 4 |
| 2016 | A Reading Comprehension Corpus for Machine Translation Evaluation
Carolina Scarton, Lucia Specia |
LREC | 1 |
| 2015 | Searching for Context: a Study on Document-Level Labels for Translation Quality Estimation
Carolina Scarton, Marcos Zampieri, Mihaela Vela, Josef van Genabith, Lucia Specia |
EAMT | 1 |
| 2015 | Discourse and Document-level Information for Evaluating Language Output TasksabstractEvaluating the quality of language output tasks such as Machine Translation (MT) and Automatic Summarisation (AS) is a challenging topic in Natural Language Processing (NLP).Recently, techniques focusing only on the use of outputs of the systems and source information have been investigated.In MT, this is referred to as Quality Estimation (QE), an approach that uses machine learning techniques to predict the quality of unseen data, generalising from a few labelled data points.Traditional QE research addresses sentencelevel QE evaluation and prediction, disregarding document-level information.Documentlevel QE requires a different set up from sentence-level, which makes the study of appropriate quality scores, features and models necessary.Our aim is to explore documentlevel QE of MT, focusing on discourse information.However, the findings of this research can improve other NLP tasks, such as AS. Carolina Scarton |
HLT-NAACL | 1 |
| 2014 | Verb Clustering for Brazilian Portuguese
Carolina Scarton, Lin Sun 0003, Karin Kipper Schuler, Magali Sanches Duran, Martha Palmer, Anna Korhonen |
CICLing (1) | 1 |
| 2014 | Document-level translation quality estimation: exploring discourse and pseudo-references
Carolina Scarton, Lucia Specia |
EAMT | 1 |