EDBT 2026 Demo / reviewers in the wild / expert
Lucia Specia
dblp:23/2900
· DBLP profile ↗
124ranked-venue papers
20as first author
27since 2021 · last 2025
0000-0002-5495-3128ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 119 · 18 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-authorTheory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DiffuseDef: Improved Robustness to Adversarial Attacks via Iterative DenoisingabstractPretrained language models have significantly advanced performance across various natural language processing tasks.However, adversarial attacks continue to pose a critical challenge to systems built using these models, as they can be exploited with carefully crafted adversarial texts.Inspired by the ability of diffusion models to predict and reduce noise in computer vision, we propose a novel and flexible adversarial defense method for language classification tasks, DiffuseDef 1 , which incorporates a diffusion layer as a denoiser between the encoder and the classifier.The diffusion layer is trained on top of the existing classifier, ensuring seamless integration with any model in a plug-and-play manner.During inference, the adversarial hidden state is first combined with sampled noise, then denoised iteratively and finally ensembled to produce a robust text representation.By integrating adversarial training, denoising, and ensembling techniques, we show that DiffuseDef improves over existing adversarial defense methods and achieves stateof-the-art performance against common blackbox and white-box adversarial attacks. Zhenhao Li 0003, Huichi Zhou, Marek Rei, Lucia Specia |
ACL (1) | 4 |
| 2025 | Discourse Features Enhance Detection of Document-Level Machine-Generated ContentabstractThe availability of high-quality APIs for Large Language Models (LLMs) has facilitated the widespread creation of Machine-Generated Content (MGC), posing challenges such as academic plagiarism and the spread of misinformation. Existing MGC detectors often focus solely on surface-level information, overlooking implicit and structural features. This makes them susceptible to deception by surface-level sentence patterns, particularly for longer texts and in texts that have been subsequently paraphrased. To overcome these challenges, we introduce novel methodologies and datasets. Besides the publicly available dataset Plagbench, we developed the paraphrased Long-Form Question and Answer (paraLFQA) and paraphrased Writing Prompts (paraWP) datasets using GPT and DIPPER, a discourse paraphrasing tool, by extending artifacts from their original versions. To better capture the structure of longer texts at document level, we propose DTransformer, a model that integrates discourse analysis through PDTB preprocessing to encode structural features. It results in substantial performance gains across both datasets – 15.5% absolute improvement on paraLFQA, 4% absolute improvement on paraWP, and 1.5% absolute improvemene on M4 compared to SOTA approaches. The data and code are available at this link1. Yupei Li, Manuel Milling, Lucia Specia, Björn W. Schuller |
IJCNN | 3 |
| 2024 | Classifying Social Media Users before and after Depression Diagnosis via Their Language Usage: A Dataset and StudyabstractMental illness can significantly impact individuals’ quality of life. Analysing social media data to uncover potential mental health issues in individuals via their posts is a popular research direction. However, most studies focus on the classification of users suffering from depression versus healthy users, or on the detection of suicidal thoughts. In this paper, we instead aim to understand and model linguistic changes that occur when users transition from a healthy to an unhealthy state. Addressing this gap could lead to better approaches for earlier depression detection when signs are not as obvious as in cases of severe depression or suicidal ideation. In order to achieve this goal, we have collected the first dataset of textual posts by the same users before and after reportedly being diagnosed with depression. We then use this data to build multiple predictive models (based on SVM, Random Forests, BERT, RoBERTa, MentalBERT, GPT-3, GPT-3.5, Bard, and Alpaca) for the task of classifying user posts. Transformer-based models achieved the best performance, while large language models used off-the-shelf proved less effective as they produced random guesses (GPT and Bard) or hallucinations (Alpaca). Falwah Alhamed, Julia Ive, Lucia Specia |
LREC/COLING | 3 |
| 2023 | A study towards contextual understanding of toxicity in online conversationsabstractAbstract Identifying and annotating toxic online content on social media platforms is an extremely challenging problem. Work that studies toxicity in online content has predominantly focused on comments as independent entities. However, comments on social media are inherently conversational, and therefore, understanding and judging the comments fundamentally requires access to the context in which they are made. We introduce a study and resulting annotated dataset where we devise a number of controlled experiments on the importance of context and other observable confounders – namely gender, age and political orientation – towards the perception of toxicity in online content. Our analysis clearly shows the significance of context and the effect of observable confounders on annotations. Namely, we observe that the ratio of toxic to non-toxic judgements can be very different for each control group, and a higher proportion of samples are judged toxic in the presence of contextual information. Pranava Swaroop Madhyastha, Antigoni Founta, Lucia Specia |
Nat. Lang. Eng. | 3 |
| 2022 | Bias Mitigation in Machine Translation Quality EstimationabstractMachine Translation Quality Estimation (QE) aims to build predictive models to assess the quality of machine-generated translations in the absence of reference translations.While state-of-the-art QE models have been shown to achieve good results, they over-rely on features that do not have a causal impact on the quality of a translation.In particular, there appears to be a partial input bias, i.e., a tendency to assign high-quality scores to translations that are fluent and grammatically correct, even though they do not preserve the meaning of the source.We analyse the partial input bias in further detail and evaluate four approaches to use auxiliary tasks for bias mitigation.Two approaches use additional data to inform and support the main task, while the other two are adversarial, actively discouraging the model from learning the bias.We compare the methods with respect to their ability to reduce the partial input bias while maintaining the overall performance.We find that training a multitask architecture with an auxiliary binary classification task that utilises additional augmented data best achieves the desired effects and generalises well to different languages and quality metrics. Hanna Behnke, Marina Fomicheva, Lucia Specia |
ACL (1) | 3 |
| 2022 | A Taxonomy and Study of Critical Errors in Machine TranslationabstractNot all machine mistranslations are equal. For example, mistranslating a date or time in an appointment, mistranslating the number or currency in a contract, or hallucinating profanity may lead to consequences for the users even when MT is just used for gisting. The severity of the errors is important, but overlooked, aspect of MT quality evaluation. In this paper, we present the result of our effort to bring awareness to the problem of critical translation errors. We study, validate and improve an initial taxonomy of critical errors with the view of providing guidance for critical error analysis, annotation and mitigation. We test the taxonomy for three different languages to examine to what extent it generalises across languages. We provide an account of factors that affect annotation tasks along with recommendations on how to improve the practice in future work. We also study the impact of the source text on generating critical errors in the translation and, based on this, propose a set of recommendations on aspects of the MT that need further scrutiny, especially for user-generated content, to avoid generating such errors, and hence improve online communication. Khetam Al Sharou, Lucia Specia |
EAMT | 2 |
| 2022 | Logically Consistent Adversarial Attacks for Soft Theorem ProversabstractRecent efforts within the AI community have yielded impressive results towards “soft theorem proving” over natural language sentences using language models. We propose a novel, generative adversarial framework for probing and improving these models’ reasoning capabilities. Adversarial attacks in this domain suffer from the logical inconsistency problem, whereby perturbations to the input may alter the label. Our Logically consistent AdVersarial Attacker, LAVA, addresses this by combining a structured generative process with a symbolic solver, guaranteeing logical consistency. Our framework successfully generates adversarial attacks and identifies global weaknesses common across multiple target models. Our analyses reveal naive heuristics and vulnerabilities in these models’ reasoning capabilities, exposing an incomplete grasp of logical deduction under logic programs. Finally, in addition to effective probing of these models, we show that training on the generated samples improves the target model’s performance. Alexander Gaskell, Yishu Miao, Francesca Toni, Lucia Specia |
IJCAI | 4 |
| 2022 | MLQE-PE: A Multilingual Quality Estimation and Post-Editing DatasetabstractWe present MLQE-PE, a new dataset for Machine Translation (MT) Quality Estimation (QE) and Automatic Post-Editing (APE). The dataset contains annotations for eleven language pairs, including both high- and low-resource languages. Specifically, it is annotated for translation quality with human labels for up to 10,000 translations per language pair in the following formats: sentence-level direct assessments and post-editing effort, and word-level binary good/bad labels. Apart from the quality-related scores, each source-translation sentence pair is accompanied by the corresponding post-edited sentence, as well as titles of the articles where the sentences were extracted from, and information on the neural MT models used to translate the text. We provide a thorough description of the data collection and annotation process as well as an analysis of the annotation distribution for each language pair. We also report the performance of baseline systems trained on the MLQE-PE dataset. The dataset is freely available and has already been used for several WMT shared tasks. Marina Fomicheva, Erick Rocha Fonseca, Chrysoula Zerva, Frédéric Blain, Vishrav Chaudhary, Francisco Guzmán, Nina Lopatina, Lucia Specia, André F. T. Martins |
LREC | 9 |
| 2022 | Leveraging Pre-trained Language Models for Gender DebiasingabstractStudying and mitigating gender and other biases in natural language have become important areas of research from both algorithmic and data perspectives. This paper explores the idea of reducing gender bias in a language generation context by generating gender variants of sentences. Previous work in this field has either been rule-based or required large amounts of gender balanced training data. These approaches are however not scalable across multiple languages, as creating data or rules for each language is costly and time-consuming. This work explores a light-weight method to generate gender variants for a given text using pre-trained language models as the resource, without any task-specific labelled data. The approach is designed to work on multiple languages with minimal changes in the form of heuristics. To showcase that, we have tested it on a high-resourced language, namely Spanish, and a low-resourced language from a different family, namely Serbian. The approach proved to work very well on Spanish, and while the results were less positive for Serbian, it showed potential even for languages where pre-trained models are less effective. Nishtha Jain, Declan Groves, Lucia Specia, Maja Popovic |
LREC | 3 |
| 2022 | Multilingual and Multimodal Learning for Brazilian PortugueseabstractHumans constantly deal with multimodal information, that is, data from different modalities, such as texts and images. In order for machines to process information similarly to humans, they must be able to process multimodal data and understand the joint relationship between these modalities. This paper describes the work performed on the VTLM (Visual Translation Language Modelling) framework from (Caglayan et al., 2021) to test its generalization ability for other language pairs and corpora. We use the multimodal and multilingual corpus How2 (Sanabria et al., 2018) in three parallel streams with aligned English-Portuguese-Visual information to investigate the effectiveness of the model for this new language pair and in more complex scenarios, where the sentence associated with each image is not a simple description of it. Our experiments on the Portuguese-English multimodal translation task using the How2 dataset demonstrate the efficacy of cross-lingual visual pretraining. We achieved a BLEU score of 51.8 and a METEOR score of 78.0 on the test set, outperforming the MMT baseline by about 14 BLEU and 14 METEOR. The good BLEU and METEOR values obtained for this new language pair, regarding the original English-German VTLM, establish the suitability of the model to other languages. Júlia Sato, Helena de Medeiros Caseli, Lucia Specia |
LREC | 3 |
| 2022 | MultiSubs: A Large-scale Multimodal and Multilingual DatasetabstractThis paper introduces a large-scale multimodal and multilingual dataset that aims to facilitate research on grounding words to images in their contextual usage in language. The dataset consists of images selected to unambiguously illustrate concepts expressed in sentences from movie subtitles. The dataset is a valuable resource as (i) the images are aligned to text fragments rather than whole sentences; (ii) multiple images are possible for a text fragment and a sentence; (iii) the sentences are free-form and real-world like; (iv) the parallel texts are multilingual. We also set up a fill-in-the-blank game for humans to evaluate the quality of the automatic image selection process of our dataset. Finally, we propose a fill-in-the-blank task to demonstrate the utility of the dataset, and present some baseline prediction models. The dataset will benefit research on visual grounding of words especially in the context of free-form sentences, and can be obtained from https://doi.org/10.5281/zenodo.5034604 under a Creative Commons licence. Josiah Wang, Josiel Figueiredo, Lucia Specia |
LREC | 3 |
| 2022 | Guiding Visual Question GenerationabstractNihir Vedd, Zixu Wang, Marek Rei, Yishu Miao, Lucia Specia. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Nihir Vedd, Marek Rei, Yishu Miao, Lucia Specia |
NAACL-HLT | 5 |
| 2022 | Supervised Visual Attention for Simultaneous Multimodal Machine TranslationabstractThere has been a surge in research in multimodal machine translation (MMT), where additional modalities such as images are used to improve translation quality of textual systems. A particular use for such multimodal systems is the task of simultaneous machine translation, where visual context has been shown to complement the partial information provided by the source sentence, especially in the early phases of translation. In this paper, we propose the first Transformer-based simultaneous MMT architecture, which has not been previously explored in simultaneous translation. Additionally, we extend this model with an auxiliary supervision signal that guides the visual attention mechanism using labelled phrase-region alignments. We perform comprehensive experiments on three language directions and conduct thorough quantitative and qualitative analyses using both automatic metrics and manual inspection. Our results show that (i) supervised visual attention consistently improves the translation quality of the simultaneous MMT models, and (ii) fine-tuning the MMT with supervision loss enabled leads to better performance than training the MMT from scratch. Compared to the state-of-the-art, our proposed model achieves improvements of up to 2.3 BLEU and 3.5 METEOR points. Veneta Haralampieva, Ozan Caglayan, Lucia Specia |
J. Artif. Intell. Res. | 3 |
| 2021 | BERTGen: Multi-task Generation through BERTabstractFaidon Mitzalis, Ozan Caglayan, Pranava Madhyastha, Lucia Specia. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Faidon Mitzalis, Ozan Caglayan, Pranava Swaroop Madhyastha, Lucia Specia |
ACL/IJCNLP (1) | 4 |
| 2021 | Cross-Modal Generative Augmentation for Visual Question Answering
Yishu Miao, Lucia Specia |
BMVC | 3 |
| 2021 | Cross-lingual Visual Pre-training for Multimodal Machine TranslationabstractOzan Caglayan, Menekse Kuyu, Mustafa Sercan Amac, Pranava Madhyastha, Erkut Erdem, Aykut Erdem, Lucia Specia. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Ozan Caglayan, Menekse Kuyu, Mustafa Sercan Amac, Pranava Swaroop Madhyastha, Erkut Erdem, Aykut Erdem, Lucia Specia |
EACL | 7 |
| 2021 | Exploiting Multimodal Reinforcement Learning for Simultaneous Machine TranslationabstractJulia Ive, Andy Mingren Li, Yishu Miao, Ozan Caglayan, Pranava Madhyastha, Lucia Specia. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Julia Ive, Andy Mingren Li, Yishu Miao, Ozan Caglayan, Pranava Swaroop Madhyastha, Lucia Specia |
EACL | 6 |
| 2021 | Exploring Supervised and Unsupervised Rewards in Machine TranslationabstractReinforcement Learning (RL) is a powerful framework to address the discrepancy between loss functions used during training and the final evaluation metrics to be used at test time. When applied to neural Machine Translation (MT), it minimises the mismatch between the cross-entropy loss and non-differentiable evaluation metrics like BLEU. However, the suitability of these metrics as reward function at training time is questionable: they tend to be sparse and biased towards the specific words used in the reference texts. We propose to address this problem by making models less reliant on such metrics in two ways: (a) with an entropy-regularised RL method that does not only maximise a reward function but also explore the action space to avoid peaky distributions; (b) with a novel RL method that explores a dynamic unsupervised reward function to balance between exploration and exploitation. We base our proposals on the Soft Actor-Critic (SAC) framework, adapting the off-policy maximum entropy model for language generation applications such as MT. We demonstrate that SAC with BLEU reward tends to overfit less to the training data and performs better on out-of-domain data. We also show that our dynamic unsupervised reward can lead to better translation of ambiguous words. Julia Ive, Marina Fomicheva, Lucia Specia |
EACL | 4 |
| 2021 | Quality Estimation without Human-labeled DataabstractYi-Lin Tuan, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary, Francisco Guzmán, Lucia Specia. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Yi-Lin Tuan, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary, Francisco Guzmán, Lucia Specia |
EACL | 6 |
| 2021 | A Generative Framework for Simultaneous Machine TranslationabstractWe propose a generative framework for simultaneous machine translation.Conventional approaches use a fixed number of source words to translate or learn dynamic policies for the number of source words by reinforcement learning.Here we formulate simultaneous translation as a structural sequence-tosequence learning problem.A latent variable is introduced to model read or translate actions at every time step, which is then integrated out to consider all the possible translation policies.A re-parameterised Poisson prior is used to regularise the policies which allows the model to explicitly balance translation quality and latency.The experiments demonstrate the effectiveness and robustness of the generative framework, which achieves the best BLEU scores given different average translation latencies on benchmark datasets. Yishu Miao, Phil Blunsom, Lucia Specia |
EMNLP (1) | 3 |
| 2021 | Classification-based Quality Estimation: Small and Efficient Models for Real-world ApplicationsabstractSentence-level Quality Estimation (QE) of machine translation is traditionally formulated as a regression task, and the performance of QE models is typically measured by Pearson correlation with human labels.Recent QE models have achieved previously-unseen levels of correlation with human judgments, but they rely on large multilingual contextualized language models that are computationally expensive and thus infeasible for many real-world applications.In this work, we evaluate several model compression techniques for QE and find that, despite their popularity in other NLP tasks, they lead to poor performance in this regression setting.We observe that a full model parameterization is required to achieve SoTA results in a regression task.However, we argue that the level of expressiveness of a model in a continuous range is unnecessary given the downstream applications of QE, and show that reframing QE as a classification problem and evaluating QE models using classification metrics would better reflect their actual performance in real-world applications. Ahmed El-Kishky, Vishrav Chaudhary, James Cross 0003, Lucia Specia, Francisco Guzmán |
EMNLP (1) | 5 |
| 2021 | SentSim: Crosslingual Semantic Evaluation of Machine TranslationabstractMachine translation (MT) is currently evaluated in one of two ways: in a monolingual fashion, by comparison with the system output to one or more human reference translations, or in a trained crosslingual fashion, by building a supervised model to predict quality scores from human-labeled data.In this paper, we propose a more cost-effective, yet well performing unsupervised alternative SentSim: relying on strong pretrained multilingual word and sentence representations, we directly compare the source with the machine translated sentence, thus avoiding the need for both reference translations and labelled training data.The metric builds on state-of-the-art embedding-based approachesnamely BERTScore and Word Mover's Distance -by incorporating a notion of sentence semantic similarity.By doing so, it achieves better correlation with human scores on different datasets.We show that it outperforms these and other metrics in the standard monolingual setting (MT-reference translation), a well as in the source-MT bilingual setting, where it performs on par with glass-box approaches to quality estimation that rely on MT model information. Yurun Song, Junchen Zhao, Lucia Specia |
NAACL-HLT | 3 |
| 2021 | Backtranslation Feedback Improves User Confidence in MT, Not QualityabstractVilém Zouhar, Michal Novák, Matúš Žilinec, Ondřej Bojar, Mateo Obregón, Robin L. Hill, Frédéric Blain, Marina Fomicheva, Lucia Specia, Lisa Yankovskaya. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Vilém Zouhar, Michal Novák 0001, Matús Zilinec, Ondrej Bojar, Mateo Obregón, Robin L. Hill, Frédéric Blain, Marina Fomicheva, Lucia Specia, Lisa Yankovskaya |
NAACL-HLT | 9 |
| 2021 | The (Un)Suitability of Automatic Evaluation Metrics for Text SimplificationabstractAbstract In order to simplify sentences, several rewriting operations can be performed, such as replacing complex words per simpler synonyms, deleting unnecessary information, and splitting long sentences. Despite this multi-operation nature, evaluation of automatic simplification systems relies on metrics that moderately correlate with human judgments on the simplicity achieved by executing specific operations (e.g., simplicity gain based on lexical replacements). In this article, we investigate how well existing metrics can assess sentence-level simplifications where multiple operations may have been applied and which, therefore, require more general simplicity judgments. For that, we first collect a new and more reliable data set for evaluating the correlation of metrics and human judgments of overall simplicity. Second, we conduct the first meta-evaluation of automatic metrics in Text Simplification, using our new data set (and other existing data) to analyze the variation of the correlation between metrics’ scores and human judgments across three dimensions: the perceived simplicity level, the system type, and the set of references used for computation. We show that these three aspects affect the correlations and, in particular, highlight the limitations of commonly used operation-specific metrics. Finally, based on our findings, we propose a set of recommendations for automatic evaluation of multi-operation simplifications, suggesting which metrics to compute and how to interpret their scores. Fernando Alva-Manchego, Carolina Scarton, Lucia Specia |
Comput. Linguistics | 3 |
| 2021 | MSVD-Turkish: a comprehensive multimodal video dataset for integrated vision and language research in Turkish
Begüm Çitamak Erdinç, Ozan Caglayan, Menekse Kuyu, Erkut Erdem, Aykut Erdem, Pranava Swaroop Madhyastha, Lucia Specia |
Mach. Transl. | 7 |
| 2021 | Read, spot and translateabstractAbstract We propose multimodal machine translation (MMT) approaches that exploit the correspondences between words and image regions. In contrast to existing work, our referential grounding method considers objects as the visual unit for grounding, rather than whole images or abstract image regions, and performs visual grounding in the source language, rather than at the decoding stage via attention. We explore two referential grounding approaches: (i) implicit grounding, where the model jointly learns how to ground the source language in the visual representation and to translate; and (ii) explicit grounding, where grounding is performed independent of the translation model, and is subsequently used to guide machine translation. We performed experiments on the Multi30K dataset for three language pairs: English–German, English–French and English–Czech. Our referential grounding models outperform existing MMT models according to automatic and human evaluation metrics. Lucia Specia, Josiah Wang, Sun Jae Lee, Alissa Ostapenko, Pranava Swaroop Madhyastha |
Mach. Transl. | 1 |
| 2021 | Leveraging auxiliary image descriptions for dense video captioning
Emre Boran, Aykut Erdem, Nazli Ikizler-Cinbis, Erkut Erdem, Pranava Swaroop Madhyastha, Lucia Specia |
Pattern Recognit. Lett. | 6 |
| 2020 | ASSET: A Dataset for Tuning and Evaluation of Sentence Simplification Models with Multiple Rewriting TransformationsabstractIn order to simplify a sentence, human editors perform multiple rewriting transformations: they split it into several shorter sentences, paraphrase words (i.e.replacing complex words or phrases by simpler synonyms), reorder components, and/or delete information deemed unnecessary.Despite these varied range of possible text alterations, current models for automatic sentence simplification are evaluated using datasets that are focused on a single transformation, such as lexical paraphrasing or splitting.This makes it impossible to understand the ability of simplification models in more realistic settings.To alleviate this limitation, this paper introduces ASSET, a new dataset for assessing sentence simplification in English.ASSET is a crowdsourced multi-reference corpus where each simplification was produced by executing several rewriting transformations.Through quantitative and qualitative experiments, we show that simplifications in ASSET are better at capturing characteristics of simplicity when compared to other standard evaluation datasets for the task.Furthermore, we motivate the need for developing better methods for automatic evaluation using ASSET, since we show that current popular metrics may not be suitable when multiple simplification transformations are performed. Fernando Alva-Manchego, Louis Martin, Antoine Bordes, Carolina Scarton, Benoît Sagot, Lucia Specia |
ACL | 6 |
| 2020 | Multi-Hypothesis Machine Translation EvaluationabstractReliably evaluating Machine Translation (MT) through automated metrics is a long-standing problem.One of the main challenges is the fact that multiple outputs can be equally valid.Attempts to minimise this issue include metrics that relax the matching of MT output and reference strings, and the use of multiple references.The latter has been shown to significantly improve the performance of evaluation metrics.However, collecting multiple references is expensive and in practice a single reference is generally used.In this paper, we propose an alternative approach: instead of modelling linguistic variation in human reference we exploit the MT model uncertainty to generate multiple diverse translations and use these: (i) as surrogates to reference translations; (ii) to obtain a quantification of translation variability to either complement existing metric scores or (iii) replace references altogether.We show that for a number of popular evaluation metrics our variability estimates lead to substantial improvements in correlation with human judgements of quality by up 15%. Marina Fomicheva, Lucia Specia, Francisco Guzmán |
ACL | 2 |
| 2020 | Multimodal Quality Estimation for Machine Translationabstract© 2020 The Authors. Published by Association for Computational Linguistics. This is an open access article available under a Creative Commons licence. \nThe published version can be accessed at the following link on the publisher’s website: http://dx.doi.org/10.18653/v1/2020.acl-main.114 Shu Okabe, Frédéric Blain, Lucia Specia |
ACL | 3 |
| 2020 | Are we Estimating or Guesstimating Translation Quality?abstractRecent advances in pre-trained multilingual language models lead to state-of-the-art results on the task of quality estimation (QE) for machine translation.A carefully engineered ensemble of such models won the QE shared task at WMT19.Our in-depth analysis, however, shows that the success of using pre-trained language models for QE is overestimated due to three issues we observed in current QE datasets: (i) The distributions of quality scores are imbalanced and skewed towards good quality scores; (ii) QE models can perform well on these datasets while looking at only source or translated sentences; (iii) They contain statistical artifacts that correlate well with human-annotated QE labels.Our findings suggest that although QE models might capture fluency of translated sentences and complexity of source sentences, they cannot model adequacy of translations effectively. Francisco Guzmán, Lucia Specia |
ACL | 3 |
| 2020 | Curious Case of Language Generation Evaluation Metrics: A Cautionary TaleabstractAutomatic evaluation of language generation systems is a well-studied problem in Natural Language Processing.While novel metrics are proposed every year, a few popular metrics remain as the de facto metrics to evaluate tasks such as image captioning and machine translation, despite their known limitations.This is partly due to ease of use, and partly because researchers expect to see them and know how to interpret them.In this paper, we urge the community for more careful consideration of how they automatically evaluate their models by demonstrating important failure cases on multiple datasets, language pairs and tasks.Our experiments show that metrics (i) usually prefer system outputs to human-authored texts, (ii) can be insensitive to correct translations of rare words, (iii) can yield surprisingly high scores when given a single sentence as system output for the entire test set. Ozan Caglayan, Pranava Swaroop Madhyastha, Lucia Specia |
COLING | 3 |
| 2020 | Quality In, Quality Out: Learning from Actual MistakesabstractApproaches to Quality Estimation (QE) of machine translation have shown promising results at predicting quality scores for translated sentences. However, QE models are often trained on noisy approximations of quality annotations derived from the proportion of post-edited words in translated sentences instead of direct human annotations of translation errors. The latter is a more reliable ground-truth but more expensive to obtain. In this paper, we present the first attempt to model the task of predicting the proportion of actual translation errors in a sentence while minimising the need for direct human annotation. For that purpose, we use transfer-learning to leverage large scale noisy annotations and small sets of high-fidelity human annotated translation errors to train QE models. Experiments on four language pairs and translations obtained by statistical and neural models show consistent gains over strong baselines. Frédéric Blain, Nikolaos Aletras, Lucia Specia |
EAMT | 3 |
| 2020 | Deciding When, How and for Whom to SimplifyabstractCurrent Automatic Text Simplification (TS) work relies on sequence-to-sequence neural models that learn simplification operations from parallel complex-simple corpora. In this paper we address three open challenges in these approaches: (i) avoiding unnecessary transformations, (ii) determining which operations to perform, and (iii) generating simplifications that are suitable for a given target audience. For (i), we propose joint and two-stage approaches where instances are marked or classified as simple or complex. For (ii) and (iii), we propose fusion-based approaches to incorporate information on the target grade level as well as the types of operation to perform in the models. While grade-level information is provided as metadata, we devise predictors for the type of operation. We study different representations for this information as well as different ways in which it is used in the models. Our approach outperforms previous work on neural TS, with our best model following the two-stage approach and using the information about grade level and type of operation to initialise the encoder and the decoder, respectively. Carolina Scarton, Pranava Swaroop Madhyastha, Lucia Specia |
ECAI | 3 |
| 2020 | Simultaneous Machine Translation with Visual ContextabstractSimultaneous machine translation (SiMT) aims to translate a continuous input text stream into another language with the lowest latency and highest quality possible.The translation thus has to start with an incomplete source text, which is read progressively, creating the need for anticipation.In this paper, we seek to understand whether the addition of visual information can compensate for the missing source context.To this end, we analyse the impact of different multimodal approaches and visual features on state-of-the-art SiMT frameworks.Our results show that visual context is helpful and that visually-grounded models based on explicit object region information are much better than commonly used global features, reaching up to 3 BLEU points improvement under low latency scenarios.Our qualitative analysis illustrates cases where only the multimodal systems are able to translate correctly from English into gender-marked languages, as well as deal with differences in word order, such as adjective-noun placement between English and French. Ozan Caglayan, Julia Ive, Veneta Haralampieva, Pranava Swaroop Madhyastha, Loïc Barrault, Lucia Specia |
EMNLP (1) | 6 |
| 2020 | FIND: Human-in-the-Loop Debugging Deep Text ClassifiersabstractSince obtaining a perfect training dataset (i.e., a dataset which is considerably large, unbiased, and well-representative of unseen cases) is hardly possible, many real-world text classifiers are trained on the available, yet imperfect, datasets.These classifiers are thus likely to have undesirable properties.For instance, they may have biases against some sub-populations or may not work effectively in the wild due to overfitting.In this paper, we propose FINDa framework which enables humans to debug deep learning text classifiers by disabling irrelevant hidden features.Experiments show that by using FIND, humans can improve CNN text classifiers which were trained under different types of imperfect datasets (including datasets with biases and datasets with dissimilar traintest distributions). Piyawat Lertvittayakumjorn, Lucia Specia, Francesca Toni |
EMNLP (1) | 2 |
| 2020 | A Post-Editing Dataset in the Legal Domain: Do we Underestimate Neural Machine Translation Quality?abstractWe introduce a machine translation dataset for three pairs of languages in the legal domain with post-edited high-quality neural machine translation and independent human references. The data was collected as part of the EU APE-QUEST project and comprises crawled content from EU websites with translation from English into three European languages: Dutch, French and Portuguese. Altogether, the data consists of around 31K tuples including a source sentence, the respective machine translation by a neural machine translation system, a post-edited version of such translation by a professional translator, and - where available - the original reference translation crawled from parallel language websites. We describe the data collection process, provide an analysis of the resulting post-edits and benchmark the data using state-of-the-art quality estimation and automatic post-editing models. One interesting by-product of our post-editing analysis suggests that neural systems built with publicly available general domain data can provide high-quality translations, even though comparison to human references suggests that this quality is quite low. This makes our dataset a suitable candidate to test evaluation metrics. The data is freely available as an ELRC-SHARE resource. Julia Ive, Lucia Specia, Sara Szoc, Tom Vanallemeersch, Joachim Van den Bogaert, Eduardo Farah, Christine Maroti, Artur Ventura, Maxim Khalilov |
LREC | 2 |
| 2020 | Data-Driven Sentence Simplification: Survey and BenchmarkabstractSentence Simplification (SS) aims to modify a sentence in order to make it easier to read and understand. In order to do so, several rewriting transformations can be performed such as replacement, reordering, and splitting. Executing these transformations while keeping sentences grammatical, preserving their main idea, and generating simpler output, is a challenging and still far from solved problem. In this article, we survey research on SS, focusing on approaches that attempt to learn how to simplify using corpora of aligned original-simplified sentence pairs in English, which is the dominant paradigm nowadays. We also include a benchmark of different approaches on common data sets so as to compare them and highlight their strengths and limitations. We expect that this survey will serve as a starting point for researchers interested in the task and help spark new ideas for future developments. Fernando Alva-Manchego, Carolina Scarton, Lucia Specia |
Comput. Linguistics | 3 |
| 2020 | Multimodal machine translation through visuals and speechabstractAbstract Multimodal machine translation involves drawing information from more than one modality, based on the assumption that the additional modalities will contain useful alternative views of the input data. The most prominent tasks in this area are spoken language translation, image-guided translation, and video-guided translation, which exploit audio and visual modalities, respectively. These tasks are distinguished from their monolingual counterparts of speech recognition, image captioning, and video captioning by the requirement of models to generate outputs in a different language. This survey reviews the major data resources for these tasks, the evaluation campaigns concentrated around them, the state of the art in end-to-end and pipeline approaches, and also the challenges in performance evaluation. The paper concludes with a discussion of directions for future research in these areas: the need for more expansive and challenging datasets, for targeted evaluations of model performance, and for multimodality in both the input and output space. Umut Sulubacak, Ozan Caglayan, Stig-Arne Grönroos, Aku Rouhe, Desmond Elliott, Lucia Specia, Jörg Tiedemann |
Mach. Transl. | 6 |
| 2020 | Unsupervised Quality Estimation for Neural Machine TranslationabstractQuality Estimation (QE) is an important component in making Machine Translation (MT) useful in real-world applications, as it is aimed to inform the user on the quality of the MT output at test time. Existing approaches require large amounts of expert annotated data, computation, and time for training. As an alternative, we devise an unsupervised approach to QE where no training or access to additional resources besides the MT system itself is required. Different from most of the current work that treats the MT system as a black box, we explore useful information that can be extracted from the MT system as a by-product of translation. By utilizing methods for uncertainty quantification, we achieve very good correlation with human judgments of quality, rivaling state-of-the-art supervised QE models. To evaluate our approach we collect the first dataset that enables work on both black-box and glass-box approaches to QE. Marina Fomicheva, Lisa Yankovskaya, Frédéric Blain, Francisco Guzmán, Mark Fishel, Nikolaos Aletras, Vishrav Chaudhary, Lucia Specia |
Trans. Assoc. Comput. Linguistics | 9 |
| 2019 | Distilling Translations with Visual AwarenessabstractPrevious work on multimodal machine translation has shown that visual information is only needed in very specific cases, for example in the presence of ambiguous words where the textual context is not sufficient.As a consequence, models tend to learn to ignore this information.We propose a translate-and-refine approach to this problem where images are only used by a second stage decoder.This approach is trained jointly to generate a good first draft translation and to improve over this draft by (i) making better use of the target language textual context (both left and right-side contexts) and (ii) making use of visual context.This approach leads to the state of the art results.Additionally, we show that it has the ability to recover from erroneous or missing words in the source language.EN: Three children in football uniforms are playing football.DE: Drei Kinder in Fußballtrikots spielen Fußball.PE: Drei Kinder in Footballtrikots spielen Football.(a) Ambiguous word football translated as soccer (Fußball) EN: A baseball player in a black shirt just tagged a player in a white shirt.DE: Ein Baseballspieler in einem schwarzen Shirt fängt einen Spieler in einem weißen Shirt.PE: Eine Baseballspielerin in einem schwarzen Shirt fängt eine Spielerin in einem weißen Shirt.(b) Gender-neutral word player translated as male player (Spieler) EN: A woman wearing a white shirt works out on an elliptical machine. Julia Ive, Pranava Swaroop Madhyastha, Lucia Specia |
ACL (1) | 3 |
| 2019 | VIFIDEL: Evaluating the Visual Fidelity of Image DescriptionsabstractWe address the task of evaluating image description generation systems.We propose a novel image-aware metric for this task: VIFIDEL.It estimates the faithfulness of a generated caption with respect to the content of the actual image, based on the semantic similarity between labels of objects depicted in images and words in the description.The metric is also able to take into account the relative importance of objects mentioned in human reference descriptions during evaluation.Even if these human reference descriptions are not available, VIFIDEL can still reliably evaluate system descriptions.The metric achieves high correlation with human judgments on two well-known datasets and is competitive with metrics that depend on and rely exclusively on human references. Pranava Swaroop Madhyastha, Josiah Wang, Lucia Specia |
ACL (1) | 3 |
| 2019 | Phrase-Level Simplification for Non-native Speakers
Gustavo Paetzold, Lucia Specia |
CICLing (1) | 2 |
| 2019 | Deep Copycat Networks for Text-to-Text GenerationabstractJulia Ive, Pranava Madhyastha, Lucia Specia. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Julia Ive, Pranava Swaroop Madhyastha, Lucia Specia |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Phrase Localization Without Paired Training ExamplesabstractLocalizing phrases in images is an important part of image understanding and can be useful in many applications that require mappings between textual and visual information. Existing work attempts to learn these mappings from examples of phrase-image region correspondences (strong supervision) or from phrase-image pairs (weak supervision). We postulate that such paired annotations are unnecessary, and propose the first method for the phrase localization problem where neither training procedure nor paired, task-specific data is required. Our method is simple but effective: we use off-the-shelf approaches to detect objects, scenes and colours in images, and explore different approaches to measure semantic similarity between the categories of detected visual elements and words in phrases. Experiments on two well-known phrase localization datasets show that this approach surpasses all weakly supervised methods by a large margin and performs very competitively to strongly supervised methods, and can thus be considered a strong baseline to the task. The non-paired nature of our method makes it applicable to any domain and where no paired phrase localization annotation is available. Josiah Wang, Lucia Specia |
ICCV | 2 |
| 2019 | APE-QUEST
Joachim Van den Bogaert, Heidi Depraetere, Sara Szoc, Tom Vanallemeersch, Koen Van Winckel, Frederic Everaert, Lucia Specia, Julia Ive, Maxim Khalilov, Christine Maroti, Eduardo Farah, Artur Ventura |
MTSummit (2) | 7 |
| 2019 | Taking MT Evaluation Metrics to Extremes: Beyond Correlation with Human JudgmentsabstractAutomatic Machine Translation (MT) evaluation is an active field of research, with a handful of new metrics devised every year. Evaluation metrics are generally benchmarked against manual assessment of translation quality, with performance measured in terms of overall correlation with human scores. Much work has been dedicated to the improvement of evaluation metrics to achieve a higher correlation with human judgments. However, little insight has been provided regarding the weaknesses and strengths of existing approaches and their behavior in different settings. In this work we conduct a broad meta-evaluation study of the performance of a wide range of evaluation metrics focusing on three major aspects. First, we analyze the performance of the metrics when faced with different levels of translation quality, proposing a local dependency measure as an alternative to the standard, global correlation coefficient. We show that metric performance varies significantly across different levels of MT quality: Metrics perform poorly when faced with low-quality translations and are not able to capture nuanced quality distinctions. Interestingly, we show that evaluating low-quality translations is also more challenging for humans. Second, we show that metrics are more reliable when evaluating neural MT than the traditional statistical MT systems. Finally, we show that the difference in the evaluation accuracy for different metrics is maintained even if the gold standard scores are based on different criteria. Marina Fomicheva, Lucia Specia |
Comput. Linguistics | 2 |
| 2018 | End-to-end Image Captioning Exploits Distributional Similarity in Multimodal Space
Pranava Swaroop Madhyastha, Josiah Wang, Lucia Specia |
BMVC | 3 |
| 2018 | deepQuest: A Framework for Neural-based Quality EstimationabstractPredicting Machine Translation (MT) quality can help in many practical tasks such as MT post-editing. The performance of Quality Estimation (QE) methods has drastically improved recently with the introduction of neural approaches to the problem. However, thus far neural approaches have only been designed for word and sentence-level prediction. We present a neural framework that is able to accommodate neural QE approaches at these fine-grained levels and generalize them to the level of documents. We test the framework with two sentence-level neural QE approaches: a state of the art approach that requires extensive pre-training, and a new light-weight approach that we propose, which employs basic encoders. Our approach is significantly faster and yields performance improvements for a range of document-level quality estimation tasks. To our knowledge, this is the first neural architecture for document-level QE. In addition, for the first time we apply QE models to the output of both statistical and neural MT systems for a series of European languages and highlight the new challenges resulting from the use of neural MT. Julia Ive, Frédéric Blain, Lucia Specia |
COLING | 3 |
| 2018 | Multi-modal Context Modelling for Machine TranslationabstractMultiMT is an European Research Council Starting Grant whose aim is to devise data, methods and algorithms to exploit multi-modal information (images, audio, metadata) for context modelling in machine translation and other cross- lingual tasks. The project draws upon different research fields including natural language processing, computer vision, speech processing and machine learning. Lucia Specia |
EAMT | 1 |
| 2018 | Multimodal Lexical Translation
Chiraag Lala, Lucia Specia |
LREC | 2 |
| 2018 | Text Simplification from Professionally Produced Corpora
Carolina Scarton, Gustavo Paetzold, Lucia Specia |
LREC | 3 |
| 2018 | SimPA: A Sentence-Level Simplification Corpus for the Public Administration Domain
Carolina Scarton, Gustavo Paetzold, Lucia Specia |
LREC | 3 |
| 2018 | Object Counts! Bringing Explicit Detections Back into Image CaptioningabstractJosiah Wang, Pranava Swaroop Madhyastha, Lucia Specia. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Josiah Wang, Pranava Swaroop Madhyastha, Lucia Specia |
NAACL-HLT | 3 |
| 2018 | Assessing multilingual multimodal image description: Studies of native speaker preferences and translator choicesabstractAbstract Two studies on multilingual multimodal image description provide empirical evidence towards two questions at the core of the task: (i) whether target language speakers prefer descriptions generated directly in their native language, as compared to descriptions translated from a different language; (ii) whether images improve human translation of descriptions. These results provide guidance for future work in multimodal natural language processing by first showing that on the whole, translations are not distinguished from native language descriptions, and second delineating and quantifying the information gained from the image during the human translation task. Stella Frank, Desmond Elliott, Lucia Specia |
Nat. Lang. Eng. | 3 |
| 2018 | The role of image representations in vision to language tasksabstractAbstract Tasks that require modeling of both language and visual information, such as image captioning, have become very popular in recent years. Most state-of-the-art approaches make use of image representations obtained from a deep neural network, which are used to generate language information in a variety of ways with end-to-end neural-network-based models. However, it is not clear how different image representations contribute to language generation tasks. In this paper, we probe the representational contribution of the image features in an end-to-end neural modeling framework and study the properties of different types of image representations. We focus on two popular vision to language problems: The task of image captioning and the task of multimodal machine translation. Our analysis provides interesting insights into the representational properties and suggests that end-to-end approaches implicitly learn a visual-semantic subspace and exploit the subspace to generate captions. Pranava Swaroop Madhyastha, Josiah Wang, Lucia Specia |
Nat. Lang. Eng. | 3 |
| 2017 | Exploring the use of acoustic embeddings in neural machine translationabstractNeural Machine Translation (NMT) has recently demonstrated improved performance over statistical machine translation and relies on an encoder-decoder framework for translating text from source to target. The structure of NMT makes it amenable to add auxiliary features, which can provide complementary information to that present in the source text. In this paper, auxiliary features derived from accompanying audio, are investigated for NMT and are compared and combined with text-derived features. These acoustic embeddings can help resolve ambiguity in the translation, thus improving the output. The following features are experimented with: Latent Dirichlet Allocation (LDA) topic vectors and GMM subspace i-vectors derived from audio. These are contrasted against: skip-gram/Word2Vec features and LDA features derived from text. The results are encouraging and show that acoustic information does help with NMT, leading to an overall 3.3% relative improvement in BLEU scores. Salil Deena, Raymond W. M. Ng, Pranava Swaroop Madhyastha, Lucia Specia, Thomas Hain |
ASRU | 4 |
| 2017 | Personalized Machine Translation: Preserving Original Author TraitsabstractElla Rabinovich, Raj Nath Patel, Shachar Mirkin, Lucia Specia, Shuly Wintner. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Ella Rabinovich, Raj Nath Patel, Shachar Mirkin, Lucia Specia, Shuly Wintner |
EACL (1) | 4 |
| 2017 | Learning How to Simplify From Explicit Labeling of Complex-Simplified Text PairsabstractCurrent research in text simplification has been hampered by two central problems: (i) the small amount of high-quality parallel simplification data available, and (ii) the lack of explicit annotations of simplification operations, such as deletions or substitutions, on existing data. While the recently introduced Newsela corpus has alleviated the first problem, simplifications still need to be learned directly from parallel text using black-box, end-to-end approaches rather than from explicit annotations. These complex-simple parallel sentence pairs often differ to such a high degree that generalization becomes difficult. End-to-end models also make it hard to interpret what is actually learned from data. We propose a method that decomposes the task of TS into its sub-problems. We devise a way to automatically identify operations in a parallel corpus and introduce a sequence-labeling approach based on these annotations. Finally, we provide insights on the types of transformations that different approaches can model. Fernando Alva-Manchego, Joachim Bingel, Gustavo Paetzold, Carolina Scarton, Lucia Specia |
IJCNLP(1) | 5 |
| 2017 | Semi-Supervised Adaptation of RNNLMs by Fine-Tuning with Domain-Specific Auxiliary FeaturesabstractRecurrent neural network language models (RNNLMs) can be augmented with auxiliary features, which can provide an extra modality on top of the words. It has been found that RNNLMs perform best when trained on a large corpus of generic text and then fine-tuned on text corresponding to the sub-domain for which it is to be applied. However, in many cases the auxiliary features are available for the sub-domain text but not for the generic text. In such cases, semi-supervised techniques can be used to infer such features for the generic text data such that the RNNLM can be trained and then fine-tuned on the available in-domain data with corresponding auxiliary features. \n \nIn this paper, several novel approaches are investigated for dealing with the semi-supervised adaptation of RNNLMs with auxiliary features as input. These approaches include: using zero features during training to mask the weights of the feature sub-network; adding the feature sub-network only at the time of fine-tuning; deriving the features using a parametric model and; back-propagating to infer the features on the generic text. These approaches are investigated and results are reported both in terms of PPL and WER on a multi-genre broadcast ASR task. Salil Deena, Raymond W. M. Ng, Pranava Swaroop Madhyastha, Lucia Specia, Thomas Hain |
INTERSPEECH | 4 |
| 2017 | Exploring Hypotheses Spaces in Neural Machine Translation
Frédéric Blain, Lucia Specia, Pranava Swaroop Madhyastha |
MTSummit (1) | 2 |
| 2017 | One-parameter models for sentence-level post-editing effort estimation
Mikel L. Forcada, Miquel Esplà-Gomis, Felipe Sánchez-Martínez, Lucia Specia |
MTSummit (1) | 4 |
| 2017 | Feature-rich NMT and SMT post-edited corpora for productivity and evaluation tasks with a subset of MQM-annotated data
Kim Harris, Lucia Specia, Aljoscha Burchardt |
MTSummit (2) | 2 |
| 2017 | Translation Quality and Productivity: A Study on Rich Morphology Languages
Lucia Specia, Kim Harris, Frédéric Blain, Aljoscha Burchardt, Vivien Macketanz, Inguna Skadina, Matteo Negri, Marco Turchi |
MTSummit (1) | 1 |
| 2017 | A Survey on Lexical Simpli cationabstractLexical Simplification is the process of replacing complex words in a given sentence with simpler alternatives of equivalent meaning. This task has wide applicability both as an assistive technology for readers with cognitive impairments or disabilities, such as Dyslexia and Aphasia, and as a pre-processing tool for other Natural Language Processing tasks, such as machine translation and summarisation. The problem is commonly framed as a pipeline of four steps: the identification of complex words, the generation of substitution candidates, the selection of those candidates that fit the context, and the ranking of the selected substitutes according to their simplicity. In this survey we review the literature for each step in this typical Lexical Simplification pipeline and provide a benchmarking of existing approaches for these steps on publicly available datasets. We also provide pointers for datasets and resources available for the task. Gustavo Paetzold, Lucia Specia |
J. Artif. Intell. Res. | 2 |
| 2016 | Unsupervised Lexical Simplification for Non-Native SpeakersabstractLexical Simplification is the task of replacing complex words with simpler alternatives. We propose a novel, unsupervised approach for the task. It relies on two resources: a corpus of subtitles and a new type of word embeddings model that accounts for the ambiguity of words. We compare the performance of our approach and many others over a new evaluation dataset, which accounts for the simplification needs of 400 non-native English speakers. The experiments show that our approach outperforms state-of-the-art work in Lexical Simplification. Gustavo Paetzold, Lucia Specia |
AAAI | 2 |
| 2016 | Understanding the Lexical Simplification Needs of Non-Native Speakers of EnglishabstractWe report three user studies in which the Lexical Simplification needs of non-native English speakers are investigated. Our analyses feature valuable new insight on the relationship between the non-natives’ notion of complexity and various morphological, semantic and lexical word properties. Some of our findings contradict long-standing misconceptions about word simplicity. The data produced in our studies consists of 211,564 annotations made by 1,100 volunteers, which we hope will guide forthcoming research on Text Simplification for non-native speakers of English. Gustavo Paetzold, Lucia Specia |
COLING | 2 |
| 2016 | Collecting and Exploring Everyday Language for Predicting Psycholinguistic Properties of WordsabstractExploring language usage through frequency analysis in large corpora is a defining feature in most recent work in corpus and computational linguistics. From a psycholinguistic perspective, however, the corpora used in these contributions are often not representative of language usage: they are either domain-specific, limited in size, or extracted from unreliable sources. In an effort to address this limitation, we introduce SubIMDB, a corpus of everyday language spoken text we created which contains over 225 million words. The corpus was extracted from 38,102 subtitles of family, comedy and children movies and series, and is the first sizeable structured corpus of subtitles made available. Our experiments show that word frequency norms extracted from this corpus are more effective than those from well-known norms such as Kucera-Francis, HAL and SUBTLEXus in predicting various psycholinguistic properties of words, such as lexical decision times, familiarity, age of acquisition and simplicity. We also provide evidence that contradict the long-standing assumption that the ideal size for a corpus can be determined solely based on how well its word frequencies correlate with lexical decision times. Gustavo Paetzold, Lucia Specia |
COLING | 2 |
| 2016 | Exploring Prediction Uncertainty in Machine Translation Quality EstimationabstractMachine Translation Quality Estimation is a notoriously difficult task, which lessens its usefulness in real-world translation environments.Such scenarios can be improved if quality predictions are accompanied by a measure of uncertainty.However, models in this task are traditionally evaluated only in terms of point estimate metrics, which do not take prediction uncertainty into account.We investigate probabilistic methods for Quality Estimation that can provide well-calibrated uncertainty estimates and evaluate them in terms of their full posterior predictive distributions.We also show how this posterior information can be useful in an asymmetric risk scenario, which aims to capture typical situations in translation workflows. Daniel Beck, Lucia Specia, Trevor Cohn |
CoNLL | 2 |
| 2016 | Semantic Textual Similarity in Quality Estimation
Hannah Béchara, Carla Parra Escartín, Constantin Orasan, Lucia Specia |
EAMT | 4 |
| 2016 | The Trouble with Machine Translation Coherence
Karin Sim Smith, Wilker Aziz, Lucia Specia |
EAMT | 3 |
| 2016 | Predicting and Using Implicit Discourse Elements in Chinese-English Translation
David Steele, Lucia Specia |
EAMT | 2 |
| 2016 | Groupwise learning for ASR k-best list reranking in spoken language translationabstractQuality estimation models are used to predict the quality of the output from a spoken language translation (SLT) system. When these scores are used to rerank a k-best list, the rank of the scores is more important than their absolute values. This paper proposes groupwise learning to model this rank. Groupwise features were constructed by grouping pairs, triplets or M-plets among the ASR k-best outputs of the same sentence. Regression and classification models were learnt and a score combination strategy was used to predict the rank among the k-best list. Regression models with pairwise features give a bigger gain over other model and feature constructions. Groupwise learning is robust to sentences with different ASR-confidence. This technique is also complementary to linear discriminant analysis feature projection. An overall BLEU score improvement of 0.80 was achieved on an in-domain English-to-French SLT task. Raymond W. M. Ng, Kashif Shah, Lucia Specia, Thomas Hain |
ICASSP | 3 |
| 2016 | Phrase Level Segmentation and Labelling of Machine Translation Errors
Frédéric Blain, Varvara Logacheva, Lucia Specia |
LREC | 3 |
| 2016 | MARMOT: A Toolkit for Translation Quality Estimation at the Word Level
Varvara Logacheva, Chris Hokamp, Lucia Specia |
LREC | 3 |
| 2016 | Benchmarking Lexical Simplification Systems
Gustavo Paetzold, Lucia Specia |
LREC | 2 |
| 2016 | A Reading Comprehension Corpus for Machine Translation Evaluation
Carolina Scarton, Lucia Specia |
LREC | 2 |
| 2016 | Cohere: A Toolkit for Local Coherence
Karin Sim Smith, Wilker Aziz, Lucia Specia |
LREC | 3 |
| 2016 | Inferring Psycholinguistic Properties of WordsabstractWe introduce a bootstrapping algorithm for regression that exploits word embedding models.We use it to infer four psycholinguistic properties of words: Familiarity, Age of Acquisition, Concreteness and Imagery and further populate the MRC Psycholinguistic Database with these properties.The approach achieves 0.88 correlation with humanproduced values and the inferred psycholinguistic features lead to state-of-the-art results when used in a Lexical Simplification task. Gustavo Paetzold, Lucia Specia |
HLT-NAACL | 2 |
| 2016 | Large-scale Multitask Learning for Machine Translation Quality Estimation
Kashif Shah, Lucia Specia |
HLT-NAACL | 2 |
| 2015 | The role of artificially generated negative data for quality estimation of machine translation
Varvara Logacheva, Lucia Specia |
EAMT | 2 |
| 2015 | Okapi+QuEst: Translation Quality Estimation within Okapi
Gustavo Paetzold, Lucia Specia, Yves Savourel |
EAMT | 2 |
| 2015 | Truly Exploring Multiple References for Machine Translation Evaluation
Lucia Specia |
EAMT | 2 |
| 2015 | Searching for Context: a Study on Document-Level Labels for Translation Quality Estimation
Carolina Scarton, Marcos Zampieri, Mihaela Vela, Josef van Genabith, Lucia Specia |
EAMT | 5 |
| 2015 | Investigating Continuous Space Language Models for Machine Translation Quality EstimationabstractWe present novel features designed with a deep neural network for Machine Translation (MT) Quality Estimation (QE).The features are learned with a Continuous Space Language Model to estimate the probabilities of the source and target segments.These new features, along with standard MT system-independent features, are benchmarked on a series of datasets with various quality labels, including postediting effort, human translation edit rate, post-editing time and METEOR.Results show significant improvements in prediction over the baseline, as well as over systems trained on state of the art feature sets for all datasets.More notably, the addition of the newly proposed features improves over the best QE systems in WMT12 and WMT14 by a significant margin. Kashif Shah, Raymond W. M. Ng, Fethi Bougares, Lucia Specia |
EMNLP | 4 |
| 2015 | Quality estimation for asr k-best list rescoring in spoken language translationabstractSpoken language translation (SLT) combines automatic speech recognition (ASR) and machine translation (MT). During the decoding stage, the best hypothesis produced by the ASR system may not be the best input candidate to the MT system, but making use of multiple sub-optimal ASR results in SLT has been shown to be too complex computationally. This paper presents a method to rescore the k-best ASR output such as to improve translation quality. A translation quality estimation model is trained on a large number of features which aim to capture complementary information from both ASR and MT on translation difficulty and adequacy, as well as syntactic properties of the SLT inputs and outputs. Based on the predicted quality score, the ASR hypotheses are rescored before they are fed to the MT system. ASR confidence is found to be crucial in guiding the rescoring step. In an English-to-French speech-to-text translation task, the coupling of ASR and MT systems led to an increase of 0.5 BLEU points in translation quality. Raymond W. M. Ng, Kashif Shah, Wilker Aziz, Lucia Specia, Thomas Hain |
ICASSP | 4 |
| 2015 | A study on the stability and effectiveness of features in quality estimation for spoken language translationabstractA quality estimation (QE) approach informed with machine translation (MT) and speech recognition (ASR) features has recently shown to improve the performance of a spoken language translation (SLT) system in an in-domain scenario. When domain mismatch is progressively introduced in the MT and ASR systems, the SLT system’s performance naturally degrades. The use of QE to improve SLT performance has not been studied in this context. In this paper we investigate the effectiveness of QE under this setting. Our experiments showed that across moderate levels of domain mismatches, QE led to consistent translation improvements of around 0.4 in BLEU score. The QE system relies on 116 features derived from the ASR and MT system input and output. Feature analysis was conducted to understand the information sources contributing the most to performance improvements. LDA dimension reduction was used to summarise effective features into sets as small as 3 without affecting the SLT performance. By inspecting the principal components, eight features including the acoustic model scores and count-based word statistics on the bilingual text were found to be critically important, leading to a further boost of around 0.1 BLEU score over the full set of features. These findings provide interesting possibilities for further work by incorporating the effective QE features in SLT system training or decoding. Raymond W. M. Ng, Kashif Shah, Lucia Specia, Thomas Hain |
INTERSPEECH | 3 |
| 2015 | A Bayesian non-linear method for feature selection in machine translation quality estimation
Kashif Shah, Trevor Cohn, Lucia Specia |
Mach. Transl. | 3 |
| 2015 | Learning Structural Kernels for Natural Language ProcessingabstractStructural kernels are a flexible learning paradigm that has been widely used in Natural Language Processing. However, the problem of model selection in kernel-based methods is usually overlooked. Previous approaches mostly rely on setting default values for kernel hyperparameters or using grid search, which is slow and coarse-grained. In contrast, Bayesian methods allow efficient model selection by maximizing the evidence on the training data through gradient-based methods. In this paper we show how to perform this in the context of structural kernels by using Gaussian Processes. Experimental results on tree kernels show that this procedure results in better prediction performance compared to hyperparameter optimization via grid search. The framework proposed in this paper can be adapted to other structures besides trees, e.g., strings and graphs, thereby extending the utility of kernel-based methods. Daniel Beck, Trevor Cohn, Christian Hardmeier, Lucia Specia |
Trans. Assoc. Comput. Linguistics | 4 |
| 2014 | Statistical Relational Learning to Recognise Textual Entailment
Miguel Ángel Ríos-Gaona, Lucia Specia, Alexander F. Gelbukh, Ruslan Mitkov |
CICLing (1) | 2 |
| 2014 | Document-level translation quality estimation: exploring discourse and pseudo-references
Carolina Scarton, Lucia Specia |
EAMT | 2 |
| 2014 | Quality estimation for translation selection
Kashif Shah, Lucia Specia |
EAMT | 2 |
| 2014 | Data selection for discriminative training in statistical machine translation
Xingyi Song, Lucia Specia, Trevor Cohn |
EAMT | 2 |
| 2014 | Exact Decoding for Phrase-Based Statistical Machine TranslationabstractThe combinatorial space of translation derivations in phrase-based statistical ma-chine translation is given by the intersec-tion between a translation lattice and a tar-get language model. We replace this in-tractable intersection by a tractable relax-ation which incorporates a low-order up-perbound on the language model. Exact optimisation is achieved through a coarse-to-fine strategy with connections to adap-tive rejection sampling. We perform ex-act optimisation with unpruned language models of order 3 to 5 and show search-error curves for beam search and cube pruning on standard test sets. This is the first work to tractably tackle exact opti-misation with language models of orders higher than 3. 1 Wilker Aziz, Marc Dymetman, Lucia Specia |
EMNLP | 3 |
| 2014 | Joint Emotion Analysis via Multi-task Gaussian ProcessesabstractWe propose a model for jointly predicting multiple emotions in natural language sentences.Our model is based on a low-rank coregionalisation approach, which combines a vector-valued Gaussian Process with a rich parameterisation scheme.We show that our approach is able to learn correlations and anti-correlations between emotions on a news headlines dataset.The proposed model outperforms both singletask baselines and other multi-task approaches. Daniel Beck, Trevor Cohn, Lucia Specia |
EMNLP | 3 |
| 2014 | A Quality-based Active Sample Selection Strategy for Statistical Machine Translation
Varvara Logacheva, Lucia Specia |
LREC | 2 |
| 2014 | An efficient and user-friendly tool for machine translation quality estimation
Kashif Shah, Marco Turchi, Lucia Specia |
LREC | 3 |
| 2013 | Modelling Annotator Bias with Multi-task Gaussian Processes: An Application to Machine Translation Quality Estimation
Trevor Cohn, Lucia Specia |
ACL (1) | 2 |
| 2013 | Multilingual WSD-like Constraints for Paraphrase Extraction
Wilker Aziz, Lucia Specia |
CoNLL | 2 |
| 2013 | An Investigation on the Effectiveness of Features for Translation Quality Estimation
Kashif Shah, Trevor Conn, Lucia Specia |
MTSummit | 3 |
| 2013 | Investigating the contribution of linguistic information to quality estimation
Mariano Felice, Lucia Specia |
Mach. Transl. | 2 |
| 2013 | Kirsten Malmkjær and Kevin Windle (eds.): The Oxford handbook of translation studies - Oxford University Press, 2011, xvii + 607 pp, ISBN: 978-0-19-923930-6
Lucia Specia |
Mach. Transl. | 1 |
| 2013 | Quality estimation for machine translation: preface
Lucia Specia, Radu Soricut |
Mach. Transl. | 1 |
| 2012 | Cross-lingual Sentence Compression for Subtitles
Wilker Aziz, Sheila C. M. de Sousa, Lucia Specia |
EAMT | 3 |
| 2012 | Relevance Ranking for Translated Texts
Marco Turchi, Josef Steinberger, Lucia Specia |
EAMT | 3 |
| 2012 | PET: a Tool for Post-editing and Assessing Machine Translation
Wilker Aziz, Sheila Castilho, Lucia Specia |
LREC | 3 |
| 2011 | Exploiting Objective Annotations for Minimising Translation Post-editing Effort
Lucia Specia |
EAMT | 1 |
| 2011 | Predicting Machine Translation Adequacy
Lucia Specia, Najeh Hajlaoui, Catalina Hallett, Wilker Aziz |
MTSummit | 1 |
| 2010 | Learning an Expert from Human Annotations in Statistical Machine Translation: the Case of Out-of-Vocabulary Words
Wilker Aziz, Marc Dymetman, Lucia Specia, Shachar Mirkin |
EAMT | 3 |
| 2010 | A Dataset for Assessing Machine Translation Evaluation Metrics
Lucia Specia, Nicola Cancedda, Marc Dymetman |
LREC | 1 |
| 2010 | Pushing the frontier of Statistical Machine Translation: Preface
Lucia Specia, Nicola Cancedda |
Mach. Transl. | 1 |
| 2010 | Machine translation evaluation versus quality estimation
Lucia Specia, Dhwaj Raj, Marco Turchi |
Mach. Transl. | 1 |
| 2009 | Source-Language Entailment Modeling for Translating Unknown Terms
Shachar Mirkin, Lucia Specia, Nicola Cancedda, Ido Dagan, Marc Dymetman, Idan Szpektor |
ACL/IJCNLP | 2 |
| 2009 | Estimating the Sentence-Level Quality of Machine Translation Systems
Lucia Specia, Marco Turchi, Nicola Cancedda, Nello Cristianini, Marc Dymetman |
EAMT | 1 |
| 2009 | Improving the Confidence of Machine Translation Quality Estimates
Lucia Specia, Marco Turqui, John Shawe-Taylor, Craig Saunders |
MTSummit | 1 |
| 2009 | An investigation into feature construction to assist word sense disambiguation
Lucia Specia, Ashwin Srinivasan 0001, Sachindra Joshi, Ganesh Ramakrishnan, Maria das Graças Volpe Nunes |
Mach. Learn. | 1 |
| 2008 | n-Best Reranking for the Efficient Integration of Word Sense Disambiguation and Statistical Machine Translation
Lucia Specia, Baskaran Sankaran, Maria das Graças Volpe Nunes |
CICLing | 1 |
| 2008 | Towards Brazilian Portuguese automatic text simplification systemsabstractIn this paper we investigate the main linguistic phenomena that can make texts complex and how they could be simplified. We focus on a corpus analysis of simple account texts available on the web for Brazilian Portuguese and propose simplification strategies for this language. This study illustrates the need for text simplification to facilitate accessibility to information by poor literacy readers and potentially by people with other cognitive disabilities. It also highlights characteristics of simplification for Portuguese, which may differ from other languages. Such study consists of the first step towards building Brazilian Portuguese text simplification systems. One of the scenarios in which these systems could be used is that of reading electronic texts produced, e.g., by the Brazilian government or by relevant news agencies. Sandra M. Aluísio, Lucia Specia, Thiago A. S. Pardo, Erick Galani Maziero, Renata Pontin de Mattos Fortes |
ACM Symposium on Document Engineering | 2 |
| 2007 | Learning Expressive Models for Word Sense Disambiguation
Lucia Specia, Mark Stevenson 0001, Maria das Graças Volpe Nunes |
ACL | 1 |
| 2007 | Integrating Folksonomies with the Semantic Web
Lucia Specia, Enrico Motta |
ESWC | 1 |
| 2006 | A Hybrid Relational Approach for WSD - First Results
Lucia Specia |
ACL | 1 |
| 2006 | Translation Context Sensitive WSD
Lucia Specia, Maria das Graças Volpe Nunes, Mark Stevenson 0001 |
EAMT | 1 |
| 2006 | A Hybrid Approach for Relation Extraction Aimed at the Semantic Web
Lucia Specia, Enrico Motta |
FQAS | 1 |
| 2006 | Word Sense Disambiguation Using Inductive Logic Programming
Lucia Specia, Ashwin Srinivasan 0001, Ganesh Ramakrishnan, Maria das Graças Volpe Nunes |
ILP | 1 |