EDBT 2026 Demo / reviewers in the wild / expert
Marta R. Costa-jussà
dblp:17/2183 · also Marta Ruiz Costa-jussà
· DBLP profile ↗
76ranked-venue papers
31as first author
28since 2021 · last 2025
0000-0002-5703-520XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 69 · 28 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Language and Modality Transfer in Translation by Character-level ModelingabstractCurrent translation systems, despite being highly multilingual, cover only 5% of the world’s languages. Expanding language coverage to the long-tail of low-resource languages requires data-efficient methods that rely on cross-lingual and cross-modal knowledge transfer. To this end, we propose a character-based approach to improve adaptability to new languages and modalities. Our method leverages SONAR, a multilingual fixed-size embedding space with different modules for encoding and decoding. We use a teacher-student approach with parallel translation data to obtain a character-level encoder. Then, using ASR data, we train a lightweight adapter to connect a massively multilingual CTC ASR model (MMS), to the character-level encoder, potentially enabling speech translation from 1,000+ languages. Experimental results in text translation for 75 languages on FLORES+ demonstrate that our character-based approach can achieve better language transfer than traditional subword-based models, especially outperforming them in low-resource settings, and demonstrating better zero-shot generalizability to unseen languages. Our speech adaptation, maximizing knowledge transfer from the text modality, achieves state-of-the-art results in speech-to-text translation on the FLEURS benchmark on 33 languages, surpassing previous supervised and cascade models, albeit being a zero-shot model with minimal supervision from ASR data. Ioannis Tsiamas, David Dale, Marta R. Costa-jussà |
ACL (1) | 3 |
| 2025 | BOUQuET : dataset, Benchmark and Open initiative for Universal Quality Evaluation in TranslationabstractPierre Andrews, Mikel Artetxe, Mariano Coria Meglioli, Marta R. Costa-jussà, Joe Chuang, David Dale, Mark Duppenthaler, Nathanial Paul Ekberg, Cynthia Gao, Daniel Edward Licht, Jean Maillard, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Eduardo Sánchez, Ioannis Tsiamas, Arina Turkatenko, Albert Ventayol-Boada, Shireen Yates. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Pierre Andrews, Mikel Artetxe, Mariano Coria Meglioli, Marta R. Costa-jussà, Joe Chuang, David Dale, Mark Duppenthaler, Nathanial Paul Ekberg, Cynthia Gao, Daniel Edward Licht, Jean Maillard, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Eduardo Sánchez, Ioannis Tsiamas, Arina Turkatenko, Albert Ventayol-Boada, Shireen Yates |
EMNLP | 4 |
| 2025 | On the Role of Speech Data in Reducing Toxicity Detection BiasabstractSamuel Bell, Mariano Coria Meglioli, Megan Richards, Eduardo Sánchez, Christophe Ropers, Skyler Wang, Adina Williams, Levent Sagun, Marta R. Costa-jussà. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Samuel J. Bell, Mariano Coria Meglioli, Megan Richards, Eduardo Sánchez, Christophe Ropers, Skyler Wang, Adina Williams, Levent Sagun, Marta R. Costa-jussà |
NAACL (Long Papers) | 9 |
| 2025 | Linguini: A benchmark for language-agnostic linguistic reasoningabstractWe propose a new benchmark to measure a language model's linguistic reasoning skills without relying on pre-existing language-specific knowledge. The test covers 894 questions grouped in 160 problems across 75 (mostly) extremely low-resource languages, extracted from the International Linguistic Olympiad corpus. To attain high accuracy on this benchmark, models don't need previous knowledge of the tested language, as all the information needed to solve the linguistic puzzle is presented in the context. We find that, while all analyzed models rank below 25% accuracy, there is a significant gap between open and closed models, with the best-performing proprietary model scoring 24.05% and the best-performing open model 8.84%. Eduardo Sánchez, Belen Alastruey, Christophe Ropers, Arina Turkatenko, Pontus Stenetorp, Mikel Artetxe, Marta R. Costa-jussà |
NeurIPS | 7 |
| 2024 | SpeechAlign: A Framework for Speech Translation Alignment EvaluationabstractSpeech-to-Speech and Speech-to-Text translation are currently dynamic areas of research. In our commitment to advance these fields, we present SpeechAlign, a framework designed to evaluate the underexplored field of source-target alignment in speech models. The SpeechAlign framework has two core components. First, to tackle the absence of suitable evaluation datasets, we introduce the Speech Gold Alignment dataset, built upon a English-German text translation gold alignment dataset. Secondly, we introduce two novel metrics, Speech Alignment Error Rate (SAER) and Time-weighted Speech Alignment Error Rate (TW-SAER), which enable the evaluation of alignment quality within speech models. While the former gives equal importance to each word, the latter assigns weights based on the length of the words in the speech signal. By publishing SpeechAlign we provide an accessible evaluation framework for model assessment, and we employ it to benchmark open-source Speech Translation models. In doing so, we contribute to the ongoing research progress within the fields of Speech-to-Speech and Speech-to-Text translation. Belen Alastruey, Aleix Sant, Gerard I. Gállego, David Dale, Marta R. Costa-jussà |
LREC/COLING | 5 |
| 2024 | Added Toxicity Mitigation at Inference Time for Multimodal and Massively Multilingual TranslationabstractMachine translation models sometimes lead to added toxicity: translated outputs may contain more toxic content that the original input. In this paper, we introduce MinTox, a novel pipeline to automatically identify and mitigate added toxicity at inference time, without further model training. MinTox leverages a multimodal (speech and text) toxicity classifier that can scale across languages.We demonstrate the capabilities of MinTox when applied to SEAMLESSM4T, a multi-modal and massively multilingual machine translation system. MinTox significantly reduces added toxicity: across all domains, modalities and language directions, 25% to95% of added toxicity is successfully filtered out, while preserving translation quality Marta R. Costa-jussà, David Dale, Maha Elbayad, Bokai Yu |
EAMT (1) | 1 |
| 2024 | ReSeTOX: Re-learning attention weights for toxicity mitigation in machine translationabstractOur proposed method, RESETOX (REdoSEarch if TOXic), addresses the issue ofNeural Machine Translation (NMT) gener-ating translation outputs that contain toxicwords not present in the input. The ob-jective is to mitigate the introduction oftoxic language without the need for re-training. In the case of identified addedtoxicity during the inference process, RE-SETOX dynamically adjusts the key-valueself-attention weights and re-evaluates thebeam search hypotheses. Experimental re-sults demonstrate that RESETOX achievesa remarkable 57% reduction in added tox-icity while maintaining an average trans-lation quality of 99.5% across 164 lan-guages. Our code is available at: https://github.com Javier García Gilabert, Carlos Escolano, Marta R. Costa-jussà |
EAMT (1) | 3 |
| 2024 | Unveiling the Role of Pretraining in Direct Speech TranslationabstractDirect speech-to-text translation systems encounter an important drawback in data scarcity.A common solution consists on pretraining the encoder on automatic speech recognition, hence losing efficiency in the training process.In this study, we compare the training dynamics of a system using a pretrained encoder, the conventional approach, and one trained from scratch.We observe that, throughout the training, the randomly initialized model struggles to incorporate information from the speech inputs for its predictions.Hence, we hypothesize that this issue stems from the difficulty of effectively training an encoder for direct speech translation.While a model trained from scratch needs to learn acoustic and semantic modeling simultaneously, a pretrained one can just focus on the latter.Based on these findings, we propose a subtle change in the decoder crossattention to integrate source information from earlier steps in training.We show that with this change, the model trained from scratch can achieve comparable performance to the pretrained one, while reducing the training time. Belen Alastruey, Gerard I. Gállego, Marta R. Costa-jussà |
EMNLP | 3 |
| 2023 | BLASER: A Text-Free Speech-to-Speech Translation Evaluation MetricabstractMingda Chen, Paul-Ambroise Duquenne, Pierre Andrews, Justine Kao, Alexandre Mourachko, Holger Schwenk, Marta R. Costa-jussà. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Mingda Chen, Paul-Ambroise Duquenne, Pierre Andrews, Justine Kao, Alexandre Mourachko, Holger Schwenk, Marta R. Costa-jussà |
ACL (1) | 7 |
| 2023 | Detecting and Mitigating Hallucinations in Machine Translation: Model Internal Workings Alone Do Well, Sentence Similarity Even BetterabstractWhile the problem of hallucinations in neural machine translation has long been recognized, so far the progress on its alleviation is very little.Indeed, recently it turned out that without artificially encouraging models to hallucinate, previously existing methods fall short and even the standard sequence log-probability is more informative.It means that internal characteristics of the model can give much more information than we expect, and before using external models and measures, we first need to ask: how far can we go if we use nothing but the translation model itself ?We propose to use a method that evaluates the percentage of the source contribution to a generated translation.Intuitively, hallucinations are translations "detached" from the source, hence they can be identified by low source contribution.This method improves detection accuracy for the most severe hallucinations by a factor of 2 and is able to alleviate hallucinations at test time on par with the previous best approach that relies on external models.Next, if we move away from internal model characteristics and allow external tools, we show that using sentence similarity from cross-lingual embeddings further improves these results.We release the code of our experiments.1 David Dale, Elena Voita, Loïc Barrault, Marta R. Costa-jussà |
ACL (1) | 4 |
| 2023 | Explaining How Transformers Use Context to Build PredictionsabstractLanguage Generation Models produce words based on the previous context.Although existing methods offer input attributions as explanations for a model's prediction, it is still unclear how prior words affect the model's decision throughout the layers.In this work, we leverage recent advances in explainability of the Transformer and present a procedure to analyze models for language generation.Using contrastive examples, we compare the alignment of our explanations with evidence of the linguistic phenomena, and show that our method consistently aligns better than gradient-based and perturbation-based baselines.Then, we investigate the role of MLPs inside the Transformer and show that they learn features that help the model predict words that are grammatically acceptable.Lastly, we apply our method to Neural Machine Translation models, and demonstrate that they generate human-like source-target alignments for building predictions. Javier Ferrando, Gerard I. Gállego, Ioannis Tsiamas, Marta R. Costa-jussà |
ACL (1) | 4 |
| 2023 | Multilingual Holistic Bias: Extending Descriptors and Patterns to Unveil Demographic Biases in Languages at ScaleabstractMarta Costa-jussà, Pierre Andrews, Eric Smith, Prangthip Hansanti, Christophe Ropers, Elahe Kalbassi, Cynthia Gao, Daniel Licht, Carleigh Wood. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Marta R. Costa-jussà, Pierre Andrews, Eric Michael Smith, Prangthip Hansanti, Christophe Ropers, Elahe Kalbassi, Cynthia Gao, Daniel Licht, Carleigh Wood |
EMNLP | 1 |
| 2023 | HalOmi: A Manually Annotated Benchmark for Multilingual Hallucination and Omission Detection in Machine TranslationabstractDavid Dale, Elena Voita, Janice Lam, Prangthip Hansanti, Christophe Ropers, Elahe Kalbassi, Cynthia Gao, Loic Barrault, Marta Costa-jussà. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. David Dale, Elena Voita, Janice Lam, Prangthip Hansanti, Christophe Ropers, Elahe Kalbassi, Cynthia Gao, Loïc Barrault, Marta R. Costa-jussà |
EMNLP | 9 |
| 2023 | Efficient Speech Translation with Dynamic Latent PerceiversabstractTransformers have been the dominant architecture for Speech Translation in recent years, achieving significant improvements in translation quality. Since speech signals are longer than their textual counterparts, and due to the quadratic complexity of the Transformer, a down-sampling step is essential for its adoption in Speech Translation. Instead, in this research, we propose to ease the complexity by using a Perceiver encoder to map the speech inputs to a fixed-length latent representation. Furthermore, we introduce a novel way of training Perceivers, with Dynamic Latent Access (DLA), unlocking larger latent spaces without any additional computational overhead. Speech-to-Text Perceivers with DLA can match the performance of Transformer baselines across three language pairs in MuST-C. Finally, a DLA-trained model is easily adaptable to DLA at inference, and can be flexibly deployed with various computational budgets, without significant drops in translation quality. Ioannis Tsiamas, Gerard I. Gállego, José A. R. Fonollosa, Marta R. Costa-jussà |
ICASSP | 4 |
| 2023 | Towards lifelong human assisted speaker diarization
Meysam Shamsi, Anthony Larcher, Loïc Barrault, Sylvain Meignier, Yevhenii Prokopalo, Marie Tahon, Ambuj Mehrish, Simon Petit-Renaud, Olivier Galibert, Samuel Gaist, André Anjos, Sébastien Marcel, Marta R. Costa-jussà |
Comput. Speech Lang. | 13 |
| 2022 | Interpreting Gender Bias in Neural Machine Translation: Multilingual Architecture MattersabstractMultilingual neural machine translation architectures mainly differ in the number of sharing modules and parameters applied among languages. In this paper, and from an algorithmic perspective, we explore whether the chosen architecture, when trained with the same data, influences the level of gender bias. Experiments conducted in three language pairs show that language-specific encoder-decoders exhibit less bias than the shared architecture. We propose two methods for interpreting and studying gender bias in machine translation based on source embeddings and attention. Our analysis shows that, in the language-specific case, the embeddings encode more gender information, and their attention is more diverted. Both behaviors help in mitigating gender bias. Marta R. Costa-jussà, Carlos Escolano, Christine Basta, Javier Ferrando, Roser Batlle, Ksenia Kharitonova |
AAAI | 1 |
| 2022 | Towards Opening the Black Box of Neural Machine Translation: Source and Target Interpretations of the TransformerabstractIn Neural Machine Translation (NMT), each token prediction is conditioned on the source sentence and the target prefix (what has been previously translated at a decoding step).However, previous work on interpretability in NMT has mainly focused solely on source sentence tokens' attributions.Therefore, we lack a full understanding of the influences of every input token (source sentence and target prefix) in the model predictions.In this work, we propose an interpretability method that tracks input tokens' attributions for both contexts.Our method, which can be extended to any encoder-decoder Transformer-based model, allows us to better comprehend the inner workings of current NMT models.We apply the proposed method to both bilingual and multilingual Transformers and present insights into their behaviour. Javier Ferrando, Gerard I. Gállego, Belen Alastruey, Carlos Escolano, Marta R. Costa-jussà |
EMNLP | 5 |
| 2022 | Measuring the Mixing of Contextual Information in the TransformerabstractThe Transformer architecture aggregates input information through the self-attention mechanism, but there is no clear understanding of how this information is mixed across the entire model.Additionally, recent works have demonstrated that attention weights alone are not enough to describe the flow of information.In this paper, we consider the whole attention block -multi-head attention, residual connection, and layer normalization-and define a metric to measure token-to-token interactions within each layer.Then, we aggregate layer-wise interpretations to provide input attribution scores for model predictions.Experimentally, we show that our method, ALTI (Aggregation of Layer-wise Token-to-token Interactions), provides more faithful explanations and increased robustness than gradient-based methods.Grad ℓ2 went here just before a movie .the service was fast but that ' s it .i ordered the mango and shrimp quesadilla .my friend ordered nachos .the food was not good .i and my friend could not finish our food and we had stomach aches immediately .IG ℓ2 went here just before a movie .the service was fast but that ' s it .i ordered the mango and shrimp quesadilla .my friend ordered nachos .the food was not good .i and my friend could not finish our food and we had stomach aches immediately .ALTI went here just before a movie .the service was fast but that ' s it .i ordered the mango and shrimp quesadilla .my friend ordered nachos .the food was not good .i and my Javier Ferrando, Gerard I. Gállego, Marta R. Costa-jussà |
EMNLP | 3 |
| 2022 | SHAS: Approaching optimal Segmentation for End-to-End Speech TranslationabstractSpeech translation models are unable to directly process long audios, like TED talks, which have to be split into shorter segments. Speech translation datasets provide manual segmentations of the audios, which are not available in real-world scenarios, and existing segmentation methods usually significantly reduce translation quality at inference time. To bridge the gap between the manual segmentation of training and the automatic one at inference, we propose Supervised Hybrid Audio Segmentation (SHAS), a method that can effectively learn the optimal segmentation from any manually segmented speech corpus. First, we train a classifier to identify the included frames in a segmentation, using speech representations from a pre-trained wav2vec 2.0. The optimal splitting points are then found by a probabilistic Divide-and-Conquer algorithm that progressively splits at the frame of lowest probability until all segments are below a pre-specified length. Experiments on MuST-C and mTEDx show that the translation of the segments produced by our method approaches the quality of the manual segmentation on 5 languages pairs. Namely, SHAS retains 95-98% of the manual segmentation's BLEU score, compared to the 87-93% of the best existing methods. Our method is additionally generalizable to different domains and achieves high zero-shot performance in unseen languages. Copyright © 2022 ISCA. Ioannis Tsiamas, Gerard I. Gállego, José A. R. Fonollosa, Marta R. Costa-jussà |
INTERSPEECH | 4 |
| 2022 | Evaluating Gender Bias in Speech TranslationabstractThe scientific community is increasingly aware of the necessity to embrace pluralism and consistently represent major and minor social groups. Currently, there are no standard evaluation techniques for different types of biases. Accordingly, there is an urgent need to provide evaluation sets and protocols to measure existing biases in our automatic systems. Evaluating the biases should be an essential step towards mitigating them in the systems. This paper introduces WinoST, a new freely available challenge set for evaluating gender bias in speech translation. WinoST is the speech version of WinoMT, an MT challenge set, and both follow an evaluation protocol to measure gender accuracy. Using an S-Transformer end-to-end speech translation system, we report the gender bias evaluation on four language pairs, and we reveal the inaccuracies in translations generating gender-stereotyped translations. Marta R. Costa-jussà, Christine Basta, Gerard I. Gállego |
LREC | 1 |
| 2022 | OccGen: Selection of Real-world Multilingual Parallel Data Balanced in Gender within OccupationsabstractThis paper describes the OCCGEN toolkit, which allows extracting multilingual parallel data balanced in gender within occupations. OCCGEN can extract datasets that reflect gender diversity (beyond binary) more fairly in society to be further used to explicitly mitigate occupational gender stereotypes. We propose two use cases that extract evaluation datasets for machine translation in four high-resourcelanguages from different linguistic families and in a low-resource African language. Our analysis of these use cases shows that translation outputs in high-resource languages tend to worsen in feminine subsets (compared to masculine). This can be explained because less attention is paid to the source sentence. Then, more attention is given to the target prefix overgeneralizing to the most frequent masculine forms. Marta R. Costa-jussà, Christine Basta, Oriol Domingo, André Rubungo |
NeurIPS | 1 |
| 2022 | Multilingual Machine Translation: Deep Analysis of Language-Specific Encoder-DecodersabstractState-of-the-art multilingual machine translation relies on a shared encoder-decoder. In this paper, we propose an alternative approach based on language-specific encoder-decoders, which can be easily extended to new languages by learning their corresponding modules. To establish a common interlingua representation, we simultaneously train N initial languages. Our experiments show that the proposed approach improves over the shared encoder-decoder for the initial languages and when adding new languages, without the need to retrain the remaining modules. All in all, our work closes the gap between shared and language-specific encoder-decoders, advancing toward modular multilingual machine translation systems that can be flexibly extended in lifelong learning settings. Carlos Escolano, Marta R. Costa-jussà, José A. R. Fonollosa |
J. Artif. Intell. Res. | 2 |
| 2021 | Enabling Zero-Shot Multilingual Spoken Language Translation with Language-Specific Encoders and DecodersabstractCurrent end-to-end approaches to Spoken Language Translation (SLT) rely on limited training resources, especially for multilingual settings. On the other hand, Multilingual Neural Machine Translation (MultiNMT) approaches rely on higher-quality and more massive data sets. Our proposed method extends a MultiNMT architecture based on language-specific encoders-decoders to the task of Multilingual SLT (Multi-SLT). Our method entirely eliminates the dependency from MultiSLT data and it is able to translate while training only on ASR and MultiNMT data. Our experiments on four different languages show that coupling the speech encoder to the MultiNMT architecture produces similar quality translations compared to a bilingual baseline (±0.2 BLEU) while effectively allowing for zero-shot MultiSLT. Additionally, we propose using an Adapter module for coupling the speech inputs. This Adapter module produces consistent improvements up to +6 BLEU points on the proposed architecture and +1 BLEU point on the end-to-end baseline. Carlos Escolano, Marta R. Costa-jussà, José A. R. Fonollosa, Carlos Segura |
ASRU | 2 |
| 2021 | Multilingual Machine Translation: Closing the Gap between Shared and Language-specific Encoder-DecodersabstractCarlos Escolano, Marta R. Costa-jussà, José A. R. Fonollosa, Mikel Artetxe. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Carlos Escolano, Marta R. Costa-jussà, José A. R. Fonollosa, Mikel Artetxe |
EACL | 2 |
| 2021 | From bilingual to multilingual neural-based machine translation by incremental trainingabstractAbstract A common intermediate language representation in neural machine translation can be used to extend bilingual systems by incremental training. We propose a new architecture based on introducing an interlingual loss as an additional training objective. By adding and forcing this interlingual loss, we can train multiple encoders and decoders for each language, sharing among them a common intermediate representation. Translation results on the low‐resource tasks (Turkish‐English and Kazakh‐English tasks) show a BLEU improvement of up to 2.8 points. However, results on a larger dataset (Russian‐English and Kazakh‐English) show BLEU losses of a similar amount. While our system provides improvements only for the low‐resource tasks in terms of translation quality, our system is capable of quickly deploying new language pairs without the need to retrain the rest of the system, which may be a game changer in some situations. Specifically, what is most relevant regarding our architecture is that it is capable of: reducing the number of production systems, with respect to the number of languages, from quadratic to linear; incrementally adding a new language to the system without retraining the languages already there; and allowing for translations from the new language to all the others present in the system. Carlos Escolano, Marta R. Costa-jussà, José A. R. Fonollosa |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2021 | Semantic and syntactic information for neural machine translationabstractAbstract Introducing factors such as linguistic features has long been proposed in machine translation to improve the quality of translations. More recently, factored machine translation has proven to still be useful in the case of sequence-to-sequence systems. In this work, we investigate whether this gains hold in the case of the state-of-the-art architecture in neural machine translation, the Transformer, instead of recurrent architectures. We propose a new model, the Factored Transformer, to introduce an arbitrary number of word features in the source sequence in an attentional system. Specifically, we suggest two variants depending on the level at which the features are injected. Moreover, we suggest two combination mechanisms for the word features and words themselves. We experiment both with classical linguistic features and semantic features extracted from a linked data database, and with two low-resource datasets. With the best-found configuration, we show improvements of 0.8 BLEU over the baseline Transformer in the IWSLT German-to-English task. Moreover, we experiment with the more challenging FLoRes English-to-Nepali benchmark, which includes both low-resource and very distant languages, and obtain an improvement of 1.2 BLEU. These improvements are achieved with linguistic and not with semantic information. Jordi Armengol-Estapé, Marta R. Costa-jussà |
Mach. Transl. | 2 |
| 2021 | Extensive study on the underlying gender bias in contextualized word embeddings
Christine Basta, Marta R. Costa-jussà, Noe Casas |
Neural Comput. Appl. | 2 |
| 2021 | Linguistic knowledge-based vocabularies for Neural Machine TranslationabstractAbstract Neural Networks applied to Machine Translation need a finite vocabulary to express textual information as a sequence of discrete tokens. The currently dominant subword vocabularies exploit statistically-discovered common parts of words to achieve the flexibility of character-based vocabularies without delegating the whole learning of word formation to the neural network. However, they trade this for the inability to apply word-level token associations, which limits their use in semantically-rich areas and prevents some transfer learning approaches e.g. cross-lingual pretrained embeddings, and reduces their interpretability. In this work, we propose new hybrid linguistically-grounded vocabulary definition strategies that keep both the advantages of subword vocabularies and the word-level associations, enabling neural networks to profit from the derived benefits. We test the proposed approaches in both morphologically rich and poor languages, showing that, for the former, the quality in the translation of out-of-domain texts is improved with respect to a strong subword baseline. Noe Casas, Marta R. Costa-jussà, José A. R. Fonollosa, Juan A. Alonso, Ramón Fanlo |
Nat. Lang. Eng. | 2 |
| 2020 | Continual Lifelong Learning in Natural Language Processing: A SurveyabstractContinual learning (CL) aims to enable information systems to learn from a continuous data stream across time.However, it is difficult for existing deep learning architectures to learn a new task without largely forgetting previously acquired knowledge.Furthermore, CL is particularly challenging for language learning, as natural language is ambiguous: it is discrete, compositional, and its meaning is context-dependent.In this work, we look at the problem of CL through the lens of various NLP tasks.Our survey discusses major challenges in CL and current methods applied in neural network models.We also provide a critical review of the existing CL evaluation methods and datasets in NLP.Finally, we present our outlook on future research directions. Magdalena Biesialska, Katarzyna Biesialska, Marta R. Costa-jussà |
COLING | 3 |
| 2020 | Refinement of Unsupervised Cross-Lingual Word EmbeddingsabstractCross-lingual word embeddings aim to bridge the gap between high-resource and low-resource languages by allowing to learn multilingual word representations even without using any direct bilingual signal.The lion's share of the methods are projectionbased approaches that map pre-trained embeddings into a shared latent space.These methods are mostly based on the orthogonal transformation, which assumes language vector spaces to be isomorphic.However, this criterion does not necessarily hold, especially for morphologically-rich languages.In this paper, we propose a selfsupervised method to refine the alignment of unsupervised bilingual word embeddings.The proposed model moves vectors of words and their corresponding translations closer to each other as well as enforces length-and center-invariance, thus allowing to better align cross-lingual embeddings.The experimental results demonstrate the effectiveness of our approach, as in most cases it outperforms stateof-the-art methods in a bilingual lexicon induction task. Magdalena Biesialska, Marta R. Costa-jussà |
ECAI | 2 |
| 2020 | Automatic Spanish Translation of SQuAD Dataset for Multi-lingual Question AnsweringabstractRecently, multilingual question answering became a crucial research topic, and it is receiving increased interest in the NLP community. However, the unavailability of large-scale datasets makes it challenging to train multilingual QA systems with performance comparable to the English ones. In this work, we develop the Translate Align Retrieve (TAR) method to automatically translate the Stanford Question Answering Dataset (SQuAD) v1.1 to Spanish. We then used this dataset to train Spanish QA systems by fine-tuning a Multilingual-BERT model. Finally, we evaluated our QA models with the recently proposed MLQA and XQuAD benchmarks for cross-lingual Extractive QA. Experimental results show that our models outperform the previous Multilingual-BERT baselines achieving the new state-of-the-art values of 68.1 F1 on the Spanish MLQA corpus and 77.6 F1 on the Spanish XQuAD corpus. The resulting, synthetically generated SQuAD-es v1.1 corpora, with almost 100% of data contained in the original English version, to the best of our knowledge, is the first large-scale QA training resource for Spanish. Casimiro Pio Carrino, Marta R. Costa-jussà, José A. R. Fonollosa |
LREC | 2 |
| 2020 | Abusive language in Spanish children and young teenager's conversations: data preparation and short text classification with contextual word embeddingsabstractAbusive texts are reaching the interests of the scientific and social community. How to automatically detect them is onequestion that is gaining interest in the natural language processing community. The main contribution of this paper is toevaluate the quality of the recently developed ”Spanish Database for cyberbullying prevention” for the purpose of trainingclassifiers on detecting abusive short texts. We compare classical machine learning techniques to the use of a more ad-vanced model: the contextual word embeddings in the particular case of classification of abusive short-texts for the Spanishlanguage. As contextual word embeddings, we use Bidirectional Encoder Representation from Transformers (BERT), pro-posed at the end of 2018. We show that BERT mostly outperforms classical techniques. Far beyond the experimentalimpact of our research, this project aims at planting the seeds for an innovative technological tool with a high potentialsocial impact and aiming at being part of the initiatives in artificial intelligence for social good. Marta R. Costa-jussà, Esther González 0002, Asunción Moreno, Eudald Cumalat |
LREC | 1 |
| 2020 | GeBioToolkit: Automatic Extraction of Gender-Balanced Multilingual Corpus of Wikipedia BiographiesabstractWe introduce GeBioToolkit, a tool for extracting multilingual parallel corpora at sentence level, with document and gender information from Wikipedia biographies. Despite the gender inequalities present in Wikipedia, the toolkit has been designed to extract corpus balanced in gender. While our toolkit is customizable to any number of languages (and different domains), in this work we present a corpus of 2,000 sentences in English, Spanish and Catalan, which has been post-edited by native speakers to become a high-quality dataset for machine translation evaluation. While GeBioCorpus aims at being one of the first non-synthetic gender-balanced test datasets, GeBioToolkit aims at paving the path to standardize procedures to produce gender-balanced datasets. Marta R. Costa-jussà, Pau Li Lin, Cristina España-Bonet |
LREC | 1 |
| 2020 | Multilingual and Interlingual Semantic Representations for Natural Language Processing: A Brief IntroductionabstractWe introduce the Computational Linguistics special issue on Multilingual and Interlingual Semantic Representations for Natural Language Processing. We situate the special issue’s five articles in the context of our fast-changing field, explaining our motivation for this project. We offer a brief summary of the work in the issue, which includes developments on lexical and sentential semantic representations, from symbolic and neural perspectives. Marta R. Costa-jussà, Cristina España-Bonet, Pascale Fung, Noah A. Smith |
Comput. Linguistics | 1 |
| 2019 | Impact of Gender Debiased Word Embeddings in Language Modeling
Christine Basta, Marta R. Costa-jussà |
CICLing (1) | 2 |
| 2019 | Chinese-Catalan: A Neural Machine Translation Approach Based on Pivoting and Attention MechanismsabstractThis article innovatively addresses machine translation from Chinese to Catalan using neural pivot strategies trained without any direct parallel data. The Catalan language is very similar to Spanish from a linguistic point of view, which motivates the use of Spanish as pivot language. Regarding neural architecture, we are using the latest state-of-the-art, which is the Transformer model, only based on attention mechanisms. Additionally, this work provides new resources to the community, which consists of a human-developed gold standard of 4,000 sentences between Catalan and Chinese and all the others United Nations official languages (Arabic, English, French, Russian, and Spanish). Results show that the standard pseudo-corpus or synthetic pivot approach performs better than cascade. Marta R. Costa-jussà, Noe Casas, Carlos Escolano, José A. R. Fonollosa |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2018 | From Feature to Paradigm: Deep Learning in Machine Translation (Extended Abstract)abstractIn the last years, deep learning algorithms have highly revolutionized several areas including speech, image and natural language processing. The specific field of Machine Translation (MT) has not remained invariant. Integration of deep learning in MT varies from re-modeling existing features into standard statistical systems to the development of a new architecture. Among the different neural networks, research works use feed-forward neural networks, recurrent neural networks and the encoder-decoder schema. These architectures are able to tackle challenges as having low-resources or morphology variations. This extended abstract focuses on describing the foundational works on the neural MT approach; mentioning its strengths and weaknesses; and including an analysis of the corresponding challenges and future work. The full manuscript [Costa-jussà, 2018] describes, in addition, how these neural networks have been integrated to enhance different aspects and models from statistical MT, including language modeling, word alignment, translation, reordering, and rescoring; and on describing the new neural MT approach together with recent approaches on using subword, characters and training with multilingual languages, among others. Marta R. Costa-jussà |
IJCAI | 1 |
| 2018 | From Feature To Paradigm: Deep Learning In Machine TranslationabstractIn the last years, deep learning algorithms have highly revolutionized several areas including speech, image and natural language processing. The specific field of Machine Translation (MT) has not remained invariant. Integration of deep learning in MT varies from re-modeling existing features into standard statistical systems to the development of a new architecture. Among the different neural networks, research works use feed-forward neural networks, recurrent neural networks and the encoder-decoder schema. These architectures are able to tackle challenges as having low-resources or morphology variations. This manuscript focuses on describing how these neural networks have been integrated to enhance different aspects and models from statistical MT, including language modeling, word alignment, translation, reordering, and rescoring. Then, we report the new neural MT approach together with a description of the foundational related works and recent approaches on using subword, characters and training with multilingual languages, among others. Finally, we include an analysis of the corresponding challenges and future work in using deep learning in MT. Marta R. Costa-jussà |
J. Artif. Intell. Res. | 1 |
| 2017 | Bridging deep and kernel methods
Lluís A. Belanche Muñoz, Marta R. Costa-jussà |
ESANN | 2 |
| 2017 | Introduction to the special issue on deep learning approaches for machine translation
Marta R. Costa-jussà, Alexandre Allauzen, Loïc Barrault, Kyunghyun Cho, Holger Schwenk |
Comput. Speech Lang. | 1 |
| 2017 | Chinese-Spanish neural machine translation enhanced with character and word bitmap fonts
Marta R. Costa-jussà, David Aldón, José A. R. Fonollosa |
Mach. Transl. | 1 |
| 2016 | Combining Phrase and Neural-Based Machine Translation: What Worked and Did Not
Marta R. Costa-jussà, José A. R. Fonollosa |
CICLing (2) | 1 |
| 2016 | Introduction to the Special Issue on Cross-Language Algorithms and ApplicationsabstractWith the increasingly global nature of our everyday interactions, the need for multilin- gual technologies to support efficient and effective information access and communication cannot be overemphasized. Computational modeling of language has been the focus of Natural Language Processing, a subdiscipline of Artificial Intelligence. One of the current challenges for this discipline is to design methodologies and algorithms that are cross- language in order to create multilingual technologies rapidly. The goal of this JAIR special issue on Cross-Language Algorithms and Applications (CLAA) is to present leading re- search in this area, with emphasis on developing unifying themes that could lead to the development of the science of multi- and cross-lingualism. In this introduction, we provide the reader with the motivation for this special issue and summarize the contributions of the papers that have been included. The selected papers cover a broad range of cross-lingual technologies including machine translation, domain and language adaptation for sentiment analysis, cross-language lexical resources, dependency parsing, information retrieval and knowledge representation. We anticipate that this special issue will serve as an invaluable resource for researchers interested in topics of cross-lingual natural language processing. Marta R. Costa-jussà, Srinivas Bangalore, Patrik Lambert, Lluís Màrquez, Elena Montiel-Ponsoda |
J. Artif. Intell. Res. | 1 |
| 2016 | Selection of correction candidates for the normalization of Spanish user-generated contentabstractAbstract We present research aiming to build tools for the normalization of User-Generated Content (UGC). We argue that processing this type of text requires the revisiting of the initial steps of Natural Language Processing, since UGC (micro-blog, blog, and, generally, Web 2.0 user-generated texts) presents a number of nonstandard communicative and linguistic characteristics – often closer to oral and colloquial language than to edited text. We present a corpus of UGC text in Spanish from three different sources: Twitter, consumer reviews, and blogs, and describe its main characteristics. We motivate the need for UGC text normalization by analyzing the problems found when processing this type of text through a conventional language processing pipeline, particularly in the tasks of lemmatization and morphosyntactic tagging. Our aim with this paper is to seize the power of already existing spell and grammar correction engines and endow them with automatic normalization capabilities in order to pave the way for the application of standard Natural Language Processing tools to typical UGC text. Particularly, we propose a strategy for automatically normalizing UGC by adding a module on top of a pre-existing spell-checker that selects the most plausible correction from an unranked list of candidates provided by the spell-checker. To build this selector module we train four language models, each one containing a different type of linguistic information in a trade-off with its generalization capabilities. Our experiments show that the models trained on truecase and lowercase word forms are more discriminative than the others at selecting the best candidate. We have also experimented with a parametrized combination of the models by both optimizing directly on the selection task and doing a linear interpolation of the models. The resulting parametrized combinations obtain results close to the best performing model but do not improve on those results, as measured on the test set. The precision of the selector module in ranking number one the expected correction proposal on the test corpora reaches 82.5% for Twitter text (baseline 57%) and 88% for non-Twitter text (baseline 64%). Maite Melero, Marta R. Costa-jussà, Patrik Lambert, Martí Quixal |
Nat. Lang. Eng. | 2 |
| 2016 | A deep source-context feature for lexical selection in statistical machine translation
Parth Gupta, Marta R. Costa-jussà, Paolo Rosso, Rafael E. Banchs |
Pattern Recognit. Lett. | 2 |
| 2016 | Description of the Chinese-to-Spanish Rule-Based Machine Translation System Developed Using a Hybrid Combination of Human Annotation and Statistical TechniquesabstractTwo of the most popular Machine Translation (MT) paradigms are rule based (RBMT) and corpus based, which include the statistical systems (SMT). When scarce parallel corpus is available, RBMT becomes particularly attractive. This is the case of the Chinese--Spanish language pair. This article presents the first RBMT system for Chinese to Spanish. We describe a hybrid method for constructing this system taking advantage of available resources such as parallel corpora that are used to extract dictionaries and lexical and structural transfer rules. The final system is freely available online and open source. Although performance lags behind standard SMT systems for an in-domain test set, the results show that the RBMT’s coverage is competitive and it outperforms the SMT system in an out-of-domain test set. This RBMT system is available to the general public, it can be further enhanced, and it opens up the possibility of creating future hybrid MT systems. Marta R. Costa-jussà, Jordi Centelles |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2015 | Latest trends in hybrid machine translation and its applicationsabstractThis survey on hybrid machine translation (MT) is motivated by the fact that hybridization techniques have become popular as they attempt to combine the best characteristics of highly advanced pure rule or corpus-based MT approaches. Existing research typically covers either simple or more complex architectures guided by either rule or corpus-based approaches. The goal is to combine the best properties of each type. This survey provides a detailed overview of the modification of the standard rule-based architecture to include statistical knowledge, the introduction of rules in corpus-based approaches, and the hybridization of approaches within this last single category. The principal aim here is to cover the leading research and progress in this field of MT and in several related applications. Marta R. Costa-jussà, José A. R. Fonollosa |
Comput. Speech Lang. | 1 |
| 2015 | Editorial
José A. R. Fonollosa, Marta R. Costa-jussà |
Comput. Speech Lang. | 2 |
| 2015 | How much hybridization does machine translation Need?abstractRule‐based and corpus‐based machine translation (MT) have coexisted for more than 20 years. Recently, boundaries between the two paradigms have narrowed and hybrid approaches are gaining interest from both academia and businesses. However, since hybrid approaches involve the multidisciplinary interaction of linguists, computer scientists, engineers, and information specialists, understandably a number of issues exist. While statistical methods currently dominate research work in MT, most commercial MT systems are technically hybrid systems. The research community should investigate the benefits and questions surrounding the hybridization of MT systems more actively. This paper discusses various issues related to hybrid MT including its origins, architectures, achievements, and frustrations experienced in the community. It can be said that both rule‐based and corpus‐ based MT systems have benefited from hybridization when effectively integrated. In fact, many of the current rule/corpus‐based MT approaches are already hybridized since they do include statistics/rules at some point. Marta R. Costa-jussà |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2014 | An IR-Based Strategy for Supporting Chinese-Portuguese Translation Services in Off-line Mode
Jordi Centelles, Marta R. Costa-jussà, Rafael E. Banchs, Alexander F. Gelbukh |
CICLing (2) | 2 |
| 2014 | CHISPA on the GO: A mobile Chinese-Spanish translation service for travellers in troubleabstractThis demo showcases a translation service that allows travelers to have an easy and convenient access to Chinese-Spanish translations via a mobile app. The system integrates a phrase-based translation system with other open source components such as Optical Character Recognition and Automatic Speech Recognition to provide a very friendly user experience. Jordi Centelles, Marta R. Costa-jussà, Rafael E. Banchs |
EACL | 2 |
| 2014 | A client mobile application for Chinese-Spanish statistical machine translation
Jordi Centelles, Marta R. Costa-jussà, Rafael E. Banchs |
INTERSPEECH | 2 |
| 2014 | Using annotations on Mechanical Turk to perform supervised polarity classification of Spanish customer comments
Marta R. Costa-jussà, Jens Grivolla, Bart Mellebeek, Francesc Benavent, Joan Codina, Rafael E. Banchs |
Inf. Sci. | 1 |
| 2013 | Evaluating Indirect Strategies for Chinese - Spanish Statistical Machine Translation: Extended Abstract
Marta R. Costa-jussà, Carlos A. Henríquez Q., Rafael E. Banchs |
IJCAI | 1 |
| 2013 | Morphological, Syntactical and Semantic Knowledge in Statistical Machine Translation
Marta R. Costa-jussà, Chris Quirk |
HLT-NAACL | 1 |
| 2012 | BUCEADOR, a multi-language search engine for digital libraries
Jordi Adell, Antonio Bonafonte, Antonio Cardenal López, Marta R. Costa-jussà, José A. R. Fonollosa, Asunción Moreno, Eva Navas, Eduardo Rodríguez Banga |
LREC | 4 |
| 2012 | A Richly Annotated, Multilingual Parallel Corpus for Hybrid Machine Translation
Eleftherios Avramidis, Marta R. Costa-jussà, Christian Federmann, Josef van Genabith, Maite Melero, Pavel Pecina |
LREC | 2 |
| 2012 | The ML4HMT Workshop on Optimising the Division of Labour in Hybrid Machine Translation
Christian Federmann, Eleftherios Avramidis, Marta R. Costa-jussà, Josef van Genabith, Maite Melero, Pavel Pecina |
LREC | 3 |
| 2012 | Holaaa!! writin like u talk is kewl but kinda hard 4 NLP
Maite Melero, Marta R. Costa-jussà, Judith Domingo, Montse Marquina, Martí Quixal |
LREC | 2 |
| 2012 | Evaluating Indirect Strategies for Chinese-Spanish Statistical Machine TranslationabstractAlthough, Chinese and Spanish are two of the most spoken languages in the world, not much research has been done in machine translation for this language pair. This paper focuses on investigating the state-of-the-art of Chinese-to-Spanish statistical machine translation (SMT), which nowadays is one of the most popular approaches to machine translation. For this purpose, we report details of the available parallel corpus which are Basic Traveller Expressions Corpus (BTEC), Holy Bible and United Nations (UN). Additionally, we conduct experimental work with the largest of these three corpora to explore alternative SMT strategies by means of using a pivot language. Three alternatives are considered for pivoting: cascading, pseudo-corpus and triangulation. As pivot language, we use either English, Arabic or French. Results show that, for a phrase-based SMT system, English is the best pivot language between Chinese and Spanish. We propose a system output combination using the pivot strategies which is capable of outperforming the direct translation strategy. The main objective of this work is motivating and involving the research community to work in this important pair of languages given their demographic impact. Marta R. Costa-jussà, Carlos A. Henríquez Q., Rafael E. Banchs |
J. Artif. Intell. Res. | 1 |
| 2012 | Study and correlation analysis of linguistic, perceptual, and automatic machine translation evaluationsabstractAbstract Evaluation of machine translation output is an important task. Various human evaluation techniques as well as automatic metrics have been proposed and investigated in the last decade. However, very few evaluation methods take the linguistic aspect into account. In this article, we use an objective evaluation method for machine translation output that classifies all translation errors into one of the five following linguistic levels: orthographic, morphological, lexical, semantic, and syntactic. Linguistic guidelines for the target language are required, and human evaluators use them in to classify the output errors. The experiments are performed on English‐to‐Catalan and Spanish‐to‐Catalan translation outputs generated by four different systems: 2 rule‐based and 2 statistical. All translations are evaluated using the 3 following methods: a standard human perceptual evaluation method, several widely used automatic metrics, and the human linguistic evaluation. Pearson and Spearman correlation coefficients between the linguistic, perceptual, and automatic results are then calculated, showing that the semantic level correlates significantly with both perceptual evaluation and automatic metrics. Mireia Farrús, Marta R. Costa-jussà, Maja Popovic |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2011 | Enhancing scarce-resource language translation through pivot combinations
Marta R. Costa-jussà, Carlos A. Henríquez Q., Rafael E. Banchs |
IJCNLP | 1 |
| 2011 | A vector-space dynamic feature for phrase-based statistical machine translation
Marta R. Costa-jussà, Rafael E. Banchs |
J. Intell. Inf. Syst. | 1 |
| 2010 | Integration of statistical collocation segmentations in a phrase-based statistical machine translation system
Marta R. Costa-jussà, Vidas Daudaravicius, Rafael E. Banchs |
EAMT | 1 |
| 2010 | Linguistic-based Evaluation Criteria to identify Statistical Machine Translation Errors
Mireia Farrús, Marta R. Costa-jussà, José B. Mariño, José A. R. Fonollosa |
EAMT | 2 |
| 2010 | Where are you From? - Tell Me HOW you Write and I Will Tell you WHO you are
Marta R. Costa-jussà, Rafael E. Banchs, Joan Codina |
ICAART (1) | 1 |
| 2010 | Using Linear Interpolation and Weighted Reordering Hypotheses in the Moses System
Marta R. Costa-jussà, José A. R. Fonollosa |
LREC | 1 |
| 2010 | Automatic and Human Evaluation Study of a Rule-based and a Statistical Catalan-Spanish Machine Translation Systems
Marta R. Costa-jussà, Mireia Farrús, José B. Mariño, José A. R. Fonollosa |
LREC | 1 |
| 2009 | Improving a Catalan-Spanish Statistical Translation System using Morphosyntactic Knowledge
Mireia Farrús, Marta R. Costa-jussà, Marc Poch, Adolfo Hernández, José B. Mariño |
EAMT | 2 |
| 2009 | An Ngram-based reordering model
Marta R. Costa-jussà, José A. R. Fonollosa |
Comput. Speech Lang. | 1 |
| 2008 | Using Reordering in Statistical Machine Translation based on Alignment Block Classification
Marta R. Costa-jussà, José A. R. Fonollosa, Enric Monte-Moreno |
LREC | 1 |
| 2007 | Smooth Bilingual N-Gram Translation
Holger Schwenk, Marta R. Costa-jussà, José A. R. Fonollosa |
EMNLP-CoNLL | 2 |
| 2006 | Statistical Machine Reordering
Marta R. Costa-jussà, José A. R. Fonollosa |
EMNLP | 1 |
| 2006 | Machine Translation System Development Based on Human LikenessabstractWe present a novel approach for parameter adjustment in empirical machine translation systems. Instead of relying on a single evaluation metric, or in an ad-hoc linear combination of metrics, our method works over metric combinations with maximum descriptive power, aiming to maximise the Human Likeness of the automatic translations. We apply it to the problem of optimising decoding stage parameters of a state- of-the-art Statistical machine translation system. By means of a rigorous manual evaluation, we show how our methodology provides more reliable and robust system configurations than a tuning strategy based on the BLEU metric alone. Patrik Lambert, Jesús Giménez, Marta R. Costa-jussà, Enrique Amigó, Rafael E. Banchs, Lluís Màrquez, José A. R. Fonollosa |
SLT | 3 |
| 2006 | N-gram-based Machine TranslationabstractThis article describes in detail an n-gram approach to statistical machine translation. This approach consists of a log-linear combination of a translation model based on n-grams of bilingual units, which are referred to as tuples, along with four specific feature functions. Translation performance, which happens to be in the state of the art, is demonstrated with Spanish-to-English and English-to-Spanish translations of the European Parliament Plenary Sessions (EPPS). José B. Mariño, Rafael E. Banchs, Josep Maria Crego, Adrià de Gispert, Patrik Lambert, José A. R. Fonollosa, Marta R. Costa-jussà |
Comput. Linguistics | 7 |
| 2005 | Bilingual N-gram Statistical Machine TranslationabstractThis paper describes a statistical machine translation system that uses a translation model which is based on bilingual n-grams. When this translation model is log-linearly combined with four specific feature functions, state of the art translations are achieved for Spanish-to-English and English-to-Spanish translation tasks. Some specific results obtained for the EPPS (European Parliament Plenary Sessions) data are presented and discussed. Finally, future research issues are depicted. José B. Mariño, Rafael E. Banchs, Josep Maria Crego, Adrià de Gispert, Patrik Lambert, José A. R. Fonollosa, Marta R. Costa-jussà |
MTSummit | 7 |