Carlos Escolano

dblp:51/7736 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0001-6657-673XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Linguistic Knowledge-Infused Fine-Tuning for Mitigating Gender Bias in Machine Translation
abstract
Large Language Models (LLMs) achieve strong performance in machine translation (MT) but often encode gender bias, particularly when translating from non-gendered into gendered languages. This paper introduces a fine-tuning strategy to mitigate such bias in English-Spanish and English-Catalan translation. Using parameter-efficient LoRA fine-tuning, we apply linguistic knowledge infusion—a reasoning-based method that trains models to identify gendered referents and syntactic cues before generating translations. Experiments with Mistral–7B and Salamandrata–7B on MT-GenEval show that linguistically infused models improve gender accuracy by 15 percentage points and reduce gender gaps by 27 points in English-Spanish translation, with comparable trends for Catalan. Gains are strongest for Mistral, suggesting that explicit linguistic reasoning particularly benefits general-purpose LLMs. Overall, these results demonstrate that structured linguistic priors can enhance fairness and referential consistency in multilingual machine translation.
Ernesto Garcia-Estrada, Audrey Mash, Carlos Escolano, Maite Melero, Christine Basta
LREC3
2026 Optimizing Multilingual LLMs via Federated Learning: A Study of Client Language Composition
abstract
Federated Learning (FL) of Large Language Models (LLMs) in multilingual environments presents significant challenges stemming from heterogeneous language distributions across clients and disparities in language resource availability. To address these challenges, we extended the FederatedScope-LLM framework to support multilingual instruction-tuning experiments with LLMs. We also introduced a novel client-specific early stopping mechanism, Local Dynamic Early Stopping (LDES-FL), which allows clients to pause and resume local training based on client-side validation performance, enhancing training efficiency and sustainability. Through a series of experiments, we studied how client language composition — from fully monolingual to increasingly multilingual clients — affects multilingual quality, fairness and training cost. Monolingual local fine-tuning remains the most effective for single-language specialization, whereas federated training is better suited to learning a single balanced multilingual model. In FL, increasing within-client multilinguality leads to stronger and fairer global models, narrows the gap to centralized multilingual fine-tuning, and yields the largest gains for lower-resource languages, albeit at the cost of more optimization steps. Overall, our results identify client language composition as a key design variable in multilingual FL, shaping performance, fairness and efficiency.
Aleix Sant, Jordi Luque, Carlos Escolano
LREC3
2025 Investigating the translation capabilities of Large Language Models trained on parallel data only
abstract
In recent years, Large Language Models (LLMs) have demonstrated exceptional proficiency across a broad spectrum of Natural Language Processing (NLP) tasks, including Machine Translation. However, previous methods predominantly relied on iterative processes such as instruction fine-tuning or continual pre-training, leaving unexplored the challenges of training LLMs solely on parallel data. In this work, we introduce Plume (Parallel Language Model), a collection of three 2B LLMs featuring varying vocabulary sizes (32k, 128k, and 256k) trained exclusively on Catalan-centric parallel examples. These models perform comparably to previous encoder-decoder architectures on 16 supervised translation directions and 56 zero-shot ones. Utilizing this set of models, we conduct a thorough investigation into the translation capabilities of LLMs, probing their performance, the role of vocabulary size, the impact of the different elements of the prompt, and their cross-lingual representation space. We find that larger vocabulary sizes improve zero-shot performance and that different layers specialize in distinct aspects of the prompt, such as language-specific tags. We further show that as the vocabulary size grows, a larger number of attention heads can be pruned with minimal loss in translation quality, achieving a reduction of over 64.7% in attention heads.
Javier García Gilabert, Carlos Escolano, Aleix Sant, Francesca de Luca Fornaciari, Audrey Mash, Xixian Liao, Maite Melero
MTSummit (1)2
2025 Culture-aware machine translation: the case study of low-resource language pair Catalan-Chinese
abstract
High-quality machine translation requires datasets that not only ensure linguistic accuracy but also capture regional and cultural nuances. While many existing benchmarks, such as FLORES-200, rely on English as a pivot language, this approach can overlook the specificity of direct language pairs, particularly for underrepresented combinations like Catalan-Chinese. In this study, we demonstrate that even with a relatively small dataset of approximately 1,000 sentences, we can significantly improve MT localization. To this end, we introduce a dataset specifically designed to enhance Catalan-to-Chinese translation by prioritizing regionally and culturally specific topics. Unlike pivot-based datasets, our data source ensures a more faithful representation of Catalan linguistic and cultural elements, leading to more accurate translations of local terms and expressions. Using this dataset, we demonstrate better performance over the English-pivot FLORES-200 dev set and achieve competitive results on the FLORES-200 devtest set when evaluated with neural-based metrics. We release this dataset as both a human-preference resource and a benchmark for Catalan-Chinese translation. Additionally, we include Spanish translations for each sentence, facilitating extensions to Spanish-Chinese translation tasks.
Xixian Liao, Carlos Escolano, Audrey Mash, Francesca de Luca Fornaciari, Javier García Gilabert, Miguel Claramunt Argote, Ella Bohman, Maite Melero
MTSummit (1)2
2024 Unmasking Biases: Exploring Gender Bias in English-Catalan Machine Translation through Tokenization Analysis and Novel Dataset
abstract
This paper presents a comprehensive evaluation of gender bias in English-Catalan machine translation, encompassing the creation of a novel language resource and an analysis of translation quality across four different tokenization models. The study introduces a new dataset derived from the MuST-SHE corpus, focusing on gender-neutral terms that necessitate gendered translations in Catalan. The results reveal noteworthy gender bias across all translation models, with a consistent preference for masculine forms. Notably, the study finds that when context is available, BPE and Sentencepiece Unigram tokenization methods outperform others, achieving higher accuracy in gender translation. However, when no context is provided, Morfessor outputs more feminine forms than other tokenization methods, albeit still a small percentage. The study also reflects that stereotypes present in the data are amplified in the translation output. Ultimately, this work serves as a valuable resource for addressing and mitigating gender bias in machine translation, emphasizing the need for improved awareness and sensitivity to gender issues in natural language processing applications.
Audrey Mash, Carlos Escolano, Aleix Sant, Maite Melero, Francesca de Luca Fornaciari
LREC/COLING2
2024 ReSeTOX: Re-learning attention weights for toxicity mitigation in machine translation
abstract
Our proposed method, RESETOX (REdoSEarch if TOXic), addresses the issue ofNeural Machine Translation (NMT) gener-ating translation outputs that contain toxicwords not present in the input. The ob-jective is to mitigate the introduction oftoxic language without the need for re-training. In the case of identified addedtoxicity during the inference process, RE-SETOX dynamically adjusts the key-valueself-attention weights and re-evaluates thebeam search hypotheses. Experimental re-sults demonstrate that RESETOX achievesa remarkable 57% reduction in added tox-icity while maintaining an average trans-lation quality of 99.5% across 164 lan-guages. Our code is available at: https://github.com
Javier García Gilabert, Carlos Escolano, Marta R. Costa-jussà
EAMT (1)2
2023 Analysis of Acoustic information in End-to-End Spoken Language Translation
abstract
End-to-End Transformer-based models are the most popular approach for Spoken Language Translation (SLT). While obtaining state-of-the-art results, we are still far from understanding how these models extract acoustic information from the data and how they are transformed into semantic representations. In this paper, we seek to provide a better understanding of the flow of acoustic information along speech-to-text translation models. By means of the Speaker Classification and Spectrogram Reconstruction tasks, this study (i) interprets the main role of the encoder with respect to the acoustic features, (ii) highlights the importance of the acoustic information throughout the model and its transfer between encoder and decoder, and (iii) reveals the significant effect of downsampling convolutional layers for learning acoustic features. (iv) Finally, we also observe the existence of a strong correlation between the semantic domain and the speakers’ labels in MuST-C.
Gerard Sant, Carlos Escolano
INTERSPEECH2
2022 Interpreting Gender Bias in Neural Machine Translation: Multilingual Architecture Matters
abstract
Multilingual neural machine translation architectures mainly differ in the number of sharing modules and parameters applied among languages. In this paper, and from an algorithmic perspective, we explore whether the chosen architecture, when trained with the same data, influences the level of gender bias. Experiments conducted in three language pairs show that language-specific encoder-decoders exhibit less bias than the shared architecture. We propose two methods for interpreting and studying gender bias in machine translation based on source embeddings and attention. Our analysis shows that, in the language-specific case, the embeddings encode more gender information, and their attention is more diverted. Both behaviors help in mitigating gender bias.
Marta R. Costa-jussà, Carlos Escolano, Christine Basta, Javier Ferrando, Roser Batlle, Ksenia Kharitonova
AAAI2
2022 Towards Opening the Black Box of Neural Machine Translation: Source and Target Interpretations of the Transformer
abstract
In Neural Machine Translation (NMT), each token prediction is conditioned on the source sentence and the target prefix (what has been previously translated at a decoding step).However, previous work on interpretability in NMT has mainly focused solely on source sentence tokens' attributions.Therefore, we lack a full understanding of the influences of every input token (source sentence and target prefix) in the model predictions.In this work, we propose an interpretability method that tracks input tokens' attributions for both contexts.Our method, which can be extended to any encoder-decoder Transformer-based model, allows us to better comprehend the inner workings of current NMT models.We apply the proposed method to both bilingual and multilingual Transformers and present insights into their behaviour.
Javier Ferrando, Gerard I. Gállego, Belen Alastruey, Carlos Escolano, Marta R. Costa-jussà
EMNLP4
2022 Multilingual Machine Translation: Deep Analysis of Language-Specific Encoder-Decoders
abstract
State-of-the-art multilingual machine translation relies on a shared encoder-decoder. In this paper, we propose an alternative approach based on language-specific encoder-decoders, which can be easily extended to new languages by learning their corresponding modules. To establish a common interlingua representation, we simultaneously train N initial languages. Our experiments show that the proposed approach improves over the shared encoder-decoder for the initial languages and when adding new languages, without the need to retrain the remaining modules. All in all, our work closes the gap between shared and language-specific encoder-decoders, advancing toward modular multilingual machine translation systems that can be flexibly extended in lifelong learning settings.
Carlos Escolano, Marta R. Costa-jussà, José A. R. Fonollosa
J. Artif. Intell. Res.1
2021 Enabling Zero-Shot Multilingual Spoken Language Translation with Language-Specific Encoders and Decoders
abstract
Current end-to-end approaches to Spoken Language Translation (SLT) rely on limited training resources, especially for multilingual settings. On the other hand, Multilingual Neural Machine Translation (MultiNMT) approaches rely on higher-quality and more massive data sets. Our proposed method extends a MultiNMT architecture based on language-specific encoders-decoders to the task of Multilingual SLT (Multi-SLT). Our method entirely eliminates the dependency from MultiSLT data and it is able to translate while training only on ASR and MultiNMT data. Our experiments on four different languages show that coupling the speech encoder to the MultiNMT architecture produces similar quality translations compared to a bilingual baseline (±0.2 BLEU) while effectively allowing for zero-shot MultiSLT. Additionally, we propose using an Adapter module for coupling the speech inputs. This Adapter module produces consistent improvements up to +6 BLEU points on the proposed architecture and +1 BLEU point on the end-to-end baseline.
Carlos Escolano, Marta R. Costa-jussà, José A. R. Fonollosa, Carlos Segura
ASRU1
2021 Multilingual Machine Translation: Closing the Gap between Shared and Language-specific Encoder-Decoders
abstract
Carlos Escolano, Marta R. Costa-jussà, José A. R. Fonollosa, Mikel Artetxe. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Carlos Escolano, Marta R. Costa-jussà, José A. R. Fonollosa, Mikel Artetxe
EACL1
2021 From bilingual to multilingual neural-based machine translation by incremental training
abstract
Abstract A common intermediate language representation in neural machine translation can be used to extend bilingual systems by incremental training. We propose a new architecture based on introducing an interlingual loss as an additional training objective. By adding and forcing this interlingual loss, we can train multiple encoders and decoders for each language, sharing among them a common intermediate representation. Translation results on the low‐resource tasks (Turkish‐English and Kazakh‐English tasks) show a BLEU improvement of up to 2.8 points. However, results on a larger dataset (Russian‐English and Kazakh‐English) show BLEU losses of a similar amount. While our system provides improvements only for the low‐resource tasks in terms of translation quality, our system is capable of quickly deploying new language pairs without the need to retrain the rest of the system, which may be a game changer in some situations. Specifically, what is most relevant regarding our architecture is that it is capable of: reducing the number of production systems, with respect to the number of languages, from quadratic to linear; incrementally adding a new language to the system without retraining the languages already there; and allowing for translations from the new language to all the others present in the system.
Carlos Escolano, Marta R. Costa-jussà, José A. R. Fonollosa
J. Assoc. Inf. Sci. Technol.1
2019 Chinese-Catalan: A Neural Machine Translation Approach Based on Pivoting and Attention Mechanisms
abstract
This article innovatively addresses machine translation from Chinese to Catalan using neural pivot strategies trained without any direct parallel data. The Catalan language is very similar to Spanish from a linguistic point of view, which motivates the use of Spanish as pivot language. Regarding neural architecture, we are using the latest state-of-the-art, which is the Transformer model, only based on attention mechanisms. Additionally, this work provides new resources to the community, which consists of a human-developed gold standard of 4,000 sentences between Catalan and Chinese and all the others United Nations official languages (Arabic, English, French, Russian, and Spanish). Results show that the standard pseudo-corpus or synthetic pivot approach performs better than cascade.
Marta R. Costa-jussà, Noe Casas, Carlos Escolano, José A. R. Fonollosa
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2012 A Telepresence Mobile Robot Controlled With a Noninvasive Brain-Computer Interface
abstract
This paper reports an electroencephalogram-based brain-actuated telepresence system to provide a user with presence in remote environments through a mobile robot, with access to the Internet. This system relies on a P300-based brain-computer interface (BCI) and a mobile robot with autonomous navigation and camera orientation capabilities. The shared-control strategy is built by the BCI decoding of task-related orders (selection of visible target destinations or exploration areas), which can be autonomously executed by the robot. The system was evaluated using five healthy participants in two consecutive steps: 1) screening and training of participants and 2) preestablished navigation and visual exploration telepresence tasks. On the basis of the results, the following evaluation studies are reported: 1) technical evaluation of the device and its main functionalities and 2) the users' behavior study. The overall result was that all participants were able to complete the designed tasks, reporting no failures, which shows the robustness of the system and its feasibility to solve tasks in real settings where joint navigation and visual exploration were needed. Furthermore, the participants showed great adaptation to the telepresence system.
Carlos Escolano, Javier Mauricio Antelis, Javier Minguez
IEEE Trans. Syst. Man Cybern. Part B1
2009 Human brain-teleoperated robot between remote places
abstract
This paper describes an EEG-based human brain-actuated robotic system, which allows performing navigation and visual exploration tasks between remote places via Internet, using only brain activity. In operation, two teleoperation modes can be combined: robot navigation and camera exploration. In both modes, the user faces a real-time video captured by the robot camera merged with augmented reality items. In this representation, the user concentrates on a target area to navigate to or visually explore; then, a visual stimulation process elicits the neurological phenomenon that enables the brain-computer system to decode the intentions of the user. In the navigation mode, the target destination is transferred to the autonomous navigation system, which drives the robot to the desired place while avoiding collisions with the obstacles detected by the laser scanner. In the camera mode, the camera is aligned with the target area to perform an active visual exploration of the remote scenario. In June 2008, within the framework of the experimental methodology, five healthy subjects performed pre-established navigation and visual exploration tasks for one week between two cities separated by 260 km. On the basis of the results, a technical evaluation of the device and its main functionalities is reported. The overall result is that all the subjects were able to successfully solve all the tasks reporting no failures, showing a high robustness of the system.
Carlos Escolano, Javier Mauricio Antelis, Javier Minguez
ICRA1