VLDB 2026 Research / reviewers in the wild / expert
Loïc Barrault
dblp:86/7823
· DBLP profile ↗
27ranked-venue papers
3as first author
11since 2021 · last 2025
0000-0002-0634-6147ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MEXMA: Token-level objectives improve sentence representationsabstractCross-lingual sentence encoders (CLSE) create fixed-size sentence representations with aligned translations.Current pre-trained CLSE approaches use sentence-level objectives only.This can lead to loss of information, especially for tokens, which then degrades the sentence representation.We propose MEXMA, a novel approach that integrates both sentence-level and token-level objectives.The sentence representation in one language is used to predict masked tokens in another language, with both the sentence representation and all tokens directly updating the encoder.We show that adding token-level objectives greatly improves the sentence representation quality across several tasks.Our approach outperforms current pre-trained cross-lingual sentence encoders on bitext mining as well as several downstream tasks.We also analyse the information encoded in our tokens, and how the sentence representation is built from them. João Maria Janeiro, Benjamin Piwowarski, Patrick Gallinari, Loïc Barrault |
ACL (1) | 4 |
| 2025 | Mixture of Languages: Improved Multilingual Encoders Through Language GroupingabstractJoão Maria Janeiro, Belen Alastruey, Francisco Massa, Maha Elbayad, Benjamin Piwowarski, Patrick Gallinari, Loic Barrault. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. João Maria Janeiro, Belen Alastruey, Francisco Massa, Maha Elbayad, Benjamin Piwowarski, Patrick Gallinari, Loïc Barrault |
EMNLP | 7 |
| 2023 | Detecting and Mitigating Hallucinations in Machine Translation: Model Internal Workings Alone Do Well, Sentence Similarity Even BetterabstractWhile the problem of hallucinations in neural machine translation has long been recognized, so far the progress on its alleviation is very little.Indeed, recently it turned out that without artificially encouraging models to hallucinate, previously existing methods fall short and even the standard sequence log-probability is more informative.It means that internal characteristics of the model can give much more information than we expect, and before using external models and measures, we first need to ask: how far can we go if we use nothing but the translation model itself ?We propose to use a method that evaluates the percentage of the source contribution to a generated translation.Intuitively, hallucinations are translations "detached" from the source, hence they can be identified by low source contribution.This method improves detection accuracy for the most severe hallucinations by a factor of 2 and is able to alleviate hallucinations at test time on par with the previous best approach that relies on external models.Next, if we move away from internal model characteristics and allow external tools, we show that using sentence similarity from cross-lingual embeddings further improves these results.We release the code of our experiments.1 David Dale, Elena Voita, Loïc Barrault, Marta R. Costa-jussà |
ACL (1) | 3 |
| 2023 | FrameBERT: Conceptual Metaphor Detection with Frame Embedding LearningabstractIn this paper, we propose FrameBERT, a RoBERTa-based model that can explicitly learn and incorporate FrameNet Embeddings for concept-level metaphor detection.FrameBERT not only achieves better or comparable performance to the state-of-the-art, but also is more explainable and interpretable compared to existing models, attributing to its ability of accounting for external knowledge of FrameNet. Yucheng Li 0001, Chenghua Lin 0002, Frank Guerin, Loïc Barrault |
EACL | 5 |
| 2023 | Metaphor Detection with Effective Context DenoisingabstractWe propose a novel RoBERTa-based model, RoPPT, which introduces a target-oriented parse tree structure in metaphor detection.Compared to existing models, RoPPT focuses on semantically relevant information and achieves the state-of-the-art on several main metaphor datasets.We also compare our approach against several popular denoising and pruning methods, demonstrating the effectiveness of our approach in context denoising.Our code and dataset can be found at https: //github.com/MajiBear000/RoPPT. Yucheng Li 0001, Chenghua Lin 0002, Loïc Barrault, Frank Guerin |
EACL | 4 |
| 2023 | HalOmi: A Manually Annotated Benchmark for Multilingual Hallucination and Omission Detection in Machine TranslationabstractDavid Dale, Elena Voita, Janice Lam, Prangthip Hansanti, Christophe Ropers, Elahe Kalbassi, Cynthia Gao, Loic Barrault, Marta Costa-jussà. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. David Dale, Elena Voita, Janice Lam, Prangthip Hansanti, Christophe Ropers, Elahe Kalbassi, Cynthia Gao, Loïc Barrault, Marta R. Costa-jussà |
EMNLP | 8 |
| 2023 | We Need to Talk About Classification Evaluation Metrics in NLPabstractPeter Vickers, Loic Barrault, Emilio Monti, Nikolaos Aletras. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Peter Vickers, Loïc Barrault, Emilio Monti, Nikolaos Aletras |
IJCNLP (1) | 2 |
| 2023 | Towards lifelong human assisted speaker diarization
Meysam Shamsi, Anthony Larcher, Loïc Barrault, Sylvain Meignier, Yevhenii Prokopalo, Marie Tahon, Ambuj Mehrish, Simon Petit-Renaud, Olivier Galibert, Samuel Gaist, André Anjos, Sébastien Marcel, Marta R. Costa-jussà |
Comput. Speech Lang. | 3 |
| 2022 | Controlling Extra-Textual Attributes about Dialogue Participants: A Case Study of English-to-Polish Neural Machine TranslationabstractUnlike English, morphologically rich languages can reveal characteristics of speakers or their conversational partners, such as gender and number, via pronouns, morphological endings of words and syntax. When translating from English to such languages, a machine translation model needs to opt for a certain interpretation of textual context, which may lead to serious translation errors if extra-textual information is unavailable. We investigate this challenge in the English-to-Polish language direction. We focus on the underresearched problem of utilising external metadata in automatic translation of TV dialogue, proposing a case study where a wide range of approaches for controlling attributes in translation is employed in a multi-attribute scenario. The best model achieves an improvement of +5.81 chrF++/+6.03 BLEU, with other models achieving competitive performance. We additionally contribute a novel attribute-annotated dataset of Polish TV dialogue and a morphological analysis script used to evaluate attribute control in models. Sebastian T. Vincent, Loïc Barrault, Carolina Scarton |
EAMT | 2 |
| 2022 | Speech Resources in the Tamasheq LanguageabstractIn this paper we present two datasets for Tamasheq, a developing language mainly spoken in Mali and Niger. These two datasets were made available for the IWSLT 2022 low-resource speech translation track, and they consist of collections of radio recordings from daily broadcast news in Niger (Studio Kalangou) and Mali (Studio Tamani). We share (i) a massive amount of unlabeled audio data (671 hours) in five languages: French from Niger, Fulfulde, Hausa, Tamasheq and Zarma, and (ii) a smaller 17 hours parallel corpus of audio recordings in Tamasheq, with utterance-level translations in the French language. All this data is shared under the Creative Commons BY-NC-ND 3.0 license. We hope these resources will inspire the speech community to develop and benchmark models using the Tamasheq language. Marcely Zanon Boito, Fethi Bougares, Florentin Barbier, Souhir Gahbiche-Braham, Loïc Barrault, Mickael Rouvier, Yannick Estève |
LREC | 5 |
| 2021 | Active Learning by Acquiring Contrastive ExamplesabstractCommon acquisition functions for active learning use either uncertainty or diversity sampling, aiming to select difficult and diverse data points from the pool of unlabeled data, respectively.In this work, leveraging the best of both worlds, we propose an acquisition function that opts for selecting contrastive examples, i.e. data points that are similar in the model feature space and yet the model outputs maximally different predictive likelihoods.We compare our approach, CAL (Contrastive Active Learning), with a diverse set of acquisition functions in four natural language understanding tasks and seven datasets.Our experiments show that CAL performs consistently better or equal than the best performing baseline across all tasks, on both in-domain and out-of-domain data.We also conduct an extensive ablation study of our method and we further analyze all actively acquired datasets showing that CAL achieves a better trade-off between uncertainty and diversity compared to other strategies. Aikaterini Margatina, Giorgos Vernikos, Loïc Barrault, Nikolaos Aletras |
EMNLP (1) | 3 |
| 2020 | Simultaneous Machine Translation with Visual ContextabstractSimultaneous machine translation (SiMT) aims to translate a continuous input text stream into another language with the lowest latency and highest quality possible.The translation thus has to start with an incomplete source text, which is read progressively, creating the need for anticipation.In this paper, we seek to understand whether the addition of visual information can compensate for the missing source context.To this end, we analyse the impact of different multimodal approaches and visual features on state-of-the-art SiMT frameworks.Our results show that visual context is helpful and that visually-grounded models based on explicit object region information are much better than commonly used global features, reaching up to 3 BLEU points improvement under low latency scenarios.Our qualitative analysis illustrates cases where only the multimodal systems are able to translate correctly from English into gender-marked languages, as well as deal with differences in word order, such as adjective-noun placement between English and French. Ozan Caglayan, Julia Ive, Veneta Haralampieva, Pranava Swaroop Madhyastha, Loïc Barrault, Lucia Specia |
EMNLP (1) | 5 |
| 2020 | Evaluation of Lifelong Learning SystemsabstractCurrent intelligent systems need the expensive support of machine learning experts to sustain their performance level when used on a daily basis. To reduce this cost, i.e. remaining free from any machine learning expert, it is reasonable to implement lifelong (or continuous) learning intelligent systems that will continuously adapt their model when facing changing execution conditions. In this work, the systems are allowed to refer to human domain experts who can provide the system with relevant knowledge about the task. Nowadays, the fast growth of lifelong learning systems development rises the question of their evaluation. In this article we propose a generic evaluation methodology for the specific case of lifelong learning systems. Two steps will be considered. First, the evaluation of human-assisted learning (including active and/or interactive learning) outside the context of lifelong learning. Second, the system evaluation across time, with propositions of how a lifelong learning intelligent system should be evaluated when including human assisted learning or not. Yevhenii Prokopalo, Sylvain Meignier, Olivier Galibert, Loïc Barrault, Anthony Larcher |
LREC | 4 |
| 2020 | Addressing data sparsity for neural machine translation between morphologically rich languages
Mercedes García-Martínez, Walid Aransa, Fethi Bougares, Loïc Barrault |
Mach. Transl. | 4 |
| 2019 | Multimodal Grounding for Sequence-to-sequence Speech RecognitionabstractHumans are capable of processing speech by making use of multiple sensory modalities. For example, the environment where a conversation takes place generally provides semantic and/or acoustic context that helps us to resolve ambiguities or to recall named entities. Motivated by this, there have been many works studying the integration of visual information into the speech recognition pipeline. Specifically, in our previous work, we propose a multistep visual adaptive training approach which improves the accuracy of an audio-based Automatic Speech Recognition (ASR) system. This approach, however, is not end-to-end as it requires fine-tuning the whole model with an adaptation layer. In this paper, we propose novel end-to-end multimodal ASR systems and compare them to the adaptive approach by using a range of visual representations obtained from state-of-the-art convolutional neural networks. We show that adaptive training is effective for S2S models leading to an absolute improvement of 1.4% in word error rate. As for the end-to-end systems, although they perform better than baseline, the improvements are slightly less than adaptive training, 0.8 absolute WER reduction in single-best models. Using ensemble decoding, end-to-end models reach a WER of 15% which is the lowest score among all systems. Ozan Caglayan, Ramon Sanabria, Shruti Palaskar, Loïc Barrault, Florian Metze |
ICASSP | 4 |
| 2018 | What you can cram into a single \$&!#* vector: Probing sentence embeddings for linguistic propertiesabstractAlthough much effort has recently been devoted to training high-quality sentence embeddings, we still have a poor understanding of what they are capturing."Downstream" tasks, often based on sentence classification, are commonly used to evaluate the quality of sentence representations.The complexity of the tasks makes it however difficult to infer what kind of information is present in the representations.We introduce here 10 probing tasks designed to capture simple linguistic features of sentences, and we use them to study embeddings generated by three different encoders trained in eight distinct ways, uncovering intriguing properties of both encoders and training methods. Alexis Conneau, Germán Kruszewski, Guillaume Lample, Loïc Barrault, Marco Baroni |
ACL (1) | 4 |
| 2017 | Very Deep Convolutional Networks for Text ClassificationabstractThe dominant approach for many NLP tasks are recurrent neural networks, in particular LSTMs, and convolutional neural networks. However, these architectures are rather shallow in comparison to the deep convolutional networks which have pushed the state-of-the-art in computer vision. We present a new architecture (VDCNN) for text processing which operates directly at the character level and uses only small convolutions and pooling operations. We are able to show that the performance of this model increases with depth: using up to 29 convolutional layers, we report improvements over the state-of-the-art on several public text classification tasks. To the best of our knowledge, this is the first time that very deep convolutional nets have been applied to text processing. Alexis Conneau, Holger Schwenk, Loïc Barrault, Yann LeCun |
EACL (1) | 3 |
| 2017 | Supervised Learning of Universal Sentence Representations from Natural Language Inference DataabstractMany modern NLP systems rely on word embeddings, previously trained in an unsupervised manner on large corpora, as base features.Efforts to obtain embeddings for larger chunks of text, such as sentences, have however not been so successful.Several attempts at learning unsupervised representations of sentences have not reached satisfactory enough performance to be widely adopted.In this paper, we show how universal sentence representations trained using the supervised data of the Stanford Natural Language Inference datasets can consistently outperform unsupervised methods like SkipThought vectors (Kiros et al., 2015) on a wide range of transfer tasks.Much like how computer vision uses ImageNet to obtain features, which can then be transferred to other tasks, our work tends to indicate the suitability of natural language inference for transfer learning to other NLP tasks.Our encoder is publicly available 1 . Alexis Conneau, Douwe Kiela, Holger Schwenk, Loïc Barrault, Antoine Bordes |
EMNLP | 4 |
| 2017 | Introduction to the special issue on deep learning approaches for machine translation
Marta R. Costa-jussà, Alexandre Allauzen, Loïc Barrault, Kyunghyun Cho, Holger Schwenk |
Comput. Speech Lang. | 3 |
| 2016 | Building and using multimodal comparable corpora for machine translationabstractAbstract In recent decades, statistical approaches have significantly advanced the development of machine translation systems. However, the applicability of these methods directly depends on the availability of very large quantities of parallel data. Recent works have demonstrated that a comparable corpus can compensate for the shortage of parallel corpora. In this paper, we propose an alternative to comparable corpora containing text documents as resources for extracting parallel data: a multimodal comparable corpus with audio documents in source language and text document in target language, built fromEuronewsandTEDweb sites. The audio is transcribed by an automatic speech recognition system, and translated with a baseline statistical machine translation system. We then use information retrieval in a large text corpus in the target language in order to extract parallel sentences/phrases. We evaluate the quality of the extracted data on an English to French translation task and show significant improvements over a state-of-the-art baseline. Haithem Afli, Loïc Barrault, Holger Schwenk |
Nat. Lang. Eng. | 2 |
| 2015 | Continuous Adaptation to User Feedback for Statistical Machine TranslationabstractFrédéric Blain, Fethi Bougares, Amir Hazem, Loïc Barrault, Holger Schwenk. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Frédéric Blain, Fethi Bougares, Amir Hazem, Loïc Barrault, Holger Schwenk |
HLT-NAACL | 4 |
| 2014 | Translation project adaptation for MT-enhanced computer assisted translation
Mauro Cettolo, Nicola Bertoldi, Marcello Federico, Holger Schwenk, Loïc Barrault, Christophe Servan |
Mach. Transl. | 5 |
| 2013 | Multimodal Comparable Corpora as Resources for Extracting Parallel Data: Parallel Phrases Extraction
Haithem Afli, Loïc Barrault, Holger Schwenk |
IJCNLP | 2 |
| 2011 | Parametric Weighting of Parallel Data for Statistical Machine Translation
Kashif Shah, Loïc Barrault, Holger Schwenk |
IJCNLP | 2 |
| 2008 | Frame-based acoustic feature integration for speech understandingabstractWith the purpose of improving spoken language understanding (SLU) performance, a combination of different acoustic speech recognition (ASR) systems is proposed. State a posteriori probabilities obtained with systems using different acoustic feature sets are combined with log-linear interpolation. In order to perform a coherent combination of these probabilities, acoustic models must have the same topology (i.e. same set of states). For this purpose, a fast and efficient twin model training protocol is proposed. By a wise choice of acoustic feature sets and log-linear interpolation of their likelihood ratios, a substantial concept error rate (CER) reduction has been observed on the test part of the French MEDIA corpus. Loïc Barrault, Christophe Servan, Driss Matrouf, Georges Linarès, Renato De Mori |
ICASSP | 1 |
| 2006 | Characterizing Feature Variability in Automatic Speech Recognition SystemsabstractA method is described for predicting acoustic feature variability by analyzing the consensus and relative entropy of phoneme posterior probability distributions obtained with different acoustic models having the same type of observations. Variability prediction is used for diagnosis of automatic speech recognition (ASR) systems. When errors are likely to occur, different feature sets are considered for correcting recognition results. Experimental results are provided on the CH1 Italian portion of AURORA3 Loïc Barrault, Driss Matrouf, Renato De Mori, Roberto Gemello, Franco Mana |
ICASSP (5) | 1 |
| 2005 | Variability of automatic speech recognition systems using different featuresabstractInternational audience Loïc Barrault, Renato De Mori, Roberto Gemello, Franco Mana, Driss Matrouf |
INTERSPEECH | 1 |