VLDB 2026 Research / reviewers in the wild / expert
David Dale
dblp:293/7322
· DBLP profile ↗
12ranked-venue papers
3as first author
12since 2021 · last 2025
0000-0003-2045-6833ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Less Mature is More Adaptable for Sentence-level Language ModelingabstractThis work investigates sentence-level models (i.e., models that operate at the sentence-level) to study how sentence representations from various encoders influence downstream task performance, and which syntactic, semantic, and discourse-level properties are essential for strong performance.Our experiments encompass encoders with diverse training regimes and pretraining domains, as well as various pooling strategies applied to multi-sentence input tasks (including sentence ordering, sentiment classification, and natural language inference) requiring coarse-to-fine-grained reasoning.We find that "less mature" representations (e.g., mean-pooled representations from BERT's first or last layer, or representations from encoders with limited fine-tuning) exhibit greater generalizability and adaptability to downstream tasks compared to representations from extensively fine-tuned models (e.g.,, SBERT or Sim-CSE).These findings are consistent across different pretraining seed initializations for BERT.Our probing analysis reveals that syntactic and discourse-level properties are stronger indicators of downstream performance than MTEB scores or decodability.Furthermore, the data and time efficiency of sentence-level models, often outperforming token-level models, underscores their potential for future research. Abhilasha Sancheti, David Dale, Artyom Kozhevnikov, Maha Elbayad |
ACL (1) | 2 |
| 2025 | Improving Language and Modality Transfer in Translation by Character-level ModelingabstractCurrent translation systems, despite being highly multilingual, cover only 5% of the world’s languages. Expanding language coverage to the long-tail of low-resource languages requires data-efficient methods that rely on cross-lingual and cross-modal knowledge transfer. To this end, we propose a character-based approach to improve adaptability to new languages and modalities. Our method leverages SONAR, a multilingual fixed-size embedding space with different modules for encoding and decoding. We use a teacher-student approach with parallel translation data to obtain a character-level encoder. Then, using ASR data, we train a lightweight adapter to connect a massively multilingual CTC ASR model (MMS), to the character-level encoder, potentially enabling speech translation from 1,000+ languages. Experimental results in text translation for 75 languages on FLORES+ demonstrate that our character-based approach can achieve better language transfer than traditional subword-based models, especially outperforming them in low-resource settings, and demonstrating better zero-shot generalizability to unseen languages. Our speech adaptation, maximizing knowledge transfer from the text modality, achieves state-of-the-art results in speech-to-text translation on the FLEURS benchmark on 33 languages, surpassing previous supervised and cascade models, albeit being a zero-shot model with minimal supervision from ASR data. Ioannis Tsiamas, David Dale, Marta R. Costa-jussà |
ACL (1) | 2 |
| 2025 | BOUQuET : dataset, Benchmark and Open initiative for Universal Quality Evaluation in TranslationabstractPierre Andrews, Mikel Artetxe, Mariano Coria Meglioli, Marta R. Costa-jussà, Joe Chuang, David Dale, Mark Duppenthaler, Nathanial Paul Ekberg, Cynthia Gao, Daniel Edward Licht, Jean Maillard, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Eduardo Sánchez, Ioannis Tsiamas, Arina Turkatenko, Albert Ventayol-Boada, Shireen Yates. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Pierre Andrews, Mikel Artetxe, Mariano Coria Meglioli, Marta R. Costa-jussà, Joe Chuang, David Dale, Mark Duppenthaler, Nathanial Paul Ekberg, Cynthia Gao, Daniel Edward Licht, Jean Maillard, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Eduardo Sánchez, Ioannis Tsiamas, Arina Turkatenko, Albert Ventayol-Boada, Shireen Yates |
EMNLP | 6 |
| 2024 | SpeechAlign: A Framework for Speech Translation Alignment EvaluationabstractSpeech-to-Speech and Speech-to-Text translation are currently dynamic areas of research. In our commitment to advance these fields, we present SpeechAlign, a framework designed to evaluate the underexplored field of source-target alignment in speech models. The SpeechAlign framework has two core components. First, to tackle the absence of suitable evaluation datasets, we introduce the Speech Gold Alignment dataset, built upon a English-German text translation gold alignment dataset. Secondly, we introduce two novel metrics, Speech Alignment Error Rate (SAER) and Time-weighted Speech Alignment Error Rate (TW-SAER), which enable the evaluation of alignment quality within speech models. While the former gives equal importance to each word, the latter assigns weights based on the length of the words in the speech signal. By publishing SpeechAlign we provide an accessible evaluation framework for model assessment, and we employ it to benchmark open-source Speech Translation models. In doing so, we contribute to the ongoing research progress within the fields of Speech-to-Speech and Speech-to-Text translation. Belen Alastruey, Aleix Sant, Gerard I. Gállego, David Dale, Marta R. Costa-jussà |
LREC/COLING | 4 |
| 2024 | Added Toxicity Mitigation at Inference Time for Multimodal and Massively Multilingual TranslationabstractMachine translation models sometimes lead to added toxicity: translated outputs may contain more toxic content that the original input. In this paper, we introduce MinTox, a novel pipeline to automatically identify and mitigate added toxicity at inference time, without further model training. MinTox leverages a multimodal (speech and text) toxicity classifier that can scale across languages.We demonstrate the capabilities of MinTox when applied to SEAMLESSM4T, a multi-modal and massively multilingual machine translation system. MinTox significantly reduces added toxicity: across all domains, modalities and language directions, 25% to95% of added toxicity is successfully filtered out, while preserving translation quality Marta R. Costa-jussà, David Dale, Maha Elbayad, Bokai Yu |
EAMT (1) | 2 |
| 2023 | Detecting and Mitigating Hallucinations in Machine Translation: Model Internal Workings Alone Do Well, Sentence Similarity Even BetterabstractWhile the problem of hallucinations in neural machine translation has long been recognized, so far the progress on its alleviation is very little.Indeed, recently it turned out that without artificially encouraging models to hallucinate, previously existing methods fall short and even the standard sequence log-probability is more informative.It means that internal characteristics of the model can give much more information than we expect, and before using external models and measures, we first need to ask: how far can we go if we use nothing but the translation model itself ?We propose to use a method that evaluates the percentage of the source contribution to a generated translation.Intuitively, hallucinations are translations "detached" from the source, hence they can be identified by low source contribution.This method improves detection accuracy for the most severe hallucinations by a factor of 2 and is able to alleviate hallucinations at test time on par with the previous best approach that relies on external models.Next, if we move away from internal model characteristics and allow external tools, we show that using sentence similarity from cross-lingual embeddings further improves these results.We release the code of our experiments.1 David Dale, Elena Voita, Loïc Barrault, Marta R. Costa-jussà |
ACL (1) | 1 |
| 2023 | HalOmi: A Manually Annotated Benchmark for Multilingual Hallucination and Omission Detection in Machine TranslationabstractDavid Dale, Elena Voita, Janice Lam, Prangthip Hansanti, Christophe Ropers, Elahe Kalbassi, Cynthia Gao, Loic Barrault, Marta Costa-jussà. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. David Dale, Elena Voita, Janice Lam, Prangthip Hansanti, Christophe Ropers, Elahe Kalbassi, Cynthia Gao, Loïc Barrault, Marta R. Costa-jussà |
EMNLP | 1 |
| 2023 | Exploring Methods for Cross-lingual Text Style Transfer: The Case of Text DetoxificationabstractDaryna Dementieva, Daniil Moskovskiy, David Dale, Alexander Panchenko. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Daryna Dementieva, Daniil Moskovskiy, David Dale, Alexander Panchenko |
IJCNLP (1) | 3 |
| 2023 | Don't Lose the Message While Paraphrasing: A Study on Content Preserving Style Transfer
Nikolay Babakov, David Dale, Ilya Gusev, Irina Krotova, Alexander Panchenko |
NLDB | 2 |
| 2022 | ParaDetox: Detoxification with Parallel DataabstractVarvara Logacheva, Daryna Dementieva, Sergey Ustyantsev, Daniil Moskovskiy, David Dale, Irina Krotova, Nikita Semenov, Alexander Panchenko. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Varvara Logacheva, Daryna Dementieva, Sergey Ustyantsev, Daniil Moskovskiy, David Dale, Irina Krotova, Nikita Semenov, Alexander Panchenko |
ACL (1) | 5 |
| 2022 | Studying the Role of Named Entities for Content Preservation in Text Style Transfer
Nikolay Babakov, David Dale, Varvara Logacheva, Irina Krotova, Alexander Panchenko |
NLDB | 2 |
| 2021 | Text Detoxification using Large Pre-trained Neural ModelsabstractDavid Dale, Anton Voronov, Daryna Dementieva, Varvara Logacheva, Olga Kozlova, Nikita Semenov, Alexander Panchenko. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. David Dale, Anton Voronov, Daryna Dementieva, Varvara Logacheva, Olga Kozlova, Nikita Semenov, Alexander Panchenko |
EMNLP (1) | 1 |