VLDB 2026 Research / reviewers in the wild / expert
Elizabeth Salesky
dblp:184/8920
· DBLP profile ↗
19ranked-venue papers
10as first author
11since 2021 · last 2024
0000-0001-6765-1447ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 9 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Evaluating the IWSLT2023 Speech Translation Tasks: Human Annotations, Automatic Metrics, and SegmentationabstractHuman evaluation is a critical component in machine translation system development and has received much attention in text translation research. However, little prior work exists on the topic of human evaluation for speech translation, which adds additional challenges such as noisy data and segmentation mismatches. We take the first steps to fill this gap by conducting a comprehensive human evaluation of the results of several shared tasks from the last International Workshop on Spoken Language Translation (IWSLT 2023). We propose an effective evaluation strategy based on automatic resegmentation and direct assessment with segment context. Our analysis revealed that: 1) the proposed evaluation strategy is robust and scores well-correlated with other types of human judgements; 2) automatic metrics are usually, but not always, well-correlated with direct assessment scores; and 3) COMET as a slightly stronger automatic metric than chrF, despite the segmentation noise introduced by the resegmentation step systems. We release the collected human-annotated data in order to encourage further investigation. Matthias Sperber, Ondrej Bojar, Barry Haddow, Dávid Javorský, Xutai Ma, Matteo Negri, Jan Niehues, Peter Polak, Elizabeth Salesky, Katsuhito Sudoh, Marco Turchi |
LREC/COLING | 9 |
| 2023 | Text Rendering Strategies for Pixel Language ModelsabstractPixel-based language models process text rendered as images, which allows them to handle any script, making them a promising approach to open vocabulary language modelling.However, recent approaches use text renderers that produce a large set of almost-equivalent input patches, which may prove sub-optimal for downstream tasks, due to redundancy in the input representations.In this paper, we investigate four approaches to rendering text in the PIXEL model (Rust et al., 2023), and find that simple character bigram rendering brings improved performance on sentence-level tasks without compromising performance on tokenlevel or multilingual tasks.This new rendering strategy also makes it possible to train a more compact model with only 22M parameters that performs on par with the original 86M parameter model.Our analyses show that character bigram rendering leads to a consistently better model but with an anisotropic patch embedding space, driven by a patch frequency bias, highlighting the connections between image patchand tokenization-based language models. Jonas F. Lotz, Elizabeth Salesky, Phillip Rust, Desmond Elliott |
EMNLP | 2 |
| 2023 | Multilingual Pixel Representations for Translation and Effective Cross-lingual TransferabstractWe introduce and demonstrate how to effectively train multilingual machine translation models with pixel representations.We experiment with two different data settings with a variety of language and script coverage, demonstrating improved performance compared to subword embeddings.We explore various properties of pixel representations such as parameter sharing within and across scripts to better understand where they lead to positive transfer.We observe that these properties not only enable seamless cross-lingual transfer to unseen scripts, but make pixel representations more data-efficient than alternatives such as vocabulary expansion.We hope this work contributes to more extensible multilingual models for all languages and scripts. Elizabeth Salesky, Neha Verma 0001, Philipp Koehn, Matt Post |
EMNLP | 1 |
| 2023 | A Holistic Cascade System, Benchmark, and Human Evaluation Protocol for Expressive Speech-to-Speech TranslationabstractExpressive speech-to-speech translation (S2ST) aims to transfer prosodic attributes of source speech to target speech while maintaining translation accuracy. Existing research in expressive S2ST is limited, typically focusing on a single expressivity aspect at a time. Likewise, this research area lacks standard evaluation protocols and well-curated benchmark datasets. In this work, we propose a holistic cascade system for expressive S2ST, combining multiple prosody transfer techniques previously considered only in isolation. We curate a benchmark expressivity test set in the TV series domain and explored a second dataset in the audiobook domain. Finally, we present a human evaluation protocol to assess multiple expressive dimensions across speech pairs. Experimental results indicate that bi-lingual annotators can assess the quality of expressive preservation in S2ST systems, and the holistic modeling approach outperforms single-aspect systems. Audio samples can be accessed through our demo webpage: https://facebookresearch.github.io/speech_translation/cascade_expressive_s2st. Wen-Chin Huang, Benjamin N. Peloquin, Justine Kao, Changhan Wang, Hongyu Gong, Elizabeth Salesky, Yossi Adi, Ann Lee 0001, Peng-Jen Chen |
ICASSP | 6 |
| 2023 | Language Modelling with Pixels
Phillip Rust, Jonas F. Lotz, Emanuele Bugliarello, Elizabeth Salesky, Miryam de Lhoneux, Desmond Elliott |
ICLR | 4 |
| 2022 | BibleTTS: a large, high-fidelity, multilingual, and uniquely African speech corpusabstractBibleTTS is a large, high-quality, open speech dataset for ten languages spoken in Sub-Saharan Africa.The corpus contains up to 86 hours of aligned, studio quality 48kHz single speaker recordings per language, enabling the development of high-quality text-to-speech models.The ten languages represented are: Akuapem Twi, Asante Twi, Chichewa, Ewe, Hausa, Kikuyu, Lingala, Luganda, Luo, and Yoruba.This corpus is a derivative work of Bible recordings made and released by the Open.Bible project from Biblica.We have aligned, cleaned, and filtered the original recordings, and additionally hand-checked a subset of the alignments for each language.We present results for text-to-speech models with Coqui TTS. David Ifeoluwa Adelani, Edresson Casanova, Alp Öktem, Daniel Whitenack, Julian Weber, Salomon Kabongo, Elizabeth Salesky, Iroro Orife, Colin Leong, Perez Ogayo, Chris C. Emezue, Jonathan Mukiibi, Salomey Osei, Apelete Agbolo, Victor Akinode, Bernard Opoku, Samuel Olanrewaju, Jesujoba O. Alabi, Shamsuddeen Hassan Muhammad |
INTERSPEECH | 8 |
| 2022 | UniMorph 4.0: Universal MorphologyabstractThe Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological inflection tables for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation, and a type-level resource of annotated data in diverse languages realizing that schema. This paper presents the expansions and improvements on several fronts that were made in the last couple of years (since McCarthy et al. (2020)). Collaborative efforts by numerous linguists have added 66 new languages, including 24 endangered languages. We have implemented several improvements to the extraction pipeline to tackle some issues, e.g., missing gender and macrons information. We have amended the schema to use a hierarchical structure that is needed for morphological phenomena like multiple-argument agreement and case stacking, while adding some missing morphological features to make the schema more inclusive. In light of the last UniMorph release, we also augmented the database with morpheme segmentation for 16 languages. Lastly, this new release makes a push towards inclusion of derivational morphology in UniMorph by enriching the data and annotation schema with instances representing derivational processes from MorphyNet. Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieras, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina J. Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Lane 0002, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóga, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer C. White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo Maria Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar 0002, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Tucker Prud'hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova |
LREC | 49 |
| 2021 | Assessing Evaluation Metrics for Speech-to-Speech TranslationabstractSpeech-to-speech translation combines machine translation with speech synthesis, introducing evaluation challenges not present in either task alone. How to automatically evaluate speech-to-speech translation is an open question which has not previously been explored. Translating to speech rather than to text is often motivated by unwritten languages or languages without standardized orthographies. However, we show that the previously used automatic metric for this task is best equipped for standardized high-resource languages only. In this work, we first evaluate current metrics for speech-to-speech translation, and second assess how translation to dialectal variants rather than to standardized languages impacts various evaluation methods. Elizabeth Salesky, Julian Mäder, Severin Klinger |
ASRU | 1 |
| 2021 | A surprisal-duration trade-off across and within the world's languagesabstractWhile there exist scores of natural languages, each with its unique features and idiosyncrasies, they all share a unifying theme: enabling human communication.We may thus reasonably predict that human cognition shapes how these languages evolve and are used.Assuming that the capacity to process information is roughly constant across human populations, we expect a surprisal-duration trade-off to arise both across and within languages.We analyse this trade-off using a corpus of 600 languages and, after controlling for several potential confounds, we find strong supporting evidence in both settings.Specifically, we find that, on average, phones are produced faster in languages where they are less surprising, and vice versa.Further, we confirm that more surprising phones are longer, on average, in 319 languages out of the 600.We thus conclude that there is strong evidence of a surprisal-duration trade-off in operation, both across and within the world's languages. Tiago Pimentel, Clara Meister, Elizabeth Salesky, Simone Teufel, Damián E. Blasi, Ryan Cotterell |
EMNLP (1) | 3 |
| 2021 | Robust Open-Vocabulary Translation from Visual Text RepresentationsabstractMachine translation models have discrete vo cabularies and commonly use subword seg mentation techniques to achieve an 'open vo cabulary.'This approach relies on consis tent and correct underlying unicode sequences, and makes models susceptible to degrada tion from common types of noise and vari ation.Motivated by the robustness of hu man language processing, we propose the use of visual text representations, which dispense with a finite set of text embeddings in favor of continuous vocabularies created by process ing visually rendered text with sliding win dows.We show that models using visual text representations approach or match per formance of traditional text models on small and larger datasets.More importantly, mod els with visual embeddings demonstrate sig nificant robustness to varied types of noise, achieving e.g., 25.9 BLEU on a character per muted German-English task where subword models degrade to 1.9. Elizabeth Salesky, David Etter, Matt Post |
EMNLP (1) | 1 |
| 2021 | The Multilingual TEDx Corpus for Speech Recognition and TranslationabstractWe present the Multilingual TEDx corpus, built to support speech recognition (ASR) and speech translation (ST) research across many non-English source languages. The corpus is a collection of audio recordings from TEDx talks in 8 source languages. We segment transcripts into sentences and align them to the source-language audio and target-language translations. The corpus is released along with open-sourced code enabling extension to new talks and languages as they become available. Our corpus creation methodology can be applied to more languages than previous work, and creates multi-way parallel evaluation sets. We provide baselines in multiple ASR and ST settings, including multilingual models to improve translation performance for low-resource language pairs. Elizabeth Salesky, Matthew Wiesner, Jacob Bremerman, Roldano Cattoni, Matteo Negri, Marco Turchi, Douglas W. Oard, Matt Post |
Interspeech | 1 |
| 2020 | Generalized Entropy Regularization or: There's Nothing Special about Label SmoothingabstractPrior work has explored directly regularizing the output distributions of probabilistic models to alleviate peaky (i.e.over-confident) predictions, a common sign of overfitting.This class of techniques, of which label smoothing is one, has a connection to entropy regularization.Despite the consistent success of label smoothing across architectures and datasets in language generation tasks, two problems remain open:(1) there is little understanding of the underlying effects entropy regularizers have on models, and (2) the full space of entropy regularization techniques is largely unexplored.We introduce a parametric family of entropy regularizers, which includes label smoothing as a special case, and use it to gain a better understanding of the relationship between the entropy of a trained model and its performance on language generation tasks.We also find that variance in model performance can be explained largely by the resulting entropy of the model.Lastly, we find that label smoothing provably does not allow for sparse distributions, an undesirable property for language generation models, and therefore advise the use of other entropy regularization methods in its place.Our code is available online at https://github.com/ rycolab/entropyRegularization.2 H(p, q) := -z∈Z p(z) log q(z) is cross-entropy and H(p) := H(p, p) = -z∈Z p(z) log p(z) is the Shannon entropy, for which log = log 2 and Z = supp(p).3 The notation used by Pereyra et al. (2017) is imprecise. Clara Meister, Elizabeth Salesky, Ryan Cotterell |
ACL | 2 |
| 2020 | Phone Features Improve Speech TranslationabstractEnd-to-end models for speech translation (ST) more tightly couple speech recognition (ASR) and machine translation (MT) than a traditional cascade of separate ASR and MT models, with simpler model architectures and the potential for reduced error propagation.Their performance is often assumed to be superior, though in many conditions this is not yet the case.We compare cascaded and end-to-end models across high, medium, and low-resource conditions, and show that cascades remain stronger baselines.Further, we introduce two methods to incorporate phone features into ST models.We show that these features improve both architectures, closing the gap between end-to-end models and cascades, and outperforming previous academic work -by up to 9 BLEU on our low-resource setting. Elizabeth Salesky, Alan W. Black |
ACL | 1 |
| 2020 | A Corpus for Large-Scale Phonetic TypologyabstractA major hurdle in data-driven research on typology is having sufficient data in many languages to draw meaningful conclusions. We present VoxClamantis v1.0, the first large-scale corpus for phonetic typology, with aligned segments and estimated phoneme-level labels in 690 readings spanning 635 languages, along with acoustic-phonetic measures of vowels and sibilants. Access to such data can greatly facilitate investigation of phonetic typology at a large scale and across many languages. However, it is non-trivial and computationally intensive to obtain such alignments for hundreds of languages, many of which have few to no resources presently available. We describe the methodology to create our corpus, discuss caveats with current methods and their impact on the utility of this data, and illustrate possible research directions through a series of case studies on the 48 highest-quality readings. Our corpus and scripts are publicly available for non-commercial use at https://voxclamantisproject.github.io. Elizabeth Salesky, Eleanor Chodroff, Tiago Pimentel, Matthew Wiesner, Ryan Cotterell, Alan W. Black, Jason Eisner |
ACL | 1 |
| 2020 | Relative Positional Encoding for Speech Recognition and Direct TranslationabstractTransformer models are powerful sequence-to-sequence architectures that are capable of directly mapping speech inputs to transcriptions or translations. However, the mechanism for modeling positions in this model was tailored for text modeling, and thus is less ideal for acoustic inputs. In this work, we adapt the relative position encoding scheme to the Speech Transformer, where the key addition is relative distance between input states in the self-attention network. As a result, the network can better adapt to the variable distributions present in speech data. Our experiments show that our resulting model achieves the best recognition result on the Switchboard benchmark in the non-augmentation condition, and the best published result in the MuST-C speech translation benchmark. We also show that this model is able to better utilize synthetic data than the Transformer, and adapts better to variable sentence segmentation quality for speech translation. Ngoc-Quan Pham, Thanh-Le Ha, Tuan-Nam Nguyen, Thai Son Nguyen, Elizabeth Salesky, Sebastian Stüker, Jan Niehues, Alex Waibel |
INTERSPEECH | 5 |
| 2020 | Optimizing segmentation granularity for neural machine translation
Elizabeth Salesky, Andrew Runge, Alex Coda, Jan Niehues, Graham Neubig |
Mach. Transl. | 1 |
| 2019 | Exploring Phoneme-Level Speech Representations for End-to-End Speech TranslationabstractPrevious work on end-to-end translation from speech has primarily used frame-level features as speech representations, which creates longer, sparser sequences than text.We show that a naïve method to create compressed phoneme-like speech representations is far more effective and efficient for translation than traditional frame-level speech features.Specifically, we generate phoneme labels for speech frames and average consecutive frames with the same label to create shorter, higher-level source sequences for translation.We see improvements of up to 5 BLEU on both our high and low resource language pairs, with a reduction in training time of 60%.Our improvements hold across multiple data sizes and two language pairs. Elizabeth Salesky, Matthias Sperber, Alan W. Black |
ACL (1) | 1 |
| 2018 | Towards Fluent Translations From Disfluent SpeechabstractWhen translating from speech, special consideration for conversational speech phenomena such as disfluencies is necessary. Most machine translation training data consists of well-formed written texts, causing issues when translating spontaneous speech. Previous work has introduced an intermediate step between speech recognition (ASR) and machine translation (MT) to remove disfluencies, making the data better-matched to typical translation text and significantly improving performance. However, with the rise of end-to-end speech translation systems, this intermediate step must be incorporated into the sequence-to-sequence architecture. Further, though translated speech datasets exist, they are typically news or rehearsed speech without many disfluencies (e.g. TED), or the disfluencies are translated into the references (e.g. Fisher). To generate clean translations from disfluent speech, cleaned references are necessary for evaluation. We introduce a corpus of cleaned target data for the Fisher Spanish-English dataset for this task. We compare how different architectures handle disfluencies and provide a baseline for removing disfluencies in end-to-end translation. Elizabeth Salesky, Susanne Burger, Jan Niehues, Alex Waibel |
SLT | 1 |
| 2016 | Operational Assessment of Keyword Search on Oral History
Elizabeth Salesky, Jessica Ray, Wade Shen |
LREC | 1 |