VLDB 2026 Research / reviewers in the wild / expert
Jason Riesa
dblp:20/9013
· DBLP profile ↗
14ranked-venue papers
4as first author
6since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Multimodal Modeling for Spoken Language IdentificationabstractSpoken language identification refers to the task of automatically predicting the spoken language in a given utterance. Conventionally, it is modeled as a speech-based language identification task. Prior techniques have been constrained to a single modality; however in the case of video data there is a wealth of other metadata that may be beneficial for this task. In this work, we propose MuSeLI, a Multimodal Spoken Language Identification method, which delves into the use of various metadata sources to enhance language identification. Our study reveals that metadata such as video title, description and geographic location provide substantial information to identify the spoken language of the multimedia recording. We conduct experiments using two diverse public datasets of YouTube videos, and obtain state-of-the-art results on the language identification task. We additionally conduct an ablation study that describes the distinct contribution of each modality for language recognition. Shikhar Bharadwaj, Shikhar Vashishth, Ankur Bapna, Sriram Ganapathy, Vera Axelrod, Siddharth Dalmia, Daan van Esch, Sandy Ritchie, Partha Talukdar, Jason Riesa |
ICASSP | 13 |
| 2024 | FLEURS-R: A Restored Multilingual Speech Corpus for Generation Tasks
Yuma Koizumi, Shigeki Karita, Heiga Zen, Jason Riesa, Haruko Ishikawa, Michiel Bacchiani |
INTERSPEECH | 5 |
| 2023 | SQuId: Measuring Speech Naturalness in Many LanguagesabstractMuch of text-to-speech research relies on human evaluation. This incurs heavy costs and slows down the development process, especially in heavily multilingual applications where recruiting and polling annotators can take weeks. We introduce SQuId (Speech Quality Identification), a multilingual naturalness prediction model trained on over a million ratings and tested in 65 locales—the largest effort of this type to date. The main insight is that training one model on many locales consistently surpasses mono-locale baselines. We show that the model outperforms a competitive baseline based on w2v-BERT and VoiceMOS by 50.0%. We then demonstrate the effectiveness of cross-locale transfer during fine-tuning and highlight its effect on zero-shot locales, for which there is no fine-tuning data. We highlight the role of non-linguistic effects such as sound artifacts in cross-locale transfer. Finally, we present the effect of model size and pre-training diversity with ablation experiments. Thibault Sellam, Ankur Bapna, Joshua Camp, Diana Mackinnon, Ankur P. Parikh, Jason Riesa |
ICASSP | 6 |
| 2023 | FRMT: A Benchmark for Few-Shot Region-Aware Machine TranslationabstractAbstract We present FRMT, a new dataset and evaluation benchmark for Few-shot Region-aware Machine Translation, a type of style-targeted translation. The dataset consists of professional translations from English into two regional variants each of Portuguese and Mandarin Chinese. Source documents are selected to enable detailed analysis of phenomena of interest, including lexically distinct terms and distractor terms. We explore automatic evaluation metrics for FRMT and validate their correlation with expert human evaluation across both region-matched and mismatched rating scenarios. Finally, we present a number of baseline models for this task, and offer guidelines for how researchers can train, evaluate, and compare their own models. Our dataset and evaluation code are publicly available: https://bit.ly/frmt-task. Parker Riley, Timothy Dozat, Jan A. Botha, Xavier Garcia, Dan Garrette, Jason Riesa, Orhan Firat, Noah Constant |
Trans. Assoc. Comput. Linguistics | 6 |
| 2022 | XTREME-S: Evaluating Cross-lingual Speech RepresentationsabstractWe introduce XTREME-S, a new benchmark to evaluate universal cross-lingual speech representations in many languages.XTREME-S covers four task families: speech recognition, classification, speech-to-text translation and retrieval.Covering 102 languages from 10+ language families, 3 different domains and 4 task families, XTREME-S aims to simplify multilingual speech representation evaluation, as well as catalyze research in "universal" speech representation learning.This paper describes the new benchmark and establishes the first speech-only and speechtext baselines using XLS-R and mSLAM on all downstream tasks.We motivate the design choices and detail how to use the benchmark.Datasets and fine-tuning scripts are made easily accessible through the HuggingFace platform. 1 Alexis Conneau, Ankur Bapna, Yu Zhang 0033, Patrick von Platen, Anton Lozhkov, Colin Cherry, Ye Jia, Clara Rivera, Mihir Kale, Daan van Esch, Vera Axelrod, Simran Khanuja, Jonathan H. Clark, Orhan Firat, Michael Auli, Sebastian Ruder, Jason Riesa, Melvin Johnson |
INTERSPEECH | 18 |
| 2022 | FLEURS: FEW-Shot Learning Evaluation of Universal Representations of SpeechabstractWe introduce FLEURS, the Few-shot Learning Evaluation of Universal Representations of Speech benchmark. FLEURS is an n-way parallel speech dataset in 102 languages built on top of the machine translation FLoRes-101 benchmark, with approximately 12 hours of speech supervision per language. FLEURS can be used for a variety of speech tasks, including Automatic Speech Recognition (ASR), Speech Language Identification (Speech LangID), Speech-Text Retrieval. In this paper, we provide baselines for the tasks based on multilingual pre-trained models like speech-only w2v-BERT [1] and speech-text multimodal mSLAM [2]. The goal of FLEURS is to enable speech technology in more languages and catalyze research in low-resource speech understanding.1. Alexis Conneau, Simran Khanuja, Yu Zhang 0033, Vera Axelrod, Siddharth Dalmia, Jason Riesa, Clara Rivera, Ankur Bapna |
SLT | 7 |
| 2020 | Evaluating the Cross-Lingual Effectiveness of Massively Multilingual Neural Machine TranslationabstractThe recently proposed massively multilingual neural machine translation (NMT) system has been shown to be capable of translating over 100 languages to and from English within a single model (Aharoni, Johnson, and Firat 2019). Its improved translation performance on low resource languages hints at potential cross-lingual transfer capability for downstream tasks. In this paper, we evaluate the cross-lingual effectiveness of representations from the encoder of a massively multilingual NMT model on 5 downstream classification and sequence labeling tasks covering a diverse set of over 50 languages. We compare against a strong baseline, multilingual BERT (mBERT) (Devlin et al. 2018), in different cross-lingual transfer learning scenarios and show gains in zero-shot transfer in 4 out of these 5 tasks. Aditya Siddhant, Melvin Johnson, Henry Tsai, Naveen Ari, Jason Riesa, Ankur Bapna, Orhan Firat, Karthik Raman 0001 |
AAAI | 5 |
| 2020 | Improving Multilingual Models with Language-Clustered VocabulariesabstractState-of-the-art multilingual models depend on vocabularies that cover all of the languages the model will expect to see at inference time, but the standard methods for generating those vocabularies are not ideal for massively multilingual applications.In this work, we introduce a novel procedure for multilingual vocabulary generation that combines the separately trained vocabularies of several automatically derived language clusters, thus balancing the trade-off between cross-lingual subword sharing and language-specific vocabularies.Our experiments show improvements across languages on key multilingual benchmark tasks TYDI QA (+2.9 F1), XNLI (+2.1%), and WikiAnn NER (+2.8 F1) and factor of 8 reduction in out-of-vocabulary rate, all without increasing the size of the model or data. Hyung Won Chung, Dan Garrette, Kiat Chuan Tan, Jason Riesa |
EMNLP (1) | 4 |
| 2019 | Small and Practical BERT Models for Sequence LabelingabstractHenry Tsai, Jason Riesa, Melvin Johnson, Naveen Arivazhagan, Xin Li, Amelia Archer. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Henry Tsai, Jason Riesa, Melvin Johnson, Naveen Arivazhagan, Amelia Archer |
EMNLP/IJCNLP (1) | 2 |
| 2018 | A Fast, Compact, Accurate Model for Language Identification of Codemixed TextabstractWe address fine-grained multilingual language identification: providing a language code for every token in a sentence, including codemixed text containing multiple languages.Such text is prevalent online, in documents, social media, and message boards.We show that a feed-forward network with a simple globally constrained decoder can accurately and rapidly label both codemixed and monolingual text in 100 languages and 100 language pairs.This model outperforms previously published multilingual approaches in terms of both accuracy and speed, yielding an 800x speed-up and a 19.5% averaged absolute gain on three codemixed datasets.It furthermore outperforms several benchmark systems on monolingual language identification. Yuan Zhang 0001, Jason Riesa, Daniel Gillick, Anton Bakalov, Jason Baldridge, David Weiss 0001 |
EMNLP | 2 |
| 2012 | Automatic Parallel Fragment Extraction from Noisy Data
Jason Riesa, Daniel Marcu |
HLT-NAACL | 1 |
| 2011 | Feature-Rich Language-Independent Syntax-Based Alignment for Statistical Machine Translation
Jason Riesa, Ann Irvine, Daniel Marcu |
EMNLP | 1 |
| 2010 | Hierarchical Search for Word Alignment
Jason Riesa, Daniel Marcu |
ACL | 1 |
| 2006 | Building an English-iraqi Arabic machine translation system for spoken utterances with limited resourcesabstractThis paper presents an English-Iraqi Arabic speech-to-speech statistical machine translation system using limited resources. In it, we explore the constraints involved, how we endeavored to mitigate such problems as a non-standard orthography and a highly inflected grammar, and discuss leveraging existing plentiful resources for Modern Standard Arabic to assist in this task. These combined techniques yield a reduction in unknown words at translation time by over 40 % and a +3.65 increase in BLEU score over a previous state-of-the-art system using the same parallel training corpus of spoken utterances. Index Terms: speech translation, limited resources, Arabic 1. Jason Riesa, Behrang Mohit, Kevin Knight, Daniel Marcu |
INTERSPEECH | 1 |