VLDB 2026 Research / reviewers in the wild / expert
Kyle Kastner
dblp:163/1843
· DBLP profile ↗
11ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Audio Diffusion with Large Language ModelsabstractIn this paper, we explore an alternate approach to the popular method of using large language models (LLMs) as a second decoder for Automated Speech Recognition (ASR) and speech understanding tasks. We propose to employ diffusion networks to generate a correction signal that can be applied on the original input audio features to improve performance. Specifically, the diffusion network is trained to predict the gradient of any ASR objective with respect to the input audio features conditioned on LLM embeddings. Our experiments are conducted on public corpora, namely, Librispeech and Common Voice. We show that the diffusion model is able to improve ASR performance on noisy and accented speech, with the addition of knowledge from the LLM, and also helps improve generalization to out-of-domain test sets. Kyle Kastner, Kartik Audhkhasi, Bhuvana Ramabhadran, Andrew Rosenberg |
ICASSP | 2 |
| 2025 | Speech Re-Painting for Robust ASRabstractSynthetic speech is a useful source for augmentation of automatic speech recognition (ASR) systems, but there is a "sim-to-real" gap between synthetic and real speech that can limit generalization. The natural variability of real speech is essential to the training of robust ASR systems. While synthetic data augmentation can be used to approximate the variability of natural speech, however, not all aspects of variation are equally relevant for augmentation. In this work, we introduce speech re-painting, a method for in-context augmented synthesis, using target training datasets to generate new utterances guided by speech and text on the fly in a zero-shot manner. We evaluate this technique using downstream ASR word error rate (WER) using the VCTK and LibriSpeech datasets. These represent unique speaker and lexical challenges that are addressed by re-painting, realizing a reduction of WER more than 50% in particular settings. Kyle Kastner, Gary Wang, Isaac Elias, Takaaki Saeki, Pedro J. Moreno 0001, Françoise Beaufays, Andrew Rosenberg, Bhuvana Ramabhadran |
ICASSP | 1 |
| 2024 | Extending Multilingual Speech Synthesis to 100+ Languages without Transcribed DataabstractCollecting high-quality studio recordings of audio is challenging, which limits the language coverage of text-to-speech (TTS) systems. This paper proposes a framework for scaling a multilingual TTS model to 100+ languages using found data without supervision. The proposed framework combines speech-text encoder pretraining with unsupervised training using untranscribed speech and unspoken text data sources, thereby leveraging massively multilingual joint speech and text representation learning. Without any transcribed speech in a new language, this TTS model can generate intelligible speech in >30 unseen languages (CER difference of <10% to ground truth). With just 15 minutes of transcribed, found data, we can reduce the intelligibility difference to 1% or less from the ground-truth, and achieve naturalness scores that match the ground-truth in several languages. Takaaki Saeki, Gary Wang, Nobuyuki Morioka, Isaac Elias, Kyle Kastner, Andrew Rosenberg, Bhuvana Ramabhadran, Heiga Zen, Françoise Beaufays, Hadar Shemtov |
ICASSP | 5 |
| 2024 | Adaptive Accompaniment with ReaLchordsabstractJamming requires coordination, anticipation, and collaborative creativity between musicians. Current generative models of music produce expressive output but are not able to generate in an online manner, meaning simultaneously with other musicians (human or otherwise). We propose ReaLchords, an online generative model for improvising chord accompaniment to user melody. We start with an online model pretrained by maximum likelihood, and use reinforcement learning to finetune the model for online use. The finetuning objective leverages both a novel reward model that provides feedback on both harmonic and temporal coherency between melody and chord, and a divergence term that implements a novel type of distillation from a teacher model that can see the future melody. Through quantitative experiments and listening tests, we demonstrate that the resulting model adapts well to unfamiliar input and produce fitting accompaniment. ReaLchords opens the door to live jamming, as well as simultaneous co-creation in other modalities. Yusong Wu, Tim Cooijmans, Kyle Kastner, Adam Roberts, Ian Simon, Alexander Scarlatos, Chris Donahue, Cassie Tarakajian, Shayegan Omidshafiei, Aaron C. Courville, Pablo Samuel Castro, Natasha Jaques, Cheng-Zhi Anna Huang |
ICML | 3 |
| 2024 | Enhancing Low-Resource Spoken Language Identification Via Cross-Modality Retrieval and Cross-Lingual Text-to-Speech SynthesisabstractSpoken language identification (SLID) for low-resource languages remains challenging due to limited data availability. In this paper, we present two novel approaches to address the issue: cross-modality retrieval-based data selection and cross-lingual text-to-speech (TTS) based data augmentation. Incorporating semi-supervised speech and synthetic speech produced by the two methods, we successfully enhance SLID on low-resource languages and on the full set of target languages, at a publicly available YouTube-derived dataset. Our best recipe reduces training data amount by 28% and ensures a more balanced distribution of training data across languages. The two general frameworks offer innovative strategies for leveraging resources to add valuable data to enhance SLID in extremely low-resource scenarios. Gary Wang, Kyle Kastner, Isaac Caswell, Charles Yoon, Andrew Rosenberg |
SLT | 3 |
| 2023 | Understanding Shared Speech-Text RepresentationsabstractRecently, a number of approaches to train speech models by incorporating text into end-to-end models have been developed, with Maestro advancing state-of-the-art automatic speech recognition (ASR) and Speech Translation (ST) performance. In this paper, we expand our understanding of the resulting shared speech-text representations with two types of analyses. First we examine the limits of speech-free domain adaptation, finding that a corpus-specific duration model for speech-text alignment is the most important component for learning a shared speech-text representation. Second, we inspect the similarities between activations of unimodal (speech or text) encoders as compared to the activations of a shared encoder. We find that the shared encoder learns a more compact and overlapping speech-text representation than the uni-modal encoders. We hypothesize that this partially explains the effectiveness of the Maestro shared speech-text representations. Gary Wang, Kyle Kastner, Ankur Bapna, Zhehuai Chen, Andrew Rosenberg, Bhuvana Ramabhadran, Yu Zhang 0033 |
ICASSP | 2 |
| 2022 | MIDI-DDSP: Detailed Control of Musical Performance via Hierarchical Modeling
Yusong Wu, Ethan Manilow, Rigel Swavely, Kyle Kastner, Tim Cooijmans, Aaron C. Courville, Cheng-Zhi Anna Huang, Jesse H. Engel |
ICLR | 5 |
| 2019 | Representation Mixing for TTS SynthesisabstractRecent character and phoneme-based parametric TTS systems using deep learning have shown strong performance in natural speech generation. However, the choice between character or phoneme input can create serious limitations for practical deployment, as direct control of pronunciation is crucial in certain cases. We demonstrate a simple method for combining multiple types of linguistic information in a single encoder, named representation mixing, enabling flexible choice between character, phoneme, or mixed representations during inference. Experiments and user studies on a public audiobook corpus show the efficacy of our approach. Kyle Kastner, João Felipe Santos, Yoshua Bengio, Aaron C. Courville |
ICASSP | 1 |
| 2017 | Learning to Discover Sparse Graphical ModelsabstractWe consider structure discovery of undirected graphical models from observational data. Inferring likely structures from few examples is a complex task often requiring the formulation of priors and sophisticated inference procedures. Popular methods rely on estimating a penalized maximum likelihood of the precision matrix. However, in these approaches structure recovery is an indirect consequence of the data-fit term, the penalty can be difficult to adapt for domain-specific knowledge, and the inference is computationally demanding. By contrast, it may be easier to generate training samples of data that arise from graphs with the desired structure properties. We propose here to leverage this latter source of information as training data to learn a function, parametrized by a neural network, that maps empirical covariance matrices to estimated graph structures. Learning this function brings two benefits: it implicitly models the desired structure or sparsity properties to form suitable priors, and it can be tailored to the specific problem of edge structure discovery, rather than maximizing data likelihood. Applying this framework, we find our learnable graph-discovery method trained on synthetic data generalizes well: identifying relevant edges in both synthetic and real data, completely unknown at training time. We find that on genetics, brain imaging, and simulation data we obtain performance generally superior to analytical methods. Eugene Belilovsky, Kyle Kastner, Gaël Varoquaux, Matthew B. Blaschko |
ICML | 2 |
| 2015 | A Recurrent Latent Variable Model for Sequential DataabstractIn this paper, we explore the inclusion of latent random variables into the hidden state of a recurrent neural network (RNN) by combining the elements of the variational autoencoder. We argue that through the use of high-level latent random variables, the variational RNN (VRNN) can model the kind of variability observed in highly structured sequential data such as natural speech. We empirically evaluate the proposed model against other related sequential models on four speech datasets and one handwriting dataset. Our results show the important roles that latent random variables can play in the RNN dynamics. Junyoung Chung, Kyle Kastner, Laurent Dinh, Kratarth Goel, Aaron C. Courville, Yoshua Bengio |
NIPS | 2 |
| 2015 | Learning Distributed Representations from Reviews for Collaborative FilteringabstractRecent work has shown that collaborative filter-based recommender systems can be improved by incorporating side information, such as natural language reviews, as a way of regularizing the derived product representations. Motivated by the success of this approach, we introduce two different models of reviews and study their effect on collaborative filtering performance. While the previous state-of-the-art approach is based on a latent Dirichlet allocation (LDA) model of reviews, the models we explore are neural network based: a bag-of-words product-of-experts model and a recurrent neural network. We demonstrate that the increased flexibility offered by the product-of-experts model allowed it to achieve state-of-the-art performance on the Amazon review dataset, outperforming the LDA-based approach. However, interestingly, the greater modeling power offered by the recurrent neural network appears to undermine the model's ability to act as a regularizer of the product representations. Amjad Almahairi, Kyle Kastner, Kyunghyun Cho, Aaron C. Courville |
RecSys | 2 |