VLDB 2026 Research / reviewers in the wild / expert
Andrew Rosenberg
dblp:21/6080
· DBLP profile ↗
95ranked-venue papers
23as first author
26since 2021 · last 2025
0000-0003-1780-4390ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 85 · 20 first-author · 26 since 2021Artificial intelligence and machine learning · 64 · 19 first-author · 15 since 2021Human-computer interaction and ubiquitous computing · 3Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Audio Diffusion with Large Language ModelsabstractIn this paper, we explore an alternate approach to the popular method of using large language models (LLMs) as a second decoder for Automated Speech Recognition (ASR) and speech understanding tasks. We propose to employ diffusion networks to generate a correction signal that can be applied on the original input audio features to improve performance. Specifically, the diffusion network is trained to predict the gradient of any ASR objective with respect to the input audio features conditioned on LLM embeddings. Our experiments are conducted on public corpora, namely, Librispeech and Common Voice. We show that the diffusion model is able to improve ASR performance on noisy and accented speech, with the addition of knowledge from the LLM, and also helps improve generalization to out-of-domain test sets. Kyle Kastner, Kartik Audhkhasi, Bhuvana Ramabhadran, Andrew Rosenberg |
ICASSP | 5 |
| 2025 | Speech Re-Painting for Robust ASRabstractSynthetic speech is a useful source for augmentation of automatic speech recognition (ASR) systems, but there is a "sim-to-real" gap between synthetic and real speech that can limit generalization. The natural variability of real speech is essential to the training of robust ASR systems. While synthetic data augmentation can be used to approximate the variability of natural speech, however, not all aspects of variation are equally relevant for augmentation. In this work, we introduce speech re-painting, a method for in-context augmented synthesis, using target training datasets to generate new utterances guided by speech and text on the fly in a zero-shot manner. We evaluate this technique using downstream ASR word error rate (WER) using the VCTK and LibriSpeech datasets. These represent unique speaker and lexical challenges that are addressed by re-painting, realizing a reduction of WER more than 50% in particular settings. Kyle Kastner, Gary Wang, Isaac Elias, Takaaki Saeki, Pedro J. Moreno 0001, Françoise Beaufays, Andrew Rosenberg, Bhuvana Ramabhadran |
ICASSP | 7 |
| 2024 | Extending Multilingual Speech Synthesis to 100+ Languages without Transcribed DataabstractCollecting high-quality studio recordings of audio is challenging, which limits the language coverage of text-to-speech (TTS) systems. This paper proposes a framework for scaling a multilingual TTS model to 100+ languages using found data without supervision. The proposed framework combines speech-text encoder pretraining with unsupervised training using untranscribed speech and unspoken text data sources, thereby leveraging massively multilingual joint speech and text representation learning. Without any transcribed speech in a new language, this TTS model can generate intelligible speech in >30 unseen languages (CER difference of <10% to ground truth). With just 15 minutes of transcribed, found data, we can reduce the intelligibility difference to 1% or less from the ground-truth, and achieve naturalness scores that match the ground-truth in several languages. Takaaki Saeki, Gary Wang, Nobuyuki Morioka, Isaac Elias, Kyle Kastner, Andrew Rosenberg, Bhuvana Ramabhadran, Heiga Zen, Françoise Beaufays, Hadar Shemtov |
ICASSP | 6 |
| 2024 | Speech Prefix-Tuning with RNNT Loss for Improving LLM Predictions
Murali Karthick Baskar, Andrew Rosenberg, Bhuvana Ramabhadran, Neeraj Gaur, Zhong Meng |
INTERSPEECH | 2 |
| 2024 | ASTRA: Aligning Speech and Text Representations for Asr without Sampling
Neeraj Gaur, Rohan Agrawal, Gary Wang, Parisa Haghani, Andrew Rosenberg, Bhuvana Ramabhadran |
INTERSPEECH | 5 |
| 2024 | Contemplative Mechanism for Speech Recognition: Speech Encoders can Think
Tien-Ju Yang, Andrew Rosenberg, Bhuvana Ramabhadran |
INTERSPEECH | 2 |
| 2024 | Speculative Speech Recognition by Audio-Prefixed Low-Rank Adaptation of Language Models
Bolaji Yusuf, Murali Karthick Baskar, Andrew Rosenberg, Bhuvana Ramabhadran |
INTERSPEECH | 3 |
| 2024 | Enhancing Low-Resource Spoken Language Identification Via Cross-Modality Retrieval and Cross-Lingual Text-to-Speech SynthesisabstractSpoken language identification (SLID) for low-resource languages remains challenging due to limited data availability. In this paper, we present two novel approaches to address the issue: cross-modality retrieval-based data selection and cross-lingual text-to-speech (TTS) based data augmentation. Incorporating semi-supervised speech and synthetic speech produced by the two methods, we successfully enhance SLID on low-resource languages and on the full set of target languages, at a publicly available YouTube-derived dataset. Our best recipe reduces training data amount by 28% and ensures a more balanced distribution of training data across languages. The two general frameworks offer innovative strategies for leveraging resources to add valuable data to enhance SLID in extremely low-resource scenarios. Gary Wang, Kyle Kastner, Isaac Caswell, Charles Yoon, Andrew Rosenberg |
SLT | 6 |
| 2023 | Mask-Conformer: Augmenting Conformer with Mask-Predict DecoderabstractMuch of the recent progress in automatic speech recognition (ASR) lies in developing an acoustic encoder, such as enlarging its capacity and designing a refined architecture for speech processing. With these highly optimized encoders, the decoder has become less influential in its role as a language model (LM). In this work, we explore an effective approach for employing the LM structure in an ASR model. The proposed Mask-Conformer augments a Conformer-based model with a mask-predict decoder, which learns output context via the masked LM objective. The mask-predict decoder is applied to stacks of encoder layers, where the decoder output explicitly conditions the subsequent layers using cross-attention. We also propose a fill-mask decoding algorithm that refines a sequence using the decoder’s linguistic information. Experimental results show that Mask-Conformer outperforms strong baselines on some tasks. In addition, our analyses validate the effectiveness of the proposed model design. Yosuke Higuchi, Andrew Rosenberg, Murali Karthick Baskar, Bhuvana Ramabhadran |
ASRU | 2 |
| 2023 | JEIT: Joint End-to-End Model and Internal Language Model Training for Speech RecognitionabstractWe propose JEIT, a joint end-to-end (E2E) model and internal language model (ILM) training method to inject large-scale unpaired text into ILM during E2E training which improves rare-word speech recognition. With JEIT, the E2E model computes an E2E loss on audio-transcript pairs while its ILM estimates a cross-entropy loss on unpaired text. The E2E model is trained to minimize a weighted sum of E2E and ILM losses. During JEIT, ILM absorbs knowledge from unpaired text while the E2E training serves as regularization. Unlike ILM adaptation methods, JEIT does not require a separate adaptation step and avoids the need for Kullback-Leibler divergence regularization of ILM. We also show that modular hybrid autoregressive transducer (MHAT) performs better than HAT in the JEIT framework, and is much more robust than HAT during ILM adaptation. To push the limit of unpaired text injection, we further propose a combined JEIT and JOIST training (CJJT) that benefits from modality matching, encoder text injection and ILM training. Both JEIT and CJJT can foster a more effective LM fusion. With 100B unpaired sentences, JEIT/CJJT improves rare-word recognition accuracy by up to 16.4% over a model trained without unpaired text. Zhong Meng, Rohit Prabhavalkar, Tara N. Sainath, Tongzhou Chen, Ehsan Variani, Yu Zhang 0033, Bo Li 0028, Andrew Rosenberg, Bhuvana Ramabhadran |
ICASSP | 9 |
| 2023 | Virtuoso: Massive Multilingual Speech-Text Joint Semi-Supervised Learning for Text-to-SpeechabstractThis paper proposes Virtuoso, a massively multilingual speech–text joint semi-supervised learning framework for text-to-speech synthesis (TTS) models. Existing multilingual TTS typically supports tens of languages, which are a small fraction of the thousands of languages in the world. One difficulty to scale multilingual TTS to hundreds of languages is collecting high-quality speech–text paired data in low-resource languages. This study extends Maestro, a speech–text joint pretraining framework for automatic speech recognition (ASR), to speech generation tasks. To train a TTS model from various types of speech and text data, different training schemes are designed to handle supervised (paired TTS and ASR data) and unsupervised (untranscribed speech and unspoken text) datasets. Experimental evaluation shows that 1) multilingual TTS models trained on Virtuoso can achieve significantly better naturalness and intelligibility than baseline ones in seen languages, and 2) they can synthesize reasonably intelligible and naturally sounding speech for unseen languages where no high-quality paired TTS data is available. Takaaki Saeki, Heiga Zen, Zhehuai Chen, Nobuyuki Morioka, Gary Wang, Yu Zhang 0033, Ankur Bapna, Andrew Rosenberg, Bhuvana Ramabhadran |
ICASSP | 8 |
| 2023 | Understanding Shared Speech-Text RepresentationsabstractRecently, a number of approaches to train speech models by incorporating text into end-to-end models have been developed, with Maestro advancing state-of-the-art automatic speech recognition (ASR) and Speech Translation (ST) performance. In this paper, we expand our understanding of the resulting shared speech-text representations with two types of analyses. First we examine the limits of speech-free domain adaptation, finding that a corpus-specific duration model for speech-text alignment is the most important component for learning a shared speech-text representation. Second, we inspect the similarities between activations of unimodal (speech or text) encoders as compared to the activations of a shared encoder. We find that the shared encoder learns a more compact and overlapping speech-text representation than the uni-modal encoders. We hypothesize that this partially explains the effectiveness of the Maestro shared speech-text representations. Gary Wang, Kyle Kastner, Ankur Bapna, Zhehuai Chen, Andrew Rosenberg, Bhuvana Ramabhadran, Yu Zhang 0033 |
ICASSP | 5 |
| 2023 | O-1: Self-training with Oracle and 1-best Hypothesis
Murali Karthick Baskar, Andrew Rosenberg, Bhuvana Ramabhadran, Kartik Audhkhasi |
INTERSPEECH | 2 |
| 2023 | Using Text Injection to Improve Recognition of Personal Identifiers in Speech
Yochai Blau, Rohan Agrawal, Lior Madmony, Gary Wang, Andrew Rosenberg, Zhehuai Chen, Zorik Gekhman, Genady Beryozkin, Parisa Haghani, Bhuvana Ramabhadran |
INTERSPEECH | 5 |
| 2023 | Improving Joint Speech-Text Representations Without Alignment
Cal Peyser, Zhong Meng, Rohit Prabhavalkar, Andrew Rosenberg, Tara N. Sainath, Michael Picheny, Kyunghyun Cho |
INTERSPEECH | 4 |
| 2022 | Tts4pretrain 2.0: Advancing the use of Text and Speech in ASR Pretraining with Consistency and Contrastive LossesabstractAn effective way to learn representations from untranscribed speech and unspoken text with linguistic/lexical representations derived from synthesized speech was introduced in tts4pretrain [1]. However, the representations learned from synthesized and real speech are likely to be different, potentially limiting the improvements from incorporating unspoken text. In this paper, we introduce learning from supervised speech earlier on in the training process with consistency-based regularization between real and synthesized speech. This allows for better learning of shared speech and text representations. Thus, we introduce a new objective, with encoder and decoder consistency and contrastive regularization between real and synthesized speech derived from the labeled corpora during the pretraining stage. We show that the new objective leads to more similar representations derived from speech and text that help downstream ASR. The proposed pretraining method yields Word Error Rate (WER) reductions of 7-21% relative on six public corpora, Librispeech, AMI, TEDLIUM, Common Voice, Switchboard, CHiME-6, over a state-of-the-art baseline pretrained with wav2vec2.0 and 2-17% over the previously proposed tts4pretrain. The proposed method outperforms the supervised SpeechStew by up to 17%. Moreover, we show that the proposed method also yields WER reductions on larger data sets by evaluating on a large resource, in-house Voice Search task and streaming ASR. Zhehuai Chen, Yu Zhang 0033, Andrew Rosenberg, Bhuvana Ramabhadran, Pedro J. Moreno 0001, Gary Wang |
ICASSP | 3 |
| 2022 | Reducing Domain mismatch in Self-supervised speech pre-training
Murali Karthick Baskar, Andrew Rosenberg, Bhuvana Ramabhadran, Yu Zhang 0033 |
INTERSPEECH | 2 |
| 2022 | A Scalable Model Specialization Framework for Training and Inference using Submodels and its Application to Speech Model PersonalizationabstractModel fine-tuning and adaptation have become a common approach for model specialization for downstream tasks or domains. Fine-tuning the entire model or a subset of the parameters using light-weight adaptation has shown considerable success across different specialization tasks. Fine-tuning a model for a large number of domains typically requires starting a new training job for every domain posing scaling limitations. Once these models are trained, deploying them also poses significant scalability challenges for inference for real-time applications. In this paper, building upon prior light-weight adaptation techniques, we propose a modular framework that enables us to substantially improve scalability for model training and inference. We introduce Submodels that can be quickly and dynamically loaded for on-the-fly inference. We also propose multiple approaches for training those Submodels in parallel using an embedding space in the same training job. We test our framework on an extreme use-case which is speech model personalization for atypical speech, requiring a Submodel for each user. We obtain 128x Submodel throughput with a fixed computation budget without a loss of accuracy. We also show that learning a speaker-embedding space can scale further and reduce the amount of personalization training data required per speaker. Fadi Biadsy, Youzheng Chen, Oleg Rybakov, Andrew Rosenberg, Pedro J. Moreno 0001 |
INTERSPEECH | 5 |
| 2022 | MAESTRO: Matched Speech Text Representations through Modality MatchingabstractWe present Maestro, a self-supervised training method to unify representations learnt from speech and text modalities.Self-supervised learning from speech signals aims to learn the latent structure inherent in the signal, while self-supervised learning from text attempts to capture lexical information.Learning aligned representations from unpaired speech and text sequences is a challenging task.Previous work either implicitly enforced the representations learnt from these two modalities to be aligned in the latent space through multitasking and parameter sharing or explicitly through conversion of modalities via speech synthesis.While the former suffers from interference between the two modalities, the latter introduces additional complexity.In this paper, we propose Maestro, a novel algorithm to learn unified representations from both these modalities simultaneously that can transfer to diverse downstream tasks such as Automated Speech Recognition (ASR) and Speech Translation (ST).Maestro learns unified representations through sequence alignment, duration prediction and matching embeddings in the learned space through an aligned masked-language model loss.We establish a new state-of-the-art (SOTA) on VoxPopuli multilingual ASR with a 8% relative reduction in Word Error Rate (WER), multidomain SpeechStew ASR (3.7% relative) and 21 languages to English multilingual ST on CoVoST 2 with an improvement of 2.8 BLEU averaged over 21 languages. Zhehuai Chen, Yu Zhang 0033, Andrew Rosenberg, Bhuvana Ramabhadran, Pedro J. Moreno 0001, Ankur Bapna, Heiga Zen |
INTERSPEECH | 3 |
| 2022 | Towards Disentangled Speech RepresentationsabstractThe careful construction of audio representations has become a dominant feature in the design of approaches to many speech tasks.Increasingly, such approaches have emphasized "disentanglement", where a representation contains only parts of the speech signal relevant to transcription while discarding irrelevant information.In this paper, we construct a representation learning task based on joint modeling of ASR and TTS, and seek to learn a representation of audio that disentangles that part of the speech signal that is relevant to transcription from that part which is not.We present empirical evidence that successfully finding such a representation is tied to the randomness inherent in training.We then make the observation that these desired, disentangled solutions to the optimization problem possess unique statistical properties.Finally, we show that enforcing these properties during training improves WER by 24.5% relative on average for our joint modeling task.These observations motivate a novel approach to learning effective audio representations. Cal Peyser, W. Ronny Huang, Andrew Rosenberg, Tara N. Sainath, Michael Picheny, Kyunghyun Cho |
INTERSPEECH | 3 |
| 2022 | Non-Parallel Voice Conversion for ASR AugmentationabstractAutomatic speech recognition (ASR) needs to be robust to speaker differences.Voice Conversion (VC) modifies speaker characteristics of input speech.This is an attractive feature for ASR data augmentation.In this paper, we demonstrate that voice conversion can be used as a data augmentation technique to improve ASR performance, even on LibriSpeech, which contains 2,456 speakers.For ASR augmentation, it is necessary that the VC model be robust to a wide range of input speech.This motivates the use of a non-autoregressive, non-parallel VC model, and the use of a pretrained ASR encoder within the VC model.This work suggests that despite including many speakers, speaker diversity may remain a limitation to ASR quality.Finally, interrogation of our VC performance has provided useful metrics for objective evaluation of VC quality. Gary Wang, Andrew Rosenberg, Bhuvana Ramabhadran, Fadi Biadsy, Jesse Emond, Pedro J. Moreno 0001 |
INTERSPEECH | 2 |
| 2022 | Maestro-U: Leveraging Joint Speech-Text Representation Learning for Zero Supervised Speech ASRabstractTraining state-of-the-art Automated Speech Recognition (ASR) models typically requires a substantial amount of transcribed speech. In this work, we demonstrate that a modality-matched joint speech and text model introduced in [1] can be leveraged to train a massively multilingual ASR model without any supervised (manually transcribed) speech for some languages. This paper explores the use of jointly learnt speech and text representations in a massively multilingual, zero supervised speech, real-world setting to expand the set of languages covered by ASR with only unlabeled speech and text in the target languages. Using the FLEURS dataset, we define the task to cover 102 languages, where transcribed speech is available in 52 of these languages and can be used to improve end-to-end ASR quality on the remaining 50. First, we show that by combining speech representations with byte-level text representations and use of language embeddings, we can dramatically reduce the Character Error Rate (CER) on languages with no supervised speech from 64.8% to 30.8%, a relative reduction of 53%. Second, using a subset of South Asian languages we show that Maestro-U can promote knowledge transfer from languages with supervised speech even when there is limited to no graphemic overlap. Overall, Maestro-U closes the gap to oracle performance by 68.5% relative and reduces the CER of 19 languages below 15%. Zhehuai Chen, Ankur Bapna, Andrew Rosenberg, Yu Zhang 0033, Bhuvana Ramabhadran, Pedro J. Moreno 0001, Nanxin Chen |
SLT | 3 |
| 2022 | G-Augment: Searching for the Meta-Structure of Data Augmentation Policies for ASRabstractData augmentation is a ubiquitous technique used to provide robustness to automatic speech recognition (ASR) training. However, even as so much of the ASR training process has become automated and more “end-to-end,” the data augmentation policy (what augmentation functions to use, and how to apply them) remains hand-crafted. We present G(raph)-Augment, a technique to define the augmentation space as directed acyclic graphs (DAGs) and search over this space to optimize the augmentation policy itself. We show that given the same computational budget, policies produced by G-Augment are able to perform better than SpecAugment policies obtained by random search on fine-tuning tasks on CHiME-6 and AMI. G-Augment is also able to establish a new state-of-the-art ASR performance on the CHiME-6 evaluation set (30.7% WER). We further demonstrate that G- Augment policies show better transfer properties across warm-start to cold-start training and model size compared to random-searched SpecAugment policies. Gary Wang, Ekin Dogus Cubuk, Andrew Rosenberg, Shuyang Cheng, Ron J. Weiss, Bhuvana Ramabhadran, Pedro J. Moreno 0001, Quoc V. Le, Daniel S. Park |
SLT | 3 |
| 2021 | Injecting Text in Self-Supervised Speech PretrainingabstractSelf-supervised pretraining for Automated Speech Recognition (ASR) has shown varied degrees of success. In this paper, we propose to jointly learn representations during pretraining from two different modalities: speech and text. The proposed method, tts4pretrain complements the power of contrastive learning in self-supervision with linguistic/lexical representations derived from synthesized speech, effectively learning from untranscribed speech and unspoken text. Lexical learning in the speech encoder is enforced through an additional sequence loss term that is coupled with contrastive loss during pretraining. We demonstrate that this novel pretraining method yields Word Error Rate (WER) reductions of 10% relative on the well-benchmarked, Librispeech task over a state-of-the-art baseline pretrained with wav2vec2.0 only. The proposed method also serves as an effective strategy to compensate for the lack of transcribed speech, effectively matching the performance of 5000 hours of transcribed speech with just 100 hours of transcribed speech on the AMI meeting transcription task. Finally, we demonstrate WER reductions of up to 15% on an inhouse Voice Search task over traditional pretraining. Incorporating text into encoder pretraining is complimentary to rescoring with a larger or in-domain language model, resulting in additional 6% relative reduction in WER. Zhehuai Chen, Yu Zhang 0033, Andrew Rosenberg, Bhuvana Ramabhadran, Gary Wang, Pedro J. Moreno 0001 |
ASRU | 3 |
| 2021 | Extending Parrotron: An End-to-End, Speech Conversion and Speech Recognition Model for Atypical SpeechabstractWe present an extended Parrotron model: a single, end-to-end network that enables voice conversion and recognition simultaneously. Input spectrograms are transformed to output spectrograms in the voice of a predetermined target speaker while also generating hypotheses in a target vocabulary. We study the performance of this novel architecture, which jointly predicts speech and text, on atypical (e.g. dysarthric) speech. We show that with as little as an hour of atypical speech, speaker adaptation can yield a 77% relative reduction in Word Error Rate (WER), measured by ASR performance on the converted speech. We also show that data augmentation using a customized synthesizer built on atypical speech can provide an additional 10% relative improvement over the best speaker-adapted model. Finally, we show how these methods generalize across 8 types of atypical speech for a range of speech impairment severities. Rohan Doshi, Youzheng Chen, Liyang Jiang, Fadi Biadsy, Bhuvana Ramabhadran, Fang Chu, Andrew Rosenberg, Pedro J. Moreno 0001 |
ICASSP | 8 |
| 2021 | Semi-Supervision in ASR: Sequential MixMatch and Factorized TTS-Based Augmentation
Zhehuai Chen, Andrew Rosenberg, Yu Zhang 0033, Heiga Zen, Mohammadreza Ghodsi, Jesse Emond, Gary Wang, Bhuvana Ramabhadran, Pedro J. Moreno 0001 |
Interspeech | 2 |
| 2020 | Generating Diverse and Natural Text-to-Speech Samples Using a Quantized Fine-Grained VAE and Autoregressive Prosody PriorabstractRecent neural text-to-speech (TTS) models with fine-grained latent features enable precise control of the prosody of synthesized speech. Such models typically incorporate a fine-grained variational autoencoder (VAE) structure, extracting latent features at each input token (e.g., phonemes). However, generating samples with the standard VAE prior often results in unnatural and discontinuous speech, with dramatic prosodic variation between tokens. This paper proposes a sequential prior in a discrete latent space which can generate more naturally sounding samples. This is accomplished by discretizing the latent features using vector quantization (VQ), and separately training an autoregressive (AR) prior model over the result. We evaluate the approach using listening tests, objective metrics of automatic speech recognition (ASR) performance, and measurements of prosody attributes. Experimental results show that the proposed model significantly improves the naturalness in random sample generation. Furthermore, initial experiments demonstrate that randomly sampling from the proposed model can be used as data augmentation to improve the ASR performance. Guangzhi Sun, Yu Zhang 0033, Ron J. Weiss, Yuan Cao 0007, Heiga Zen, Andrew Rosenberg, Bhuvana Ramabhadran |
ICASSP | 6 |
| 2020 | Improving Speech Recognition Using Consistent Predictions on Synthesized SpeechabstractSpeech synthesis has advanced to the point of being close to indistinguishable from human speech. However, efforts to train speech recognition systems on synthesized utterances have not been able to show that synthesized data can be effectively used to augment or replace human speech. In this work, we demonstrate that promoting consistent predictions in response to real and synthesized speech enables significantly improved speech recognition performance. We also find that training on 460 hours of LibriSpeech augmented with 500 hours of transcripts (without audio) performance is within 0.2% WER of a system trained on 960 hours of transcribed audio. This suggests that with this approach, when there is sufficient text available, reliance on transcribed audio can be cut nearly in half. Gary Wang, Andrew Rosenberg, Zhehuai Chen, Yu Zhang 0033, Bhuvana Ramabhadran, Pedro J. Moreno 0001 |
ICASSP | 2 |
| 2020 | Improving Speech Recognition Using GAN-Based Speech Synthesis and Contrastive Unspoken Text Selection
Zhehuai Chen, Andrew Rosenberg, Yu Zhang 0033, Gary Wang, Bhuvana Ramabhadran, Pedro J. Moreno 0001 |
INTERSPEECH | 2 |
| 2020 | SCADA: Stochastic, Consistent and Adversarial Data Augmentation to Improve ASR
Gary Wang, Andrew Rosenberg, Zhehuai Chen, Yu Zhang 0033, Bhuvana Ramabhadran, Pedro J. Moreno 0001 |
INTERSPEECH | 2 |
| 2019 | Speech Recognition with Augmented Synthesized SpeechabstractRecent success of the Tacotron speech synthesis architecture and its variants in producing natural sounding multi-speaker synthesized speech has raised the exciting possibility of replacing expensive, manually transcribed, domain-specific, human speech that is used to train speech recognizers. The multi-speaker speech synthesis architecture can learn latent embedding spaces of prosody, speaker and style variations derived from input acoustic representations thereby allowing for manipulation of the synthesized speech. In this paper, we evaluate the feasibility of enhancing speech recognition performance using speech synthesis using two corpora from different domains. We explore algorithms to provide the necessary acoustic and lexical diversity needed for robust speech recognition. Finally, we demonstrate the feasibility of this approach as a data augmentation strategy for domain-transfer. We find that improvements to speech recognition performance is achievable by augmenting training data with synthesized material. However, there remains a substantial gap in performance between recognizers trained on human speech those trained on synthesized speech. Andrew Rosenberg, Yu Zhang 0033, Bhuvana Ramabhadran, Ye Jia, Pedro J. Moreno 0001, Zelin Wu |
ASRU | 1 |
| 2019 | Comparison of Data Augmentation and Adaptation Strategies for Code-switched Automatic Speech RecognitionabstractCode-switching occurs when the speaker alternates between two or more languages or dialects. It is a pervasive phenomenon in most Indic spoken languages. Code-switching poses a challenge in language modeling as it complicates the orthographic realization of text, and generally, there is a shortage of code-switched data. In this paper, we investigate data augmentation and adaptation strategies for language modeling. Using Bengali and English as an example, we study augmenting the code-switched transcripts with separate transliterated Bengali and English corpora. We present results on two speech recognition tasks, namely, voice search and dictation. We show improvements on both tasks with Maximum Entropy (MaxEnt) and Long Short-Term Memory (LSTM) language models (LMs). We also explore different adaptation strategies for MaxEnt LM and LSTM LM, demonstrating that the transliteration-based data-augmented LSTM LM matches the adapted MaxEnt LM which is trained on more Bengali-English data. Bhuvana Ramabhadran, Jesse Emond, Andrew Rosenberg, Fadi Biadsy |
ICASSP | 4 |
| 2019 | Learning to Speak Fluently in a Foreign Language: Multilingual Speech Synthesis and Cross-Language Voice CloningabstractWe present a multispeaker, multilingual text-to-speech (TTS) synthesis model based on Tacotron that is able to produce high quality speech in multiple languages.Moreover, the model is able to transfer voices across languages, e.g.synthesize fluent Spanish speech using an English speaker's voice, without training on any bilingual or parallel examples.Such transfer works across distantly related languages, e.g.English and Mandarin.Critical to achieving this result are: 1. using a phonemic input representation to encourage sharing of model capacity across languages, and 2. incorporating an adversarial loss term to encourage the model to disentangle its representation of speaker identity (which is perfectly correlated with language in the training data) from the speech content.Further scaling up the model by training on multiple speakers of each language, and incorporating an autoencoding input to help stabilize attention during training, results in a model which can be used to consistently synthesize intelligible speech for training speakers in all languages seen during training, and in native or foreign accents. Yu Zhang 0033, Ron J. Weiss, Heiga Zen, R. J. Skerry-Ryan, Ye Jia, Andrew Rosenberg, Bhuvana Ramabhadran |
INTERSPEECH | 8 |
| 2018 | Measuring the Effect of Linguistic Resources on Prosody Modeling for Speech SynthesisabstractThe generation of natural and expressive prosodic contours is an important component of a text-to-speech (TTS) system which, in most classical architectures, relies on the existence of a text-analysis processor that can extract prosody-predictive features and pass them to a statistical learning model. These features can range from basic properties of the input string to rich high-level features which may not be always available when developing a TTS system in a new language with sparse computational resources. In this work we investigate how the prosody model of a speech-synthesis system performs as a function of different predictive feature sets that assume access to a certain amount of rich resources. We investigate, using objective metrics, the effect of relaxing the assumptions on input representations for prosody prediction for 5 languages, and evaluate the perceptual implications for US English. Andrew Rosenberg, Raul Fernandez, Bhuvana Ramabhadran |
ICASSP | 1 |
| 2018 | Joint Modeling of Accents and Acoustics for Multi-Accent Speech RecognitionabstractThe performance of automatic speech recognition systems degrades with increasing mismatch between the training and testing scenarios. Differences in speaker accents are a significant source of such mismatch. The traditional approach to deal with multiple accents involves pooling data from several accents during training and building a single model in multi-task fashion, where tasks correspond to individual accents. In this paper, we explore an alternate model where we jointly learn an accent classifier and a multi-task acoustic model. Experiments on the American English Wall Street Journal and British English Cambridge corpora demonstrate that our joint model outperforms the strong multi-task acoustic model baseline. We obtain a 5.94% relative improvement in word error rate on British English, and 9.47% relative improvement on American English. This illustrates that jointly modeling with accent information improves acoustic model performance. Xuesong Yang, Kartik Audhkhasi, Andrew Rosenberg, Samuel Thomas 0001, Bhuvana Ramabhadran, Mark Hasegawa-Johnson |
ICASSP | 3 |
| 2018 | Data Augmentation Improves Recognition of Foreign Accented Speech
Takashi Fukuda, Raul Fernandez, Andrew Rosenberg, Samuel Thomas 0001, Bhuvana Ramabhadran, Alexander Sorin, Gakuto Kurata |
INTERSPEECH | 3 |
| 2018 | Interpersonal Relationship Labels for the CALLHOME Corpus
Denys Katerenchuk, David Guy Brizan, Andrew Rosenberg |
LREC | 3 |
| 2018 | Comparing Prosodic Frameworks: Investigating the Acoustic-Symbolic Relationship in ToBI and RaPabstractToBI is the dominant tool for symbolically describing prosodic content in American English speech material. This is due to its descriptive power and its theoretical grounding, but also to the amount of available annotated data. Recently, a modest amount of material annotated with the Rhythm and Pitch (RaP) framework was released publicly. In this paper, we investigate the acoustic-symbolic relationship under these two systems. We present experiments looking at this relationship in both directions. From acoustic to symbolic, we compare the automatic prediction of prosodic prominence as defined under the two systems. From symbolic to acoustic, we examine the utility of these annotation standards to correctly prescribe the acoustics of a given utterance from their symbolic sequences. We find RaP to be promising, showing a somewhat stronger acoustic-symbolic relationship than ToBI given a comparable amount of data for some aspects of these tasks. While with more annotated data ToBI results are stronger, it remains to be shown whether RaP performance can scale up. Raul Fernandez, Andrew Rosenberg |
SLT | 2 |
| 2017 | Investigating native and non-native English classification and transfer effects using Legendre polynomial coefficient clusteringabstractIn this paper, we investigate similarities and differences in pitch contours among native English speakers and non-native English speakers (whose first language is Mandarin). In particular, we investigate if there are particular prosodic contours that are predictive of native and non-native English speech in the area of question intonation contours. We also look to see if we find evidence of negative transfer effects or second language learning effects around native Mandarin speakers who may be using Mandarin prosody when speaking English. To investigate these questions, we explore prosodic contour modeling techniques for native and non-native English speech by clustering Legendre polynomial coefficients. Our results show evidence of non-native English speakers using unexpected contours in the place of expected English prosody. We additionally find support that speakers in our corpus may be experiencing negative language transfer effects, as well as second language learning effects. Rachel Rakov, Andrew Rosenberg |
ASRU | 2 |
| 2017 | End-to-end ASR-free keyword search from speechabstractEnd-to-end (E2E) systems have achieved competitive results compared to conventional hybrid hidden Markov model (HMM)-deep neural network based automatic speech recognition (ASR) systems. Such E2E systems are attractive due to the lack of dependence on alignments between input acoustic and output grapheme or HMM state sequence during training. This paper explores the design of an ASR-free end-to-end system for text query-based keyword search (KWS) from speech trained with minimal supervision. Our E2E KWS system consists of three sub-systems. The first sub-system is a recurrent neural network (RNN)-based acoustic auto-encoder trained to reconstruct the audio through a finite-dimensional representation. The second sub-system is a character-level RNN language model using embeddings learned from a convolutional neural network. Since the acoustic and text query embeddings occupy different representation spaces, they are input to a third feed-forward neural network that predicts whether the query occurs in the acoustic utterance or not. This E2E ASR-free KWS system performs respectably despite lacking a conventional ASR system and trains much faster. Kartik Audhkhasi, Andrew Rosenberg, Abhinav Sethy, Bhuvana Ramabhadran, Brian Kingsbury |
ICASSP | 2 |
| 2017 | Knowledge distillation across ensembles of multilingual models for low-resource languagesabstractThis paper investigates the effectiveness of knowledge distillation in the context of multilingual models. We show that with knowledge distillation, Long Short-Term Memory(LSTM) models can be used to train standard feed-forward Deep Neural Network (DNN) models for a variety of low-resource languages. We then examine how the agreement between the teacher's best labels and the original labels affects the student model's performance. Next, we show that knowledge distillation can be easily applied to semi-supervised learning to improve model performance. We also propose a promising data selection method to filter un-transcribed data. Then we focus on knowledge transfer among DNN models with multilingual features derived from CNN+DNN, LSTM, VGG, CTC and attention models. We show that a student model equipped with better input features not only learns better from the teacher's labels, but also outperforms the teacher. Further experiments suggest that by learning from each other, the original ensemble of various models is able to evolve into a new ensemble with even better combined performance. Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, Tom Sercu, Kartik Audhkhasi, Abhinav Sethy, Markus Nußbaum-Thom, Andrew Rosenberg |
ICASSP | 9 |
| 2017 | Voice-transformation-based data augmentation for prosodic classificationabstractIn this work we explore data-augmentation techniques for the task of improving the performance of a supervised recurrent-neural-network classifier tasked with predicting prosodic-boundary and pitch-accent labels. The technique is based on applying voice transformations to the training data that modify the pitch baseline and range, as well as the vocal-tract and vocal-source characteristics of the speakers to generate further training examples. We demonstrate the validity of the approach by improving performance when the amount of base labeled examples is small (showing reductions in the range of 7%–12% for reduced-data conditions) as well as in terms of its generalization to speakers unseen in the training set (showing a relative reduction in the error rate of 8.74% and 4.75%, on the average, for boundaries and accent tasks respectively, in leave-one-speaker-out validation). Raul Fernandez, Andrew Rosenberg, Alexander Sorin, Bhuvana Ramabhadran, Ron Hoory |
ICASSP | 2 |
| 2017 | End-to-end speech recognition and keyword search on low-resource languagesabstractIn recent years, so-called, “end-to-end” speech recognition systems have emerged as viable alternatives to traditional ASR frameworks. Keyword search, localizing an orthographic query in a speech corpus, is typically performed by using automatic speech recognition (ASR) to generate an index. Previous work has evaluated the use of end-to-end systems for ASR on well known corpora (WSJ, Switchboard, TIMIT, etc.) in high-resource languages like English and Mandarin. In this work, we investigate the use of Connectionist Temporal Classification (CTC) networks, recurrent encoder-decoders with attention, two end-to-end ASR systems for keyword search and speech recognition on low resource languages. We find end-to-end systems can generate high quality 1-best transcripts on low-resource languages, but, because they generate very sharp posteriors, their utility is limited for KWS. We explore a number of ways to address this limitation with modest success. Experimental results reported are based on the IARPA BABEL OP3 languages and evaluation framework. This paper represents the first results using “end-to-end” techniques for speech recognition and keyword search on low-resource languages. Andrew Rosenberg, Kartik Audhkhasi, Abhinav Sethy, Bhuvana Ramabhadran, Michael Picheny |
ICASSP | 1 |
| 2017 | Active learning for low-resource speech recognition: Impact of selection size and language modeling dataabstractActive learning aims to reduce the time and cost of developing speech recognition systems by selecting for transcription highly informative subsets from large pools of audio data. Previous evaluations at OpenKWS and IARPA BABEL have investigated data selection for low-resource languages in very constrained scenarios with 2-hour data selections given a 1-hour seed set. We expand on this to investigate what happens with larger selections and fewer constraints on language modeling data. Our results, on four languages from the final BABEL OP3 period, show that active learning is helpful at larger selections with consistent gains up to 14 hours. We also find that the impact of additional language model data is orthogonal to the impact of the active learning selection criteria. Ali Raza Syed, Andrew Rosenberg, Michael I. Mandel |
ICASSP | 2 |
| 2017 | Weakly-Supervised Phrase Assignment from Text in a Speech-Synthesis System Using Noisy Labels
Asaf Rendel, Raul Fernandez, Zvi Kons, Andrew Rosenberg, Ron Hoory, Bhuvana Ramabhadran |
INTERSPEECH | 4 |
| 2017 | Bias and Statistical Significance in Evaluating Speech Synthesis with Mean Opinion Scores
Andrew Rosenberg, Bhuvana Ramabhadran |
INTERSPEECH | 1 |
| 2017 | Utilizing overt and latent linguistic structure to improve keystroke-based authentication
Adam Goodkind, David Guy Brizan, Andrew Rosenberg |
Image Vis. Comput. | 3 |
| 2016 | Hierarchy Prediction in Online CommunitiesabstractWith the development of the Internet, a big part of social interactions have moved online, and people have unconsciously brought their daily communicational habits to the web. Understanding these communications is important because it will lead to a better understanding of online communities, and can improve areas such as e-commerce, advertisement, topic modeling, security, and others. We propose to develop a natural language based ranking algorithm to predict user influence levels in online communication groups. Denys Katerenchuk, Andrew Rosenberg |
AAAI | 2 |
| 2016 | Supervised and unsupervised active learning for automatic speech recognition of low-resource languagesabstractAutomatic speech recognition (ASR) systems rely on large quantities of transcribed acoustic data. The collection of audio data is relatively cheap, whereas the transcription of that data is relatively expensive. Thus there is an interest in the ASR community in active learning, in which only a small subset of highly representative data chosen from a large pool of untranscribed audio need be transcribed in order to approach the performance of the system trained with much larger amounts of transcribed audio. In this paper, we compare two basic approaches to active learning: a supervised approach in which we build a speech recognition system from a small amount of seed data in order to make the selection of a limited amount of additional audio for transcription, and an unsupervised approach in which no intermediate system recognition system built from seed data is necessary. Our best unsupervised approach performs quite close to our supervised approach, with both outperforming a random selection scheme. Ali Raza Syed, Andrew Rosenberg, Ellen Eide |
ICASSP | 2 |
| 2016 | Automatically Classifying Self-Rated Personality Scores from Speech
Guozhen An, Sarah Ita Levitan, Rivka Levitan, Andrew Rosenberg, Michelle Levine, Julia Hirschberg |
INTERSPEECH | 4 |
| 2016 | Combining Acoustic-Prosodic, Lexical, and Phonotactic Features for Automatic Deception Detection
Sarah Ita Levitan, Guozhen An, Rivka Levitan, Andrew Rosenberg, Julia Hirschberg |
INTERSPEECH | 5 |
| 2016 | RankDCG: Rank-Ordering Evaluation Measure
Denys Katerenchuk, Andrew Rosenberg |
LREC | 2 |
| 2015 | Automatic recognition of unified parkinson's disease rating from speech with acoustic, i-vector and phonotactic features
Guozhen An, David Guy Brizan, Michelle Morales, Ali Raza Syed, Andrew Rosenberg |
INTERSPEECH | 6 |
| 2015 | Modeling phrasing and prominence using deep recurrent learning
Andrew Rosenberg, Raul Fernandez, Bhuvana Ramabhadran |
INTERSPEECH | 1 |
| 2015 | Utilizing linguistically enhanced keystroke dynamics to predict typist cognition and demographics
David Guy Brizan, Adam Goodkind, Patrick Koch, Kiran S. Balagani, Vir V. Phoha, Andrew Rosenberg |
Int. J. Hum. Comput. Stud. | 6 |
| 2014 | Using word burst analysis to rescore keyword search candidates on low-resource languagesabstractFor low-resource languages, keyword search (KWS) remains challenging due to the lack of training data. This work aims to bolster KWS performance in low-resource languages by incorporating word burst information into the decision process. We find that this information can improve performance when we focus analysis on particularly problematic KWS candidates: low-scoring correct hits, and high-scoring false alarms. Justin Richards, Andrew Rosenberg |
ICASSP | 3 |
| 2014 | Rescoring Confusion Networks for Keyword SearchabstractWe introduce a two-stage cascaded scheme to rescore Confusion Networks (CNs) for Keyword Search in the context of Low-Resource Languages. In the first stage we rescore the CN to improve the error rate of the 1-best hypothesis using a large number of lexical, phonetic, false alarms and structural features. Using a rank learning Support Vector Machine classifier, we obtain WER gains between 0.54% and 2.84% on Cantonese, Tagalog, Turkish, Pashto and Vietnamese. In the second stage we generate keyword hits from the rescored CN and use logistic regression to detect true hits and false alarms. We compare these to hits generated from the unrescored CN and obtain gains between 0.45% and 0.9% on the MTWV metric by using the mentioned features and including acoustic and prosodic features on Tagalog, Turkish and Pashto. Victor Soto, Erica Cooper, Lidia Mangu, Andrew Rosenberg, Julia Hirschberg |
ICASSP | 4 |
| 2014 | Continuous authentication with cognition-centric text production and revision featuresabstractMost continuous user authentication techniques based on typing behavior rely on the keystroke dynamics or on the linguistic style of the user. However, there is a rich spectrum of cognition-centric behavioral traits that a typist exhibits during different stages of text production (e.g., composition, translation, and revision), which to our knowledge, have not been considered for continuous authentication. We study the continuous authentication performance of 123 behavioral traits extracted from discrete cognitive units called bursts. We performed experiments on typing data collected from 486 volunteer subjects. Our findings include: (1) features from bursts delimited by pause events have significantly higher availability and authentication performance compared to bursts delimited by revision events; (2) bursts with pause durations of at least one second provide the best authentication accuracy and availability; and (3) fusing our features with traditional keystroke dynamics features reduced authentication error rates. We achieved an equal error rate between 13.37 and 4.55 percent for authentication windows as low as 30 seconds to 3.5 minutes. Hilbert Locklear, Sathya Govindarajan, Zdenka Sitova, Adam Goodkind, David Guy Brizan, Andrew Rosenberg, Vir V. Phoha, Paolo Gasti, Kiran S. Balagani |
IJCB | 6 |
| 2014 | Improving deep neural network acoustic modeling for audio corpus indexing under the IARPA babel program
Brian Kingsbury, Jia Cui, Bhuvana Ramabhadran, Andrew Rosenberg, Mohammad Sadegh Rasooli, Owen Rambow, Nizar Habash, Vaibhava Goel |
INTERSPEECH | 5 |
| 2014 | Recent improvements in neural network acoustic modeling for LVCSR in low resource languages
Jia Cui, Bhuvana Ramabhadran, Andrew Rosenberg, Brian Kingsbury, Abhinav Sethy |
INTERSPEECH | 4 |
| 2014 | Exploiting vocal-source features to improve ASR accuracy for low-resource languages
Raul Fernandez, Jia Cui, Andrew Rosenberg, Bhuvana Ramabhadran |
INTERSPEECH | 3 |
| 2014 | "was that your mother on the phone?": classifying interpersonal relationships between dialog participants with lexical and acoustic propertiesabstractUnderstanding interpersonal relationships provides important context in understanding spoken communication. In addition to increasing knowledge of the social indicators in spoken communication, the automatic recognition of interpersonal relationships has an application in providing structure to social networks. This paper presents exploratory work on the challenging problem of distinguishing family from friends in spontaneous dialogs drawn from the CALLHOME English corpus. We find both acoustic/prosodic and lexical features useful in classifying these relationships. In binary classification experiments, we achieve accuracy of 10.71% absolute improvement over chance (50%) assignment. Index Terms: speech understanding, interpersonal relationships, information extraction Denys Katerenchuk, David Guy Brizan, Andrew Rosenberg |
INTERSPEECH | 3 |
| 2014 | Improving named entity recognition with prosodic featuresabstractIn natural language processing (NLP) the problem of named entity (NE) recognition in speech is well known, yet remains a challenge where performance is dependent on automatic speech recognition (ASR) system error rates. NEs are often foreign or out-of-vocabulary (OOV) words, leaving conventional ASR systems unable to recognize them. In our research, we improve a CRF-based NE recognition system by incorporating two styles of prosodic features, hypothesized ToBI labels and unsupervised clusters of acoustic features. ToBI-based features improve NE recognition by 6% absolute (F1:0.39 v.s. F1: 0.45) on automatically recognized spontaneous speech from ACE’05. Denys Katerenchuk, Andrew Rosenberg |
INTERSPEECH | 2 |
| 2014 | Strategies for rescoring keyword search results using word-burst and acoustic featuresabstractThe identification of keyword queries in speech data from lowresources languages poses a challenge for current methods as speech recognition algorithms lack sufficient training data to produce high accuracy transcript. To compensate for these shortcomings, we extract signals from the data that are useful in keyword identification but are not being used by the speech recognizer. These signals take multiple forms — word burstiness, rescored confusion network posteriors and acoustic/prosodic qualities. The former denotes the tendency for keywords to occur in bursts within a conversational topic. We employ three different strategies to exploit this information: 1) a four-way classification of keyword hypotheses that targets low-scoring correct hits and high-scoring false alarms, 2) ranking algorithms, and 3) a direct adjustment of keyword hit scores based on hypothesized repetition. We find that interpolating the results of these three strategies in an ensemble provides a reliable way to improve the results of keyword search. Justin Richards, Victor Soto, Julia Hirschberg, Andrew Rosenberg |
INTERSPEECH | 5 |
| 2014 | A comparison of multiple methods for rescoring keyword search lists for low resource languagesabstractWe review the performance of a new two-stage cascaded machine learning approach for rescoring keyword search output for low resource languages. In the first stage Confusion Networks (CNs) are rescored for improved Automatic Speech Recognition (ASR) by reranking the arcs of each confusion bin. In the second stage we generate keyword search hypotheses from the rescored ASR output and rescore them using logistic regression classifiers to detect true hits and false alarms. We compare the performance of our system with state of the art rescoring techniques, including probability of false alarm normalization, exponential normalization, rank-normalized posterior scores and sum-to-one normalization and show promising results. Experimental validation is performed using the Term Weighted Value (TWV) metric on four corpora from the IARPA-Babel program for keyword search on low resource languages, including Assamese, Bengali, Lao and Zulu. Victor Soto, Lidia Mangu, Andrew Rosenberg, Julia Hirschberg |
INTERSPEECH | 3 |
| 2013 | Cross-language phrase boundary detectionabstractWe describe models of prosodic phrasing trained on multiple languages to identify boundaries in an unseen language. Our goal is to create models from High Resource languages, in which hand-annotated prosodic phrase boundaries are available, to use in identifying boundaries in a Low Resource language, with little or no training material. We train models on American English, Italian, Mandarin, and German and test on each of these languages. We find that, while pause is the most important feature for phrase boundary prediction in all languages examined, the role of pause in boundary identification varies by annotator and the relative importance of other features varies significantly by language. We also find that different acoustic correlates of prosodic boundaries characterize different languages. In some, the relative importance of features is silence > pitch > intensity > duration, while for other languages intensity is more important than pitch. These differences do not appear to be attributable to language family, since, e.g. English and German display different patterns. Victor Soto, Erica Cooper, Andrew Rosenberg, Julia Hirschberg |
ICASSP | 3 |
| 2013 | Detecting laughter and filled pauses using syllable-based featuresabstractIdentifying laughter and filled pauses is important to understanding spontaneous human speech. These are two common vocal expressions that are non-lexical and incredibly communicative. In this paper, we use a two-tiered system for identifying laughter and filled pauses. We first generate frame level hypotheses and subsequently rescore these based on features derived from acoustic syllable segmentation. Using Interspeech 2013 ComParE challenge corpus, SVC, we find that these rescoring experiments and inclusion of syllable based acoustic/prosodic features allow for the detection of laughter and filled pauses by at 89.3 % UAAUC on the development set, an improvement of 1.7 % over the challenge baseline. Index Terms: laughter detection, filled pause detection, prosodic analysis Guozhen An, David Guy Brizan, Andrew Rosenberg |
INTERSPEECH | 3 |
| 2013 | Let me finish: automatic conflict detection using speaker overlap
Félix Grèzes, Justin Richards, Andrew Rosenberg |
INTERSPEECH | 3 |
| 2013 | "sure, i did the right thing": a system for sarcasm detection in speechabstractWhile a fair amount of work has been done on automatically detecting emotion in human speech, there has been little research on sarcasm detection. Although sarcastic speech acts are inherently subjective, humans have relatively clear intuitions as to what constitutes sarcastic speech. In this paper, we present a system for automatic sarcasm detection. Using a new acted speech corpus that is annotated for sarcastic and sincere speech, we examine a number of features that are indicative of sarcasm. The first set of features looks at a baseline of basic acoustic features that have been found to be helpful in human sarcasm identification. We then present an effective way of modeling and applying prosodic contours to the task of automatic sarcasm detection. This approach applies sequential modeling to categorical representations of pitch and intensity contours obtained via k-means clustering. Using a SimpleLogistic (LogitBoost) classifier, we are able to predict sarcasm with 81.57% accuracy. This result suggests that certain pitch and intensity contours are predictive of sarcastic speech. Rachel Rakov, Andrew Rosenberg |
INTERSPEECH | 2 |
| 2013 | Modeling prosodic sequences with k-means and dirichlet process GMMsabstractIn this paper we describe two unsupervised representations of prosodic sequences based on k-means and Dirichlet Process Gaussian Mixture Model (DPGMM) clustering. The clustering algorithms are used to infer an inventory of prosodic categories over automatically segmented syllables. A tri-gram model is trained over these sequences to characterize speech. We find that DPGMM clusters show a greater correspondence with manual ToBI labels than k-means clusters. However, sequence models trained on k-means clusters significantly outperform DPGMM sequences in classifying speaking style, nativeness and speakers. We also investigate the use of these sequence models in the detection of outliers regarding these three tasks. Non-parametric Bayesian techniques have the advantage of being able to learn a clustering solution and infer the number of clusters directly from data. While it is attractive to avoid specifying k before clustering, on the tasks of characterizing prosodic sequences we find that effective use of DPGMMs still requires a significant amount of parameter tuning, and performance fails to reach the level of k-means. Andrew Rosenberg |
INTERSPEECH | 1 |
| 2013 | Automatic detection of speaker state: Lexical, prosodic, and phonetic approaches to level-of-interest and intoxication classification
William Yang Wang, Fadi Biadsy, Andrew Rosenberg, Julia Hirschberg |
Comput. Speech Lang. | 3 |
| 2012 | Power Mean Pyramid Scores for Summarization EvaluationabstractWe present Power Mean Pyramid Scores (PMP), an evaluation metric that extends the Pyramid evaluation scheme for summarization by combining Sentence Content Units (SCU) scores using Power Mean. The Pyramid method generates a summarization score by linearly combining component SCU scores. We find that by combining SCU scores using Power Mean, we can optimize a single parameter, α, leading to significantly improved correlation with human judgements. We demonstrate this result through an empirical study based on TAC-08 evaluation. 1. Sameer Maskey, Andrew Rosenberg |
INTERSPEECH | 2 |
| 2012 | Rethinking The Corpus: Moving towards Dynamic Linguistic Resources
Andrew Rosenberg |
INTERSPEECH | 1 |
| 2012 | Using Prominence and Phrasing Predictions to Improve Weighted Dictionary Pronunciation ModelsabstractProsody impacts the pronunciation of lexical items in a number of ways. Accented syllables tend to be pronounced with their canonical (dictionary) vowel. Deaccented vowels are more likely to be reduced. Coarticulatory influences rarely span intonational phrase boundaries. In this work, we investigate the use of automatically generated prosodic hypotheses to improve a weighted dictionary pronunciation model. We use the phonemically transcribed, Buckeye Corpus for this investigation. We find that predictions of pitch accent and intonational phrase boundaries can be used to lower pronunciation model perplexity. Andrew Rosenberg |
INTERSPEECH | 1 |
| 2012 | Classifying Skewed Data: Importance Weighting to Optimize Average RecallabstractPromoted in part by its use in the Interspeech Challenges in 2009-2012, Average Recall has emerged as an attractive evaluation measure of classifier performance where the data has a skewed class distribution. In this paper, we show that importance weighting can be used to optimize Average Recall directly. We compare this approach to sampling techniques that have been previously used to classify skewed data. We demonstrate the use of this approach on the Interspeech 2009 Emotion Challenge tasks, and prosodic analysis tasks. Andrew Rosenberg |
INTERSPEECH | 1 |
| 2012 | Phrase Boundary Assignment from Text in Multiple DomainsabstractDetecting and modeling proper phrasing from an input text string is an important aspect when producing synthesis that sounds intelligible and natural. Knowledge of proper phrase structure influences, e.g., the placement and length of pauses, and the realization of phrase-final boundary contours, both of which can have an effect in a listener’s percepts ranging from naturalness to semantic interpretation. In this work, we look at modeling the occurrence, and types, of phrase breaks from purely textual features, paying close attention to how the performance of the systems generalizes inand out-of-domain for corpora of various types (such as broadcast news, spontaneous speech, and synthesis databases), and as a function of various subsets of syntactical and lexical features investigated. Andrew Rosenberg, Raul Fernandez, Bhuvana Ramabhadran |
INTERSPEECH | 1 |
| 2012 | Analysis of speech transcripts to predict winners of U.S. Presidential and Vice-Presidential debatesabstractIn this paper, we describe investigations into the speech used in American Presidential and Vice-Presidential debates. We explore possible transcript-based features that may correlate with personally appealing or politically persuasive language. We identify, with chi-squared analysis, features that correlate with success in the debates. We find that with a set of surface-level features from historical debates, we can predict the winners of presidential debates with success moderately above chance. Ian Kaplan, Andrew Rosenberg |
SLT | 2 |
| 2012 | Modeling intensity contours and the interaction of pitch and intensity to improve automatic prosodic event detection and classificationabstractProsody, or the way words are spoken, carries important information to understanding a speaker's communicative intention. Many studies on automatic prosodic analysis focus on parameterizing pitch content. In this work, we extend previous pitch contour modeling features to intensity contours, and develop a set of features based on the interaction of pitch and intensity. These new features improve the state-of-the-art on all prosodic event detection and classification tasks related to automatic ToBI labeling. Andrew Rosenberg |
SLT | 1 |
| 2011 | Evaluating importance of facial expression in american sign language and pidgin signed english animationsabstractAnimations of American Sign Language (ASL) and Pidgin Signed English (PSE) have accessibility benefits for many signers with lower levels of written language literacy. In prior experimental studies we conducted evaluating animations of ASL, native signers gave informal feedback in which they critiqued the insufficient and inaccurate facial expressions of the virtual human character. While face movements are important for conveying grammatical and prosodic information in human ASL signing, no empirical evaluation of their impact on the understandability and perceived quality of ASL animations had previously been conducted. To quantify the suggestions of deaf participants in our prior studies, we experimentally evaluated ASL and PSE animations with and without various types of facial expressions, and we found that their inclusion does lead to measurable benefits for the understandability and perceived quality of the animations. This finding provides motivation for our future work on facial expressions in ASL and PSE animations, and it lays a novel methodological groundwork for evaluating the quality of facial expressions for conveying prosodic or grammatical information. Matt Huenerfauth, Andrew Rosenberg |
ASSETS | 3 |
| 2011 | Multi-objective Genetic Programming for Visual Analytics
Ilknur Icke, Andrew Rosenberg |
EuroGP | 2 |
| 2011 | Intoxication Detection Using Phonetic, Phonotactic and Prosodic Cues
Fadi Biadsy, William Yang Wang, Andrew Rosenberg, Julia Hirschberg |
INTERSPEECH | 3 |
| 2011 | Symbolic and Direct Sequential Modeling of Prosody for Classification of Speaking-Style and NativenessabstractIn this paper, we explore the differences between direct and symbolic sequential modeling of prosody. We use sequential models to characterize speech in two tasks, classifying speaking-style and distinguishing native from non-native speech. We explore the use of a spike-and-slab model to directly model pitch contour data. We find in both of these tasks that sequences of symbolic prosodic events to lead to improved performance over approaches that model pitch contours directly. We also explore the use of hypothesized prosodic events in both tasks. We find the speaking-style results to be robust to automatic annotation, while, when classifying nativeness, the spike-and-slab model leads to better performance. Andrew Rosenberg |
INTERSPEECH | 1 |
| 2011 | Using Mutual Information to Identify Regions of Analysis for Prosodic AnalysisabstractThis paper presents a novel technique for empirically identifying regions of analysis for time/value information. The technique relies on analysis of mutual information between the contour, and some variable of interest. We present the use of this technique in the analysis of prosody in American English speech, where we identify valuable regions of analysis for the classification of phrase ending intonation. We also use the technique to investigate the most informative region of analysis for pitch accent detection. Andrew Rosenberg |
INTERSPEECH | 1 |
| 2011 | "What is... Dengue Fever?" - Modeling and Predicting Pronunciation Errors in a Text-to-Speech SystemabstractWe propose a system to predict baseform-generation errors in a text-to-speech (TTS) front-end, and aid in the process of cus-tomizing the synthesis engine to a novel application with a large, open-ended vocabulary. We motivate the use of the sys-tem by using data collected during the deployment of the IBM TTS engine in the Watson Deep Question-Answering system customized to play a game of Jeopardy!. We propose a set of features derived from a lexeme’s orthography and candidate baseform, and use a variety of learning schemes and data sam-pling algorithms to address the issue of skewed class priors in the training data. We show that 1) these different approaches provide complementary information that can then be exploited by fusion schemes to improve on the baseline performances, and 2) it is possible to use these techniques to retrieve a list of likely incorrect lexemes so as to reduce the number of tokens that must be vetted before finding and fixing an error. Index Terms: front-end error modeling, speech synthesis 1. Andrew Rosenberg, Raul Fernandez, Bhuvana Ramabhadran |
INTERSPEECH | 1 |
| 2010 | AutoBI - a tool for automatic toBI annotationabstractThis paper describes the AuToBI tool for automatic gener-ation of hypothesized ToBI labels. While research on automatic prosodic annotation has been conducted for many years, Au-ToBI represents the first publicly available tool to automatically detect and classify the breaks and tones that make up the ToBI annotation standard. This paper describes the feature extraction routines as well as the classifiers used to detect and classify the prosodic events of the ToBI standard. Additionally, we report performance evaluating AuToBI models trained on the Boston Directions Corpus on the Columbia Games Corpus. By evaluat-ing on distinct speakers domains and recording conditions, this evaluation represents an accurate representation of the perfor-mance of the system when applied to novel spoken material. Index Terms: prosody, automatic prosody annotation, tools 1. Andrew Rosenberg |
INTERSPEECH | 1 |
| 2010 | Classification of Prosodic Events using Quantized Contour Modeling
Andrew Rosenberg |
HLT-NAACL | 1 |
| 2009 | Charisma perception from text and speech
Andrew Rosenberg, Julia Hirschberg |
Speech Commun. | 1 |
| 2008 | Intonational phrases for speech summarizationabstractExtractive speech summarization approaches select relevant segments of spoken documents and concatenate them to generate a summary. The extraction unit chosen, whether a sentence, syntactic constituent, or other segment, has a significant impact on the overall quality and fluency of the summary. Even though sentences tend to be the choice of most the extractive speech summarizers, in this paper, we present the results of an empirical study indicating that intonational phrases are better units of extraction for summarization. Our study compared four types of input segmentation: sentences, two pause-based segmentation, and intonational phrases (IP). We found that IPs are the best candidates for extractive summarization, improving over the second highest-performing approach, sentence-based summarization, by 8.2% F-measure. Sameer Maskey, Andrew Rosenberg, Julia Hirschberg |
INTERSPEECH | 2 |
| 2007 | V-Measure: A Conditional Entropy-Based External Cluster Evaluation Measure
Andrew Rosenberg, Julia Hirschberg |
EMNLP-CoNLL | 1 |
| 2007 | Comparing american and palestinian perceptions of charisma using acoustic-prosodic and lexical analysisabstractCharisma, the ability to lead by virtue of personality alone, is difficult to define but relatively easy to identify. However, cultural factors clearly affect perceptions of charisma. In this paper we compare results from parallel perception studies investigating charismatic speech in Palestinian Arabic and American English. We examine acoustic/prosodic and lexical correlates of charisma ratings to determine how the two cultures differ with respect to their views of charismatic speech. Fadi Biadsy, Julia Hirschberg, Andrew Rosenberg, Wisam Dakka |
INTERSPEECH | 3 |
| 2007 | Detecting pitch accent using pitch-corrected energy-based predictorsabstractPrevious work has shown that the energy components of frequency subbands with a variety of frequencies and bandwidths predict pitch accent with various degrees of accuracy, and produce correct predictions for distinct subsets of data points. In this paper, we describe a series of experiments exploring techniques to leverage the predictive power of these energy components by including pitch and duration features – other known correlates to pitch accent. We perform these experiments on Standard American English read, spontaneous and broadcast news speech, each corpus containing at least four speakers. Using an approach by which we correct energy-based predictions using pitch and duration information prior to using a majority voting classifier, we were able to detect pitch accent in read, spontaneous and broadcast news speech at 84.0%, 88.3 % and 88.5 % accuracy, respectively. Human performance at pitch accent detection is generally taken to be between 85 % and 90%. Andrew Rosenberg, Julia Hirschberg |
INTERSPEECH | 1 |
| 2007 | Varying input segmentation for story boundary detection in English, Arabic and Mandarin broadcast newsabstractStory segmentation of news broadcasts has been shown to improve the accuracy of the subsequent processes such as question answering and information retrieval. In previous work, a decision tree trained on automatically extracted lexical and acoustic features was trained to predict story boundaries, using hypothesized sentence boundaries to define potential story boundaries. In this paper, we empirically evaluate several alternatives to choice of segmentation on three languages: English, Mandarin and Arabic. Our results suggest that the best performance can be achieved by using 250ms pause-based segmentation or sentence boundaries determined using a very low confidence score threshold. Andrew Rosenberg, Mehrbod Sharifi, Julia Hirschberg |
INTERSPEECH | 1 |
| 2006 | On the correlation between energy and pitch accent in read English speechabstractIn this paper, we describe a set of experiments that examine the correlation between energy and pitch accent.We tested the discriminative power of the energy component of frequency subbands with a variety of frequencies and bandwidths on read speech spoken by four native speakers of Standard American English, using an analysis by classification approach.We found that the frequency region most robust to speaker differences is between 2 and 20 bark.Across all speakers, using only energy features we were able to predict pitch accent in read speech with accuracy of 81.9%. Andrew Rosenberg, Julia Hirschberg |
INTERSPEECH | 1 |
| 2006 | Story Segmentation of Broadcast News in English, Mandarin and Arabic
Andrew Rosenberg, Julia Hirschberg |
HLT-NAACL | 1 |
| 2005 | Acoustic/prosodic and lexical correlates of charismatic speechabstractCharisma, the ability to command authority on the basis of personal qualities, is more difcult to dene than to identify. How do charismatic leaders such as Fidel Castro or Pope John Paul II attract and retain their followers? We present results of an analysis of subjective ratings of charisma from a corpus of American political speech. We identify the associations be- tween charisma ratings and ratings of other personal attributes. We also examine acoustic/prosodic and lexical features of this speech and correlate these with charisma ratings. Andrew Rosenberg, Julia Hirschberg |
INTERSPEECH | 1 |