EDBT 2026 Demo / reviewers in the wild / expert
Samuel Thomas 0001
dblp:99/1552
· DBLP profile ↗
100ranked-venue papers
23as first author
32since 2021 · last 2025
0000-0001-7573-0620ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 97 · 23 first-author · 31 since 2021Artificial intelligence and machine learning · 56 · 11 first-author · 18 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?abstractWe propose Omni-R1 which fine-tunes a recent multi-modal LLM, Qwen2.5-Omni, on an audio question answering dataset with the reinforcement learning method GRPO. This leads to new State-of-the-Art performance on the recent MMAU and MMAR benchmarks. On MMAU, Omni-R1 achieves the highest accuracies on the sounds, music, speech, and overall average categories, both on the Test-mini and Test-full splits. To understand the performance improvement, we tested models both with and without audio and found that much of the performance improvement from GRPO could be attributed to better text-based reasoning. We also made a surprising discovery that fine-tuning without audio on a text-only dataset was effective at improving the audio-based performance. Andrew Rouditchenko, Saurabhchand Bhati, Edson Araujo, Samuel Thomas 0001, Hilde Kuehne, Rogério Feris, James R. Glass |
ASRU | 4 |
| 2025 | Granite-speech: open-source speech-aware LLMs with strong English ASR capabilitiesabstractGranite-speech LLMs are compact and efficient speech language models specifically designed for English ASR1and automatic speech translation (AST). The models were trained by modality aligning granite-3.3-instruct to speech on publicly available open-source corpora. Comprehensive benchmarking on English ASR shows that they outperform several competitors’ models that were trained on orders of magnitude more proprietary data, and they keep pace on English-to-X AST for major European languages, Japanese, and Mandarin. The speech-specific components are: a conformer acoustic encoder using block attention and self-conditioning trained with connectionist temporal classification, a windowed query-transformer speech modality adapter used to do temporal downsampling of the acoustic embeddings and map them to the LLM text embedding space, and LoRA adapters to further fine-tune the text LLM. The models are freely available on HuggingFace2under a permissive Apache 2.0 license.1The latest models (revision 3.3.2) support multilingual ASR in English, French, German, Spanish and Portuguese and bidirectional speech translation to and from English. This paper covers the initial English-only release.2https://huggingface.co/ibm-granite/granite-speech-3.3-2b (and…-8b). George Saon, Avihu Dekel, Alexi Brooks, Tohru Nagano, Abraham Daniels, Aharon Satt, Ashish R. Mittal, Brian Kingsbury, David Haws, Edmilson da Silva Morais, Gakuto Kurata, Hagai Aronowitz, Ibrahim Ibrahim, Hong-Kwang Jeff Kuo, Kate Soule, Luis A. Lastras, Masayuki Suzuki, Ron Hoory, Samuel Thomas 0001, Sashi Novitasari, Takashi Fukuda, Vishal Sunder, Zvi Kons |
ASRU | 19 |
| 2025 | CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained AlignmentabstractRecent advances in audio-visual learning have shown promising results in learning representations across modalities. However, most approaches rely on global audio representations that fail to capture fine-grained temporal correspondences with visual frames. Additionally, existing methods often struggle with conflicting optimization objectives when trying to jointly learn reconstruction and cross-modal alignment. In this work, we propose CAV-MAE Sync as a simple yet effective extension of the original CAV-MAE [14] framework for self-supervised audio-visual learning. We address three key challenges: First, we tackle the granularity mismatch between modalities by treating audio as a temporal sequence aligned with video frames, rather than using global representations. Second, we resolve conflicting optimization goals by separating contrastive and reconstruction objectives through dedicated global tokens. Third, we improve spatial localization by introducing learnable register tokens that reduce the semantic load on patch tokens. We evaluate the proposed approach on AudioSet, VGG Sound, and the ADE20K Sound dataset on zero-shot retrieval, classification, and localization tasks demonstrating state-of-the-art performance and outperforming more complex architectures. Code is available at https://github.com/edsonroteia/cav-mae-sync. Edson Araujo, Andrew Rouditchenko, Yuan Gong 0001, Saurabhchand Bhati, Samuel Thomas 0001, Brian Kingsbury, Leonid Karlinsky, Rogério Feris, James R. Glass, Hilde Kuehne |
CVPR | 5 |
| 2025 | LLM based Text Generation for Improved Low-resource Speech Recognition ModelsabstractLimited transcribed spoken style data is a critical bottleneck in building automatic speech recognition (ASR) systems for low-resource languages. Prompting a large language model (LLM) to paraphrase input text can generate novel text data that is constrained to be semantically similar to the source data. We leverage this capability of LLMs to improve the performance of low-resource ASR systems by increasing the limited text training data while keeping the same spoken style. Since word sequences in the training data are now more diverse and the vocabulary of the ASR model is also expanded, this approach allows for building general purpose ASR without prior knowledge of various domains in the low-resource language. In our experiments with Brazilian Portuguese as a low-resource language, paraphrased data enhanced the n-gram language model (LM) used to build the weighted finite state transducer (WFST) for decoding with a Conformer-CTC speech recognition model, resulting in improvement of word error rate (WER) by 15.6% over the baseline model. Synthesizing the paraphrased text into speech and using it to fine-tune the acoustic model (AM) component helped to further improve the WER by 2.9%, achieving a combined improvement of 18.5%. We also demonstrate the usefulness of our proposed approach for high-resource languages like English. Tohru Nagano, Gakuto Kurata, Samuel Thomas 0001, Hong-Kwang Jeff Kuo, Daniel Bolaños, Hyun Jung, George Saon |
ICASSP | 3 |
| 2025 | A Non-autoregressive Model for Joint STT and TTSabstractIn this paper, we take a step towards jointly modeling automatic speech recognition (STT) and speech synthesis (TTS) in a fully non-autoregressive way. We develop a novel multimodal framework capable of handling the speech and text modalities as input either individually or together. The proposed model can also be trained with unpaired speech or text data owing to its multimodal nature. We further propose an iterative refinement strategy to improve the STT and TTS performance of our model such that the partial hypothesis at the output can be fed back to the input of our model, thus iteratively improving both STT and TTS predictions. We show that our joint model can effectively perform both STT and TTS tasks, outperforming the STT-specific baseline in all tasks and performing competitively with the TTS-specific baseline across a wide range of evaluation metrics. Vishal Sunder, Brian Kingsbury, George Saon, Samuel Thomas 0001, Slava Shechtman, Hagai Aronowitz, Eric Fosler-Lussier, Luis A. Lastras |
ICASSP | 4 |
| 2025 | mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech RecognitionabstractAudio-Visual Speech Recognition (AVSR) combines lip-based video with audio and can improve performance in noise, but most methods are trained only on English data. One limitation is the lack of large-scale multilingual video data, which makes it hard to train models from scratch. In this work, we propose mWhisper-Flamingo for multilingual AVSR which combines the strengths of a pre-trained audio model (Whisper) and video model (AV-HuBERT). To enable better multi-modal integration and improve the noisy multilingual performance, we introduce decoder modality dropout where the model is trained both on paired audio-visual inputs and separate audio/visual inputs. mWhisper-Flamingo achieves state-of-the-art WER on MuAViC, an AVSR dataset of 9 languages. Audio-visual mWhisper-Flamingo consistently outperforms audio-only Whisper on all languages in noisy conditions. Andrew Rouditchenko, Samuel Thomas 0001, Hilde Kuehne, Rogério Feris, James R. Glass |
IEEE Signal Process. Lett. | 2 |
| 2024 | What, When, and Where? Self-Supervised Spatio- Temporal Grounding in Untrimmed Multi-Action Videos from Narrated InstructionsabstractSpatio-temporal grounding describes the task of localizing events in space and time, e.g., in video data, based on verbal descriptions only. Models for this task are usually trained with human-annotated sentences and bounding box supervision. This work addresses this task from a multimodal supervision perspective, proposing a framework for spatio-temporal action grounding trained on loose video and subtitle supervision only, without human annotation. To this end, we combine local representation learning, which focuses on leveraging fine-grained spatial information, with a global representation encoding that captures higher-level representations and incorporates both in a joint approach. To evaluate this challenging task in a real-life setting, a new benchmark dataset is proposed, providing dense spatio-temporal grounding annotations in long, untrimmed, multi-action instructional videos for over 5K events. We evaluate the proposed approach and other methods on the proposed and standard downstream tasks, showing that our method improves over current baselines in various settings, including spatial, temporal, and untrimmed multi-action spatio-temporal grounding. Brian Chen 0001, Nina Shvetsova, Andrew Rouditchenko, Daniel Kondermann, Samuel Thomas 0001, Shih-Fu Chang, Rogério Feris, James R. Glass, Hilde Kuehne |
CVPR | 5 |
| 2024 | Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
Andrew Rouditchenko, Yuan Gong 0001, Samuel Thomas 0001, Leonid Karlinsky, Hilde Kuehne, Rogério Feris, James R. Glass |
INTERSPEECH | 3 |
| 2023 | Effective Training of RNN Transducer Models on Diverse Sources of Speech and Text DataabstractThis paper proposes a novel modeling framework for effective training of end-to-end automatic speech recognition (ASR) models on various sources of data from diverse domains: speech paired with clean ground truth transcripts, speech with noisy pseudo transcripts from semi-supervised decodes and unpaired text-only data. In our proposed approach, we build a recurrent neural network transducer (RNN-T) model with a shared multimodal encoder, multi-branch prediction networks and a shared common joint network. To train on unpaired text-only data sets along with transcribed speech data, the shared encoder is trained to process both speech and text modalities. Differences in data from multiple domains are effectively handled by training a multi-branch prediction network on various different data sets before an interpolation step combines the multi-branch prediction networks back into a computationally-efficient single branch. We show the benefit of our proposed technique on several ASR test sets by comparing our models to those trained by simple data mixing. The technique provides a significant relative improvement of up to 6% over baseline systems operating at a similar decoding cost. Takashi Fukuda, Samuel Thomas 0001 |
ICASSP | 2 |
| 2023 | C2KD: Cross-Lingual Cross-Modal Knowledge Distillation for Multilingual Text-Video RetrievalabstractMultilingual text-video retrieval methods have improved significantly in recent years, but the performance for languages other than English still lags. We propose a Cross-Lingual Cross-Modal Knowledge Distillation method to improve multilingual text-video retrieval. Inspired by the fact that English text-video retrieval outperforms other languages, we train a student model using input text in different languages to match the cross-modal predictions from teacher models using input text in English. We propose a cross entropy based objective which forces the distribution over the student’s text-video similarity scores to be similar to those of the teacher models. We introduce a new multilingual video dataset, Multi-YouCook2, by translating the English captions in the YouCook2 video dataset to 8 other languages. Our method improves multilingual text-video retrieval performance on Multi-YouCook2 and several other datasets such as Multi-MSRVTT and VATEX. We also conducted an analysis on the effectiveness of different multilingual text models as teachers. Andrew Rouditchenko, Yung-Sung Chuang, Nina Shvetsova, Samuel Thomas 0001, Rogério Feris, Brian Kingsbury, Leonid Karlinsky, David F. Harwath, Hilde Kuehne, James R. Glass |
ICASSP | 4 |
| 2023 | Fine-Grained Textual Knowledge Transfer to Improve RNN Transducers for Speech Recognition and UnderstandingabstractRNN Tranducer (RNN-T) technology is very popular for building deployable models for end-to-end (E2E) automatic speech recognition (ASR) and spoken language understanding (SLU). Since these are E2E models operating on speech directly, there remains a potential to improve their performance using purely text based models like BERT, which have strong language understanding capabilities. In this paper, we propose a new training criteria for RNN-T based E2E ASR and SLU to transfer BERT’s knowledge into these systems. In the first stage of our proposed mechanism, we improve ASR performance by using a fine-grained, tokenwise knowledge transfer from BERT. In the second stage, we fine-tune the ASR model for SLU such that the above knowledge is explicitly utilized by the RNN-T model for improved performance. Our techniques improve ASR performance on the Switchboard and CallHome test sets of the NIST Hub5 2000 evaluation and on the recently released SLURP dataset on which we achieve a new state-of-the-art performance. For SLU, we show significant improvements on the SLURP slot filling task, outperforming HuBERT-base and reaching a performance close to HuBERTlarge. Compared to large transformer based speech models like HuBERT, our model is significantly more compact and uses only 300 hours of speech pretraining data. Vishal Sunder, Samuel Thomas 0001, Hong-Kwang Jeff Kuo, Brian Kingsbury, Eric Fosler-Lussier |
ICASSP | 2 |
| 2023 | Multi-Speaker Data Augmentation for Improved end-to-end Automatic Speech RecognitionabstractPublicly available datasets traditionally used to train E2E ASR models for conversational telephone speech recognition are based on clean, short duration, single speaker utterances collected on separate channels. While E2E ASR models achieve state-of-the-art performance on recognition tasks that match well with such training data, they are observed to fail on test recordings that contain multiple speakers, significant channel or background noise or span longer durations than training data utterances. To mitigate these issues, we propose an on-the-fly data augmentation strategy that transforms single speaker training data into multiple speaker data by appending together multiple single speaker utterances. The proposed technique encourages the E2E model to become robust to speaker changes and also process longer utterances effectively. During training, the model is also guided by a teacher model trained on single speaker utterances to map its multi-speaker encoder embeddings to better performing single speaker representations. With the proposed technique we obtain 7-14% relative improvement on various single speaker and multiple speaker test sets. We also show how this technique is able to improve recognition performance by up to 14% by capturing useful information from preceding spoken utterances used as dialog history. Samuel Thomas 0001, Hong-Kwang Jeff Kuo, George Saon, Brian Kingsbury |
ICASSP | 1 |
| 2023 | Comparison of Multilingual Self-Supervised and Weakly-Supervised Speech Pre-Training for Adaptation to Unseen Languages
Andrew Rouditchenko, Sameer Khurana, Samuel Thomas 0001, Rogério Feris, Leonid Karlinsky, Hilde Kuehne, David F. Harwath, Brian Kingsbury, James R. Glass |
INTERSPEECH | 3 |
| 2023 | ConvKT: Conversation-Level Knowledge Transfer for Context Aware End-to-End Spoken Language Understanding
Vishal Sunder, Eric Fosler-Lussier, Samuel Thomas 0001, Hong-Kwang Jeff Kuo, Brian Kingsbury |
INTERSPEECH | 3 |
| 2022 | Everything at Once - Multi-modal Fusion Transformer for Video RetrievalabstractMulti-modal learning from video data has seen increased attention recently as it allows training of semantically meaningful embeddings without human annotation, enabling tasks like zero-shot retrieval and action localization. In this work, we present a multi-modal, modality agnostic fusion transformer that learns to exchange information between multiple modalities, such as video, audio, and text, and integrate them into a fused representation in a joined multi-modal embedding space. We propose to train the system with a combinatorial loss on everything at once – any combination of input modalities, such as single modalities as well as pairs of modalities, explicitly leaving out any add-ons such as position or modality encoding. At test time, the resulting model can process and fuse any number of input modalities. Moreover, the implicit properties of the transformer allow to process inputs of different lengths. To evaluate the proposed approach, we train the model on the large scale HowTo100M dataset and evaluate the resulting embedding space on four challenging benchmark datasets obtaining state-of-the-art results in zero-shot video retrieval and zero-shot video action localization. Our code for this work is also available.11https://github.com/ninatu/everything_at_once Nina Shvetsova, Brian Chen 0001, Andrew Rouditchenko, Samuel Thomas 0001, Brian Kingsbury, Rogério Feris, David F. Harwath, James R. Glass, Hilde Kuehne |
CVPR | 4 |
| 2022 | A New Data Augmentation Method for Intent Classification Enhancement and its Application on Spoken Conversation Datasets
Zvi Kons, Aharon Satt, Hong-Kwang Jeff Kuo, Samuel Thomas 0001, Boaz Carmeli, Ron Hoory, Brian Kingsbury |
ICASSP | 4 |
| 2022 | Improving End-to-end Models for Set Prediction in Spoken Language UnderstandingabstractThe goal of spoken language understanding (SLU) systems is to determine the meaning of the input speech signal, unlike speech recognition which aims to produce verbatim transcripts. Advances in end-to-end (E2E) speech modeling have made it possible to train solely on semantic entities, which are far cheaper to collect than verbatim transcripts. We focus on this set prediction problem, where entity order is unspecified. Using two classes of E2E models, RNN transducers and attention based encoder-decoders, we show that these models work best when the training entity sequence is arranged in spoken order. To improve E2E SLU models when entity spoken order is unknown, we propose a novel data augmentation technique along with an implicit attention based alignment method to infer the spoken order. F1 scores significantly increased by more than 11% for RNN-T and about 2% for attention based encoder-decoder SLU models, outperforming previously reported results. Hong-Kwang Jeff Kuo, Zoltán Tüske, Samuel Thomas 0001, Brian Kingsbury, George Saon |
ICASSP | 3 |
| 2022 | Towards End-to-End Integration of Dialog History for Improved Spoken Language UnderstandingabstractDialog history plays an important role in spoken language understanding (SLU) performance in a dialog system. For end-to-end (E2E) SLU, previous work has used dialog history in text form, which makes the model dependent on a cascaded automatic speech recognizer (ASR). This rescinds the benefits of an E2E system which is intended to be compact and robust to ASR errors. In this paper, we propose a hierarchical conversation model that is capable of directly using dialog history in speech form, making it fully E2E. We also distill semantic knowledge from the available gold conversation transcripts by jointly training a similar text-based conversation model with an explicit tying of acoustic and semantic embeddings. We also propose a novel technique that we call DropFrame to deal with the long training time incurred by adding dialog history in an E2E manner. On the HarperValleyBank dialog dataset, our E2E history integration outperforms a history independent baseline by 7.7% absolute F1 score on the task of dialog action recognition. Our model performs competitively with the state-of-the-art history based cascaded baseline, but uses 48% fewer parameters. In the absence of gold transcripts to fine-tune an ASR model, our model outperforms this baseline by a significant margin of 10% absolute F1 score. Vishal Sunder, Samuel Thomas 0001, Hong-Kwang Jeff Kuo, Jatin Ganhotra, Brian Kingsbury, Eric Fosler-Lussier |
ICASSP | 2 |
| 2022 | Towards Reducing the Need for Speech Training Data to Build Spoken Language Understanding SystemsabstractThe lack of speech data annotated with labels required for spoken language understanding (SLU) is often a major hurdle in building end-to-end (E2E) systems that can directly process speech inputs. In contrast, large amounts of text data with suitable labels are usually available. In this paper, we propose a novel text representation and training methodology that allows E2E SLU systems to be effectively constructed using these text resources. With very limited amounts of additional speech, we show that these models can be further improved to perform at levels close to similar systems built on the full speech datasets. The efficacy of our proposed approach is demonstrated on both intent and entity tasks using three different SLU datasets. With text-only training, the proposed system achieves up to 90% of the performance possible with full speech training. With just an additional 10% of speech data, these models significantly improve further to 97% of full performance. Samuel Thomas 0001, Hong-Kwang Jeff Kuo, Brian Kingsbury, George Saon |
ICASSP | 1 |
| 2022 | Integrating Text Inputs for Training and Adapting RNN Transducer ASR ModelsabstractCompared to hybrid automatic speech recognition (ASR) systems that use a modular architecture in which each component can be in-dependently adapted to a new domain, recent end-to-end (E2E) ASR system are harder to customize due to their all-neural monolithic construction. In this paper, we propose a novel text representation and training framework for E2E ASR models. With this approach, we show that a trained RNN Transducer (RNN-T) model’s internal LM component can be effectively adapted with text-only data. An RNN-T model trained using both speech and text inputs improves over a baseline model trained on just speech with close to 13% word error rate (WER) reduction on the Switchboard and CallHome test sets of the NIST Hub5 2000 evaluation. The usefulness of the proposed approach is further demonstrated by customizing this general purpose RNN-T model to three separate datasets. We observe 20-45% relative word error rate (WER) reduction in these settings with this novel LM style customization technique using only unpaired text data from the new domains. Samuel Thomas 0001, Brian Kingsbury, George Saon, Hong-Kwang Jeff Kuo |
ICASSP | 1 |
| 2022 | Global RNN Transducer Models For Multi-dialect Speech Recognition
Takashi Fukuda, Samuel Thomas 0001, Masayuki Suzuki, Gakuto Kurata, George Saon, Brian Kingsbury |
INTERSPEECH | 2 |
| 2022 | Extending RNN-T-based speech recognition systems with emotion and language classificationabstractSpeech transcription, emotion recognition, and language identification are usually considered to be three different tasks.Each one requires a different model with a different architecture and training process.We propose using a recurrent neural network transducer (RNN-T)-based speech-to-text (STT) system as a common component that can be used for emotion recognition and language identification as well as for speech recognition.Our work extends the STT system for emotion classification through minimal changes, and shows successful results on the IEMOCAP and MELD datasets.In addition, we demonstrate that by adding a lightweight component to the RNN-T module, it can also be used for language identification.In our evaluations, this new classifier demonstrates state-of-the-art accuracy for the NIST-LRE-07 dataset. Zvi Kons, Hagai Aronowitz, Edmilson da Silva Morais, Matheus Damasceno, Hong-Kwang Jeff Kuo, Samuel Thomas 0001, George Saon |
INTERSPEECH | 6 |
| 2022 | Tokenwise Contrastive Pretraining for Finer Speech-to-BERT Alignment in End-to-End Speech-to-Intent SystemsabstractRecent advances in End-to-End (E2E) Spoken Language Understanding (SLU) have been primarily due to effective pretraining of speech representations.One such pretraining paradigm is the distillation of semantic knowledge from state-of-the-art text-based models like BERT to speech encoder neural networks.This work is a step towards doing the same in a much more efficient and fine-grained manner where we align speech embeddings and BERT embeddings on a token-by-token basis.We introduce a simple yet novel technique that uses a cross-modal attention mechanism to extract token-level contextual embeddings from a speech encoder such that these can be directly compared and aligned with BERT based contextual embeddings.This alignment is performed using a novel tokenwise contrastive loss.Fine-tuning such a pretrained model to perform intent recognition using speech directly yields state-of-the-art performance on two widely used SLU datasets.Our model improves further when fine-tuned with additional regularization using SpecAugment especially when speech is noisy, giving an absolute improvement as high as 8% over previous results. Vishal Sunder, Eric Fosler-Lussier, Samuel Thomas 0001, Hong-Kwang Jeff Kuo, Brian Kingsbury |
INTERSPEECH | 3 |
| 2021 | RNN Transducer Models for Spoken Language UnderstandingabstractWe present a comprehensive study on building and adapting RNN transducer (RNN-T) models for spoken language understanding (SLU). These end-to-end (E2E) models are constructed in three practical settings: a case where verbatim transcripts are available, a constrained case where the only available annotations are SLU labels and their values, and a more restrictive case where transcripts are available but not corresponding audio. We show how RNN-T SLU models can be developed starting from pre-trained automatic speech recognition (ASR) systems, followed by an SLU adaptation step. In settings where real audio data is not available, artificially synthesized speech is used to successfully adapt various SLU models. When evaluated on two SLU data sets, the ATIS corpus and a customer call center data set, the proposed models closely track the performance of other E2E models and achieve state-of-the-art results. Samuel Thomas 0001, Hong-Kwang Jeff Kuo, George Saon, Zoltán Tüske, Brian Kingsbury, Gakuto Kurata, Zvi Kons, Ron Hoory |
ICASSP | 1 |
| 2021 | End-to-End Spoken Language Understanding Using Transformer Networks and Self-Supervised Pre-Trained FeaturesabstractTransformer networks and self-supervised pre-training have consistently delivered state-of-art results in the field of natural language processing (NLP); however, their merits in the field of spoken language understanding (SLU) still need further investigation. In this paper we introduce a modular End-to-End (E2E) SLU transformer network based architecture which allows the use of self-supervised pre- trained acoustic features, pre-trained model initialization and multi-task training. Several SLU experiments for predicting intent and entity labels/values using the ATIS dataset are performed. These experiments investigate the interaction of pre-trained model initialization and multi-task training with either traditional filterbank or self-supervised pre-trained acoustic features. Results show not only that self-supervised pre-trained acoustic features outperform filterbank features in almost all the experiments, but also that when these features are used in combination with multi-task training, they almost eliminate the necessity of pre-trained model initialization. Edmilson da Silva Morais, Hong-Kwang Jeff Kuo, Samuel Thomas 0001, Zoltán Tüske, Brian Kingsbury |
ICASSP | 3 |
| 2021 | Multimodal Clustering Networks for Self-supervised Learning from Unlabeled VideosabstractMultimodal self-supervised learning is getting more and more attention as it allows not only to train large networks without human supervision but also to search and retrieve data across various modalities. In this context, this paper proposes a framework that, starting from a pre-trained backbone, learns a common multimodal embedding space that, in addition to sharing representations across different modalities, enforces a grouping of semantically similar instances. To this end, we extend the concept of instance-level contrastive learning with a multimodal clustering step in the training pipeline to capture semantic similarities across modalities. The resulting embedding space enables retrieval of samples across all modalities, even from unseen datasets and different domains. To evaluate our approach, we train our model on the HowTo100M dataset and evaluate its zero-shot retrieval capabilities in two challenging domains, namely text-to-video retrieval, and temporal action localization, showing state-of-the-art results on four different datasets. Brian Chen 0001, Andrew Rouditchenko, Kevin Duarte, Hilde Kuehne, Samuel Thomas 0001, Angie W. Boggust, Rameswar Panda, Brian Kingsbury, Rogério Feris, David F. Harwath, James R. Glass, Michael Picheny, Shih-Fu Chang |
ICCV | 5 |
| 2021 | Speak or Chat with Me: End-to-End Spoken Language Understanding System with Flexible InputsabstractA major focus of recent research in spoken language understanding (SLU) has been on the end-to-end approach where a single model can predict intents directly from speech inputs without intermediate transcripts.However, this approach presents some challenges.First, since speech can be considered as personally identifiable information, in some cases only automatic speech recognition (ASR) transcripts are accessible.Second, intent-labeled speech data is scarce.To address the first challenge, we propose a novel system that can predict intents from flexible types of inputs: speech, ASR transcripts, or both.We demonstrate strong performance for either modality separately, and when both speech and ASR transcripts are available, through system combination, we achieve better results than using a single input modality.To address the second challenge, we leverage a semantically robust pre-trained BERT model and adopt a cross-modal system that co-trains text embeddings and acoustic embeddings in a shared latent space.We further enhance this system by utilizing an acoustic module pre-trained on LibriSpeech and domain-adapting the text module on our target datasets.Our experiments show significant advantages for these pre-training and fine-tuning strategies, resulting in a system that achieves competitive intent-classification performance on Snips SLU and Fluent Speech Commands datasets. Sujeong Cha, Wangrui Hou, Hyun Jung, My Phung, Michael Picheny, Hong-Kwang Jeff Kuo, Samuel Thomas 0001, Edmilson da Silva Morais |
Interspeech | 7 |
| 2021 | Knowledge Distillation Based Training of Universal ASR Source Models for Cross-Lingual Transfer
Takashi Fukuda, Samuel Thomas 0001 |
Interspeech | 2 |
| 2021 | Integrating Dialog History into End-to-End Spoken Language Understanding SystemsabstractEnd-to-end spoken language understanding (SLU) systems that process human-human or human-computer interactions are often context independent and process each turn of a conversation independently. Spoken conversations on the other hand, are very much context dependent, and dialog history contains useful information that can improve the processing of each conversational turn. In this paper, we investigate the importance of dialog history and how it can be effectively integrated into end-to-end SLU systems. While processing a spoken utterance, our proposed RNN transducer (RNN-T) based SLU model has access to its dialog history in the form of decoded transcripts and SLU labels of previous turns. We encode the dialog history as BERT embeddings, and use them as an additional input to the SLU model along with the speech features for the current utterance. We evaluate our approach on a recently released spoken dialog data set, the HarperValleyBank corpus. We observe significant improvements: 8% for dialog action and 30% for caller intent recognition tasks, in comparison to a competitive context independent end-to-end baseline system. Jatin Ganhotra, Samuel Thomas 0001, Hong-Kwang Jeff Kuo, Sachindra Joshi, George Saon, Zoltán Tüske, Brian Kingsbury |
Interspeech | 2 |
| 2021 | Cascaded Multilingual Audio-Visual Learning from VideosabstractIn this paper, we explore self-supervised audio-visual models that learn from instructional videos.Prior work has shown that these models can relate spoken words and sounds to visual content after training on a large-scale dataset of videos, but they were only trained and evaluated on videos in English.To learn multilingual audio-visual representations, we propose a cascaded approach that leverages a model trained on English videos and applies it to audio-visual data in other languages, such as Japanese videos.With our cascaded approach, we show an improvement in retrieval performance of nearly 10x compared to training on the Japanese videos solely.We also apply the model trained on English videos to Japanese and Hindi spoken captions of images, achieving state-of-the-art performance. Andrew Rouditchenko, Angie W. Boggust, David F. Harwath, Samuel Thomas 0001, Hilde Kuehne, Brian Chen 0001, Rameswar Panda, Rogério Feris, Brian Kingsbury, Michael Picheny, James R. Glass |
Interspeech | 4 |
| 2021 | AVLnet: Learning Audio-Visual Language Representations from Instructional VideosabstractCurrent methods for learning visually grounded language from videos often rely on text annotation, such as human generated captions or machine generated automatic speech recognition (ASR) transcripts. In this work, we introduce the Audio-Video Language Network (AVLnet), a self-supervised network that learns a shared audio-visual embedding space directly from raw video inputs. To circumvent the need for text annotation, we learn audio-visual representations from randomly segmented video clips and their raw audio waveforms. We train AVLnet on HowTo100M, a large corpus of publicly available instructional videos, and evaluate on image retrieval and video retrieval tasks, achieving state-of-the-art performance. We perform analysis of AVLnet's learned representations, showing our model utilizes speech and natural sounds to learn audio-visual concepts. Further, we propose a tri-modal model that jointly processes raw audio, video, and text captions from videos to learn a multi-modal semantic embedding space useful for text-video retrieval. Our code, data, and trained models will be released at avlnet.csail.mit.edu Andrew Rouditchenko, Angie W. Boggust, David F. Harwath, Brian Chen 0001, Dhiraj Joshi, Samuel Thomas 0001, Kartik Audhkhasi, Hilde Kuehne, Rameswar Panda, Rogério Feris, Brian Kingsbury, Michael Picheny, Antonio Torralba 0001, James R. Glass |
Interspeech | 6 |
| 2021 | Auxiliary Networks for Joint Speaker Adaptation and Speaker Change DetectionabstractSpeaker adaptation and speaker change detection have both been studied extensively to improve automatic speech recognition (ASR). In many cases, these two problems are investigated separately: speaker change detection is implemented first to obtain single-speaker regions, and speaker adaptation is then performed using the derived speaker segments for improved ASR. However, in an online setting, we want to achieve both goals in a single pass. In this study, we propose a neural network architecture that learns a speaker embedding from which it can perform both speaker adaptation for ASR and speaker change detection. The proposed speaker embedding is computed using self-attention based on an auxiliary network attached to a main ASR network. ASR adaptation is then performed by subtracting, from the main network activations, a segment dependent affine transformation of the learned speaker embedding. In experiments on a broadcast news dataset and the Switchboard conversational dataset, we test our system on utterances with a change point in them and show that the proposed method achieves significantly better performance as compared to the unadapted main network (10-14% relative reduction in word error rate (WER)). The proposed architecture also outperforms three different speaker segmentation methods followed by ASR (around 10% relative reduction in WER). Leda Sari, Mark Hasegawa-Johnson, Samuel Thomas 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Leveraging Unpaired Text Data for Training End-To-End Speech-to-Intent SystemsabstractTraining an end-to-end (E2E) neural network speech-to-intent (S2I) system that directly extracts intents from speech requires large amounts of intent-labeled speech data, which is time consuming and expensive to collect. Initializing the S2I model with an ASR model trained on copious speech data can alleviate data sparsity. In this paper, we attempt to leverage NLU text resources. We implemented a CTC-based S2I system that matches the performance of a state-of-the-art, traditional cascaded SLU system. We performed controlled experiments with varying amounts of speech and text training data. When only a tenth of the original data is available, intent classification accuracy degrades by 7.6% absolute. Assuming we have additional text-to-intent data (without speech) available, we investigated two techniques to improve the S2I system: (1) transfer learning, in which acoustic embeddings for intent classification are tied to fine-tuned BERT text embeddings; and (2) data augmentation, in which the text-to-intent data is converted into speech-to-intent data using a multi-speaker text-to-speech system. The proposed approaches recover 80% of performance lost due to using limited intent-labeled speech. Hong-Kwang Jeff Kuo, Samuel Thomas 0001, Zvi Kons, Kartik Audhkhasi, Brian Kingsbury, Ron Hoory, Michael Picheny |
ICASSP | 3 |
| 2020 | Audio-Assisted Image Inpainting for Talking FacesabstractThe goal of our work is to complete missing areas of images of talking faces, exploiting information from both the visual and audio modalities. Existing image inpainting methods rely solely on visual content that doesn't always provide sufficient information for the task. To counter this, we propose a neural network that employs an encoder-decoder architecture with a bimodal fusion mechanism, thus taking into account both visual and audio content. Our proposed method demonstrates consistently superior performance over a baseline visual-only model, reaching for example up to 17% relative improvement in mean absolute error. The presented model is applicable to practical video editing tasks, such as object and overlay-text removal from talking faces, where existing lip and face generation works are not applicable as they require clean input. Alexandros Koumparoulis, Gerasimos Potamianos, Samuel Thomas 0001, Edmilson da Silva Morais |
ICASSP | 3 |
| 2020 | Training Spoken Language Understanding Systems with Non-Parallel Speech and TextabstractEnd-to-end spoken language understanding (SLU) systems are typically trained on large amounts of data. In many practical scenarios, the amount of labeled speech is often limited as opposed to text. In this study, we investigate the use of non-parallel speech and text to improve the performance of dialog act recognition as an example SLU task. We propose a multiview architecture that can handle each modality separately. To effectively train on such data, this model enforces the internal speech and text encodings to be similar using a shared classifier. On the Switchboard Dialog Act corpus, we show that pretraining the classifier using large amounts of text helps learning better speech encodings, resulting in up to 40% relatively higher classification accuracies. We also show that when the speech embeddings from an automatic speech recognition (ASR) system are used in this framework, the speech-only accuracy exceeds the performance of ASR-text based tests up to 15% relative and approaches the performance of using true transcripts. Leda Sari, Samuel Thomas 0001, Mark Hasegawa-Johnson |
ICASSP | 2 |
| 2020 | Transliteration Based Data Augmentation for Training Multilingual ASR Acoustic Models in Low Resource Settings
Samuel Thomas 0001, Kartik Audhkhasi, Brian Kingsbury |
INTERSPEECH | 1 |
| 2020 | Implicit Transfer of Privileged Acoustic Information in a Generalized Knowledge Distillation Framework
Takashi Fukuda, Samuel Thomas 0001 |
INTERSPEECH | 2 |
| 2020 | Resource-Adaptive Deep Learning for Visual Speech RecognitionabstractWe focus on the problem of efficient architectures for lipreading that allow trading-off computational resources for visual speech recognition accuracy. In particular, we make two contributions: First, we introduce MobiLipNetV3, an efficient and accurate lipreading model, based on our earlier work on MobiLipNetV2 and incorporating recent advances in convolutional neural network architectures. Second, we propose a novel recognition paradigm, called MultiRate Ensemble (MRE), that combines a “lean” and a “full” MobiLipNetV3 in the lipreading pipeline, with the latter applied at a lower frame rate. This architecture yields a family of systems offering multiple accuracy vs. efficiency operating points depending on the frame-rate decimation of the “full” model, thus allowing adaptation to the available device resources. We evaluate our approach on the TCD-TIMIT corpus, popular in speaker-independent lipreading of continuous speech. The proposed MRE family of systems can be up to 73 times more efficient compared to residual neural network based lipreading, and up to twice as MobiLipNetV2, while in both cases reaching up to 8% absolute WER reduction, depending on the MRE chosen operating point. For example, a temporal decimation of three yields a 7% absolute WER reduction and a 26% relative decrease in computations over MobiLipNetV2. © 2020 ISCA Alexandros Koumparoulis, Gerasimos Potamianos, Samuel Thomas 0001, Edmilson da Silva Morais |
INTERSPEECH | 3 |
| 2020 | End-to-End Spoken Language Understanding Without Full TranscriptsabstractAn essential component of spoken language understanding (SLU) is slot filling: representing the meaning of a spoken utterance using semantic entity labels. In this paper, we develop end-to-end (E2E) spoken language understanding systems that directly convert speech input to semantic entities and investigate if these E2E SLU models can be trained solely on semantic entity annotations without word-for-word transcripts. Training such models is very useful as they can drastically reduce the cost of data collection. We created two types of such speech-to-entities models, a CTC model and an attention-based encoder-decoder model, by adapting models trained originally for speech recognition. Given that our experiments involve speech input, these systems need to recognize both the entity label and words representing the entity value correctly. For our speech-to-entities experiments on the ATIS corpus, both the CTC and attention models showed impressive ability to skip non-entity words: there was little degradation when trained on just entities versus full transcripts. We also explored the scenario where the entities are in an order not necessarily related to spoken order in the utterance. With its ability to do re-ordering, the attention model did remarkably well, achieving only about 2% degradation in speech-to-bag-of-entities F1 score. Hong-Kwang Jeff Kuo, Zoltán Tüske, Samuel Thomas 0001, Kartik Audhkhasi, Brian Kingsbury, Gakuto Kurata, Zvi Kons, Ron Hoory, Luis A. Lastras |
INTERSPEECH | 3 |
| 2019 | Mixed Bandwidth Acoustic Modeling Leveraging Knowledge DistillationabstractTraining of mixed bandwidth acoustic models have recently been realized by incorporating special Mel filterbanks. To fit information into every filterbank bin available across both narrowband and wideband data, these filterbanks pad zeros at high frequency ranges of narrowband data. Although these methods succeed in decreasing word error rates (WER) on broadband data, they fail to improve on narrowband signals. In this paper, we propose methods to mitigate these effects with generalized knowledge distillation. In our method, specialized teacher networks are first trained on lossless acoustic features with full scale Mel filterbanks. While training student networks, privileged knowledge from these teacher networks is then used to compensate for missing information at high frequencies introduced by the special Mel filterbanks. We show the benefit of the proposed technique for both narrowband (10% relative WER improvement) and wideband data (7.5% relative WER improvement) on the Aurora 4 task over traditional methods. Takashi Fukuda, Samuel Thomas 0001 |
ASRU | 2 |
| 2019 | Semi-Supervised Training and Data Augmentation for Adaptation of Automatic Broadcast News Captioning SystemsabstractIn this paper we present a comprehensive study on building and adapting deep neural network based speech recognition systems for automatic closed captioning. We develop the proposed systems by first building base automatic speech recognition (ASR) systems that are not specific to any particular show or station. These models are trained on nearly 6000 hours of broadcast news data using conventional hybrid and more recent attention based end-to-end acoustic models. We then employ various adaptation and data augmentation strategies to further improve the trained base models. We use 535 hours of data from two independent BN sources to study how the base models can be customized. We observe up to 32% relative improvement using the proposed techniques on test sets related to, but independent of the adaptation data. At these low word error rates (WERs), we believe the customized BN ASR systems can be used effectively for automatic closed captioning. Samuel Thomas 0001, Masayuki Suzuki, Zoltán Tüske, Larry Sansone, Michael Picheny |
ASRU | 2 |
| 2019 | Simplified LSTMS for Speech RecognitionabstractIn this paper we explore new variants of Long Short-Term Memory (LSTM) networks for sequential modeling of acoustic features. In particular, we show that: (i) removing the output gate, (ii) replacing the hyperbolic tangent nonlinearity at the cell output with hard tanh, and (iii) collapsing the cell and hidden state vectors leads to a model that is conceptually simpler than and comparable in effectiveness to a regular LSTM for speech recognition. The proposed model has 25% fewer parameters than an LSTM with the same number of cells, trains faster because it has larger gradients leading to larger steps in weight space, and reaches a better optimum because there are fewer nonlinearities to traverse across layers. We report experimental results for both hybrid and CTC acoustic models on three publicly available English datasets: Switchboard 300 hours telephone conversations, 400 hours broadcast news transcription, and the MALACH 176 hours corpus of Holocaust survivor testimonies. In all cases the proposed models achieve similar or better accuracy than regular LSTMs while being conceptually simpler. George Saon, Zoltán Tüske, Kartik Audhkhasi, Brian Kingsbury, Michael Picheny, Samuel Thomas 0001 |
ASRU | 6 |
| 2019 | Pre-training of Speaker Embeddings for Low-latency Speaker Change Detection in Broadcast NewsabstractIn this work, we investigate pre-training of neural network based speaker embeddings for low-latency speaker change detection. Our proposed system takes two speech segments, generates embeddings using shared Siamese layers and then classifies the concatenated embeddings depending on whether they are spoken by the same speaker. We investigate gender classification, contrastive loss and triplet loss based pre-training of the embedding layers and also joint training of the embedding layers along with a same/different classifier. Training is performed on 2-second single speaker segments based on ground truth speaker segmentation of broadcast news data. However, during test, we use the detection system in a practical low-latency setting for annotating automatic closed captions. In contrast to training, test pairs are now created around automatic speech recognition (ASR) based segmentation boundaries. The ASR segments are often shorter than 2 seconds causing duration mismatch during testing. In our experiments, although the baseline i-vector based classifier performs well, the proposed triplet loss based pre-training followed by joint training provides 7-50% relative F-measure improvement in matched and mismatched conditions. In addition, the degradation in performance is less severe for network based embeddings as compared to using i-vectors in the variable duration test conditions. Leda Sari, Samuel Thomas 0001, Mark Hasegawa-Johnson, Michael Picheny |
ICASSP | 2 |
| 2019 | Improvements to N-gram Language Model Using Text Generated from Neural Language ModelabstractAlthough neural language models have emerged, n-gram language models are still used for many speech recognition tasks. This paper proposes four methods to improve n-gram language models using text generated from a recurrent neural network language model (RNNLM). First, we use multiple RNNLMs from different domains instead of a single RNNLM. The final n-gram language model is obtained by interpolating generated n-gram models from each domain. Second, we use subwords instead of words for RNNLM to reduce the out-of-vocabulary rate. Third, we generate text templates using an RNNLM for template-based data augmentation for named entities. Fourth, we use both forward RNNLM and backward RNNLM to generate text. We found that these four methods improved performance of speech recognition up to 4% relative in various tasks. Masayuki Suzuki, Nobuyasu Itoh, Tohru Nagano, Gakuto Kurata, Samuel Thomas 0001 |
ICASSP | 5 |
| 2019 | English Broadcast News Speech Recognition by Humans and MachinesabstractWith recent advances in deep learning, considerable attention has been given to achieving automatic speech recognition performance close to human performance on tasks like conversational telephone speech (CTS) recognition. In this paper we evaluate the usefulness of these proposed techniques on broadcast news (BN), a similar challenging task. We also perform a set of recognition measurements to understand how close the achieved automatic speech recognition results are to human performance on this task. On two publicly available BN test sets, DEV04F and RT04, our speech recognition system using LSTM and residual network based acoustic models with a combination of n-gram and neural network language models performs at 6.5% and 5.9% word error rate. By achieving new performance milestones on these test sets, our experiments show that techniques developed on other related tasks, like CTS, can be transferred to achieve similar performance. In contrast, the best measured human recognition performance on these test sets is much lower, at 3.6% and 2.8% respectively, indicating that there is still room for new techniques and improvements in this space, to reach human performance levels. Samuel Thomas 0001, Masayuki Suzuki, Gakuto Kurata, Zoltán Tüske, George Saon, Brian Kingsbury, Michael Picheny, Tom Dibert, Alice Kaiser-Schatzlein, Bern Samko |
ICASSP | 1 |
| 2019 | Learning Speaker Aware Offsets for Speaker Adaptation of Neural Networks
Leda Sari, Samuel Thomas 0001, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2019 | Detection and Recovery of OOVs for Improved English Broadcast News Captioning
Samuel Thomas 0001, Kartik Audhkhasi, Zoltán Tüske, Michael Picheny |
INTERSPEECH | 1 |
| 2018 | Joint Modeling of Accents and Acoustics for Multi-Accent Speech RecognitionabstractThe performance of automatic speech recognition systems degrades with increasing mismatch between the training and testing scenarios. Differences in speaker accents are a significant source of such mismatch. The traditional approach to deal with multiple accents involves pooling data from several accents during training and building a single model in multi-task fashion, where tasks correspond to individual accents. In this paper, we explore an alternate model where we jointly learn an accent classifier and a multi-task acoustic model. Experiments on the American English Wall Street Journal and British English Cambridge corpora demonstrate that our joint model outperforms the strong multi-task acoustic model baseline. We obtain a 5.94% relative improvement in word error rate on British English, and 9.47% relative improvement on American English. This illustrates that jointly modeling with accent information improves acoustic model performance. Xuesong Yang, Kartik Audhkhasi, Andrew Rosenberg, Samuel Thomas 0001, Bhuvana Ramabhadran, Mark Hasegawa-Johnson |
ICASSP | 4 |
| 2018 | Data Augmentation Improves Recognition of Foreign Accented Speech
Takashi Fukuda, Raul Fernandez, Andrew Rosenberg, Samuel Thomas 0001, Bhuvana Ramabhadran, Alexander Sorin, Gakuto Kurata |
INTERSPEECH | 4 |
| 2018 | Inference-Invariant Transformation of Batch Normalization for Domain Adaptation of Acoustic Models
Masayuki Suzuki, Tohru Nagano, Gakuto Kurata, Samuel Thomas 0001 |
INTERSPEECH | 4 |
| 2018 | A Recorded Debating Dataset
Shachar Mirkin, Michal Jacovi, Tamar Lavee, Hong-Kwang Jeff Kuo, Samuel Thomas 0001, Leslie Sager, Lili Kotlerman, Elad Venezian, Noam Slonim |
LREC | 5 |
| 2017 | Effective joint training of denoising feature space transforms and Neural Network based acoustic modelsabstractNeural Network (NN) based acoustic frontends, such as denoising autoencoders, are actively being investigated to improve the robustness of NN based acoustic models to various noise conditions. In recent work the joint training of such frontends with backend NNs has been shown to significantly improve speech recognition performance. In this paper, we propose an effective algorithm to jointly train such a denoising feature space transform and a NN based acoustic model with various kinds of data. Our proposed method first pretrains a Convolutional Neural Network (CNN) based denoising frontend and then jointly trains this frontend with a NN backend acoustic model. In the unsupervised pretraining stage, the frontend is designed to estimate clean log Mel-filterbank features from noisy log-power spectral input features. A subsequent multi-stage training of the proposed frontend, with the dropout technique applied only at the joint layer between the frontend and backend NNs, leads to significant improvements in the overall performance. On the Aurora-4 task, our proposed system achieves an average WER of 9.98%. This is a 9.0% relative improvement over one of the best reported speaker independent baseline system's performance. A final semi-supervised adaptation of the frontend NN, similar to feature space adaptation, reduces the average WER to 7.39%, a further relative WER improvement of 25%. Takashi Fukuda, Osamu Ichikawa, Gakuto Kurata, Ryuki Tachibana, Samuel Thomas 0001, Bhuvana Ramabhadran |
ICASSP | 5 |
| 2017 | Efficient Knowledge Distillation from an Ensemble of Teachers
Takashi Fukuda, Masayuki Suzuki, Gakuto Kurata, Samuel Thomas 0001, Jia Cui, Bhuvana Ramabhadran |
INTERSPEECH | 4 |
| 2017 | English Conversational Telephone Speech Recognition by Humans and MachinesabstractOne of the most difficult speech recognition tasks is accurate recognition of human to human communication. Advances in deep learning over the last few years have produced major speech recognition improvements on the representative Switchboard conversational corpus. Word error rates that just a few years ago were 14% have dropped to 8.0%, then 6.6% and most recently 5.8%, and are now believed to be within striking range of human performance. This then raises two issues - what IS human performance, and how far down can we still drive speech recognition error rates? A recent paper by Microsoft suggests that we have already achieved human performance. In trying to verify this statement, we performed an independent set of human performance measurements on two conversational tasks and found that human performance may be considerably better than what was earlier reported, giving the community a significantly harder goal to achieve. We also report on our own efforts in this area, presenting a set of acoustic and language modeling techniques that lowered the word error rate of our own English conversational telephone LVCSR system to the level of 5.5%/10.3% on the Switchboard/CallHome subsets of the Hub5 2000 evaluation, which - at least at the writing of this paper - is a new performance milestone (albeit not at what we measure to be human performance!). On the acoustic side, we use a score fusion of three models: one LSTM with multiple feature inputs, a second LSTM trained with speaker-adversarial multi-task learning and a third residual net (ResNet) with 25 convolutional layers and time-dilated convolutions. On the language modeling side, we use word and character LSTMs and convolutional WaveNet-style language models. George Saon, Gakuto Kurata, Tom Sercu, Kartik Audhkhasi, Samuel Thomas 0001, Dimitrios Dimitriadis, Bhuvana Ramabhadran, Michael Picheny, Lynn-Li Lim, Bergul Roomi, Phil Hall |
INTERSPEECH | 5 |
| 2016 | On the importance of event detection for ASRabstractThe performance of modern large vocabulary continuous speech recognition (LVCSR) systems is heavily affected by segment boundaries, proper speaker identification of the segments, as well as removal of spurious data. We propose to use Long Short Term Memory (LSTM) recurrent neural networks to partition audio into speech segments as well as track speaker turns. Additionally, we train an LSTM to also identify music segments. We show that the accurate detection of events, along with removal of silence and music, using our LSTM yields a 9-10% relative improvement in ASR performance. Secondary processing by speaker clustering provides an additional boost in accuracy. Event detection accuracy of the LSTM approach is also described. David Haws, Dimitrios Dimitriadis, George Saon, Samuel Thomas 0001, Michael Picheny |
ICASSP | 4 |
| 2016 | CNMF-based acoustic features for noise-robust ASRabstractWe present an algorithm using convolutive non-negative matrix factorization (CNMF) to create noise-robust features for automatic speech recognition (ASR). Typically in noise-robust ASR, CNMF is used to remove noise from noisy speech prior to feature extraction. However, we find that denoising introduces distortion and artifacts, which can degrade ASR performance. Instead, we propose using the time-activation matrices from CNMF as acoustic model features. In this paper, we describe how to create speech and noise dictionaries that generate noise-robust time-activation matrices from noisy speech. Using the time-activation matrices created by our proposed algorithm, we achieve a 11.8% relative improvement in the word error rate on the Aurora 4 corpus compared to using log-mel filterbank energies. Furthermore, we attain a 13.8% relative improvement over log-mel filterbank energies when we combine them with our proposed features, indicating that our features contain complementary information to log-mel features. Colin Vaz, Dimitrios Dimitriadis, Samuel Thomas 0001, Shri Narayanan |
ICASSP | 3 |
| 2016 | An Investigation on the Use of i-Vectors for Robust ASRabstractIn this paper we propose two different i-vector representations that improve the noise robustness of automatic speech recognition (ASR). The first kind of i-vectors is derived from ``noise only'' components of speech provided by an adaptive denoising algorithm, the second variant is extracted from mel filterbank energies containing both speech and noise. The effectiveness of both these representations is shown by combining them with two different kinds of spectral features - the commonly used log-mel filterbank energies and Teager energy spectral coefficients (TESCs). Using two different DNN architectures for acoustic modeling - a standard state-of-the-art sigmoid-based DNN and an advanced architecture using leaky ReLUs, dropout and resealing, we demonstrate the benefit of the proposed representations. On the Aurora-4 multi-condition training task the proposed front-end improves ASR performance by 4%. Dimitrios Dimitriadis, Samuel Thomas 0001, Sriram Ganapathy |
INTERSPEECH | 2 |
| 2016 | Domain Adaptation of CNN Based Acoustic Models Under Limited Resource Settings
Masayuki Suzuki, Ryuki Tachibana, Samuel Thomas 0001, Bhuvana Ramabhadran, George Saon |
INTERSPEECH | 3 |
| 2016 | Multilingual Data Selection for Low Resource Speech RecognitionabstractAbstract : Feature representations extracted from deep neural network-based multilingual frontends provide significant improvements to speech recognition systems in low resource settings. To effectively train these frontends, we introduce a data selection technique that discovers language groups from an available set of training languages. This data selection method reduces the required amount of training data and training time by approximately 40 , with minimal performance degradation. We present speech recognition results on 7 very limited language pack (VLLP) languages from the second option period of the IARPA Babel program using multilingual features trained on up to 10 languages. The proposed multilingual features provide up to 15 relative improvement over baseline acoustic features on the VLLP languages. Samuel Thomas 0001, Kartik Audhkhasi, Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran |
INTERSPEECH | 1 |
| 2015 | Improvements to the IBM speech activity detection system for the DARPA RATS programabstractIn this paper we describe improvements to the IBM speech activity detection (SAD) system for the third phase of the DARPA RATS program. The progress during this final phase comes from jointly training convolutional and regular deep neural networks with rich time-frequency representations of speech. With these additions, the phase 3 system reduces the equal error rate (EER) significantly on both of the program's development sets (relative improvements of 20% on dev1 and 7% on dev2) compared to an earlier phase 2 system. For the final program evaluation, the newly developed system also performs well past the program target of 3% Pmissat 1% Pfawith a performance of 1.2% Pmissat 1% Pfaand 0.3% Pfaat 3% Pmiss. Samuel Thomas 0001, George Saon, Maarten Van Segbroeck, Shri Narayanan |
ICASSP | 1 |
| 2015 | Investigating factor analysis features for deep neural networks in noisy speech recognitionabstractThe problem of speaker and channel adaptation in deep neural network (DNN) based automatic speech recognition (ASR) sys-tems is of substantial interest in advancing the performance of these systems. Recently, the speaker identity vectors (i-vectors) have shown improvements for ASR systems in matched condi-tions. In this paper, we propose the application of the general factor analysis framework for noisy speech recognition tasks. Several methods for deriving speaker and channel factors are explored including joint factor analysis (JFA) and i-vectors de-rived from DNN posteriors instead of the traditional Universal background model (UBM) approach. We also experiment with the late fusion of i-vector features with bottleneck (BN) fea-tures obtained from a previously trained convolutional neural network (CNN) system. The ASR experiments are performed on the Aspire challenge test data which contains noisy far-field speech while the acoustic models are trained with conversa-tional telephone speech (CTS) data from the Fisher corpus. In these experiments, we show that the factor analysis based meth-ods provide significant improvements in the word error rate (relative improvements of about 11 % compared to the baseline DNN system trained with speaker adapted features). Sriram Ganapathy, Samuel Thomas 0001, Dimitrios Dimitriadis, Steven J. Rennie |
INTERSPEECH | 2 |
| 2015 | The IBM BOLT speech transcription systemabstractWe describe the IBM automatic speech recognition (ASR) sys-tem for the DARPA Broad Operational Language Translation (BOLT) program. The system is used to transcribe conversa-tional telephone speech (CTS) prior to machine translation for Phase 3 of the program’s Activity A. The ASR system is a com-bination of novel sequence trained ensemble deep neural net-work acoustic models on speaker adapted features and convolu-tional neural network models on two kinds of spectro-temporal representations of speech, in conjunction with a variety of class, neural network and n-gram based language models. Acoustic and language models for the recognition system are built on transcribed audio released under the program and further opti-mized for the final machine translation task as well. The evalua-tion system has a word error rate of 32.7 % on a 2 hour Egyptian Arabic development set for this task. Index Terms: Automatic speech recognition, conversational telephone speech, deep neural networks, machine translation Samuel Thomas 0001, George Saon, Hong-Kwang Jeff Kuo, Lidia Mangu |
INTERSPEECH | 1 |
| 2014 | Analyzing convolutional neural networks for speech activity detection in mismatched acoustic conditionsabstractConvolutional neural networks (CNN) are extensions to deep neural networks (DNN) which are used as alternate acoustic models with state-of-the-art performances for speech recognition. In this paper, CNNs are used as acoustic models for speech activity detection (SAD) on data collected over noisy radio communication channels. When these SAD models are tested on audio recorded from radio channels not seen during training, there is severe performance degradation. We attribute this degradation to mismatches between the two dimensional filters learnt in the initial CNN layers and the novel channel data. Using a small amount of supervised data from the novel channels, the filters can be adapted to provide significant improvements in SAD performance. In mismatched acoustic conditions, the adapted models provide significant improvements (about 10-25%) relative to conventional DNN-based SAD systems. These results illustrate that CNNs have a considerable advantage in fast adaptation for acoustic modeling in these settings. Samuel Thomas 0001, Sriram Ganapathy, George Saon, Hagen Soltau |
ICASSP | 1 |
| 2014 | Robust language identification using convolutional neural network featuresabstractThe language identification (LID) task in the Robust Automatic Transcription of Speech (RATS) program is challenging due to the noisy nature of the audio data collected over highly degraded radio communication channels as well as the use of short duration speech segments for testing. In this paper, we report the recent advances made in the RATS LID task by using bottleneck features from a convolutional neural network (CNN). The CNN, which is trained with labelled data from one of target languages, generates bottleneck features which are used in a Gaussian mixture model (GMM)-ivector LID system. The CNN bottleneck features provide substantial complimentary information to the conventional acoustic features even on languages not seen in its training. Using these bottleneck features in conjunction with acoustic features, we obtain significant improvements (average relative improvements of 25% in terms of equal error rate (EER) compared to the corresponding acoustic system) for the LID task. Furthermore, these improvements are consistent for various choices of acoustic features as well as speech segment durations. Sriram Ganapathy, Kyu Jeong Han, Samuel Thomas 0001, Mohamed Kamal Omar, Maarten Van Segbroeck, Shri Narayanan |
INTERSPEECH | 3 |
| 2014 | Deep Order Statistic NetworksabstractRecently, Maxout networks have demonstrated state-of-the-art performance on several machine learning tasks, which has fueled aggressive research on Maxout networks and generalizations thereof. In this work, we propose the utilization of order statistics as a generalization of the max non-linearity. A particularly general example of an order-statistic non-linearity is the “sortout” non-linearity, which outputs all input activations, but in sorted order. Such Order-statistic networks (OSNs), in contrast with other recently proposed generalizations of Maxout networks, leave the determination of the interpolation weights on the activations to the network, and remain conditionally linear given the input, and so are well suited for powerful model aggregation techniques such as dropout, drop connect, and annealed dropout. Experimental results demonstrate that the use of order statistics rather than Maxout networks can lead to substantial improvements in the word error rate (WER) performance of automatic speech recognition systems. Steven J. Rennie, Vaibhava Goel, Samuel Thomas 0001 |
SLT | 3 |
| 2014 | Annealed dropout training of deep networksabstractRecently it has been shown that when training neural networks on a limited amount of data, randomly zeroing, or “dropping out” a fixed percentage of the outputs of a given layer for each training case can improve test set performance significantly. Dropout training discourages the detectors in the network from co-adapting, which limits the capacity of the network and prevents overfitting. In this paper we show that annealing the dropout rate from a high initial value to zero over the course of training can substantially improve the quality of the resulting model. As dropout (approximately) implements model aggregation over an exponential number of networks, this procedure effectively initializes the ensemble of models that will be learned during a given iteration of training with an enemble of models that has a lower average number of neurons per network, and higher variance in the number of neurons per network-which regularizes the structure of the final model toward models that avoid unnecessary co-adaptation between neurons. Importantly, this regularization procedure is stochastic, and so promotes the learning of “balanced” networks with neurons that have high average entropy, and low variance in their entropy, by smoothly transitioning from “exploration” with high learning rates to “fine tuning” with full support for co-adaptation between neurons where necessary. Experimental results demonstrate that annealed dropout leads to significant reductions in word error rate over standard dropout training. Steven J. Rennie, Vaibhava Goel, Samuel Thomas 0001 |
SLT | 3 |
| 2013 | A summary of the 2012 JHU CLSP workshop on zero resource speech technologies and models of early language acquisitionabstractWe summarize the accomplishments of a multi-disciplinary workshop exploring the computational and scientific issues surrounding zero resource (unsupervised) speech technologies and related models of early language acquisition. Centered around the tasks of phonetic and lexical discovery, we consider unified evaluation metrics, present two new approaches for improving speaker independence in the absence of supervision, and evaluate the application of Bayesian word segmentation algorithms to automatic subword unit tokenizations. Finally, we present two strategies for integrating zero resource techniques into supervised settings, demonstrating the potential of unsupervised methods to improve mainstream technologies. Aren Jansen, Emmanuel Dupoux, Sharon Goldwater, Mark Johnson 0001, Sanjeev Khudanpur, Kenneth Church 0001, Naomi Feldman, Hynek Hermansky, Florian Metze, Richard C. Rose, Mike Seltzer, Pascal Clark, Ian McGraw, Balakrishnan Varadarajan, Erin D. Bennett, Benjamin Börschinger, Justin T. Chiu, Ewan Dunbar, Abdellah Fourtassi, David F. Harwath, Chia-ying Lee, Keith D. Levin, Atta Norouzian, Vijayaditya Peddinti, Rachael Richardson, Thomas Schatz, Samuel Thomas 0001 |
ICASSP | 27 |
| 2013 | Weak top-down constraints for unsupervised acoustic model trainingabstractTypical supervised acoustic model training relies on strong top-down constraints provided by dynamic programming alignment of the input observations to phonetic sequences derived from orthographic word transcripts and pronunciation dictionaries. This paper investigates a much weaker form of top-down supervision for use in place of transcripts and dictionaries in the zero resource setting. Our proposed constraints, which can be produced using recent spoken term discovery systems, come in the form of pairs of isolated word examples that share the same unknown type. For each pair, we perform a dynamic programming alignment of the acoustic observations of the two constituent examples, generating an inventory of cross-speaker frame pairs that each provide evidence that the same subword unit model should account for them. We find these weak top-down constraints are capable of improving model speaker independence by up to 57% relative over bottom-up training alone. Aren Jansen, Samuel Thomas 0001, Hynek Hermansky |
ICASSP | 2 |
| 2013 | Developing a speaker identification system for the DARPA RATS projectabstractThis paper describes the speaker identification (SID) system developed by the Patrol team for the first phase of the DARPA RATS (Robust Automatic Transcription of Speech) program, which seeks to advance state of the art detection capabilities on audio from highly degraded communication channels. We present results using multiple SID systems differing mainly in the algorithm used for voice activity detection (VAD) and feature extraction. We show that (a) unsupervised VAD performs as well supervised methods in terms of downstream SID performance, (b) noise-robust feature extraction methods such as CFCCs out-perform MFCC front-ends on noisy audio, and (c) fusion of multiple systems provides 24% relative improvement in EER compared to the single best system when using a novel SVM-based fusion algorithm that uses side information such as gender, language, and channel id. Oldrich Plchot, Spyridon Matsoukas, Pavel Matejka, Najim Dehak, Jeff Z. Ma, Sandro Cumani, Ondrej Glembek, Hynek Hermansky, Sri Harish Reddy Mallidi, Nima Mesgarani, Richard M. Schwartz, Mehdi Soufifar, Zheng-Hua Tan, Samuel Thomas 0001, Bing Zhang 0004, Xinhui Zhou |
ICASSP | 14 |
| 2013 | Deep neural network features and semi-supervised training for low resource speech recognitionabstractWe propose a new technique for training deep neural networks (DNNs) as data-driven feature front-ends for large vocabulary continuous speech recognition (LVCSR) in low resource settings. To circumvent the lack of sufficient training data for acoustic modeling in these scenarios, we use transcribed multilingual data and semi-supervised training to build the proposed feature front-ends. In our experiments, the proposed features provide an absolute improvement of 16% in a low-resource LVCSR setting with only one hour of in-domain training data. While close to three-fourths of these gains come from DNN-based features, the remaining are from semi-supervised training. Samuel Thomas 0001, Michael L. Seltzer, Kenneth Church 0001, Hynek Hermansky |
ICASSP | 1 |
| 2013 | The IBM speech activity detection system for the DARPA RATS program
George Saon, Samuel Thomas 0001, Hagen Soltau, Sriram Ganapathy, Brian Kingsbury |
INTERSPEECH | 2 |
| 2012 | The UMD-JHU 2011 speaker recognition systemabstractIn recent years, there have been significant advances in the field of speaker recognition that has resulted in very robust recognition systems. The primary focus of many recent developments have shifted to the problem of recognizing speakers in adverse conditions, e.g in the presence of noise/reverberation. In this paper, we present the UMD-JHU speaker recognition system applied on the NIST 2010 SRE task. The novel aspects of our systems are: 1) Improved performance on trials involving different vocal effort via the use of linear-scale features; 2) Expected improved recognition performance in the presence of reverberation and noise via the use of frequency domain perceptual linear predictor and cortical features; 3) A new discriminative kernel partial least squares (KPLS) framework that complements state-of-the-art back-end systems JFA and PLDA to aid in better overall recognition; and 4) Acceleration of JFA, PLDA and KPLS back-ends via distributed computing. The individual components of the system and the fused system are compared against a baseline JFA system and results reported by SRI and MIT-LL on SRE2010. Daniel Garcia-Romero, Xinhui Zhou, Dmitry N. Zotkin, Balaji Vasan Srinivasan, Yuancheng Luo, Sriram Ganapathy, Samuel Thomas 0001, Sridhar Krishna Nemala, Garimella S. V. S. Sivaram, Majid Mirbagheri, Sri Harish Reddy Mallidi, Thomas Janu, Padmanabhan Rajan, Nima Mesgarani, Mounya Elhilali, Hynek Hermansky, Shihab A. Shamma, Ramani Duraiswami |
ICASSP | 7 |
| 2012 | Multilingual MLP features for low-resource LVCSR systemsabstractWe introduce a new approach to training multilayer perceptrons (MLPs) for large vocabulary continuous speech recognition (LVCSR) in new languages which have only few hours of annotated in-domain training data (for example, 1 hour of data). In our approach, large amounts of annotated out-of-domain data from multiple languages are used to train multilingual MLP systems without dealing with the different phoneme sets for these languages. Features extracted from these MLP systems are used to train LVCSR systems in the low-resource language similar to the Tandem approach. In our experiments, the proposed features provide a relative improvement of about 30% in an low-resource LVCSR setting with only one hour of training data. Samuel Thomas 0001, Sriram Ganapathy, Hynek Hermansky |
ICASSP | 1 |
| 2012 | Intrinsic Spectral Analysis for Zero and High Resource Speech RecognitionabstractThe constraints of the speech production apparatus imply that our vocalizations are approximately restricted to a lowdimensional manifold embedded in a high-dimensional space. Manifold learning algorithms provide a means to recover the approximate embedding from untranscribed data and enable use of the manifold’s intrinsic distance metric to characterize acoustic similarity for downstream automatic speech applications. In this paper, we consider a previously unevaluated nonlinear outof-sample extension for intrinsic spectral analysis (ISA), investigating its performance in both unsupervised and supervised tasks. In the zero resource regime, where the lack of transcribed resources forces us to rely solely on the phonetic salience of the acoustic features themselves, ISA provides substantial gains relative to canonical acoustic front-ends. When large amounts of transcribed speech for supervised acoustic model training are also available, we find that the data-driven intrinsic spectrogram matches the performance of and is complementary to these signal processing derived counterparts. Index Terms: intrinsic spectral analysis, manifold learning, speech recognition, zero resource Aren Jansen, Samuel Thomas 0001, Hynek Hermansky |
INTERSPEECH | 2 |
| 2012 | Exploiting Discriminative Point Process Models for Spoken Term DetectionabstractState-of-the-art spoken term detection (STD) systems are built on top of large vocabulary speech recognition engines, which generate lattices that encode candidate occurrences of each invocabulary query. These lattices specifiy start and stop times of hypothesized term occurrences, providing a clear opportunity to return to the acoustics to incorporate novel confidence measures for verification. In this paper, we introduce a novel exemplar distance metric to the recently proposed discriminative point process modeling (DPPM) framework and use the resulting whole word models to generate STD confidence scores. In doing so, we introduce STD to a completely distinct acoustic modeling pipeline, trading Gaussian mixture models (GMM) for multi-layer perceptrons and replacing dictionary-derived hidden Markov models (HMM) with exemplar-based point process models. We find that whole word DPPM scores both perform comparably and are complementary to lattice posterior scores produced by a state-of-the-art speech recognition engine. Index Terms: spoken term detection, point process model, discriminative training, whole word model Atta Norouzian, Aren Jansen, Richard C. Rose, Samuel Thomas 0001 |
INTERSPEECH | 4 |
| 2012 | Data-driven Posterior Features for Low Resource Speech Recognition ApplicationsabstractIn low resource settings, with very few hours of training data, state-of-the-art speech recognition systems that require large amounts of task specific training data perform very poorly. We address this issue by building data-driven speech recognition front-ends on significant amounts of task independent data from different languages and genres collected in similar acoustic conditions as the data in the low resource scenario. We show that features derived from these trained front-ends perform significantly better and can alleviate the effect of reduced task specific training data in low resource settings. The proposed features provide a absolute improvement of about 12 % (18 % relative) in an low-resource LVCSR setting with only one hour of training data. We also demonstrate the usefulness of these features for zero-resource speech applications like spoken term discovery, which operate without any transcribed speech to train systems. The proposed features provide significant gains over conventional acoustic features on various information retrieval metrics for this task. Index Terms: Low-resource speech recognition, spoken term discovery, posterior features. Samuel Thomas 0001, Sriram Ganapathy, Aren Jansen, Hynek Hermansky |
INTERSPEECH | 1 |
| 2012 | Acoustic and Data-driven Features for Robust Speech Activity DetectionabstractIn this paper we evaluate different features for speech activity detection (SAD). Several signal processing techniques are used to derive acoustic features that capture attributes of speech useful in differentiating speech segments in noise. The acoustic features include short-term spectral features, long-term modulation features both derived using Frequency Domain Linear Prediction (FDLP), and joint spectro-temporal features extracted using 2D filters on a cortical representation of speech. Posteriors of speech and non-speech from a trained multi-layer perceptron are also used as data-driven features for this task. These feature extraction techniques form part of an elaborate feature extraction front-end where information spanning several hundreds of milliseconds of the signal are used along with heteroscedastic linear discriminant analysis for dimensionality reduction. Processed feature outputs from the proposed front-end are used to train SAD systems based on Gaussian mixture models for processing of speech from multiple languages transmitted over noisy radio communication channels under the ongoing DARPA Robust Automatic Transcription of Speech (RATS) program. The proposed front-end performs significantly better than standard acoustic feature extraction techniques in these noisy conditions. Samuel Thomas 0001, Sri Harish Reddy Mallidi, Thomas Janu, Hynek Hermansky, Nima Mesgarani, Xinhui Zhou, Shihab A. Shamma, Tim Ng, Bing Zhang 0004, Long Nguyen 0001, Spyridon Matsoukas |
INTERSPEECH | 1 |
| 2011 | MLP based phoneme detectors for Automatic Speech RecognitionabstractPhoneme posterior probabilities estimated using Multi-Layer Perceptrons (MLPs) are extensively used both as acoustic scores and features for speech recognition. In this paper we explore a different application of these posteriors as phonetic event detectors for speech recognition. We show how these detectors can be built to reliably capture phonetic events in the acoustic signal by integrating both acoustic and phonetic information about sound classes. These event detectors are used along with Segmental Conditional Random Fields (SCRFs) to improve the performance of speech recognition systems on the Broadcast News task. Samuel Thomas 0001, Patrick Nguyen, Geoffrey Zweig, Hynek Hermansky |
ICASSP | 1 |
| 2011 | Speech recognitionwith segmental conditional random fields: A summary of the JHU CLSP 2010 Summer WorkshopabstractThis paper summarizes the 2010 CLSP Summer Workshop on speech recognition at Johns Hopkins University. The key theme of the workshop was to improve on state-of-the-art speech recognition systems by using Segmental Conditional Random Fields (SCRFs) to integrate multiple types of information. This approach uses a state of-the-art baseline as a springboard from which to add a suite of novel features including ones derived from acoustic templates, deep neural net phoneme detections, duration models, modulation features, and whole word point-process models. The SCRF framework is able to appropriately weight these different information sources to produce significant gains on both die Broadcast News and Wall Street Journal tasks. Geoffrey Zweig, Patrick Nguyen, Dirk Van Compernolle, Kris Demuynck, Les E. Atlas, Pascal Clark, Gregory Sell, Meihong Wang, Fei Sha, Hynek Hermansky, Damianos Karakos, Aren Jansen, Samuel Thomas 0001, Sivaram G. S. V. S., Samuel R. Bowman, Justine T. Kao |
ICASSP | 13 |
| 2011 | Rapid Evaluation of Speech Representations for Spoken Term DiscoveryabstractAcoustic front-ends are typically developed for supervised learning tasks and are thus optimized to minimize word error rate, phone error rate, etc. However, in recent efforts to develop zero-resource speech technologies, the goal is not to use transcribed speech to train systems but instead to discover the acoustic structure of the spoken language automatically. For this new setting, we require a framework for evaluating the quality of speech representations without coupling to a particular recognition architecture. Motivated by the spoken term discovery task, we present a dynamic time warping-based framework for quantifying how well a representation can associate words of the same type spoken by different speakers. We benchmark the quality of a wide range of speech representations using multiple frame-level distance metrics and demonstrate that our performance metrics can also accurately predict phone recognition accuracies. Index Terms: evaluation methods, acoustic front-end, spoken term discovery, zero resource Michael A. Carlin, Samuel Thomas 0001, Aren Jansen, Hynek Hermansky |
INTERSPEECH | 2 |
| 2011 | Adaptive Stream Fusion in Multistream Recognition of SpeechabstractA new method to deal with variable distortions of speech during the operation of the system is proposed. First, multiple processing streams are formed by extracting different spectral and temporal modulation components from the speech signal. Information in each stream is used to estimate posterior probabilities of phonemes. Initial values for a weighted integration of these individual estimates are found by normalized cross-correlation of the estimates with the actual phoneme labels on the training data. A statistical model of the final estimated posterior probabilities is used to characterize the system performance. During the operation, the weights in the linear fusion are adapted using particle filtering to optimize the performance. Results on phoneme recognition from noisy speech indicate the effectiveness of the proposed method. Index Terms: multistream speech recognition, spectrotemporal modulations Nima Mesgarani, Samuel Thomas 0001, Hynek Hermansky |
INTERSPEECH | 2 |
| 2011 | Mixture of Auto-Associative Neural Networks for Speaker VerificationabstractThe paper introduces a mixture of auto-associative neural networks for speaker verification. A new objective function based on posterior probabilities of phoneme classes is used for training the mixture. This objective function allows each component of the mixture to model part of the acoustic space corresponding to a broad phonetic class. This paper also proposes how factor analysis can be applied in this setting. The proposed techniques show promising results on a subset of NIST-08 speaker recognition evaluation (SRE) and yield about 10% relative improvement when combined with the state-of-the-art Gaussian Mixture Model i-vector system. Garimella S. V. S. Sivaram, Samuel Thomas 0001, Hynek Hermansky |
INTERSPEECH | 2 |
| 2011 | The subspace Gaussian mixture model - A structured model for speech recognition
Daniel Povey, Lukás Burget, Mohit Agarwal 0005, Pinar Akyazi, Arnab Ghoshal, Ondrej Glembek, Nagendra K. Goel, Martin Karafiát, Ariya Rastrow, Richard C. Rose, Petr Schwarz, Samuel Thomas 0001 |
Comput. Speech Lang. | 13 |
| 2010 | Multilingual acoustic modeling for speech recognition based on subspace Gaussian Mixture ModelsabstractAlthough research has previously been done on multilingual speech recognition, it has been found to be very difficult to improve over separately trained systems. The usual approach has been to use some kind of “universal phone set” that covers multiple languages. We report experiments on a different approach to multilingual speech recognition, in which the phone sets are entirely distinct but the model has parameters not tied to specific states that are shared across languages. We use a model called a “Subspace Gaussian Mixture Model” where states' distributions are Gaussian Mixture Models with a common structure, constrained to lie in a subspace of the total parameter space. The parameters that define this subspace can be shared across languages. We obtain substantial WER improvements with this approach, especially with very small amounts of in-language training data. Lukás Burget, Petr Schwarz, Mohit Agarwal 0005, Pinar Akyazi, Arnab Ghoshal, Ondrej Glembek, Nagendra K. Goel, Martin Karafiát, Daniel Povey, Ariya Rastrow, Richard C. Rose, Samuel Thomas 0001 |
ICASSP | 13 |
| 2010 | Robust spectro-temporal features based on autoregressive models of Hilbert envelopesabstractIn this paper, we present a robust spectro-temporal feature extraction technique using autoregressive models (AR) of sub-band Hilbert envelopes. AR models of Hilbert envelopes are derived using frequency domain linear prediction (FDLP). From the sub-band Hilbert envelopes, spectral features are derived by integrating these envelopes in short-term frames and the temporal features are formed by converting these envelopes into modulation frequency components. The spectral and temporal feature streams are then combined at the phoneme posterior level and are used as the input features for a recognition system. For the proposed features, robustness is achieved by using novel techniques of noise compensation and gain normalization. Phoneme recognition experiments on telephone speech in the HTIMIT database show significant performance improvements for the proposed features when compared to other robust feature techniques (average relative reduction of 10.6 % in phoneme error rate). In addition to the overall phoneme recognition rates, the performance with broad phonetic classes is also reported. Sriram Ganapathy, Samuel Thomas 0001, Hynek Hermansky |
ICASSP | 2 |
| 2010 | Comparison of modulation features for phoneme recognitionabstractIn this paper, we compare several approaches for the extraction of modulation frequency features from speech signal using a phoneme recognition system. The general framework in these approaches is to decompose the speech signal into a set of sub-bands. Amplitude modulations (AM) in the sub-band signal are used to derive features for automatic speech recognition (ASR). Then, we propose a feature extraction technique which uses autoregressive models (AR) of sub-band Hilbert envelopes in relatively long segments of speech signal. AR models of Hilbert envelopes are derived using frequency domain linear prediction (FDLP). Features are formed by converting the FDLP envelopes into static and dynamic modulation frequency components. In the phoneme recognition experiments using the TIMIT database, the FDLP based modulation frequency features provide significant improvements compared to other techniques (average relative improvement of 7.5% over the base-line features). Furthermore, a detailed analysis is performed to determine the relative contribution of various processing stages in the proposed technique. Sriram Ganapathy, Samuel Thomas 0001, Hynek Hermansky |
ICASSP | 2 |
| 2010 | A novel estimation of feature-space MLLR for full-covariance modelsabstractIn this paper we present a novel approach for estimating feature-space maximum likelihood linear regression (fMLLR) transforms for full-covariance Gaussian models by directly maximizing the likelihood function by repeated line search in the direction of the gradient. We do this in a pre-transformed parameter space such that an approximation to the expected Hessian is proportional to the unit matrix. The proposed algorithm is as efficient or more efficient than standard approaches, and is more flexible because it can naturally be combined with sets of basis transforms and with full covariance and subspace precision and mean (SPAM) models. Arnab Ghoshal, Daniel Povey, Mohit Agarwal 0005, Pinar Akyazi, Lukás Burget, Ondrej Glembek, Nagendra K. Goel, Martin Karafiát, Ariya Rastrow, Richard C. Rose, Petr Schwarz, Samuel Thomas 0001 |
ICASSP | 13 |
| 2010 | Approaches to automatic lexicon learning with limited training examplesabstractPreparation of a lexicon for speech recognition systems can be a significant effort in languages where the written form is not exactly phonetic. On the other hand, in languages where the written form is quite phonetic, some common words are often mispronounced. In this paper, we use a combination of lexicon learning techniques to explore whether a lexicon can be learned when only a small lexicon is available for boot-strapping. We discover that for a phonetic language such as Spanish, it is possible to do that better than what is possible from generic rules or hand-crafted pronunciations. For a more complex language such as English, we find that it is still possible but with some loss of accuracy. Nagendra K. Goel, Samuel Thomas 0001, Mohit Agarwal 0005, Pinar Akyazi, Lukás Burget, Arnab Ghoshal, Ondrej Glembek, Martin Karafiát, Daniel Povey, Ariya Rastrow, Richard C. Rose, Petr Schwarz |
ICASSP | 2 |
| 2010 | Subspace Gaussian Mixture Models for speech recognitionabstractWe describe an acoustic modeling approach in which all phonetic states share a common Gaussian Mixture Model structure, and the means and mixture weights vary in a subspace of the total parameter space. We call this a Subspace Gaussian Mixture Model (SGMM). Globally shared parameters define the subspace. This style of acoustic model allows for a much more compact representation and gives better results than a conventional modeling approach, particularly with smaller amounts of training data. Daniel Povey, Lukás Burget, Mohit Agarwal 0005, Pinar Akyazi, Arnab Ghoshal, Ondrej Glembek, Nagendra K. Goel, Martin Karafiát, Ariya Rastrow, Richard C. Rose, Petr Schwarz, Samuel Thomas 0001 |
ICASSP | 13 |
| 2010 | A multistream multiresolution framework for phoneme recognitionabstractSpectrotemporal representation of speech has already shown promising results in speech processing technologies, however, many inherent issues of such representation, such as high dimensionality have limited their use in speech and speaker recognition. Multistream framework fits very well to such representation where different regions can be separately mapped into posterior probabilities of classes before merging. In this study, we investigated the effective ways of forming streams out of this representation for robust phoneme recognition. We also investigated multiple ways of fusing the posteriors of different streams based on their individual confidence or interactions between them. We observed 8.6% relative improvement in clean and 4 % in noise. We developed a simple yet effective linear combination technique that provides intuitive understanding of stream combinations and how even systematic errors can be learned to reduce confusions. Index Terms: speech recognition, spectrotemporal modulations, multistream Nima Mesgarani, Samuel Thomas 0001, Hynek Hermansky |
INTERSPEECH | 2 |
| 2010 | Cross-lingual and multi-stream posterior features for low resource LVCSR systemsabstractWe investigate approaches for large vocabulary continuous speech recognition (LVCSR) system for new languages or new domains using limited amounts of transcribed training data. In these low resource conditions, the performance of conventional LVCSR systems degrade significantly. We propose to train low resource LVCSR system with additional sources of information like annotated data from other languages (German and Spanish) and various acoustic feature streams (short-term and modulation features). We train multilayer perceptrons (MLPs) on these sources of information and use Tandem features derived from the MLPs for low resource LVCSR. In our experiments, the proposed system trained using only one hour of English conversational telephone speech (CTS) provides a relative improvement of 11% over the baseline system. Index Terms: Cross-lingual posterior features, Multi-stream features, Low resource ASR, Tandem features Samuel Thomas 0001, Sriram Ganapathy, Hynek Hermansky |
INTERSPEECH | 1 |
| 2010 | A phoneme recognition framework based on auditory spectro-temporal receptive fieldsabstractWe propose to incorporate features derived using spectrotemporal receptive fields (STRFs) of neurons in the auditory cortex for phoneme recognition. Each of these STRFs is tuned to different auditory frequencies, scales and modulation rates. We select different sets of STRFs which are specific for phonemes in different broad phonetic classes (BPC) of sounds. These STRFs are then used as spectro-temporal filters on spectrograms of speech to extract features for phoneme recognition. For the phoneme recognition task on the TIMIT database, the proposed features show a relative improvement of about 5% over conventional feature extraction techniques. Samuel Thomas 0001, Kailash Patil, Sriram Ganapathy, Nima Mesgarani, Hynek Hermansky |
INTERSPEECH | 1 |
| 2009 | Temporal envelope subtraction for robust speech recognition using modulation spectrumabstractIn this paper, we present a new noise compensation technique for modulation frequency features derived from syllable length segments of subband temporal envelopes. The subband temporal envelopes are estimated using frequency domain linear prediction (FDLP). We propose a technique for noise compensation in FDLP where an estimate of the noise envelope is subtracted from the noisy speech envelope. The noise compensated FDLP envelopes are compressed with static (logarithmic) and dynamic (adaptive loops) compression and are transformed into modulation spectral features. Experiments are performed on a phoneme recognition task as well as a connected digit recognition task where the test data is corrupted with variety of noise types at different signal to noise ratios. In these experiments with mismatched train and test conditions, the proposed features provide considerable improvements compared to other state of the art noise robust feature extraction techniques (average relative improvement of 25% and 35% over the baseline PLP features for phoneme and word recognition tasks respectively). Sriram Ganapathy, Samuel Thomas 0001, Hynek Hermansky |
ASRU | 2 |
| 2009 | Phoneme recognition using spectral envelope and modulation frequency featuresabstractWe present a new feature extraction technique for phoneme recognition that uses short-term spectral envelope and modulation frequency features. These features are derived from sub-band temporal envelopes of speech estimated using frequency domain linear prediction (FDLP). While spectral envelope features are obtained by the short-term integration of the sub-band envelopes, the modulation frequency components are derived from the long-term evolution of the sub-band envelopes. These features are combined at the phoneme posterior level and used as features for a hybrid HMM-ANN phoneme recognizer. For the phoneme recognition task on the TIMIT database, the proposed features show an improvement of 4.7% over the other feature extraction techniques. Samuel Thomas 0001, Sriram Ganapathy, Hynek Hermansky |
ICASSP | 1 |
| 2009 | Static and dynamic modulation spectrum for speech recognitionabstractWe present a feature extraction technique based on static and dynamic modulation spectrum derived from long-term envelopes in sub-bands. Estimation of the sub-band temporal envelopes is done using Frequency Domain Linear Prediction (FDLP). These sub-band envelopes are compressed with a static (logarithmic) and dynamic (adaptive loops) compression. The compressed sub-band envelopes are transformed into modulation spectral components which are used as features for speech recognition. Experiments are performed on a phoneme recognition task using a hybrid HMM-ANN phoneme recognition system and an ASR task using the TANDEM speech recognition system. The proposed features provide a relative improvements of 3.8 % and 11.5 % in phoneme recognition accuracies for TIMIT and conversation telephone speech (CTS) respectively. Further, these improvements are found to be consistent for ASR tasks on OGI-Digits database (relative improvement of 13.5 %). Sriram Ganapathy, Samuel Thomas 0001, Hynek Hermansky |
INTERSPEECH | 2 |
| 2009 | Tandem representations of spectral envelope and modulation frequency features for ASRabstractWe present a feature extraction technique for automatic speech recognition that uses Tandem representation of short-term spectral envelope and modulation frequency features. These features, derived from sub-band temporal envelopes of speech estimated using frequency domain linear prediction, are combined at the phoneme posterior level. Tandem representations derived from these phoneme posteriors are used along with HMM based ASR systems for both small and large vocabulary continuous speech recognition (LVCSR) tasks. For a small vocabulary continuous digit task on the OGI Digits database, the proposed features reduce the word error rate (WER) by 13 % relative to other feature extraction techniques. We obtain a relative reduction of about 14 % in WER for an LVCSR task using the NIST RT05 evaluation data. For phoneme recognition tasks on the TIMIT database these features provide a relative improvement of 13% compared to other techniques. Samuel Thomas 0001, Sriram Ganapathy, Hynek Hermansky |
INTERSPEECH | 1 |
| 2008 | Front-end for far-field speech recognition based on frequency domain linear predictionabstractAbstract. Automatic Speech Recognition (ASR) systems usually fail when they encounter speech from far-field microphone in reverberant environments. This is due to the application of short-term feature extraction techniques which do not compensate for the artifacts introduced by long room impulse responses. In this paper, we propose a front-end, based on Frequency Domain Linear Prediction (FDLP), that tries to remove reverberation artifacts present in far-field speech. Long temporal segments of far-field speech are analyzed in narrow frequency sub-bands to extract FDLP envelopes and residual signals. Filtering the residual signals with gain normalized inverse FDLP filters result in a set of sub-band signals which are synthesized to reconstruct the signal back. ASR experiments on far-field speech data processed by the proposed front-end show significant improvements (relative reduction of 30 % in word error rate) compared to other robust feature extraction techniques. 2 IDIAP–RR 08-17 1 Sriram Ganapathy, Samuel Thomas 0001, Hynek Hermansky |
INTERSPEECH | 2 |
| 2008 | Hilbert envelope based spectro-temporal features for phoneme recognition in telephone speechabstractAbstract. In this paper, we present a spectro-temporal feature extraction technique using sub-band Hilbert envelopes of relatively long segments of speech signal. Hilbert envelopes of the sub-bands are estimated using Frequency Domain Linear Prediction (FDLP). Spectral features are derived by integrating the sub-band Hilbert envelopes in short-term frames and the temporal features are formed by converting the FDLP envelopes into modulation frequency components. These are then combined at the phoneme posterior level and are used as the input features for a phoneme recognition system. In order to improve the robustness of the proposed features to telephone speech, the sub-band temporal envelopes are gain normalized prior to feature extraction. Phoneme recognition experiments on telephone speech in the HTIMIT database show significant performance improvements for the proposed features when compared to other robust feature techniques (average relative reduction of 11 % in phoneme error rate).2 IDIAP–RR 08-18 n this paper, we present a spectro-temporal feature extraction technique using sub-band Hilbert envelopes of relatively long segments of speech signal. Hilbert envelopes of the sub-bands are estimated using Frequency Domain Linear Prediction (FDLP). Spectral features are derived by integrating the sub-band Hilbert envelopes in short-term frames and the temporal features are formed by converting Samuel Thomas 0001, Sriram Ganapathy, Hynek Hermansky |
INTERSPEECH | 1 |
| 2008 | Recognition of Reverberant Speech Using Frequency Domain Linear PredictionabstractPerformance of a typical automatic speech recognition (ASR) system severely degrades when it encounters speech from reverberant environments. Part of the reason for this degradation is the feature extraction techniques that use analysis windows which are much shorter than typical room impulse responses. We present a feature extraction technique based on modeling temporal envelopes of the speech signal in narrow subbands using frequency domain linear prediction (FDLP). FDLP provides an all-pole approximation of the Hilbert envelope of the signal obtained by linear prediction on cosine transform of the signal. ASR experiments on speech data degraded with a number of room impulse responses (with varying degrees of distortion) show significant performance improvements for the proposed FDLP features when compared to other robust feature extraction techniques (average relative reduction of 24% in word error rate). Similar improvements are also obtained for far-field data which contain natural reverberation in background noise. These results are achieved without any noticeable degradation in performance for clean speech. Samuel Thomas 0001, Sriram Ganapathy, Hynek Hermansky |
IEEE Signal Process. Lett. | 1 |
| 2007 | Language identification of person names using CF-IOF based weighing functionabstractInformation about the language of origin helps in generating pronunciation for foreign words, specially person names, in a text-to-speech synthesis system. It can be used to apply language specific letter-to-sound (LTS) rules to these words during synthesis. In this paper, we propose a novel approach for using substrings of a person name (called letter N-grams) to identify the language of its origin. We use a weight for the letter N-grams that is motivated by the techniques used in text document classification, different from the usual N-gram probabilities used in earlier approaches. We also propose a tree based approach to select the letter N-grams of different lengths for language identification. Several experiments have been conducted to evaluate the performance of the proposed approach and compare it with those of the earlier proposed approaches based on N-gram probabilities. We show an improvement in classification results over the earlier approaches without using any language specific rules. Samuel Thomas 0001, Ashish Verma 0001 |
INTERSPEECH | 1 |