EDBT 2026 Demo / reviewers in the wild / expert
James R. Glass
dblp:37/6580 · also James Glass 0001, Jim Glass 0001
· DBLP profile ↗
336ranked-venue papers
14as first author
72since 2021 · last 2026
0000-0002-3097-360XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 242 · 12 first-author · 41 since 2021Artificial intelligence and machine learning · 236 · 9 first-author · 57 since 2021Human-computer interaction and ubiquitous computing · 4Systems, architecture and hardware · 2Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CMKD: CNN/Transformer-Based Cross-Model Knowledge Distillation for Audio ClassificationabstractAudio classification is an active research area with a wide range of applications. Over the past decade, convolutional neural networks (CNNs) have been the de-facto standard building block for end-to-end audio classification models. Recently, neural networks based solely on self-attention mechanisms such as the Audio Spectrogram Transformer (AST) have been shown to outperform CNNs. In this paper, we find an intriguing interaction between the two very different models - CNN and AST models are good teachers for each other. When we use either of them as the teacher and train the other model as the student via knowledge distillation (KD), the performance of the student model noticeably improves, and in many cases, is better than the teacher model. In our experiments with this CNN/Transformer Cross-Model Knowledge Distillation (CMKD) method we achieve new state-of-the-art performance on FSD50 K, AudioSet, and ESC-50. Yuan Gong 0001, Sameer Khurana, Andrew Rouditchenko, James R. Glass |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Towards unsupervised speech recognition without pronunciation models
Junrui Ni, Liming Wang 0003, Yang Zhang 0001, Kaizhi Qian, Heting Gao, Mark Hasegawa-Johnson, James R. Glass, Chang Dong Yoo |
Speech Commun. | 7 |
| 2025 | Decoding on Graphs: Faithful and Sound Reasoning on Knowledge Graphs through Generation of Well-Formed ChainsabstractKnowledge Graphs (KGs) can serve as reliable knowledge sources for question answering (QA) due to their structured representation of knowledge.Existing research on the utilization of KG for large language models (LLMs) prevalently relies on subgraph retriever or iterative prompting, overlooking the potential synergy of LLMs' step-wise reasoning capabilities and KGs' structural nature.In this paper, we present DoG (Decoding on Graphs), a novel framework that facilitates a deep synergy between LLMs and KGs.We first define a concept, well-formed chain, which consists of a sequence of interrelated fact triplets on the KGs, starting from question entities and leading to answers.We argue that this concept can serve as a principle for making faithful and sound reasoning for KGQA.To enable LLMs to generate well-formed chains, we propose graph-aware constrained decoding, in which a constraint derived from the topology of the KG regulates the decoding process of the LLMs.This constrained decoding method ensures the generation of well-formed chains while making full use of the step-wise reasoning capabilities of LLMs.Based on the above, DOG, a trainingfree approach, is able to provide faithful and sound reasoning trajectories grounded on the KGs.Experiments across various KGQA tasks with different background KGs demonstrate that DOG achieves superior and robust performance.DOG also shows general applicability with various open-source LLMs 1 .* Equal contribution. 1 The code is available here. Kun Li 0003, Tianhua Zhang, Xixin Wu, Hongyin Luo, James R. Glass, Helen M. Meng |
ACL (1) | 5 |
| 2025 | USAD: Universal Speech and Audio Representation via DistillationabstractSelf-supervised learning (SSL) has revolutionized audio representations, yet models often remain domain-specific, focusing on either speech or non-speech tasks. In this work, we present Universal Speech and Audio Distillation (USAD), a unified approach to audio representation learning that integrates diverse audio types—speech, sound, and music—into a single model. USAD employs efficient layer-to-layer distillation from domain-specific SSL models to train a student on a comprehensive audio dataset. USAD offers competitive performance across various benchmarks and datasets, including frame and instance-level speech processing tasks, audio tagging, and sound classification, achieving near state-of-the-art results with a single encoder on SUPERB and HEAR benchmarks.11Models: https://huggingface.co/MIT-SLS/USAD-Base Heng-Jui Chang, Saurabhchand Bhati, James R. Glass, Alexander H. Liu |
ASRU | 3 |
| 2025 | Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?abstractWe propose Omni-R1 which fine-tunes a recent multi-modal LLM, Qwen2.5-Omni, on an audio question answering dataset with the reinforcement learning method GRPO. This leads to new State-of-the-Art performance on the recent MMAU and MMAR benchmarks. On MMAU, Omni-R1 achieves the highest accuracies on the sounds, music, speech, and overall average categories, both on the Test-mini and Test-full splits. To understand the performance improvement, we tested models both with and without audio and found that much of the performance improvement from GRPO could be attributed to better text-based reasoning. We also made a surprising discovery that fine-tuning without audio on a text-only dataset was effective at improving the audio-based performance. Andrew Rouditchenko, Saurabhchand Bhati, Edson Araujo, Samuel Thomas 0001, Hilde Kuehne, Rogério Feris, James R. Glass |
ASRU | 7 |
| 2025 | Recognizing Dementia from Neuropsychological Tests with State Space ModelsabstractEarly detection of dementia is critical for timely medical intervention and improved patient outcomes. Neuropsychological tests are widely used for cognitive assessment but have traditionally relied on manual scoring. Automatic dementia classification (ADC) systems aim to infer cognitive decline directly from speech recordings of such tests. We propose Demenba, a novel ADC framework based on state space models, which scale linearly in memory and computation with sequence length. Trained on over 1,000 hours of cognitive assessments administered to Framingham Heart Study participants, some of whom were diagnosed with dementia through adjudicated review, our method outperforms prior approaches in fine-grained dementia classification by 21%, while using fewer parameters. We further analyze its scaling behavior and demonstrate that our model gains additional improvement when fused with large language models, paving the way for more transparent and scalable dementia assessment tools12.1Code: https://github.com/lwang114/Demenba2This work was supported by the Framingham Heart Study’s National Heart, Lung, and Blood Institute contract N01-HC-25195; National Institutes of Health grants U19-AG068753, R01- AG016495, R01-AG008122, R01AG033040. The authors would also like to thank the staff and participants of the Framingham Heart Study. Liming Wang 0003, Saurabhchand Bhati, Cody Karjadi, Rhoda Au, James R. Glass |
ASRU | 5 |
| 2025 | CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained AlignmentabstractRecent advances in audio-visual learning have shown promising results in learning representations across modalities. However, most approaches rely on global audio representations that fail to capture fine-grained temporal correspondences with visual frames. Additionally, existing methods often struggle with conflicting optimization objectives when trying to jointly learn reconstruction and cross-modal alignment. In this work, we propose CAV-MAE Sync as a simple yet effective extension of the original CAV-MAE [14] framework for self-supervised audio-visual learning. We address three key challenges: First, we tackle the granularity mismatch between modalities by treating audio as a temporal sequence aligned with video frames, rather than using global representations. Second, we resolve conflicting optimization goals by separating contrastive and reconstruction objectives through dedicated global tokens. Third, we improve spatial localization by introducing learnable register tokens that reduce the semantic load on patch tokens. We evaluate the proposed approach on AudioSet, VGG Sound, and the ADE20K Sound dataset on zero-shot retrieval, classification, and localization tasks demonstrating state-of-the-art performance and outperforming more complex architectures. Code is available at https://github.com/edsonroteia/cav-mae-sync. Edson Araujo, Andrew Rouditchenko, Yuan Gong 0001, Saurabhchand Bhati, Samuel Thomas 0001, Brian Kingsbury, Leonid Karlinsky, Rogério Feris, James R. Glass, Hilde Kuehne |
CVPR | 9 |
| 2025 | RAG-Zeval: Enhancing RAG Responses Evaluator through End-to-End Reasoning and Ranking-Based Reinforcement LearningabstractRobust evaluation is critical for deploying trustworthy retrieval-augmented generation (RAG) systems.However, current LLM-based evaluation frameworks predominantly rely on directly prompting resource-intensive models with complex multi-stage prompts, underutilizing models' reasoning capabilities and introducing significant computational cost.In this paper, we present RAG-Zeval (RAG-Zero Evaluator), a novel end-to-end framework that formulates faithfulness and correctness evaluation of RAG systems as a rule-guided reasoning task.Our approach trains evaluators with reinforcement learning, facilitating compact models to generate comprehensive and sound assessments with detailed explanation in onepass.We introduce a ranking-based outcome reward mechanism, using preference judgments rather than absolute scores, to address the challenge of obtaining precise pointwise reward signals.To this end, we synthesize the ranking references by generating quality-controlled responses with zero human annotation.Experiments demonstrate RAG-Zeval's superior performance, achieving the strongest correlation with human judgments and outperforming baselines that rely on LLMs with 10 -100× more parameters.Our approach also exhibits superior interpretability in response evaluation 1 . Kun Li 0003, Tianhua Zhang, Hongyin Luo, Xixin Wu, James R. Glass, Helen M. Meng |
EMNLP | 6 |
| 2025 | Teaching VLMs to Localize Specific Objects from In-Context ExamplesabstractVision-Language Models (VLMs) have shown remarkable capabilities across diverse visual tasks, including image recognition, video understanding, and Visual Question Answering (VQA) when explicitly trained for these tasks. Despite these advances, we find that present-day VLMs (including the proprietary GPT-4o) lack a fundamental cognitive ability: learning to localize specific objects in a scene by taking into account the context. In this work, we focus on the task of few-shot personalized localization, where a model is given a small set of annotated images (in-context examples) -- each with a category label and bounding box -- and is tasked with localizing the same object type in a query image. Personalized localization can be particularly important in cases of ambiguity of several related objects that can respond to a text or an object that is hard to describe with words. To provoke personalized localization abilities in models, we present a data-centric solution that fine-tunes them using carefully curated data from video object tracking datasets. By leveraging sequences of frames tracking the same object across multiple shots, we simulate instruction-tuning dialogues that promote context awareness. To reinforce this, we introduce a novel regularization technique that replaces object labels with pseudo-names, ensuring the model relies on visual context rather than prior knowledge. Our method significantly enhances the few-shot localization performance of recent VLMs ranging from 7B to 72B in size, without sacrificing generalization, as demonstrated on several benchmarks tailored towards evaluating personalized localization abilities. This work is the first to explore and benchmark personalized few-shot localization for VLMs -- exposing critical weaknesses in present-day VLMs, and laying a foundation for future research in context-driven vision-language applications. Sivan Doveh, Nimrod Shabtay, Eli Schwartz, Hilde Kuehne, Raja Giryes, Rogério Feris, Leonid Karlinsky, James R. Glass, Assaf Arbelle, Shimon Ullman, Muhammad Jehanzeb Mirza |
ICCV | 8 |
| 2025 | Self-MoE: Towards Compositional Large Language Models with Self-Specialized ExpertsabstractWe present Self-MoE, an approach that transforms a monolithic LLM into a compositional, modular system of self-specialized experts, named MiXSE (MiXture of Self-specialized Experts). Our approach leverages self-specialization, which constructs expert modules using self-generated synthetic data, each equipping a shared base LLM with distinct domain-specific capabilities, activated via self-optimized routing. This allows for dynamic and capability-specific handling of various target tasks, enhancing overall capabilities, without extensive human-labeled data and added parameters. Our empirical results reveal that specializing LLMs may exhibit potential trade-offs in performances on non-specialized tasks. On the other hand, our Self-MoE demonstrates substantial improvements (6.5%p on average) over the base LLM across diverse benchmarks such as knowledge, reasoning, math, and coding. It also consistently outperforms other methods, including instance merging and weight merging, while offering better flexibility and interpretability by design with semantic experts and routing. Our findings highlight the critical role of modularity, the applicability of Self-MoE to multiple base LLMs, and the potential of self-improvement in achieving efficient, scalable, and adaptable systems. Junmo Kang, Leonid Karlinsky, Hongyin Luo, Zhen Wang 0041, Jacob A. Hansen, James R. Glass, David D. Cox, Rameswar Panda, Rogério Feris, Alan Ritter |
ICLR | 6 |
| 2025 | UniWav: Towards Unified Pre-training for Speech Representation Learning and GenerationabstractPre-training and representation learning have been playing an increasingly important role in modern speech processing. Nevertheless, different applications have been relying on different foundation models, since predominant pre-training techniques are either designed for discriminative tasks or generative tasks. In this work, we make the first attempt at building a unified pre-training framework for both types of tasks in speech. We show that with the appropriate design choices for pre-training, one can jointly learn a representation encoder and generative audio decoder that can be applied to both types of tasks. We propose UniWav, an encoder-decoder framework designed to unify pre-training representation learning and generative tasks. On speech recognition, text-to-speech, and speech tokenization, UniWav achieves comparable performance to different existing foundation models, each trained on a specific task. Our findings suggest that a single general-purpose foundation model for speech can be built to replace different foundation models, reducing the overhead and cost of pre-training. Alexander H. Liu, Sang-gil Lee, Chao-Han Huck Yang, Yuan Gong 0001, Yu-Chiang Frank Wang, James R. Glass, Rafael Valle, Bryan Catanzaro |
ICLR | 6 |
| 2025 | Quantifying Generalization Complexity for Large Language ModelsabstractWhile large language models (LLMs) have shown exceptional capabilities in understanding complex queries
and performing sophisticated tasks, their generalization abilities are often deeply entangled with
memorization, necessitating more precise evaluation.
To address this challenge, we introduce Scylla, a dynamic evaluation framework that quantitatively measures the generalization abilities of LLMs. Scylla disentangles generalization from memorization via assessing model performance on both in-distribution (ID) and out-of-distribution (OOD) data through 20 tasks across 5 levels of complexity.
Through extensive experiments, we uncover a non-monotonic relationship between task complexity and the performance gap between ID and OOD
data, which we term the generalization valley.
Specifically, this phenomenon reveals a critical threshold---referred to
as critical complexity---where reliance on non-generalizable behavior peaks, indicating the
upper bound of LLMs' generalization capabilities.
As model size increases, the critical complexity shifts toward higher levels of task complexity,
suggesting that larger models can handle more complex reasoning tasks before over-relying on
memorization.
Leveraging Scylla and the concept of critical complexity, we benchmark 28 LLMs including
both open-sourced models such as LLaMA and Qwen families, and closed-sourced models like Claude and
GPT, providing a more robust evaluation and establishing a clearer
understanding of LLMs' generalization capabilities. Zhenting Qi, Hongyin Luo, Xuliang Huang, Zhuokai Zhao, Yibo Jiang, Xiangjun Fan, Himabindu Lakkaraju, James R. Glass |
ICLR | 8 |
| 2025 | SelfCite: Self-Supervised Alignment for Context Attribution in Large Language ModelsabstractWe introduce SelfCite, a novel self-supervised approach that aligns LLMs to generate high-quality, fine-grained, sentence-level citations for the statements in their generated responses. Instead of only relying on costly and labor-intensive annotations, SelfCite leverages a reward signal provided by the LLM itself through context ablation: If a citation is necessary, removing the cited text from the context should prevent the same response; if sufficient, retaining the cited text alone should preserve the same response. This reward can guide the inference-time best-of-N sampling strategy to improve citation quality significantly, as well as be used in preference optimization to directly fine-tune the models for generating better citations. The effectiveness of SelfCite is demonstrated by increasing citation F1 up to 5.3 points on the LongBench-Cite benchmark across five long-form question answering tasks. The source code is available at https://github.com/facebookresearch/SelfCite. Yung-Sung Chuang, Benjamin Cohen-Wang, Shannon Shen 0001, Zhaofeng Wu, Hu Xu 0001, Xi Victoria Lin, James R. Glass, Shang-Wen Li 0001, Scott Yih |
ICML | 7 |
| 2025 | DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models
Heng-Jui Chang, Hongyu Gong, Changhan Wang, James R. Glass, Yu-An Chung |
INTERSPEECH | 4 |
| 2025 | THREAD: Thinking Deeper with Recursive SpawningabstractPhilip Schroeder, Nathaniel W. Morgan, Hongyin Luo, James R. Glass. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Philip Schroeder, Nathaniel Morgan, Hongyin Luo, James R. Glass |
NAACL (Long Papers) | 4 |
| 2025 | Meta CLIP 2: A Worldwide Scaling RecipeabstractContrastive Language-Image Pretraining (CLIP) is a popular foundation model, supporting from zero-shot classification, retrieval to encoders for multimodal large language models (MLLMs). Although CLIP is successfully trained on billion-scale image-text pairs from the English world, scaling CLIP's training further to learning from the worldwide web data is still challenging: (1) no curation method is available to handle data points from non-English world; (2) the English performance from existing multilingual CLIP is worse than its English-only counterpart, i.e., "curse of multilinguality" that is common in LLMs. Here, we present Meta CLIP 2, the first recipe training CLIP from scratch on worldwide web-scale image-text pairs. To generalize our findings, we conduct rigorous ablations with minimal changes that are necessary to address the above challenges and present a recipe enabling mutual benefits from English and non-English world data.
In zero-shot ImageNet classification, Meta CLIP 2 ViT-H/14 surpasses its English-only counterpart by 0.8% and mSigLIP by 0.7%, and surprisingly sets new state-of-the-art without system-level confounding factors (e.g., translation, bespoke architecture changes) on multilingual benchmarks, such as CVQA with 57.4%, Babel-ImageNet with 50.2% and XM3600 with 64.3% on image-to-text retrieval. Code and model are available at https://github.com/facebookresearch/MetaCLIP. Yung-Sung Chuang, Ching-Feng Yeh, Kehan Lyu, Ramya Raghavendra, James R. Glass, Lifei Huang, Jason Weston, Luke Zettlemoyer, Xinlei Chen, Zhuang Liu 0003, Saining Xie, Scott Yih, Shang-Wen Li 0001, Hu Xu 0001 |
NeurIPS | 7 |
| 2025 | ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied TasksabstractVision-language models (VLMs) have exhibited impressive capabilities across diverse image understanding tasks, but still struggle in settings that require reasoning over extended sequences of camera frames from a video. This limits their utility in embodied settings, which require reasoning over long frame sequences from a continuous stream of visual input at each moment of a task attempt. To address this limitation, we propose ROVER (Reasoning Over VidEo Recursively), a framework that enables the model to recursively decompose long-horizon video trajectories into segments corresponding to shorter subtasks within the trajectory. In doing so, ROVER facilitates more focused and accurate reasoning over temporally localized frame sequences without losing global context. We evaluate ROVER, implemented using an in-context learning approach, on diverse OpenX Embodiment videos and on a new dataset derived from RoboCasa that consists of 543 videos showing both expert and perturbed non-expert trajectories across 27 manipulation tasks. ROVER outperforms strong baselines across three video reasoning tasks: task progress estimation, frame-level natural language reasoning, and video question answering. We observe that, by reducing the number of frames the model reasons over at each timestep, ROVER mitigates model hallucinations, especially during unexpected or non-optimal moments of a trajectory. In addition, by enabling the implementation of a subtask-specific sliding context window, ROVER's time complexity scales linearly with video length, an asymptotic improvement over baselines. Philip Schroeder, Ondrej Biza, Thomas Weng, Hongyin Luo, James R. Glass |
NeurIPS | 5 |
| 2025 | Can Diffusion Models Disentangle? A Theoretical PerspectiveabstractThis paper presents a novel theoretical framework for understanding how diffusion models can learn disentangled representations with commonly used weak supervision such as partial labels and multiple views. Within this framework, we establish identifiability conditions for diffusion models to disentangle latent variable models with \emph{stochastic}, \emph{non-invertible} mixing processes. We also prove \emph{finite-sample global convergence} for diffusion models to disentangle independent subspace models. To validate our theory, we conduct extensive disentanglement experiments on subspace recovery in latent subspace Gaussian mixture models, image colorization, denoising, and voice conversion for speech classification. Our experiments show that training strategies inspired by our theory, such as style guidance regularization, consistently enhance disentanglement performance. Liming Wang 0003, Muhammad Jehanzeb Mirza, Yishu Gong, Yuan Gong 0001, Jiaqi Zhang 0006, Brian Tracey, Katerina Placek, Marco Vilela, James R. Glass |
NeurIPS | 9 |
| 2025 | mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech RecognitionabstractAudio-Visual Speech Recognition (AVSR) combines lip-based video with audio and can improve performance in noise, but most methods are trained only on English data. One limitation is the lack of large-scale multilingual video data, which makes it hard to train models from scratch. In this work, we propose mWhisper-Flamingo for multilingual AVSR which combines the strengths of a pre-trained audio model (Whisper) and video model (AV-HuBERT). To enable better multi-modal integration and improve the noisy multilingual performance, we introduce decoder modality dropout where the model is trained both on paired audio-visual inputs and separate audio/visual inputs. mWhisper-Flamingo achieves state-of-the-art WER on MuAViC, an AVSR dataset of 9 languages. Audio-visual mWhisper-Flamingo consistently outperforms audio-only Whisper on all languages in noisy conditions. Andrew Rouditchenko, Samuel Thomas 0001, Hilde Kuehne, Rogério Feris, James R. Glass |
IEEE Signal Process. Lett. | 5 |
| 2024 | What, When, and Where? Self-Supervised Spatio- Temporal Grounding in Untrimmed Multi-Action Videos from Narrated InstructionsabstractSpatio-temporal grounding describes the task of localizing events in space and time, e.g., in video data, based on verbal descriptions only. Models for this task are usually trained with human-annotated sentences and bounding box supervision. This work addresses this task from a multimodal supervision perspective, proposing a framework for spatio-temporal action grounding trained on loose video and subtitle supervision only, without human annotation. To this end, we combine local representation learning, which focuses on leveraging fine-grained spatial information, with a global representation encoding that captures higher-level representations and incorporates both in a joint approach. To evaluate this challenging task in a real-life setting, a new benchmark dataset is proposed, providing dense spatio-temporal grounding annotations in long, untrimmed, multi-action instructional videos for over 5K events. We evaluate the proposed approach and other methods on the proposed and standard downstream tasks, showing that our method improves over current baselines in various settings, including spatial, temporal, and untrimmed multi-action spatio-temporal grounding. Brian Chen 0001, Nina Shvetsova, Andrew Rouditchenko, Daniel Kondermann, Samuel Thomas 0001, Shih-Fu Chang, Rogério Feris, James R. Glass, Hilde Kuehne |
CVPR | 8 |
| 2024 | Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention MapsabstractWhen asked to summarize articles or answer questions given a passage, large language models (LLMs) can hallucinate details and respond with unsubstantiated answers that are inaccurate with respect to the input context.This paper describes a simple approach for detecting such contextual hallucinations.We hypothesize that contextual hallucinations are related to the extent to which an LLM attends to information in the provided context versus its own generations.Based on this intuition, we propose a simple hallucination detection model whose input features are given by the ratio of attention weights on the context versus newly generated tokens (for each attention head).We find that a linear classifier based on these lookback ratio features is as effective as a richer detector that utilizes the entire hidden states of an LLM or a text-based entailment model.The lookback ratio-based detector-Lookback Lens-is found to transfer across tasks and even models, allowing a detector that is trained on a 7B model to be applied (without retraining) to a larger 13B model.We further apply this detector to mitigate contextual hallucinations, and find that a simple classifier-guided decoding approach is able to reduce the amount of hallucination, for example by 9.6% in the XSum summarization task. 1 Yung-Sung Chuang, Linlu Qiu, Cheng-Yu Hsieh, Ranjay Krishna, James R. Glass |
EMNLP | 6 |
| 2024 | Adaptive Query Rewriting: Aligning Rewriters through Marginal Probability of Conversational AnswersabstractQuery rewriting is a crucial technique for passage retrieval in open-domain conversational question answering (CQA).It decontexualizes conversational queries into self-contained questions suitable for off-the-shelf retrievers.Existing methods attempt to incorporate retriever's preference during the training of rewriting models.However, these approaches typically rely on extensive annotations such as in-domain rewrites and/or relevant passage labels, limiting the models' generalization and adaptation capabilities.In this paper, we introduce AdaQR (Adaptive Query Rewriting), a framework for training query rewriting models with limited rewrite annotations from seed datasets and completely no passage label.Our approach begins by fine-tuning compact large language models using only ~10% of rewrite annotations from the seed dataset training split.The models are then utilized to self-sample rewrite candidates for each query instance, further eliminating the expense for human labeling or larger language model prompting often adopted in curating preference data.A novel approach is then proposed to assess retriever's preference for these candidates with the probability of answers conditioned on the conversational query by marginalizing the Top-K passages.This serves as the reward for optimizing the rewriter further using Direct Preference Optimization (DPO), a process free of rewrite and retrieval annotations.Experimental results on four open-domain CQA datasets demonstrate that AdaQR not only enhances the in-domain capabilities of the rewriter with limited annotation requirement, but also adapts effectively to out-of-domain datasets. Tianhua Zhang, Kun Li 0003, Hongyin Luo, Xixin Wu, James R. Glass, Helen M. Meng |
EMNLP | 5 |
| 2024 | Revisiting Self-supervised Learning of Speech Representation from a Mutual Information PerspectiveabstractExisting studies on self-supervised speech representation learning have focused on developing new training methods and applying pre-trained models for different applications. However, the quality of these models is often measured by the performance of different downstream tasks. How well the representations access the information of interest is less studied. In this work, we take a closer look into existing self-supervised methods of speech from an information-theoretic perspective. We aim to develop metrics using mutual information to help practical problems such as model design and selection. We use linear probes to estimate the mutual information between the target information and learned representations, showing another insight into the accessibility to the target information from speech representations. Further, we explore the potential of evaluating representations in a self-supervised fashion, where we estimate the mutual information between different parts of the data without using any labels. Finally, we show that both supervised and unsupervised measures echo the performance of the models on layer-wise linear probing and speech recognition. Alexander H. Liu, Sung-Lin Yeh, James R. Glass |
ICASSP | 3 |
| 2024 | Listen, Think, and UnderstandabstractThe ability of artificial intelligence (AI) systems to perceive and comprehend audio signals is crucial for many applications. Although significant progress has been made in this area since the development of AudioSet, most existing models are designed to map audio inputs to pre-defined, discrete sound label sets. In contrast, humans possess the ability to not only classify sounds into general categories, but also to listen to the finer details of the sounds, explain the reason for the predictions, think about what the sound infers, and understand the scene and what action needs to be taken, if any. Such capabilities beyond perception are not yet present in existing audio models. On the other hand, modern large language models (LLMs) exhibit emerging reasoning ability but they lack audio perception capabilities. Therefore, we ask the question: can we build a model that has both audio perception and reasoning ability?
In this paper, we propose a new audio foundation model, called LTU (Listen, Think, and Understand). To train LTU, we created a new OpenAQA-5M dataset consisting of 1.9 million closed-ended and 3.7 million open-ended, diverse (audio, question, answer) tuples, and have used an autoregressive training framework with a perception-to-understanding curriculum. LTU demonstrates strong performance and generalization ability on conventional audio tasks such as classification and captioning. More importantly, it exhibits emerging audio reasoning and comprehension abilities that are absent in existing audio models. To the best of our knowledge, LTU is the first multimodal large language model that focuses on general audio (rather than just speech) understanding. Yuan Gong 0001, Hongyin Luo, Alexander H. Liu, Leonid Karlinsky, James R. Glass |
ICLR | 5 |
| 2024 | DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language ModelsabstractDespite their impressive capabilities, large language models (LLMs) are prone to hallucinations, i.e., generating content that deviates from facts seen during pretraining. We propose a simple decoding strategy for reducing hallucinations with pretrained LLMs that does not require conditioning on retrieved external knowledge nor additional fine-tuning. Our approach obtains the next-token distribution by contrasting the differences in logits obtained from projecting the later layers versus earlier layers to the vocabulary space, exploiting the fact that factual knowledge in an LLMs has generally been shown to be localized to particular transformer layers. We find that this **D**ecoding by C**o**ntrasting **La**yers (DoLa) approach is able to better surface factual knowledge and reduce the generation of incorrect facts. DoLa consistently improves the truthfulness across multiple choices tasks and open-ended generation tasks, for example improving the performance of LLaMA family models on TruthfulQA by 12-17% absolute points, demonstrating its potential in making LLMs reliably generate truthful facts. Yung-Sung Chuang, Yujia Xie, Hongyin Luo, James R. Glass |
ICLR | 5 |
| 2024 | Curiosity-driven Red-teaming for Large Language ModelsabstractLarge language models (LLMs) hold great potential for many natural language applications but risk generating incorrect or toxic content. To probe when an LLM generates unwanted content, the current paradigm is to recruit a $\textit{red team}$ of human testers to design input prompts (i.e., test cases) that elicit undesirable responses from LLMs.
However, relying solely on human testers is expensive and time-consuming. Recent works automate red teaming by training a separate red team LLM with reinforcement learning (RL) to generate test cases that maximize the chance of eliciting undesirable responses from the target LLM. However, current RL methods are only able to generate a small number of effective test cases resulting in a low coverage of the span of prompts that elicit undesirable responses from the target LLM.
To overcome this limitation, we draw a connection between the problem of increasing the coverage of generated test cases and the well-studied approach of curiosity-driven exploration that optimizes for novelty.
Our method of curiosity-driven red teaming (CRT) achieves greater coverage of test cases while mantaining or increasing their effectiveness compared to existing methods.
Our method, CRT successfully provokes toxic responses from LLaMA2 model that has been heavily fine-tuned using human preferences to avoid toxic outputs. Code is available at https://github.com/Improbable-AI/curiosity_redteam. Zhang-Wei Hong, Idan Shenfeld, Tsun-Hsuan Wang, Yung-Sung Chuang, Aldo Pareja, James R. Glass, Akash Srivastava, Pulkit Agrawal 0001 |
ICLR | 6 |
| 2024 | Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
Andrew Rouditchenko, Yuan Gong 0001, Samuel Thomas 0001, Leonid Karlinsky, Hilde Kuehne, Rogério Feris, James R. Glass |
INTERSPEECH | 7 |
| 2024 | Automatic Prediction of Amyotrophic Lateral Sclerosis Progression using Longitudinal Speech Transformer
Liming Wang 0003, Yuan Gong 0001, Nauman Dawalatabad, Marco Vilela, Katerina Placek, Brian Tracey, Yishu Gong, Alan Premasiri, Fernando Vieira, James R. Glass |
INTERSPEECH | 10 |
| 2024 | R-Spin: Efficient Speaker and Noise-invariant Representation Learning with Acoustic PiecesabstractHeng-Jui Chang, James Glass. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Heng-Jui Chang, James R. Glass |
NAACL-HLT | 2 |
| 2024 | DASS: Distilled Audio State Space Models are Stronger and More Duration-Scalable LearnersabstractState-space models (SSMs) have emerged as an alternative to Transformers for audio modeling due to their high computational efficiency with long inputs. While recent efforts on Audio SSMs have reported encouraging results, two main limitations remain: First, in 10 -second short audio tagging tasks, Audio SSMs still underperform compared to Transformer-based models such as Audio Spectrogram Transformer (AST). Second, although Audio SSMs theoretically support long audio inputs, their actual performance with long audio has not been thoroughly evaluated. To address these limitations, in this paper, 1) We applied knowledge distillation in audio space model training, resulting in a model called Knowledge Distilled Audio SSM (DASS). To the best of our knowledge, it is the first SSM that outperforms the Transformers on AudioSet and achieves an mAP of 47.6; and 2) We designed a new test called Audio Needle In A Haystack (Audio NIAH). We find that DASS, trained with only 10 -second audio clips, can retrieve sound events in audio recordings up to 2.5 hours long, while the AST model fails when the input is just 50 seconds, demonstrating state-space models are indeed more duration scalable. Saurabhchand Bhati, Yuan Gong 0001, Leonid Karlinsky, Hilde Kuehne, Rogério Feris, James R. Glass |
SLT | 6 |
| 2024 | Codec-Superb @ SLT 2024: A Lightweight Benchmark For Neural Audio Codec ModelsabstractNeural audio codec models are becoming increasingly important as they serve as tokenizers for audio, enabling efficient transmission or facilitating speech language modeling. The ideal neural audio codec should maintain content, paralinguistics, speaker characteristics, and audio information even at low bitrates. Recently, numerous advanced neural codec models have been proposed. However, codec models are often tested under varying experimental conditions. As a result, we introduce the Codec-SUPERB challenge at SLT 20241, designed to facilitate fair and lightweight comparisons among existing codec models and inspire advancements in the field. This challenge brings together representative speech applications and objective metrics, and carefully selects license-free datasets, sampling them into small sets to reduce evaluation computation costs. This paper presents the challenge’s rules, datasets, participant systems, results, and findings.1https://codecsuperb.github.io/ Xuanjun Chen, Yi-Cheng Lin, Kai-Wei Chang 0001, Jiawei Du 0003, Ke-Han Lu, Alexander H. Liu, Ho-Lam Chung, Yuan-Kuei Wu, Dongchao Yang, Songxiang Liu, Yi-Chiao Wu, Xu Tan 0003, James R. Glass, Shinji Watanabe 0001, Hung-yi Lee |
SLT | 14 |
| 2023 | Entailment as Robust Self-LearnerabstractEntailment has been recognized as an important metric for evaluating natural language understanding (NLU) models, and recent studies have found that entailment pretraining benefits weakly supervised fine-tuning.In this work, we design a prompting strategy that formulates a number of different NLU tasks as contextual entailment.This approach improves the zero-shot adaptation of pretrained entailment models.Secondly, we notice that self-training entailment-based models with unlabeled data can significantly improve the adaptation performance on downstream tasks.To achieve more stable improvement, we propose the Simple Pseudo-Label Editing (SimPLE) algorithm for better pseudo-labeling quality in self-training.We also found that both pretrained entailmentbased models and the self-trained models are robust against adversarial evaluation data.Experiments on binary and multi-class classification tasks show that SimPLE leads to more robust self-training results, indicating that the self-trained entailment models are more efficient and trustworthy than large language models on language understanding tasks. Jiaxin Ge, Hongyin Luo, James R. Glass |
ACL (1) | 4 |
| 2023 | On the Blind Spots of Model-Based Evaluation Metrics for Text GenerationabstractTianxing He, Jingyu Zhang, Tianle Wang, Sachin Kumar, Kyunghyun Cho, James Glass, Yulia Tsvetkov. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Tianxing He, Tianle Wang 0003, Sachin Kumar 0009, Kyunghyun Cho, James R. Glass, Yulia Tsvetkov |
ACL (1) | 6 |
| 2023 | Joint Audio and Speech UnderstandingabstractHumans are surrounded by audio signals that include both speech and non-speech sounds. The recognition and understanding of speech and non-speech audio events, along with a profound comprehension of the relationship between them, constitute fundamental cognitive capabilities. For the first time, we build a machine learning model, called LTU-AS, that has a conceptually similar universal audio perception and advanced reasoning ability. Specifically, by integrating Whisper [1] as a perception module and LLaMA [2] as a reasoning module, LTU-AS can simultaneously recognize and jointly understand spoken text, speech paralinguistics, and non-speech audio events - almost everything perceivable from audio signals. Yuan Gong 0001, Alexander H. Liu, Hongyin Luo, Leonid Karlinsky, James R. Glass |
ASRU | 5 |
| 2023 | Audio-Visual Neural Syntax AcquisitionabstractWe study phrase structure induction from visually-grounded speech. The core idea is to first segment the speech waveform into sequences of word segments, and subsequently induce phrase structure using the inferred segment-level continuous representations. We present the Audio-Visual Neural Syntax Learner (AV-NSL) that learns phrase structure by listening to audio and looking at images, without ever being exposed to text. By training on paired images and spoken captions, AV-NSL exhibits the capability to infer meaningful phrase structures that are comparable to those derived by naturally-supervised text parsers, for both English and German. Our findings extend prior work in unsupervised language acquisition from speech and grounded grammar induction, and present one approach to bridge the gap between the two topics. Cheng-I Lai, Freda Shi, Puyuan Peng, Kevin Gimpel, Shiyu Chang, Yung-Sung Chuang, Saurabhchand Bhati, David D. Cox, David F. Harwath, Yang Zhang 0001, Karen Livescu, James R. Glass |
ASRU | 13 |
| 2023 | Logic Against Bias: Textual Entailment Mitigates Stereotypical Sentence ReasoningabstractDue to their similarity-based learning objectives, pretrained sentence encoders often internalize stereotypical assumptions that reflect the social biases that exist within their training corpora.In this paper, we describe several kinds of stereotypes concerning different communities that are present in popular sentence representation models, including pretrained next sentence prediction and contrastive sentence representation models.We compare such models to textual entailment models that learn language logic for a variety of downstream language understanding tasks.By comparing strong pretrained models based on text similarity with textual entailment learning, we conclude that the explicit logic learning with textual entailment can significantly reduce bias and improve the recognition of social communities, without an explicit de-biasing process.The code, model, and data associated with this work are publicly available at https: //github.com/luohongyin/ESP.git. Hongyin Luo, James R. Glass |
EACL | 2 |
| 2023 | On Unsupervised Uncertainty-Driven Speech Pseudo-Label Filtering and Model CalibrationabstractPseudo-label (PL) filtering forms a crucial part of Self-Training (ST) methods for unsupervised domain adaptation. Dropout-based Uncertainty-driven Self-Training (DUST) proceeds by first training a teacher model on source domain labeled data. Then, the teacher model is used to provide PLs for the unlabeled target domain data. Finally, we train a student on augmented labeled and pseudo-labeled data. The process is iterative, where the student becomes the teacher for the next DUST iteration. A crucial step that precedes the student model training in each DUST iteration is filtering out noisy PLs that could lead the student model astray. In DUST, we proposed a simple, effective, and theoretically sound PL filtering strategy based on the teacher model’s uncertainty about its predictions on unlabeled speech utterances. We estimate the model’s uncertainty by computing disagreement amongst multiple samples drawn from the teacher model during inference by injecting noise via dropout. In this work, we show that DUST’s PL filtering, as initially used, fail under severe source and target domain mismatch. We suggest several approaches to eliminate or alleviate this issue. Further, we bring insights from the research in neural network model calibration to DUST and show that a well-calibrated model correlates strongly with a positive outcome of the DUST PL filtering step. We demonstrate effectiveness of our methods on target domain GigaSpeech YouTube dataset. Nauman Dawalatabad, Sameer Khurana, Antoine Laurent, James R. Glass |
ICASSP | 4 |
| 2023 | C2KD: Cross-Lingual Cross-Modal Knowledge Distillation for Multilingual Text-Video RetrievalabstractMultilingual text-video retrieval methods have improved significantly in recent years, but the performance for languages other than English still lags. We propose a Cross-Lingual Cross-Modal Knowledge Distillation method to improve multilingual text-video retrieval. Inspired by the fact that English text-video retrieval outperforms other languages, we train a student model using input text in different languages to match the cross-modal predictions from teacher models using input text in English. We propose a cross entropy based objective which forces the distribution over the student’s text-video similarity scores to be similar to those of the teacher models. We introduce a new multilingual video dataset, Multi-YouCook2, by translating the English captions in the YouCook2 video dataset to 8 other languages. Our method improves multilingual text-video retrieval performance on Multi-YouCook2 and several other datasets such as Multi-MSRVTT and VATEX. We also conducted an analysis on the effectiveness of different multilingual text models as teachers. Andrew Rouditchenko, Yung-Sung Chuang, Nina Shvetsova, Samuel Thomas 0001, Rogério Feris, Brian Kingsbury, Leonid Karlinsky, David F. Harwath, Hilde Kuehne, James R. Glass |
ICASSP | 10 |
| 2023 | Contrastive Audio-Visual Masked Autoencoder
Yuan Gong 0001, Andrew Rouditchenko, Alexander H. Liu, David F. Harwath, Leonid Karlinsky, Hilde Kuehne, James R. Glass |
ICLR | 7 |
| 2023 | Whisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong General Audio Event TaggersabstractIn this paper, we focus on Whisper [1], a recent automatic speech recognition model trained with a massive 680k hour labeled speech corpus recorded in diverse conditions.We first show an interesting finding that while Whisper is very robust against real-world background sounds (e.g., music), its audio representation is actually not noise-invariant, but is instead highly correlated to non-speech sounds, indicating that Whisper recognizes speech conditioned on the noise type.With this finding, we build a unified audio tagging and speech recognition model Whisper-AT by freezing the backbone of Whisper, and training a lightweight audio tagging model on top of it.With <1% extra computational cost, Whisper-AT can recognize audio events, in addition to spoken text, in a single forward pass. Yuan Gong 0001, Sameer Khurana, Leonid Karlinsky, James R. Glass |
INTERSPEECH | 4 |
| 2023 | Self-supervised Fine-tuning for Improved Content Representations by Speaker-invariant Clustering
Heng-Jui Chang, Alexander H. Liu, James R. Glass |
INTERSPEECH | 3 |
| 2023 | Comparison of Multilingual Self-Supervised and Weakly-Supervised Speech Pre-Training for Adaptation to Unseen Languages
Andrew Rouditchenko, Sameer Khurana, Samuel Thomas 0001, Rogério Feris, Leonid Karlinsky, Hilde Kuehne, David F. Harwath, Brian Kingsbury, James R. Glass |
INTERSPEECH | 9 |
| 2023 | DinoSR: Self-Distillation and Online Clustering for Self-supervised Speech Representation LearningabstractIn this paper, we introduce self-distillation and online clustering for self-supervised speech representation learning (DinoSR) which combines masked language modeling, self-distillation, and online clustering. We show that these concepts complement each other and result in a strong representation learning model for speech. DinoSR first extracts contextualized embeddings from the input audio with a teacher network, then runs an online clustering system on the embeddings to yield a machine-discovered phone inventory, and finally uses the discretized tokens to guide a student network. We show that DinoSR surpasses previous state-of-the-art performance in several downstream tasks, and provide a detailed analysis of the model and the learned discrete units. Alexander H. Liu, Heng-Jui Chang, Michael Auli, Wei-Ning Hsu, James R. Glass |
NeurIPS | 5 |
| 2022 | SSAST: Self-Supervised Audio Spectrogram TransformerabstractRecently, neural networks based purely on self-attention, such as the Vision Transformer (ViT), have been shown to outperform deep learning models constructed with convolutional neural networks (CNNs) on various vision tasks, thus extending the success of Transformers, which were originally developed for language processing, to the vision domain. A recent study showed that a similar methodology can also be applied to the audio domain. Specifically, the Audio Spectrogram Transformer (AST) achieves state-of-the-art results on various audio classification benchmarks. However, pure Transformer models tend to require more training data compared to CNNs, and the success of the AST relies on supervised pretraining that requires a large amount of labeled data and a complex training pipeline, thus limiting the practical usage of AST. This paper focuses on audio and speech classification, and aims to reduce the need for large amounts of labeled data for the AST by leveraging self-supervised learning using unlabeled data. Specifically, we propose to pretrain the AST model with joint discriminative and generative masked spectrogram patch modeling (MSPM) using unlabeled audio from AudioSet and Librispeech. We evaluate our pretrained models on both audio and speech classification tasks including audio event classification, keyword spotting, emotion recognition, and speaker identification. The proposed self-supervised framework significantly boosts AST performance on all tasks, with an average improvement of 60.9%, leading to similar or even better results than a supervised pretrained AST. To the best of our knowledge, it is the first patch-based self-supervised learning framework in the audio and speech domain, and also the first self-supervised learning framework for AST. Yuan Gong 0001, Cheng-I Lai, Yu-An Chung, James R. Glass |
AAAI | 4 |
| 2022 | Cross-Modal Discrete Representation LearningabstractAlexander Liu, SouYoung Jin, Cheng-I Lai, Andrew Rouditchenko, Aude Oliva, James Glass. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Alexander H. Liu, SouYoung Jin, Cheng-I Lai, Andrew Rouditchenko, Aude Oliva, James R. Glass |
ACL (1) | 6 |
| 2022 | Everything at Once - Multi-modal Fusion Transformer for Video RetrievalabstractMulti-modal learning from video data has seen increased attention recently as it allows training of semantically meaningful embeddings without human annotation, enabling tasks like zero-shot retrieval and action localization. In this work, we present a multi-modal, modality agnostic fusion transformer that learns to exchange information between multiple modalities, such as video, audio, and text, and integrate them into a fused representation in a joined multi-modal embedding space. We propose to train the system with a combinatorial loss on everything at once – any combination of input modalities, such as single modalities as well as pairs of modalities, explicitly leaving out any add-ons such as position or modality encoding. At test time, the resulting model can process and fuse any number of input modalities. Moreover, the implicit properties of the transformer allow to process inputs of different lengths. To evaluate the proposed approach, we train the model on the large scale HowTo100M dataset and evaluate the resulting embedding space on four challenging benchmark datasets obtaining state-of-the-art results in zero-shot video retrieval and zero-shot video action localization. Our code for this work is also available.11https://github.com/ninatu/everything_at_once Nina Shvetsova, Brian Chen 0001, Andrew Rouditchenko, Samuel Thomas 0001, Brian Kingsbury, Rogério Feris, David F. Harwath, James R. Glass, Hilde Kuehne |
CVPR | 8 |
| 2022 | Transformer-Based Multi-Aspect Multi-Granularity Non-Native English Speaker Pronunciation AssessmentabstractAutomatic pronunciation assessment is an important technology to help self-directed language learners. While pronunciation quality has multiple aspects including accuracy, fluency, completeness, and prosody, previous efforts typically only model one aspect (e.g., accuracy) at one granularity (e.g., at the phoneme-level). In this work, we explore modeling multi-aspect pronunciation assessment at multiple granularities. Specifically, we train a Goodness Of Pronunciation feature-based Transformer (GOPT) with multi-task learning. Experiments show that GOPT achieves the best results on speechocean762 with a public automatic speech recognition (ASR) acoustic model trained on Librispeech. Yuan Gong 0001, Ziyi Chen 0005, Iek-Heng Chu, Peng Chang 0002, James R. Glass |
ICASSP | 5 |
| 2022 | Vocalsound: A Dataset for Improving Human Vocal Sounds RecognitionabstractRecognizing human non-speech vocalizations is an important task and has broad applications such as automatic sound transcription and health condition monitoring. However, existing datasets have a relatively small number of vocal sound samples or noisy labels. As a consequence, state-of-the-art audio event classification models may not perform well in detecting human vocal sounds. To support research on building robust and accurate vocal sound recognition, we have created a VocalSound dataset consisting of over 21,000 crowdsourced recordings of laughter, sighs, coughs, throat clearing, sneezes, and sniffs from 3,365 unique subjects. Experiments show that the vocal sound recognition performance of a model can be significantly improved by 41.9% by adding VocalSound dataset to an existing dataset as training material. In addition, different from previous datasets, the VocalSound dataset contains meta information such as speaker age, gender, native language, country, and health condition. Yuan Gong 0001, James R. Glass |
ICASSP | 3 |
| 2022 | Repetition Assessment for Speech and Language Disorders: A Study of the Logopenic Variant of Primary Progressive AphasiaabstractImpaired repetition is a characteristic of several speech and language disorders, including certain variants of Primary Progressive Aphasia (PPA). People with the logopenic variant of PPA (lvPPA) can present with impaired repetition abilities and repetition tasks can be used to distinguish lvPPA speakers from healthy controls. In this paper, we propose a novel technique for quantifying the quality of repetition in speech recordings and demonstrate the utility of the technique by using it to distinguish between healthy speakers and lvPPA speakers. We train several classifiers on features extracted from the repetition recordings. The best classifier distinguishes the lvPPA speakers with impaired repetition from the healthy speakers with 85.7% accuracy and classifies all healthy speakers with perfect accuracy. Although we evaluate the method on lvPPA detection, we believe that the method has potential utility for a range of tasks and speech disorders where repetition occurs. R'mani Haulcy, Katerina Placek, Brian Tracey, Adam P. Vogel, James R. Glass |
ICASSP | 5 |
| 2022 | Magic Dust for Cross-Lingual Adaptation of Monolingual Wav2vec-2.0abstractWe propose a simple and effective cross-lingual transfer learning method to adapt monolingual wav2vec-2.0 models for Automatic Speech Recognition (ASR) in resource-scarce languages. We show that a monolingual wav2vec-2.0 is a good few-shot ASR learner in several languages. We improve its performance further via several iterations of Dropout Uncertainty-Driven Self-Training (DUST) by using a moderate-sized unlabeled speech dataset in the target language. A key finding of this work is that the adapted monolingual wav2vec-2.0 achieves similar performance as the topline multilingual XLSR model, which is trained on fifty-three languages, on the target language ASR task. Sameer Khurana, Antoine Laurent, James R. Glass |
ICASSP | 3 |
| 2022 | On the Interplay between Sparsity, Naturalness, Intelligibility, and Prosody in Speech SynthesisabstractAre end-to-end text-to-speech (TTS) models over-parametrized? To what extent can these models be pruned, and what happens to their synthesis capabilities? This work serves as a starting point to explore pruning both spectrogram prediction networks and vocoders. We thoroughly investigate the tradeoffs between sparsity and its subsequent effects on synthetic speech. Additionally, we explore several aspects of TTS pruning: amount of finetuning data versus sparsity, TTS-Augmentation to utilize unspoken text, and combining knowledge distillation and pruning. Our findings suggest that not only are end-to-end TTS models highly prunable, but also, perhaps surprisingly, pruned TTS models can produce synthetic speech with equal or higher naturalness and intelligibility, with similar prosody. All of our experiments are conducted on publicly available models, and findings in this work are backed by large-scale subjective tests and objective measures. Code and 200 pruned models are made available to facilitate future research on efficiency in TTS1. Cheng-I Lai, Erica Cooper, Yang Zhang 0001, Shiyu Chang, Kaizhi Qian, Yi-Lun Liao, Yung-Sung Chuang, Alexander H. Liu, Junichi Yamagishi, David D. Cox, James R. Glass |
ICASSP | 11 |
| 2022 | Simple and Effective Unsupervised Speech SynthesisabstractWe introduce the first unsupervised speech synthesis system based on a simple, yet effective recipe.The framework leverages recent work in unsupervised speech recognition as well as existing neural-based speech synthesis.Using only unlabeled speech audio and unlabeled text as well as a lexicon, our method enables speech synthesis without the need for a human-labeled corpus.Experiments demonstrate the unsupervised system can synthesize speech similar to a supervised counterpart in terms of naturalness and intelligibility measured by human evaluation. Alexander H. Liu, Cheng-I Lai, Wei-Ning Hsu, Michael Auli, Alexei Baevski, James R. Glass |
INTERSPEECH | 6 |
| 2022 | Speak: A Toolkit Using Amazon Mechanical Turk to Collect and Validate Speech Audio RecordingsabstractWe present Speak, a toolkit that allows researchers to crowdsource speech audio recordings using Amazon Mechanical Turk (MTurk). Speak allows MTurk workers to submit speech recordings in response to a task prompt and stimulus (e.g. image, text excerpt, audio file) defined by researchers, a functionality that is not natively offered by MTurk at the time of writing this paper. Importantly, the toolkit employs numerous measures to ensure that speech recordings collected are of adequate quality, in order to avoid accepting unusable data and prevent abuse/fraud. Speak has demonstrated utility, having collected over 600,000 recordings to date. The toolkit is open-source and available for download. Christopher Song, David F. Harwath, Tuka Al Hanai, James R. Glass |
LREC | 4 |
| 2022 | DiffCSE: Difference-based Contrastive Learning for Sentence EmbeddingsabstractYung-Sung Chuang, Rumen Dangovski, Hongyin Luo, Yang Zhang, Shiyu Chang, Marin Soljacic, Shang-Wen Li, Scott Yih, Yoon Kim, James Glass. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yung-Sung Chuang, Rumen Dangovski, Hongyin Luo, Yang Zhang 0001, Shiyu Chang, Marin Soljacic, Shang-Wen Li 0001, Scott Yih, James R. Glass |
NAACL-HLT | 10 |
| 2022 | Cooperative Self-training of Machine Reading ComprehensionabstractHongyin Luo, Shang-Wen Li, Mingye Gao, Seunghak Yu, James Glass. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Hongyin Luo, Shang-Wen Li 0001, Mingye Gao, Seunghak Yu, James R. Glass |
NAACL-HLT | 5 |
| 2022 | UAVM: Towards Unifying Audio and Visual ModelsabstractConventional audio-visual models have independent audio and video branches. In this work, weunifythe audio and visual branches by designing aUnifiedAudio-VisualModel (UAVM). The UAVM achieves a new state-of-the-art audio-visual event classification accuracy of 65.8% on VGGSound. More interestingly, we also find a few intriguing properties of UAVM that the modality-independent counterparts do not have. Yuan Gong 0001, Alexander H. Liu, Andrew Rouditchenko, James R. Glass |
IEEE Signal Process. Lett. | 4 |
| 2021 | Text-Free Image-to-Speech Synthesis Using Learned Segmental UnitsabstractWei-Ning Hsu, David Harwath, Tyler Miller, Christopher Song, James Glass. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Wei-Ning Hsu, David F. Harwath, Tyler Miller, Christopher Song, James R. Glass |
ACL/IJCNLP (1) | 5 |
| 2021 | Spoken Moments: Learning Joint Audio-Visual Representations From Video DescriptionsabstractWhen people observe events, they are able to abstract key information and build concise summaries of what is happening. These summaries include contextual and semantic information describing the important high-level details (what, where, who and how) of the observed event and exclude background information that is deemed unimportant to the observer. With this in mind, the descriptions people generate for videos of different dynamic events can greatly improve our understanding of the key information of interest in each video. These descriptions can be captured in captions that provide expanded attributes for video labeling (e.g. actions/objects/scenes/sentiment/etc.) while allowing us to gain new insight into what people find important or necessary to summarize specific events. Existing caption datasets for video understanding are either small in scale or restricted to a specific domain. To address this, we present the Spoken Moments (S-MiT) dataset of 500k spoken captions each attributed to a unique short video depicting a broad range of different events. We collect our descriptions using audio recordings to ensure that they remain as natural and concise as possible while allowing us to scale the size of a large classification dataset. In order to utilize our proposed dataset, we present a novel Adaptive Mean Margin (AMM) approach to contrastive learning and evaluate our models on video/caption retrieval on multiple datasets. We show that our AMM approach consistently improves our results and that models trained on our Spoken Moments dataset generalize better than those trained on other video-caption datasets.http://moments.csail.mit.edu/spoken.html Mathew Monfort, SouYoung Jin, Alexander H. Liu, David F. Harwath, Rogério Feris, James R. Glass, Aude Oliva |
CVPR | 6 |
| 2021 | Analyzing the Forgetting Problem in Pretrain-Finetuning of Open-domain Dialogue Response ModelsabstractTianxing He, Jun Liu, Kyunghyun Cho, Myle Ott, Bing Liu, James Glass, Fuchun Peng. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Tianxing He, Kyunghyun Cho, Myle Ott, Bing Liu 0024, James R. Glass, Fuchun Peng |
EACL | 6 |
| 2021 | Exposure Bias versus Self-Recovery: Are Distortions Really Incremental for Autoregressive Text Generation?abstractExposure bias has been regarded as a central problem for auto-regressive language models (LM).It claims that teacher forcing would cause the test-time generation to be incrementally distorted due to the training-generation discrepancy.Although a lot of algorithms have been proposed to avoid teacher forcing and therefore alleviate exposure bias, there is little work showing how serious the exposure bias problem actually is.In this work, we focus on the task of open-ended language generation, propose metrics to quantify the impact of exposure bias in the aspects of quality, diversity, and consistency.Our key intuition is that if we feed ground-truth data prefixes (instead of prefixes generated by the model itself) into the model and ask it to continue the generation, the performance should become much better because the training-generation discrepancy in the prefix is removed.Both automatic and human evaluations are conducted in our experiments.On the contrary to the popular belief in exposure bias, we find that the the distortion induced by the prefix discrepancy is limited, and does not seem to be incremental during the generation.Moreover, our analysis reveals an interesting self-recovery ability of the LM, which we hypothesize to be countering the harmful effects from exposure bias. Tianxing He, Jingzhao Zhang, Zhiming Zhou 0001, James R. Glass |
EMNLP (1) | 4 |
| 2021 | Similarity Analysis of Self-Supervised Speech RepresentationsabstractSelf-supervised speech representation learning has recently been a prosperous research topic. Many algorithms have been proposed for learning useful representations from large-scale unlabeled data, and their applications to a wide range of speech tasks have also been investigated. However, there has been little research focusing on understanding the properties of existing approaches. In this work, we aim to provide a comparative study of some of the most representative self-supervised algorithms. Specifically, we quantify the similarities between different self-supervised representations using existing similarity measures. We also design probing tasks to study the correlation between the models’ pre-training loss and the amount of specific speech information contained in their learned representations. In addition to showing how various self-supervised models behave differently given the same input, our study also finds that the training objective has a higher impact on representation similarity than architectural choices such as building blocks (RNN/Transformer/CNN) and directionality (uni/bidirectional). Our results also suggest that there exists a strong correlation between pre-training loss and downstream performance for some self-supervised algorithms. Yu-An Chung, Yonatan Belinkov, James R. Glass |
ICASSP | 3 |
| 2021 | Semi-Supervised Spoken Language Understanding via Self-Supervised Speech and Language Model PretrainingabstractMuch recent work on Spoken Language Understanding (SLU) is limited in at least one of three ways: models were trained on oracle text input and neglected ASR errors, models were trained to predict only intents without the slot values, or models were trained on a large amount of in-house data. In this paper, we propose a clean and general framework to learn semantics directly from speech with semi-supervision from transcribed or untranscribed speech to address these issues. Our framework is built upon pretrained end-to-end (E2E) ASR and self-supervised language models, such as BERT, and fine-tuned on a limited amount of target SLU data. We study two semi-supervised settings for the ASR component: supervised pretraining on transcribed speech, and unsupervised pretraining by replacing the ASR encoder with self-supervised speech representations, such as wav2vec. In parallel, we identify two essential criteria for evaluating SLU models: environmental noise-robustness and E2E semantics evaluation. Experiments on ATIS show that our SLU framework with speech as input can perform on par with those using oracle text as input in semantics understanding, even though environmental noise is present and a limited amount of labeled semantics data is available for training. Cheng-I Lai, Yung-Sung Chuang, Hung-yi Lee, Shang-Wen Li 0001, James R. Glass |
ICASSP | 5 |
| 2021 | Multimodal Clustering Networks for Self-supervised Learning from Unlabeled VideosabstractMultimodal self-supervised learning is getting more and more attention as it allows not only to train large networks without human supervision but also to search and retrieve data across various modalities. In this context, this paper proposes a framework that, starting from a pre-trained backbone, learns a common multimodal embedding space that, in addition to sharing representations across different modalities, enforces a grouping of semantically similar instances. To this end, we extend the concept of instance-level contrastive learning with a multimodal clustering step in the training pipeline to capture semantic similarities across modalities. The resulting embedding space enables retrieval of samples across all modalities, even from unseen datasets and different domains. To evaluate our approach, we train our model on the HowTo100M dataset and evaluate its zero-shot retrieval capabilities in two challenging domains, namely text-to-video retrieval, and temporal action localization, showing state-of-the-art results on four different datasets. Brian Chen 0001, Andrew Rouditchenko, Kevin Duarte, Hilde Kuehne, Samuel Thomas 0001, Angie W. Boggust, Rameswar Panda, Brian Kingsbury, Rogério Feris, David F. Harwath, James R. Glass, Michael Picheny, Shih-Fu Chang |
ICCV | 11 |
| 2021 | AST: Audio Spectrogram TransformerabstractIn the past decade, convolutional neural networks (CNNs) have been widely adopted as the main building block for endto-end audio classification models, which aim to learn a direct mapping from audio spectrograms to corresponding labels.To better capture long-range global context, a recent trend is to add a self-attention mechanism on top of the CNN, forming a CNN-attention hybrid model.However, it is unclear whether the reliance on a CNN is necessary, and if neural networks purely based on attention are sufficient to obtain good performance in audio classification.In this paper, we answer the question by introducing the Audio Spectrogram Transformer (AST), the first convolution-free, purely attention-based model for audio classification.We evaluate AST on various audio classification benchmarks, where it achieves new state-of-the-art results of 0.485 mAP on AudioSet, 95.6% accuracy on ESC-50, and 98.1% accuracy on Speech Commands V2. Yuan Gong 0001, Yu-An Chung, James R. Glass |
Interspeech | 3 |
| 2021 | CLAC: A Speech Corpus of Healthy English Speakers
R'mani Haulcy, James R. Glass |
Interspeech | 2 |
| 2021 | Non-Autoregressive Predictive Coding for Learning Speech Representations from Local DependenciesabstractSelf-supervised speech representations have been shown to be effective in a variety of speech applications. However, existing representation learning methods generally rely on the autoregressive model and/or observed global dependencies while generating the representation. In this work, we propose Non-Autoregressive Predictive Coding (NPC), a self-supervised method, to learn a speech representation in a non-autoregressive manner by relying only on local dependencies of speech. NPC has a conceptually simple objective and can be implemented easily with the introduced Masked Convolution Blocks. NPC offers a significant speedup for inference since it is parallelizable in time and has a fixed inference time for each time step regardless of the input sequence length. We discuss and verify the effectiveness of NPC by theoretically and empirically comparing it with other methods. We show that the NPC representation is comparable to other methods in speech experiments on phonetic and speaker classification while being more efficient. Alexander H. Liu, Yu-An Chung, James R. Glass |
Interspeech | 3 |
| 2021 | Joint Retrieval-Extraction Training for Evidence-Aware Dialog Response Selection
Hongyin Luo, James R. Glass, Garima Lalwani, Shang-Wen Li 0001 |
Interspeech | 2 |
| 2021 | Spoken ObjectNet: A Bias-Controlled Spoken Caption DatasetabstractVisually-grounded spoken language datasets can enable models to learn cross-modal correspondences with very weak supervision.However, modern audio-visual datasets contain biases that undermine the real-world performance of models trained on that data.We introduce Spoken ObjectNet, which is designed to remove some of these biases and provide a way to better evaluate how effectively models will perform in real-world scenarios.This dataset expands upon ObjectNet, which is a biascontrolled image dataset that features similar image classes to those present in ImageNet.We detail our data collection pipeline, which features several methods to improve caption quality, including automated language model checks.Lastly, we show baseline results on image retrieval and audio retrieval tasks.These results show that models trained on other datasets and then evaluated on Spoken ObjectNet tend to perform poorly due to biases in other datasets that the models have learned.We also show evidence that the performance decrease is due to the dataset controls, and not the transfer setting. Ian Palmer, Andrew Rouditchenko, Andrei Barbu, Boris Katz, James R. Glass |
Interspeech | 5 |
| 2021 | Cascaded Multilingual Audio-Visual Learning from VideosabstractIn this paper, we explore self-supervised audio-visual models that learn from instructional videos.Prior work has shown that these models can relate spoken words and sounds to visual content after training on a large-scale dataset of videos, but they were only trained and evaluated on videos in English.To learn multilingual audio-visual representations, we propose a cascaded approach that leverages a model trained on English videos and applies it to audio-visual data in other languages, such as Japanese videos.With our cascaded approach, we show an improvement in retrieval performance of nearly 10x compared to training on the Japanese videos solely.We also apply the model trained on English videos to Japanese and Hindi spoken captions of images, achieving state-of-the-art performance. Andrew Rouditchenko, Angie W. Boggust, David F. Harwath, Samuel Thomas 0001, Hilde Kuehne, Brian Chen 0001, Rameswar Panda, Rogério Feris, Brian Kingsbury, Michael Picheny, James R. Glass |
Interspeech | 11 |
| 2021 | AVLnet: Learning Audio-Visual Language Representations from Instructional VideosabstractCurrent methods for learning visually grounded language from videos often rely on text annotation, such as human generated captions or machine generated automatic speech recognition (ASR) transcripts. In this work, we introduce the Audio-Video Language Network (AVLnet), a self-supervised network that learns a shared audio-visual embedding space directly from raw video inputs. To circumvent the need for text annotation, we learn audio-visual representations from randomly segmented video clips and their raw audio waveforms. We train AVLnet on HowTo100M, a large corpus of publicly available instructional videos, and evaluate on image retrieval and video retrieval tasks, achieving state-of-the-art performance. We perform analysis of AVLnet's learned representations, showing our model utilizes speech and natural sounds to learn audio-visual concepts. Further, we propose a tri-modal model that jointly processes raw audio, video, and text captions from videos to learn a multi-modal semantic embedding space useful for text-video retrieval. Our code, data, and trained models will be released at avlnet.csail.mit.edu Andrew Rouditchenko, Angie W. Boggust, David F. Harwath, Brian Chen 0001, Dhiraj Joshi, Samuel Thomas 0001, Kartik Audhkhasi, Hilde Kuehne, Rameswar Panda, Rogério Feris, Brian Kingsbury, Michael Picheny, Antonio Torralba 0001, James R. Glass |
Interspeech | 14 |
| 2021 | PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech RecognitionabstractSelf-supervised speech representation learning (speech SSL) has demonstrated the benefit of scale in learning rich representations for Automatic Speech Recognition (ASR) with limited paired data, such as wav2vec 2.0. We investigate the existence of sparse subnetworks in pre-trained speech SSL models that achieve even better low-resource ASR results. However, directly applying widely adopted pruning methods such as the Lottery Ticket Hypothesis (LTH) is suboptimal in the computational cost needed. Moreover, we show that the discovered subnetworks yield minimal performance gain compared to the original dense network.We present Prune-Adjust-Re-Prune (PARP), which discovers and finetunes subnetworks for much better performance, while only requiring a single downstream ASR finetuning run. PARP is inspired by our surprising observation that subnetworks pruned for pre-training tasks need merely a slight adjustment to achieve a sizeable performance boost in downstream ASR tasks. Extensive experiments on low-resource ASR verify (1) sparse subnetworks exist in mono-lingual/multi-lingual pre-trained speech SSL, and (2) the computational advantage and performance gain of PARP over baseline pruning methods.In particular, on the 10min Librispeech split without LM decoding, PARP discovers subnetworks from wav2vec 2.0 with an absolute 10.9%/12.6% WER decrease compared to the full model. We further demonstrate the effectiveness of PARP via: cross-lingual pruning without any phone recognition degradation, the discovery of a multi-lingual subnetwork for 10 spoken languages in 1 finetuning run, and its applicability to pre-trained BERT/XLNet for natural language tasks1. Cheng-I Lai, Yang Zhang 0001, Alexander H. Liu, Shiyu Chang, Yi-Lun Liao, Yung-Sung Chuang, Kaizhi Qian, Sameer Khurana, David D. Cox, James R. Glass |
NeurIPS | 10 |
| 2021 | PSLA: Improving Audio Tagging With Pretraining, Sampling, Labeling, and AggregationabstractAudio tagging is an active research area and has a wide range of applications. Since the release of AudioSet, great progress has been made in advancing model performance, which mostly comes from the development of novel model architectures and attention modules. However, we find that appropriate training techniques are equally important for building audio tagging models with AudioSet, but have not received the attention they deserve. To fill the gap, in this work, we present PSLA, a collection of training techniques that can noticeably boost the model accuracy including ImageNet pretraining, balanced sampling, data augmentation, label enhancement, model aggregation and their design choices. By training an EfficientNet with these techniques, we obtain a single model (with 13.6M parameters) and an ensemble model that achieve mean average precision (mAP) scores of 0.444 and 0.474 on AudioSet, respectively, outperforming the previous best system of 0.439 with 81M parameters. In addition, our model also achieves a new state-of-the-art mAP of 0.567 on FSD50K. Yuan Gong 0001, Yu-An Chung, James R. Glass |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | What Was Written vs. Who Read It: News Media Profiling Using Text Analysis and Social Media ContextabstractPredicting the political bias and the factuality of reporting of entire news outlets are critical elements of media profiling, which is an understudied but an increasingly important research direction.The present level of proliferation of fake, biased, and propagandistic content online, has made it impossible to fact-check every single suspicious claim, either manually or automatically.Alternatively, we can profile entire news outlets and look for those that are likely to publish fake or biased content.This approach makes it possible to detect likely "fake news" the moment they are published, by simply checking the reliability of their source.From a practical perspective, political bias and factuality of reporting have a linguistic aspect but also a social context.Here, we study the impact of both, namely (i) what was written (i.e., what was published by the target medium, and how it describes itself on Twitter) vs. (ii) who read it (i.e., analyzing the readers of the target medium on Facebook, Twitter, and YouTube).We further study (iii) what was written about the target medium on Wikipedia.The evaluation results show that what was written matters most, and that putting all information sources together yields huge improvements over the current state-of-the-art. Ramy Baly, Georgi Karadzhov, Jisun An, Haewoon Kwak, Yoan Dinkov, Ahmed Ali 0002, James R. Glass, Preslav Nakov |
ACL | 7 |
| 2020 | Improved Speech Representations with Multi-Target Autoregressive Predictive CodingabstractTraining objectives based on predictive coding have recently been shown to be very effective at learning meaningful representations from unlabeled speech.One example is Autoregressive Predictive Coding (Chung et al., 2019), which trains an autoregressive RNN to generate an unseen future frame given a context such as recent past frames.The basic hypothesis of these approaches is that hidden states that can accurately predict future frames are a useful representation for many downstream tasks.In this paper we extend this hypothesis and aim to enrich the information encoded in the hidden states by training the model to make more accurate future predictions.We propose an auxiliary objective that serves as a regularization to improve generalization of the future frame prediction task.Experimental results on phonetic classification, speech recognition, and speech translation not only support the hypothesis, but also demonstrate the effectiveness of our approach in learning representations that contain richer phonetic content. Yu-An Chung, James R. Glass |
ACL | 2 |
| 2020 | Negative Training for Neural Dialogue Response GenerationabstractAlthough deep learning models have brought tremendous advancements to the field of opendomain dialogue response generation, recent research results have revealed that the trained models have undesirable generation behaviors, such as malicious responses and generic (boring) responses.In this work, we propose a framework named "Negative Training" to minimize such behaviors.Given a trained model, the framework will first find generated samples that exhibit the undesirable behavior, and then use them to feed negative training signals for fine-tuning the model.Our experiments show that negative training can significantly reduce the hit rate of malicious responses, or discourage frequent responses and improve response diversity. Tianxing He, James R. Glass |
ACL | 2 |
| 2020 | Similarity Analysis of Contextual Word Representation ModelsabstractThis paper investigates contextual word representation models from the lens of similarity analysis.Given a collection of trained models, we measure the similarity of their internal representations and attention.Critically, these models come from vastly different architectures.We use existing and novel similarity measures that aim to gauge the level of localization of information in the deep models, and facilitate the investigation of which design factors affect model similarity, without requiring any external linguistic annotation.The analysis reveals that models within the same family are more similar to one another, as may be expected.Surprisingly, different architectures have rather similar representations, but different individual neurons.We also observed differences in information localization in lower and higher layers and found that higher layers are more affected by fine-tuning on downstream tasks. 1 John M. Wu, Yonatan Belinkov, Hassan Sajjad 0001, Nadir Durrani, Fahim Dalvi, James R. Glass |
ACL | 6 |
| 2020 | We Can Detect Your Bias: Predicting the Political Ideology of News ArticlesabstractWe explore the task of predicting the leading political ideology or bias of news articles.First, we collect and release a large dataset of 34,737 articles that were manually annotated for political ideology -left, center, or right-, which is well-balanced across both topics and media.We further use a challenging experimental setup where the test examples come from media that were not seen during training, which prevents the model from learning to detect the source of the target news article instead of predicting its political ideology.From a modeling perspective, we propose an adversarial media adaptation, as well as a specially adapted triplet loss.We further add background information about the source, and we show that it is quite helpful for improving article-level prediction.Our experimental results show very sizable improvements over using state-of-the-art pre-trained Transformers in this challenging setup. Ramy Baly, Giovanni Da San Martino, James R. Glass, Preslav Nakov |
EMNLP (1) | 3 |
| 2020 | Generative Pre-Training for Speech with Autoregressive Predictive CodingabstractLearning meaningful and general representations from unannotated speech that are applicable to a wide range of tasks remains challenging. In this paper we propose to use autoregressive predictive coding (APC), a recently proposed self-supervised objective, as a generative pre-training approach for learning meaningful, non-specific, and transferable speech representations. We pre-train APC on large-scale unlabeled data and conduct transfer learning experiments on three speech applications that require different information about speech characteristics to perform well: speech recognition, speech translation, and speaker identification. Extensive experiments show that APC not only outperforms surface features (e.g., log Mel spectrograms) and other popular representation learning methods on all three tasks, but is also effective at reducing downstream labeled data size and model parameters. We also investigate the use of Transformers for modeling APC and find it superior to RNNs. Yu-An Chung, James R. Glass |
ICASSP | 2 |
| 2020 | Learning a Subword Inventory Jointly with End-to-End Automatic Speech RecognitionabstractRecent work has demonstrated the promise of using subword units as output targets for sequence-to-sequence automatic speech recognition (ASR) models. Our work builds on the latent sequence decomposition (LSD) framework, in which the use of subword units in ASR is dependent on both the speech input and text output. In this paper, we follow the LSD method for using subword units but introduce an updated loss function that allows the ASR model to explicitly perform unit discovery, as well. We show that our n-gram loss function outperforms standard maximum likelihood loss within the LSD framework. We also show that uniform greedy sampling of subword units, which is much faster than LSD, is also an effective decomposition strategy when combined with the n-gram loss. Along with quantitative results on the Wall Street Journal Corpus, we present an analysis of the subword inventory learned by our model. Jennifer Drexler Fox, James R. Glass |
ICASSP | 2 |
| 2020 | Audio-Visual Calibration with Polynomial Regression for 2-D Projection Using SVD-PHATabstractThis paper proposes a straightforward 2-D method to spatially calibrate the visual field of a camera with the auditory field of an array microphone by generating and overlaying an acoustic image over an optical image. Using a low-cost microphone array and an off-the-shelf camera, we show that polynomial regression can deal efficiently with non-linear camera distortion, and that a recently proposed sound source localization method for real-time processing, SVD-PHAT, can be adapted for this task. François Grondin, Hao Tang 0002, James R. Glass |
ICASSP | 3 |
| 2020 | Trilingual Semantic Embeddings of Visually Grounded Speech with Self-Attention MechanismsabstractWe propose a trilingual semantic embedding model that associates visual objects in images with segments of speech signals corresponding to spoken words in an unsupervised manner. Unlike the existing models, our model incorporates three different languages, namely, English, Hindi, and Japanese. To build the model, we used the existing English and Hindi datasets and collected a new corpus of Japanese speech captions. These spoken captions are spontaneous descriptions by individual speakers, rather than readings based on prepared transcripts. Therefore, we introduce a self-attention mechanism into the model to better map the spoken captions associated with the same image into the embedding space. We hope that the self-attention mechanism efficiently captures relationships between widely separated word-like segments. Experimental results show that the introduction of a third language improves the average performance in terms of cross-modal and cross-lingual retrieval accuracy, and that the self-attention mechanism added to the model works effectively. Yasunori Ohishi, Akisato Kimura, Takahito Kawanishi, Kunio Kashino, David F. Harwath, James R. Glass |
ICASSP | 6 |
| 2020 | ADI17: A Fine-Grained Arabic Dialect Identification DatasetabstractIn this paper, we describe a method to collect dialectal speech from YouTube videos to create a large-scale Dialect Identification (DID) dataset. Using this method, we collected dialectal Arabic from known YouTube channels from 17 Arabic speaking countries in the Middle East and Northern Africa. After a refinement process, a total of 3,000 hours of speech was available for training DID systems, with an additional 57 hours of speech for development and testing. For detailed evaluations, the DID data was divided into three sub-categories based on the segment duration: short (less than 5s), medium (5-20s), and long (over 20s). We compare state-of-the-art DID techniques on these data, and also analyze a DID system trained on these data. Since the training and test data share the same channel domain, we also used the Multi-Genre Broadcast 3 (MGB-3) test set to evaluate on domain mismatched condition. Suwon Shon, Ahmed Ali 0002, Younes Samih, Hamdy Mubarak, James R. Glass |
ICASSP | 5 |
| 2020 | Learning Hierarchical Discrete Linguistic Units from Visually-Grounded Speech
David F. Harwath, Wei-Ning Hsu, James R. Glass |
ICLR | 3 |
| 2020 | What Does an End-to-End Dialect Identification Model Learn About Non-Dialectal Information?
Shammur Absar Chowdhury, Ahmed Ali 0002, Suwon Shon, James R. Glass |
INTERSPEECH | 4 |
| 2020 | Vector-Quantized Autoregressive Predictive CodingabstractAutoregressive Predictive Coding (APC), as a self-supervised objective, has enjoyed success in learning representations from large amounts of unlabeled data, and the learned representations are rich for many downstream tasks.However, the connection between low self-supervised loss and strong performance in downstream tasks remains unclear.In this work, we propose Vector-Quantized Autoregressive Predictive Coding (VQ-APC), a novel model that produces quantized representations, allowing us to explicitly control the amount of information encoded in the representations.By studying a sequence of increasingly limited models, we reveal the constituents of the learned representations.In particular, we confirm the presence of information with probing tasks, while showing the absence of information with mutual information, uncovering the model's preference in preserving speech information as its capacity becomes constrained.We find that there exists a point where phonetic and speaker information are amplified to maximize a selfsupervised objective.As a byproduct, the learned codes for a particular model capacity correspond well to English phones. Yu-An Chung, Hao Tang 0002, James R. Glass |
INTERSPEECH | 3 |
| 2020 | Unsupervised Methods for Evaluating Speech Representations
Michael Gump, Wei-Ning Hsu, James R. Glass |
INTERSPEECH | 3 |
| 2020 | A Convolutional Deep Markov Model for Unsupervised Speech Representation LearningabstractProbabilistic Latent Variable Models (LVMs) provide an alternative to self-supervised learning approaches for linguistic representation learning from speech. LVMs admit an intuitive probabilistic interpretation where the latent structure shapes the information extracted from the signal. Even though LVMs have recently seen a renewed interest due to the introduction of Variational Autoencoders (VAEs), their use for speech representation learning remains largely unexplored. In this work, we propose Convolutional Deep Markov Model (ConvDMM), a Gaussian state-space model with non-linear emission and transition functions modelled by deep neural networks. This unsupervised model is trained using black box variational inference. A deep convolutional neural network is used as an inference network for structured variational approximation. When trained on a large scale speech dataset (LibriSpeech), ConvDMM produces features that significantly outperform multiple self-supervised feature extracting methods on linear phone classification and recognition on the Wall Street Journal dataset. Furthermore, we found that ConvDMM complements self-supervised methods like Wav2Vec and PASE, improving on the results achieved with any of the methods alone. Lastly, we find that ConvDMM features enable learning better phone recognizers than any other features in an extreme low-resource regime with few labeled training examples. Sameer Khurana, Antoine Laurent, Wei-Ning Hsu, Jan Chorowski, Adrian Lancucki, Ricard Marxer, James R. Glass |
INTERSPEECH | 7 |
| 2020 | Prototypical Q Networks for Automatic Conversational Diagnosis and Few-Shot New Disease AdaptionabstractSpoken dialog systems have seen applications in many domains, including medical for automatic conversational diagnosis.State-of-the-art dialog managers are usually driven by deep reinforcement learning models, such as deep Q networks (DQNs), which learn by interacting with a simulator to explore the entire action space since real conversations are limited.However, the DQN-based automatic diagnosis models do not achieve satisfying performances when adapted to new, unseen diseases with only a few training samples.In this work, we propose the Prototypical Q Networks (ProtoQN) as the dialog manager for the automatic diagnosis systems.The model calculates prototype embeddings with real conversations between doctors and patients, learning from them and simulator-augmented dialogs more efficiently.We create both supervised and few-shot learning tasks with the Muzhi corpus.Experiments showed that the ProtoQN significantly outperformed the baseline DQN model in both supervised and few-shot learning scenarios, and achieves state-of-the-art few-shot learning performances. Hongyin Luo, Shang-Wen Li 0001, James R. Glass |
INTERSPEECH | 3 |
| 2020 | Pair Expansion for Learning Multilingual Semantic Embeddings Using Disjoint Visually-Grounded Speech Audio Datasets
Yasunori Ohishi, Akisato Kimura, Takahito Kawanishi, Kunio Kashino, David F. Harwath, James R. Glass |
INTERSPEECH | 6 |
| 2020 | Multimodal Association for Speaker Verification
Suwon Shon, James R. Glass |
INTERSPEECH | 2 |
| 2020 | On the Linguistic Representational Power of Neural Machine Translation ModelsabstractDespite the recent success of deep neural networks in natural language processing and other spheres of artificial intelligence, their interpretability remains a challenge. We analyze the representations learned by neural machine translation (NMT) models at various levels of granularity and evaluate their quality through relevant extrinsic properties. In particular, we seek answers to the following questions: (i) How accurately is word structure captured within the learned representations, which is an important aspect in translating morphologically rich languages? (ii) Do the representations capture long-range dependencies, and effectively handle syntactically divergent languages? (iii) Do the representations capture lexical semantics? We conduct a thorough investigation along several parameters: (i) Which layers in the architecture capture each of these linguistic phenomena; (ii) How does the choice of translation unit (word, character, or subword unit) impact the linguistic properties captured by the underlying representations? (iii) Do the encoder and decoder learn differently and independently? (iv) Do the representations learned by multilingual NMT models capture the same amount of linguistic information as their bilingual counterparts? Our data-driven, quantitative evaluation illuminates important aspects in NMT models and their ability to capture various linguistic phenomena. We show that deep NMT models trained in an end-to-end fashion, without being provided any direct supervision during the training process, learn a non-trivial amount of linguistic information. Notable findings include the following observations: (i) Word morphology and part-of-speech information are captured at the lower layers of the model; (ii) In contrast, lexical semantics or non-local syntactic and semantic dependencies are better represented at the higher layers of the model; (iii) Representations learned using characters are more informed about word-morphology compared to those learned using subword units; and (iv) Representations learned by multilingual models are richer compared to bilingual models. Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad 0001, James R. Glass |
Comput. Linguistics | 5 |
| 2020 | Jointly Discovering Visual Objects and Spoken Words from Raw Sensory Input
David F. Harwath, Adrià Recasens, Didac Suris, Galen Chuang, Antonio Torralba 0001, James R. Glass |
Int. J. Comput. Vis. | 6 |
| 2019 | What Is One Grain of Sand in the Desert? Analyzing Individual Neurons in Deep NLP ModelsabstractDespite the remarkable evolution of deep neural networks in natural language processing (NLP), their interpretability remains a challenge. Previous work largely focused on what these models learn at the representation level. We break this analysis down further and study individual dimensions (neurons) in the vector representation learned by end-to-end neural models in NLP tasks. We propose two methods: Linguistic Correlation Analysis, based on a supervised method to extract the most relevant neurons with respect to an extrinsic task, and Cross-model Correlation Analysis, an unsupervised method to extract salient neurons w.r.t. the model itself. We evaluate the effectiveness of our techniques by ablating the identified neurons and reevaluating the network’s performance for two tasks: neural machine translation (NMT) and neural language modeling (NLM). We further present a comprehensive analysis of neurons with the aim to address the following questions: i) how localized or distributed are different linguistic properties in the models? ii) are certain neurons exclusive to some properties and not others? iii) is the information more or less distributed in NMT vs. NLM? and iv) how important are the neurons identified through the linguistic correlation method to the overall task? Our code is publicly available as part of the NeuroX toolkit (Dalvi et al. 2019a). This paper is a non-archived version of the paper published at AAAI (Dalvi et al. 2019b). Fahim Dalvi, Nadir Durrani, Hassan Sajjad 0001, Yonatan Belinkov, Anthony Bau, James R. Glass |
AAAI | 6 |
| 2019 | NeuroX: A Toolkit for Analyzing Individual Neurons in Neural NetworksabstractWe present a toolkit to facilitate the interpretation and understanding of neural network models. The toolkit provides several methods to identify salient neurons with respect to the model itself or an external task. A user can visualize selected neurons, ablate them to measure their effect on the model accuracy, and manipulate them to control the behavior of the model at the test time. Such an analysis has a potential to serve as a springboard in various research directions, such as understanding the model, better architectural choices, model distillation and controlling data biases. The toolkit is available for download.1 Fahim Dalvi, Avery Nortonsmith, Anthony Bau, Yonatan Belinkov, Hassan Sajjad 0001, Nadir Durrani, James R. Glass |
AAAI | 7 |
| 2019 | Improving Neural Language Models by Segmenting, Attending, and Predicting the FutureabstractCommon language models typically predict the next word given the context.In this work, we propose a method that improves language modeling by learning to align the given context and the following phrase.The model does not require any linguistic annotation of phrase segmentation.Instead, we define syntactic heights and phrase segmentation rules, enabling the model to automatically induce phrases, recognize their task-specific heads, and generate phrase embeddings in an unsupervised learning manner.Our method can easily be applied to language models with different network architectures since an independent module is used for phrase induction and context-phrase alignment, and no change is required in the underlying language modeling network.Experiments have shown that our model outperformed several strong baseline models on different data sets.We achieved a new state-of-the-art performance of 17.4 perplexity on the Wikitext-103 dataset.Additionally, visualizing the outputs of the phrase induction module showed that our model is able to learn approximate phrase-level structural knowledge without any annotation. Hongyin Luo, Yonatan Belinkov, James R. Glass |
ACL (1) | 4 |
| 2019 | The MGB-5 Challenge: Recognition and Dialect Identification of Dialectal Arabic SpeechabstractThis paper describes the fifth edition of the Multi-Genre Broadcast Challenge (MGB-5), an evaluation focused on Arabic speech recognition and dialect identification. MGB-5 extends the previous MGB-3 challenge in two ways: first it focuses on Moroccan Arabic speech recognition; second the granularity of the Arabic dialect identification task is increased from 5 dialect classes to 17, by collecting data from 17 Arabic speaking countries. Both tasks use YouTube recordings to provide a multi-genre multi-dialectal challenge in the wild. Moroccan speech transcription used about 13 hours of transcribed speech data, split across training, development, and test sets, covering 7-genres: comedy, cooking, family/kids, fashion, drama, sports, and science (TEDx). The fine-grained Arabic dialect identification data was collected from known YouTube channels from 17 Arabic countries. 3,000 hours of this data was released for training, and 57 hours for development and testing. The dialect identification data was divided into three sub-categories based on the segment duration: short (under 5 s), medium (5-20 s), and long (>20 s). Overall, 25 teams registered for the challenge, and 9 teams submitted systems for the two tasks. We outline the approaches adopted in each system and summarize the evaluation results. Ahmed Ali 0002, Suwon Shon, Younes Samih, Hamdy Mubarak, Ahmed Abdelali, James R. Glass, Steve Renals, Khalid Choukri |
ASRU | 6 |
| 2019 | Explicit Alignment of Text and Speech Encodings for Attention-Based End-to-End Speech RecognitionabstractIn this work, we present a novel training procedure for attention-based end-to-end automatic speech recognition. Our goal is to push the encoder network to output only linguistic information, improving generalization performance particularly in low-resource scenarios. We accomplish this with the addition of a text encoder network, which the speech encoder is encouraged to mimic. Our main innovation is the comparison of the attention-weighted speech encoder outputs to the outputs of the text encoder - this guarantees two sequences of the same length that can be directly aligned. We show that our training procedure significantly decreases word error rates in all experiments and has the biggest absolute impact in the lowest resource scenarios. Jennifer Drexler Fox, James R. Glass |
ASRU | 2 |
| 2019 | Learning Words by Drawing ImagesabstractWe propose a framework for learning through drawing. Our goal is to learn the correspondence between spoken words and abstract visual attributes, from a dataset of spoken descriptions of images. Building upon recent findings that GAN representations can be manipulated to edit semantic concepts in the generated output, we propose a new method to use such GAN-generated images to train a model using a triplet loss. To apply the method, we develop Audio CLEVRGAN, a new dataset of audio descriptions of GAN-generated CLEVR images, and we describe a training procedure that creates a curriculum of GAN-generated images that focuses training on image pairs that differ in a specific, informative way. Training is done without additional supervision beyond the spoken captions and the GAN. We find that training that takes advantage of GAN-generated edited examples results in improvements in the model's ability to learn attributes compared to previous results. Our proposed learning framework also results in models that can associate spoken words with some abstract visual concepts such as color and size. Didac Suris, Adrià Recasens, David Bau, David F. Harwath, James R. Glass, Antonio Torralba 0001 |
CVPR | 5 |
| 2019 | Contrastive Language Adaptation for Cross-Lingual Stance DetectionabstractMitra Mohtarami, James Glass, Preslav Nakov. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Mitra Mohtarami, James R. Glass, Preslav Nakov |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Towards Unsupervised Speech-to-text TranslationabstractWe present a framework for building speech-to-text translation (ST) systems using only monolingual speech and text corpora, in other words, speech utterances from a source language and independent text from a target language. As opposed to traditional cascaded systems and end-to-end architectures, our system does not require any labeled data (i.e., transcribed source audio or parallel source and target text corpora) during training, making it especially applicable to language pairs with very few or even zero bilingual resources. The framework initializes the ST system with a cross-modal bilingual dictionary inferred from the monolingual corpora, that maps every source speech segment corresponding to a spoken word to its target text translation. For unseen source speech utterances, the system first performs word-by-word translation on each speech segment in the utterance. The translation is improved by leveraging a language model and a sequence denoising autoencoder to provide prior knowledge about the target language. Experimental results show that our unsupervised system achieves comparable BLEU scores to supervised end-to-end models despite the lack of supervision. We also provide an ablation analysis to examine the utility of each component in our system. Yu-An Chung, Wei-Hung Weng, Schrasing Tong, James R. Glass |
ICASSP | 4 |
| 2019 | Subword Regularization and Beam Search Decoding for End-to-end Automatic Speech RecognitionabstractIn this paper, we experiment with the recently introduced subword regularization technique [1] in the context of end-to-end automatic speech recognition (ASR). We present results from both attention-based and CTC-based ASR systems on two common benchmark datasets, the 80 hour Wall Street Journal corpus and 1,000 hour Librispeech corpus. We also introduce a novel subword beam search decoding algorithm that significantly improves the final performance of the CTC-based systems. Overall, we find that subword regularization improves the performance of both types of ASR systems, with the regularized attention-based model performing best overall. Jennifer Drexler Fox, James R. Glass |
ICASSP | 2 |
| 2019 | SVD-PHAT: A Fast Sound Source Localization MethodabstractThis paper introduces a new localization method called SVD-PHAT. The SVD-PHAT method relies on Singular Value De-composition of the SRP-PHAT projection matrix. A k-d tree is also proposed to speed up the search for the most likely direction of arrival of sound. We show that this method per-forms as accurately as SRP-PHAT, while reducing significantly the amount of computation required. François Grondin, James R. Glass |
ICASSP | 2 |
| 2019 | Towards Visually Grounded Sub-word Speech Unit DiscoveryabstractIn this paper, we investigate the manner in which interpretable sub-word speech units emerge within a convolutional neural network model trained to associate raw speech waveforms with semantically related natural image scenes. We show how diphone boundaries can be superficially extracted from the activation patterns of intermediate layers of the model, suggesting that the model may be leveraging these events for the purpose of word recognition. We present a series of experiments investigating the information encoded by these events. David F. Harwath, James R. Glass |
ICASSP | 2 |
| 2019 | Disentangling Correlated Speaker and Noise for Speech Synthesis via Data Augmentation and Adversarial FactorizationabstractTo leverage crowd-sourced data to train multi-speaker text-to-speech (TTS) models that can synthesize clean speech for all speakers, it is essential to learn disentangled representations which can independently control the speaker identity and background noise in generated signals. However, learning such representations can be challenging, due to the lack of labels describing the recording conditions of each training example, and the fact that speakers and recording conditions are often correlated, e.g. since users often make many recordings using the same equipment. This paper proposes three components to address this problem by: (1) formulating a conditional generative model with factorized latent variables, (2) using data augmentation to add noise that is not correlated with speaker identity and whose label is known during training, and (3) using adversarial factorization to improve disentanglement. Experimental results demonstrate that the proposed method can disentangle speaker and noise attributes even if they are correlated in the training data, and can be used to consistently synthesize clean speech for all speakers. Ablation studies verify the importance of each proposed component. Wei-Ning Hsu, Yu Zhang 0033, Ron J. Weiss, Yu-An Chung, Yuxuan Wang 0002, James R. Glass |
ICASSP | 7 |
| 2019 | A Factorial Deep Markov Model for Unsupervised Disentangled Representation Learning from SpeechabstractWe present the Factorial Deep Markov Model (FDMM) for representation learning of speech. The FDMM learns disentangled, interpretable and lower dimensional latent representations from speech without supervision. We use a static and dynamic latent variable to exploit the fact that information in a speech signal evolves at different time scales. Latent representations learned by the FDMM outperform a baseline i-vector system on speaker verification and dialect identification while also reducing the error rate of a phone recognition system in a domain mismatch scenario. Sameer Khurana, Shafiq R. Joty, Ahmed Ali 0002, James R. Glass |
ICASSP | 4 |
| 2019 | Dialogue State Tracking with Convolutional Semantic TaggersabstractIn this paper, we present our novel approach to the 6th Dialogue State Tracking Challenge (DSTC6) track for end-to-end goal-oriented dialogue, in which the goal is to select the best system response from among a list of candidates in a restaurant booking conversation. Our model uses a convolutional neural network (CNN) for semantic tagging of each utterance in the dialogue history to update the dialogue state, and another CNN for predicting the best system action template. Our model is competitive with the top two submissions to the challenge, achieving 100% precision on subtasks 1 and 2 with a CNN rather than an LSTM for action selection, and a CNN for slot-value tagging, instead of an LSTM or CRF. Mandy Korpusik, James R. Glass |
ICASSP | 2 |
| 2019 | Domain Attentive Fusion for End-to-end Dialect Identification with Unknown Target DomainabstractEnd-to-end deep learning language or dialect identification systems operate on the spectrogram or other acoustic feature and directly generate identification scores for each class. An important issue for end-to-end systems is to have some knowledge of the application domain, because the system can be vulnerable to use cases that were not seen in the training phase; such a scenario is often referred to as a domain mismatched condition. In general, we assume that there is enough variation in the training dataset to expose the system to multiple domains. In this work, we study how to best make use a training dataset in order to have maximum effectiveness on unknown target domains. Our goal is to process the input without any knowledge of the target domain while preserving robust performance on other domains as well. To accomplish this objective, we propose a domain attentive fusion approach for end-to-end dialect/language identification systems. To help with experimentation, we collect a dataset from three different domains, and create experimental protocols for a domain mismatched condition. The results of our proposed approach, which were tested on a variety of broadcast and YouTube data, shows significant performance gain compared to traditional approaches, even without any prior target domain information. Suwon Shon, Ahmed Ali 0002, James R. Glass |
ICASSP | 3 |
| 2019 | Noise-tolerant Audio-visual Online Person Verification Using an Attention-based Neural Network FusionabstractIn this paper, we present a multi-modal online person verification system using both speech and visual signals. Inspired by neuroscientific findings on the association of voice and face, we propose an attention-based end-to-end neural network that learns multi-sensory association for the task of person verification. The attention mechanism in our proposed network learns to conditionally select a salient modality between speech and facial representations that provides a balance between complementary inputs. By virtue of this capability, the network is robust to missing or corrupted data from either modality. In the VoxCeleb2 dataset, we show that our method performs favorably against competing multi-modal methods. Even for extreme cases of large corruption or missing data on either modality, our method demonstrates robustness over other unimodal methods. Suwon Shon, Tae-Hyun Oh, James R. Glass |
ICASSP | 3 |
| 2019 | Identifying and Controlling Important Neurons in Neural Machine Translation
Anthony Bau, Yonatan Belinkov, Hassan Sajjad 0001, Nadir Durrani, Fahim Dalvi, James R. Glass |
ICLR (Poster) | 6 |
| 2019 | Detecting Egregious Responses in Neural Sequence-to-sequence Models
Tianxing He, James R. Glass |
ICLR (Poster) | 2 |
| 2019 | Towards Bilingual Lexicon Discovery From Visually Grounded Speech Audio
Emmanuel Azuh, David F. Harwath, James R. Glass |
INTERSPEECH | 3 |
| 2019 | Analyzing Phonetic and Graphemic Representations in End-to-End Automatic Speech RecognitionabstractEnd-to-end neural network systems for automatic speech recognition (ASR) are trained from acoustic features to text transcriptions. In contrast to modular ASR systems, which contain separately-trained components for acoustic modeling, pronunciation lexicon, and language modeling, the end-to-end paradigm is both conceptually simpler and has the potential benefit of training the entire system on the end task. However, such neural network models are more opaque: it is not clear how to interpret the role of different parts of the network and what information it learns during training. In this paper, we analyze the learned internal representations in an end-to-end ASR model. We evaluate the representation quality in terms of several classification tasks, comparing phonemes and graphemes, as well as different articulatory features. We study two languages (English and Arabic) and three datasets, finding remarkable consistency in how different properties are represented in different layers of the deep neural network. Yonatan Belinkov, Ahmed Ali 0002, James R. Glass |
INTERSPEECH | 3 |
| 2019 | An Unsupervised Autoregressive Model for Speech Representation LearningabstractThis paper proposes a novel unsupervised autoregressive neural model for learning generic speech representations.In contrast to other speech representation learning methods that aim to remove noise or speaker variabilities, ours is designed to preserve information for a wide range of downstream tasks.In addition, the proposed model does not require any phonetic or word boundary labels, allowing the model to benefit from large quantities of unlabeled data.Speech representations learned by our model significantly improve performance on both phone classification and speaker verification over the surface features and other supervised and unsupervised approaches.Further analysis shows that different levels of speech information are captured by our model at different layers.In particular, the lower layers tend to be more discriminative for speakers, while the upper layers provide more phonetic content. Yu-An Chung, Wei-Ning Hsu, Hao Tang 0002, James R. Glass |
INTERSPEECH | 4 |
| 2019 | A Deep Residual Network for Large-Scale Acoustic Scene Analysis
Logan Ford, Hao Tang 0002, François Grondin, James R. Glass |
INTERSPEECH | 4 |
| 2019 | Multiple Sound Source Localization with SVD-PHATabstractThis paper introduces a modification of phase transform on singular value decomposition (SVD-PHAT) to localize multiple sound sources. This work aims to improve localization accuracy and keeps the algorithm complexity low for real-time applications. This method relies on multiple scans of the search space, with projection of each low-dimensional observation onto orthogonal subspaces. We show that this method localizes multiple sound sources more accurately than discrete SRP-PHAT, with a reduction in the Root Mean Square Error up to 0.0395 radians. François Grondin, James R. Glass |
INTERSPEECH | 2 |
| 2019 | Transfer Learning from Audio-Visual Grounding to Speech RecognitionabstractTransfer learning aims to reduce the amount of data required to excel at a new task by re-using the knowledge acquired from learning other related tasks. This paper proposes a novel transfer learning scenario, which distills robust phonetic features from grounding models that are trained to tell whether a pair of image and speech are semantically correlated, without using any textual transcripts. As semantics of speech are largely determined by its lexical content, grounding models learn to preserve phonetic information while disregarding uncorrelated factors, such as speaker and channel. To study the properties of features distilled from different layers, we use them as input separately to train multiple speech recognition models. Empirical results demonstrate that layers closer to input retain more phonetic information, while following layers exhibit greater invariance to domain shift. Moreover, while most previous studies include training data for speech recognition for feature extractor training, our grounding models are not trained on any of those data, indicating more universal applicability to new domains. Wei-Ning Hsu, David F. Harwath, James R. Glass |
INTERSPEECH | 3 |
| 2019 | A Comparison of Deep Learning Methods for Language Understanding
Mandy Korpusik, Zoe Liu, James R. Glass |
INTERSPEECH | 3 |
| 2019 | Integrating Video Retrieval and Moment Detection in a Unified Corpus for Video Question Answering
Hongyin Luo, Mitra Mohtarami, James R. Glass, Karthik Krishnamurthy, Brigitte Richardson |
INTERSPEECH | 3 |
| 2019 | MCE 2018: The 1st Multi-Target Speaker Detection and Identification Challenge EvaluationabstractThe Multi-target Challenge aims to assess how well current speech technology is able to determine whether or not a recorded utterance was spoken by one of a large number of blacklisted speakers. It is a form of multi-target speaker detection based on real-world telephone conversations. Data recordings are generated from call center customer-agent conversations. The task is to measure how accurately one can detect 1) whether a test recording is spoken by a blacklisted speaker, and 2) which specific blacklisted speaker was talking. This paper outlines the challenge and provides its baselines, results, and discussions. Suwon Shon, Najim Dehak, Douglas A. Reynolds, James R. Glass |
INTERSPEECH | 4 |
| 2019 | VoiceID Loss: Speech Enhancement for Speaker VerificationabstractIn this paper, we propose VoiceID loss, a novel loss function for training a speech enhancement model to improve the robustness of speaker verification. In contrast to the commonly used loss functions for speech enhancement such as the L2 loss, the VoiceID loss is based on the feedback from a speaker verification model to generate a ratio mask. The generated ratio mask is multiplied pointwise with the original spectrogram to filter out unnecessary components for speaker verification. In the experiments, we observed that the enhancement network, after training with the VoiceID loss, is able to ignore a substantial amount of time-frequency bins, such as those dominated by noise, for verification. The resulting model consistently improves the speaker verification system on both clean and noisy conditions. Suwon Shon, Hao Tang 0002, James R. Glass |
INTERSPEECH | 3 |
| 2019 | Fast and Robust 3-D Sound Source Localization with DSVD-PHATabstractThis paper introduces a variant of the Singular Value Decomposition with Phase Transform (SVD-PHAT), named Difference SVD-PHAT (DSVD-PHAT), to achieve robust Sound Source Localization (SSL) in noisy conditions. Experiments are performed on a Baxter robot with a four-microphone planar array mounted on its head. Results show that this method offers similar robustness to noise as the state-of-the-art Multiple Signal Classification based on Generalized Singular Value Decomposition (GSVD-MUSIC) method, and considerably reduces the computational load by a factor of 250. This performance gain thus makes DSVD-PHAT appealing for real-time application on robots with limited on-board computing power. François Grondin, James R. Glass |
IROS | 2 |
| 2019 | Language processing and learning models for community question answering in Arabic
Salvatore Romeo, Giovanni Da San Martino, Yonatan Belinkov, Alberto Barrón-Cedeño, Mohamed Eldesouki, Kareem Darwish, Hamdy Mubarak, James R. Glass, Alessandro Moschitti |
Inf. Process. Manag. | 8 |
| 2019 | Analysis Methods in Neural Language Processing: A SurveyabstractAbstract The field of natural language processing has seen impressive progress in recent years, with neural network models replacing many of the traditional systems. A plethora of new models have been proposed, many of which are thought to be opaque compared to their feature-rich counterparts. This has led researchers to analyze, interpret, and evaluate neural networks in novel and more fine-grained ways. In this survey paper, we review analysis methods in neural language processing, categorize them according to prominent research trends, highlight existing limitations, and point to potential directions for future work. Yonatan Belinkov, James R. Glass |
Trans. Assoc. Comput. Linguistics | 2 |
| 2019 | Deep Learning for Database Mapping and Asking Clarification Questions in Dialogue SystemsabstractA dialogue system will often ask followup clarification questions when interacting with a user if the agent is unsure how to respond. In this new study, we explore deep reinforcement learning RL for asking followup questions when a user records a meal description, and the system needs to narrow down the options for which foods the person has eaten. We build off of prior work in which we use novel convolutional neural network models to bypass the standard feature engineering used in dialogue systems to handle the text mismatch between natural language user queries and structured database entries, demonstrating that our model learns semantically meaningful embedding representations of natural language. In this new nutrition domain, the followup clarification questions consist of possible attributes for each food that was consumed; for example, if the user drinks a cup of milk, the system should ask about the percent milkfat. We investigate an RL agent to dynamically follow up with the user, which we compare to rule-based and entropy-based methods. On a held-out test set, assuming the followup questions are answered correctly, deep RL significantly boosts top five food recall from 54.9% without followup to 89.0%. We also demonstrate that a hybrid RL model achieves the best perceived naturalness ratings in a human evaluation. Mandy Korpusik, James R. Glass |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Time-Contrastive Learning Based Deep Bottleneck Features for Text-Dependent Speaker VerificationabstractThere are a number of studies about extraction of bottleneck (BN) features from deep neural networks (DNNs) trained to discriminate speakers, pass-phrases, and triphone states for improving the performance of text-dependent speaker verification (TD-SV). However, a moderate success has been achieved. A recent study presented a time contrastive learning (TCL) concept to explore the non-stationarity of brain signals for classification of brain states. Speech signals have similar non-stationarity property, and TCL further has the advantage of having no need for labeled data. We therefore present a TCL based BN feature extraction method. The method uniformly partitions each speech utterance in a training dataset into a predefined number of multi-frame segments. Each segment in an utterance corresponds to one class, and class labels are shared across utterances. DNNs are then trained to discriminate all speech frames among the classes to exploit the temporal structure of speech. In addition, we propose a segment-based unsupervised clustering algorithm to re-assign class labels to the segments. TD-SV experiments were conducted on the RedDots challenge database. The TCL-DNNs were trained using speech data of fixed pass-phrases that were excluded from the TD-SV evaluation set, so the learned features can be considered phrase-independent. We compare the performance of the proposed TCL BN feature with those of short-time cepstral features and BN features extracted from DNNs discriminating speakers, pass-phrases, speaker+pass-phrase, as well as monophones whose labels and boundaries are generated by three different automatic speech recognition (ASR) systems. Experimental results show that the proposed TCL-BN outperforms cepstral features and speaker+pass-phrase discriminant BN features, and its performance is on par with those of ASR derived BN features. Moreover, the clustering method improves the TD-SV performance of TCL-BN and ASR derived BN features with respect to their standalone counterparts. We further study the TD-SV performance of fusing cepstral and BN features. Achintya Kumar Sarkar, Zheng-Hua Tan, Hao Tang 0002, Suwon Shon, James R. Glass |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2018 | Fact Checking in Community ForumsabstractCommunity Question Answering (cQA) forums are very popular nowadays, as they represent effective means for communities around particular topics to share information. Unfortunately, this information is not always factual. Thus, here we explore a new dimension in the context of cQA, which has been ignored so far: checking the veracity of answers to particular questions in cQA forums. As this is a new problem, we create a specialized dataset for it. We further propose a novel multi-faceted model, which captures information from the answer content (what is said and how), from the author profile (who says it), from the rest of the community forum (where it is said), and from external authoritative sources of information (external support). Evaluation results show a MAP value of 86.54, which is 21 points absolute above the baseline. Tsvetomila Mihaylova, Preslav Nakov, Lluís Màrquez, Alberto Barrón-Cedeño, Mitra Mohtarami, Georgi Karadzhov, James R. Glass |
AAAI | 7 |
| 2018 | Jointly Discovering Visual Objects and Spoken Words from Raw Sensory Input
David F. Harwath, Adrià Recasens, Didac Suris, Galen Chuang, Antonio Torralba 0001, James R. Glass |
ECCV (6) | 6 |
| 2018 | Predicting Factuality of Reporting and Bias of News Media SourcesabstractWe present a study on predicting the factuality of reporting and bias of news media.While previous work has focused on studying the veracity of claims or documents, here we are interested in characterizing entire news media.These are under-studied but arguably important research problems, both in their own right and as a prior for fact-checking systems.We experiment with a large list of news websites and with a rich set of features derived from (i) a sample of articles from the target news medium, (ii) its Wikipedia page, (iii) its Twitter account, (iv) the structure of its URL, and (v) information about the Web traffic it attracts.The experimental results show sizable performance gains over the baselines, and confirm the importance of each feature type. Ramy Baly, Georgi Karadzhov, Dimitar Alexandrov, James R. Glass, Preslav Nakov |
EMNLP | 4 |
| 2018 | Learning Word Representations with Cross-Sentence Dependencyfor End-to-End Co-reference ResolutionabstractIn this work, we present a word embedding model that learns cross-sentence dependency for improving end-to-end co-reference resolution (E2E-CR).While the traditional E2E-CR model generates word representations by running long short-term memory (LSTM) recurrent neural networks on each sentence of an input article or conversation separately, we propose linear sentence linking and attentional sentence linking models to learn crosssentence dependency.Both sentence linking strategies enable the LSTMs to make use of valuable information from context sentences while calculating the representation of the current input word.With this approach, the LSTMs learn word embeddings considering knowledge not only from the current sentence but also from the entire input document.Experiments show that learning cross-sentence dependency enriches information contained by the word representations, and improves the performance of the co-reference resolution model compared with our baseline. Hongyin Luo, James R. Glass |
EMNLP | 2 |
| 2018 | Vision as an Interlingua: Learning Multilingual Semantic Embeddings of Untranscribed SpeechabstractIn this paper, we explore the learning of neural network embeddings for natural images and speech waveforms describing the content of those images. These embeddings are learned directly from the waveforms without the use of linguistic transcriptions or conventional speech recognition technology. While prior work has investigated this setting in the monolingual case using English speech data, this work represents the first effort to apply these techniques to languages beyond English. Using spoken captions collected in English and Hindi, we show that the same model architecture can be successfully applied to both languages. Further, we demonstrate that training a multilingual model simultaneously on both languages offers improved performance over the monolingual models. Finally, we show that these models are capable of performing semantic cross-lingual speech-to-speech retrieval. David F. Harwath, Galen Chuang, James R. Glass |
ICASSP | 3 |
| 2018 | Extracting Domain Invariant Features by Unsupervised Learning for Robust Automatic Speech RecognitionabstractThe performance of automatic speech recognition (ASR) systems can be significantly compromised by previously unseen conditions, which is typically due to a mismatch between training and testing distributions. In this paper, we address robustness by studying domain invariant features, such that domain information becomes transparent to ASR systems, resolving the mismatch problem. Specifically, we investigate a recent model, called the Factorized Hierarchical Variational Autoencoder (FHVAE). FHVAEs learn to factorize sequence-level and segment-level attributes into different latent variables without supervision. We argue that the set of latent variables that contain segment -level information is our desired domain invariant feature for ASR. Experiments are conducted on Aurora-4 and CHiME-4, which demonstrate 41 % and 27% absolute word error rate reductions respectively on mismatched domains. Wei-Ning Hsu, James R. Glass |
ICASSP | 2 |
| 2018 | Energy-Efficient Speaker Identification with Low-Precision NetworksabstractPower-consumption in small devices is dominated by off-chip memory accesses, necessitating small models that can fit in on-chip memory. In the task of text-dependent speaker identification, we demonstrate a 16× byte-size reduction for state-of-art small-footprint LCN/CNN/DNN speaker identification models. We achieve this by using ternary quantization that constrains the weights to {-1, 0, 1}. Our model comfortably fits in the 1 MB on-chip BRAM of most off-the-shelf FPGAs, allowing for a power-efficient speaker ID implementation with 100× fewer floating point multiplications, and a 1000× decrease in estimated energy cost. Additionally, we explore the use of depth-wise separable convolutions for speaker identification, and show while significantly reducing multiplications in full-precision networks, they perform poorly when ternarized. We simulate hardware designs for inference on our model, the first hardware design targeted for efficient evaluation of ternary networks and end-to-end neural network-based speaker identification. Skanda Koppula, James R. Glass, Anantha P. Chandrakasan |
ICASSP | 2 |
| 2018 | Convolutional Neural Networks and Multitask Strategies for Semantic Mapping of Natural Language Input to a Structured DatabaseabstractIn this work, we investigate mapping both natural language food and quantity descriptions to matching USDA database entries. We demonstrate that a convolutional neural network (CNN) model with a softmax layer on top to directly predict the most likely database matches outperforms our previous state-of-the-art approach of learning binary classification and subsequently ranking database entries using similarity scores with the learned embeddings. The softmax model achieves 97.3% top-5 USDA quantity and 91.1 % food recall over the full database, compared to only 70.0% quantity and 46.4% food recall with a sigmoid model, where top-5 recall indicates the percentage of test cases in which the correct quantity or food is in the top-5 hits. Evaluated on 9,600 spoken meals over all foods, the softmax model achieves 91.6% top-5 quantity and 80.1 % food recall. We also explore jointly learning both mappings with a single CNN to boost quantity mapping, and improve food mapping by reranking the food database entries using the predicted quantity matches. Mandy Korpusik, James R. Glass |
ICASSP | 2 |
| 2018 | Exploiting Convolutional Neural Networks for Phonotactic Based Dialect IdentificationabstractIn this paper, we investigate different approaches for Dialect Identification (DID) in Arabic broadcast speech. Dialects differ in their inventory of phonological segments. This paper proposes a new phonotactic based feature representation approach which enables discrimination among different occurrences of the same phone n-grams with different phone duration and probability statistics. To achieve further gain in accuracy we used multi-lingual phone recognizers, trained separately on Arabic, English, Czech, Hungarian and Russian languages. We use Support Vector Machines (SVMs), and Convolutional Neural Networks (CNN s) as backend classifiers throughout the study. The final system fusion results in 24.7% and 19.0% relative error rate reduction compared to that of a conventional phonotactic DID, and i-vectors with bottleneck features. Maryam Najafian, Sameer Khurana, Suwon Shon, Ahmed Ali 0002, James R. Glass |
ICASSP | 5 |
| 2018 | A Noise-Robust Self-Adaptive Multitarget Speaker Detection SystemabstractWe describe a multitarget speaker detection system that provides a robust way to classify the utterance of a speaker in noisy environments. The multitarget detection problem is known to be much more difficult to tackle than single target speaker verification tasks, especially when the target set is large and the data is corrupted by noise. In this work we aim to improve the performance of our multitarget speaker detection system in real-world settings, where complicated background noise and unpredictable speaker behavior are present. We make three major improvements that contribute to our goal. First, we discover an effective noise-filtering method using GMM-based voice activity detector followed by unsupervised bottom-up clustering. Second, we incorporate a Highway-LSTM network to estimate posterior distributions of senones, replacing the traditional GMM-UBM with senone posteriors. Finally, we apply a self-adaptive approach on the classifier back-end so that our PLDA parameters and S-normalization subsets can be updated online. Jianzong Wang, Jing Xiao 0006, Wei-Ning Hsu, James R. Glass |
ICPR | 5 |
| 2018 | Speech2Vec: A Sequence-to-Sequence Framework for Learning Word Embeddings from SpeechabstractIn this paper, we propose a novel deep neural network architecture, Speech2Vec, for learning fixed-length vector representations of audio segments excised from a speech corpus, where the vectors contain semantic information pertaining to the underlying spoken words, and are close to other vectors in the embedding space if their corresponding underlying spoken words are semantically similar.The proposed model can be viewed as a speech version of Word2Vec [1].Its design is based on a RNN Encoder-Decoder framework, and borrows the methodology of skipgrams or continuous bag-of-words for training.Learning word embeddings directly from speech enables Speech2Vec to make use of the semantic information carried by speech that does not exist in plain text.The learned word embeddings are evaluated and analyzed on 13 widely used word similarity benchmarks, and outperform word embeddings learned by Word2Vec from the transcriptions. Yu-An Chung, James R. Glass |
INTERSPEECH | 2 |
| 2018 | Detecting Depression with Audio/Text Sequence Modeling of Interviews
Tuka Al Hanai, Mohammad M. Ghassemi, James R. Glass |
INTERSPEECH | 3 |
| 2018 | Scalable Factorized Hierarchical Variational Autoencoder TrainingabstractDeep generative models have achieved great success in unsupervised learning with the ability to capture complex nonlinear relationships between latent generating factors and observations. Among them, a factorized hierarchical variational autoencoder (FHVAE) is a variational inference-based model that formulates a hierarchical generative process for sequential data. Specifically, an FHVAE model can learn disentangled and interpretable representations, which have been proven useful for numerous speech applications, such as speaker verification, robust speech recognition, and voice conversion. However, as we will elaborate in this paper, the training algorithm proposed in the original paper is not scalable to datasets of thousands of hours, which makes this model less applicable on a larger scale. After identifying limitations in terms of runtime, memory, and hyperparameter optimization, we propose a hierarchical sampling training algorithm to address all three issues. Our proposed method is evaluated comprehensively on a wide variety of datasets, ranging from 3 to 1,000 hours and involving different types of generating factors, such as recording conditions and noise types. In addition, we also present a new visualization method for qualitatively evaluating the performance with respect to the interpretability and disentanglement. Models trained with our proposed algorithm demonstrate the desired characteristics on all the datasets. Wei-Ning Hsu, James R. Glass |
INTERSPEECH | 2 |
| 2018 | Unsupervised Adaptation with Interpretable Disentangled Representations for Distant Conversational Speech RecognitionabstractThe current trend in automatic speech recognition is to leverage large amounts of labeled data to train supervised neural network models. Unfortunately, obtaining data for a wide range of domains to train robust models can be costly. However, it is relatively inexpensive to collect large amounts of unlabeled data from domains that we want the models to generalize to. In this paper, we propose a novel unsupervised adaptation method that learns to synthesize labeled data for the target domain from unlabeled in-domain data and labeled out-of-domain data. We first learn without supervision an interpretable latent representation of speech that encodes linguistic and nuisance factors (e.g., speaker and channel) using different latent variables. To transform a labeled out-of-domain utterance without altering its transcript, we transform the latent nuisance variables while maintaining the linguistic variables. To demonstrate our approach, we focus on a channel mismatch setting, where the domain of interest is distant conversational speech, and labels are only available for close-talking speech. Our proposed method is evaluated on the AMI dataset, outperforming all baselines and bridging the gap between unadapted and in-domain models by over 77% without using any parallel data. Wei-Ning Hsu, Hao Tang 0002, James R. Glass |
INTERSPEECH | 3 |
| 2018 | A Study of Enhancement, Augmentation and Autoencoder Methods for Domain Adaptation in Distant Speech RecognitionabstractSpeech recognizers trained on close-talking speech do not generalize to distant speech and the word error rate degradation can be as large as 40% absolute.Most studies focus on tackling distant speech recognition as a separate problem, leaving little effort to adapting close-talking speech recognizers to distant speech.In this work, we review several approaches from a domain adaptation perspective.These approaches, including speech enhancement, multi-condition training, data augmentation, and autoencoders, all involve a transformation of the data between domains.We conduct experiments on the AMI data set, where these approaches can be realized under the same controlled setting.These approaches lead to different amounts of improvement under their respective assumptions.The purpose of this paper is to quantify and characterize the performance gap between the two domains, setting up the basis for studying adaptation of speech recognizers from close-talking speech to distant speech.Our results also have implications for improving distant speech recognition. Hao Tang 0002, Wei-Ning Hsu, François Grondin, James R. Glass |
INTERSPEECH | 4 |
| 2018 | Supervised and Unsupervised Transfer Learning for Question AnsweringabstractYu-An Chung, Hung-Yi Lee, James Glass. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Yu-An Chung, Hung-yi Lee, James R. Glass |
NAACL-HLT | 3 |
| 2018 | Automatic Stance Detection Using End-to-End Memory NetworksabstractMitra Mohtarami, Ramy Baly, James Glass, Preslav Nakov, Lluís Màrquez, Alessandro Moschitti. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Mitra Mohtarami, Ramy Baly, James R. Glass, Preslav Nakov, Lluís Màrquez, Alessandro Moschitti |
NAACL-HLT | 3 |
| 2018 | Unsupervised Cross-Modal Alignment of Speech and Text Embedding SpacesabstractRecent research has shown that word embedding spaces learned from text corpora of different languages can be aligned without any parallel data supervision. Inspired by the success in unsupervised cross-lingual word embeddings, in this paper we target learning a cross-modal alignment between the embedding spaces of speech and text learned from corpora of their respective modalities in an unsupervised fashion. The proposed framework learns the individual speech and text embedding spaces, and attempts to align the two spaces via adversarial training, followed by a refinement procedure. We show how our framework could be used to perform the tasks of spoken word classification and translation, and the experimental results on these two tasks demonstrate that the performance of our unsupervised alignment approach is comparable to its supervised counterpart. Our framework is especially useful for developing automatic speech recognition (ASR) and speech-to-text translation systems for low- or zero-resource languages, which have little parallel audio-text data for training modern supervised ASR and speech-to-text translation models, but account for the majority of the languages spoken across the world. Yu-An Chung, Wei-Hung Weng, Schrasing Tong, James R. Glass |
NeurIPS | 4 |
| 2018 | Combining End-to-End and Adversarial Training for Low-Resource Speech RecognitionabstractIn this paper, we develop an end-to-end automatic speech recognition (ASR) model designed for a common low-resource scenario: no pronunciation dictionary or phonemic transcripts, very limited transcribed speech, and much larger non-parallel text and speech corpora. Our semi-supervised model is built on top of an encoder-decoder model with attention and takes advantage of non-parallel speech and text corpora in several ways: a denoising text autoencoder that shares parameters with the ASR decoder, a speech autoencoder that shares parameters with the ASR encoder, and adversarial training that encourages the speech and text encoders to use the same embedding space. We show that a model with this architecture significantly outperforms the baseline in this low-resource condition. We additionally perform an ablation evaluation, demonstrating that all of our added components contribute substantially to the overall performance of our model. We propose several avenues for further work, noting in particular that a model with this architecture could potentially enable fully unsupervised speech recognition. Jennifer Drexler Fox, James R. Glass |
SLT | 2 |
| 2018 | Convolutional Neural Networks for Dialogue State Tracking without Pre-Trained Word Vectors or Semantic DictionariesabstractA crucial step in task-oriented dialogue systems is tracking the user's goal over the course of the conversation. This involves maintaining a probability distribution over possible values for each slot (e.g., the foodslot might map to the value Turkish), which gets updated at each turn of the dialogue. Previously, rule-based methods were applied to dialogue systems, or models that required hand-crafted semantic dictionaries mapping phrases to those that are similar in meaning (e.g., area might map to part of town). However, these are expensive to design for each domain, limiting the generalizability. In addition, often a spoken language understanding (SLU) component precedes the dialogue state update mechanism; however, this leads to compounded errors as the output from one module is passed to the next. Instead, more recent work has explored deep learning models for directly updating dialogue state, bypassing the need for SLU or expert-engineered rules. We demonstrate that a novel convolutional neural architecture without any pre-trained word vectors or semantic dictionaries achieves 86.9% joint goal accuracy and 95.4% requested slot accuracy on WOZ 2.0. Mandy Korpusik, James R. Glass |
SLT | 2 |
| 2018 | Unsupervised Representation Learning of Speech for Dialect IdentificationabstractIn this paper, we explore the use of a factorized hierarchical variational autoencoder (FHVAE) model to learn an unsupervised latent representation for dialect identification (DID). An FHVAE can learn a latent space that separates the more static attributes within an utterance from the more dynamic attributes by encoding them into two different sets of latent variables. Useful factors for dialect identification, such as phonetic or linguistic content, are encoded by a segmental latent variable, while irrelevant factors that are relatively constant within a sequence, such as a channel or a speaker information, are encoded by a sequential latent variable. The disentanglement property makes the segmental latent variable less susceptible to channel and speaker variation, and thus reduces degradation from channel domain mismatch. We demonstrate that on fully-supervised DID tasks, an end-to-end model trained on the features extracted from the FHVAE model achieves the best performance, compared to the same model trained on conventional acoustic features and an i-vector based system. Moreover, we also show that the proposed approach can leverage a large amount of unlabeled data for FHVAE training to learn domain-invariant features for DID, and significantly improve the performance in a low-resource condition, where the labels for the in-domain data are not available. Suwon Shon, Wei-Ning Hsu, James R. Glass |
SLT | 3 |
| 2018 | Frame-Level Speaker Embeddings for Text-Independent Speaker Recognition and Analysis of End-to-End ModelabstractIn this paper, we propose a Convolutional Neural Network (CNN) based speaker recognition model for extracting robust speaker embeddings. The embedding can be extracted efficiently with linear activation in the embedding layer. To understand how the speaker recognition model operates with text-independent input, we modify the structure to extract frame-level speaker embeddings from each hidden layer. We feed utterances from the TIMIT dataset to the trained network and use several proxy tasks to study the networks ability to represent speech input and differentiate voice identity. We found that the networks are better at discriminating broad phonetic classes than individual phonemes. In particular, frame-level embeddings that belong to the same phonetic classes are similar (based on cosine distance) for the same speaker. The frame level representation also allows us to analyze the networks at the frame level, and has the potential for other analyses to improve speaker recognition. Suwon Shon, Hao Tang 0002, James R. Glass |
SLT | 3 |
| 2018 | On Training Recurrent Networks with Truncated Backpropagation Through time in Speech RecognitionabstractRecurrent neural networks have been the dominant models for many speech and language processing tasks. However, we understand little about the behavior and the class of functions recurrent networks can realize. Moreover, the heuristics used during training complicate the analyses. In this paper, we study recurrent networks' ability to learn long-term dependency in the context of speech recognition. We consider two decoding approaches, online and batch decoding, and show the classes of functions to which the decoding approaches correspond. We then draw a connection between batch decoding and a popular training approach for recurrent networks, truncated backpropagation through time. Changing the decoding approach restricts the amount of past history recurrent networks can use for prediction, allowing us to analyze their ability to remember. Empirically, we utilize long-term dependency in subphonetic states, phonemes, and words, and show how the design decisions, such as the decoding approach, lookahead, context frames, and consecutive prediction, characterize the behavior of recurrent networks. Finally, we draw a connection between Markov processes and vanishing gradients. These results have implications for studying the long-term dependency in speech data and how these properties are learned by recurrent networks. Hao Tang 0002, James R. Glass |
SLT | 2 |
| 2017 | What do Neural Machine Translation Models Learn about Morphology?abstractNeural machine translation (MT) models obtain state-of-the-art performance while maintaining a simple, end-to-end architecture.However, little is known about what these models learn about source and target languages during the training process.In this work, we analyze the representations learned by neural MT models at various levels of granularity and empirically evaluate the quality of the representations for learning morphology through extrinsic part-of-speech and morphological tagging tasks.We conduct a thorough investigation along several parameters: word-based vs. character-based representations, depth of the encoding layer, the identity of the target language, and encoder vs. decoder representations.Our data-driven, quantitative evaluation sheds light on important aspects in the neural MT system and its ability to capture word structure.1 Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad 0001, James R. Glass |
ACL (1) | 5 |
| 2017 | Learning Word-Like Units from Joint Audio-Visual AnalysisabstractGiven a collection of images and spoken audio captions, we present a method for discovering word-like acoustic units in the continuous speech signal and grounding them to semantically relevant image regions.For example, our model is able to detect spoken instances of the words "lighthouse" within an utterance and associate them with image regions containing lighthouses.We do not use any form of conventional automatic speech recognition, nor do we use any text transcriptions or conventional linguistic annotations.Our model effectively implements a form of spoken language acquisition, in which the computer learns not only to recognize word categories by sound, but also to enrich the words it learns with semantics by grounding them in images. David F. Harwath, James R. Glass |
ACL (1) | 2 |
| 2017 | Spoken language biomarkers for detecting cognitive impairmentabstractIn this study we developed an automated system that evaluates speech and language features from audio recordings of neuropsychological examinations of 92 subjects in the Framingham Heart Study. A total of 265 features were used in an elastic-net regularized binomial logistic regression model to classify the presence of cognitive impairment, and to select the most predictive features. We compared performance with a demographic model from 6,258 subjects in the greater study cohort (0.79 AUC), and found that a system that incorporated both audio and text features performed the best (0.92 AUC), with a True Positive Rate of 29% (at 0% False Positive Rate) and a good model fit (Hosmer-Lemeshow test > 0.05). We also found that decreasing pitch and jitter, shorter segments of speech, and responses phrased as questions were positively associated with cognitive impairment. Tuka Al Hanai, Rhoda Au, James R. Glass |
ASRU | 3 |
| 2017 | Unsupervised domain adaptation for robust speech recognition via variational autoencoder-based data augmentationabstractDomain mismatch between training and testing can lead to significant degradation in performance in many machine learning scenarios. Unfortunately, this is not a rare situation for automatic speech recognition deployments in real-world applications. Research on robust speech recognition can be regarded as trying to overcome this domain mismatch issue. In this paper, we address the unsupervised domain adaptation problem for robust speech recognition, where both source and target domain speech are available, but word transcripts are only available for the source domain speech. We present novel augmentation-based methods that transform speech in a way that does not change the transcripts. Specifically, we first train a variational autoencoder on both source and target domain data (without supervision) to learn a latent representation of speech. We then transform nuisance attributes of speech that are irrelevant to recognition by modifying the latent representations, in order to augment labeled training data with additional data whose distribution is more similar to the target domain. The proposed method is evaluated on the CHiME-4 dataset and reduces the absolute word error rate (WER) by as much as 35% compared to the non-adapted baseline. Wei-Ning Hsu, Yu Zhang 0033, James R. Glass |
ASRU | 3 |
| 2017 | Learning modality-invariant representations for speech and imagesabstractIn this paper, we explore the unsupervised learning of a semantic embedding space for co-occurring sensory inputs. Specifically, we focus on the task of learning a semantic vector space for both spoken and handwritten digits using the TIDIGITs and MNIST datasets. Current techniques encode image and audio/textual inputs directly to semantic embeddings. In contrast, our technique maps an input to the mean and log variance vectors of a diagonal Gaussian from which sample semantic embeddings are drawn. In addition to encouraging semantic similarity between co-occurring inputs, our loss function includes a regularization term borrowed from variational autoencoders (VAEs) which drives the posterior distributions over embeddings to be unit Gaussian. We can use this regularization term to filter out modality information while preserving semantic information. We speculate this technique may be more broadly applicable to other areas of cross-modality/domain information retrieval and transfer learning. Kenneth Leidal, David F. Harwath, James R. Glass |
ASRU | 3 |
| 2017 | Automatic speech recognition of Arabic multi-genre broadcast mediaabstractThis paper describes an Arabic Automatic Speech Recognition system developed on 15 hours of Multi-Genre Broadcast (MGB-3) data from YouTube, plus 1,200 hours of Multi-Dialect and Multi-Genre MGB-2 data recorded from the Aljazeera Arabic TV channel. In this paper, we report our investigations of a range of signal pre-processing, data augmentation, topic-specific language model adaptation, accent specific re-training, and deep learning based acoustic modeling topologies, such as feed-forward Deep Neural Networks (DNNs), Time-delay Neural Networks (TDNNs), Long Short-term Memory (LSTM) networks, Bidirectional LSTMs (BLSTMs), and a Bidirectional version of the Prioritized Grid LSTM (BPGLSTM) model. We propose a system combination for three purely sequence trained recognition systems based on lattice-free maximum mutual information, 4-gram language model re-scoring, and system combination using the minimum Bayes risk decoding criterion. The best word error rate we obtained on the MGB-3 Arabic development set using a 4-gram re-scoring strategy is 42.25% for a chain BLSTM system, compared to 65.44% baseline for a DNN system. Maryam Najafian, Wei-Ning Hsu, Ahmed Ali 0002, James R. Glass |
ASRU | 4 |
| 2017 | MIT-QCRI Arabic dialect identification system for the 2017 multi-genre broadcast challengeabstractIn order to successfully annotate the Arabic speech content found in open-domain media broadcasts, it is essential to be able to process a diverse set of Arabic dialects. For the 2017 Multi-Genre Broadcast challenge (MGB-3) there were two possible tasks: Arabic speech recognition, and Arabic Dialect Identification (ADI). In this paper, we describe our efforts to create an ADI system for the MGB-3 challenge, with the goal of distinguishing amongst four major Arabic dialects, as well as Modern Standard Arabic. Our research focused on dialect variability and domain mismatches between the training and test domain. In order to achieve a robust ADI system, we explored both Siamese neural network models to learn similarity and dissimilarities among Arabic dialects, as well as i-vector post-processing to adapt domain mismatches. Both Acoustic and linguistic features were used for the final MGB-3 submissions, with the best primary system achieving 75% accuracy on the official 10hr test set. Suwon Shon, Ahmed Ali 0002, James R. Glass |
ASRU | 3 |
| 2017 | Semantic mapping of natural language input to database entries via convolutional neural networksabstractNatural language processing research has made major advances with the concept of representing words, sentences, paragraphs, and even documents by embedded vector representations. We apply this idea to the problem of relating foods, as expressed in natural language meal descriptions, to corresponding database entries. We generate fixed-length embeddings for U.S. Department of Agriculture (USDA) food database entries, as well as vector-based representations of natural language meal descriptions, through a convolutional neural network (CNN) architecture that predicts whether or not a USDA food item is present in the meal description. We compute dot products between each token in a meal description and a USDA food entry. By ranking the network's predicted average dot product between each possible database food entry and a meal description, we show it is possible to directly predict the USDA foods mentioned in a meal without requiring intermediate steps that would be used in a conventional database access application. We report the performance of this model on a binary verification task of over 48k meal descriptions, and show that this approach, when integrated with a Markov model, substantially outperforms our previous best multistage approach involving a conditional random field tagger, probabilistic segmentation, and database lookup. Mandy Korpusik, Zachary Collins, James R. Glass |
ICASSP | 3 |
| 2017 | Evaluating Layers of Representation in Neural Machine Translation on Part-of-Speech and Semantic Tagging TasksabstractWhile neural machine translation (NMT) models provide improved translation quality in an elegant framework, it is less clear what they learn about language. Recent work has started evaluating the quality of vector representations learned by NMT models on morphological and syntactic tasks. In this paper, we investigate the representations learned at different layers of NMT encoders. We train NMT systems on parallel data and use the models to extract features for training a classifier on two tasks: part-of-speech and semantic tagging. We then measure the performance of the classifier as a proxy to the quality of the original NMT model for the given task. Our quantitative analysis yields interesting insights regarding representation learning in NMT models. For instance, we find that higher layers are better at learning semantics while lower layers tend to be better for part-of-speech tagging. We also observe little effect of the target language on source-side representations, especially in higher quality models. Yonatan Belinkov, Lluís Màrquez, Hassan Sajjad 0001, Nadir Durrani, Fahim Dalvi, James R. Glass |
IJCNLP(1) | 6 |
| 2017 | An Environmental Feature Representation for Robust Speech Recognition and for Environment Identification
Xue Feng 0002, Brigitte Richardson, Scott Amman, James R. Glass |
INTERSPEECH | 4 |
| 2017 | Learning Latent Representations for Speech Generation and TransformationabstractAn ability to model a generative process and learn a latent representation for speech in an unsupervised fashion will be crucial to process vast quantities of unlabelled speech data. Recently, deep probabilistic generative models such as Variational Autoencoders (VAEs) have achieved tremendous success in modeling natural images. In this paper, we apply a convolutional VAE to model the generative process of natural speech. We derive latent space arithmetic operations to disentangle learned latent representations. We demonstrate the capability of our model to modify the phonetic content or the speaker identity for speech segments using the derived operations, without the need for parallel supervisory data. Wei-Ning Hsu, Yu Zhang 0033, James R. Glass |
INTERSPEECH | 3 |
| 2017 | QMDIS: QCRI-MIT Advanced Dialect Identification System
Sameer Khurana, Maryam Najafian, Ahmed Ali 0002, Tuka Al Hanai, Yonatan Belinkov, James R. Glass |
INTERSPEECH | 6 |
| 2017 | Character-Based Embedding Models and Reranking Strategies for Understanding Natural Language Meal Descriptions
Mandy Korpusik, Zachary Collins, James R. Glass |
INTERSPEECH | 3 |
| 2017 | Analyzing Hidden Representations in End-to-End Automatic Speech Recognition SystemsabstractNeural networks have become ubiquitous in automatic speech recognition systems. While neural networks are typically used as acoustic models in more complex systems, recent studies have explored end-to-end speech recognition systems based on neural networks, which can be trained to directly predict text from input acoustic features. Although such systems are conceptually elegant and simpler than traditional systems, it is less obvious how to interpret the trained models. In this work, we analyze the speech representations learned by a deep end-to-end model that is based on convolutional and recurrent layers, and trained with a connectionist temporal classification (CTC) loss. We use a pre-trained model to generate frame-level features which are given to a classifier that is trained on frame classification into phones. We evaluate representations from different layers of the deep model and compare their quality for predicting phone labels. Our experiments shed light on important aspects of the end-to-end model such as layer depth, model complexity, and other design choices. Yonatan Belinkov, James R. Glass |
NIPS | 2 |
| 2017 | Unsupervised Learning of Disentangled and Interpretable Representations from Sequential DataabstractWe present a factorized hierarchical variational autoencoder, which learns disentangled and interpretable representations from sequential data without supervision. Specifically, we exploit the multi-scale nature of information in sequential data by formulating it explicitly within a factorized hierarchical graphical model that imposes sequence-dependent priors and sequence-independent priors to different sets of latent variables. The model is evaluated on two speech corpora to demonstrate, qualitatively, its ability to transform speakers or linguistic content by manipulating different sets of latent variables; and quantitatively, its ability to outperform an i-vector baseline for speaker verification and reduce the word error rate by as much as 35% in mismatched train/test scenarios for automatic speech recognition tasks. Wei-Ning Hsu, Yu Zhang 0033, James R. Glass |
NIPS | 3 |
| 2017 | Spoken Language Understanding for a Nutrition Dialogue SystemabstractFood logging is recommended by dieticians for prevention and treatment of obesity, but currently available mobile applications for diet tracking are often too difficult and time-consuming for patients to use regularly. For this reason, we propose a novel approach to food journaling that uses speech and language understanding technology in order to enable efficient self-assessment of energy and nutrient consumption. This paper presents ongoing language understanding experiments conducted as part of a larger effort to create a nutrition dialogue system that automatically extracts food concepts from a user's spoken meal description. We first summarize the data collection and annotation of food descriptions performed via Amazon Mechanical Turk (AMT), for both a written corpus and spoken data from an in-domain speech recognizer. We show that the addition of word vector features improves conditional random field (CRF) performance for semantic tagging of food concepts, achieving an average F1 test score of 92.4 on written data; we also demonstrate that a convolutional neural network (CNN) with no hand-crafted features outperforms the best CRF on spoken data, achieving an F1 test score of 91.3. We illustrate two methods for associating foods with properties: segmenting meal descriptions with a CRF, and a complementary method that directly predicts associations with a feed-forward neural network. Finally, we conduct an end-to-end system evaluation through an AMT user study with worker ratings of 83% semantic tagging accuracy. Mandy Korpusik, James R. Glass |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Neural Attention for Learning to Rank Questions in Community Question AnsweringabstractIn real-world data, e.g., from Web forums, text is often contaminated with redundant or irrelevant content, which leads to introducing noise in machine learning algorithms. In this paper, we apply Long Short-Term Memory networks with an attention mechanism, which can select important parts of text for the task of similar question retrieval from community Question Answering (cQA) forums. In particular, we use the attention weights for both selecting entire sentences and their subparts, i.e., word/chunk, from shallow syntactic trees. More interestingly, we apply tree kernels to the filtered text representations, thus exploiting the implicit features of the subtree space for learning question reranking. Our results show that the attention-based pruning allows for achieving the top position in the cQA challenge of SemEval 2016, with a relatively large gap from the other participants while greatly decreasing running time. Salvatore Romeo, Giovanni Da San Martino, Alberto Barrón-Cedeño, Alessandro Moschitti, Yonatan Belinkov, Wei-Ning Hsu, Yu Zhang 0033, Mitra Mohtarami, James R. Glass |
COLING | 9 |
| 2016 | Multilingual data selection for training stacked bottleneck featuresabstractDeep Neural Networks (DNNs) trained on multilingual data have proven useful for improving speech recognition in languages with limited resources. In this framework, data from rich resource languages are pooled together to train a single system and then adapted to a new language. However, data from a rich language that are similar to the target language are generally more helpful. We explore methods of training bottleneck features by using data that are more similar to the target language. Our experiments on speech recognition and keyword spotting tasks with IARPA-Babel languages show that our proposed methods outperform typical multilingual DNNs. Ekapol Chuangsuwanich, Yu Zhang 0033, James R. Glass |
ICASSP | 3 |
| 2016 | Distributional semantics for understanding spoken meal descriptionsabstractThis paper presents ongoing language understanding experiments conducted as part of a larger effort to create a nutrition dialogue system that automatically extracts food concepts from a user's spoken meal description. We first discuss the technical approaches to understanding, including three methods for incorporating word vector features into conditional random field (CRF) models for semantic tagging, as well as classifiers for directly associating foods with properties. We report experiments on both text and spoken data from an in-domain speech recognizer. On text data, we show that the addition of word vector features significantly improves performance, achieving an F1 test score of 90.8 for semantic tagging and 86.3 for food-property association. On speech, the best model achieves an F1 test score of 87.5 for semantic tagging and 86.0 for association. Finally, we conduct an end-to-end system evaluation through a user study with human ratings of 83% semantic tagging accuracy. Mandy Korpusik, Calvin Huang, Michael Price 0001, James R. Glass |
ICASSP | 4 |
| 2016 | Personalized mispronunciation detection and diagnosis based on unsupervised error pattern discoveryabstractIn this work, we introduce two improvements to our previously proposed mispronunciation detection framework. The framework focuses on each learner individually and consists of two main procedures: unsupervised error pattern discovery and pronunciation error decoding. First, we propose nbest filtering to disambiguate uncertain error candidate hypotheses obtained from acoustic similarity clustering. Second, we propose personalized template-based rescoring to refine the mispronunciation detection results. The second contribution of the paper is that we demonstrate the portability of the framework to a new target language. Experimental results on the iCALL corpus, a nonnative Mandarin corpus consisting of speakers of European origin, show that the new error pattern discovery process significantly reduces the size and increases the coverage of the error candidate set. Also, the rescoring technique effectively improves system performance on mispronunciation detection and diagnosis. Ann Lee 0001, Nancy F. Chen, James R. Glass |
ICASSP | 3 |
| 2016 | Prediction-adaptation-correction recurrent neural networks for low-resource language speech recognitionabstractIn this paper, we investigate the use of prediction-adaptation-correction recurrent neural networks (PAC-RNNs) for low-resource speech recognition. A PAC-RNN is comprised of a pair of neural networks in which a correction network uses auxiliary information given by a prediction network to help estimate the state probability. The information from the correction network is also used by the prediction network in a recurrent loop. Our model outperforms other state-of-the-art neural networks (DNNs, LSTMs) on IARPA-Babel tasks. Moreover, transfer learning from a language that is similar to the target language can help improve performance further. Yu Zhang 0033, Ekapol Chuangsuwanich, James R. Glass, Dong Yu 0001 |
ICASSP | 3 |
| 2016 | Highway long short-term memory RNNS for distant speech recognitionabstractIn this paper, we extend the deep long short-term memory (DL-STM) recurrent neural networks by introducing gated direct connections between memory cells in adjacent layers. These direct links, called highway connections, enable unimpeded information flow across different layers and thus alleviate the gradient vanishing problem when building deeper LSTMs. We further introduce the latency-controlled bidirectional LSTMs (BLSTMs) which can exploit the whole history while keeping the latency under control. Efficient algorithms are proposed to train these novel networks using both frame and sequence discriminative criteria. Experiments on the AMI distant speech recognition (DSR) task indicate that we can train deeper LSTMs and achieve better improvement from sequence training with highway LSTMs (HLSTMs). Our novel model obtains 43.9/47.7% WER on AMI (SDM) dev and eval sets, outperforming all previous works. It beats the strong DNN and DLSTM baselines with 15.7% and 5.3% relative improvement respectively. Yu Zhang 0033, Guoguo Chen, Dong Yu 0001, Kaisheng Yao, Sanjeev Khudanpur, James R. Glass |
ICASSP | 6 |
| 2016 | Automatic Dialect Detection in Arabic Broadcast SpeechabstractWe investigate different approaches for dialect identification in Arabic broadcast speech, using phonetic, lexical features obtained from a speech recognition system, and acoustic features using the i-vector framework. We studied both generative and discriminate classifiers, and we combined these features using a multi-class Support Vector Machine (SVM). We validated our results on an Arabic/English language identification task, with an accuracy of 100%. We used these features in a binary classifier to discriminate between Modern Standard Arabic (MSA) and Dialectal Arabic, with an accuracy of 100%. We further report results using the proposed method to discriminate between the five most widely used dialects of Arabic: namely Egyptian, Gulf, Levantine, North African, and MSA, with an accuracy of 52%. We discuss dialect identification errors in the context of dialect code-switching between Dialectal Arabic and MSA, and compare the error pattern between manually labeled data, and the output from our classifier. We also release the train and test data as standard corpus for dialect identification. Ahmed Ali 0002, Najim Dehak, Patrick Cardinal, Sameer Khurana, Sree Harsha Yella, James R. Glass, Peter Bell 0001, Steve Renals |
INTERSPEECH | 6 |
| 2016 | Exploiting Depth and Highway Connections in Convolutional Recurrent Deep Neural Networks for Speech Recognition
Wei-Ning Hsu, Yu Zhang 0033, Ann Lee 0001, James R. Glass |
INTERSPEECH | 4 |
| 2016 | Memory-Efficient Modeling and Search Techniques for Hardware ASR Decoders
Michael Price 0001, Anantha P. Chandrakasan, James R. Glass |
INTERSPEECH | 3 |
| 2016 | Unsupervised Learning of Spoken Language with Visual ContextabstractHumans learn to speak before they can read or write, so why can’t computers do the same? In this paper, we present a deep neural network model capable of rudimentary spoken language acquisition using untranscribed audio training data, whose only supervision comes in the form of contextually relevant visual images. We describe the collection of our data comprised of over 120,000 spoken audio captions for the Places image dataset and evaluate our model on an image search and annotation task. We also provide some visualizations which suggest that our model is learning to recognize meaningful words within the caption spectrograms. David F. Harwath, Antonio Torralba 0001, James R. Glass |
NIPS | 3 |
| 2016 | The MGB-2 challenge: Arabic multi-dialect broadcast media recognitionabstractThis paper describes the Arabic Multi-Genre Broadcast (MGB-2) Challenge for SLT-2016. Unlike last year's English MGB Challenge, which focused on recognition of diverse TV genres, this year, the challenge has an emphasis on handling the diversity in dialect in Arabic speech. Audio data comes from 19 distinct programmes from the Aljazeera Arabic TV channel between March 2005 and December 2015. Programmes are split into three groups: conversations, interviews, and reports. A total of 1,200 hours have been released with lightly supervised transcriptions for the acoustic modelling. For language modelling, we made available over 110M words crawled from Aljazeera Arabic website Aljazeera.net for a 10 year duration 2000-2011. Two lexicons have been provided, one phoneme based and one grapheme based. Finally, two tasks were proposed for this year's challenge: standard speech transcription, and word alignment. This paper describes the task data and evaluation process used in the MGB challenge, and summarises the results obtained. Ahmed Ali 0002, Peter Bell 0001, James R. Glass, Yacine Messaoui, Hamdy Mubarak, Steve Renals |
SLT | 3 |
| 2016 | Development of the MIT ASR system for the 2016 Arabic Multi-genre Broadcast ChallengeabstractThe Arabic language, with over 300 million speakers, has significant diversity and breadth. This proves challenging when building an automated system to understand what is said. This paper describes an Arabic Automatic Speech Recognition system developed on a 1,200 hour speech corpus that was made available for the 2016 Arabic Multi-genre Broadcast (MGB) Challenge. A range of Deep Neural Network (DNN) topologies were modeled including; Feed-forward, Convolutional, Time-Delay, Recurrent Long Short-Term Memory (LSTM), Highway LSTM (H-LSTM), and Grid LSTM (GLSTM). The best performance came from a sequence discriminatively trained G-LSTM neural network. The best overall Word Error Rate (WER) was 18.3% (p <; 0:001) on the development set, after combining hypotheses of 3 and 5 layer sequence discriminatively trained G-LSTM models that had been rescored with a 4-gram language model. Tuka Al Hanai, Wei-Ning Hsu, James R. Glass |
SLT | 3 |
| 2016 | A prioritized grid long short-term memory RNN for speech recognitionabstractRecurrent neural networks (RNNs) are naturally suitable for speech recognition because of their ability of utilizing dynamically changing temporal information. Deep RNNs have been argued to be able to model temporal relationships at different time granularities, but suffer vanishing gradient problems. In this paper, we extend stacked long short-term memory (LSTM) RNNs by using grid LSTM blocks that formulate computation along not only the temporal dimension, but also the depth dimension, in order to alleviate this issue. Moreover, we prioritize the depth dimension over the temporal one to provide the depth dimension more updated information, since the output from it will be used for classification. We call this model the prioritized Grid LSTM (pGLSTM). Extensive experiments on four large datasets (AMI, HKUST, GALE, and MGB) indicate that the pGLSTM outperforms alternative deep LSTM models, beating stacked LSTMs with 4% to 7% relative improvement, and achieve new benchmarks among uni-directional models on all datasets. Wei-Ning Hsu, Yu Zhang 0033, James R. Glass |
SLT | 3 |
| 2016 | Look, listen, and decode: Multimodal speech recognition with imagesabstractIn this paper, we introduce a multimodal speech recognition scenario, in which an image provides contextual information for a spoken caption to be decoded. We investigate a lattice rescoring algorithm that integrates information from the image at two different points: the image is used to augment the language model with the most likely words, and to rescore the top hypotheses using a word-level RNN. This rescoring mechanism decreases the word error rate by 3 absolute percentage points, compared to a baseline speech recognizer operating with only the speech recording. Felix Sun, David F. Harwath, James R. Glass |
SLT | 3 |
| 2016 | On the Use of Acoustic Unit Discovery for Language RecognitionabstractIn this paper, we explore the use of large-scale acoustic unit discovery for language recognition. The deep neural network-based approaches that have achieved recent success in this task require transcribed speech and pronunciation dictionaries, which may be limited in availability and expensive to obtain. We aim to replace the need for such supervision via the unsupervised discovery of acoustic units. In this work, we present a parallelized version of a Bayesian nonparametric model from previous work and use it to learn acoustic units from a few hundred hours of multilingual data. These unit (or senone) sequences are then used as targets to train a deep neural network-based i-vector language recognition system. We find that a score-level fusion of our unsupervised system with an acoustic baseline can shrink the gap significantly between the baseline and a supervised benchmark system built using transcribed English. Subsequent experiments also show that an improved acoustic representation of the data can yield substantial performance gains and that language specificity is important for discovering meaningful acoustic units. We validate the generalizability of our proposed approach by presenting state-of-the-art results that exhibit similar trends on the NIST Language Recognition Evaluations from 2011 and 2015. Stephen H. Shum, David F. Harwath, Najim Dehak, James R. Glass |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2015 | Deep multimodal semantic embeddings for speech and imagesabstractIn this paper, we present a model which takes as input a corpus of images with relevant spoken captions and finds a correspondence between the two modalities. We employ a pair of convolutional neural networks to model visual objects and speech signals at the word level, and tie the networks together with an embedding and alignment model which learns a joint semantic space over both modalities. We evaluate our model using image search and annotation tasks on the Flickr8k dataset, which we augmented by collecting a corpus of 40,000 spoken captions using Amazon Mechanical Turk. David F. Harwath, James R. Glass |
ASRU | 2 |
| 2015 | Wait-Learning: Leveraging Wait Time for Second Language EducationabstractCompeting priorities in daily life make it difficult for those with a casual interest in learning to set aside time for regular practice. In this paper, we explore wait-learning: leveraging brief moments of waiting during a person's existing conversations for second language vocabulary practice, even if the conversation happens in the native language. We present an augmented version of instant messaging, WaitChatter, that supports the notion of wait-learning by displaying contextually relevant foreign language vocabulary and micro-quizzes just-in-time while the user awaits a response from her conversant. Through a two week field study of WaitChatter with 20 people, we found that users were able to learn 57 new words on average during casual instant messaging. Furthermore, we found that users were most receptive to learning opportunities immediately after sending a chat message, and that this timing may be critical given user tendency to multi-task during waiting periods. Carrie J. Cai, Philip J. Guo, James R. Glass, Rob Miller 0001 |
CHI | 3 |
| 2015 | Arabic Diacritization with Recurrent Neural NetworksabstractArabic, Hebrew, and similar languages are typically written without diacritics, leading to ambiguity and posing a major challenge for core language processing tasks like speech recognition.Previous approaches to automatic diacritization employed a variety of machine learning techniques.However, they typically rely on existing tools like morphological analyzers and therefore cannot be easily extended to new genres and languages.We develop a recurrent neural network with long shortterm memory layers for predicting diacritics in Arabic text.Our language-independent approach is trained solely from diacritized text without relying on external tools.We show experimentally that our model can rival state-of-the-art methods that have access to additional resources. Yonatan Belinkov, James R. Glass |
EMNLP | 2 |
| 2015 | On using heterogeneous data for vehicle-based speech recognition: A DNN-based approachabstractMost automatic speech recognition (ASR) systems incorporate a single source of information about their input, namely, features and transformations derived from the speech signal. However, in many applications, e.g., vehicle-based speech recognition, sensor data and environmental information are often available to complement audio information. In this paper, we show how these data can be used to improve hybrid DNN-HMM ASR systems for a vehicle-based speech recognition task. Feature fusion is accomplished by augmenting acoustic features with additional side information before being presented to the DNN acoustic model. The additional features are extracted from the vehicle speed, HVAC status, windshield wiper status, and vehicle type. This supplementary information improves the DNNs ability to discriminate phonetic events in an environment-aware way without having to make any modification to the DNN training algorithms. Experimental results show that heterogeneous data are effective irrespective of whether cross-entropy or sequence training is used. For CE training, a WER reduction of 6.3% is obtained, while sequential training reduces it by 5.5%. Xue Feng 0002, Brigitte Richardson, Scott Amman, James R. Glass |
ICASSP | 4 |
| 2015 | Speaker adaptation using the i-vector technique for bottleneck featuresabstractDeep Neural Networks (DNN) have been largely used and successfully applied in the context of speaker independent Automatic Speech Recognition (ASR). However, these models are not easily adapted to model a specific speaker characteristic. Recently, one approach was proposed to address this issue, which consists of using the I-vector representation as input to the DNN. The I-vector is playing the role of providing information about the speaker as well as the environmental conditions for a given recording. This approach achieved a significant improvement in the context of a hybrid system of DNN combined with Hidden Markov Model (HMM). In this paper, we study the effect of speaker adaptation based on the I-vector framework in the context of stacked bottleneck features. These features, extracted from a second level of DNNs, are modelled by a classical Gaussian Mixture Model (GMM) ASR system. The proposed approach achieved an absolute WER improvement of 1.2% on an Arabic Broadcast news task. Index Terms: DNN, I-Vector, Bottleneck Features, Speech Recognition Patrick Cardinal, Najim Dehak, Yu Zhang 0033, James R. Glass |
INTERSPEECH | 4 |
| 2015 | Mispronunciation detection without nonnative training dataabstractConventional mispronunciation detection systems that have the capability of providing corrective feedback typically require a set of common error patterns that are known beforehand, obtained either by consulting with experts, or from a humanannotated nonnative corpus. In this paper, we propose a mispronunciation detection framework that does not rely on nonnative training data. We first discover an individual learner’s possible pronunciation error patterns by analyzing the acoustic similarities across their utterances. With the discovered error candidates, we iteratively compute forced alignments and decode learner-specific context-dependent error patterns in a greedy manner. We evaluate the framework on a Chinese University of Hong Kong (CUHK) corpus containing both Cantonese and Mandarin speakers reading English. Experimental results show that the proposed framework effectively detects mispronunciations and also has a good ability to prioritize feedback. Index Terms: Computer-Assisted Pronunciation Training (CAPT), Gaussian mixture model (GMM), Extended Recognition Network (ERN) Ann Lee 0001, James R. Glass |
INTERSPEECH | 2 |
| 2015 | Unsupervised Lexicon Discovery from Acoustic InputabstractWe present a model of unsupervised phonological lexicon discovery—the problem of simultaneously learning phoneme-like and word-like units from acoustic input. Our model builds on earlier models of unsupervised phone-like unit discovery from acoustic data (Lee and Glass, 2012), and unsupervised symbolic lexicon discovery using the Adaptor Grammar framework (Johnson et al., 2006), integrating these earlier approaches using a probabilistic model of phonological variation. We show that the model is competitive with state-of-the-art spoken term discovery systems, and present analyses exploring the model’s behavior and the kinds of linguistic structures it learns. Chia-ying Lee, Timothy J. O'Donnell, James R. Glass |
Trans. Assoc. Comput. Linguistics | 3 |
| 2015 | Spoken Content Retrieval - Beyond Cascading Speech Recognition with Text RetrievalabstractSpoken content retrieval refers to directly indexing and retrieving spoken content based on the audio rather than text descriptions. This potentially eliminates the requirement of producing text descriptions for multimedia content for indexing and retrieval purposes, and is able to precisely locate the exact time the desired information appears in the multimedia. Spoken content retrieval has been very successfully achieved with the basic approach of cascading automatic speech recognition (ASR) with text information retrieval: after the spoken content is transcribed into text or lattice format, a text retrieval engine searches over the ASR output to find desired information. This framework works well when the ASR accuracy is relatively high, but becomes less adequate when more challenging real-world scenarios are considered, since retrieval performance depends heavily on ASR accuracy. This challenge leads to the emergence of another approach to spoken content retrieval: to go beyond the basic framework of cascading ASR with text retrieval in order to have retrieval performances that are less dependent on ASR accuracy. This overview article is intended to provide a thorough overview of the concepts, principles, approaches, and achievements of major technical contributions along this line of investigation. This includes five major directions: 1) Modified ASR for Retrieval Purposes: cascading ASR with text retrieval, but the ASR is modified or optimized for spoken content retrieval purposes; 2) Exploiting the Information not present in ASR outputs: to try to utilize the information in speech signals inevitably lost when transcribed into phonemes and words; 3) Directly Matching at the Acoustic Level without ASR: for spoken queries, the signals can be directly matched at the acoustic level, rather than at the phoneme or word levels, bypassing all ASR issues; 4) Semantic Retrieval of Spoken Content: trying to retrieve spoken content that is semantically related to the query, but not necessarily including the query terms themselves; 5) Interactive Retrieval and Efficient Presentation of the Retrieved Objects: with efficient presentation of the retrieved objects, an interactive retrieval process incorporating user actions may produce better retrieval results and user experiences. Lin-Shan Lee, James R. Glass, Hung-yi Lee, Chun-an Chan |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | One-shot learning of generative speech concepts
Brenden M. Lake, Chia-ying Lee, James R. Glass, Josh Tenenbaum |
CogSci | 3 |
| 2014 | A Study of using Syntactic and Semantic Structures for Concept Segmentation and Labeling
Iman Saleh 0001, D. Scott Cyphers, James R. Glass, Shafiq R. Joty, Lluís Màrquez, Alessandro Moschitti, Preslav Nakov |
COLING | 3 |
| 2014 | Speech feature denoising and dereverberation via deep autoencoders for noisy reverberant speech recognitionabstractDenoising autoencoders (DAs) have shown success in generating robust features for images, but there has been limited work in applying DAs for speech. In this paper we present a deep denoising autoencoder (DDA) framework that can produce robust speech features for noisy reverberant speech recognition. The DDA is first pre-trained as restricted Boltzmann machines (RBMs) in an unsupervised fashion. Then it is unrolled to autoencoders, and fine-tuned by corresponding clean speech features to learn a nonlinear mapping from noisy to clean features. Acoustic models are re-trained using the reconstructed features from the DDA, and speech recognition is performed. The proposed approach is evaluated on the CHiME-WSJ0 corpus, and shows a 16-25% absolute improvement on the recognition accuracy under various SNRs. Xue Feng 0002, James R. Glass |
ICASSP | 3 |
| 2014 | Extracting deep neural network bottleneck features using low-rank matrix factorizationabstractIn this paper, we investigate the use of deep neural networks (DNNs) to generate a stacked bottleneck (SBN) feature representation for low-resource speech recognition. We examine different SBN extraction architectures, and incorporate low-rank matrix factorization in the final weight layer. Experiments on several low-resource languages demonstrate the effectiveness of the SBN configurations when compared to state-of-the-art hybrid DNN approaches. Yu Zhang 0033, Ekapol Chuangsuwanich, James R. Glass |
ICASSP | 3 |
| 2014 | Recent advances in ASR applied to an Arabic transcription system for Al-JazeeraabstractThis paper describes a detailed comparison of several state-of-the-art speech recognition techniques applied to a limited Ara-bic broadcast news dataset. The different approaches were all trained on 50 hours of transcribed audio from the Al-Jazeera news channel. The best results were obtained using i-vector-based speaker adaptation in a training scenario using the Min-imum Phone Error (MPE) criteria combined with sequential Deep Neural Network (DNN) training. We report results for two different types of test data: broadcast news reports, with a best word error rate (WER) of 17.86%, and a broadcast conver-sations with a best WER of 29.85%. The overall WER on this test set is 25.6%. Index Terms: Arabic, ASR system, Kaldi 1. Patrick Cardinal, Ahmed Ali 0002, Najim Dehak, Yu Zhang 0033, Tuka Al Hanai, James R. Glass, Stephan Vogel |
INTERSPEECH | 7 |
| 2014 | Language ID-based training of multilingual stacked bottleneck featuresabstractIn this paper, we explore multilingual feature-level data sharing via Deep Neural Network (DNN) stacked bottleneck features. Given a set of available source languages, we apply language identification to pick the language most similar to the target language, for more efficient use of multilingual resources. Our experiments with IARPA-Babel languages show that bottleneck features trained on the most similar source language perform better than those trained on all available source languages. Further analysis suggests that only data similar to the target language is useful for multilingual training. Index Terms: Multilingual, Bottleneck features, DNN Anne Cutler, Yu Zhang 0033, Ekapol Chuangsuwanich, James R. Glass |
INTERSPEECH | 4 |
| 2014 | Lexical modeling for Arabic ASR: a systematic approachabstractArabic has an ambiguous mapping between words and pronunciations, making it a deep orthographic system. This ambiguity can be resolved through diacritics, which if displayed, would compose 30% of characters in a text. We investigate the different dimensions of lexical modeling, covering diacritics, pronunciation rules, and acoustic based pronunciation modeling. We show the impact of explicitly modeling the different classes of diacritics (short vowels, geminates, nunnations). We further show that a phonetic lexicon, derived by applying simple pronunciation rules to diacritized words, offers the best gains in ASR performance. Finally, deriving pronunciations from acoustics, yields improvements, beyond a canonical lexicon. Index Terms: automatic speech recognition, Arabic, diacritics, pronunciation rules, language model, lexical model, joint sequence model, pronunciation mixture model. Tuka Al Hanai, James R. Glass |
INTERSPEECH | 2 |
| 2014 | Speech recognition without a lexicon - bridging the gap between graphemic and phonetic systemsabstractModern speech recognizers rely on three core components: an acoustic model, a language model, and a pronunciation lexicon. In order to expand speech recognition capability to lowresource languages and domains, techniques to peel away the expert knowledge required to craft these three components have been growing in popularity. In this paper, we present a method for automatically learning a weighted pronunciation lexicon in a data-driven fashion without assuming the existence of any phonetic lexicon whatsoever. Given an initial grapheme acoustic model, our method utilizes a novel technique for semiconstrained acoustic unit decoding, which is used to help train a letter to sound (L2S) model. The L2S model is then used in conjunction with a Pronunciation Mixture Model (PMM) to infer a pronunciation lexicon. We evaluate our method on English as well as Lao and Haitian, two low-resource languages featured in the IARPA Babel program. Index Terms: lexicon learning, pronunciation modeling David F. Harwath, James R. Glass |
INTERSPEECH | 2 |
| 2014 | Context-dependent pronunciation error pattern discovery with limited annotationsabstractA Computer-Assisted Pronunciation Training (CAPT) system can provide greater benefit to language learners if it provides not only scoring but also corrective feedback. However, the process of deriving pronunciation error patterns usually requires linguistic knowledge, or large quantities of expensive, annotated, corpora from nonnative speakers. In this paper we explore the possibility of deriving context-dependent error patterns with limited human annotations. A two-stage labeling mechanism is proposed, which first selects a set of templates for human annotation, and then propagates the labels. To deal with the imbalanced number of correct and incorrect phone-level pronunciations in nonnative speech, pronunciation patterns on an individual learner-level are first summarized, and then corpuslevel clustering is done for template selection. The concept of contextual similarity based on a phonemic broad class definition is also proposed for label propagation. For evaluation, we view the task as an information retrieval task, and take advantage of metrics that consider both the importance and the ranking of an error type. Experimental results on a Chinese University of Hong Kong (CUHK) nonnative corpus show that the proposed framework can effectively discover prominent error patterns. Index Terms: Computer-Assisted Language Learning, unsupervised clustering, graph-based label propagation Ann Lee 0001, James R. Glass |
INTERSPEECH | 2 |
| 2014 | Graph-based re-ranking using acoustic feature similarity between search results for spoken term detection on low-resource languagesabstractAcoustic feature similarity between search results has been shown to be very helpful for the task of spoken term detection (STD). A graph-based re-ranking approach for STD has been proposed based on the concept that search results, which are acoustically similar to other results with higher confidence scores, should have higher scores themselves. In this approach, the similarity between all search results of a given term are considered as a graph, and the confidence scores of the search results propagate through this graph. Since this approach can improve STD results without any additional labelled data, it is especially suitable for STD on languages with limited amounts of annotated data. However, its performance has not been widely studied on benchmark corpora. In this paper, we investigate the effectiveness of the graph-based reranking approach on limited language data from the IARPA Babel program. Experiments on the low-resource languages, Assamese, Bengali and Lao, show that graph-based re-ranking improves STD systems using fuzzy matching, and lattices based on different kinds of units including words, subwords, and hybrids. Index Terms: Random Walk, Spoken Term Detection Hung-yi Lee, Yu Zhang 0033, Ekapol Chuangsuwanich, James R. Glass |
INTERSPEECH | 4 |
| 2014 | Limited labels for unlimited data: active learning for speaker recognitionabstractIn this paper, we attempt to quantify the amount of labeled data necessary to build a state-of-the-art speaker recognition system. We begin by using i-vectors and the cosine similarity metric to represent an unlabeled set of utterances, then obtain labels from a noiseless oracle in the form of pairwise queries. Finally, we use the resulting speaker clusters to train a PLDA scoring function, which is assessed on the 2010 NIST Speaker Recognition Evaluation. After presenting the initial results of an algorithm that sorts queries based on nearest-neighbor pairs, we develop techniques that further minimize the number of queries needed to obtain state-of-the-art performance. We show the generalizability of our methods in anecdotal fashion by applying our methods to two different distributions of utterances-per-speaker and, ultimately, find that the actual number of pairwise labels needed to obtain state-of-the-art results may be a mere fraction of the queries required to fully label the entire set of utterances. Index Terms: speaker recognition, i-vectors, active learning Stephen H. Shum, Najim Dehak, James R. Glass |
INTERSPEECH | 3 |
| 2014 | A complete KALDI recipe for building Arabic speech recognition systemsabstractIn this paper we present a recipe and language resources for training and testing Arabic speech recognition systems using the KALDI toolkit. We built a prototype broadcast news system using 200 hours GALE data that is publicly available through LDC. We describe in detail the decisions made in building the system: using the MADA toolkit for text normalization and vowelization; why we use 36 phonemes; how we generate pronunciations; how we build the language model. We report results using state-of-the-art modeling and decoding techniques. The scripts are released through KALDI and resources are made available on QCRI's language resources web portal. This is the first effort to share reproducible sizable training and testing results on MSA system. Ahmed Ali 0002, Patrick Cardinal, Najim Dehak, Stephan Vogel, James R. Glass |
SLT | 6 |
| 2014 | Data collection and language understanding of food descriptionsabstractThis paper presents initial data collection and language understanding experiments conducted as part of a larger effort to create a nutrition dialogue system that automatically extracts food concepts from a user's spoken meal description. We first summarize the data collection and annotation of food descriptions performed via Amazon Mechanical Turk. We then present semantic labeling experiments using a semi-Markov conditional random field (CRF) that obtains an F1 test score of 85.1. Finally, we report food segmentation experiments that explored three methods for associating foods with their corresponding attributes: a generative Markov model, transformation-based learning, and a CRF classifier. The CRF performed best, achieving an F1 test score of 87.1. Mandy Korpusik, Nicole Schmidt, Jennifer Drexler Fox, D. Scott Cyphers, James R. Glass |
SLT | 5 |
| 2014 | Non-Negative Factor Analysis of Gaussian Mixture Model Weight Adaptation for Language and Dialect RecognitionabstractRecent studies show that Gaussian mixture model (GMM) weights carry less, yet complimentary, information to GMM means for language and dialect recognition. However, state-of-the-art language recognition systems usually do not use this information. In this research, a non-negative factor analysis (NFA) approach is developed for GMM weight decomposition and adaptation. This modeling, which is conceptually simple and computationally inexpensive, suggests a new low-dimensional utterance representation method using a factor analysis similar to that of the i-vector framework. The obtained subspace vectors are then applied in conjunction with i-vectors to the language/dialect recognition problem. The suggested approach is evaluated on the NIST 2011 and RATS language recognition evaluation (LRE) corpora and on the QCRI Arabic dialect recognition evaluation (DRE) corpus. The assessment results show that the proposed adaptation method yields more accurate recognition results compared to three conventional weight adaptation approaches, namely maximum likelihood re-estimation, non-negative matrix factorization, and a subspace multinomial model. Experimental results also show that the intermediate-level fusion of i-vectors and NFA subspace vectors improves the performance of the state-of-the-art i-vector framework especially for the case of short utterances. Mohamad Hasan Bahari, Najim Dehak, Hugo Van hamme, Lukás Burget, Ahmed Ali 0002, James R. Glass |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2013 | Query understanding enhanced by hierarchical parsing structuresabstractQuery understanding has been well studied in the areas of information retrieval and spoken language understanding (SLU). There are generally three layers of query understanding: domain classification, user intent detection, and semantic tagging. Classifiers can be applied to domain and intent detection in real systems, and semantic tagging (or slot filling) is commonly defined as a sequence-labeling task - mapping a sequence of words to a sequence of labels. Various statistical features (e.g., n-grams) can be extracted from annotated queries for learning label prediction models; however, linguistic characteristics of queries, such as hierarchical structures and semantic relationships, are usually neglected in the feature extraction process. In this work, we propose an approach that leverages linguistic knowledge encoded in hierarchical parse trees for query understanding. Specifically, for natural language queries, we extract a set of syntactic structural features and semantic dependency features from query parse trees to enhance inference model learning. Experiments on real natural language queries show that augmenting sequence labeling models with linguistic knowledge can improve query understanding performance in various domains. Panupong Pasupat, D. Scott Cyphers, James R. Glass |
ASRU | 5 |
| 2013 | Joint Learning of Phonetic Units and Word Pronunciations for ASRabstractThe creation of a pronunciation lexicon remains the most inefficient process in developing an Automatic Speech Recognizer (ASR).In this paper, we propose an unsupervised alternative -requiring no language-specific knowledge -to the conventional manual approach for creating pronunciation dictionaries.We present a hierarchical Bayesian model, which jointly discovers the phonetic inventory and the Letter-to-Sound (L2S) mapping rules in a language using only transcribed data.When tested on a corpus of spontaneous queries, the results demonstrate the superiority of the proposed joint learning scheme over its sequential counterpart, in which the latent phonetic inventory and L2S mappings are learned separately.Furthermore, the recognizers built with the automatically induced lexicon consistently outperform grapheme-based recognizers and even approach the performance of recognition systems trained using conventional supervised procedures. Chia-ying Lee, Yu Zhang 0033, James R. Glass |
EMNLP | 3 |
| 2013 | Zero resource spoken audio corpus analysisabstractZero-resource speech processing involves the automatic analysis of a collection of speech data in a completely unsupervised fashion without the benefit of any transcriptions or annotations of the data. In this paper, our zero-resource system seeks to automatically discover important words, phrases and topical themes present in an audio corpus. This system employs a segmental dynamic time warping (S-DTW) algorithm for acoustic pattern discovery in conjunction with a probabilistic model which treats the topic and pseudo-word identity of each discovered pattern as hidden variables. By applying an Expectation-Maximization (EM) algorithm, our system estimates the latent probability distributions over the pseudo-words and topics associated with the discovered patterns. Using this information, we produce acoustic summaries of the dominant topical themes of the audio document collection. David F. Harwath, Timothy J. Hazen, James R. Glass |
ICASSP | 3 |
| 2013 | Mispronunciation detection via dynamic time warping on deep belief network-based posteriorgramsabstractIn this paper, we explore the use of deep belief network (DBN) posteriorgrams as input to our previously proposed comparison-based system for detecting word-level mispronunciation. The system works by aligning a nonnative utterance with at least one native utterance and extracting features that describe the degree of mis-alignment from the aligned path and the distance matrix. We report system performance under different DBN training scenarios: pre-training and fine-tuning with either native data only or both native and nonnative data. Experimental results have shown that by substituting the system input from MFCC or Gaussian posteriorgrams obtained in a fully unsupervised manner to DBN posteriorgrams, the system performance can be improved by at least 10.4% relatively. Moreover, the system performance remains steady when only 30% of the annotations being used. Ann Lee 0001, James R. Glass |
ICASSP | 3 |
| 2013 | Asgard: A portable architecture for multilingual dialogue systemsabstractSpoken dialogue systems have been studied for years, yet portability is still one of the biggest challenges in terms of language extensibility, domain scalability, and platform compatibility. In this work, we investigate the portability issue from the language understanding perspective and present the Asgard architecture, a CRF-based (Conditional Random Fields) and crowd-sourcing-centered framework, which supports expert-free development of multilingual dialogue systems and seamless deployment to mobile platforms. Combinations of linguistic and statistical features are employed for multilingual semantic understanding, such as n-grams, tokenization and part-of-speech. English and Mandarin systems in various domains (movie, flight and restaurant) are implemented with the proposed framework and ported to mobile platforms as well, which sheds lights on large-scale speech App development. Panupong Pasupat, D. Scott Cyphers, James R. Glass |
ICASSP | 4 |
| 2013 | Bayesian distance metric learning on i-vector for speaker verificationabstractThesis (S.M.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 2013. Najim Dehak, James R. Glass |
INTERSPEECH | 3 |
| 2013 | Learning Lexicons From Speech Using a Pronunciation Mixture ModelabstractIn many ways, the lexicon remains the Achilles heel of modern automatic speech recognizers. Unlike stochastic acoustic and language models that learn the values of their parameters from training data, the baseform pronunciations of words in a recognizer's lexicon are typically specified manually, and do not change, unless they are edited by an expert. Our work presents a novel generative framework that uses speech data to learn stochastic lexicons, thereby taking a step towards alleviating the need for manual intervention and automatically learning high-quality pronunciations for words. We test our model on continuous speech in a weather information domain. In our experiments, we see significant improvements over a manually specified “expert-pronunciation” lexicon. We then analyze variations of the parameter settings used to achieve these gains. Ian McGraw, Ibrahim Badr, James R. Glass |
IEEE Trans. Speech Audio Process. | 3 |
| 2013 | Unsupervised Methods for Speaker Diarization: An Integrated and Iterative ApproachabstractIn speaker diarization, standard approaches typically perform speaker clustering on some initial segmentation before refining the segment boundaries in a re-segmentation step to obtain a final diarization hypothesis. In this paper, we integrate an improved clustering method with an existing re-segmentation algorithm and, in iterative fashion, optimize both speaker cluster assignments and segmentation boundaries jointly. For clustering, we extend our previous research using factor analysis for speaker modeling. In continuing to take advantage of the effectiveness of factor analysis as a front-end for extracting speaker-specific features (i.e., i-vectors), we develop a probabilistic approach to speaker clustering by applying a Bayesian Gaussian Mixture Model (GMM) to principal component analysis (PCA)-processed i-vectors. We then utilize information at different temporal resolutions to arrive at an iterative optimization scheme that, in alternating between clustering and re-segmentation steps, demonstrates the ability to improve both speaker cluster assignments and segmentation boundaries in an unsupervised manner. Our proposed methods attain results that are comparable to those of a state-of-the-art benchmark set on the multi-speaker CallHome telephone corpus. We further compare our system with a Bayesian nonparametric approach to diarization and attempt to reconcile their differences in both methodology and performance. Stephen H. Shum, Najim Dehak, Réda Dehak, James R. Glass |
IEEE Trans. Speech Audio Process. | 4 |
| 2012 | A Nonparametric Bayesian Approach to Acoustic Model Discovery
Chia-ying Lee, James R. Glass |
ACL (1) | 2 |
| 2012 | Evaluation of multi-level context-dependent acoustic model for large vocabulary speaker adaptation tasksabstractIn this paper, we investigate the ability of a recently proposed discriminatively trained, multi-level context-dependent acoustic model to adapt to a new speaker in both supervised and unsupervised adaptation scenarios. Speaker adaptive speech recognition experiments performed on a large-vocabulary spoken lecture task show that the multi-level model reduces word error rates by more than 10% in both cases as compared to the conventional clustering-based decision-tree context-dependent acoustic model approach. Hung-An Chang, James R. Glass |
ICASSP | 2 |
| 2012 | Handling uncertain observations in unsupervised topic-mixture language model adaptationabstractWe propose an extension to the recent approaches in topic-mixture modeling such as Latent Dirichlet Allocation and Topic Tracking Model for the purpose of unsupervised adaptation in speech recognition. Instead of using the 1-best input given by the speech recognizer, the proposed model takes confusion network as an input to alleviate recognition errors. We incorporate a selection variable which helps reweight the recognition output, thus creating a more accurate latent topic estimate. Compared to adapting based on just one recognition hypothesis, the proposed model show WER improvements on two different tasks. Ekapol Chuangsuwanich, Shinji Watanabe 0001, Takaaki Hori, Tomoharu Iwata, James R. Glass |
ICASSP | 5 |
| 2012 | Fast spoken query detection using lower-bound Dynamic Time Warping on Graphical Processing UnitsabstractIn this paper we present a fast unsupervised spoken term detection system based on lower-bound Dynamic Time Warping (DTW) search on Graphical Processing Units (GPUs). The lower-bound estimate and the K nearest neighbor DTW search are carefully designed to fit the GPU parallel computing architecture. In a spoken term detection task on the TIMIT corpus, a 55x speed-up is achieved compared to our previous implementation on a CPU without affecting detection performance. On large, artificially created corpora, measurements show that the total computation time of the entire spoken term detection system grows linearly with corpus size. On average, searching a keyword on a single desktop computer with modern GPUs requires 2.4 seconds/corpus hour. Kiarash Adl, James R. Glass |
ICASSP | 3 |
| 2012 | Resource configurable spoken query detection using Deep Boltzmann MachinesabstractIn this paper we present a spoken query detection method based on posteriorgrams generated from Deep Boltzmann Machines (DBMs). The proposed method can be deployed in both semi-supervised and unsupervised training scenarios. The DBM-based posteriorgrams were evaluated on a series of keyword spotting tasks using the TIMIT speech corpus. In unsupervised training conditions, the DBM-approach improved upon our previous best unsupervised keyword detection performance using Gaussian mixture model-based posteriorgrams by over 10%. When limited amounts of labeled data were incorporated into training, the DBM-approach required less than one third of the annotated data in order to achieve a comparable performance of a system that used all of the annotated data for training. Ruslan Salakhutdinov, Hung-An Chang, James R. Glass |
ICASSP | 4 |
| 2012 | Sentence Detection Using Multiple AnnotationsabstractIn this paper, we develop a sentence boundary detection system which incorporates a prosodic model, word and preterminal-level language models, and a global sentence-length model. An important aspect of this research was the investigation of crowdsourced punctuation annotations as a source of multiple references for evaluation purposes. In order to evaluate the system we propose a BLEU-like metric which compares a hypothesis to multiple references. Experiments on both transcription and ASR output show that the global sentence length model can improve the performance by 7.2 % on reference transcripts and 3.8 % on ASR output. Index Terms: sentence boundary detection, prosody, finite-state transducer, amazon mechanical turk Ann Lee 0001, James R. Glass |
INTERSPEECH | 2 |
| 2012 | A Conversational Movie Search System Based on Conditional Random FieldsabstractOnline streaming companies such as Netflix have become dominant in the media distribution sector. However, such media delivery services often support very rudimentary search, especially for natural language queries. To provide a more natural search interface, we have developed a conversational movie search system, which parses the recognition hypothesis of a spoken query into semantic classes using conditional random fields (CRFs), and then searches an indexed database with the identified semantics. Topic modeling on user-generated content (e.g., movie reviews) is employed for query expansion. Thirteen searching schemas are supported (such as genre, plot, character and soundtrack search). A crowd-sourcing platform was utilized to automatically collect large-scale annotated data for incremental CRF training. Index Terms: conditional random fields, spoken dialogue system 1. D. Scott Cyphers, Panupong Pasupat, Ian McGraw, James R. Glass |
INTERSPEECH | 5 |
| 2012 | Automating Crowd-supervised Learning for Spoken Language SystemsabstractSpoken language systems often rely on static speech recognizers. When the underlying models are improved on-the-fly, training is usually performed using unsupervised methods. In this work, we explore an alternative approach that uses human computation to provide crowd-supervised training of a deployed system. Although the framework we describe is applicable to any stochastic model for which the training data can be generated by non-experts, we demonstrate its utility on the lexicon and language model of a speech recognizer in a cinema voicesearch domain. We show how an initially shaky system can achieve over a 10 % absolute improvement in word error rate (WER) – entirely without expert intervention. We then analyze how these gains were made. 1. Ian McGraw, D. Scott Cyphers, Panupong Pasupat, James R. Glass |
INTERSPEECH | 5 |
| 2012 | On the Use of Spectral and Iterative Methods for Speaker DiarizationabstractThis paper extends upon our previous work using i-vectors for speaker diarization. We examine the effectiveness of spectral clustering as an alternative to our previous approach using K-means clustering and adapt a previously-used heuristic to es-timate the number of speakers. Additionally, we consider an iterative optimization scheme and experiment with its ability to improve both cluster assignments and segmentation boundaries in an unsupervised manner. Our proposed methods attain re-sults similar to those of a state-of-the-art benchmark set on the multi-speaker CallHome telephone corpus. Stephen H. Shum, Najim Dehak, James R. Glass |
INTERSPEECH | 3 |
| 2012 | A comparison-based approach to mispronunciation detectionabstractThe task of mispronunciation detection for language learning is typically accomplished via automatic speech recognition (ASR). Unfortunately, less than 2% of the world's languages have an ASR capability, and the conventional process of creating an ASR system requires large quantities of expensive, annotated data. In this paper we report on our efforts to develop a comparison-based framework for detecting word-level mispronunciations in nonnative speech. Dynamic time warping (DTW) is carried out between a student's (non-native speaker) utterance and a teacher's (native speaker) utterance, and we focus on extracting word-level and phone-level features that describe the degree of mis-alignment in the warping path and the distance matrix. Experimental results on a Chinese University of Hong Kong (CUHK) nonnative corpus show that the proposed framework improves the relative performance on a mispronounced word detection task by nearly 50% compared to an approach that only considers DTW alignment scores. Ann Lee 0001, James R. Glass |
SLT | 2 |
| 2011 | Multi-level context-dependent acoustic modeling for automatic speech recognitionabstractIn this paper, we propose a multi-level, context-dependent acoustic modeling framework for automatic speech recognition. For each context-dependent unit considered by the recognizer, we construct a set of classifiers that target different amounts of contextual resolution, and then combine them for scoring. Since information from multiple levels of contexts is appropriately combined, the proposed modeling framework provides reasonable scores for units with few or no training examples, while maintaining an ability to distinguish between different context-dependent units. On a large vocabulary lecture transcription task, the proposed modeling framework outperforms a traditional clustering-based context-dependent acoustic model by 3.5% (11.4% relative) in terms of word error rate. Hung-An Chang, James R. Glass |
ASRU | 2 |
| 2011 | A channel-blind system for speaker verificationabstractThe majority of speaker verification systems proposed in the NIST speaker recognition evaluation are conditioned on the type of data to be processed: telephone or microphone. In this paper, we propose a new speaker verification system that can be applied to both types of data. This system, named blind system, is based on an extension of the total variability framework. Recognition results with the pro posed channel-independent system are comparable to state of the art systems that require conditioning on the channel type. Another ad vantage of our proposed system is that it allows for combining data from multiple channels in the same visualization in order to explore the effects of different microphones and collection environments. Najim Dehak, Zahi N. Karam, Douglas A. Reynolds, Réda Dehak, William M. Campbell, James R. Glass |
ICASSP | 6 |
| 2011 | An inner-product lower-bound estimate for dynamic time warpingabstractIn this paper, we present a lower-bound estimate for dynamic time warping (DTW) on time series consisting of multi-dimensional posterior probability vectors known as posteriorgrams. We develop a lower-bound estimate based on the inner-product distance that has been found to be an effective metric for computing similarities between posteriorgrams. In addition to deriving the lower-bound estimate, we show how it can be efficiently used in an admissible K nearest neighbor (KNN) search for spotting matching sequences. We quantify the amount of computational savings achieved by performing a set of unsupervised spoken keyword spotting experiments using Gaussian mixture model posteriorgrams. In these experiments the proposed lower-bound estimate eliminates 89% of the DTW previously required calculations without affecting overall keyword detection performance. James R. Glass |
ICASSP | 2 |
| 2011 | Pronunciation Learning from Continuous SpeechabstractThis paper explores the use of continuous speech data to learn stochastic lexicons. Building on previous work in which we augmented graphones with acoustic examples of isolated words, we extend our pronunciation mixture model framework to two domains containing spontaneous speech: a weather information retrieval spoken dialogue system and the academic lectures domain. We find that our learned lexicons out-perform expert, hand-crafted lexicons in each domain. Index Terms: grapheme-to-phoneme conversion, pronunciation models, lexical representation Ibrahim Badr, Ian McGraw, James R. Glass |
INTERSPEECH | 3 |
| 2011 | Robust Voice Activity Detector for Real World Applications Using Harmonicity and Modulation FrequencyabstractThe task of robustly detecting distant speech in low SNR environments for automatic speech recognition is examined using a two-stage approach based on two distinguishing features of speech, namely harmonicity and modulation frequency (MF). A modified metric for harmonicity is used as a gating function to a set of parallel classifiers that incorporate MFs computed on different frequency bands. Performance is evaluated on both the frame-level discriminative power and also the system level ASR results on a real-world robotic forklift task. Compared to other previously proposed features such as relative spectral entropy, and classification strategies involving MFs, the combined approach shows good generalization across different kinds of dynamic noise conditions, and obtains a significant improvement on the false alarm rate at low speech miss rate settings. The overall ASR results also improved significantly compared to the ESTI AMR-VAD2, while reducing the number of false alarms by a factor of two. Index Terms: voice activity detection, modulation frequency, harmonicity, human-robot interaction. Ekapol Chuangsuwanich, James R. Glass |
INTERSPEECH | 2 |
| 2011 | A Transcription Task for Crowdsourcing with Automatic Quality ControlabstractIn this paper, we propose a two-stage transcription task design for crowdsourcing with an automatic quality control mechanism embedded in each stage. For the first stage, a support vector machine (SVM) classifier is utilized to quickly filter poor quality transcripts based on acoustic cues and language patterns in the transcript. In the second stage, word level confidence scores are used to estimate a transcription quality and provide instantaneous feedback to the transcriber. The proposed design was evaluated using Amazon Mechanical Turk (MTurk) and tested on seven hours of academic lecture speech, which is typically conversational in nature and contains technical material. Compared to baseline transcripts which were also collected from MTurk using a ROVER-based method, we observed that the new method resulted in higher quality transcripts while requiring less transcriber effort. Index Terms: Transcription, crowdsourcing, quality control Chia-ying Lee, James R. Glass |
INTERSPEECH | 2 |
| 2011 | An Efferent-Inspired Auditory Model Front-End for Speech RecognitionabstractIn this paper, we investigate a closed-loop auditory model and explore its potential as a feature representation for speech recognition. The closed-loop representation consists of an auditory-based, efferent-inspired feedback mechanism that regulates the operating point of a filter bank, thus enabling it to dynamically adapt to changing background noise. With dynamic adaptation, the closed-loop representation demonstrates an ability to compensate for the effects of noise on speech, and generates a consistent feature representation for speech when contaminated by different kinds of noises. Our preliminary experimental results indicate that the efferent-inspired feedback mechanism enables the closed-loop auditory model to consistently improve word recognition accuracies, when compared with an open-loop representation, for mismatched training and test noise conditions in a connected digit recognition task. Index Terms: efferent, auditory model, feature extraction Chia-ying Lee, James R. Glass, Oded Ghitza |
INTERSPEECH | 2 |
| 2011 | Growing a Spoken Language Interface on Amazon Mechanical TurkabstractTypically data collection, transcription, language model generation, and deployment are separate phases of creating a spoken language interface. An unfortunate consequence of this is that the recognizer usually remains a static element of systems often deployed in dynamic environments. By providing an API for human intelligence, Amazon Mechanical Turk changes the way system developers can construct spoken language systems. In this work, we describe an architecture that automates and connects these four phases, effectively allowing the developer to grow a spoken language interface. In particular, we show that a human-in-the-loop programming paradigm, in which workers transcribe utterances behind the scenes, can alleviate the need for expert guidance in language model construction. We demonstrate the utility of these organic language models in a voice-search interface for photographs. Ian McGraw, James R. Glass, Stephanie Seneff |
INTERSPEECH | 2 |
| 2011 | Exploiting Intra-Conversation Variability for Speaker DiarizationabstractIn this paper, we propose a new approach to speaker diariza-tion based on the Total Variability approach to speaker verifica-tion. Drawing on previous work done in applying factor anal-ysis priors to the diarization problem, we arrive at a simplified approach that exploits intra-conversation variability in the To-tal Variability space through the use of Principal Component Analysis (PCA). Using our proposed methods, we demonstrate the ability to achieve state-of-the-art performance (0.9 % DER) in the diarization of summed-channel telephone data from the NIST 2008 SRE. Stephen H. Shum, Najim Dehak, Ekapol Chuangsuwanich, Douglas A. Reynolds, James R. Glass |
INTERSPEECH | 5 |
| 2011 | A Piecewise Aggregate Approximation Lower-Bound Estimate for Posteriorgram-Based Dynamic Time WarpingabstractIn this paper, we propose a novel lower-bound estimate for dynamic time warping (DTW) methods that use an inner product distance on multi-dimensional posterior probability vectors known as posteriorgrams. Compared to our previous work, the new lower-bound estimate uses piecewise aggregate approximation (PAA) to reduce the time required for calculating the lower-bound estimate. We describe the PAA lower-bound construction process and prove that it can be efficiently used in an admissible K nearest neighbor (KNN) search. The amount of computational savings is quantified by a set of unsupervised spoken keyword spotting experiments. The results show that the newly proposed PAA lower-bound is able to speed up DTW-KNN search by 28 % without affecting the keyword spotting performance. Index Terms: dynamic time warping, lower-bound, posteriorgram 1. James R. Glass |
INTERSPEECH | 2 |
| 2010 | Multimodal interaction with an autonomous forkliftabstractWe describe a multimodal framework for interacting with an autonomous robotic forklift. A key element enabling effective interaction is a wireless, handheld tablet with which a human supervisor can command the forklift using speech and sketch. Most current sketch interfaces treat the canvas as a blank slate. In contrast, our interface uses live and synthesized camera images from the forklift as a canvas, and augments them with object and obstacle information from the world. This connection enables users to "draw on the world," enabling a simpler set of sketched gestures. Our interface supports commands that include summoning the forklift and directing it to lift, transport, and place loads of palletized cargo. We describe an exploratory evaluation of the system designed to identify areas for detailed study. Andrew Correa, Matthew R. Walter, Luke Fletcher, James R. Glass, Seth J. Teller, Randall Davis |
HRI | 4 |
| 2010 | Towards multi-speaker unsupervised speech pattern discoveryabstractIn this paper, we explore the use of a Gaussian posteriorgram based representation for unsupervised discovery of speech patterns. Compared with our previous work, the new approach provides significant improvement towards speaker independence. The framework consists of three main procedures: a Gaussian posteriorgram generation procedure which learns an unsupervised Gaussian mixture model and labels each speech frame with a Gaussian posteriorgram representation; a segmental dynamic time warping procedure which locates pairs of similar sequences of Gaussian posteriorgram vectors; and a graph clustering procedure which groups similar sequences into clusters. We demonstrate the viability of using the posteriorgram approach to handle many talkers by finding clusters of words in the TIMIT corpus. James R. Glass |
ICASSP | 2 |
| 2010 | A voice-commandable robotic forklift working alongside humans in minimally-prepared outdoor environmentsabstractOne long-standing challenge in robotics is the realization of mobile autonomous robots able to operate safely in existing human workplaces in a way that their presence is accepted by the human occupants. We describe the development of a multi-ton robotic forklift intended to operate alongside human personnel, handling palletized materials within existing, busy, semi-structured outdoor storage facilities. The system has three principal novel characteristics. The first is a multimodal tablet that enables human supervisors to use speech and pen-based gestures to assign tasks to the forklift, including manipulation, transport, and placement of palletized cargo. Second, the robot operates in minimally-prepared, semi-structured environments, in which the forklift handles variable palletized cargo using only local sensing (and no reliance on GPS), and transports it while interacting with other moving vehicles. Third, the robot operates in close proximity to people, including its human supervisor, other pedestrians who may cross or block its path, and forklift operators who may climb inside the robot and operate it manually. This is made possible by novel interaction mechanisms that facilitate safe, effective operation around people. We describe the architecture and implementation of the system, indicating how real-world operational requirements motivated the development of the key subsystems, and provide qualitative and quantitative descriptions of the robot operating in real settings. Seth J. Teller, Matthew R. Walter, Matthew E. Antone, Andrew Correa, Randall Davis, Luke Fletcher, Emilio Frazzoli, James R. Glass, Jonathan P. How, Albert S. Huang, Jeong hwan Jeon, Sertac Karaman, Brandon Luders, Nicholas Roy, Tara N. Sainath |
ICRA | 8 |
| 2010 | Learning new word pronunciations from spoken examplesabstractA lexicon containing explicit mappings between words and pronunciations is an integral part of most automatic speech recognizers (ASRs). While many ASR components can be trained or adapted using data, the lexicon is one of the few that typically remains static until experts make manual changes. This work takes a step towards alleviating the need for manual intervention by integrating a popular grapheme-to-phoneme conversion technique with acoustic examples to automatically learn highquality baseform pronunciations for unknown words. We explore two models in a Bayesian framework, and discuss their individual advantages and shortcomings. We show that both are able to generate better-than-expert pronunciations with respect to word error rate on an isolated word recognition task. Index Terms: grapheme-to-phoneme conversion, pronunciation models, lexical representation Ibrahim Badr, Ian McGraw, James R. Glass |
INTERSPEECH | 3 |
| 2010 | Collecting Voices from the Cloud
Ian McGraw, Chia-ying Lee, I. Lee Hetherington, Stephanie Seneff, James R. Glass |
LREC | 5 |
| 2010 | Spoken command of large mobile robots in outdoor environmentsabstractWe describe a speech system for commanding robots in human-occupied outdoor military supply depots. To operate in such environments, the robots must be as easy to interact with as are humans, i.e. they must reliably understand ordinary spoken instructions, such as orders to move supplies, as well as commands and warnings, spoken or shouted from distances of tens of meters. These design goals preclude close-talking microphones and “push-to-talk” buttons that are typically used to isolate commands from the sounds of vehicles, machinery and non-relevant speech. We used multiple microphones to provide omnidirectional coverage. A novel voice activity detector was developed to detect speech and select the appropriate microphone to listen to. Finally, we developed a recognizer model that could successfully recognize commands when heard amidst other speech within a noisy environment. When evaluated on speech data in the field, this system performed significantly better than a more computationally intensive baseline system, reducing the effective false alarm rate by a factor of 40, while maintaining the same level of precision. Ekapol Chuangsuwanich, D. Scott Cyphers, James R. Glass, Seth J. Teller |
SLT | 3 |
| 2010 | A collective data generation method for speech language modelsabstractRecently we began using Amazon Mechanical Turk (AMT), an Internet marketplace, to deploy our spoken dialogue systems to large audiences for user testing and data collection purposes. This crowdsourcing method of collecting data contrasts with the time- and labor- intensive developer annotation methods. In this paper, we compare these data in various combinations with traditionally-collected corpora for training our speech recognizer's language model. Our results show that AMT text queries are effective for initial language model training for spoken dialogue systems, and that crowd-sourced speech collection within the context of a spoken dialogue framework provides significant improvement. Sean Liu, Stephanie Seneff, James R. Glass |
SLT | 3 |
| 2010 | Combining missing-feature theory, speech enhancement, and speaker-dependent/-independent modeling for speech separation
Ji Ming, Timothy J. Hazen, James R. Glass |
Comput. Speech Lang. | 3 |
| 2009 | Unsupervised spoken keyword spotting via segmental DTW on Gaussian posteriorgramsabstractIn this paper, we present an unsupervised learning framework to address the problem of detecting spoken keywords. Without any transcription information, a Gaussian Mixture Model is trained to label speech frames with a Gaussian posteriorgram. Given one or more spoken examples of a keyword, we use segmental dynamic time warping to compare the Gaussian posteriorgrams between keyword samples and test utterances. The keyword detection result is then obtained by ranking the distortion scores of all the test utterances. We examine the TIMIT corpus as a development set to tune the parameters in our system, and the MIT Lecture corpus for more substantial evaluation. The results demonstrate the viability and effectiveness of our unsupervised learning framework on the keyword spotting task. James R. Glass |
ASRU | 2 |
| 2009 | Syntactic Phrase Reordering for English-to-Arabic Statistical Machine Translation
Ibrahim Badr, Rabih Zbib, James R. Glass |
EACL | 3 |
| 2009 | Discriminative training of hierarchical acoustic models for large vocabulary continuous speech recognitionabstractIn this paper we propose discriminative training of hierarchical acoustic models for large vocabulary continuous speech recognition tasks. After presenting our hierarchical modeling framework, we describe how the models can be generated with either minimum classification error or large-margin training. Experiments on a large vocabulary lecture transcription task show that the hierarchical model can yield more than 1.0% absolute word error rate reduction over non-hierarchical models for both kinds of discriminative training. Hung-An Chang, James R. Glass |
ICASSP | 2 |
| 2009 | Language model parameter estimation using user transcriptionsabstractIn limited data domains, many effective language modeling techniques construct models with parameters to be estimated on an in-domain development set. However, in some domains, no such data exist beyond the unlabeled test corpus. In this work, we explore the iterative use of the recognition hypotheses for unsupervised parameter estimation. We also evaluate the effectiveness of supervised adaptation using varying amounts of user-provided transcripts of utterances selected via multiple strategies. While unsupervised adaptation obtains 80% of the potential error reductions, it is outperformed by using only 300 words of user transcription. By transcribing the lowest confidence utterances first, we further obtain an effective word error rate reduction of 0.6%. Bo-June Paul Hsu, James R. Glass |
ICASSP | 2 |
| 2009 | On the phonetic information in ultrasonic microphone signalsabstractWe study the phonetic information in the signal from an ultrasonic ldquomicrophonerdquo, a device that emits an ultrasonic wave toward a speaker and receives the reflected, Doppler-shifted signal. This can be used in addition to audio to improve automatic speech recognition. This work is an effort to better understand the ultrasonic signal, and potentially to determine a set of natural sub-word units. We present classification and clustering experiments on CVC and VCV sequences in speaker-dependent and multi-speaker settings. Using a set of ultrasonic spectral features and diagonal Gaussian models, it is possible to distinguish all consonants and most vowels. When clustering the confusion data, the consonant clusters mostly correspond to places and manners of articulation; the vowel data roughly clusters into high, low, and rounded vowels. Karen Livescu, Bo Zhu 0004, James R. Glass |
ICASSP | 3 |
| 2009 | Speech rhythm guided syllable nuclei detectionabstractIn this paper, we present a novel speech-rhythm-guided syllable-nuclei location detection algorithm. As a departure from conventional methods, we introduce an instantaneous speech rhythm estimator to predict possible regions where syllable nuclei can appear. Within a possible region, a simple slope based peak counting algorithm is used to get the exact location of each syllable nucleus. We verify the correctness of our method by investigating the syllable nuclei interval distribution in TIMIT dataset, and evaluate the performance by comparing with a state-of-the-art syllable nuclei based speech rate detection approach. James R. Glass |
ICASSP | 2 |
| 2009 | A back-off discriminative acoustic model for automatic speech recognitionabstractIn this paper we propose a back-off discriminative acoustic model for Automatic Speech Recognition (ASR). We use a set of broad phonetic classes to divide the classification problem originating from context-dependent modeling into a set of sub-problems. By appropriately combining the scores from clas-sifiers designed for the sub-problems, we can guarantee that the back-off acoustic score for different context-dependent units will be different. The back-off model can be combined with discriminative training algorithms to further improve the per-formance. Experimental results on a large vocabulary lecture transcription task show that the proposed back-off discrimina-tive acoustic model has more than a 2.0 % absolute word error rate reduction compared to clustering-based acoustic model. Index Terms: context-dependent acoustic modeling, back-off acoustic models, discriminative training, Hung-An Chang, James R. Glass |
INTERSPEECH | 2 |
| 2009 | Multistream Articulatory Feature-Based Models for Visual Speech RecognitionabstractWe study the problem of automatic visual speech recognition (VSR) using dynamic Bayesian network (DBN)-based models consisting of multiple sequences of hidden states, each corresponding to an articulatory feature (AF) such as lip opening (LO) or lip rounding (LR). A bank of discriminative articulatory feature classifiers provides input to the DBN, in the form of either virtual evidence (VE) (scaled likelihoods) or raw classifier margin outputs. We present experiments on two tasks, a medium-vocabulary word-ranking task and a small-vocabulary phrase recognition task. We show that articulatory feature-based models outperform baseline models, and we study several aspects of the models, such as the effects of allowing articulatory asynchrony, of using dictionary-based versus whole-word models, and of incorporating classifier outputs via virtual evidence versus alternative observation models. Kate Saenko, Karen Livescu, James R. Glass, Trevor Darrell |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2008 | N-gram Weighting: Reducing Training Data Mismatch in Cross-Domain Language Model Estimation
Bo-June Paul Hsu, James R. Glass |
EMNLP | 2 |
| 2008 | A turbo-style algorithm for lexical baseforms estimationabstractIn this research, an iterative and unsupervised Turbo-style algorithm is presented and implemented for the task of automatic lexical acquisition. The algorithm makes use of spoken examples of both spellings and words and fuses information from letter and subword recognizers to boost the overall lexical learning performance. The algorithm is tested on a challenging lexicon of restaurant and street names and evaluated in terms of spelling accuracy and letter error rate. Absolute improvements of 7.2% and 3% (15.5% relative improvement) are obtained in the spelling accuracy and the letter error rate respectively following only 2 iterations of the algorithm. Ghinwa F. Choueiter, Mesrob I. Ohannessian, Stephanie Seneff, James R. Glass |
ICASSP | 4 |
| 2008 | Iterative language model estimation: efficient data structure & algorithms
Bo-June Paul Hsu, James R. Glass |
INTERSPEECH | 2 |
| 2008 | Unsupervised Pattern Discovery in SpeechabstractWe present a novel approach to speech processing based on the principle of pattern discovery. Our work represents a departure from traditional models of speech recognition, where the end goal is to classify speech into categories defined by a prespecified inventory of lexical units (i.e., phones or words). Instead, we attempt to discover such an inventory in an unsupervised manner by exploiting the structure of repeating patterns within the speech signal. We show how pattern discovery can be used to automatically acquire lexical entities directly from an untranscribed audio stream. Our approach to unsupervised word acquisition utilizes a segmental variant of a widely used dynamic programming technique, which allows us to find matching acoustic patterns between spoken utterances. By aggregating information about these matching patterns across audio streams, we demonstrate how to group similar acoustic sequences together to form clusters corresponding to lexical entities such as words and short multiword phrases. On a corpus of academic lecture material, we demonstrate that clusters found using this technique exhibit high purity and that many of the corresponding lexical identities are relevant to the underlying audio stream. A. S. Park, James R. Glass |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | Making Sense of Sound: Unsupervised Topic Segmentation over Acoustic Input
Igor Malioutov, Alex Park 0001, Regina Barzilay, James R. Glass |
ACL | 4 |
| 2007 | Hierarchical large-margin Gaussian mixture models for phonetic classificationabstractIn this paper we present a hierarchical large-margin Gaussian mixture modeling framework and evaluate it on the task of phonetic classification. A two-stage hierarchical classifier is trained by alternately updating parameters at different levels in the tree to maximize the joint margin of the overall classification. Since the loss function required in the training is convex to the parameter space the problem of spurious local minima is avoided. The model achieves good performance with fewer parameters than single-level classifiers. In the TIMIT benchmark task of context-independent phonetic classification, the proposed modeling scheme achieves a state-of-the-art phonetic classification error of 16.7% on the core test set. This is an absolute reduction of 1.6% from the best previously reported result on this task, and 4-5% lower than a variety of classifiers that have been recently examined on this task. Hung-An Chang, James R. Glass |
ASRU | 2 |
| 2007 | Automatic lexical pronunciations generation and updateabstractMost automatic speech recognizers use a dictionary that maps words to one or more canonical pronunciations. Such entries are typically hand-written by lexical experts. In this research, we investigate a new approach for automatically generating lexical pronunciations using a linguistically motivated subword model, and refining the pronunciations with spoken examples. The approach is evaluated on an isolated word recognition task with a 2 k lexicon of restaurant and street names. A letter-to-sound model is first used to generate seed baseforms for the lexicon. Then spoken utterances of words in the lexicon are presented to a subword recognizer and the top hypotheses are used to update the lexical base-forms. The spelling of each word is also used to constrain the subword search space and generate spelling-constrained baseforms. The results obtained are quite encouraging and indicate that our approach can be successfully used to learn valid pronunciations of new words. Ghinwa F. Choueiter, Stephanie Seneff, James R. Glass |
ASRU | 3 |
| 2007 | Speech recognition with localized time-frequency pattern detectorsabstractA method for acoustic modeling of speech is presented which is based on learning and detecting the occurrence of localized time-frequency patterns in a spectrogram. A boosting algorithm is applied to both build classifiers and perform feature selection from a large set of features derived by filtering spectrograms. Initial experiments are performed to discriminate digits in the Aurora database. The system succeeds in learning sequences of localized time-frequency patterns which are highly interpretable from an acoustic-phonetic viewpoint. While the work and the results are preliminary, they suggest that pursuing these techniques further could lead to new approaches to acoustic modeling for ASR which are more noise robust and offer better encoding of temporal dynamics than typical features such as frame-based cepstra. Ken Schutte, James R. Glass |
ASRU | 2 |
| 2007 | Open-Vocabulary Spoken Utterance Retrieval using Confusion NetworksabstractThis paper presents a novel approach to open-vocabulary spoken utterance retrieval using confusion networks. If out-of-vocabulary (OOV) words are present in queries and the corpus, word-based indexing will not be sufficient. For this problem, we apply phone confusion networks and combine them with word confusion networks. With this approach, we can generate a more compact index table that enables robust keyword matching compared with typical lattice-based methods. In the retrieval experiments with speech recordings in MIT lecture corpus, our method using phone confusion networks outperformed lattice-based methods especially for OOV queries. Takaaki Hori, I. Lee Hetherington, Timothy J. Hazen, James R. Glass |
ICASSP (4) | 4 |
| 2007 | Noise Robust Phonetic Classificationwith Linear Regularized Least Squares and Second-Order FeaturesabstractWe perform phonetic classification with an architecture whose elements are binary classifiers trained via linear regularized least squares (RLS). RLS is a simple yet powerful regularization algorithm with the desirable property that a good value of the regularization parameter can be found efficiently by minimizing leave-one-out error on the training set. Our system achieves state-of-the-art single classifier performance on the TIMIT phonetic classification task, (slightly) beating other recent systems. We also show that in the presence of additive noise, our model is much more robust than a well-trained Gaussian mixture model. Ryan Rifkin, Ken Schutte, Michelle Saad, Jake V. Bouvrie, James R. Glass |
ICASSP (4) | 5 |
| 2007 | New word acquisition using subword modelingabstractIn this paper, we use subword modeling to learn the pronunciations and spellings of new words. The subwords are generated with a context-free grammar, and are intermediate units between phonemes and syllables. We first evaluate the effectiveness of the subword model in automatically generating the spelling and pronunciation of new words. Then the subword model is embedded in a multi-stage recognizer which consists of word, subword, and letter recognizers. In a preliminary set of experiments, the hybrid system outperforms a large-vocabulary isolated word recognizer. The subword model is also used to improve the performance of the letter recognizer by generating a spelling cohort which is used to train a small letter n-gram. The small letter n-gram has a reduced perplexity compared to a much larger n-gram, and can be used by the letter recognizer for the spoken spelling mode. This could translate to an improved letter error rate in future letter recognition experiments. Index Terms: subword modeling, new word acquisition Ghinwa F. Choueiter, Stephanie Seneff, James R. Glass |
INTERSPEECH | 3 |
| 2007 | Recent progress in the MIT spoken lecture processing projectabstractIn this paper we discuss our research activities in the area of spoken lecture processing. Our goal is to improve the access to on-line audio/visual recordings of academic lectures by developing tools for the processing, transcription, indexing, segmentation, summarization, retrieval and browsing of this media. In this paper, we provide an overview of the technology components and systems that have been developed as part of this project, present some experimental results, and discuss our ongoing and future research plans. Index Terms:spoken lecture processing, spoken document retrieval, audio browsing James R. Glass, Timothy J. Hazen, D. Scott Cyphers, Igor Malioutov, David Huynh, Regina Barzilay |
INTERSPEECH | 1 |
| 2007 | Multimodal speech recognition with ultrasonic sensorsabstractThesis (M. Eng.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 2008. Bo Zhu 0004, Timothy J. Hazen, James R. Glass |
INTERSPEECH | 3 |
| 2007 | An Implementation of Rational Wavelets and Filter Design for Phonetic ClassificationabstractAlthough wavelet analysis has been proposed for speech processing as an alternative to Fourier analysis, most approaches make use of off-the-shelf wavelets and dyadic tree-structured filter banks. In this paper, we extend previous wavelet-based frameworks in two ways. First, we increase the flexibility in wavelet selection by taking advantage of the relationship between wavelets and filter banks and by designing new wavelets using filter design methods. We adopt two filter design techniques that we refer to as filter matching and attenuation minimization. Second, we improve the flexibility in frequency partitioning by implementing rational as well as dyadic filter banks. Rational filter banks naturally incorporate the critical-band effect in the human auditory system. To test our extensions, we implement an energy-based measurement which we also compare in performance to the mel-frequency cepstral coefficients (MFCCs) in a phonetic classification task. We show that the designed wavelets outperform off-the-shelf wavelets as well as an MFCC baseline Ghinwa F. Choueiter, James R. Glass |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | Robust Speaker Recognition in Noisy ConditionsabstractThis paper investigates the problem of speaker identification and verification in noisy conditions, assuming that speech signals are corrupted by environmental noise, but knowledge about the noise characteristics is not available. This research is motivated in part by the potential application of speaker recognition technologies on handheld devices or the Internet. While the technologies promise an additional biometric layer of security to protect the user, the practical implementation of such systems faces many challenges. One of these is environmental noise. Due to the mobile nature of such systems, the noise sources can be highly time-varying and potentially unknown. This raises the requirement for noise robustness in the absence of information about the noise. This paper describes a method that combines multicondition model training and missing-feature theory to model noise with unknown temporal-spectral characteristics. Multicondition training is conducted using simulated noisy data with limited noise variation, providing a ldquocoarserdquo compensation for the noise, and missing-feature theory is applied to refine the compensation by ignoring noise variation outside the given training conditions, thereby reducing the training and testing mismatch. This paper is focused on several issues relating to the implementation of the new model for real-world applications. These include the generation of multicondition training data to model noisy speech, the combination of different training data to optimize the recognition performance, and the reduction of the model's complexity. The new algorithm was tested using two databases with simulated and realistic noisy speech data. The first database is a redevelopment of the TIMIT database by rerecording the data in the presence of various noise types, used to test the model for speaker identification with a focus on the varieties of noise. The second database is a handheld-device database collected in realistic noisy conditions, used to further validate the model for real-world speaker verification. The new model is compared to baseline systems and is found to achieve lower error rates. Ji Ming, Timothy J. Hazen, James R. Glass, Douglas A. Reynolds |
IEEE Trans. Speech Audio Process. | 3 |
| 2006 | Style & Topic Language Model Adaptation Using HMM-LDA
Bo-June Paul Hsu, James R. Glass |
EMNLP | 2 |
| 2006 | Flexible Multi-Stream Framework for Speech Recognition using Multi-Tape Finite-State TransducersabstractWe present an approach to general multi-stream recognition utilizing multi-tape finite-state transducers (FSTs). The approach is novel in that each of the multiple "streams'" of features can represent either a sequence (e.g., fixed- or variable-rate frames) or a directed acyclic graph (e.g., containing hypothesized phonetic segmentations). Each transition of the multi-tape FST specifies the models to be applied to each stream and the degree of feature stream asynchrony to be allowed. We show how this framework can easily represent the 2-stream variable-rate landmark and segment modeling utilised by our baseline SUMMIT speech recognizer. We present experiments merging standard hidden Markov models (HMMs) with landmark models on the Wall Street Journal speech recognition task, and find that some degree of asynchrony can be critical when combining different types of models. We also present experiments performing audio-visual speech recognition on the AV-TIMIT task I. Lee Hetherington, Han Shu, James R. Glass |
ICASSP (1) | 3 |
| 2006 | Speaker Verification Over Handheld Devices with Realistic Noisy Speech DataabstractWe study speaker verification for handheld devices assuming realistic, noisy test conditions and assuming no prior knowledge of the noise characteristics. Data were recorded in office ("quiet") and street intersection ("noisy") environments, with the use of an internal microphone and an external headset. We assume that the speaker models are trained using the office data and tested in matched and mismatched environment/microphone conditions. Two approaches were studied, both built upon a subband feature framework: 1) a posterior union model (PUM) that focuses verification on matching subbands thereby reducing the effect of the training and testing mismatch, and 2) universal compensation (UC) that combines multi-condition training and the PUM to provide robustness to noises of arbitrary temporal-spectral characteristics. Multi-condition training using simulated noise data of different characteristics provides a "coarse" compensation for the noise, and the PUM refines the compensation by ignoring noise variations outside the given training conditions. These two models were compared to baseline systems and have shown improved robustness for realistic noisy speech data. Ji Ming, Timothy J. Hazen, James R. Glass |
ICASSP (1) | 3 |
| 2006 | Unsupervised Word Acquisition from Speech using Pattern DiscoveryabstractIn this paper, we present an unsupervised method for automatically discovering words from speech using a combination of acoustic pattern discovery, graph clustering, and baseform searching. The algorithm we propose represents an alternative to traditional methods of speech recognition and makes use of the acoustic similarity of multiple realizations of the same words or phrases. On a set of three academic lectures on different subjects, we show that the clustering component of the algorithm is able to successfully generate word clusters that have good coverage of subject-relevant words. Moreover, we illustrate how to use the cluster nodes to retrieve the word identity of each cluster from a large baseform dictionary. Results indicate that this algorithm may prove useful for applications such as vocabulary initialization, speech summarization, or augmentation of existing recognition systems Alex Park 0001, James R. Glass |
ICASSP (1) | 2 |
| 2006 | Combining missing-feature theory, speech enhancement and speaker-dependent/-independent modeling for speech separationabstractThis paper considers the separation and recognition of overlapped speech sentences assuming single-channel observation. A system based on a combination of several different techniques is proposed. The system uses a missing-feature approach for improving crosstalk/noise robustness, a Wiener filter for speech enhancement, hidden Markov models for speech reconstruction, and speaker-dependent/-independent modeling for speaker and speech recognition. We develop the system on the Speech Separation Challenge database, involving a task of separating and recognizing two mixing sentences without assuming advanced knowledge about the identity of the speakers nor about the signal-to-noise ratio. The paper is an extended version of a previous conference paper submitted for the challenge. 1 1 Ji Ming, Timothy J. Hazen, James R. Glass |
INTERSPEECH | 3 |
| 2006 | A Novel DTW-Based Distance Measure for speaker SegmentationabstractWe present a novel distance measure for comparing two speech segments that uses a local version of the well-known DTW algorithm. Our approach is based on the idea of finding word-level speech patterns that are repeated by the same speaker. Using this distance measure, we develop a speaker segmentation procedure and apply it to the task of segmenting multi-speaker lectures. We demonstrate that our approach is able to generate segmentations that correlate well to independently generated human segmentations. In experiments performed on over ten hours of multi-speaker lecture data, we were able to find speaker change points with precision and recall rates of 80% and 100%, respectively. Alex Park 0001, James R. Glass |
SLT | 2 |
| 2005 | A Wavelet and Filter Bank Framework For Phonetic ClassificationabstractWe present a wavelet and filter bank framework for context-independent phonetic classification with the aim of extending the work towards a full speech recognition system. The framework addresses the limitations of the Fourier analysis stage commonly used for short-time spectral representation of speech signals. Also, previous research pertaining to wavelet analysis for speech processing mostly makes use of off-the-shelf wavelets and dyadic-based signal decomposition. Our framework provides more flexibility by taking advantage of the relationship between wavelet transforms and filter banks, and using two filter design techniques as well as 'rational' wavelets. On the standard 39 phone TIMIT classification task, we achieve 22.9% error rate on the core test set using rational filter banks and 4-fold aggregation. This is improved to 18.5% when combined with multiple classifiers defined over non-wavelet acoustic measurements. Ghinwa F. Choueiter, James R. Glass |
ICASSP (1) | 2 |
| 2005 | Automatic Processing of Audio Lectures for Information Retrieval: Vocabulary Selection and Language ModelingabstractThis paper describes our initial progress towards developing a system for automatically transcribing and indexing audio-visual academic lectures for audio information retrieval. We investigate the problem of how to combine generic spoken data sources with subject-specific text sources for processing lecture speech. In addition to word recognition experiments, we perform audio information retrieval simulations to characterize retrieval performance when using errorful automatic transcriptions. Given an appropriately selected vocabulary, we observe that good retrieval performance can be obtained even with high recognition error rates. For language model training, we observe that the addition of spontaneous speech data to subject-specific written material results in more accurate transcriptions, but has a marginal effect on retrieval performance. 1. Alex Park 0001, Timothy J. Hazen, James R. Glass |
ICASSP (1) | 3 |
| 2005 | Production domain modeling of pronunciation for visual speech recognitionabstractArticulatory feature models have been proposed in the automatic speech recognition community as an alternative to phone-based models of speech. In this paper, we extend this approach to the visual modality. Specifically, we adapt a recently proposed feature-based model of pronunciation variation to visual speech recognition (VSR) using a set of visually-salient features. The model uses a dynamic Bayesian network (DBN) to represent the evolution of the feature streams. A bank of SVM feature classifiers, with outputs converted to likelihoods, provides input to the DBN. We present preliminary experiments on an isolated-word VSR task, comparing feature-based and viseme-based units and studying the effects of modeling inter-feature asynchrony. Kate Saenko, Karen Livescu, James R. Glass, Trevor Darrell |
ICASSP (5) | 3 |
| 2005 | Visual Speech Recognition with Loosely Synchronized Feature StreamsabstractWe present an approach to detecting and recognizing spoken isolated phrases based solely on visual input. We adopt an architecture that first employs discriminative detection of visual speech and articulate features, and then performs recognition using a model that accounts for the loose synchronization of the feature streams. Discriminative classifiers detect the subclass of lip appearance corresponding to the presence of speech, and further decompose it into features corresponding to the physical components of articulate production. These components often evolve in a semi-independent fashion, and conventional viseme-based approaches to recognition fail to capture the resulting co-articulation effects. We present a novel dynamic Bayesian network with a multi-stream structure and observations consisting of articulate feature classifier scores, which can model varying degrees of co-articulation in a principled way. We evaluate our visual-only recognition system on a command utterance task. We show comparative results on lip detection and speech/non-speech classification, as well as recognition performance against several baseline systems Kate Saenko, Karen Livescu, Michael Siracusa, Kevin W. Wilson, James R. Glass, Trevor Darrell |
ICCV | 5 |
| 2005 | Morphing spectral envelopes using audio flowabstractWe present a method for morphing between smooth spectral magnitude envelopes of speech. An important element of our method is the notion of audio flow, which is inspired by similar notions of optical flow computed between images in computer vision applications. Audio flow defines the correspondence between two smooth spectral magnitude envelopes, and encodes the formant shifting that occurs from one sound to another. We present several algorithms for the automatic computation of audio flow from a small 20 second corpus of speech. In addition, we present an algorithm for morphing smoothly between any two spectral magnitude envelopes, given the computed audio flow between them. Tony Ezzat, Ethan Meyers, James R. Glass, Tomaso A. Poggio |
INTERSPEECH | 3 |
| 2005 | Robust detection of sonorant landmarksabstractA sonorant detection scheme using Mel-frequency cepstral coefficients and support vector machines (SVMs) is presented and tested in a variety of noise conditions. Adapting the classifier threshold using an estimate of the noise level is used to bias the classifier to effectively compensate for mismatched training and testing conditions. The adaptive threshold classifier achieves low frame error rates using only clean training data without requiring specifically designed features or learning algorithms. The frame-by-frame SVM output is analyzed over longer time periods to uncover temporal modulations related to syllable structure which may aid in landmark-based speech recognition and speech detection. Appropriate filtering of this signal leads to a representation which is stable over a wide range of noise conditions. Using the smoothed output for landmark detection results in a high precision rate, enabling confident pruning of the search-space used by landmark-based speech recognizers. Ken Schutte, James R. Glass |
INTERSPEECH | 2 |
| 2004 | A segment-based audio-visual speech recognizer: data collection, development, and initial experimentsabstractThis paper presents the development and evaluation of a speaker-independent audio-visual speech recognition (AVSR) system that utilizes a segment-based modeling strategy. To support this research, we have collected a new video corpus, called Audio-Visual TIMIT (AV-TIMIT), which consists of 4 total hours of read speech collected from 223 different speakers. This new corpus was used to evaluate our new AVSR system which incorporates a novel audio-visual integration scheme using segment-constrained Hidden Markov Models (HMMs). Preliminary experiments have demonstrated improvements in phonetic recognition performance when incorporating visual information into the speech recognition process. Timothy J. Hazen, Kate Saenko, Chia-Hao La, James R. Glass |
ICMI | 4 |
| 2004 | Articulatory features for robust visual speech recognitionabstractVisual information has been shown to improve the performance of speech recognition systems in noisy acoustic environments. However, most audio-visual speech recognizers rely on a clean visual signal. In this paper, we explore a novel approach to visual speech modeling, based on articulatory features, which has potential benefits under visually challenging conditions. The idea is to use a set of parallel classifiers to extract different articulatory attributes from the input images, and then combine their decisions to obtain higher-level units, such as visemes or words. We evaluate our approach in a preliminary experiment on a small audio-visual database, using several image noise conditions, and compare it to the standard viseme-based modeling approach. Kate Saenko, Trevor Darrell, James R. Glass |
ICMI | 3 |
| 2004 | Feature-based pronunciation modeling with trainable asynchrony probabilities
Karen Livescu, James R. Glass |
INTERSPEECH | 2 |
| 2003 | Hidden feature models for speech recognition using dynamic Bayesian networksabstractIn this paper, we investigate the use of dynamic Bayesian networks (DBNs) to explicitly represent models of hidden features, such as articulatory or other phonological features, for automatic speech recognition. In previous work using the idea of hidden features, the representation has typically been implicit, relying on a single hidden state to represent a combination of features. We present a class of DBN-based hidden feature models, and show that such a representation can be not only more expressive but also more parsimonious. We also describe a way of representing the acoustic observation model with fewer distributions using a product of models, each corresponding to a subset of the features. Finally, we describe our recent experiments using hidden feature models on the Aurora 2.0 corpus. 1. Karen Livescu, James R. Glass, Jeff A. Bilmes |
INTERSPEECH | 2 |
| 2003 | A probabilistic framework for segment-based speech recognition
James R. Glass |
Comput. Speech Lang. | 1 |
| 2002 | A multi-class approach for modelling out-of-vocabulary wordsabstractIn this paper we present a multi-class extension to our approach for modelling out-of-vocabulary (OOV) words [1]. Instead of augmenting the word search space with a single OOV model, we add several OOV models, one for each class of words. We present two approaches for designing the OOV word classes. The first approach relies on using common part-of-speech tags. The second approach is a data-driven two-step clustering procedure, where the first step uses agglomerative clustering to derive an initial class assignment, while the second step uses iterative clustering to move words from one class to another in order to reduce the model perplexity. We present experiments within the JUPITER weather information domain. Results show that the multi-class model significantly improves performance over using a single OOV class. For an OOV detection rate of 70%, the false alarm rate is reduced from 5.3% for a single class to 2.9% for an eight-class model. Issam Bazzi, James R. Glass |
INTERSPEECH | 2 |
| 2002 | Information-theoretic criteria for unit selection synthesisabstractIn our recent work on concatenative speech synthesis, we have devised an efficient, graph-based search to perform unit selection given symbolic information. By encapsulating concatenation and substitution costs defined at the class level, the graph expands only linearly with respect to corpus size. To date, these costs were manually tuned over pre-specified classes, which was a knowledgeintensive engineering process. In this research paper, we turn to information-theoretic metrics for automatically learning the costs from data. These costs can be analyzed in a minimum description length (MDL) framework. The performance of these automatically determined weights is compared against that of manually tuned weights in a perceptual evaluation. Jon R. W. Yi, James R. Glass |
INTERSPEECH | 2 |
| 2001 | Learning units for domain-independent out-of- vocabulary word modellingabstractThis paper describes our recent work on detecting and recognizing out-of-vocabulary (OOV) words for robust speech recognition and understanding. To allow for OOV recognition within a word-based recognizer, the in-vocabulary (IV) word network is augmented with an OOV word model so that OOV words are considered simultaneously with IV words during recognition. We explore several configurations for the OOV model, the best of which utilizes a set of domain-independent, automatically derived, variable-length units. The units are created using an iterative bottom-up procedure where, at each iteration, the unit pairs with maximum mutual information are merged. When evaluating this method on a weather information domain, the false alarm rate of our baseline OOV model [1] is reduced by over 60%. For example, with an OOV detection rate of 70%, the OOV false alarm rate is reduced from 8.5% to 3.2%. At these settings the addition of the OOV model degrades the word error rate on IV data by only 0.3% absolute (3% relative). 1. Issam Bazzi, James R. Glass |
INTERSPEECH | 2 |
| 2001 | Speechbuilder: facilitating spoken dialogue system developmentabstractAbstract. In this paper we report our attempts to facilitate the creation of mixed-initiative spoken dialogue systems for both novice and experienced developers of human language technology. Our efforts have resulted in the creation of a utility called SpeechBuilder, whichallows developers to specify linguistic information about their domains, and rapidly create spoken dialogue interfaces to them. SpeechBuilder has been used to create domains providing access to structured information contained in a relational database, as well as toprovide human language interfaces to control or transaction-based applications. 1 James R. Glass, Eugene Weinstein |
INTERSPEECH | 1 |
| 2001 | Segment-based recognition on the phonebook task: initial results and observations on duration modelingabstractThis paper describes preliminary recognition experiments on PhoneBook [1], a corpus of isolated, telephone-bandwidth, read words from a large (almost 8,000-word) vocabulary. We have chosen this corpus as a testbed for experiments on the language model-independent parts of a segment-based recognizer. We present results showing that a segment-based recognizer performs well on this task, and that a simple Gaussian mixture phone duration model significantly reduces the error rate. We compare context-independent, stress-dependent, and word position-dependent duration models and obtain relative error rate reductions of up to 12% on the test set. Finally, we make some observations regarding the effects of stress and word position in this isolated-word task and discuss our plans for further research using PhoneBook. 1. Karen Livescu, James R. Glass |
INTERSPEECH | 2 |
| 2001 | Mokusei: a telephone-based Japanese conversational system in the weather domainabstractThis paper describes MOKUSEI, an end-to-end Japanese version of our JUPITER weather information system. MOKUSEI delivers weather information over the phone through natural conversation with the user. For the most part, MOKUSEI uses the same components for recognition, understanding, and generation that JUPITER uses, and the database and the semantic frames for the weather information content are also shared. However, MOKUSEI motivated us to redesign our GENESIS generation system, in order to improve the quality of translations of weather reports into Japanese. We also had to develop new ways to transcribe user utterances through morphological analysis. MOKU- SEI is fully functional and has already been used for data collection with about 700 naive users. These data have been used for improvement and evaluation of MOKUSEI. This paper also presents the result of evaluating the current version of MOKUSEI. Mikio Nakano, Yasuhiro Minami, Stephanie Seneff, Timothy J. Hazen, D. Scott Cyphers, James R. Glass, Joseph Polifroni, Victor Zue |
INTERSPEECH | 6 |
| 2000 | Heterogeneous lexical units for automatic speech recognition: preliminary investigationsabstractThis paper explores the use of the phone and syllable as primary units of representation in the first stage of a two-stage recognizer. A finite-state transducer speech recognizer is utilized to configure the recognition as a two-stage process, where either phone or syllable graphs are computed in the first stage, and passed to the second stage to determine the most likely word hypotheses. Preliminary experiments in a weather information speech understanding domain show that a syllable representation with either bigram or trigram language models provides more constraint than a phonetic representation with a higher-order n-gram language model (up to a 6-gram), and approaches the performance of a more conventional single-stage word-based configuration. Issam Bazzi, James R. Glass |
ICASSP | 2 |
| 2000 | Lexical modeling of non-native speech for automatic speech recognitionabstractThe paper examines the recognition of non-native speech in JUPITER, a speaker-independent, spontaneous-speech conversational system. Because the non-native speech in this domain is limited and varied, speaker- and accent-specific methods are impractical. We therefore chose to model all of the non-native data with a single model. In particular, the paper describes an attempt to better model non-native lexical patterns. These patterns are incorporated by applying context-independent phonetic confusion rules, whose probabilities are estimated from training data. Using this approach, the word error rate on a non-native test set is reduced from 20.9% to 18.8%. Karen Livescu, James R. Glass |
ICASSP | 2 |
| 2000 | Modeling out-of-vocabulary words for robust speech recognitionabstractIn this paper we present an approach for modeling and recognizing out-of-vocabulary (OOV) words in a single stage recognizer. A word-based recognizer is augmented with an extra OOV word model, which enables the OOV word to be predicted by a wordbased language model. The OOV model itself is phone-based, so that an OOV word can be realized as an arbitrary sequence of phones. A phone bigram is used to provide phonotactic constraints within the OOV model. A recognizer with this configuration can recognize words in the original vocabulary as well as any potential new words of arbitrary pronunciation. In our preliminary investigation of this framework, we have evaluated the recognizer on a weather information domain with one test set containing only in-vocabulary (IV) data, and another containing OOV words. On the IV test set, the recognizer had an OOV insertion rate of only 1.3%, and degraded the baseline WER from 10.4% to 10.7%. On the OOV test set, the recognizer was able to detect nearly... Issam Bazzi, James R. Glass |
INTERSPEECH | 2 |
| 2000 | Data collection and performance evaluation of spoken dialogue systems: the MIT experienceabstractIn this paper we report our efforts in data collection and performance evaluation in support of spoken dialogue system development. We describe two understanding metrics called query density and concept efficiency which can be interpreted on a perutterance basis, but which are measured over the course of a dialogue. We also describe the evaluation infrastructure we have developed to support off-line data processing using our GALAXY client-server architecture [8]. We show how we have used these metrics and mechanisms as part of the development of a spoken dialogue system for air-travel information. James R. Glass, Joseph Polifroni, Stephanie Seneff, Victor Zue |
INTERSPEECH | 1 |
| 2000 | A flexible, scalable finite-state transducer architecture for corpus-based concatenative speech synthesisabstractIn this paper we describe our work involving the conversion of our phonologically-based synthesizer into a finite-state transducer (FST) representation which can be used for real-time natural-sounding synthesis. We have designed a transducer structure to efficiently perform the common task of unit selection in concatenative speech synthesis. By encapsulating domainindependent concatenative synthesis costs into a constraint kernel, we have obtained a topology that scales linearly with the size of the synthesis corpus. The FST representation provides a flexible, unified framework in which we can leverage our previous work in speech recognition in areas such as pronunciation modelling and search. The FST synthesizer has been incorporated into two servers which operate within our conversational system architecture to convert meaning representations into waveforms. We have had preliminary success with the new FST-based synthesis in several constrained spoken dialogue applications. 1. INTRO... Jon R. W. Yi, James R. Glass, I. Lee Hetherington |
INTERSPEECH | 2 |
| 2000 | Conversational interfaces: advances and challengesabstractThe past decade has witnessed the emergence of a new breed of human-computer interfaces that combines several human language technologies to enable humans to converse with computers using spoken dialogue for information access, creation and processing. In this paper, we introduce the nature of these conversational interfaces and describe the underlying human language technologies on which they are based. After summarizing some of the recent progress in this area around the world, we discuss development issues faced by researchers creating these kinds of systems and present some of the ongoing and unmet research challenges in this field. Victor Zue, James R. Glass |
Proc. IEEE | 2 |
| 2000 | Guest editorial introduction to the special issue on language modeling and dialogue systems
James R. Glass, Ronald Rosenfeld |
IEEE Trans. Speech Audio Process. | 1 |
| 2000 | JUPlTER: a telephone-based conversational interface for weather informationabstractIn early 1997, our group initiated a project to develop JUPITER, a conversational interface that allows users to obtain worldwide weather forecast information over the telephone using spoken dialogue. It has served as the primary research platform for our group on many issues related to human language technology, including telephone-based speech recognition, robust language understanding, language generation, dialogue modeling, and multilingual interfaces. Over a two year period since coming online in May 1997, JUPITER has received, via a toll-free number in North America, over 30000 calls (totaling over 180000 utterances), mostly from naive users. The purpose of this paper is to describe our development effort in terms of the underlying human language technologies as well as other system-related issues such as utterance rejection and content harvesting. We also present some evaluation results on the system and its components. Victor Zue, Stephanie Seneff, James R. Glass, Joseph Polifroni, Christine Pao, Timothy J. Hazen, I. Lee Hetherington |
IEEE Trans. Speech Audio Process. | 3 |
| 1999 | Real-time telephone-based speech recognition in the Jupiter domainabstractThis paper describes our experiences with developing a real-time telephone-based speech recognizer as part of a conversational system in the weather information domain. This system has been used to collect spontaneous speech data which has proven to be extremely valuable for research in a number of different areas. After describing the corpus we have collected, we describe the development of the recognizer vocabulary, pronunciations, language and acoustic models for this system, the new weighted finite-state transducer-based lexical access component, and report on the current performance of the recognizer under several different conditions. We also analyze recognition latency to verify that the system performs in real-time. James R. Glass, Timothy J. Hazen, I. Lee Hetherington |
ICASSP | 1 |
| 1998 | Telephone-based conversational speech recognition in the JUPITER domainabstractThis paper describes our experiences with developing a telephone-based speech recognizer as part of a conversational system in the weather information domain. This system has been used to collect spontaneous speech data which has proven to be extremely valuable for research in a number of different areas. After describing the corpus we have collected, we describe the development of the recognizer vocabulary, pronunciations, language and acoustic models for this system, and report on the current performance of the recognizer under several different conditions. 1. INTRODUCTION Over the past year and a half, we have developed a telephonebased, weather information system called JUPITER [11], which is available via a toll-free number for users to query a relational database of current weather conditions using natural, conversational speech 2 . Using information obtained from several different internet sites, JUPITER can provide weather forecasts for approximately 500 cities around the ... James R. Glass, Timothy J. Hazen |
ICSLP | 1 |
| 1998 | Heterogeneous measurements and multiple classifiers for speech recognitionabstractThis paper addresses the problem of acoustic phonetic modeling. First, heterogeneous acoustic measurements are chosen in order to maximize the acoustic-phonetic information extracted from the speech signal in preprocessing. Second, classifier systems are presented for successfully utilizing high-dimensional acoustic measurement spaces. The techniques used for achieving these two goals can be broadly categorized as hierarchical, committeebased, or a hybrid of these two. This paper presents committeebased and hybrid approaches. In context-independent classification and context-dependent recognition on the TIMIT core test set using 39 classes, the system achieved error rates of 18.3% and 24.4%, respectively. These error rates are the lowest we have seen reported on these tasks. In addition, experiments with a telephone-based weather information word recognition task led to word error rate reductions of 10–16%. Andrew K. Halberstadt, James R. Glass |
ICSLP | 2 |
| 1998 | Real-time probabilistic segmentation for segment-based speech recognitionabstractIn this work, we investigate modifications to a probabilistic segmentation algorithm to achieve a real-time, and pipelined capability for our segment-based speech recognizer [4]. The existing algorithm used a Viterbi and backwards A search to hypothesize phonetic segments [2]. We were able to reduce the computational requirements of this algorithm by reducing the effective search space to acoustic landmarks, and were able to achieve pipelined capability by executing theA search in blocks defined by reliably detected phonetic boundaries. The new algorithm produces 30% fewer segments, and improves TIMIT phonetic recognition performance by 2.4% over an acoustic segmentation baseline. We were also able to produce 30% fewer segments on a word recognition task in a weather information domain [11]. Steven C. Lee, James R. Glass |
ICSLP | 2 |
| 1998 | Confidence scoring for speech understanding systemsabstractThis research investigates the use of utterance-level features for confidence scoring. Confidence scores are used to accept or reject user utterances in our conversational weather information system [10]. We have developed an automatic labeling algorithm based on a semantic frame comparison between recognized and transcribed orthographies. We explore recognition-based features along with semantic, linguistic, and application-specific features for utterance rejection. Discriminant analysis is used in an iterative process to select the best set of classification features for our utterance rejection sub-system. Experiments show that we can correctly reject over 60% of incorrectly understood utterances while accepting 98% of all correctly understood utterances. Christine Pao, Philipp Schmid 0001, James R. Glass |
ICSLP | 3 |
| 1998 | Natural-sounding speech synthesis using variable-length unitsabstractThe goal of this work was to develop a speech synthesis system which concatenates variable-length units to create naturalsounding speech. Our initial work in this area showed that by careful design of system responses to ensure consistent intonation contours, natural-sounding speech synthesis was achievable with word- and phrase-level concatenation. In order to extend the flexibility of this framework, we focused on the problem of generating novel words from a corpus of sub-word units. The design of the sub-word units was motivated by perceptual studies that investigated where speech could be spliced with minimal audible distortion and what contextual constraints were necessary to maintain in order to produce natural sounding speech. The sub-word corpus is searched during synthesis using a Viterbi search which selects a sequence of units based on how well they individually match the input specification and on how well they sound as an ensemble. This concatenative speech synthesis system, ENVOICE, has been used in a conversational information retrieval system in two application domains to convert meaning representations into speech waveforms. Jon R. W. Yi, James R. Glass |
ICSLP | 2 |
| 1998 | Evaluation methodology for a telephone-based conversational system
Joseph Polifroni, Stephanie Seneff, James R. Glass, Timothy J. Hazen |
LREC | 3 |
| 1997 | Segmentation and modeling in segment-based recognitionabstractRecently, we have developed a probabilistic framework for segment-based speech recognition that represents the speech signal as a network of segments and associated feature vectors [2]. Although in general, each path through the network does not traverse all segments, we ar-gued that each path must account for all feature vectors in the network. We then demonstrated an efficient search algorithm that uses a single additional model to account for segments that are not traversed. In this paper, we present two new extensions to our framework. First, we replace our acoustic segmentation algorithm with “segmentation by recognition, ” a probabilistic algorithm that can combine multiple contextual constraints towards hypothesizing only the most likely seg-ments. Second, we generalize our framework to “near-miss modeling” and describe a search algorithm that can efficiently use multiple mod-els to enforce contextual constraints across all segments in a network. We report experiments in phonetic recognition on the TIMIT corpus in which we achieve a diphone context-dependent error rate of 26.6% on the NIST core test set over 39 classes. This is a 12.8 % reduction in error rate from our best previously reported result. 1. Jane W. Chang, James R. Glass |
EUROSPEECH | 2 |
| 1997 | Heterogeneous acoustic measurements for phonetic classification 1abstractIn this paper we describe our recent efforts to improve acousticphonetic modeling by developing sets of heterogeneous, phoneclass -specific measurements, and combining these diverse measurements into a probabilistic classification framework. We first describe a baseline classifier using homogeneous measurements. After comparing selected sub-tasks to known human performance, we define sets of phone-class-specific measurements which improve within-class classification performance. Subsequently, we combine these heterogeneous measurements into an overall context-independent classification framework. We report on a series of phonetic classification experiments using the TIMIT acoustic-phonetic corpus. Our overall framework achieves 79.0% accuracy on the NIST core test set. 1. INTRODUCTION Over the past several years, our group has pursued a segmentbased approach to speech recognition. One of the potential advantages of this approach over conventional frame-based methods is that it provides... Andrew K. Halberstadt, James R. Glass |
EUROSPEECH | 2 |
| 1997 | A comparison of novel techniques for instantaneous speaker adaptationabstractThis paper introduces two novel techniques for instantaneous speaker adaptation, reference speaker weighting and consistency modeling. An approach to hierarchical speaker clustering using gender and speaking rate as the clustering criteria is also presented. All three methods attempt to utilize the underlying within-speaker correlations that are present between the acoustic realizations of different phones. By accounting for these correlations a limited amount of adaptation data can be used to adapt the models of every phonetic acoustic model including those for phones which havenot been observed in the adaptation data. In instantaneous adaptation experiments using the DARPA Resource Management corpus, a reduction in word error rate of 20 % has been achieved using a combination of these new techniques. Timothy J. Hazen, James R. Glass |
EUROSPEECH | 2 |
| 1997 | MUSE: a scripting language for the development of interactive speech analysis and recognition toolsabstractSpeech research is a complex endeavor, as reflected in the numerous tools and specialized languages the modern researcher needs to learn. These tools, while adequate for what they have been designed for, are difficult to customize or extend in new directions, even though this is often required. We feel this situation can be improved and propose a new scripting language, MUSE, designed explicitly for speech research, in order to facilitate exploration of new ideas. MUSE is designed to support many modes of research from interactive speech analysis through compute-intensive speech understanding systems, and has facilities for automating some of the more difficult requirements of speech tools: user interactivity, distributed computation, and caching. In this paper we describe the design of the MUSE language and our current prototype MUSE interpreter. Michael K. McCandless, James R. Glass |
EUROSPEECH | 2 |
| 1997 | YINHE: a Mandarin Chinese version of the GALAXY systemabstractThe galaxy system is a human-computer conversational system providing a spoken language interface for accessing on-line information. It was initially implemented for English in travel-related domains, including air travel, local city navigation, and weather. We began an effort to develop multilingual systems within the framework of galaxy several years ago. This paper describes our recent work on porting the system to Mandarin Chinese, including speech recognition, language understanding, and language generation components. Overall, the system produced reasonable responses nearly 70% of the time for spontaneous test data collected in a wizard environment. 1. INTRODUCTION The galaxy system is a client/server architecture for computer conversational systems [1]. In designing galaxy, we drew heavily on experience gained in the development of galaxy's predecessor, voyager [2]. Voyager was not initially designed to easily support multiple languages, but through a trial-and-error process... Chao Wang 0018, James R. Glass, Helen M. Meng, Joseph Polifroni, Stephanie Seneff, Victor Zue |
EUROSPEECH | 2 |
| 1997 | From interface to content: translingual access and delivery of on-line information
Victor Zue, Stephanie Seneff, James R. Glass, I. Lee Hetherington, Edward Hurley, Helen M. Meng, Christine Pao, Joseph Polifroni, Rafael Schloming, Philipp Schmid 0001 |
EUROSPEECH | 3 |
| 1996 | A probabilistic framework for feature-based speech recognitionabstractMost current speech recognizers use an observation space which is based on a temporal sequence of "frames" (e.g., Mel-cepstra).There is another class of recognizer which further processes these frames to produce a segment-based network, and represents each segment by fixed-dimensional "features."In such feature-based recognizers the observation space takes the form of a temporal network of feature vectors, so that a single segmentation of an utterance will use a subset of all possible feature vectors.In this work we examine a maximum a posteriori decoding strategy for feature-based recognizers and develop a normalization criterion useful for a segmentbased Viterbi or A search.We report experimental results for the task of phonetic recognition on the TIMIT corpus where we achieved context-independent and context-dependent (using diphones) results on the core test set of 64.1% and 69.5% respectively. James R. Glass, Jane W. Chang, Michael K. McCandless |
ICSLP | 1 |
| 1996 | Telephone data collection using the world wide webabstractOver the past year our group has begun development of telephonebased speech understanding capability for our GALAXY conversational system.An important part of this process has been the collection of telephone speech which was used for training and evaluation.In the first phase of data collection our goal was to collect read speech from a wide variety of talkers, telephone handsets, and noise/channel conditions.In the second phase of data collection our additional goal was to collect spontaneous telephone speech from subjects actually using the system.In order to maximize variation in telephone conditions, as well as ease of use for subjects, the data collection software was designed to telephone subjects at their specified phone numbers around North America.Subjects initiate the data collection session by submitting an electronic form accessible by a WWW browser.For read speech collection, a set of prompts is automatically generated for the subject.This paper describes the design of the data collection system we are using for these purposes.To date we have collected over 9,000 utterances from over 270 subjects. Edward Hurley, Joseph Polifroni, James R. Glass |
ICSLP | 3 |
| 1996 | WHEELS: a conversational system in the automobile classifieds domain
Helen M. Meng, Senis Busayapongchai, James R. Glass, David Goddeau, I. Lee Hetherington, Edward Hurley, Christine Pao, Joseph Polifroni, Stephanie Seneff, Victor Zue |
ICSLP | 3 |
| 1996 | Multilingual human-computer interactions: from information access to language learning
Victor Zue, Stephanie Seneff, Joseph Polifroni, Helen M. Meng, James R. Glass |
ICSLP | 5 |
| 1995 | Multilingual spoken-language understanding in the MIT Voyager system
James R. Glass, Giovanni Flammia, David Goodine, Michael S. Phillips 0001, Joseph Polifroni, Shinsuke Sakai, Stephanie Seneff, Victor Zue |
Speech Commun. | 1 |
| 1994 | Porting the bilingual voyager system to Italian
Giovanni Flammia, James R. Glass, Michael S. Phillips 0001, Joseph Polifroni, Stephanie Seneff, Victor Zue |
ICSLP | 2 |
| 1994 | Multilingual language generation across multiple domains
James R. Glass, Joseph Polifroni, Stephanie Seneff |
ICSLP | 1 |
| 1994 | GALAXY: a human-language interface to on-line travel information
David Goddeau, Eric Brill, James R. Glass, Christine Pao, Michael S. Phillips 0001, Joseph Polifroni, Stephanie Seneff, Victor Zue |
ICSLP | 3 |
| 1994 | Statistical trajectory models for phonetic recognition
William Goldenthal, James R. Glass |
ICSLP | 2 |
| 1994 | Empirical acquisition of language models for speech recognition
Michael K. McCandless, James R. Glass |
ICSLP | 2 |
| 1994 | PEGASUS: A spoken dialogue interface for on-line air travel planning
Victor Zue, Stephanie Seneff, Joseph Polifroni, Michael S. Phillips 0001, Christine Pao, David Goodine, David Goddeau, James R. Glass |
Speech Communication | 8 |
| 1993 | A comparative study of signal representations and classification techniques for speech recognition
Hong C. Leung, Benjamin Chigier, James R. Glass |
ICASSP (2) | 3 |
| 1993 | A bilingual Voyager system
James R. Glass, David Goodine, Michael S. Phillips 0001, Shinsuke Sakai, Stephanie Seneff, Victor Zue |
EUROSPEECH | 1 |
| 1993 | Modelling spectral dynamics for vowel classification
William Goldenthal, James R. Glass |
EUROSPEECH | 2 |
| 1993 | A* word network search for continuous speech recognition
I. Lee Hetherington, Michael S. Phillips 0001, James R. Glass, Victor Zue |
EUROSPEECH | 3 |
| 1993 | Empirical acquisition of word and phrase classes in the atis domain
Michael K. McCandless, James R. Glass |
EUROSPEECH | 2 |
| 1992 | Vowel classification based on analysis-by-synthesisabstractIn this paper, we report on a sequence of experiments designed to explore the use of analysis-by-synthesis methods for speech recognition and speech analysis in general. An intermediate Rolf Carlson, James R. Glass |
ICSLP | 2 |
| 1992 | Collection and analyses of WSJ-CSR corpus at MIT
Michael S. Phillips 0001, James R. Glass, Joseph Polifroni, Victor Zue |
ICSLP | 2 |
| 1991 | Integration of speech recognition and natural language processing in the MIT VOYAGER systemabstractThe MIT VOYAGER speech understanding system is an urban exploration and navigation system that interacts with the user through spoken dialogue, text, and graphics. The authors describe recent attempts at improving the integration between the speech recognition and natural language components. They used the generation capability of the natural language component to produce a word-pair language model to constrain the recognizer's search space, thus improving the coverage of the overall system. They also implemented a strategy in which the recognizer generates the top N word strings and passes them along to the natural language component for filtering. Results on performance evaluation are presented.> Victor Zue, James R. Glass, David Goodine, Hong C. Leung, Michael S. Phillips 0001, Joseph Polifroni, Stephanie Seneff |
ICASSP | 2 |
| 1991 | Automatic learning of lexical representations for sub-word unit based speech recognition systems
Michael S. Phillips 0001, James R. Glass, Victor Zue |
EUROSPEECH | 2 |
| 1991 | The MIT ATIS system; preliminary development, spontaneous speech data collection, and performance evaluation
Victor Zue, James R. Glass, David Goodine, Lynette Hirschman, Hong C. Leung, Michael S. Phillips 0001, Joseph Polifroni, Stephanie Seneff |
EUROSPEECH | 2 |
| 1990 | The VOYAGER speech understanding system: preliminary development and evaluationabstractEarly experience with the development of the MIT VOYAGER spoken language system is described, and its current performance is documented. The three components of VOYAGER, the speech recognition component, the natural language component, and the application back-end, are described.> Victor Zue, James R. Glass, David Goodine, Hong C. Leung, Michael S. Phillips 0001, Joseph Polifroni, Stephanie Seneff |
ICASSP | 2 |
| 1990 | The SUMMIT speech recognition system: phonological modelling and lexical accessabstractThe phonological modeling and lexical access components of the SUMMIT speech recognition system are described in detail. SUMMIT makes explicit use of acoustic-phonetic knowledge, embedded in a segmental framework that can be trained automatically. Performance results for the complete system on the DARPA 1000-word Naval Resource Management task are presented.> Victor Zue, James R. Glass, David Goodine, Michael Philips, Stephanie Seneff |
ICASSP | 2 |
| 1990 | Detection and classification of phonemes using context-independent error back-propagation
Hong C. Leung, James R. Glass, Michael S. Phillips 0001, Victor Zue |
ICSLP | 2 |
| 1990 | Recent progress on the MIT VOYAGER spoken language system
Victor Zue, James R. Glass, David Goddeau, David Goodine, Hong C. Leung, Michael K. McCandless, Michael S. Phillips 0001, Joseph Polifroni, Stephanie Seneff, Dave Whitney |
ICSLP | 2 |
| 1990 | Phonetic Classification and Recognition Using the Multi-Layer Perceptron
Hong C. Leung, James R. Glass, Michael S. Phillips 0001, Victor Zue |
NIPS | 2 |
| 1990 | From Speech Recognition to Spoken Language Understanding
Victor Zue, James R. Glass, David Goodine, Lynette Hirschman, Hong C. Leung, Michael S. Phillips 0001, Joseph Polifroni, Stephanie Seneff |
NIPS | 2 |
| 1990 | Speech database development at MIT: Timit and beyond
Victor Zue, Stephanie Seneff, James R. Glass |
Speech Commun. | 3 |
| 1989 | Acoustic segmentation and phonetic classification in the SUMMIT systemabstractRecently, the authors initiated a project to develop a phonetically-based spoken-language-understanding system called SUMMIT. In contrast to many of the past efforts that make use of heuristic rules whose development requires intense knowledge engineering, their approach attempts to express the speech knowledge within a formal framework using well-defined mathematical tools. In the authors' system, features and decision strategies are discovered and trained automatically, using a large body of speech data. The authors describe those parts of the system dealing with acoustic segmentation and phonetic classification and document its current performance.> Victor Zue, James R. Glass, Michael Philips, Stephanie Seneff |
ICASSP | 2 |
| 1988 | Multi-level acoustic segmentation of continuous speechabstractAs part of the goal to better understand the relationship between the speech signal and the underlying phonemic representation, the authors have developed a procedure that describes the acoustic structure of the signal. Acoustic events are embedded in a multi-level structure in which information ranging from coarse to fine is represented in an organized fashion. An analysis of the acoustic structure, using 500 utterances from 100 different talkers, show that it captures over 96% of the acoustic-phonetic events of interest with an insertion rate of less than 5%. The signal representation, and the algorithms for determining the acoustic segments and the multi-level structure are described. Performance results and a comparison with scale-space filtering is also included. Possible use of this segmental description for automatic speech recognition is discussed.> James R. Glass, Victor Zue |
ICASSP | 1 |
| 1986 | Detection and recognition of nasal consonants in American EnglishabstractThis paper deals with the recognition of nasal consonants, /m, n, η/, in American English. From an acoustic study conducted earlier, parameters found to be useful in distinguishing nasal murmurs from a set of phonemically defined impostors were used in a recognition experiment. The algorithm assumes that the boundaries of the nasals, and their broad phonetic context, are known. Our results, based on 600 sentences spoken by 60 speakers, show that nasal consonants can be distinguished from the impostors with an accuracy of 83.5%. In parallel, a nasal detection algorithm based on a local decision criterion is being developed, using the outputs of an auditory model. Evaluation of its performance on 600 sentences indicates that the nasals are found 96% of the time, with a 2 to 1 impostor-to-nasal ratio. The final step of merging the detection and recognition components has yet to be completed. James R. Glass, Victor Zue |
ICASSP | 1 |
| 1985 | Detection of nasalized vowels in American EnglishabstractThis study is concerned with the quantification of acoustic measures that characterize nasalized vowels. It is motivated by the fact that the detection of nasalized vowels is useful in speech recognition, since regions can be identified in the speech signal where a nasal consonant may be present, and where the vocal tract resonances are distorted. Our study consists of several steps. First, an acoustic study was performed using utterances from a large database in order to propose potential measures of nasality. Next, automatic algorithms were developed to extract these measures, and their utilities were established through examination of a large amount of data. Finally, recognition experiments were performed using these measures. The system detected nasalized vowels with an accuracy of approximately 74%, when tested on one speaker at a time, and trained on the speech of the remaining speakers in the database. James R. Glass, Victor Zue |
ICASSP | 1 |