VLDB 2026 Research / reviewers in the wild / expert
Mark Hasegawa-Johnson
dblp:70/3186 · also Mark A. Hasegawa-Johnson
· DBLP profile ↗
238ranked-venue papers
10as first author
68since 2021 · last 2026
0000-0002-5631-2893ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 186 · 9 first-author · 47 since 2021Artificial intelligence and machine learning · 143 · 5 first-author · 41 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 since 2021Databases, data management, data science and information retrieval · 3Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Fair Speech Recognition: Mitigating Demographic Bias in End-to-End ASR Systems
Maliha Jahan, Thomas Thebaud, Zsuzsanna Fagyal, Jesús Villalba 0001, Mark Hasegawa-Johnson, Laureano Moro-Velázquez, Najim Dehak |
LREC | 5 |
| 2026 | Towards unsupervised speech recognition without pronunciation models
Junrui Ni, Liming Wang 0003, Yang Zhang 0001, Kaizhi Qian, Heting Gao, Mark Hasegawa-Johnson, James R. Glass, Chang Dong Yoo |
Speech Commun. | 6 |
| 2025 | Unveiling Performance Bias in ASR Systems: A Study on Gender, Age, Accent, and MoreabstractWith the recent advancements in speech recognition, it is crucial to ensure these systems are free from performance biases against any speaker subgroups. This study examined the performance of twenty variants of seven Automatic Speech Recognition models across four datasets in English language: L2 Arctic, Speech Accent Archive, CORAAL, and SBCSAE. We employed Poisson regression and drop-in-deviance tests to identify which attributes significantly contribute to the Word Error Rate. Our analysis revealed biases related to attributes such as native language, location, occupation, and birthplace. Most systems did not exhibit bias related to factors like gender and age. Additionally, we conducted an experiment to detect bias related to "variant" (accent and dialect) by combining the CORAAL (African American Vernacular English (AAVE)) and SBCSAE (General American English (GAE)) datasets, aiming to identify the sources of any observed bias. We found that both speaker variability and dialectal difference contribute to observed bias for variant. Maliha Jahan, Priyam Mazumdar, Thomas Thebaud, Mark Hasegawa-Johnson, Jesús Villalba 0001, Najim Dehak, Laureano Moro-Velázquez |
ICASSP | 4 |
| 2025 | Improved Recognition of the Speech of People with Parkinson's Who StutterabstractStuttering is a speech disorder often associated with neurological conditions, including Parkinson’s disease (PD). Despite advancements in modern automatic speech recognition (ASR) technologies, today’s systems still face challenges in accurately recognizing dysarthric speech, particularly when stuttering is present. In this study, we propose a novel stuttered speech data augmentation approach to improve dysarthric speech recognition. We utilize typical speech data from LibriSpeech to generate artificial stuttered speech by applying Voice Activity Detection and Forced Alignment techniques to accurately identify word boundaries, and integrating an adaptive stuttering filter to simulate severe stuttering patterns. Additionally, dysarthric speech data from individuals with PD, collected by the Speech Accessibility Project (SAP), is integrated into the model. Our experimental results demonstrate that the proposed augmentation approach outperforms existing methods in enhancing the recognition of stuttered speech. Furthermore, fine-tuning the ASR systems with SAP data yields additional performance improvements for both stuttering and non-stuttering individuals with PD. Jonghwan Na, Xiuwen Zheng 0003, Bowon Lee, Mark Hasegawa-Johnson |
ICASSP | 4 |
| 2025 | Cohort-Sensitive Labeling: An Effective Approach for Enhancing ASR PerformanceabstractThis paper proposes a cohort-sensitive labeling (CSL) for automatic speech recognition (ASR). CSL is a method that distinguishes data labels based on cohorts, allowing models to learn cohort-specific information. For evaluation, we applied CSL using gender information in the training data of LibriSpeech dataset. Experimental results demonstrate that the CSL-based approach outperforms methods without CSL, given sufficient training data. Specifically, our method achieved average word error rate reduction (WERR) of 1.81% on the LibriSpeech test-clean and 5.76% on test-other datasets, when more than 100 hours of data were used for training. Moreover, on TIMIT and Common Voice test sets, it achieved WERR of up to 11.52% and 2.91%, respectively demonstrating its robustness and generalizability to unseen data. Additionally, the proposed method reached up to 97.21% accuracy in classifying the gender cohort, suggesting that ASR models trained with the CSL effectively leverage the cohort information. Jonghwan Na, Mark Hasegawa-Johnson, Bowon Lee |
ICASSP | 2 |
| 2025 | LIMMITS'25: Multilingual Streaming TTS With Neural Codecs for Indian LanguagesabstractThis work provides a summary of the Multilingual streaming TTS with neural codecs for Indian languages challenge (LIMMITS’25), organized as part of the ICASSP 2025 signal processing grand challenge. Towards this, 278 hours of TTS data in 4 Indian languages - Gujarati, Indian English, Bhojpuri, and Kannada got released. The challenge focuses on advancing research in neural codec-based and streaming TTS systems. The top teams in the challenge attained high subjective scores on naturalness and similarity, thus contributing to the progress in text-to-speech generation systems. Philipp Olbrich, Hema A. Murthy, Pranaw Kumar, Shinji Watanabe 0001, Sheng Zhao 0002, Mark Hasegawa-Johnson |
ICASSP | 6 |
| 2025 | Robust Cross-Etiology and Speaker-Independent Dysarthric Speech RecognitionabstractIn this paper, we present a speaker-independent dysarthric speech recognition system, with a focus on evaluating the recently released Speech Accessibility Project (SAP-1005) dataset, which includes speech data from individuals with Parkinson’s disease (PD). Despite the growing body of research in dysarthric speech recognition, many existing systems are speaker-dependent and adaptive, limiting their generalizability across different speakers and etiologies. Our primary objective is to develop a robust speaker-independent model capable of accurately recognizing dysarthric speech, irrespective of the speaker. Additionally, as a secondary objective, we aim to test the cross-etiology performance of our model by evaluating it on the TORGO dataset, which contains speech samples from individuals with cerebral palsy (CP) and amyotrophic lateral sclerosis (ALS). By leveraging the Whisper model, our speaker-independent system achieved a CER of 6.99% and a WER of 10.71% on the SAP-1005 dataset. Further, in cross-etiology settings, we achieved a CER of 25.08% and a WER of 39.56% on the TORGO dataset. These results highlight the potential of our approach to generalize across unseen speakers and different etiologies of dysarthria. Satwinder Singh, Zihan Zhong, Clarion Mendes, Mark Hasegawa-Johnson, Waleed Abdullah, Seyed Reza Shahamiri |
ICASSP | 5 |
| 2025 | Dysarthric Speech Conformer: Adaptation for Sequence-to-Sequence Dysarthric Speech RecognitionabstractAutomatic Speech Recognition (ASR) holds immense potential to provide an effective interface for assistive technologies, but its performance remains unsatisfactory for people with speech impairments such as dysarthria. Existing ASR systems struggle to accurately recognize dysarthric speech due to the significant speaker variability in dysarthric speech and the scarcity of dysarthric datasets. In this study, we propose a two-phase adaptation pipeline based on the Conformer architecture that leverages typical speech to transfer to individualized ASR models for dysarthric speakers. ASR performance is evaluated for isolated words and continuous sentences, yielding an average Word Error Rate of 21.5% on the UASpeech dataset and 12.7% on the TORGO dataset. Selectively freezing decoder layers was more often successful than selectively freezing encoder layers, suggesting that optimal performance is achieved by focusing the adaptation on the acoustic information contained in the encoder. Zihan Zhong, Satwinder Singh, Clarion Mendes, Mark Hasegawa-Johnson, Waleed Abdullah, Seyed Reza Shahamiri |
ICASSP | 5 |
| 2025 | Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language ModelsabstractIn the broader context of deep learning, Multimodal Large Language Models have achieved significant breakthroughs by leveraging powerful Large Language Models as a backbone to align different modalities into the language space. A prime exemplification is the development of Video Large Language Models (Video-LLMs). While numerous advancements have been proposed to enhance the video understanding capabilities of these models, they are predominantly trained on questions generated directly from video content. However, in real-world scenarios, users often pose questions that extend beyond the informational scope of the video, highlighting the need for Video-LLMs to assess the relevance of the question. We demonstrate that even the best-performing Video-LLMs fail to reject unfit questions-not necessarily due to a lack of video understanding, but because they have not been trained to identify and refuse such questions. To address this limitation, we propose alignment for answerability, a framework that equips Video-LLMs with the ability to evaluate the relevance of a question based on the input video and appropriately decline to answer when the question exceeds the scope of the video, as well as an evaluation framework with a comprehensive set of metrics designed to measure model behavior before and after alignment. Furthermore, we present a pipeline for creating a dataset specifically tailored for alignment for answerability, leveraging existing video-description paired datasets. Eunseop Yoon, Hee Suk Yoon, Mark Hasegawa-Johnson, Chang Dong Yoo |
ICLR | 3 |
| 2025 | ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference OptimizationabstractWe introduce ConfPO, a method for preference learning in Large Language Models (LLMs) that identifies and optimizes preference-critical tokens based solely on the training policy’s confidence, without requiring any auxiliary models or compute. Unlike prior Direct Alignment Algorithms (DAAs) such as Direct Preference Optimization (DPO), which uniformly adjust all token probabilities regardless of their relevance to preference, ConfPO focuses optimization on the most impactful tokens. This targeted approach improves alignment quality while mitigating overoptimization (i.e., reward hacking) by using the KL divergence budget more efficiently. In contrast to recent token-level methods that rely on credit-assignment models or AI annotators, raising concerns about scalability and reliability, ConfPO is simple, lightweight, and model-free. Experimental results on challenging alignment benchmarks, including AlpacaEval 2 and Arena-Hard, demonstrate that ConfPO consistently outperforms uniform DAAs across various LLMs, delivering better alignment with zero additional computational overhead. Hee Suk Yoon, Eunseop Yoon, Mark Hasegawa-Johnson, Sungwoong Kim, Chang Dong Yoo |
ICML | 3 |
| 2025 | The Interspeech 2025 Speech Accessibility Project Challenge
Xiuwen Zheng 0003, Bornali Phukan, Jonghwan Na, Edward Cutrell, Kyu J. Han, Mark Hasegawa-Johnson, Pan-Pan Jiang, Aadhrik Kuila, Colin Lea, Bob MacDonald, Gautam Varma Mantena, Venkatesh Ravichandran, Leda Sari, Katrin Tomanek, Chang Dong Yoo, Chris Zwilling |
INTERSPEECH | 6 |
| 2025 | SiamCTC: Learning Speech Representations through Monotonic Temporal AlignmentabstractSelf-supervised speech representation learning has made significant progress through Siamese networks, which leverage different views of the same input. However, existing methods often require frame-wise alignment between these views, overlooking the broader linguistic context invariance across different speaking styles. We introduce SiamCTC, a framework that integrates Siamese networks with Connectionist Temporal Classification (CTC) to learn speech representations without strict frame-level correspondence. By employing CTC loss to establish flexible, monotonic alignments between differing temporal realizations of the same content, SiamCTC accommodates speed perturbations and other temporal augmentations. This design relaxes frame-wise constraints while preserving temporal coherence and enhancing robustness to speaking-rate variations in downstream tasks. Our experiments demonstrate that SiamCTC leads to more adaptable speech representations, particularly at diverse speaking rates. SooHwan Eom, Mark Hasegawa-Johnson, Chang Dong Yoo |
INTERSPEECH | 2 |
| 2025 | Band-Split Self-supervised Mamba for Infant-centered Audio Analysis
Xulin Fan, Jialu Li 0002, Mark Hasegawa-Johnson, Nancy McElwain |
INTERSPEECH | 3 |
| 2025 | FaiST: A Benchmark Dataset for Fairness in Speech Technology
Maliha Jahan, Yinglun Sun, Priyam Mazumdar, Zsuzsanna Fagyal, Thomas Thebaud, Jesús Villalba 0001, Mark Hasegawa-Johnson, Najim Dehak, Laureano Moro-Velázquez |
INTERSPEECH | 7 |
| 2025 | Aligning ASR Evaluation with Human and LLM Judgments: Intelligibility Metrics Using Phonetic, Semantic, and NLI Approaches
Bornali Phukan, Xiuwen Zheng 0003, Mark Hasegawa-Johnson |
INTERSPEECH | 3 |
| 2025 | The Speech Accessibility Project: Best Practices for Collection and Curation of Disordered Speech
Chris Zwilling, Mark Hasegawa-Johnson, Heather Hodges, Lorraine O. Ramig, Adina Bradshaw, Clarion Mendes, Alexandria Barkhimer, Laura Mattie, Meg Dickinson, Shawnise Carter, Marie Moore Channell |
INTERSPEECH | 2 |
| 2025 | SyncDiff: Diffusion-Based Talking Head Synthesis with Bottlenecked Temporal Visual Prior for Improved SynchronizationabstractTalking head synthesis, also known as speech-to-lip synthesis, reconstructs the facial motions that align with the given audio tracks. The synthesized videos are evaluated on mainly two aspects, lip-speech synchronization and image fidelity. Recent studies demonstrate that GAN-based and diffusion-based models achieve state-of-the-art (SOTA) performance on this task, with diffusion-based models achieving superior image fidelity but experiencing lower synchronization compared to their GAN-based counterparts. To this end, we propose SYNcDIFF, a simple yet effective approach to improve diffusion-based models using a temporal pose frame with information bottleneck and facial-informative audio features extracted from AVHuBERT, as conditioning input into the diffusion process. We evaluate SYNcDIFF on two canonical talking head datasets, LRS2 and LRS3 for direct comparison with other SOTA models. Experiments on LRS2/LRS3 datasets show that SYNcDIFF achieves a synchronization score 27.7%/62.3% relatively higher than previous diffusion-based methods, while preserving their high-fidelity characteristics. Xulin Fan, Heting Gao, Ziyi Chen 0005, Peng Chang 0002, Mark Hasegawa-Johnson |
WACV | 6 |
| 2024 | InfantMotion2Vec: Unlabeled Data-Driven Infant Pose Estimation Using a Single Chest IMUabstractEarly identification of neuro-developmental risks in infants is crucial for timely intervention and improved quality of life. Current screening methods are costly, intrusive, and limited by artificial environments or require the infant to wear multiple sensors. To address these challenges, we propose a novel approach leveraging inertial measurement units (IMUs) to monitor infants' spontaneous motor abilities in natural settings. Our method introduces a hierarchical semi-supervised classifier and the InfantMotion2Vec embedding to capture detailed motion patterns, accommodating a wide age range (up to 36 months) while minimizing reliance on labeled data and cumbersome sensor setups. We collected labeled IMU data from 25 families and unlabeled data from 42 families using a single wearable sensor. Pretraining an embedding network using unlabeled data with a hierarchical pose estimator resulted in a 26% increase in F1-score and a 77.7% increase in Cohen's Kappa score compared to using only labeled data. The InfantMotion2Vec embedding adequately handles highly unbalanced labeled data, demonstrating its effectiveness in infant posture classification. Mohammad Nur Hossain Khan, Nancy McElwain, Mark Hasegawa-Johnson, Bashima Islam |
BSN | 3 |
| 2024 | Finding Spoken Identifications: Using GPT-4 Annotation for an Efficient and Fast Dataset Creation PipelineabstractThe growing emphasis on fairness in speech-processing tasks requires datasets with speakers from diverse subgroups that allow training and evaluating fair speech technology systems. However, creating such datasets through manual annotation can be costly. To address this challenge, we present a semi-automated dataset creation pipeline that leverages large language models. We use this pipeline to generate a dataset of speakers identifying themself or another speaker as belonging to a particular race, ethnicity, or national origin group. We use OpenaAI’s GPT-4 to perform two complex annotation tasks- separating files relevant to our intended dataset from the irrelevant ones (filtering) and finding and extracting information on identifications within a transcript (tagging). By evaluating GPT-4’s performance using human annotations as ground truths, we show that it can reduce resources required by dataset annotation while barely losing any important information. For the filtering task, GPT-4 had a very low miss rate of 6.93%. GPT-4’s tagging performance showed a trade-off between precision and recall, where the latter got as high as 97%, but precision never exceeded 45%. Our approach reduces the time required for the filtering and tagging tasks by 95% and 80%, respectively. We also present an in-depth error analysis of GPT-4’s performance. Maliha Jahan, Helin Wang, Thomas Thebaud, Yinglun Sun, Giang Ha Le, Zsuzsanna Fagyal, Odette Scharenborg, Mark Hasegawa-Johnson, Laureano Moro-Velázquez, Najim Dehak |
LREC/COLING | 8 |
| 2024 | AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech RecognitionabstractIn Automatic Speech Recognition (ASR) systems, a recurring obstacle is the generation of narrowly focused output distributions. This phenomenon emerges as a side effect of Connectionist Temporal Classification (CTC), a robust sequence learning tool that utilizes dynamic programming for sequence mapping. While earlier efforts have tried to combine the CTC loss with an entropy maximization regularization term to mitigate this issue, they employed a constant weighting term on the regularization during the training, which we find may not be optimal. In this work, we introduce Adaptive Maximum Entropy Regularization (AdaMER), a technique that can modulate the impact of entropy regularization throughout the training process. This approach not only refines ASR model training but ensures that as training proceeds, predictions display the desired model confidence. SooHwan Eom, Eunseop Yoon, Hee Suk Yoon, Chanwoo Kim 0001, Mark Hasegawa-Johnson, Chang Dong Yoo |
ICASSP | 5 |
| 2024 | G2PU: Grapheme-To-Phoneme Transducer with Speech UnitsabstractMost phoneme transcripts are generated using forced alignment: typically a grapheme-to-phoneme transducer (G2P) is applied to text sequences to generate candidate phoneme transcripts, which are then time-aligned to the waveform using an acoustic model. This paper demonstrates, for the first time, simultaneous optimization of the G2P, the acoustic model, and the acoustic alignment to a corpus. To this end, we propose G2PU, a joint CTC-attention model consisting of an encoder-decoder G2P network and an encoder-CTC unit-to-phoneme (U2P) network, where the units are extracted from speech. We demonstrate that the G2P and U2P, operating in parallel, produce lower phone error rates than those of state-of-the-art open-source G2P and forced alignment systems. Furthermore, although the G2P and U2P are trained using parallel speech and text, their synergy can be generalized to text-only test corpora if we also train a grapheme-to-unit (G2U) network that generates speech units from text in the absence of parallel speech. Our G2PU model is trained using phoneme transcripts generated by a teacher G2P tool. Our experiments on Chinese and Japanese show that G2PU reduces phoneme error rate by 7% to 29% relative compared to its teacher. Finally, we include case studies to provide insights into the system’s workings. Heting Gao, Mark Hasegawa-Johnson, Chang Dong Yoo |
ICASSP | 2 |
| 2024 | Unsupervised Speech Recognition with N-skipgram and Positional Unigram MatchingabstractTraining unsupervised speech recognition systems presents challenges due to GAN-associated instability, misalignment between speech and text, and significant memory demands. To tackle these challenges, we introduce a novel ASR system, ESPUM. This system harnesses the power of lower-order N-skipgrams (up to N = 3) combined with positional unigram statistics gathered from a small batch of samples. Evaluated on the TIMIT benchmark, our model showcases competitive performance in ASR and phoneme segmentation tasks. Access our publicly available code at https://github.com/lwang114/GraphUnsupASR. Liming Wang 0003, Mark Hasegawa-Johnson, Chang Dong Yoo |
ICASSP | 2 |
| 2024 | C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature DispersionabstractIn deep learning, test-time adaptation has gained attention as a method for model fine-tuning without the need for labeled data. A prime exemplification is the recently proposed test-time prompt tuning for large-scale vision-language models such as CLIP. Unfortunately, these prompts have been mainly developed to improve accuracy, overlooking the importance of calibration, which is a crucial aspect for quantifying prediction uncertainty. However, traditional calibration methods rely on substantial amounts of labeled data, making them impractical for test-time scenarios. To this end, this paper explores calibration during test-time prompt tuning by leveraging the inherent properties of CLIP. Through a series of observations, we find that the prompt choice significantly affects the calibration in CLIP, where the prompts leading to higher text feature dispersion result in better-calibrated predictions. Introducing the Average Text Feature Dispersion (ATFD), we establish its relationship with calibration error and present a novel method, Calibrated Test-time Prompt Tuning (C-TPT), for optimizing prompts during test-time with enhanced calibration. Through extensive experiments on different CLIP architectures and datasets, we show that C-TPT can effectively improve the calibration of test-time prompt tuning without needing labeled data. The code is publicly accessible at https://github.com/hee-suk-yoon/C-TPT. Hee Suk Yoon, Eunseop Yoon, Joshua Tian Jin Tee, Mark Hasegawa-Johnson, Yingzhen Li, Chang Dong Yoo |
ICLR | 4 |
| 2024 | Speech Self-Supervised Learning Using Diffusion Model Synthetic DataabstractWhile self-supervised learning (SSL) in speech has greatly reduced the reliance of speech processing systems on annotated corpora, the success of SSL still hinges on the availability of a large-scale unannotated corpus, which is still often impractical for many low-resource languages or under privacy concerns. Some existing work seeks to alleviate the problem by data augmentation, but most works are confined to introducing perturbations to real speech and do not introduce new variations in speech prosody, speakers, and speech content, which are important for SSL. Motivated by the recent finding that diffusion models have superior capabilities for modeling data distributions, we propose DiffS4L, a pretraining scheme that augments the limited unannotated data with synthetic data with different levels of variations, generated by a diffusion model trained on the limited unannotated data. Finally, an SSL model is pre-trained on the real and the synthetic speech. Our experiments show that DiffS4L can significantly improve the performance of SSL models, such as reducing the WER of the HuBERT pretrained model by 6.26 percentage points in the English ASR task. Notably, we find that the synthetic speech with all levels of variations, i.e. new prosody, new speakers, and even new content (despite the new content being mostly babble), accounts for significant performance improvement. The code is available at github.com/Hertin/DiffS4L. Heting Gao, Kaizhi Qian, Junrui Ni, Chuang Gan 0001, Mark Hasegawa-Johnson, Shiyu Chang, Yang Zhang 0001 |
ICML | 5 |
| 2024 | Enhancing Child Vocalization Classification with Phonetically-Tuned Embeddings for Assisting Autism Diagnosis
Jialu Li 0002, Mark Hasegawa-Johnson, Karrie Karahalios |
INTERSPEECH | 2 |
| 2024 | Fine-Tuning Automatic Speech Recognition for People with Parkinson's: An Effective Strategy for Enhancing Speech Technology AccessibilityabstractThis paper enhances dysarthric and dysphonic speech recognition by fine-tuning pretrained automatic speech recognition (ASR) models on the 2023-10-05 data package of the Speech Accessibility Project (SAP), which contains the speech of 253 people with Parkinson's disease.Experiments tested methods that have been effective for Cerebral Palsy, including the use of speaker clustering and severity-dependent models, weighted fine-tuning, and multi-task learning.Best results were obtained using a multi-task learning model, in which the ASR is trained to produce an estimate of the speaker's impairment severity as an auxiliary output.The resulting word error rates are considerably improved relative to a baseline model fine-tuned using only Librispeech data, with word error rate improvements of 37.62% and 26.97% compared to fine-tuning on 100h and 960h of LibriSpeech data, respectively. Xiuwen Zheng 0003, Bornali Phukan, Mark Hasegawa-Johnson |
INTERSPEECH | 3 |
| 2024 | Visualization for improving foreign language pronunciation
Charlotte Yoder, Karrie Karahalios, Mark Hasegawa-Johnson, Shreyansh Agrawal |
INTERSPEECH | 3 |
| 2024 | LI-TTA: Language Informed Test-Time Adaptation for Automatic Speech Recognition
Eunseop Yoon, Hee Suk Yoon, John B. Harvill, Mark Hasegawa-Johnson, Chang Dong Yoo |
INTERSPEECH | 4 |
| 2024 | Just ASR + LLM? A Study on Speech Large Language Models' Ability to Identify And Understand Speaker in Spoken DialogueabstractIn recent years, we have observed a rapid advancement in speech language models (SpeechLLMs), catching up with humans’ listening and reasoning abilities. SpeechLLMs have demonstrated impressive spoken dialog question-answering (SQA) performance in benchmarks like Gaokao, the English listening test of the college entrance exam in China, which seemingly requires understanding both the spoken content and voice characteristics of speakers in a conversation. However, after carefully examining Gaokao’s questions, we find the correct answers to many questions can be inferred from the conversation transcript alone, i.e. without speaker segmentation and identification. Our evaluation of state-of-the-art models Qwen-Audio and WavLLM on both Gaokao and our proposed “What Do You Like?” dataset shows a significantly higher accuracy in these context-based questions than in identity-critical questions, which can only be answered reliably with correct speaker identification. The results and analysis suggest that when solving SQA, the current SpeechLLMs exhibit limited speaker awareness from the audio and behave similarly to an LLM reasoning from the conversation transcription without sound. We propose that tasks focused on identity-critical questions could offer a more accurate evaluation framework of SpeechLLMs in SQA. Junkai Wu, Xulin Fan, Bo-Ru Lu, Xilin Jiang, Nima Mesgarani, Mark Hasegawa-Johnson, Mari Ostendorf |
SLT | 6 |
| 2023 | A Theory of Unsupervised Speech RecognitionabstractUnsupervised speech recognition (ASR-U) is the problem of learning automatic speech recognition (ASR) systems from unpaired speech-only and text-only corpora.While various algorithms exist to solve this problem, a theoretical framework is missing to study their properties and address such issues as sensitivity to hyperparameters and training instability.In this paper, we proposed a general theoretical framework to study the properties of ASR-U systems based on random matrix theory and the theory of neural tangent kernels.Such a framework allows us to prove various learnability conditions and sample complexity bounds of ASR-U.Extensive ASR-U experiments on synthetic languages with three classes of transition graphs provide strong empirical evidence for our theory (code available at cactuswith- thoughts/UnsupASRTheory.git). Liming Wang 0003, Mark Hasegawa-Johnson, Chang Dong Yoo |
ACL (1) | 2 |
| 2023 | Lightweight, Multi-Speaker, Multi-Lingual Indic Text-to-SpeechabstractThe Lightweight, Multi-speaker, Multi-lingual Indic Text-to-Speech (LIMMITS’23) challenge is organized as part of the ICASSP 2023 signal processing grand challenge. LIMMITS’23 aims at the development of a lightweight, multi-speaker, multi-lingual Text to Speech (TTS) model using datasets in Marathi, Hindi, and Telugu. The challenge encourages the advancement of TTS in Indian Languages as well as the development of techniques involved in TTS data selection and model compression. The 3 tracks of LIMMITS’23 have provided an opportunity for various researchers and practitioners around the world to explore the state of the art in TTS research. Abhayjeet Singh, Amala Nagireddi, Deekshitha G, Jesuraja Bandekar, Roopa R., Sandhya Badiger, Sathvik Udupa, Prasanta Kumar Ghosh, Hema A. Murthy, Heiga Zen, Pranaw Kumar, Kamal Kant, Amol Bole, Bira Chandra Singh, Keiichi Tokuda, Mark Hasegawa-Johnson, Philipp Olbrich |
ICASSP | 16 |
| 2023 | Dual-Path Cross-Modal Attention for Better Audio-Visual Speech ExtractionabstractAudiovisual target speaker extraction is the task of separating, from an audio mixture, the speaker whose face is visible in an accompanying video. Published approaches typically upsample the video or downsample the audio, then fuse the two streams using concatenation, multiplication, or cross-modal attention. This paper proposes, instead, to use a dual-path attention architecture in which the audio chunk length is comparable to the duration of a video frame. Audio is transformed by intra-chunk attention, concatenated to video features, then transformed by inter-chunk attention. Because of residual connections, the audio and video features remain logically distinct across multiple network layers, therefore dual-path audiovisual feature fusion can be performed repeatedly across multiple layers. When given 2-5-speaker mixtures constructed from the challenging LRS3 test set, results are about 7dB better than Con-vTasNet or AV-ConvTasNet, with the performance gap widening slightly as the number of speakers increases. Zhongweiyang Xu, Xulin Fan, Mark Hasegawa-Johnson |
ICASSP | 3 |
| 2023 | End-to-End Zero-Shot Voice Conversion with Location-Variable ConvolutionsabstractZero-shot voice conversion is becoming an increasingly popular research topic, as it promises the ability to transform speech to sound like any speaker.However, relatively little work has been done on end-to-end methods for this task, which are appealing because they remove the need for a separate vocoder to generate audio from intermediate features.In this work, we propose LVC-VC, an end-to-end zero-shot voice conversion model that uses location-variable convolutions (LVCs) to jointly model the conversion and speech synthesis processes.LVC-VC utilizes carefully designed input features that have disentangled content and speaker information, and it uses a neural vocoder-like architecture that utilizes LVCs to efficiently combine them and perform voice conversion while directly synthesizing time domain audio.Experiments show that our model achieves especially well balanced performance between voice style transfer and speech intelligibility compared to several baselines. Wonjune Kang, Mark Hasegawa-Johnson, Deb Roy |
INTERSPEECH | 2 |
| 2023 | Towards Robust Family-Infant Audio Analysis Based on Unsupervised Pretraining of Wav2vec 2.0 on Large-Scale Unlabeled Family AudioabstractTo perform automatic family audio analysis, past studies have collected recordings using phone, video, or audio-only recording devices like LENA, investigated supervised learning methods, and used or fine-tuned general-purpose embeddings learned from large pretrained models. In this study, we advance the audio component of a new infant wearable multi-modal device called LittleBeats (LB) by learning family audio representation via wav2vec 2.0 (W2V2) pretraining. We show given a limited number of labeled LB home recordings, W2V2 pretrained using 1k-hour of unlabeled home recordings outperforms oracle W2V2 pretrained on 52k-hour unlabeled audio in terms of parent/infant speaker diarization (SD) and vocalization classifications (VC) at home. Extra relevant external unlabeled and labeled data further benefit W2V2 pretraining and fine-tuning. With SpecAug and environmental speech corruptions, we obtain 12% relative gain on SD and moderate boost on VC. Code and model weights are available. Jialu Li 0002, Mark Hasegawa-Johnson, Nancy McElwain |
INTERSPEECH | 2 |
| 2023 | Mitigating the Exposure Bias in Sentence-Level Grapheme-to-Phoneme (G2P) Transduction
Eunseop Yoon, Hee Suk Yoon, Dhananjaya Gowda, SooHwan Eom, Daehyeok Kim, John B. Harvill, Heting Gao, Mark Hasegawa-Johnson, Chanwoo Kim 0001, Chang Dong Yoo |
INTERSPEECH | 8 |
| 2023 | Wav2ToBI: a new approach to automatic ToBI transcription
Wanyue Zhai, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2023 | Automated morphological phenotyping using learned shape descriptors and functional maps: A novel approach to geometric morphometricsabstractThe methods of geometric morphometrics are commonly used to quantify morphology in a broad range of biological sciences. The application of these methods to large datasets is constrained by manual landmark placement limiting the number of landmarks and introducing observer bias. To move the field forward, we need to automate morphological phenotyping in ways that capture comprehensive representations of morphological variation with minimal observer bias. Here, we present Morphological Variation Quantifier (morphVQ), a shape analysis pipeline for quantifying, analyzing, and exploring shape variation in the functional domain. morphVQ uses descriptor learning to estimate the functional correspondence between whole triangular meshes in lieu of landmark configurations. With functional maps between pairs of specimens in a dataset we can analyze and explore shape variation. morphVQ uses Consistent ZoomOut refinement to improve these functional maps and produce a new representation of shape variation, area-based and conformal (angular) latent shape space differences (LSSDs). We compare this new representation of shape variation to shape variables obtained via manual digitization and auto3DGM, an existing approach to automated morphological phenotyping. We find that LSSDs compare favorably to modern 3DGM and auto3DGM while being more computationally efficient. By characterizing whole surfaces, our method incorporates more morphological detail in shape analysis. We can classify known biological groupings, such as Genus affiliation with comparable accuracy. The shape spaces produced by our method are similar to those produced by modern 3DGM and to auto3DGM, and distinctiveness functions derived from LSSDs show us how shape variation differs between groups. morphVQ can capture shape in an automated fashion while avoiding the limitations of manually digitized landmarks, and thus represents a novel and computationally efficient addition to the geometric morphometrics toolkit. Oshane O. Thomas, Hongyu Shen, Ryan L. Raaum, William E. H. Harcourt-Smith, John D. Polk, Mark Hasegawa-Johnson |
PLoS Comput. Biol. | 6 |
| 2022 | Fast and Efficient MMD-Based Fair PCA via Optimization over Stiefel ManifoldabstractThis paper defines fair principal component analysis (PCA) as minimizing the maximum mean discrepancy (MMD) between the dimensionality-reduced conditional distributions of different protected classes. The incorporation of MMD naturally leads to an exact and tractable mathematical formulation of fairness with good statistical properties. We formulate the problem of fair PCA subject to MMD constraints as a non-convex optimization over the Stiefel manifold and solve it using the Riemannian Exact Penalty Method with Smoothing (REPMS). Importantly, we provide a local optimality guarantee and explicitly show the theoretical effect of each hyperparameter in practical settings, extending previous results. Experimental comparisons based on synthetic and UCI datasets show that our approach outperforms prior work in explained variance, fairness, and runtime. Gwangsu Kim, Mahbod Olfat, Mark Hasegawa-Johnson, Chang Dong Yoo |
AAAI | 4 |
| 2022 | Self-supervised Semantic-driven Phoneme Discovery for Zero-resource Speech RecognitionabstractPhonemes are defined by their relationship to words: changing a phoneme changes the word.Learning a phoneme inventory with little supervision has been a longstanding challenge with important applications to underresourced speech technology.In this paper, we bridge the gap between the linguistic and statistical definition of phonemes and propose a novel neural discrete representation learning model for self-supervised learning of phoneme inventory with raw speech and word labels.Given the availability of phoneme segmentation and some mild conditions, we prove that the phoneme inventory learned by our approach converges to the true one with an exponentially low error rate.Moreover, in experiments on TIMIT and Mboshi benchmarks, our approach consistently learns a better phonemelevel representation and achieves a lower error rate in a zero-resource phoneme recognition task than previous state-of-the-art selfsupervised representation learning algorithms. Liming Wang 0003, Siyuan Feng 0003, Mark Hasegawa-Johnson, Chang Dong Yoo |
ACL (1) | 3 |
| 2022 | Equivariance Discovery by Learned Parameter-SharingabstractDesigning equivariance as an inductive bias into deep-nets has been a prominent approach to build effective models, e.g., a convolutional neural network incorporates translation equivariance. However, incorporating these inductive biases requires knowledge about the equivariance properties of the data, which may not be available, e.g., when encountering a new domain. To address this, we study how to "discover interpretable equivariances" from data. Specifically, we formulate this discovery process as an optimization problem over a model’s parameter-sharing schemes. We propose to use the partition distance to empirically quantify the accuracy of the recovered equivariance. Also, we theoretically analyze the method for Gaussian data and provide a bound on the mean squared gap between the studied discovery scheme and the oracle scheme. Empirically, we show that the approach recovers known equivariances, such as permutations and shifts, on sum of numbers and spatially-invariant data. Raymond A. Yeh, Yuan-Ting Hu, Mark Hasegawa-Johnson, Alexander G. Schwing |
AISTATS | 3 |
| 2022 | SpeechSplit2.0: Unsupervised Speech Disentanglement for Voice Conversion without Tuning Autoencoder BottlenecksabstractSpeechSplit can perform aspect-specific voice conversion by disentangling speech into content, rhythm, pitch, and timbre using multiple autoencoders in an unsupervised manner. However, SpeechSplit requires careful tuning of the autoencoder bottlenecks, which can be time-consuming and less robust. This paper proposes SpeechSplit2.0, which constrains the information flow of the speech component to be disentangled on the autoencoder input using efficient signal processing methods instead of bottleneck tuning. Evaluation results show that SpeechSplit2.0 achieves comparable performance to SpeechSplit in speech disentanglement and superior robustness to the bottleneck size variations. Our code is available at https://github.com/biggytruck/SpeechSplit2. Chak Ho Chan, Kaizhi Qian, Yang Zhang 0001, Mark Hasegawa-Johnson |
ICASSP | 4 |
| 2022 | Detection of Covid-19 from Joint Time and Frequency Analysis of Speech, Breathing and Cough AudioabstractThe distinct cough sounds produced by a variety of respiratory diseases suggest the potential for the development of a new class of audio bio-markers for the detection of COVID-19. Accurate audio biomarker-based COVID-19 tests would be inexpensive, readily scalable, and non-invasive. Audio biomarker screening could also be utilized in resource-limited settings prior to traditional diagnostic testing. Here we explore the possibility of leveraging three audio modalities: cough, breathing, and speech to determine COVID-19 status. We train a separate neural classification system on each modality, as well as a fused classification system on all three modalities together. Ablation studies are performed to understand the relationship between individual and collective performance of the modalities. Additionally, we analyze the extent to which temporal and spectral features contribute to COVID-19 status information contained in the audio signals. John B. Harvill, Yash R. Wani, Moitreya Chatterjee, Mustafa Alam, David G. Beiser, David Chestek, Mark Hasegawa-Johnson, Narendra Ahuja |
ICASSP | 7 |
| 2022 | Forget-free Continual Learning with Winning SubnetworksabstractInspired by Lottery Ticket Hypothesis that competitive subnetworks exist within a dense network, we propose a continual learning method referred to as Winning SubNetworks (WSN), which sequentially learns and selects an optimal subnetwork for each task. Specifically, WSN jointly learns the model weights and task-adaptive binary masks pertaining to subnetworks associated with each task whilst attempting to select a small set of weights to be activated (winning ticket) by reusing weights of the prior subnetworks. The proposed method is inherently immune to catastrophic forgetting as each selected subnetwork model does not infringe upon other subnetworks. Binary masks spawned per winning ticket are encoded into one N-bit binary digit mask, then compressed using Huffman coding for a sub-linear increase in network capacity with respect to the number of tasks. Haeyong Kang, Rusty Mina, Sultan Rizky Hikmawan Madjid, Jaehong Yoon, Mark Hasegawa-Johnson, Sung Ju Hwang, Chang Dong Yoo |
ICML | 5 |
| 2022 | ContentVec: An Improved Self-Supervised Speech Representation by Disentangling SpeakersabstractSelf-supervised learning in speech involves training a speech representation network on a large-scale unannotated speech corpus, and then applying the learned representations to downstream tasks. Since the majority of the downstream tasks of SSL learning in speech largely focus on the content information in speech, the most desirable speech representations should be able to disentangle unwanted variations, such as speaker variations, from the content. However, disentangling speakers is very challenging, because removing the speaker information could easily result in a loss of content as well, and the damage of the latter usually far outweighs the benefit of the former. In this paper, we propose a new SSL method that can achieve speaker disentanglement without severe loss of content. Our approach is adapted from the HuBERT framework, and incorporates disentangling mechanisms to regularize both the teacher labels and the learned representations. We evaluate the benefit of speaker disentanglement on a set of content-related downstream tasks, and observe a consistent and notable performance advantage of our speaker-disentangled representations. Kaizhi Qian, Yang Zhang 0001, Heting Gao, Junrui Ni, Cheng-I Lai, David D. Cox, Mark Hasegawa-Johnson, Shiyu Chang |
ICML | 7 |
| 2022 | WavPrompt: Towards Few-Shot Spoken Language Understanding with Frozen Language Models
Heting Gao, Junrui Ni, Kaizhi Qian, Yang Zhang 0001, Shiyu Chang, Mark Hasegawa-Johnson |
INTERSPEECH | 6 |
| 2022 | Frame-Level Stutter Detection
John B. Harvill, Mark Hasegawa-Johnson, Chang Dong Yoo |
INTERSPEECH | 2 |
| 2022 | Cross-lingual articulatory feature information transfer for speech recognition using recurrent progressive neural networks
Mahir Morshed, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2022 | Unsupervised Text-to-Speech Synthesis by Unsupervised Automatic Speech Recognition
Junrui Ni, Liming Wang 0003, Heting Gao, Kaizhi Qian, Yang Zhang 0001, Shiyu Chang, Mark Hasegawa-Johnson |
INTERSPEECH | 7 |
| 2022 | Syn2Vec: Synset Colexification Graphs for Lexical Semantic SimilarityabstractIn this paper we focus on patterns of colexification (co-expressions of form-meaning mapping in the lexicon) as an aspect of lexicalsemantic organization, and use them to build large scale synset graphs across BabelNet's typologically diverse set of 499 world languages.We introduce and compare several approaches: monolingual and cross-lingual colexification graphs, popular distributional models, and fusion approaches.The models are evaluated against human judgments on a semantic similarity task for nine languages.Our strong empirical findings also point to the importance of universality of our graph synset embedding representations with no need for any language-specific adaptation when evaluated on the lexical similarity task.The insights of our exploratory investigation of largescale colexification graphs could inspire significant advances in NLP across languages, especially for tasks involving languages which lack dedicated lexical resources, and can benefit from language transfer from large shared cross-lingual semantic spaces. John B. Harvill, Roxana Girju, Mark Hasegawa-Johnson |
NAACL-HLT | 3 |
| 2022 | Discovering phonetic inventories with crosslingual automatic speech recognition
Piotr Zelasko, Siyuan Feng 0001, Laureano Moro-Velázquez, Ali Abavisani, Saurabhchand Bhati, Odette Scharenborg, Mark Hasegawa-Johnson, Najim Dehak |
Comput. Speech Lang. | 7 |
| 2022 | Seamless equal accuracy ratio for inclusive CTC speech recognitionabstractConcerns have been raised regarding performance disparity in automatic speech recognition (ASR) systems as they provide unequal transcription accuracy for different user groups defined by different attributes that include gender, dialect, and race. In this paper, we propose “equal accuracy ratio”, a novel inclusiveness measure for ASR systems that can be seamlessly integrated into the standard connectionist temporal classification (CTC) training pipeline of an end-to-end neural speech recognizer to increase the recognizer’s inclusiveness. We also create a novel multi-dialect benchmark dataset to study the inclusiveness of ASR, by combining data from existing corpora in seven dialects of English (African American, General American, Latino English, British English, Indian English, Afrikaaner English, and Xhosa English). Experiments on this multi-dialect corpus show that using the equal accuracy ratio as a regularization term along with CTC loss, succeeds in lowering the accuracy gap between user groups and reduces the recognition error rate compared with a non-regularized baseline. Experiments on additional speech corpora that have different user groups also confirm our findings. Heting Gao, Sunghun Kang, Rusty Mina, Dias Issa, John B. Harvill, Leda Sari, Mark Hasegawa-Johnson, Chang Dong Yoo |
Speech Commun. | 8 |
| 2022 | Autosegmental Neural Nets 2.0: An Extensive Study of Training Synchronous and Asynchronous Phones and Tones for Under-Resourced Tonal LanguagesabstractPhones, the segmental units in the International Phonetic Alphabet (IPA), include isolated consonants or vowels; tones, the suprasegemental units, represent pitch and voice quality movements that may span many phones. The timings of tones and phones are loosely connected, e.g., tones may be synchronized with their associated vowels, syllable finals, or a sequence of two or three syllables depending on the language. Many past studies have investigated cross-lingual adaptation in an automatic speech recognition (ASR) tone-marked phone model, yet very few studied the interaction between cross-lingual adaptation and tone-phone synchronization. In this study, we perform an extensive study by multilingual training on four tonal languages and cross-lingual testing on the fifth, in a five-fold cross-validation framework, using four CTC-based systems that impose different degrees of synchronization between tones and phones. We discover that multilingual and cross-lingual training benefit from different training architectures. In multilingual training, when a large corpus of test-language training data is part of the training corpus, a system that requires synchronization of tones with phones produces significantly lower tone error rates than any of the systems that score tones and phones asynchronously. In cross-lingual training, however, when only limited adaptation data are available in the test language, jointly training synchronous tone-marked phones together with asynchronous phones and tones, as three separate system outputs jointly optimized using a multi-task learning framework, consistently and significantly outperforms the system that requires tone-phone synchrony. Jialu Li 0002, Mark Hasegawa-Johnson |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | How Phonotactics Affect Multilingual and Zero-Shot ASR PerformanceabstractThe idea of combining multiple languages’ recordings to train a single automatic speech recognition (ASR) model brings the promise of the emergence of universal speech representation. Recently, a Transformer encoder-decoder model has been shown to leverage multilingual data well in IPA transcriptions of languages presented during training. However, the representations it learned were not successful in zero-shot transfer to unseen languages. Because that model lacks an explicit factorization of the acoustic model (AM) and language model (LM), it is unclear to what degree the performance suffered from differences in pronunciation or the mismatch in phono-tactics. To gain more insight into the factors limiting zero-shot ASR transfer, we replace the encoder-decoder with a hybrid ASR system consisting of a separate AM and LM. Then, we perform an extensive evaluation of monolingual, multilingual, and crosslingual (zero-shot) acoustic and language models on a set of 13 phonetically diverse languages. We show that the gain from modeling crosslingual phonotactics is limited, and imposing a too strong model can hurt the zero-shot transfer. Furthermore, we find that a multilingual LM hurts a multilingual ASR system’s performance, and retaining only the target language’s phonotactic data in LM training is preferable. Siyuan Feng 0001, Piotr Zelasko, Laureano Moro-Velázquez, Ali Abavisani, Mark Hasegawa-Johnson, Odette Scharenborg, Najim Dehak |
ICASSP | 5 |
| 2021 | Synthesis of New Words for Improved Dysarthric Speech Recognition on an Expanded VocabularyabstractDysarthria is a condition where people experience a reduction in speech intelligibility due to a neuromotor disorder. Previous works in dysarthric speech recognition have focused on accurate recognition of words encountered in training data. Due to the rarity of dysarthria in the general population, a relatively small amount of publicly-available training data exists for dysarthric speech. The number of unique words in these datasets is small, so ASR systems trained with existing dysarthric speech data are limited to recognition of those words. In this paper, we propose a data augmentation method using voice conversion that allows dysarthric ASR systems to accurately recognize words outside of the training set vocabulary. We demonstrate that a small amount of dysarthric speech data can be used to capture the relevant vocal characteristics of a speaker with dysarthria through a parallel voice conversion system. We show that it’s possible to synthesize utterances of new words that were never recorded by speakers with dysarthria, and that these synthesized utterances can be used to train a dysarthric ASR system. John B. Harvill, Dias Issa, Mark Hasegawa-Johnson, Chang Dong Yoo |
ICASSP | 3 |
| 2021 | Continuous Cnn For Nonuniform Time SeriesabstractCNN for time series data implicitly assumes that the data are uniformly sampled, whereas many event-based and multi-modal data are nonuniform or have heterogeneous sampling rates. Directly applying regular CNN to nonuniform time series is ungrounded, because it is unable to recognize and extract common patterns from the nonuniform input signals. In this paper, we propose the Continuous CNN (CCNN), which estimates the inherent continuous inputs by interpolation, and performs continuous convolution on the continuous input. The interpolation and convolution kernels are learned in an end-to-end manner, and are able to learn useful patterns despite the nonuniform sampling rate. Results of several experiments verify that CCNN achieves a better performance on nonuniform data, and learns meaningful continuous kernels.1 Yang Zhang 0001, Shiyu Chang, Kaizhi Qian, Mark Hasegawa-Johnson, Jishen Zhao |
ICASSP | 6 |
| 2021 | Show and Speak: Directly Synthesize Spoken Description of ImagesabstractThis paper proposes a new model, referred to as the show and speak (SAS) model that, for the first time, is able to directly synthesize spoken descriptions of images, bypassing the need for any text or phonemes. The basic structure of SAS is an encoder-decoder architecture that takes an image as input and predicts the spectrogram of speech that describes this image. The final speech audio is obtained from the predicted spectrogram via WaveNet. Extensive experiments on the public benchmark database Flickr8k demonstrate that the proposed SAS is able to synthesize natural spoken descriptions for images, indicating that synthesizing spoken descriptions for images while bypassing text and phonemes is feasible. Siyuan Feng 0001, Jihua Zhu, Mark Hasegawa-Johnson, Odette Scharenborg |
ICASSP | 4 |
| 2021 | Align or attend? Toward More Efficient and Accurate Spoken Word Discovery Using Speech-to-Image RetrievalabstractMultimodal word discovery (MWD) is often treated as a byproduct of the speech-to-image retrieval problem. However, our theoretical analysis shows that some kind of alignment/attention mechanism is crucial for a MWD system to learn meaningful word-level representation. We verify our theory by conducting retrieval and word discovery experiments on MSCOCO and Flickr8k, and empirically demonstrate that both neural MT with self-attention and statistical MT achieve word discovery scores that are superior to those of a state-of-the-art neural retrieval system, outperforming it by 2% and 5% alignment F1 scores respectively. Liming Wang 0003, Mark Hasegawa-Johnson, Odette Scharenborg, Najim Dehak |
ICASSP | 3 |
| 2021 | A Comparison Study on Infant-Parent Voice Diarization
Junzhe Zhu, Mark Hasegawa-Johnson, Nancy McElwain |
ICASSP | 2 |
| 2021 | Multi-Decoder Dprnn: Source Separation for Variable Number of SpeakersabstractWe propose an end-to-end trainable approach to single-channel speech separation with unknown number of speakers. Our approach extends the MulCat source separation backbone with additional output heads: a count-head to infer the number of speakers, and decoder-heads for reconstructing the original signals. Beyond the model, we also propose a metric on how to evaluate source separation with variable number of speakers. Specifically, we clear up the issue on how to evaluate the quality when the ground-truth has more or less speakers than the ones predicted by the model. We evaluate our approach on the WSJ0-mix datasets, with mixtures up to five speakers. We demonstrate that our approach outperforms state-of-the-art in counting the number of speakers and remains competitive in quality of reconstructed signals. Junzhe Zhu, Raymond A. Yeh, Mark Hasegawa-Johnson |
ICASSP | 3 |
| 2021 | Interpretable Visual Reasoning via Induced Symbolic SpaceabstractWe study the problem of concept induction in visual reasoning, i.e., identifying concepts and their hierarchical relationships from question-answer pairs associated with images; and achieve an interpretable model via working on the induced symbolic concept space. To this end, we first design a new framework named object-centric compositional attention model (OCCAM) to perform the visual reasoning task with object-level visual features. Then, we come up with a method to induce concepts of objects and relations using clues from the attention patterns between objects’ visual features and question words. Finally, we achieve a higher level of interpretability by imposing OCCAM on the objects represented in the induced symbolic concept space. Experiments on the CLEVR and GQA datasets demonstrate: 1) our OCCAM achieves a new state of the art without human-annotated functional programs; 2) our induced concepts are both accurate and sufficient as OCCAM achieves an on-par performance on objects represented either in visual features or in the induced symbolic concept space. Zhonghao Wang 0001, Kai Wang 0058, Mo Yu, Jinjun Xiong, Wen-Mei W. Hwu, Mark Hasegawa-Johnson, Humphrey Shi |
ICCV | 6 |
| 2021 | Global Prosody Style Transfer Without Text TranscriptionsabstractProsody plays an important role in characterizing the style of a speaker or an emotion, but most non-parallel voice or emotion style transfer algorithms do not convert any prosody information. Two major components of prosody are pitch and rhythm. Disentangling the prosody information, particularly the rhythm component, from the speech is challenging because it involves breaking the synchrony between the input speech and the disentangled speech representation. As a result, most existing prosody style transfer algorithms would need to rely on some form of text transcriptions to identify the content information, which confines their application to high-resource languages only. Recently, SpeechSplit has made sizeable progress towards unsupervised prosody style transfer, but it is unable to extract high-level global prosody style in an unsupervised manner. In this paper, we propose AutoPST, which can disentangle global prosody style from speech without relying on any text transcriptions. AutoPST is an Autoencoder-based Prosody Style Transfer framework with a thorough rhythm removal module guided by the self-expressive representation learning. Experiments on different style transfer tasks show that AutoPST can effectively convert prosody that correctly reflects the styles of the target domains. Kaizhi Qian, Yang Zhang 0001, Shiyu Chang, Jinjun Xiong, Chuang Gan 0001, David D. Cox, Mark Hasegawa-Johnson |
ICML | 7 |
| 2021 | Zero-Shot Cross-Lingual Phonetic Recognition with External Language Embedding
Heting Gao, Junrui Ni, Yang Zhang 0001, Kaizhi Qian, Shiyu Chang, Mark Hasegawa-Johnson |
Interspeech | 6 |
| 2021 | Classification of COVID-19 from Cough Using Autoregressive Predictive Coding Pretraining and Spectral Data AugmentationabstractSerum and saliva-based testing methods have been crucial to slowing the COVID-19 pandemic, yet have been limited by slow throughput and cost.A system able to determine COVID-19 status from cough sounds alone would provide a low cost, rapid, and remote alternative to current testing methods.We explore the applicability of recent techniques such as pre-training and spectral augmentation in improving the performance of a neural cough classification system.We use Autoregressive Predictive Coding (APC) to pre-train a unidirectional LSTM on the COUGHVID dataset.We then generate our final model by finetuning added BLSTM layers on the DiCOVA challenge dataset.We perform various ablation studies to see how each component impacts performance and improves generalization with a small dataset.Our final system achieves an AUC of 85.35 and places third out of 29 entries in the DiCOVA challenge. John B. Harvill, Yash R. Wani, Mark Hasegawa-Johnson, Narendra Ahuja, David G. Beiser, David Chestek |
Interspeech | 3 |
| 2021 | Worldly Wise (WoW) - Cross-Lingual Knowledge Fusion for Fact-based Visual Spoken-Question AnsweringabstractKiran Ramnath, Leda Sari, Mark Hasegawa-Johnson, Chang Yoo. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Kiran Ramnath, Leda Sari, Mark Hasegawa-Johnson, Chang Dong Yoo |
NAACL-HLT | 3 |
| 2021 | Analysis of acoustic and voice quality features for the classification of infant and mother vocalizationsabstractClassification of infant and parent vocalizations, particularly emotional vocalizations, is critical to understanding how infants learn to regulate emotions in social dyadic processes. This work is an experimental study of classifiers, features, and data augmentation strategies applied to the task of classifying infant and parent vocalization types. Our data were recorded both in the home and in the laboratory. Infant vocalizations were manually labeled as cry, fus (fuss), lau (laugh), bab (babble) or scr (screech), while parent (mostly mother) vocalizations were labeled as ids (infant-directed speech), ads (adult-directed speech), pla (playful), rhy (rhythmic speech or singing), lau (laugh) or whi (whisper). Linear discriminant analysis (LDA) was selected as a baseline classifier, because it gave the highest accuracy in a previously published study covering part of this corpus. LDA was compared to two neural network architectures: a two-layer fully-connected network (FCN), and a convolutional neural network with self-attention (CNSA). Baseline features extracted using the OpenSMILE toolkit were augmented by extra voice quality, phonetic, and prosodic features, each targeting perceptual features of one or more of the vocalization types. Three web data augmentation and transfer learning methods were tested: pre-training of network weights for a related task (adult emotion classification), augmentation of under-represented classes using data uniformly sampled from other corpora, and augmentation of under-represented classes using data selected by a minimum cross-corpus information difference criterion. Feature selection using Fisher scores and experiments of using weighted and unweighted samplers were also tested. Two datasets were evaluated: a benchmark dataset (CRIED) and our own corpus. In terms of unweighted-average recall of CRIED dataset, the CNSA achieved the best UAR compared with previous studies. In terms of classification accuracy, weighted F1, and macro F1 of our own dataset, the neural networks both significantly outperformed LDA; the FCN slightly (but not significantly) outperformed the CNSA. Cross-examining features selected by different feature selection algorithms permits a type of post-hoc feature analysis, in which the most important acoustic features for each binary type discrimination are listed. Examples of each vocalization type of overlapped features were selected, and their spectrograms are presented, and discussed with respect to the type-discriminative acoustic features selected by various algorithms. MFCC, log Mel Frequency Band Energy, LSP frequency, and F1 are found to be the most important spectral envelope features; F0 is found to be the most important prosodic feature. Jialu Li 0002, Mark Hasegawa-Johnson, Nancy McElwain |
Speech Commun. | 2 |
| 2021 | Auxiliary Networks for Joint Speaker Adaptation and Speaker Change DetectionabstractSpeaker adaptation and speaker change detection have both been studied extensively to improve automatic speech recognition (ASR). In many cases, these two problems are investigated separately: speaker change detection is implemented first to obtain single-speaker regions, and speaker adaptation is then performed using the derived speaker segments for improved ASR. However, in an online setting, we want to achieve both goals in a single pass. In this study, we propose a neural network architecture that learns a speaker embedding from which it can perform both speaker adaptation for ASR and speaker change detection. The proposed speaker embedding is computed using self-attention based on an auxiliary network attached to a main ASR network. ASR adaptation is then performed by subtracting, from the main network activations, a segment dependent affine transformation of the learned speaker embedding. In experiments on a broadcast news dataset and the Switchboard conversational dataset, we test our system on utterances with a change point in them and show that the proposed method achieves significantly better performance as compared to the unadapted main network (10-14% relative reduction in word error rate (WER)). The proposed architecture also outperforms three different speaker segmentation methods followed by ASR (around 10% relative reduction in WER). Leda Sari, Mark Hasegawa-Johnson, Samuel Thomas 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Counterfactually Fair Automatic Speech RecognitionabstractWidely used automatic speech recognition (ASR) systems have been empirically demonstrated in various studies to be unfair, having higher error rates for some groups of users than others. One way to define fairness in ASR is to require that changing the demographic group affiliation of any individual (e.g., changing their gender, age, education or race) should not change the probability distribution across possible speech-to-text transcriptions. In the paradigm of counterfactual fairness, all variables independent of group affiliation (e.g., the text being read by the speaker) remain unchanged, while variables dependent on group affiliation (e.g., the speaker's voice) are counterfactually modified. Hence, we approach the fairness of ASR by training the ASR to minimize change in its outcome probabilities despite a counterfactual change in the individual's demographic attributes. Starting from the individualized counterfactual equal odds criterion, we provide relaxations to it and compare their performances for connectionist temporal classification (CTC) based end-to-end ASR systems. We perform our experiments on the Corpus of Regional African American Languages (CORAAL) and the LibriSpeech dataset to accommodate for differences due to gender, age, education, and race. We show that with counterfactual training, we can reduce average character error rates while achieving lower performance gap between demographic groups, and lower error standard deviation among individuals. Leda Sari, Mark Hasegawa-Johnson, Chang Dong Yoo |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Synthesizing Spoken Descriptions of ImagesabstractImage captioning technology has great potential in many scenarios. However, current text-based image captioning methods cannot be applied to approximately half of the world's languages due to these languages lack of a written form. To solve this problem, recently the image-to-speech task was proposed, which generates spoken descriptions of images bypassing any text via an intermediate representation consisting of phonemes (image-to-phoneme). Here, we present a comprehensive study on the image-to-speech task in which, 1) several representative image-to-text generation methods are implemented for the image-to-phoneme task, 2) objective metrics are sought to evaluate the image-to-phoneme task, and 3) an end-to-end image-to-speech model that is able to synthesize spoken descriptions of images bypassing both text and phonemes is proposed. Extensive experiments are conducted on the public benchmark database Flickr8k. Results of our experiments demonstrate that 1) State-of-the-art image-to-text models can perform well on the image-to-phoneme task, and 2) several evaluation metrics, including BLEU3, BLEU4, BLEU5, and ROUGE-L can be used to evaluate image-to-phoneme performance. Finally, 3) end-to-end image-to-speech bypassing text and phonemes is feasible. Justin van der Hout, Jihua Zhu, Mark Hasegawa-Johnson, Odette Scharenborg |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | F0-Consistent Many-To-Many Non-Parallel Voice Conversion Via Conditional AutoencoderabstractNon-parallel many-to-many voice conversion remains an interesting but challenging speech processing task. Many style-transfer-inspired methods such as generative adversarial networks (GANs) and variational autoencoders (VAEs) have been proposed. Recently, AutoVC, a conditional autoencoders (CAEs) based method achieved state-of-the-art results by disentangling the speaker identity and speech content using information-constraining bottlenecks, and it achieves zero-shot conversion by swapping in a different speaker's identity embedding to synthesize a new voice. However, we found that while speaker identity is disentangled from speech content, a significant amount of prosodic information, such as source F0, leaks through the bottleneck, causing target F0 to fluctuate unnaturally. Furthermore, AutoVC has no control of the converted F0 and thus unsuitable for many applications. In the paper, we modified and improved autoencoder-based voice conversion to disentangle content, F0, and speaker identity at the same time. Therefore, we can control the F0 contour, generate speech with F0 consistent with the target speaker, and significantly improve quality and similarity. We support our improvement through quantitative and qualitative analysis. Kaizhi Qian, Zeyu Jin, Mark Hasegawa-Johnson, Gautham J. Mysore |
ICASSP | 3 |
| 2020 | Training Spoken Language Understanding Systems with Non-Parallel Speech and TextabstractEnd-to-end spoken language understanding (SLU) systems are typically trained on large amounts of data. In many practical scenarios, the amount of labeled speech is often limited as opposed to text. In this study, we investigate the use of non-parallel speech and text to improve the performance of dialog act recognition as an example SLU task. We propose a multiview architecture that can handle each modality separately. To effectively train on such data, this model enforces the internal speech and text encodings to be similar using a shared classifier. On the Switchboard Dialog Act corpus, we show that pretraining the classifier using large amounts of text helps learning better speech encodings, resulting in up to 40% relatively higher classification accuracies. We also show that when the speech embeddings from an automatic speech recognition (ASR) system are used in this framework, the speech-only accuracy exceeds the performance of ASR-text based tests up to 15% relative and approaches the performance of using true transcripts. Leda Sari, Samuel Thomas 0001, Mark Hasegawa-Johnson |
ICASSP | 3 |
| 2020 | Unsupervised Speech Decomposition via Triple Information BottleneckabstractSpeech information can be roughly decomposed into four components: language content, timbre, pitch, and rhythm. Obtaining disentangled representations of these components is useful in many speech analysis and generation applications. Recently, state-of-the-art voice conversion systems have led to speech representations that can disentangle speaker-dependent and independent information. However, these systems can only disentangle timbre, while information about pitch, rhythm and content is still mixed together. Further disentangling the remaining speech components is an under-determined problem in the absence of explicit annotations for each component, which are difficult and expensive to obtain. In this paper, we propose SpeechSplit, which can blindly decompose speech into its four components by introducing three carefully designed information bottlenecks. SpeechSplit is among the first algorithms that can separately perform style transfer on timbre, pitch and rhythm without text labels. Our code is publicly available at https://github.com/auspicious3000/SpeechSplit. Kaizhi Qian, Yang Zhang 0001, Shiyu Chang, Mark Hasegawa-Johnson, David D. Cox |
ICML | 4 |
| 2020 | Automatic Estimation of Intelligibility Measure for Consonants in SpeechabstractIn this article, we provide a model to estimate a real-valued measure of the intelligibility of individual speech segments. We trained regression models based on Convolutional Neural Networks (CNN) for stop consonants \textipa{/p,t,k,b,d,g/} associated with vowel \textipa{/A/}, to estimate the corresponding Signal to Noise Ratio (SNR) at which the Consonant-Vowel (CV) sound becomes intelligible for Normal Hearing (NH) ears. The intelligibility measure for each sound is called SNR$_{90}$, and is defined to be the SNR level at which human participants are able to recognize the consonant at least 90\% correctly, on average, as determined in prior experiments with NH subjects. Performance of the CNN is compared to a baseline prediction based on automatic speech recognition (ASR), specifically, a constant offset subtracted from the SNR at which the ASR becomes capable of correctly labeling the consonant. Compared to baseline, our models were able to accurately estimate the SNR$_{90}$~intelligibility measure with less than 2 [dB$^2$] Mean Squared Error (MSE) on average, while the baseline ASR-defined measure computes SNR$_{90}$~with a variance of 5.2 to 26.6 [dB$^2$], depending on the consonant. Ali Abavisani, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2020 | Evaluating Automatically Generated Phoneme Captions for ImagesabstractImage2Speech is the relatively new task of generating a spoken description of an image. This paper presents an investigation into the evaluation of this task. For this, first an Image2Speech system was implemented which generates image captions consisting of phoneme sequences. This system outperformed the original Image2Speech system on the Flickr8k corpus. Subsequently, these phoneme captions were converted into sentences of words. The captions were rated by human evaluators for their goodness of describing the image. Finally, several objective metric scores of the results were correlated with these human ratings. Although BLEU4 does not perfectly correlate with human ratings, it obtained the highest correlation among the investigated metrics, and is the best currently existing metric for the Image2Speech task. Current metrics are limited by the fact that they assume their input to be words. A more appropriate metric for the Image2Speech task should assume its input to be parts of words, i.e. phonemes, instead. Justin van der Hout, Zoltán D'Haese, Mark Hasegawa-Johnson, Odette Scharenborg |
INTERSPEECH | 3 |
| 2020 | Autosegmental Neural Nets: Should Phones and Tones be Synchronous or Asynchronous?abstractPhones, the segmental units of the International Phonetic Alphabet (IPA), are used for lexical distinctions in most human languages; Tones, the suprasegmental units of the IPA, are used in perhaps 70%. Many previous studies have explored cross-lingual adaptation of automatic speech recognition (ASR) phone models, but few have explored the multilingual and cross-lingual transfer of synchronization between phones and tones. In this paper, we test four Connectionist Temporal Classification (CTC)-based acoustic models, differing in the degree of synchrony they impose between phones and tones. Models are trained and tested multilingually in three languages, then adapted and tested cross-lingually in a fourth. Both synchronous and asynchronous models are effective in both multilingual and cross-lingual settings. Synchronous models achieve lower error rate in the joint phone+tone tier, but asynchronous training results in lower tone error rate. Jialu Li 0002, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2020 | Deep F-Measure Maximization for End-to-End Speech UnderstandingabstractSpoken language understanding (SLU) datasets, like many other machine learning datasets, usually suffer from the label imbalance problem. Label imbalance usually causes the learned model to replicate similar biases at the output which raises the issue of unfairness to the minority classes in the dataset. In this work, we approach the fairness problem by maximizing the F-measure instead of accuracy in neural network model training. We propose a differentiable approximation to the F-measure and train the network with this objective using standard backpropagation. We perform experiments on two standard fairness datasets, Adult, and Communities and Crime, and also on speech-to-intent detection on the ATIS dataset and speech-to-image concept classification on the Speech-COCO dataset. In all four of these tasks, F-measure maximization results in improved micro-F1 scores, with absolute improvements of up to 8% absolute, as compared to models trained with the cross-entropy loss function. In the two multi-class SLU tasks, the proposed approach significantly improves class coverage, i.e., the number of classes with positive recall. Leda Sari, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2020 | A DNN-HMM-DNN Hybrid Model for Discovering Word-Like Units from Spoken Captions and Image Regions
Liming Wang 0003, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2020 | That Sounds Familiar: An Analysis of Phonetic Representations Transfer Across LanguagesabstractOnly a handful of the world's languages are abundant with the resources that enable practical applications of speech processing technologies. One of the methods to overcome this problem is to use the resources existing in other languages to train a multilingual automatic speech recognition (ASR) model, which, intuitively, should learn some universal phonetic representations. In this work, we focus on gaining a deeper understanding of how general these representations might be, and how individual phones are getting improved in a multilingual setting. To that end, we select a phonetically diverse set of languages, and perform a series of monolingual, multilingual and crosslingual (zero-shot) experiments. The ASR is trained to recognize the International Phonetic Alphabet (IPA) token sequences. We observe significant improvements across all languages in the multilingual setting, and stark degradation in the crosslingual setting, where the model, among other errors, considers Javanese as a tone language. Notably, as little as 10 hours of the target language training data tremendously reduces ASR error rates. Our analysis uncovered that even the phones that are unique to a single language can benefit greatly from adding training data from other languages - an encouraging result for the low-resource speech community. Piotr Zelasko, Laureano Moro-Velázquez, Mark Hasegawa-Johnson, Odette Scharenborg, Najim Dehak |
INTERSPEECH | 3 |
| 2020 | Identify Speakers in Cocktail Parties with End-to-End AttentionabstractIn scenarios where multiple speakers talk at the same time, it is important to be able to identify the talkers accurately.This paper presents an end-to-end system that integrates speech source extraction and speaker identification, and proposes a new way to jointly optimize these two parts by max-pooling the speaker predictions along the channel dimension.Residual attention permits us to learn spectrogram masks that are optimized for the purpose of speaker identification, while residual forward connections permit dilated convolution with a sufficiently large context window to guarantee correct streaming across syllable boundaries.End-to-end training results in a system that recognizes one speaker in a two-speaker broadcast speech mixture with 99.9% accuracy and both speakers with 93.9% accuracy, and that recognizes all speakers in three-speaker scenarios with 81.2% accuracy. Junzhe Zhu, Mark Hasegawa-Johnson, Leda Sari |
INTERSPEECH | 2 |
| 2020 | Speech Technology for Unwritten LanguagesabstractSpeech technology plays an important role in our everyday life. Among others, speech is used for human-computer interaction, for instance for information retrieval and on-line shopping. In the case of an unwritten language, however, speech technology is unfortunately difficult to create, because it cannot be created by the standard combination of pre-trained speech-to-text and text-to-speech subsystems. The research presented in this article takes the first steps towards speech technology for unwritten languages. Specifically, the aim of this work was 1) to learn speech-to-meaning representations without using text as an intermediate representation, and 2) to test the sufficiency of the learned representations to regenerate speech or translated text, or to retrieve images that depict the meaning of an utterance in an unwritten language. The results suggest that building systems that go directly from speech-to-meaning and from meaning-to-speech, bypassing the need for text, is possible. Odette Scharenborg, Lucas Ondel Yang, Shruti Palaskar, Philip Arthur, Francesco Ciannella, Mingxing Du, Elin Larsen, Danny Merkx, Rachid Riad, Liming Wang 0003, Emmanuel Dupoux, Laurent Besacier, Alan W. Black, Mark Hasegawa-Johnson, Florian Metze, Graham Neubig, Sebastian Stüker, Pierre Godard, Markus Müller 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 14 |
| 2020 | Multimodal Word Discovery and Retrieval With Spoken Descriptions and Visual ConceptsabstractIn the absence of dictionaries, translators, or grammars, it is still possible to learn some of the words of a new language by listening to spoken descriptions of images. If several images, each containing a particular visually salient object, each co-occur with a particular sequence of speech sounds, we can infer that those speech sounds are a word whose definition is the visible object. A multimodal word discovery system accepts, as input, a database of spoken descriptions of images (or a set of corresponding phone transcriptions) and learns a mapping from waveform segments (or phone strings) to their associated image concepts. In this article, four multimodal word discovery systems are demonstrated: three models based on statistical machine translation (SMT) and one based on neural machine translation (NMT). The systems are trained with phonetic transcriptions, MFCC and multilingual bottleneck features (MBN). On the phone-level, the SMT outperforms the NMT model, achieving a 61.6% F1 score in the phone-level word discovery task on Flickr30k. On the audio-level, we compared our models with the existing ES-KMeans algorithm for word discovery and present some of the challenges in multimodal spoken word discovery. Liming Wang 0003, Mark Hasegawa-Johnson |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | When CTC Training Meets Acoustic LandmarksabstractConnectionist temporal classification (CTC) provides an end-to-end acoustic model (AM) training strategy. CTC learns accurate AMs without time-aligned phonetic transcription, but sometimes fails to converge, especially in resource-constrained scenarios. In this paper, the convergence properties of CTC are improved by incorporating acoustic landmarks. We tailored a new set of acoustic landmarks to help CTC training converge more rapidly and smoothly while also reducing recognition error rates. We leveraged new target label sequences mixed with both phone and manner changes to guide CTC training. Experiments on TIMIT demonstrated that CTC based acoustic models converge significantly faster and smoother when they are augmented by acoustic landmarks. The models pretrained with mixed target labels can be further finetuned, resulting in phone error rates 8.72% below baseline on TIMIT. Consistent performance gain is also observed on WSJ (a larger corpus) and reduced TIMIT (smaller). With WSJ, we are the first to succeed in verifying the effectiveness of acoustic landmark theory on a mid-sized ASR task. Di He 0004, Xuesong Yang, Boon Pang Lim, Mark Hasegawa-Johnson, Deming Chen |
ICASSP | 5 |
| 2019 | Dimensional Analysis of Laughter in Female Conversational SpeechabstractHow do people hear laughter in expressive, unprompted speech? What is the range of expressivity and function of laughter in this speech, and how can laughter inform the recognition of higher-level expressive dimensions in a corpus? This paper presents a scalable method for collecting natural human description of laughter, transforming the description to a vector of quantifiable laughter dimensions, and deriving baseline classifiers for the different dimensions of expressive laughter. Then, it explores the impact of leveraging nuances of laughter in the recognition of higher-level, general expressive dimensions, discovered in the same way, such as genuine happiness, sarcasm, nervous reflection, and more. The performance of the low-level laughter classifiers is presented, along with the performance of the high-level laughter-aware and laughter-unaware classifiers. Mary Pietrowicz, Carla Agurto, Jonah Casebeer, Mark Hasegawa-Johnson, Karrie Karahalios, Guillermo A. Cecchi |
ICASSP | 4 |
| 2019 | Pre-training of Speaker Embeddings for Low-latency Speaker Change Detection in Broadcast NewsabstractIn this work, we investigate pre-training of neural network based speaker embeddings for low-latency speaker change detection. Our proposed system takes two speech segments, generates embeddings using shared Siamese layers and then classifies the concatenated embeddings depending on whether they are spoken by the same speaker. We investigate gender classification, contrastive loss and triplet loss based pre-training of the embedding layers and also joint training of the embedding layers along with a same/different classifier. Training is performed on 2-second single speaker segments based on ground truth speaker segmentation of broadcast news data. However, during test, we use the detection system in a practical low-latency setting for annotating automatic closed captions. In contrast to training, test pairs are now created around automatic speech recognition (ASR) based segmentation boundaries. The ASR segments are often shorter than 2 seconds causing duration mismatch during testing. In our experiments, although the baseline i-vector based classifier performs well, the proposed triplet loss based pre-training followed by joint training provides 7-50% relative F-measure improvement in matched and mismatched conditions. In addition, the degradation in performance is less severe for network based embeddings as compared to using i-vectors in the variable duration test conditions. Leda Sari, Samuel Thomas 0001, Mark Hasegawa-Johnson, Michael Picheny |
ICASSP | 3 |
| 2019 | AutoVC: Zero-Shot Voice Style Transfer with Only Autoencoder LossabstractDespite the progress in voice conversion, many-to-many voice conversion trained on non-parallel data, as well as zero-shot voice conversion, remains under-explored. Deep style transfer algorithms, generative adversarial networks (GAN) in particular, are being applied as new solutions in this field. However, GAN training is very sophisticated and difficult, and there is no strong evidence that its generated speech is of good perceptual quality. In this paper, we propose a new style transfer scheme that involves only an autoencoder with a carefully designed bottleneck. We formally show that this scheme can achieve distribution-matching style transfer by training only on self-reconstruction loss. Based on this scheme, we proposed AutoVC, which achieves state-of-the-art results in many-to-many voice conversion with non-parallel data, and which is the first to perform zero-shot voice conversion. Kaizhi Qian, Yang Zhang 0001, Shiyu Chang, Xuesong Yang, Mark Hasegawa-Johnson |
ICML | 5 |
| 2019 | Study of the Performance of Automatic Speech Recognition Systems in Speakers with Parkinson's DiseaseabstractParkinson’s Disease (PD) affects motor capabilities of patients, who in some cases need to use human-computer assistive technologies to regain independence. The objective of this work is to study in detail the differences in error patterns from state-of-the-art Automatic Speech Recognition (ASR) systems on speech from people with and without PD. Two different speech recognizers (attention-based end-to-end and Deep Neural Network - Hidden Markov Models hybrid systems) were trained on a Spanish language corpus and subsequently tested on speech from 43 speakers with PD and 46 without PD. The differences related to error rates, substitutions, insertions and deletions of characters and phonetic units between the two groups were analyzed, showing that the word error rate is 27% higher in speakers with PD than in control speakers, with a moderated correlation between that rate and the developmental stage of the disease. The errors were related to all manner classes, and were more pronounced in the vowel /u/. This study is the first to evaluate ASR systems’ responses to speech from patients at different stages of PD in Spanish. The analyses showed general trends but individual speech deficits must be studied in the future when designing new ASR systems for this population. Laureano Moro-Velázquez, Shinji Watanabe 0001, Mark Hasegawa-Johnson, Odette Scharenborg, Najim Dehak |
INTERSPEECH | 4 |
| 2019 | Learning Speaker Aware Offsets for Speaker Adaptation of Neural Networks
Leda Sari, Samuel Thomas 0001, Mark Hasegawa-Johnson |
INTERSPEECH | 3 |
| 2019 | The Neural Correlates Underlying Lexically-Guided Perceptual LearningabstractThere is ample evidence showing that listeners are able to quickly adapt their phoneme classes to ambiguous sounds using a process called lexically-guided perceptual learning. This paper presents the first attempt to examine the neural correlates underlying this process. Specifically, we compared the brain’s responses to ambiguous [f/s] sounds in Dutch non-native listeners of English (N=36) before and after exposure to the ambiguous sound to induce learning, using Event-Related Potentials (ERPs). We identified a group of participants who showed lexically-guided perceptual learning in their phonetic categorization behavior as observed by a significant difference in /s/ responses between pretest and posttest and a group who did not. Moreover, we observed differences in mean ERP amplitude to ambiguous phonemes at pretest and posttest, shown by a reliable reduction in amplitude of a positivity over medial central channels from 250 to 550 ms. However, we observed no significant correlation between the size of behavioral and neural pre/posttest effects. Possibly, the observed behavioral and ERP differences between pretest and posttest link to different aspects of the sound classification task. In follow-up research, these differences will be further investigated by assessing their relationship to neural responses to the ambiguous sounds in the exposure phase. Odette Scharenborg, Jiska Koemans, Cybelle Smith, Mark Hasegawa-Johnson, Kara D. Federmeier |
INTERSPEECH | 4 |
| 2019 | Multimodal Word Discovery and Retrieval with Phone Sequence and Image Concepts
Liming Wang 0003, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2018 | Using Conversational Agents to Explain Medication Instructions to Older Adults
Renato Ferreira Leitão Azevedo, Daniel G. Morrow, James Graumlich, Ann Willemsen-Dunlap, Mark Hasegawa-Johnson, Thomas S. Huang, Kuangxiao Gu, Suma Bhat, Tarek Sakakini, Victor Sadauskas, Donald Halpin |
AMIA | 5 |
| 2018 | Recognizing Zero-Resourced Languages Based on Mismatched Machine TranscriptionsabstractMismatched crowdsourcing based probabilistic human transcription has been proposed recently for training and adapting acoustic models for zero-resourced languages where we do not have any native transcriptions. This paper describes a machine transcription based phone recognition system for recognizing zero-resourced languages and compares it with baseline systems of MAP adaptation and semi-supervised self training. With a set of available speech recognizers in source languages that cover all the basic phonetic features, this work shows that we can use mismatched machine transcriptions from these source languages to achieve human level transcriptions, bypassing the laborious efforts of obtaining human transcriptions. We also present a fully automated unsupervised approach for zero-resourced speech recognition using mismatched machine transcriptions for transfer learning of phone models. Wenda Chen, Mark Hasegawa-Johnson, Nancy F. Chen |
ICASSP | 2 |
| 2018 | Time-Frequency Networks for Audio Super-ResolutionabstractAudio super-resolution (a.k.a. bandwidth extension) is the challenging task of increasing the temporal resolution of audio signals. Recent deep networks approaches achieved promising results by modeling the task as a regression problem in either time or frequency domain. In this paper, we introduced Time-Frequency Network (TFNet), a deep network that utilizes supervision in both the time and frequency domain. We proposed a novel model architecture which allows the two domains to be jointly optimized. Results demonstrate that our method outperforms the state-of-the-art both quantitatively and qualitatively. Teck-Yian Lim, Raymond A. Yeh, Yijia Xu, Minh N. Do, Mark Hasegawa-Johnson |
ICASSP | 5 |
| 2018 | Bayesian Models for Unit Discovery on a Very Low Resource LanguageabstractDeveloping speech technologies for low-resource languages has become a very active research field over the last decade. Among others, Bayesian models have shown some promising results on artificial examples but still lack of in situ experiments. Our work applies state-of-the-art Bayesian models to unsupervised Acoustic Unit Discovery (AUD) in a real low-resource language scenario. We also show that Bayesian models can naturally integrate information from other resourceful languages by means of informative prior leading to more consistent discovered units. Finally, discovered acoustic units are used, either as the I-best sequence or as a lattice, to perform word segmentation. Word segmentation results show that this Bayesian approach clearly outperforms a Segmental-DTW baseline on the same corpus. Lucas Ondel Yang, Pierre Godard, Laurent Besacier, Elin Larsen, Mark Hasegawa-Johnson, Odette Scharenborg, Emmanuel Dupoux, Lukás Burget, François Yvon, Sanjeev Khudanpur |
ICASSP | 5 |
| 2018 | Deep Learning Based Speech BeamformingabstractMulti-channel speech enhancement with ad-hoc sensors has been a challenging task. Speech model guided beamforming algorithms are able to recover natural sounding speech, but the speech models tend to be oversimplified or the inference would otherwise be too complicated. On the other hand, deep learning based enhancement approaches are able to learn complicated speech distributions and perform efficient inference, but they are unable to deal with variable number of input channels. Also, deep learning approaches introduce a lot of errors, particularly in the presence of unseen noise types and settings. We have therefore proposed an enhancement framework called DEEPBEAM, which combines the two complementary classes of algorithms. DEEPBEAM introduces a beamforming filter to produce natural sounding speech, but the filter coefficients are determined with the help of a monaural speech enhancement neural network. Experiments on synthetic and real-world data show that DEEPBEAM is able to produce clean, dry and natural sounding speech, and is robust against unseen noise. Kaizhi Qian, Yang Zhang 0001, Shiyu Chang, Xuesong Yang, Dinei A. F. Florêncio, Mark Hasegawa-Johnson |
ICASSP | 6 |
| 2018 | Linguistic Unit Discovery from Multi-Modal Inputs in Unwritten Languages: Summary of the "Speaking Rosetta" JSALT 2017 WorkshopabstractWe summarize the accomplishments of a multi-disciplinary workshop exploring the computational and scientific issues surrounding the discovery of linguistic units (subwords and words) in a language without orthography. We study the replacement of orthographic transcriptions by images and/or translated text in a well-resourced language to help unsupervised discovery from raw speech. Odette Scharenborg, Laurent Besacier, Alan W. Black, Mark Hasegawa-Johnson, Florian Metze, Graham Neubig, Sebastian Stüker, Pierre Godard, Markus Müller 0001, Lucas Ondel Yang, Shruti Palaskar, Philip Arthur, Francesco Ciannella, Mingxing Du, Elin Larsen, Danny Merkx, Rachid Riad, Liming Wang 0003, Emmanuel Dupoux |
ICASSP | 4 |
| 2018 | Joint Modeling of Accents and Acoustics for Multi-Accent Speech RecognitionabstractThe performance of automatic speech recognition systems degrades with increasing mismatch between the training and testing scenarios. Differences in speaker accents are a significant source of such mismatch. The traditional approach to deal with multiple accents involves pooling data from several accents during training and building a single model in multi-task fashion, where tasks correspond to individual accents. In this paper, we explore an alternate model where we jointly learn an accent classifier and a multi-task acoustic model. Experiments on the American English Wall Street Journal and British English Cambridge corpora demonstrate that our joint model outperforms the strong multi-task acoustic model baseline. We obtain a 5.94% relative improvement in word error rate on British English, and 9.47% relative improvement on American English. This illustrates that jointly modeling with accent information improves acoustic model performance. Xuesong Yang, Kartik Audhkhasi, Andrew Rosenberg, Samuel Thomas 0001, Bhuvana Ramabhadran, Mark Hasegawa-Johnson |
ICASSP | 6 |
| 2018 | Image Restoration with Deep Generative ModelsabstractMany image restoration problems are ill-posed in nature, hence, beyond the input image, most existing methods rely on a carefully engineered image prior, which enforces some local image consistency in the recovered image. How tightly the prior assumptions are fulfilled has a big impact on the resulting task performance. To obtain more flexibility, in this work, we proposed to design the image prior in a data-driven manner. Instead of explicitly defining the prior, we learn it using deep generative models. We demonstrate that this learned prior can be applied to many image restoration problems using an unified framework. Raymond A. Yeh, Teck-Yian Lim, Chen Chen 0003, Alexander G. Schwing, Mark Hasegawa-Johnson, Minh N. Do |
ICASSP | 5 |
| 2018 | Topic and Keyword Identification for Low-resourced Speech Using Cross-Language Transfer Learning
Wenda Chen, Mark Hasegawa-Johnson, Nancy F. Chen |
INTERSPEECH | 2 |
| 2018 | Improving DNNs Trained with Non-Native Transcriptions Using Knowledge Distillation and Target Interpolation
Amit Das 0007, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2018 | Improved ASR for Under-resourced Languages through Multi-task Learning with Acoustic LandmarksabstractFurui first demonstrated that the identity of both consonant and vowel can be perceived from the C-V transition; later, Stevens proposed that acoustic landmarks are the primary cues for speech perception, and that steady-state regions are secondary or supplemental. Acoustic landmarks are perceptually salient, even in a language one doesn't speak, and it has been demonstrated that non-speakers of the language can identify features such as the primary articulator of the landmark. These factors suggest a strategy for developing language-independent automatic speech recognition: landmarks can potentially be learned once from a suitably labeled corpus and rapidly applied to many other languages. This paper proposes enhancing the cross-lingual portability of a neural network by using landmarks as the secondary task in multi-task learning (MTL). The network is trained in a well-resourced source language with both phone and landmark labels (English), then adapted to an under-resourced target language with only word labels (Iban). Landmark-tasked MTL reduces source-language phone error rate by 2.9% relative, and reduces target-language word error rate by 1.9%-5.9% depending on the amount of target-language training data. These results suggest that landmark-tasked MTL causes the DNN to learn hidden-node features that are useful for cross-lingual adaptation. Di He 0004, Boon Pang Lim, Xuesong Yang, Mark Hasegawa-Johnson, Deming Chen |
INTERSPEECH | 4 |
| 2018 | Speaker Adaptive Audio-Visual Fusion for the Open-Vocabulary Section of AVICAR
Leda Sari, Mark Hasegawa-Johnson, Kumaran S, Georg Stemmer, Krishnakumar N. Nair |
INTERSPEECH | 2 |
| 2018 | Visualizing Phoneme Category Adaptation in Deep Neural NetworksabstractBoth human listeners and machines need to adapt their sound categories whenever a new speaker is encountered. This perceptual learning is driven by lexical information. The aim of this paper is two-fold: investigate whether a deep neural network-based (DNN) ASR system can adapt to only a few examples of ambiguous speech as humans have been found to do; investigate a DNN’s ability to serve as a model of human perceptual learning. Crucially, we do so by looking at intermediate levels of phoneme category adaptation rather than at the output level. We visualize the activations in the hidden layers of the DNN during perceptual learning. The results show that, similar to humans, DNN systems learn speaker-adapted phone category boundaries from a few labeled examples. The DNN adapts its category boundaries not only by adapting the weights of the output layer, but also by adapting the implicit feature maps computed by the hidden layers, suggesting the possibility that human perceptual learning might involve a similar nonlinear distortion of a perceptual space that is intermediate between the acoustic input and the phonological categories. Comparisons between DNNs and humans can thus provide valuable insights into the way humans process speech and improve ASR technology. Odette Scharenborg, Sebastian Tiesmeyer, Mark Hasegawa-Johnson, Najim Dehak |
INTERSPEECH | 3 |
| 2018 | Infant Emotional Outbursts Detection in Infant-parent Spoken Interactions
Yijia Xu, Mark Hasegawa-Johnson, Nancy McElwain |
INTERSPEECH | 2 |
| 2018 | Multitask Learning for Phone Recognition of Underresourced Languages Using Mismatched TranscriptionabstractIt is challenging to obtain large amounts of native (matched) labels for speech audio in underresourced languages. This challenge is often due to a lack of literate speakers of the language, or in extreme cases, a lack of universally acknowledged orthography as well. One solution is to increase the amount of labeled data by using mismatched transcription, which employs transcribers who do not speak the underresourced language of interest called the target language (in place of native speakers), to transcribe what they hear as nonsense speech in their own annotation language (≠ target language). Previous uses of mismatched transcription converted it to a probabilistic transcription (PT), but PT is limited by the errors of nonnative perception. This paper proposes, instead, a multitask learning framework in which one deep neural network (DNN) is trained to optimize two separate tasks: acoustic modeling of a small number of matched transcription with matched target-language graphemes; and acoustic modeling of a large number of mismatched transcription with mismatched annotation-language graphemes. We find that: first, the multitask learning framework gives significant improvement over monolingual, semisupervised learning, multilingual DNN training, and transfer learning baselines; second, a Gaussian Mixture Model-Hidden-Markov Model (GMM-HMM) model adapted using PT improves alignments, thereby improving training; and third, bottleneck features trained on the mismatched transcriptions lead to even better alignments, resulting in further performance gains of the multitask DNN. Our experiments are conducted on the IARPA Georgian and Vietnamese BABEL corpora as well as on our newly collected speech corpus of Singapore Hokkien, an underresourced language with no standard written form. Van Hai Do, Nancy F. Chen, Boon Pang Lim, Mark Hasegawa-Johnson |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2017 | Using Computer Agents to Explain Clinical Test Results
Renato Ferreira Leitão Azevedo, Kuangxiao Gu, Yang Zhang 0001, Victor Sadauskas, Tarek Sakakini, Daniel G. Morrow, Mark Hasegawa-Johnson, Thomas S. Huang, Suma Bhat, Ann Willemsen-Dunlap, Donald Halpin, James Graumlich, William Schuh |
AMIA | 7 |
| 2017 | Dr. Babel Fish: A Machine Translator to Simplify Providers' Language
Tarek Sakakini, Renato Ferreira Leitão Azevedo, Victor Sadauskas, Kuangxiao Gu, Yang Zhang 0001, Suma Bhat, Daniel G. Morrow, Mark Hasegawa-Johnson, Thomas S. Huang, Ann Willemsen-Dunlap, Donald Halpin, James Graumlich |
AMIA | 8 |
| 2017 | Semantic Image Inpainting with Deep Generative ModelsabstractSemantic image inpainting is a challenging task where large missing regions have to be filled based on the available visual data. Existing methods which extract information from only a single image generally produce unsatisfactory results due to the lack of high level context. In this paper, we propose a novel method for semantic image inpainting, which generates the missing content by conditioning on the available data. Given a trained generative model, we search for the closest encoding of the corrupted image in the latent image manifold using our context and prior losses. This encoding is then passed through the generative model to infer the missing content. In our method, inference is possible irrespective of how the missing content is structured, while the state-of-the-art learning based method requires specific information about the holes in the training phase. Experiments on three datasets show that our method successfully predicts information in large missing regions and achieves pixel-level photorealism, significantly outperforming the state-of-the-art methods. Raymond A. Yeh, Chen Chen 0003, Teck-Yian Lim, Alexander G. Schwing, Mark Hasegawa-Johnson, Minh N. Do |
CVPR | 5 |
| 2017 | Low-resource grapheme-to-phoneme conversion using recurrent neural networksabstractGrapheme-to-phoneme (G2P) conversion is an important problem for many speech and language processing applications. G2P models are particularly useful for low-resource languages that do not have well-developed pronunciation lexicons. Prominent G2P paradigms are based on initial alignments between grapheme and phoneme sequences. In this work, we devise new alignment strategies that work effectively with recurrent neural network based models when only a small number of pronunciations are available to train the models. In a small data setting, we build G2P models for Pashto, Tagalog and Lithuanian that significantly outperform a joint sequence model and a baseline recurrent neural network based model, giving up to 14% and 9% relative reductions in phone and word error rates when trained on a dataset of 250 words. Preethi Jyothi, Mark Hasegawa-Johnson |
ICASSP | 2 |
| 2017 | Discovering dimensions of perceived vocal expression in semi-structured, unscripted oral history accountsabstractWhat do people hear in expressive, unprompted speech? And how can their descriptions be transformed into a representative set of dimensions of vocal expression? This paper presents a methodology for collecting user description of vocal expression, transforms the user descriptions into a set of measurable expressive dimensions, and derives a representative feature set and baseline classifiers across these dimensions. The resulting classifiers recognized the top 13 dimensions over an oral history corpus, with a maximum unweighted recall score of 80.5%. Mary Pietrowicz, Mark Hasegawa-Johnson, Karrie Karahalios |
ICASSP | 2 |
| 2017 | Mismatched Crowdsourcing from Multiple Annotator Languages for Recognizing Zero-Resourced Languages: A Nullspace Clustering Approach
Wenda Chen, Mark Hasegawa-Johnson, Nancy F. Chen, Boon Pang Lim |
INTERSPEECH | 2 |
| 2017 | Deep Auto-Encoder Based Multi-Task Learning Using Probabilistic Transcriptions
Amit Das 0007, Mark Hasegawa-Johnson, Karel Veselý |
INTERSPEECH | 2 |
| 2017 | Multi-Task Learning Using Mismatched Transcription for Under-Resourced Speech Recognition
Van Hai Do, Nancy F. Chen, Boon Pang Lim, Mark Hasegawa-Johnson |
INTERSPEECH | 4 |
| 2017 | Using Approximated Auditory Roughness as a Pre-Filtering Feature for Human Screaming and Affective Speech AED
Di He 0004, Zuofu Cheng, Mark Hasegawa-Johnson, Deming Chen |
INTERSPEECH | 3 |
| 2017 | Team ELISA System for DARPA LORELEI Speech Evaluation 2016
Pavlos Papadopoulos, Ruchir Travadi, Colin Vaz, Nikos Malandrakis, Ulf Hermjakob, Nima Pourdamghani, Michael Pust, Boliang Zhang, Xiaoman Pan, Di Lu 0003, Ondrej Glembek, Murali Karthick Baskar, Martin Karafiát, Lukás Burget, Mark Hasegawa-Johnson, Heng Ji 0001, Jonathan May, Kevin Knight, Shri Narayanan |
INTERSPEECH | 16 |
| 2017 | Speech Enhancement Using Bayesian Wavenet
Kaizhi Qian, Yang Zhang 0001, Shiyu Chang, Xuesong Yang, Dinei A. F. Florêncio, Mark Hasegawa-Johnson |
INTERSPEECH | 6 |
| 2017 | Glottal Model Based Speech Beamforming for ad-hoc Microphone Arrays
Yang Zhang 0001, Dinei A. F. Florêncio, Mark Hasegawa-Johnson |
INTERSPEECH | 3 |
| 2017 | Dilated Recurrent Neural NetworksabstractLearning with recurrent neural networks (RNNs) on long sequences is a notoriously difficult task. There are three major challenges: 1) complex dependencies, 2) vanishing and exploding gradients, and 3) efficient parallelization. In this paper, we introduce a simple yet effective RNN connection structure, the DilatedRNN, which simultaneously tackles all of these challenges. The proposed architecture is characterized by multi-resolution dilated recurrent skip connections and can be combined flexibly with diverse RNN cells. Moreover, the DilatedRNN reduces the number of parameters needed and enhances training efficiency significantly, while matching state-of-the-art performance (even with standard RNN cells) in tasks involving very long-term dependencies. To provide a theory-based quantification of the architecture's advantages, we introduce a memory capacity measure, the mean recurrent length, which is more suitable for RNNs with long skip connections than existing measures. We rigorously prove the advantages of the DilatedRNN over other recurrent neural architectures. The code for our method is publicly available at https://github.com/code-terminator/DilatedRNN. Shiyu Chang, Yang Zhang 0001, Wei Han 0002, Mo Yu, Wei Tan 0001, Michael Witbrock, Mark Hasegawa-Johnson, Thomas S. Huang |
NIPS | 9 |
| 2017 | Streaming Recommender SystemsabstractThe increasing popularity of real-world recommender systems produces data continuously and rapidly, and it becomes more realistic to study recommender systems under streaming scenarios. Data streams present distinct properties such as temporally ordered, continuous and high-velocity, which poses tremendous challenges to traditional recommender systems. In this paper, we investigate the problem of recommendation with stream inputs. In particular, we provide a principled framework termed sRec, which provides explicit continuous-time random process models of the creation of users and topics, and of the evolution of their interests. A variational Bayesian approach called recursive meanfield approximation is proposed, which permits computationally efficient instantaneous on-line inference. Experimental results on several real-world datasets demonstrate the advantages of our sRec over other state-of-the-arts. Shiyu Chang, Yang Zhang 0001, Jiliang Tang, Dawei Yin 0001, Yi Chang 0001, Mark Hasegawa-Johnson, Thomas S. Huang |
WWW | 6 |
| 2017 | A multidisciplinary approach to designing and evaluating Electronic Medical Record portal messages that support patient self-care
Daniel G. Morrow, Mark Hasegawa-Johnson, Thomas S. Huang, William Schuh, Renato Ferreira Leitão Azevedo, Kuangxiao Gu, Yang Zhang 0001, Bidisha Roy, Rocío García-Retamero |
J. Biomed. Informatics | 2 |
| 2017 | ASR for Under-Resourced Languages From Probabilistic TranscriptionabstractIn many under-resourced languages it is possible to find text, and it is possible to find speech, but transcribed speech suitable for training automatic speech recognition (ASR) is unavailable. In the absence of native transcripts, this paper proposes the use of a probabilistic transcript: A probability mass function over possible phonetic transcripts of the waveform. Three sources of probabilistic transcripts are demonstrated. First, self-training is a well-established semisupervised learning technique, in which a cross-lingual ASR first labels unlabeled speech, and is then adapted using the same labels. Second, mismatched crowdsourcing is a recent technique in which nonspeakers of the language are asked to write what they hear, and their nonsense transcripts are decoded using noisy channel models of second-language speech perception. Third, EEG distribution coding is a new technique in which nonspeakers of the language listen to it, and their electrocortical response signals are interpreted to indicate probabilities. ASR was trained in four languages without native transcripts. Adaptation using mismatched crowdsourcing significantly outperformed self-training, and both significantly outperformed a cross-lingual baseline. Both EEG distribution coding and text-derived phone language models were shown to improve the quality of probabilistic transcripts derived from mismatched crowdsourcing. Mark Hasegawa-Johnson, Preethi Jyothi, Daniel McCloy, Majid Mirbagheri, Giovanni M. Di Liberto, Amit Das 0007, Bradley Ekin, Chunxi Liu, Vimal Manohar, Hao Tang 0002, Edmund C. Lalor, Nancy F. Chen, Paul Hager, Tyler Kekona, Rose Sloan, Adrian K. C. Lee |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2016 | Adapting ASR for under-resourced languages using mismatched transcriptionsabstractMismatched transcriptions of speech in a target language refers to transcriptions provided by people unfamiliar with the language, using English letter sequences. In this work, we demonstrate the value of such transcriptions in building an ASR system for the target language. For different languages, we use less than an hour of mismatched transcriptions to successfully adapt baseline multilingual models built with no access to native transcriptions in the target language. The adapted models provide up to 25% relative improvement in phone error rates on an unseen evaluation set. Chunxi Liu, Preethi Jyothi, Hao Tang 0002, Vimal Manohar, Rose Sloan, Tyler Kekona, Mark Hasegawa-Johnson, Sanjeev Khudanpur |
ICASSP | 7 |
| 2016 | Landmark of Mandarin nasal codas and its application in pronunciation error detectionabstractL2 learners of Mandarin have difficulty learning native-like pronunciation of nasal codas. In order to help them learn native-like pronunciation, we propose to develop targeted classifiers for automatic pronunciation error detection. In this paper, perceptual experiments with modified speech are designed to analyze the exact position of the landmark of a nasal coda. Based on perceptual results from isolated words, we propose that information about nasal coda place of articulation is most dense near a landmark at the center of the nasalized vowel. Landmarks detected in a database of Japanese learners of Mandarin, and classified as correct vs. incorrect using an SVM. The result shows that the detection performance of the SVM+Landmark system is similar to that of a DNN-HMM+MFCC system. When the two systems are combined, an FRR of 4.6% is achieved at DA of 83.9%. This performance is comparable to that of previously developed classifiers for 16 common Mandarin pronunciation errors. Yanlu Xie, Mark Hasegawa-Johnson, Leyuan Qu, Jinsong Zhang 0001 |
ICASSP | 2 |
| 2016 | Stable and symmetric filter convolutional neural networkabstractFirst we present a proof that convolutional neural networks (CNN) with max-norm regularization, max-pooling, and Relu non-linearity are stable to additive noise. Second, we explore the use of symmetric and antisymmetric filters in a baseline CNN model on digit classification, which enjoys the stability to additive noise. Experimental results indicate that the symmetric CNN outperforms the baseline model for nearly all training sizes and matches the state-of-the-art deep-net in the cases of limited training examples. Raymond A. Yeh, Mark Hasegawa-Johnson, Minh N. Do |
ICASSP | 2 |
| 2016 | An Investigation on Training Deep Neural Networks Using Probabilistic Transcriptions
Amit Das 0007, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2016 | Automatic Speech Recognition Using Probabilistic Transcriptions in Swahili, Amharic, and Dinka
Amit Das 0007, Preethi Jyothi, Mark Hasegawa-Johnson |
INTERSPEECH | 3 |
| 2016 | Analysis of Mismatched Transcriptions Generated by Humans and Machines for Under-Resourced Languages
Van Hai Do, Nancy F. Chen, Boon Pang Lim, Mark Hasegawa-Johnson |
INTERSPEECH | 4 |
| 2016 | Positive-Unlabeled Learning in Streaming NetworksabstractData of many problems in real-world systems such as link prediction and one-class recommendation share common characteristics. First, data are in the form of positive unlabeled (PU) measurements (e.g. Twitter "following", Facebook "like", etc.) that do not provide negative information, which can be naturally represented as networks. Second, in the era of big data, such data are generated temporally-ordered, continuously and rapidly, which determines its streaming nature. These common characteristics allow us to unify many problems into a novel framework -- PU learning in streaming networks. In this paper, a principled probabilistic approach SPU is proposed to leverage the characteristics of the streaming PU inputs. In particular, SPU captures temporal dynamics and provides real-time adaptations and predictions by identifying the potential negative signals concealed in unlabeled data. Our empirical results on various real-world datasets demonstrate the effectiveness of the proposed framework over other state-of-the-art methods in both link prediction and recommendation. Shiyu Chang, Yang Zhang 0001, Jiliang Tang, Dawei Yin 0001, Yi Chang 0001, Mark Hasegawa-Johnson, Thomas S. Huang |
KDD | 6 |
| 2016 | Speech Production in Speech Technologies: Introduction to the CSL Special Issue
Karen Livescu, Frank Rudzicz, Eric Fosler-Lussier, Mark Hasegawa-Johnson, Jeff A. Bilmes |
Comput. Speech Lang. | 4 |
| 2015 | Acquiring Speech Transcriptions Using Mismatched CrowdsourcingabstractTranscribed speech is a critical resource for building statistical speech recognition systems. Recent work has looked towards soliciting transcriptions for large speech corpora from native speakers of the language using crowdsourcing techniques. However, native speakers of the target language may not be readily available for crowdsourcing. We examine the following question: can humans unfamiliar with the target language help transcribe? We follow an information-theoretic approach to this problem: (1) We learn the characteristics of a noisy channel that models the transcribers' systematic perception biases. (2) We use an error-correcting code, specifically a repetition code, to encode the inputs to this channel, in conjunction with a maximum-likelihood decoding rule. To demonstrate the feasibility of this approach, we transcribe isolated Hindi words with the help of Mechanical Turk workers unfamiliar with Hindi. We successfully recover Hindi words with an accuracy of over 85% (and 94% in a 4-best list) using a 15-fold repetition code. We also estimate the conditional entropy of the input to this channel (Hindi words) given the channel output (transcripts from crowdsourced workers) to be less than 2 bits; this serves as a theoretical estimate of the average number of bits of auxiliary information required for errorless recovery. Preethi Jyothi, Mark Hasegawa-Johnson |
AAAI | 2 |
| 2015 | Multichannel transient acoustic signal classification using task-driven dictionary with joint sparsity and beamformingabstractWe are interested in a multichannel transient acoustic signal classification task which suffers from additive/convolutionary noise corruption. To address this problem, we propose a double-scheme classifier that takes the advantage of multichannel data to improve noise robustness. Both schemes adopt task-driven dictionary learning as the basic framework, and exploit multichannel data at different levels - scheme 1 imposes joint sparsity constraint while learning the dictionary and classifier; scheme 2 adopts beamforming at signal formation level. In addition, matched filter and robust ceptral coefficients are applied to improve noise robustness of the input feature. Experiments show that the proposed classifier significantly outperforms the baseline algorithms. Yang Zhang 0001, Nasser M. Nasrabadi, Mark Hasegawa-Johnson |
ICASSP | 3 |
| 2015 | Cross-lingual transfer learning during supervised training in low resource scenariosabstractIn this study, transfer learning techniques are presented for cross-lingual speech recognition to mitigate the effects of limited availability of data in a target language using data from richly resourced source languages. First, a maximum likelihood (ML) based regularization criterion is used to learn context-dependent Gaussian mixture model (GMM) based hidden Markov model (HMM) parameters for phones in target language using data from both target and source languages. Recognition results indicate improved HMM state alignments. Second, the hidden layers of a deep neural network (DNN) are initialized using unsupervised pre-training of a multilingual deep belief network (DBN). The DNN is fine-tuned jointly using a modified cross entropy criterion that uses HMM state alignments from both target and source languages. Third, another DNN fine-tuning technique is explored where the training is performed in a sequential manner source language followed by the target language. Experiments conducted using varying amounts of target data indicate further improvements in performance can be obtained using joint and sequential training of the DNN compared to existing techniques. Turkish and English were chosen to be the target and source languages respectively. Amit Das 0007, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2015 | Transcribing continuous speech using mismatched crowdsourcingabstractMismatched crowdsourcing was recently proposed as a poten-tial approach to deriving moderately accurate speech transcrip-tions using crowd workers unfamiliar with the language be-ing spoken. In introducing this approach, we demonstrated its promise with the help of an isolated word recovery task for Hindi. However, it remained open whether mismatched crowd sourcing can yield non-trivial accuracy in a continuous speech task. In this work, we focus on this question and demonstrate a word error rate of under 45 % in a large-vocabulary task (again for Hindi). In achieving this, we develop several new tech-niques capable of scaling effectively to continuous speech. We also provide an information theoretic analysis and estimate the amount of information lost in transcription by the mismatched crowd workers to be under 5 bits. Preethi Jyothi, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2015 | Improved hindi broadcast ASR by adapting the language model and pronunciation model using a priori syntactic and morphophonemic knowledgeabstractIn this work, we present a new large-vocabulary, broadcast news ASR system for Hindi. Since Hindi has a largely phone-mic orthography, the pronunciation model was automatically generated from text. We experiment with several variants of this model and study the effect of incorporating word bound-ary information with these models. We also experiment with knowledge-based adaptations to the language model in Hindi, derived in an unsupervised manner, that lead to small im-provements in word error rate (WER). Our experiments were conducted on a new corpus assembled from publicly-available Hindi news broadcasts. We evaluate our techniques on an open-vocabulary task and obtain competitive WERs on an unseen test set. Index Terms: Hindi LVCSR system, Broadcast news ASR, Grapheme and phoneme-based models, Knowledge-based language-model adaptation 1. Preethi Jyothi, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2015 | Acoustic correlates for perceived effort levels in expressive speechabstractActors and other vocal performers vary their speech across the continuum of vocal effort to express ideas, emphasize thoughts, communicate emotions, and create drama. They are experts at vocal expression. To analyze this range of expression across effort levels, we curated a corpus of professional actors ’ Hamlet soliloquy performances and present an acoustic feature set and classification model suitable for tracking actors ’ expressive speech from extreme to extreme – from whispered, to breathy, through modal, to resonant speech. Mary Pietrowicz, Mark Hasegawa-Johnson, Karrie Karahalios |
INTERSPEECH | 2 |
| 2015 | Joint Optimization of Masks and Deep Recurrent Neural Networks for Monaural Source SeparationabstractMonaural source separation is important for many real world applications. It is challenging because, with only a single channel of information available, without any constraints, an infinite number of solutions are possible. In this paper, we explore joint optimization of masking functions and deep recurrent neural networks for monaural source separation tasks, including speech separation, singing voice separation, and speech denoising. The joint optimization of the deep recurrent neural networks with an extra masking layer enforces a reconstruction constraint. Moreover, we explore a discriminative criterion for training neural networks to further enhance the separation performance. We evaluate the proposed system on the TSP, MIR-1K, and TIMIT datasets for speech separation, singing voice separation, and speech denoising tasks, respectively. Our approaches achieve 2.30-4.98 dB SDR gain compared to NMF models in the speech separation task, 2.30-2.48 dB GNSDR gain and 4.32-5.42 dB GSIR gain compared to existing models in the singing voice separation task, and outperform NMF and DNN baselines in the speech denoising task. Po-Sen Huang, Minje Kim 0001, Mark Hasegawa-Johnson, Paris Smaragdis |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | A PAC-Bayesian Approach to Minimum Perplexity Language Modeling
Sujeeth Bharadwaj, Mark Hasegawa-Johnson |
COLING | 2 |
| 2014 | Deep learning for monaural speech separationabstractMonaural source separation is useful for many real-world applications though it is a challenging problem. In this paper, we study deep learning for monaural speech separation. We propose the joint optimization of the deep learning models (deep neural networks and recurrent neural networks) with an extra masking layer, which enforces a reconstruction constraint. Moreover, we explore a discriminative training criterion for the neural networks to further enhance the separation performance. We evaluate our approaches using the TIMIT speech corpus for a monaural speech separation task. Our proposed models achieve about 3.8∼4.9 dB SIR gain compared to NMF models, while maintaining better SDRs and SARs. Po-Sen Huang, Minje Kim 0001, Mark Hasegawa-Johnson, Paris Smaragdis |
ICASSP | 3 |
| 2014 | Improvement of Probabilistic Acoustic Tube model for speech decompositionabstractCurrent model-based speech analysis tends to be incomplete - only a part of parameters of interest (e.g. only the pitch or vocal tract) are modeled, while the rest that might as well be important are disregarded. The drawback is that without joint modeling of parameters that are correlated, the analysis on speech parameters may be inaccurate or even incorrect. Under this motivation, we have proposed such a model called PAT (Probabilistic Acoustic Tube), where pitch, vocal tract and energy are jointly modeled. This paper proposes an improved version of PAT model, named PAT2, where both signal and probabilistic modeling are tremendously renovated. Compared to related works, PAT2 is much more comprehensive, which incorporates mixed excitation, glottal wave and phase modeling. Experimental results show its ability in decomposing speech into desirable parameters and its potential for speech synthesis. Yang Zhang 0001, Zhijian Ou, Mark Hasegawa-Johnson |
ICASSP | 3 |
| 2014 | Foreground object detection in highly dynamic scenes using saliencyabstractIn this paper, we propose a novel saliency-based algorithm to detect foreground regions in highly dynamic scenes. We first convert input video frames to multiple patch-based feature maps. Then, we apply temporal saliency analysis to the pixels of each feature map. For each temporal set of co-located pixels, the feature distance of a point from its kthnearest neighbor is used to compute the temporal saliency. By computing and combining temporal saliency maps of different features, we obtain foreground likelihood maps. A simple segmentation method based on adaptive thresholding is applied to detect the foreground objects. We test our algorithm on images sequences of dynamic scenes, including public datasets and a new challenging wildlife dataset we constructed. The experimental results demonstrate the proposed algorithm achieves state-of-the-art results. Kai-Hsiang Lin, Pooya Khorrami, Jiangping Wang, Mark Hasegawa-Johnson, Thomas S. Huang |
ICIP | 4 |
| 2014 | An iterative approach to decision tree training for context dependent speech synthesisabstractIn speech synthesis with sparse training data, phonetic decision trees are frequently used for balance between model complexity and avail-able data. The traditional training procedure is that decision trees are constructed after parameters for each phones optimized in the EM al-gorithm. This paper proposes an iterative re-optimization algorithm in which the decision tree is re-learned after every iteration of the EM algorithm. The performance of the new procedure is compared with the original procedure by training parameters for MFCC and F0 features using an EDHMM model with data from The Boston Uni-versity Radio Speech corpus. A convergence proof is presented, and experimental tests demonstrate that iterative re-optimization gener-ates statistically significant test corpus log-likelihood improvements. Index Terms — speech synthesis, speech clustering, EM algo-rithm, decision tree 1. Xiayu Chen, Yang Zhang 0001, Mark Hasegawa-Johnson |
INTERSPEECH | 3 |
| 2014 | Detecting articulatory compensation in acoustic data through linear regression modelingabstractExamining articulatory compensation has been important in understanding how the speech production system is organized, and how it relates to the acoustic and ultimately phonological levels. This paper offers a method that detects articulatory compensation in the acoustic signal, which is based on linear regression modeling of co-variation patterns between acoustic cues. We demonstrate the method on selected acoustic cues for spontaneously produced American English stop consonants. Compensatory patterns of cue variation were observed for voiced stops in some cue pairs, while uniform patterns of cue variation were found for stops as a function of place of articulation or position in the word. Overall, the results suggest that this method can be useful for observing articulatory strategies indirectly from acoustic data and testing hypotheses about the conditions under which articulatory compensation is most likely. Alina Khasanova, Jennifer Cole 0001, Mark Hasegawa-Johnson |
INTERSPEECH | 3 |
| 2014 | Development of a TV Broadcasts Speech Recognition System for Qatari Arabic
Mohamed Elmahdy 0001, Mark Hasegawa-Johnson, Eiman Mustafawi |
LREC | 2 |
| 2014 | Automatic Long Audio Alignment and Confidence Scoring for Conversational Arabic Speech
Mohamed Elmahdy 0001, Mark Hasegawa-Johnson, Eiman Mustafawi |
LREC | 2 |
| 2014 | Automatic detection of auditory salience with optimized linear filters derived from human annotation
Kyungtae Kim, Kai-Hsiang Lin, Dirk Bernhardt-Walther, Mark Hasegawa-Johnson, Thomas S. Huang |
Pattern Recognit. Lett. | 4 |
| 2014 | Mixed stereo audio classification using a stereo-input mixed-to-panned level featureabstractMany past studies have been conducted on speech/music discrimination due to the potential applications for broadcast and other media; however, it remains possible to expand the experimental scope to include samples of speech with varying amounts of background music. This paper focuses on the development and evaluation of two measures of the ratio between speech energy and music energy: a reference measure called speech-to-music ratio (SMR), which is known objectively only prior to mixing, and a feature called the stereo-input mix-to-peripheral level feature (SIMPL), which is computed from the stereo mixed signal as an imprecise estimate of SMR. SIMPL is an objective signal measure calculated by taking advantage of broadcast mixing techniques in which vocals are typically placed at stereo center, unlike most instruments. Conversely, SMR is a hidden variable defined by the relationship between the powers of portions of audio attributed to speech and music. It is shown that SIMPL is predictive of SMR and can be combined with state-of-the-art features in order to improve performance. For evaluation, this new metric is applied in speech/music (binary) classification, speech/music/mixed (trinary) classification, and a new speech-to-music ratio estimation problem. Promising results are achieved, including 93.06% accuracy for trinary classification and 3.86 dB RMSE for estimation of the SMR. Austin Chen, Mark Hasegawa-Johnson |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2013 | Sparse hidden Markov models for purer clustersabstractThe hidden Markov model (HMM) is widely popular as the de facto tool for representing temporal data; in this paper, we add to its utility in the sequence clustering domain - we describe a novel approach that allows us to directly control purity in HMM-based clustering algorithms. We show that encouraging sparsity in the observation probabilities increases cluster purity and derive an algorithm based on lpregularization; as a corollary, we also provide a different and useful interpretation of the value of p in Renyi p-entropy. We test our method on the problem of clustering non-speech audio events from the BBC sound effects corpus. Experimental results confirm that our approach does learn purer clusters, with (unweighted) average purity as high as 0.88 - a considerable improvement over both the baseline HMM (0.72) and k-means clustering (0.69). Sujeeth Bharadwaj, Mark Hasegawa-Johnson, Jitendra Ajmera, Om Deshmukh, Ashish Verma 0001 |
ICASSP | 2 |
| 2013 | Random features for Kernel Deep Convex NetworkabstractThe recently developed deep learning architecture, a kernel version of the deep convex network (K-DCN), is improved to address the scalability problem when the training and testing samples become very large. We have developed a solution based on the use of random Fourier features, which possess the strong theoretical property of approximating the Gaussian kernel while rendering efficient computation in both training and evaluation of the K-DCN with large training samples. We empirically demonstrate that just like the conventional K-DCN exploiting rigorous Gaussian kernels, the use of random Fourier features also enables successful stacking of kernel modules to form a deep architecture. Our evaluation experiments on phone recognition and speech understanding tasks both show the computational efficiency of the K-DCN which makes use of random features. With sufficient depth in the K-DCN, the phone recognition accuracy and slot-filling accuracy are shown to be comparable or slightly higher than the K-DCN with Gaussian kernels while significant computational saving has been achieved. Po-Sen Huang, Li Deng 0001, Mark Hasegawa-Johnson, Xiaodong He 0001 |
ICASSP | 3 |
| 2013 | Accurate speech segmentation by mimicking human auditory processingabstractThis paper addresses the problem of locating phone boundaries without prior knowledge of the text of an utterance. A biomimetic model of human auditory processing is used to calculate the neural features of frequency synchrony and average signal level. Frequency synchrony and average signal level are used as input to a two-layered support vector machine (SVM)-based system to detect phone boundaries. Phone boundaries are detected with 87.0% precision and 84.8% recall when the automatic segmentation system has no prior knowledge of the phone sequence in the utterance. Sarah King, Mark Hasegawa-Johnson |
ICASSP | 2 |
| 2013 | Acoustic model adaptation using in-domain background models for dysarthric speech recognition
Harsh Vardhan Sharma, Mark Hasegawa-Johnson |
Comput. Speech Lang. | 2 |
| 2013 | Saliency-maximized audio visualization and efficient audio-visual browsing for faster-than-real-time human acoustic event detectionabstractBrowsing large audio archives is challenging because of the limitations of human audition and attention. However, this task becomes easier with a suitable visualization of the audio signal, such as a spectrogram transformed to make unusual audio events salient. This transformation maximizes the mutual information between an isolated event's spectrogram and an estimate of how salient the event appears in its surrounding context. When such spectrograms are computed and displayed with fluid zooming over many temporal orders of magnitude, sparse events in long audio recordings can be detected more quickly and more easily. In particular, in a 1/10-real-time acoustic event detection task, subjects who were shown saliency-maximized rather than conventional spectrograms performed significantly better. Saliency maximization also improves the mutual information between the ground truth of nonbackground sounds and visual saliency, more than other common enhancements to visualization. Kai-Hsiang Lin, Xiaodan Zhuang, Camille Goudeseune, Sarah King, Mark Hasegawa-Johnson, Thomas S. Huang |
ACM Trans. Appl. Percept. | 5 |
| 2012 | Singing-voice separation from monaural recordings using robust principal component analysisabstractSeparating singing voices from music accompaniment is an important task in many applications, such as music information retrieval, lyric recognition and alignment. Music accompaniment can be assumed to be in a low-rank subspace, because of its repetition structure; on the other hand, singing voices can be regarded as relatively sparse within songs. In this paper, based on this assumption, we propose using robust principal component analysis for singing-voice separation from music accompaniment. Moreover, we examine the separation result by using a binary time-frequency masking method. Evaluations on the MIR-1K dataset show that this method can achieve around 1~1.4 dB higher GNSDR compared with two state-of-the-art approaches without using prior training or requiring particular features. Po-Sen Huang, Scott Deeann Chen, Paris Smaragdis, Mark Hasegawa-Johnson |
ICASSP | 4 |
| 2012 | How to put it into words - using random forests to extract symbol level descriptions from audio content for concept detectionabstractThis paper presents a system that uses symbolic representations of audio concepts as words for the descriptions of audio tracks, that enable it to go beyond the state of the art, which is audio event classification of a small number of audio classes in constrained settings, to large-scale classification in the wild. These audio words might be less meaningful for an annotator but they are descriptive for computer algorithms. We devise a random-forest vocabulary learning method with an audio word weighting scheme based on TF-IDF and TD-IDD, so as to combine the computational simplicity and accurate multi-class classification of the random forest with the data-driven discriminative power of the TF-IDF/TD-IDD methods. The proposed random forest clustering with text-retrieval methods significantly outperforms two state-of-the-art methods on the dry-run set and the full set of the TRECVID MED 2010 dataset. Po-Sen Huang, Robert Mertens 0003, Ajay Divakaran, Gerald Friedland, Mark Hasegawa-Johnson |
ICASSP | 5 |
| 2012 | Improving faster-than-real-time human acoustic event detection by saliency-maximized audio visualizationabstractWe propose a saliency-maximized audio spectrogram as a representation that lets human analysts quickly search for and detect events in audio recordings. By rendering target events as visually salient patterns, this representation minimizes the time and effort needed to examine a recording. In particular, we propose a transformation of a conventional spectrogram that maximizes the mutual information between the spectrograms of isolated target events and the estimated saliency of the overall visual representation. When subjects are shown spectrograms that are saliency-maximized, they perform significantly better in a 1/10-real-time acoustic event detection task. Kai-Hsiang Lin, Xiaodan Zhuang, Camille Goudeseune, Sarah King, Mark Hasegawa-Johnson, Thomas S. Huang |
ICASSP | 5 |
| 2012 | Pooling Robust Shift-Invariant Sparse Representations of Acoustic SignalsabstractIn recent years, designing the coding and pooling structures in layered networks has been shown to be a useful method for learning high-level feature representations for visual data. Yet, such learning structures have not been extensively studied for audio signals. In this paper, we investigate the different pooling strategies based on the sparse coding scheme and propose a temporal pyramid pooling method to extract discriminative and shift-invariant feature representations. We demonstrate the superiority of our new feature representation over traditional features on the acoustic event classification task. Index Terms: sparse coding, pooling, acoustic event classification 1. Po-Sen Huang, Jianchao Yang, Mark Hasegawa-Johnson, Feng Liang 0002, Thomas S. Huang |
INTERSPEECH | 3 |
| 2012 | F0 and the Perception of ProminenceabstractThis study investigates the role F0 plays in the perception of prominence in American English. Raw, log and locally normalized measures of F0 were extracted from words in a 35K word corpus of spontaneous speech. Linear regression analyses were conducted to test the strength of these measures as cues to prominence, with prominence based on judgments made by ordinary listeners in real-time auditory perception. The Bayesian Information Criterion was used to further investigate whether these F0 measures cue prominence in a linear or piecewise linear function, corresponding to a linguistic model of prominence as a gradient or discrete feature. The results of this study show that F0 measures are similar to intensity measures in both their strength as cues to perceived prominence, and in signaling a discrete prominence distinction that distinguishes nonor weaklyprominent words from words with greater prominence. Our finding that F0 and intensity cue a discrete prominence distinction is compared with our prior finding that duration and word frequency signal gradient prominence distinctions. This apparent discrepancy is discussed in terms of the dual nature of prominence in English, as an expression of layered metrical (stress) structure in phonology, and as an expression of pragmatic focus. Tim Mahrt, Jennifer Cole 0001, Margaret M. Fleck, Mark Hasegawa-Johnson |
INTERSPEECH | 4 |
| 2012 | Partially Supervised Speaker ClusteringabstractContent-based multimedia indexing, retrieval, and processing as well as multimedia databases demand the structuring of the media content (image, audio, video, text, etc.), one significant goal being to associate the identity of the content to the individual segments of the signals. In this paper, we specifically address the problem of speaker clustering, the task of assigning every speech utterance in an audio stream to its speaker. We offer a complete treatment to the idea of partially supervised speaker clustering, which refers to the use of our prior knowledge of speakers in general to assist the unsupervised speaker clustering process. By means of an independent training data set, we encode the prior knowledge at the various stages of the speaker clustering pipeline via 1) learning a speaker-discriminative acoustic feature transformation, 2) learning a universal speaker prior model, and 3) learning a discriminative speaker subspace, or equivalently, a speaker-discriminative distance metric. We study the directional scattering property of the Gaussian mixture model (GMM) mean supervector representation of utterances in the high-dimensional space, and advocate exploiting this property by using the cosine distance metric instead of the euclidean distance metric for speaker clustering in the GMM mean supervector space. We propose to perform discriminant analysis based on the cosine distance metric, which leads to a novel distance metric learning algorithm—linear spherical discriminant analysis (LSDA). We show that the proposed LSDA formulation can be systematically solved within the elegant graph embedding general dimensionality reduction framework. Our speaker clustering experiments on the GALE database clearly indicate that 1) our speaker clustering methods based on the GMM mean supervector representation and vector-based distance metrics outperform traditional speaker clustering methods based on the “bag of acoustic features” representation and statistical model-based distance metrics, 2) our advocated use of the cosine distance metric yields consistent increases in the speaker clustering performance as compared to the commonly used euclidean distance metric, 3) our partially supervised speaker clustering concept and strategies significantly improve the speaker clustering performance over the baselines, and 4) our proposed LSDA algorithm further leads to state-of-the-art speaker clustering performance. Hao Tang 0001, Stephen M. Chu, Mark Hasegawa-Johnson, Thomas S. Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | On Improving Dynamic State Space Approaches to Articulatory Inversion With MAP-Based Parameter EstimationabstractThis paper presents a complete framework for articulatory inversion based on jump Markov linear systems (JMLS). In the model, the acoustic measurements and the position of each articulator are considered as observable measurement and continuous-valued hidden state of the system, respectively, and discrete regimes of the system are represented by the use of a discrete-valued hidden modal state. Articulatory inversion based on JMLS involves learning the model parameter set of the system and making inference about the state (position of each articulator) of the system using acoustic measurements. Iterative learning algorithms based on maximum-likelihood (ML) and maximum a posteriori (MAP) criteria are proposed to learn the model parameter set of the JMLS. It is shown that the learning procedure of the JMLS is a generalized version of hidden Markov model (HMM) training when both acoustic and articulatory data are given. In this paper, it is shown that the MAP-based learning algorithm improves modeling performance of the system and gives significantly better results compared to ML. The inference stage of the proposed algorithm is based on an interacting multiple models (IMM) approach, and done online (filtering), and/or offline (smoothing). Formulas are provided for IMM-based JMLS smoothing. It is shown that smoothing significantly improves the performance of articulatory inversion compared to filtering. Several experiments are conducted with the MOCHA database to show the performance of the proposed method. Comparison of the performance of the proposed method with the ones given in the literature shows that the proposed method improves the performance of state space approaches, making state space approaches comparable to the best published results. I. Yücel Özbek, Mark Hasegawa-Johnson, Mübeccel Demirekler |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Multi-sensory features for personnel detection at border crossings
Po-Sen Huang, Thyagaraju Damarla, Mark Hasegawa-Johnson |
FUSION | 3 |
| 2011 | Improving acoustic event detection using generalizable visual features and multi-modality modelingabstractAcoustic event detection (AED) aims to identify both timestamps and types of multiple events and has been found to be very challenging. The cues for these events often times exist in both audio and vision, but not necessarily in a synchronized fashion. We study improving the detection and classification of the events using cues from both modalities. We propose optical flow based spatial pyramid histograms as a generalizable visual representation that does not require training on labeled video data. Hidden Markov models (HMMs) are used for audio-only modeling, and multi-stream HMMs or coupled HMMs (CHMM) are used for audio-visual joint modeling. To allow the flexibility of audio-visual state asynchrony, we explore effective CHMM training via HMM state-space mapping, parameter tying and different initialization schemes. The proposed methods successfully improve acoustic event classification and detection on a multimedia meeting room dataset containing eleven types of general non-speech events without using extra data resource other than the video stream accompanying the audio observations. Our systems perform favorably compared to previously reported systems leveraging ad-hoc visual cue detectors and localization information obtained from multiple microphones. Po-Sen Huang, Xiaodan Zhuang, Mark Hasegawa-Johnson |
ICASSP | 3 |
| 2011 | Optimal Models of Prosodic Prominence Using the Bayesian Information CriterionabstractThis study investigated the relation between various acoustic features and prominence. Past research has suggested that duration, pitch, and intensity all play a role in the perception of prominence. In our past work, we found a correlation between these acoustic features and speaker agreement over the placement of prominence. The current study was motivated by a need to enrich our understanding of this correlation. Using the Bayesian information criterion, we show that the best model for a feature that cues prosody is not necessarily a single Gaussian. Rather, the best model depends on the feature. This finding has consequences for our understanding of the role of these features in the perception of prosody and for prosody recognition systems. Tim Mahrt, Jui-Ting Huang, Yoonsook Mo, Margaret M. Fleck, Mark Hasegawa-Johnson, Jennifer Cole 0001 |
INTERSPEECH | 5 |
| 2011 | Intelligibility predictors and neural representation of speech
Bryce E. Lobdell, Jont B. Allen, Mark Hasegawa-Johnson |
Speech Commun. | 3 |
| 2011 | Estimation of Articulatory Trajectories Based on Gaussian Mixture Model (GMM) With Audio-Visual Information Fusion and Dynamic Kalman SmoothingabstractThis paper presents a detailed framework for Gaussian mixture model (GMM)-based articulatory inversion equipped with special postprocessing smoothers, and with the capability to perform audio-visual information fusion. The effects of different acoustic features on the GMM inversion performance are investigated and it is shown that the integration of various types of acoustic (and visual) features improves the performance of the articulatory inversion process. Dynamic Kalman smoothers are proposed to adapt the cutoff frequency of the smoother to data and noise characteristics; Kalman smoothers also enable the incorporation of auxiliary information such as phonetic transcriptions to improve articulatory estimation. Two types of dynamic Kalman smoothers are introduced: global Kalman (GK) and phoneme-based Kalman (PBK). The same dynamic model is used for all phonemes in the GK smoother; it is shown that GK improves the performance of articulatory inversion better than the conventional low-pass (LP) smoother. However, the PBK smoother, which uses one dynamic model for each phoneme, gives significantly better results than the GK smoother. Different methodologies to fuse the audio and visual information are examined. A novel modified late fusion algorithm, designed to consider the observability degree of the articulators, is shown to give better results than either the early or the late fusion methods. Extensive experimental studies are conducted with the MOCHA database to illustrate the performance gains obtained by the proposed algorithms. The average RMS error and correlation coefficient between the true (measured) and the estimated articulatory trajectories are 1.227 mm and 0.868 using audiovisual information fusion and GK smoothing, and 1.199 mm and 0.876 using audiovisual information fusion together with PBK smoothing based on a phonetic transcription of the utterance. I. Yücel Özbek, Mark Hasegawa-Johnson, Mübeccel Demirekler |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Joint estimation of DOA and speech based on EM beamformingabstractIn this paper, we propose a multi-microphone joint optimal estimation of the direction of arrival (DOA) and the source speech signal through newly introduced EM beamforming. This produces a posterior PDF for the DOA, based only on the reliable speech spectrum. By maximizing over the posterior PDF of the DOA, we achieve maximum a posteriori DOA estimation. After convergence, the estimated source spectrum through weighted sum in the Bayesian sense is a maximum likelihood estimate (MLE). This is a sufficient statistic for minimum mean square error (MMSE) optimal estimation using a subsequent single channel MMSE filter. Lae-Hoon Kim, Mark Hasegawa-Johnson, Gerasimos Potamianos, Vit Libal |
ICASSP | 2 |
| 2010 | Toward robust learning of the Gaussian mixture state emission densities for hidden Markov modelsabstractOne important class of state emission densities of the hiddenMarkov model (HMM) is the Gaussian mixture densities. The classical Baum-Welch algorithm often fails to reliably learn the Gaussian mixture densities when there is insufficient training data, due to the large number of free parameters present in the model. In this paper, we propose a novel strategy for robustly and accurately learning the Gaussian mixture state emission densities of the HMM. The strategy is based on an ensemble framework for probability density estimation in which the learning of the Gaussian mixture densities is formulated as a gradient descent search in a function space. The resulting learning algorithm is called “the boosting Baum-Welch algorithm.” Our preliminary experiment results on emotion recognition from speech show that the proposed algorithm outperforms the original Baum-Welch algorithm on this task. Hao Tang 0001, Mark Hasegawa-Johnson, Thomas S. Huang |
ICASSP | 2 |
| 2010 | Non-frontal view facial expression recognition based on ergodic hidden Markov model supervectorsabstractAutomatic facial expression recognition from non-frontal views is a challenging research topic which has recently started to attract the attention of the research community. In this paper, we propose a novel approach to tackling this problem based on the ergodic hidden Markov model (EHMM) supervector representation of facial images. First, the scale-invariant feature transform (SIFT) feature vectors are extracted from a dense grid of every facial images. Next, an EHMM is trained over all facial images in the training set and is referred to as the universal background model (UBM). The UBM is then maximum a posteriori adapted to each facial image in the training and test sets to produce the image-specific EHMMs. Based on these EHMMs, we derive a supervector representation of the facial images by means of an upper bound approximation of the Kullback-Leibler divergence rate between two EHMMs. Finally, facial expression recognition is performed in the linear discriminant subspace of the EHMM supervectors using the k-nearest-neighbor classification algorithm. Our experiments of recognizing six universal facial expressions over extensive multiview facial images with seven pan angles (-45° ~ +45°) and five tilt angles (-30° ~ +30°), which are synthesized from the BU-3DFE facial expression database, show promising results compared to the state of the arts recently reported. Hao Tang 0001, Mark Hasegawa-Johnson, Thomas S. Huang |
ICME | 2 |
| 2010 | FSM-based pronunciation modeling using articulatory phonological codeabstractAccording to articulatory phonology, the gestural score is an invariant speech representation. Though the timing schemes, i.e., the onsets and offsets, of \nthe gestural activations may vary, the ensemble of these activations tends to remain unchanged, informing the speech content. "Gestural pattern vector" \n(GPV) has been proposed to encode the instantaneous gestural activations that exist across all tract variables at each time. Therefore, a gestural score with a particular timing scheme can be approximated using a GPV sequence. \n \nIn this work, we propose a pronunciation modeling method that uses a finite state machine (FSM) to represent the invariance of a gestural score. Given the "canonical" gestural score of a word with a known activation timing \nscheme, the plausible activation onsets and offsets are recursively generated and encoded as a weighted FSM. An empirical measure is used to prune out gestural activation timing schemes that deviate too much from the "canonical" gestural score. Speech recognition is achieved by matching the recovered \ngestural activations to the FSM-encoded gestural scores of different speech contents. In particular, the observation distribution of each GPV is modeled \nby an artificial neural network and Gaussian mixture tandem model. These models are used together with the FSM-based pronunciation models in a Bayesian framework. \n \nWe carry out pilot word classification experiments using synthesized data from one speaker. The proposed pronunciation modeling achieves over 90% accuracy for a vocabulary of 139 words with no training observations, outperforming direct use of the "canonical" gestural score. Chi Hu, Xiaodan Zhuang, Mark Hasegawa-Johnson |
INTERSPEECH | 3 |
| 2010 | Semi-supervised training of Gaussian mixture models by conditional entropy minimizationabstractIn this paper, we propose a new semi-supervised training method for Gaussian Mixture Models. We add a conditional entropy minimizer to the maximum mutual information criteria, which enables to incorporate unlabeled data in a discriminative training fashion. The training method is simple but surprisingly effective. The preconditioned conjugate gradient method provides a reasonable convergence rate for parameter update. The phonetic classification experiments on the TIMIT corpus demonstrate significant improvements due to unlabeled data via our training criteria. Index Terms: semi-supervised learning, conditional entropy, Gaussian Mixture Models, phonetic classification Jui-Ting Huang, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2010 | Robust automatic speech recognition with decoder oriented ideal binary mask estimationabstractIn this paper, we propose a joint optimal method for automatic speech recognition (ASR) and ideal binary mask (IBM) estimation in transformed into the cepstral domain through a newly derived generalized expectation maximization algorithm. First, cepstral domain missing feature marginalization is established using a linear transformation, after tying the mean and variance of non-existing cepstral coefficients. Second, IBM estimation is formulated using a generalized expectation maximization algorithm directly to optimize the ASR performance. Experimental results show that even in highly non-stationary mismatch condition (dance music as background noise), the proposed method achieves much higher absolute ASR accuracy improvement ranging from 14.69 % at 0 dB SNR to 40.10 % at 15 dB SNR compared with the conventional noise suppression method. Index Terms: robust speech recognition, ideal binary mask classification, missing feature 1. Lae-Hoon Kim, Kyung-Tae Kim, Mark Hasegawa-Johnson |
INTERSPEECH | 3 |
| 2010 | Kinematic analysis of tongue movement control in spastic dysarthriaabstractThis study provided a quantitative analysis of the kinematic deviances in dysarthria associated with spastic cerebral palsy. Of particular interest were tongue tip movements during alveolar consonant release. Our analysis based on EMA measures indicated that speakers with spastic dysarthria had a restricted range of articulation and disturbances in articulatory-voicing coordination. The degree of kinematic deviances was greater for lower intelligibility speakers, supporting an association between articulatory dysfunctions and intelligibility in spastic dysarthria. Index Terms: dysarthria, kinematic analysis, EMA 1. Panying Rong, Torrey M. Loucks, Mark Hasegawa-Johnson |
INTERSPEECH | 4 |
| 2010 | A procedure for estimating gestural scores from natural speechabstractAbstract * Speech can be represented as a constellation of constricting events, gestures, , which are defined at distinct vocal tract sites, in the form of a gestural score.. Gestures and their output trajectories, tract variables, , which are available only in synthetic speech, have recently been shown to improve automatic speech recognition (ASR) performance. In this paper we propose an iterative analysis-by-synthesis synthesis landmark based time-warping architecture to obtain gestural scores for natural speech. Given an utterance, the Haskins Laboratories Task Dynamics and Application (TADA) model was used to generate its prototype gestural score and the corresponding synthetic acoustic output. An optimal gestural score was estimated through iterative time-warping processes such that the distance between original and TADA-synthesized synthesized speech is minimized. We compared the performance of our approach to that of a conventional dynamic time warping procedure using Log-Spectral and Itakura Distance measures. We also performed a word recognition experiment using the gestural annotations to show that the gestural scores are suitable for word recognition. Hosung Nam, Vikramjit Mitra, Mark K. Tiede, Elliot Saltzman, Louis Goldstein, Carol Y. Espy-Wilson, Mark Hasegawa-Johnson |
INTERSPEECH | 7 |
| 2010 | Landmark-based automated pronunciation error detectionabstractWe present a pronunciation error detection method for second language learners of English (L2 learners). The method is a combination of confidence scoring and landmark-based Support Vector Machines (SVMs). Landmark-based SVMs were implemented to specialize the method for the specific phonemes with which L2 learners make frequent errors. The method was trained for the difficult phonemes for Korean learners and tested on intermediate Korean learners. In the data where distortion errors (non-phonemic errors) occupied high proportion, SVM method achieved significantly higher F-score (0.67) than confidence scoring (0.60). However, the combination of two methods without the appropriate training data did not lead to improvement. Even for intermediate learners, a high proportion of errors (40%) was related to these difficult phonemes. Therefore, the method specialized for these phonemes will be beneficial for both beginners and intermediate learners. Index Terms: automated pronunciation error detection, computer aided pronunciation training system, confidence score, Su-Youn Yoon, Mark Hasegawa-Johnson, Richard Sproat |
INTERSPEECH | 2 |
| 2010 | A minimum converted trajectory error (MCTE) approach to high quality speech-to-lips conversionabstractHigh quality speech-to-lips conversion, investigated in this work, ren-ders realistic lips movement (video) consistent with input speech (audio) without knowing its linguistic content. Instead of memoryless frame-based conversion, we adopt maximum likelihood estimation of the vi-sual parameter trajectories using an audio-visual joint Gaussian Mixture Model (GMM). We propose a minimum converted trajectory error ap-proach (MCTE) to further refine the converted visual parameters. First, we reduce the conversion error by training the joint audio-visual GMM with weighted audio and visual likelihood. Then MCTE uses the gen-eralized probabilistic descent algorithm to minimize a conversion error of the visual parameter trajectories defined on the optimal Gaussian ker-nel sequence according to the input speech. We demonstrate the effec-tiveness of the proposed methods using the LIPS 2009 Visual Speech Synthesis Challenge dataset, without knowing the linguistic (phonetic) content of the input speech. Index Terms: visual speech synthesis, speech-to-lips conversion, mini-mum conversion error, minimum generation error Xiaodan Zhuang, Frank K. Soong, Mark Hasegawa-Johnson |
INTERSPEECH | 4 |
| 2010 | Novel Gaussianized vector representation for improved natural scene categorization
Xiaodan Zhuang, Hao Tang 0001, Mark Hasegawa-Johnson, Thomas S. Huang |
Pattern Recognit. Lett. | 4 |
| 2010 | Real-world acoustic event detection
Xiaodan Zhuang, Mark Hasegawa-Johnson, Thomas S. Huang |
Pattern Recognit. Lett. | 3 |
| 2010 | A Novel Vector Representation of Stochastic Signals Based on Adapted Ergodic HMMsabstractIn this letter, we propose a novel vector representation of stochastic signals for pattern recognition (PR) based on adapted ergodic hidden Markov models (HMMs). This vector representation is generic in nature and may be used with various types of stochastic signals (e.g., image, speech, etc.) and applied to a broad range of PR tasks (e.g., classification, regression, etc.). More importantly, by combining the vector representation with optimal distance metric learning (e.g., linear discriminant analysis) directly from the data, the performance of a PR system may be significantly improved. Our experiments on an image-based recognition task clearly demonstrate the effectiveness of the proposed vector representation of stochastic signals for potential use in many PR systems. Hao Tang 0001, Mark Hasegawa-Johnson, Thomas S. Huang |
IEEE Signal Process. Lett. | 2 |
| 2009 | Kernel metric learning for phonetic classificationabstractWhile a sound spoken is described by a handful of frame-level spectral vectors, not all frames have equal contribution for either human perception or machine classification. In this paper, we introduce a novel framework to automatically emphasize important speech frames relevant to phonetic information. We jointly learn the importance of speech frames by a distance metric across the phone classes, attempting to satisfy a large margin constraint: the distance from a segment to its correct label class should be less than the distance to any other phone class by the largest possible margin. Furthermore, an universal background model structure is proposed to give the correspondence between statistical models of phone types and tokens, allowing us to use statistical models of each phone token in a large margin speech recognition framework. Experiments on TIMIT database demonstrated the effectiveness of our framework. Jui-Ting Huang, Mark Hasegawa-Johnson, Thomas S. Huang |
ASRU | 3 |
| 2009 | Acoustic fall detection using Gaussian mixture models and GMM supervectorsabstractWe present a system that detects human falls in the home environment, distinguishing them from competing noise, by using only the audio signal from a single far-field microphone. The proposed system models each fall or noise segment by means of a Gaussian mixture model (GMM) supervector, whose Euclidean distance measures the pairwise difference between audio segments. A support vector machine built on a kernel between GMM supervectors is employed to classify audio segments into falls and various types of noise. Experiments on a dataset of human falls, collected as part of the Netcarity project, show that the method improves fall classification F-score to 67% from 59% of a baseline GMM classifier. The approach also effectively addresses the more difficult fall detection problem, where audio segment boundaries are unknown. Specifically, we employ it to reclassify confusable segments produced by a dynamic programming scheme based on traditional GMMs. Such post-processing improves a fall detection accuracy metric by 5% relative. Xiaodan Zhuang, Jing Huang 0019, Gerasimos Potamianos, Mark Hasegawa-Johnson |
ICASSP | 4 |
| 2009 | Emotion recognition from speech VIA boosted Gaussian mixture modelsabstractGaussian mixture models (GMMs) and the minimum error rate classifier (i.e. Bayesian optimal classifier) are popular and effective tools for speech emotion recognition. Typically, GMMs are used to model the class-conditional distributions of acoustic features and their parameters are estimated by the expectation maximization (EM) algorithm based on a training data set. Then, classification is performed to minimize the classification error w.r.t. the estimated class-conditional distributions. We call this method the EM-GMM algorithm. In this paper, we introduce a Boosting algorithm for reliably and accurately estimating the class-conditional GMMs. The resulting algorithm is named the Boosted-GMM algorithm. Our speech emotion recognition experiments show that the emotion recognition rates are effectively and significantly "Boosted" by the Boosted-GMM algorithm as compared to the EM-GMM algorithm. This is due to the fact that the Boosting algorithm can lead to more accurate estimates of the class-conditional GMMs, namely the class-conditional distributions of acoustic features. Hao Tang 0001, Stephen M. Chu, Mark Hasegawa-Johnson, Thomas S. Huang |
ICME | 3 |
| 2009 | Prosodic effects on vowel production: evidence from formant structureabstractSpeakers communicate pragmatic and discourse meaning through the prosodic form assigned to an utterance, and listeners must attend to the acoustic cues to prosodic form to fully recover the speaker’s intended meaning. While much of the research on prosody examines supra-segmental cues such as F0 and temporal patterns, prosody is also known to affect the phonetic properties of segments as well. This paper reports on the effect of prosodic prominence on the formant patterns of vowels using speech data from the Buckeye corpus of spontaneous American English. A prosody annotation was obtained for a subset of this corpus based on the auditory perception of 97 ordinary, untrained listeners. To understand the relationship between prominence perception and formant structure, as a measure of the ‘strength ’ of the vowel Yoonsook Mo, Jennifer Cole 0001, Mark Hasegawa-Johnson |
INTERSPEECH | 3 |
| 2009 | Formant trajectories for acoustic-to-articulatory inversionabstractThis work examines the utility of formant frequencies and their energies in acoustic-to-articulatory inversion. For this purpose, formant frequencies and formant spectral amplitudes are automatically estimated from audio, and are treated as observations for the purpose of estimating electromagnetic articulography (EMA) coil positions. A mixture Gaussian regression model with mel-frequency cepstral (MFCC) observations is modified by using formants and energies to either replace or augment the MFCC observation vector. The augmented observation results in 3.4 % lower RMS error, and 2.7 % higher correlation coefficient, than the baseline MFCC observation. Improvement is especially good for plosive consonants, possibly because formant tracking provides information about the acoustic resonances that would be otherwise unavailable during plosive closure and release. Index Terms: acoustic-to-articulatory inversion, formant tracking, GMM regression I. Yücel Özbek, Mark Hasegawa-Johnson, Mübeccel Demirekler |
INTERSPEECH | 2 |
| 2009 | Universal access: speech recognition for talkers with spastic dysarthria
Harsh Vardhan Sharma, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2009 | Automated pronunciation scoring using confidence scoring and landmark-based SVMabstractIn this study, we present a pronunciation scoring method for second language learners of English (hereafter, L2 learners). This study presents a method using both confidence scoring and classifiers. Classifiers have an advantage over confidence scoring for specialization in the specific phonemes where L2 learners make frequent errors. Classifiers (Landmark-based Support Vector Machines) were trained in order to distinguish L2 phonemes from their frequent substitution patterns. In this study, the method was evaluated on the specific English phonemes where L2 English learners make frequent errors. The results suggest that the automated pronunciation scoring method can be improved consistently by combining the two methods. Su-Youn Yoon, Mark Hasegawa-Johnson, Richard Sproat |
INTERSPEECH | 2 |
| 2009 | Articulatory phonological code for word classificationabstractWe propose a framework that leverages articulatory phonology for speech recognition. “Gestural pattern vectors ” (GPV) encode the instantaneous gestural activations that exist across all tract variables at each time. Given a speech observation, recognizing the sequence of GPV recovers the ensemble of gestural activations, i.e., the gestural score. For each word in the vocabulary, we use a task dynamic model of inter-articulator speech coordination to generate the “canonical ” gestural score. Speech recognition is achieved by matching the ensemble of gestural activations. In particular, we estimate the likelihood of the recognized GPV sequence on word-dependent GPV sequence models trained using the “canonical” gestural scores. These likelihoods, weighted by confidence score of the recognized GPVs, are used in a Bayesian speech recognizer. Pilot gestural score recovery and word classification experiments are carried out using synthesized data from one speaker. The observation distribution of each GPV is modeled by an artificial neural network and Gaussian mixture tandem model. Bigram GPV sequence models are used to distinguish gestural scores of different words. Given the tract variable time functions, about 80 % of the instantaneous gestural activation is correctly recovered. Word recognition accuracy is over 85 % for a vocabulary of 139 words with no training observations. These results suggest that the proposed framework might be a viable alternative to the classic sequence-of-phones model. Index Terms: speech production, speech gesture, tandem model, artificial neural network, Gaussian mixture model Xiaodan Zhuang, Hosung Nam, Mark Hasegawa-Johnson, Louis Goldstein, Elliot Saltzman |
INTERSPEECH | 3 |
| 2008 | Regression from patch-kernelabstractIn this paper, we present a patch-based regression framework for addressing the human age and head pose estimation problems. Firstly, each image is encoded as an ensemble of orderless coordinate patches, the global distribution of which is described by Gaussian Mixture Models (GMM), and then each image is further expressed as a specific distribution model by Maximum a Posteriori adaptation from the global GMM. Then the patch-kernel is designed for characterizing the Kullback-Leibler divergence between the derived models for any two images, and its discriminating power is further enhanced by a weak learning process, called inter-modality similarity synchronization. Finally, kernel regression is employed for ultimate human age or head pose estimation. These three stages are complementary to each other, and jointly minimize the regression error. The effectiveness of this regression framework is validated by three experiments: 1) on the YAMAHA aging database, our solution brings a more than 50% reduction in age estimation error compared with the best reported results; 2) on the FG-NET aging database, our solution based on raw image features performs even better than the state-of-the-art algorithms which require fine face alignment for extracting warped appearance features; and 3) on the CHIL head pose database, our solution significantly outperforms the best one reported in the CLEAR07 evaluation. Shuicheng Yan, Ming Liu 0009, Mark Hasegawa-Johnson, Thomas S. Huang |
CVPR | 4 |
| 2008 | Optimal speech estimator considering room response as well as additive noise: Different approaches in low and high frequency rangeabstractThis paper proposes minimum mean squared error (MMSE) speech signal estimation in a reverberant space using different optimal estimators in the low and high frequency ranges. At low frequencies, an MMSE spectral amplitude estimator divided by the spectral amplitude of a representative impulse response produces optimal performance. In the high frequency range, the MMSE estimator is computed based on its sufficient statistic: the maximum likelihood (ML) estimate. Inference is factored using a two-step algorithm: the maximum likelihood value of the source spectrum is first estimated using expectation-maximization (EM) under the assumption of the hidden room response with complex Gaussian pdf, then the MMSE source spectral estimate is computed. Lae-Hoon Kim, Mark Hasegawa-Johnson |
ICASSP | 2 |
| 2008 | Feature analysis and selection for acoustic event detectionabstractSpeech perceptual features, such as Mel-frequency Cepstral Coefficients (MFCC), have been widely used in acoustic event detection. However, the different spectral structures between speech and acoustic events degrade the performance of the speech feature sets. We propose quantifying the discriminative capability of each feature component according to the approximated Bayesian accuracy and deriving a discriminative feature set for acoustic event detection. Compared to MFCC, feature sets derived using the proposed approaches achieve about 30% relative accuracy improvement in acoustic event detection. Xiaodan Zhuang, Thomas S. Huang, Mark Hasegawa-Johnson |
ICASSP | 4 |
| 2008 | Real-time conversion from a single 2D face image to a 3D text-driven emotive audio-visual avatarabstractIn this paper, we propose a complete pipeline of efficient and low-cost techniques to construct a realistic 3D text-driven emotive audio-visual avatar from a single 2D frontal-view face image of any person on the fly. This real-time conversion is achieved through three steps. First, a personalized 3D face model is built based on the 2D face image using a fully automatic 3D face shape and texture reconstruction framework. Second, using standard MPEG-4 FAPs (Facial Animation Parameters), the face model is animated by the viseme and expression channels and is complemented by the visual prosody channel that controls head, eye and eyelid movements. Finally, the facial animation is combined and synchronized with the emotive synthetic speech generated by incorporating an emotion transformer into a Festival-MBROLA text to neutral speech synthesizer. Hao Tang 0001, Yuxiao Hu 0001, Yun Fu 0001, Mark Hasegawa-Johnson, Thomas S. Huang |
ICME | 4 |
| 2008 | A novel Gaussianized vector representation for natural scene categorizationabstractThis paper presents a novel Gaussianized vector representation for scene images by an unsupervised approach. First, each image is encoded as an ensemble of orderless bag of features, and then a global Gaussian Mixture Model (GMM) learned from all images is used to randomly distribute each feature into one Gaussian component by a multinomial trial. The parameters of the multinomial distribution are defined by the posteriors of the feature on all the Gaussian components. Finally, the normalized means of the features distributed in every Gaussian component are concatenated to form a supervector, which is a compact representation for each scene image. We prove that these super-vectors observe the standard normal distribution. Our experiments on scene categorization tasks using this vector representation show significantly improved performance compared with the bag-of-features representation. Xiaodan Zhuang, Hao Tang 0001, Mark Hasegawa-Johnson, Thomas S. Huang |
ICPR | 4 |
| 2008 | Face age estimation using patch-based hidden Markov model supervectorsabstractRecent studies in patch-based Gaussian Mixture Model (GMM) approaches for face age estimation present promising results. We propose using a hidden Markov model (HMM) supervector to represent face image patches, to improve from the previous GMM supervector approach by capturing the spatial structure of human faces and loosening the assumption of identical face patch distribution within a face image. The Euclidean distance of HMM supervectors constructed from two face images measures the similarity of the human faces, derived from the approximated Kullback-Leibler divergence between the joint distributions of patches with implicit unsupervised alignment of different regions in two human faces. The proposed HMM supervector approach compares favorably with the GMM supervector approach in face age estimation on a large face dataset. Xiaodan Zhuang, Mark Hasegawa-Johnson, Thomas S. Huang |
ICPR | 3 |
| 2008 | Maximum mutual information estimation with unlabeled data for phonetic classificationabstractThis paper proposes a new training framework for mixed labeled and unlabeled data and evaluates it on the task of binary phonetic classification. Our training objective function combines Maximum Mutual Information (MMI) for labeled data and Maximum Likelihood (ML) for unlabeled data. Through the modified training objective, MMI estimates are smoothed with ML estimates obtained from unlabeled data. On the other hand, our training criterion can also help the existing model adapt to new speech characteristics from unlabeled speech. In our experiments of phonetic classification, there is a consistent reduction of error rate from MLE to MMIE with I-smoothing, and then to MMIE with unlabeled-smoothing. Error rates can be further reduced by transductive-MMIE. We also experimented with the gender-mismatched case, in which the best improvement shows MMIE with unlabeled data has a 9.3 % absolute lower error rate than MLE and a 2.35 % absolute lower error rate than MMIE with I-smoothing. Index Terms: unlabeled speech, Maximum mutual information, Gaussian mixture models Jui-Ting Huang, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2008 | Dysarthric speech database for universal access researchabstractThis paper describes a database of dysarthric speech produced by 19 speakers with cerebral palsy. Speech materials consist of 765 isolated words per speaker: 300 distinct uncommon words and 3 repetitions of digits, computer commands, radio alphabet and common words. Data is recorded through an 8-microphone array and one digital video camera. Our database provides a fundamental resource for automatic speech recognition development for people with neuromotor disability. Research on articulation errors in dysarthria will benefit clinical treatments and contribute to our knowledge of neuromotor mechanisms in speech production. Data files are available via secure ftp upon request. Index Terms: speech recognition, dysarthria, cerebral palsy 1. Mark Hasegawa-Johnson, Adrienne Perlman, Jon R. Gunderson, Thomas S. Huang, Kenneth L. Watkin, Simone Frame |
INTERSPEECH | 2 |
| 2008 | Human speech perception and feature extractionabstractSpeech perception experiments tell us a great deal about which factors affect human performance and behavior. In particular many experiments indicate that the signal-to-noise ratio spectrum is an important factor, indeed the signal-to-noise ratio spectrum is the basis of the Articulation Index, a standard measure of “speech channel capacity. ” In this paper we compare speech recognition performance for features based on the Articulation Index with two alternatives typically used in speech recognition. The experimental conditions vary the spectrum and level of noise distorting the speech in the training and test set. The perceptually inspired features generally perform better when there is a mismatch between the training and test noise spectrum and level, but worse when the test and training noises match. 1. Bryce E. Lobdell, Mark Hasegawa-Johnson, Jont B. Allen |
INTERSPEECH | 2 |
| 2008 | Two-stage prosody prediction for emotional text-to-speech synthesisabstractIn this paper, we adopt a difference approach to prosody prediction for emotional text-to-speech synthesis, where the prosodic variations between emotional and neutral speech are decomposed into the global and local prosodic variations and predicted using a two-stage model. The global prosodic variations are modeled by the means and standard deviations of the prosodic parameters, while the local prosodic variations are modeled by the classification and regression tree (CART) and dynamic programming. The proposed two-stage prosody prediction model has been successfully implemented as a prosodic module in a Festival-MBROLA architecture based emotional text-to-speech synthesis system, which is able to synthesize highly intelligible, natural and expressive speech. Index Terms: TTS, speech synthesis, prosody prediction, CART, dynamic programming Hao Tang 0001, Matthias Odisio, Mark Hasegawa-Johnson, Thomas S. Huang |
INTERSPEECH | 4 |
| 2008 | The entropy of the articulatory phonological code: recognizing gestures from tract variablesabstractWe propose an instantaneous “gestural pattern vector ” to encode the instantaneous pattern of gesture activations across tract variables in the gestural score. The design of these gestural pattern vectors is the first step towards an automatic speech recognizer motivated by articulatory phonology, which is expected to be more invariant to speech coarticulation and reduction than conventional speech recognizers built with the sequenceof-phones assumption. We use a tandem model to recover the instantaneous gestural pattern vectors from tract variable time functions in local time windows, and achieve classification accuracy up to 84.5% for synthesized data from one speaker. Recognizing all gestural pattern vectors is equivalent to recognizing the ensemble of gestures. This result suggests that the proposed gestural pattern vector might be a viable unit in statistical models for speech recognition. Index Terms: speech production, speech gesture, tandem model, artificial neural network, Gaussian mixture model Xiaodan Zhuang, Hosung Nam, Mark Hasegawa-Johnson, Louis Goldstein, Elliot Saltzman |
INTERSPEECH | 3 |
| 2008 | SIFT-Bag kernel for video event analysisabstractIn this work, we present a SIFT-Bag based generative-todiscriminative framework for addressing the problem of video event recognition in unconstrained news videos. In the generative stage, each video clip is encoded as a bag of SIFT feature vectors, the distribution of which is described by a Gaussian Mixture Models (GMM). In the discriminative stage, the SIFT-Bag Kernel is designed for characterizing the property of Kullback-Leibler divergence between the specialized GMMs of any two video clips, and then this kernel is utilized for supervised learning in two ways. On one hand, this kernel is further refined in discriminating power for centroid-based video event classification by using the Within-Class Covariance Normalization approach, which depresses the kernel components with high-variability for video clips of the same event. On the other hand, the SIFT-Bag Kernel is used in a Support Vector Machine for margin-based video event classification. Finally, the outputs from these two classifiers are fused together for final decision. The experiments on the TRECVID 2005 corpus demonstrate that the mean average precision is boosted from the best reported 38.2 % in [36] to 60.4 % based on our new framework. Xiaodan Zhuang, Shuicheng Yan, Shih-Fu Chang, Mark Hasegawa-Johnson, Thomas S. Huang |
ACM Multimedia | 5 |
| 2008 | EAVA: A 3D Emotive Audio-Visual AvatarabstractEmotive audio-visual avatars have the potential of significantly improving the quality of Human-Computer Interaction (HCI). In this paper, the various technical approaches of a novel framework leading to a text-driven 3D Emotive Audio-Visual Avatar (EAVA) are proposed. Primary work is focused on 3D face modeling, realistic emotional facial expression animation, emotive speech synthesis, and the co-articulation of speech gestures (i.e., lip movements due to speech production) and facial expressions. Experimental results clearly indicate that a certain degree of naturalness and expressiveness has been achieved by EAVA in both audio and visual aspects. Promising potential improvements can be expected by incorporating various data-driven statistical learning models into the framework. Hao Tang 0001, Yun Fu 0001, Jilin Tu, Thomas S. Huang, Mark Hasegawa-Johnson |
WACV | 5 |
| 2008 | Humanoid Audio-Visual Avatar With Emotive Text-to-Speech SynthesisabstractEmotive audio-visual avatars are virtual computer agents which have the potential of improving the quality of human-machine interaction and human-human communication significantly. However, the understanding of human communication has not yet advanced to the point where it is possible to make realistic avatars that demonstrate interactions with natural-sounding emotive speech and realistic-looking emotional facial expressions. In this paper, We propose the various technical approaches of a novel multimodal framework leading to a text-driven emotive audio-visual avatar. Our primary work is focused on emotive speech synthesis, realistic emotional facial expression animation, and the co-articulation between speech gestures (i.e., lip movements) and facial expressions. A general framework of emotive text-to-speech (TTS) synthesis using a diphone synthesizer is designed and integrated into a generic 3-D avatar face model. Under the guidance of this framework, we therefore developed a realistic 3-D avatar prototype. A rule-based emotive TTS synthesis system module based on the Festival-MBROLA architecture has been designed to demonstrate the effectiveness of the framework design. Subjective listening experiments were carried out to evaluate the expressiveness of the synthetic talking avatar. Hao Tang 0001, Yun Fu 0001, Jilin Tu, Mark Hasegawa-Johnson, Thomas S. Huang |
IEEE Trans. Multim. | 4 |
| 2007 | Articulatory Feature-Based Methods for Acoustic and Audio-Visual Speech Recognition: Summary from the 2006 JHU Summer workshopabstractWe report on investigations, conducted at the 2006 Johns Hopkins Workshop, into the use of articulatory features (AFs) for observation and pronunciation models in speech recognition. In the area of observation modeling, we use the outputs of AF classifiers both directly, in an extension of hybrid HMM/neural network models, and as part of the observation vector, an extension of the "tandem" approach. In the area of pronunciation modeling, we investigate a model having multiple streams of AF states with soft synchrony constraints, for both audio-only and audio-visual recognition. The models are implemented as dynamic Bayesian networks, and tested on tasks from the small-vocabulary switchboard (SVitchboard) corpus and the CUAVE audio-visual digits corpus. Finally, we analyze AF classification and forced alignment using a newly collected set of feature-level manual transcriptions. Karen Livescu, Özgür Çetin, Mark Hasegawa-Johnson, Simon King 0001, Chris D. Bartels, Nash M. Borges, Arthur Kantor, Partha Lal, Lisa Yung, Ari Bezman, Stephen Dawson-Haggerty, Bronwyn Woods, Joe Frankel, Mathew Magimai-Doss, Kate Saenko |
ICASSP (4) | 3 |
| 2007 | Lipreading by Locality Discriminant GraphabstractThe major problem in building a good lipreading system is to extract effective visual features from the enormous quantity of video sequences data. For appearance-based feature analysis in lipreading, classical methods, e.g. DCT, PCA and LDA, are usually applied to dimensionality reduction. We present a new pattern classification algorithm, called locality discriminant graph (LDG), and develop a novel lipreading framework to successfully apply LDG to the problem. LDG takes the advantages of both manifold learning and Fisher criteria to seek the linear embedding which preserves the local neighborhood affinity withinsameclasswhile discriminating the neighborhood amongdifferentclasses. The LDG embedding is computed in closed-form and tuned by the only open parameter of k-NN number. Experiments on AVICAR corpus provide evidence that the graph-based pattern classification methods can outperform classical ones for lipreading. Yun Fu 0001, Ming Liu 0009, Mark Hasegawa-Johnson, Thomas S. Huang |
ICIP (3) | 4 |
| 2007 | Exploring Discriminative Learning for Text-Independent Speaker RecognitionabstractSpeaker verification is a technology of verifying the claimed identity of a speaker based on the speech signal from the speaker (voice print). To learn the score of similarity between each pair of target and trial utterances, we investigated two different discriminative learning frameworks: Fisher mapping followed by SVM learning and utterance transform followed by iterative cohort modeling (ICM). In both methods, a mapping is applied to map speech utterance from a variable-length acoustic feature sequence into a fixed dimensional vector. SVM learning constructs a classifier in the mapped vector space for speaker verification. ICM learns a metric in this vector space by incorporating discriminative learning methods. The obtained metric is then used by a nearest neighbor classifier for speaker verification. The experiments conducted on NIST02 corpus show that both discriminative learning methods outperform the baseline GMM-UBM system. Furthermore, we observe that the ICM-based method is more effective than the SVM-based method, indicating that the metric learning scheme is more powerful in constructing a better metric in the mapped vector space. Ming Liu 0009, Zhengyou Zhang, Mark Hasegawa-Johnson, Thomas S. Huang |
ICME | 3 |
| 2007 | Robust Analysis and Weighting on MFCC Components for Speech Recognition and Speaker IdentificationabstractMismatch between training and testing data is a major error source for both automatic speech recognition (ASR) and automatic speaker identification (ASI). In this paper, we first present a statistical weighting concept to exploit the unequal sensitivity of mel-frequency cepstral coefficients (MFCC) components to against the mismatch, such as ambient noise, recording equipment, transmission channels, and inter-speaker variations. We further design a new Kullback-Leibler (KL) distance based weighting algorithm according to the proposed weighting concept to real-world problems in which the label information is often not provided. We examine our algorithm in ASR with mismatch by different speakers and also in ASI with mismatch by channel noises. Experimental results demonstrate the effectiveness and robustness of our proposed method. Yun Fu 0001, Ming Liu 0009, Mark Hasegawa-Johnson, Thomas S. Huang |
ICME | 4 |
| 2007 | Frequency domain correspondence for speaker normalizationabstractDue to physiology and linguistic difference between speakers, the spectrum pattern for the same phoneme of two speakers can be quite dissimilar. Without appropriate alignment on the frequency axis, the misalignment will reduce the modeling efficiency resutling in performance degradation. In this paper, a novel data-driven framework is proposed to build the alignment of the frequency axes of two speakers. This alignment between two frequency axes is essentially a frequency domain correspondence of the two speakers. To establish the correspondence, we formulate the task as a global optimal matching problem. The local matching of frequency bins is achieved by comparing the local feature of the spectrogram along the frequency bins. The local feature is actually capturing the local pattern in the spectrogram. Given the local matching score, a dynamic programming is then applied to find the optimal correspondence. Experiments on TIMIT corpus and TIDIGITS corpus clearly show the effectiveness of this method. 1. Ming Liu 0009, Mark Hasegawa-Johnson, Thomas S. Huang, Zhengyou Zhang |
INTERSPEECH | 3 |
| 2007 | A Multi-Stream Approach to Audiovisual Automatic Speech RecognitionabstractThis paper proposes a multi-stream approach to automatic audiovisual speech recognition, based in part on Hickok and Poeppel's dual-stream model of human speech processing. The dual-stream model proposes that semantic networks may be accessed by at least three parallel neural streams: at least two ventral streams that map directly from acoustics to words (with different time scales), and at least one dorsal stream that maps from acoustics to articulation. Our implementation represents each of these streams by a dynamic Bayesian network; disagreements between the three streams are resolved using a voting scheme. The proposed algorithm was tested using the CUAVE audiovisual speech corpus. Results indicate that the ventral stream model tends to make fewer mistakes in the labeling of vowels, while the dorsal stream model tends to make fewer mistakes in the labeling of consonants; the recognizer voting scheme takes advantage of these differences to reduce overall word error rate. Mark Hasegawa-Johnson |
MMSP | 1 |
| 2006 | Hmm-Based and Svm-Based Recognition of the Speech of Talkers With Spastic DysarthriaabstractThis paper studies the speech of three talkers with spastic dysarthria caused by cerebral palsy. All three subjects share the symptom of low intelligibility, but causes differ. First, all subjects tend to reduce or delete word-initial consonants; one subject deletes all consonants. Second, one subject exhibits a painstaking stutter. Two algorithms were used to develop automatic isolated digit recognition systems for these subjects. HMM-based recognition was successful for two subjects, but failed for the subject who deletes all consonants. Conversely, digit recognition experiments assuming a fixed word length (using SVMs) were successful for two subjects, but failed for the subject with the stutter. Mark Hasegawa-Johnson, Jon R. Gunderson, Adrienne Perlman, Thomas S. Huang |
ICASSP (3) | 1 |
| 2006 | Generalized Optimal Multi-Microphone Speech Enhancement Using Sequential Minimum Variance Distortionless Response(MVDR) Beamforming and PostfilteringabstractA theoretical basis for optimal multichannel speech enhancementis presented, sufficient, flexible to be used with any assumed statistical model and optimality criterion. Any Bayesian optimal one-channel estimator for speech enhancement can be generalized to the multichannel case as a sequentially constructed minimum variance distortionless response (MVDR) beamformer followed by an optimal one-channel postfilter. We present experimental results using the minimum mean-square error log-spectral amplitude (MMSE-logSA) optimality criterion, applied to a statistical model with simplified channel but realistic inter-microphone noise coherence. Word error rate in the audio-visual speech in a car (AVICAR) corpus (moving car, windows open) is reduced from 18% to 9%. Lae-Hoon Kim, Mark Hasegawa-Johnson, Koeng-Mo Sung |
ICASSP (3) | 2 |
| 2006 | Novel time domain multi-class SVMs for landmark detectionabstractThe training of precise speech recognition models depends on accurate segmentation of the phonemes in a training corpus. Segmentation is typically performed using HMMs, but recent speech recognition work suggests that the transient acoustic features characteristic of manner-class phoneme boundaries (landmarks) may be more precisely localized using acoustic classifiers specifically designed for the task of landmark detection. This paper makes an empirical exploration of new features which suit Landmark Detection and the application of Multi-class SVMs that are capable of improving the time alignment of phoneme boundaries proposed by Binary SVMs and HMM-based speech recognizer. On a standard benchmark data set (A database of Telugu- Official Indian Language, spoken by 75 million people), we achieve a new state-of-the-art performance, reducing RMS phone boundary alignment error from 32ms to 22ms. Rahul Chitturi, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2006 | Novel entropy based moving average refiners for HMM landmarksabstractThe training of precise speech recognition models depends on accurate segmentation of the phonemes in a training corpus. Segmentation is typically performed using HMMs, but recent speech recognition work suggests that the transient acoustic features characteristic of manner-class phoneme boundaries (landmarks) may be more precisely localized using acoustic classifiers specifically designed for the task of landmark detection. This paper makes an empirical exploration of entropy based moving average techniques that are capable of improving the time alignment of phoneme boundaries proposed by an HMM-based speech recognizer. On a standard benchmark data set (A database of Hindi – National Language of India), we achieve new state-of-the-art performance, reducing RMS phone boundary alignment error from 28ms to 15ms. Rahul Chitturi, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2006 | Extraction of pragmatic and semantic salience from spontaneous spoken English
Tong Zhang 0005, Mark Hasegawa-Johnson, Stephen E. Levinson |
Speech Commun. | 2 |
| 2006 | Cognitive state classification in a spoken tutorial dialogue system
Tong Zhang 0005, Mark Hasegawa-Johnson, Stephen E. Levinson |
Speech Commun. | 2 |
| 2006 | Prosody dependent speech recognition on radio news corpus of American EnglishabstractDoes prosody help word recognition? This paper proposes a novel probabilistic framework in which word and phoneme are dependent on prosody in a way that reduces word error rates (WER) relative to a prosody-independent recognizer with comparable parameter count. In the proposed prosody-dependent speech recognizer, word and phoneme models are conditioned on two important prosodic variables: the intonational phrase boundary and the pitch accent. An information-theoretic analysis is provided to show that prosody dependent acoustic and language modeling can increase the mutual information between the true word hypothesis and the acoustic observation by exciting the interaction between prosody dependent acoustic model and prosody dependent language model. Empirically, results indicate that the influence of these prosodic variables on allophonic models are mainly restricted to a small subset of distributions: the duration PDFs (modeled using an explicit duration hidden Markov model or EDHMM) and the acoustic-prosodic observation PDFs (normalized pitch frequency). Influence of prosody on cepstral features is limited to a subset of phonemes: for example, vowels may be influenced by both accent and phrase position, but phrase-initial and phrase-final consonants are independent of accent. Leveraging these results, effective prosody dependent allophonic models are built with minimal increase in parameter count. These prosody dependent speech recognizers are able to reduce word error rates by up to 11% relative to prosody independent recognizers with comparable parameter count, in experiments based on the prosodically-transcribed Boston Radio News corpus. Ken Chen 0001, Mark Hasegawa-Johnson, Aaron Cohen, Sarah Borys, Sung-Suk Kim, Jennifer Cole 0001, Jeung-Yoon Choi |
IEEE Trans. Speech Audio Process. | 2 |
| 2005 | Landmark-Based Speech Recognition: Report of the 2004 Johns Hopkins Summer WorkshopabstractThree research prototype speech recognition systems are described, all of which use recently developed methods from artificial intelligence (specifically support vector machines, dynamic Bayesian networks, and maximum entropy classification) in order to implement, in the form of an automatic speech recognizer, current theories of human speech perception and phonology (specifically landmark-based speech perception, nonlinear phonology, and articulatory phonology). All three systems begin with a high-dimensional multiframe acoustic-to-distinctive feature transformation, implemented using support vector machines trained to detect and classify acoustic phonetic landmarks. Distinctive feature probabilities estimated by the support vector machines are then integrated using one of three pronunciation models: a dynamic programming algorithm that assumes canonical pronunciation of each word, a dynamic Bayesian network implementation of articulatory phonology, or a discriminative pronunciation model trained using the methods of maximum entropy classification. Log probability scores computed by these models are then combined, using log-linear combination, with other word scores available in the lattice output of a first-pass recognizer, and the resulting combination score is used to compute a second-pass speech recognition output. Mark Hasegawa-Johnson, James Baker, Sarah Borys, Ken Chen 0001, Emily Coogan, Steven Greenberg, Amit Juneja, Katrin Kirchhoff, Karen Livescu, Srividya Mohan, Jennifer Muller, M. Kemal Sönmez |
ICASSP (1) | 1 |
| 2005 | Distinctive feature based SVM discriminant features for improvements to phone recognition on telephone band speech
Sarah Borys, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2005 | Simultaneous recognition of words and prosody in the Boston University Radio Speech Corpus
Mark Hasegawa-Johnson, Ken Chen 0001, Jennifer Cole 0001, Sarah Borys, Sung-Suk Kim, Aaron Cohen, Tong Zhang 0005, Jeung-Yoon Choi, Taejin Yoon |
Speech Commun. | 1 |
| 2004 | An automatic prosody labeling system using ANN-based syntactic-prosodic model and GMM-based acoustic-prosodic modelabstractAutomatic prosody labeling is important for both speech synthesis and automatic speech understanding. Humans use both syntactic cues and acoustic cues to develop their prediction of prosody for a given utterance. This process can be effectively modeled by an ANN-based syntactic-prosodic model that predicts prosody from syntax and a GMM-based acoustic-prosodic model that predicts prosody from acoustic-prosodic observations. Our experiments on the Radio News Corpus show that ANN is effective in learning the stochastic mapping from the syntactic representation of word strings to prosody labels, with an accuracy of 82.7% for pitch accent labeling and 90.5% for intonational phrase boundary (IPB) labeling. When acoustic observations and reasonably accurate phoneme transcriptions are given, a GMM-based acoustic-prosodic model, coupled with the syntactical-prosodic model, can achieve 84% pitch accent recognition accuracy and 93% IPB recognition accuracy. These results are obtained using different speakers for training and testing and have considerably exceeded all previously reported results on the same corpus, especially for the task of IPB detection. Ken Chen 0001, Mark Hasegawa-Johnson, Aaron Cohen |
ICASSP (1) | 2 |
| 2004 | A factorial HMM approach to simultaneous recognition of isolated digits spoken by multiple talkers on one audio channelabstractThis paper addresses the novel problem of recognizing digits spoken simultaneously by two different talkers. A factorial hidden Markov model architecture is proposed to accurately model the simultaneous utterance of two digits. Nadas' (1999) MIXMAX approximation is extended to a mixture of Gaussians observation PDF which enables the implementation of the proposed system. The multiple digit recognizer is found to successfully recognize pairs of simultaneous utterances of digits at 0db SNR with up to 89% accuracy. Ameya N. Deoras, Mark Hasegawa-Johnson |
ICASSP (1) | 2 |
| 2004 | Formant tracking by mixture state particle filterabstractThis paper presents a mixture state particle filter method for formant tracking during both vowels and consonants. We show that the mixture state particle filter model is able to incorporate prior information about phoneme class into the system, which helps the system to find global optimal solutions. Formant frequencies are defined as eigenfrequencies of the vocal tract in this paper, and by exploring this fact using spectral estimation techniques, the observation PDF of the particle filter can be simplified. We show that by using this likelihood function in the importance weights, the system is able to track the formants using a small number of particles. Yanli Zheng, Mark Hasegawa-Johnson |
ICASSP (1) | 2 |
| 2004 | Modeling and recognition of phonetic and prosodic factors for improvements to acoustic speech recognition modelsabstractThis paper examines the usefulness of including prosodic and phonetic context information in the phoneme model of a speech recognizer. This is done by creating a series of prosodic and phonetic models and then comparing the mutual information between the observations and each possible context variable. Prosodic variables show improvement less often than phone context variables, however, prosodic variables generally show a larger increase in mutual information. A recognizer with allophones defined using the maximum mutual information prosodic and phonetic variables outperforms a recognizer with allophones defined exclusively using phonetic variables. 1. Sarah Borys, Aaron Cohen, Mark Hasegawa-Johnson, Jennifer Cole 0001 |
INTERSPEECH | 3 |
| 2004 | Modeling pronunciation variation using artificial neural networks for English spontaneous speechabstractPronunciation variation in conversational speech has caused significant amount of word errors in large vocabulary automatic speech recognition. Rule-based approaches and decision-tree based approaches have been previously proposed to model pronunciation variation. In this paper, we report our work on modeling pronunciation variation using artificial neural networks (ANN). The results we achieved are significantly better than previously published ones on two different corpora, indicating that ANN may be better suited for modeling pronunciation variation than other statistical models that have been previously investigated. Our experiments indicate that binary distinctive features can be used to effectively represent the phonological context. We also find that including pitch accent feature in input improves the prediction of pronunciation variation on a ToBI-labeled subset of the Switchboard corpus. Ken Chen 0001, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2004 | Source separation using particle filtersabstractOur goal is to study the statistical methods for source separation based on temporal and frequency specific features by using particle filtering. Particle filtering is an advanced state-space Bayesian estimation technique that supports non-Gaussian and nonlinear models along with time-varying noise, allowing for a more accurate model of the underlying system dynamics. We present a system that combines standard speech processing techniques in a novel method to separate two noisy speech sources. The system models the pitch and amplitude over time separately, and adopts particle filtering to reduce complexity by generating a discrete distribution that approximates well the desired continuous distribution. Preliminary results that demonstrate the separation of two noisy sources using this system are presented. Mital Gandhi, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2004 | A factorial HMM aproach to robust isolated digit recognition in background musicabstractThis paper presents a novel solution to the problem of isolated digit recognition in background music. A Factorial Hidden Markov Model (FHMM) architecture is proposed to accurately model the simultaneous occurrence of two independent processes, such as an utterance of a digit and an excerpt of music. The FHMM is implemented with its equivalent HMM by extending Nadas’ MIXMAX algorithm to a mixture of Gaussians PDF. At around 0 dB SNR, the proposed system shows an average relative reduction in word error rate of 57% in the recognition of isolated digits in background music. Mark Hasegawa-Johnson, Ameya N. Deoras |
INTERSPEECH | 1 |
| 2004 | Automatic detection of contrast for speech understandingabstractContrast is a very popular phenomenon in spoken language, and carries very important information to help understanding contents and structures of spoken language. In this paper, we propose an idea of automatic contrast detection as an effort for better speech understanding. We study the automatic tagging of three specific types of contrast: symmetric contrast, contrastive focus, and contrastive topic. We label the three types of contrasted words as contrast (C), and other words as noncontrast (¬C). The classification of contrast events is based on prosodic, spectral, and part-of-speech (POS) information sources. The integration of different knowledge sources is realized by a time-delay recursive neural network (TDRNN). The approach we proposed was testified on 235 spontaneous utterances consisting of 3500 words (samples). The contrast detection was speaker independent. The tests yielded an average of 87.9% classification rate. Mark Hasegawa-Johnson, Stephen E. Levinson, Tong Zhang 0005 |
INTERSPEECH | 1 |
| 2004 | Children's emotion recognition in an intelligent tutoring scenarioabstractThis paper presents an approach to automatically recognize emotion which children exhibit in an intelligent tutoring system. Emotion recognition can assist the computer agent to adapt its tutorial strategies to improve the efficiency of knowledge transmission. In this study, we detect three emotional classes: confidence, puzzle, and hesitation. Emotion is detected by means of lexical, prosodic, spectral, and syntactic analyses of users’ speech. An automatic speech recognition system serves as the fundamental constituent of the system. A robust classification and regression tree (CART) integrates the various information sources together for final decision. The effectiveness of the proposed approach has been tested on data collected by Wizard-of-Oz (WoZ) experiments. Our emotion recognition was speaker-independent, and yielded 91.3% accuracy. The test results showed that the spectral and duration-related prosodic features played very important roles in emotion recognition. Mark Hasegawa-Johnson, Stephen E. Levinson, Tong Zhang 0005 |
INTERSPEECH | 1 |
| 2004 | AVICAR: audio-visual speech corpus in a car environmentabstractWe describe a large audio-visual speech corpus recorded in a car environment, as well as the equipment and procedures used to build this corpus. Data are collected through a multi-sensory array consisting of eight microphones on the sun visor and four video cameras on the dashboard. The script for the corpus consists of four categories: isolated digits, isolated letters, phone numbers, and sentences, all in English. Speakers from various language backgrounds are included, 50 male and 50 female. In order to vary the signal-to-noise ratio, each script has five different noise conditions: idling, driving at 35 mph with windows open and closed, and driving at 55 mph with windows open and closed. The corpus is available through Bowon Lee, Mark Hasegawa-Johnson, Camille Goudeseune, Suketu Kamdar, Sarah Borys, Ming Liu 0009, Thomas S. Huang |
INTERSPEECH | 2 |
| 2004 | Intertranscriber reliability of prosodic labeling on telephone conversation using toBIabstractTwo transcribers have labeled prosodic events indepen-dently on a subset of Switchboard corpus using adapted ToBI (TOnes and Break Indices) system. Transcriptions of two types of pitch accents (H * and L*), phrasal accents (H- and L-) and boundary tones (H % and L%) encoded independently by two transcribers are compared for intertranscriber reliabil-ity. Two commonly used methods of reliability measurement, ‘transcriber-pair-word ’ comparison and kappa statistic, are used for comparison with previous reports on the intertranscriber consistency. The results obtained from transcriber-pair-word comparison are: The overall agreement on the presence or ab-sence and choice of pitch accent is 86.57%. The agreement on the presence or absence and the choice of phrasal accent is 85.63%. The presence and choice of boundary tone is 89.33%. When both transcribers agreed that there is at least a phrasal tone, the agreement on the choice of the type of either phrasal accent or boundary tone is 73.86%. The kappa coefficient of agreement (K) of 0.7 to 1 indicates the degree of reliability to be from good to perfect. A kappa coefficient of 0.75 is obtained for agreement on the presence or absence of pitch accents, 0.67 for the presence of phrasal accents, and 0.61 for the strength of disjuncture between phrasal accent and boundary tone. Com-parison of the present results with those of previous reliability studies [1][2][3][4] suggests that some higher agreement rates for this study may result from our adoption of fewer labeling distinctions in the transcription of pitch accent events. The results for phrase boundary labeling suggest that spontaneous speech of the type found in the Switchboard corpus is harder to code for the degree of disjuncture between prosodic domains than is read speech. 1. Taejin Yoon, Sandra Chavarria, Jennifer Cole 0001, Mark Hasegawa-Johnson |
INTERSPEECH | 4 |
| 2004 | Stop consonant classification by dynamic formant trajectoryabstractLPC analysis is one of the most powerful techniques in speech analysis. Spectral zeros during consonant or consonant-vowel transition regions introduce difficulties in estimating LPC parameters. In this paper, we propose to estimate formant frequencies from LPC model by MUSIC (Multiple Signal Classification) and ES-PRIT (Estimation of Signal Parameters via Rotational Invariance Techniques). Formant candidates estimated by LS (Least Square), MUSIC and ESPRIT are combined to find an optimal solution. The effectiveness of this algorithm is verified by place classification task of stop consonants. 1. OVERVIEW Classification of stop consonants remains one of the most challenging problems in speech recognition. Halberstadt (1998) [3] reported classification of phones in the TIMIT database using heterogeneous Yanli Zheng, Mark Hasegawa-Johnson, Sarah Borys |
INTERSPEECH | 2 |
| 2004 | Semantic analysis for a speech user interface in an intelligent tutoring systemabstractIn this paper, we describe the strategy of semantic analysis for a speech user interface that is designed for a multimodal intelligent tutoring system. The semantic analysis involves three phases: semantic parsing, salient words/phrases spotting, and accented word detection. Semantic parsing attempts to represent the recognized sentence with a well-formed semantic frame. The recognized sentence consists of the a posterior most probably hypothesized words given the acoustic evidence, and is compliant with the grammatical knowledge that is represented by a semantic language model. The salient words/phrases are useful when semantic parsing fails. The accented words are useful when the user response is out of our expectations, and assist to make the computer agent smarter and smarter. Yuexi Ren, Mark Hasegawa-Johnson, Stephen E. Levinson |
IUI | 2 |
| 2004 | Automatic recognition of pitch movements using multilayer perceptron and time-Delay Recursive neural networkabstractThis letter demonstrates hidden Markov model (HMM), multilayer perceptron (MLP), and time-delay recursive neural network (TDRNN) architectures for the purpose of recognizing pitch accents given observation of the F0 and energy trajectories. At an insertion error rate of 25%, the deletion error rates of the MLP, TDRNN, and HMM are 13.2%, 7.9%, and 32.7%, respectively, despite the fact that both MLP and TDRNN have 70% fewer trainable parameters than the HMM. Error analysis suggests that low-pitch accents may require long-term context to correctly recognize, while high-pitch accents may be recognizable based on local pitch contour. Sung-Suk Kim, Mark Hasegawa-Johnson, Ken Chen 0001 |
IEEE Signal Process. Lett. | 2 |
| 2003 | Acoustic segmentation using switching state Kalman filterabstractSegmenting the acoustic signal in the TIMIT database by a switching state Kalman filter model is reported in this paper. According to the assumption that the high dimensional acoustic feature vector of the LSF (line spectrum frequency) of the speech signal is probably embedded in a low dimensional space, a two dimensional vector is used to represent the continuous state vector in this model. The parameters of the model are initialized by PPCA (probabilistic principal component analysis) and first order vector autoregression, and are re-estimated by the EM algorithm. We show that this model can be used to classify vowels, nasals, frication and silence by an approximate Viterbi inference. Yanli Zheng, Mark Hasegawa-Johnson |
ICASSP (1) | 2 |
| 2003 | Prosody dependent speech recognition with explicit duration modelling at intonational phrase boundariesabstractDoes prosody help word recognition? In this paper, we propose a novel probabilistic framework in which word and phoneme are dependent on prosody in a way that improves word recognition. The prosody attribute that we investigate in this study is the duration lengthening effects of the speech segments in the vicinity of intonational phrase boundaries. Explicit Duration Hidden Markov Model (EDHMM) is implemented to provide an accurate phoneme duration model. This study is conducted on Boston University Radio New Corpus with prosodic boundaries marked using ToBI labelling system. We found that lengthening of the phrase final rhymes can be reliably modelled by EDHMM, which significantly improves the prosody dependent acoustic modelling. Conversely, no systematic duration variation is found at phrase initial position. With prosody dependence implemented in acoustic model, pronunciation model and language model, both word recognition accuracy and boundary recognition accuracy are improved by 1% over systems without prosody dependence. Ken Chen 0001, Sarah Borys, Mark Hasegawa-Johnson, Jennifer Cole 0001 |
INTERSPEECH | 3 |
| 2003 | Maximum conditional mutual information projection for speech recognitionabstractLinear discriminant analysis (LDA) in its original modelfree formulation is best suited to classification problems with equal-covariance classes. Heteroscedastic discriminant analysis (HDA) removes this equal covariance constraint, and therefore is more suitable for automatic speech recognition (ASR) systems. However, maximizing HDA objective function does not correspond directly to minimizing the recognition error. In its original formulation, HDA solves a maximum likelihood estimation problem in the original feature space to calculate the HDA transformation matrix. Since the dimension of the original feature space in ASR problems is usually high, the estimation of the HDA transformation matrix becomes computationally expensive and requires a large amount of training data. This paper presents a generalization of LDA that solves these two problems. We start with showing that the calculation of the LDA projection matrix is a maximum mutual information estimation problem in the lower-dimensional space with some constraints on the model of the joint conditional and unconditional probability density functions (PDF) of the features, and then, by relaxing these constraints, we develop a dimensionality reduction approach that maximizes the conditional mutual information between the class identity and the feature vector in the lower-dimensional space given the recognizer model. Mohamed Kamal Omar, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2003 | Non-linear maximum likelihood feature transformation for speech recognitionabstractMost automatic speech recognition systems (ASR) use Hidden Markov model (HMM) with a diagonal-covariance Gaussian mixture model for the state-conditional probability density function. The diagonal-covariance Gaussian mixture can model discrete sources of variability like speaker variations, gender variations, or local dialect, but can not model continuous types of variability that account for correlation between the elements of the feature vector. In this paper, we present a transformation of the acoustic feature vector that minimize an empirical estimate of the relative entropy between the likelihood based on the diagonal-covriance Gaussian mixture HMM model and the true likelihood. We show that this minimization is equivalent to maximizing the likelihood in the original feature space. Based on this formulation, we provide a computationally efficient solution to the problem based on volume-preserving maps; existing linear feature transform designs are shown to be special cases of the proposed solution. Since most of the acoustic features used in ASR are not linear functions of the sources of correlation in the speech signal, we use a non-linear transformation of the features to minimize this objective function. We describe an iterative algorithm to estimate the parameters of both the volume-preserving feature transformation and the hidden Markov models (HMM) that jointly optimize the objective function for an HMM-based speech recognizer. Using this algorithm, we achieved 2% improvement in phoneme recognition accuracy compared to the original system that uses the original Mel-frequency cepstral coeeficients (MFCC) acoustic features. Our approach is compared also to previous similar linear approaches like MLLT and ICA. Mohamed Kamal Omar, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2003 | Approximately independent factors of speech using nonlinear symplectic transformationabstractThis paper addresses the problem of representing the speech signal using a set of features that are approximately statistically independent. This statistical independence simplifies building probabilistic models based on these features that can be used in applications like speech recognition. Since there is no evidence that the speech signal is a linear combination of separate factors or sources, we use a more general nonlinear transformation of the speech signal to achieve our approximately statistically independent feature set. We choose the transformation to be symplectic to maximize the likelihood of the generated feature set. In this paper, we describe applying this nonlinear transformation to the speech time-domain data directly and to the Mel-frequency cepstrum coefficients (MFCC). We discuss also experiments in which the generated feature set is transformed into a more compact set using a maximum mutual information linear transformation. This linear transformation is used to generate the acoustic features that represent the distinctions among the phonemes. The features resulted from this transformation are used in phoneme recognition experiments. The best results achieved show about 2% improvement in recognition accuracy compared to results based on MFCC features. Mohamed Kamal Omar, Mark Hasegawa-Johnson |
IEEE Trans. Speech Audio Process. | 2 |
| 2002 | Auditory-modeling inspired methods of feature extraction for robust automatic speech recognitionabstractThis paper proposes a technique of extracting robust feature vectors for ASR. The technique is inspired by work related to auditory modeling. It involves first filtering the speech signal through a bank of band-pass filters, which are based on a model of the human cochlea. Autocorrelation functions (ACF) are computed on the filters' outputs. Then the individual ACFs are scaled by their corresponding voice indices (VIs), which use information related to the pitch. A summed ACF is then obtained by summing the individual ACFs across the bands. Feature vectors are then computed using standard cepstral analysis, by treating the summed ACF as a regular ACF. Finally, frame indices (FIs) weigh the feature vectors in the time domain. The effectiveness of the proposed techniques, compared to LPCC and MFCC, are demonstrated by comparing the results obtained from simple recognition experiments. Zhinian Jing, Mark Hasegawa-Johnson |
ICASSP | 2 |
| 2002 | Maximum mutual information based acoustic-features representation of phonological features for speech recognitionabstractThis paper addresses the problem of finding a subset of the acoustic feature space that best represents a set of phonological features. A maximum mutual information approach is presented for selecting acoustic features to be combined together to represent the distinctions coded by a set of correlated phonological features. Each set of phonological features is chosen on the basis of acoustic phonetic similarity, so the sets can be considered approximately independent. This means that the output of recognizers that recognize these sets independently using the acoustic representation achieved by an algorithm presented in this paper can be combined together to increase efficiency and robustness of speech recognition systems. The mutual information between the phonological feature sets and their achieved acoustic representation is increased by up to 220% over the best single-type acoustic representation in the feature space of the same length. Mohamed Kamal Omar, Mark Hasegawa-Johnson |
ICASSP | 2 |
| 2002 | An evaluation of using mutual information for selection of acoustic-features representation of phonemes for speech recognitionabstractThis paper addresses the problem of finding a subset of the acoustic feature space that best represents the phoneme set used in a speech recognition system. A maximum mutual information approach is presented for selecting acoustic features to be combined together to represent the distinctions among the phonemes. Mohamed Kamal Omar, Ken Chen 0001, Mark Hasegawa-Johnson, Yigal Brandman |
INTERSPEECH | 3 |
| 2001 | PLP coefficients can be quantized at 400 bpsabstractPrevious work in wireless speech recognition has focused on two methods, namely, quantizing recognition features (e.g. MFCC) or performing recognition using speech coding parameters (e.g. LPC). All of this previous research assumes that the communication channel is only large enough to transmit either speech coding parameters or speech recognition parameters. By contrast, we propose that the speech recognition parameters can be quantized at a rate sufficiently low to allow transmission of both speech coding and speech recognition parameters over a standard cellular channel. In particular, the paper shows that the perceptual LPC (PLP) coefficients can be transmitted at 400 bps with an insignificant loss of digit recognition accuracy. Wira Gunawan, Mark Hasegawa-Johnson |
ICASSP | 2 |
| 2000 | Multivariate-state hidden Markov models for simultaneous transcription of phones and formantsabstractA multivariate-state HMM-an HMM with a vector state variable-can be used to find jointly optimal phonetic and formant transcriptions of an utterance. The complexity of searching a multivariate state space using the Baum-Welch algorithm is substantial, but may be significantly reduced if the formant frequencies are assumed to be conditionally independent given knowledge of the phone. Operating with a known phonetic transcription, the multivariate-state model can provide a maximum a posteriori formant trajectory, complete with confidence limits on each of the formant frequency measurements. The model can also be used as a phonetic classifier by adding the probabilities of all possible formant trajectories. A test system is described which requires only nine trainable parameters per formant per phonetic state: five parameters to model formant transitions, and four to model spectral observations. Further simplifications were achieved through parameter tying. Mark Hasegawa-Johnson |
ICASSP | 1 |
| 2000 | Time-frequency distribution of partial phonetic information measured using mutual information
Mark Hasegawa-Johnson |
INTERSPEECH | 1 |
| 2000 | Signal approximation in Hilbert space and its application on articulatory speech synthesis
Stephen E. Levinson, Mark Hasegawa-Johnson |
INTERSPEECH | 3 |