John H. L. Hansen

dblp:85/2699 · DBLP profile ↗
← Back
491ranked-venue papers
39as first author
55since 2021 · last 2026
0000-0003-1382-9929ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 406 · 29 first-author · 42 since 2021Artificial intelligence and machine learning · 305 · 23 first-author · 34 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 A Multi-Modal Driving Dataset with Event-Level Ground Truth for Behavior Evaluation
John H. L. Hansen
IV2
2026 Advancing automatic speech recognition using feature fusion with self-supervised learning features: A case study on Fearless Steps Apollo corpus
Szu-Jui Chen, John H. L. Hansen
Speech Commun.2
2025 Bridging the Modality Gap: Softly Discretizing Audio Representation for LLM-based Automatic Speech Recognition
abstract
One challenge of integrating speech input with large language models (LLMs) stems from the discrepancy between the continuous nature of audio data and the discrete tokenbased paradigm of LLMs. To mitigate this gap, we propose a method for integrating vector quantization (VQ) into LLM-based automatic speech recognition (ASR). Using the LLM embedding table as the VQ codebook, the VQ module aligns the continuous representations from the audio encoder with the discrete LLM inputs, enabling the LLM to operate on a discretized audio representation that better reflects the linguistic structure. We further create a “soft discretization” of the audio representation by updating the codebook and performing a weighted sum over the codebook embeddings. Empirical results demonstrate that our proposed method significantly improves upon the LLMbased ASR baseline, particularly in out-of-domain conditions. This work highlights the potential of soft discretization as a modality bridge in LLM-based ASR.
Mu Yang, Szu-Jui Chen, Jiamin Xie, John H. L. Hansen
ASRU4
2025 Situational Signal Processing with Ecological Momentary Assessment: Advancing Speech Vocoder Implementation for Naturalistic Cochlear Implant Scenarios
abstract
Cochlear implants (CIs) are surgically implanted medical devices that rely on real-time digital signal processing (DSP) strategies for acoustic-to-sound conversion. Because most fixed strategies have been implemented and tested only in clinical and laboratory settings, the ability for CI systems to adapt to varied feedback in spontaneous environments is limited. To help allocate real-time CI feedback in naturalistic spaces, this study proposes the first CI framework for situational signal processing: “Emaging”, and considers CI vocoded testing approaches to help record and document collected data when CI users are often difficult to recruit for experimental testing. This unprecedented application implements ecological momentary assessment (EMA), an “on-the-go” data collection method for instantaneous feedback from CI subjects. The “Emaging” algorithm solution runs on portable devices alongside CCi-MOBILE, a customized portable CI signal processing platform. This study evaluates two parameters of EMA for the CI participant: sound source localization (SSL) and sound source identification (SSI) for non-spoken sounds. With “Emaging”, CI users document and “tag” situational data from their naturalistic environments in real-time. Due to the many constraints with CI subject recruitment and testing, vocoded simulations with normal hearing (NH) participants can contribute valuable information and considerations aptly integrated with CI algorithm development. “Emaging” and its collected responses from CI, NH, and vocoded (V) subjects provides a unique opportunity for next generational CI processing design that integrates effective sound coding strategies for non-linguistic sound intelligibility and source localization.
Taylor Lawson, John H. L. Hansen
BSN2
2025 Semi-Supervised Speaker Diarization Using Graph Transformers and LLMs on Naturalistic Apollo 11 Data
abstract
Speaker diarization is the process of segmenting and tagging audio streams based on speaker identity. Traditional methods face significant challenges in real-world scenarios when applied to spontaneous multi-speaker conversational speech. The Fearless Steps Apollo 11 corpus (FS-A11) presents real-world challenges such as inconsistent number of speakers per channel, varying speaker utterance duration, and diverse acoustic environments. In this study, we introduce a novel speaker diarization framework that leverages Large Language Models (LLMs) to generate initial speaker change labels by analyzing both speech content and conversational dynamics. These labels are further refined using a proposed multi-segmentation system to obtain refined speaker turns. Finally, we propose a Graph Transformer architecture to build robust speaker embeddings by modeling relationships between audio segments. The embeddings are then clustered using agglomerative hierarchical clustering (AHC) to produce the final diarization output. Evaluation of the proposed system demonstrates a relative improvement in Diarization Error Rate (DER) of +12.8% over baseline systems on the FS-A11 dataset. Furthermore, we analyze and track 5 key speaker roles over the entire Apollo-11 mission to analyze primary speaker engagement and conversational dynamics across mission communication channels.
Meena Chandra Shekar, John H. L. Hansen
ICASSP2
2025 DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification
abstract
Being a form of biometric identification, the security of the speaker identification (SID) system is of utmost importance. To better understand the robustness of SID systems, we aim to perform more realistic attacks in SID, which are challenging for humans and machines to detect. In this study, we propose DiffAttack, a novel timbre-reserved adversarial attack approach, that exploits the capability of a diffusion-based voice conversion (DiffVC) model to generate adversarial fake audio with distinct target speaker attribution. By introducing adversarial constraints into the diffusion-based voice conversion model’s generative process, we aim to craft fake samples that effectively mislead target models while preserving the speaker-wised characteristics. Specifically, inspired by the utilization of randomly sampled Gaussian noise in conventional adversarial attack and diffusion processes, we incorporate adversarial constraints into the reverse diffusion process. As a result, these adversarial constraints subtly guide the reverse diffusion process toward aligning with the target speaker distribution. Our experiments on the LibriTTS dataset indicate that our proposed DiffAttack significantly improves the attack success rate compared to vanilla DiffVC or other methods. Furthermore, objective and subjective evaluations demonstrate that introducing adversarial constraints does not compromise the speech quality generated by the DiffVC model.
Qing Wang 0039, Jixun Yao, Zhaokai Sun, Lei Xie 0001, John H. L. Hansen
ICASSP6
2025 A Deformable Convolution GAN Approach for Speech Dereverberation in Cochlear Implant Users
Hsin-Tien Chiang, John H. L. Hansen
INTERSPEECH2
2025 A Neural Codec Approach for Noise-Robust Bandwidth Expansion
Mu Yang, Szu-Jui Chen, John H. L. Hansen
INTERSPEECH4
2025 Exploring discrete speech units for privacy-preserving and efficient speech recognition for school-aged and preschool children
Satwik Dutta, Dwight Irvin, John H. L. Hansen
Int. J. Hum. Comput. Stud.3
2024 T-EnFP: An Efficient Transformer Encoder-Based System for Driving Behavior Classification
abstract
Recently, Transformer-based architectures have been explored for classifying driving behavior. Although the Transformer effectively employs self-attention for global temporal learning, the presence of redundant modules can detrimentally affect task-specific performance and overall efficiency. In this study, we optimize the method for classification task using Transformer encoder-based network in two steps. First, we propose the feature embedding and encoding modules for time series data. Second, we propose a Time Series Transformer encoder-based system (T-EnFP) that is suitable for and performs time series data classification tasks. We evaluate the proposed approaches on two naturalistic driving behavior datasets, UAH-Drivest and UTDrive dataset. The proposed models achieve 0.97, 0.94 and 0.99 F1 score for different subsets of UAH-Driveset, outperforming the previously proposed Transformer-based models and LSTM-based models. On the UTDrive dataset, the proposed model achieves the best result with a 0.98 F1 score compared to other baseline models.
John H. L. Hansen
ICASSP2
2024 Fearless Steps Apollo: Team Communications Based Community Resource Development for Science, Technology, Education, and Historical Preservation
abstract
The Fearless Steps Apollo (FS-APOLLO) resource is a collection of 150,000 hours of audio, associated meta-data, and supplemental speech technology infrastructure intended to benefit the (i) speech processing technology, (ii) communication science, team-based psychology, and (iii) education/STEM, history/preservation/archival communities. The FS-APOLLO initiative which started in 2014 has since resulted in the preservation of over 75,000 hours of NASA Apollo Missions audio. Systems created for this audio collection have led to the emergence of several new Speech and Language Technologies (SLT). This paper seeks to provide an overview of the latest advancements in the FS-Apollo effort and explore upcoming strategies in big-data deployment, outreach, and novel avenues of K-12 and STEM education facilitated through this resource.
John H. L. Hansen, Aditya Joglekar, Meena Chandra Shekar, Szu-Jui Chen
ICASSP1
2024 Situational Signal Processing with Ecological Momentary Assessment: Leveraging Environmental Context for Cochlear Implant Users
abstract
Technological advancements for biomedical interfaces and devices, such as cochlear implants (CIs), depend on the integration of novel signal processing strategies enhanced by situational real-time feedback in naturalistic spontaneous environments. This study proposes the first CI framework for situational signal processing, "Emaging", which incorporates ecological momentary assessment (EMA) on portable and wearable devices. The original "Emaging" application operates simultaneously with the CCi-MOBILE platform, a customized portable signal processing CI platform. EMA is an integrative variant of behavioral medicine research that evaluates physiological and psychological outcomes of individuals in natural contexts and incorporates real-time signaling that prompts the subject to report on their current state in a non-clinical environment. "Emaging" considers two parameters of EMA for the CI subject: sound source localization (SSL) and sound source identification (SSI) for non-spoken sounds. SSL involves neurological coding of interaural time and level differences (ITDs and ILDs), which enable the listener to localize auditory stimuli. SSI, which entails identifying what sound was heard, is especially challenging for CI users due to commercial CI algorithms predominantly optimized for speech perception. With "Emaging", CI users document and "tag" their respective experiences with SSL and SSI in a natural real-time environment on their portable and wearable devices. The "tagging" can run simultaneously with regular CCi-MOBILE processing in online or offline scenarios. CCi-MOBILE’s synchronization with "Emaging" provides a unique opportunity for next generational algorithm design based on situational signal processing algorithms. Such algorithms are conducive to more effective non-linguistic sound coding strategies for next-generation biomedical hearing devices.
Taylor Lawson, John H. L. Hansen
ICASSP2
2024 Dual-Path Minimum-Phase and All-Pass Decomposition Network for Single Channel Speech Dereverberation
abstract
With the development of deep neural networks (DNN), many DNN-based speech dereverberation approaches have been proposed to achieve significant improvement over the traditional methods. However, most deep learning-based dereverberation methods solely focus on suppressing time-frequency domain reverberations without utilizing cepstral domain features which are potentially useful for dereverberation. In this paper, we propose a dual-path neural network structure to separately process minimum-phase and all-pass components of single channel speech. First, we decompose speech signal into minimum-phase and all-pass components in cepstral domain, then Conformer embedded U-Net is used to remove reverberations of both components. Finally, we combine these two processed components together to synthesize the enhanced output. The performance of proposed method is tested on REVERB-Challenge evaluation dataset in terms of commonly used objective metrics. Experimental results demonstrate that our method outperforms other compared methods.
Szu-Jui Chen, John H. L. Hansen
ICASSP3
2024 Efficient Adapter Tuning of Pre-Trained Speech Models for Automatic Speaker Verification
abstract
With excellent generalization ability, self-supervised speech models have shown impressive performance on various downstream speech tasks in the pre-training and fine-tuning paradigm. However, as the growing size of pre-trained models, fine-tuning becomes practically unfeasible due to heavy computation and storage overhead, as well as the risk of overfitting. Adapters are lightweight modules inserted into pre-trained models to facilitate parameter-efficient adaptation. In this paper, we propose an effective adapter framework designed for adapting self-supervised speech models to the speaker verification task. With a parallel adapter design, our proposed framework inserts two types of adapters into the pre-trained model, allowing the adaptation of latent features within intermediate Transformer layers and output embeddings from all Transformer layers. We conduct comprehensive experiments to validate the efficiency and effectiveness of the proposed framework. Experimental results on the VoxCeleb1 dataset demonstrate that the proposed adapters surpass fine-tuning and other parameter-efficient transfer learning methods, achieving superior performance while updating only 5% of the parameters.
Mufan Sang, John H. L. Hansen
ICASSP2
2024 Apollo's Unheard Voices: Graph Attention Networks for Speaker Diarization and Clustering for Fearless Steps Apollo Collection
abstract
Speaker diarization has traditionally been explored using datasets that are either clean, feature a limited number of speakers, or have a large volume of data but lack the complexities of real-world scenarios. This study takes a unique approach by focusing on the Fearless Steps APOLLO audio resource, a challenging data that contains over 70,000 hours of audio data (A-11: 10k hrs), the majority of which remains unlabeled. This corpus presents considerable challenges such as diverse acoustic conditions, high levels of background noise, overlapping speech, data imbalance, and a variable number of speakers with varying utterance duration. To address these challenges, we propose a robust speaker diarization framework built on dynamic Graph Attention Network optimized using data augmentation. Our proposed framework attains a Diarization Error Rate (DER) of 19.6% when evaluated using ground truth speech segments. Notably, our work is the first to recognize, track, and perform conversational analysis on the entire Apollo-11 mission for speakers who were unidentified until now. This work stands as a significant contribution to both historical archiving and the development of robust diarization systems, particularly relevant for challenging real-world scenarios.
Meena Chandra Shekar, John H. L. Hansen
ICASSP2
2024 DNN-based monaural speech enhancement using alternate analysis windows for phase and magnitude modification
John H. L. Hansen
INTERSPEECH2
2024 Monaural Speech Dereverberation Using Deformable Convolutional Networks
abstract
Reverberation and background noise can degrade speech quality and intelligibility when captured by a distant microphone. In recent years, researchers have developed several deep learning (DL)-based single-channel speech dereverberation systems that aim to minimize distortions introduced into speech captured in naturalistic environments. A majority of these DL-based systems enhance an unseen distorted speech signal by applying a predetermined set of weights to regions of the speech spectrogram, regardless of the degree of distortion within the respective regions. Such a system might not be an ideal solution for dereverberation task. To address this, we present a DL-based end-to-end single-channel speech dereverberation system that uses deformable convolution networks (DCN) that dynamically adjusts its receptive field based on the degree of distortions within an unseen speech signal. The proposed system includes the following components to simultaneously enhance the magnitude and phase responses of speech, which leads to improved perceptual quality: (i) a complex spectrum enhancement module that uses multi-frame filtering technique to implicitly correct the phase response, (ii) a magnitude enhancement module that suppresses dominant reflections and recovers the formant structure using deep filtering (DF) technique, and (iii) a speech activity detection (SAD) estimation module that predicts frame-wise speech activity to suppress residuals in non-speech regions. We assess the performance of the proposed system by employing objective speech quality metrics on both simulated and real speech recordings from the REVERB challenge corpus. The experimental results demonstrate the benefits of using DCNs and multi-frame filtering for speech dereverberation task. We compare the performance of our proposed system against other signal processing (SP) and DL-based systems and observe that it consistently outperforms other approaches across all speech quality metrics.
Vinay Kothapally, John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 Speech Enhancement for Cochlear Implant Recipients Using Deep Complex Convolution Transformer With Frequency Transformation
abstract
The presence of background noise or competing talkers is one of the main communication challenges for cochlear implant (CI) users in speech understanding in naturalistic spaces. These external factors distort the time-frequency (T-F) content including magnitude spectrum and phase of speech signals. While most existing speech enhancement (SE) solutions focus solely on enhancing the magnitude response, recent research highlights the importance of phase in perceptual speech quality. Motivated by multi-task machine learning, this study proposes a deep complex convolution transformer network (DCCTN) for complex spectral mapping, which simultaneously enhances the magnitude and phase responses of speech. The proposed network leverages a complex-valued U-Net structure with a transformer within the bottleneck layer to capture sufficient low-level detail of contextual information in the T-F domain. To capture the harmonic correlation in speech, DCCTN incorporates a frequency transformation block in the encoder structure of the U-Net architecture. The DCCTN learns a complex transformation matrix to accurately recover speech in the T-F domain from a noisy input spectrogram. Experimental results demonstrate that the proposed DCCTN outperforms existing model solutions such as the convolutional recurrent network (CRN), deep complex convolutional recurrent network (DCCRN), and gated convolutional recurrent network (GCRN) in terms of objective speech intelligibility and quality, both for seen and unseen noise conditions. To evaluate the effectiveness of the proposed SE solution, a formal listener evaluation involving four CI recipients was conducted. Results indicate a significant improvement in speech intelligibility performance for CI recipients in noisy environments. Additionally, DCCTN demonstrates the capability to suppress highly non-stationary noise without introducing musical artifacts commonly observed in conventional SE methods.
Nursadul Mamun, John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Filterbank Learning for Noise-Robust Small-Footprint Keyword Spotting
abstract
In the context of keyword spotting (KWS), the replacement of handcrafted speech features by learnable features has not yielded superior KWS performance. In this study, we demonstrate that filterbank learning outperforms handcrafted speech features for KWS whenever the number of filterbank channels is severely decreased. Reducing the number of channels might yield certain KWS performance drop, but also a substantial energy consumption reduction, which is key when deploying common always-on KWS on low-resource devices. Experimental results on a noisy version of the Google Speech Commands Dataset show that filterbank learning adapts to noise characteristics to provide a higher degree of robustness to noise, especially when dropout is integrated. Thus, switching from typically used 40-channel log-Mel features to 8channel learned features leads to a relative KWS accuracy loss of only 3.5% while simultaneously achieving a 6.3× energy consumption reduction.
Iván López-Espejo, Ram C. M. C. Shekar, Zheng-Hua Tan, Jesper Jensen 0001, John H. L. Hansen
ICASSP5
2023 Improving Transformer-Based Networks with Locality for Automatic Speaker Verification
abstract
Recently, Transformer-based architectures have been explored for speaker embedding extraction. Although the Transformer employs the self-attention mechanism to efficiently model the global interaction between token embeddings, it is inadequate for capturing short-range local context, which is essential for the accurate extraction of speaker information. In this study, we enhance the Transformer with the enhanced locality modeling in two directions. First, we propose the Locality-Enhanced Conformer (LE-Confomer) by introducing depth-wise convolution and channel-wise attention into the Conformer blocks. Second, we present the Speaker Swin Transformer (SST) by adapting the Swin Transformer, originally proposed for vision tasks, into speaker embedding network. We evaluate the proposed approaches on the VoxCeleb datasets and a large-scale Microsoft internal multilingual (MS-internal) dataset. The proposed models achieve 0.75% EER on VoxCeleb 1 test set, outperforming the previously proposed Transformer-based models and CNN-based models, such as ResNet34 and ECAPA-TDNN. When trained on the MS-internal dataset, the proposed models achieve promising results with 14.6% relative reduction in EER over the Res2Net50 model.
Mufan Sang, Yong Zhao 0008, Gang Liu 0001, John H. L. Hansen, Jian Wu 0027
ICASSP4
2023 CFTNet: Complex-valued Frequency Transformation Network for Speech Enhancement
Nursadul Mamun, John H. L. Hansen
INTERSPEECH2
2023 Speaker Tracking using Graph Attention Networks with Varying Duration Utterances across Multi-Channel Naturalistic Data: Fearless Steps Apollo-11 Audio Corpus
Meena Chandra Shekar, John H. L. Hansen
INTERSPEECH2
2023 Assessment of Non-Native Speech Intelligibility using Wav2vec2-based Mispronunciation Detection and Multi-level Goodness of Pronunciation Transformer
Ram C. M. C. Shekar, Mu Yang, Kevin Hirschi, Stephen D. Looney, Okim Kang, John H. L. Hansen
INTERSPEECH6
2023 MixRep: Hidden Representation Mixup for Low-Resource Speech Recognition
abstract
In this paper, we present MixRep, a simple and effective data augmentation strategy based on mixup for low-resource ASR. MixRep interpolates the feature dimensions of hidden representations in the neural network that can be applied to both the acoustic feature input and the output of each layer, which generalizes the previous MixSpeech method. Further, we propose to combine the mixup with a regularization along the time axis of the input, which is shown as complementary. We apply MixRep to a Conformer encoder of an E2E LAS architecture trained with a joint CTC loss. We experiment on the WSJ dataset and subsets of the SWB dataset, covering reading and telephony conversational speech. Experimental results show that MixRep consistently outperforms other regularization methods for low-resource ASR. Compared to a strong SpecAugment baseline, MixRep achieves a +6.5\% and a +6.7\% relative WER reduction on the eval92 set and the Callhome part of the eval'2000 set.
Jiamin Xie, John H. L. Hansen
INTERSPEECH2
2023 What Can an Accent Identifier Learn? Probing Phonetic and Prosodic Information in a Wav2vec2-based Accent Identification Model
Mu Yang, Ram C. M. C. Shekar, Okim Kang, John H. L. Hansen
INTERSPEECH4
2023 Single-channel speech separation using soft-minimum permutation invariant training
Midia Yousefi, John H. L. Hansen
Speech Commun.2
2023 DeepComboSAD: Spectro-Temporal Correlation Based Speech Activity Detection for Naturalistic Audio Streams
abstract
Speech activity detection (SAD) serves as a crucial front-end system to several downstream Speech and Language Technology (SLT) tasks such as speaker diarization (SD), speaker identification (SID), and speech recognition (ASR). Recent years have seen deep learning (DL)-based SAD systems designed to improve robustness against static background noise and interfering speakers. However, SAD performance can be severely limited for conversations recorded in naturalistic environments due to dynamic acoustic scenarios and previously unseen nonspeech artifacts. In this study, we propose an end-to-end deep learning framework designed to be robust to time-varying noise profiles observed in naturalistic audio. We develop a novel SAD solution for the UTDallas Fearless Steps Apollo Corpus based on NASA's Apollo missions. The proposed system leverages spectrotemporal correlations with a threshold optimization mechanism to adjust to acoustic variabilities across multiple channels and missions. This system is trained and evaluated on the Fearless Steps Challenge (FSC) corpus (a subset of the Apollo corpus). Experimental results indicate a high degree of adaptability to out-of-domain data, achieving a relative Detection Cost Function (DCF) performance improvement of over 50% compared to the previous FSC baselines and state-of-the-art (SOTA) SAD systems. The proposed model also outperforms the most recent DL-based SOTA systems from FSC Phase-4. Ablation analysis validates the effectiveness of the combined spectro-temporal features computed in the DeepComboSAD network.
Aditya Joglekar, John H. L. Hansen
IEEE Signal Process. Lett.2
2023 Domain Expansion for End-to-End Speech Recognition: Applications for Accent/Dialect Speech
abstract
Training Automatic Speech Recognition (ASR) systems with sequentially incoming data from alternate domains is an essential milestone in order to reach human intelligibility level in speech recognition. The main challenge of sequential learning is that current adaptation techniques result in significant performance degradation for previously-seen domains. To mitigate the catastrophic forgetting problem, this study proposes effective domain expansion techniques for two scenarios: 1) where only new domain data is available, and 2) where both prior and new domain data are available. We examine the efficacy of the approaches through experiments on adapting a model trained with native English to different English accents. For the first scenario, we study several existing and proposed regularization-based approaches to mitigate performance loss of initial data. The experiments demonstrate the superior performance of our proposed Soft KL-Divergence (SKLD)-Model Averaging (MA) approach. In this approach, SKLD first alleviates the forgetting problem during adaptation; next, MA makes the final efficient compromise between the two domains by averaging parameters of the initial and adapted models. For the second scenario, we explore several rehearsal-based approaches, which leverage initial data to maintain the original model performance. We propose Gradient Averaging (GA) as well as an approach which operates by averaging gradients computed for both initial and new domains. Experiments demonstrate that GA outperforms retraining and specifically designed continual learning approaches, such as Averaged Gradient Episodic Memory (AGEM). Moreover, GA significantly improves computational costs over the complete retraining approach.
Shahram Ghorbani, John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Bilateral Cochlear Implant Processing of Coding Strategies With CCi-MOBILE, an Open-Source Research Platform
abstract
While speech understanding for cochlear implant (CI) users in quiet is relatively effective, listeners experience difficulty in identification of speaker and sound location. To assist for better residual hearing abilities and speech intelligibility support, bilateral and bimodal forms of assisted hearing is becoming popular among CI users. Effective bilateral processing calls for testing precise algorithm synchronization and fitting between both left and right ear channels in order to capture interaural time and level difference cues (ITD and ILDs). This work demonstrates bilateral implant algorithm processing using a custom-made CI research platform - CCi-MOBILE, which is capable of capturing precise source localization information and supports researchers in testing bilateral CI processing in real-time naturalistic environments. Simulation-based, objective, and subjective testing has been performed to validate the accuracy of the platform. The subjective test results produced an RMS error of ±8.66° for source localization, which is comparable to the performance of commercial CI processors.
Ria Ghosh, John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Attention and DCT Based Global Context Modeling for Text-Independent Speaker Recognition
abstract
Learning an effective speaker representation is crucial for achieving reliable performance in speaker verification tasks. Speech signals are high-dimensional, long, and variable-length sequences that entail a complex hierarchical structure. Signals may contain diverse information at each time-frequency (TF) location. The standard convolutional layer that operates on neighboring local regions often fails to capture the complex TF global information. Our motivation stems from the need to alleviate these challenges by increasing the modeling capacity, emphasizing significant information, and suppressing possible redundancies in the speaker representation. We aim to design a more robust and efficient speaker recognition system by incorporating the benefits of attention mechanisms and Discrete Cosine Transform (DCT) based signal processing techniques, to effectively represent the global information in speech signals. To achieve this, we propose a general global time-frequency context modeling block for speaker modeling. First, an attention-based context model is introduced to capture the long-range and non-local relationship across different time-frequency locations. Second, a 2D-DCT based context model is proposed to improve model efficiency and examine the benefits of signal modeling. A multi-DCT attention mechanism is presented to improve modeling power with alternate DCT base forms. Finally, the global context information is used to recalibrate salient time-frequency locations by computing the similarity between the global context and local features. The proposed lightweight blocks can be easily incorporated into a speaker model with little additional computational costs. This effectively improves the speaker verification performance compared to the standard ResNet model and Squeeze&Excitation block by a large margin. Detailed ablation studies are also performed to analyze various factors that may impact performance of the proposed individual modules. Our experimental results show that the proposed global context modeling method can efficiently improve the learned speaker representations by achieving channel-wise and time-frequency feature recalibration.
John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Challenges in Metadata Creation for Massive Naturalistic Team-Based Audio Data
Chelzy Belitz, John H. L. Hansen
INTERSPEECH2
2022 Speaker Trait Enhancement for Cochlear Implant Users: A Case Study for Speaker Emotion Perception
Avamarie Brueggeman, John H. L. Hansen
INTERSPEECH2
2022 FeaRLESS: Feature Refinement Loss for Ensembling Self-Supervised Learning Features in Robust End-to-end Speech Recognition
abstract
Self-supervised learning representations (SSLR) have resulted in robust features for downstream tasks in many fields.Recently, several SSLRs have shown promising results on automatic speech recognition (ASR) benchmark corpora.However, previous studies have only shown performance for solitary SSLRs as an input feature for ASR models.In this study, we propose to investigate the effectiveness of diverse SSLR combinations using various fusion methods within end-to-end (E2E) ASR models.In addition, we will show there are correlations between these extracted SSLRs.As such, we further propose a feature refinement loss for decorrelation to efficiently combine the set of input features.For evaluation, we show that the proposed "FeaRLESS learning features" perform better than systems without the proposed feature refinement loss for both the WSJ and Fearless Steps Challenge (FSC) corpora.
Szu-Jui Chen, Jiamin Xie, John H. L. Hansen
INTERSPEECH3
2022 Challenges remain in Building ASR for Spontaneous Preschool Children Speech in Naturalistic Educational Environments
Satwik Dutta, Sarah Anne Tao, Jacob C. Reyna, Rebecca Elizabeth Hacker, Dwight Irvin, Jay Buzhardt, John H. L. Hansen
INTERSPEECH7
2022 Audio Anti-spoofing Using Simple Attention Module and Joint Optimization Based on Additive Angular Margin Loss and Meta-learning
abstract
Automatic speaker verification systems are vulnerable to a variety of access threats, prompting research into the formulation of effective spoofing detection systems to act as a gate to filter out such spoofing attacks. This study introduces a simple attention module to infer 3-dim attention weights for the feature map in a convolutional layer, which then optimizes an energy function to determine each neuron's importance. With the advancement of both voice conversion and speech synthesis technologies, unseen spoofing attacks are constantly emerging to limit spoofing detection system performance. Here, we propose a joint optimization approach based on the weighted additive angular margin loss for binary classification, with a meta-learning training framework to develop an efficient system that is robust to a wide range of spoofing attacks for model generalization enhancement. As a result, when compared to current state-of-the-art systems, our proposed approach delivers a competitive result with a pooled EER of 0.99% and min t-DCF of 0.0289.
John H. L. Hansen, Zhenyu Wang 0011
INTERSPEECH1
2022 Complex-Valued Time-Frequency Self-Attention for Speech Dereverberation
abstract
Several speech processing systems have demonstrated considerable performance improvements when deep complex neural networks (DCNN) are coupled with self-attention (SA) networks. However, the majority of DCNN-based studies on speech dereverberation that employ self-attention do not explicitly account for the inter-dependencies between real and imaginary features when computing attention. In this study, we propose a complex-valued T-F attention (TFA) module that models spectral and temporal dependencies by computing two-dimensional attention maps across time and frequency dimensions. We validate the effectiveness of our proposed complex-valued TFA module with the deep complex convolutional recurrent network (DCCRN) using the REVERB challenge corpus. Experimental findings indicate that integrating our complex-TFA module with DCCRN improves overall speech quality and performance of back-end speech applications, such as automatic speech recognition, compared to earlier approaches for self-attention.
Vinay Kothapally, John H. L. Hansen
INTERSPEECH2
2022 Speech Modification for Intelligibility in Cochlear Implant Listeners: Individual Effects of Vowel- and Consonant-Boosting
Juliana N. Saba, John H. L. Hansen
INTERSPEECH2
2022 Multi-Frequency Information Enhanced Channel Attention Module for Speaker Representation Learning
abstract
Recently, attention mechanisms have been applied successfully in neural network-based speaker verification systems.Incorporating the Squeeze-and-Excitation block into convolutional neural networks has achieved remarkable performance.However, it uses global average pooling (GAP) to simply average the features along time and frequency dimensions, which is incapable of preserving sufficient speaker information in the feature maps.In this study, we show that GAP is a special case of a discrete cosine transform (DCT) on time-frequency domain mathematically using only the lowest frequency component in frequency decomposition.To strengthen the speaker information extraction ability, we propose to utilize multi-frequency information and design two novel and effective attention modules, called Single-Frequency Single-Channel (SFSC) attention module and Multi-Frequency Single-Channel (MFSC) attention module.The proposed attention modules can effectively capture more speaker information from multiple frequency components on the basis of DCT.We conduct comprehensive experiments on the VoxCeleb datasets and a probe evaluation on the 1 st 48-UTD forensic corpus.Experimental results demonstrate that our proposed SFSC and MFSC attention modules can efficiently generate more discriminative speaker representations and outperform ResNet34-SE and ECAPA-TDNN systems with relative 20.9% and 20.2% reduction in EER, without adding extra network parameters.
Mufan Sang, John H. L. Hansen
INTERSPEECH2
2022 DEFORMER: Coupling Deformed Localized Patterns with Global Context for Robust End-to-end Speech Recognition
abstract
Convolutional neural networks (CNN) have improved speech recognition performance greatly by exploiting localized timefrequency patterns.But these patterns are assumed to appear in symmetric and rigid kernels by the conventional CNN operation.It motivates the question: What about asymmetric kernels?In this study, we illustrate adaptive views can discover local features which couple better with attention than fixed views of the input.We replace depthwise CNNs in the Conformer architecture with a deformable counterpart, dubbed this "Deformer".By analyzing our best-performing model, we visualize both local receptive fields and global attention maps learned by the Deformer and show increased feature associations on the utterance level.The statistical analysis of learned kernel offsets provides an insight into the change of information in features with the network depth.Finally, replacing only half of the layers in the encoder, the Deformer improves +5.6% relative WER without a LM and +6.4% relative WER with a LM over the Conformer baseline on the WSJ eval92 set.
Jiamin Xie, John H. L. Hansen
INTERSPEECH2
2022 Improving Mispronunciation Detection with Wav2vec2-based Momentum Pseudo-Labeling for Accentedness and Intelligibility Assessment
abstract
Current leading mispronunciation detection and diagnosis (MDD) systems achieve promising performance via end-to-end phoneme recognition.One challenge of such end-to-end solutions is the scarcity of human-annotated phonemes on natural L2 speech.In this work, we leverage unlabeled L2 speech via a pseudo-labeling (PL) procedure and extend the fine-tuning approach based on pre-trained self-supervised learning (SSL) models.Specifically, we use Wav2vec 2.0 as our SSL model, and fine-tune it using original labeled L2 speech samples plus the created pseudo-labeled L2 speech samples.Our pseudo labels are dynamic and are produced by an ensemble of the online model on-the-fly, which ensures that our model is robust to pseudo label noise.We show that fine-tuning with pseudo labels achieves a 5.35% phoneme error rate reduction and 2.48% MDD F1 score improvement over a labeled-samples-only finetuning baseline.The proposed PL method is also shown to outperform conventional offline PL methods.Compared to the state-of-the-art MDD systems, our MDD solution produces a more accurate and consistent phonetic error diagnosis.In addition, we conduct an open test on a separate UTD-4Accents dataset, where our system recognition outputs show a strong correlation with human perception, based on accentedness and intelligibility.
Mu Yang, Kevin Hirschi, Stephen D. Looney, Okim Kang, John H. L. Hansen
INTERSPEECH5
2022 Assessing child communication engagement and statistical speech patterns for American English via speech recognition in naturalistic active learning spaces
Rasa Lileikyte, Dwight Irvin, John H. L. Hansen
Speech Commun.3
2022 SkipConvGAN: Monaural Speech Dereverberation Using Generative Adversarial Networks via Complex Time-Frequency Masking
abstract
With the advancements in deep learning approaches, the performance of speech enhancing systems in the presence of background noise have shown significant improvements. However, improving the system’s robustness against reverberation is still a work in progress, as reverberation tends to cause loss of formant structure due to smearing effects in time and frequency. A wide range of deep learning-based systems either enhance the magnitude response and reuse the distorted phase or enhance complex spectrogram using a complex time-frequency mask. Though these approaches have demonstrated satisfactory performance, they do not directly address the lost formant structure caused by reverberation. We believe that retrieving the formant structure can help improve the efficiency of existing systems. In this study, we propose SkipConvGAN - an extension of our prior work SkipConvNet. The proposed system’s generator network tries to estimate an efficient complex time-frequency mask, while the discriminator network aids in driving the generator to restore the lost formant structure. We evaluate the performance of our proposed system on simulated and real recordings of reverberant speech from the single-channel task of the REVERB challenge corpus. The proposed system shows a consistent improvement across multiple room configurations over other deep learning-based generative adversarial frameworks.
Vinay Kothapally, John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Multi-Source Domain Adaptation for Text-Independent Forensic Speaker Recognition
abstract
Adapting speaker recognition systems to new environments is a widely-used technique to improve a well-performing model learned from large-scale data towards a task-specific small-scale data scenarios. However, previous studies focus on single domain adaptation, which neglects a more practical scenario where training data are collected from multiple acoustic domains needed in forensic scenarios. Audio analysis for forensic speaker recognition offers unique challenges in model training with multi-domain training data due to location/scenario uncertainty and diversity mismatch between reference and naturalistic field recordings. It is also difficult to directly employ small-scale domain-specific data to train complex neural network architectures due to domain mismatch and performance loss. Fine-tuning is a commonly-used method for adaptation in order to retrain the model with weights initialized from a well-trained model. Alternatively, in this study, three novel adaptation methods based on domain adversarial training, discrepancy minimization, and moment-matching approaches are proposed to further promote adaptation performance across multiple acoustic domains. A comprehensive set of experiments are conducted to demonstrate that: 1) diverse acoustic environments do impact speaker recognition performance, which could advance research in audio forensics, 2) domain adversarial training learns the discriminative features which are also invariant to shifts between domains, 3) discrepancy-minimizing adaptation achieves effective performance simultaneously across multiple acoustic domains, and 4) moment-matching adaptation along with dynamic distribution alignment also significantly promotes speaker recognition performance on each domain, especially for the LENA-field domain with noise compared to all other systems. Advancements shown here in adaptation therefore helper ensure more consistent performance for field operational data in audio forensics.
Zhenyu Wang 0011, John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 Scenario Aware Speech Recognition: Advancements for Apollo Fearless Steps & CHiME-4 Corpora
abstract
In this study, we propose to investigate triplet loss for the purpose of an alternative feature representation for ASR. We consider a general non-semantic speech representation, which is trained with a self-supervised criteria based on triplet loss called TRILL, for acoustic modeling to represent the acoustic characteristics of each audio. This strategy is then applied to the CHiME-4 corpus and CRSS-UTDallas Fearless Steps Corpus, with emphasis on the 100-hour challenge corpus which consists of 5 selected NASA Apollo-11 channels. An analysis of the extracted embeddings provides the foundation needed to characterize training utterances into distinct groups based on acoustic distinguishing properties. Moreover, we also demonstrate that triplet-loss based embedding performs better than i-Vector in acoustic modeling, confirming that the triplet loss is more effective than a speaker feature. With additional techniques such as pronunciation and silence probability modeling, plus multi-style training, we achieve a +5.42% and +3.18% relative WER improvement for the development and evaluation sets of the Fearless Steps Corpus. To explore generalization, we further test the same technique on the 1 channel track of CHiME-4 and observe a +11.90% relative WER improvement for real test data.
Szu-Jui Chen, John H. L. Hansen
ASRU3
2021 Speaker Conditioning of Acoustic Models Using Affine Transformation for Multi-Speaker Speech Recognition
abstract
This study addresses the problem of single-channel Automatic Speech Recognition of a target speaker within an overlap speech scenario. In the proposed method, the hidden representations in the acoustic model are modulated by speaker auxiliary information to recognize only the desired speaker. Affine transformation layers are inserted into the acoustic model network to integrate speaker information with the acoustic features. The speaker conditioning process allows the acoustic model to perform computation in the context of target-speaker auxiliary information. The proposed speaker conditioning method is a general approach and can be applied to any acoustic model architecture. Here, we employ speaker conditioning on a ResNet acoustic model. Experiments on the WSJ corpus show that the proposed speaker conditioning method is an effective solution to fuse speaker auxiliary information with acoustic features for multi-speaker speech recognition, achieving +9% and +20% relative WER reduction for clean and overlap speech scenarios, respectively, compared to the original ResNet acoustic model baseline.
Midia Yousefi, John H. L. Hansen
ASRU2
2021 DEAAN: Disentangled Embedding and Adversarial Adaptation Network for Robust Speaker Representation Learning
abstract
Despite speaker verification has achieved significant performance improvement with the development of deep neural networks, do-main mismatch is still a challenging problem in this field. In this study, we propose a novel framework to disentangle speaker-related and domain-specific features and apply domain adaptation on the speaker-related feature space solely. Instead of performing domain adaptation directly on the feature space where domain information is not removed, using disentanglement can efficiently boost adaptation performance. To be specific, our model’s input speech from the source and target domains is first encoded into different latent feature spaces. The adversarial domain adaptation is conducted on the shared speaker-related feature space to encourage the property of domain-invariance. Further, we minimize the mutual information between speaker-related and domain-specific features for both do-mains to enforce the disentanglement. Experimental results on the VOiCES dataset demonstrate that our proposed framework can effectively generate more speaker-discriminative and domain-invariant speaker representations with a relative 20.3% reduction of EER com-pared to the original ResNet-based system.
Mufan Sang, John H. L. Hansen
ICASSP3
2021 Fearless Steps Challenge Phase-3 (FSC P3): Advancing SLT for Unseen Channel and Mission Data Across NASA Apollo Audio
Aditya Joglekar, Seyed Omid Sadjadi, Meena Chandra Shekar, Christopher Cieri, John H. L. Hansen
Interspeech5
2021 Real-Time Speaker Counting in a Cocktail Party Scenario Using Attention-Guided Convolutional Neural Network
abstract
Most current speech technology systems are designed to operate well even in the presence of multiple active speakers. However, most solutions assume that the number of co-current speakers is known. Unfortunately, this information might not always be available in real-world applications. In this study, we propose a real-time, single-channel attention-guided Convolutional Neural Network (CNN) to estimate the number of active speakers in overlapping speech. The proposed system extracts higher-level information from the speech spectral content using a CNN model. Next, the attention mechanism summarizes the extracted information into a compact feature vector without losing critical information. Finally, the active speakers are classified using a fully connected network. Experiments on simulated overlapping speech using WSJ corpus show that the attention solution is shown to improve the performance by almost 3% absolute over conventional temporal average pooling. The proposed Attention-guided CNN achieves 76.15% for both Weighted Accuracy and average Recall, and 75.80% Precision on speech segments as short as 20 frames (i.e., 200 ms). All the classification metrics exceed 92% for the attention-guided model in offline scenarios where the input signal is more than 100 frames long (i.e., 1s).
Midia Yousefi, John H. L. Hansen
Interspeech2
2021 Development of CNN-Based Cochlear Implant and Normal Hearing Sound Recognition Models Using Natural and Auralized Environmental Audio
abstract
Restoration of auditory function among hearing impaired individuals using Cochlear Implant (CI) technology has contributed significantly towards an improved quality of life. CI users experience greater challenges in recognizing speech effectively in noisy, reverberant, or time-varying diverse environments. Most CI research efforts focus on enhancing speech perception and environmental sound awareness has received little or no attention. This study focuses on a comparative analysis of normal hearing (NH) vs. CI environmental sound recognition using classifiers trained on learned sound representations using a CNN-based sound event model. Sounds experienced by CI listeners are recreated by auralizing electrical stimuli. CCi-MOBILE is used to generate electrical stimuli and Braecker Vocoder is used for auralization. Natural and auralized sound representations are then applied in order to develop NH and CI sound recognition models. Comparative assessment of environmental sound recognition is carried out by analyzing f1-scores and other performance characteristics. Benefits stemming from this research can help CI researchers improve sound recognition performance, develop novel sound processing algorithms, exclusively for environmental sounds, and identify optimal CI electrical stimulation characteristics to enhance sound perception. Among CI users, improvement in environmental sound awareness contributes to improved quality of life.
Ram C. M. C. Shekar, Chelzy Belitz, John H. L. Hansen
SLT3
2021 An investigation of domain adaptation in speaker embedding space for speaker recognition
Fahimeh Bahmaninezhad, John H. L. Hansen
Speech Commun.3
2021 Nonlinear waveform distortion: Assessment and detection of clipping on speech data and systems
abstract
Speech, speaker, and language systems have traditionally relied on carefully collected speech material for training acoustic models. There is an enormous amount of freely accessible audio content. A major challenge, however, is that such data is not professionally recorded, and therefore may contain a wide diversity of background noise, nonlinear distortions, or other unknown environmental or technology-based contamination or mismatch. There is a crucial need for automatic analysis to screen such unknown datasets before acoustic model development training, or to perform input audio purity screening prior to classification. In this study, we propose a waveform based clipping detection algorithm for naturalistic audio streams and examine the impact of clipping at different severities on speech quality measurements and automatic speaker recognition systems. We use the TIMIT and NIST SRE08 corpora as case studies. The results show, as expected, that clipping introduces a nonlinear distortion into clean speech data, which reduces speech quality and performance for speaker recognition. We also investigate what degree of clipping can be present to sustain effective speech system performance. The proposed detection system, which will be released, could contribute to massive new audio collections for speech and language technology development (e.g. Google Audioset (Gemmeke et al., 2017), CRSS-UTDallas Apollo Fearless-Steps (Yu et al., 2014) (19,000 h naturalistic audio from NASA Apollo missions)).
John H. L. Hansen, Allen R. Stauffer
Speech Commun.1
2021 Curriculum Learning based approaches for robust end-to-end far-field speech recognition
Shivesh Ranjan, John H. L. Hansen
Speech Commun.2
2021 Guided Generative Adversarial Neural Network for Representation Learning and Audio Generation Using Fewer Labelled Audio Data
abstract
The Generation power of Generative Adversarial Neural Networks (GANs) has shown great promise to learn representations from unlabelled data while guided by a small amount of labelled data. We aim to utilise the generation power of GANs to learn Audio Representations. Most existing studies are, however, focused on images. Some studies use GANs for speech generation, but they are conditioned on text or acoustic features, limiting their use for other audio, such as instruments, and even for speech where transcripts are limited. This paper proposes a novel GAN-based model that we named Guided Generative Adversarial Neural Network (GGAN), which can learn powerful representations and generate good-quality samples using a small amount of labelled data as guidance. Experimental results based on a speech [Speech Command Dataset (S09)] and a non-speech [Musical Instrument Sound dataset (Nsyth)] dataset demonstrate that using only 5% of labelled data as guidance, GGAN learns significantly better representations than the state-of-the-art models.
Kazi Nazmul Haque, Rajib Rana, Jiajun Liu 0013, John H. L. Hansen, Nicholas Cummins, Carlos Busso, Björn W. Schuller
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 Analysis and Calibration of Lombard Effect and Whisper for Speaker Recognition
abstract
Variations in vocal effort can create challenges for speaker recognition systems that are optimized for use with neutral speech. The Lombard effect and whisper are two commonly-occurring forms of vocal effort variation that result in non-neutral speech, the first due to noise exposure and the second due to intentional adjustment on the part of the speaker. In this article, a comparative evaluation of speaker recognition performance in non-neutral conditions is presented using multiple Lombard effect and whisper corpora. The detrimental impact of these vocal effort variations on discrimination and calibration performance on global, per-corpus, and per-speaker levels is explored using conventional error metrics, along with visual representations of the model and score spaces. A non-neutral speech detector is subsequently introduced and used to inform score calibration in several ways. Two calibration approaches are proposed and shown to reduce error to the same level as an optimal calibration approach that relies on ground-truth vocal effort information. This article contributes a generalizable methodology towards detecting vocal effort variation and using this knowledge to inform and advance speaker recognition system behavior.
Finnian Kelly, John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 Block-Based High Performance CNN Architectures for Frame-Level Overlapping Speech Detection
abstract
Speech technology systems such as Automatic Speech Recognition (ASR), speaker diarization, speaker recognition, and speech synthesis have advanced significantly by the emergence of deep learning techniques. However, none of these voice-enabled systems perform well in natural environmental circumstances, specifically in situations where one or more potential interfering talkers are involved. Therefore, overlapping speech detection has become an important front-end triage step for speech technology applications. This is crucial for large-scale datasets where manual labeling in not possible. A block-based CNN architecture is proposed to address modeling overlapping speech in audio streams with frames as short as 25 ms. The proposed architecture is robust to both: (i) shifts in distribution of network activations due to the change in network parameters during training, (ii) local variations from the input features caused by feature extraction, environmental noise, or room interference. We also investigate the effect of alternate input features including spectral magnitude, MFCC, MFB, and pyknogram on both computational time and classification performance. Evaluation is performed on simulated overlapping speech signals based on the GRID corpus. The experimental results highlight the capability of the proposed system in detecting overlapping speech frames with 90.5% accuracy, 93.5% precision, 92.7% recall, and 92.8% Fscore on same gender overlapped speech. For opposite gender cases, the network scores exceed 95% in all the classification metrics.
Midia Yousefi, John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 A multi-view approach for Mandarin non-native mispronunciation verification
abstract
Traditionally, the performance of non-native mispronunciation verification systems relied on effective phone-level labelling of non-native corpora. In this study, a multi-view approach is proposed to incorporate discriminative feature representations which requires less annotation for non-native mispronunciation verification of Mandarin. Here, models are jointly learned to embed acoustic sequence and multi-source information for speech attributes and bottleneck features. Bidirectional LSTM embedding models with contrastive losses are used to map acoustic sequences and multi-source information into fixed-dimensional embeddings. The distance between acoustic embeddings is taken as the similarity between phones. Accordingly, examples of mispronounced phones are expected to have a small similarity score with their canonical pronunciations. The approach shows improvement over GOP-based approach by +11.23% and single-view approach by +1.47% in diagnostic accuracy for a mispronunciation verification task.
Zhenyu Wang 0011, John H. L. Hansen, Yanlu Xie
ICASSP2
2020 Frame-Based Overlapping Speech Detection Using Convolutional Neural Networks
abstract
Naturalistic speech recordings usually contain speech signals from multiple speakers. This phenomenon can degrade the performance of speech technologies due to the complexity of tracing and recognizing individual speakers. In this study, we investigate the detection of overlapping speech on segments as short as 25 ms using Convolutional Neural Networks. We evaluate the detection performance using different spectral features, and show that pyknogram features outperforms other commonly used speech features. The proposed system can predict overlapping speech with an accuracy of 84% and Fs-core of 88% on a dataset of mixed speech generated based on the GRID dataset.
Midia Yousefi, John H. L. Hansen
ICASSP2
2020 Effect of Spectral Complexity Reduction and Number of Instruments on Musical Enjoyment with Cochlear Implants
Avamarie Brueggeman, John H. L. Hansen
INTERSPEECH2
2020 Mobile-Assisted Prosody Training for Limited English Proficiency: Learner Background and Speech Learning Pattern
abstract
Contains fulltext : 228191.pdf (Publisher’s version ) (Open Access)
Kevin Hirschi, Okim Kang, Catia Cucchiarini, John H. L. Hansen, Keelan Evanini, Helmer Strik
INTERSPEECH4
2020 FEARLESS STEPS Challenge (FS-2): Supervised Learning with Massive Naturalistic Apollo Data
abstract
The Fearless Steps Initiative by UTDallas-CRSS led to the digitization, recovery, and diarization of 19,000 hours of original analog audio data, as well as the development of algorithms to extract meaningful information from this multi-channel naturalistic data resource. The 2020 FEARLESS STEPS (FS-2) Challenge is the second annual challenge held for the Speech and Language Technology community to motivate supervised learning algorithm development for multi-party and multi-stream naturalistic audio. In this paper, we present an overview of the challenge sub-tasks, data, performance metrics, and lessons learned from Phase-2 of the Fearless Steps Challenge (FS-2). We present advancements made in FS-2 through extensive community outreach and feedback. We describe innovations in the challenge corpus development, and present revised baseline results. We finally discuss the challenge outcome and general trends in system development across both phases (Phase FS-1 Unsupervised, and Phase FS-2 Supervised) of the challenge, and its continuation into multi-channel challenge tasks for the upcoming Fearless Steps Challenge Phase-3.
Aditya Joglekar, John H. L. Hansen, Meena Chandra Shekhar, Abhijeet Sangwan
INTERSPEECH2
2020 SkipConvNet: Skip Convolutional Neural Network for Speech Dereverberation Using Optimally Smoothed Spectral Mapping
abstract
The reliability of using fully convolutional networks (FCNs) has been successfully demonstrated by recent studies in many speech applications. One of the most popular variants of these FCNs is the `U-Net', which is an encoder-decoder network with skip connections. In this study, we propose `SkipConvNet' where we replace each skip connection with multiple convolutional modules to provide decoder with intuitive feature maps rather than encoder's output to improve the learning capacity of the network. We also propose the use of optimal smoothing of power spectral density (PSD) as a pre-processing step, which helps to further enhance the efficiency of the network. To evaluate our proposed system, we use the REVERB challenge corpus to assess the performance of various enhancement approaches under the same conditions. We focus solely on monitoring improvements in speech quality and their contribution to improving the efficiency of back-end speech systems, such as speech recognition and speaker verification, trained on only clean speech. Experimental findings show that the proposed system consistently outperforms other approaches.
Vinay Kothapally, Shahram Ghorbani, John H. L. Hansen, Jing Huang 0019
INTERSPEECH4
2020 Open-Set Short Utterance Forensic Speaker Verification Using Teacher-Student Network with Explicit Inductive Bias
abstract
In forensic applications, it is very common that only small naturalistic datasets consisting of short utterances in complex or unknown acoustic environments are available. In this study, we propose a pipeline solution to improve speaker verification on a small actual forensic field dataset. By leveraging large-scale out-of-domain datasets, a knowledge distillation based objective function is proposed for teacher-student learning, which is applied for short utterance forensic speaker verification. The objective function collectively considers speaker classification loss, Kullback-Leibler divergence, and similarity of embeddings. In order to advance the trained deep speaker embedding network to be robust for a small target dataset, we introduce a novel strategy to fine-tune the pre-trained student model towards a forensic target domain by utilizing the model as a finetuning start point and a reference in regularization. The proposed approaches are evaluated on the 1st48-UTD forensic corpus, a newly established naturalistic dataset of actual homicide investigations consisting of short utterances recorded in uncontrolled conditions. We show that the proposed objective function can efficiently improve the performance of teacher-student learning on short utterances and that our fine-tuning strategy outperforms the commonly used weight decay method by providing an explicit inductive bias towards the pre-trained model.
Mufan Sang, John H. L. Hansen
INTERSPEECH3
2020 Cross-Domain Adaptation with Discrepancy Minimization for Text-Independent Forensic Speaker Verification
abstract
Forensic audio analysis for speaker verification offers unique challenges due to location/scenario uncertainty and diversity mismatch between reference and naturalistic field recordings. The lack of real naturalistic forensic audio corpora with ground-truth speaker identity represents a major challenge in this field. It is also difficult to directly employ small-scale domain-specific data to train complex neural network architectures due to domain mismatch and loss in performance. Alternatively, cross-domain speaker verification for multiple acoustic environments is a challenging task which could advance research in audio forensics. In this study, we introduce a CRSS-Forensics audio dataset collected in multiple acoustic environments. We pre-train a CNN-based network using the VoxCeleb data, followed by an approach which fine-tunes part of the high-level network layers with clean speech from CRSS-Forensics. Based on this fine-tuned model, we align domain-specific distributions in the embedding space with the discrepancy loss and maximum mean discrepancy (MMD). This maintains effective performance on the clean set, while simultaneously generalizes the model to other acoustic domains. From the results, we demonstrate that diverse acoustic environments affect the speaker verification performance, and that our proposed approach of cross-domain adaptation can significantly improve the results in this scenario.
Zhenyu Wang 0011, John H. L. Hansen
INTERSPEECH3
2020 Speaker Representation Learning Using Global Context Guided Channel and Time-Frequency Transformations
abstract
In this study, we propose the global context guided channel and time-frequency transformations to model the long-range, non-local time-frequency dependencies and channel variances in speaker representations. We use the global context information to enhance important channels and recalibrate salient time-frequency locations by computing the similarity between the global context and local features. The proposed modules, together with a popular ResNet based model, are evaluated on the VoxCeleb1 dataset, which is a large scale speaker verification corpus collected in the wild. This lightweight block can be easily incorporated into a CNN model with little additional computational costs and effectively improves the speaker verification performance compared to the baseline ResNet-LDE model and the Squeeze&Excitation block by a large margin. Detailed ablation studies are also performed to analyze various factors that may impact the performance of the proposed modules. We find that by employing the proposed L2-tf-GTFC transformation block, the Equal Error Rate decreases from 4.56% to 3.07%, a relative 32.68% reduction, and a relative 27.28% improvement in terms of the DCF score. The results indicate that our proposed global context guided transformation modules can efficiently improve the learned speaker representations by achieving time-frequency and channel-wise feature recalibration.
John H. L. Hansen
INTERSPEECH2
2020 Sensor Fusion of Camera and Cloud Digital Twin Information for Intelligent Vehicles
abstract
With the rapid development of intelligent vehicles and Advanced Driving Assistance Systems (ADAS), a mixed level of human driver engagements is involved in the transportation system. Visual guidance for drivers is essential under this situation to prevent potential risks. To advance the development of visual guidance systems, we introduce a novel sensor fusion methodology, integrating camera image and Digital Twin knowledge from the cloud. Target vehicle bounding box is drawn and matched by combining results of object detector running on ego vehicle and position information from the cloud. The best matching result, with a 79.2% accuracy under 0.7 Intersection over Union (IoU) threshold, is obtained with depth image served as an additional feature source. Game engine-based simulation results also reveal that the visual guidance system could improve driving safety significantly cooperate with the cloud Digital Twin system.
Yongkang Liu 0005, Ziran Wang, Kyungtae Han, Zhenyu Shou, Prashant Tiwari, John H. L. Hansen
IV6
2019 Domain Expansion in DNN-Based Acoustic Models for Robust Speech Recognition
abstract
Training acoustic models with sequentially incoming data - while both leveraging new data and avoiding the forgetting effect - is an essential obstacle to achieving human intelligence level in speech recognition. An obvious approach to leverage data from a new domain (e.g., new accented speech) is to first generate a comprehensive dataset of all domains, by combining all available data, and then use this dataset to retrain the acoustic models. However, as the amount of training data grows, storing and retraining on such a large-scale dataset becomes practically impossible. To deal with this problem, in this study, we study several domain expansion techniques which exploit only the data of the new domain to build a stronger model for all domains. These techniques are aimed at learning the new domain with a minimal forgetting effect (i.e., they maintain original model performance). These techniques modify the adaptation procedure by imposing new constraints including (1) weight constraint adaptation (WCA): keeping the model parameters close to the original model parameters; (2) elastic weight consolidation (EWC): slowing down training for parameters that are important for previously established domains; (3) soft KL-divergence (SKLD): restricting the KL-divergence between the original and the adapted model output distributions; and (4) hybrid SKLD-EWC: incorporating both SKLD and EWC constraints. We evaluate these techniques in an accent adaptation task in which we adapt a deep neural network (DNN) acoustic model trained with native English to three different English accents: Australian, Hispanic, and Indian. The experimental results show that SKLD significantly outperforms EWC, and EWC works better than WCA. The hybrid SKLD-EWC technique results in the best overall performance.
Shahram Ghorbani, Soheil Khorram, John H. L. Hansen
ASRU3
2019 Analyzing Large Receptive Field Convolutional Networks for Distant Speech Recognition
abstract
Despite significant efforts over the last few years to build a robust automatic speech recognition (ASR) system for different acoustic settings, the performance of the current state-of-the-art technologies significantly degrades in noisy reverberant environments. Convolutional Neural Networks (CNNs) have been successfully used to achieve substantial improvements in many speech processing applications including distant speech recognition (DSR). However, standard CNN architectures were not efficient in capturing long-term speech dynamics, which are essential in the design of a robust DSR system. In the present study, we address this issue by investigating variants of large receptive field CNNs (LRF-CNNs) which include deeply recursive networks, dilated convolutional neural networks, and stacked hourglass networks. To compare the efficacy of the aforementioned architectures with the standard CNN for Wall Street Journal (WSJ) corpus, we use a hybrid DNN-HMM based speech recognition system. We extend the study to evaluate the system performances for distant speech simulated using realistic room impulse responses (RIRs). Our experiments show that with fixed number of parameters across all architectures, the large receptive field networks show consistent improvements over the standard CNNs for distant speech. Amongst the explored LRF-CNNs, stacked hourglass network has shown improvements with a 8.9% relative reduction in word error rate (WER) and 10.7% relative improvement in frame accuracy compared to the standard CNNs for distant simulated speech signals.
Salar Jafarlou, Soheil Khorram, Vinay Kothapally, John H. L. Hansen
ASRU4
2019 Transfer Learning Using Raw Waveform Sincnet for Robust Speaker Diarization
abstract
Speaker diarization tells who spoke and when? in an audio stream. SincNet is a recently developed novel convolutional neural network (CNN) architecture where the first layer consists of parameterized sinc filters. Unlike conventional CNNs, SincNet take raw speech waveform as input. This paper leverages SincNet in vanilla transfer learning (VTL) setup. Out-domain data is used for training SincNet-VTL to perform frame-level speaker classification. Trained SincNet-VTL is later utilized as feature extractor for in-domain data. We investigated pooling (max, avg) strategies for deriving utterance-level embedding using frame-level features extracted from trained network. These utterance/segment level embedding are adopted as speaker models during clustering stage in diarization pipeline. We compared the proposed SincNet-VTL embedding with baseline i-vector features. We evaluated our approaches on two corpora, CRSS-PLTL and AMI. Results show the efficacy of trained SincNet-VTL for speaker-discriminative embedding even when trained on small amount of data. Proposed features achieved relative DER improvements of 19.12% and 52.07% for CRSS-PLTL and AMI data, respectively over baseline i-vectors.
Harishchandra Dubey, Abhijeet Sangwan, John H. L. Hansen
ICASSP3
2019 Cross-lingual Text-independent Speaker Verification Using Unsupervised Adversarial Discriminative Domain Adaptation
abstract
Speaker verification systems often degrade significantly when there is a language mismatch between training and testing data. Being able to improve cross-lingual speaker verification system using unlabeled data can greatly increase the robustness of the system and reduce human labeling costs. In this study, we introduce an unsupervised Adversarial Discriminative Domain Adaptation (ADDA) method to effectively learn an asymmetric mapping that adapts the target domain encoder to the source domain, where the target domain and source domain are speech data from different languages. ADDA, together with a popular Domain Adversarial Training (DAT) approach, are evaluated on a cross-lingual speaker verification task: the training data is in English from NIST SRE04-08, Mixer 6 and Switchboard, and the test data is in Chinese from AISHELL-I. We show that with the ADDA adaptation, Equal Error Rate (EER) of the x-vector system decreases from 9.331% to 7.645%, relatively 18.07% reduction of EER, and 6.32% reduction from DAT as well. Further data analysis of ADDA adapted speaker embedding shows that the learned speaker embeddings can perform well on speaker classification for the target domain data, and are less dependent with respect to the shift in language.
Jing Huang 0019, John H. L. Hansen
ICASSP3
2019 UTD-CRSS Systems for 2018 NIST Speaker Recognition Evaluation
abstract
In this study, we present systems submitted by the Center for Robust Speech Systems (CRSS) from UTDallas to NIST SRE 2018 (SRE18). Three alternative front-end speaker embedding frameworks are investigated, that includes: (i) i-vector, (ii) x-vector, (iii) and a modified triplet speaker embedding system (t-vector). Similar to the previous SRE, language mismatch between training and enrollment/test data, the so-called domain mismatch, remains as a major challenge in this evaluation. In addition, SRE18 also introduces a small portion of audio from an unstructured video corpus in which speaker detection/diarization is supposedly needed to be effectively integrated into speaker recognition for system robustness. In our system development, we focused on: (i) building novel deep neural network based speaker discriminative embedding systems as utterance level feature representations, (ii) exploring alternative dimension reduction methods, back-end classifiers, score normalization techniques which can incorporate unlabeled in-domain data for domain adaptation, (iii) finding an improved data set configurations for the speaker embedding network, LDA/PLDA, and score calibration training (v) and finally, investigating effective score calibration and fusion strategies. The final resulting systems are shown to be both complementary and effective in achieving overall improved speaker recognition performance.
Fahimeh Bahmaninezhad, Shivesh Ranjan, Harishchandra Dubey, John H. L. Hansen
ICASSP6
2019 Semi-supervised Learning with Generative Adversarial Networks for Arabic Dialect Identification
abstract
Dialect Identification (DID) refers to the process of identifying different dialects within the same language class. Compared with more general language identification (LID), DID is a more challenging task because of the substantial similarity between dialects. For an i-vector based LID/DID, prior studies have shown advancements with deep neural networks (DNNs) over Gaussian Mixture Models (GMMs) in acoustic modeling. In this study, a novel i-vector representation which is based on unsupervised bottleneck features is examined as the feature to identify dialects from Arabic broadcast speech. To utilize the unlabeled training data, semi-supervised learning with generative adversarial networks (GANs) are incorporated in the back-end classifier development. Experiments with the proposed method in the third release version of the Multi-Genre Broadcast (MGB-3) Challenge yields the best single system performance among all submitted systems. An overall classification accuracy of 73.8% achieves a +28.8% relative improvement over the MGB-3 baseline with an accuracy of 57.3%, which is the state-of-the-art performance in this DID task. The fused system further achieves an improvement of +39.4% in accuracy.
Qian Zhang 0019, John H. L. Hansen
ICASSP3
2019 A Machine Learning Based Clustering Protocol for Determining Hearing Aid Initial Configurations from Pure-Tone Audiograms
abstract
Of the nearly 35 million people in the USA who are hearing impaired, only an estimated 25% use hearing aids (HA). A good number of HAs are prescribed but not used partially because of the time to convergence for best operation between the audiologist and user. To improve HA retention, it is suggested that a machine learning (ML) protocol could be established which improves initial HA configurations given a user's pure-tone audiogram. This study examines a ML clustering method to predict the best initial HA fitting from a corpus of over 90,000 audiogram-fitting pairs collected from hearing centers throughout the USA. We first examine the final HA comfort targets to determine a limited number of preset configurations using several multi-dimensional clustering methods (Birch, Ward, and k-means). The goal is to reduce the amount of adjustments between the centroid, selected as a fitting configuration to represent the cluster, and the final HA configurations. This may be used to reduce the adjustment cycles for HAs or as preset starting configurations for personal sound amplification products (PSAPs). Using various classification methods, audiograms are mapped to a limited number of potential preset configurations. Finally, the average adjustment between the preset fitting targets and the final fitting targets is examined.
Chelzy Belitz, Hussnain Ali, John H. L. Hansen
INTERSPEECH3
2019 Toeplitz Inverse Covariance Based Robust Speaker Clustering for Naturalistic Audio Streams
abstract
Speaker diarization determines who spoke and when? in an audio stream. In this study, we propose a model-based approach for robust speaker clustering using i-vectors. The ivectors extracted from different segments of same speaker are correlated. We model this correlation with a Markov Random Field (MRF) network. Leveraging the advancements in MRF modeling, we used Toeplitz Inverse Covariance (TIC) matrix to represent the MRF correlation network for each speaker. This approaches captures the sequential structure of i-vectors (or equivalent speaker turns) belonging to same speaker in an audio stream. A variant of standard Expectation Maximization (EM) algorithm is adopted for deriving closed-form solution using dynamic programming (DP) and the alternating direction method of multiplier (ADMM). Our diarization system has four steps: (1) ground-truth segmentation; (2) i-vector extraction; (3) post-processing (mean subtraction, principal component analysis, and length-normalization) ; and (4) proposed speaker clustering. We employ cosine K-means and movMF speaker clustering as baseline approaches. Our evaluation data is derived from: (i) CRSS-PLTL corpus, and (ii) two meetings subset of the AMI corpus. Relative reduction in diarization error rate (DER) for CRSS-PLTL corpus is 43.22% using the proposed advancements as compared to baseline. For AMI meetings IS1000a and IS1003b, relative DER reduction is 29.37% and 9.21%, respectively.
Harishchandra Dubey, Abhijeet Sangwan, John H. L. Hansen
INTERSPEECH3
2019 The 2019 Inaugural Fearless Steps Challenge: A Giant Leap for Naturalistic Audio
John H. L. Hansen, Aditya Joglekar, Meena Chandra Shekhar, Vinay Kothapally, Chengzhu Yu, Lakshmish Kaushik, Abhijeet Sangwan
INTERSPEECH1
2019 Quantifying Cochlear Implant Users' Ability for Speaker Identification Using CI Auditory Stimuli
abstract
Speaker recognition is a biometric modality that uses underlying speech information to determine the identity of the speaker. Speaker Identification (SID) under noisy conditions is one of the challenging topics in the field of speech processing, specifically when it comes to individuals with cochlear implants (CI). This study analyzes and quantifies the ability of CI-users to perform speaker identification based on direct electric auditory stimuli. CI users employ a limited number of frequency bands (8 ∼ 22) and use electrodes to directly stimulate the Basilar Membrane/Cochlear in order to recognize the speech signal. The sparsity of electric stimulation within the CI frequency range is a prime reason for loss in human speech recognition, as well as SID performance. Therefore, it is assumed that CI-users might be unable to recognize and distinguish a speaker given dependent information such as formant frequencies, pitch etc. which are lost to un-simulated electrodes. To quantify this assumption, the input speech signal is processed using a CI Advanced Combined Encoder (ACE) signal processing strategy to construct the CI auditory electrodogram. The proposed study uses 50 speakers from each of three different databases for training the system using two different classifiers under quiet, and tested under both quiet and noisy conditions. The objective result shows that, the CI users can effectively identify a limited number of speakers. However, their performance decreases when more speakers are added in the system, as well as when noisy conditions are introduced. This information could therefore be used for improving CI-user signal processing techniques to improve human SID.
Nursadul Mamun, Ria Ghosh, John H. L. Hansen
INTERSPEECH3
2019 Convolutional Neural Network-Based Speech Enhancement for Cochlear Implant Recipients
abstract
Attempts to develop speech enhancement algorithms with improved speech intelligibility for cochlear implant (CI) users have met with limited success. To improve speech enhancement methods for CI users, we propose to perform speech enhancement in a cochlear filter-bank feature space, a feature-set specifically designed for CI users based on CI auditory stimuli. We leverage a convolutional neural network (CNN) to extract both stationary and non-stationary components of environmental acoustics and speech. We propose three CNN architectures: (1) vanilla CNN that directly generates the enhanced signal; (2) spectral-subtraction-style CNN (SS-CNN) that first predicts noise and then generates the enhanced signal by subtracting noise from the noisy signal; (3) Wiener-style CNN (Wiener-CNN) that generates an optimal mask for suppressing noise. An important problem of the proposed networks is that they introduce considerable delays, which limits their real-time application for CI users. To address this, this study also considers causal variations of these networks. Our experiments show that the proposed networks (both causal and non-causal forms) achieve significant improvement over existing baseline systems. We also found that causal Wiener-CNN outperforms other networks, and leads to the best overall envelope coefficient measure (ECM). The proposed algorithms represent a viable option for implementation on the CCi-MOBILE research platform as a pre-processor for CI users in naturalistic environments.
Nursadul Mamun, Soheil Khorram, John H. L. Hansen
INTERSPEECH3
2019 Adversarial Regularization for End-to-End Robust Speaker Verification
Qing Wang 0039, Sining Sun, Lei Xie 0001, John H. L. Hansen
INTERSPEECH5
2019 Probabilistic Permutation Invariant Training for Speech Separation
abstract
Single-microphone, speaker-independent speech separation is normally performed through two steps: (i) separating the specific speech sources, and (ii) determining the best output-label assignment to find the separation error. The second step is the main obstacle in training neural networks for speech separation. Recently proposed Permutation Invariant Training (PIT) addresses this problem by determining the output-label assignment which minimizes the separation error. In this study, we show that a major drawback of this technique is the overconfident choice of the output-label assignment, especially in the initial steps of training when the network generates unreliable outputs. To solve this problem, we propose Probabilistic PIT (Prob-PIT) which considers the output-label permutation as a discrete latent random variable with a uniform prior distribution. Prob-PIT defines a log-likelihood function based on the prior distributions and the separation errors of all permutations; it trains the speech separation networks by maximizing the log-likelihood function. Prob-PIT can be easily implemented by replacing the minimum function of PIT with a soft-minimum function. We evaluate our approach for speech separation on both TIMIT and CHiME datasets. The results show that the proposed method significantly outperforms PIT in terms of Signal to Distortion Ratio and Signal to Interference Ratio.
Midia Yousefi, Soheil Khorram, John H. L. Hansen
INTERSPEECH3
2019 Multi-domain adversarial training of neural network acoustic models for distant speech recognition
Seyedmahdad Mirsamadi, John H. L. Hansen
Speech Commun.2
2018 Robust Feature Clustering for Unsupervised Speech Activity Detection
abstract
In certain applications such as zero-resource speech processing or very-low resource speech-language systems, it might not be feasible to collect speech activity detection (SAD) annotations. However, the state-of-the-art supervised SAD techniques based on neural networks or other machine learning methods require annotated training data matched to the target domain. This paper establish a clustering approach for fully unsupervised SAD useful for cases where SAD annotations are not available. The proposed approach leverages Harti-gan dip test in a recursive strategy for segmenting the feature space into prominent modes. Statistical dip is invariant to distortions that lends robustness to the proposed method. We evaluate the method on NIST OpenSAD 2015 and NIST OpenSAT 2017 public safety communications data. The results showed the superiority of proposed approach over the two-component GMM baseline.
Harishchandra Dubey, Abhijeet Sangwan, John H. L. Hansen
ICASSP3
2018 Compensation for Domain Mismatch in Text-independent Speaker Recognition
Fahimeh Bahmaninezhad, John H. L. Hansen
INTERSPEECH2
2018 Robust Speaker Clustering using Mixtures of von Mises-Fisher Distributions for Naturalistic Audio Streams
abstract
Speaker Diarization (i.e. determining who spoke and when?) for multi-speaker naturalistic interactions such as Peer-Led Team Learning (PLTL) sessions is a challenging task. In this study, we propose robust speaker clustering based on mixture of multivariate von Mises-Fisher distributions. Our diarization pipeline has two stages: (i) ground-truth segmentation; (ii) proposed speaker clustering. The ground-truth speech activity information is used for extracting i-Vectors from each speechsegment. We post-process the i-Vectors with principal component analysis for dimension reduction followed by lengthnormalization. Normalized i-Vectors are high-dimensional unit vectors possessing discriminative directional characteristics. We model the normalized i-Vectors with a mixture model consisting of multivariate von Mises-Fisher distributions. K-means clustering with cosine distance is chosen as baseline approach. The evaluation data is derived from: (i) CRSS-PLTL corpus; and (ii) three-meetings subset of AMI corpus. The CRSSPLTL data contain audio recordings of PLTL sessions which is student-led STEM education paradigm. Proposed approach is consistently better than baseline leading to upto 44.48% and 53.68% relative improvements for PLTL and AMI corpus, respectively. Index Terms: Speaker clustering, von Mises-Fisher distribution, Peer-led team learning, i-Vector, Naturalistic Audio.
Harishchandra Dubey, Abhijeet Sangwan, John H. L. Hansen
INTERSPEECH3
2018 Leveraging Native Language Information for Improved Accented Speech Recognition
abstract
Recognition of accented speech is a long-standing challenge for automatic speech recognition (ASR) systems, given the increasing worldwide population of bi-lingual speakers with English as their second language. If we consider foreign-accented speech as an interpolation of the native language (L1) and English (L2), using a model that can simultaneously address both languages would perform better at the acoustic level for accented speech. In this study, we explore how an end-to-end recurrent neural network (RNN) trained system with English and native languages (Spanish and Indian languages) could leverage data of native languages to improve performance for accented English speech. To this end, we examine pre-training with native languages, as well as multi-task learning (MTL) in which the main task is trained with native English and the secondary task is trained with Spanish or Indian Languages. We show that the proposed MTL model performs better than the pre-training approach and outperforms a baseline model trained simply with English data. We suggest a new setting for MTL in which the secondary task is trained with both English and the native language, using the same output set. This proposed scenario yields better performance with +11.95% and +17.55% character error rate gains over baseline for Hispanic and Indian accents, respectively.
Shahram Ghorbani, John H. L. Hansen
INTERSPEECH2
2018 Fearless Steps: Apollo-11 Corpus Advancements for Speech Technologies from Earth to the Moon
John H. L. Hansen, Abhijeet Sangwan, Aditya Joglekar, Ahmet Emin Bulut, Lakshmish Kaushik, Chengzhu Yu
INTERSPEECH1
2018 Fusing Text-dependent Word-level i-Vector Models to Screen 'at Risk' Child Speech
Prasanna V. Kothalkar, Johanna Rudolph, Christine Dollaghan, Jennifer McGlothlin, Thomas F. Campbell, John H. L. Hansen
INTERSPEECH6
2018 Testing Paradigms for Assistive Hearing Devices in Diverse Acoustic Environments
abstract
Many individuals worldwide are at risk of hearing loss due to unsafe acoustical exposure and chronic listening experience using personal audio devices. Assistive hearing devices(AHD), such as hearing-aids(HAs) and cochlear-implants(CIs) are a common choice for the restoration and rehabilitation of the auditory function. Audio sound processors in CIs and HAs operate within limits, prescribed by audiologists, not only for acceptable sound perception but also for safety reasons. Signal processing(SP) engineers follow best design practices to ensure reliable performance and incorporate necessary safety checks within the design of SP strategies to ensure safety limits are never exceeded irrespective of acoustic environments. This paper proposes a comprehensive testing and evaluation paradigm to investigate the behavior of audio devices that addresses the safety concerns in diverse acoustic conditions. This is achieved by characterizing the performance of devices with large amounts of acoustic inputs and monitoring the output behavior. The CCi-MOBILE Research-Interface(RI) (used for CI/HA research) is used in this study as the testing paradigm. Factors such as pulse-width(PW), inter-phase gap(IPG) and a number of other parameters are estimated to evaluate the impact of AHDs on hearing comfort, subjective sound quality and characterize audio devices in terms of listening perception and biological safety.
Ram Charan Chandra Shekar, Hussnain Ali, John H. L. Hansen
INTERSPEECH3
2018 Speaker Recognition with Nonlinear Distortion: Clipping Analysis and Impact
John H. L. Hansen
INTERSPEECH2
2018 Assessing Speaker Engagement in 2-Person Debates: Overlap Detection in United States Presidential Debates
Midia Yousefi, Navid Shokouhi, John H. L. Hansen
INTERSPEECH3
2018 Advancing Multi-Accented Lstm-CTC Speech Recognition Using a Domain Specific Student-Teacher Learning Paradigm
abstract
Non-native speech causes automatic speech recognition systems to degrade in performance. Past strategies to address this challenge have considered model adaptation, accent classification with a model selection, alternate pronunciation lexicon, etc. In this study, we consider a recurrent neural network (RNN) with connectionist temporal classification (CTC) cost function trained on multi-accent English data including US (Native), Indian and Hispanic accents. We exploit dark knowledge from a model trained with the multi-accent data to train student models under the guidance of both a teacher model and CTC cost of target transcription. We show that transferring knowledge from a single RNN-CTC trained model toward a student model, yields better performance than the stand-alone teacher model. Since the outputs of different trained CTC models are not necessarily aligned, it is not possible to simply use an ensemble of CTC teacher models. To address this problem, we train accent specific models under the guidance of a single multi-accent teacher, which results in having multiple aligned and trained CTC models. Furthermore, we train a student model under the supervision of the accent-specific teachers, resulting in an even further complementary model, which achieves +20.1% relative Character Error Rate (CER) reduction compared to the baseline trained without any teacher. Having this effective multi-accent model, we can achieve further improvement for each accent by adapting the model to each accent. Using the accent specific model's outputs to regularize the adapting process (i.e., a knowledge distillation version of Kullback-Leibler (KL) divergence) results in even superior performance compared to the conventional approach using general teacher models.
Shahram Ghorbani, Ahmet Emin Bulut, John H. L. Hansen
SLT3
2018 Detection and Calibration of Whisper for Speaker Recognition
abstract
Whisper is a commonly encountered form of speech that differs significantly from modal speech. As speaker recognition technology becomes more ubiquitous, it is important to assess the abilities and limitations of systems in the presence of variability such as whisper. In this paper, a comparative evaluation of whispered speaker recognition performance across two independent datasets is presented. Whisper-neutral speech comparisons are observed to consistently degrade performance relative to both neutral-neutral and whisper-whisper comparisons. An i-vector-based approach to whisper detection is introduced, and is shown to perform accurately across datasets even at short durations. The output of the whisper detector is subsequently used to select score calibration parameters for whispered speech comparisons, leading to a reduction in global calibration and discrimination error.
Finnian Kelly, John H. L. Hansen
SLT2
2018 On the issues of intra-speaker variability and realism in speech, speaker, and language recognition tasks
John H. L. Hansen, Hynek Boril
Speech Commun.1
2018 Modelling and compensation for language mismatch in speaker verification
Abhinav Misra, John H. L. Hansen
Speech Commun.2
2018 Speech Activity Detection in Naturalistic Audio Environments: Fearless Steps Apollo Corpus
abstract
Speech activity detection (SAD) is a fundamental building block for most spoken language technology systems. Developing efficient SAD systems in highly naturalist data scenarios is a challenge. In this study, we investigate the SAD problem on NASAs Apollo space mission data [1]. Apollo data consists of long-term naturalistic audio recordings (i.e., 6-12 day missions). The Apollo data poses various challenges like: 1) noise distortion with variable SNR, 2) channel distortion, 3) very high density of speech, 4) foreground versus background speech, and 5) extended periods of nonspeech activity. In this study, we use threshold optimized combo-SAD [21] as our baseline unsupervised system. This technique was developed to address variable speech/nonspeech density issues in long-term audio data. To mitigate issues related to Apollo audio loops, multispeaker scenarios including foreground versus background conversations within loops, and highly noisy background, a new curriculum learning (CL) based convolutional neural network (CNN) model is developed. This efficient method leverages the long-term learning capability of CNN and CL strategies where data are trained in a manner that improves the efficiency during the learning process. Here, we use signal-to-noise ratio as the learning parameter. Our experiments on free flowing Apollo audio data show that the proposed approach provides a significant improvement in SAD performance (> 10%).
Lakshmish Kaushik, Abhijeet Sangwan, John H. L. Hansen
IEEE Signal Process. Lett.3
2018 Leveraging Frequency-Dependent Kernel and DIP-Based Clustering for Robust Speech Activity Detection in Naturalistic Audio Streams
abstract
Speech activity detection (SAD) is front-end in most speech systems, e.g., speaker verification, speech recognition etc. Supervised SAD typically leverages machine learning models trained on annotated data. For applications like zero-resource speech processing and NIST-OpenSAT-2017 public safety communications task, it might not be feasible to collect SAD annotations. SAD is challenging for naturalistic audio streams containing multiple noise-sources simultaneously. We propose a novel frequency-dependent kernel (FDK) based SAD features. FDK provides enhanced spectral decomposition from which several statistical descriptors are derived. FDK statistical descriptors are combined by principal component analysis into one-dimensional FDK-SAD features. We further proposed two decision backends: First, variable model-size Gaussian mixture model (VMGMM); and second, Hartigan dip-based robust feature clustering. While VMGMM is a model-based approach, the DipSAD is nonparametric. We used both backends for comparative evaluations in two phases: first, standalone SAD performance; and second, the effect of SAD on text-dependent speaker verification using RedDots data. The NIST-OpenSAD-2015 and NIST-OpenSAT-2017 corpora are used for standalone SAD evaluations. We establish two Center for Robust Speech Systems (CRSS) corpora namely CRSS-PLTL-II and CRSS long-duration naturalistic noise corpus. The CRSS corpora facilitate standalone SAD evaluations on naturalistic audio streams. We performed comparative studies of the proposed approaches with multiple baselines including SohnSAD, rSAD, semisupervised Gaussian mixture model, and Gammatone spectrogram features.
Harishchandra Dubey, Abhijeet Sangwan, John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.3
2018 Maximum-Likelihood Linear Transformation for Unsupervised Domain Adaptation in Speaker Verification
abstract
Recent advances in front-end factor analysis through development of i-Vectors have led to significant gains in speaker recognition technology. However, the problem of mismatch between the domains of system development and evaluation data remains a challenging one. This domain mismatch occurs primarily because of the variability in the sources of development and evaluation data. In this study, we propose a novel method of unsupervised probabilistic feature transformation (UPFT) to reduce this domain mismatch by transforming an out-of-domain development data toward in-domain development data. We formulate the alignment of two different domains as a probability density estimation problem. We first train a Gaussian mixture model (GMM) using the out-of-domain i-Vectors. Next, we employ an expectation-maximization (EM) algorithm to fit the means of the GMM to the in-domain i-Vectors by maximizing the overall likelihood. At the optimum, the two domains become closer to each other in the i-Vector space. While reaching the optimum through multiple iterations of the EM, we reparameterize the centroid locations using the following set of transformation parameters: rotation, translation, and scaling. These transformation parameters, which are obtained during the optimization process, are later used to transform the out-of-domain i-Vectors toward in-domain i-Vectors. We observe that such a transformation leads to an improvement in performance of the out-of-domain speaker recognition system. Our proposed method has an added advantage of being completely unsupervised, and thus does not rely on any tuning parameters. We conduct experiments on both 2013 domain adaptation challenge corpus as well as National Institute of Standards and Technology Speaker Recognition Evaluation (SRE)-2016 corpus. On both corpora, we obtain significant improvements using the proposed UPFT solution. Specifically for the SRE-2016 corpus, using a cosine distance scoring based system, we are able to recover almost 90% of the performance gap between an in-domain and out-of-domain system.
Abhinav Misra, John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Curriculum Learning Based Approaches for Noise Robust Speaker Recognition
abstract
Performance of speaker identification (SID) systems is known to degrade rapidly in the presence of mismatch such as noise and channel degradations. This study introduces a novel class of curriculum learning (CL) based algorithms for noise robust speaker recognition. We introduce CL-based approaches at two stages within a state-of-the-art speaker verification system: at the i-Vector extractor estimation and at the probabilistic linear discriminant (PLDA) back-end. Our proposed CL-based approaches operate by categorizing the available training data into progressively more challenging subsets using a suitable difficulty criterion. Next, the corresponding training algorithms are initialized with a subset that is closest to a clean noise-free set, and progressively moving to subsets that are more challenging for training as the algorithms progress. We evaluate the performance of our proposed approaches on the noisy and severely degraded data from the DARPA RATS SID task, and show consistent and significant improvement across multiple test sets over a baseline SID framework with a standard i-Vector extractor and multisession PLDA-based back-end. We also construct a very challenging evaluation set by adding noise to the NIST SRE 2010 C5 extended condition trials, where our proposed CL-based PLDA is shown to offer significant improvements over a traditional PLDA based back-end.
Shivesh Ranjan, John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Language/Dialect Recognition Based on Unsupervised Deep Learning
abstract
Over the past decade, bottleneck features within an i-Vector framework have been used for state-of-the-art language/dialect identification (LID/DID). However, traditional bottleneck feature extraction requires additional transcribed speech information. Alternatively, two types of unsupervised deep learning methods are introduced in this study. To address this limitation, an unsupervised bottleneck feature extraction approach is proposed, which is derived from the traditional bottleneck structure but trained with estimated phonetic labels. In addition, based on a generative modeling autoencoder, two types of latent variable learning algorithms are introduced for speech feature processing, which have been previous considered for image processing/reconstruction. Specifically, a variational autoencoder and adversarial autoencoder are utilized on alternative phase of speech processing. To demonstrate the effectiveness of the proposed methods, three corpora are evaluated: 1) a four Chinese dialect dataset, 2) a five Arabic dialect corpus, and 3) multigenre broadcast challenge corpus (MGB-3) for arabic DID. The proposed features are shown to outperform traditional acoustic feature MFCCs consistently across three corpora. Taken collectively, the proposed features achieve up to a relative +58% improvement in Cavgfor LID/DID without the need of any secondary speech corpora.
Qian Zhang 0019, John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Text-Independent Speaker Verification Based on Triplet Convolutional Neural Network Embeddings
abstract
The effectiveness of introducing deep neural networks into conventional speaker recognition pipelines has been broadly shown to benefit system performance. A novel text-independent speaker verification (SV) framework based on the triplet loss and a very deep convolutional neural network architecture (i.e., Inception-Resnet-v1) are investigated in this study, where a fixed-length speaker discriminative embedding is learned from sparse speech features and utilized as a feature representation for the SV tasks. A concise description of the neural network based speaker discriminative training with triplet loss is presented. An Euclidean distance similarity metric is applied in both network training and SV testing, which ensures the SV system to follow an end-to-end fashion. By replacing the final max/average pooling layer with a spatial pyramid pooling layer in the Inception-Resnet-v1 architecture, the fixed-length input constraint is relaxed and an obvious performance gain is achieved compared with the fixed-length input speaker embedding system. For datasets with more severe training/test condition mismatches, the probabilistic linear discriminant analysis (PLDA) back end is further introduced to replace the distance based scoring for the proposed speaker embedding system. Thus, we reconstruct the SV task with a neural network based front-end speaker embedding system and a PLDA that provides channel and noise variabilities compensation in the back end. Extensive experiments are conducted to provide useful hints that lead to a better testing performance. Comparison with the state-of-the-art SV frameworks on three public datasets (i.e., a prompt speech corpus, a conversational speech Switchboard corpus, and NIST SRE10 10 s-10 s condition) justifies the effectiveness of our proposed speaker embedding system.
Kazuhito Koishida, John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.3
2017 UTD-CRSS submission for MGB-3 Arabic dialect identification: Front-end and back-end advancements on broadcast speech
abstract
This study presents systems submitted by the University of Texas at Dallas, Center for Robust Speech Systems (UTD-CRSS) to the MGB-3 Arabic Dialect Identification (ADI) subtask. This task is defined to discriminate between five dialects of Arabic, including Egyptian, Gulf, Levantine, North African, and Modern Standard Arabic. We develop multiple single systems with different front-end representations and back-end classifiers. At the front-end level, feature extraction methods such as Mel-frequency cepstral coefficients (MFCCs) and two types of bottleneck features (BNF) are studied for an i-Vector framework. As for the back-end level, Gaussian back-end (GB), and Generative Adversarial Networks (GANs) classifiers are applied alternately. The best submission (contrastive) is achieved for the ADI subtask with an accuracy of 76.94% by augmenting the randomly chosen part of the development dataset. Further, with a post evaluation correction in the submitted system, final accuracy is increased to 79.76%, which represents the best performance achieved so far for the challenge on the test dataset.
Ahmet Emin Bulut, Qian Zhang 0019, Fahimeh Bahmaninezhad, John H. L. Hansen
ASRU5
2017 i-Vector/PLDA speaker recognition using support vectors with discriminant analysis
abstract
i-Vector feature representation with probabilistic linear discriminant analysis (PLDA) scoring in speaker recognition system has recently achieved effective permanence even on channel mismatch conditions. In general, experiments carried out using this combined strategy employ linear discriminant analysis (LDA) after the i-Vector extraction phase to suppress irrelevant directions, such as those introduced by noise or channel distortions. However, speaker-related and -non-related variability present in the data may prevent LDA from finding the best projection matrix. In this study, we exclusively use support vectors of each class to find the optimum linear transformation. Post-processing of the i-Vectors by discriminant analysis via support vectors (SVDA) and traditional LDA is evaluated on NIST2010 speaker recognition evaluation (SRE) core and extended core (coreext) conditions. In addition, truncated coreext test data is used to examine the performance of the system for both long and short duration test segments. Computed equal error rate (EER) and minimum detection cost function (minDCF) criteria confirm consistent improvement of SVDA over traditional LDA. The relative improvement in terms of EER and minDCF with SVDA are about 32% and 9%, respectively.
Fahimeh Bahmaninezhad, John H. L. Hansen
ICASSP2
2017 Environment aware speaker diarization for moving targets using parallel DNN-based recognizers
abstract
Current diarization algorithms are commonly applied to the outputs of single non-moving microphones. They do not explicitly identify the content of overlapped segments from multiple speakers or acoustic events. This paper presents an acoustic environment aware child-adult diarization applied to the audio recorded by a single microphone attached to moving targets under realistic high noise conditions. The proposed system exploits a parallel deep neural network and hidden Markov model based approach which enables tracking of rapid turn changes in audio segments as well as capturing the cross talk labels for overlapped speech. It outperforms the state-of-the-art diarization systems without the need to prior clustering or front-end speech activity detection.
Maryam Najafian, John H. L. Hansen
ICASSP2
2017 A study of speaker verification performance with expressive speech
abstract
Expressive speech introduces variations in the acoustic features affecting the performance of speech technology such as speaker verification systems. It is important to identify the range of emotions for which we can reliably estimate speaker verification tasks. This paper studies the performance of a speaker verification system as a function of emotions. Instead of categorical classes such as happiness or anger, which have important intra-class variability, we use the continuous attributes arousal, valence, and dominance which facilitate the analysis. We evaluate an speaker verification system trained with the i-vector framework with a probabilistic linear discriminant analysis (PLDA) back-end. The study relies on a subset of the MSP-PODCAST corpus, which has naturalistic recordings from 40 speakers. We train the system with neutral speech, creating mismatches on the testing set. The results show that speaker verification errors increase when the values of the emotional attributes increase. For neutral/moderate values of arousal, valence and dominance, the speaker verification performance are reliable. These results are also observed when we artificially force the sentences to have the same duration.
Srinivas Parthasarathy, John H. L. Hansen, Carlos Busso
ICASSP3
2017 Acoustic Scene Classification Using a CNN-SuperVector System Trained with Auditory and Spectrogram Image Features
Rakib Hyder, Shabnam Ghaffarzadegan, Zhe Feng 0003, John H. L. Hansen, Taufiq Hasan
INTERSPEECH4
2017 Multi-Channel Apollo Mission Speech Transcripts Calibration
Lakshmish Kaushik, Abhijeet Sangwan, John H. L. Hansen
INTERSPEECH3
2017 Speech Detection and Enhancement Using Single Microphone for Distant Speech Applications in Reverberant Environments
Vinay Kothapally, John H. L. Hansen
INTERSPEECH2
2017 The I4U Mega Fusion and Collaboration for NIST Speaker Recognition Evaluation 2016
abstract
18th Annual Conference of the International Speech Communication Association, INTERSPEECH 2017, Stockholm, Sweden, 20-24 August 2017
Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Anthony Larcher, Andreas Nautsch, Themos Stafylakis, Gang Liu 0001, Mickael Rouvier, Wei Rao 0002, Federico Alegre, Man-Wai Mak, Achintya Kumar Sarkar, Héctor Delgado, Rahim Saeidi, Hagai Aronowitz, Aleksandr Sizov, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Bin Ma 0001, Ville Vestman, Md. Sahidullah, M. Halonen, Anssi Kanervisto, Gaël Le Lan, Fahimeh Bahmaninezhad, Sergey Isadskiy, Christian Rathgeb, Christoph Busch 0001, Georgios Tzimiropoulos, Q. Qian, Q. Zhao, J. Xue, R. Jin, T. Zhao, Pierre-Michel Bousquet, Moez Ajili, Waad Ben Kheder, Driss Matrouf, Zhi Hao Lim, Chenglin Xu, Haihua Xu 0001, Chng Eng Siong, Benoit G. B. Fauve, Kaavya Sriskandaraja, Vidhyasaharan Sethu, W. W. Lin, Dennis Alexander Lehmann Thomsen, Zheng-Hua Tan, Massimiliano Todisco, Nicholas W. D. Evans, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Eliathamby Ambikairajah
INTERSPEECH60
2017 On Multi-Domain Training and Adaptation of End-to-End RNN Acoustic Models for Distant Speech Recognition
Seyedmahdad Mirsamadi, John H. L. Hansen
INTERSPEECH2
2017 Locally Weighted Linear Discriminant Analysis for Robust Speaker Verification
Abhinav Misra, Shivesh Ranjan, John H. L. Hansen
INTERSPEECH3
2017 Improved Gender Independent Speaker Recognition Using Convolutional Neural Network Based Bottleneck Features
Shivesh Ranjan, John H. L. Hansen
INTERSPEECH2
2017 Curriculum Learning Based Probabilistic Linear Discriminant Analysis for Noise Robust Speaker Recognition
Shivesh Ranjan, Abhinav Misra, John H. L. Hansen
INTERSPEECH3
2017 Speech Enhancement Based on Harmonic Estimation Combined with MMSE to Improve Speech Intelligibility for Cochlear Implant Recipients
Dongmei Wang, John H. L. Hansen
INTERSPEECH2
2017 UTD-CRSS Systems for 2016 NIST Speaker Recognition Evaluation
abstract
This document briefly describes the systems submitted by the Center for Robust Speech Systems (CRSS) from The University of Texas at Dallas (UTD) to the 2016 National Institute of Standards and Technology (NIST) Speaker Recognition Evaluation (SRE). We developed several UBM and DNN i-Vector based speaker recognition systems with different data sets and feature representations. Given that the emphasis of the NIST SRE 2016 is on language mismatch between training and enrollment/test data, so-called domain mismatch, in our system development we focused on: (1) using unlabeled in-domain data for centralizing data to alleviate the domain mismatch problem, (2) finding the best data set for training LDA/PLDA, (3) using newly proposed dimension reduction technique incorporating unlabeled in-domain data before PLDA training, (4) unsupervised speaker clustering of unlabeled data and using them alone or with previous SREs for PLDA training, (5) score calibration using only unlabeled data and combination of unlabeled and development (Dev) data as separate experiments.
Fahimeh Bahmaninezhad, Shivesh Ranjan, Chengzhu Yu, Navid Shokouhi, John H. L. Hansen
INTERSPEECH6
2017 Dialect Recognition Based on Unsupervised Bottleneck Features
Qian Zhang 0019, John H. L. Hansen
INTERSPEECH2
2017 Navigation-orientated natural spoken language understanding for intelligent vehicle dialogue
abstract
Voice-based human-machine interfaces are becoming a key feature for next generation intelligent vehicles. For the navigation dialogue systems, it is desired to understand a driver's spoken language in a natural way. This study proposes a two-stage framework, which first converts the audio streams into text sentences through Automatic Speech Recognition (ASR), followed by Natural Language Processing (NLP) to retrieve the navigation-associated information. The NLP stage is based on a Deep Neural Network (DNN) framework, which contains sentence-level sentiment analysis and word/phrase-level context extraction. Experiments are conducted using the CU-Move in-vehicle speech corpus. Results indicate that the DNN architecture is effective for navigation dialog language understanding, whereas the NLP performances are affected by ASR errors. Overall, it is expected that the proposed RNN-based NLP approach, with the corresponding reduced vocabulary designed for navigation-oriented tasks, will benefit the development of advanced intelligent vehicle human-machine interfaces.
Yongkang Liu 0005, John H. L. Hansen
Intelligent Vehicles Symposium3
2017 Assessment and classification of singing quality based on audio-visual features
abstract
The process of speech production changes between speaking and singing due to excitation, vocal tract articulatory positioning, and cognitive motor planning while singing. Singing does not only deviate from typical spoken speech, but it varies across various styles of singing. This is due to alternative genres of music, singing quality of an individual, as well as different languages and cultures. Because of this variation, it is important to establish a baseline system for differentiating between certain aspects of singing. In this study, we establish a classification system that automatically estimates singing quality of candidates from an American TV singing show based on their singing speech acoustics, lip and eye movements. We employ three classifiers that include: Logistic Regression, Naive Bayes and K-nearest neighbor (k-NN) and compare performance of each using unimodal and multimodal features. We also compare performance based on different modalities (speech, lip, eye structure). The results show that audio content performs the best, with modest gains when lip and eye content are fused. An interesting outcome is that lip and eye content achieve an 82% quality assessment while audio achieves 95%. The ability to assess singing quality from lip and eye content at this level is remarkable.
Mangona Bokshi, Fei Tao 0003, Carlos Busso, John H. L. Hansen
VCIP4
2017 Using speech technology for quantifying behavioral characteristics in peer-led team learning sessions
Harishchandra Dubey, Abhijeet Sangwan, John H. L. Hansen
Comput. Speech Lang.3
2017 Automatic Sentiment Detection in Naturalistic Audio
abstract
Audio sentiment analysis using automatic speech recognition is an emerging research area where opinion or sentiment exhibited by a speaker is detected from natural audio. It is relatively underexplored when compared to text based sentiment detection. Extracting speaker sentiment from natural audio sources is a challenging problem. Generic methods for sentiment extraction generally use transcripts from a speech recognition system, and process the transcript using text-based sentiment classifiers. In this study, we show that this baseline system is suboptimal for audio sentiment extraction. Alternatively, new architecture using keyword spotting (KWS) is proposed for sentiment detection. In the new architecture, a text-based sentiment classifier is utilized to automatically determine the most useful and discriminative sentiment-bearing keyword terms, which are then used as a term list for KWS. In order to obtain a compact yet discriminative sentiment term list, iterative feature optimization for maximum entropy sentiment model is proposed to reduce model complexity while maintaining effective classification accuracy. A new hybrid ME-KWS joint scoring methodology is developed to model both text and audio based parameters in a single integrated formulation. For evaluation, two new databases are developed for audio based sentiment detection, namely, YouTube sentiment database and another newly developed corpus called UT-Opinion Opinion audio archive. These databases contain naturalistic opinionated audio collected in real-world conditions. The proposed solution is evaluated on audio obtained from videos in youtube.com and UT-Opinion corpus. Our experimental results show that the proposed KWS based system significantly outperforms the traditional ASR architecture in detecting sentiment for challenging practical tasks.
Lakshmish Kaushik, Abhijeet Sangwan, John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.3
2017 Teager-Kaiser Energy Operators for Overlapped Speech Detection
abstract
Overlapped speech is referred to a monophonic audio signal in which at least two speakers are present at the same time. In this study, the focus is on distinguishing overlapped from single-speaker speech, i.e., overlapped speech detection. We develop an overlap detection algorithm using an enhanced time-frequency representation, called Pyknogram, estimated directly from the input audio signal. Pyknograms use the Teager-Kaiser energy operator to detect resonant time-frequency units and thereby suppress nonharmonic structures. We show how the resulting Pyknograms provide high separability in terms of detecting the presence of interfering speech. Our proposed unsupervised Pyknogram-based detection results in over 30% relative improvement in overlap detection error rates across different signal-to-interference ratios (SIR) compared to baseline systems. In addition, a case study is presented where we evaluate speaker verification performance under different overlap conditions using the GRID database and observe that speaker verification equal error rates (EER) vary from 2% to 30%, depending on the average SIR values introduced to train and test sets. In order to estimate the reliability of speaker verification scores across different trials, overlap detection results are interpreted as low-level information and stacked alongside verification outputs. The resulting high-dimensional space is passed through a support vector machine classifier to find the separating hyperplane between target and imposter scores. Combining overlap detection scores with speaker verification on average yields 20% relative decrease in EER. We also provide an upper bound for this approach using existing overlap labels, which yields 23% relative improvement.
Navid Shokouhi, John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Robust Harmonic Features for Classification-Based Pitch Estimation
abstract
Pitch estimation in diverse naturalistic audio streams remains a challenge for speech processing and spoken language technology. In this study, we investigate the use of robust harmonic features for classification-based pitch estimation. The proposed pitch estimation algorithm is composed of two stages: pitch candidate generation and target pitch selection. Based on energy intensity and spectral envelope shape, five types of robust harmonic features are proposed to reflect pitch associated harmonic structure. A neural network is adopted for modeling the relationship between input harmonic features and output pitch salience for each specific pitch candidate. In the test stage, each pitch candidate is assessed with an output salience that indicates the potential as a true pitch value, based on its input feature vector processed through the neural network. Finally, according to the temporal continuity of pitch values, pitch contour tracking is performed using a hidden Markov model (HMM), and the Viterbi algorithm is used for HMM decoding. Experimental results show that the proposed algorithm outperforms several state-of-the-art pitch estimation methods in terms of accuracy in both high and low levels of additive noise.
Dongmei Wang, Chengzhu Yu, John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.3
2017 Single Sideband Frequency Offset Estimation and Correction for Quality Enhancement and Speaker Recognition
abstract
Communication system mismatch represents a major influence in the losses of both speech quality and speaker recognition system performance. Although microphone and handset differences have been considered for speaker recognition (e.g., NIST SRE), nonlinear communication system differences, such as modulation/demodulation (Mod/DeMod) carrier mismatch, have yet to be explored. While such mismatch was common in traditional analog communications, today, with the diversity and blending of communication technologies, it is reconsidered as a major distortion. This paper is focused on estimating and correcting the frequency-shift distortion resulting from Mod/DeMod carrier frequency mismatch in high-frequency single sideband (HF-SSB) speech. To overcome the drawbacks of existing solutions, a two-step algorithm is proposed to improve estimation performance. In the first step, the offset of speech is scaled to a small frequency interval, which eliminates or reduces the nonuniqueness issue due to the periodicity within the spectrum; the second step performs fine tuning within the estimated predetermined uniqueness interval (UI). For the first time, a statistical framework is developed for UI detection, where an innovative acoustic feature is proposed to represent alternative frequency shifts. Additionally, in the estimation process, statistical techniques such as GMM-SVM, i-Vector, and deep neural networks are applied in the first step to improve the estimation accuracy. An evaluation using DARPA RATS HF-SSB data shows that the proposed algorithm achieves a significant improvement in the estimation performance (up to +35.6% improvement in accuracy), speech quality measurement (up to +27.3% relative improvement in the PESQ score), and speaker verification (up to +59.9% relative improvement in equal error rate).
Hua Xing, John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Active Learning Based Constrained Clustering For Speaker Diarization
abstract
Most speaker diarization research has focused on unsupervised scenarios, where no human supervision is available. However, in many real-world applications, a certain amount of human input could be expected, especially when minimal human supervision brings significant performance improvement. In this study, we propose an active learning based bottom-up speaker clustering algorithm to effectively improve speaker diarization performance with limited human input. Specifically, the proposed active learning based speaker clustering has two different stages: explore and constrained clustering. The explore stage is to quickly discover at least one sample for each speaker for boosting speaker clustering process with reliable initial speaker clusters. After discovering all, or a majority, of the involved speakers during explore stage, the constrained clustering is performed. Constrained clustering is similar to traditional bottom-up clustering process with an important difference that the clusters created during explore stage are restricted from merging with each other. Constrained clustering continues until only the clusters generated from the explore stage are left. Since the objective of active learning based speaker clustering algorithm is to provide good initial speaker models, performance saturates as soon as sufficient examples are ensured for each cluster. To further improve diarization performance with increasing human input, we propose a second method which actively select speech segments that account for the largest expected speaker error from existing cluster assignments for human evaluation and reassignment. The algorithms are evaluated on our recently created Apollo Mission Control Center dataset as well as augmented multiparty interaction meeting corpus. The results indicate that the proposed active learning algorithms are able to reduce diarization error rate significantly with a relatively small amount of human supervision.
Chengzhu Yu, John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.2
2016 Language recognition using deep neural networks with very limited training data
abstract
This study proposes a novel deep neural network (DNN) based approach to language identification (LID) for the NIST 2015 Language Recognition (LRE) i-Vector Machine Learning Challenge. State-of-the-art DNN based LID systems utilize large amounts of labeled training data. The 2015 LRE i-Vector Machine Learning Challenge limits the access to only ready-to-use i-Vectors for LID system training and testing. This poses unique challenges in designing DNN based LID systems, since optimized front-ends and network architectures can no longer be used. We propose to use the training i-Vectors to train an initial DNN for LID. Next, we present a novel strategy to use this initial DNN to estimate out-of-set language labels from the development data. The final DNN for LID is trained using the original training data, and the estimated out-of-set language data. We show that augmenting the training set with out-of-set labels leads to significant improvement in the LID performance. Our approach obtains very competitive costs (defined by NIST) of 26.56, and 25.98 respectively, on the progress and evaluation subsets of the challenge. Since the amount of training data is very limited (300 i-Vectors per language), this study outlines a successful recipe for DNN based LID using very limited resources.
Shivesh Ranjan, Chengzhu Yu, Finnian Kelly, John H. L. Hansen
ICASSP5
2016 F0 estimation for noisy speech by exploring temporal harmonic structures in local time frequency spectrum segment
abstract
In this paper, we propose a noise robust F0 estimation approach by exploring the temporal harmonic structures in local time-frequency (TF) spectrum segment. Since the speech energy is sparsely distributed on the TF plane, the speech harmonic structures occupied in the higher speech energy TF segment are tending to dominate over noise. Thus, we attempt to derive F0 from such high (signal to noise ratio) SNR TF segments rather than full band signal. Our algorithm comprises of two stages: i) F0 candidate estimation for a series of TF segments; ii) F0 tracking based on the acoustic features of each TF segment as well as the F0 temporal continuity constraints. Experimental results show that our approach outperforms the compared methods in terms of F0 estimation accuracy.
Dongmei Wang, John H. L. Hansen
ICASSP2
2016 UTD-CRSS system for the NIST 2015 language recognition i-vector machine learning challenge
abstract
In this paper, we present the system developed by the Center for Robust Speech Systems (CRSS), University of Texas at Dallas, for the NIST 2015 language recognition i-vector machine learning challenge. Our system includes several subsystems, based on Linear Discriminant Analysis - Support Vector Machine (LDA-SVM) and deep neural network (DNN) approaches. An important feature of this challenge is the emphasis on out-of-set language detection. As a result, our system development focuses mainly on the evaluation and comparison of two different out-of-set language detection strategies: direct out-of-set detection and indirect out-of-set detection. These out-of-set detection strategies differ mainly on whether the unlabeled development data are used or not. The experimental results indicate that indirect out-of-set detection strategies used in our system could efficiently exploit the unlabeled development data, and therefore consistently outperform the direct out-of-set detection approach. Finally, by fusing four variants of indirect out-of-set detection based subsystems, our system achieves a relative performance gain of up to 45%, compared to the baseline cosine distance scoring (CDS) system provided by organizer.
Chengzhu Yu, Shivesh Ranjan, Qian Zhang 0019, Abhinav Misra, Finnian Kelly, John H. L. Hansen
ICASSP7
2016 Joint information from nonlinear and linear features for spoofing detection: An i-vector/DNN based approach
abstract
Sustaining automatic speaker verification(ASV) systems from spoofing attacks remains an essential challenge, even if significant progress in ASV has been achieved in recent years. In this study, an automatic spoofing detection approach using an i-vector framework is proposed. Two approaches are used for frame-level feature extraction: cepstral-based Perceptual Minimum Variance Distortionless Response (PMVDR), and non-linear speech-production-motivated Teager Energy Operator (TEO) Critical Band (CB) Autocorrelation Envelope (Auto-Env). An utterance-level i-vector for each recording is formed by concatenating PMVDR and TEO-CB-Auto-Envi-vectors, followed by linear discriminative analysis (LDA) for maximizing the ratio of between-class to within-class scatterings. A Gaussian classifier and DNN are also investigated for back-end scoring. Experiments using the ASVspoof 2015 corpus show that our proposed method successfully detects spoofing attacks. By combining the TEO-CB-Auto-Env and PMVDR features, a relative 76.7% improvement in terms of EER is obtained compared with the best single-feature system.
Shivesh Ranjan, Mahesh Kumar Nandwana, Qian Zhang 0019, Abhinav Misra, Gang Liu 0001, Finnian Kelly, John H. L. Hansen
ICASSP8
2016 Generalized Discriminant Analysis (GDA) for Improved i-Vector Based Speaker Recognition
Fahimeh Bahmaninezhad, John H. L. Hansen
INTERSPEECH2
2016 A Speaker Diarization System for Studying Peer-Led Team Learning Groups
abstract
Peer-led team learning (PLTL) is a model for teaching STEM courses where small student groups meet periodically to collaboratively discuss coursework. Automatic analysis of PLTL sessions would help education researchers to get insight into how learning outcomes are impacted by individual participation, group behavior, team dynamics, etc.. Towards this, speech and language technology can help, and speaker diarization technology will lay the foundation for analysis. In this study, a new corpus is established called CRSS-PLTL, that contains speech data from 5 PLTL teams over a semester (10 sessions per team with 5-to-8 participants in each team). In CRSS-PLTL, every participant wears a LENA device (portable audio recorder) that provides multiple audio recordings of the event. Our proposed solution is unsupervised and contains a new online speaker change detection algorithm, termed G 3 algorithm in conjunction with Hausdorff-distance based clustering to provide improved detection accuracy. Additionally, we also exploit cross channel information to refine our diarization hypothesis. The proposed system provides good improvements in diarization error rate (DER) over the baseline LIUM system. We also present higher level analysis such as the number of conversational turns taken in a session, and speaking-time duration (participation) for each speaker.
Harishchandra Dubey, Lakshmish Kaushik, Abhijeet Sangwan, John H. L. Hansen
INTERSPEECH4
2016 Robustness in Speech, Speaker, and Language Recognition: "You've Got to Know Your Limitations"
John H. L. Hansen, Hynek Boril
INTERSPEECH1
2016 Fusion Strategies for Robust Speech Recognition and Keyword Spotting for Channel- and Noise-Degraded Speech
Vikramjit Mitra, Julien van Hout, Wen Wang 0001, Chris Bartels, Horacio Franco, Dimitra Vergyri, Abeer Alwan, Adam Janin, John H. L. Hansen, Richard M. Stern, Abhijeet Sangwan, Nelson Morgan
INTERSPEECH9
2016 Discussion
Dayana Ribas González, Emmanuel Vincent 0001, John H. L. Hansen, Emma Jokinen, Mirco Ravanelli, Hannes Gamper, Fred Richardson
INTERSPEECH3
2016 Improving Boundary Estimation in Audiovisual Speech Activity Detection Using Bayesian Information Criterion
Fei Tao 0003, John H. L. Hansen, Carlos Busso
INTERSPEECH2
2016 Text-Available Speaker Recognition System for Forensic Applications
Chengzhu Yu, Finnian Kelly, Abhijeet Sangwan, John H. L. Hansen
INTERSPEECH5
2016 A robust diarization system for measuring dominance in Peer-Led Team Learning groups
abstract
Peer-Led Team Learning (PLTL) is a structured learning model where a team leader is appointed to facilitate collaborative problem solving among students for Science, Technology, Engineering and Mathematics (STEM) courses. This paper presents an informed HMM-based speaker diarization system. The minimum duration of short conversational-turns and number of participating students were fed as side information to the HMM system. A modified form of Bayesian Information Criterion (BIC) was used for iterative merging and re-segmentation. Finally, we used the diarization output to compute a novel dominance score based on unsupervised acoustic analysis.
Harishchandra Dubey, Abhijeet Sangwan, John H. L. Hansen
SLT3
2016 Evaluation and calibration of Lombard effects in speaker verification
abstract
The Lombard effect is the involuntary tendency of speakers to increase their vocal effort in noisy environments in order to maintain intelligible communication. This study assesses the impact of the Lombard effect on the performance of a current speaker verification system. Lombard speech produced in the presence of several noise types and noise levels is drawn from the UT-Scope corpus. The performance of an i-vector PLDA (Probabilistic Linear Discriminant Analysis) system is observed to degrade significantly with Lombard speech. The resulting error rates are found to be dependent on the noise type and noise level. A score calibration scheme based on Quality Measure Functions (QMFs) is adopted, allowing noise information to be incorporated into calibration. This approach leads to a reduction in discrimination error relative to conventional calibration.
Finnian Kelly, John H. L. Hansen
SLT2
2016 Speaker independent diarization for child language environment analysis using deep neural networks
abstract
Large-scale monitoring of the child language environment through measuring the amount of speech directed to the child by other children and adults during a vocal communication is an important task. Using the audio extracted from a recording unit worn by a child within a childcare center, at each point in time our proposed diarization system can determine the content of the child's language environment, by categorizing the audio content into one of the four major categories, namely (1) speech initiated by the child wearing the recording unit, speech originated by other (2) children or (3) adults and directed at the primary child, and (4) non-speech contents. In this study, we exploit complex Hidden Markov Models (HMMs) with multiple states to model the temporal dependencies between different sources of acoustic variability and estimate the HMM state output probabilities using deep neural networks as a discriminative modeling approach. The proposed system is robust against common diarization errors caused by rapid turn takings, between class similarities, and background noise without the need to prior clustering techniques. The experimental results confirm that this approach outperforms the state-of-the-art Gaussian mixture model based diarization without the need for bottom-up clustering and leads to 22.24% relative error reduction.
Maryam Najafian, John H. L. Hansen
SLT2
2016 Unsupervised k-means clustering based out-of-set candidate selection for robust open-set language recognition
abstract
Research in open-set language identification (LID) generally focuses more on accurate in-set language modeling versus improved out-of-set (OOS) language rejection. The main reason for this is the increased cost/resources in collecting sufficient OOS data, versus the in-set languages of interest. Therefore, unknown or OOS language rejection is a challenge. To address this through efficient data collection, we propose a flexible OOS candidate selection method for universal OOS language coverage. Since state-of-the-art i-vector system followed by generative Gaussian back-end achieves effective performance for LID, the selected K candidates are expected to be general enough to represent the entire OOS language space. Therefore, an unsupervised k-means clustering approach is proposed for effective OOS candidate selection. This method is evaluated on a dataset derived from a large-scale corpus (LRE-09) which contains 40 languages. With the proposed selection method, the total OOS training diversity can be reduced by 89% and still achieve better performance on both OOS rejection and overall classification. The proposed method also shows clear benefits for greater data enhancement. Therefore, the proposed solution achieves sustained performance with the advantage of employing a minimum number of OOS language candidates efficiently.
Qian Zhang 0019, John H. L. Hansen
SLT2
2016 Unsupervised accent classification for deep data fusion of accent and language information
John H. L. Hansen, Gang Liu 0001
Speech Commun.1
2016 Effective word count estimation for long duration daily naturalistic audio recordings
Ali Ziaei, Abhijeet Sangwan, John H. L. Hansen
Speech Commun.3
2016 Microphone Array Processing Strategies for Distant-Based Automatic Speech Recognition
abstract
Robust distant speech recognition (DSR) is necessary in many speech technology applications using multiple microphones but has received only limited treatment in the literature. In this paper, we work on communicating with vehicle voice-controlled system which is one of the applications of DSR. Two approaches for DSR are i) signal-level combination using beamforming followed by automatic speech recognition (ACR), and ii) word hypothesis-level combination using several speech recognition engines followed by confusion network combination or followed by recognizer output voting error reduction (ROVER). In addition to these approaches, it is possible to examine training-level combination by training the recognizer on audio signals from multiple channels (microphones). In this paper, the authors investigate how these methods can be leveraged for in-vehicle ACR using the CU-Move corpus. The authors propose various combinations of these three methods to find an optimum structure for in-vehicle ACR. The authors also investigate the effect of speaker adaptation (SA). The author's experience shows that applying SA on individual channels and merging the results with ROVER reduces the negative effects of SA reported by others in the field, and illustrates the overall improvement obtained with front-end enhancement techniques in DSR.
Soudeh A. Khoubrouy, John H. L. Hansen
IEEE Signal Process. Lett.2
2016 Generative Modeling of Pseudo-Whisper for Robust Whispered Speech Recognition
abstract
Whisper is a common means of communication used to avoid disturbing individuals or to exchange private information. As a vocal style, whisper would be an ideal candidate for human-handheld/computer interactions in open-office or public area scenarios. Unfortunately, current speech technology is predominantly focused on modal (neutral) speech and completely breaks down when exposed to whisper. One of the major barriers for successful whisper recognition engines is the lack of available large transcribed whispered speech corpora. This study introduces two strategies that require only a small amount of untranscribed whisper samples to produce excessive amounts of whisper-like (pseudo-whisper) utterances from easily accessible modal speech recordings. Once generated, the pseudo-whisper samples are used to adapt modal acoustic models of a speech recognizer toward whisper. The first strategy is based on Vector Taylor Series (VTS) where a whisper “background” model is first trained to capture a rough estimate of global whisper characteristics from a small amount of actual whisper data. Next, that background model is utilized in the VTS to establish specific broad phone classes' (unvoiced/voiced phones) transformations from each input modal utterance to its pseudo-whispered version. The second strategy generates pseudo-whisper samples by means of denoising autoencoders (DAE). Two generative models are investigated-one produces pseudo-whisper cepstral features on a frame-by-frame basis, while the second generates pseudo-whisper statistics for whole phone segments. It is shown that word error rates of a TIMIT-trained speech recognizer are considerably reduced for a whisper recognition task with a constrained lexicon after adapting the acoustic model toward the VTS or DAE pseudo-whisper samples, compared to model adaptation on an available small whisper set.
Shabnam Ghaffarzadegan, Hynek Boril, John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.3
2016 Score-Aging Calibration for Speaker Verification
abstract
The gradual changes that occur in the human voice due to aging create challenges for speaker verification. This study presents an approach to calibrating the output scores of a speaker verification system using the time interval between comparison samples as additional information. Several functions are proposed for the incorporation of this time information, which is viewed as aging information, in a conventional linear score calibration transformation. Experiments are presented on data with shortterm aging intervals ranging between 2 months and 3 years, and long-term aging intervals of up to 30 years. The aging calibration proposal is shown to offset the decreased discrimination and calibration performance for both shortand long-term intervals, and to extrapolate well to unseen aging intervals. Relative reductions in Cuℓℓr(cost of log-likelihood ratio) of 1-4% and 10-43% are obtained at shortand long-term intervals, respectively. Assuming that a system has knowledge of the time interval between samples under comparison, this approach represents a straightforward means of compensating for the detrimental impact of aging on speaker verification performance.
Finnian Kelly, John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.2
2016 A Generalized Nonnegative Tensor Factorization Approach for Distant Speech Recognition With Distributed Microphones
abstract
Automatic speech recognition (ASR) using distant (far-field) microphones is a challenging task, in which room reverberation is one of the primary causes of performance degradation. This study proposes a multichannel spectral enhancement method for reverberation-robust ASR using distributed microphones. The proposed method uses the techniques of nonnegative tensor factorization in order to identify the clean speech component from a set of observed reverberant spectrograms from the different channels. The general family of alpha-beta divergences is used for the tensor decomposition task which provides increased flexibility for the algorithm and is shown to provide improvements in highly reverberant scenarios. Unlike many conventional array processing solutions, the proposed method does not require closely-spaced microphones and is independent of source and microphone locations. The algorithm can automatically adapt to unbalanced direct-to-reverberation ratios among different channels, which is useful in blind scenarios in which no information is available about source-to-microphone distances. For a medium vocabulary distant ASR task based on TIMIT utterances, and using clean-trained deep neural network acoustic models, absolute WER improvements of +17.2%, +20.7%, and +23.2% are achieved in single-channel, two-channel, and four-channel scenarios.
Seyedmahdad Mirsamadi, John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 An i-Vector PLDA based gender identification approach for severely distorted and multilingual DARPA RATS data
abstract
This study proposes an i-Vector based approach to gender identification. Gender-labeled utterances from the Fisher English (FE) corpus are used to formulate an i-Vector extraction framework, and a Probabilistic Linear Discriminant Analysis (PLDA) back-end is employed to compute the scores for gender identification. A novel duration mismatch compensation strategy is also presented that offers very little degradation in identification accuracy even with a large reduction in the duration of the test-segment. The proposed method is shown to consistently outperform a GMM-UBM based gender-identification scheme on several test-sets created from a held-out portion of the FE corpus, and is able to achieve an identification accuracy of up to 97.63%. On the severely distorted and multilingual DARPA-RATS (Robust Automatic Transcription of Speech) corpora, the proposed approach achieves an identification accuracy of 76.48% using only the FE data in training. Next, a novel unsupervised domain adaptation strategy is also presented that utilizes only unlabeled RATS data to adapt the out-of-domain PLDA parameters derived from the FE training data. The strategy is able to offer a 6.8% relative improvement in identification accuracy, and a 14.75% relative reduction in Equal Error Rate (EER) compared to using the out-of-domain PLDA model on the RATS test-utterances. These improvements are significant since: 1) RATS test-utterances are severely distorted, 2) No labeled data of any kind is used for 4 of the 5 languages present in the test-utterances.
Shivesh Ranjan, Gang Liu 0001, John H. L. Hansen
ASRU3
2015 Image-guided customization of frequency-place mapping in cochlear implants
abstract
Multi-channel cochlear implants (CI) leverage frequency based cochlear tonotopic mapping to map acoustic information to the cochlear place of stimulation which is primarily determined by electrode locations. Despite the fact that electrode locations within the cochlea are unique to each patient, the acoustic frequencies assigned to the electrodes by the CI processor are determined generically, resulting in a mismatch between intended and actual pitch perception. This is known to be a limiting factor for hearing outcomes with CIs. In this study, we propose a novel, image-guided CI processor programming strategy to select more optimal, patient-customized frequency assignments. The performance of the proposed strategy was evaluated using vocoder-based simulations with ten normal hearing listeners. In our simulations, our strategy results in significantly better speech recognition scores than the standard clinical strategy.
Hussnain Ali, Jack H. Noble, René H. Gifford, Robert F. Labadie, Benoit M. Dawant, John H. L. Hansen, Emily Tobey
ICASSP6
2015 Generative modeling of pseudo-target domain adaptation samples for whispered speech recognition
abstract
The lack of available large corpora of transcribed whispered speech is one of the major roadblocks for development of successful whisper recognition engines. Our recent study has introduced a Vector Taylor Series (VTS) approach to pseudo-whisper sample generation which requires availability of only a small number of real whispered utterances to produce large amounts of whisper-like samples from easily accessible transcribed neutral recordings. The pseudo-whisper samples were found particularly effective in adapting a neutral-trained recognizer to whisper. Our current study explores the use of denoising autoencoders (DAE) for pseudo-whisper sample generation. Two types of generative models are investigated - one which produces pseudo-whispered cepstral vectors on a frame basis and another which generates pseudo-whisper statistics of whole phone segments. It is shown that the DAE approach considerably reduces word error rates of the baseline system as well as the system adapted on real whisper samples. The DAE approach provides competitive results to the VTS-based method while cutting its computational overhead nearly in half.
Shabnam Ghaffarzadegan, Hynek Boril, John H. L. Hansen
ICASSP3
2015 Analysis of speech and language communication for cochlear implant users in noisy lombard conditions
abstract
Acoustic/linguistic modification of speech production with respect to auditory feedback is an important research domain for robust human-to-human and human-to-machine communication. For instance, in the presence of environmental noise, a speaker experiences the well-known phenomenon termed as Lombard effect. Lombard effect has been well studied for normal hearing listeners as well as for automatic speech/speaker recognition systems. However, limited effort has been employed to study if the speech production of cochlear implant (CI) users is influenced by the auditory feedback. The purpose of this study is to analyze the speech production and natural language model of CI users with respect to environmental changes. A mobile personal audio recording from continuous single-session audio streams collected over an individual's daily life was used for our study. The findings from this study will provide fundamental knowledge on the characteristics of speech production under Lombard effect in CI users. These specific variations in speech production can be leveraged in new algorithm development and further applications in speech systems to benefit cochlear implant users.?
Hussnain Ali, Ali Ziaei, John H. L. Hansen
ICASSP4
2015 Robust unsupervised detection of human screams in noisy acoustic environments
abstract
This study is focused on an unsupervised approach for detection of human scream vocalizations from continuous recordings in noisy acoustic environments. The proposed detection solution is based on compound segmentation, which employs weighted mean distance, T2-statistics and Bayesian Information Criteria for detection of screams. This solution also employs an unsupervised threshold optimized Combo-SAD for removal of non-vocal noisy segments in the preliminary stage. A total of five noisy environments were simulated for noise levels ranging from -20dB to +20dB for five different noisy environments. Performance of proposed system was compared using two alternative acoustic front-end features (i) Mel-frequency cepstral coefficients (MFCC) and (ii) perceptual minimum variance distortionless response (PMVDR). Evaluation results show that the new scream detection solution works well for clean, +20, +10 dB SNR levels, with performance declining as SNR decreases to -20dB across a number of the noise sources considered.
Mahesh Kumar Nandwana, Ali Ziaei, John H. L. Hansen
ICASSP3
2015 Weighted training for speech under Lombard Effect for speaker recognition
abstract
The presence of Lombard Effect in speech is proven to have severe effects on the performance of speech systems, especially speaker recognition. Varying kinds of Lombard speech are produced by speakers under influence of varying noise types [1]. This study proposes a high-accuracy classifier using deep neural networks for detecting various kinds of Lombard speech against neutral speech, independent of the noise levels causing the Lombard Effect. Lombard Effect detection accuracies as high as 95.7% are achieved using this novel model. The deep neural network based classification is further exploited by validation based weighted training of robust i-Vector based speaker identification systems. The proposed weighted training achieves a relative EER improvement of 28.4% over an i-Vector baseline system, confirming the effectiveness of deep neural networks in modeling Lombard Effect.
Muhammad Muneeb Saleem, Gang Liu 0001, John H. L. Hansen
ICASSP3
2015 Robust overlapped speech detection and its application in word-count estimation for Prof-Life-Log data
abstract
The ability to estimate the number of words spoken by an individual over a certain period of time is valuable in second language acquisition, healthcare, and assessing language development. However, establishing a robust automatic framework to achieve high accuracy is non-trivial in realistic/naturalistic scenarios due to various factors such as different styles of conversation or types of noise that appear in audio recordings, especially in multi-party conversations. In this study, we propose a noise robust overlapped speech detection algorithm to estimate the likelihood of overlapping speech in a given audio file in the presence of environment noise. This information is embedded into a word-count estimator, which uses a linear minimum mean square estimator (LMMSE) to predict the number of words from the syllable rate. Syllables are detected using a modified version of the mrate algorithm. The proposed word-count estimator is tested on long duration files from the Prof-Life-Log corpus. Data is recorded using a LENA recording device, worn by a primary speaker in various environments and under different noise conditions. The overlap detection system significantly outperforms baseline performance in noisy conditions. Furthermore, applying overlap detection results to word-count estimation achieves 35% relative improvement over our previous efforts, which included speech enhancement using spectral subtraction and silence removal.
Navid Shokouhi, Ali Ziaei, Abhijeet Sangwan, John H. L. Hansen
ICASSP4
2015 Leveraging automatic speech recognition in cochlear implants for improved speech intelligibility under reverberation
abstract
Despite recent advancements in digital signal processing technology for cochlear implant (CI) devices, there still remains a significant gap between speech identification performance of CI users in reverberation compared to that in anechoic quiet conditions. Alternatively, automatic speech recognition (ASR) systems have seen significant improvements in recent years resulting in robust speech recognition in a variety of adverse environments, including reverberation. In this study, we exploit advancements seen in ASR technology for alternative formulated solutions to benefit CI users. Specifically, an ASR system is developed using multicondition training on speech data with different reverberation characteristics (e.g., T60values), resulting in low word error rates (WER) in reverberant conditions. A speech synthesizer is then utilized to generate speech waveforms from the output of the ASR system, from which the synthesized speech is presented to CI listeners. The effectiveness of this hybrid recognition-synthesis CI strategy is evaluated under moderate to highly reverberant conditions (i.e., T60= 0.3, 0.6, 0.8, and 1.0s) using speech material extracted from the TIMIT corpus. Experimental results confirm the effectiveness of multi-condition training on performance of the ASR system in reverberation, which consequently results in substantial speech intelligibility gains for CI users in reverberant environments.
Oldooz Hazrati Yadkoori, Shabnam Ghaffarzadegan, John H. L. Hansen
ICASSP3
2015 Prof-Life-Log: Analysis and classification of activities in daily audio streams
abstract
A new method to analyze and classify daily activities in personal audio recordings (PARs) is presented. The method employs speech activity detection (SAD) and speaker diarization systems to provide high level semantic segmentation of the audio file. Subsequently, a number of audio, speech and lexical features are computed in order to characterize events in daily audio streams. The features are selected to capture the statistical properties of conversations, topics and turn-taking behavior, which creates a classification space that allows us to capture the differences in interactions. The proposed system is evaluated on 9 days of data from Prof-Life-Log corpus, which contains naturalistic long duration audio recordings (each file is collected continuously and lasts between 8-to-16 hours). Our experimental results show that the proposed system achieves good classification accuracy on a difficult real-world dataset.
Ali Ziaei, Abhijeet Sangwan, Lakshmish Kaushik, John H. L. Hansen
ICASSP4
2015 Laughter and filler detection in naturalistic audio
abstract
Laughter and fillers are common phenomenon in speech, and play an important role in communication. In this study, we present Deep Neural Network (DNN) and Convolutional Neural Network (CNN) based systems to classify non-verbal cues (laughter and fillers) from verbal speech in naturalistic audio. We propose improvements over a deep learning system proposed in [1]. Particularly, we propose a simple method to combine spectral features with pitch information to capture prosodic and spectral cues for filler/laughter. Additionally, we propose using a wider time context for feature extraction so that the time evolution of the spectral and prosodic structure can also be exploited for classification. Furthermore, we propose to use CNN for classification. The new method is evaluated on conversational telephony speech (CTS, drawn from Switchboard and Fisher) data and UT-Opinion corpus. Our results shows that the new system improves the AUC (area under the curve) metric by 8.15% and 11.9% absolute for laughters, and 4.85% and 6.01% absolute for fillers, over the baseline system, for CTS and UT-Opinion data, respectively. Finally, we analyze the results to explain the difference in performance between traditional CTS data and naturalistic audio (UT-Opinion), and identify challenges that need to be addressed to make systems perform better for practical data.
Lakshmish Kaushik, Abhijeet Sangwan, John H. L. Hansen
INTERSPEECH3
2015 Automatic audio sentiment extraction using keyword spotting
Lakshmish Kaushik, Abhijeet Sangwan, John H. L. Hansen
INTERSPEECH3
2015 Evaluation and calibration of short-term aging effects in speaker verification
Finnian Kelly, John H. L. Hansen
INTERSPEECH2
2015 A study on deep neural network acoustic model adaptation for robust far-field speech recognition
abstract
Even though deep neural network acoustic models provide an increased degree of robustness in automatic speech recognition, there is still a large performance drop in the task of far-field speech recognition in reverberant and noisy environments. In this study, we explore DNN adaptation techniques to achieve improved robustness to environmental mismatch for far-field speech recognition. In contrast to many recent studies investigating the role of feature processing in DNN-HMM systems, we focus on adaptation of a clean-trained DNN model to speech data captured by a distant-talking microphone in a target environment with substantial reverberation and noise. We show that significant performance gains can be obtained by discriminatively estimating a set of adaptation parameters to compensate the mismatch between a clean-trained model and a small set of noisy and reverberant adaptation data. Using various adaptation strategies, relative word error rate improvements of up to 16% could be obtained on the single-channel task of the recent Aspire challenge.
Seyedmahdad Mirsamadi, John H. L. Hansen
INTERSPEECH2
2015 Anti-spoofing system: an investigation of measures to detect synthetic and human speech
Abhinav Misra, Shivesh Ranjan, John H. L. Hansen
INTERSPEECH4
2015 A new front-end for classification of non-speech sounds: a study on human whistle
abstract
Speech/non-speech sound classification is an important problem in audio diarization, audio document retrieval and advanced human interfaces. The focus of this study is on the development of spectral and temporal acoustic features for speech/non-speech sound classification based on production differences in speech versus whistle. Seven time- and frequency-domain based features are investigated. Performance of the proposed feature set for the task of speech/whistle classification is evaluated at frame level. This evaluation utilizes support vector machine (SVM) models and Gaussian mixture models (GMM) for back-end classifiers. At the frame-level, the proposed front-end fusion gives an absolute performance gain of +15.0% and +3.1% over MFCC with SVM and GMM based classifiers, respectively. This research will benefit the development of intelligent speech interfaces for identification, recognition, and speech coding, as a preprocessing step for real world audio streams.
Mahesh Kumar Nandwana, Hynek Boril, John H. L. Hansen
INTERSPEECH3
2015 Probabilistic linear discriminant analysis for robust speaker identification in co-channel speech
Navid Shokouhi, John H. L. Hansen
INTERSPEECH2
2015 An unsupervised visual-only voice activity detection approach using temporal orofacial features
abstract
Detecting the presence or absence of speech is an important step toward building robust speech-based interfaces. While previous studies have made progress on voice activity detection (VAD), the performance of these systems significantly degrades when subjects employ challenging speech modes that deviate from normal acoustic patterns (e.g., whisper speech), or in noisy/adverse conditions. An appealing approach under these conditions is visual voice activity detection (VVAD), which detects speech using features characterizing the orofacial activity. This study proposes an unsupervised approach that relies only on visual features, and, therefore, is insensitive to vocal style or time-varying background noise. This study proposes an unsupervised approach that relies on visual features. We estimate optical flow variance and geometrical features around lips, extracting the short-time zero crossing rates, short-time variances, and delta features over a small temporal window. These variables are fused using principal component analysis (PCA) to obtain a “combo” feature, which displays a bimodal distributions (speech versus silence). A threshold is automatically determine using the expectation-maximization (EM) algorithm. The approach can be easily transformed into a supervised VVAD, if needed. We evaluate the system in neutral and whisper speech. While speech based VADs generally fail to detect speech activity in whisper speech, given its important acoustic differences, the proposed VVAD achieves near 80% accuracy in both neutral and whisper speech, highlighting the benefits of the system. Index Terms: Visual voice activity detection, whisper speech
Fei Tao 0003, John H. L. Hansen, Carlos Busso
INTERSPEECH2
2015 Frequency offset correction in single sideband (SSB) speech by deep neural network for speaker verification
abstract
Communication system mismatch represents a major influence for loss in speaker recognition performance. This paper considers a type of nonlinear communication system mismatch- modulation/demodulation (Mod/DeMod) carrier drift in single sideband (SSB) speech signals. We focus on the problem of estimating frequency offset in SSB speech in order to improve speaker verification performance of the drifted speech. Based on a two-step framework from previous work, we propose using a multi-layered neural network architecture, stacked denoising autoencoder (SDA), to determine the unique interval of the offset value in the first step. Experimental results demonstrate that the SDA based system can produce up to a +16.1% relative improvement in frequency offset estimation accuracy. A speaker verification evaluation shows a +65.9% relative improvement in EER when SSB speech signal is compensated with the frequency offset value estimated by the proposed method. Index Terms: frequency offset, single sideband, speaker verification, denoising autoencoder
Hua Xing, Gang Liu 0001, John H. L. Hansen
INTERSPEECH3
2015 Robust i-vector extraction for neural network adaptation in noisy environment
Chengzhu Yu, Atsunori Ogawa, Marc Delcroix, Takuya Yoshioka, Tomohiro Nakatani, John H. L. Hansen
INTERSPEECH6
2015 I-vector based physical task stress detection with different fusion strategies
abstract
It is common for subjects to produce speech while performing a physical task where speech technology may be used. Variabilities are introduced to speech since physical task can influence human speech production. These variabilities degrade the performance of most speech systems. It is vital to detect speech under physical stress variabilities for subsequent algorithm processsing. This study presents a method for detecting physical task stress from speech. Inspired by the fact that i-vectors can generally model total factors from speech, a state-of-the-art ivector framework is investigated with MFCCs and our previously formulated TEO-CB-Auto-Env features for neutral/physical task stress detection. Since MFCCs are derived from a linear speech production model and TEO-CB-Auto-Env features employ a nonlinear operator, these two features are believed to have complementary effects on physical task stress detection. Two alternative fusion strategies (feature-level and score-level fusion) are investigated to validate this hypothesis. Experiments over the UT-Scope Physical Corpus demonstrate that a relative accuracy gain of 2.68% is obtained when fusing different feature based i-vectors. An additional relative performance boost with of 6.52% in accuracy is achieved using score level fusion.
Gang Liu 0001, Chengzhu Yu, John H. L. Hansen
INTERSPEECH4
2015 In-vehicle speech recognition and tutorial keywords spotting for novice drivers' performance evaluation
abstract
Novice young drivers are more frequently involved in traffic accidents, and studies have shown that effective supervised driver training is the key in reducing young drivers' risks. Using our previously developed Mobile-UTDrive in-vehicle data acquisition platform, two 16-age novice drivers participated in naturalistic drive training data collection. This paper focuses on analysis of novice driver training signals from an audio processing perspective. Specifically, analysis of supervised driver instruction audio and resulting CAN-Bus maneuver operation is performed. Following a procedure which consists of noise suppression, speech recognition and keyword spotting, five tutorial keywords - Brake, Gas, Left, Right and Stop - are spotted at an overall accuracy rate of 40% versus all spontaneous continuous speech. The time stamps of these keywords are then used as indications of driving maneuvers. As examples of driving performance evaluation, the case of making Left-Turn maneuvers for the two novice drivers are assessed and compared, and the increase of driving skills over experiences are analyzed.
Xian Shi, Amardeep Sathyanarayana, Navid Shokouhi, John H. L. Hansen
Intelligent Vehicles Symposium5
2015 Advanced parallel combined Gaussian mixture model based feature compensation integrated with iterative channel estimation
Wooil Kim, John H. L. Hansen
Speech Commun.2
2015 Mean Hilbert envelope coefficients (MHEC) for robust speaker and language identification
Seyed Omid Sadjadi, John H. L. Hansen
Speech Commun.2
2015 An advanced entropy-based feature with a frame-level vocal effort likelihood space modeling for distant whisper-island detection
John H. L. Hansen
Speech Commun.2
2015 A Hybrid Coherence Model for Noise Reduction in Reverberant Environments
abstract
In this letter, we propose a novel dual-microphone technique for enhancement of speech degraded by background noise and reverberation. Our algorithm is based on a prediction of the coherence function between the noisy input signals, considering both direct and reverberant speech and noise components received by the sensors, and therefore, is capable of dealing with both coherent and diffuse noise. After predicting the coherence function, the signal to noise ratio (SNR) can be estimated by solving a quadratic equation obtained from the real and imaginary parts of the function. Objective evaluation in a room with reverberation time T60= 465 ms, demonstrated noticeable improvements in SNR and quality of the outputs processed with the proposed algorithm over the baseline (front microphone), as well as a recently proposed coherence-based noise reduction algorithm.
Nima Yousefian, John H. L. Hansen, Philipos C. Loizou
IEEE Signal Process. Lett.2
2015 Howling Detection in Hearing Aids Based on Generalized Teager-Kaiser Operator
abstract
With the ongoing miniaturization in the hearing aid industry, acoustical coupling between the loudspeaker and the microphone(s) of the hearing aid causes a major problem to users. Howling is one of the most severe and annoying consequences of this acoustical coupling. This study presents a howling detection method using the Generalized Teager-Kaiser Operator (GTKO). Since the GTKO is both time and frequency sensitive, its resolution parameter must be assigned properly to ensure satisfactory performance of this operator in the frequency range of the input signal for the hearing aid. In order to cover the entire band of the input signal with appropriate resolution parameters, the input signal is decomposed into a filterbank (i.e., uniform and nonuniform filterbanks). GTKO is applied to the output of each band to detect the howling, and the resolution parameter of the GTKO block is selected depending on the central frequency of that particular band. Experimental results compare the performance of each proposed method with two known howling detection approaches, Peak-to-harmonic power ratio (PHPR) approach and a multiple-feature approach. The proposed method has high detection probability and short detection time. It is also shown that considering a hybrid algorithm which includes the PHPR approach with each of the proposed methods (i.e. combination of GTKO blocks with different types of filterbanks) results in lower false alarm probability.
Soudeh A. Khoubrouy, Issa M. S. Panahi, John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.3
2014 Improving channel selection of sound coding algorithms in cochlear implants
abstract
Spectral maxima sound coding algorithms, for example n-of-m strategies, used in commercial cochlear implant devices rely on selecting channels with the highest energy in each frequency band. This technique works well in quiet, but is inherently problematic in noisy conditions when noise dominates the target, and noise-dominant channels are mistakenly selected for stimulation. A new channel selection criterion is proposed to addresses this shortcoming which adaptively assigns weights to each time-frequency unit based on the formant location of speech and instantaneous signal to noise ratio. The performance of the proposed technique is evaluated acutely with three cochlear implant users in different noise scenarios. Results indicate that the proposed technique improves speech intelligibility and perception quality, particularly at low signal-to-noise ratio. Significance of the proposed technique lies in its ability to be integrated with the existing sound coding framework employed within commercial cochlear implant processors, making it easier to adapt for resource-limited and time critical devices.
Hussnain Ali, John H. L. Hansen, Emily Tobey
ICASSP3
2014 UT-Vocal Effort II: Analysis and constrained-lexicon recognition of whispered speech
abstract
This study focuses on acoustic variations in speech introduced by whispering, and proposes several strategies to improve robustness of automatic speech recognition of whispered speech with neutral-trained acoustic models. In the analysis part, differences in neutral and whispered speech captured in the UT-Vocal Effort II corpus are studied in terms of energy, spectral slope, and formant center frequency and bandwidth distributions in silence, voiced, and unvoiced speech signal segments. In the part dedicated to speech recognition, several strategies involving front-end filter bank redistribution, cepstral dimensionality reduction, and lexicon expansion for alternative pronunciations are proposed. The proposed neutral-trained system employing redistributed filter bank and reduced features provides a 7.7 % absolute WER reduction over the baseline system trained on neutral speech, and a 1.3 % reduction over a baseline system with whisper-adapted acoustic models.
Shabnam Ghaffarzadegan, Hynek Boril, John H. L. Hansen
ICASSP3
2014 Frequency offset correction in single sideband speech for speaker verification
abstract
Communication system mismatch represents a major influence for loss in speaker recognition performance. While microphone and handset differences have been considered in the NIST SRE, nonlinear communication system differences, such as modulation/demodulation (Mod/DeMod) carrier drift, have yet to be considered. In this study, an algorithm for estimating and correcting Mod/DeMod frequency offsets distortion in signal sideband modulation (SSB) speech is formulated based on two processing steps. In the first step, the offset of speech can be roughly scaled to a small frequency interval, which eliminates the ambiguity caused by periodicity of the spectrum. The second step performs fine-tuning within the pre-determined interval. For the first time, a statistical framework is developed for unique interval detection, where an innovative acoustic feature is proposed to represent different offsets and state-of-the-art techniques, the total variety method and PLDA, are applied. Speaker recognition experiments on SSB speech obtained from DAPPA RATS corpus show that a significant performance improvement (up to 50% relative improvement in EER) for speaker verification in SSB speech can be obtained by the proposed estimation and compensation method.
Hua Xing, Philipos C. Loizou, John H. L. Hansen
ICASSP3
2014 Robust and efficient environment detection for adaptive speech enhancement in cochlear implants
abstract
Cochlear implant (CI) recipients require alternative signal processing for speech enhancement, since the quantities needed for intelligibility and quality improvement differ significantly when direct stimulation of the basilar membrane is employed for CIs. Here, a robust feature vector is proposed for environment classification in CI devices. The feature vector is directly computed from the output of the advanced combination encoder (ACE), which is a sound coding strategy commonly used in CIs. Performance of the proposed feature vector is evaluated in the context of environment classification tasks under anechoic quiet, noisy, reverberant, and noisy reverberant conditions. Speech material taken from the IEEE corpus are used to simulate different environmental acoustic conditions with: 1) three measured room impulse responses (RIR) with distinct reverberation times (T60) for generating reverberant environments, and 2) car, train, white Gaussian, multi-talker babble, and speech-shaped noise (SSN) samples for creating noisy conditions at 4 different signal-to-noise ratio (SNR) levels. We investigate 3 different classifiers for environment detection, namely Gaussian mixture models (GMM), support vector machines (SVM), and neural networks (NN). Experimental results illustrate the effectiveness of the proposed features for environment classification.
Oldooz Hazrati Yadkoori, Seyed Omid Sadjadi, John H. L. Hansen
ICASSP3
2014 Uncertainty propagation in front end factor analysis for noise robust speaker recognition
abstract
In this study, we explore the propagation of uncertainty in the state-of-the-art speaker recognition system. Specifically, we incorporate the uncertainty associated with observation features into the i-Vector extraction framework. To prove the concept, both the oracle and practically estimated uncertainty are used for evaluation. The oracle uncertainty is calculated assuming the knowledge of clean speech features, while the estimated uncertainties are obtained using SPLICE and joint-GMM based methods. We evaluate the proposed framework on both YOHO and NIST 2010 Speaker Recognition Evaluation (SRE) corpora by artificially introducing noise at different SNRs. In the speaker verification experiments, we confirmed that the proposed uncertainty based i-Vector extraction framework shows significant robustness against noise.
Chengzhu Yu, Gang Liu 0001, Seongjun Hahm, John H. L. Hansen
ICASSP4
2014 Model and feature based compensation for whispered speech recognition
abstract
This study proposes model and feature based strategies for au-tomatic whispered speech recognition. Our goal is to compensate for the mismatch between neutral-trained recognizer models and parameters of whispered speech. We propose a pseudo-whisper generation from neutral speech samples for efficient acoustic model adaptation. The scheme is based on the popular Vector Tay-lor Series (VTS) algorithm. In the first step, a ‘background ’ model capturing a rough estimate of the target whispered speech charac-teristics from a small amount of whispered data is trained. Second, the target background model is utilized in the VTS strategy to es-tablish broad phone classes (consonants and vowels) transforma-tions for individual neutral utterances and transform them towards whisper. Finally, these pseudo-whisper samples are used to adapt neutral recognizer models towards whisper. This approach is eval-uated together with Vocal Tract Length Normalization (VTLN) and Shift frequency transforms and show to greatly benefit recog-nition performance compared to a traditional whisper-adaptation approach. The absolute WER on the closed speakers whisper sce-nario has been reduced from 17.3 % to 8.4 % and the open speakers scenario from 27.7 % to 17.5 %. Index Terms: whispered speech recognition, Vector Taylor Series, vocal length normalization
Shabnam Ghaffarzadegan, Hynek Boril, John H. L. Hansen
INTERSPEECH3
2014 Multichannel speech dereverberation based on convolutive nonnegative tensor factorization for ASR applications
abstract
Room reverberation is a primary cause of failure in distant speech recognition (DSR) systems. In this study, we present a multichannel spectrum enhancement method for reverberant speech recognition, which is an extension of a single-channel dereverberation algorithm based on convolutive nonnegative matrix factorization (NMF). The generalization to a multichannel scenario is shown to be a special case of convolutive nonnegative tensor factorization (NTF). The presented algorithm integrates information from across different channels in the magnitude short time Fourier transform (STFT) domain. By doing so, it eliminates any limitations on the array geometry or a need for information concerning the source location, making the algorithm particularly suitable for distributed microphone arrays. Experiments are performed on speech data using actual room impulse responses from AIR database. Relative WER improvements using a clean-trained ASR system vary from +7.1% to +30.1% based on the number of channels and the source to microphone distances (1 to 3 meters).
Seyedmahdad Mirsamadi, John H. L. Hansen
INTERSPEECH2
2014 Analysis and identification of human scream: implications for speaker recognition
Mahesh Kumar Nandwana, John H. L. Hansen
INTERSPEECH2
2014 Co-channel speech detection via spectral analysis of frequency modulated sub-bands
Navid Shokouhi, Seyed Omid Sadjadi, John H. L. Hansen
INTERSPEECH3
2014 Investigation of the relative perceptual importance of temporal envelope and temporal fine structure between tonal and non-tonal languages
Dongmei Wang, James M. Kates, John H. L. Hansen
INTERSPEECH3
2014 Noisy speech enhancement based on long term harmonic model to improve speech intelligibility for hearing impaired listeners
abstract
This study proposes a speech enhancement algorithm to improve speech intelligibility for hearing impaired listeners in adverse conditions. The proposed algorithm is based on a long term harmonic model, where the harmonics of target speech are more distinguished from noise spectrum interference. Our method consists of two stages: i) Prominent pitch estimation based on long term harmonic feature analysis and neural network classification. ii) Target speech spectrum estimation with pitch information based on long term noise spectrum extraction. The listening experiment with EAS vocoder speech shows that our algorithm is substantially beneficial for cochlear implant recipients to perceive speech in noisy environment in terms of word recognition rate.
Dongmei Wang, Philipos C. Loizou, John H. L. Hansen
INTERSPEECH3
2014 F0 estimation in noisy speech based on long-term harmonic feature analysis combined with neural network classification
Dongmei Wang, Philipos C. Loizou, John H. L. Hansen
INTERSPEECH3
2014 'houston, we have a solution': a case study of the analysis of astronaut speech during NASA apollo 11 for long-term speaker modeling
abstract
Speech and language processing technology has the potential of playing an important role in future deep space missions. To be able to replicate the success of speech technologies from ground to space, it is important to understand how astronaut’s speech production mechanism changes when they are in space. In this study, we investigate the variations of astronaut’s voice charac-teristic during NASA Apollo 11 mission. While the focus is constrained to analysis of the three astronauts voices who par-ticipated in the Apollo 11 mission, it is the first step towards our long term objective of automating large components of space missions with speech and language technology. The result of this study is also significant from an historical point of view as it provides a new perspective of understanding the key mo-ment of human history- landing a man on the moon, as well as employed for future advancement in speech and language tech-nology in “non-neutral”conditions.
Chengzhu Yu, John H. L. Hansen, Douglas W. Oard
INTERSPEECH2
2014 Acoustic feature transformation using UBM-based LDA for speaker recognition
abstract
In state-of-the-art speaker recognition system, universal background model (UBM) plays a role of acoustic space division. Each Gaussian mixture of trained UBM represents one distinct acoustic region. The posterior probabilities of features belonging to each region are further used as core components of Baum-Welch statistics. Therefore, the quality of estimated Baum-Welch statistics depends highly on how acoustic regions are separable with each other. In this paper, we propose to transform the front end acoustical features into a space where the separability of mixtures of trained UBM can be optimized. To achieve this, an UBM was first trained from the acoustical features and a transformation matrix is estimated using linear discriminant analysis (LDA) by treating each mixture of trained UBM as independent class. Therefore, the proposed method named as UBM-based LDA (uLDA) does not require any speaker labels or other supervised information. The obtained transformation matrix is then applied to acoustic features for i-Vector extraction. Experimental results on the male part of core conditions of NIST SRE 2010 dataset confirmed the improved performance using proposed method. Index Terms: Speaker recognition, i-Vector, universal background model (UBM), Baum-Welch statistic, LDA.
Chengzhu Yu, Gang Liu 0001, John H. L. Hansen
INTERSPEECH3
2014 Speech activity detection for NASA apollo space missions: challenges and solutions
abstract
Speech Activity Detection(SAD) is a well researched problem for communication, command and control applications, where audio segments are short duration and solution proposed for noisy as well as clean environments. In this study, we inves-tigate the SAD problem using NASA’s Apollo space mission data [1]. Unlike traditional speech corpora, the audio recordings in Apollo are extensive from a longitudinal perspective (i.e., 6-12 days each). From SAD perspective, the data offers many challenges: (i) noise distortion with variable SNR, (ii) chan-nel distortion, and (iii) extended periods of non-speech activity. Here, we use the recently proposed Combo-SAD, which has performed remarkably well in DARPA RATS evaluations, as our baseline system [2]. Our analysis reveals that the Combo-SAD performs well when speech-pause durations are balanced in the audio segment, but deteriorates significantly when speech is sparse or absent. In order to mitigate this problem, we pro-pose a simple yet efficient technique which builds an alternative model of speech using data from a separate corpora, and em-beds this new information within the Combo-SAD framework. Our experiments show that the proposed approach has a major impact on SAD performance (i.e., +30 % absolute), especially in audio segments that contain sparse or no speech information.
Ali Ziaei, Lakshmish Kaushik, Abhijeet Sangwan, John H. L. Hansen, Douglas W. Oard
INTERSPEECH4
2014 A speech system for estimating daily word counts
Ali Ziaei, Abhijeet Sangwan, John H. L. Hansen
INTERSPEECH3
2014 Utilization of unlabeled development data for speaker verification
abstract
State-of-the-art speaker verification systems model speaker identity by mapping i-Vectors onto a probabilistic linear discriminant analysis (PLDA) space. Compared to other modeling approaches (such as cosine distance scoring), PLDA provides a more efficient mechanism to separate speaker information from other sources of undesired variabilities and offers superior speaker verification performance. Unfortunately, this efficiency is obtained at the cost of a required large corpus of labeled development data, which is too expensive/unrealistic in many cases. This study investigates a potential solution to resolve this challenge by effectively utilizing unlabeled development data with universal imposter clustering. The proposed method offers +21.9% and +34.6% relative gains versus the baseline system on two public available corpora, respectively. This significant improvement proves the effectiveness of the proposed method.
Gang Liu 0001, Chengzhu Yu, Navid Shokouhi, Abhinav Misra, Hua Xing, John H. L. Hansen
SLT6
2014 Multichannel feature enhancement in distributed microphone arrays for robust distant speech recognition in smart rooms
abstract
Room reverberation and environmental noise present challenges for integration of speech recognition technology in smart room applications. We present a multichannel enhancement framework for distributed microphone arrays to mitigate the effects of both additive noise and reverberation on distant-talking microphones. The proposed approach uses techniques of nonnegative matrix and tensor factorization to achieve both noise suppression (through sparse representation of speech spectra) and dereverberation (through decomposition of magnitude spectra into convolutive components). Results of ASR experiments on the DIRHA-GRID corpus confirm that the proposed approach can achieve relative improvements of up to +20% in recognition accuracy in highly reverberant and noisy conditions using clean-trained models.
Seyedmahdad Mirsamadi, John H. L. Hansen
SLT2
2014 Spoken language mismatch in speaker verification: An investigation with NIST-SRE and CRSS Bi-Ling corpora
abstract
Compensation for mismatch between acoustic conditions in automatic speaker recognition has been widely addressed in recent years. However, performance degradation due to language mismatch has yet to be thoroughly addressed. In this study, we address langauge mismatch for speaker verification. We select bilingual speaker data from the NIST SRE 04-08 corpora and develop train/test-trials for language matched and mismatched conditions. We first show that language variability significantly degrades speaker recognition performance even with a state-of-the-art i-vector system. Next, we consider two ideas to improve performance: i) we introduce small amounts of multi-lingual speech data to the Probabilistic Linear Discriminant Analysis (PLDA) development set, and ii) explore phoneme level analysis to investigate the effect of language mismatch. It is shown that introducing small amounts of multi-lingual seed data within PLDA training has a significant improvement in speaker verification performance. Also, using data from the CRSS Bi-Ling corpus, we show how various phoneme classes affect speaker verification in language mismatch. This speech corpus consists of bilingual speakers who speak either Hindi or Mandarin, in addition to English. Using this corpus, we propose a novel phoneme histogram normalization technique to match the phonetic spaces of two different languages and show a +16.6% relative improvement in speaker verification performance in the presence of language mismatch.
Abhinav Misra, John H. L. Hansen
SLT2
2014 Training candidate selection for effective rejection in open-set language identification
abstract
Research in open-set language identification (LID) generally focuses more on accurate in-set modeling versus improved out-of-set (OOS) rejection. Unknown or OOS language rejection is a challenge, since research developers seldom commit equivalent OOS corpus development effort versus the desired in-set languages. To address this, we propose an OOS candidate selection method for universal OOS language coverage. Since effective selection always requires abundant knowledge of inter-language relationships, three broad measurements across world languages are considered. Finally, the advanced OOS selection method is evaluated on a database derived from a large-scale corpus (LRE-09) with a state-of-the-art i-Vector system followed by two back-ends. The baseline system is realized using a random selection of OOS candidates. With the proposed selection method and probabilistic linear discriminative analysis (PLDA) back-end, the OOS rejection performance is improved with false alarm and miss rates achieving a relative reduction of 32.6% and 4.4%, respectively. In addition, the overall classification performance are relatively improved 8.4% and 7.5% according to the two back-ends based on an average cost function.
Qian Zhang 0019, John H. L. Hansen
SLT2
2014 A coherence-based noise reduction algorithm for binaural hearing aids
Nima Yousefian, Philipos C. Loizou, John H. L. Hansen
Speech Commun.3
2014 Maximum Likelihood Acoustic Factor Analysis Models for Robust Speaker Verification in Noise
abstract
Recent speaker recognition/verification systems generally utilize an utterance dependent fixed dimensional vector as features to Bayesian classifiers. These vectors, known as i-Vectors, are lower dimensional representations of Gaussian Mixture Model (GMM) mean super-vectors adapted from a Universal Background Model (UBM) using speech utterance features, and extracted utilizing a Factor Analysis (FA) framework. This method is based on the assumption that the speaker dependent information resides in a lower dimensional sub-space. In this study, we utilize a mixture of Acoustic Factor Analyzers (AFA) to model the acoustic features instead of a GMM-UBM. Following our previously proposed AFA technique (“Acoustic factor analysis for robust speaker verification,” by Hasan and Hansen, IEEE Trans. Audio, Speech, Lang. Process., vol. 21, no. 4, April 2013), this model is based on the assumption that the speaker relevant information lies in a lower dimensional subspace in the multi-dimensional feature space localized by the mixture components. Unlike our previous method, here we train the AFA-UBM model directly from the data using an Expectation-Maximization (EM) algorithm. This method shows improved robustness to noise as the nuisance dimensions are removed in each EM iteration. Two variants of the AFA model are considered utilizing an isotropic and diagonal covariance residual term. The method is integrated within a standard i-Vector system where the hidden variables of the model, termed as acoustic factors, are utilized as the input for total variability modeling. Experimental results obtained on the 2012 National Institute of Standards and Technology (NIST) Speaker Recognition Evaluation (SRE) core-extended trials indicate the effectiveness of the proposed strategy in both clean and noisy conditions.
Taufiq Hasan, John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 An investigation into back-end advancements for speaker recognition in multi-session and noisy enrollment scenarios
abstract
This study aims to explore the case of robust speaker recognition with multi-session enrollments and noise, with an emphasis on optimal organization and utilization of speaker information presented in the enrollment and development data. This study has two core objectives. First, we investigate more robust back-ends to address noisy multi-session enrollment data for speaker recognition. This task is achieved by proposing novel back-end algorithms. Second, we construct a highly discriminative speaker verification framework. This task is achieved through intrinsic and extrinsic back-end algorithm modification, resulting in complementary sub-systems. Evaluation of the proposed framework is performed on the NIST SRE2012 corpus. Results not only confirm individual sub-system advancements over an established baseline, the final grand fusion solution also represents a comprehensive overall advancement for the NIST SRE2012 core tasks. Compared with state-of-the-art SID systems on the NIST SRE2012, the novel parts of this study are: 1) exploring a more diverse set of solutions for low-dimensional i-Vector based modeling; and 2) diversifying the information configuration before modeling. All these two parts work together, resulting in very competitive performance with reasonable computational cost.
Gang Liu 0001, John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 Blind Spectral Weighting for Robust Speaker Identification under Reverberation Mismatch
abstract
Room reverberation poses various deleterious effects on performance of automatic speech systems. Speaker identification (SID) performance, in particular, degrades rapidly as reverberation time increases. Reverberation causes two forms of spectro-temporal distortions on speech signals: i) self-masking which is due to early reflections and ii) overlap-masking which is due to late reverberation. Overlap-masking effect of reverberation has been shown to have a greater adverse impact on performance of speech systems. Motivated by this fact, this study proposes a blind spectral weighting (BSW) technique for suppressing the reverberation overlap-masking effect on SID systems. The technique is blind in the sense that prior knowledge of neither the anechoic signal nor the room impulse response is required. Performance of the proposed technique is evaluated on speaker verification tasks under simulated and actual reverberant mismatched conditions. Evaluations are conducted in the context of the conventional GMM-UBM as well as the state-of-the-art i-vector based systems. The GMM-UBM experiments are performed using speech material from a new data corpus well suited for speaker verification experiments under actual reverberant mismatched conditions, entitled MultiRoom8. The i-vector experiments are carried out with microphone (interview and phonecall) data from the NIST SRE 2010 extended evaluation set which are digitally convolved with three different measured room impulse responses extracted from the Aachen impulse response (AIR) database. Experimental results prove that incorporating the proposed blind technique into the standard MFCC feature extraction framework yields significant improvement in SID performance under reverberation mismatch.
Seyed Omid Sadjadi, John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.2
2013 Automatic sentiment extraction from YouTube videos
abstract
Extracting speaker sentiment from natural audio streams such as YouTube is challenging. A number of factors contribute to the task difficulty, namely, Automatic Speech Recognition (ASR) of spontaneous speech, unknown background environments, variable source and channel characteristics, accents, diverse topics, etc. In this study, we build upon our previous work [5], where we had proposed a system for detecting sentiment in YouTube videos. Particularly, we propose several enhancements including (i) better text-based sentiment model due to training on larger and more diverse dataset, (ii) an iterative scheme to reduce sentiment model complexity with minimal impact on performance accuracy, (iii) better speech recognition due to superior acoustic modeling and focused (domain dependent) vocabulary/language models, and (iv) a larger evaluation dataset. Collectively, our enhancements provide an absolute 10% improvement over our previous system in terms of sentiment detection accuracy. Additionally, we also present analysis that helps understand the impact of WER (word error rate) on sentiment detection accuracy. Finally, we investigate the relative importance of different Parts-of-Speech (POS) tag features towards sentiment detection. Our analysis reveals the practicality of this technology and also provides several potential directions for future work.
Lakshmish Kaushik, Abhijeet Sangwan, John H. L. Hansen
ASRU3
2013 Duration mismatch compensation for i-vector based speaker recognition systems
abstract
Speaker recognition systems trained on long duration utterances are known to perform significantly worse when short test segments are encountered. To address this mismatch, we analyze the effect of duration variability on phoneme distributions of speech utterances and i-vector length. We demonstrate that, as utterance duration is decreased, number of detected unique phonemes and i-vector length approaches zero in a logarithmic and non-linear fashion, respectively. Assuming duration variability as an additive noise in the i-vector space, we propose three different strategies for its compensation: i) multi-duration training in Probabilistic Linear Discriminant Analysis (PLDA) model, ii) score calibration using log duration as a Quality Measure Function (QMF), and iii) multi-duration PLDA training with synthesized short duration i-vectors. Experiments are designed based on the 2012 National Institute of Standards and Technology (NIST) Speaker Recognition Evaluation (SRE) protocol with varying test utterance duration. Experimental results demonstrate the effectiveness of the proposed schemes on short duration test conditions, especially with the QMF calibration approach.
Taufiq Hasan, Rahim Saeidi, John H. L. Hansen, David A. van Leeuwen
ICASSP3
2013 CRSS systems for 2012 NIST Speaker Recognition Evaluation
abstract
This paper describes the systems developed by the Center for Robust Speech Systems (CRSS), for the 2012 National Institute of Standards and Technology (NIST) Speaker Recognition Evaluation (SRE). Given that the emphasis of SRE'12 is on noisy and short duration test conditions, our system development focused on: (i) novel robust acoustic features, (ii) new feature normalization schemes, (iii) various back-end strategies utilizing multi-session and multi-condition training, and (iv) quality measure based system fusion. Noisy and short duration training/test conditions are artificially generated and effectively utilized. Active speech duration and signal-to-noise-ratio (SNR) estimates are successfully employed as quality measures for system calibration and fusion. Overall system performance was very successful for the given test conditions.
Taufiq Hasan, Seyed Omid Sadjadi, Gang Liu 0001, Navid Shokouhi, Hynek Boril, John H. L. Hansen
ICASSP6
2013 Sentiment extraction from natural audio streams
abstract
Automatic sentiment extraction for natural audio streams containing spontaneous speech is a challenging area of research that has received little attention. In this study, we propose a system for automatic sentiment detection in natural audio streams such as those found in YouTube. The proposed technique uses POS (part of speech) tagging and Maximum Entropy modeling (ME) to develop a text-based sentiment detection model. Additionally, we propose a tuning technique which dramatically reduces the number of model parameters in ME while retaining classification capability. Finally, using decoded ASR (automatic speech recognition) transcripts and the ME sentiment model, the proposed system is able to estimate the sentiment in the YouTube video. In our experimental evaluation, we obtain encouraging classification accuracy given the challenging nature of the data. Our results show that it is possible to perform sentiment analysis on natural spontaneous speech data despite poor WER (word error rates).
Lakshmish Kaushik, Abhijeet Sangwan, John H. L. Hansen
ICASSP3
2013 An advanced feature compensation method employing acoustic model with phonetically constrained structure
abstract
This study proposes an effective model-based feature compensation method for robust speech recognition in background noise conditions. In the proposed scheme, an acoustic model with a phonetically constrained structure is employed for the Parallel Combined Gaussian Mixture Model (PCGMM [1]) based feature compensation method. The structure of the acoustic model includes a collection of context independent phone models. A phonetically constrained prior probability is formulated by integrating transition probability of phone models into the reconstruction procedure. Experimental results show that the PCGMM-based feature compensation employing the proposed phonetically constrained structure of acoustic model consistently outperforms the case of employing the conventional Gaussian mixture model. This demonstrates that the proposed configuration of the acoustic model is effective at improving the intelligibility of the speech reconstructed by the feature compensation method for speech recognition under diverse background noise conditions.
Wooil Kim, John H. L. Hansen
ICASSP2
2013 An investigation on back-end for speaker recognition in multi-session enrollment
abstract
This study explores various back-end classifiers for robust speaker recognition in multi-session enrollment, with emphasis on optimal utilization and organization of speaker information present in the development data. Our objective is to construct a highly discriminative back-end framework by fusing several back-ends on an i-vector system framework. It is demonstrated that, by using different information/data configuration and modeling schemes, performance of the fused system can be significantly improved compared to an individual system using a single front-end and back-end. Averaged across both genders, we obtain a relative improvement in EER and minDCF by 56.5% and 49.4%, respectively. Consistent performance gains obtained using the proposed strategy validates its effectiveness. This system is part of the CRSS' NIST SRE 2012 submission system.
Gang Liu 0001, Taufiq Hasan, Hynek Boril, John H. L. Hansen
ICASSP4
2013 Robust front-end processing for speaker identification over extremely degraded communication channels
abstract
Effective front-end processing, which often involves feature extraction and speech activity detection (SAD), is essential for robustness in speech systems. In this study, we propose an unsupervised SAD scheme based on four different speech voicing measures which are combined with a perceptual spectral flux feature. Effectiveness of the proposed scheme is evaluated and compared against several commonly adopted unsupervised SAD methods under actual harsh acoustic conditions. As an example application, we also evaluate performance of the proposed SAD in the context of an i-vector based speaker identification (SID) system, where the recently introduced mean Hilbert envelope coefficients (MHEC) are benchmarked against conventional MFCCs. Long and spontaneous conversational audio recordings from DARPA program RATS (Phase-I) are used in our evaluations. Experimental results indicate that the proposed SAD solution is highly effective and provides superior performance compared to other unsupervised SAD techniques considered. In addition, it is shown that MHECs are effective alternatives to MFCCs for SID tasks under severe degraded channel conditions.
Seyed Omid Sadjadi, John H. L. Hansen
ICASSP2
2013 Overlapped-speech detection with applications to driver assessment for in-vehicle active safety systems
abstract
In this study we propose a system for overlapped-speech detection. Spectral harmonicity and envelope features are extracted to represent overlapped and single-speaker speech using Gaussian mixture models (GMM). The system is shown to effectively discriminate the single and overlapped speech classes. We further increase the discrimination by proposing a phoneme selection scheme to generate more reliable artificial overlapped data for model training. Evaluations on artificially generated co-channel data show that the novelty in feature selection and phoneme omission results in a relative improvement of 10% in the detection accuracy compared to baseline. As an example application, we evaluate the effectiveness of overlapped-speech detection for vehicular environments and its potential in assessing driver alertness. Results indicate a good correlation between driver performance and the amount and location of overlapped-speech segments.
Navid Shokouhi, Amardeep Sathyanarayana, Seyed Omid Sadjadi, John H. L. Hansen
ICASSP4
2013 Speaker height estimation combining GMM and linear regression subsystems
abstract
There are both scientific and technology based motivations for establishing effective speech processing algorithms that estimate speaker traits. Estimating speaker height can assist in voice forensic analysis [1], as well as provide additional side knowledge to improve speaker ID systems, or acoustic model selection for improved speech recognition. In this study, two distinct approaches for height estimation are explored. The first approach is statistical based and incorporates acoustic models within a GMM structure, while the second is a direct speech analysis approach that employs linear regression to obtain the height directly. The accuracy and trade-offs of these systems are explored as well a fusion of the two systems using data from the TIMIT corpus (which includes ground truth on speaker height).
Keri A. Williams, John H. L. Hansen
ICASSP2
2013 A new mask-based objective measure for predicting the intelligibility of binary masked speech
abstract
Mask-based objective speech-intelligibility measures have been successfully proposed for evaluating the performance of binary masking algorithms. These objective measures were computed directly by comparing the estimated binary mask against the ground truth ideal binary mask (IdBM). Most of these objective measures, however, assign equal weight to all time-frequency (T-F) units. In this study, we propose to improve the existing mask-based objective measures by weighting each T-F unit according to its target or masker loudness. The proposed objective measure shows significantly better performance than two other existing mask-based objective measures.
Chengzhu Yu, Kamil K. Wójcicki, Philipos C. Loizou, John H. L. Hansen
ICASSP4
2013 Supervector pre-processing for PRSVM-based Chinese and Arabic dialect identification
abstract
Phonotactic modeling has become a widely used means for speaker, language, and dialect recognition. This paper explores variations to supervector pre-processing for phone recognition-support vector machines (PRSVM) based dialect identification. The aspects studied are: (i) normalization of supervector dimensions in the pre-squashing stage, (ii) impact of alternative squashing functions, and (iii) N-gram selection for supervector dimensionality reduction. In (i) and (ii), we find that several alternatives to commonly used approaches can provide moderate, yet consistent performance improvements. In (iii), a newly proposed dialect salience measure is applied in supervector dimension selection and compared to a common N-gram frequency based selection. The results show a strong correlation between dialect-salience and frequency of occurrence in N-grams. The evaluations in this study are conducted on a corpus of Chinese dialects, a Pan-Arabic corpus, and a set of Arabic CTS corpora.
Qian Zhang 0019, Hynek Boril, John H. L. Hansen
ICASSP3
2013 Prof-Life-Log: Personal interaction analysis for naturalistic audio streams
abstract
Analysis of personal audio recordings is a challenging and interesting subject. Using contemporary speech and language processing techniques, it is possible to mine personal audio recordings for a wealth of information that can be used to measure a person's engagement with their environment as well as other people. In this study, we propose an analysis system that uses personal audio recordings to automatically estimate the number of unique people and environments which encompass the total engagement within the recording. The proposed system uses speech activity detection (SAD), speaker diarization and environmental sniffing techniques, and is evaluated on naturalistic audio streams from the Prof-Life-Log corpus. We also report performance of the individual systems, and also present a combined analysis which reveals the interaction of the subject with both people and environment. Hence, this study establishes the efficacy and novelty of using contemporary speech technology for life logging applications.
Ali Ziaei, Abhijeet Sangwan, John H. L. Hansen
ICASSP3
2013 A preliminary study of child vocalization on a parallel corpus of US and shanghainese toddlers
abstract
This paper studies various aspects of child vocalization as captured in a newly established parallel corpus of sixteen 18–31 months old US and Shanghainese toddlers. The recordings were acquired in 16-hour sessions during an ‘ordinary’ day in the child’s natural environment and manually labeled. The vocalization characteristics are studied by means of phonotactic and prosodic analysis with emphasis on automatic processing. In the phonotactic domain, a Gaussian mixture model (GMM) tokenizer, a bank of phone recognizers, and formant tracking are used to analyze the movements in the acoustic-phonetic space. In the prosodic domain, pitch patterns, duration, and rhythm are analyzed. Besides strong individual-specific characteristics of the subjects in some of the domains considered, the two language groups show differences in the occupation of the F1–F2 formant space, choice of pitch pattern durations, and consistency in producing complex phonetic patterns. Index Terms: children vocalization, speech acquisition, phonotactic modeling, pitch patterns, rhythmicity parameters
Hynek Boril, Qian Zhang 0019, Pongtep Angkititrakul, John H. L. Hansen, Dongxin Xu, Jill Gilkerson, Jeffrey A. Richards
INTERSPEECH4
2013 Impact of noise reduction and spectrum estimation on noise robust speaker identification
abstract
Many spectrum estimation methods and speech enhancement algorithms have previously been evaluated for noise-robust speaker identification (SID). However, these techniques have mostly been evaluated over artificially noised, mismatched training tasks with GMM-UBM speaker models. It is therefore unclear whether performance improvements observed with these methods translate to a broader range of noisy SID tasks. This study compares selected spectrum estimation methods from three classes: cochlear filterbanks, alternative time-domain windowing, and linear prediction-based techniques, as well as a set of frequencydomain noise reduction algorithms, across a suite of 8 evaluation tasks. The evaluation tasks are designed to expand upon the limited tasks addressed in past evaluations by exploring three research questions: performance on real noise versus artificial noise, performance on matched training tasks versus mismatched tasks, and performance when paired with an i-vector backend versus a GMM-UBM backend. We find that noise-robust spectrum estimation methods can improve the performance of SID systems over the range of noise tasks evaluated, including real noisy tasks, matched training tasks, and i-vector backends. However, performance on the typical GMM-UBM mismatched artificially noised case did not predict performance on other tasks. Finally, the matched enrollment case is a significantly different problem than the mismatched enrollment case. Index Terms: mismatched condition, noise robustness, robust features, speaker identification, speech enhancement
Keith W. Godin, Seyed Omid Sadjadi, John H. L. Hansen
INTERSPEECH3
2013 All for one: feature combination for highly channel-degraded speech activity detection
abstract
Speech activity detection (SAD) on channel transmissions is a critical preprocessing task for speech, speaker and language recognition or for further human analysis. This paper presents a feature combination approach to improve SAD on highly channel degraded speech as part of the Defense Advanced
Martin Graciarena, Abeer Alwan, Daniel P. W. Ellis, Horacio Franco, Luciana Ferrer, John H. L. Hansen, Adam Janin, Yun Lei, Vikramjit Mitra, Nelson Morgan, Seyed Omid Sadjadi, T. J. Tsai 0001, Nicolas Scheffer, Lee Ngee Tan
INTERSPEECH6
2013 Acoustic factor analysis based universal background model for robust speaker verification in noise
abstract
The Universal Background Model (UBM) is known as a speaker independent Gaussian Mixture Model (GMM) trained on a large speech corpus containing many speakers’ recordings in various conditions. When noisy test data is involved, UBM trained on clean data is generally not optimal. Using noisy data for UBM training, however, creates a bias towards the specific development noise samples resulting in degraded speaker recognition performance in unseen noise types. In this study, we utilize an Acoustic Factor Analysis (AFA) based UBM that iteratively learns the dominant feature sub-spaces in each mixture component, resulting in a more robust model. We explore two variants of the model: one with an isotropic and the other with a diagonal residual noise. The Maximum-Likelihood (ML) training formulations of the models are provided. The latent variables of the model, termed acoustic factors, are used as features to train the second stage of factor analysis parameters, i.e., the traditional i-vector extractor. Experiments performed on the 2012 National Institute of Standards and Technology (NIST) Speaker Recognition Evaluation (SRE) indicate the effectiveness of the proposed strategy in both clean and noisy conditions. Index Terms: speaker verification, NIST SRE 2012, noisy data, acoustic factor analysis
Taufiq Hasan, John H. L. Hansen
INTERSPEECH2
2013 Automatic regularization of cross-entropy cost for speaker recognition fusion
abstract
\n Contains fulltext :\n 116325.pdf (author's version ) (Open Access)\n
Ville Hautamäki, Kong-Aik Lee, David A. van Leeuwen, Rahim Saeidi, Anthony Larcher, Tomi Kinnunen, Taufiq Hasan, Seyed Omid Sadjadi, Gang Liu 0001, Hynek Boril, John H. L. Hansen, Benoit G. B. Fauve
INTERSPEECH11
2013 Dimensionality analysis of singing speech based on locality preserving projections
abstract
In this study, we expand the question of ”what is the intrinsic dimensionality of speech?” to ”how does the intrinsic dimensionality of speech change from speaking to singing?”. Our focus is on dimensionality of the vowel space regarding spectral features, which is important in acoustic modeling applications. Locality Preserving Projection (LPP) is applied for dimensionality reduction of the spectral feature vectors, and vowel classification performance is studied in low-dimensional subspaces. Performance analysis of singing and speaking vowel classification based on reducing the dimension shows that compared to speaking, a higher number of dimensions is required for effective representation of singing vowels. The results are also explained in terms of differences in the formant spaces of singing and speaking, and vowel classification performance is analyzed based on feature vectors consisting of formant frequencies. The formant analysis results are shown to be consistent with LPP dimensionality analysis, which verifies the inherent dimensionality increase of the vowel space from speaking to singing.
Mahnoosh Mehrabani, John H. L. Hansen
INTERSPEECH2
2013 'houston, we have a solution': using NASA apollo program to advance speech and language processing technology
abstract
NASA’s Apollo program stands as one of mankind’s greatest achievements in the 20th century. During a span of 4 years (from 1968 to 1972), a total of 9 lunar missions were launched and 12 astronauts walked on the surface of the moon. It was one the most complex operations executed from scientific, techno-logical and operational perspectives. In this paper, we describe our recent efforts in gathering and organizing the Apollo pro-gram data. It is important to note that the audio content captured during the 7-10 day missions represent the coordinated efforts of hundreds of individuals within NASA Mission Control, re-sulting in well over 100k hours of data for the entire program. It is our intention to make the material stemming from this ef-fort available to the research community to further research ad-vancements in speech and language processing. Particularly, we describe the speech and text aspects of the Apollo data while pointing out its applicability to several classical speech pro-cessing and natural language processing problems such as au-dio processing, speech and speaker recognition, information re-trieval, document linking and a range of other processing tasks which enable knowledge search, retrieval, and understanding.. We also highlight some of the outstanding opportunities and challenges associated with this dataset. Finally, we also present initial results for speech recognition, document linking, and au-dio processing systems.
Abhijeet Sangwan, Lakshmish Kaushik, Chengzhu Yu, John H. L. Hansen, Douglas W. Oard
INTERSPEECH4
2013 I4u submission to NIST SRE 2012: a large-scale collaborative effort for noise-robust speaker verification
abstract
I4U is a joint entry of nine research Institutes and Universities across 4 continents to NIST SRE 2012. It started with a brief discussion during the Odyssey 2012 workshop in Singapore. An online discussion group was soon set up, providing a discussion platform for different issues surrounding NIST SRE’12. Noisy test segments, uneven multi-session training, variable enrollment duration, and the issue of open-set identification were actively discussed leading to various solutions integrated to the I4U submission. The joint submission and several of its 17 sub-systems were among top-performing systems. We summarize the lessons learnt from this large-scale effort.
Rahim Saeidi, Kong-Aik Lee, Tomi Kinnunen, Tawfik Hasan, Benoit G. B. Fauve, Pierre-Michel Bousquet, Elie Khoury 0001, Pablo Luis Sordo Martinez, Jia Min Karen Kua, Chang Huai You, Hanwu Sun, Anthony Larcher, Padmanabhan Rajan, Ville Hautamäki, Cemal Hanilçi, Billy Braithwaite, Rosa González Hautamäki, Seyed Omid Sadjadi, Gang Liu 0001, Hynek Boril, Navid Shokouhi, Driss Matrouf, Laurent El Shafey, Pejman Mowlaee, Julien Epps, Tharmarajah Thiruvaran, David A. van Leeuwen, Bin Ma 0001, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Sébastien Marcel, John S. D. Mason, Eliathamby Ambikairajah
INTERSPEECH30
2013 Belt Up: Investigating the impact of in-vehicular conversation on driving performance
abstract
In-vehicle conversations are typical when there is more than one person in the car. Although many conversations are beneficial in keeping the driver alert and active, there are also instances where a competitive conversation may adversely influence driving performance. Identifying such scenarios can improve vehicle safety systems by fusing the knowledge obtained from conversational speech analysis and vehicle dynamic signals. In this study we incorporate the use of smart portable devices to create a unified platform for recording in-vehicular conversations as well as the vehicle dynamic signals required to evaluate the driving performance. Results show that turn taking rate and overlapping speech segments under certain conditions correlate with deviations from normal driving patterns. The conversational speech analysis can thus be utilized as a component in driver assistance systems such that the impact of in-vehicle speech activity on driving performance is controlled or minimized.
Amardeep Sathyanarayana, Navid Shokouhi, Seyed Omid Sadjadi, John H. L. Hansen
Intelligent Vehicles Symposium4
2013 Linking transcribed conversational speech
abstract
As large collections of historically significant recorded speech become increasingly available, scholars are faced with the challenge of making sense of what they hear. This paper proposes automatically linking conversational speech to related resources as one way of supporting that sense-making task. Experiment results with transcribed conversations suggest that this kind of linking has promise for helping to contextualize recordings of detail-oriented conversations, and that simple sliding-window bag-of-words techniques can identify some useful links.
Joseph Malionek, Douglas W. Oard, Abhijeet Sangwan, John H. L. Hansen
SIGIR4
2013 Acoustic analysis and feature transformation from neutral to whisper for speaker identification within whispered speech audio streams
John H. L. Hansen
Speech Commun.2
2013 In-set/out-of-set speaker recognition in sustained acoustic scenarios using sparse data
John H. L. Hansen, Jun-Won Suh, Matthew R. Leonard
Speech Commun.1
2013 Singing speaker clustering based on subspace learning in the GMM mean supervector space
abstract
In this study, we propose algorithms based on subspace learning in the GMM mean supervector space to improve performance of speaker clustering with speech from both reading and singing. As a speaking style, singing introduces changes in the time-frequency structure of a speaker’s voice. The purpose of this study is to introduce advancements for speech systems such as speech indexing and retrieval which improve robustness to intrinsic variations in speech production. Speaker clustering techniques such as k-means and hierarchical are explored for analysis of acoustic space differences of a corpus consisting of reading and singing of lyrics for each speaker. Furthermore, a distance based on fuzzy c-means membership degrees is proposed to more accurately measure clustering difficulty or speaker confusability. Two categories of subspace learning methods are studied: unsupervised based on LPP, and supervised based on PLDA. Our proposed clustering method based on PLDA is a two stage algorithm: where first, initial clusters are obtained using full dimension supervectors, and next, each cluster is refined in a PLDA subspace resulting in a more speaker dependent representation that is less sensitive to speaking style. It is shown that LPP improves average clustering accuracy by 5.1% absolute versus a hierarchical baseline for a mixture of reading and singing, and PLDA based clustering increases accuracy by 9.6% absolute versus a k-means baseline. The advancements offer novel techniques to improve model formulation for speech applications including speaker ID, audio search, and audio content analysis.
Mahnoosh Mehrabani, John H. L. Hansen
Speech Commun.2
2013 Unsupervised Speech Activity Detection Using Voicing Measures and Perceptual Spectral Flux
abstract
Effective speech activity detection (SAD) is a necessary first step for robust speech applications. In this letter, we propose a robust and unsupervised SAD solution that leverages four different speech voicing measures combined with a perceptual spectral flux feature, for audio-based surveillance and monitoring applications. Effectiveness of the proposed technique is evaluated and compared against several commonly adopted unsupervised SAD methods under simulated and actual harsh acoustic conditions with varying distortion levels. Experimental results indicate that the proposed SAD scheme is highly effective and provides superior and consistent performance across various noise types and distortion levels.
Seyed Omid Sadjadi, John H. L. Hansen
IEEE Signal Process. Lett.2
2013 Acoustic Factor Analysis for Robust Speaker Verification
abstract
Factor analysis based channel mismatch compensation methods for speaker recognition are based on the assumption that speaker/utterance dependent Gaussian Mixture Model (GMM) mean super-vectors can be constrained to reside in a lower dimensional subspace. This approach does not consider the fact that conventional acoustic feature vectors also reside in a lower dimensional manifold of the feature space, when feature covariance matrices contain close to zero eigenvalues. In this study, based on observations of the covariance structure of acoustic features, we propose a factor analysis modeling scheme in the acoustic feature space instead of the super-vector space and derive a mixture dependent feature transformation. We demonstrate how this single linear transformation performs feature dimensionality reduction, de-correlation, normalization and enhancement, at once. The proposed transformation is shown to be closely related to signal subspace based speech enhancement schemes. In contrast to traditional front-end mixture dependent feature transformations, where feature alignment is performed using the highest scoring mixture, the proposed transformation is integrated within the speaker recognition system using a probabilistic feature alignment technique, which nullifies the need for regenerating the features/retraining the Universal Background Model (UBM). Incorporating the proposed method with a state-of-the-art i-vector and Gaussian Probabilistic Linear Discriminant Analysis (PLDA) framework, we perform evaluations on National Institute of Science and Technology (NIST) Speaker Recognition Evaluation (SRE) 2010 core telephone and microphone tasks. The experimental results demonstrate the superiority of the proposed scheme compared to both full-covariance and diagonal covariance UBM based systems. Simple equal-weight fusion of baseline and proposed systems also yield significant performance gains.
Taufiq Hasan, John H. L. Hansen
IEEE Trans. Speech Audio Process.2
2013 Automatic Accent Assessment Using Phonetic Mismatch and Human Perception
abstract
In this study, a new algorithm for automatic accent evaluation of native and non-native speakers is presented. The proposed system consists of two main steps: alignment and scoring. In the alignment step, the speech utterance is processed using a Weighted Finite State Transducer (WFST) based technique to automatically estimate the pronunciation mismatches (substitutions, deletions, and insertions). Subsequently, in the scoring step, two scoring systems which utilize the pronunciation mismatches from the alignment phase are proposed: (i) a WFST-scoring system to measure the degree of accentedness on a scale from -1 (non-native like) to +1 (native like), and a (ii) Maximum Entropy (ME) based technique to assign perceptually motivated scores to pronunciation mismatches. The accent scores provided from the WFST-scoring system as well as the ME scoring system are termed as the WFST and P-WFST (perceptual WFST) accent scores, respectively. The proposed systems are evaluated on American English (AE) spoken by native and non-native (native speakers of Mandarin-Chinese) speakers from the CU-Accent corpus. A listener evaluation of 50 Native American English (N-AE) was employed to assist in validating the performance of the proposed accent assessment systems. The proposed P-WFST algorithm shows higher and more consistent correlation with human evaluated accent scores, when compared to the Goodness Of Pronunciation (GOP) measure. The proposed solution for accent classification and assessment based on WFST and P-WFST scores show that an effective advancement is possible which correlates well with human perception.
Freddy William, Abhijeet Sangwan, John H. L. Hansen
IEEE Trans. Speech Audio Process.3
2012 A multi-modal highlight extraction scheme for sports videos using an information-theoretic excitability measure
abstract
A generic method for sports video highlight selection is presented in this study. Processing begins where the video is divided into short segments and several multi-modal features are extracted from each video segment. Excitability is computed based on the likelihood of the features lying in certain regions of their probability density functions that are exciting and rare. The proposed measure is used to rank order the partitioned segment stream to compress the overall video sequence and produce a contiguous set of highlights. Experiments are performed on baseball videos using excitement in the commentators' speech, audio energy, slow motion replay, scene cut density, and motion activity as features. Subjective evaluation of excitability and ranking of video segments yield a higher correlation with the proposed measure compared to well-established techniques indicating the effectiveness of the approach.
Taufiq Hasan, Hynek Boril, Abhijeet Sangwan, John H. L. Hansen
ICASSP4
2012 Feature compensation employing online GMM adaptation for speech recognition in unknown severely adverse environments
abstract
This study proposes an effective feature compensation-method to improve speech recognition in real-life speech conditions, where (i) severe background noise and channel distortion simultaneously exist, (ii) no development data is available, and (iii) clean data for ASR training and the latent clean speech in the test data are mismatched in the acoustic structure. The proposed feature compensation method employs an online GMM adaptation procedure which is based on MLLR, and a minimum statistics replacement technique for non-speech segments. The DARPA Tank corpus is used for performance evaluation, which includes severe real-life noisy conditions. The clean Broadcast News (BN) corpus is used for training the speech recognition system in this study. Experimental results show that the proposed feature compensation scheme outperforms GMM-based FMLLR and the ETSI AFE for DARPA Tank data, achieving a +5.56% relative improvement compared to FMLLR. These results demonstrate that the proposed feature compensation scheme is effective at improving speech recognition performance in unknown real-life adverse environments.
Wooil Kim, John H. L. Hansen
ICASSP2
2012 Robust feature front-end for speaker identification
abstract
One important challenge for speaker identification (SID) system is sustained performance in diverse conditions. This study presents a novel front-end feature extraction method for SID in clean, noisy, and channel-mismatched acoustic conditions. To address the problem, the perceptual minimum variance distortionless response (PMVDR) feature is employed. While PMVDR has been successfully used for noisy ASR, it has not been considered for SID. We also incorporate longer temporal speaker knowledge based on the shifted delta cepstral (SDC) algorithm. The evaluation over YOHO and another new diversified Robust Open-Set Speaker Identification (ROSSI) database show that both PMVDR and the union with SDC can improve performance significantly. Compared with traditional feature extraction, PMVDR and PMVDR-SDC always give improvement across diverse adverse conditions. Also, PMVDR-SDC can contribute additional improvement in the presence of noise and channel mismatch.
Gang Liu 0001, Yun Lei, John H. L. Hansen
ICASSP3
2012 A fast speaker verification with universal background support data selection
abstract
In this study, a fast universal background support imposter data selection method is proposed, which is integrated within a support vector machine (SVM) based speaker verification system. Selection of an informative background dataset is crucial in constructing a discriminative decision super-plane between the enrollment and imposter speakers. Previous studies generally derive the optimal number of imposter examples from development data and apply to the evaluation data, which cannot guarantee consistent performance and often necessitate expensive searching. In the proposed method, the universal background dataset is derived so as to embed imposter knowledge in a more balanced way. Next, the derived dataset is taken as the imposter set in the SVM modeling process for each enrollment speaker. By using imposter adaptation, a more detailed subspace per target speaker can be constructed. Compared to the popular support-vector frequency based method, the proposed method can not only avoid parameter searching but offers a significant improvement and generalizes better on the unseen data.
Gang Liu 0001, Jun-Won Suh, John H. L. Hansen
ICASSP3
2012 A comparison of front-end compensation strategies for robust LVCSR under room reverberation and increased vocal effort
abstract
Automatic speech recognition is known to deteriorate in the presence of room reverberation and variation of vocal effort in speakers. This study considers robustness of several state-of-the-art front-end feature extraction and normalization strategies to these sources of speech signal variability in the context of large vocabulary continuous speech recognition (LVCSR). A speech database recorded in an anechoic room, capturing modal speech and speech produced at different levels of vocal effort, is reverberated using measured room impulse responses and utilized in the evaluations. It is shown that the combination of recently introduced mean Hilbert envelope coefficients (MHEC) and a normalization strategy combining cepstral gain normalization and modified RASTA filtering (CGN RASTALP) provides considerable recognition performance gains for reverberant modal and high vocal effort speech.
Seyed Omid Sadjadi, Hynek Boril, John H. L. Hansen
ICASSP3
2012 Blind reverberation mitigation for robust speaker identification
abstract
Reverberation poses detrimental effects on performance of automatic speaker identification (SID) systems. This paper proposes a blind spectral weighting technique for combating the late reverberation effect (aka overlap-masking effect) on SID systems. The technique is blind in the sense that prior knowledge of neither the anechoic signal nor the room impulse response is required. Performance of the proposed technique is evaluated in terms of: 1) accuracy obtained from closed-set SID experiments, using speech material from the TIMIT corpus and four different measured room impulse responses from Aachen impulse response (AIR) database, and 2) equal-error rate (EER) obtained from experiments on a new data corpus well suited for speaker verification experiments under actual reverberant mismatched conditions, entitled MultiRoom8. Results prove that incorporating the proposed blind technique into the standard MFCC feature extraction framework yields significant improvement in SID performance.
Seyed Omid Sadjadi, John H. L. Hansen
ICASSP2
2012 ProfLifeLog: Environmental analysis and keyword recognition for naturalistic daily audio streams
abstract
This study presents keyword recognition evaluation on a new corpus named ProfLifeLog. ProfLifeLog is a collection of data captured on a portable audio recording device called the LENA unit. Each session in ProfLifeLog consists of 10+ hours of continuous audio recording that captures the work day of the speaker (person wearing the LENA unit). This study presents keyword spotting evaluation on the ProfLifeLog corpus using the PCN-KWS (phone confusion network-keyword spotting) algorithm [2]. The ProfLifeLog corpus contains speech data in a variety of noise backgrounds which is challenging for keyword recognition. In order to improve keyword recognition, this study also develops a front-end environment estimation strategy that uses the knowledge of speech-pause decisions and SNR (signal-to-noise ratio) to provide noise robustness. The combination of the PCN-KWS and the proposed front-end technique is evaluated on 1 hour of ProfLifeLog corpus. Our evaluation experiments demonstrate the effectiveness of the proposed technique as the number of false alarms in keyword recognition are reduced considerably.
Abhijeet Sangwan, Ali Ziaei, John H. L. Hansen
ICASSP3
2012 Arabic Dialect Identification - 'Is the Secret in the Silence?' and Other Observations
abstract
Conversational telephone speech (CTS) collections of Arabic dialects distributed trough the Linguistic Data Consortium (LDC) provide an invaluable resource for the development of robust speech systems including speaker and speech recognition, translation, spoken dialogue modeling, and information summarization. They are frequently relied on also in language (LID) and dialect identification (DID) evaluations. The first part of this study attempts to identify the source of the relatively high DID performance on LDC’s Arabic CTS corpora seen in recent literature. It is found that recordings of each dialect exhibit unique channel and noise characteristics and that silence regions are sufficient for performing reasonably accurate DID. The second part focuses on phonotactic dialect modeling that utilizes phone recognizers and support vector machines (PRSVM). A simple N-gram normalization of PRSVM input supervectors utilizing hard limiting is introduced and shown to outperform the standard approach used in current LID and DID systems.
Hynek Boril, Abhijeet Sangwan, John H. L. Hansen
INTERSPEECH3
2012 Glottal Waveform Analysis of Physical Task Stress Speech
abstract
Physical task stress affects the acoustic speech wave in various ways. Motivated by observations that fundamental frequency and open quotient are affected by physical task stress, this study examines the effects of physical task stress on a set of glottal features. It is shown that a set of six glottal features can be used for physical stress detection, implying that physical task stress affects vocal fold behavior. It is also shown that the distributions of these six glottal features, across all available speech data, are not affected by physical task stress, leading to the conclusion that covariation in the features due to physical task stress represents different behavioral responses to physical stress. Age and exertion level are explored and rejected as explanatory variables for behavioral types. It is shown, however, that the glottal measurements show changes in distribution when restricted to one recording session of one speaker, suggesting that a given speaker may adopt and retain a particular response to physical task stress for the duration of a task.
Keith W. Godin, Taufiq Hasan, John H. L. Hansen
INTERSPEECH3
2012 Front-end Channel Compensation using Mixture-dependent Feature Transformations for i-Vector Speaker Recognition
abstract
State-of-the-art session variability compensation for speaker recognition are generally based on various linear statistical models of the Gaussian Mixture Model (GMM) mean super-vectors, while frontend features are only processed by standard normalization techniques. In this study, we propose a front-end channel compensation frame-work using mixture-localized linear transforms that operate before super-vector domain modeling begins. In this approach, local linear transforms are trained for each Gaussian component of a Universal Background Model (UBM), and are applied to acoustic features according to their mixture-wise probabilistic alignment, yielding an operation that is globally non-linear. We examine Principal Component Analysis (PCA), whitening, Linear Discriminant Analysis (LDA) and Nuisance Attribute Projection (NAP) as frontend feature transformations. We also propose a method, Nuisance Attribute Elimination (NAE), which is similar to NAP but performs dimensionality reduction in addition to channel compensation. We show that the proposed frame-work can be readily integrated with a standard i-Vector system by simply applying the transformations on the first order Baum-Welch statistics and transforming the UBM. Experiments performed on the telephone trials of the NIST SRE 2010 demonstrate significant performance gain from the proposed frame-work, especially using LDA as the front-end transformation.
Taufiq Hasan, John H. L. Hansen
INTERSPEECH2
2012 Integrated Feature Normalization and Enhancement for robust Speaker Recognition using Acoustic Factor Analysis
abstract
Abstract : State-of-the-art factor analysis based channel compensation methods for speaker recognition are based on the assumption that speaker/utterance dependent Gaussian Mixture Model (GMM) mean super-vectors can be constrained to lie in a lower dimensional subspace, which does not consider the fact that conventional acoustic features may also be constrained in a similar way in the feature space. In this study, motivated by the low-rank covariance structure of cepstral features, we propose a factor analysis model in the acoustic feature space instead of the super-vector domain and derive a mixture of dependent feature transformation. We demonstrate that, the proposed Acoustic Factor Analysis (AFA) transformation performs feature dimensionality reduction, decorrelation, variance normalization and enhancement at the same time. The transform applies a square-root Wiener gain on the acoustic feature eigenvector directions, and is similar to the signal sub-space based speech enhancement schemes. We also propose several methods of adaptively selecting the AFA parameter for each mixture. The proposed feature transformation is applied using a probabilistic mixture alignment, and is integrated with a conventional i-Vector system. Experimental results on the telephone trials of the NIST SRE 2010 demonstrate the effectiveness of the proposed scheme.
Taufiq Hasan, John H. L. Hansen
INTERSPEECH2
2012 Gaussian Map based Acoustic Model Adaptation Using Untranscribed Data for Speech Recognition in Severely Adverse Environments
Wooil Kim, John H. L. Hansen
INTERSPEECH2
2012 Speaker Clustering for a Mixture of Singing and Reading
abstract
In this study, we propose a speaker clustering algorithm based on reading and singing speech samples for each speaker. As a speaking style, singing introduces changes in the time-frequency structure of a speaker’s voice. The purpose of this study is to introduce advancements into speech systems such as speech indexing and retrieval which improve robustness to intrinsic variations in speech production. Clustering is performed within a GMM mean supervector space. The proposed method includes two stages. First, initial clusters are obtained using traditional clustering techniques such as k-means, and hierarchical. Next, each cluster is refined in a PLDA subspace resulting in a more speaker dependent representation that is less sensitive to speaking style. The proposed algorithm improves the average clustering accuracy of the k-means baseline by +9.3% absolute.
Mahnoosh Mehrabani, John H. L. Hansen
INTERSPEECH2
2012 Mean Hilbert Envelope Coefficients (MHEC) for Robust Speaker Recognition
abstract
The recently introduced mean Hilbert envelope coefficients (MHEC) have been shown to be an effective alternative to MFCCs for robust speaker identification under noisy and reverberant conditions in relatively small tasks. In this study, we investigate the effectiveness of these acoustic features in the context of a state-of-the-art speaker recognition system. The i-vectors are used to represent the acoustic space of speakers, while modeling is performed via probabilistic linear discriminant analysis (PLDA). We report speaker verification performance on the NIST SRE-2010 extended telephone and microphone trials for both female and male genders. Experimental results confirm consistent superiority of MHECs to traditional MFCCs within i-vector speaker verification, particularly under microphone and telephone training-test mismatch conditions. In addition, fusion of subsystems trained with the individual front-ends proves that the two acoustic features (i.e., MHEC and MFCC) provide complimentary information for recognizing speakers.
Seyed Omid Sadjadi, Taufiq Hasan, John H. L. Hansen
INTERSPEECH3
2012 Phoneme Class Based Adaptation for Mismatch Acoustic Modeling of Distant Noisy Speech
abstract
Abstract : A new adaptation strategy for distant noisy speech is created by phoneme class based approaches for context independent acoustic models. Unlike the previous approaches such as MLLR-MAP adaptation which adapts acoustic model to the features, our phoneme-class based adaptation (PCBA) adapts the distant data features to our acoustic model which has trained on close microphone TIMIT sentences. The essence of PCBA is to create a transformation strategy which makes the distribution of phoneme-classes of distant noisy speech be similar to those of close microphone acoustic model in thirteen dimensional MFCC space (mostly in c0-c1 plane). It creates a mean, orientation and variance adaptation scheme for each phoneme class to compensate the mismatch. New adapted features and new and improved acoustic models which are produced by PCBA are outperforming those created by MLLR-MAP adaptation for ASR and KWS. And PCBA offers a new powerful understanding in acoustic-modeling of distant speech.
Seçkin Uluskan, John H. L. Hansen
INTERSPEECH2
2012 Prof-Life-Log: Audio Environment Detection for Naturalistic Audio Streams
abstract
In this study, we develop a new system for real world audio environment matching. Environment detection within unknown audio streams requires a system that operates in an unsupervised manner since it will be faced with unknown environments without prior information. In addition, the overall solution should be computationally efficient for large audio collection. In the proposed approach, a Gaussian mixture model(GMM) is trained on large amounts of unlabeled audio data and used as a background acoustic model. Subsequently, an acoustic signature vector (ASV) is computed for each environment. Here, the ASV vector is designed to capture the unique acoustic characteristics of an environment. Using the ASV vectors, we demonstrate that it is possible to compute an effective similarity measure between two acoustic environments. We demonstrate the performance of the proposed system on real-world audio data, and compare it to a traditional GMM-UBM (Universal Background Model) system. Experiments show that our system achieves an equal error rate (EER) that is +35% better than a baseline GMM-UBM system.
Ali Ziaei, Abhijeet Sangwan, John H. L. Hansen
INTERSPEECH3
2012 Automatic analysis of Mandarin accented English using phonological features
Abhijeet Sangwan, John H. L. Hansen
Speech Commun.2
2012 Constrained Iterative Speech Enhancement Using Phonetic Classes
abstract
The degree of influence of noise over phonemes is not uniform since it is dependent on their distinct acoustic properties. In this study, the problem of selectively enhancing speech based on broad phoneme classes is addressed using Auto-(LSP), a constrained iterative speech enhancement algorithm. Multiple enhanced utterances are generated for every noisy utterance by varying the Auto-LSP parameters. The noisy utterance is then partitioned into segments based on broad level phoneme classes, and constraints are applied on each segment using a hard decision solution. To alleviate the effect of hard decision errors, a Gaussian mixture model (GMM)-based maximum-likelihood (ML) soft decision solution is also presented. The resulting utterances are evaluated over the TIMIT speech corpus using the Itakura-Saito, segmental signal-to-noise ratio (SNR) and perceptual evaluation of speech quality (PESQ) metrics over four noise types at three SNR levels. Comparative assessment over baseline enhancement algorithms like Auto-LSP, log-minimum mean squared error (log-MMSE), and log-MMSE with speech presence uncertainty (log-MMSE-SPU) demonstrate that the proposed solution exhibits greater consistency in improving speech quality over most phoneme classes and noise types considered in this study.
John H. L. Hansen
IEEE Trans. Speech Audio Process.2
2012 Phoneme Selective Speech Enhancement Using Parametric Estimators and the Mixture Maximum Model: A Unifying Approach
abstract
This study presents a ROVER speech enhancement algorithm that employs a series of prior enhanced utterances, each customized for a specific broad level phoneme class, to generate a single composite utterance which provides overall improved objective quality across all classes. The noisy utterance is first partitioned into speech and non-speech regions using a voice activity detector, followed by a mixture maximum (MIXMAX) model which is used to make probabilistic decisions in the speech regions to determine phoneme class weights. The prior enhanced utterances are weighted by these decisions and combined to form the final composite utterance. The enhancement system that generates the prior enhanced utterances comprises of a family of parametric gain functions whose parameters are flexible and can be varied to achieve high enhancement levels per phoneme class. These parametric gain functions are derived using 1) a weighted Euclidean distortion cost function, and 2) by modeling clean speech spectral magnitudes or discrete Fourier transform coefficients by Chi or two-sided Gamma priors, respectively. The special case estimators of these gain functions are the generalized spectral subtraction (GSS), minimum mean square error (MMSE), two-sided Gamma or joint maximum a posteriori (MAP) estimators. Performance evaluations performed over two noise types and signal-to-noise ratios (SNRs) ranging from${-}$5 dB to 10 dB suggest that the proposed ROVER algorithm not only outperforms the special case estimators but also the family of parametric estimators when all phoneme classes are jointly considered.
John H. L. Hansen
IEEE Trans. Speech Audio Process.2
2011 UT-Scope: Towards LVCSR under Lombard effect induced by varying types and levels of noisy background
abstract
Adverse environments impact the performance of automatic speech recognition systems in two ways - directly by introducing acoustic mismatch between the speech signal and acoustic models, and indirectly by affecting the way speakers communicate to maintain intelligible communication over noise (Lombard effect). Currently, an increasing number of studies have analyzed Lombard effect with respect to speech production and perception, yet limited attention has been paid to its impact on speech systems, especially within a larger vocabulary context. This study presents a large vocabulary speech material captured in the recently acquired portion of UT-Scope database, produced in several types and levels of simulated background noise (highway, crowd, pink). The impact of noisy background variations on speech parameters is studied together with the effects on automatic speech recognition. Front-end cepstral normalization utilizing a modified RASTA filter is proposed and shown to improve recognition performance in a side-by-side evaluation with several common and state-of-the-art normalization algorithms.
Hynek Boril, John H. L. Hansen
ICASSP2
2011 Phoneme selective speech enhancement using the generalized parametric spectral subtraction estimator
abstract
In this study, the generalized parametric spectral subtraction estimator is employed in the context of a ROVER speech enhancement framework to develop a robust phoneme class selective enhancement algorithm. The parametric estimator is derived by a) optimizing the weighted Euclidean distortion cost function and b) by modeling clean speech spectral magnitudes as Rayleigh distributed priors. A set of enhanced utterances are generated from a single noisy utterance by tuning the parameters of the parametric estimator for different phoneme classes. The speech and non-speech segments are segregated using a voice activity detector. Thereafter, the mixture maximum model is used to make soft decisions on these segments to determine their phoneme class weights. The segments from the enhanced utterances are weighted by these decisions and combined to form the final composite utterance. Using segmental SNR and Itakura-Saito metrics over two noise types and four SNR levels, it was demonstrated that the composite utterance exhibited better phoneme class improvement than the individual utterances enhanced from the parametric estimator.
John H. L. Hansen
ICASSP2
2011 Language identification for singing
abstract
In spoken language processing, considerable research has been accomplished on language identification. Singing language identification is an important yet challenging area that has attracted only a few researchers in music processing. As one information source that can be extracted from music, the language of vocal music is useful for song classification, recognition, and retrieval based on the singing language, specially in unlabeled or mislabeled music collections. In addition, consideration of singing as a speaking style introduces new challenges to existing language identification systems. Our objective in this paper, as one of the first attempts for singing language identification, is to evaluate successful LID systems, specifically PPRLM with singing speech. Further more, we propose a prosodic approach based on pitch contour approximation and compare the results to PPRLM system. Language identification performance for singing and read speech are compared in both systems. Finally, we combine the PPRLM and prosodic systems which achieves an average performance improvement of 4.7% for singing, and 8.7% for read speech compared to the baseline PPRLM system. Our evaluations are based on a multilingual singing corpus that we have collected for this study.
Mahnoosh Mehrabani, John H. L. Hansen
ICASSP2
2011 Hilbert envelope based features for robust speaker identification under reverberant mismatched conditions
abstract
It is well known that MFCC based speaker identification (SID) systems easily break down under mismatched training and test conditions. One such mismatch occurs when a SID system is trained on anechoic speech data, while test is carried out using reverberant data collected via a distant microphone. In this study, a new set of feature parameters based on the Hilbert envelope of Gammatone filterbank outputs is proposed to improve SID performance in the presence of room reverberation. Considering two distinct perceptual effects of reverberation on speech signals, i.e., coloration and long-term reverberation, two different compensation strategies are integrated within the feature extraction framework to effectively suppress the effects of reverberation. Experimental evaluation is performed using speech material from the TIMIT, four different measured room impulse responses (RIR) from Aachen impulse response (AIR) database, and a GMM-based SID system. Obtained results indicate significant improvement over the baseline system with MFCCs plus cepstral mean subtraction (CMS), confirming the effectiveness of the proposed feature parameters for SID under reverberant mismatched conditions.
Seyed Omid Sadjadi, John H. L. Hansen
ICASSP2
2011 Language identification using a combined articulatory prosody framework
abstract
This study presents new advancements in our articulatory-based language identification (LID) system. Our LID system automatically identifies language-features (LFs) from a phonological features (PFs) based representation of speech. While our baseline system uses a static PF-representation for extracting LFs, die new system is based on a dynamic PF representation for feature extraction. Interestingly, the new LFs outperform our baseline system by 11.8% absolute in a difficult 5-way classification task of South Indian Languages. Additionally, we incorporate pitch and energy based features in our new system to leverage prosody in classification. In particular, we employ a Legendre polynomial based contour-estimation to capture shape parameters which are used in classification. Additionally, die fusion of PF and prosody-based LFs further improves die overall classification result by 16.5% absolute over die baseline system Finally, die proposed articulatory language ID system is combined with a PPRLM (parallel phone recognition language model) system to obtain an overall classification accuracy of 86.6%.
Abhijeet Sangwan, Mahnoosh Mehrabani, John H. L. Hansen
ICASSP3
2011 Effective background data selection in SVM speaker recognition for unseen test environment: More is not always better
abstract
This study focuses on determining a procedure to select effective negative examples for development of improved Support Vector Machine (SVM) based speaker recognition. Selection of a background dataset, comprising of a group of negative examples, is critical in development of an effective decision surface between the primary speaker and outside speaker rejection space. Previous studies generally fix the number of examples based on development data for system performance evaluation, while for real applications this does not guarantee sustained performance for unseen data. In the proposed method, the error is estimated on the support vector to select the background dataset, thereby by customizing the back ground dataset for each enrollment speaker instead of training models with a fixed background data. The proposed method finds the equivalent or improved EER and DCF compared with the previous SVM-based studies, and provides consistent performance for unseen data. The method improves the 6% relative improvement on EER and DCF for NIST SRE 2010.
Jun-Won Suh, Yun Lei, Wooil Kim, John H. L. Hansen
ICASSP4
2011 Relative proportionate NLMS: Improving convergence for acoustic channel identification
abstract
It is known that the proportionate normalized least mean square (PNLMS) algorithm outperforms traditional normalized least mean square(NLMS) algorithm, in terms of fast initial convergence rate. However, the PNLMS has been widely observed to not be optimal. This study presents a new perspective into the "proportionate" gain (step-size) allocation scheme. A relative proportionate scheme is established and shows better performance than the original absolute proportionate scheme. Although the correspondingly derived relative proportional LMS (R-PNLMS) algorithm is similar to PNLMS, it differs greatly in terms of conception and convergence behavior. Simulation results for the problem of acoustic channels identification, show improved performance over existing methods.
John H. L. Hansen
ICASSP2
2011 Front-End Compensation Methods for LVCSR Under Lombard Effect
abstract
This study analyzes the impact of noisy background variations and Lombard effect (LE) on large vocabulary continuous speech recognition (LVCSR). Robustness of several front-end feature extraction strategies combined with state-of-the-art feature distribution normalizations is tested on neutral and Lombard speech from the UT-Scope database presented in two types of background noise at various levels of SNR. An extension of a bottleneck (BN) front-end utilizing normalization of both critical band energies (CRBE) and BN outputs is proposed and shown to provide a competitive performance compared to the best MFCC-based system. A novel MFCC-based BN front-end is introduced and shown to outperform all other systems in all conditions considered (average 4.1% absolute WER reduction over the second best system). Additionally, two phenomena are observed: (i) combination of cepstral mean subtraction and recently established RASTALP filtering significantly reduces transient effects of RASTA band-pass filtering and increases ASR robustness to noise and LE; (ii) histogram equalization may benefit from utilizing reference distributions derived from pre-normalized rather than raw training features, and also from adopting distributions from different front-ends. Index Terms: speech recognition, Lombard effect, UT-Scope database, bottleneck features, quantile-based cepstral distribution normalization, histogram equalization.
Hynek Boril, Frantisek Grézl, John H. L. Hansen
INTERSPEECH3
2011 Acoustic Analysis of Whispered Speech for Phoneme and Speaker Dependency
abstract
Whisper is used by speakers in certain circumstances to protect personal information. Due to the differences in production mechanisms between neutral and whispered speech, there are considerable differences between the spectral structure of neutral and whispered speech, such as formant shifts and shifts in spectral slope. This study analyzes the dependency of these differences on speakers and phonemes by applying a Vector Taylor Series (VTS) approximation to a model of the transformation of neutral speech into whispered speech, and estimating the parameters of this model using an Expectation Maximization (EM) algorithm. The results from this study shed light on the speaker and phoneme dependency of the shifts of neutral to whisper speech, and suggest that similarly derived model adaptation or compensation schemes for whisper speech/speaker recognition will be highly speaker dependent. Index Terms: whispered speech, speech analysis
Keith W. Godin, John H. L. Hansen
INTERSPEECH3
2011 Speaker Identification for Whispered Speech Using a Training Feature Transformation from Neutral to Whisper
abstract
A number of research studies in speaker recognition have recently focused on robustness due to microphone and channel mismatch(e.g., NIST SRE). However, changes in vocal effort, especially whispered speech, present significant challenges in maintaining system performance. Due to the mismatch spectral structure resulting from the different production mechanisms, performance of speaker identification systems trained with neutral speech degrades significantly when tested with whispered speech. This study considers a feature transformation method in the training phase that leads to a more robust speaker model for speaker ID with whispered speech. In the proposed system, a Speech Mode Independent (SMI) Universal Background Model (UBM) is built using collected real neutral features and pseudo whispered features generated with Vector Taylor Series (VTS), or via Constrained Maximum Likelihood Linear Regression (CMLLR) model adaptation. Text-independent closed set speaker ID results using the UT-VocalEffort II corpus show an accuracy of 88.87% using the proposed method, which represents a relative improvement of 46.26% compared with the 79.29% accuracy of the baseline system. This result confirms a viable approach to improving speaker ID performance for neutral and whispered speech mismatched conditions. Index Terms: whispered speech, speech identification
John H. L. Hansen
INTERSPEECH2
2011 Vowel Context and Speaker Interactions Influencing Glottal Open Quotient and Formant Frequency Shifts in Physical Task Stress
abstract
Physical task stress is known to affect the fundamental frequency of speech. This study of two American English vowels /IY/ and /AH/ investigates whether physical task stress affects the center frequencies of formants F1 and F2, and whether it affects the glottal open quotient, and whether these effects are different for different speakers, the different vowels, and two different vowel contexts. Formant center frequencies are measured from the acousticwaveform, and the glottal openquotient is measured from the electroglottograph signal. The study finds in general that the production of vowels is affected by physical task stress. In particular, the study finds that F1, F2, and the glottal open quotient are affected by physicaltask stress. It also finds that the effects of stress on F1 vary for different speakers, and that the effects of stress on the glottal open quotient vary for different combinations of speakerand vowel. Index Terms: physical task stress, open quotient, electroglottograph
Keith W. Godin, John H. L. Hansen
INTERSPEECH2
2011 Robust Speaker Recognition in Non-Stationary Room Environments Based on Empirical Mode Decomposition
abstract
In this study, we consider the problem of speaker recognition in a non-stationary room/channel mismatched condition. In such circumstances, cepstral coefficients are affected in a way that the short-term stationarity assumption, on which conventional feature normalization methods are based on, may not be valid. We observe that the empirical mode decomposition (EMD) applied to the cepstral feature stream can partially separate out the non-stationary channel components, if present, into its residual signal and other lower order intrinsic mode functions (IMFs), which leads us to develop a filtering scheme based on this decomposition. The proposed method works in the time domain making use of the instantaneous frequency function obtained through Hilbert spectral analysis of the IMFs. Experimental evaluations on the TIMIT database with added non-stationary room channels in test demonstrate the superiority of the proposed scheme compared to conventional feature normalization schemes. Additional experiments performed on the newly released noisy robust open set speaker identification (ROSSI) and NIST SRE corpora also confirm the effectiveness of the proposed method in stationary room/channel mismatched conditions. Index Terms: Speaker verification, non-stationary room channel, empirical mode decomposition
Taufiq Hasan, John H. L. Hansen
INTERSPEECH2
2011 Feature Compensation for Speech Recognition in Severely Adverse Environments Due to Background Noise and Channel Distortion
abstract
This paper proposes an effective feature compensation scheme to address severely adverse environments for robust speech recognition, where background noise and channel distortion are simultaneously involved. An iterative channel estimation method is integrated into the framework of our Parallel Combined Gaussian Mixture Model (PCGMM) based feature compensation algorithm [1]. A new speech corpus is generated which reflects both additive and convolutional noise corruption. The channel distortion effects are obtained from the NTIMIT and CTIMIT corpora. Evaluation of objective speech quality measures including STNR, PESQ, and speech recognition shows that the generated speech corpus represents highly challenging acoustic conditions for speech recognition. Performance evaluation of the proposed system over the obtained speech corpus demonstrates that the proposed feature compensation scheme is significantly effective at improving speech recognition performance with presence of both background noise and channel distortion, comparing to the conventional methods including the ETSI AFE. Index Terms: channel estimation, feature compensation, corpus generation, PCGMM, robust speech recognition.
Wooil Kim, John H. L. Hansen
INTERSPEECH2
2011 Detecting Sleepiness by Fusing Classifiers Trained with Novel Acoustic Features
abstract
Automatic sleepiness detection is a challenging task that can lead to advances in various domains including traffic safety, medicine and human-machine interaction. This paper analyzes the discriminative power of different acoustic features to detect sleepiness. The study uses the sleepy language corpus (SLC). Along with standard acoustic features, novel features are proposed including functionals across voiced segment statistics in the F0 contour, likelihoods of reference models used to contrast non-neutral speech, and a set of robust to noise spectral features. These feature sets, which have performed well in other paralinguistic tasks such as emotion recognition, are used to train classifiers that are combined at the feature and decision levels. The best unweighted accuracy (UA) is obtained by combining the classifiers at the decision level under a maximum likelihood framework (UA = 70.97%). This performance is higher than the best results reported in the corpus. Index Terms: Speaker State Recognition, Paralinguistics, Affective Computing, Sleepiness
Tauhidur Rahman, Soroosh Mariooryad, Shalini Keshavamurthy, Gang Liu 0001, John H. L. Hansen, Carlos Busso
INTERSPEECH5
2011 Phone Impact Based Speech Transmission Technique for Reliable Speech Recognition in Poor Wireless Network Conditions
abstract
This paper presents a preliminary study on an effective differentiable network service technique to achieve improved speech recognition under severely poor wireless channel conditions, by leveraging multiple priority levels applied to speech classes. Each speech class is assigned a different priority level based on its level of impact on speech recognition performance. Based on their priority level, frames of each speech class are given distinct levels of network quality of service (QoS) to satisfy the delay requirement and enable speech recognition at the receiver. This proposed Phone Impact (PI) based priority class is compared to the Voiced/Unvoiced (VU) based priority class in this study. The experimental results prove that the proposed scheme is effective at providing wireless network service for robust speech recognition under poor channel conditions, showing up to 2.67 dB and 5.93 dB lower Signal to Noise Ratio (SNR) operating regions compared to the VU based and plain protocols respectively. The PI based method also shows acceptable WERs at lower SNRs where VU and plain systems significantly degrade in speech recognition performance in case of retry limit of 6. Index Terms: Phone Impact, Priority Class, Speech Recognition, IEEE 802.11, Differentiated Maximum Retry Limit.
Azar Taufique, Kumaran Vijayasankar, Wooil Kim, John H. L. Hansen, Marco Tacca, Andrea Fumagalli
INTERSPEECH4
2011 Using Human Perception for Automatic Accent Assessment
abstract
In this study, a new algorithm for automatic accent evaluation of native and non-native speakers is presented. The proposed system consists of two main steps: alignment and scoring. At the alignment step, the speech utterance is processed using a Weighted Finite State Transducer (WFST) based technique to automatically estimate the pronunciation errors. Subsequently, in the scoring step a Maximum Entropy (ME) based technique is employed to assign perceptually motivated scores to pronunciation errors. The combination of the two steps yields an approach that measures accent based on perceptual impact of pronunciation errors, and is termed as the Perceptual WFST (P-WFST). The P-WFST is evaluated on American English (AE) spoken by native and non-native (native speakers of Mandarin-Chinese) speakers from the CUAccent corpus. The proposed P-WFST algorithm shows higher and more consistent correlation with human evaluated accent scores, when compared to the Goodness Of Pronunciation (GOP) algorithm. Index Terms: automatic accent assessment, pronunciation scoring, Finite State Transducers, Maximum Entropy
Freddy William, Abhijeet Sangwan, John H. L. Hansen
INTERSPEECH3
2011 Frame-Level Vocal Effort Likelihood Space Modeling for Improved Whisper-Island Detection
abstract
In this study, a frame-based vocal effort likelihood space modeling framework for improved whisper-island detection within normally phonated audio streams is proposed. The proposed method is based on first training a traditional Gaussian mixture model for whisper and neutral speech, which is then employed to extract a newly proposed discriminative feature set entitled Vocal Effort Likelihood (VEL), for whisper-island detection. The VEL feature set is integrated within a BIC/T 2 -BIC segmentation scheme for vocal effort change point(VECP) detection. With the dimension-reduced VEL 2-D feature set, the proposed framework has reduced computational costs versus prior method [1]. Experimental results using the UT-VocalEffort II corpus for whisper-island detection using the proposed framework are presented and compared with a previous algorithm introduced in [1]. The proposed algorithm is shown to improve performance in VECP detection with the lowest MultiError Score(MES) of 6.33. Furthermore, very accurate whisperisland detection was obtained using proposed algorithm, which is useful for sustained performance in speech systems (ASR, Speaker-ID, etc.)which might experience whisper speech. Finally, experimental performance achieves a 100% detection rate for the proposed algorithm, which represents the best whisperisland detection performance with lowest computational costs available in the literature to date. Index Terms: Vocal Effort Likelihood, Vocal Effort, WhisperIsland Detection, GMM Classifier
John H. L. Hansen
INTERSPEECH2
2011 Variational noise model composition through model perturbation for robust speech recognition with time-varying background noise
Wooil Kim, John H. L. Hansen
Speech Commun.2
2011 Mismatch modeling and compensation for robust speaker verification
Yun Lei, John H. L. Hansen
Speech Commun.2
2011 Speaker Identification Within Whispered Speech Audio Streams
abstract
Whisper is an alternative speech production mode used by subjects in natural conversation to protect the privacy. Due to the profound differences between whisper and neutral speech in both excitation and vocal tract function, the performance of speaker identification systems trained with neutral speech degrades significantly. In this paper, a seamless neutral/whisper mismatched closed-set speaker recognition system is developed. First, performance characteristics of a neutral trained closed-set speaker ID system based on an Mel-frequency cepstral coefficient-Gaussian mixture model (MFCC-GMM) framework is considered. It is observed that for whisper speaker recognition, performance degradation is concentrated for only a subset of speakers. Next, it is shown that the performance loss for speaker identification in neutral/whisper mismatched conditions is focused on phonemes other than low-energy unvoiced consonants. In order to increase system performance for unvoiced consonants, an alternative feature extraction algorithm based on linear and exponential frequency scales is applied. The acoustic properties of misrecognized and correctly recognized whisper are analyzed in order to develop more effective processing schemes. A two-dimensional feature space is proposed in order to predict on which whispered utterances the system will perform poorly, with evaluations conducted to measure the quality of whispered speech. Finally, a system for seamless neutral/whisper speaker identification is proposed, resulting in an absolute improvement of 8.85%-10.30% for speaker recognition, with the best closed set speaker ID performance of 88.35% obtained for a total of 961 read whisper test utterances, and 83.84% using a total of 495 spontaneous whisper test utterances.
John H. L. Hansen
IEEE Trans. Speech Audio Process.2
2011 A Study on Universal Background Model Training in Speaker Verification
abstract
State-of-the-art Gaussian mixture model (GMM)-based speaker recognition/verification systems utilize a universal background model (UBM), which typically requires extensive resources, especially if multiple channel and microphone categories are considered. In this study, a systematic analysis of speaker verification system performance is considered for which the UBM data is selected and purposefully altered in different ways, including variation in the amount of data, sub-sampling structure of the feature frames, and variation in the number of speakers. An objective measure is formulated from the UBM covariance matrix which is found to be highly correlated with system performance when the data amount was varied while keeping the UBM data set constant, and increasing the number of UBM speakers while keeping the data amount constant. The advantages of feature sub-sampling for improving UBM training speed is also discussed, and a novel and effective phonetic distance-based frame selection method is developed. The sub-sampling methods presented are shown to retain baseline equal error rate (EER) system performance using only 1% of the original UBM data, resulting in a drastic reduction in UBM training computation time. This, in theory, dispels the myth of “There's no data like more data” for the purpose of UBM construction. With respect to the UBM speakers, the effect of systematically controlling the number of training (UBM) speakers versus overall system performance is analyzed. It is shown experimentally that increasing the inter-speaker variability in the UBM data while maintaining the overall total data size constant gradually improves system performance. Finally, two alternative speaker selection methods based on different speaker diversity measures are presented. Using the proposed schemes, it is shown that by selecting a diverse set of UBM speakers, the baseline system performance can be retained using less than 30% of the original UBM speakers.
Taufiq Hasan, John H. L. Hansen
IEEE Trans. Speech Audio Process.2
2011 A Novel Mask Estimation Method Employing Posterior-Based Representative Mean Estimate for Missing-Feature Speech Recognition
abstract
This paper proposes a novel mask estimation method for missing-feature reconstruction to improve speech recognition performance in various types of background noise conditions. A conventional mask estimation method based on spectral subtraction degrades performance, due to incorrect estimation of the noise signal which fails to accurately represent the variations of background noise during the incoming speech utterance. The proposed mask estimation method utilizes a Posterior-based Representative Mean (PRM) estimate for determining the reliability of the input speech spectral components, which is obtained as a weighted sum of the mean parameters of the speech model using the posterior probability. To obtain the noise-corrupted speech model, a model combination method is employed, which was proposed in our previous study for a feature compensation method. Experimental results demonstrate that the proposed mask estimation method provides more separable distributions for the reliable/unreliable component classifier compared to the conventional mask estimation method. The recognition performance is evaluated using the Aurora 2.0 framework over various types of background noise conditions and the CU-Move real-life in-vehicle corpus. The performance evaluation shows that the proposed mask estimation method is considerably more effective at increasing speech recognition performance in various types of background noise conditions, compared to the conventional mask estimation method which is based on spectral subtraction. By employing the proposed PRM-based mask estimation for missing-feature reconstruction, we obtain +23.41% and +9.45% average relative improvements in word error rate for all four types of noise conditions and CU-Move corpus, respectively, compared to conventional mask estimation methods.
Wooil Kim, John H. L. Hansen
IEEE Trans. Speech Audio Process.2
2011 Dialect Classification via Text-Independent Training and Testing for Arabic, Spanish, and Chinese
abstract
Automatic dialect classification has emerged as an important area in the speech research field. Effective dialect classification is useful in developing robust speech systems, such as speech recognition and speaker identification. In this paper, two novel algorithms are proposed to improve dialect classification for text-independent spontaneous speech in Arabic and Spanish languages, along with probe results for Chinese. The problem considers the case where no transcripts but dialect labels are available for training and test data, and speakers are speaking spontaneously, which is defined as text-independent dialect classification. The Gaussian mixture model (GMM) is used as the baseline system for text-independent dialect classification. The major motivation is to suppress confused/distractive regions from the dialect language space and emphasize discriminative/sensitive information of the available dialects. In the training phase, a symmetric version of the Kullback-Leibler divergence is used to find the most discriminative GMM mixtures (KLD-GMM), where the confused acoustic GMM region is suppressed. For testing, the more discriminative frames are detected and used via the location of where the frames are in the GMM mixture feature space, which is termed frame selection decoding (FSD-GMM). The first KLD-GMM and second FSD-GMM techniques, are shown to improve dialect classification performance for three-way dialect tasks. The two algorithms and their combination are evaluated on dialects of Arabic and Spanish corpora. Measurable improvement is achieved in both two cases, over a generalized maximum-likelihood estimation GMM baseline (MLE-GMM).
Yun Lei, John H. L. Hansen
IEEE Trans. Speech Audio Process.2
2011 Whisper-Island Detection Based on Unsupervised Segmentation With Entropy-Based Speech Feature Processing
abstract
Whisper island detection is a challenging research problem which has received little attention in the research community. Effective whisper-island detection is the first step necessary to ensure engagement of effective subsequent speech processing steps to address mismatch between whisper and neutral speech production. In this paper, we propose an effective approach for detecting whisper-islands embedded within normally phonated speech via BIC/T2-BIC using a proposed 4-D feature set. Performance is assessed using our proposed multi-error score (MES), which shows that the new proposed algorithm achieves the lowest MES (11.51) to date and along with a perfect 100% correct whisper/neutral vocal effort labeling. The results show that we can correctly and precisely detect vocal effort change points (VECP) between whisper-islands and neutral speech as well as label the vocal effort of the whisper-island. The proposed feature is sensitive to the vocal effort change between whisper and neutral speech and is gender independent. The result suggests that the proposed algorithm is effective and precise for the whisper-island detection.
John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.2
2011 International Large-Scale Vehicle Corpora for Research on Driver Behavior on the Road
abstract
This paper considers a comprehensive and collaborative project to collect large amounts of driving data on the road for use in a wide range of areas of vehicle-related research centered on driving behavior. Unlike previous data collection efforts, the corpora collected here contain both human and vehicle sensor data, together with rich and continuous transcriptions. While most efforts on in-vehicle research are generally focused within individual countries, this effort links a collaborative team from three diverse regions (i.e., Asia, American, and Europe). Details relating to the data collection paradigm, such as sensors, driver information, routes, and transcription protocols, are discussed, and a preliminary analysis of the data across the three data collection sites from the U.S. (Dallas), Japan (Nagoya), and Turkey (Istanbul) is provided. The usability of the corpora has been experimentally verified with a Cohen's kappa coefficient of 0.74 for transcription reliability, as well as being successfully exploited for several in-vehicle applications. Most importantly, the corpora are publicly available for research use and represent one of the first multination efforts to share resources and understand driver characteristics. Future work on distributing the corpora to the wider research community is also discussed.
Kazuya Takeda, John H. L. Hansen, Pinar Boyraz Baykas, Lucas Malta, Chiyomi Miyajima, Hüseyin Abut
IEEE Trans. Intell. Transp. Syst.2
2010 Limited resource speech recognition for Nigerian English
abstract
In this study, we introduce the UISpeech corpus which consists of Nigerian-Accented English audio-visual data. The corpus captures the linguistic diversity of Nigeria with data collected from native-speakers of Yoruba, Hausa, Igbo, Tiv, Funali and others. The UIS-peech corpus comprises isolated word recordings and read speech utterances. The new corpus is intended to provide a unique opportunity to apply and expand speech processing techniques to a limited resource language. Acoustic-phonetic differences between American English (AE) and Nigerian English (NE) are studied in terms of pronunciation variations, vowel locations in the formant space, and distances between AE-trained acoustic models and models adapted to NE. A strong impact of the AE-NE acoustic mismatch on automatic speech recognition (ASR) is observed. A combination of model adaptation and extension of AE lexicon for newly established NE pronunciation variants is shown to substantially improve performance of the AE-trained ASR system in the new NE task. This study represents the first step towards incorporating speech technology in Nigerian English.
Sulyman Amuda, Hynek Boril, Abhijeet Sangwan, John H. L. Hansen
ICASSP4
2010 Broad phoneme class based speech enhancement using mixture maximum model
abstract
This study develops a speech enhancement technique that uses a series of prior enhanced speech utterances, each optimized for a specific broad phoneme class, to generate a single, composite utterance to improve objective quality scores over all phoneme classes. The noisy utterance is partitioned into phoneme class segments using probabilistic decisions made from the mixture maximum model (MIXMAX). Based on these phoneme class decisions, the composite segment is constructed using a combination of the prior enhanced utterances. The enhancement system that generates multiple enhanced utterances is assumed to belong to the class of short-time spectral magnitude estimators which either minimizes the weighted Euclidean distortion (WED) between clean speech and clean speech estimate spectral magnitudes or which finds the joint MAP(JMAP) estimate of clean speech spectral magnitude and phase. Performance evaluations of the composite utterance exhibit better performance than the individual utterances over all phoneme classes in most cases of the noise types and SNR levels considered.
John H. L. Hansen
ICASSP2
2010 Acoustic analysis for speaker identification of whispered speech
abstract
Whisper is an alternative speech production mode from neutral speech, which is used by talkers intentionally in natural conversational scenarios to protect personal privacy and avoid being overheard. Due to differences between whispered and neutral speech in vocal excitation and vocal tract function, the performance of speaker ID systems trained with neutral speech degrades significantly. In this study, a neutral trained closed-set speaker ID task based on MFCC-GMM is considered. It is observed that for whisper speaker recognition, the degradation is concentrated for a certain number of speakers. Next, an acoustic analysis is conducted in order to determine the reason affecting the degradation for those speakers. Finally, a confidence space is proposed to measure the quality of whispered speech for the task of speaker ID. Experimental evaluations demonstrate the effectiveness of this method in searching whispered utterances with poor speaker information for a neutral/whisper mismatch speaker ID system. The proposed method makes it possible to compensate for those poor utterances, meanwhile avoiding any harm to other utterances that remain the performance of neutral speaker ID task.
John H. L. Hansen
ICASSP2
2010 A novel feature sub-sampling method for efficient universal background model training in speaker verification
abstract
Speaker recognition/verification systems require an extensive universal background model (UBM), which typically requires extensive resources, especially if new channel domains are considered. In this study we propose an effective and computationally efficient algorithm for training the UBM for speaker verification. A novel method based on Euclidean distance between features is developed for effective sub-sampling of potential training feature vectors. Using only about 1.5 seconds of data from each development utterance, the proposed UBM training method drastically reduces the computation time, while improving, or at least retaining original speaker verification system performance. While methods such as factor analysis can mitigate some of the issues associated with channel/microphone/environmental mismatch, the proposed rapid UBM training scheme offers a viable alternative for rapid environment dependent UBMs.
Taufiq Hasan, Yun Lei, Aravind Chandrasekaran, John H. L. Hansen
ICASSP4
2010 Angry emotion detection from real-life conversational speech by leveraging content structure
abstract
This study proposes an effective angry speech detection approach by leveraging content structure within the input speech. A classifier based on an “emotional” language model score is formulated and combined with acoustic feature based classifiers including TEO-based feature and conventional Mel frequency cepstral coefficients (MFCC). The proposed detection algorithm is evaluated on real-life conversational speech which was recorded between customers and call center operators over a telephone network. Analysis on the conversational speech corpus presents a distinctive property between neutral and angry speech in word distribution and frequently occurring words. An improvement of up to 6.23% in Equal Error Rate (EER) is obtained by combining the TEO-based and MFCC features, and emotional language model score based classifiers.
Wooil Kim, John H. L. Hansen
ICASSP2
2010 A kernel mean matching approach for environment mismatch compensation in speech recognition
abstract
The mismatch between training and test environmental conditions presents a challenge to speech recognition systems. In this paper, we investigate an approach for matching the distributions of training and test data in the feature space. This approach uses the property of reproducing kernel Hilbert space (RKHS) with a universal kernel for the task of distribution matching. The approach is unsupervised, requiring no transcripts of data for compensation, and can be employed either with explicit adaptation data or with live test data. The approach is evaluated on two real car environments - CU-Move and UTDrive. Relative improvements of between 10-25% are obtained for different experimental setups.
John H. L. Hansen
ICASSP2
2010 Dialect distance assessment method based on comparison of pitch pattern statistical models
abstract
Dialect variations of a language have a severe impact on the performance of speech systems. Therefore, knowing how close or diverse dialects are in a given language space provides useful information to predict, or improve, system performance when there is a mismatch between train and test data. Distance measures have been used in several applications of speech processing. However, apart from phonetic measures, little if any work has been done on dialect distance measurement. This study explores differences in pitch movement microstructure among dialects. A method of dialect distance assessment based on pitch patterns modeled progressively from pitch contour primitives is proposed. The presented method does not require any manual labeling and is text-independent. The KL divergence is employed to compare the resulting statistical models. The proposed scheme is evaluated on a corpus of Arabic dialects, and shown to be consistent with the results from the spectral-based dialect classification system. Finally, it is also shown using a perceptive evaluation that the proposed objective approach correlates well with subjective distances.
Mahnoosh Mehrabani, Hynek Boril, John H. L. Hansen
ICASSP3
2010 Speech under physical stress: A production-based framework
abstract
This paper examines the impact of physical stress on speech. The methodology adopted here identifies inter-utterance breathing (IUB) patterns as a key intermediate variable while studying the relationship between physical stress and speech. Additionally, this work connects high-level prosodic changes in the speech signal (energy, pitch, and duration) to the corresponding breathing patterns. Our results demonstrate the diversity of breathing and articulation patterns that speakers employ in order to compensate for the increased body oxygen demand. Here, we identify the normalized value of breathing energy rate (proportional to minute volume) acquired from a conventional as well as physiological microphone as a reliable and accurate estimator of physical stress. Additionally, we also show that the prosodic patterns (pitch, energy, and duration) of high-level speech structure shows good correlation with the normalized-breathing energy rate. In this manner, the study establishes the interconnection between temporal speech structure and physical stress through breathing.
Sanjay A. Patil, Abhijeet Sangwan, John H. L. Hansen
ICASSP3
2010 Towards more intelligible physiological microphone speech: A probabilistic transformation approach
abstract
The non-acoustic physiological microphone (PMIC) has been shown to be useful for speech systems under adverse noisy conditions. However, the signal is not a true speech for the listener, therefore appears muffled and metallic with variations to the speaker dependent structure. This study presents a probabilistic transformation approach to improve the perceptual quality and intelligibility of PMIC speech not only by mapping the non-acoustic signal into the conventional speech production space, but also by minimizing distortions arising from alternative pickup location. Performance of the proposed approach is assessed based on five distinct objective metrics. Obtained results indicate that incorporating the probabilistic transformation yields significant improvement in overall PMIC speech quality and intelligibility. This technique along with the PMIC can thus find applications in noise robust human-to-human speech communication.
Seyed Omid Sadjadi, Sanjay A. Patil, John H. L. Hansen
ICASSP3
2010 Automatic language analysis and identification based on speech production knowledge
abstract
In this paper, a language analysis and classification system that leverages knowledge of speech production is proposed. The proposed scheme automatically extracts key production traits (or “hot-spots”) that are strongly tied to the underlying language structure. Particularly, the speech utterance is first parsed into consonant and vowel clusters. Subsequently, the production traits for each cluster is represented by the corresponding temporal evolution of speech articulatory states. It is hypothesized that a selection of these production traits are strongly tied to the underlying language, and can be exploited for language ID. The new scheme is evaluated on our South Indian Languages (SInL) corpus which consists of 5 closely related languages spoken in India, namely, Kannada, Tamil, Telegu, Malayalam, and Marathi. Good accuracy is achieved with a rate of 65% obtained in a difficult 5-way classification task with about 4sec of train and test speech data per utterance. Furthermore, the proposed scheme is also able to automatically identify key production traits of each language (e.g., dominant vowels, stop-consonants, fricatives etc.).
Abhijeet Sangwan, Mahnoosh Mehrabani, John H. L. Hansen
ICASSP3
2010 An efficient microphone array based voice activity detector for driver's speech in noise and music rich in-vehicle environments
abstract
In this study, an efficient microphone array based voice activity detector (VAD) is proposed, specifically for the driver's speech in a vehicle. Two pre-fixed beampatterns are designed to explore the knowledge of the in-vehicle spatial acoustic power distributions, based on which a simple voice active/inacitve classifier is also designed. Through real data experiment, the proposed VAD presents a novel and robust performance against various in-vehicle noisy scenarios.
John H. L. Hansen
ICASSP2
2010 Advancements in whisper-island detection using the linear predictive residual
abstract
In this study, we consider the use of a new entropy-based feature extracted from linear predictive residual for whisper-island detection within normally phonated audio streams. The proposed feature, which is sensitive to vocal effort changes between whisper and neutral speech, is integrated within a BIC/T2-BIC segmentation for vocal effort change point(VECP) detection and utilized for vocal effort classification. Evaluation is based on the proposed multi-error score(MES), where the improved feature is shown to improve performance in VECP detection with the lowest MES of 20.70. Furthermore, more accurate whisper-island detection was obtained using the proposed feature and algorithm. Finally, the experimental detection rate results of 97.37% represents the best whisper-island detection performance available in the literature to date.
John H. L. Hansen
ICASSP2
2010 Automatic excitement-level detection for sports highlights generation
abstract
The problem of automatic excitement detection in baseball videos is considered and applied for highlight generation. This paper focuses on detecting exciting events in video using complementary information from the audio and video domains. First, a new measure for non-stationarity which is extremely effective in separating background from speech is proposed. This new feature is employed in an unsupervised GMM-based segmentation algorithm that identifies the sports commentators speech within the crowd background. Thereafter, the “level-of-excitement” is measured using features such as pitch, F1‐F3 center frequencies, and spectral center of gravity extracted from the commentators speech. Our experiments using actual baseball videos show that these features are well correlated with human assessment of excitability. Furthermore, slow-motion replay and baseball pitching-scenes from the video are also detected to estimate scene end-points. Finally, audio/video information is fused to rank-order scenes by “excitability” in order to generate highlights of user-defined timelengths. The techniques described in this paper are generic and applicable to a variety of topic and video/acoustic domains. Index Terms: Video Segmentation, Multimodal Signal Processing
Hynek Boril, Abhijeet Sangwan, Taufiq Hasan, John H. L. Hansen
INTERSPEECH4
2010 Analysis and detection of cognitive load and frustration in drivers' speech
abstract
Non-driving related cognitive load and variations of emotional state may impact a driver’s capability to control a vehicle and introduces driving errors. Availability of reliable cognitive load and emotion detection in drivers would benefit the design of active safety systems and other intelligent in-vehicle interfaces. In this study, speech produced by 68 subjects while driving in urban areas is analyzed. A particular focus is on speech production differences in two secondary cognitive tasks, interactions with a co-driver and calls to automated spoken dialog systems (SDS), and two emotional states during the SDS interactions - neutral/negative. A number of speech parameters are found to vary across the cognitive/emotion classes. Suitability of selected cepstral- and production-based features for automatic cognitive task/emotion classification is investigated. A fusion of GMM/SVM classifiers yields an accuracy of 94.3% in cognitive task and 81.3% in emotion classification.
Hynek Boril, Seyed Omid Sadjadi, Tristan Kleinschmidt, John H. L. Hansen
INTERSPEECH4
2010 Session variability contrasts in the MARP corpus
abstract
Intra-session and inter-session variability in the Multi-session Audio Research Project (MARP) corpus are contrasted in two experiments that exploit the long-term nature of the corpus. In the first experiment, Gaussian Mixture Models (GMMs) model 30-second session chunks, clustering chunks using the Kullback-Leibler (KL) divergence. Cross-session relationships are found to dominate the clusters. Secondly, session detection with 3 variations in training subsets is performed. Results showed that small changes in long-term characteristics are observed throughout the sessions. These results enhance understanding of the relationship between long-term and short-term variability in speech and will find application in speaker and speech recognition systems. Index Terms: speaker identification, session variability
Keith W. Godin, John H. L. Hansen
INTERSPEECH2
2010 An effective feature compensation scheme tightly matched with speech recognizer employing SVM-based GMM generation
Wooil Kim, Jun-Won Suh, John H. L. Hansen
INTERSPEECH3
2010 Speaker recognition using supervised probabilistic principal component analysis
abstract
In this study, a supervised probabilistic principal component analysis (SPPCA) model is proposed in order to integrate the speaker label information into a factor analysis approach using the well-known probabilistic principal component analysis (PPCA) model under a support vector machine (SVM) framework. The latent factor from the proposed model is believed to be more discriminative than one from the PPCA model. The proposed model, combined with different types of intersession compensation techniques in the back-end, is evaluated using the National Institute of Standards and Technology (NIST) Speaker Recognition Evaluation (SRE) 2008 data corpus, along with a comparison to the PPCA model. Index Terms: speaker recognition, factor analysis, supervised modeling
Yun Lei, John H. L. Hansen
INTERSPEECH2
2010 A novel feature extraction strategy for multi-stream robust emotion identification
abstract
We investigate an effective feature extraction front-end for speech emotion recognition, which performs well in clean and noisy conditions. First, we explore the use of perceptual minimum variance distortionless response (PMVDR). These features, originally proposed for accent/dialect and language identification (LID), can better approximate the perceptual scales and are less sensitive to noise and speaker variation. Also developed for LID, shifted delta cepstral (SDC) approach can be used to incorporate additional temporal information. It is known that supra-segmental speech characteristics, such as pitch and intensity, provide better discriminative information for emotion recognition by fusing with other emotion dependent features. Combined PMVDR and SDC together, the system outperforms the baseline system (MFCC based) by 10.3% (absolute). Furthermore, we find both PMVDR and SDC offer much better robustness in noisy condition, which is critical for real applications. All the evaluation the proposed features using the Berlin database of emotion speech.
Gang Liu 0001, Yun Lei, John H. L. Hansen
INTERSPEECH3
2010 Assessment of single-channel speech enhancement techniques for speaker identification under mismatched conditions
abstract
It is well known that MFCC based speaker identification (SID) systems easily break down under mismatched training and test conditions. In this paper, we report on a study that considers four different single-channel speech enhancement front-ends for robust SID under such conditions. Speech files from the YOHO database are corrupted with four types of noise including babble, car, factory, and white Gaussian at five SNR levels (0–20 dB), and processed using four speech enhancement techniques representing distinct classes of algorithms: spectral subtraction, statistical model-based, subspace, and Wiener filtering. Both processed and unprocessed files are submitted to a SID system trained on clean data. In addition, a new set of acoustic feature parameters based on Hilbert envelope of gammatone filterbank outputs are proposed and evaluated for SID task. Experimental results indicate that: (i) depending on the noise type and SNR level, the enhancement front-ends may help or hurt SID performance, (ii) the proposed feature significantly achieves higher SID accuracy compared to MFCCs under mismatched conditions.
Seyed Omid Sadjadi, John H. L. Hansen
INTERSPEECH2
2010 Quality conversion of non-acoustic signals for facilitating human-to-human speech communication under harsh acoustic conditions
abstract
Harsh acoustic conditions limit the effectiveness of human speech communication to a great extent. There is a consensus that even at moderate SNR levels, traditional speech enhancement techniques tend to improve the perceptual quality of speech rather than its intelligibility. As an alternative, non-acoustic contact sensors have recently been developed for noise-robust signal capture. Although relatively immune to ambient noise, due to alternative pickup location and non-acoustic principle of operation, signals measured from these sensors are of lower speech quality and intelligibility when compared to those obtained from a conventional microphone in clean conditions. To facilitate human-to-human speech communication under acoustically adverse environments, in this study we present and evaluate a probabilistic transformation framework to improve perceptual quality and intelligibility of signals acquired from one such sensor entitled: physiological microphone (PMIC). Results from both objective and subjective tests confirm that incorporating this framework as a post-processing stage yields significant improvement in overall quality and intelligibility of the PMIC signals.
Seyed Omid Sadjadi, Sanjay A. Patil, John H. L. Hansen
INTERSPEECH3
2010 A Bayesian approach to voice activity detection using multiple statistical models and discriminative training
John H. L. Hansen
INTERSPEECH2
2010 Driver adaptive and context aware active safety systems using CAN-bus signals
abstract
Increasing stress levels in drivers, along with their ability to multi task with infotainment systems cause the drivers to deviate their attention from the primary task of driving. With the rapid advancements in technology, along with the development of infotainment systems, much emphasis is being given to occupant safety. Modern vehicles are equipped with many sensors and ECUs (Embedded Control Units) and CAN-bus (Controller Area Network) plays a significant role in handling the entire communication between the sensors, ECUs and actuators. Most of the mechanical links are replaced by intelligent processing units (ECU) which take in signals from the sensors and provide measurements for proper functioning of engine and vehicle functionalities along with several active safety systems such as ABS (Anti-lock Brake System) and ESP (Electronic Stability program). Current active safety systems utilize the vehicle dynamics (using signals on CAN-bus) but are unaware of context and driver status, and do not adapt to the changing mental and physical conditions of the driver. The traditional engine and active safety systems use a very small time window (t<;2sec) of the CAN-bus to operate. On the contrary, the implementation of driver adaptive and context aware systems require longer time windows and different methods for analysis. The long-term history and trends in the CAN-bus signals contain important information on driving patterns and driver characteristics. In this paper, a summary of systems that can be built on this type of analysis is presented. The CAN-bus signals are acquired and analyzed to recognize driving sub-tasks, maneuvers and routes. Driver inattention is assessed and an overall system which acquires, analyses and warns the driver in real-time while the driver is driving the car is presented showing that an optimal human-machine cooperative system can be designed to achieve improved overall safety.
Amardeep Sathyanarayana, Pinar Boyraz Baykas, Zelam Purohit, Rosarita Lubag, John H. L. Hansen
Intelligent Vehicles Symposium5
2010 The "UTDrive" in-vehicle voice activity detection system
abstract
In this study, we specifically address the problem of in-vehicle voice activity detection (VAD), which has a significant importance for the speech controlled intelligent vehicle. A novel VAD system is proposed based on microphone array beam- forming and discriminative Gaussian mixture model. As a binary classification problem, the features and classifiers are explored under the in-vehicle acoustic environment. Using microphone array, we show that the spatial power ratio can serve as an effective feature for speech activity detection. Further, a discriminative training based Gaussian mixture model (GMM) classifier is employed to enhance the VAD performance in terms of receiver operating characteristics (ROC). Compared to the conventional VAD systems, the proposed VAD system presents a novel and robust performance against various in-vehicle noisy scenarios from the UTDrive project.
John H. L. Hansen
Intelligent Vehicles Symposium2
2010 Automatic voice onset time detection for unvoiced stops (/p/, /t/, /k/) with application to accent classification
John H. L. Hansen, Sharmistha S. Gray, Wooil Kim
Speech Commun.1
2010 Analysis of CFA-BF: Novel combined fixed/adaptive beamforming for robust speech recognition in real car environments
John H. L. Hansen, Xianxian Zhang
Speech Commun.1
2010 The physiological microphone (PMIC): A competitive alternative for speaker assessment in stress detection and speaker verification
Sanjay A. Patil, John H. L. Hansen
Speech Commun.2
2010 Phonetic Distance Based Confidence Measure
abstract
This letter presents a novel confidence measure for the purpose of improving user performance in Spoken Document Retrieval (SDR). The proposed confidence measure is based on the phonetic distance between subword models, employing an anti-model which is determined to be discriminative to a target model using offline training data. As an advancement from our previous work, the proposed method employs separate phonetic similarity knowledge for vowels and consonants, resulting in more reliable performance over diverse SDR recorded speech conditions. A transcript reliability estimator is also presented, with evaluation as an application of the proposed confidence measure. Analysis on a variety of corpora including background noise, frequency band-restrictions, and a range of real-life conditions, shows that the proposed confidence measure is more reliable in detecting corrupted speech due to acoustic conditions or an unarticulated speaking style, providing a higher correlation to word error rate (WER). The proposed confidence measure is effective in increasing transcript reliability estimation performance with a 16.21% relative improvement.
Wooil Kim, John H. L. Hansen
IEEE Signal Process. Lett.2
2010 Discriminative Training for Multiple Observation Likelihood Ratio Based Voice Activity Detection
abstract
It is possible to show that the likelihood ratio (LR) test from multiple observations can enhance the performance of a statically modeled voice actively detection (VAD) system. However, the combination weights for the likelihood ratios (LRs) in each observation are rather empirical and heuristical. In this study, the optimal combination weights from two discriminative training methods are studied to directly improve VAD performance, in terms of reduced misclassification errors and improved receiver operating characteristics (ROC) curves. As shown in the evaluations, VAD performance, both in terms of absolute performance and consistency across noise types, can be significantly improved using the proposed method.
John H. L. Hansen
IEEE Signal Process. Lett.2
2010 Spoken Proper Name Retrieval for Limited Resource Languages Using Multilingual Hybrid Representations
abstract
Research in multilingual speech recognition has shown that current speech recognition technology generalizes across different languages, and that similar modeling assumptions hold, provided that linguistic knowledge (e.g., phone inventory, pronunciation dictionary, etc.) and transcribed speech data are available for the target language. Linguists make a very conservative estimate that 4000 languages are spoken today in the world, and in many of these languages, very limited linguistic knowledge and speech data/resources are available. Rapid transition to a new target language becomes a practical concern within the concept of tiered resources (e.g., different amounts of acoustically matched/mismatched data). In this paper, we present our research efforts towards multilingual spoken information retrieval with limitations in acoustic training data. We propose different retrieval algorithms to leverage existing resources from resource-rich languages as well as the target language. Proposed algorithms employ confusion-embedded hybrid pronunciation networks, and lattice-based phonetic search within a proper name retrieval task. We use Latin-American Spanish as the target language by intentionally limiting available resources for this language. After searching for queries consisting of Spanish proper names in Spanish Broadcast News data, we demonstrate that retrieval performance degradations (due to data sparseness during automatic speech recognition (ASR) deployment in the target language) are compensated by employing English acoustic models. It is shown that the proposed algorithms for developing rapid transition of rich languages to underrepresented languages are able to achieve comparable retrieval performance using 25% of the available training data.
Murat Akbacak, John H. L. Hansen
IEEE Trans. Speech Audio Process.2
2010 Unsupervised Equalization of Lombard Effect for Speech Recognition in Noisy Adverse Environments
abstract
In the presence of environmental noise, speakers tend to adjust their speech production in an effort to preserve intelligible communication. The noise-induced speech adjustments, called Lombard effect (LE), are known to severely impact the accuracy of automatic speech recognition (ASR) systems. The reduced performance results from the mismatch between the ASR acoustic models trained typically on noise-clean neutral (modal) speech and the actual parameters of noisy LE speech. In this study, novel unsupervised frequency domain and cepstral domain equalizations that increase ASR resistance to LE are proposed and incorporated in a recognition scheme employing a codebook of noisy acoustic models. In the frequency domain, short-time speech spectra are transformed towards neutral ASR acoustic models in a maximum-likelihood fashion. Simultaneously, dynamics of cepstral samples are determined from the quantile estimates and normalized to a constant range. A codebook decoding strategy is applied to determine the noisy models best matching the actual mixture of speech and noisy background. The proposed algorithms are evaluated side by side with conventional compensation schemes on connected Czech digits presented in various levels of background car noise. The resulting system provides an absolute word error rate (WER) reduction on 10-dB signal-to-noise ratio data of 8.7% and 37.7% for female neutral and LE speech, respectively, and of 8.7% and 32.8% for male neutral and LE speech, respectively, when compared to the baseline recognizer employing perceptual linear prediction (PLP) coefficients and cepstral mean and variance normalization.
Hynek Boril, John H. L. Hansen
IEEE Trans. Speech Audio Process.2
2010 Missing-Feature Reconstruction by Leveraging Temporal Spectral Correlation for Robust Speech Recognition in Background Noise Conditions
abstract
This paper proposes a novel missing-feature reconstruction method to improve speech recognition in background noise environments. The existing missing-feature reconstruction method utilizes log-spectral correlation across frequency bands. In this paper, we propose to employ a temporal spectral feature analysis to improve the missing-feature reconstruction performance by leveraging temporal correlation across neighboring frames. In a similar manner with the conventional method, a Gaussian mixture model is obtained by training over the obtained temporal spectral feature set. The final estimates for missing-feature reconstruction are obtained by a selective combination of the original frequency correlation based method and the proposed temporal correlation-based method. Performance of the proposed method is evaluated on the TIMIT speech corpus using various types of background noise conditions and the CU-Move in-vehicle speech corpus. Experimental results demonstrate that the proposed method is more effective at increasing speech recognition performance in adverse conditions. By employing the proposed temporal-frequency based reconstruction method, a$+17.71\%$average relative improvement in word error rate (WER) is obtained for white, car, speech babble, and background music conditions over 5-, 10-, and 15-dB SNR, compared to the original frequency correlation-based method. We also obtain a$+16.72\%$relative improvement in real-life in-vehicle conditions using data from the CU-Move corpus.
Wooil Kim, John H. L. Hansen
IEEE Trans. Speech Audio Process.2
2009 Mask estimation employing Posterior-based Representative Mean for missing-feature speech recognition with time-varying background noise
abstract
This paper proposes a novel mask estimation method for missing-feature reconstruction to improve speech recognition performance in time-varying background noise conditions. Conventional mask estimation methods based on noise estimates and spectral subtraction fail to reliably estimate the mask. The proposed mask estimation method utilizes a posterior-based representative mean (PRM) vector for determining the reliability of the input speech spectrum, which is obtained as a weighted sum of the mean parameters of the speech model with posterior probabilities. To obtain the noise-corrupted speech model, a model combination method is employed, which was proposed in our previous study for a feature compensation method. Experimental results demonstrate that the proposed mask estimation method is considerably more effective at increasing speech recognition performance in time-varying background noise conditions. By employing the proposed PRM-based mask estimation for missing-feature reconstruction, we obtain +36.29% and +30.45% average relative improvements in WER for speech babble and background music conditions respectively, compared to conventional mask estimation methods.
Wooil Kim, John H. L. Hansen
ASRU2
2009 Leveraging speech production knowledge for improved speech recognition
abstract
This study presents a novel phonological methodology for speech recognition based on phonological features (PFs) which leverages the relationship between speech phonology and phonetics. In particular, the proposed scheme estimates the likelihood of observing speech phonology given an associative lexicon. In this manner, the scheme is capable of choosing the most likely hypothesis (word candidate) among a group of competing alternative hypotheses. The framework employs the maximum entropy (ME) model to learn the relationship between phonetics and phonology. Subsequently, we extend the ME model to a ME-HMM (maximum entropy-hidden Markov model) which captures the speech production and linguistic relationship between phonology and words. The proposed ME-HMM model is applied to the task of re-processing N-best lists where an absolute WRA (word recognition rate) increase of 1.7%, 1.9% and 1% are reported for TIMIT, NTIMIT, and the SPINE (speech in noise) corpora (15.5% and 22.5% relative reduction in word error rate for TIMIT and NTIMIT).
Abhijeet Sangwan, John H. L. Hansen
ASRU2
2009 Unsupervised equalization of Lombard effect for speech recognition in noisy adverse environment
abstract
When exposed to environmental noise, speakers adjust their speech production to maintain intelligible communication. This phenomenon, called Lombard effect (LE), is known to considerably impact the performance of automatic speech recognition (ASR) systems. In this study, novel frequency and cepstral domain equalizations that reduce the impact of LE on ASR are proposed. Short-time spectra of LE speech are transformed towards neutral ASR models in a maximum likelihood fashion. Dynamics of cepstral coefficients are normalized to a constant range using quantile estimations. The algorithms are incorporated in a recognizer employing a codebook of noisy acoustic models. In a recognition task on connected Czech digits presented in various levels of background car noise, the resulting system provides an absolute reduction in word error rate (WER) on 10 dB SNR data of 8.7% and 37.7% for female neutral and LE speech, and of 8.7% and 32.8% for male neutral and LE speech when compared to the baseline system employing perceptual linear prediction (PLP) coefficients and cepstral mean and variance normalization.
Hynek Boril, John H. L. Hansen
ICASSP2
2009 Speaker identification with whispered speech based on modified LFCC parameters and feature mapping
abstract
Much research recently in speaker recognition has been devoted to robustness due to microphone and channel effects. However, changes in vocal effort, especially whispered speech, present significant challenges in maintaining system performance. Due to the absence of any periodic excitation in whisper, the spectral structure in whisper and neutral speech will differ. Therefore, performance of speaker ID systems, trained mainly with high energy voiced phonemes, degrades when tested with whisper. This study considers a front-end feature compensation method for whispered speech to improve speaker recognition using a neutral trained system. First, an alternative feature vector with linear frequency cepstral coefficients (LFCC) is introduced based on spectral analysis from both speech modes. Next, for the first time a feature mapping is proposed for reducing whisper/neutral mismatch in speaker ID. Feature mapping is applied on a frame-by-frame basis between two speaker independent GMMs (Gaussian Mixture Models) of whispered and neutral speech. Text independent closed set speaker ID results show an absolute 20% improvement in accuracy when compared with a traditional MFCC feature based system. This result confirms a viable approach to improving speaker ID performance between neutral and whispered speech conditions.
John H. L. Hansen
ICASSP2
2009 Factor analysis-based information integration for Arabic dialect identification
abstract
In this study, we propose a new factor analysis-based modeling technique to more clearly describe the composition of the supervector defined by the GMM model for dialect identification. The method utilizes knowledge types of information contained in the transcript file of the data. We evaluate the effects of the proposed modeling algorithm on a GMM-based Arabic dialect identification system. In particular, we compare eigenchannel modeling and our proposed information integration modeling. We show that the proposed modeling can obtain a 4.23% relative EER reduction with the same total number of factors, and a 9.37% relative EER reduction with the same number of channel/session factors versus eigenchannel modeling.
Yun Lei, John H. L. Hansen
ICASSP2
2009 A speech presence microphone array beamformer using model based speech presence probability estimation
abstract
The purpose of this study is to investigate the performance of speech presence (SP) microphone array beamforming. When the presence uncertainty of the desired speech is considered, noise reduction is greatly achieved while preserving low speech distortion level. Furthermore, we propose a novel model based speech presence probability (SPP) estimator, exploring both the sinusoid structure of speech and signal-to-noise ratio (SNR). Finally, experiments verify the effectiveness of the proposed SP-beamformer, resulting in a better trade-off between speech distortion and noise leakage, and a corresponding higher output segmental SNR, when compared with the classical beamformers.
John H. L. Hansen
ICASSP2
2009 Reduced complexity equalization of lombard effect for speech recognition in noisy adverse environments
abstract
In real-world adverse environments, speech signal corruption by background noise, microphone channel variations, and speech production adjustments introduced by speakers in an effort to communicate efficiently over noise (Lombard effect) severely impact automatic speech recognition (ASR) performance. Recently, a set of unsupervised techniques reducing ASR sensitivity to these sources of distortion have been presented, with the main focus on equalization of Lombard effect (LE). The algorithms performing maximum-likelihood spectral transformation, cepstral dynamics normalization, and decoding with a codebook of noisy speech models have been shown to outperform conventional methods, however, at a cost of considerable increase in computational complexity due to required numerous decoding passes through the ASR models. In this study, a scheme utilizing a set of speech-in-noise Gaussian mixture models and a neutral/LE classifier is shown to substantially decrease the computational load (from 14 to 2‐4 ASR decoding passes) while preserving overall system performance. In addition, an extended codebook capturing multiple environmental noises is introduced and shown to improve ASR in changing environments (8.2‐49.2% absolute WER improvement). The evaluation is performed on the Czech Lombard Speech Database (CLSD‘05). The task is to recognize neutral/LE connected digit strings presented in different levels of background car noise and Aurora 2 noises. Index Terms: Lombard effect, speech recognition, codebook decoding, frequency transformation, cepstral normalization
Hynek Boril, John H. L. Hansen
INTERSPEECH2
2009 Speech enhancement minimizing generalized euclidean distortion using supergaussian priors
John H. L. Hansen
INTERSPEECH2
2009 Speaker identification for whispered speech using modified temporal patterns and MFCCs
John H. L. Hansen
INTERSPEECH2
2009 Variational model composition for robust speech recognition with time-varying background noise
Wooil Kim, John H. L. Hansen
INTERSPEECH2
2009 Robust angry speech detection employing a TEO-based discriminative classifier combination
abstract
This study proposes an effective angry speech detection approach employing the TEO-based feature extraction. Decorrelation processing is applied to the TEO-based feature to increase model training ability by decreasing the correlation between feature elements and vector size. Minimum classification error training is employed to increase the discrimination between the angry speech model and other stressed speech models. Combination with the conventional Mel frequency cepstral coefficients (MFCC) is also employed to leverage the effectiveness of MFCC to characterize the spectral envelope of speech signals. Experimental results over the SUSAS corpus demonstrate the proposed angry speech detection scheme is effective at increasing detection accuracy on an open-speaker and open-vocabulary task. An improvement of up to 7.78% in classification accuracy is obtained by combination of the proposed methods including decorrelation of TEO-based feature vector, discriminative training, and classifier combination.
Wooil Kim, John H. L. Hansen
INTERSPEECH2
2009 The role of age in factor analysis for speaker identification
abstract
The speaker acoustic space described by a factor analysis model is assumed to reflect a majority of the speaker variations using a reduced number of latent factors. In this study, the age factor, as an observable important factor of a speaker’s voice, is analyzed and employed in the description of the speaker acoustic space, using a factor analysis approach. An age dependent acoustic space is developed for speakers, and the effect of the age dependent space in eigenvoice is evaluated using the NIST SRE08 corpus. In addition, the data pool with different age distributions are evaluated based on joint factor analysis model to assess age influence from the data pool.
Yun Lei, John H. L. Hansen
INTERSPEECH2
2009 On the use of phonological features for automatic accent analysis
abstract
In this paper, we present an automatic accent analysis system that is based on phonological features (PFs). The proposed system exploits the knowledge of articulation embedded in phonology by rapidly build Markov models (MMs) of PFs extracted from accented speech. The Markov models capture information in the PF space along two dimensions of articulation: PF state-transitions and state-durations. Furthermore, by utilizing MMs of native and non-native accents a new statistical measure of “accentedness” is developed which rates the articulation of a word on a scale of native-like ( 1) to non-native like (+1. The proposed methodology is then used to perform an automatic cross-sectional study of accented English spoken by native speakers of Mandarin Chinese (N-MC). The experimental results demonstrate the capability of the proposed system to rapidly perform quantitative as well as qualitative analysis of foreign accents. The work developed in this paper is easily assimilated into language learning systems, and has impact in the areas of speaker and speech recognition. Index Terms: automatic accent analysis, phonological features
Abhijeet Sangwan, John H. L. Hansen
INTERSPEECH2
2009 Robust minimal variance distortionless speech power spectra enhancement using order statistic filter for microphone array
John H. L. Hansen
INTERSPEECH2
2009 Advancements in whisper-island detection within normally phonated audio streams
abstract
In this study, several improvements are proposed for improved whisper-island detection within normally phonated audio streams. Based on our previous study, an improved feature, which is more sensitive to vocal effort change points between whisper and neutral speech, is developed and utilized in vocal effort change point(VECP) detection and vocal effort classification. Evaluation is based on the proposed multi-error score, where the improved feature showed better performance in VECPs detection with the lowest MES of 19.08. Furthermore, a more accurate whisper-island detection was obtained using the improved algorithm. Finally, the experimental detection rate results of 95.33% reflects better whisper-island detection performance for the improved algorithm versus that of the original baseline algorithm.
John H. L. Hansen
INTERSPEECH2
2009 Feature compensation in the cepstral domain employing model combination
Wooil Kim, John H. L. Hansen
Speech Commun.2
2009 Analysis and Compensation of Lombard Speech Across Noise Type and Levels With Application to In-Set/Out-of-Set Speaker Recognition
abstract
Speech production in the presence of noise results in the Lombard effect, which is known to have a serious impact on speech system performance. In this study, Lombard speech produced under different types and levels of noise is analyzed in terms of duration, energy histogram, and spectral tilt. Acoustic-phonetic differences are shown to exist between different ldquoflavorsrdquo of Lombard speech based on analysis of trends from a Gaussian mixture model (GMM)-based Lombard speech type classifier. For the first time, the dependence of Lombard speech on noise type and noise level is established for the purposes of speech processing systems. Also, the impact of the different flavors of Lombard effect on speech system performance is shown with respect to an in-set/out-of-set speaker recognition task. System performance is shown to degrade from an equal error rate (EER) of7.0%under matched neutral training and testing conditions, to an average EER of26.92%when trained with neutral and tested with Lombard effect speech. Furthermore, improvement in the performance of in-set/out-of-set speaker recognition is demonstrated by adapting neutral speaker models with Lombard speech data of limited duration. Improved average EERs of4.75%and12.37%were achieved for matched and mismatched adaptation and testing conditions, respectively. At the highest noise levels, an EER as low as1.78%was obtained by adapting neutral speaker models with Lombard speech of limited duration. The study therefore illustrates the impact of Lombard effect on speaker recognition, and effective methods to improve system performance for speaker recognition when train/test conditions are mismatched for neutral versus Lombard effect speech.
John H. L. Hansen, Vaishnevi S. Varadarajan
IEEE Trans. Speech Audio Process.1
2009 Time-Frequency Correlation-Based Missing-Feature Reconstruction for Robust Speech Recognition in Band-Restricted Conditions
abstract
Band-limited speech represents one of the most challenging factors for robust speech recognition. This is especially true in supporting audio corpora from sources that have a range of conditions in spoken document retrieval requiring effective automatic speech recognition. The missing-feature reconstruction method has a problem when applied to band-limited speech reconstruction, since it assumes the observations in the unreliable regions are always greater than the latent original clean speech. The approach developed here depends only on reliable components to calculate the posterior probability to mitigate the problem. This study proposes an advanced method to effectively utilize the correlation information of the spectral components across time and frequency axes in an effort to increase the performance of missing-feature reconstruction in band-limited conditions. We employ an F1 area window and cutoff border window in order to include more knowledge on reliable components which are highly correlated with the cutoff frequency band. To detect the cutoff regions for missing-feature reconstruction, blind mask estimation is also presented, which employs the synthesized band-limited speech model without secondary training data. Experiments to evaluate the performance of the proposed methods are accomplished using the SPHINX3 speech recognition engine and the TIMIT corpus. Experimental results demonstrate that the proposed time-frequency (TF) correlation based missing-feature reconstruction method is significantly more effective in improving band-limited speech recognition accuracy. By employing the proposed TF-missing feature reconstruction method, we obtain up to 14.61% of average relative improvement in word error rate (WER) for four available bandwidths with cutoff frequencies 1.0, 1.5, 2.0, and 2.5 kHz, respectively, compared to earlier formulated methods. Experimental results on the National Gallery of the Spoken Word (NGSW) corpus also show the proposed method is effective in improving band-limited speech recognition in real-life spoken document conditions.
Wooil Kim, John H. L. Hansen
IEEE Trans. Speech Audio Process.2
2009 Babble Noise: Modeling, Analysis, and Applications
abstract
Speech babble is one of the most challenging noise interference for all speech systems. Here, a systematic approach to model its underlying structure is proposed to further the existing knowledge of speech processing in noisy environments. This paper establishes a working foundation for the analysis and modeling of babble speech. We first address the underlying model for multiple speaker babble speech - considering the number of conversations versus the number of speakers contributing to babble. Next, based on this model, we develop an algorithm to detect the range of the number of speakers within an unknown babble speech sequence. Evaluation is performed using 110 h of data from the Switchboard corpus. The number of simultaneous conversations ranges from one to nine, or one to 18 subjects speaking. A speaker conversation stream detection rate in excess of 80% is achieved with a speaker window size of plusmn1 speakers. Finally, the problem of in-set/out-of-set speaker recognition is considered in the context of interfering babble speech noise. Results are shown for test durations from 2-8 s, with babble speaker groups ranging from two to nine subjects. It is shown that by choosing the correct number of speakers in the background babble an overall average performance gain of 6.44% equal error rate can be obtained. This study represents effectively the first effort in developing an overall model for speech babble, and with this, contributions are made for speech system robustness in noise.
Nitish Krishnamurthy, John H. L. Hansen
IEEE Trans. Speech Audio Process.2
2009 Feature Compensation Techniques for ASR on Band-Limited Speech
abstract
Band-limited speech (speech for which parts of the spectrum are completely lost) is a major cause for accuracy degradation of automatic speech recognition (ASR) systems particularly when acoustic models have been trained with data with a different spectral range. In this paper, we present an extensive study of the problem of ASR of band-limited speech with full-bandwidth acoustic models. Our focus is mainly on band-limited feature compensation, covering even the case of time-varying band-limiting distortions, but we also compare this approach to more common model-side techniques (adaptation and retraining) and explore the combination of feature-based and model-side approaches. The feature compensation algorithms proposed are organized in a unified framework supported by a novel mathematical model of the impact of such distortions on Mel-frequency cepstral coefficient (MFCC) features. A crucial and novel contribution is the analysis made of the relative correlation of different elements in the MFCC feature vector for the cases of full-bandwidth and limited-bandwidth speech, which justifies an important modification in the feature compensation scheme. Furthermore, an intensive experimental analysis is provided. Experiments are conducted on real telephone channels, as well as artificial low-pass and bandpass filters applied over TIMIT data, and results are given for different experimental constraints and variations of the feature compensation method. Results for other well-known robustness approaches, such as cepstral mean normalization (CMN), model retraining, and model adaptation are also given for comparison. ASR performance with our approach is similar or even better than model adaptation, and we argue that in particular cases such as rapidly varying distortions, or limited computational or memory resources, feature compensation is more convenient. Furthermore, we show that feature-side and model-side approaches may be combined, outperforming any of those approaches alone.
Nicolás Morales, Doroteo T. Toledano, John H. L. Hansen, Javier Garrido Salas
IEEE Trans. Speech Audio Process.3
2008 Speech babble: Analysis and modeling for speech systems
abstract
Speech babble represents the most challenging noise interference in all speech systems, yet no research has been performed at a systematic level to model the underlying structure. For the first time, this study establishes a working foundation for the analysis and modeling of babble speech. We first address the underlying model for multiple speaker babble speech - considering the number of conversations versus the number of speakers. Next, based on this model, we develop an algorithm to detect the range of speakers within an unknown babble speech sequence. Evaluation is performed using 110 hours of data from the SWITCHBOARD corpus. The number of simultaneous conversations ranges from 1-9, or 1 to 18 subjects speaking. A speaker conversation stream detection rate in excess of 80% is achieved with a speaker window size of plusmn 1 speaker. This study is the first in developing an effective speaker babble model to contribute to robust speech systems.
Nitish Krishnamurthy, John H. L. Hansen
ICASSP2
2008 In-set/out-of-set speaker recognition: leverging the speaker and noise balance
abstract
This study addresses the problem of identifying in-set versus out-of-set speakers in noise for limited train/test durations in situations where rapid detection and tracking is required. The objective is to form a decision as to whether the current input speaker is accepted as a member of the enrolled in-set group or rejected as an outside speaker. A new scoring algorithm that combines scores across an energy-frequency grid is developed where high-energy speaker dependent frames are fused with weighted scores from low-energy noise dependent frames. By leveraging the balance between the speaker versus the background noise environment, it is possible to see an improvement in equal error rate performance. Using an initial form of the algorithm with speakers from the TIMIT database with 5 seconds of train and 2 seconds of test, the average relative EER performance improvement is 27.4%. The results confirm that for situations in which the background environment type remains constant between train and test, an in-set/out-of-set speaker recognition system that takes advantage of information gathered from the environmental noise can be formulated which realizes significant improvement.
Matthew R. Leonard, John H. L. Hansen
ICASSP2
2008 Generalized parametric spectral subtraction using weighted Euclidean distortion
John H. L. Hansen
INTERSPEECH2
2008 Speaker identification for whispered speech based on frequency warping and score competition
John H. L. Hansen
INTERSPEECH2
2008 Analysis and perception of speech under physical task stress
abstract
It is known that speech under physical task stress degrades speech system performance. Therefore, an analysis of speech under physical task stress is performed across several parame-ters to identify acoustic correlates. Formal listener tests are also performed to determine the relationship between acoustic corre-lates and perception. To verify the statistical significance of all results, student-t statistical tests are applied. It was found that fundamental frequency decreases for many speakers, that ut-terance duration increases for some speakers and decreases for others, and that the glottal waveform is quantifiably different for many speakers. Perturbation of two speech features, funda-mental frequency and the glottal waveform, is applied in listener tests to quantify the degree to which these features convey phys-ical stress content in speech. Finally, the enhanced understand-ing of physical task stress speech provided here is discussed in the context of speech systems. Index Terms: physical task stress, stress analysis 1.
Keith W. Godin, John H. L. Hansen
INTERSPEECH2
2008 Missing-feature method for speaker recognition in band-restricted conditions
Wooil Kim, John H. L. Hansen
INTERSPEECH2
2008 Babble speech: acoustic and perceptual variability
Nitish Krishnamurthy, Ayako Ikeno, John H. L. Hansen
INTERSPEECH3
2008 Environment mismatch compensation using average eigenspace for speech recognition
abstract
The performance of speech recognition systems is adversely affected by mismatch in training and testing environmental conditions. In addition to test data from noisy environments, there are scenarios where the training data itself is noisy. Speech enhancement techniques which solely focus on finding a clean speech estimate from the noisy signal are not effective here. Model adaptation techniques may also be ineffective due to the dynamic nature of the environment. In this paper, we propose a method for mismatch compensation between training and testing environments using the ”average eigenspace ” approach when the mismatch is non-stationary. There is no need for explicit adaptation data as the method works on incoming test data to find the compensatory transform. This method is different from traditional signal-noise subspace filtering techniques where the dimensionality of the clean signal space is assumed to be less than the noise space and noise affects all dimensions to the same extent. We evaluate this approach on two corpora which are collected from real car environments: CU-Move and UTDrive. Using Sphinx, a relative reduction of 40-50 % is achieved in WER compared to the baseline system. The method also results in a reduction in the dimensionality of the feature vectors allowing for a more compact set of acoustic models in the phoneme space. Index Terms: speech recognition, feature adaptation, eigenvector, simultaneous diagonalization
John H. L. Hansen
INTERSPEECH2
2008 Dialect classification via discriminative training
Yun Lei, John H. L. Hansen
INTERSPEECH2
2008 Dialect separation assessment using log-likelihood score distributions
Mahnoosh Mehrabani, John H. L. Hansen
INTERSPEECH2
2008 Detection of speech under physical stress: model development, sensor selection, and feature fusion
Sanjay A. Patil, John H. L. Hansen
INTERSPEECH2
2008 Evidence of coarticulation in a phonological feature detection system
abstract
In this study, we investigate the capability of phonological features (PFs) in capturing the fine variational structure in speech which arise due to natural phenomenon such as coarticulation. The PF theory provides a framework in which a far more richer description of speech is possible when compared to traditional phonetic representations. However, current approaches toward training PF detectors do not explicitly expose the statistical system to patterns of coarticulation. The analysis presented here shows that despite this handicap, our PF system still learns to capture these variants in speech. In fact, it is noted that the use of phone-based transcriptions to judge the performance of PF systems erroneously labels such variants as errors. Our result show that a large proportion of speech frames that are deemed errors by phone-transcriptions are actually coarticulated as is evidenced by their phonetic context. These findings offer important knowledge in analyzing and improving the utility of PFs in ASR (automatic speech recognition) for spontaneous conversational speech.
Abhijeet Sangwan, Ayako Ikeno, John H. L. Hansen
INTERSPEECH3
2008 Filling acoustic holes through leveraged uncorellated GMMs for in-set/out-of-set speaker recognition
Jun-Won Suh, Pongtep Angkititrakul, John H. L. Hansen
INTERSPEECH3
2008 An entropy based feature for whisper-island detection within audio streams
John H. L. Hansen
INTERSPEECH2
2008 A new perceptually motivated MVDR-based acoustic front-end (PMVDR) for robust automatic speech recognition
Umit H. Yapanel, John H. L. Hansen
Speech Commun.2
2007 Multi-stream dialect classification using SVM-GMM hybrid classifiers
abstract
In this paper, we investigate two important issues that influence dialect classification: (i) exploring dialect dependent features, and (ii) an effective way of combining spectral, excitation, and vocal tract information to improve dialect classification. The motivation is that dialect dependent features such as formants, LSP (line spectral pairs) and MEPZ (MFCCs + energy + pitch) span a wider range of speech production traits and are therefore better suited than traditional MFCCs for characterizing dialects. After establishing the proposed algorithm, we compare individual performances of each feature on a corpus of three dialects of Spanish. Next, we present a method for combining these features using GMM-SVM hybrid classifiers. The final combined system achieves a 30% relative improvement in dialect classification accuracy, confirming that the proposed advances significantly outperform conventional methods for dialect classification.
Rahul Chitturi, John H. L. Hansen
ASRU2
2007 Speechfind for CDP: Advances in spoken document retrieval for the U. S. collaborative digitization program
abstract
This paper presents our recent advances for SpeechFind, a CRSS-UTD designed spoken document retrieval system for the U.S. based Collaborative Digitization Program (CDP1). A proto-type of SpeechFind for the CDP is currently serving as the search engine for 1,300 hours of CDP audio content which contain a wide range of acoustic conditions, vocabulary and period selection, and topics. In an effort to determine the amount of user corrected transcripts needed to impact automatic speech recognition (ASR) and audio search, a webbased online interface for veri..cation of ASR-generated transcripts was developed. The procedure for enhancing the transcription performance for SpeechFind is also presented. A selection of adaptation methods for language and acoustic models are employed depending on the acoustics of the corpora under test. Experimental results on the CDP corpus demonstrate that the employed model adaptation scheme using the veri..ed transcripts is effective in improving recognition accuracy. Through a combination of feature/acoustic model enhancement and language model selection, up to 24.8% relative improvement in ASR was obtained. The SpeechFind system, employing automatic transcript generation, online CDP transcript correction, and our transcript reliability estimator, demonstrates a comprehensive support mechanism to ensure reliable transcription and search for U.S. libraries with limited speech technology experience.
Wooil Kim, John H. L. Hansen
ASRU2
2007 Phonological feature based variable frame rate scheme for improved speech recognition
abstract
In this paper, we propose a new scheme for variable frame rate (VFR) feature processing based on high level segmentation (HLS) of speech into broad phone classes. Traditional fixed-rate processing is not capable of accurately reflecting the dynamics of continuous speech. On the other hand, the proposed VFR scheme adapts the temporal representation of the speech signal by tying the framing strategy with the detected phone class sequence. The phone classes are detected and segmented by using appropriately trained phonological features (PFs). In this manner, the proposed scheme is capable of tracking the evolution of speech due to the underlying phonetic content, and exploiting the non-uniform information flow-rate of speech by using a variable framing strategy. The new VFR scheme is applied to automatic speech recognition of TIMIT and NTIMIT corpora, where it is compared to a traditional fixed window-size/frame-rate scheme. Our experiments yield encouraging results with relative reductions of 24% and 8% in WER (word error rate) for TIMIT and NTIMIT tasks, respectively.
Abhijeet Sangwan, John H. L. Hansen
ASRU2
2007 Language Normalization for Bilingual Speaker Recognition Systems
abstract
In this study, we focus on the problem of removing/normalizing the impact of spoken language variation in bilingual speaker recognition (BSR) systems. In addition to environment, recording, and channel mismatches, spoken language mismatch is an additional factor resulting in performance degradation in speaker recognition systems. In today's world, the number of bilingual speakers is increasing with English becoming the universal second language. Data sparseness is becoming an important research issue to deploy speaker recognition systems with limited resources (e.g., short train/test durations). Therefore, leveraging existing resources from different languages becomes a practical concern in limited-resource BSR applications, and effective language normalization schemes are required to achieve more robust speaker recognition systems. Here, we propose two novel algorithms to address the spoken language mismatch problem: normalization at the utterance-level via language identification (LID), and normalization at the segment-level via multilingual phone recognition (PR). We evaluated our algorithms using a bilingual (Spanish-English) speaker set of 80 speakers. Experimental results show improvements over a baseline system which employs fusion of language-dependent speaker models with fixed weights.
Murat Akbacak, John H. L. Hansen
ICASSP (4)2
2007 Dialect Classification on Printed Text using Perplexity Measure and Conditional Random Fields
abstract
Studies have shown that dialect variation has a significant impact in speech recognition performance, and therefore it is important to be able to perform effective dialect classification to improve speech systems. Dialects differ at the acoustic, grammar, and vocabulary levels. In this study, topic-specific printed text dialect data are collected from the ten major newspapers in Australia, United Kingdom, and United States. An n-gram language model is trained for each topic in each country/dialect. The perplexity measure is applied to classify the dialect-dependent documents. In addition to the n-gram information, further features can be extracted from text structure. Conditional random fields (CRF) is such a model which can extract different levels of features and is still mathematically tractable. The CRF is applied to train the language model and classify documents. Significant improvement on dialect classification is achieved by using the CRF based classifier, especially on the small size documents (10% to 22% relative error reduction). Text classification based on variable size documents is explored and a document with several hundred words is shown to be sufficient for dialect classification. The vocabulary difference among the text documents from different countries are explored and the dialect difference is smoothly connected with the vocabulary difference. Five document topics are evaluated and performance for cross topic dialect classification is explored.
Rongqing Huang, John H. L. Hansen
ICASSP (4)2
2007 Getting start with UTDrive: driver-behavior modeling and assessment of distraction for in-vehicle speech systems
abstract
This paper describes our first step for advances in humanmachine interactive systems for in-vehicle environments of the UTDrive project. UTDrive is part of an on-going international collaboration to collect and research rich multi-modal data recorded for modeling behavior while the driver is interacting with speech-activated systems or performing other secondary tasks. A simultaneous second goal is to better understand speech characteristics of the driver undergoing additional cognitive load since dialog systems are generally not formulated for high task-stress environment (e.g., driving a vehicle). The corpus consists of audio, video, brake/gas pedal pressure, forward distance, GPS information, and CAN-Bus information. The resulting corpus, analysis, and modeling will contribute to more effective speech systems which are able to sense driver cognitive distraction/stress and adapt itself to the driver’s cognitive capacity and driving situations for improved safety while driving. Index Terms: in-vehicle speech system, driver distraction, multi-modal resource,driver behavior modeling, safety driving
Pongtep Angkititrakul, DongGu Kwak, SangJo Choi, JeongHee Kim, Anh PhucPhan, Amardeep Sathyanarayana, John H. L. Hansen
INTERSPEECH7
2007 Class constrained ROVER based speech enhancement
John H. L. Hansen
INTERSPEECH2
2007 Lombard speech impact on perceptual speaker recognition
Ayako Ikeno, John H. L. Hansen
INTERSPEECH2
2007 Advances in speechfind: transcript reliability estimation employing confidence measure based on discriminative sub-word model for SDR
Wooil Kim, John H. L. Hansen
INTERSPEECH2
2007 Noise tracking for speech systems in adverse environments
Nitish Krishnamurthy, John H. L. Hansen
INTERSPEECH2
2007 Score distribution scaling for speaker recognition
Vinod Prakash, John H. L. Hansen
INTERSPEECH2
2007 Environmentally aware voice activity detector
abstract
Abstract Traditional voice activity detectors (VADs) tend to be deaf to theacoustical background noise, as they (i) utilize a single operatingpoint for all SNRs (signal-to-noise ratios) and noise types, and(ii) attempt to learn the background noise model online from fi-nite data length. In this paper, we address the aforementionedissues by designing an environmentally aware (EA) VAD. TheEA VAD scheme builds prior offline knowledge of commonlyencountered acoustical backgrounds, and also combines the re-cently proposed competitive Neyman-Pearson (CNP) VAD witha SVM (support vector machine) based noise classifier. In oper-ation, the EA VAD obtains accurate noise models of the acous-tical background by employing the noise classifier and its priorknowledge of the noise type, and thereafter uses this informa-tion to set the best operating point and initialization parametersfor the CNP VAD. The superior performance of the EA VADscheme over the standard AMR (adaptive multi-rate) VADs inlow SNR is confirmed in a simulation study, where speech andnoise data were drawn from the SWITCHBOARD and NOISEXdatabases. We report an absolute improvement of 10-15% in de-tection rates over AMR VADs in low SNR for different noisetypes.Index Terms: noise modeling, voice activity detector, environ-mental sniffing
Abhijeet Sangwan, Nitish Krishnamurthy, John H. L. Hansen
INTERSPEECH3
2007 Analysis and classification of speech mode: whispered through shouted
abstract
Variation in vocal effort represents one of the most challenging problems in maintaining speech system performance for coding, speech and speaker recognition. Changes in vocal effort (or mode) result in a fundamental change in speech production which is not simply a change in volume. This is the first study to collectively consider the five speech modes: whispered, soft, neutral, loud and shouted. After corpus development, analysis is performed for i) sound intensity level, ii) duration and silence percentage, iii) frame energy distribution and iv) spectral tilt. The analysis shows vocal effort dependent traits which are used to investigate speaker recognition. Matched vocal mode conditions result in a closed-set speaker ID rate of 97.62%, with mismatch vocal conditions producing 54.02%. Finally, a speech mode classification system is developed, which has a range of classification rate from 44.5 % to 98.5 % confusing with adjacent vocal modes. These advancements can provide improved speech/speaker modeling information, as well as classified vocal mode knowledge to improve speech and language technology in real scenarios.
John H. L. Hansen
INTERSPEECH2
2007 Blind Feature Compensation for Time-Variant Band-Limited Speech Recognition
abstract
Mismatch in speech bandwidth between training and real operation greatly affects automatic speech recognition. This letter extends previous work on feature compensation of band-limited speech to establish a framework for blind compensation of speech data of unknown bandwidth, valid even when the distortion (band-limiting channel) changes rapidly and continuously in time. The available bandwidth of the input speech signal is automatically detected, and the band-limited feature vectors are compensated prior to being compared against full-bandwidth acoustic models. For a fixed bandwidth limitation, phoneme recognition performance using the proposed method is similar to that achieved by models adapted to match the distortion. However, compared to model adaptation approaches, this new approach can seamlessly be extended to rapidly time-varying conditions while maintaining low computational and memory costs
Nicolás Morales, Doroteo T. Toledano, John H. L. Hansen, José Colás Pasamontes
IEEE Signal Process. Lett.3
2007 Environmental Sniffing: Noise Knowledge Estimation for Robust Speech Systems
abstract
Automatic speech recognition systems work reasonably well under clean conditions but become fragile in practical applications involving real-world environments. To date, most approaches dealing with environmental noise in speech systems are based on assumptions concerning the noise, or differences in collecting and training on a specific noise condition, rather than exploring the nature of the noise. As such, speech recognition, speaker ID, or coding systems are typically retrained when new acoustic conditions are to be encountered. In this paper, we propose a new framework entitled Environmental Sniffing to detect, classify, and track acoustic environmental conditions. The first goal of the framework is to seek out detailed information about the environmental characteristics instead of just detecting environmental changes. The second goal is to organize this knowledge in an effective manner to allow smart decisions to direct subsequent speech processing systems. Our current framework uses a number of speech processing modules including a hybrid algorithm with T2-BIC segmentation, Gaussian mixture model/hidden Markov model (GMM/HMM)-based classification and noise language modeling to achieve effective noise knowledge estimation. We define a new information criterion that incorporates the impact of noise into Environmental Sniffing performance. We use an in-vehicle speech and noise environment as a test platform for our evaluations and investigate the integration of Environmental Sniffing for automatic speech recognition (ASR) in this environment. Noise sniffing experiments show that our proposed hybrid algorithm achieves a classification error rate of 25.51%, outperforming our baseline system by 7.08%. The sniffing framework is compared to a ROVER solution for automatic speech recognition (ASR) using different noise conditioned recognizers in terms of word error rate (WER) and CPU usage. Results show that the model matching scheme using the knowledge extracted from the audio stream by Environmental Sniffing achieves better performance than a ROVER solution both in accuracy and computation. A relative 11.1% WER improvement is achieved with a relative 75% reduction in CPU resources
Murat Akbacak, John H. L. Hansen
IEEE Trans. Speech Audio Process.2
2007 Discriminative In-Set/Out-of-Set Speaker Recognition
abstract
In this paper, the problem of identifying in-set versus out-of-set speakers for limited training/test data durations is addressed. The recognition objective is to form a decision regarding an input speaker as being a legitimate member of a set of enrolled speakers or outside speakers. The general goal is to perform rapid speaker model construction from limited enrollment and test size resources for in-set testing for input audio streams. In-set detection can help ensure security and proper access to private information, as well as detecting and tracking input speakers. Areas of applications of these concepts include rapid speaker tagging and tracking for information retrieval, communication networks, personal device assistants, and location access. We propose an integrated system with emphasis on short-enrollment data (about 5 s of speech for each enrolled speaker) and test data (2-8 s) within a text-independent mode. We present a simple and yet powerful decision rule to accept or reject speakers using a discriminative vector in the decision score space, together with statistical hypothesis testing based on the conventional likelihood ratio test. Discriminative training is introduced to further improve system performance for both decision techniques, by employing minimum classification error and minimum verification error frameworks. Experiments are performed using three separate corpora. Using the YOHO speaker recognition database, the alternative decision rule achieves measurable improvement over the likelihood ratio test, and discriminative training consistently enhances overall system performance with relative improvements ranging from 11.26%-28.68%. A further extended evaluation using the TIMIT (CORPUS1) and actual noisy aircraft communications data (CORPUS2) shows measurable improvement over the traditional MAP based scheme using the likelihood ratio test (MAP-LRT), with average EERs of 9%-23% for TIMIT and 13%-32% for noisy aircraft communications. The results confirm that an effective in-set/out-of-set speaker recognition system can be formulated using discriminative training for rapid tagging of input speakers from limited training and test data sizes
Pongtep Angkititrakul, John H. L. Hansen
IEEE Trans. Speech Audio Process.2
2007 Unsupervised Discriminative Training With Application to Dialect Classification
abstract
Automatic dialect classification has gained interest in the field of speech research because of its importance in characterizing speaker traits and knowledge estimation which could improve integrated speech technology (e.g., speech recognition, speaker recognition). This study addresses novel advances in unsupervised spontaneous dialect classification in English and Spanish. The problem considers the case where no transcripts are available for training and test data, and speakers are talking spontaneously. The Gaussian mixture model (GMM) is used for unsupervised dialect classification in our study. Techniques which aim to deal with confused acoustic regions in the GMMs are proposed, where confused regions in the GMMs are identified through data driven methods. The first technique excludes confused regions by finding dialect dependence in the untranscribed audio by selecting the most discriminative Gaussian mixtures [mixture selection (MS)]. The second technique includes the confused regions in the model, but the confused regions are balanced over all classes. This technique is implemented by identifying discriminative frames and confused frames in the audio data [frame selection (FS)]. The new confused regions contribute to model representation but does not impact classification performance. The third technique is to reduce the confused regions in the original model. Minimum classification error (MCE) is applied to achieve this objective. All three techniques implement discriminative training for GMM-based classification. Both the first technique (MS-GMM, GMM trained with mixture selection) and the second technique (FS-GMM, GMM trained with frame selection) improve dialect classification performance. Further improvement is achieved after applying the third technique (MCE training) before the first or second techniques. The system is evaluated using British English dialects and Latin American Spanish dialects. Measurable improvement is achieved in both corpora. Finally, the system is compared with human listener performance, and shown to outperform human listeners in terms of classification accuracy.
Rongqing Huang, John H. L. Hansen
IEEE Trans. Speech Audio Process.2
2007 Dialect/Accent Classification Using Unrestricted Audio
abstract
This study addresses novel advances in English dialect/accent classification. A word-based modeling technique is proposed that is shown to outperform a large vocabulary continuous speech recognition (LVCSR)-based system with significantly less computational costs. The new algorithm, which is named Word-based Dialect Classification (WDC), converts the text-independent decision problem into a text-dependent decision problem and produces multiple combination decisions at the word level rather than making a single decision at the utterance level. The basic WDC algorithm also provides options for further modeling and decision strategy improvement. Two sets of classifiers are employed for WDC: a word classifier DW(k)and an utterance classifier Du. DW(k)is boosted via the AdaBoost algorithm directly in the probability space instead of the traditional feature space. Duis boosted via the dialect dependency information of the words. For a small training corpus, it is difficult to obtain a robust statistical model for each word and each dialect. Therefore, a context adapted training (CAT) algorithm is formulated, which adapts the universal phoneme Gaussian mixture models (GMMs) to dialect-dependent word hidden Markov models (HMMs) via linear regression. Three separate dialect corpora are used in the evaluations that include the Wall Street Journal (American and British English), NATO N4 (British, Canadian, Dutch, and German accent English), and IViE (eight British dialects). Significant improvement in dialect classification is achieved for all corpora tested
Rongqing Huang, John H. L. Hansen, Pongtep Angkititrakul
IEEE Trans. Speech Audio Process.2
2007 In-Set/Out-of-Set Speaker Recognition Under Sparse Enrollment
abstract
In this paper, the problem of identifying in-set versus out-of-set speakers using extremely limited enrollment data is addressed. The recognition objective is to form a binary decision regarding an input speaker as being a legitimate member of a set of enrolled speakers or not. Here, the emphasis is on low enrollment (about 5 sec of speech for each enrolled speaker) and test data durations (2-8 sec), in a text-independent scenario. In order to overcome the limited enrollment, data from speakers that are acoustically close to a given in-set speaker are used to form an informative prior (base model) for speaker adaptation. Score normalization for in-set systems is addressed, and the difficulty of using conventional score normalization schemes for in-set speaker recognition is highlighted. Distribution scaling based score normalization techniques are developed specifically for the in-set/out-of-set problem and compared against existing score normalization schemes used in open-set speaker recognition. Experiments are performed using the following three separate corpora: (1) noise-free TIMIT; (2) noisy in-vehicle CU-move; and (3) the NIST-SRE-2006 database. Experimental results show a consistent increase in system performance for the proposed techniques.
Vinod Prakash, John H. L. Hansen
IEEE Trans. Speech Audio Process.2
2006 Spoken Proper Name Retrieval in Audio Streams for Limited-Resource Languages Via Lattice Based Search Using Hybrid Representations
abstract
Research in multilingual speech recognition has shown that current speech recognition technology generalizes across different languages, and that similar modeling assumptions hold, provided that linguistic knowledge (e.g., phoneme inventory, pronunciation dictionary, etc.) and transcribed speech data are available for the target language. Linguists make a very conservative estimate that 4000 languages are spoken today in the world, and in many of these languages, very limited linguistic knowledge and speech data/resources are available. Rapid transition to a new target language becomes a practical concern within the concept of tiered resources. In this study, we present our research efforts towards multilingual spoken information retrieval with limitations in acoustic training data. We propose different retrieval algorithms to leverage existing resources from resource-rich languages as well as the target language using a lattice-based search. We use Latin-American Spanish as the target language. After searching for queries consisting of Spanish proper names in Spanish Broadcast News data, we obtain performance (max-F value of 28.3%) close to that of a Spanish based system (trained on speech data from 36 speakers) using only 25% of all the available speech data from the original target language
Murat Akbacak, John H. L. Hansen
ICASSP (1)2
2006 Perceptual Recognition Cues in Native English Accent Variation: "Listener Accent, Perceived Accent, and Comprehension"
abstract
There are many aspects of speech that can provide information about a particular speaker's characteristics. Accent is a linguistic trait of speaker identity. It indicates the speaker's language and social background. The goal of this study is to provide perceptual human recognition of English native accent variation for accent and dialect identification applications. To examine relationships of the listener's accent background with perceived accent and comprehension of the speech, perceptual experiments are conducted with three types of listeners - US and British native English listeners, and nonnative English listeners. The tasks are accent detection and classification, and transcription of the speech. The results from the study show that listeners' accent background significantly impacts accent perception. The results also indicate that listeners use perceptual cues differently based on the task. Our analysis also suggests that comprehensibility of the speech affects accuracy of accent detection and classification. These observations point to the complex nature of the cognitive process involved in accent perception, which is bidirectional (bottom-up and top-down processing) and multi-dimensional (speech perception, language comprehension, etc.). This suggests the importance of understanding accent variation from a cognitive perspective for further development of accent and dialect identification systems as well as speech processing algorithms in general
Ayako Ikeno, John H. L. Hansen
ICASSP (1)2
2006 Unsupervised Class-Based Feature Compensation for Time-Variable Bandwidth-Limited Speech
abstract
This paper deals with the problem of speech recognition on band-limited speech. In our previous work we showed how a simple polynomial correction framework could be used for compensation of band-limited speech to minimize the mismatch using full-bandwidth acoustic models. This paper extends this approach to time-varying multiple-channel environments. The compensation framework is extended to perform automatic channel classification prior to compensation, thus allowing for unsupervised multi-channel compensation without the need for an explicit channel classifier. Performance is demonstrated on a wide range of channel bandwidth conditions. This extension makes our compensation approach potentially applicable in a much wider range of scenarios with only very limited performance degradation compared to the supervised approach
Nicolás Morales, Doroteo T. Toledano, John H. L. Hansen, Javier Garrido Salas, José Colás Pasamontes
ICASSP (1)3
2006 Stress Level Classification of Speech Using Euclidean Distance Metrics in a Novel Hybrid Multi-Dimensional Feature Space
abstract
Presently, automatic stress detection methods for speech employ a binary decision approach, deciding whether the speaker is or is not under stress. Since the amount of stress a speaker is under varies and can change gradually, a reliable stress level detection scheme becomes necessary to accurately assess the condition of the speaker. Such a capability is pertinent to a number of applications, such as for those personnel in law enforcement positions. Using speech and biometric data collected from a realworld, variable-stress level law enforcement training scenario, this study illustrates two methods for automatically assessing stress levels in speech using a hybrid multi-dimensional feature space comprised of frequency-based and Teager Energy Operator-based features. The first approach uses a nearest neighbor-type clustering scheme at the vowel token level to classify speech data into one of three levels of stress, yielding an overall error rate of 50.5%. The second approach employs accumulated Euclidean distance metric weighting at the sentence-level to yield a relative improvement of 12.1% in performance.
Evan Ruzanski, John H. L. Hansen, James Meyerhoff, George Saviolakis, William Norris 0001, Terry Wollert
ICASSP (1)2
2006 A robust fusion method for multilingual spoken document retrieval systems employing tiered resources
Murat Akbacak, John H. L. Hansen
INTERSPEECH2
2006 Decision directed constrained iterative speech enhancement
John H. L. Hansen
INTERSPEECH2
2006 Unsupervised Spanish dialect classification
Rongqing Huang, John H. L. Hansen
INTERSPEECH2
2006 The role of prosody in the perception of US native English accents
Ayako Ikeno, John H. L. Hansen
INTERSPEECH2
2006 Missing-feature reconstruction for band-limited speech recognition in spoken document retrieval
Wooil Kim, John H. L. Hansen
INTERSPEECH2
2006 Noise update modeling for speech enhancement: when do we do enough?
Nitish Krishnamurthy, John H. L. Hansen
INTERSPEECH2
2006 A cohort - UBM approach to mitigate data sparseness for in-set/out-of-set speaker recognition
Vinod Prakash, John H. L. Hansen
INTERSPEECH2
2006 Analysis of lombard effect under different types and levels of noise with application to in-set speaker ID systems
Vaishnevi S. Varadarajan, John H. L. Hansen
INTERSPEECH2
2006 Advances in phone-based modeling for automatic accent classification
abstract
It is suggested that algorithms capable of estimating and characterizing accent knowledge would provide valuable information in the development of more effective speech systems such as speech recognition, speaker identification, audio stream tagging in spoken document retrieval, channel monitoring, or voice conversion. Accent knowledge could be used for selection of alternative pronunciations in a lexicon, engage adaptation for acoustic modeling, or provide information for biasing a language model in large vocabulary speech recognition. In this paper, we propose a text-independent automatic accent classification system using phone-based models. Algorithm formulation begins with a series of experiments focused on capturing the spectral evolution information as potential accent sensitive cues. Alternative subspace representations using principal component analysis and linear discriminant analysis with projected trajectories are considered. Finally, an experimental study is performed to compare the spectral trajectory model framework to a traditional hidden Markov model recognition framework using an accent sensitive word corpus. System evaluation is performed using a corpus representing five English speaker groups with native American English, and English spoken with Mandarin Chinese, French, Thai, and Turkish accents for both male and female speakers.
Pongtep Angkititrakul, John H. L. Hansen
IEEE Trans. Speech Audio Process.2
2006 Speech Enhancement Based on Generalized Minimum Mean Square Error Estimators and Masking Properties of the Auditory System
abstract
In this paper, the family of conditional minimum mean square error (MMSE) spectral estimators is studied which take on the form (E(XEalphap/|Xp+ Dp|))1alpha/, where Xpis the clean speech spectrum, and Dpis the noise spectrum, resulting in a generalized MMSE estimator (GMMSE). The degree of noise suppression versus musical tone artifacts of these estimators is studied. The tradeoffs in selection of (alpha), across noise spectral structure and signal-to-noise ratio (SNR) level, are also considered. Members of this family of estimators include the Ephraim-Malah (EM) amplitude estimator and, for high SNRs, the Wiener Filter. It is shown that the colorless residual noise observed in the EM estimator is a characteristic of this general family of estimators. An application of these estimators in an auditory enhancement scheme using the masking threshold of the human auditory system is formulated, resulting in the GMMSE-auditory masking threshold (AMT) enhancement method. Finally, a detailed evaluation of the proposed algorithms is performed over the phonetically balanced TIMIT database and the National Gallery of the Spoken Word (NGSW) audio archive using subjective and objective speech quality measures. Results show that the proposed GMMSE-AMT outperforms MMSE and log-MMSE enhancement methods using a detailed phoneme-based objective quality analysis
John H. L. Hansen, V. Radhakrishnan, Kathryn Hoberg Arehart
IEEE Trans. Speech Audio Process.1
2006 Advances in unsupervised audio classification and segmentation for the broadcast news and NGSW corpora
abstract
The problem of unsupervised audio classification and segmentation continues to be a challenging research problem which significantly impacts automatic speech recognition (ASR) and spoken document retrieval (SDR) performance. This paper addresses novel advances in 1) audio classification for speech recognition and 2) audio segmentation for unsupervised multispeaker change detection. A new algorithm is proposed for audio classification, which is based on weighted GMM Networks (WGN). Two new extended-time features: variance of the spectrum flux (VSF) and variance of the zero-crossing rate (VZCR) are used to preclassify the audio and supply weights to the output probabilities of the GMM networks. The classification is then implemented using weighted GMM networks. Since historically there have been no features specifically designed for audio segmentation, we evaluate 16 potential features including three new proposed features: perceptual minimum variance distortionless response (PMVDR), smoothed zero-crossing rate (SZCR), and filterbank log energy coefficients (FBLC) in 14 noisy environments to determine the best robust features on the average across these conditions. Next, a new distance metric, T/sup 2/-mean, is proposed which is intended to improve segmentation for short segment turns (i.e., 1-5 s). A new false alarm compensation procedure is implemented, which can compensate the false alarm rate significantly with little cost to the miss rate. Evaluations on a standard data set-Defense Advanced Research Projects Agency (DARPA) Hub4 Broadcast News 1997 evaluation data-show that the WGN classification algorithm achieves over a 50% improvement versus the GMM network baseline algorithm, and the proposed compound segmentation algorithm achieves 23%-10% improvement in all metrics versus the baseline Mel-frequency cepstral coefficients (MFCC) and traditional Bayesian information criterion (BIC) algorithm. The new classification and segmentation algorithms also obtain very satisfactory results on the more diverse and challenging National Gallery of the Spoken Word (NGSW) corpus.
Rongqing Huang, John H. L. Hansen
IEEE Trans. Speech Audio Process.2
2005 Dialect/Accent Classification via Boosted Word Modeling
abstract
The paper addresses novel advances in English dialect/accent classification/identification. A word level based modeling technique is proposed that is shown to outperform a LVCSR based system with significantly less computational cost. The new algorithm, which is named WDC (word-based dialect classification), converts the text independent decision problem into a text dependent problem and produces multiple combination decisions at the word level rather than make a single decision at the utterance level. There are two sets of classifiers employed for WDC, word classifier, D/sub W(k)/, and utterance classifier, D/sub u/. D/sub W(k)/ is boosted via the real AdaBoost.MH algorithm in the probability space directly instead of the feature space. D/sub u/ is boosted via the dialect dependency information of the words. Two dialect corpora are used in the evaluation. Significant improvement in dialect classification is achieved for both corpora.
Rongqing Huang, John H. L. Hansen
ICASSP (1)2
2005 MFCC Compensation for Improved Recognition of Filtered and Band-Limited Speech
abstract
The paper addresses the problem of bandwidth expansion for the purpose of robust speech recognition. We show that an HMM-based ASR engine trained with full spectrum range data (0-8 kHz) can successfully perform speech recognition tasks over band-filtered test data compensated by means of a series of simple MFCC parameter corrector functions. The problem is important when ASR is employed for audio streams of unknown frequency bandwidth, common in spoken document retrieval. Evaluation is based on recognition rates. Accuracy varies depending on the width and spectral regions eliminated, but the system shows great advantages over the use of uncompensated filtered test data. The theoretical maximum recognition rates using corrector functions over filtered test data are very close to the base rate (unfiltered data) even when the greatest part of the spectrum of the original data is suppressed. These rates are even better than those obtained in the matched train/test HMMs with filtered data.
Nicolás Morales, John H. L. Hansen, Doroteo T. Toledano
ICASSP (1)2
2005 Effects of Phoneme Characteristics on TEO Feature-based Automatic Stress Detection in Speech
abstract
A major challenge of automatic speech recognition systems found in many areas of today's society is the ability to overcome natural phoneme conditions that potentially degrade performance. In this study, we discuss the effects of two critical phoneme characteristics, decreased vowel duration and mismatched vowel type, on the performance of automatic stress detection in speech using Teager energy operator features. We determine the scope and magnitude of these effects on stress detection performance and propose an algorithm to compensate for vowel type and duration shortening on stress detection performance using a composite phoneme decision scheme, which results in relative error reductions of 24% and 39% in the non-stress and stress conditions, respectively.
Evan Ruzanski, John H. L. Hansen, James Meyerhoff, George Saviolakis, Michael Koenig
ICASSP (1)2
2005 Towards an Intelligent Acoustic Front-End for Automatic Speech Recognition: Built-In Speaker Normalization (BISN)
abstract
Much effort has transpired over the past three decades in the formulation of "ideal" acoustic features which represent the speech signal in a discriminative and compact manner while being robust to adverse conditions and invariant to speaker differences. A good way of making ASR systems invariant to speaker differences is to perform speaker normalization on the input features. The most popular speaker normalization technique is the vocal tract length normalization (VTLN). However, its implementation requires immense computational resources and is not practically applicable in real-time/embedded ASR systems. In this paper, we propose a new speaker normalization algorithm entitled built-in speaker normalization (BISN) which is performed on-the-fly within the newly proposed PMVDR acoustic front-end and reduces computational resources significantly enabling its use within contemporary ASR systems. Evaluations using an in-car extended digit recognition task showed that on-the-fly implementation of the BISN algorithm produced a relative word error rate (WER) reduction of 24% compared to a no speaker normalization baseline.
Umit H. Yapanel, John H. L. Hansen
ICASSP (1)2
2005 Collaborative voice activity detection for hearing aids
Louisa Busca Grisoni, John H. L. Hansen
INTERSPEECH2
2005 Advances in word based dialect/accent classification
Rongqing Huang, John H. L. Hansen
INTERSPEECH2
2005 Statistical class-based MFCC enhancement of filtered and band-limited speech for robust ASR
abstract
Proceedings of Interspeech-Eurospeech 2005, Lisbon (Portugal)
Nicolás Morales, Doroteo T. Toledano, John H. L. Hansen, José Colás Pasamontes, Javier Garrido Salas
INTERSPEECH3
2005 Improved "TEO" feature-based automatic stress detection using physiological and acoustic speech sensors
Evan Ruzanski, John H. L. Hansen, Don Finan, James Meyerhoff, William Norris 0001, Terry Wollert
INTERSPEECH2
2005 In-set/out-of-set speaker identification based on discriminative speech frame selection
Xianxian Zhang, John H. L. Hansen
INTERSPEECH2
2005 Speaker verification using Gaussian mixture models within changing real car environments
Xianxian Zhang, John H. L. Hansen, Pongtep Angkititrakul, Kazuya Takeda
INTERSPEECH2
2005 SpeechFind: Advances in Spoken Document Retrieval for a National Gallery of the Spoken Word
abstract
Advances in formulating spoken document retrieval for a new National Gallery of the Spoken Word (NGSW) are addressed. NGSW is the first large-scale repository of its kind, consisting of speeches, news broadcasts, and recordings from the 20th century. After presenting an overview of the audio stream content of the NGSW, with sample audio files from U.S. Presidents from 1893 to the present, an overall system diagram is proposed with a discussion of critical tasks associated with effective audio information retrieval. These include advanced audio segmentation, speech recognition model adaptation for acoustic background noise and speaker variability, and information retrieval using natural language processing for text query requests that include document and query expansion. For segmentation, a new evaluation criterion entitled fused error score (FES) is proposed, followed by application of the CompSeg segmentation scheme on DARPA Hub4 Broadcast News (30.5% relative improvement in FES) and NGSW data. Transcript generation is demonstrated for a six-decade portion of the NGSW corpus. Novel model adaptation using structure maximum likelihood eigenspace mapping shows a relative 21.7% improvement. Issues regarding copyright assessment and metadata construction are also addressed for the purposes of a sustainable audio collection of this magnitude. Advanced parameter-embedded watermarking is proposed with evaluations showing robustness to correlated noise attacks. Our experimental online system entitled "SpeechFind" is presented, which allows for audio retrieval from a portion of the NGSW corpus. Finally, a number of research challenges such as language modeling and lexicon for changing time periods, speaker trait and identification tracking, as well as new directions, are discussed in order to address the overall task of robust phrase searching in unrestricted audio corpora.
John H. L. Hansen, Rongqing Huang, Michael Seadle, John R. Deller Jr., Aparna Gurijala, Mikko Kurimo, Pongtep Angkititrakul
IEEE Trans. Speech Audio Process.1
2005 Efficient audio stream segmentation via the combined T2 statistic and Bayesian information criterion
abstract
In many speech and audio applications, it is first necessary to partition and classify acoustic events prior to voice coding for communication or speech recognition for spoken document retrieval. In this paper, we propose an efficient approach for unsupervised audio stream segmentation and clustering via the Bayesian Information Criterion (BIC). The proposed method extends an earlier formulation by Chen and Gopalakrishnan. In our formulation, Hotelling's T/sup 2/-Statistic is used to pre-select candidate segmentation boundaries followed by BIC to perform the segmentation decision. The proposed algorithm also incorporates a variable-size increasing window scheme and a skip-frame test. Our experiments show that we can improve the final algorithm speed by a factor of 100 compared to that in Chen and Gopalakrishnan's while achieving a 6.7% reduction in the acoustic boundary miss rate at the expense of a 5.7% increase in false alarm rate using DARPA Hub4 1997 evaluation data. The approach is particularly successful for short segment turns of less than 2 s in duration. The results suggest that the proposed algorithm is sufficiently effective and efficient for audio stream segmentation applications.
John H. L. Hansen
IEEE Trans. Speech Audio Process.2
2005 Rapid discriminative acoustic model based on eigenspace mapping for fast speaker adaptation
abstract
It is widely believed that strong correlations exist across an utterance as a consequence of time-invariant characteristics of speaker and acoustic environments. It is verified in this paper that the first primary eigendirections of the utterance covariance matrix are speaker dependent. Based on this observation, a novel family of fast speaker adaptation algorithms entitled Eigenspace Mapping (EigMap) is proposed. The proposed algorithms are applied to continuous density Hidden Markov Model (HMM) based speech recognition. The EigMap algorithm rapidly constructs discriminative acoustic models in the test speaker's eigenspace by preserving discriminative information learned from baseline models in the directions of the test speaker's eigenspace. Moreover, the adapted models are compressed by discarding model parameters that are assumed to contain no discrimination information. The core idea of EigMap can be extended in many ways, and a family of algorithms based on EigMap is described in this paper. Unsupervised adaptation experiments show that EigMap is effective in improving baseline models using very limited amounts of adaptation data with superior performance to conventional adaptation techniques such as MLLR and block diagonal MLLR. A relative improvement of 18.4% over a baseline recognizer is achieved using EigMap with only about 4.5 s of adaptation data. Furthermore, it is also demonstrated that EigMap is additive to MLLR by encompassing important speaker dependent discriminative information. A significant relative improvement of 24.6% over baseline is observed using 4.5 s of adaptation data by combining MLLR and EigMap techniques.
John H. L. Hansen
IEEE Trans. Speech Audio Process.2
2004 Identifying in-set and out-of-set speakers using neighborhood information
abstract
We study the problem of identifying in-set and out-of-set speakers. The goal is to identify whether an unknown input speaker belongs to either a group of in-set speakers or an unseen out-of-set group. A state-of-the-art GMM classifier, with universal background model (UBM) and standard likelihood ratio test, is used as our baseline system. We propose an alternative hypothesis testing method that employs neighborhood information with respect to each in-set speaker model in the model space based on the Kullback-Leibier divergence. The Bayes factor is used in the verification stage (accept/reject hypothesis). We evaluate the proposed procedure on a clean CORPUS 1 set, and a noisy CORPUS 2 set which contains session-to-session variability. Experiments show an improvement in equal error rate for the system even when in-set speaker models are acoustically close in the model space, and as the size of the in-set speaker group increases.
Pongtep Angkititrakul, John H. L. Hansen
ICASSP (1)2
2004 Advances in unsupervised audio segmentation for the broadcast news and NGSW corpora
abstract
The problem of unsupervised audio segmentation continues to be a challenging research problem which significantly impacts automatic speech recognition (ASR) and spoken document retrieval (SDR) performance. This paper addresses novel advances in audio segmentation for unsupervised multi-speaker change detection. First, we investigate new features which are intended to be more appropriate for segmentation that include: PMVDR (perceptual minimum variance distortionless response), SZCR ( smoothed zero crossing rate), and FBLC (filterbank log coefficients); next we consider a new distance metric, T/sup 2/-mean which is intended to improve segmentation for short segments (<5s). A novel false alarm compensation procedure is also developed and used after the segmentation phase. We establish a more effective evaluation procedure for segmentation versus the more traditional EER and frame accuracy approaches. Employing these advances within our new scheme, results in more than a 30% improvement in segmentation performance using the 3-hour Hub4 broadcast news 1997 evaluation data. Evaluations are also presented for audio from the NGSW corpus.
Rongqing Huang, John H. L. Hansen
ICASSP (1)2
2004 Speech enhancement based on a combined multi-channel array with constrained iterative and auditory masked processing
abstract
While a number of studies have investigated various speech enhancement and noise suppression schemes, most consider either a single channel or array processing framework. Clearly there are potential advantages in leveraging the strengths of array processing solutions in suppressing noise from a direction other than the speaker, with that seen in single channel methods that include speech spectral constraints or psychoacoustically motivated processing. In this paper, we propose to integrate a combined fixed/adaptive beamforming algorithm (CFA-BF) for speech enhancement with two single channel methods based on speech spectral constrained iterative processing (Auto-LSP), and an auditory masked threshold based method using equivalent rectangular bandwidth filtering (GMMSE-AMTERB). After formulating the method, we evaluate performance on a subset of the TIMIT corpus with four real noise sources. We demonstrate a consistent level of noise suppression and voice communication quality improvement using the proposed method as reflected by an overall average 26dB increase in SegSNR from the original degraded audio corpus.
Xianxian Zhang, John H. L. Hansen, Kathryn Hoberg Arehart
ICASSP (1)2
2004 Cluster-dependent modeling and confidence measure processing for in-set/out-of-set speaker identification
abstract
In this paper, we propose an approach to address the problem of text-independent open-set speaker identification. The in-set speakers are clustered into smaller subsets without merging speaker models. The Anti-Speaker or Background Model is then adapted for each subset which minimizes the identification errors of the pseudo impostors during the training stage. Score normalization is applied to align all the in-set speaker score distributions to share a single scale. Finally, confidence measure processing is used to identify in-set versus out-of-set speakers. Experiments with TIMIT and the CU-Accent corpora show an improvement in Equal Error Rate on the average of 20.28 and 8.35 over the baseline performance respectively. Finally, a probe experiment is also included that considers prosody for in-set speaker detection.
Pongtep Angkititrakul, Sepideh Baghaii, John H. L. Hansen
INTERSPEECH3
2004 Dialect analysis and modeling for automatic classification
abstract
In this paper, we present our recent work in the analysis and modeling of speech under dialect. Dialect and accent significantly influence automatic speech recognition performance, and therefore it is critical to detect and classify non-native speech. In this study, we consider three areas that include: (i) prosodic structure (normalized f0, syllable rate, and sentence duration), (ii) phoneme acoustic space modeling and sub-word classification, and (iii) word-level based modeling using large vocabulary data. The corpora used in this study include: the NATO N-4 corpus (2 accents, 2 dialects of English), TIMIT (7 dialect regions), and American and British English versions of the WSJ corpus. These corpora were selected because the contained audio material from specific dialects/accents of English (N-4), were phonetically balanced and organized across U.S. (TIMIT), or contained significant amounts of read audio material from distinct dialects (WSJ). The results show that significant changes occur at the prosodic, phoneme space, and word levels for dialect analysis, and that effective dialect classification can be achieved using processing strategies from each domain.
John H. L. Hansen, Umit H. Yapanel, Rongqing Huang, Ayako Ikeno
INTERSPEECH1
2004 High-level feature weighted GMM network for audio stream classification
abstract
The problem of unsupervised audio classification con-tinuous to be a challenging research problem which sig-nificantly impacts ASR and Spoken Document Retrieval (SDR) performance. This paper addresses novel ad-vances in audio classification for speech recognition. A new algorithm is proposed for audio classification, which is based on Weighted GMM Network (WGN). Two new high-level features: VSF (Variance of the Spectrum Flux) and VZCR (Variance of the Zero-Crossing Rate) are used to pre-classify the audio and supply weights to the output probabilities of the GMM networks. The classification is then implemented using weighted GMM networks. Eval-uations on a standard data set — DARPA Hub4 Broadcast News 1997 evaluation data, shows that the WGN classi-fication algorithm achieves over a 50 % improvement ver-sus the GMM network baseline algorithm. The WGN also obtains very satisfactory results on the more diverse and challenging NGSW (National Gallery of the Spoken Word [8]) corpus. Classification based on segmentation method is also explored. 1.
Rongqing Huang, John H. L. Hansen
INTERSPEECH2
2004 In-vehicle based speech processing for hearing impaired subjects
Xianxian Zhang, John H. L. Hansen, Kathryn Hoberg Arehart, Jessica Rossi-Katz
INTERSPEECH2
2004 Audio-visual SPeaker localization for car navigation systems
abstract
Human-computer interaction for in-vehicle information and navigation systems is a challenging problem because of the diverse and changing acoustic environments. It is proposed that the integration of video and audio information can significantly improve dialog system performance, since the visual modality is not impacted by acoustic noise. In this paper, we propose a robust audio-visual integration system for source tracking and speech enhancement for an in-vehicle speech dialog system. The proposed system integrates both audio and visual information to locate the desired speaker source. Using real data collected in car environments, the proposed system can improve speech accuracy by up to 40.75% compared with audio data alone.
Xianxian Zhang, Kazuya Takeda, John H. L. Hansen, Toshiki Maeno
INTERSPEECH3
2003 Environmental sniffing: noise knowledge estimation for robust speech systems
abstract
We propose a framework for extracting knowledge about environmental noise from an input audio sequence and organizing this knowledge for use by other speech systems. To date, most approaches dealing with environmental noise in speech systems are based on assumptions about the noise, or differences in the collection of and training on a specific noise condition, rather than exploring the nature of the noise. We are interested in constructing a new speech framework, entitled environmental sniffing, to detect, classify and track acoustic environmental conditions. The first goal of the framework is to seek out detailed information about the environmental characteristics instead of just detecting environmental changes. The second goal is to organize this knowledge in an effective manner to allow smart decisions to direct other speech systems. Our current framework uses a number of speech processing modules including the Teager energy operator (TEO) and a hybrid algorithm with T/sup 2/-BIC segmentation, noise language modeling and GMM classification in noise knowledge estimation. We define a new information criterion that incorporates the impact of noise on environmental sniffing performance. We use an in-vehicle speech and noise environment as a test platform for our evaluations and investigate the integration of environmental sniffing into an automatic speech recognition (ASR) engine in this environment. Noise classification experiments show that the hybrid algorithm achieves an error rate of 25.51%, outperforming a baseline system by an absolute 7.08%.
Murat Akbacak, John H. L. Hansen
ICASSP (2)2
2003 CSA-BF: novel constrained switched adaptive beamforming for speech enhancement & recognition in real car environments
abstract
While a number of studies have investigated various speech enhancement and processing schemes for in-vehicle speech systems, little research has been performed using actual voice data collected in noisy car environments. We propose a new constrained switched adaptive beamforming algorithm (CSA-BF) for speech enhancement and recognition in real moving car environments. The proposed algorithm consists of a speech/noise constraint section, a speech adaptive beamformer, and a noise adaptive beamformer. We investigate CSA-BF performance with a comparison to classic delay-and-sum beamforming (DASB) in realistic car environments using a large quantity of data recorded in various car noise environments from across the United States. After analyzing the experimental results and considering the range of complex noise situations in the car environment using the CU-Move corpus, we formulate the CSA-BF algorithm. This method is shown to decrease WER (word error rate) for speech recognition by up to 31% and improve speech quality via the SEGSNR (segment signal-to-noise ratio) by up to 5.5 dB on the average, simultaneously.
Xianxian Zhang, John H. L. Hansen
ICASSP (2)2
2003 Discriminative acoustic model using eigenspace mapping for rapid speaker adaptation
abstract
It is widely believed that strong correlations exist across an utterance as a consequence of time-invariant characteristics of speaker and acoustic environments. It is verified in this paper that the first primary eigendirections of the utterance covariance matrix are speaker dependent. Based on this observation, a fast speaker adaptation algorithm entitled Eigenspace Mapping (EigMap) is proposed and described. EigMap rapidly adapts the speaker independent models by constructing discriminative acoustic models in the test speaker's eigenspace. Unsupervised adaptation experiments show that EigMap is effective in improving baseline models using very limited amounts of adaptation data with superior performance to conventional adaptation technique such as block diagonal MLLR. A relative improvement of 18.4% over baseline recognizer is achieved using EigMap with only about 4.5 seconds of adaptation data. It is also demonstrated that EigMap is additive to MLLR by encompassing the speaker dependent discrimination information. A significant relative improvement of 24.6% over baseline is observed by combining MLLR and EigMap techniques.
John H. L. Hansen
ICASSP (1)2
2003 Environmental sniffing: robust digit recognition for an in-vehicle environment
Murat Akbacak, John H. L. Hansen
INTERSPEECH2
2003 Use of trajectory models for automatic accent classification
Pongtep Angkititrakul, John H. L. Hansen
INTERSPEECH2
2003 Perceptual based speech enhancement for normal-hearing and hearing-impaired individuals
Ajay Natarajan, John H. L. Hansen, Kathryn Hoberg Arehart, Jessica Rossi-Katz
INTERSPEECH2
2003 Frequency distribution based weighted sub-band approach for classification of emotional/stressful content in speech
Mandar Rahurkar, John H. L. Hansen
INTERSPEECH2
2003 Perceptual MVDR-based cepstral coefficients (PMCCs) for high accuracy speech recognition
Umit H. Yapanel, Satya Dharanipragada, John H. L. Hansen
INTERSPEECH3
2003 A new perspective on feature extraction for robust in-vehicle speech recognition
abstract
The problem of reliable speech recognition for in-vehicle applications has recently emerged as a challenging research domain. This study focuses on the feature extraction stage of this problem. The approach is based on MinimumVariance Distortionless Response (MVDR) spectrum estimation. MVDR is used for robustly estimating the envelope of the speech signal and shown to be very accurate and relatively less sensitive to additive noise. The proposed feature estimation process removes the traditional Mel-scaled filterbank as a perceptually motivated frequency partitioning. Instead, we directly warp the FFT power spectrum of speech. The word error rate (WER) is shown to decrease by 27.3 % with respect to the MFCCs and 18.8 % with respect to recently proposed PMCCs on an extended digit recognition task in real car environments. The proposed feature estimation approach is called PMVDR and conclusively shown to be a better speech representation in real environments with emphasis on time-varying car noise. 1.
Umit H. Yapanel, John H. L. Hansen
INTERSPEECH2
2003 CFA-BF: a novel combined fixed/adaptive beamforming for robust speech recognition in real car environments
Xianxian Zhang, John H. L. Hansen
INTERSPEECH2
2003 Evaluation of an auditory masked threshold noise suppression algorithm in normal-hearing and hearing-impaired listeners
Kathryn Hoberg Arehart, John H. L. Hansen, Stephen Gallant, Laura Kalstein
Speech Commun.2
2003 CSA-BF: a constrained switched adaptive beamformer for speech enhancement and recognition in real car environments
abstract
While a number of studies have investigated various speech enhancement and processing schemes for in-vehicle speech systems, little research has been performed using actual voice data collected in noisy car environments. In this paper, we propose a new constrained switched adaptive beamforming algorithm (CSA-BF) for speech enhancement and recognition in real moving car environments. The proposed algorithm consists of a speech/noise constraint section, a speech adaptive beamformer, and a noise adaptive beamformer. We investigate CSA-BF performance with a comparison to classic delay-and-sum beamforming (DASB) in realistic car conditions using a corpus of data recorded in various car noise environments from across the U.S. After analyzing the experimental results and considering the range of complex noise situations in the car environment using the CU-Move corpus, we formulate the three specific processing stages of the CSA-BF algorithm. This method is evaluated and shown to simultaneously decrease word-error-rate (WER) for speech recognition by up to 31% and improve speech quality via the SEGSNR measure by up to +5.5 dB on the average.
Xianxian Zhang, John H. L. Hansen
IEEE Trans. Speech Audio Process.2
2002 Application of automatic speech recognition in call classification
abstract
Call classification is the process of characterizing the audio signals encountered when a phone call is placed. The signal can be a live person, an answering machine, a call progress tone (ringback, busy, etc.), a fax modem tone or an announcement message. Call processing system uses the output of the call classifier to perform automated call routing. Traditionally, in-band signaling between communication system endpoints and users are in the form of audible tones. With the advent of technology, these signals have been augmented by synthesized and recorded human speech. A “busy tone”, for example, may be replaced by recorded speech. Traditional call classifiers, which are mainly tone detectors, have become increasingly ineffective. This paper proposes a new generation call classifier employing Automatic Speech Recognition (ASR). Modifications of the ASR paradigm are suggested to incorporate signal recognition, and results presented for two call classifier scenarios.
Sharmistha Sarkar Das, Norman Chan, Danny Wages, John H. L. Hansen
ICASSP4
2002 Rapid speaker adaptation using multi-stream Structural Maximum Likelihood Eigenspace Mapping
abstract
In this paper, we extend our previously proposed algorithm entitled Structural Maximum Likelihood Eigenspace Mapping (SMLEM) for rapid speaker adaptation. The SMLEM algorithm directly adapts Speaker Independent (SI) acoustic models to a test speaker by mapping the mixture Gaussian components from a SI eigenspace to Speaker Dependent (SD) eigenspaces in a maximum likelihood manner, with very limited adaptation data. In previous SMLEM paper, we presented encouraging results for SMLEM by adapting only the static feature components. In this paper, we propose a multi-stream approach where the static and dynamic feature streams are adapted. For small amounts of adaptation data ranging from 15 to 50 seconds, superior performance is demonstrated over both standard MLLR and block diagonal MLLR.
John H. L. Hansen
ICASSP2
2002 Stochastic trajectory model analysis for accent classification
abstract
This paper presents recent results using statistics generated by a MMI-supervised vector quantizer as a measure of audio similarity. Such a measure has proved successful for talker identification, and the extension from speech to general audio, such as music, is straightforward. A classifier that distinguishes speech from music and non-vocal sounds is presented, as well as experimental results showing how perfect classification accuracy may be achieved on a small corpus using substantially less than two seconds per test audio file. The techniques a presented here may be extended to other applications and domains, such as audio retrieval-by-similarity, musical genre classification, and automatic segmentation of continuous audio. Introduction This paper presents a method of rapidly determining the characteristics of audio samples, using a supervised tree-based vector quantizer trained to maximise mutual information (MMI). Unlike other approaches based on perceptual criteria (Pfeiffer, F...
Pongtep Angkititrakul, John H. L. Hansen
INTERSPEECH2