EDBT 2026 Demo / reviewers in the wild / expert
Abeer Alwan
dblp:40/6586 · also Abeer A. Alwan
· DBLP profile ↗
196ranked-venue papers
6as first author
33since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 168 · 6 first-author · 28 since 2021Artificial intelligence and machine learning · 117 · 2 first-author · 19 since 2021Systems, architecture and hardware · 3Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Compositional domain adaptation for automatic speech recognition with headwise selective attention merging
Natarajan Balaji Shankar, Zilai Wang, Eray Eren, Abeer Alwan |
Comput. Speech Lang. | 4 |
| 2025 | Selective Attention Merging for low resource tasks: A case study of Child ASRabstractWhile Speech Foundation Models (SFMs) excel in various speech tasks, their performance for low-resource tasks such as child Automatic Speech Recognition (ASR) is hampered by limited pretraining data. To address this, we explore different model merging techniques to leverage knowledge from models trained on larger, more diverse speech corpora. This paper also introduces Selective Attention (SA) Merge, a novel method that selectively merges task vectors from attention matrices to enhance SFM performance on low-resource tasks. Experiments on the MyST database show significant reductions in relative word error rate of up to 14%, outperforming existing model merging and data augmentation techniques. By combining data augmentation techniques with SA Merge, we achieve a new state-of-the-art WER of 8.69 on the MyST database for the Whisper-small model, highlighting the potential of SA Merge for improving low-resource ASR. Natarajan Balaji Shankar, Zilai Wang, Eray Eren, Abeer Alwan |
ICASSP | 4 |
| 2025 | Enhancing Age-Related Robustness in Children Speaker VerificationabstractOne of the main challenges in children’s speaker verification (C-SV) is the significant change in children’s voices as they grow. In this paper, we propose two approaches to improve age-related robustness in C-SV. We first introduce a Feature Transform Adapter (FTA) module that integrates local patterns into higher-level global representations, reducing overfitting to specific local features and improving the inter-year SV performance of the system. We then employ Synthetic Audio Augmentation (SAA) to increase data diversity and size, thereby improving robustness against age-related changes. Since the lack of longitudinal speech datasets makes it difficult to measure age-related robustness of C-SV systems, we introduce a longitudinal dataset to assess inter-year verification robustness of C-SV systems. By integrating both of our proposed methods, the average equal error rate was reduced by 19.4%, 13.0%, and 6.1% in the one-year, two-year, and three-year gap inter-year evaluation sets, respectively, compared to the baseline. Vishwas M. Shetty, Jiusi Zheng, Steven M. Lulich, Abeer Alwan |
ICASSP | 4 |
| 2025 | ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs
Eray Eren, Qingju Liu, Hyeongwoo Kim, Pablo Garrido 0001, Abeer Alwan |
INTERSPEECH | 5 |
| 2025 | CHSER: A Dataset and Case Study on Generative Speech Error Correction for Child ASR
Natarajan Balaji Shankar, Zilai Wang, Mohan Shi, Abeer Alwan |
INTERSPEECH | 5 |
| 2025 | The Ohio Child Speech CorpusabstractThis paper reports on the creation and composition of a new corpus of children's speech, the Ohio Child Speech Corpus, which is publicly available on the Talkbank-CHILDES website. The audio corpus contains speech samples from 303 children ranging in age from 4 – 9 years old, all of whom participated in a seven-task elicitation protocol conducted in a science museum lab. In addition, an interactive social robot controlled by the researchers joined the sessions for approximately 60% of the children, and the corpus itself was collected in the peri‑pandemic period. Two analyses are reported that highlighted these last two features. One set of analyses found that the children spoke significantly more in the presence of the robot relative to its absence, but no effects of speech complexity (as measured by MLU) were found for the robot's presence. Another set of analyses compared children tested immediately post-pandemic to children tested a year later on two school-readiness tasks, an Alphabet task and a Reading Passages task. This analysis showed no negative impact on these tasks for our highly-educated sample of children just coming off of the pandemic relative to those tested later. These analyses demonstrate just two possible types of questions that this corpus could be used to investigate. Sharifa Alghowinem, Abeer Alwan, Kristina Bowdrie, Cynthia Breazeal, Cynthia G. Clopper, Eric Fosler-Lussier, Izabela A. Jamsek, Devan Lander, Rajiv Ramnath, Jory Ross |
Speech Commun. | 3 |
| 2024 | CORAAL QA: A Dataset and Framework for Open Domain Spontaneous Speech Question Answering from Long Audio FilesabstractThis paper presents a novel dataset (CORAAL QA) and framework for audio question-answering from long audio recordings containing spontaneous speech. The dataset introduced here provides sets of questions that can be factually answered from short spans of a long audio files (typically 30min to 1hr) from the Corpus of Regional African American Language. Using this dataset, we divide the audio recordings into 60 second segments, automatically transcribe each segment, and use PLDA scoring of BERT-based semantic embeddings to rank the relevance of ASR transcript segments in answering the target question. In order to improve this framework through data augmentation, we use large language models including ChatGPT and Llama 2 to automatically generate further training examples and show how prompt engineering can be optimized for this process. By creatively leveraging knowledge from large-language models, we achieve state-of-the-art question-answering performance in this information retrieval task. Natarajan Balaji Shankar, Alexander Johnson, Christina Chance, Hariram Veeramani, Abeer Alwan |
ICASSP | 5 |
| 2024 | Benchmarking Children's ASR with Supervised and Self-supervised Speech Foundation Models
Ruchao Fan, Natarajan Balaji Shankar, Abeer Alwan |
INTERSPEECH | 3 |
| 2024 | Enhancing accuracy and privacy in speech-based depression detection through speaker disentanglementabstractSpeech signals are valuable biomarkers for assessing an individual’s mental health, including identifying Major Depressive Disorder (MDD) automatically. A frequently used approach in this regard is to employ features related to speaker identity, such as speaker-embeddings. However, over-reliance on speaker identity features in mental health screening systems can compromise patient privacy. Moreover, some aspects of speaker identity may not be relevant for depression detection and could serve as a bias factor that hampers system performance. To overcome these limitations, we propose disentangling speaker-identity information from depression-related information. Specifically, we present four distinct disentanglement methods to achieve this - adversarial speaker identification (SID)-loss maximization (ADV), SID-loss equalization with variance (LEV), SID-loss equalization using Cross-Entropy (LECE) and SID-loss equalization using KL divergence (LEKLD). Our experiments, which incorporated diverse input features and model architectures, have yielded improved F1 scores for MDD detection and voice-privacy attributes, as quantified by Gain in Voice Distinctiveness (GVD) and De-Identification Scores (DeID). On the DAIC-WOZ dataset (English), LECE using ComparE16 features results in the best F1-Scores of 80% which represents the audio-only SOTA depression detection F1-Score along with a GVD of −1.1 dB and a DeID of 85%. On the EATD dataset (Mandarin), ADV using raw-audio signal achieves an F1-Score of 72.38% surpassing multi-modal SOTA along with a GVD of −0.89 dB dB and a DeID of 51.21%. By reducing the dependence on speaker-identity-related features, our method offers a promising direction for speech-based depression detection that preserves patient privacy. Vijay Ravi, Jinhan Wang, Jonathan Flint, Abeer Alwan |
Comput. Speech Lang. | 4 |
| 2024 | Speechformer-CTC: Sequential modeling of depression detection with speech temporal classificationabstractSpeech-based automatic depression detection systems have been extensively explored over the past few years. Typically, each speaker is assigned a single label (Depressive or Non-depressive), and most approaches formulate depression detection as a speech classification task without explicitly considering the non-uniformly distributed depression pattern within segments, leading to low generalizability and robustness across different scenarios. However, depression corpora do not provide fine-grained labels (at the phoneme or word level) which makes the dynamic depression pattern in speech segments harder to track using conventional frameworks. To address this, we propose a novel framework, Speechformer-CTC, to model non-uniformly distributed depression characteristics within segments using a Connectionist Temporal Classification (CTC) objective function without the necessity of input-output alignment. Two novel CTC-label generation policies, namely the Expectation-One-Hot and the HuBERT policies, are proposed and incorporated in objectives on various granularities. Additionally, experiments using Automatic Speech Recognition (ASR) features are conducted to demonstrate the compatibility of the proposed method with content-based features. Our results show that the performance of depression detection, in terms of Macro F1-score, is improved on both DAIC-WOZ (English) and CONVERGE (Mandarin) datasets. On the DAIC-WOZ dataset, the system with HuBERT ASR features and a CTC objective optimized using HuBERT policy for label generation achieves 83.15% F1-score, which is close to state-of-the-art without the need for phoneme-level transcription or data augmentation. On the CONVERGE dataset, using Whisper features with the HuBERT policy improves the F1-score by 9.82% on CONVERGE1 (in-domain test set) and 18.47% on CONVERGE2 (out-of-domain test set). These findings show that depression detection can benefit from modeling non-uniformly distributed depression patterns and the proposed framework can be potentially used to determine significant depressive regions in speech utterances. Jinhan Wang, Vijay Ravi, Jonathan Flint, Abeer Alwan |
Speech Commun. | 4 |
| 2024 | UniEnc-CASSNAT: An Encoder-Only Non-Autoregressive ASR for Speech SSL ModelsabstractNon-autoregressive automatic speech recognition (NASR) models have gained attention due to their parallelism and fast inference. The encoder-based NASR, e.g. connectionist temporal classification (CTC), can be initialized from the speech foundation models (SFM) but does not account for any dependencies among intermediate tokens. The encoder-decoder-based NASR, like CTC alignment-based single-step non-autoregressive transformer (CASS-NAT), can mitigate the dependency problem but is not able to efficiently integrate SFM. Inspired by the success of recent work of speech-text joint pre-training with a shared transformer encoder, we propose a new encoder-based NASR, UniEnc-CASSNAT, to combine the advantages of CTC and CASS-NAT. UniEnc-CASSNAT consists of only an encoder as the major module, which can be the SFM. The encoder plays the role of both the CASS-NAT encoder and decoder by two forward passes. The first pass of the encoder accepts the speech signal as input, while the concatenation of the speech signal and the token-level acoustic embedding is used as the input for the second pass. Examined on the Librispeech 100h, MyST, and Aishell1 datasets, the proposed UniEnc-CASSNAT achieves state-of-the-art NASR results and is better or comparable to CASS-NAT with only an encoder and hence, fewer model parameters. Our codes1are publicly available. Ruchao Fan, Natarajan Balaji Shankar, Abeer Alwan |
IEEE Signal Process. Lett. | 3 |
| 2023 | Leveraging Multiple Sources in Automatic African American English Dialect Detection for Adults and ChildrenabstractThis paper1presents a novel system which utilizes acoustic, phonological, morphosyntactic, and prosodic information for binary automatic dialect detection of African American English. We train this system utilizing adult speech data and then evaluate on both children’s and adults’ speech with unmatched training and testing scenarios. The proposed system combines novel and state-of-the-art architectures, including a multi-source transformer language model pre-trained on Twitter text data and fine-tuned on ASR transcripts as well as an LSTM acoustic model trained on self-supervised learning representations, in order to learn a comprehensive view of dialect. We show robust, explainable performance across recording conditions for different features for adult speech, but fusing multiple features is important for good results on children’s speech. Alexander Johnson, Vishwas M. Shetty, Mari Ostendorf, Abeer Alwan |
ICASSP | 4 |
| 2023 | FusedF0: Improving DNN-based F0 Estimation by Fusion of Summary-Correlograms and Raw Waveform Representations of Speech Signals
Eray Eren, Lee Ngee Tan, Abeer Alwan |
INTERSPEECH | 3 |
| 2023 | An Equitable Framework for Automatically Assessing Children's Oral Narrative Language Abilities
Alexander Johnson, Hariram Veeramani, Natarajan Balaji Shankar, Abeer Alwan |
INTERSPEECH | 4 |
| 2023 | Developmental Articulatory and Acoustic Features for Six to Ten Year Old Children
Vishwas M. Shetty, Steven M. Lulich, Abeer Alwan |
INTERSPEECH | 3 |
| 2023 | Non-uniform Speaker Disentanglement For Depression Detection From Raw Speech SignalsabstractWhile speech-based depression detection methods that use speaker-identity features, such as speaker embeddings, are popular, they often compromise patient privacy. To address this issue, we propose a speaker disentanglement method that utilizes a non-uniform mechanism of adversarial SID loss maximization. This is achieved by varying the adversarial weight between different layers of a model during training. We find that a greater adversarial weight for the initial layers leads to performance improvement. Our approach using the ECAPA-TDNN model achieves an F1-score of 0.7349 (a 3.7% improvement over audio-only SOTA) on the DAIC-WoZ dataset, while simultaneously reducing the speaker-identification accuracy by 50%. Our findings suggest that identifying depression through speech signals can be accomplished without placing undue reliance on a speaker's identity, paving the way for privacy-preserving approaches of depression detection. Jinhan Wang, Vijay Ravi, Abeer Alwan |
INTERSPEECH | 3 |
| 2023 | A CTC Alignment-Based Non-Autoregressive Transformer for End-to-End Automatic Speech RecognitionabstractRecently, end-to-end models have been widely used in automatic speech recognition (ASR) systems. Two of the most representative approaches are connectionist temporal classification (CTC) and attention-based encoder-decoder (AED) models. Autoregressive transformers, variants of AED, adopt an autoregressive mechanism for token generation and thus are relatively slow during inference. In this paper, we present a comprehensive study of a CTC Alignment-based Single-Step Non-Autoregressive Transformer (CASS-NAT) for end-to-end ASR. In CASS-NAT, word embeddings in the autoregressive transformer (AT) are substituted with token-level acoustic embeddings (TAE) that are extracted from encoder outputs with the acoustical boundary information offered by the CTC alignment. TAE can be obtained in parallel, resulting in a parallel generation of output tokens. During training, Viterbi-alignment is used for TAE generation, and multiple training strategies are further explored to improve the word error rate (WER) performance. During inference, an error-based alignment sampling method is investigated in depth to reduce the alignment mismatch in the training and testing processes. Experimental results show that the CASS-NAT has a WER that is close to AT on various ASR tasks, while providing a$\sim$24x inference speedup. With and without self-supervised learning, we achieve new state-of-the-art results for non-autoregressive models on several datasets. We also analyze the behavior of the CASS-NAT decoder to explain why it can perform similarly to AT. We find that TAEs have similar functionality to word embeddings for grammatical structures, which might indicate the possibility of learning some semantic information from TAEs without a language model. Ruchao Fan, Peng Chang 0002, Abeer Alwan |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | Can Social Robots Effectively Elicit Curiosity in STEM Topics from K-1 Students During Oral Assessments?abstractThis paper presents the results of a pilot study that introduces social robots into kindergarten and first-grade classroom tasks. This study aims to understand 1) how effective social robots are in administering educational activities and assessments, and 2) if these interactions with social robots can serve as a gateway into learning about robotics and STEM for young children. We administered a commonly-used assessment (GFTA3) of speech production using a social robot and compared the quality of recorded responses to those obtained with a human assessor. In a comparison done between 40 children, we found no significant differences in the student responses between the two conditions over the three metrics used: word repetition accuracy, number of times additional help was needed, and similarity of prosody to the assessor. We also found that interactions with the robot were successfully able to stimulate curiosity in robotics, and therefore STEM, from a large number of the 164 student participants. Alexander Johnson, Alejandra Martin, Marlen Quintero, Alison L. Bailey, Abeer Alwan |
EDUCON | 5 |
| 2022 | LPC Augment: an LPC-based ASR Data Augmentation Algorithm for Low and Zero-Resource Children's DialectsabstractThis paper proposes a novel linear prediction coding-based data augmentation method for children’s low and zero resource dialect ASR. The data augmentation procedure consists of perturbing the formant peaks of the LPC spectrum during LPC analysis and reconstruction. The method is evaluated on two novel children’s speech datasets with one containing California English from the Southern California Area and the other containing a mix of Southern American English and African American English from the Atlanta, Georgia area. We test the proposed method in training both an HMM-DNN system and an end-to-end system to show model-robustness and demonstrate that the algorithm improves ASR performance, especially for zero resource dialect children’s task, as compared to common data augmentation methods such as VTLP, Speed Perturbation, and SpecAugment. Alexander Johnson, Ruchao Fan, Robin Morris, Abeer Alwan |
ICASSP | 4 |
| 2022 | Fraug: A Frame Rate Based Data Augmentation Method for Depression Detection from Speech SignalsabstractIn this paper, a data augmentation method is proposed for depression detection from speech signals. Samples for data augmentation were created by changing the frame-width and the frame-shift parameters during the feature extraction process. Unlike other data augmentation methods (such as VTLP, pitch perturbation, or speed perturbation), the proposed method does not explicitly change acoustic parameters but rather the time-frequency resolution of frame-level features. The proposed method was evaluated using two different datasets, models, and input acoustic features. For the DAIC-WOZ (English) dataset when using the DepAudioNet model and mel-Spectrograms as input, the proposed method resulted in an improvement of 5.97% (validation) and 25.13% (test) when compared to the baseline. The improvements for the CONVERGE (Mandarin) dataset when using the x-vector embeddings with CNN as the backend and MFCCs as input features were 9.32% (validation) and 12.99% (test). Baseline systems do not incorporate any data augmentation. Further, the proposed method outperformed commonly used data-augmentation methods such as noise augmentation, VTLP, Speed, and Pitch Perturbation. All improvements were statistically significant. Vijay Ravi, Jinhan Wang, Jonathan Flint, Abeer Alwan |
ICASSP | 4 |
| 2022 | Towards Better Meta-Initialization with Task Augmentation for Kindergarten-Aged Speech RecognitionabstractChildren’s automatic speech recognition (ASR) is always difficult due to, in part, the data scarcity problem, especially for kindergarten-aged kids. When data are scarce, the model might overfit to the training data, and hence good starting points for training are essential. Recently, meta-learning was proposed to learn model initialization (MI) for ASR tasks of different languages. This method leads to good performance when the model is adapted to an unseen language. How-ever, MI is vulnerable to overfitting on training tasks (learner overfitting). It is also unknown whether MI generalizes to other low-resource tasks. In this paper, we validate the effectiveness of MI in children’s ASR and attempt to alleviate the problem of learner overfitting. To achieve model-agnostic meta-learning (MAML), we regard children’s speech at each age as a different task. In terms of learner overfitting, we propose a task-level augmentation method by simulating new ages using frequency warping techniques. Detailed experiments are conducted to show the impact of task augmentation on each age for kindergarten-aged speech. As a result, our approach achieves a relative word error rate (WER) improvement of 51% over the baseline system with no augmentation or initialization. Yunzheng Zhu, Ruchao Fan, Abeer Alwan |
ICASSP | 3 |
| 2022 | Attention-based conditioning methods using variable frame rate for style-robust speaker verificationabstractWe propose an approach to extract speaker embeddings that are robust to speaking style variations in text-independent speaker verification.Typically, speaker embedding extraction includes training a DNN for speaker classification and using the bottleneck features as speaker representations.Such a network has a pooling layer to transform frame-level to utterance-level features by calculating statistics over all utterance frames, with equal weighting.However, self-attentive embeddings perform weighted pooling such that the weights correspond to the importance of the frames in a speaker classification task.Entropy can capture acoustic variability due to speaking style variations.Hence, an entropy-based variable frame rate vector is proposed as an external conditioning vector for the self-attention layer to provide the network with information that can address style effects.This work explores five different approaches to conditioning.The best conditioning approach, concatenation with gating, provided statistically significant improvements over the x-vector baseline in 12/23 tasks and was the same as the baseline in 11/23 tasks when using the UCLA speaker variability database.It also significantly outperformed self-attention without conditioning in 9/23 tasks and was worse in 1/23.The method also showed significant improvements in multi-speaker scenarios of SITW. Amber Afshan, Abeer Alwan |
INTERSPEECH | 2 |
| 2022 | Learning from human perception to improve automatic speaker verification in style-mismatched conditionsabstractOur prior experiments show that humans and machines seem to employ different approaches to speaker discrimination, especially in the presence of speaking style variability.The experiments examined read versus conversational speech.Listeners focused on speaker-specific idiosyncrasies while "telling speakers together", and on relative distances in a shared acoustic space when "telling speakers apart".However, automatic speaker verification (ASV) systems use the same loss function irrespective of target or non-target trials.To improve ASV performance in the presence of style variability, insights learnt from human perception are used to design a new training loss function that we refer to as "CllrCE loss".CllrCE loss uses both speaker-specific idiosyncrasies and relative acoustic distances between speakers to train the ASV system.When using the UCLA speaker variability database, in the x-vector and conditioning setups, CllrCE loss results in significant relative improvements in EER by 1-66%, and minDCF by 1-31% and 1-56%, respectively, when compared to the x-vector baseline.Using the SITW evaluation tasks, which involve different conversational speech tasks, the proposed loss combined with selfattention conditioning results in significant relative improvements in EER by 2-5% and minDCF by 6-12% over baseline.In the SITW case, performance improvements were consistent only with conditioning. Amber Afshan, Abeer Alwan |
INTERSPEECH | 2 |
| 2022 | DRAFT: A Novel Framework to Reduce Domain Shifting in Self-supervised Learning and Its Application to Children's ASRabstractSelf-supervised learning (SSL) in the pretraining stage using un-annotated speech data has been successful in low-resource automatic speech recognition (ASR) tasks. However, models trained through SSL are biased to the pretraining data which is usually different from the data used in finetuning tasks, causing a domain shifting problem, and thus resulting in limited knowledge transfer. We propose a novel framework, domain responsible adaptation and finetuning (DRAFT), to reduce domain shifting in pretrained speech models through an additional adaptation stage. In DRAFT, residual adapters (RAs) are inserted in the pretrained model to learn domain-related information with the same SSL loss as the pretraining stage. Only RA parameters are updated during the adaptation stage. DRAFT is agnostic to the type of SSL method used and is evaluated with three widely used approaches: APC, Wav2vec2.0, and HuBERT. On two child ASR tasks (OGI and MyST databases), using SSL models trained with un-annotated adult speech data (Librispeech), relative WER improvements of up to 19.7% are observed when compared to the pretrained models without adaptation. Additional experiments examined the potential of cross knowledge transfer between the two datasets and the results are promising, showing a broader usage of the proposed DRAFT framework. Ruchao Fan, Abeer Alwan |
INTERSPEECH | 2 |
| 2022 | Automatic Dialect Density Estimation for African American EnglishabstractIn this paper, we explore automatic prediction of dialect density of the African American English (AAE) dialect, where dialect density is defined as the percentage of words in an utterance that contain characteristics of the non-standard dialect.We investigate several acoustic and language modeling features, including the commonly used X-vector representation and Com-ParE feature set, in addition to information extracted from ASR transcripts of the audio files and prosodic information.To address issues of limited labeled data, we use a weakly supervised model to project prosodic and X-vector features into lowdimensional task-relevant representations.An XGBoost model is then used to predict the speaker's dialect density from these features and show which are most significant during inference.We evaluate the utility of these features both alone and in combination for the given task.This work, which does not rely on hand-labeled transcripts, is performed on audio segments from the CORAAL database.We show a significant correlation between our predicted and ground truth dialect density measures for AAE speech in this database and propose this work as a tool for explaining and mitigating bias in speech technology. Alexander Johnson, Kevin Everson, Vijay Ravi, Anissa Gladney, Mari Ostendorf, Abeer Alwan |
INTERSPEECH | 6 |
| 2022 | A Step Towards Preserving Speakers' Identity While Detecting Depression Via Speaker DisentanglementabstractPreserving a patient's identity is a challenge for automatic, speech-based diagnosis of mental health disorders. In this paper, we address this issue by proposing adversarial disentanglement of depression characteristics and speaker identity. The model used for depression classification is trained in a speaker-identity-invariant manner by minimizing depression prediction loss and maximizing speaker prediction loss during training. The effectiveness of the proposed method is demonstrated on two datasets - DAIC-WOZ (English) and CONVERGE (Mandarin), with three feature sets (Mel-spectrograms, raw-audio signals, and the last-hidden-state of Wav2vec2.0), using a modified DepAudioNet model. With adversarial training, depression classification improves for every feature when compared to the baseline. Wav2vec2.0 features with adversarial learning resulted in the best performance (F1-score of 69.2% for DAIC-WOZ and 91.5% for CONVERGE). Analysis of the class-separability measure (J-ratio) of the hidden states of the DepAudioNet model shows that when adversarial learning is applied, the backend model loses some speaker-discriminability while it improves depression-discriminability. These results indicate that there are some components of speaker identity that may not be useful for depression detection and minimizing their effects provides a more accurate diagnosis of the underlying disorder and can safeguard a speaker's identity. Vijay Ravi, Jinhan Wang, Jonathan Flint, Abeer Alwan |
INTERSPEECH | 4 |
| 2022 | Unsupervised Instance Discriminative Learning for Depression Detection from Speech Signalsabstract-value 0.0015 and 0.05, respectively, are observed using PIS in the detection of MDD relative to the baseline without pre-training. Jinhan Wang, Vijay Ravi, Jonathan Flint, Abeer Alwan |
INTERSPEECH | 4 |
| 2022 | Spoken language interaction with robots: Recommendations for future researchabstractWith robotics rapidly advancing, more effective human–robot interaction is increasingly needed to realize the full potential of robots for society. While spoken language must be part of the solution, our ability to provide spoken language interaction capabilities is still very limited. In this article, based on the report of an interdisciplinary workshop convened by the National Science Foundation, we identify key scientific and engineering advances needed to enable effective spoken language interaction with robotics. We make 25 recommendations, involving eight general themes: putting human needs first, better modeling the social and interactive aspects of language, improving robustness, creating new methods for rapid adaptation, better integrating speech and language with other communication modalities, giving speech and language components access to rich representations of the robot’s current knowledge and state, making all components operate in real time, and improving research infrastructure and resources. Research and development that prioritizes these topics will, we believe, provide a solid foundation for the creation of speech-capable robots that are easy and effective for humans to work with. Matthew Marge, Carol Y. Espy-Wilson, Nigel G. Ward, Abeer Alwan, Yoav Artzi, Mohit Bansal, Gilmer L. Blankenship, Joyce Y. Chai, Hal Daumé III, Debadeepta Dey, Mary P. Harper, Thomas Howard, Casey Kennington, Ivana Kruijff-Korbayová, Dinesh Manocha, Cynthia Matuszek, Ross Mead, Raymond J. Mooney, Roger K. Moore, Mari Ostendorf, Heather Pon-Barry, Alexander I. Rudnicky, Matthias Scheutz, Robert St. Amant, Stefanie Tellex, David R. Traum, Zhou Yu 0005 |
Comput. Speech Lang. | 4 |
| 2021 | Bi-APC: Bidirectional Autoregressive Predictive Coding for Unsupervised Pre-Training and its Application to Children's ASRabstractWe present a bidirectional unsupervised model pre-training (UPT) method and apply it to children’s automatic speech recognition (ASR). An obstacle to improving child ASR is the scarcity of child speech databases. A common approach to alleviate this problem is model pre-training using data from adult speech. Pre-training can be done using supervised (SPT) or unsupervised methods, depending on the availability of annotations. Typically, SPT performs better. In this paper, we focus on UPT to address the situations when pre-training data are unlabeled. Autoregressive predictive coding (APC), a UPT method, predicts frames from only one direction, limiting its use to uni-directional pre-training. Conventional bidirectional UPT methods, however, predict only a small portion of frames. To extend the benefits of APC to bi-directional pre-training, Bi-APC is proposed. We then use adaptation techniques to transfer knowledge learned from adult speech (using the Librispeech corpus) to child speech (OGI Kids corpus). LSTM-based hybrid systems are investigated. For the uni-LSTM structure, APC obtains similar WER improvements to SPT over the baseline. When applied to BLSTM, however, APC is not as competitive as SPT, but our proposed Bi-APC has comparable improvements to SPT. Ruchao Fan, Amber Afshan, Abeer Alwan |
ICASSP | 3 |
| 2021 | Fundamental Frequency Feature Normalization and Data Augmentation for Child Speech RecognitionabstractAutomatic speech recognition (ASR) systems for young children are needed due to the importance of age-appropriate educational technology. Because of the lack of publicly available young child speech data, feature extraction strategies such as feature normalization and data augmentation must be considered to successfully train child ASR systems. This study proposes a novel technique for child ASR using both feature normalization and data augmentation methods based on the relationship between formants and fundamental frequency (fo). Both the fofeature normalization and data augmentation techniques are implemented as a frequency shift in the Mel domain. These techniques are evaluated on a child read speech ASR task. Child ASR systems are trained by adapting a BLSTM-based acoustic model trained on adult speech. Using both fonormalization and data augmentation results in a relative word error rate (WER) improvement of 19.3% over the baseline when tested on the OGI Kids’ Speech Corpus, and the resulting child ASR system achieves the best WER currently reported on this corpus. Gary Yeung, Ruchao Fan, Abeer Alwan |
ICASSP | 3 |
| 2021 | An Improved Single Step Non-Autoregressive Transformer for Automatic Speech RecognitionabstractNon-autoregressive mechanisms can significantly decrease inference time for speech transformers, especially when the single step variant is applied. Previous work on CTC alignment-based single step non-autoregressive transformer (CASS-NAT) has shown a large real time factor (RTF) improvement over autoregressive transformers (AT). In this work, we propose several methods to improve the accuracy of the end-to-end CASS-NAT, followed by performance analyses. First, convolution augmented self-attention blocks are applied to both the encoder and decoder modules. Second, we propose to expand the trigger mask (acoustic boundary) for each token to increase the robustness of CTC alignments. In addition, iterated loss functions are used to enhance the gradient update of low-layer parameters. Without using an external language model, the WERs of the improved CASS-NAT, when using the three methods, are 3.1%/7.2% on Librispeech test clean/other sets and the CER is 5.4% on the Aishell1 test set, achieving a 7%~21% relative WER/CER improvement. For the analyses, we plot attention weight distributions in the decoders to visualize the relationships between token-level acoustic embeddings. When the acoustic embeddings are visualized, we find that they have a similar behavior to word embeddings, which explains why the improved CASS-NAT performs similarly to AT. Ruchao Fan, Peng Chang 0002, Jing Xiao 0006, Abeer Alwan |
Interspeech | 5 |
| 2021 | Low Resource German ASR with Untranscribed Data Spoken by Non-Native Children - INTERSPEECH 2021 Shared Task SPAPL SystemabstractThis paper describes the SPAPL system for the INTER-SPEECH 2021 Challenge: Shared Task on Automatic Speech Recognition for Non-Native Children's Speech in German.∼ 5 hours of transcribed data and ∼ 60 hours of untranscribed data are provided to develop a German ASR system for children.For the training of the transcribed data, we propose a non-speech state discriminative loss (NSDL) to mitigate the influence of long-duration non-speech segments within speech utterances.In order to explore the use of the untranscribed data, various approaches are implemented and combined together to incrementally improve the system performance.First, bidirectional autoregressive predictive coding (Bi-APC) is used to learn initial parameters for acoustic modelling using the provided untranscribed data.Second, incremental semi-supervised learning is further used to iteratively generate pseudo-transcribed data.Third, different data augmentation schemes are used at different training stages to increase the variability and size of the training data.Finally, a recurrent neural network language model (RNNLM) is used for rescoring.Our system achieves a word error rate (WER) of 39.68% on the evaluation data, an approximately 12% relative improvement over the official baseline (45.21%). Jinhan Wang, Yunzheng Zhu, Ruchao Fan, Abeer Alwan |
Interspeech | 5 |
| 2021 | Fundamental frequency feature warping for frequency normalization and data augmentation in child automatic speech recognition
Gary Yeung, Ruchao Fan, Abeer Alwan |
Speech Commun. | 3 |
| 2020 | Variable Frame Rate-Based Data Augmentation to Handle Speaking-Style Variability for Automatic Speaker VerificationabstractThe effects of speaking-style variability on automatic speaker verification were investigated using the UCLA Speaker Variability database which comprises multiple speaking styles per speaker. An x-vector/PLDA (probabilistic linear discriminant analysis) system was trained with the SRE and Switchboard databases with standard augmentation techniques and evaluated with utterances from the UCLA database. The equal error rate (EER) was low when enrollment and test utterances were of the same style (e.g., 0.98% and 0.57% for read and conversational speech, respectively), but it increased substantially when styles were mismatched between enrollment and test utterances. For instance, when enrolled with conversation utterances, the EER increased to 3.03%, 2.96% and 22.12% when tested on read, narrative, and pet-directed speech, respectively. To reduce the effect of style mismatch, we propose an entropy-based variable frame rate technique to artificially generate style-normalized representations for PLDA adaptation. The proposed system significantly improved performance. In the aforementioned conditions, the EERs improved to 2.69% (conversation -- read), 2.27% (conversation -- narrative), and 18.75% (pet-directed -- read). Overall, the proposed technique performed comparably to multi-style PLDA adaptation without the need for training data in different speaking styles per speaker. Amber Afshan, Jinxi Guo, Soo Jin Park, Vijay Ravi, Alan McCree, Abeer Alwan |
INTERSPEECH | 6 |
| 2020 | Speaker Discrimination in Humans and Machines: Effects of Speaking Style VariabilityabstractDoes speaking style variation affect humans' ability to distinguish individuals from their voices? How do humans compare with automatic systems designed to discriminate between voices? In this paper, we attempt to answer these questions by comparing human and machine speaker discrimination performance for read speech versus casual conversations. Thirty listeners were asked to perform a same versus different speaker task. Their performance was compared to a state-of-the-art x-vector/PLDA-based automatic speaker verification system. Results showed that both humans and machines performed better with style-matched stimuli, and human performance was better when listeners were native speakers of American English. Native listeners performed better than machines in the style-matched conditions (EERs of 6.96% versus 14.35% for read speech, and 15.12% versus 19.87%, for conversations), but for style-mismatched conditions, there was no significant difference between native listeners and machines. In all conditions, fusing human responses with machine results showed improvements compared to each alone, suggesting that humans and machines have different approaches to speaker discrimination tasks. Differences in the approaches were further confirmed by examining results for individual speakers which showed that the perception of distinct and confused speakers differed between human listeners and machines. Amber Afshan, Jody Kreiman, Abeer Alwan |
INTERSPEECH | 3 |
| 2020 | Exploring the Use of an Unsupervised Autoregressive Model as a Shared Encoder for Text-Dependent Speaker VerificationabstractIn this paper, we propose a novel way of addressing text-dependent automatic speaker verification (TD-ASV) by using a shared-encoder with task-specific decoders. An autoregressive predictive coding (APC) encoder is pre-trained in an unsupervised manner using both out-of-domain (LibriSpeech, VoxCeleb) and in-domain (DeepMine) unlabeled datasets to learn generic, high-level feature representation that encapsulates speaker and phonetic content. Two task-specific decoders were trained using labeled datasets to classify speakers (SID) and phrases (PID). Speaker embeddings extracted from the SID decoder were scored using a PLDA. SID and PID systems were fused at the score level. There is a 51.9% relative improvement in minDCF for our system compared to the fully supervised x-vector baseline on the cross-lingual DeepMine dataset. However, the i-vector/HMM method outperformed the proposed APC encoder-decoder system. A fusion of the x-vector/PLDA baseline and the SID/PLDA scores prior to PID fusion further improved performance by 15% indicating complementarity of the proposed approach to the x-vector system. We show that the proposed approach can leverage from large, unlabeled, data-rich domains, and learn speech patterns independent of downstream tasks. Such a system can provide competitive performance in domain-mismatched scenarios where test data is from data-scarce domains. Vijay Ravi, Ruchao Fan, Amber Afshan, Huanhua Lu, Abeer Alwan |
INTERSPEECH | 5 |
| 2020 | Analysis of Disfluency in Children's SpeechabstractDisfluencies are prevalent in spontaneous speech, as shown in many studies of adult speech. Less is understood about children's speech, especially in pre-school children who are still developing their language skills. We present a novel dataset with annotated disfluencies of spontaneous explanations from 26 children (ages 5--8), interviewed twice over a year-long period. Our preliminary analysis reveals significant differences between children's speech in our corpus and adult spontaneous speech from two corpora (Switchboard and CallHome). Children have higher disfluency and filler rates, tend to use nasal filled pauses more frequently, and on average exhibit longer reparandums than repairs, in contrast to adult speakers. Despite the differences, an automatic disfluency detection system trained on adult (Switchboard) speech transcripts performs reasonably well on children's speech, achieving an F1 score that is 10\% higher than the score on an adult out-of-domain dataset (CallHome). Trang Tran 0001, Morgan Tinkler, Gary Yeung, Abeer Alwan, Mari Ostendorf |
INTERSPEECH | 4 |
| 2019 | Target and Non-target Speaker Discrimination by Humans and MachinesabstractThe manner in which acoustic features contribute to perceiving speaker identity remains unclear. In an attempt to better understand speaker perception, we investigated human and machine speaker discrimination with utterances shorter than 2 seconds. Sixty-five listeners performed a same vs. different task. Machine performance was estimated with i-vector/PLDA-based automatic speaker verification systems, one using mel-frequency cepstral coefficients (MFCCs) and the other using voice quality features (VQual2) inspired by a psychoacoustic model of voice quality. Machine performance was measured in terms of the detection and log-likelihood-ratio cost functions. Humans showed higher confidence for correct target decisions compared to correct non-target decisions, suggesting that they rely on different features and/or decision making strategies when identifying a single speaker compared to when distinguishing between speakers. For non-target trials, responses were highly correlated between humans and the VQual2-based system, especially when speakers were perceptually marked. Fusing human responses with an MFCC-based system improved performance over human-only or MFCC-only results, while fusing with the VQual2-based system did not. The study is a step towards understanding human speaker discrimination strategies and suggests that automatic systems might be able to supplement human decisions especially when speakers are marked. Soo Jin Park, Amber Afshan, Jody Kreiman, Gary Yeung, Abeer Alwan |
ICASSP | 5 |
| 2019 | Voice Quality and Between-Frame Entropy for Sleepiness Estimation
Vijay Ravi, Soo Jin Park, Amber Afshan, Abeer Alwan |
INTERSPEECH | 4 |
| 2019 | A Frequency Normalization Technique for Kindergarten Speech Recognition Inspired by the Role of fo in Vowel Perception
Gary Yeung, Abeer Alwan |
INTERSPEECH | 2 |
| 2018 | Effectiveness of Voice Quality Features in Detecting Depression
Amber Afshan, Jinxi Guo, Soo Jin Park, Vijay Ravi, Jonathan Flint, Abeer Alwan |
INTERSPEECH | 6 |
| 2018 | Filter Sampling and Combination CNN (FSC-CNN): A Compact CNN Model for Small-footprint ASR Acoustic Modeling Using Raw Waveforms
Jinxi Guo, Ning Xu 0010, Abeer Alwan |
INTERSPEECH | 6 |
| 2018 | Using Voice Quality Supervectors for Affect Identification
Soo Jin Park, Amber Afshan, Zhi Ming Chua, Abeer Alwan |
INTERSPEECH | 4 |
| 2018 | On the Difficulties of Automatic Speech Recognition for Kindergarten-Aged Children
Gary Yeung, Abeer Alwan |
INTERSPEECH | 2 |
| 2018 | Deep neural network based i-vector mapping for speaker verification using short utterances
Jinxi Guo, Ning Xu 0010, Kailun Qian, Ying Nian Wu, Abeer Alwan |
Speech Commun. | 7 |
| 2017 | CNN-Based Joint Mapping of Short and Long Utterance i-Vectors for Speaker Verification Using Short Utterances
Jinxi Guo, Usha Amrutha Nookala, Abeer Alwan |
INTERSPEECH | 3 |
| 2017 | Attention Based CLDNNs for Short-Duration Acoustic Scene Classification
Jinxi Guo, Ning Xu 0010, Li-Jia Li 0001, Abeer Alwan |
INTERSPEECH | 4 |
| 2017 | Using Voice Quality Features to Improve Short-Utterance, Text-Independent Speaker Verification Systems
Soo Jin Park, Gary Yeung, Jody Kreiman, Patricia A. Keating, Abeer Alwan |
INTERSPEECH | 5 |
| 2016 | Speaker Verification Using Short Utterances with DNN-Based Estimation of Subglottal Acoustic Features
Jinxi Guo, Gary Yeung, Deepak Muralidharan, Harish Arsikere, Amber Afshan, Abeer Alwan |
INTERSPEECH | 6 |
| 2016 | Noise-Robust Hidden Markov Models for Limited Training Data for Within-Species Bird Phrase Classification
Kantapon Kaewtip, Charles E. Taylor, Abeer Alwan |
INTERSPEECH | 3 |
| 2016 | Fusion Strategies for Robust Speech Recognition and Keyword Spotting for Channel- and Noise-Degraded Speech
Vikramjit Mitra, Julien van Hout, Wen Wang 0001, Chris Bartels, Horacio Franco, Dimitra Vergyri, Abeer Alwan, Adam Janin, John H. L. Hansen, Richard M. Stern, Abhijeet Sangwan, Nelson Morgan |
INTERSPEECH | 7 |
| 2016 | Speaker Identity and Voice Quality: Modeling Human Responses and Automatic Speaker Recognition
Soo Jin Park, Caroline Sigouin, Jody Kreiman, Patricia A. Keating, Jinxi Guo, Gary Yeung, Fang-Yu Kuo, Abeer Alwan |
INTERSPEECH | 8 |
| 2015 | Bird-phrase segmentation and verification: A noise-robust template-based approachabstractIn this paper, we present a birdsong-phrase segmentation and verification algorithm that is robust to limited training data, class variability, and noise. The algorithm comprises a noise-robust, Dynamic-Time-Warping (DTW)-based segmentation and a discriminative classifier for outlier rejection. The algorithm utilizes DTW and prominent (high energy) time-frequency regions of training spectrograms to derive a reliable noise-robust template for each phrase class. The resulting template is then used for segmenting continuous recordings to obtain segment candidates whose spectrogram amplitudes in the prominent regions are used as features to a Support Vector Machine (SVM). The algorithm is evaluated on the Cassin's Vireo recordings; our proposed system yields low Equal Error Rates (EER) and segment boundaries that are close to those obtained from manual annotations and, is better than energy or entropy-based birdsong segmentation algorithms. In the presence of additive noise (-10 to 10 dB SNR), the proposed phrase detection system does not degrade as significantly as the other algorithms do. Kantapon Kaewtip, Lee Ngee Tan, Charles E. Taylor, Abeer Alwan |
ICASSP | 4 |
| 2015 | Age-dependent height estimation and speaker normalization for children's speech using the first three subglottal resonancesabstractThis paper proposes an age-dependent scheme for automatic height estimation and speaker normalization of children’s speech, using the first three subglottal resonances (SGRs). Similar to previous work, our analysis indicates that children above the age of 11 years show different acoustic properties from those under 11. Therefore, an age-dependent model is investigated. The estimation algorithms for the first three SGRs are motivated by our previous research for adults. The algorithms for the first two SGRs have been applied to children’s speech before. This paper proposes a similar approach to estimate Sg3 for children. The algorithm is trained and evaluated on 46 children, aged between 6-17 years, using cross-validation. Average RMS errors in estimating Sg1, Sg2 and Sg3 using the age-dependent model are 51, 128 and 168 Hz, respectively. The height estimation algorithm employs a negative correlation between SGRs and height, and the mean absolute height estimation error was found to be less than 3.8cm for the younger children and 4.9cm for the older children. In addition, using TIDIGITS, a linear frequency warping scheme using age-dependent Sg3 gives statisticallysignificant word error rate reductions (up to 26%) relative to conventional VTLN. Jinxi Guo, Rohit Paturi, Gary Yeung, Steven M. Lulich, Harish Arsikere, Abeer Alwan |
INTERSPEECH | 6 |
| 2015 | The relationship between acoustic and perceived intraspeaker variability in voice qualityabstractLittle is known about intraspeaker changes in voice across changing speaking situations in everyday life. In this study, we examined acoustic variations between and within 5 talkers and their effect on the likelihood that voice samples would not be identified as coming from the same talker. Talkers were drawn from a large database recorded to capture everyday variations in vocal characteristics. Nine samples of /a/, recorded on three different days, were examined for each talker. Acoustic characteristics were estimated using VoiceSauce and analysis-by-synthesis, and listeners judged whether pairs of voices came from the same or two different talkers. Results indicate that interspeaker variability in voice quality exceeds intraspeaker variability, but differences are smaller than expected. As predicted by models that treat voice quality as an auditory pattern, the acoustic attributes associated with incorrect “different speaker ” responses varied from talker to talker, depending on the particular characteristics of the voice in question. Index Terms: voice quality, speaker recognition, intra-speaker variability Jody Kreiman, Soo Jin Park, Patricia A. Keating, Abeer Alwan |
INTERSPEECH | 4 |
| 2014 | Frequency warping using subglottal resonances: Complementarity with VTLN and robustness to additive noiseabstractBased on our recently-proposed frequency-warping scheme using subglottal resonances (SGRs), this paper addresses two well-known limitations of conventional vocal-tract length normalization (VTLN): (1) its sub-optimal nature owing to the lack of frequency-dependent scaling, and (2) sensitivity to noise. Based on the idea of filter-bank interpolation, a novel approach is proposed to realize the combined effect of VTLN and SGR-based warping (which provides frequency-dependent scaling). Using the Wall Street Journal database and the conventional MFCC front end, SGR warping is shown to be complementary to VTLN in performance. Since SGR warping depends more on the given signal and less on models trained a priori, we argue that SGR warping is less sensitive to noise than VTLN. Through experiments on the AURORA-4 database with power-normalized cepstral coefficients as noise-robust front-end features, we show that SGR warping is better than VTLN, in clean as well as multi-conditional training. Harish Arsikere, Abeer Alwan |
ICASSP | 2 |
| 2014 | Non-linear dimension reduction of Gabor features for noise-robust ASRabstractIt has been shown that Gabor filters closely resemble the spectro-temporal response fields of neurons in the primary auditory cortex. A filter bank of 2-D Gabor filters can be applied to either the mel-spectrogram or power normalized spectrogram to obtain a set of physiologically inspired Gabor Filter Bank Features. The high dimensionality and the correlated nature of these features pose an issue for ASR. In the past, dimension reduction was performed through (1) feature selection, (2) channel selection, (3) linear dimension reduction or (4) tandem acoustic modelling. In this paper, we propose a novel solution to this issue based on channel selection and non-linear dimension reduction using Laplacian Eigenmaps. These features are concatenated with Power Normalized Cepstral Coefficients (PNCC) to evaluate if the two are complementary and provide an improvement in performance. We show a relative reduction of 12.66% in the WER compared to the PNCC baseline, when applied to the Aurora 4 database. Hitesh Anand Gupta, Anirudh Raju, Abeer Alwan |
ICASSP | 3 |
| 2014 | Feature enhancement using sparse reference and estimated soft-mask exemplar-pairs for noisy speech recognitionabstractA feature enhancement technique for noise-robust speech recognition is proposed. Existing sparse exemplar-based feature enhancement methods use clean speech and pure noise Mel-spectral exemplars, or clean and noisy speech log-Mel-spectral exemplar-pairs, in their dictionaries. In contrast, the proposed technique constructs its dictionaries using reference soft-mask (SMref) and estimated soft-mask (SMest) exemplar-pairs derived from the training data. The sparse linear combination of SMestdictionary exemplars that best represents the test utterance's SMestis obtained by solving an L1-minimization problem. This sparse linear combination is applied to the SMrefexemplar dictionary to generate an enhanced soft-mask for denoising the utterance's Mel-spectra before MFCC extraction. On the Aurora-2 noisy speech recognition task, the proposed algorithm outperforms other sparse Mel-spectral exemplar-based feature enhancement schemes when mismatch exists between the dictionary exemplars and the test set. A preliminary experiment on Aurora-4 shows similar trends. Lee Ngee Tan, Abeer Alwan |
ICASSP | 2 |
| 2014 | Investigating the effect of F0 and vocal intensity on harmonic magnitudes: data from high-speed laryngeal videoendoscopyabstractThe relative magnitude of the first two harmonics of the voice source (H1*-H2*) is an important measure and is assumed to be one exponent of changes in vocal quality along a breathyto-pressed continuum. H1*-H2* is often associated with glottal open quotient (OQ) and glottal pulse skewness (as quantified by speed quotient, SQ), but may also covary with fundamental frequency (F0) and vocal intensity. We examined the relationship between H1*-H2*, F0, and vocal intensity using phonations in which vocal qualities varied continuously in F0 and intensity. Glottal area measures (OQ and SQ) and acoustic measures (F0, intensity, and H1*-H2*) were studied using simultaneously-collected laryngeal high-speed videoendoscopy and audio recordings from 9 subjects. Analyses of individual speakers showed that H1*-H2* may sometimes vary as a function of F0 alone, with OQ and SQ remaining rather constant, hypothetically when nonlinear source-filter interaction is strong. Although conventionally H1*-H2* is assumed to decrease with increasing vocal intensity due to a decrease in OQ, results showed examples where H1*-H2* increased with increasing vocal intensity, hypothetically when the effect of decreasing pulse skewness exceeds the effect of decreasing OQ. In some phonatory modes, the relationship between SQ and H1*H2* may not be as monotonic as previously assumed. Gang Chen 0009, Soo Jin Park, Jody Kreiman, Abeer Alwan |
INTERSPEECH | 4 |
| 2014 | Speaker recognition via fusion of subglottal features and MFCCsabstractMotivated by the speaker-specificity and stationarity of subglot-tal acoustics, this paper investigates the utility of subglottal cep-stral coefficients (SGCCs) for speaker identification (SID) and verification (SV). SGCCs can be computed using accelerom-eter recordings of subglottal acoustics, but such an approach is infeasible in real-world scenarios. To estimate SGCCs from speech signals, we adopt the Bayesian minimum mean squared error (MMSE) estimator proposed in the speech-to-articulatory inversion literature. The joint distribution of SGCCs and speech MFCCs is modeled using the WashU-UCLA corpus (containing simultaneous recordings of speech and subglottal acoustics), and the resulting model is used to obtain an MMSE estimate of SGCCs from unseen (test) MFCCs. Cross-validation experi-ments on the WashU-UCLA corpus show that the estimation ef-ficacy, on average, is speaker dependent. A score-level fusion of MFCC and SGCC systems outperforms the MFCC-only base-line in both SID and SV tasks. On the TIMIT database (SID), the relative reduction in identification error is 16, 40 and 51% for G.712-filtered (300–3400 Hz), narrowband (0–4000 Hz) and wideband (0–8000 Hz) speech, respectively. On the NIST 2008 database (SV), the relative reduction in equal error rate is 4 and 11 % for 10 and 5 second utterances, respectively. Index Terms: speaker recognition, subglottal acoustics, cep-stral coefficients, score combination, MMSE estimation Harish Arsikere, Hitesh Anand Gupta, Abeer Alwan |
INTERSPEECH | 3 |
| 2014 | The relationship between the second subglottal resonance and vowel class, standing height, trunk length, and F0 variation for Mandarin speakersabstractThe relationship between vowel formants and the second subglottal resonance (Sg2) has previously been explored in English, German, Hungarian and Korean. Results from these studies indicate that vowel space is categorically divided by Sg2 and that Sg2 correlates well with standing height. One of the goals of this work is to verify if the above findings hold true in Mandarin as well. The correlation between Sg2 and sitting height (trunk length) is also studied. Further, since Mandarin is a tonal language (with more pitch variations compared to English), we study the relationship between Sg2 and fundamental frequency (F0). A new corpus of simultaneous recordings of speech and subglottal acoustics was collected from 20 native Mandarin speakers. Results on this corpus indicate that Sg2 divides vowel space in Mandarin as well, and that it is more correlated with sitting height than standing height. Paired t-tests are conducted on the Sg2 measurements from different vowel parts, which represent different F0 regions. Preliminary results show that there is no statistically-significant variation of Sg2 with F0 within a tone. Index Terms: second subglottal resonance, Mandarin, vowel space, sitting height, tonal language Jinxi Guo, Angli Liu, Harish Arsikere, Abeer Alwan, Steven M. Lulich |
INTERSPEECH | 4 |
| 2014 | The glottaltopogram: A method of analyzing high-speed images of the vocal folds
Gang Chen 0009, Jody Kreiman, Abeer Alwan |
Comput. Speech Lang. | 3 |
| 2014 | Glottal source processing: From analysis to applications
Thomas Drugman, Paavo Alku, Abeer Alwan, Bayya Yegnanarayana |
Comput. Speech Lang. | 3 |
| 2013 | Non-linear frequency warping for VTLN using subglottal resonances and the third formant frequencyabstractThis paper proposes a non-linear frequency warping scheme for VTLN. It is based on mapping the subglottal resonances (SGRs) and the third formant frequency (F3) of a given utterance to those of a reference speaker. SGRs are used because they relate to formants in specific ways while remaining phonetically invariant, and F3 is used because it is somewhat correlated to vocal-tract length. Given an utterance, the warping parameters (SGRs and F3) are determined by obtaining initial estimates from the signal, and refining the estimates with respect to a speaker-independent model. For children (TIDIGITS), the proposed method yields statistically-significant word error rate (WER) reductions (up to 15%) relative to conventional VTLN (linear warping) when: (1) speakers show poor baseline performance, and/or (2) training data are limited. For adults (Wall Street Journal), the WER reduction relative to conventional VTLN is 4-5%. Comparison with other non-linear warping techniques is also reported. Harish Arsikere, Steven M. Lulich, Abeer Alwan |
ICASSP | 3 |
| 2013 | A robust automatic bird phrase classifier using dynamic time-warping with prominent region identificationabstractIn this paper, we present a novel approach to birdsong phase classification using template-based techniques suitable even for limited training data and noisy environments. The algorithm utilizes dynamic time-warping and prominent (high-energy) time-frequency regions of training spectrograms to derive templates. The algorithm is evaluated on 32 classes of Cassin's Vireo bird phrases. Using only three training examples per class, our algorithm yields a phrase accuracy of 96.23%, outperforming other classifiers (e.g. 85.21% classification accuracy of SVM). In the presence of additive noise (10 dB SNR degradation), the proposed classifier does not degrade significantly, compared to others. Kantapon Kaewtip, Lee Ngee Tan, Abeer Alwan, Charles E. Taylor |
ICASSP | 3 |
| 2013 | A sparse representation-based classifier for in-set bird phrase verification and classification with limited training dataabstractThe performance of a sparse representation-based (SR) classifier for in-set bird phrase verification and classification is studied. The database contains phrases segmented from songs of the Cassin's Vireo (Vireo cassinii). Each test phrase belongs to one of 33 phrase classes - 32 in-set categories, and 1 collective out-of-set category. Only in-set phrases are used for training. From each phrase segment, spectrographic features were extracted, followed by dimension reduction using PCA. A threshold is applied on the sparsity concentration index (SCI) computed by the SR classifier, for in-set bird phrase verification using a limited number of training tokens (3 - 7) per phrase class. When evaluated against the nearest subspace (NS) and support vector machine (SVM) classifiers using the same framework, the SR classifier has the highest classification accuracy, due to its good performances in both the verification and classification tasks. Lee Ngee Tan, George Kossan, Martin L. Cody, Charles E. Taylor, Abeer Alwan |
ICASSP | 5 |
| 2013 | Bird phrase segmentation by entropy-driven change point detectionabstractA bird phrase segmentation method using entropy-based change point detection is proposed. Spectrograms of bird calls are usually sparse while the background noise is relatively white. Therefore, considering the entropy of a sliding time-frequency block on the spectrogram, the entropy dips when detecting a signal and rises when the signal ends. Rather than applying a hard threshold on the entropy to determine the beginning and ending of a signal, a Bayesian change point detection is used to detect the statistical changes in the entropy sequence. Tests on a database of Cassin's Vireo (Vireo cassinii), our proposed segmentation method with spectral subtraction or a novel spectral whitening method as the front-end generates more accurate time labels, lower the false alarm rate than the conventional time-domain energy detection method and achieves high phrase classification rate. Ni-Chun Wang, Ralph E. Hudson, Lee Ngee Tan, Charles E. Taylor, Abeer Alwan |
ICASSP | 5 |
| 2013 | A perceptually and physiologically motivated voice source modelabstractMany glottal source models have been proposed, but none has been systematically validated perceptually.Our previous work showed that model fitting of the negative peak of the flow derivative is the most important predictor of perceptual similarity to the target voice.In this study, a new voice source model is proposed to capture perceptually-important source shape aspects.This new model, along with four other source models, was fitted to 40 voice sources (20 male and 20 female) obtained by inverse filtering and analysis-by-synthesis (AbS) of samples of natural speech.We generated synthetic copies of the voices using each modeled source pulse, with all other synthesis parameters held constant, and then conducted a visual sort-andrate task in which listeners assessed the extent of perceived similarity between the target voice samples and each copy.Results showed that the proposed model provided a more accurate fit and a better perceptual match to the target than did the other models. Gang Chen 0009, Marc Garellek, Jody Kreiman, Bruce R. Gerratt, Abeer Alwan |
INTERSPEECH | 5 |
| 2013 | Investigating the relationship between glottal area waveform shape and harmonic magnitudes through computational modeling and laryngeal high-speed videoendoscopyabstractThe glottal open quotient (OQ) is often associated with the amplitude of the first source harmonic relative to the second (H1*H2*), which is assumed to be one cause of a change in vocal quality along a breathy-to-pressed continuum. The association between OQ and H1*-H2* was investigated in a group of 5 human subjects and also in a computational voice production simulation. The simulation incorporated a parametric voice source model into a nonlinear source-filter framework. H1*-H2* and OQ were measured synchronously from audio recordings and high-speed laryngeal videoendoscopy of “glide” phonations in which quality varied continuously from breathy to pressed. Analyses of individual speakers showed large differences in the relationship between OQ and H1*-H2*. The variability in laryngeal high-speed data was consistent with simulation results, which showed that the relationship between OQ and H1*-H2* depended on mean glottal area, a parameter associated with the degree of source-filter interaction and not directly measurable from high-speed video of the vocal folds. In addition, H1*-H2* may change with increasing glottal gap size; this change contributes to the observed variability in the relationship between H1*-H2* and OQ. Index Terms: harmonic magnitudes, laryngeal high-speed videoendoscopy, glottal area waveform, open quotient Gang Chen 0009, Robin A. Samlan, Jody Kreiman, Abeer Alwan |
INTERSPEECH | 4 |
| 2013 | All for one: feature combination for highly channel-degraded speech activity detectionabstractSpeech activity detection (SAD) on channel transmissions is a critical preprocessing task for speech, speaker and language recognition or for further human analysis. This paper presents a feature combination approach to improve SAD on highly channel degraded speech as part of the Defense Advanced Martin Graciarena, Abeer Alwan, Daniel P. W. Ellis, Horacio Franco, Luciana Ferrer, John H. L. Hansen, Adam Janin, Yun Lei, Vikramjit Mitra, Nelson Morgan, Seyed Omid Sadjadi, T. J. Tsai 0001, Nicolas Scheffer, Lee Ngee Tan |
INTERSPEECH | 2 |
| 2013 | A pitch-based spectral enhancement technique for robust speech processingabstractThis paper presents a new pitch-based spectral enhancement algorithm on voiced frames for speech analysis and noiserobust speech processing. The proposed algorithm determines a time-warping function (TWF) and the speaker’s pitch with high precision, simultaneously. This technique reduces the smearing effect in between harmonics when the fundamental frequency is not constant within the analysis window. To do so, we propose a metric called the harmonic residual which measures the difference between the actual spectrum and the resynthesized spectrum derived from the linear model of speech production with various combinations of TWF and high-precision pitch values as parameters. The TWF and pitch pair that yields the minimum harmonic residual is selected and the enhanced spectrum is obtained accordingly. We show how this new representation can be used for automatic speech recognition by proposing a robust spectral representation derived from harmonic amplitude interpolation. Kantapon Kaewtip, Lee Ngee Tan, Abeer Alwan |
INTERSPEECH | 3 |
| 2013 | Automatic estimation of the first three subglottal resonances from adults' speech signals with application to speaker height estimation
Harish Arsikere, Gary K. F. Leung, Steven M. Lulich, Abeer Alwan |
Speech Commun. | 4 |
| 2013 | Multi-band summary correlogram-based pitch detection for noisy speech
Lee Ngee Tan, Abeer Alwan |
Speech Commun. | 2 |
| 2012 | Automatic height estimation using the second subglottal resonanceabstractThis paper presents an algorithm for automatically estimating speaker height. It is based on: (1) a recently-proposed model of the subglottal system that explains the inverse relation observed between subglottal resonances and height, and (2) an improved version of our previous algorithm for automatically estimating the second subglottal resonance (Sg2). The improved Sg2 estimation algorithm was trained and evaluated on recently-collected data from 30 and 20 adult speakers, respectively. Sg2 estimation error was found to reduce by 29%, on average, as compared to the previous algorithm. The height estimation algorithm, employing the inverse relation between Sg2 and height, was trained on data from the above-mentioned 50 adults. It was evaluated on 563 adult speakers in the TIMIT corpus, and the mean absolute height estimation error was found to be less than 5.6cm. Harish Arsikere, Gary K. F. Leung, Steven M. Lulich, Abeer Alwan |
ICASSP | 4 |
| 2012 | The glottaltopograph: A method of analyzing high-speed images of the vocal foldsabstractHigh-speed video provides a way to record physiological vibrational patterns of the vocal folds. Due to the large amount of data it produces, many methods have been proposed to reduce raw images to the underlying vibrational patterns. Previous methods either focus on a certain location on the vocal folds, or a certain frequency of vibrational activity. In this paper, we propose the “glottaltopograph” which is based on principal component analysis of pixels' gray scale time courses. This method reveals the overall synchronization of the vibrational patterns of the vocal folds. Experimental results showed that this method is effective in visualizing pathological and normal vocal-fold vibrational patterns. Gang Chen 0009, Jody Kreiman, Abeer Alwan |
ICASSP | 3 |
| 2012 | FBEM: A filter bank EM algorithm for the joint optimization of features and acoustic model parameters in bird call classificationabstractThis paper extends the expectation-maximization (EM) algorithm to estimate not only optimal acoustic model parameters, but also optimal center frequencies and bandwidths of the filter bank used in cepstral feature extraction for bird call classification. The search is done using the gradient ascent method. Filter bank and model parameters are optimized iteratively. Experiments are conducted on a large noisy corpus containing Antbird calls from 5 species. It is shown that features extracted using the optimized filter bank result in a lower classification error rate than those extracted using a Melscaled filter bank. Abeer Alwan |
ICASSP | 2 |
| 2012 | A novel approach to soft-mask estimation and Log-Spectral enhancement for robust speech recognitionabstractThis paper describes a technique for enhancing the Mel-filtered log spectra of noisy speech, with application to noise robust speech recognition. We first compute an SNR-based soft-decision mask in the Mel-spectral domain as an indicator of speech presence. Then, we exploit the known time-frequency correlation of speech by treating this mask as an image, and performing median filtering and blurring to remove the outliers and to smooth the decision regions. This mask constitutes a set of multiplicative coefficients (ranging in [0,1]) that are used to discard the unreliable parts of the Mel-filtered log-spectrum of noisy speech. Finally, we apply Log-Spectral Flooring [1] on the liftered spectra of both clean and noisy speech so as to match their respective dynamic ranges and to emphasize the information in the spectral peaks. The noisy MFCCs computed on these modified log-spectra show an increased similarity with their corresponding clean MFCCs. Evaluation on the Aurora-2 corpus shows that the proposed approach competes with state-of-the-art front-ends, like ETSI-AFE, MVA or PNCC. Julien van Hout, Abeer Alwan |
ICASSP | 2 |
| 2012 | Automatic estimation of the first two subglottal resonances in children's speech with application to speaker normalization in limited-data conditionsabstractThis paper proposes an automatic algorithm for estimating the first two subglottal resonances (SGRs)—Sg1 and Sg2— from continuous speech of children, and applies it to automatic speaker normalization in mismatched, limited-data conditions. The proposed algorithm is based on the observation that Sg1 and Sg2 form phonological vowel feature boundaries, and is motivated by our recent SGR estimation algorithm for adults. The algorithm is trained and evaluated, respectively, on 25 and 9 children, aged between 7 and 18 years. The average RMS errors incurred in estimating Sg1 and Sg2 are 55 and 144 Hz, respectively. By applying the proposed algorithm to a connected digits speech recognition task, it is shown that: 1) a linear frequency warping using Sg1 or Sg2 is comparable to or better than maximum likelihood-based vocal tract length normalization (MLVTLN), 2) the performance of SGR-based frequency warping is less content dependent than that of ML-VTLN, and 3) SGRbased frequency warping can be integrated into ML-VTLN to yield a statistically-significant improvement in performance. Harish Arsikere, Gary K. F. Leung, Steven M. Lulich, Abeer Alwan |
INTERSPEECH | 4 |
| 2012 | Estimating the voice source in noiseabstractEstimation of the glottal source has applications in many areas of speech processing. Therefore, a noise-robust automatic source estimation algorithm is proposed in this paper. The source signal is estimated using a codebook search approach. The glottal area waveforms extracted from high-speed recordings of the glottis is converted to the glottal flow signals in order to evaluate the performance of the proposed source estimation algorithm. Results in clean and noisy conditions, on average, show that the proposed algorithm provides more accurate estimation than the software toolkit Aparat [1] as well as an earlier approach [2]. Gang Chen 0009, Yen-Liang Shue, Jody Kreiman, Abeer Alwan |
INTERSPEECH | 4 |
| 2012 | Evaluation of a Sparse Representation-Based Classifier For Bird Phrase Classification Under Limited Data ConditionsabstractThis paper evaluates the performance of a sparse representation-based (SR) classifier for a limited data, bird phrase classification task. The evaluation database contains 32 unique phrases segmented from songs of the Cassin’s Vireo (Vireo cassinii). Spectrographic features were extracted from each phrase-segmented audio file, followed by dimension reduction using principal component analysis (PCA). A performance comparison to the nearest subspace (NS) and support vector machine (SVM) classifiers was conducted. The SR classifier outperforms the NS and SVM classifiers, with a maximum absolute improvement of 3.4% observed when there are only four tokens per phrase in the training set. Lee Ngee Tan, Kantapon Kaewtip, Martin L. Cody, Charles E. Taylor, Abeer Alwan |
INTERSPEECH | 5 |
| 2012 | SAFE: A Statistical Approach to F0 Estimation Under Clean and Noisy ConditionsabstractA novel Statistical Algorithm for F0 Estimation (SAFE) is proposed to improve the accuracy of F0 estimation under both clean and noisy conditions. Prominent signal-to-noise ratio (SNR) peaks in speech spectra constitute a robust information source from which F0 can be inferred. A probabilistic framework is proposed to model the effect of noise on voiced speech spectra. Prominent SNR peaks in the low-frequency band (0 - 1000 Hz) are important to F0 estimation, and prominent SNR peaks in the middle and high-frequency bands (1000-3000 Hz) are also useful supplemental information to F0 estimation under noisy conditions, especially the babble noise condition. Experiments show that the SAFE algorithm has the lowest gross pitch errors (GPEs) compared to prevailing F0 trackers in white and babble noise conditions at low SNRs. Experimental results also show that SAFE is robust in maintaining a low mean and standard deviation of the fine pitch errors (MFPE and SDFPE) in noise. The code of SAFE is available at http://www.ee.ucla.edu/~weichu/safe. Abeer Alwan |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Automatic estimation of the second subglottal resonance from natural speechabstractThis paper deal s with the automatic estimation of the second subglottal resonance (Sg2) from natural speech spoken by adults, since our previous work focused only on estimating Sg2 from isolated diphthongs. A new database comprising speech and subglottal data of native American English (AE) speakers and bilingual Spanish/English speakers was used for the analysis. Data from 11 speakers (6 females and 5 males) were used to derive an empirical relation among the second and third formant frequencies (F2 and F3) and Sg2. Using the derived relation, Sg2 was automatically estimated from voiced sounds in English and Spanish sentences spoken by 20 different speakers (10 males and 10 females). On average, the error in estimating Sg2 was less than 100 Hz in at least 9 isolated AE vowels and less than 40 Hz in continuous speech consisting of English or Spanish sentences. Harish Arsikere, Steven M. Lulich, Abeer Alwan |
ICASSP | 3 |
| 2011 | Log-spectral amplitude estimation with Generalized Gamma distributions for speech enhancementabstractThis paper presents a family of log-spectral amplitude (LSA) estimators for speech enhancement. Generalized Gamma distributed (GGD) priors are assumed for speech short-time spectral amplitudes (STSAs), providing mathematical flexibility in capturing the statistical behavior of speech. Although solutions are not obtainable in closed-form, estimators are expressed as limits, and can be efficiently approximated. When applied to the Noizeus database [1], proposed estimators are shown to provide improvements in segmental signal-to-noise ratio (SSNR) and COSH distance [2], relative to the LSA estimator proposed by Ephraim and Malah [3]. Bengt J. Borgstrom, Abeer Alwan |
ICASSP | 2 |
| 2011 | Noise-robust F0 estimation using SNR-weighted summary correlograms from multi-band comb filtersabstractA noise-robust, signal-to-noise ratio (SNR)-weighted correlogram-based pitch estimation algorithm (PEA) in which a bank of comb filters operates in each of the low, mid, and high frequency bands is proposed. Correlograms are obtained by applying autocorrelations directly on the low-freq filterbank (FBK) output, and the out put envelopes of all 3 FBKs. An SNR-weighting scheme is used for channel selection to yield a summary correlogram for each FBK. These summary correlograms are averaged to obtain an overall summary correlogram, which is time-smoothed before peak extraction is performed. The final pitch contour is obtained via dynamic programming. The proposed PEA is evaluated on the Keele corpus with additive white or babble noises. In comparison with widely-used PEAs, the proposed PEA has the lowest overall gross pitch error (GPE), especially in low SNR cases. Lee Ngee Tan, Abeer Alwan |
ICASSP | 2 |
| 2011 | Acoustic Correlates of Glottal GapsabstractDuring speech production, the vocal folds may not close completely. The resulting glottal gap (GG) or incomplete glottal closure has not been systematically studied in terms of GG acoustic and/or perceptual consequences. This paper uses high-speed imaging to investigate the relationship between GG area, source parameters, acoustic measures, and voice quality for 6 subjects. Results showed that the cepstral peak prominence (CPP) and the harmonics-to-noise ratio (HNR) are affected by GG area, indicating the presence of more spectral noise with increasing GG area. Analysis of a glide phonation from breathy to pressed for one female speaker showed that measures H ∗ − H ∗ 2 and H ∗ 1 −A ∗ 3 were positively correlated with GG area under a steady fundamental frequency (F 0). In some phonatory modes, increasing F 0 may reduce the amplitude of vocal folds vibration, increase GG area, and produce a lower spectral tilt due to significant aspiration noise, leading to a negative correlation between GG area and the spectral tilt measure H ∗ − A ∗. Gang Chen 0009, Jody Kreiman, Yen-Liang Shue, Abeer Alwan |
INTERSPEECH | 4 |
| 2011 | Joint Robust Voicing Detection and Pitch Estimation Based on Residual HarmonicsabstractThis paper focuses on the problem of pitch tracking in noisy conditions. A method using harmonic information in the residual signal is presented. The proposed criterion is used both for pitch estimation, as well as for determining the voicing segments of speech. In the experiments, the method is compared to six state-of-the-art pitch trackers on the Keele and CSTR databases. The proposed technique is shown to be particularly robust to additive noise, leading to a significant improvement in adverse conditions. Index Terms: fundamental frequency, pitch tracking, pitch estimation, voicing decisions Thomas Drugman, Abeer Alwan |
INTERSPEECH | 2 |
| 2011 | Analysis and Automatic Estimation of Children's Subglottal ResonancesabstractModels and measurements of subglottal resonances are gener-ally made from adult data, but there are several applications in which it would be useful to know about subglottal resonances in children. We therefore conducted an analysis of both new and old recordings of children’s subglottal acoustics in order 1) to produce a fuller picture of the variability of children’s subglottal resonances, and 2) to confirm that existing models of subglottal acoustics can be reasonably applied to children. We also tested the effectiveness of recent algorithms for estimating children’s subglottal resonances from speech formants and the fundamen-tal frequency, which were originally formulated based on adult data. It was found that these algorithms are effective for chil-dren at least 150cm tall. Index Terms: subglottal resonances, child speech, speech pro-duction, speaker normalization Steven M. Lulich, Harish Arsikere, John R. Morton, Gary K. F. Leung, Abeer Alwan, Mitchell Sommers |
INTERSPEECH | 5 |
| 2011 | Perception of place of articulation for plosives and fricatives in noise
Abeer Alwan, Jintao Jiang, Willa S. Chen |
Speech Commun. | 1 |
| 2011 | A Unified Framework for Designing Optimal STSA Estimators Assuming Maximum Likelihood Phase Equivalence of Speech and NoiseabstractIn this paper, we present a stochastic framework for designing optimal short-time spectral amplitude (STSA) estimators for speech enhancement assuming phase equivalence of speech and noise. By assuming additive superposition of speech and noise, which is implied by the maximum-likelihood (ML) phase estimate, we effectively project the optimal spectral amplitude estimation problem onto a 1-D subspace of the complex spectral plane, thus simplifying the problem formulation. Assuming generalized Gamma distributions (GGDs) for a priori distributions of both speech and noise STSAs, we derive separate families of novel estimators according to either the maximum-likelihood (ML), the minimum mean-square error (MMSE), or the maximum a posteriori (MAP) criterion. The use of GGDs allows optimal estimators to be determined in a generalized form, so that particular solutions can be obtained by substituting statistical shape parameters corresponding to expected speech and noise priors. It is interesting to note that several of the proposed estimators exhibit strong similarities to well-known STSA solutions. For example, the magnitude spectral subtracter (MSS) and Wiener filter (WF) are obtained for specific cases of GGD shape parameters. Quantitative analysis of a selected subset of the proposed estimators shows improvement over the traditional log-spectral MMSE estimator of Ephraim and Malah, in terms of segmental signal-to-noise ratio (SNR) and the COSH distance measure, when applied to the Noizeus database. Although single-channel speech enhancement is offered as an illustrative example, the theory presented here could be applicable to other signals, such as music and images. Bengt Jonas Borgstrom, Abeer Alwan |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2011 | A Generative Student Model for Scoring Word Reading SkillsabstractThis paper presents a novel student model intended to automate word-list-based reading assessments in a classroom setting, specifically for a student population that includes both native and nonnative speakers of English. As a Bayesian Network, the model is meant to conceive of student reading skills as a conscientious teacher would, incorporating cues based on expert knowledge of pronunciation variants and their cognitive or phonological sources, as well as prior knowledge of the student and the test itself. Alongside a hypothesized structure of conditional dependencies, we also propose an automatic method for refining the Bayes Net to eliminate unnecessary arcs. Reading assessment baselines that use strict pronunciation scoring alone (without other prior knowledge) achieve 0.7 correlation of their automatic scores with human assessments on the TBALL dataset. Our proposed structure significantly outperforms this baseline, and a simpler data-driven structure achieves 0.87 correlation through the use of novel features, surpassing the lower range of inter-annotator agreement. Scores estimated by this new model are also shown to exhibit the same biases along demographic lines as human listeners. Though used here for reading assessment, this model paradigm could be used in other pedagogical applications like foreign language instruction, or for inferring abstract cognitive states like categorical emotions. Joseph Tepperman, Sungbok Lee, Shri Narayanan, Abeer Alwan |
IEEE Trans. Speech Audio Process. | 4 |
| 2010 | A new voice source model based on high-speed imaging and its application to voice source estimationabstractThere are numerous models of varying complexities which seek to efficiently represent the voice source signal. These models are typically based on data and observations which can come from air-flow masks, electroglottographs, mechanical systems, and the inverse-filtering of speech signals. The first part of this study examines observations from the high-speed imaging of the larynx and proposes a new source model, which is shown to provide a better fit for the observed data than existing models. The proposed source model is then used in an automatic source estimation application, based on methods introduced in an earlier study [1]. Results, on average, show that the proposed model provides a more accurate estimation of the source signal compared with the Liljencrants-Fant model. Yen-Liang Shue, Abeer Alwan |
ICASSP | 2 |
| 2010 | Voice activity detection using harmonic frequency components in likelihood ratio testabstractThis paper proposes a new statistical model-based likelihood ratio test (LRT) VAD to obtain reliable speech / non-speech decisions. In the proposed method, the likelihood ratio (LR) is calculated differently for voiced frames, as opposed to unvoiced frames: only DFT bins containing harmonic spectral peaks are selected for LR computation. To evaluate the new VAD's effectiveness in improving the noise-robustness of ASR, its decisions are applied to pre-processing techniques such as non-linear spectral subtraction, minimum mean square error short-time spectral amplitude estimator, and frame dropping. From the ASR experiments conducted on the Aurora2 database, the proposed harmonic frequency-based LRTs give better results than conventional LRT-based VADs and the standard G.729B and ETSI AMR VADs. Lee Ngee Tan, Bengt J. Borgstrom, Abeer Alwan |
ICASSP | 3 |
| 2010 | Efficient HMM-based estimation of missing features, with applications to packet loss concealmentabstractIn this paper, we present efficient HMM-based techniques for estimating missing features. By assuming speech features to be observations of hidden Markov processes, we derive a minimum mean-square error (MMSE) solution. We increase the computational efficiency of HMM-based methods by downsampling underlying Markov models, and by enforcing symmetry in transitional probability matrices. When applied to features generally utilized in parametric speech coding, namely line spectral frequencies (LSFs), the proposed methods provide significant improvement over the baseline repetition scheme, in terms of Itakura-Saito distortion and peak SNR. Bengt J. Borgstrom, Per Henrik Borgstrom, Abeer Alwan |
INTERSPEECH | 3 |
| 2010 | On using voice source measures in automatic gender classification of children's speechabstractAcoustic characteristics of speech signals differ with gender due to physiological differences of the glottis and the vocal tract. Previous research [1] showed that adding the voice-source related measures H ∗ 1 − H ∗ 2 and H ∗ 1 − A ∗ improved gender classification accuracy compared to using only the fundamental frequency (F0) and formant frequencies. H ∗ i refers to the i–th source spectral harmonic magnitude, and A ∗ refers to the magnitude of the source spectrum at the i–th formant. In this paper, three other voice source related measures: CPP, HNR and H ∗ 2 − H ∗ 4 are used in gender classification of children’s voices. CPP refers to the Cepstral Peak Prominence [2], HNR refers to the harmonic-to-noise ratio [3], and H ∗ 2 − H ∗ 4 refers to the difference between the 2nd and the 4th source spectral harmonic magnitudes. Results show that using these three features improves gender classification accuracy compared with [1]. Index Terms: gender classification, gender identification, voice source Gang Chen 0009, Xue Feng 0002, Yen-Liang Shue, Abeer Alwan |
INTERSPEECH | 4 |
| 2010 | SAFE: a statistical algorithm for F0 estimation for both clean and noisy speechabstractA novel Statistical Approach for F0 Estimation, SAFE, is proposed to improve the accuracy of F0 tracking under both clean and additive noise conditions. Prominent Signal-to-Noise Ratio (SNR) peaks in speech spectra are robust information source from which F0 can be inferred. A probabilistic framework is proposed to model the effect of additive noise on voiced speech spectra. It is observed that prominent SNR peaks located in the low frequency band are important to F0 estimation, and prominent SNR peaks in the middle and high frequency bands are also useful supplemental information to F0 estimation under noisy conditions, especially babble noise condition. Experiments show that the SAFE algorithm has the lowest Gross Pitch Errors (GPE) compared to prevailing F0 trackers: Get F0, Praat, TEMPO, and YIN, in white and babble noise conditions at low SNRs. Abeer Alwan |
INTERSPEECH | 2 |
| 2010 | On the interdependencies between voice quality, glottal gaps, and voice-source related acoustic measuresabstractIn human speech production, the voice source contains important non-lexical information, especially relating to a speaker’s voice quality. In this study, direct measurements of the glottal area waveforms were used to examine the effects of voice quality and glottal gaps on voice source model parameters and various acoustic measures. Results showed that the open quotient parameter, cepstral peak prominence (CPP) and most spectral tilt measures were affected by both voice quality and glottal gaps, while the asymmetry parameter was predominantly affected by voice quality, especially of the breathy type. This was also the case with the harmonic-to-noise ratio measures, indicating the presence of more spectral noise for breathy phonations. Analysis showed that the acoustic measure H1 − H2 was correlated with both the open quotient and asymmetry source parameters, which agrees with existing theoretical studies. Index Terms: voice source, voice quality, acoustic measures Yen-Liang Shue, Gang Chen 0009, Abeer Alwan |
INTERSPEECH | 3 |
| 2010 | On the acoustic correlates of high and low nuclear pitch accents in American English
Yen-Liang Shue, Stefanie Shattuck-Hufnagel, Markus Iseli, Sun-Ah Jun, Nanette Veilleux, Abeer Alwan |
Speech Commun. | 6 |
| 2010 | A Statistical Approach to Mel-Domain Mask Estimation for Missing-Feature ASRabstractIn this letter, we present a statistical approach to Mel-domain mask estimation for missing feature (MF)-based automatic speech recognition (ASR). Mel-domain time-frequency masks are of interest, since MF systems have been shown successful in that domain. Time- and channel-specific reliability measures are derived as posterior probabilities of active speech using a 2-state speech model. Since closed form distributions for Mel-domain spectra do not exist, they are instead modeled as χ2processes with empirically-determined degrees of freedom. Additionally, we present HMM-based decoding to exploit temporal correlation of spectral speech data. The proposed mask estimation algorithm is integrated with an example MF-based ASR front-end from, and is shown to outperform the spectral subtraction (SS)-based method from in terms of word-accuracy, when applied to the Aurora-2 database. Bengt J. Borgstrom, Abeer Alwan |
IEEE Signal Process. Lett. | 2 |
| 2010 | HMM-Based Reconstruction of Unreliable Spectrographic Data for Noise Robust Speech RecognitionabstractThis paper presents a framework for efficient HMM-based estimation of unreliable spectrographic speech data. It discusses the role of hidden Markov models (HMMs) during minimum mean-square error (MMSE) spectral reconstruction. We develop novel HMM-based reconstruction algorithms which exploit intra-channel (across-time) correlation and/or inter-channel (across-frequency) correlation. For the sake of computational efficiency, this paper utilizes approximations to HMM-based decoding methods by developing models constructed from lower resolution quantizers. State configurations for lower resolution models are obtained through a tree-structured mapping of quantizer centroids, and model parameters are adapted accordingly. HMM downsampling avoids expensive retraining of models, and eliminates unnecessary memory requirements. Explicit general formulae are presented for the adaptation of steady-state and transitional statistics. Adaptation of observation statistics are derived from stochastic models of noise spectral magnitude estimation accuracies. The proposed estimation methods are applied in combination with oracle masks, which provide an upper performance bound, as well as masks derived from speech presence probability, which represent a more realistic scenario. Both methods are shown to boost noise robust recognition accuracies significantly relative to the Mel-frequency cepstral coefficient (MFCC) baseline system. Furthermore, HMM downsampling greatly reduces the complexity of the HMM-based reconstruction method while negligibly affecting results. Bengt J. Borgstrom, Abeer Alwan |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Reducing F0 Frame Error of F0 tracking algorithms under noisy conditions with an unvoiced/voiced classification frontendabstractIn this paper, we propose an F0 Frame Error (FFE) metric which combines Gross Pitch Error (GPE) and Voicing Decision Error (VDE) to objectively evaluate the performance of fundamental frequency (F0) tracking methods. A GPE-VDE curve is then developed to show the trade-off between GPE and VDE. In addition, we introduce a model-based Unvoiced/Voiced (U/V) classification frontend which can be used by any F0 tracking algorithm. In the U/V classification, we train speaker independent U/V models, and then adapt them to speaker dependent models in an unsupervised fashion. The U/V classification result is taken as a mask for F0 tracking. Experiments using the KEELE corpus with additive noise show that our statistically-based U/V classifier can reduce VDE and FFE for the pitch tracker TEMPO in both white and babble noise conditions, and that minimizing FFE instead of VDE results in a reduction in error rates for a number of F0 tracking algorithms, especially in babble noise. Abeer Alwan |
ICASSP | 2 |
| 2009 | A correlation-maximization denoising filter used as an enhancement frontend for noise robust bird call classificationabstractIn this paper, we propose a Correlation-Maximization denois-ing filter which utilizes periodicity information to remove addi-tive noise in bird calls. We also developed a statistically-based noise robust bird-call classification system which uses the de-noising filter as a frontend. Enhanced bird calls which are the output of the denoising filter are used for feature extraction. Gaussian Mixture Models (GMM) and Hidden Markov Mod-els (HMM) are used for classification. Experiments on a large noisy corpus containing bird calls from 5 species have shown that the Correlation-Maximization filter is more effective than the Wiener filter in improving the classification error rate of bird calls which have a quasi-periodic structure. This improvement results in a 4.1 % classification error rate which is better than the system without a denoising frontend and a system with a Wiener filter denoising frontend. Index Terms: Correlation-Maximization filter, speech en-hancement, bird call classification Abeer Alwan |
INTERSPEECH | 2 |
| 2009 | A noise-type and level-dependent MPO-based speech enhancement architecture with variable frame analysis for noise-robust speech recognitionabstractIn previous work, a speech enhancement algorithm based on phase opponency and a periodicity measure (MPO-APP) was developed for speech recognition. Axiomatic thresholds were used in the MPO-APP regardless of the signal-to-noise ratio (SNR) of the corrupted speech or any characterization of the noise. The current work developed an algorithm for adjusting the threshold in the MPO-APP based on the SNR and whether the speech signal is clean, corrupted by aperiodic noise or corrupted with noise with periodic components. In addition, variable frame rate (VFR) analysis has been incorporated so that dynamic regions in the speech signal are more heavily sampled than steady-state regions. The result is a 2-stage algorithm that gives superior performance to the previous MPO-APP, and to several other state-of-the-art speech enhancement algorithms. Index Terms: Speech enhancement, robust speech recognition, SNR estimation, variable frame rate analysis, phase opponency. Vikramjit Mitra, Bengt J. Borgstrom, Carol Y. Espy-Wilson, Abeer Alwan |
INTERSPEECH | 4 |
| 2009 | A novel codebook search technique for estimating the open quotientabstractThe open quotient (OQ), loosely defined as the proportion of time the glottis is open during phonation, is an important parameter in many source models. Accurate estimation of OQ from acoustic signals is a non-trivial process as it involves the separation of the source signal from the vocal-tract transfer function. Often this process is hampered by the lack of direct physiological data with which to calibrate algorithms. In this paper, an analysis-by-synthesis method using a codebook of harmonically-based Liljencrants-Fant (LF) source models in conjunction with a constrained optimizer was used to obtain estimates of OQ from four subjects. The estimates were compared with physiological measurements from high-speed imaging. Results showed relatively high correlations between the estimated and measured values for only two of the speakers, suggesting that existing source models may be unable to accurately represent some source signals. Yen-Liang Shue, Jody Kreiman, Abeer Alwan |
INTERSPEECH | 3 |
| 2009 | Bark-shift based nonlinear speaker normalization using the second subglottal resonanceabstractIn this paper, we propose a Bark-scale shift based piecewise nonlinear warping function for speaker normalization, and a joint frequency discontinuity and energy attenuation detection algorithm to estimate the second subglottal resonance (Sg2). We then apply Sg2 for rapid speaker normalization. Experimental results on children’s speech recognition show that the proposed nonlinear warping function is more effective for speaker normalization than linear frequency warping. Compared to maximum likelihood based grid search methods, Sg2 normalization is more efficient and achieves comparable or better performance, especially for limited normalization data. Shizhen Wang, Yi-Hui Lee, Abeer Alwan |
INTERSPEECH | 3 |
| 2009 | Temporal modulation processing of speech signals for noise robust ASRabstractIn this paper, we analyze the temporal modulation char-acteristics of speech and noise from a speech/non-speech discrimination point of view. Although previous psychoacous-tic studies [3][10] have shown that low temporal modulation components are important for speech intelligibility, there is no reported analysis on modulation components from the point of view of speech/noise discrimination. Our data-driven analysis of modulation components of speech and noise reveals that speech and noise is more accurately classified by low-passed modulation frequencies than band-passed ones. Effects of additive noise on the modulation characteristics of speech signals are also analyzed. Based on the analysis, we propose a frequency adaptive modulation processing algorithm for a noise robust ASR task. The algorithm is based on speech channel classification and modulation pattern denoising. Speech recognition experiments are performed to compare the proposed algorithm with other noise robust frontends, including RASTA and ETSI AFE. Recognition results show that the frequency adaptive modulation processing is promising. Index Terms: noise robust speech recognition, temporal mod-ulation processing, spectro-temporal processing 1. Hong You, Abeer Alwan |
INTERSPEECH | 2 |
| 2009 | Frequency warping for VTLN and speaker adaptation by linear transformation of standard MFCC
Sankaran Panchapagesan, Abeer Alwan |
Comput. Speech Lang. | 2 |
| 2009 | Assessment of emerging reading skills in young native speakers and language learners
Patti Price, Joseph Tepperman, Markus Iseli, Thao Duong, Matthew Black, Shizhen Wang, Christy Kim Boscardin, Margaret Heritage, P. David Pearson, Shri Narayanan, Abeer Alwan |
Speech Commun. | 11 |
| 2009 | Utilizing Compressibility in Reconstructing Spectrographic Data, With Applications to Noise Robust ASRabstractIn this letter, we propose a novel algorithm for reconstructing unreliable spectrographic data, a method applicable to missing feature-based automatic speech recognition (ASR). We provide quantitative analysis illustrating the high compressibility of spectrographic speech data. The existence of sparse representations for spectrographic data motivates the spectral reconstruction solution to be posed as an optimization problem minimizing the$\ell_{1}$-norm. When applied to the Aurora-2 database, the proposed missing feature estimation algorithm is shown to provide significant improvements in recognition accuracy relative to the baseline MFCC system. Even without an oracle mask, performance approaches that of the ETSI advanced front end (AFE), with less complexity. Bengt J. Borgstrom, Abeer Alwan |
IEEE Signal Process. Lett. | 2 |
| 2008 | An efficient approximation of the forward-backward algorithm to deal with packet loss, with applications to remote speech recognitionabstractThis paper proposes an efficient approximation of the forward-backward (FB) algorithm, for the purpose of estimating missing features, based on downsampling statistical models. The paper discusses the role of hidden Markov models (HMMs) in the estimation process, and presents an approximation to the FB method by developing HMMs based on lower resolution quantizers, which are obtained through a tree-structure mapping of quantizer centroids. To illustrate the effectiveness of the proposed method, we apply it to the problem of error concealment in remote speech recognition, using the Aurora-2 database. The FB approximation provides comparable word recognition accuracy results relative to the standard FB method, while reducing the computational load by a large factor (> 250 in this case). Bengt J. Borgstrom, Abeer Alwan |
ICASSP | 2 |
| 2008 | Speaker normalization based on subglottal resonancesabstractSpeaker normalization typically focuses on variabilities of the supra-glottal (vocal tract) resonances, which constitute a major cause of spectral mismatch. Recent studies show that the subglottal airways also affect spectral properties of speech sounds. This paper presents a speaker normalization method based on estimating the second and third subglottal resonances. Since the subglottal airways do not change for a specific speaker, the subglottal resonances are independent of the sound type (i.e., vowel, consonant, etc.) and remain constant for a given speaker. This context-free property makes the proposed method suitable for limited data speaker adaptation. This method is computationally more efficient than maximum-likelihood based VTLN, with performance better than VTLN especially for limited adaptation data. Experimental results confirm that this method performs well in a variety of testing conditions and tasks. Shizhen Wang, Abeer Alwan, Steven M. Lulich |
ICASSP | 2 |
| 2008 | Dealing with limited and noisy data in ASR: a hybrid knowledge-based and statistical approachabstractIn this talk, I will focus on the importance of integrating knowledge of human speech production and speech perception mechanisms, and language-specific information with statisticallybased, data-driven approaches to develop robust and scalable automatic speech recognition (ASR) systems. As we will demonstrate, the need for such hybrid systems is especially critical when the ASR system is dealing with noisy data, when adaptation data are limited (for the case of speaker normalization and adaptation), and when dealing with accents. Index Terms: noise-robust ASR, speaker normalization, speaker adaptation, accented English, limited data, knowledgebased Abeer Alwan |
INTERSPEECH | 1 |
| 2008 | HMM-based estimation of unreliable spectral components for noise robust speech recognitionabstractThis paper presents a novel approach for reconstructing unreliable spectral components, which utilizes HMM-based missing feature algorithms, and applies them to noise robust speech recognition. The proposed technique uses the forwardbackward algorithm to estimate corrupt spectrographic data based on nearby reliable features, noisy observations, and on an underlying statistical model. The estimation process can be applied based on intra-channel information, intra-feature information, or a combination of both. The overall system is shown to provide vast improvements for the Consonant Challenge Database [1], for both MFCCs and PLP features, when using an oracle mask. Moreover, through downsampling of statistical models [2], the required complexity of the system is greatly reduced with negligible effects on results. Bengt J. Borgstrom, Abeer Alwan |
INTERSPEECH | 2 |
| 2008 | Vocal tract inversion by cepstral analysis-by-synthesis using chain matricesabstractAcoustic-to-articulatory inversion for vowels is performed by cepstral analysis-by-synthesis, using chain-matrix calculation of vocal tract (VT) acoustics and the Maeda articulatory model. The derivative of the VT chain matrix with respect to the area function was calculated in a novel efficient manner, and used in the BFGS quasi-Newton method for optimizing a distance measure between input and synthesized cepstral features over the entire articulatory trajectory. The optimization is initialized by a fast search of an articulatory codebook with a bin structure in formant space and the cost function also includes regularization and continuity terms to obtain realistic inverted VT shapes and smooth articulatory trajectories. Inversion is evaluated on the three diphthongs /ai/, /oi/ and /au/ of two speakers, one male and one female, from the University of Wisconsin X-ray microbeam (XRMB) database, and good agreement was achieved between inverted midsagittal vocal tract outlines and measured XRMB tongue and lip pellet positions, with an average relative error of less than 3% in the first three formants. Sankaran Panchapagesan, Abeer Alwan |
INTERSPEECH | 2 |
| 2008 | Effects of intonational phrase boundaries on pitch-accented syllables in american EnglishabstractRecent studies of the acoustic correlates of various prosodic elements in American English, such as prominence (in the form of phrase-level pitch accents and word-level lexical stress) and boundaries (in the form of boundary-marking tones), have begun to clarify the nature of the acoustic cues to different types and levels of these prosodic markers. This study focuses on the importance of controlling for context in such investigations, illustrating the effects of adjacent context by examining the cues to H* and L* pitch accent in early and late position in the Intonational Phrase, and how these cues vary when the accented syllable is followed immediately by boundary tones. Results show that F0 peaks for H* accents occur significantly earlier in words that also carry boundary tones, and that energy patterns are also affected; some effects on voice quality measures were also noted. Such findings highlight the caveat that the context of a particular prosodic target may significantly influence its acoustic correlates. Yen-Liang Shue, Stefanie Shattuck-Hufnagel, Markus Iseli, Sun-Ah Jun, Nanette Veilleux, Abeer Alwan |
INTERSPEECH | 6 |
| 2008 | A reliable technique for detecting the second subglottal resonance and its use in cross-language speaker adaptationabstractIn previous work [1], we proposed a speaker adaptation technique based on the second subglottal resonance (Sg2), which showed good performance relative to vocal tract length normalization (VTLN). In this paper, we propose a more reliable algorithm for automatically estimating Sg2 from speech signals. The algorithm is calibrated on children’s speech data collected simultaneously with accelerometer recordings from which Sg2 frequencies can be directly measured. To investigate whether Sg2 frequencies are independent of speech content and language, we perform a cross-language study with bilingual Spanish-English children. The study verifies that Sg2 is approximately constant for a given speaker and thus can be a good candidate for limited data speaker normalization and cross-language adaptation. We then present a cross-language speaker normalization method based on Sg2, which is computationally more efficient than maximum-likelihood based VTLN, and performs more robustly than VTLN. Shizhen Wang, Steven M. Lulich, Abeer Alwan |
INTERSPEECH | 3 |
| 2008 | A Low-Complexity Parabolic Lip Contour Model With Speaker Normalization for High-Level Feature Extraction in Noise-Robust Audiovisual Speech RecognitionabstractThis paper proposes a novel low-complexity lip contour model for high-level optic feature extraction in noise-robust audiovisual (AV) automatic speech recognition systems. The model is based on weighted least-squares parabolic fitting of the upper and lower lip contours, does not require the assumption of symmetry across the horizontal axis of the mouth, and is therefore realistic. The proposed model does not depend on the accurate estimation of specific facial points, as do other high-level models. Also, we present a novel low-complexity algorithm for speaker normalization of the optic information stream, which is compatible with the proposed model and does not require parameter training. The use of the proposed model with speaker normalization results in noise robustness improvement in AV isolated-word recognition relative to using the baseline high-level model. Bengt J. Borgstrom, Abeer Alwan |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 2007 | A Statistical Acoustic Confusability Metric Between Hidden Markov ModelsabstractWith the wide application of hidden Markov models (HMMs) in speech recognition, a statistical acoustic confusability metric is of increasing importance to many components of a speech recognition system. Although distance metrics between HMMs have been studied in the past, they didn't include a way of accounting for speaking rate and durational variations. In order to account for the underlying speech signal's properties when computing such a metric between HMMs, we propose a dynamically-aligned Kullback Leibler (KL) divergence measurement and discuss a cost-efficient implementation of the metric. The proposed approach outperforms existing metrics in predicting phonemic confusions. Hong You, Abeer Alwan |
ICASSP (4) | 2 |
| 2007 | A packetization and variable bitrate interframe compression scheme for vector quantizer-based distributed speech recognitionabstractWe propose a novel packetization and variable bitrate compression scheme for DSR source coding, based on the Group of Pictures concept from video coding. The proposed algorithm simultaneously packetizes and further compresses source coded features using the high interframe correlation of speech, and is compatible with a variety of VQ-based DSR source coders. The algorithm approximates vector quantizers as Markov Chains, and empirically trains the corresponding probability parameters. Feature frames are then compressed as I-frames, P-frames, or B-frames, using Huffman tables. The proposed scheme can perform lossless compression, but is also robust to lossy compression through VQ pruning or frame puncturing. To illustrate its effectiveness, we applied the proposed algorithm to the ETSI DSR source coder. The algorithm provided compression rates of up to 31.60% with negligible recognition accuracy degradation, and rates of up to 71.15% with performance degradation under 1.0%. Bengt J. Borgstrom, Abeer Alwan |
INTERSPEECH | 2 |
| 2007 | Pitch accent versus lexical stress: quantifying acoustic measures related to the voice sourceabstractIn this paper, we explore acoustic correlates of pitch accent and main lexical stress in American English, and the interaction of these cues with other factors that affect prosody. In a controlled study, we varied presence or absence and type of pitch accent (L ∗ vs H∗), boundary-related tone sequence (L-L % vs. H-H%) and gender of the talker, for the sentence “Dagada gave Bobby doodads”. The measures were duration, F0 (fundamen-tal frequency), H∗1−H∗2 (related to open quotient), and H∗1−A∗3 (related to spectral tilt). Contour approximations were used to analyze time-course movements of these measures. For “Da-gada ” we found that, consistent with earlier literature, a) H∗ and L ∗ pitch accents showed different F0 contours, b) pitch-accented syllables were longer than unaccented ones, c) stressed “ga ” syllables had lower H∗1 −H∗2 values than surrounding un-stressed syllables, and for male talkers, lower H∗1 −A∗3 values, indicating lesser spectral tilt. Unexpectedly, F0 maxima asso-ciated with an H ∗ accent occurred most of the time later in the accented syllable than F0 minima associated with L∗. The cues to lexical stress were consistent with or without pitch ac-cent (e.g. lower H∗1 −H∗2), but they sometimes interacted with gender and/or boundary tones: for example, lower H∗1 − A∗3 in stressed “ga ” syllables was only found for female talkers in unaccented cases, and some cues of both accent and stress were less pronounced in the final word “doodads”, which also carried boundary-related tones. Index Terms: voice source, prosody, voice quality 1. Yen-Liang Shue, Markus Iseli, Nanette Veilleux, Abeer Alwan |
INTERSPEECH | 4 |
| 2007 | A Bayesian network classifier for word-level reading assessmentabstractTo automatically assess young children’s reading skills as demonstrated by isolated words read aloud, we propose a novel structure for a Bayesian Network classifier. Our network models the generative story among speech recognition-based features, treating pronunciation variants and reading mistakes as distinct but not independent cues to a qualitative perception of reading ability. This Bayesian approach allows us to estimate the probabilistic dependencies among many highly-correlated features, and to calculate soft decision scores based on the posterior probabilities for each class. With all proposed features, the best version of our network outperforms the C4.5 decision tree classifier by 17% and a Naive Bayes classifier by 8%, in terms of correlation with speaker-level reading scores on the Tball data set. This best correlation of 0.92 approaches the expert inter-evaluator correlation, 0.95. Joseph Tepperman, Matthew Black, Patti Price, Sungbok Lee, Abe Kazemzadeh, Matteo Gerosa, Margaret Heritage, Abeer Alwan, Shri Narayanan |
INTERSPEECH | 8 |
| 2007 | A System for Technology Based Assessment of Language and Literacy in Young Children: the Role of Multiple Information SourcesabstractThis paper describes the design and realization of an automatic system for assessing and evaluating the language and literacy skills of young children. This system was developed in the context of the TBALL (technology based assessment of language and literacy) project and aims at automatically assessing the English literacy skills of both native talkers of American English and Mexican-American children in grades K-2. The automatic assessments were carried out employing appropriate speech recognition and understanding techniques. In this paper, we describe the system focusing on the role of the multiple sources of information at our disposal. We present the content of the assessment system, discuss some issues in creating a child-friendly interface, and how to provide a suitable feedback to the teachers. In addition, we will discuss the different assessment modules and the different algorithms used for speech analysis. Abeer Alwan, Yijian Bai, Matthew Black, Larry Casey, Matteo Gerosa, Margaret Heritage, Markus Iseli, Abe Kazemzadeh, Sungbok Lee, Shri Narayanan, Patti Price, Joseph Tepperman, Shizhen Wang |
MMSP | 1 |
| 2007 | Rate Allocation for Noncollaborative Multiuser Speech Communication Systems Based on Bargaining TheoryabstractWe propose a novel rate allocation algorithm for multiuser speech communication systems based on bargaining theory. Specifically, we apply the generalized Kalai-Smorodinsky bargaining solution since it allows varying bargaining powers to match the dynamic nature of speech signals. We propose a novel method to derive bargaining powers based on the short-time energy of the input speech signals, and subsequently allocate rates accordingly to the users. An important merit of the proposed framework is that it is general and can be applicable for resource allocation across a variety of multirate speech coders, and it is robust to a variety of speech quality metrics. The proposed system is also shown to involve a quick and low-complexity training process. We generalize the algorithm to scenarios in which users have unequally weighted priorities. These scenarios might arise in emergency situations, in which certain users are more important than others. The proposed rate allocation system is shown to increase the utility measures for both the Itakura and segmental signal-to-noise ratio (SNR) functions relative to the baseline system that performs uniform rate allocation. Additionally, although the instantaneous bitrate resolution of the speech encoder is not changed, the proposed system is shown to increase the short-time average bitrate resolution, and therefore provides a greater number of operational rate modes for the network Bengt J. Borgstrom, Mihaela van der Schaar, Abeer Alwan |
IEEE Trans. Speech Audio Process. | 3 |
| 2007 | Robust Speaker Adaptation by Weighted Model Averaging Based on the Minimum Description Length CriterionabstractThe maximum likelihood linear regression (MLLR) technique is widely used in speaker adaptation due to its effectiveness and computational advantages. When the adaptation data are sparse, MLLR performance degrades because of unreliable parameter estimation. In this paper, a robust MLLR speaker adaptation approach via weighted model averaging is investigated. A variety of transformation structures is first chosen and a general form of maximum likelihood (ML) estimation of the structures is given. The minimum description length (MDL) principle is applied to account for the compromise between transformation granularity and descriptive ability regarding the tying patterns of structured transformations with a regression tree. Weighted model averaging across the candidate structures is then performed based on the normalized MDL scores. Experimental results show that this kind of model averaging in combination with regression tree tying gives robust and consistent performance across various amounts of adaptation data Abeer Alwan |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | Speaker Adaptation With Limited Data Using Regression-Tree-Based Spectral Peak AlignmentabstractSpectral mismatch between training and testing utterances can cause significant degradation in the performance of automatic speech recognition (ASR) systems. Speaker adaptation and speaker normalization techniques are usually applied to address this issue. One way to reduce spectral mismatch is to reshape the spectrum by aligning corresponding formant peaks. There are various levels of mismatch in formant structures. In this paper, regression-tree-based phoneme- and state-level spectral peak alignment is proposed for rapid speaker adaptation using linearization of the vocal tract length normalization (VTLN) technique. This method is investigated in a maximum-likelihood linear regression (MLLR)-like framework, taking advantage of both the efficiency of frequency warping (VTLN) and the reliability of statistical estimations (MLLR). Two different regression classes are investigated: one based on phonetic classes (using combined knowledge and data-driven techniques) and the other based on Gaussian mixture classes. Compared to MLLR, VTLN, and global peak alignment, improved performance can be obtained for both supervised and unsupervised adaptations for both medium vocabulary (the RM1 database) and connected digits recognition (the TIDIGITS database) tasks. Performance improvements are largest with limited adaptation data which is often the case for ASR applications, and these improvements are shown to be statistically significant. Shizhen Wang, Abeer Alwan |
IEEE Trans. Speech Audio Process. | 3 |
| 2006 | A Database of Vocal Tract Resonance Trajectories for Research in Speech ProcessingabstractWhile vocal tract resonances (VTRs, or formants that are defined as such resonances) are known to play a critical role in human speech perception and in computer speech processing, there has been a lack of standard databases needed for the quantitative evaluation of automatic VTR extraction techniques. We report in this paper on our recent effort to create a publicly available database of the first three VTR frequency trajectories. The database contains a representative subset of the TIMIT corpus with respect to speaker, gender, dialect and phonetic context, with a total of 538 sentences. A Matlab-based labeling tool is developed, with high-resolution wideband spectrograms displayed to assist in visual identification of VTR frequency values which are then recorded via mouse clicks and local spline interpolation. Special attention is paid to VTR values during consonantto-vowel (CV) and vowel-to-consonant (VC) transitions, and to speech segments with vocal tract anti-resonances. Using this database, we quantitatively assess two common automatic VTR tracking techniques in terms of their average tracking errors analyzed within each of the six major broad phonetic classes as well as during CV and VC transitions. The potential use of the VTR database for research in several areas of speech processing is discussed. Robert Pruvenok, Yanyi Chen, Safiyy Momen, Abeer Alwan |
ICASSP (1) | 6 |
| 2006 | Age-and Gender-Dependent Analysis of Voice Source CharacteristicsabstractThe effects of age, gender, and vocal tract configurations on the glottal excitation signal are still only partially understood. In this paper we examine some of these effects, and show that the voice source parameters, such as fundamental frequency (F0), open quotient (related to H*1- H*2), and spectral tilt (related to H*1-A*3), are not only affected by age and gender but are also intercorrelated (the asterisk superscript denotes correction for the influence of various formants). Recordings of 92 male and female speakers from three age groups (8, 15, 20-39) are analyzed. The main observations are: for lowpitched talkers H*1- H*2(hepce, the open quotient) is proportional to F0, while for high-pitched talkers H*1- H*2is proportional to F1(high to low vowels) for F1*1- A*3showed a strong dependence on F2and F3for all talkers and age groups: increasing F2or F3yielded an increase in H*1- A*3. Spectral tilt was seen to be vowel dependent and for male talkers, spectral tilt changed dramatically with age. A better understanding of the dependencies of voice source parameters on age and gender will help improve voice source parameter estimation and analysis for a variety of speech processing and medical applications. Markus Iseli, Yen-Liang Shue, Abeer Alwan |
ICASSP (1) | 3 |
| 2006 | Multi-Parameter Frequency Warping for Vtln by Gradient SearchabstractThe current method for estimating frequency warping (FW) functions for vocal tract length normalization (VTLN) is by maximizing the ASR likelihood score by an exhaustive search over a grid of FW parameters. Exhaustive search is inefficient when estimating multi-parameter FWs, which have been shown to give improvements in recognition accuracy over single parameter FWs (J.W. McDonough, 2000). Here we develop a gradient search algorithm to obtain the optimal FW parameters for MFCC features, since previous work focussed on PLP cepstral features (J.W. McDonough, 2000). The novel calculation involved was that of the gradient of the Mel filterbank with respect to the FW parameters. Even for a single parameter, the gradient search method was more efficient than grid search by a factor of around 1.6 on the average for male children speakers tested on models trained from adult males. When used to estimate multi-parameter sine-log allpass transform (SLAPT, (J.W. McDonough, 2000)) FWs for VTLN, more than 50% reduction in word error rate was obtained with five parameter SLAPT compared to single-parameter piecewise linear FW Sankaran Panchapagesan, Abeer Alwan |
ICASSP (1) | 2 |
| 2006 | Acoustically-Driven Talking Face Synthesis using Dynamic Bayesian NetworksabstractDynamic Bayesian networks (DBNs) have been widely studied in multi-modal speech recognition applications. Here, we introduce DBNs into an acoustically-driven talking face synthesis system. Three prototypes of DBNs, namely independent, coupled, and product HMMs were studied. Results showed that the DBN methods were more effective in this study than a multilinear regression baseline. Coupled and product HMMs performed similarly better than independent HMMs in terms of motion trajectory accuracy. Audio and visual speech asynchronies were represented differently for coupled HMMs versus product HMMs Jianxia Xue, Jonas Borgstrom, Jintao Jiang, Lynne E. Bernstein, Abeer Alwan |
ICME | 5 |
| 2006 | Voice source correlates of prosodic features in american English: a pilot studyabstractIn this paper, we examine the dependencies of voice source pa-rameters F0(fundamental frequency), Ee(maximal glottal flow change), RK(glottal symmetry/skew), LIN (value related to source spectral tilt) and H∗1 −H∗2 (difference of formant-corrected magnitudes of the first two source spectral harmonics) on prosodic features such as pitch accents, stress, and sentence type and the interdependencies of some of these measures. A small, carefully designed corpus containing a sentence in different prosodic config-urations was used in this study. Statistical analysis was performed using two-way ANOVAs to test for the voice source parameter de-pendencies. Results show that F0 is positively correlated with Ee andLIN, and negatively correlated withH∗1−H∗2. Stressed sylla-bles showed lower values ofRK andH∗1−H∗2 compared to stress-less syllables. The effect of pitch accent can be seen as a combi-nation of its F0, and stress. Phrase-final syllables for interrogative sentences yielded a higher F0 and lower RK and H∗1 −H∗2 com-pared to declarative sentences. It was found that it is important to differentiate between tones when analyzing prosodic features that involve tones, such as pitch accent and probably boundary. Index Terms: voice source, prosody, voice quality. 1. Markus Iseli, Yen-Liang Shue, Melissa A. Epstein, Patricia A. Keating, Jody Kreiman, Abeer Alwan |
INTERSPEECH | 6 |
| 2006 | Automatic detection of voice onset time contrasts for use in pronunciation assessmentabstractThis study examines methods for recognizing different classes of phones from accented speech based on voice on-set time (VOT). These methods are tested on data from the Tball corpus of Los Angeles-area elementary school chil-dren [1]. The methods proposed and tested are: 1) to train models based on standard English VOT contrasts and then extract the VOT characteristics of the phones by measuring the duration of phone-level and sub-phone-level alignments, 2) to train phone models with explicit aspiration, and 3) to train different models for different phoneme classes of VOT times. Error rates of 23-53 % for different phone classes are reported for the rst method, 5-57 % for the second method, and 0-36 % for the third. The results show that different methods work better on different phone classes. We inter-pret these results in relation to past research on VOT, explain possible uses for these ndings, and propose directions for future research. Index Terms: recognition of speech variation, pronuncia-tion variation, voice onset time. Abe Kazemzadeh, Joseph Tepperman, Jorge F. Silva, Hong You, Sungbok Lee, Abeer Alwan, Shri Narayanan |
INTERSPEECH | 6 |
| 2006 | Pronunciation verification of children²s speech for automatic literacy assessmentabstractArguably the most important part of automatically assessing a new reader's literacy is in verifying his pronunciation of read-aloud target words.But the pronunciation evaluation task is especially difficult in children, non-native speakers, and pre-literates.Traditional likelihood ratio thresholding methods do not generalize easily, and even expert human evaluators do not always agree on what constitutes an acceptable pronunciation.We propose new recognition-and alignment-based features in a decision tree classification framework, along with the use of prior linguistic information and human perceptual evaluations.Our classification methods demonstrate a 91% agreement with the voted results of 20 human evaluators who agree among themselves 85% of the time. Joseph Tepperman, Jorge F. Silva, Abe Kazemzadeh, Hong You, Sungbok Lee, Abeer Alwan, Shri Narayanan |
INTERSPEECH | 6 |
| 2006 | Rapid speaker adaptation using regression-tree based spectral peak alignmentabstractIn this paper, regression-tree based spectral peak alignment is proposed for rapid speaker adaptation using the linearization of VTLN. Two different regression classes are investigated: phonetic classes (using combined knowledge and data-driven techniques) and mixture classes. Compared to MLLR and VTLN, improved performance can be obtained for both supervised and unsuper-vised adaptations on both medium vocabulary and connected dig-its recognition tasks. To further improve the performance, MLLR was integrated into this regression-tree based peak alignment. Ex-perimental results show that the performance improvements can be achieved even with limited adaptation data. Index Terms: speaker adaptation, peak alignment, regression tree 1. Shizhen Wang, Abeer Alwan |
INTERSPEECH | 3 |
| 2006 | Adaptation of children's speech with limited data based on formant-like peak alignment
Abeer Alwan |
Comput. Speech Lang. | 2 |
| 2005 | MLLR-like speaker adaptation based on linearization of VTLN with MFCC featuresabstractIn this paper, an MLLR-like adaptation approach is proposed whereby the transformation of the means is performed deter-ministically based on linearization of VTLN. Biases and adap-tation of the variances are estimated statistically by the EM al-gorithm. In the discrete frequency domain, we show that un-der certain approximations, frequency warping with Mel-£lter-bank-based MFCCs equals a linear transformation in the cep-stral domain. Utilizing the deduced linear relationship, the transformation matrix is generated by formant-like peak align-ment. Experimental results using children’s speech show im-provements over traditional MLLR and VTLN. The improve-ments occur even with limited amounts of adaptation data. 1. Abeer Alwan |
INTERSPEECH | 2 |
| 2005 | TBALL data collection: the making of a young children's speech corpusabstractIn this paper we describe the data collection for the TBALL project (Technology Based Assessment of Language and Literacy) and report the results of our efforts. We focus on aspects of our corpus that distinguish it from currently available corpora. The speakers are children (grades K-4), largely nonnative speakers of English, and from diverse socio-economic backgrounds, who are learning to read. We also describe how we adapted our methodology to accommodate these differences: our recording setup, data collection methodology, and transcription scheme. We also discuss the task this corpus was designed to serve and our research approach. Abe Kazemzadeh, Hong You, Markus Iseli, Margaret Heritage, Patti Price, Elaine Andersen, Shri Narayanan, Abeer Alwan |
INTERSPEECH | 10 |
| 2005 | Pronunciation variations of Spanish-accented English spoken by young childrenabstractWhen learning to speak English, non-native speakers may pronounce some English phonemes differently from na-tive speakers. These pronunciation variations can de-grade an automatic speech recognition system’s perfor-mance on accented English. This paper is a first attempt to find common pronunciation variations in Spanish-accented English as spoken by young children. The analysis of pronunciation variation is performed using dynamic programming-based transcription alignment on 4500 words spoken by children 5-7 years old whose first language is Spanish. The findings are then compared with linguistic hypotheses. 1. Hong You, Abeer Alwan, Abe Kazemzadeh, Shri Narayanan |
INTERSPEECH | 2 |
| 2005 | Noise robust speech recognition using feature compensation based on polynomial regression of utterance SNRabstractA feature compensation (FC) algorithm based on polynomial regression of utterance signal-to-noise ratio (SNR) for noise robust automatic speech recognition (ASR) is proposed. In this algorithm, the bias between clean and noisy speech features is approximated by a set of polynomials which are estimated from adaptation data from the new environment by the expectation-maximization (EM) algorithm under the maximum likelihood (ML) criterion. In ASR, the utterance SNR for the speech signal is first estimated and noisy speech features are then compensated for by regression polynomials. The compensated speech features are decoded via acoustic HMMs trained with clean data. Comparative experiments on the Aurora 2 (English) and the German part of the Aurora 3 databases are performed between FC and maximum likelihood linear regression (MLLR). With the Aurora2 experiments, there are two MLLR implementations: pooling adaptation data across all SNRs, and using three distinct SNR clusters. For each type of noise, FC achieves, on average, a word error rate reduction of 16.7% and 16.5% for Set A, and 20.5% and 14.6% for Set B compared to the first and second MLLR implementations, respectively. For each SNR condition, FC achieves, on average, a word error rate reduction of 33.1% and 34.5% for Set A, and 23.6% and 21.4% for Set B. Results using the Aurora3 database show that, the best FC performance outperforms MLLR by 15.9%, 3.0% and 14.6% for well-matched, medium-mismatched and high-mismatched conditions, respectively. Abeer Alwan |
IEEE Trans. Speech Audio Process. | 2 |
| 2004 | Combining feature compensation and weighted Viterbi decoding for noise robust speech recognition with limited adaptation dataabstractAcoustic models trained with clean speech signals suffer in the presence of background noise. In some situations, only a limited amount of noisy data of the new environment is available based on which the clean models could be adapted. A feature compensation approach employing polynomial regression of the signal-to-noise ratio (SNR) is proposed in this paper. While clean acoustic models remain unchanged, a bias which is a polynomial function of utterance SNR is estimated and removed from the noisy feature. Depending on the amount of noisy data available, the algorithm could be flexibly carried out at different levels of granularity. Based on the Euclidean distance, the similarity between the residual distribution and the clean models are estimated and used as the confidence factor in a back-end weighted Viterbi decoding (WVD) algorithm. With limited amounts of noisy data, the feature compensation algorithm outperforms maximum likelihood linear regression (MLLR) for the Aurora2 database. Weighted Viterbi decoding further improves recognition accuracy. Abeer Alwan |
ICASSP (1) | 2 |
| 2004 | An improved correction formula for the estimation of harmonic magnitudes and its application to open quotient estimationabstractMany voice quality parameters, such as the open quotient (OQ), depend on an accurate estimate of the source spectrum. It is known that OQ, for example, is correlated with the magnitude difference of the first two harmonics (H/sub 1/-H/sub 2/) of the speech source spectrum. In order to compare OQ estimates across different vocal tract configurations a magnitude correction is achieved by removing the influence of vocal tract resonances. The improved correction described in this paper is inspired by a correction formula of H. M. Hanson (see Glottal characteristics of female speakers, PhD. dissertation, Harvard university, Cambridge, MA, 1995). The new correction formula accounts for the bandwidths of all the vocal tract resonances, and most importantly, is not limited to the analysis of non-high vowels. H/sub 1/-H/sub 2/ estimates, using the proposed technique with synthesized vowels generated with the Liljencrants-Fant (LF) and the KLGLOTT88 models, are very accurate. Markus Iseli, Abeer Alwan |
ICASSP (1) | 2 |
| 2004 | Entropy-based variable frame rate analysis of speech signals and its application to ASRabstractMost speech processing algorithms analyze speech signals frame by frame with a fixed frame rate. Fixed-rate analysis is inconsistent with human speech perception and effectively assigns the same importance or 'weight' to all equi-duration frames. In Zhu et al. (2000), we proposed a variable frame rate (VFR) analysis technique that is based on a Euclidian distance measure. In this paper, we propose another approach for VFR based on the entropy of the signal. We compare entropy and Euclidian distance measures for VFR in ASR experiments using the Aurora2 and T146 databases. Better performance is observed for the entropy-based VFR over our earlier VFR approach and over the fixed-rate system. Hong You, Qifeng Zhu 0001, Abeer Alwan |
ICASSP (1) | 3 |
| 2003 | Speech recognition over bluetooth wireless channelsabstractThis paper studies the effect of Bluetooth wireless channels on distributed speech recognition. An approach for implementing speech recognition over Bluetooth is described. We simulate a Bluetooth environment and then incorporate its performance, in the form of packet loss ratio, into the speech recognition system. We show how intelligent framing of speech feature vectors, extracted by a fixed-point arithmetic front-end, together with an interpolation technique for lost vectors, can lead to a 50.48% relative improvement in recognition accuracy. This is achieved at a distance of 10 meters, around the maximum operating distance between a Bluetooth transmitter and a Bluetooth receiver. 1. Ziad Al Bawab, Ivo Locher, Jianxia Xue, Abeer Alwan |
INTERSPEECH | 4 |
| 2003 | A noise-robust ASR back-end technique based on weighted viterbi recognitionabstractThe performance of speech recognition systems trained in quiet degrades signicantly under noisy conditions. To address this problem, a Weighted Viterbi Recognition (WVR) algorithm that is a function of the SNR of each speech frame is proposed. Acoustic models trained on clean data, and the acoustic front-end features are kept unchanged in this approach. Instead, a condence/rob ustness factor is assigned to the output observation probability of each speech frame according to its SNR estimate during the Viterbi decoding stage. Comparative experiments are conducted with Weighted Viterbi Recognition with different front-end features such as MFCC, LPCC and PLP. Results show consistent improvements with all three feature vectors. For a reasonable size of adaptation data, WVR outperforms environment adaptation using MLLR. Alexis Bernard, Abeer Alwan |
INTERSPEECH | 3 |
| 2003 | Non-linear feature extraction for robust speech recognition in stationary and non-stationary noise
Qifeng Zhu 0001, Abeer Alwan |
Comput. Speech Lang. | 2 |
| 2003 | Editorial
Abeer Alwan |
Speech Commun. | 1 |
| 2003 | Band-limited feedback cancellation with a modified filtered-X LMS algorithm for hearing aids
Hsiang-Feng Chi, Shawn X. Gao, Sigfrid D. Soli, Abeer Alwan |
Speech Commun. | 4 |
| 2003 | A psychoacoustic-masking model to predict the perception of speech-like stimuli in noise
James J. Hant, Abeer Alwan |
Speech Commun. | 2 |
| 2002 | Efficient adaptation text design based on the Kullback-Leibler measureabstractThis paper proposes an efficient algorithm for the automatic selection of sentences given a desired phoneme distribution. The algorithm is based on the Kullback-Leiblermeasure under the criterion of minimum cross-entropy. One application of this algorithm is the design of adaptation text for automatic speech recognition with a particular phoneme distribution. The algorithm is efficient and flexible, especially in the case of limited text size. Experimental results verify the advantage of this approach. Abeer Alwan |
ICASSP | 2 |
| 2002 | Analysis by synthesis of FM modulation and aspiration noise components in pathological voicesabstractFM and source noise characteristics of pathological voices are analyzed and modeled using precision interpolating pitch tracking. Detailed tracking data allows segregation of pitch variations into low frequency (tremor) and high frequency pitch variation (HFPV) time series. Tremor data is used to resample the original voice into a quasi-constant pitch signal, which results in a more accurate source noise estimate using the noise analysis algorithm described by de Krom [1]. Gaussian distributions are used for both source HFPV and aspiration noise models. Combined analysis parameters are used to drive a formant synthesizer, resulting in improved perceived fidelity. Brian Gabelman, Abeer Alwan |
ICASSP | 2 |
| 2002 | Similarity structure in perceptual and physical measures for visual Consonants across talkersabstractThis paper investigates the relationship between visual confusion matrices and physical (facial) measures. The similarity structure in perceptual and physical measures for visual consonants was examined across four talkers. Four talkers, spanning a wide range of rated visual intelligibility, were recorded producing 69 Consonant-Vowel (CV) syllables. Audio, video, and 3-D face motion were recorded. Each talker's CV productions were presented for identification in a visual-only condition to six viewers with average or better lipreading ability. The obtained visual confusion matrices demonstrated that phonemic equivalence classes were related to visual intelligibility and were talker and vowel context dependent. Physical measures accounted for about 63% of the variance of visual consonant perception, with C/u/ syllables yielding higher correlations than C/a/ and C/i/ syllables. Jintao Jiang, Abeer Alwan, Lynne E. Bernstein, Edward T. Auer, Patricia A. Keating |
ICASSP | 2 |
| 2002 | Predicting face movements from speech acoustics using spectral dynamicsabstractThe paper introduces a new dynamical model which enhances the relationship between face movements and speech acoustics. Based on the autocorrelation of the acoustics and of the face movements, a causal and a non-causal filter are proposed to approximate dynamic features in the speech signals. The database consists of sentences recorded acoustically, and a Qualisys system is used to capture face movements, with 20 reflectors put on the face, simultaneously. Speech signals are represented by 16/sup th/-order LSPs and log-energy. With the filtered dynamic features, the acoustic features account for more than 80% of the variance of face movements. Jintao Jiang, Abeer Alwan, Lynne E. Bernstein, Edward T. Auer, Patricia A. Keating |
ICME (1) | 2 |
| 2002 | Channel noise robustness for low-bitrate remote speech recognitionabstractIn remote (or distributed) speech recognition , the recognition features are quantized at the client, and transmitted to the server via wireless or packet-based communication for recognition. In this paper, we investigate the issue of robustness of remote speech recognition applications against channel noise. The techniques presented include: 1) optimal soft decision channel decoding allowing for error detection, 2) weighted Viterbi recognition (WVR) with weighting coefficients based on the channel decoding reliability, 3) frame erasure concealment, and 4) WVR with weighting coefficients based on the quality of the erasure concealment operation. The techniques presented are implemented at the receiver (server), which limit the complexity for the client, and significantly extend the range of channel conditions for which remote recognition can be sustained. As a case study, we illustrate that remote recognition based on perceptual linear prediction (PLP) coefficients is able to provide at less than 500 bps, good recognition accuracy over a wide range of channel conditions. Alexis Bernard, Abeer Alwan |
INTERSPEECH | 2 |
| 2002 | Evaluation of noise robust features on the Aurora databasesabstracton the Aurora 2 and the German part of Aurora 3. Several algorithms are introduced and evaluated to deal with the noisy speech signals including our previous noise robust techniques used with Aurora 2, and new approaches evaluated with Aurora 3. Since there exist some differences between the two databases, modifications of front-end modules are needed. For Aurora 2, the average error rate reduction is 47% for clean training and 12% for multicondition training compared with the new baseline with endpoint detection. In Aurora 3, we obtain 17%, 27% and 53% error rate reduction for the well-matched, medium-mismatched and high-mismatched cases, respectively. Markus Iseli, Qifeng Zhu 0001, Abeer Alwan |
INTERSPEECH | 4 |
| 2002 | The effect of additive noise on speech amplitude spectra: a quantitative analysisabstractThis article analyzes the effect of additive noise on speech amplitude spectra, and introduces a method to estimate speech spectra from noisy observations. Estimated spectra are used to compute the Mel-frequency cepstral coefficients as a recognition front-end. Compared to linear spectral subtraction, this technique improves the performance of digit recognition in noise. Qifeng Zhu 0001, Abeer Alwan |
IEEE Signal Process. Lett. | 2 |
| 2002 | Low-bitrate distributed speech recognition for packet-based and wireless communicationabstractWe present a framework for developing source coding, channel coding and decoding as well as erasure concealment techniques adapted for distributed (wireless or packet-based) speech recognition. It is shown that speech recognition as opposed to speech coding, is more sensitive to channel errors than channel erasures, and appropriate channel coding design criteria are determined. For channel decoding, we introduce a novel technique for combining at the receiver soft decision decoding with error detection. Frame erasure concealment techniques are used at the decoder to deal with unreliable frames. At the recognition stage, we present a technique to modify the recognition engine itself to take into account the time-varying reliability of the decoded feature after channel transmission. The resulting engine, referred to as weighted Viterbi recognition, further improves the recognition accuracy. Together, source coding, channel coding and the modified recognition engine are shown to provide good recognition accuracy over a wide range of communication channels with bit rates of 1.2 kbps or less. Alexis Bernard, Abeer Alwan |
IEEE Trans. Speech Audio Process. | 2 |
| 2002 | Speech transmission using rate-compatible trellis codes and embedded source codingabstractThis paper presents bandwidth-efficient speech transmission systems using rate-compatible channel coders and variable bitrate embedded source coders. Rate-compatible punctured convolutional codes (RCPC) are often used to provide unequal error protection (UEP) via progressive bit puncturing. RCPC codes are well suited for constellations for which Euclidean and Hamming distances are equivalent (BPSK and 4-PSK). This paper introduces rate-compatible punctured trellis codes (RCPT) where rate compatibility and UEP are provided via progressive puncturing of symbols in a trellis. RCPT codes constitute a special class of codes designed to maximize residual Euclidean distances (RED) after symbol puncturing. They can be designed for any constellation, allowing for higher throughput than when restricted to using 4-PSK. We apply RCPC and RCPT to two embedded source coders: a perceptual subband coder and the ITU embedded ADPCM G.727 standard. Different operating modes with distinct source/channel bit allocation and UEP are defined. Each mode is optimal for a certain range of AWGN channel SNRs. Performance results using an 8-PSK constellation clearly illustrate the wide range of channel conditions at which the adaptive scheme using RCPT can operate. For an 8-PSK constellation, RCPT codes are compared to RCPC with bit interleaved coded modulation codes (RCPC-BICM). We also compare performance to RCPC codes used with a 4-PSK constellation. Alexis Bernard, Xueting Liu 0002, Richard D. Wesel, Abeer Alwan |
IEEE Trans. Commun. | 4 |
| 2001 | Source and channel coding for remote speech recognition over error-prone channelsabstractThis paper presents source and channel coding techniques for remote automatic speech recognition (ASR) systems. As a case study, line spectral pairs (LSP) extracted from the 6th order all-pole perceptual linear prediction (PLP) spectrum are transmitted and speech recognition features are then obtained. The LSPs, quantized using first-order predictive vector quantization (VQ) at 300 bps, provide recognition accuracy comparable to that of the baseline system with no quantization. A new soft decision channel decoding scheme appropriate for remote recognition is presented. The scheme outperforms commonly-used hard decision decoding in terms of error correction and error detection. The source and channel coding system operates at 500 bps and provides good digit recognition performance over a wide range of channel conditions. Alexis Bernard, Abeer Alwan |
ICASSP | 2 |
| 2001 | An efficient and scalable 2D DCT-based feature coding scheme for remote speech recognitionabstractA 2D DCT-based approach to compressing acoustic features for remote speech recognition applications is presented. The coding scheme involves computing a 2D DCT on blocks of feature vectors followed by uniform scalar quantization, runlength and Huffman coding. Digit recognition experiments were conducted in which training was done with unquantized cepstral features from clean speech and testing used the same features after coding and decoding with 2D DCT and entropy coding and in various levels of acoustic noise. The coding scheme results in recognition performance comparable to that obtained with unquantized features at low bitrates. 2D DCT coding of MFCCs (mel-frequency cepstral coefficients) together with a method for variable frame rate analysis (Zhu and Alwan, 2000) and peak isolation (Strope and Alwan, 1997) maintains the noise robustness of these algorithms at low SNRs even at 624 bps. The low-complexity scheme is scalable resulting in graceful degradation in performance with decreasing bit rate. Qifeng Zhu 0001, Abeer Alwan |
ICASSP | 2 |
| 2001 | Joint channel decoding - Viterbi recognition for wireless applicationsabstractWe introduce the concept of joint channel decoding and Viterbi recognition, by which the Viterbi recognizer is modified to take into account the confidence in the decoded feature after channel transmission. We present a metric for evaluating such confidence based on soft decision decoding. As a case study, we quantize MFCCs using predictive VQ. The overall source-channel coding scheme operating at a combined rate of 1 kbps is shown to provide good recognition accuracy over a wide range of Rayleigh fading channels. 1. Alexis Bernard, Abeer Alwan |
INTERSPEECH | 2 |
| 2001 | On the perception of voicing for plosives in noiseabstractPrevious research has shown that the VOT and first formant transition are primary perceptual cues for the voicing distinction for syllable‐initial plosives (SIP) in quiet environments. This study seeks to determine which cues are important for the perception of voicing for SIP in the presence of noise. Stimuli for the perceptual experiments consisted of naturally spoken CV syllables (six plosives in three vowel contexts) in varying levels of additive white Gaussian noise. In each experiment, plosives which share the same place of articulation (e.g., /p, b/) were presented to subjects in identification tasks. For each voiced/voiceless pair, a threshold SNR value was calculated. It was found that the perception of voicing has a strong dependence on vowel context with /Ca/ syllables being significantly better discriminated than /Ci/ and /Cu/ syllables. In addition, labials consistently had a higher threshold SNR (or more easily confusable) than alveolars and velars. Threshold SNR values were then correlated ... Marcia Chen, Abeer Alwan |
INTERSPEECH | 2 |
| 2001 | Predicting visual consonant perception from physical measuresabstractThe long term goal of our work is to predict visual confusion matrices from physical measurements. In this paper, four talkers were chosen to record 69 American-English Consonant-Vowel syllables with audio, video, and facial movements captured. During the recording, 20 markers were put on the face and an optical Qualisys system was used to track three-dimensional facial movements. The videotapes (with markers on the face and without sound) were presented to normal hearing viewers with average or above average lipreading ability, and visual confusion matrices were obtained. Results showed that the facial measurements were correlated with visual perception data by about 0.79 and account for about 63% of the variance. Jintao Jiang, Abeer Alwan, Edward T. Auer, Lynne E. Bernstein |
INTERSPEECH | 2 |
| 2001 | Noise robust feature extraction for ASR using the Aurora 2 databaseabstractFour front-end processing techniques developed for noise robust speech recognition are tested with the Aurora 2 database. These techniques include three previously published algorithms: variable frame rate analysis [Zhu and Alwan, 2000], peak isolation [Strope and Alwan, 1997], and harmonic demodulation [Zhu and Alwan, 2000], and a new technique for peak-to-valley ratio locking. Our previous work has focused on isolated digit recognition. In this paper, these algorithms are modified for recognition of connected digits. Qifeng Zhu 0001, Markus Iseli, Abeer Alwan |
INTERSPEECH | 4 |
| 2000 | On the use of variable frame rate analysis in speech recognitionabstractChanges in spectral characteristics are important cues for discriminating and identifying speech sounds. These changes can occur over very short time intervals. Computing frames every 10 ms, as commonly done in recognition systems, is not sufficient to capture such dynamic changes. In this paper, we propose a variable frame rate (VFR) algorithm. The algorithm results in an increased number of frames for rapidly-changing segments with relatively high energy and less frames for steady-state segments. The current implementation used an average data rate which is less than 100 frames per second. For an isolated word recognition task, and using an HMM-based speech recognition system, the proposed technique results in significant improvements in recognition accuracy especially at low signal-to-noise ratios. The technique was evaluated with mel frequency cepstral coefficient (MFCC) vectors and MFCC vectors with enhanced peak isolation. Qifeng Zhu 0001, Abeer Alwan |
ICASSP | 2 |
| 2000 | Place of articulation cues for voiced and voiceless plosives and fricatives in syllable-initial positionabstractIn this paper, the acoustic correlates of the labial and alveolar place of articulation for both plosive and fricative consonants are investigated, and the results are analyzed in terms of vowel context, voicing and manner of articulation. Several measurements, including formant and noise measurements, are reported for CVs spoken by two male and two female talkers. It was found that the spectral amplitude of frication noise relative to F1 at vowel onset results in 84% or better correct classification for the fricatives in 3 vowel contexts. For plosives, a measure which quantifies the amplitude of noise at high frequencies relative to F1 at vowel onset (Av-Ahi [8]) resulted in 81 % or better correct classification in the three vowel contexts. Formant frequency cues, on the other hand, were not reliable measures for all vowel contexts. 1. INTRODUCTION Various studies have attempted to find invariant acoustic cues for the place of articulation feature for plosives and fricatives. For pl... Willa S. Chen, Abeer Alwan |
INTERSPEECH | 2 |
| 2000 | Predicting the perceptual confusion of synthetic plosive consonants in noise
James J. Hant, Abeer Alwan |
INTERSPEECH | 2 |
| 2000 | Inter- and intra-speaker variability of glottal flow derivative using the LF modelabstractThe vowels /a, i, u/ spoken by American English talkers with non-pathological voices are described by means of voice source model parameters using the Liljencrants-Fant (LF) model. The sampling frequency of the data is 8 kHz which matches approximately telephone bandwidth. After inverse filtering, trends of voice source characteristics depending on the LF parameters are analyzed and compared to literature and listening results. Keywords: voice source, LF model, LF parameters. 1. INTRODUCTION Non-pathological voice source characteristics have been studied by inverse filtering the speech waveform [11], analyzing the speech spectra [6], or by measuring the airflow at the mouth [10]. Knowing the voice source parameters can be beneficial for many speech processing applications, such as speaker identification [8], and speech synthesis. In [6], individual and gender variations in source parameters have been analyzed using measures from speech spectra and taking into account the influence o... Markus Iseli, Abeer Alwan |
INTERSPEECH | 2 |
| 2000 | On the correlation between facial movements, tongue movements and speech acousticsabstractThis study is a first step in a large-scale study that aims at quantifying the relationship between external facial movements, tongue movements, and the acoustics of speech sounds. The database analyzed consisted of 69 CV syllables spoken by two males and two females; each utterance was repeated four times. A Qualysis (optical motion capture system) and an EMA (electromagnetic midsaggital articulography) system were used to characterize facial and tongue movements, respectively. Acoustic features were represented by linear spectral pairs (LSP). To quantify the correlation between them, a multilinear regression technique was applied. The results were analyzed in terms of vowel context, place of articulation, and individual articulatory (EMA or Optical) or acoustic (LSP) channel. 1. INTRODUCTION This study is a first step in a large-scale study that aims at quantifying the relationship between external facial movements, tongue movements, and the acoustics of speech sounds. A recent stu... Jintao Jiang, Abeer Alwan, Lynne E. Bernstein, Patricia A. Keating, Edward T. Auer |
INTERSPEECH | 2 |
| 2000 | AM-demodulation of speech spectra and its application io noise robust speech recognitionabstractIn this paper, a novel algorithm that resembles amplitude demodulation in the frequency domain is introduced, and its application to automatic speech recognition (ASR) is studied. Speech production can be regarded as a result of amplitude modulation (AM) with the source (excitation) spectrum being the carrier and the vocal tract transfer function (VTTF) being the modulating signal. From this point of view, the VTTF can be recovered by amplitude demodulation. Amplitude demodulation of the speech spectrum is achieved by a novel nonlinear technique, which effectively performs envelope detection by using amplitudes of the harmonics and discarding inter-harmonic valleys. The technique is noise robust since frequency bands of low energy are discarded. The same principle is used to reshape the detected envelope. The algorithm is then used to construct an ASR feature extraction module. It is shown that this technique achieves superior performance to MFCCs in the presence of additive noise. Recognition accuracy is further improved if peak isolation [1] is also performed. Qifeng Zhu 0001, Abeer Alwan |
INTERSPEECH | 2 |
| 2000 | Noise source models for fricative consonantsabstractHybrid source models for fricative consonants are derived based on aeroacoustic principles of sound generation employing vocal tract area functions obtained from magnetic resonance imaging data of voiced and unvoiced English fricatives. Results based on data from a male and a female subject indicate that a linear source-filter model is fairly adequate for capturing essential spectral characteristics of sustained voiced and unvoiced strident fricatives below 10 kHz. The hybrid source models employ a combination of acoustic monopole and distributed dipole sources, and a voice source in the case of the voiced fricatives. The number of sources, source locations, and spectral characteristics and the relative source levels are chosen based on an analysis-by-synthesis approach and are motivated by aeroacoustic theory of speech production. The resulting model is computationally efficient and can be readily used for synthesis. Shri Narayanan, Abeer Alwan |
IEEE Trans. Speech Audio Process. | 2 |
| 2000 | Steady-state analysis of continuous adaptation in acoustic feedback reduction systems for hearing-aidsabstractAcoustic feedback is a problem in hearing aids that contain a substantial amount of gain, hearing aids that are used in conjunction with vented or open molds, and in-the-ear hearing aids. Acoustic feedback is both annoying and reduces the maximum usable gain of hearing-aid devices. This paper studies analytically the steady-state convergence behavior of LMS-based adaptive algorithms when used in continuous adaptation to reduce acoustic feedback. A bias is found in the adaptive filter's estimate of the hearing-aid acoustic feedback path. Methods for reducing this bias and producing an improved estimate of the acoustic feedback path are analyzed and compared. It is shown that by the use of a delay in the forward or cancellation paths of the hearing aid plant, and for representative feedback paths, it is possible to reduce this bias by more than 15 dB. Marcio G. Siqueira, Abeer Alwan |
IEEE Trans. Speech Audio Process. | 2 |
| 1999 | Embedded joint source-channel coding of speech using symbol puncturing of trellis codesabstractThis paper presents an embedded joint source-channel coding scheme of speech. The source coder is an embedded variable bit rate perceptually based sub-band coder producing bits with different error sensitivities. The channel encoder is a rate compatible punctured trellis code (RCPT) which permits rate variability and unequal error protection by puncturing symbols. Furthermore, RCPT code design naturally incorporates large constellations, allowing high information rate per symbol. The embedded speech coder and the rate compatible puncturing of symbols provide the embeddibility of the joint coding scheme. The coder is robust to acoustic noise and produces good quality speech for a wide range of channel conditions (AWGN or fading), allowing digital transmission of speech with analog-like graceful degradation. Alexis Bernard, Xueting Liu 0002, Richard D. Wesel, Abeer Alwan |
ICASSP | 4 |
| 1999 | Bias analysis in continuous adaptation systems for hearing aidsabstractThis paper studies analytically the steady-state convergence behavior of adaptive algorithms that approximate the Wiener solution when operating in continuous adaptation to reduce acoustic feedback in hearing aids. A bias is found in the adaptive filter's estimate of the hearing-aid feedback path when the input signal is not white. Delays in the forward and cancellation paths are shown to reduce the magnitude of the bias. Equations for the bias transfer function are obtained. A discussion about properties of the bias when delays are placed in the forward and cancellation paths follows. Marcio G. Siqueira, Abeer Alwan |
ICASSP | 2 |
| 1999 | Perceptually based and embedded wideband CELP coding of speech
Alexis Bernard, Abeer Alwan |
EUROSPEECH | 2 |
| 1999 | Modeling the masking of formant transitions in noiseabstractThe Computer Aided Learning (CAL) working group of the SOCRATES thematic network in Speech Communication Science have studied how the Internet is being used and could be used for the provision of self-study materials for education. In this paper we follow up previous recommendations for the design of Internet tutorials with recommendations for their evaluation. The paper proposes that evaluation should be seen as a necessary quality assurance mechanism operating within the life-cycle of CAL materials development. We propose a structured set of criteria for evaluation, based on the features of good tutorials, against which a tutorial might be judged. Since evaluation against fixed criteria is only one possible approach, we outline how evaluation could also be performed by student users and in controlled trials. James J. Hant, Abeer Alwan |
EUROSPEECH | 2 |
| 1998 | Robust word recognition using threaded spectral peaksabstractA novel technique which characterizes the position and motion of dominant spectral peaks in speech, significantly reduces the error-rate of an HMM-based word-recognition system. The technique includes approximate auditory filtering, temporal adaptation, identification of local spectral peaks in each frame, grouping of neighboring peaks into threads, estimation of frequency derivatives, and slowly updating approximations of the threads and their derivatives. This processing provides a frame-based speech representation which is both dependent on perceptually salient aspects of the frame's immediate context, and is well-suited to segmentally-stationary statistical characterization. In noise, the representation reduces the error-rate obtained with standard Mel-filter-based feature vectors by as much as a factor of 4, and provides improvements over other common feature-vector manipulations. Brian Strope, Abeer Alwan |
ICASSP | 2 |
| 1997 | Acoustic modelling of American English /r/
Carol Y. Espy-Wilson, Shri Narayanan, Suzanne Boyce, Abeer Alwan |
EUROSPEECH | 4 |
| 1997 | New results in vowel production: MRI, EPG, and acoustic data
Shri Narayanan, Abeer Alwan |
EUROSPEECH | 2 |
| 1997 | Towards articulatory speech recognition: learning smooth maps to recover articulator informationabstractWe present a novel method for recovering articulator movements from speech acoustics based on a constrained form [9] of a hidden Markov model. The model attempts to explain sequences of high dimensional data using smooth and slow trajectories in a latent variable space. The key insight is that this continuity constraint when applied to speech helps to solve the \ill-posed problem of acoustic to articulatory mapping. By working with sequences of spectra rather than looking only at individual spectra, it is possible to choose between competing articulatory con gurations for any given spectrum by selecting the con guration \closest to those at nearby times. We present results of applying this algorithm to recover articulator movements from acoustics using data from the Wisconsin X-ray microbeam project [3]. We nd that the recovered traces are highly correlated with the measured articulator movements under a single linear transform. Such recovered traces have the potential to be used for speech recognition, an application we are currently investigating. Sam T. Roweis, Abeer Alwan |
EUROSPEECH | 2 |
| 1997 | Analysis by synthesis of pathological voices using the Klatt synthesizer
Philbert Bangayan, Abeer Alwan, Jody Kreiman, Bruce R. Gerratt |
Speech Commun. | 3 |
| 1997 | A model of dynamic auditory perception and its application to robust word recognitionabstractThis paper describes two mechanisms that augment the common automatic speech recognition (ASR) front end and provide adaptation and isolation of local spectral peaks. A dynamic model consisting of a linear filterbank with a novel additive logarithmic adaptation stage after each filter output is proposed. An extensive series of perceptual forward masking experiments, together with previously reported forward masking data, determine the model's dynamic parameters. Once parameterized, the simple exponential dynamic mechanism predicts the nature of forward masking data from several studies across wide ranging frequencies, input levels, and probe delay times. An initial evaluation of the dynamic model together with a local peak isolation mechanism as a front end for dynamic time warp (DTW) and hidden Markov model (HMM) word recognition systems shows an improvement in robustness to background noise when compared to Mel-frequency cepstral coefficients (MFCC), linear prediction cepstral coefficients (LPCC), and relative spectra (RASTA) based front ends. Brian Strope, Abeer Alwan |
IEEE Trans. Speech Audio Process. | 2 |
| 1997 | A perceptually based embedded subband speech coderabstractA new scheme for robust, high-quality, embedded speech coding based on subband decomposition and perceptually optimized bit allocation and prioritization is presented. An infinite impulse response (IIR) quadrature mirror filterbank (QMF) performs subband decomposition. A perceptual model, computed using subband spectral analysis, optimizes the coder's perceptual quality. Dynamic bit allocation and prioritization is combined with embedded quantization resulting in little performance degradation relative to a nonembedded implementation. The coder output is scalable from high quality at higher bit rates to lower quality at lower bit rates, supporting a wide range of service and resource utilization. The lower bit-rate representation is obtained simply through truncation of the higher bit-rate representation. Since source-rate adaptation is performed through truncation of the encoded stream, interaction with the coder is not required, making the embedded coder ideally suited for rate-adaptive communication systems. Performance for both speech and music was verified through subjective listening tests. Benjamim Tang, Albert Shen, Abeer Alwan, Gregory J. Pottie |
IEEE Trans. Speech Audio Process. | 3 |
| 1996 | Parametric hybrid source models for voiced and voiceless fricative consonantsabstractSource models for fricative consonants are derived based on aerodynamic principles of sound generation, in conjunction with vocal-tract models obtained from MRI data. Results indicate that a linear source-filter model is adequate for capturing essential spectral characteristics of sustained fricatives below 10 kHz. The hybrid source models employ a combination of acoustic monopole and dipole sources, and a voiced source in the case of the voiced fricatives. The number of sources, source locations and spectral characteristics are chosen based on an analysis-by-synthesis approach and are motivated by aeroacoustic theory. The resulting model is computationally efficient and can be readily used for synthesis. Shri Narayanan, Abeer Alwan |
ICASSP | 2 |
| 1996 | A model of dynamic auditory perception and its application to robust speech recognitionabstractThis paper derives a non-linear model of dynamic auditory perception. The model consists of a linear filter bank with carefully-parameterized logarithmic additive adaptation after each filter output. An extensive series of perceptual forward masking experiments, together with previously reported forward masking data, determine the model's dynamic parameters. The model's prediction error of forward masking data has a standard deviation of less than 3.3 dB across wide ranging frequencies, input levels, and probe delay times. We present an initial evaluation of the dynamic model as a front end for an isolated word recognition system, and show an improvement in the robustness to background noise when compared to MFCC and LPCC front ends. Brian Strope, Abeer Alwan |
ICASSP | 2 |
| 1996 | From MRI and acoustic data to articulatory synthesis: a case study of the lateral approximants in american English
Philbert Bangayan, Abeer Alwan, Shri Narayanan |
ICSLP | 2 |
| 1996 | A psychoacoustic model for the noise masking of voiceless plosive burstsabstractA model for predicting the masked thresholds of the voiceless plosive bursts /k,t,p/ in background noise is proposed.Because plosive bursts are brief, are generated by a noise source, and have different spectral characteristics, the modeling approach must account for duration, center frequency, signal bandwidth and type.To achieve this goal, noise-in-noise masking experiments are conducted using a broad band masker and bandpass noise signals of varying bandwidth (1-8 CB), duration (10-300 ms), and center frequency (0.4-4 kHz).The results of these experiments are used to parameterize an auditory filter model in which the effective bandwidths of the filters and the signal-to-noise ratio at threshold are frequency and durationdependent.The duration-dependent filter model is then used to predict the thresholds of both synthetic and naturally-spoken plosive bursts in background noise. James J. Hant, Brian Strope, Abeer Alwan |
ICSLP | 3 |
| 1996 | Liquids in tamil
Shri Narayanan, Abigail Kaun, Dani Byrd, Peter Ladefoged, Abeer Alwan |
ICSLP | 5 |
| 1995 | A robust variable-rate speech coderabstractThe goal of this study is to develop a robust and high-quality speech coder for wireless communication. The proposed coder is a perceptually-based variable-rate subband coder. The perceptual metric ensures that encoding is optimized to the human listener and is based on calculating the signal-to-mask ratio in short-time frames of the input signal. An adaptive bit allocation scheme is employed and the subband energies are then quantized using a Max-Lloyd quantizer. The coder is fully scalable-increasing the bit rates, improves the quality of encoded speech. Subjective listening tests, using quiet and noisy input signals, indicate that the proposed coder produces high-quality speech when operating at 12 kbps or higher. In error-free conditions, our coder has comparable performance to that of QCELP or GSM coders. For speech in background noise, however, our coder, at 12 kbps, outperforms QCELP significantly, and for music, it outperforms both QCELP and GSM. Albert Shen, Benjamim Tang, Abeer Alwan, Gregory J. Pottie |
ICASSP | 3 |
| 1995 | A novel structure to compensate for frequency-dependent loudness recruitment of sensorineural hearing lossabstractA simple structure that compensates for frequency-dependent loudness recruitment of sensorineural hearing loss is presented. The non-linear structure estimates total input signal energy, and then weights and combines the output of two parallel filters based on this energy estimation. Preliminary evaluation of this structure with noise-masked normal hearing listeners, and speech recorded in naturally noisy environments, shows a 15-20% performance increase in word recognition scores when compared to a linear structure. Brian Strope, Abeer Alwan |
ICASSP | 2 |
| 1995 | Spectral analysis of subband filtered signalsabstractA methodology for transform-based spectral analysis of subband filtered signals is developed. The methodology is based on performing the analysis on subband samples instead of on the input signal directly. Aliasing due to decimation is eliminated by including the effects of the adjacent subband in the analysis of frequencies near the filterbank transition regions. The frequency resolution and spectral leakage is nearly the same as if the transform had been performed on the input directly. In an M band filterbank, the analysis block length is reduced by a factor of M. This reduces the complexity of source compression techniques based on subband decomposition and spectral analysis. Benjamim Tang, Albert Shen, Gregory J. Pottie, Abeer Alwan |
ICASSP | 4 |
| 1995 | Finite Precision Analysis of the Fast QRD-RLS Lattice Algorithm
Marcio G. Siqueira, Abeer Alwan, Paulo S. R. Diniz |
ISCAS | 2 |
| 1994 | New adaptive-filtering techniques applied to speech echo cancellationabstractDeveloping high-quality echo cancelers is an important area of research and is becoming increasingly so with the wide use of hands-free telephones. Although LS algorithms have a faster convergence rate than the more widely-used LMS algorithms, the LS algorithms have not been popular in echo cancelers either because of their bad numerical properties and/or because of large complexity overhead. The authors investigate the performance of a fast RLS algorithm (FQRD-RLS) in the context of speech echo cancellation. The FQRD algorithm is not computationally complex and appears to possess good numerical properties. The computer simulations involve both direct and subband-filtering implementations. The results indicate that the FQRD-RLS is a powerful option for speech echo cancellation.> Marcio G. Siqueira, Abeer Alwan |
ICASSP (2) | 2 |
| 1994 | An MRI study of fricative consonants
Shri Narayanan, Abeer Alwan, Katherine Haker |
ICSLP | 2 |
| 1994 | High-Performance IIR QMF Banks for Speech Subband CodingabstractIn this paper, two high-performance implementations of IIR QMF banks for speech coding are proposed. The first implementation involves a perfect reconstruction (PR) QMF bank while the second is a tree-structured filter bank. In contrast to existing implementations of PR QMF banks, our approach does not require the transmission of initial conditions nor the design of synthesis filters which are more complex than their analysis counterparts. In systems with delay constraints, the tree structured IIR QMF bank can be used; a 4-channel scheme shows a low delay of about 1.9 ms (15 samples) at an 8 kHz sampling rate. No existing FIR QMF banks can achieve this low delay. Moreover, the phase distortion of the low-delay filter bank does not appear to affect the perceptual quality of the processed speech signals. Subjective tests were conducted to evaluate the speech quality.> Zhongnong Jiang, Abeer Alwan, Alan N. Willson Jr. |
ISCAS | 2 |
| 1994 | Infinite Precision Analysis of the Fast QR Decomposition RLS AlgorithmabstractThis work develops relations for the mean squared value of internal variables in the fast QRD-RLS. The objective is to derive relations based on known characteristics of input signals that predict the behavior of the internal quantities of the algorithm. It is shown that the fast and conventional QRD-RLS algorithms have some variables in common, and thus previous results of the infinite precision analysis of the conventional algorithm remain valid for the fast version. Conditions for avoiding over flow in fixed-point implementations are presented. Simulation results are also shown.> Marcio G. Siqueira, Paulo S. R. Diniz, Abeer Alwan |
ISCAS | 3 |
| 1993 | A perceptual metric for masking
Abeer Alwan |
ICASSP (2) | 1 |
| 1993 | Strange attractors and chaotic dynamics in the production of voiced and voiceless fricatives
Shri Narayanan, Abeer Alwan |
EUROSPEECH | 2 |
| 1992 | The role of F3 and F4 in identifying place of articulation for stop consonants
Abeer Alwan |
ICSLP | 1 |