EDBT 2026 Demo / reviewers in the wild / expert
Thomas F. Quatieri
dblp:67/4478
· DBLP profile ↗
129ranked-venue papers
30as first author
13since 2021 · last 2025
0000-0003-1925-6340ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 100 · 25 first-author · 8 since 2021Artificial intelligence and machine learning · 56 · 9 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards the Objective Characterisation of Major Depressive Disorder Using Speech Data from a 12-week Observational Study with Daily Measurements
Robert Lewis 0001, Szymon Fedor, Nelson Hidalgo Julia, Joshua Curtiss, Jiyeon Kim, Noah Jones, David Mischoulon, Thomas F. Quatieri, Nicholas Cummins, Paola Pedrelli, Rosalind W. Picard |
INTERSPEECH | 8 |
| 2025 | An exploratory characterization of speech- and fine-motor coordination in verbal children with Autism spectrum disorderabstractAutism spectrum disorder (ASD) is a neurodevelopmental disorder often associated with difficulties in speech production and fine-motor tasks. Thus, there is a need to develop objective measures to assess and understand speech production and other fine-motor challenges in individuals with ASD. In addition, recent research suggests that difficulties with speech production and fine-motor tasks may contribute to language difficulties in ASD. In this paper, we explore the utility of an off-body recording platform, from which we administer a speech- and fine-motor protocol to verbal children with ASD and neurotypical controls. We utilize a correlation-based analysis technique to develop proxy measures of motor coordination from signals derived from recordings of speech- and fine-motor behaviors. Eigenvalues of the resulting correlation matrix are inputs to Gaussian Mixture Models to discriminate between highly-verbal children with ASD and neurotypical controls. These eigenvalues also characterize the complexity (underlying dimensionality) of representative signals of speech- and fine-motor movement dynamics, and form the feature basis to estimate scores on an expressive vocabulary measure. Based on a pilot dataset (15 ASD and 15 controls), features derived from an oral story reading task are used in discriminating between the two groups with AUCs > 0.80, and highlight lower complexity of coordination in children with ASD. Features derived from handwriting and maze tracing tasks led to AUCs of 0.86 and 0.91, however features derived from ocular tasks did not aid in discrimination between the ASD and neurotypical groups. In addition, features derived from free speech and sustained vowel tasks are strongly correlated with expressive vocabulary scores. These results indicate the promise of a correlation-based analysis in elucidating motor differences between individuals with ASD and neurotypical controls. Tanya Talkar, James R. Williamson, Sophia Yuditskaya, Daniel J. Hannon, Hrishikesh Rao 0002, Lisa Nowinski, Hannah Saro, Maria Mody, Christopher J. McDougle, Thomas F. Quatieri |
Comput. Speech Lang. | 10 |
| 2024 | A Neurophysiological-Auditory "Listen Receipt" for Communication EnhancementabstractInformation overload, and specifically auditory overload, is common in critical situations and detrimental to communication. Currently, there is no auditory equivalent of an email read receipt to know if a person has heard a message, other than waiting for a reply. This work hypothesizes that it may be possible to decode whether a person has indeed heard a message, or in other words, create an an auditory “listen receipt,” through use of non-invasive physiological or neural monitoring. We extracted a variety of features derived from Electrodermal activity (EDA), Electroencephalography (EEG), and the correlations between the acoustic envelope of the radio message and EEG to use in the decoder. We were able to classify the cases in which the subject responded correctly to the question in the message, versus the cases where they missed or heard the message incorrectly, with an accuracy of 79% and a receiver operating characteristic (ROC) area under the curve (AUC) of 0.83. This work suggests that the concept of a “listen receipt” may be possible, and future wearable machine-brain interface technologies may be able to automatically determine if an important radio message has been missed for both human-to-human and human-to-machine communication. Christine Beauchene, Michael S. Brandstein, Thomas F. Quatieri, Eric Thompson, Christopher J. Smalt |
ICASSP | 3 |
| 2024 | Variability of speech timing features across repeated recordings: a comparison of open-source extraction techniquesabstractVariations in speech timing features have been reliably linked to symptoms of various health conditions, demonstrating clinical potential. However, replication challenges hinder their translation; extracted speech features are susceptible to methodological variations in the recording and processing pipeline. Investigating this, we compared exemplar timing features extracted via three different techniques from recordings of healthy speech. Our results show that features extracted via an intensity-based method differ from those produced by forced alignment. Different extraction methods also led to differing estimates of within-speaker feature variability over time in an analysis of recordings repeated systematically over three sessions in one day (n=26) and in one week (n=28). Our findings highlight the importance of feature extraction in study design and interpretation, and the need for consistent, accurate extraction techniques for clinical research. Index Terms: speech timing, feature extraction, reproducibility, longitudinal monitoring Judith Dineley, Ewan Carr, Lauren L. White, Catriona Lucas, Zahia Rahman, Tian Pan 0004, Faith Matcham, Johnny Downs, Richard J. B. Dobson, Thomas F. Quatieri, Nicholas Cummins |
INTERSPEECH | 10 |
| 2024 | Analyzing Speech Motor Movement using Surface Electromyography in Minimally Verbal Adults with Autism Spectrum Disorder
Wazeer Zulfikar, Nishat Protyasha, Camila Canales, Heli Patel, James Williamson, Laura Sarnie, Lisa Nowinski, Nataliya Kos'myna, Paige Townsend, Sophia Yuditskaya, Tanya Talkar, Utkarsh Oggy Sarawgi, Christopher J. McDougle, Thomas F. Quatieri, Pattie Maes, Maria Mody |
INTERSPEECH | 14 |
| 2023 | A Vocal Model to Predict Readiness under Sleep DeprivationabstractA variety of factors can affect cognitive readiness and influence human performance in tasks that are mission critical. Sleep deprivation is one of the most prevalent factors that degrade performance. One risk mitigation approach is to use vocal biomarkers to detect cognitive fatigue and resulting performance decrements [1], [2]. In this study, a group of 20 subjects were deprived of sleep for a period of 24 hours. Every two hours, they performed a battery of both speech tasks and cognitive performance tasks, including the psychomotor vigilance test (PVT). Performance on the PVT declined dramatically during nighttime hours between 2 AM and 8 AM. We demonstrate that a model using vocal biomarkers from read speech and free speech can be successfully trained to detect performance decrements on the PVT. We also demonstrate that the vocal model successfully generalizes to other outcomes at a similar level as PVT, detecting sleep deprivation (AUC=0.79) and cognitive performance declines on a battery of cognitive tasks (AUC=0.79). In comparison, using PVT as the basis for detecting sleep deprivation and performance declines resulted in AUC=0.75 and AUC=0.80, respectively.Clinical Relevance—This paper provides evidence to suggest that vocal models trained to predict changes in reaction times due to sleep deprivation can generalize to detect sleep deprivation and associated performance declines on more complex cognitive tasks James R. Williamson, Elizabeth Godoy, Thomas F. Quatieri |
BSN | 3 |
| 2023 | Subject-Specific Adaptation for a Causally-Trained Auditory-Attention Decoding SystemabstractFuture hearing-aid technology may allow a listener to isolate a single talker of interest from a mixture by shifting their attention as measured by Electroencephalography (EEG). Such decoding algorithms are often trained with data from a single individual or a pool of several participants (i.e., group model). Performance in either approach is limited: group models suffer due to the variability across subjects and time, while individual models are constrained by the limited data samples available. To overcome this challenge, we introduce a subject-specific adaptive form of auditory attention decoding (AAD) over short time windows to account for the variability across EEG recording sessions. Our subject-specific augmented model, adapts a group model to an individual, significantly improving decoding accuracy by approximately 10% as compared to an individual model. This result has implications for real-time applications of neuro-steered hearing aids, where causal-training data and real-time algorithms are necessary. Christine Beauchene, Michael S. Brandstein, Stephanie Haro, Thomas F. Quatieri, Christopher J. Smalt |
ICASSP | 4 |
| 2023 | Towards robust paralinguistic assessment for real-world mobile health (mHealth) monitoring: an initial study of reverberation effects on speechabstractSpeech is promising as an objective, convenient tool to monitor health remotely over time using mobile devices. Numerous paralinguistic features have been demonstrated to contain salient information related to an individual’s health. However, mobile device specification and acoustic environments vary widely, risking the reliability of the extracted features. In an initial step towards quantifying these effects, we report the variability of 13 exemplar paralinguistic features commonly reported in the speech-health literature and extracted from the speech of 42 healthy volunteers recorded consecutively in rooms with low and high reverberation with one budget and two higher-end smartphones, and a condenser microphone. Our results show reverberation has a clear effect on several features, in particular voice quality markers. They point to new research directions investigating how best to record and process in-the-wild speech for reliable longitudinal health state assessment. Judith Dineley, Ewan Carr, Faith Matcham, Johnny Downs, Richard J. B. Dobson, Thomas F. Quatieri, Nicholas Cummins |
INTERSPEECH | 6 |
| 2022 | Affective Ratings of Nonverbal Vocalizations Produced by Minimally-Speaking Individuals: What Do Naive Listeners Perceive?abstractIndividuals who produce few spoken words (“minimally-speaking” individuals) often convey rich affective and communicative information through nonverbal vocalizations, such as grunts, yells, babbles, and monosyllabic expressions. Yet, little data exists on the affective content of the vocal expressions of this population. Here, we present 78,624 arousal and valence ratings of nonverbal vocalizations from the online ReCANVo (Real-World Communicative and Affective Nonverbal Vocalizations) database. This dataset contains over 7,000 vocalizations that have been labeled with their expressive functions (delight, frustration, etc.) from eight minimally-speaking individuals. Our results suggest that raters who have no knowledge of the context or meaning of a nonverbal vocalization are still able to detect arousal and valence differences between different types of vocalizations based on Likert-scale ratings. Moreover, these ratings are consistent with hypothesized arousal and valence rankings for the different vocalization types. Raters are also able to detect arousal and valence differences between different vocalization types within individual speakers. To our knowledge, this is the first large-scale analysis of affective content within nonverbal vocalizations from minimally verbal individuals. These results complement affective computing research of nonverbal vocalizations that occur within typical verbal speech (e.g., grunts, sighs) and serve as a foundation for further understanding of how humans perceive emotions in sounds. Kristina T. Johnson, Amanda O'Brien, Ayelet M. Kershenbaum, Jaya Narain, Simon Radhakrishnan, Thomas F. Quatieri, Rosalind W. Picard |
ACII | 6 |
| 2022 | Speech Acoustics in Mild Cognitive Impairment and Parkinson's Disease With and Without Concurrent Drawing Tasks
Tanya Talkar, Christina Manxhari, James J. Williamson, Kara M. Smith, Thomas F. Quatieri |
INTERSPEECH | 5 |
| 2022 | Modeling Real-World Affective and Communicative Nonverbal Vocalizations From Minimally Speaking IndividualsabstractNonverbal vocalizations from non- and minimally speaking individuals who speak fewer than 20 words (mv* individuals) convey important communicative and affective information. While nonverbal vocalizations that occur amidst typical speech and infant vocalizations have been studied extensively in the literature, there is limited prior work on vocalizations by mv* individuals. Our work is among the first studies of the communicative and affective information expressed in nonverbal vocalizations by mv* children and adults. We collected labeled vocalizations in real-world settings with eight mv* communicators, with communicative and affective labels provided in-the-moment by a close family member. Using evaluation strategies suitable for messy, real-world data, we show that nonverbal vocalizations can be classified by function (with 4- and 5-way classifications) with F1 scores above chance for all participants. We analyze labeling and data collection practices for each participating family, and discuss the classification results in the context of our novel real-world data collection protocol. The presented work includes results from the largest classification experiments with nonverbal vocalizations from mv* communicators to date. Jaya Narain, Kristina T. Johnson, Thomas F. Quatieri, Rosalind W. Picard, Pattie Maes |
IEEE Trans. Affect. Comput. | 3 |
| 2021 | Acoustic Indicators of Speech Motor Coordination in Adults With and Without Traumatic Brain Injury
Tanya Talkar, Nancy Pearl Solomon, Douglas Brungart, Stefanie E. Kuchinsky, Megan M. Eitel, Sara M. Lippa, Tracey A. Brickell, Louis M. French, Rael T. Lange, Thomas F. Quatieri |
Interspeech | 10 |
| 2021 | Speaker separation in realistic noise environments with applications to a cognitively-controlled hearing aidabstractFuture wearable technology may provide for enhanced communication in noisy environments and for the ability to pick out a single talker of interest in a crowded room simply by the listener shifting their attentional focus. Such a system relies on two components, speaker separation and decoding the listener's attention to acoustic streams in the environment. To address the former, we present a system for joint speaker separation and noise suppression, referred to as the Binaural Enhancement via Attention Masking Network (BEAMNET). The BEAMNET system is an end-to-end neural network architecture based on self-attention. Binaural input waveforms are mapped to a joint embedding space via a learned encoder, and separate multiplicative masking mechanisms are included for noise suppression and speaker separation. Pairs of output binaural waveforms are then synthesized using learned decoders, each capturing a separated speaker while maintaining spatial cues. A key contribution of BEAMNET is that the architecture contains a separation path, an enhancement path, and an autoencoder path. This paper proposes a novel loss function which simultaneously trains these paths, so that disabling the masking mechanisms during inference causes BEAMNET to reconstruct the input speech signals. This allows dynamic control of the level of suppression applied by BEAMNET via a minimum gain level, which is not possible in other state-of-the-art approaches to end-to-end speaker separation. This paper also proposes a perceptually-motivated waveform distance measure. Using objective speech quality metrics, the proposed system is demonstrated to perform well at separating two equal-energy talkers, even in high levels of background noise. Subjective testing shows an improvement in speech intelligibility across a range of noise levels, for signals with artificially added head-related transfer functions and background noise. Finally, when used as part of an auditory attention decoder (AAD) system using existing electroencephalogram (EEG) data, BEAMNET is found to maintain the decoding accuracy achieved with ideal speaker separation, even in severe acoustic conditions. These results suggest that this enhancement system is highly effective at decoding auditory attention in realistic noise environments, and could possibly lead to improved speech perception in a cognitively controlled hearing aid. Bengt J. Borgstrom, Michael S. Brandstein, Gregory A. Ciccarelli, Thomas F. Quatieri, Christopher J. Smalt |
Neural Networks | 4 |
| 2020 | Personalized Modeling of Real-World Vocalizations from Nonverbal IndividualsabstractNonverbal vocalizations contain important affective and communicative information, especially for those who do not use traditional speech, including individuals who have autism and are non- or minimally verbal (nv/mv). Although these vocalizations are often understood by those who know them well, they can be challenging to understand for the community-at-large. This work presents (1) a methodology for collecting spontaneous vocalizations from nv/mv individuals in natural environments, with no researcher present, and personalized in-the-moment labels from a family member; (2) speaker-dependent classification of these real-world sounds for three nv/mv individuals; and (3) an interactive application to translate the nonverbal vocalizations in real time. Using support-vector machine and random forest models, we achieved speaker-dependent unweighted average recalls (UARs) of 0.75, 0.53, and 0.79 for the three individuals, respectively, with each model discriminating between 5 nonverbal vocalization classes. We also present first results for real-time binary classification of positive- and negative-affect nonverbal vocalizations, trained using a commercial wearable microphone and tested in real time using a smartphone. This work informs personalized machine learning methods for non-traditional communicators and advances real-world interactive augmentative technology for an underserved population. Jaya Narain, Kristina T. Johnson, Craig Ferguson, Amanda O'Brien, Tanya Talkar, Yue Zhang 0014, Peter Wofford, Thomas F. Quatieri, Rosalind W. Picard, Pattie Maes |
ICMI | 8 |
| 2020 | Domain Adaptation for Enhancing Speech-Based Depression Detection in Natural Environmental Conditions Using Dilated CNNsabstractDepression disorders are a major growing concern worldwide, especially given the unmet need for widely deployable depression screening for use in real-world environments. Speech-based depression screening technologies have shown promising results, but primarily in systems that are trained using laboratory-based recorded speech. They do not generalize well on data from more naturalistic settings. This paper addresses the generalizability issue by proposing multiple adaptation strategies that update pre-trained models based on a dilated convolutional neural network (CNN) framework, which improve depression detection performance in both clean and naturalistic environments. Experimental results on two depression corpora show that feature representations in CNN layers need to be adapted to accommodate environmental changes, and that increases in data quantity and quality are helpful for pre-training models for adaptation. The cross-corpus adapted systems produce relative improvements of 29.4% and 17.2% in unweighted average recall over non-adapted systems for both clean and naturalistic corpora, respectively. Zhaocheng Huang, Julien Epps, Dale Joachim, Brian Stasak, James R. Williamson, Thomas F. Quatieri |
INTERSPEECH | 6 |
| 2020 | Extended Study on the Use of Vocal Tract Variables to Quantify Neuromotor Coordination in Depression
Nadee Seneviratne, James R. Williamson, Adam C. Lammert, Thomas F. Quatieri, Carol Y. Espy-Wilson |
INTERSPEECH | 4 |
| 2020 | Detection of Subclinical Mild Traumatic Brain Injury (mTBI) Through Speech and Gait
Tanya Talkar, Sophia Yuditskaya, James R. Williamson, Adam C. Lammert, Hrishikesh Rao 0002, Daniel J. Hannon, Anne T. O'Brien, Gloria Vergara-Diaz, Richard DeLaura, Douglas E. Sturim, Gregory A. Ciccarelli, Ross Zafonte, Jeff Palmer, Paolo Bonato, Thomas F. Quatieri |
INTERSPEECH | 15 |
| 2020 | Generalized Two-Stage Rank Regression Framework for Depression Score Prediction from SpeechabstractThis paper introduces a novel speech-based depression score prediction paradigm, the 2-stage ranking prediction framework, and highlights the benefits it brings to depression prediction. Conventional regression approaches aim to discern a single functional relationship between speech features and depression scores, making an implicit assumption about the existence of a single fixed relationship between the features and scores. However, as the relationship between severity of depression and the clinical score may vary over the range of the assessment scale, this style of analysis may not be suited to depression prediction. The proposed framework on the other hand, imposes a series of partitions on the feature space, with each partition corresponding to a distinct predefined range of depression scores, and predicts the score based on measures of membership to each partition. This approach provides additional flexibility by allowing different rankings to be learnt for different depression scores, and relaxes assumptions made by conventional regression approaches. Results demonstrate the framework's suitability for depression score prediction: different 2-stage implementations, based on heterogeneous feature extraction and modelling approaches, produce state-of-the-art results on the AVEC-2013 dataset. It is also demonstrated that, unlike fusion of conventional regression systems, the fusion of two-stage systems consistently improves prediction performance. Nicholas Cummins, Vidhyasaharan Sethu, Julien Epps, James R. Williamson, Thomas F. Quatieri, Jarek Krajewski |
IEEE Trans. Affect. Comput. | 5 |
| 2019 | Correlating an Ambulatory Voice Measure to Electrodermal Activity in Patients with Vocal HyperfunctionabstractWe investigate the connection between the autonomic nervous system and the voice in patients with vocal hyperfunction and healthy-control groups. We present a methodology and preliminary results of two multi-modal measurement streams that capture this relationship. Subjects were instrumented for daily, ambulatory collection of their voice and wrist-based electrodermal activity. Measures of vocal function (e.g., fundamental frequency) were computed, as well as measures of autonomic function (e.g., skin conductance response). Spearman correlation coefficients were calculated to measure the relationship between vocal and autonomic function over sliding windows throughout each observation day. We found preliminary evidence that patients with a subtype of vocal hyperfunction (non-phonotraumatic vocal hyperfunction) exhibit a coupling between the autonomic nervous system and the vocal system. Understanding how the autonomic nervous system interacts with the voice may provide new insights into the etiology/pathophysiology of vocal hyperfunction and improve prevention, diagnosis and treatment of these disorders. Gregory A. Ciccarelli, Daryush D. Mehta, Andrew Ortiz, Jarrad H. Van Stan, Laura Toles, Katherine L. Marks, Robert E. Hillman, Thomas F. Quatieri |
BSN | 8 |
| 2019 | On-Body Monitoring of Voice-Based Cognitive Load Features in an Auditory Working Memory TaskabstractThe ability to monitor an individual's cognitive load in operational and naturalistic settings is of great importance in improving military performance and readiness. This paper describes preliminary results of using an on-body, multimodal voice monitoring system for assessing cognitive load based on vocal characteristics extracted from noise-robust sensors. The main components of the system include a commercial wired electroglottograph (EGG) and a lightweight, flex circuit that houses two sensors: an acoustic MEMS microphone (MIC) and a non-acoustic neck-surface accelerometer (ACC). We conducted human subject experiments using a cognitive load protocol under quiet laboratory conditions and computed a previously investigated vocal biomarker (creaky voice quality) from the MIC, ACC, and EGG signals. Results demonstrated the potential of discriminating low versus high cognitive load using the creaky voice correlation structure within each of the three sensor domains. Combining MIC, ACC, and EGG hardware into an integrated system would provide complementary information for quantifying voice and speech biomarkers with a high degree of robustness to competing noise sources. Daryush D. Mehta, Rohan Deshpande, Luke Letter, Edward Froehlich, Andrew M. Siegel, Thomas F. Quatieri, Laura J. Brattain |
BSN | 6 |
| 2019 | Assessing Neuromotor Coordination in Depression Using Inverted Vocal Tract Variables
Carol Y. Espy-Wilson, Adam C. Lammert, Nadee Seneviratne, Thomas F. Quatieri |
INTERSPEECH | 4 |
| 2019 | Vocal Biomarker Assessment Following Pediatric Traumatic Brain Injury: A Retrospective Cohort Study
Camille Noufi, Adam C. Lammert, Daryush D. Mehta, James R. Williamson, Gregory A. Ciccarelli, Douglas E. Sturim, Jordan R. Green, Thomas F. Campbell, Thomas F. Quatieri |
INTERSPEECH | 9 |
| 2019 | Tracking depression severity from audio and video based on speech articulatory coordination
James R. Williamson, Diana Young, Andrew A. Nierenberg, James Niemi, Brian S. Helfer, Thomas F. Quatieri |
Comput. Speech Lang. | 6 |
| 2018 | Lightweight, on-body, wireless system for ambulatory voice and ambient noise monitoringabstractIn this paper, we present a lightweight, on-body, wireless system designed for monitoring real-world, ambulatory voice characteristics. The system has the potential to provide important assessments of voice and speech disorders and the impact of environmental sound levels as individuals go about their daily life. The system's transmitter is positioned on the neck and synchronously streams dual-channel sensor data from an on-board MEMS microphone and a high-bandwidth accelerometer, which acts as a noise-robust and confidential contact microphone. These data are recorded to a receiver that can store the data locally and stream a real-time feed to a computer. We also report on the design considerations of this novel system and discuss progress leading up to the latest iteration, especially of the transmitter components on a flexible circuit. Pilot data are shown from an in-field, ambulatory recording during an individual's daily activities that included settings in quiet and with naturalistic ambient noise. Patrick Chwalek, Daryush D. Mehta, Brendon Welsh, Catherine Wooten, Kate Byrd, Edward Froehlich, David Maurer, Joseph Lacirignola, Thomas F. Quatieri, Laura J. Brattain |
BSN | 9 |
| 2018 | Vocal Biomarkers for Cognitive Performance Estimation in a Working Memory Task
Jennifer Sloboda, Adam C. Lammert, James R. Williamson, Christopher J. Smalt, Daryush D. Mehta, C. O. L. Ian Curry, Kristin Heaton, Jeff Palmer, Thomas F. Quatieri |
INTERSPEECH | 9 |
| 2018 | The Effect of Exposure to High Altitude and Heat on Speech Articulatory Coordination
James R. Williamson, Thomas F. Quatieri, Adam C. Lammert, Katherine Mitchell, Katherine Finkelstein, Nicole Ekon, Caitlin Dillon, Robert Kenefick, Kristin Heaton |
INTERSPEECH | 2 |
| 2017 | Noninvasive estimation of cognitive status in mild traumatic brain injury using speech production and facial expressionabstractIn civilian and military populations, there is strong need for objective, noninvasive technologies in monitoring cognitive status associated with mild traumatic brain injury (mTBI). Previous work has shown that monitoring technologies based upon motor control characteristics in speech production provide sensitive indication of cognitive impairments resulting from neurotraumatic injury. Here, this approach is generalized to biomarkers from both speech and facial expression during speaking and preliminary analysis is presented for noninvasively estimating cognitive status associated with mTBI. High-quality audio and video recordings were obtained from 20 subjects in conjunction with cognitive task performance, measured as Processing Speed Index (PSI). Of the 20 subjects, five had a documented history of mild traumatic brain injury, and 15 were control subjects. Models were trained on the control subjects to estimate PSI, and then used to estimate PSI scores from the mTBI cases. Pearson's correlation coefficient between the estimates and the recorded PSI scores revealed r = 0.98 for the speech features and r = 0.92 for the facial features. These results demonstrate the promise of cognitive assessment technologies based on motor timing and coordination underlying vocal and facial expression during speaking in the context of mTBI. Adam C. Lammert, James R. Williamson, Austin R. Hess, Tejash Patel, Thomas F. Quatieri, HuiJun Liao, Alexander P. Lin, Kristin Heaton |
ACII | 5 |
| 2017 | Vocal markers of motor, cognitive, and depressive symptoms in Parkinson's diseaseabstractPatients with Parkinson's disease (PD) often suffer from cognitive impairment and depression in addition to motor dysfunction. These non-motor symptoms may be challenging to diagnose and disentangle from the effects of motor impairment. Analysis of vocal acoustics may improve detection and differentiation of motor, cognitive, and depressive symptom domains simultaneously. Certain vocal markers may be distinctly correlated with specific symptom domains, while other vocal markers may overlap across domains. In this paper, a joint multi-domain characterization of PD symptoms is presented. Speech recordings from 35 PD patients were analyzed for speech markers characterizing articulatory coordination based on resonant (formant) frequencies and delta-mel cepstral coefficients (dMFCC), as well as phonemic timing based on phoneme-dependent speaking rates. Moderate correlations were found between vocal markers and the motor and cognitive symptoms of PD, and weaker correlations with depressive symptoms. Notable differences were identified in the correlation patterns for each symptom domain. Of particular interest, the durations of certain phonemes were correlated with cognitive compared with motor symptom severity. Statistical models, developed based on the vocal markers, achieved moderate accuracy in predicting motor severity (r=0.42) and global cognition (r=0.52) but not depression (r=-0.21). This work suggests it may be possible to distinguish the impact of non-motor PD symptoms on speech. Future study is warranted to further develop symptom-specific vocal marker models in PD. Kara M. Smith, James R. Williamson, Thomas F. Quatieri |
ACII | 3 |
| 2017 | Canonical Correlation Analysis and Prediction of Perceived Rhythmic Prominences and Pitch Tones in Speech
Elizabeth Godoy, James R. Williamson, Thomas F. Quatieri |
INTERSPEECH | 3 |
| 2017 | Wireless Neck-Surface Accelerometer and Microphone on Flex Circuit with Application to Noise-Robust Monitoring of Lombard Speech
Daryush D. Mehta, Patrick Chwalek, Thomas F. Quatieri, Laura J. Brattain |
INTERSPEECH | 3 |
| 2017 | Evaluation of Glottal Inverse Filtering Algorithms Using a Physiologically Based Articulatory Speech SynthesizerabstractGlottal inverse filtering aims to estimate the glottal airflow signal from a speech signal for applications such as speaker recognition and clinical voice assessment. Nonetheless, evaluation of inverse filtering algorithms has been challenging due to the practical difficulties of directly measuring glottal airflow. Apart from this, it is acknowledged that the performance of many methods degrade in voice conditions that are of great interest, such as breathiness, high pitch, soft voice, and running speech. This paper presents a comprehensive, objective, and comparative evaluation of state-of-the-art inverse filtering algorithms that takes advantage of speech and glottal airflow signals generated by a physiological speech synthesizer. The synthesizer provides a physics-based simulation of the voice production process and thus an adequate test bed for revealing the temporal and spectral performance characteristics of each algorithm. Included in the synthetic data are continuous speech utterances and sustained vowels, which are produced with multiple voice qualities (pressed, slightly pressed, modal, slightly breathy, and breathy), fundamental frequencies, and subglottal pressures to simulate the natural variations in real speech. In evaluating the accuracy of a glottal flow estimate, multiple error measures are used, including an error in the estimated signal that measures overall waveform deviation, as well as an error in each of several clinically relevant features extracted from the glottal flow estimate. Waveform errors calculated from glottal flow estimation experiments exhibited mean values around 30% for sustained vowels, and around 40% for continuous speech, of the amplitude of true glottal flow derivative. Closed-phase approaches showed remarkable stability across different voice qualities and subglottal pressures. The algorithms of choice, as suggested by significance tests, are closed-phase covariance analysis for the analysis of sustained vowels, and sparse linear prediction for the analysis of continuous speech. Results of data subset analysis suggest that analysis of close rounded vowels is an additional challenge in glottal flow estimation. Yu-Ren Chien, Daryush D. Mehta, Jón Guðnason, Matías Zanartu, Thomas F. Quatieri |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2016 | A multimodal sensor system for automated marmoset behavioral analysisabstractThe common marmoset is emerging as an important transgenic model for improving the understanding of the underlying neurological basis of many brain disorders. Automated systems for quantitative monitoring of marmoset behaviors in naturalist settings over long period of time are needed to facilitate this process. This paper presents the preliminary work toward building a novel multimodal acquisition system for the automated marmoset behavior analysis in home cage. In addition to integrating commercial available devices such as Microsoft Kinect sensors and microphones of different characteristics, we also developed a wireless flexible neck collar with acoustic and non-acoustic sensors onboard for marmoset vocalization recording and caller identification. Our initial effort has been focused on the real-time synchronization of multiple sensor outputs, the engineering design of the wireless collar, and algorithms for global 3D position and local head movement from a Microsoft Kinect sensor. With limited preliminary data, we are able to estimate 3D trajectories of two marmosets with a RMSE of ~3.2 mm and track colored ear tufts with an accuracy of RMSE ~1.8 mm. A larger dataset is needed for a complete assessment and validation. Our system architecture is modular and flexible, and can be extended to include more sensors and devices if needed. Laura J. Brattain, Rogier Landman, Kerry A. Johnson, Patrick Chwalek, Julia Hyman, Jitendra Sharma, Charles Jennings, Robert Desimone, Guoping Feng, Thomas F. Quatieri |
BSN | 10 |
| 2016 | A vocal modulation model with application to predicting depression severityabstractSpeech provides a potential simple and noninvasive “on-body” means to identify and monitor neurological diseases. Here we develop a model for a class of vocal biomarkers exploiting modulations in speech, focusing on Major Depressive Disorder (MDD) as an application area. Two model components contribute to the envelope of the speech waveform: amplitude modulation (AM) from respiratory muscles, and AM from interaction between vocal tract resonances (formants) and frequency modulation in vocal fold harmonics. Based on the model framework, we test three methods to extract envelopes capturing these modulations of the third formant for synthesized sustained vowels. Using subsequent modulation features derived from the model, we predict MDD severity scores with a Gaussian Mixture Model. Performing global optimization over classifier parameters and number of principal components, we evaluate performance of the features by examining the root-mean-squared error (RMSE), mean absolute error (MAE), and Spearman correlation between the actual and predicted MDD scores. We achieved RMSE and MAE values 10.32 and 8.46, respectively (Spearman correlation=0.487, p<;0.001), relative to a baseline RMSE of 11.86 and MAE of 10.05, obtained by predicting the mean MDD severity score. Ultimately, our model provides a framework for detecting and monitoring vocal modulations that could also be applied to other neurological diseases. Rachelle L. Horwitz-Martin, Thomas F. Quatieri, Elizabeth Godoy, James R. Williamson |
BSN | 2 |
| 2016 | Using collar-worn sensors to forecast thermal strain in military working dogsabstractMilitary working dogs (MWDs) are at high risk of heat strain both during training and missions. Body heat in a MWD increases due to work, and the primary means for reducing this heat are resting and panting. Body-worn sensors can enable monitoring of work level and respiratory rate in real time. They can thereby provide real-time objective indicators of thermal strain in MWDs. In this paper a system is proposed for using collar-worn accelerometer, global positioning system (GPS), and audio recorder sensors to provide real-time estimates of work level and respiration (breathing and panting) rate. Automated methods are demonstrated for using a collar-worn accelerometer and GPS sensor to estimate work levels during multiple short-duration activities, and for estimating respiration rates from a collar-worn audio recorder. The potential utility of these estimates for forecasting and monitoring thermal strain is assessed based on performance in out of sample prediction of core temperature (Tc) statistics, which are obtained from ingestible sensors. Using cross-validation, regression models are trained from accelerometer- and GPS-based activity estimates to predict rate of change in Tc, obtaining a correlation of r=0.59 between actual and predicted Tc change rates. Regression models are also trained from audio-based respiration rate estimates during recovery to predict the Tc values immediately prior to recovery, obtaining a correlation of r=0.49 between actual and predicted Tc. James R. Williamson, Austin R. Hess, Christopher J. Smalt, Delsey M. Sherrill, Thomas F. Quatieri, Catherine O'Brien |
BSN | 5 |
| 2016 | Neurophysiological Vocal Source Modeling for Biomarkers of DiseaseabstractSpeech is potentially a rich source of biomarkers for detecting and monitoring neuropsychological disorders. Current biomarkers typically comprise acoustic descriptors extracted from behavioral measures of source, filter, prosodic and linguistic cues. In contrast, in this paper, we extract vocal features based on a neurocomputational model of speech production, reflecting latent or internal motor control parameters that may be more sensitive to individual variation under neuropsychological disease. These features, which are constrained by neurophysiology, may be resilient to artifacts and provide an articulatory complement to acoustic features. Our features represent a mapping from a low-dimensional acoustics-based feature space to a high-dimensional space that captures the underlying neural process including articulatory commands and auditory and somatosensory feedback errors. In particular, we demonstrate a neurophysiological vocal source model that generates biomarkers of disease by modeling vocal source control. By using the fundamental frequency contour and a biophysical representation of the vocal source, we infer two neuromuscular time series whose coordination provides vocal features that are applied to depression and Parkinson’s disease as examples. These vocal source coordination features alone, on a single held vowel, outperform or are comparable to other features sets and reflect a significant compression of the feature space. Gregory A. Ciccarelli, Thomas F. Quatieri, Satrajit S. Ghosh |
INTERSPEECH | 2 |
| 2016 | Relating Estimated Cyclic Spectral Peak Frequency to Measured Epilarynx Length Using Magnetic Resonance Imaging
Elizabeth Godoy, Andrew Dumas, Jennifer Melot, Nicolas Malyska, Thomas F. Quatieri |
INTERSPEECH | 5 |
| 2016 | Relation of Automatically Extracted Formant Trajectories with Intelligibility Loss and Speaking Rate Decline in Amyotrophic Lateral Sclerosis
Rachelle L. Horwitz-Martin, Thomas F. Quatieri, Adam C. Lammert, James R. Williamson, Yana Yunusova, Elizabeth Godoy, Daryush D. Mehta, Jordan R. Green |
INTERSPEECH | 2 |
| 2016 | Investigation of Speed-Accuracy Tradeoffs in Speech Production Using Real-Time Magnetic Resonance Imaging
Adam C. Lammert, Christine H. Shadle, Shri Narayanan, Thomas F. Quatieri |
INTERSPEECH | 4 |
| 2016 | A Framework for Automated Marmoset Vocalization Detection and Classification
Alan Wisler, Laura J. Brattain, Rogier Landman, Thomas F. Quatieri |
INTERSPEECH | 4 |
| 2015 | Evaluation of speech inverse filtering techniques using a physiologically based synthesizerabstractGlottal inverse filtering methods are designed to derive a glottal flow waveform from a speech signal. In this paper, we evaluate and compare such methods using a speech synthesizer that simulates voice production in a physiologically-based manner that includes complexities such as nonlinear source-tract coupling. Five inverse filtering techniques are evaluated on 90 synthesized speech waveforms generated by setting six vowel configurations, three glottal models, and five fundamental frequencies. Using normalized mean square error as the primary performance metric of the estimated glottal flow derivative, results show that the accuracy of all methods depends on the configuration of the vocal tract, glottis and the fundamental frequency. Averaged over these conditions, the closed phase covariance and one weighted covariance algorithm yield lower error rates (0.41 ± 0.2) than iterative and adaptive inverse filtering (0.49 ± 0.1) and complex cepstrum decomposition (0.76 ± 0.1). Jón Guðnason, Daryush D. Mehta, Thomas F. Quatieri |
ICASSP | 3 |
| 2015 | Estimating lower vocal tract features with closed-open phase spectral analysesabstractPrevious studies have shown that, in addition to being speaker-dependent yet context-independent, lower vocal tract acoustics significantly impact the speech spectrum at mid-tohigh frequencies (e.g 3-6kHz). The present work automatically estimates spectral features that exhibit acoustic properties of the lower vocal tract. Specifically aiming to capture the cyclicity property of the epilarynx tube, a novel multi-resolution approach to spectral analyses is presented that exploits significant differences between the closed and open phases of a glottal cycle. A prominent null linked to the piriform fossa is also estimated. Examples of the feature estimation on natural speech of the VOICES multi-speaker corpus illustrate that a salient spectral pattern indeed emerges between 3-6kHz across all speakers. Moreover, the observed pattern is consistent with that canonically shown for the lower vocal tract in previous works. Additionally, an instance of a speaker’s formant (i.e. spectral peak around 3kHz that has been well-established as a characteristic of voice projection) is quantified here for the VOICES template speaker in relation to epilarynx acoustics. The corresponding peak is shown to be double the power on average compared to the other speakers (20 vs 10 dB). Elizabeth Godoy, Nicolas Malyska, Thomas F. Quatieri |
INTERSPEECH | 3 |
| 2015 | Vocal biomarkers to discriminate cognitive load in a working memory task
Thomas F. Quatieri, James R. Williamson, Christopher J. Smalt, Tejash Patel, Joseph Perricone, Daryush D. Mehta, Brian S. Helfer, Gregory A. Ciccarelli, Darrell O. Ricke, Nicolas Malyska, Jeff Palmer, Kristin Heaton, Marianna Eddy, Joseph Moran |
INTERSPEECH | 1 |
| 2015 | Segment-dependent dynamics in predicting parkinson's disease
James R. Williamson, Thomas F. Quatieri, Brian S. Helfer, Joseph Perricone, Satrajit S. Ghosh, Gregory A. Ciccarelli, Daryush D. Mehta |
INTERSPEECH | 2 |
| 2015 | Cognitive impairment prediction in the elderly based on vocal biomarkers
Bea Yu, Thomas F. Quatieri, James R. Williamson, James C. Mundt |
INTERSPEECH | 2 |
| 2015 | A review of depression and suicide risk assessment using speech analysis
Nicholas Cummins, Stefan Scherer, Jarek Krajewski, Sebastian Schnieder, Julien Epps, Thomas F. Quatieri |
Speech Commun. | 6 |
| 2014 | Closed phase estimation for inverse filtering the oral airflow waveformabstractGlottal closed phase estimation during speech production is critical to inverse filtering and, although addressed for radiated acoustic pressure analysis, must be better understood for the analysis of the oral airflow volume velocity signal that provides important properties of healthy and disordered voices. This paper compares the estimation of the closed phase from the acoustic speech signal and the oral airflow waveform recorded using a pneumotachograph mask. Results are presented for ten adult speakers with normal voices who sustained a set of vowels at a comfortable pitch and loudness. With electroglottography as reference, the identification rate and accuracy of glottal closure instants for the oral airflow are 96.8 % and 0.28 ms, whereas these metrics are 99.4 % and 0.10 ms for the acoustic signal. We conclude that glottal closure detection is adequate for close phase inverse filtering but that improvements to detection of glottal opening instants on the oral airflow signal are warranted. Jón Guðnason, Daryush D. Mehta, Thomas F. Quatieri |
ICASSP | 3 |
| 2014 | Articulatory dynamics and coordination in classifying cognitive change with preclinical mTBI
Brian S. Helfer, Thomas F. Quatieri, James R. Williamson, Laurel Keyes, Benjamin Evans, W. Nicholas Greene, Trina Vian, Joseph Lacirignola, Trey E. Shenk, Thomas M. Talavage, Jeff Palmer, Kristin Heaton |
INTERSPEECH | 2 |
| 2014 | Prediction of cognitive performance in an animal fluency task based on rate and articulatory markers
Bea Yu, Thomas F. Quatieri, James R. Williamson, James C. Mundt |
INTERSPEECH | 2 |
| 2013 | On the relative importance of vocal source, system, and prosody in human depressionabstractIn Major Depressive Disorder (MDD), neurophysiologic changes can alter motor control [1][2] and therefore alter speech production by influencing vocal fold motion (source), the vocal tract (system), and melody (prosody). In this paper, we use a database of voice recordings from 28 depressed subjects treated over a 6-week period [3] to compare correlations between features from each of the three speech-production components and clinical assessments of MDD. Toward biomarkers for audio-based continuous monitoring of depression severity, we explore the contextual dependence of these correlations with free-response and read speech, and show tradeoffs across categories of features in these two example contexts. Likewise, we also investigate the context-and speech component-dependence of correlations between our vocal features and assessment of individual symptoms of MDD (e.g., depressed mood, agitation, energy). Finally, motivated by our initial findings, we describe how context may be useful in “on-body” monitoring of MDD to facilitate identification of depression and evaluation of its treatment. Rachelle L. Horwitz-Martin, Thomas F. Quatieri, Brian S. Helfer, Bea Yu, James R. Williamson, James C. Mundt |
BSN | 2 |
| 2013 | Classification of depression state based on articulatory precisionabstractNeurophysiological changes in the brain associated with major depression disorder can disrupt articulatory precision in speech production. Motivated by this observation, we address the hypothesis that articulatory features, as manifested through formant frequency tracks, can help in automatically classifying depression state. Specifically, we investigate the relative importance of vocal tract formant frequencies and their dynamic features from sustained vowels and conversational speech. Using a database consisting of audio from 35 subjects with clinical measures of depression severity, we explore the performance of Gaussian mixture model (GMM) and support vector machine (SVM) classifiers. With only formant frequencies and their dynamics given by velocity and acceleration, we show that depression state can be classified with an optimal sensitivity/specificity/area under the ROC curve of 0.86/0.64/0.70 and 0.77/0.77/0.73 for GMMs and SVMs, respectively. Future work will involve merging our formant-based characterization with vocal source and prosodic features. Index Terms: major depressive disorder, motor coordination, articulatory control, vocal biomarkers, formant frequencies Brian S. Helfer, Thomas F. Quatieri, James R. Williamson, Daryush D. Mehta, Rachelle L. Horwitz-Martin, Bea Yu |
INTERSPEECH | 2 |
| 2012 | Speech Enhancement Using Sparse Convolutive Non-negative Matrix Factorization with Basis AdaptationabstractWe introduce a framework for speech enhancement based on convolutive non-negative matrix factorization that leverages available speech data to enhance arbitrary noisy utterances with no a priori knowledge of the speakers or noise types present. Previous approaches have shown the utility of a sparse recon-struction of the speech-only components of an observed noisy utterance. We demonstrate that an underlying speech represen-tation which, in addition to applying sparsity, also adapts to the noisy acoustics improves overall enhancement quality. The proposed system performs comparably to a traditional Wiener filtering approach, and the results suggest that the proposed framework is most useful in moderate- to low-SNR scenarios. Index Terms: speech enhancement, convolutive non-negative matrix factorization, basis adaptation, sparsity Michael Carlin, Nicolas Malyska, Thomas F. Quatieri |
INTERSPEECH | 3 |
| 2012 | Vocal-Source Biomarkers for Depression: A Link to Psychomotor Activityabstract1 A hypothesis in characterizing human depression is that change in the brain‟s basal ganglia results in a decline of motor coordination [6][8][14]. Such a neuro-physiological change may therefore affect laryngeal control and dynamics. Under this hypothesis, toward the goal of objective monitoring of depression severity, we investigate vocal-source biomarkers for depression; specifically, source features that may relate to precision in motor control, including vocal-fold shimmer and jitter, degree of aspiration, fundamental frequency dynamics, and frequency-dependence of variability and velocity of energy. We use a 35-subject database collected by Mundt et al. [1] in which subjects were treated over a six-week period, and investigate correlation of our features with clinical (HAMD), as well as self-reported (QIDS) Total subject assessment scores. To explicitly address the motor aspect of depression, we compute correlations with the Psychomotor Retardation component of clinical and self-reported Total assessments. For our longitudinal database, most correlations point to statistical relationships of our vocal-source biomarkers with psychomotor activity, as well as with depression severity. Thomas F. Quatieri, Nicolas Malyska |
INTERSPEECH | 1 |
| 2012 | Two-Dimensional Speech-Signal ModelingabstractTraditional approaches in speech-signal processing analyze short-time frames of the signal (e.g., the short-time Fourier transform). Findings from auditory neurophysiology coupled with image processing principles, however, have motivated an alternative 2-D processing framework in which 2-D analysis is performed on the time-frequency distribution itself. This paper develops a 2-D model of speech in local time-frequency regions of narrowband spectrograms using sinusoidal-series-based modulation. Our model is shown to distribute vocal tract and onset/offset content based on source information (e.g., noise and voicing) in a transformed 2-D space, thereby explicitly representing different classes of energy modulations commonly observed in spectrograms. We demonstrate the model's ability to represent speech sounds by developing and evaluating algorithms for analysis/synthesis of spectrograms. As an example application, we demonstrate the utility of the model for co-channel speaker separation using prior pitch information of two overlapping speakers. Finally, our separation scheme based on 2-D modeling is compared against a reference (frame-based) sinusoidal separation system using both prior and estimated pitch. Tianyu T. Wang, Thomas F. Quatieri |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | A time-warping framework for speech turbulence-noise component estimation during aperiodic phonationabstractThe accurate estimation of turbulence noise affects many areas of speech processing including separate modification of the noise component, analysis of degree of speech aspiration for treating pathological voice, the automatic labeling of speech voicing, as well as speaker characterization and recognition. Previous work in the literature has provided methods by which such a high-quality noise component may be estimated in near-periodic speech, but it is known that these methods tend to leak aperiodic phonation (with even slight deviations from periodicity) into the noise-component estimate. In this paper, we improve upon existing algorithms in conditions of aperiodicity by introducing a time-warping based approach to speech noise-component estimation, demonstrating the results on both natural and synthetic speech examples. Nicolas Malyska, Thomas F. Quatieri |
ICASSP | 2 |
| 2011 | Sinewave Representations of NonmodalityabstractRegions of nonmodal phonation, exhibiting deviations from uniform glottal-pulse periods and amplitudes, occur often and convey information about speaker- and linguistic-dependent factors. Such waveforms pose challenges for speech modeling, analysis/synthesis, and processing. In this paper, we investigate the representation of nonmodal pulse trains as a sum of harmonically-related sinewaves with time-varying amplitudes, phases, and frequencies. We show that a sinewave representation of any impulsive signal is not unique and also the converse, i.e., frame-based measurements of the underlying sinewave representation can yield different impulse trains. Finally, we argue how this ambiguity may explain addition, deletion, and movement of pulses in sinewave synthesis and a specific illustrative example of time-scale modification of a nonmodal case of diplophonia. Index Terms: sinewave analysis, sinewave synthesis, nonmodal phonation Nicolas Malyska, Thomas F. Quatieri, Robert B. Dunn |
INTERSPEECH | 2 |
| 2011 | Automatic Detection of Depression in Speech Using Gaussian Mixture Modeling with Factor AnalysisabstractAbstract 1 Of increasing importance in the civilian and military population is the recognition of Major Depressive Disorder at its earliest stages and intervention before the onset of severe symptoms. Toward the goal of more effective monitoring of depression severity, we investigate automatic classifiers of depression state, that have the important property of mitigating nuisances due to data variability, such as speaker and channel effects, unrelated to levels of depression. To assess our measures, we use a 35-speaker free-response speech database of subjects treated for depression over a six-week duration, along with standard clinical HAMD depression ratings. Preliminary experiments indicate that by mitigating nuisances, thus focusing on depression severity as a class, we can significantly improve classification accuracy over baseline Gaussian-mixture-model-based classifiers. Index Terms: major depressive disorder, Gaussian-mixture models, joint factor analysis, nuisance mitigation Douglas E. Sturim, Pedro A. Torres-Carrasquillo, Thomas F. Quatieri, Nicolas Malyska, Alan McCree |
INTERSPEECH | 3 |
| 2011 | Time-Varying Autoregressions in Speech: Detection Theory and ApplicationsabstractThis paper develops a general detection theory for speech analysis based on time-varying autoregressive models, which themselves generalize the classical linear predictive speech analysis framework. This theory leads to a computationally efficient decision-theoretic procedure that may be applied to detect the presence of vocal tract variation in speech waveform data. A corresponding generalized likelihood ratio test is derived and studied both empirically for short data records, using formant-like synthetic examples, and asymptotically, leading to constant false alarm rate hypothesis tests for changes in vocal tract configuration. Two in-depth case studies then serve to illustrate the practical efficacy of this procedure across different time scales of speech dynamics: first, the detection of formant changes on the scale of tens of milliseconds of data, and second, the identification of glottal opening and closing instants on time scales below ten milliseconds. Daniel Rudoy, Thomas F. Quatieri, Patrick J. Wolfe |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2010 | Preserving the character of perturbations in scaled pitch contoursabstractThe global and fine dynamic components of a pitch contour in voice production, as in the speaking and singing voice, are important for both the meaning and character of an utterance. In speech, for example, slow pitch inflections, rapid pitch accents, and irregular regions all comprise the pitch contour. In applications where all components of a pitch contour are stretched or compressed in the same way, as for example in time-scale modification, an unnatural scaled contour may result. In this paper, we develop a framework for scaling pitch contours, motivated by the goal of maintaining naturalness in time-scale modification of voice. Specifically, we develop a multi-band algorithm to independently modify the slow trajectory and fast perturbation components of a contour for a more natural synthesis, and we present examples where pitch contours representative of speaking and singing voice are lengthened. In the speaking voice, the frequency content of flutter or irregularity is maintained, while slow pitch inflection is simply stretched or compressed. In the singing voice, rapid vibrato is preserved while slower note-to-note variation is scaled as desired. Thomas A. Baran, Nicolas Malyska, Thomas F. Quatieri |
ICASSP | 3 |
| 2010 | Multi-pitch estimation by a joint 2-d representation of pitch and pitch dynamicsabstractMulti-pitch estimation of co-channel speech is especially challenging when the underlying pitch tracks are close in pitch value (e.g., when pitch tracks cross). Building on our previous work in [1], we demonstrate the utility of a two-dimensional (2-D) analysis method of speech for this problem by exploiting its joint representation of pitch and pitch-derivative information from distinct speakers. Specifically, we propose a novel multi-pitch estimation method consisting of 1) a datadriven classifier for pitch candidate selection, 2) local pitch and pitch-derivative estimation by k-means clustering, and 3) a Kalman filtering mechanism for pitch tracking and assignment. We evaluate our method on a database of allvoiced speech mixtures and illustrate its capability to estimate pitch tracks in cases where pitch tracks are separate and when they are close in pitch value (e.g., at crossings). Tianyu T. Wang, Thomas F. Quatieri |
INTERSPEECH | 2 |
| 2010 | High-Pitch Formant Estimation by Exploiting Temporal Change of PitchabstractThis paper considers the problem of obtaining an accurate spectral representation of speech formant structure when the voicing source exhibits a high fundamental frequency. Our work is inspired by auditory perception and physiological studies implicating the use of pitch dynamics in speech by humans. We develop and assess signal processing schemes aimed at exploiting temporal change of pitch to address the high-pitch formant frequency estimation problem. Specifically, we propose a 2-D analysis framework using 2-D transformations of the time-frequency space. In one approach, we project changing spectral harmonics over time to a 1-D function of frequency. In a second approach, we draw upon previous work of Quatieri and Ezzat , , with similarities to the auditory modeling efforts of Chi , where localized 2-D Fourier transforms of the time-frequency space provide improved source-filter separation when pitch is changing. Our methods show quantitative improvements for synthesized vowels with stationary formant structure in comparison to traditional and homomorphic linear prediction. We also demonstrate the feasibility of applying our methods on stationary vowel regions of natural speech spoken by high-pitch females of the TIMIT corpus. Finally, we show improvements afforded by the proposed analysis framework in formant tracking on examples of stationary and time-varying formant structure. Tianyu T. Wang, Thomas F. Quatieri |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Time-varying autoregressive tests for multiscale speech analysisabstractIn this paper we develop hypothesis tests for speech waveform nonstationarity based on time-varying autoregressive models, and demonstrate their efficacy in speech analysis tasks at both segmental and sub-segmental scales. Key to the successful synthesis of these ideas is our employment of a generalized likelihood ratio testing framework tailored to autoregressive coefficient evolutions suitable for speech. After evaluating our framework on speech-like synthetic signals, we present preliminary results for two distinct analysis tasks using speech waveform data. At the segmental level, we develop an adaptive short-time segmentation scheme and evaluate it on whispered speech recordings, while at the sub-segmental level, we address the problem of detecting the glottal flow closed phase. Results show that our hypothesis testing framework can reliably detect changes in the vocal tract parameters across multiple scales, thereby underscoring its broad applicability to speech analysis. Index Terms: TVAR models, hypothesis testing, GLRT, adaptive speech segmentation, glottal flow analysis Daniel Rudoy, Thomas F. Quatieri, Patrick J. Wolfe |
INTERSPEECH | 2 |
| 2009 | 2-d processing of speech for multi-pitch analysisabstractAbstract : This paper introduces a two-dimensional (2-D) processing approach for the analysis of multi-pitch speech sounds. Our framework invokes the short-space 2-D Fourier transform magnitude of a narrowband spectrogram, mapping harmonicallyrelated signal components to multiple concentrated entities in a new 2-D space. First, localized time-frequency regions of the spectrogram are analyzed to extract pitch candidates. These candidates are then combined across multiple regions for obtaining separate pitch estimates of each speech-signal component at a single point in time. We refer to this as multi-region analysis (MRA). By explicitly accounting for pitch dynamics within localized time segments, this separability is distinct from that which can be obtained using short-time autocorrelation methods typically employed in state-of-the-art multi-pitch tracking algorithms. We illustrate the feasibility of MRA for multi-pitch estimation on mixtures of synthetic and real speech. Tianyu T. Wang, Thomas F. Quatieri |
INTERSPEECH | 2 |
| 2008 | Multisensor very lowbit rate speech coding using segment quantizationabstractWe present two approaches to noise robust very low bit rate speech coding using wideband MELP analysis/synthesis. Both methods exploit multiple acoustic and non-acoustic input sensors, using our previously-presented dynamic waveform fusion algorithm to simultaneously perform waveform fusion, noise suppression, and cross-channel noise cancellation. One coder uses a 600 bps scalable phonetic vocoder, with a phonetic speech recognizer followed by joint predictive vector quantization of the error in wideband MELP parameters. The second coder operates at 300 bps with fixed 80 ms segments, using novel variable-rate multistage matrix quantization techniques. Formal test results show that both coders achieve equivalent intelligibility to the 2.4 kbps NATO standard MELPe coder in harsh acoustic noise environments, at much lower bit rates, with only modest quality loss. Alan McCree, Kevin Brady 0001, Thomas F. Quatieri |
ICASSP | 3 |
| 2008 | Adaptive short-time analysis-synthesis for speech enhancementabstractIn this paper we present a new adaptive short-time Fourier analysis-synthesis scheme and demonstrate its efficacy in speech enhancement. While a number of adaptive analyses have previously been proposed to overcome the limitations of fixed-resolution schemes, we propose here a modified overlap-add procedure that enables efficient resynthesis. Our adaptation scheme extends earlier work using local measures of time-frequency concentration, and is applicable to power spectral density estimation for the case of noisy speech. We provide evidence of increased gains in signal-to-noise ratios for synthetic signals as well as empirical evidence of reduced musical noise based on expert listening tests for voiced and phonetically balanced utterances observed in noise, relative to a standard baseline speech enhancement system whose time-frequency resolution is fixed. Daniel Rudoy, Prabahan Basu, Thomas F. Quatieri, Robert B. Dunn, Patrick J. Wolfe |
ICASSP | 3 |
| 2008 | Exploiting temporal change of pitch in formant estimationabstractThis paper considers the problem of obtaining an accurate spectral representation of speech formant structure when the voicing source exhibits a high fundamental frequency. Our work is inspired by auditory perception and physiological modeling studies implicating the use of temporal changes in speech by humans. Specifically, we develop and assess signal processing schemes aimed at exploiting temporal change of pitch as a basis for formant estimation. Our methods are cast in a generalized framework of two-dimensional processing of speech and show quantitative improvements under certain conditions over representations derived from traditional and homomorphic linear prediction. We conclude by highlighting potential benefits of our framework in the particular application of speaker recognition with preliminary results indicating a performance gender-gap closure on subsets of the TIMIT corpus. Tao T. Wang, Thomas F. Quatieri |
ICASSP | 2 |
| 2008 | Spectral Representations of Nonmodal PhonationabstractRegions of nonmodal phonation, which exhibit deviations from uniform glottal-pulse periods and amplitudes, occur often in speech and convey information about linguistic content, speaker identity, and vocal health. Some aspects of these deviations are random, including small perturbations, known as jitter and shimmer, as well as more significant aperiodicities. Other aspects are deterministic, including repeating patterns of fluctuations such as diplophonia and triplophonia. These deviations are often the source of misinterpretation of the spectrum. In this paper, we introduce a general signal-processing framework for interpreting the effects of both stochastic and deterministic aspects of nonmodality on the short-time spectrum. As an example, we show that the spectrum is sensitive to even small perturbations in the timing and amplitudes of glottal pulses. In addition, we illustrate important characteristics that can arise in the spectrum, including apparent shifting of the harmonics and the appearance of multiple pitches. For stochastic perturbations, we arrive at a formulation of the power-spectral density as the sum of a low-pass line spectrum and a high-pass noise floor. Our findings are relevant to a number of speech-processing areas including linear-prediction analysis, sinusoidal analysis-synthesis, spectrally derived features, and the analysis of disordered voices. Nicolas Malyska, Thomas F. Quatieri |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | An Evaluation of Audio-Visual Person Recognition on the XM2VTS Corpus using the Lausanne ProtocolsabstractA multimodal person recognition architecture has been developed for the purpose of improving overall recognition performance and for addressing channel-specific performance shortfalls. This multimodal architecture includes the fusion of a face recognition system with the MIT/LL GMM/UBM speaker recognition architecture. This architecture exploits the complementary and redundant nature of the face and speech modalities. The resulting multimodal architecture has been evaluated on the XM2VTS corpus using the Lausanne open set verification protocols, and demonstrates excellent recognition performance. The multimodal architecture also exhibits strong recognition performance gains over the performance of the individual modalities. Kevin Brady 0001, Michael S. Brandstein, Thomas F. Quatieri, Robert B. Dunn |
ICASSP (4) | 3 |
| 2007 | Multisensor Dynamic Waveform FusionabstractSpeech communication is significantly more difficult in severe acoustic background noise environments, especially when low-rate speech coders are used. Non-acoustic sensors, such as radar sensors, vibrometers, and bone-conduction microphones, offer significant potential in these situations. We extend previous work on fixed waveform fusion from multiple sensors to an optimal dynamic waveform fusion algorithm that minimizes both additive noise and signal distortion in the estimated speech signal. We show that a minimum mean squared error (MMSE) waveform matching criterion results in a generalized multichannel Wiener filter, and that this filter will simultaneously perform waveform fusion, noise suppression, and crosschannel noise cancellation. Formal intelligibility and quality testing demonstrate significant improvement from this approach. Alan McCree, Kevin Brady 0001, Thomas F. Quatieri |
ICASSP (4) | 3 |
| 2007 | Robust Speaker Recognition with Cross-Channel Data: MIT-LL Results on the 2006 NIST SRE Auxiliary Microphone TaskabstractOne particularly difficult challenge for cross-channel speaker verification is the auxiliary microphone task introduced in the 2005 and 2006 NIST Speaker Recognition Evaluations, where training uses telephone speech and verification uses speech from multiple auxiliary microphones. This paper presents two approaches to compensate for the effects of auxiliary microphones on the speech signal. The first compensation method mitigates session effects through Latent Factor Analysis (LFA) and Nuisance Attribute Projection (NAP). The second approach operates directly on the recorded signal with noise reduction techniques. Results are presented that show a reduction in the performance gap between telephone and auxiliary microphone data. Douglas E. Sturim, William M. Campbell, Douglas A. Reynolds, Robert B. Dunn, Thomas F. Quatieri |
ICASSP (4) | 5 |
| 2006 | Analysis of nonmodal phonation using minimum entropy deconvolutionabstractNonmodal phonation occurs when glottal pulses exhibit nonuniform pulse-to-pulse characteristics such as irregular spacings, amplitudes, and/or shapes. The analysis of regions of such nonmodality has application to automatic speech, speaker, language, and dialect recognition. In this paper, we examine the usefulness of a technique called minimum-entropy deconvolution, or MED [1], for the analysis of pulse events in nonmodal speech. Our study presents evidence for both natural and synthetic speech that MED decomposes nonmodal phonation into a series of sharp pulses and a set of mixedphase impulse responses. We show that the estimated impulse responses are quantitatively similar to those in our synthesis model. A hybrid method incorporating aspects of both MED and linear prediction is also introduced. We show preliminary evidence that the hybrid method has benefit over MED alone for composite impulse-response estimation by being more robust to short-time windowing effects as well as a speech aspiration noise component. Index Terms: inverse filtering, nonmodal speech, glottal pulse, minimum entropy Nicolas Malyska, Thomas F. Quatieri |
INTERSPEECH | 2 |
| 2006 | Pitch-scale modification using the modulated aspiration noise sourceabstractSpectral harmonic/noise component analysis of spoken vowels shows evidence of noise modulations with peaks in the estimated noise source component synchronous with both the open phase of the periodic source and with time instants of glottal closure. Inspired by this observation of natural modulations and of fullband energy in the aspiration noise source, we develop an alternate approach to high-quality pitchscale modification of continuous speech. Our strategy takes a dual processing approach, in which the harmonic and noise components of the speech signal are separately analyzed, modified, and re-synthesized. The periodic component is modified using standard modification techniques, and the noise component is handled by modifying characteristics of its source waveform. Since we have modeled an inherent coupling between the periodic and aspiration noise sources, the modification algorithm is designed to preserve the synchrony between temporal modulations of the two sources. The reconstructed modified signal is perceived in informal listening to be natural-sounding and typically reduces artifacts that occur in standard modification techniques. Index Terms: pitch modification, aspiration noise, modulated noise, breathiness, voice quality Daryush D. Mehta, Thomas F. Quatieri |
INTERSPEECH | 2 |
| 2006 | Missing feature theory with soft spectral subtraction for speaker verificationabstractThis paper considers the problem of training/testing mismatch in the context of speaker verification and, in particular, explores the application of missing feature theory in the case of additive white Gaussian noise corruption in testing. Missing feature theory allows for corrupted features to be removed from scoring, the initial step of which is the detection of these features. One method of detection, employing spectral subtraction, is studied in a controlled manner and it is shown that with missing feature compensation the resulting verification performance is improved as long as a minimum number of features remain. Finally, a blending of “soft ” spectral subtraction for noise mitigation and missing feature compensation is presented. The resulting performance improves on the constituent techniques alone, reducing the equal error rate by about 15 % over an SNR range of5-25dB. Index Terms: speaker verification, GMM, spectral subtraction, missing features. Michael T. Padilla, Thomas F. Quatieri, Douglas A. Reynolds |
INTERSPEECH | 2 |
| 2006 | Exploiting nonacoustic sensors for speech encodingabstractThe intelligibility of speech transmitted through low-rate coders is severely degraded when high levels of acoustic noise are present in the acoustic environment. Recent advances in nonacoustic sensors, including microwave radar, skin vibration, and bone conduction sensors, provide the exciting possibility of both glottal excitation and, more generally, vocal tract measurements that are relatively immune to acoustic disturbances and can supplement the acoustic speech waveform. We are currently investigating methods of combining the output of these sensors for use in low-rate encoding according to their capability in representing specific speech characteristics in different frequency bands. Nonacoustic sensors have the ability to reveal certain speech attributes lost in the noisy acoustic signal; for example, low-energy consonant voice bars, nasality, and glottalized excitation. By fusing nonacoustic low-frequency and pitch content with acoustic-microphone content, we have achieved significant intelligibility performance gains using the DRT across a variety of environments over the government standard 2400-bps MELPe coder. By fusing quantized high-band 4-to-8-kHz speech, requiring only an additional 116 bps, we obtain further DRT performance gains by exploiting the ear's insensitivity to fine spectral detail in this frequency region. Thomas F. Quatieri, Kevin Brady 0001, D. Messing, Joseph P. Campbell, William M. Campbell, Michael S. Brandstein, Clifford J. Weinstein, John D. Tardelli, Paul D. Gatewood |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | Automatic Dysphonia Recognition using Biologically-Inspired Amplitude-Modulation FeaturesabstractA dysphonia, or disorder of the mechanisms of phonation in the larynx, can create time-varying amplitude fluctuations in the voice. A model for band-dependent analysis of this amplitude modulation (AM) phenomenon in dysphonic speech is developed from a traditional communications engineering perspective. This perspective challenges current dysphonia analysis methods that analyze AM in the time-domain signal. An automatic dysphonia recognition system is designed to exploit AM in voice using a biologically inspired model of the inferior colliculus. This system, built upon a Gaussian-mixture-model (GMM) classification backend, recognizes the presence of dysphonia in the voice signal. Recognition experiments using data obtained from the Kay elemetrics voice disorders database suggest that the system provides complementary information to state-of-the-art mel-cepstral features. We present dysphonia recognition as an approach to developing features that capture glottal source differences in normal speech. Nicolas Malyska, Thomas F. Quatieri, Douglas E. Sturim |
ICASSP (1) | 2 |
| 2004 | Multisensor MELPe using parameter substitutionabstractThe estimation of speech parameters and the intelligibility of speech transmitted through low-rate coders, such as MELP (mixed excitation linear prediction), are severely degraded when there are high levels of acoustic noise in the speaking environment. The application of nonacoustic and nontraditional sensors, which are less sensitive to acoustic noise than the standard microphone, is being investigated as a means to address this problem. Sensors being investigated include the general electromagnetic motion sensor (GEMS) and the physiological microphone (P-mic). As an initial effort in this direction, a multisensor MELPe coder (MELP coder with the addition of a noise preprocessor) using parameter substitution has been developed, where pitch and voicing parameters are obtained from GEMS and P-Mic sensors, respectively, and the remaining parameters are obtained as usual from a standard acoustic microphone. This parameter substitution technique is shown to produce significant and promising DRT (diagnostic rhyme test) intelligibility improvements over the standard 2400 bps MELPe coder in several high-noise military environments. Further work is in progress aimed at utilizing the nontraditional sensors for additional intelligibility improvements and for more effective lower-rate coding in noise. Kevin Brady 0001, Thomas F. Quatieri, Joseph P. Campbell, William M. Campbell, Michael S. Brandstein, Clifford J. Weinstein |
ICASSP (1) | 2 |
| 2004 | Automated lip-reading for improved speech intelligibilityabstractVarious psycho-acoustical experiments have concluded that visual features strongly affect the perception of speech. This contribution is most pronounced in noisy environments where the intelligibility of audio-only speech is quickly degraded. The paper explores the effectiveness of using extracted visual features, such as lip height and width, for improving speech intelligibility in noisy environments. The intelligibility content of these extracted visual features is investigated through an intelligibility test on an animated rendition of the video generated from the extracted visual features, as well as on the original video. These experiments demonstrate that the extracted video features do contain important aspects of intelligibility that may be utilized in augmenting speech enhancement and coding applications. Alternatively, these extracted visual features can be transmitted in a bandwidth effective way to augment speech coders. Matthew McClain, Kevin Brady 0001, Michael S. Brandstein, Thomas F. Quatieri |
ICASSP (1) | 4 |
| 2004 | A comparison of soft and hard spectral subtraction for speaker verificationabstractAn important concern in speaker recognition is the performance degradation that occurs when speaker models trained with speech from one type of channel are subsequently used to score speech from another type of channel, known as channel mismatch. This paper investigates the relative performance of two different spectral subtraction methods for additive noise compensation in the context of speaker verification. The first method, termed “soft ” spectral subtraction, is performed in the spectral domain on the |DF T | 2 values of the speech frames while the second method, termed “hard ” spectral subtraction, is performed on the Mel-filter energy features. It is shown through both an analytical argument as well as a simulation that soft spectral subtraction results in a higher signal-to-noise ratio in the resulting Mel-filter energy features. In the context of Gaussian mixture model-based speaker verification with additive noise in testing utterances, this is shown to result in an equal error rate improvement over a system without spectral subtraction of approximately 7 % in absolute terms, 21 % in relative terms, over an additive white Gaussian noise range of 5-25 dB. 1. Michael T. Padilla, Thomas F. Quatieri |
INTERSPEECH | 2 |
| 2002 | Speech enhancement based on auditory spectral changeabstractIn this paper, an adaptive approach to the enhancement of speech signals is developed based On auditory spectral change. The algorithm is motivated by sensitivity of aural, biologic systems to signal dynamics, by evidence that noise is aurally masked by rapid changes in a signal, and by analogies to these two aural phenomena in biologic visual processing. Emphasis is on preserving nonstationarity, i.e., speech transient and time-varying components, such as plosive bursts, formant transitions, and vowel onsets, while suppressing additive noise. The essence of the enhancement technique is a Wiener :filter that uses a desired signal spectrum whose estimation adapts to stationarity of the measured signal. The degree of stationarity is derived from a signal change measurement, based on an auditory spectrum that accentuates change in spectral bands, The adaptive filter is applied in an unconventional overlap-add analysis/synthesis framework, using a very short 4-ms analysis Window and a 1-ms frame interval. In informal listening the reconstructions are judged to be “crisp” corresponding to good temporal resolution of transient and rapidly-moving speech events. Thomas F. Quatieri, Robert B. Dunn |
ICASSP | 1 |
| 2002 | Speaker verification using text-constrained Gaussian Mixture ModelsabstractIn this paper we present an approach to close the gap between text-dependent and text-independent speaker verification performance. Text-constrained GMM-UBM systems are created using word segmentations produced by a LVCSR system on conversational speech allowing the system to focus on speaker differences over a constrained set of acoustic units. Results on the 2001 NIST extended data task show this approach can be used to produce an equal error rate of < 1 %. Douglas E. Sturim, Douglas A. Reynolds, Robert B. Dunn, Thomas F. Quatieri |
ICASSP | 4 |
| 2002 | 2-d processing of speech with application to pitch estimationabstractIn this paper, we introduce a new approach to two-dimensional (2-D) processing of the one-dimensional (1-D) speech signal in the time-frequency plane. Specifically, we obtain the shortspace 2-D Fourier transform magnitude of a narrowband spectrogram of the signal and show that this 2-D transformation maps harmonically-related signal components to a concentrated entity in the new 2-D plane. We refer to this series of operations as the “grating compression transform” (GCT), consistent with sine-wave grating patterns in the spectrogram reduced to smeared impulses. The GCT forms the basis of a speech pitch estimator that uses the radial distance to the largest peak in the GCT plane. Using an average magnitude difference between pitch-contour estimates, the GCT-based pitch estimator is shown to compare favorably to a sine-wave-based pitch estimator for all-voiced speech in additive white noise. An extension to a basis for two-speaker pitch estimation is also proposed. Thomas F. Quatieri |
INTERSPEECH | 1 |
| 2000 | Speaker recognition using G.729 speech codec parametersabstractExperiments in Gaussian-mixture-model speaker recognition from mel-cepstra, derived from mel-filter bank energies (MFBs) of the G.729 codec all-pole spectral envelope, showed significant performance loss relative to the standard mel-cepstral coefficients of G.729 synthesized (coded) speech (Quatieri et al. 1999). In this paper, we investigate two approaches to recover speaker recognition performance from G.729 parameters. The first is a parametric approach that makes explicit use of G.729 parameters, rather than deriving cepstra from MFBs of an all-pole spectrum. Specifically, the G.729 LSFs are converted to "direct" cepstral coefficients for which there exists a one-to-one correspondence with the LSFs. The G.729 residual is also considered; in particular, appending G.729 pitch as a single parameter to the direct cepstral coefficients gives further performance gain. The second nonparametric approach uses the original MFB paradigm, but adds harmonic striations to the G.729 all-pole spectral envelope. Although obtaining considerable performance gains with these methods, we have yet to match the performance of G.729 synthesized speech, motivating the need for representing additional fine structure of the G.729 residual. Thomas F. Quatieri, Robert B. Dunn, Douglas A. Reynolds, Joseph P. Campbell, Elliot Singer |
ICASSP | 1 |
| 2000 | On the influence of rate, pitch, and spectrum on automatic speaker recognition performance
Thomas F. Quatieri, Robert B. Dunn, Douglas A. Reynolds |
INTERSPEECH | 1 |
| 2000 | Estimation of handset nonlinearity with application to speaker recognitionabstractA method is described for estimating telephone handset nonlinearity by matching the spectral magnitude of the distorted signal to the output of a nonlinear channel model, driven by an undistorted reference. This "magnitude only" representation allows the model to directly match unwanted speech formants that arise over nonlinear channels and that are a potential source of degradation in speaker and speech recognition algorithms. As such, the method is particularly suited to algorithms that use only spectral magnitude information. The distortion model consists of a memoryless nonlinearity sandwiched between two finite-length linear filters. Nonlinearities considered include arbitrary finite-order polynomials and parametric sigmoidal functionals derived from a carbon-button handset model. Minimization of a mean-squared spectral magnitude distance with respect to model parameters relies on iterative estimation via a gradient descent technique. Initial work has demonstrated the importance of addressing handset nonlinearity, in addition to linear distortion, in speaker recognition over telephone channels. A nonlinear handset "mapping," applied to training or testing data to reduce mismatch between different types of handset microphone outputs, improves speaker verification performance relative to linear compensation only. Finally, a method is proposed to merge the mapper strategy with a method of likelihood score normalization (hnorm) for further mismatch reduction and speaker verification performance improvement. Thomas F. Quatieri, Douglas A. Reynolds, Gerald C. O'Leary |
IEEE Trans. Speech Audio Process. | 1 |
| 1999 | 'Perfect reconstruction' time-scaling filterbanksabstractA filterbank-based method of time-scale modification is analyzed for elemental signals including clicks, sines, and AM-FM sines. It is shown that with the use of some basic properties of linear systems, as well as FM-to-AM filter transduction, "perfect reconstruction" time-scaling filterbanks can be constructed for these elemental signal classes under certain conditions on the filterbank. Conditions for perfect reconstruction time-scaling are shown analytically for the uniform filterbank case, while empirically for the nonuniform constant-Q (gammatone) case. Extension of perfect reconstruction to multi-components signals is shown to require both filterbank and signal-dependent conditions and indicates the need for a more complete theory of "perfect reconstruction" time-scaling filterbanks. Thomas F. Quatieri, Thomas E. Hanna |
ICASSP | 1 |
| 1999 | Implications of glottal source for speaker and dialect identificationabstractIn this paper we explore the importance of speaker specific information carried in the glottal source. We time align utterances of two speakers speaking the same sentence from the TIMIT database of American English. We then extract the glottal flow derivative from each speaker and interchange them. Through time alignment and this glottal flow transformation, we can make a speaker of a northern dialect sound more like his southern counterpart. We also time align the utterances of two speakers of Spanish dialects speaking the same sentence and then perform the glottal waveform transformation. Through these processes a Peruvian speaker is made to sound more Cuban-like. From these experiments we conclude that significant speaker and dialect specific information, such as noise, breathiness or aspiration, and vocalization, is carried in the glottal signal. Lisa Yanguas, Thomas F. Quatieri |
ICASSP | 2 |
| 1999 | Speaker and language recognition using speech codec parametersabstractThis paper proposes our new text-to-speech (TTS) system that concatenates large numbers of speech segments to produce very natural and intelligible synthetic speech. One novel point of our system is its new synthesis unit, which is has three remarkable characteristics as follows; The synthesis units contain all Japanese syllables together with all possible vowel sequences, so very smooth synthetic speech is produced. Both previous and succeeding phoneme environments are considered when speech segments are concatenated, so natural sounding transients from a vowel to a consonant, which is the only concatenation point with the proposed unit, are present in the synthetic speech. Each unit has various fundamental frequency (F0) contours. Therefore, F0 modification rates are very small in any synthesis event, and the F0 modification process causes only minor distortion. To develop a unit database efficiently and effectively, we analyzed 4,850,000 Japanese phrases (breath-group) containing 87,810,000 phonemes and ranked them in order of appearance frequency. Listening tests confirm the high intelligibility and naturalness of speech produced by our new TTS system. It uses the 50,000 highest frequency units that cover over 77% of Japanese texts. Thomas F. Quatieri, Elliot Singer, Robert B. Dunn, Douglas A. Reynolds, Joseph P. Campbell |
EUROSPEECH | 1 |
| 1999 | Modeling of the glottal flow derivative waveform with application to speaker identificationabstractAn automatic technique for estimating and modeling the glottal flow derivative source waveform from speech, and applying the model parameters to speaker identification, is presented. The estimate of the glottal flow derivative is decomposed into coarse structure, representing the general flow shape, and fine structure, comprising aspiration and other perturbations in the flow, from which model parameters are obtained. The glottal flow derivative is estimated using an inverse filter determined within a time interval of vocal-fold closure that is identified through differences in formant frequency modulation during the open and closed phases of the glottal cycle. This formant motion is predicted by Ananthapadmanabha and Fant (1982) to be a result of time-varying and nonlinear source/vocal tract coupling within a glottal cycle. The glottal flow derivative estimate is modeled using the Liljencrants-Fant (1986) model to capture its coarse structure, while the fine structure of the flow derivative is represented through energy and perturbation measures. The model parameters are used in a Gaussian mixture model speaker identification (SID) system. Both coarse- and fine-structure glottal features are shown to contain significant speaker-dependent information. For a large TIMIT database subset, averaging over male and female SID scores, the coarse-structure parameters achieve about 60% accuracy, the fine-structure parameters give about 40% accuracy, and their combination yields about 70% correct identification. Finally, in preliminary experiments on the counterpart telephone-degraded NTIMIT database, about a 5% error reduction in SID scores is obtained when source features are combined with traditional mel-cepstral measures. Michael David Plumpe, Thomas F. Quatieri, Douglas A. Reynolds |
IEEE Trans. Speech Audio Process. | 2 |
| 1998 | Magnitude-only estimation of handset nonlinearity with application to speaker recognitionabstractA method is described for estimating telephone handset nonlinearity by matching the spectral magnitude of the distorted signal to the output of a nonlinear channel model, driven by an undistorted reference. This "magnitude-only" representation allows the model to directly match unwanted speech formants that arise over nonlinear channels and that are a potential source of degradation in speaker and speech recognition algorithms. As such, the method is particularly suited to algorithms that use only spectral magnitude information. The distortion model consists of a memoryless polynomial nonlinearity sandwiched between two finite-length linear filters. Minimization of a mean-squared spectral magnitude error, with respect to model parameters, relies on iterative estimation via a gradient descent technique, using a Jacobian in the iterative correction term with gradients calculated by finite-element approximation. Initial work has demonstrated the algorithm's usefulness in speaker recognition over telephone channels by reducing mismatch between high- and low-quality handset conditions. Thomas F. Quatieri, Douglas A. Reynolds, Gerald C. O'Leary |
ICASSP | 1 |
| 1997 | AM-FM separation using auditory-motivated filtersabstractAn approach to the joint estimation of sine-wave amplitude modulation (AM) and frequency modulation (FM) is described based on the transduction of frequency modulation into amplitude modulation by linear filters, being motivated by the hypothesis that the auditory system uses a similar transduction mechanism in measuring sine-wave FM. An AM-FM estimation algorithm is described that uses the amplitude envelope of the output of two transduction filters of piecewise-linear spectral shape. The piecewise-linear constraint is then relaxed, allowing a wider class of transduction-filter pairs for AM-FM separation under a monotonicity constraint on the filters' quotient. The particular case of Gaussian filters is shown to yield a closed-form solution to AM-FM estimation while gammatone filters, used as a simplified model of auditory filters, and measured auditory filters, although not leading to a solution in a closed form, provide for iterative AM-FM estimation. Solution stability analysis and error evaluation are performed and the FM transduction method is compared with the energy separation algorithm, based on the Teager energy operator, and the Hilbert transform method for AM-FM estimation. Finally, a generalization to two-dimensional (2-D) filters is described. Thomas F. Quatieri, Thomas E. Hanna, Gerald C. O'Leary |
IEEE Trans. Speech Audio Process. | 1 |
| 1996 | Fine structure features for speaker identificationabstractThe performance of speaker identification (SID) systems can be improved by the addition of the rapidly varying "fine structure" features of formant amplitude and/or frequency modulation and multiple excitation pulses. This paper shows how the estimation of such fine structure features can be improved further by obtaining better estimates of formant frequency locations and uncovering various sources of error in the feature extraction systems. Most female telephone speech showed "spurious" formants, due to distortion in the telephone network. Nevertheless, SID performance was greatest with these spurious formants as formant estimates. A new feature has also been identified which can increase SID performance: cepstral coefficients from noise in the estimated excitation waveform. Finally, statistical tools have been developed to explore the relative importance of features used for SID, with the ultimate goal of uncovering the source of the features that provide SID performance improvement. Charles R. Jankowski Jr., Thomas F. Quatieri, Douglas A. Reynolds |
ICASSP | 2 |
| 1996 | AM-FM separation using auditory-motivated filtersabstractA new approach to the joint estimation of sine-wave amplitude modulation (AM) and frequency modulation (FM) is presented based on the transduction of frequency modulation into amplitude modulation by digital filters, being motivated by the hypothesis that the auditory front end uses a similar transduction mechanism in measuring sine-wave FM. An AM-FM estimation algorithm is derived which uses the amplitude envelope of the output of two transduction filters of piecewise-linear spectral shape. The piecewise-linear constraint is then relaxed, allowing a wider class of transduction filters for AM-FM separation. The particular class of Gaussian filters is shown to yield a closed form solution to AM-FM estimation. Solution stability analysis and error evaluation are performed, and the FM transduction method is shown to typically compare favorably with the Teager energy operator and the Hilbert transform for AM-FM estimation. Finally, the amenability of cochlear gammatone-filter pairs to separate AM and FM is demonstrated. Thomas F. Quatieri, Thomas E. Hanna, Gerald C. O'Leary |
ICASSP | 1 |
| 1996 | Low rate coding of the spectral envelope using channel gainsabstractA dual rate embedded sinusoidal transform coder is described in which a core 14th order allpole coder operating at 2400 b/s is augmented with a set of channel gain residuals in order to operate at the higher 4800 b/s rate. The channel. Gains are a set of non-uniformly spaced samples of the spline envelope and constitute a lowpass estimate of the short-time vocal tract magnitude spectrum. The channel gain residuals represent the difference between the spline envelope and the quantized 14th order allpole spectrum at the channel gain frequencies. The channel gain residuals are coded using pitch dependent scalar quantization. Informal listening indicates that the quality of the embedded coder at 4800 b/s is comparable to that of an existing high quality 4800 b/s allpole coder. Elliot Singer, Robert J. McAulay, Robert B. Dunn, Thomas F. Quatieri |
ICASSP | 4 |
| 1995 | Measuring fine structure in speech: application to speaker identificationabstractThe performance of systems for speaker identification (SID) can be quite good with clean speech, though much lower with degraded speech. Thus it is useful to search for new features for SID, particularly features that are robust over a degraded channel. This paper investigates features that are based on amplitude and frequency modulations of speech formants, high resolution measurement of fundamental frequency and location of "secondary pulses", measured using a high-resolution energy operator. When these features are added to traditional features using an existing SID system with a 168 speaker telephone speech database, SID performance improved by as much as 4% for male speakers and 8.2% for female speakers. Charles R. Jankowski Jr., Thomas F. Quatieri, Douglas A. Reynolds |
ICASSP | 2 |
| 1995 | The effects of telephone transmission degradations on speaker recognition performanceabstractThe two largest factors affecting automatic speaker identification performance are the size of the population and the degradations introduced by noisy communication channels (e.g., telephone transmission). To examine experimentally these two factors, this paper presents text-independent speaker identification results for varying speaker population sizes up to 630 speakers for both clean, wideband speech and telephone speech. A system based on Gaussian mixture speaker models is used for speaker identification and experiments are conducted on the TIMIT and NTIMIT databases. This is believed to be the first speaker identification experiments on the complete 630 speaker TIMIT and NTIMIT databases and the largest text-independent speaker identification task reported to date. Identification accuracies of 99.5% and 60.7% are achieved on the TIMIT and NTIMIT databases, respectively. This paper also presents experiments which examine and attempt to quantify the performance loss associated with various telephone degradations by systematically degrading the TIMIT speech in a manner consistent with measured NTIMIT degradations and measuring the performance loss at each step. It is found that the standard degradations of filtering and additive noise do not account for all of the performance gap between the TIMIT and NTIMIT data. Measurements of nonlinear microphone distortions are also described which may explain the additional performance loss. Douglas A. Reynolds, Marc A. Zissman, Thomas F. Quatieri, Gerald C. O'Leary, Beth A. Carlson |
ICASSP | 3 |
| 1995 | A subband approach to time-scale expansion of complex acoustic signalsabstractA new approach to time-scale expansion of short-duration complex acoustic signals is introduced. Using a subband signal representation, channel phases are selected to preserve a desired time-scaled temporal envelope. The phase representation is derived from locations of events that occur within filter bank outputs. A frame-based generalization of the method imposes phase consistency across consecutive synthesis frames. The method is applied to synthetic and actual complex acoustic signals consisting of closely spaced rapidly damped sine waves. Time-frequency resolution limitations are discussed. Thomas F. Quatieri, Robert B. Dunn, Thomas E. Hanna |
IEEE Trans. Speech Audio Process. | 1 |
| 1994 | High-order allpole modelling of the spectral envelopeabstractThe sinusoidal transform coder (STC) is a vocoding technique that has been evolving over the past few years which represents speech as a harmonic set of sinewaves and has demonstrated synthetic speech of good quality at rates from 2400 b/s to 4800 b/s. Previously, a cepstral model was used to code the sinewave amplitudes. This paper reports on a new approach that fits an allpole model to the sine-wave amplitudes and coding is done in terms of both scalar- and vector-quantization of the line spectral frequencies (LSFs). A new measure is developed to objectively quantify the quantizer design performance.> Terrence G. Champion, Robert J. McAulay, Thomas F. Quatieri |
ICASSP (1) | 3 |
| 1994 | Energy onset times for speaker identificationabstractOnset times of resonant energy pulses are measured with the high-resolution Teager operator and used as features in the Reynolds Gaussian-mixture speaker identification algorithm. Feature sets are constructed with primary pitch and secondary pulse locations derived from low and high speech formants. Preliminary testing was performed with a confusable 40-speaker subset from the NTIMIT (telephone channel) database. Speaker identification improved from 55 to 70% correct classification when the full set of new resonant energy-based features were added as an independent stream to conventional mel-cepstra.> Thomas F. Quatieri, Charles R. Jankowski Jr., Douglas A. Reynolds |
IEEE Signal Process. Lett. | 1 |
| 1993 | Detection of transient signals using the energy operator
Robert B. Dunn, Thomas F. Quatieri, James F. Kaiser |
ICASSP (3) | 2 |
| 1993 | The application of subband coding to improve quality and robustness of the sinusoidal transform coder
Robert J. McAulay, Thomas F. Quatieri |
ICASSP (2) | 2 |
| 1993 | Time-scale modification of complex acoustic signals
Thomas F. Quatieri, Robert B. Dunn, Thomas E. Hanna |
ICASSP (1) | 1 |
| 1992 | On separating amplitude from frequency modulations using energy operatorsabstractTo estimate the amplitude envelope and instantaneous frequency of an AM-FM signal the authors developed a novel approach that uses nonlinear combinations of instantaneous signal outputs from an energy-tracking operator to separate its output energy product into its amplitude modulation and frequency modulation components. This energy separation algorithm is then applied to search for modulations in speech resonances, which the authors model using AM-FM signals. The theoretical and experimental results demonstrate that the energy separation algorithm, due to its low computational complexity and instantaneously adapting nature, is very useful in detecting modulation patterns in speech and other time-varying signals.> Petros Maragos, James F. Kaiser, Thomas F. Quatieri |
ICASSP | 3 |
| 1991 | Speech nonlinearities, modulations, and energy operatorsabstractAn AM-FM model for representing modulations in speech resonances is investigated. Specifically, an FM model is proposed for the time-varying formants whose amplitude varies as the envelope of an AM signal. To detect the modulations the energy operator Psi ( chi ) and its discrete counterpart are applied. It is found that Psi can approximately track the envelope of AM signals, the instantaneous frequency of FM signals, and the product of these two functions in the general case of AM-FM signals. Several experiments on the application of this AM-FM modeling to speech signals, band pass filtered via Gabor filters are reported.> Petros Maragos, Thomas F. Quatieri, James F. Kaiser |
ICASSP | 2 |
| 1991 | Sine-wave phase coding at low data ratesabstractIn the context of a sinusoidal representation for speech waveforms, it is shown that synthetic speech of high quality can be obtained using a parametric model for the sine-wave phases, hence obviating the need to code the phases at low data rates. It was found that if a synthetic linear phase term was computed based on the time of occurrence of an artificially generated sequence of pitch pulses, then high-quality voiced speech reconstruction was possible. For unvoiced speech, the modeling study showed that the sine-wave phases were essentially uniformly distributed random variables.> Robert J. McAulay, Thomas F. Quatieri |
ICASSP | 2 |
| 1990 | Pitch estimation and voicing detection based on a sinusoidal speech modelabstractA technique for estimating the pitch of a speech waveform is developed. It fits a harmonic set of sine waves to the input data using a mean-squared-error (MSE) criterion. By exploiting a sinusoidal model for the input speech waveform, a pitch estimation criterion is derived that is inherently unambiguous, uses pitch-adaptive resolution, uses small-signal suppression to provide enhanced discrimination, and uses amplitude compression to eliminate the effects of pitch-formant interaction. The normalized minimum mean squared error proves to be a powerful discriminant for estimating the likelihood that a given frame of speech is voiced.> Robert J. McAulay, Thomas F. Quatieri |
ICASSP | 2 |
| 1990 | Short-time signal representation by nonlinear difference equationsabstractMethods are investigated for short-time analysis and synthesis of signals from a class of second-order difference equations with a cubic nonlinearity. In analysis, two methods are explored for estimating equation coefficients: (1) prediction error minimization (a linear estimation problem) and (2) waveform error minimization (a nonlinear estimation problem). In the latter case, which improves on the prediction error solution, an iterative analysis-by-synthesis method is derived which allows as free variables initial conditions, as well as equation coefficients. Parameter estimates from these techniques are used in sequential short-time synthesis procedures. Possible application to modeling quasi-periodic behavior in speech waveforms is discussed.> Thomas F. Quatieri, Edward M. Hofstetter |
ICASSP | 1 |
| 1990 | Noise reduction using a soft-decision sine-wave vector quantizerabstractNoise reduction is performed in the context of a high-quality harmonic zero-phase sine-wave analysis/synthesis system which is characterized by sine-wave amplitudes, a voicing probability, and a fundamental frequency. Least-squared error estimation of a harmonic sine-wave representation leads to a soft decision template estimate consisting of sine-wave amplitudes and a voicing probability. The least-squares solution is modified to use template-matching with nearest neighbors. The reconstruction is improved by using the modified least-squares solution only in spectral regions with low signal-to-noise ratio. The results, although preliminary, provide evidence that harmonic zero-phase sine-wave analysis/synthesis, combined with effective estimation of sine-wave amplitudes and probability of voicing, offers a promising approach to noise reduction. > Thomas F. Quatieri, Robert J. McAulay |
ICASSP | 1 |
| 1989 | Phase coherence in speech reconstruction for enhancement and coding applicationsabstractIt has been shown that an analysis-synthesis system based on a sinusoidal representation leads to synthetic speech that is essentially perceptually indistinguishable from the original. A change in speech quality has been observed, however, when the phase relation of the sine waves is altered. This occurs in practice when sine waves are processed for speech enhancement and for speech coding. A description is given of a zero-phase sinusoidal analysis-synthesis system which generates natural-sounding speech without the requirement of vocal tract phase. The method provides a basis for improving sound quality by providing different levels of phase coherence in speech reconstruction for time-scale modification, for a baseline system for coding, and for reducing the peak-to-RMS ratio by dispersion.> Thomas F. Quatieri, Robert J. McAulay |
ICASSP | 1 |
| 1989 | Far-echo cancellation in the presence of frequency offset [full duplex modem]abstractA design is presented for a full-duplex echo-cancelling data modem based on a combined adaptive reference echo canceller and adaptive channel equalizer. The adaptive reference algorithm has the advantage that interference to the echo canceller caused by the far-end signal can be eliminated by subtracting an estimate of the far-end signal based on receiver decisions. This technique provides a novel approach for full-duplex far-echo cancellation in which the far echo can be cancelled in spite of carrier-frequency offset. To estimate the frequency offset, the system uses a separate receiver structure for the far echo which provides equalization of the far echo channel and tracks the frequency offset in the far echo. The feasibility of the echo-cancelling algorithms is demonstrated by computer simulation with realistic channel distortions and with 4800-b/s data transmission, at which rate the frequency offset in the far echo becomes important.> Thomas F. Quatieri, Gerald C. O'Leary |
IEEE Trans. Commun. | 1 |
| 1988 | Computationally efficient sine-wave synthesis and its application to sinusoidal transform codingabstractA technique for sine-wave synthesis is described that uses the fast Fourier transform overlap-add method at a 100 Hz rate based on sine-wave parameter coded at a 50 Hz rate. This technique leads to an implementation requiring less than one-half the computational power of a digital-signal-processor chip. The synthesis method implicitly introduces a frequency jitter which renders the encoded synthetic speech more natural. For speech computed by additive acoustic noise, the synthesizer, in conjunction with straightforward noise suppression, greatly improve the quality of the synthetic speech, rendering the sinusoidal transform coder (STC) algorithm a truly robust system. More recent architecture studies of the STC algorithm suggests that an entire implementation requires no more than two ADSP2 100 chips.> Robert J. McAulay, Thomas F. Quatieri |
ICASSP | 2 |
| 1988 | An approach to co-channel talker interference suppression using a sinusoidal model for speechabstractThe technique fits a sinusoidal model to additive vocalic speech segments such that the least mean square error between the model and the summed waveforms is obtained. Enhancement is achieved by synthesizing a waveform from the sine waves attributed to the desired speaker. Least squares estimation is applied to obtain sine-wave amplitudes and phases for both talkers, based on either a priori sine-wave frequencies or a priori fundamental frequency contours. When the frequencies of the two waveforms are closely spaced, the performance is significantly improved by exploiting the time evolution of the sinusoidal parameters across multiple analysis frames. The least squares error approach is also extended to estimate fundamental frequency contours of both speakers from the summed waveform. The results obtained, though limited in their scope, provide evidence that the sinusoidal analysis/synthesis model with effective parameter estimation techniques offers a promising approach to the problem of cochannel talker interference suppression over a range of conditions.> Thomas F. Quatieri, Ronald G. Danisewicz |
ICASSP | 1 |
| 1988 | Sinewave-based phase dispersion for audio preprocessingabstractA sinusoidal-based analysis/synthesis system is used to apply radar signal design solutions to the problem of dispersing the phase of a speech waveform. Integrated with dynamic range compression, the resulting system can give a significant reduction in peak/RMS ratio with acceptable loss in quality. The spectral information in the resulting processed speech waveform is embedded primarily within the zero crossings of the modified waveform, rather than the waveform shape. Consequently, this dispersion technique also serves as a preprocessor for waveform clipping, allowing considerably deeper thresholding than can be tolerated on the original waveform.> Thomas F. Quatieri, Robert J. McAulay |
ICASSP | 1 |
| 1987 | "Multirate sinusoidal transform coding at rates from 2.4 kbps to 8 kbps"abstractIt has been shown [1] that an analysis/synthesis system based on a sinusoidal representation leads to synthetic speech that is essentially indistinguishable from the original. By exploiting the peak-to-peak correlation of the sine-wave amplitudes [2], a harmonic model for the sine-wave frequencies, and a predictive model for the sine-wave phases [3], it has also been shown that the sine-wave parameters can be coded at 8 kbps. In this paper a new technique is described for coding the sine-wave amplitudes based on the idea of a pitch-adaptive channel vocoder. Using this amplitude-coding strategy and operating at a total bit rate of 4.8 kbps, it was possible to code and transmit enough phase information so that very intelligible, natural sounding speech could be synthesized. This 4.8 kbps system has been implemented in real-time and has achieved a Diagnostic Rhyme Test (DRT) score of 95. At 2.4 kbps no explicit phase information could be coded, but by phase-locking all of the sine waves to the fundamental, by adding a pitch-adaptive quadratic phase, and by adding a voicing dependent random phase to each sine wave, natural sounding synthetic speech could be obtained. This new system is currently being implemented in real-time so that intelligibility tests can be performed. Robert J. McAulay, Thomas F. Quatieri |
ICASSP | 2 |
| 1987 | Mixed-phase deconvolution of speech based on a sine-wave modelabstractThis paper describes a new method of deconvolving the vocal cord excitation and vocal tract system response. The technique relies on a sine-wave representation of the speech waveform and forms the basis of an analysis-synthesis method which yields synthetic speech essentially indistinguishable from the original. Unlike an earlier sinusoidal analysis-synthesis technique that used a minimum-phase system estimate, the approach in this paper generates a "mixed-phase" system estimate and thus an improved decomposition of excitation and system components. Since a mixed-phase system estimate is removed from the speech waveform, the resulting excitation residual is less dispersed than the previous sinusoidal-based excitation estimate or the more commonly used linear prediction residual. A method of time-varying linear filtering is given as an alternative to sinusoidal reconstruction, similar to conventional time-domain synthesis used in certain vocoders, but without the requirement of pitch and voicing decisions. Finally, speech modification with a mixed-phase system estimate is shown to be capable of more closely preserving waveform shape in time-scale and pitch transformations than the earlier approach. Thomas F. Quatieri, Robert J. McAulay |
ICASSP | 1 |
| 1986 | Phase modelling and its application to sinusoidal transform codingabstractIn the context of a sinusoidal representation for speech waveforms [1-4], it is shown that the excitation sine-wave phases can be modelled as a linear function of frequency, the slope being determined by the onset time of the underlying pitch pulse. An algorithm for estimating the pitch onset time from the original sine-wave parameters is derived. The new phase model has been used to improve the coding efficiency of the sinusoidal transform coder (STC) and for compensating phase errors introduced by the quantization of the sine-wave frequencies. Robert J. McAulay, Thomas F. Quatieri |
ICASSP | 2 |
| 1986 | Decision-directed echo cancellation for full-duplex data transmission at 4800bpsabstractIn this paper, we present a design for a full-duplex near-echo-cancelling data modem based on a cambined "decision-directed" (DD) adaptive echo canceller and adaptive channel equalizer. This algorithm has the important feature that the interference to the canceller caused by the far-end signal is eliminated by subtracting an estimate of the far-end signal based on receiver decisions. This technique is extended to provide a new approach to full-duplex far-echo cancellation in which frequency offset of the far echo is adaptively estimated with little interference from the far-end signal and near echo. The approach uses receiver decisions on the far-end signal, but in addition, relies on a separate receiver structure devoted to joint estimation of the far-echo channel and the frequency offset in the far echo. The feasibility of both echo-cancelling modem algorithms is demonstrated by computer simulation with realistic channel distortions and with 4800 bps data transmission at which rate frequency offset in the far echo becomes important. Thomas F. Quatieri, Gerald C. O'Leary |
ICASSP | 1 |
| 1986 | Statistical model-based algorithms for image analysisabstractIn this paper, two-dimensional stochastic linear models are used in developing algorithms for image analysis such as classification, segmentation, and object detection in images characterized by textured backgrounds. These models generate two-dimensional random processes as outputs to which statistical inference procedures can naturally be applied. A common thread throughout our algorithms is the interpretation of the inference procedures in terms of linear prediction residuals. This interpretation leads to statistical tests more insightful than the original tests and makes the procedures computationally tractable. This paper also examines a computational structure tailored to one of the algorithms. In particular, we describe a processor based on systolic arrays that realizes the object detection algorithm developed in the paper. Charles W. Therrien, Thomas F. Quatieri, Dan E. Dudgeon |
Proc. IEEE | 2 |
| 1985 | Mid-rate coding based on a sinusoidal representation of speechabstractIn this paper a sinusoidal model for the speech waveform is used to develop a new analysis/synthesis technique that is characterized by the amplitudes, frequencies, and phases of the component sine waves. The resulting synthetic waveform preserves the waveform shape and is essentially perceptually indistinguishable from the original speech. Furthermore, in the presence of noise the perceptual characteristics of the speech and the noise are maintained. Based on this system, a coder operating at 8 kbps is developed that codes the amplitudes and phases of each of the sine wave components and uses a harmonic model to code all of the frequencies. Since not all of the phases can be coded, a high frequency regeneration technique is developed that exploits the properties of the sinusoidal representation of the coded baseband signal. Based on a relatively limited data base, computer simulation has demonstrated that coded speech of good quality can be achieved. A real-time simulation is being developed to provide a more thorough evaluation of the algorithm. Robert J. McAulay, Thomas F. Quatieri |
ICASSP | 2 |
| 1985 | Speech transformations based on a sinusoidal representationabstractThis paper presents a new speech analysis/synthesis technique based on a sinusoidal representation of the speech production mechanism but which is independent of pitch and the voiced/unvoiced speech state. The resulting synthetic speech preserves the waveform shape and is essentially perceptually indistinguishable from the original. The method provides the basis for a general class of speech transformations and is successfully applied to time-scale modification, frequency scaling, and scaling of pitch. Furthermore, these modifications can be performed with a time-varying rate of change, allowing, for example, continuous adjustment of a speaker's fundamental frequency and rate of articulation. Although the analysis/synthesis system was originally designed for single-speaker signals, it is equally capable of recovering and modifying nonspeech signals such as music, multi-speakers, marine biologic sounds, and speech in the presence of interferences such as noise and musical backgrounds. Thomas F. Quatieri, Robert J. McAulay |
ICASSP | 1 |
| 1984 | Magnitude-only reconstruction using a sinusoidal speech modelMagnitude-only reconstruction using a sinusoidal speech modelabstractIn this paper a sinusoidal model for the speech waveform is used to develop a new synthesis technique that requires specification of only the amplitudes and frequencies of the component sine waves. These parameters are estimated from the short-time spectral magnitude. The resulting synthetic waveform preserves the short-time spectral magnitude during rapid movements of spectral energy such as voiced/unvoiced transitions, and yields speech of very high quality and intelligibility. The approach is sufficiently flexible to also allow for high-quality time-scale modification with the option of time-varying scaling. Finally, results are given for some initial experiments that explore the possibility of magnitude-only waveform coding at 8 kbps. Robert J. McAulay, Thomas F. Quatieri |
ICASSP | 2 |
| 1984 | Homomorphic restoration of images degraded by light cloud coverabstractIn this paper, we demonstrate the use of adaptive homomorphic filtering in exposing objects under light cloud cover. In particular, the homomorphic filter invoked is space-varying and is parameterized by the local mean level of the degraded image. The local mean serves as an indication of the extent of local cloud cover degradation. We show that this adaptive procedure has greater potential than the long-space methods in exposing objects beneath light cloud cover. In addition, adaptive homomorphic filtering compares favorably with an iterative homomorphic enhancement procedure which is an extension of the one-pass nonadaptive homomorphic filter. Tamar Peli, Thomas F. Quatieri |
ICASSP | 2 |
| 1983 | Algorithms for signal reconstruction from short-time Fourier transform magnitudeabstractWe have previously established a number of conditions under which a real signal can be uniquely reconstructed from its STFT magnitude. For the STFT magnitude to be a practical signal representation, we need robust reconstruction algorithms. In this paper, we discuss such a class of algorithms within the framework of sequential extrapolation techniques. Such techniques reconstruct the short-time sections of a signal in an order determined by their positions on the time axis. We find that compared to direct reconstruction, the robust algorithms are less sensitive to roundoff errors. To further test the robustness of these algorithms, we applied them to STFT magnitudes which were purposely modified for accomplishing signal processing tasks such as noise reduction and time-scale modification of speech. S. Hamid Nawab, Thomas F. Quatieri, Jae S. Lim |
ICASSP | 2 |
| 1983 | Object detection by two-dimensional linear predictionabstractThis paper addresses the problem of detecting small areas of textured images which differ from their immediate surroundings. A significance test is described which adapts itself to the generally changing background statistics so that a constant false alarm rate is maintained. A detection algorithm is derived from the fact that this significance test can be expressed in terms of the error residuals of an adaptive two-dimensional linear predictor whose coefficients are estimated from the background. The algorithm has been successfully demonstrated with both synthetic and real-world images. Thomas F. Quatieri |
ICASSP | 1 |
| 1982 | The importance of boundary conditions in the phase retrieval problemabstractIn this paper, the importance of the boundary values of a two dimensional sequence in the phase retrieval problem is investigated. Specifically, it is shown that, given the boundary values, phase retrieval becomes a linear problem. Therefore, the interior points of the two-dimensional sequence may easily be deduced from the boundary values of the sequence and the magnitude of its Fourier transform or, equivalently, the autocorrelation function of the sequence. Furthermore, although the determination of the boundary values from only Fourier transform magnitude information is, in general, a non-trivial problem, it is shown that the shape of the region of support of the two-dimensional sequence determines the ease with which the boundary values may be determined. Monson H. Hayes III, Thomas F. Quatieri |
ICASSP | 2 |
| 1982 | Signal reconstruction from the short-time Fourier transform magnitudeabstractThis paper presents various conditions that are sufficient for reconstructing a discrete-time signal from samples of its short-time Fourier transform magnitude. For applications such as speech processing, these conditions place very mild restrictions on the signal as well as the analysis window of the transform. Examples of such reconstruction for speech signals are included in the paper. S. Hamid Nawab, Thomas F. Quatieri, Jae S. Lim |
ICASSP | 2 |
| 1981 | A nested algorithm for improving the accuracy of chirp-Fourier transform implementationsabstractThe chirp Z-transform permits large bandwidth (i.e. high speed) analog devices to be used to compute the Discrete Fourier Transform (DFT). Such Implementations are often faster and/or smaller than all digital Fast Fourier Transforms (FFT's), but suffer from accuracy limitations imposed by the analog devices. In this paper, we present a nested algorithm that permits excess bandwidth of the analog device to be exploited in order to compute the DFT more accurately. The nested structure can be realized with a single analog device, and can be trained to the errors associated with the device. Algorithm performance is verified through computer simulations in which DFT accuracy is limited due to fixed tap weight errors in the FIR filter portion of the chirp Z-transform. Stephen C. Pohlig, Gary A. Shaw, Theodore Bially, Thomas F. Quatieri |
ICASSP | 4 |
| 1981 | Extensions of 2-D iterative digital filtersabstractA 2-D digital filter with a rational frequency response, as previously shown, can be implemented with an infinite sequence of finite-extent impulse response (FIR) filtering operations, provided a certain convergence criterion is satisfied. In this paper, the iterative procedure is generalized so that the convergence requirement is always met, thus allowing an iterative implementation of any stable rational 2-D infinite impulse response filter. The modification requires the pre-filtering of the numerator and denominator polynomials of the rational frequency response and a relaxed formulation of the original iteration. A frequency-varying relaxation parameter is considered for increasing the rate of convergence over the entire frequency band of interest. Tradeoffs between the size of the 2-D FIR filters used within the iteration and convergence rates are demonstrated by applying the relaxed form of the iteration to the problem of image deblurring. Thomas F. Quatieri, Dan E. Dudgeon |
ICASSP | 1 |
| 1981 | Convergence of iterative signal reconstruction algorithmsabstractThis paper is concerned with the development of a general convergence proof which is applicable to a particular class of iterative signal reconstruction algorithms. The proof relies on the concept of a nonexpansive mapping and the uniqueness of the desired signal. Two specific iterative reconstruction algorithms which are considered to illustrate this method of proof are bandlimited extrapolation and phase-only signal reconstruction, the convergence proof of the later being a new result. The generality of the approach allows for the incorporation of nonlinear constraints such as positivity or minimum and maximum value constraints. Thomas F. Quatieri, Victor T. Tom, Monson H. Hayes III, James H. McClellan |
ICASSP | 1 |
| 1979 | A mixed-phase homomorphic vocoderabstractIn this paper short-time homomorphic analysis and the harmonic representation of voiced speech are explored with the result of a mixed-phase homomorphic vocoder of somewhat higher quality than its minimum-phase counterpart. A theoretical framework is presented for unwrapped phase estimation from harmonic spectra through smoothing real and imaginary spectral components. The short-time harmonic model leads to pitch adaptive duration and alignment requirements on time-domain windowing. The underlying phase envelope is consequently preserved so that cepstral windowing can he applied. In addition, two alternative vocoders with mixed-phase are considered: the first is based on linear interpolation of complex harmonic peaks, and the second on Lim's spectral root deconvolution scheme. Thomas F. Quatieri |
ICASSP | 1 |
| 1978 | CCD CZT Spectral analysis applied to real time homomorphics speech analysis-synthesisabstractIn this paper we discuss real time spectral analysis using charge-coupled device (CCD) nonrecursive filters with an application to homomorphic speech analysis-synthesis. We consider the conventional chirp-z-transform (CZT) realization of the discrete Fourier transform (DFT) and an analogous CZT realization of a sliding transform, a modified DFT, better suited to CCD technology than the DFT. Novel CCD configurations based on both transforms are proposed for improved spectral resolution and frame rate; such schemes utilize parallel spectral processing and interpolation. Stationarity conditions and windowing techniques required by these CCD structures are discussed within the context of short-time spectral analysis. A number of homomorphic schemes and experimental results are presented which illustrate the trade-offs that CCD technology imposes between implementational complexity and synthetic speech quality. Thomas F. Quatieri |
ICASSP | 1 |