Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Daryush D. Mehta

dblp:56/9879 · DBLP profile ↗
← Back
24ranked-venue papers
8as first author
1since 2021 · last 2024
0000-0002-6535-573XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 5 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Audio and music processing · 83% Multimedia systems and quality of experience · 17%
Artificial intelligence
1 paper
Speech recognition and synthesis · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis › speech synthesis
articulatory speech synthesis
0.312017
Evaluation of Glottal Inverse Filtering Algorithms Using a Physiologically Based Articulatory Speech Synthesizer · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Natural language and speech › Speech recognition and synthesis
speech analysis
0.312017
Evaluation of Glottal Inverse Filtering Algorithms Using a Physiologically Based Articulatory Speech Synthesizer · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Natural language and speech › Speech recognition and synthesis
speech synthesis
0.312017
Evaluation of Glottal Inverse Filtering Algorithms Using a Physiologically Based Articulatory Speech Synthesizer · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Audio and music processing
audio feature extraction
0.312017
Modal and Nonmodal Voice Quality Classification Using Acoustic and Electroglottographic Features · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Audio and music processing › speech analysis
glottal source features
0.312017
Modal and Nonmodal Voice Quality Classification Using Acoustic and Electroglottographic Features · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Audio and music processing
speech analysis
0.312017
Modal and Nonmodal Voice Quality Classification Using Acoustic and Electroglottographic Features · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Multimedia systems and quality of experience › multimedia quality assessment
voice quality assessment
0.312017
Modal and Nonmodal Voice Quality Classification Using Acoustic and Electroglottographic Features · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Audio and music processing › speech analysis
voice analysis
0.212016
Relationships Between Vocal Function Measures Derived from an Acoustic Microphone and a Subglottal Neck-Surface Accelerometer · IEEE ACM Trans. Audio Speech Lang. Process. 2016
Audio and music processing
speech production
0.212013
Subglottal Impedance-Based Inverse Filtering of Voiced Sounds Using Neck Surface Acceleration · IEEE Trans. Speech Audio Process. 2013

Methods — techniques the papers use, named apart from their topics

support vector machine · 0.3random forest · 0.3linear prediction · 0.3gaussian mixture model · 0.3deep neural network · 0.3covariance analysis · 0.3mechano-acoustic impedance modeling · 0.2inverse filtering · 0.2
YearPublicationVenuePosition
2024 Comparing ambulatory voice measures during daily life with brief laboratory assessments in speakers with and without vocal hyperfunction
abstract
The most common types of voice disorders are associated with hyperfunctional voice use in daily life. Although current clinical practice uses measures from brief laboratory recordings to assess vocal function, it is unclear how these relate to an individual's habitual voice use. The purpose of this study was to quantify the correlation and offset between voice features computed from laboratory and ambulatory recordings in speakers with and without vocal hyperfunction. Features derived from a neck-surface accelerometer included estimates of sound pressure level, fundamental frequency, cepstral peak prominence, and spectral tilt. Whereas some measures from laboratory recordings correlated significantly with those captured during daily life, only approximately 6-52% of the actual variance was accounted for. Thus, brief voice assessments are quite limited in the extent to which they can accurately characterize the daily voice use of speakers with and without vocal hyperfunction.
Daryush D. Mehta, Jarrad H. Van Stan, Hamzeh Ghasemzadeh, Robert E. Hillman
INTERSPEECH1
2019 Correlating an Ambulatory Voice Measure to Electrodermal Activity in Patients with Vocal Hyperfunction
abstract
We investigate the connection between the autonomic nervous system and the voice in patients with vocal hyperfunction and healthy-control groups. We present a methodology and preliminary results of two multi-modal measurement streams that capture this relationship. Subjects were instrumented for daily, ambulatory collection of their voice and wrist-based electrodermal activity. Measures of vocal function (e.g., fundamental frequency) were computed, as well as measures of autonomic function (e.g., skin conductance response). Spearman correlation coefficients were calculated to measure the relationship between vocal and autonomic function over sliding windows throughout each observation day. We found preliminary evidence that patients with a subtype of vocal hyperfunction (non-phonotraumatic vocal hyperfunction) exhibit a coupling between the autonomic nervous system and the vocal system. Understanding how the autonomic nervous system interacts with the voice may provide new insights into the etiology/pathophysiology of vocal hyperfunction and improve prevention, diagnosis and treatment of these disorders.
Gregory A. Ciccarelli, Daryush D. Mehta, Andrew Ortiz, Jarrad H. Van Stan, Laura Toles, Katherine L. Marks, Robert E. Hillman, Thomas F. Quatieri
BSN2
2019 On-Body Monitoring of Voice-Based Cognitive Load Features in an Auditory Working Memory Task
abstract
The ability to monitor an individual's cognitive load in operational and naturalistic settings is of great importance in improving military performance and readiness. This paper describes preliminary results of using an on-body, multimodal voice monitoring system for assessing cognitive load based on vocal characteristics extracted from noise-robust sensors. The main components of the system include a commercial wired electroglottograph (EGG) and a lightweight, flex circuit that houses two sensors: an acoustic MEMS microphone (MIC) and a non-acoustic neck-surface accelerometer (ACC). We conducted human subject experiments using a cognitive load protocol under quiet laboratory conditions and computed a previously investigated vocal biomarker (creaky voice quality) from the MIC, ACC, and EGG signals. Results demonstrated the potential of discriminating low versus high cognitive load using the creaky voice correlation structure within each of the three sensor domains. Combining MIC, ACC, and EGG hardware into an integrated system would provide complementary information for quantifying voice and speech biomarkers with a high degree of robustness to competing noise sources.
Daryush D. Mehta, Rohan Deshpande, Luke Letter, Edward Froehlich, Andrew M. Siegel, Thomas F. Quatieri, Laura J. Brattain
BSN1
2019 Vocal Biomarker Assessment Following Pediatric Traumatic Brain Injury: A Retrospective Cohort Study
Camille Noufi, Adam C. Lammert, Daryush D. Mehta, James R. Williamson, Gregory A. Ciccarelli, Douglas E. Sturim, Jordan R. Green, Thomas F. Campbell, Thomas F. Quatieri
INTERSPEECH3
2018 Lightweight, on-body, wireless system for ambulatory voice and ambient noise monitoring
abstract
In this paper, we present a lightweight, on-body, wireless system designed for monitoring real-world, ambulatory voice characteristics. The system has the potential to provide important assessments of voice and speech disorders and the impact of environmental sound levels as individuals go about their daily life. The system's transmitter is positioned on the neck and synchronously streams dual-channel sensor data from an on-board MEMS microphone and a high-bandwidth accelerometer, which acts as a noise-robust and confidential contact microphone. These data are recorded to a receiver that can store the data locally and stream a real-time feed to a computer. We also report on the design considerations of this novel system and discuss progress leading up to the latest iteration, especially of the transmitter components on a flexible circuit. Pilot data are shown from an in-field, ambulatory recording during an individual's daily activities that included settings in quiet and with naturalistic ambient noise.
Patrick Chwalek, Daryush D. Mehta, Brendon Welsh, Catherine Wooten, Kate Byrd, Edward Froehlich, David Maurer, Joseph Lacirignola, Thomas F. Quatieri, Laura J. Brattain
BSN2
2018 Vocal Biomarkers for Cognitive Performance Estimation in a Working Memory Task
Jennifer Sloboda, Adam C. Lammert, James R. Williamson, Christopher J. Smalt, Daryush D. Mehta, C. O. L. Ian Curry, Kristin Heaton, Jeff Palmer, Thomas F. Quatieri
INTERSPEECH5
2017 Classification of voice modes using neck-surface accelerometer data
abstract
This study analyzes signals recorded using a neck-surface accelerometer from subjects producing speech with different voice modes. The purpose is to explore if the recorded waveforms can capture the glottal vibratory patterns which can be related to the movement of the vocal folds and thus voice quality. The accelerometer waveforms do not contain the supraglottal resonances, and these characteristics make the proposed method suitable for real-life voice quality assessment and monitoring as it does not breach patient privacy. The experiments with a Gaussian mexture model classifier demonstrate that different voice qualities produce distinctly different accelerometer waveforms. The system achieved 80.2% and 89.5% for frame- and utterance-level accuracy, respectively, for classifying among modal, breathy, pressed, and rough voice modes using a speaker-dependent classifier. Finally, the article presents characteristic waveforms for each modality and discusses their attributes.
Michal Borsky, Marion Cocude, Daryush D. Mehta, Matías Zanartu, Jón Guðnason
ICASSP3
2017 Wireless Neck-Surface Accelerometer and Microphone on Flex Circuit with Application to Noise-Robust Monitoring of Lombard Speech
Daryush D. Mehta, Patrick Chwalek, Thomas F. Quatieri, Laura J. Brattain
INTERSPEECH1
2017 Modal and Nonmodal Voice Quality Classification Using Acoustic and Electroglottographic Features
abstract
The goal of this study was to investigate the performance of different feature types for voice quality classification using multiple classifiers. The study compared the COVAREP feature set; which included glottal source features, frequency warped cepstrum and harmonic model features; against the mel-frequency cepstral coefficients (MFCCs) computed from the acoustic voice signal, acoustic-based glottal inverse filtered (GIF) waveform, and electroglottographic (EGG) waveform. Our hypothesis was that MFCCs can capture the perceived voice quality from either of these three voice signals. Experiments were carried out on recordings from 28 participants with normal vocal status who were prompted to sustain vowels with modal and non-modal voice qualities. Recordings were rated by an expert listener using the Consensus Auditory-Perceptual Evaluation of Voice (CAPE-V), and the ratings were transformed into a dichotomous label (presence or absence) for the prompted voice qualities of modal voice, breathiness, strain, and roughness. The classification was done using support vector machines, random forests, deep neural networks and Gaussian mixture model classifiers, which were built as speaker independent using a leave-one-speaker-out strategy. The best classification accuracy of 79.97% was achieved for the full COVAREP set. The harmonic model features were the best performing subset, with 78.47% accuracy, and the static+dynamic MFCCs scored at 74.52%. A closer analysis showed that MFCC and dynamic MFCC features were able to classify modal, breathy, and strained voice quality dimensions from the acoustic and GIF waveforms. Reduced classification performance was exhibited by the EGG waveform.
Michal Borsky, Daryush D. Mehta, Jarrad H. Van Stan, Jón Guðnason
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Evaluation of Glottal Inverse Filtering Algorithms Using a Physiologically Based Articulatory Speech Synthesizer
abstract
Glottal inverse filtering aims to estimate the glottal airflow signal from a speech signal for applications such as speaker recognition and clinical voice assessment. Nonetheless, evaluation of inverse filtering algorithms has been challenging due to the practical difficulties of directly measuring glottal airflow. Apart from this, it is acknowledged that the performance of many methods degrade in voice conditions that are of great interest, such as breathiness, high pitch, soft voice, and running speech. This paper presents a comprehensive, objective, and comparative evaluation of state-of-the-art inverse filtering algorithms that takes advantage of speech and glottal airflow signals generated by a physiological speech synthesizer. The synthesizer provides a physics-based simulation of the voice production process and thus an adequate test bed for revealing the temporal and spectral performance characteristics of each algorithm. Included in the synthetic data are continuous speech utterances and sustained vowels, which are produced with multiple voice qualities (pressed, slightly pressed, modal, slightly breathy, and breathy), fundamental frequencies, and subglottal pressures to simulate the natural variations in real speech. In evaluating the accuracy of a glottal flow estimate, multiple error measures are used, including an error in the estimated signal that measures overall waveform deviation, as well as an error in each of several clinically relevant features extracted from the glottal flow estimate. Waveform errors calculated from glottal flow estimation experiments exhibited mean values around 30% for sustained vowels, and around 40% for continuous speech, of the amplitude of true glottal flow derivative. Closed-phase approaches showed remarkable stability across different voice qualities and subglottal pressures. The algorithms of choice, as suggested by significance tests, are closed-phase covariance analysis for the analysis of sustained vowels, and sparse linear prediction for the analysis of continuous speech. Results of data subset analysis suggest that analysis of close rounded vowels is an additional challenge in glottal flow estimation.
Yu-Ren Chien, Daryush D. Mehta, Jón Guðnason, Matías Zanartu, Thomas F. Quatieri
IEEE ACM Trans. Audio Speech Lang. Process.2
2016 Classification of Voice Modality Using Electroglottogram Waveforms
Michal Borsky, Daryush D. Mehta, Julius P. Gudjohnsen, Jón Guðnason
INTERSPEECH2
2016 Relation of Automatically Extracted Formant Trajectories with Intelligibility Loss and Speaking Rate Decline in Amyotrophic Lateral Sclerosis
Rachelle L. Horwitz-Martin, Thomas F. Quatieri, Adam C. Lammert, James R. Williamson, Yana Yunusova, Elizabeth Godoy, Daryush D. Mehta, Jordan R. Green
INTERSPEECH7
2016 Relationships Between Vocal Function Measures Derived from an Acoustic Microphone and a Subglottal Neck-Surface Accelerometer
abstract
Monitoring subglottal neck-surface acceleration has received renewed attention due to the ability of low-profile accelerometers to confidentially and noninvasively track properties related to normal and disordered voice characteristics and behavior. This study investigated the ability of subglottal neck-surface acceleration to yield vocal function measures traditionally derived from the acoustic voice signal and help guide the development of clinically functional accelerometer-based measures from a physiological perspective. Results are reported for 82 adult speakers with voice disorders and 52 adult speakers with normal voices who produced the sustained vowels /a/, /i/, and /u/ at a comfortable pitch and loudness during the simultaneous recording of radiated acoustic pressure and subglottal neck-surface acceleration. As expected, timing-related measures of jitter exhibited the strongest correlation between acoustic and neck-surface acceleration waveforms (r≤0.99), whereas amplitude-based measures of shimmer correlated less strongly (r≤0.74). Additionally, weaker correlations were exhibited by spectral measures of harmonics-to-noise ratio (r≤0.69) and tilt (r≤0.57), whereas the cepstral peak prominence correlated more strongly (r≤0.90). These empirical relationships provide evidence to support the use of accelerometers as effective complements to acoustic recordings in the assessment and monitoring of vocal function in the laboratory, clinic, and during an individual's daily activities.
Daryush D. Mehta, Jarrad H. Van Stan, Robert E. Hillman
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 Evaluation of speech inverse filtering techniques using a physiologically based synthesizer
abstract
Glottal inverse filtering methods are designed to derive a glottal flow waveform from a speech signal. In this paper, we evaluate and compare such methods using a speech synthesizer that simulates voice production in a physiologically-based manner that includes complexities such as nonlinear source-tract coupling. Five inverse filtering techniques are evaluated on 90 synthesized speech waveforms generated by setting six vowel configurations, three glottal models, and five fundamental frequencies. Using normalized mean square error as the primary performance metric of the estimated glottal flow derivative, results show that the accuracy of all methods depends on the configuration of the vocal tract, glottis and the fundamental frequency. Averaged over these conditions, the closed phase covariance and one weighted covariance algorithm yield lower error rates (0.41 ± 0.2) than iterative and adaptive inverse filtering (0.49 ± 0.1) and complex cepstrum decomposition (0.76 ± 0.1).
Jón Guðnason, Daryush D. Mehta, Thomas F. Quatieri
ICASSP2
2015 Vocal biomarkers to discriminate cognitive load in a working memory task
Thomas F. Quatieri, James R. Williamson, Christopher J. Smalt, Tejash Patel, Joseph Perricone, Daryush D. Mehta, Brian S. Helfer, Gregory A. Ciccarelli, Darrell O. Ricke, Nicolas Malyska, Jeff Palmer, Kristin Heaton, Marianna Eddy, Joseph Moran
INTERSPEECH6
2015 Segment-dependent dynamics in predicting parkinson's disease
James R. Williamson, Thomas F. Quatieri, Brian S. Helfer, Joseph Perricone, Satrajit S. Ghosh, Gregory A. Ciccarelli, Daryush D. Mehta
INTERSPEECH7
2015 Three-Dimensional Optical Reconstruction of Vocal Fold Kinematics Using High-Speed Video With a Laser Projection System
abstract
Vocal fold kinematics and its interaction with aerodynamic characteristics play a primary role in acoustic sound production of the human voice. Investigating the temporal details of these kinematics using high-speed videoendoscopic imaging techniques has proven challenging in part due to the limitations of quantifying complex vocal fold vibratory behavior using only two spatial dimensions. Thus, we propose an optical method of reconstructing the superior vocal fold surface in three spatial dimensions using a high-speed video camera and laser projection system. Using stereo-triangulation principles, we extend the camera-laser projector method and present an efficient image processing workflow to generate the three-dimensional vocal fold surfaces during phonation captured at 4000 frames per second. Initial results are provided for airflow-driven vibration of an ex vivo vocal fold model in which at least 75% of visible laser points contributed to the reconstructed surface. The method captures the vertical motion of the vocal folds at a high accuracy to allow for the computation of three-dimensional mucosal wave features such as vibratory amplitude, velocity, and asymmetry.
Georg Luegmair, Daryush D. Mehta, James B. Kobler, Michael Döllinger
IEEE Trans. Medical Imaging2
2014 Closed phase estimation for inverse filtering the oral airflow waveform
abstract
Glottal closed phase estimation during speech production is critical to inverse filtering and, although addressed for radiated acoustic pressure analysis, must be better understood for the analysis of the oral airflow volume velocity signal that provides important properties of healthy and disordered voices. This paper compares the estimation of the closed phase from the acoustic speech signal and the oral airflow waveform recorded using a pneumotachograph mask. Results are presented for ten adult speakers with normal voices who sustained a set of vowels at a comfortable pitch and loudness. With electroglottography as reference, the identification rate and accuracy of glottal closure instants for the oral airflow are 96.8 % and 0.28 ms, whereas these metrics are 99.4 % and 0.10 ms for the acoustic signal. We conclude that glottal closure detection is adequate for close phase inverse filtering but that improvements to detection of glottal opening instants on the oral airflow signal are warranted.
Jón Guðnason, Daryush D. Mehta, Thomas F. Quatieri
ICASSP2
2013 Smartphone-based detection of voice disorders by long-term monitoring of neck acceleration features
abstract
Many common voice disorders are chronic or recurring conditions that are likely to result from inefficient and/or abusive patterns of vocal behavior, termed vocal hyperfunction. Thus an ongoing goal in clinical voice assessment is the long-term monitoring of noninvasively derived measures to track hyperfunction. This paper reports on a smartphone-based voice health monitor that records the high-bandwidth accelerometer signal from the neck skin above the collarbone. Data collection is under way from patients with vocal hyperfunction and matched-control subjects to create a dataset designed to identify the best set of diagnostic measures for hyperfunctional patterns of vocal behavior. Vocal status is tracked from neck acceleration using previously-developed vocal dose measures and novel model-based features of glottal airflow estimates. Clinically, the treatment of hyperfunctional disorders would be greatly enhanced by the ability to unobtrusively monitor and quantify detrimental behaviors and, ultimately, to provide real-time biofeedback that could facilitate healthier voice use.
Daryush D. Mehta, Matías Zanartu, Jarrad H. Van Stan, Shengran W. Feng, Harold A. Cheyne II, Robert E. Hillman
BSN1
2013 Classification of depression state based on articulatory precision
abstract
Neurophysiological changes in the brain associated with major depression disorder can disrupt articulatory precision in speech production. Motivated by this observation, we address the hypothesis that articulatory features, as manifested through formant frequency tracks, can help in automatically classifying depression state. Specifically, we investigate the relative importance of vocal tract formant frequencies and their dynamic features from sustained vowels and conversational speech. Using a database consisting of audio from 35 subjects with clinical measures of depression severity, we explore the performance of Gaussian mixture model (GMM) and support vector machine (SVM) classifiers. With only formant frequencies and their dynamics given by velocity and acceleration, we show that depression state can be classified with an optimal sensitivity/specificity/area under the ROC curve of 0.86/0.64/0.70 and 0.77/0.77/0.73 for GMMs and SVMs, respectively. Future work will involve merging our formant-based characterization with vocal source and prosodic features. Index Terms: major depressive disorder, motor coordination, articulatory control, vocal biomarkers, formant frequencies
Brian S. Helfer, Thomas F. Quatieri, James R. Williamson, Daryush D. Mehta, Rachelle L. Horwitz-Martin, Bea Yu
INTERSPEECH4
2013 Subglottal Impedance-Based Inverse Filtering of Voiced Sounds Using Neck Surface Acceleration
abstract
A model-based inverse filtering scheme is proposed for an accurate, non-invasive estimation of the aerodynamic source of voiced sounds at the glottis. The approach, referred to as subglottal impedance-based inverse filtering (IBIF), takes as input the signal from a lightweight accelerometer placed on the skin over the extrathoracic trachea and yields estimates of glottal airflow and its time derivative, offering important advantages over traditional methods that deal with the supraglottal vocal tract. The proposed scheme is based on mechano-acoustic impedance representations from a physiologically-based transmission line model and a lumped skin surface representation. A subject-specific calibration protocol is used to account for individual adjustments of subglottal impedance parameters and mechanical properties of the skin. Preliminary results for sustained vowels with various voice qualities show that the subglottal IBIF scheme yields comparable estimates with respect to current aerodynamics-based methods of clinical vocal assessment. A mean absolute error of less than 10% was observed for two glottal airflow measures -maximum flow declination rate and amplitude of the modulation component- that have been associated with the pathophysiology of some common voice disorders caused by faulty and/or abusive patterns of vocal behavior (i.e., vocal hyperfunction). The proposed method further advances the ambulatory assessment of vocal function based on the neck acceleration signal, that previously have been limited to the estimation of phonation duration, loudness, and pitch. Subglottal IBIF is also suitable for other ambulatory applications in speech communication, in which further evaluation is underway.
Matías Zanartu, Julio C. Ho, Daryush D. Mehta, Robert E. Hillman, George R. Wodicka
IEEE Trans. Speech Audio Process.3
2012 Duration of ambulatory monitoring needed to accurately estimate voice use
abstract
Voice use is considered to play a major role in the development of many voice disorders, and clinicians focus on evaluating and modifying how patients use their voices throughout the day. Some voice monitoring devices have used neck-mounted accelerometers to unobtrusively and confidentially track voice use–related measures, such as phonation time, fundamental frequency, and sound intensity. Guidelines for the clinical use of such monitoring devices have yet to be established. This is a preliminary investigation into establishing initial benchmarks for obtaining robust estimates of long-term average voice use that may be used to begin examining basic relationships between vocal loading and voice use–related pathology. As expected, adequate monitoring durations depend on the inherent variability of the parameter of interest, with much of the error decreased after 26 hours of monitoring. Investigations are currently under way to take advantage of a smartphone-based voice monitoring system that is designed to enhance device wearability and enable the derivation of new clinically relevant measures. Index Terms: ambulatory voice monitoring, voice use, voice disorders, accelerometer 1.
Daryush D. Mehta, Rebecca Woodbury Listfield, Harold A. Cheyne II, James T. Heaton, Shengran W. Feng, Matías Zanartu, Robert E. Hillman
INTERSPEECH1
2011 Joint source-filter modeling using flexible basis functions
abstract
Improving on recent work on joint source-filter analysis of speech waveforms, we explore improvements to an autoregressive model with exogenous inputs represented by flexible basis functions. Following a brief review of the maximum likelihood estimators of the model parameters, the Cramer-Rao bounds are derived to provide evidence for the challenging nature of estimating source and filter characteristics with overlapping spectra. Wavelet expansion of the exogenous inputs is employed, and the selection of an appropriate subset of wavelets is described as an online, signal-adaptive approach. Results from synthesized and real vowel analysis illustrate the promise of iterative wavelet shrinkage using soft and hard thresholding and an alternative regularization method.
Daryush D. Mehta, Daniel Rudoy, Patrick J. Wolfe
ICASSP1
2006 Pitch-scale modification using the modulated aspiration noise source
abstract
Spectral harmonic/noise component analysis of spoken vowels shows evidence of noise modulations with peaks in the estimated noise source component synchronous with both the open phase of the periodic source and with time instants of glottal closure. Inspired by this observation of natural modulations and of fullband energy in the aspiration noise source, we develop an alternate approach to high-quality pitchscale modification of continuous speech. Our strategy takes a dual processing approach, in which the harmonic and noise components of the speech signal are separately analyzed, modified, and re-synthesized. The periodic component is modified using standard modification techniques, and the noise component is handled by modifying characteristics of its source waveform. Since we have modeled an inherent coupling between the periodic and aspiration noise sources, the modification algorithm is designed to preserve the synchrony between temporal modulations of the two sources. The reconstructed modified signal is perceived in informal listening to be natural-sounding and typically reduces artifacts that occur in standard modification techniques. Index Terms: pitch modification, aspiration noise, modulated noise, breathiness, voice quality
Daryush D. Mehta, Thomas F. Quatieri
INTERSPEECH1