Brad H. Story

dblp:34/7156 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
0since 2021 · last 2020
0000-0002-6530-8781ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Audio and music processing · 100%
Artificial intelligence
1 paper
Speech recognition and synthesis · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing › speech analysis
vocal tract estimation
0.622020
Time-Varying Quasi-Closed-Phase Analysis for Accurate Formant Tracking in Speech Signals · IEEE ACM Trans. Audio Speech Lang. Process. 2020
Joint Source-Filter Optimization for Accurate Vocal Tract Estimation Using Differential Evolution · IEEE Trans. Speech Audio Process. 2013
Audio and music processing › speech analysis
formant tracking
0.412020
Time-Varying Quasi-Closed-Phase Analysis for Accurate Formant Tracking in Speech Signals · IEEE ACM Trans. Audio Speech Lang. Process. 2020
Audio and music processing
speech processing
0.412020
Time-Varying Quasi-Closed-Phase Analysis for Accurate Formant Tracking in Speech Signals · IEEE ACM Trans. Audio Speech Lang. Process. 2020
Audio and music processing › speech analysis
glottal inverse filtering
0.212014
Quasi Closed Phase Glottal Inverse Filtering Analysis With Weighted Linear Prediction · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Audio and music processing
speech production
0.212013
Joint Source-Filter Optimization for Accurate Vocal Tract Estimation Using Differential Evolution · IEEE Trans. Speech Audio Process. 2013
Natural language and speech › Speech recognition and synthesis › speech production
speech production modeling
0.112006
Vocal-tract modeling: fractional elongation of segment lengths in a waveguide model with half-sample delays · IEEE Trans. Speech Audio Process. 2006
Audio and music processing
speech synthesis
0.112014
Quasi Closed Phase Glottal Inverse Filtering Analysis With Weighted Linear Prediction · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Audio and music processing › speech coding
vocoder
0.112014
Quasi Closed Phase Glottal Inverse Filtering Analysis With Weighted Linear Prediction · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Audio and music processing
speech analysis
0.012013
Joint Source-Filter Optimization for Accurate Vocal Tract Estimation Using Differential Evolution · IEEE Trans. Speech Audio Process. 2013

Methods — techniques the papers use, named apart from their topics

time-varying linear prediction · 0.4quasi-closed-phase analysis · 0.4l1 optimization · 0.4weighted linear prediction · 0.2attenuated main excitation · 0.2liljencrants-fant model · 0.2differential evolution · 0.2auto-regressive model with exogenous input · 0.2spatial interpolation · 0.1fractional-delay filter · 0.1
YearPublicationVenuePosition
2020 Time-Varying Quasi-Closed-Phase Analysis for Accurate Formant Tracking in Speech Signals
abstract
In this paper, we propose a new method for the accurate estimation and tracking of formants in speech signals using time-varying quasi-closed-phase (TVQCP) analysis. Conventional formant tracking methods typically adopt a two-stage estimateand-track strategy wherein an initial set of formant candidates are estimated using short-time analysis (e.g., 10-50 ms), followed by a tracking stage based on dynamic programming or a linear state-space model. One of the main disadvantages of these approaches is that the tracking stage, however good it may be, cannot improve upon the formant estimation accuracy of the first stage. The proposed TVQCP method provides a single-stage formant tracking that combines the estimation and tracking stages into one. TVQCP analysis combines three approaches to improve formant estimation and tracking: (1) it uses temporally weighted quasi-closed-phase analysis to derive closed-phase estimates of the vocal tract with reduced interference from the excitation source, (2) it increases the residual sparsity by using the L1 optimization and (3) it uses time-varying linear prediction analysis overlong time windows (e.g., 100-200 ms) to impose a continuity constraint on the vocal tract model and hence on the formant trajectories. Formant tracking experiments with a wide variety of synthetic and natural speech signals show that the proposed TVQCP method performs better than conventional and popular formant tracking tools, such as Wavesurfer and Praat (based on dynamic programming), the KARMA algorithm (based on Kalman filtering), and DeepFormants (based on deep neural networks trained in a supervised manner). Matlab scripts for the proposed method can be found at:
Dhananjaya Gowda, Sudarsana Reddy Kadiri, Brad H. Story, Paavo Alku
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 OPENGLOT - An open environment for the evaluation of glottal inverse filtering
Paavo Alku, Tiina Murtola, Jarmo Malinen, Juha Kuortti, Brad H. Story, Manu Airaksinen, Mika Salmi, Erkki Vilkman, Ahmed Geneid
Speech Commun.5
2019 Estimation of the glottal source from coded telephone speech using deep neural networks
N. P. Narendra, Manu Airaksinen, Brad H. Story, Paavo Alku
Speech Commun.3
2018 Estimation of the glottal flow from speech pressure signals: Evaluation of three variants of iterative adaptive inverse filtering using computational physical modelling of voice production
abstract
The aim of this study is to comparatively review and evaluate three variants of the glottal inverse filtering algorithm based on iterative adaptive inverse filtering (IAIF): the Standard algorithm, and two recently proposed variants that use iterative optimal preemphasis (IOP) and a glottal flow model (GFM), respectively. To enable an objective evaluation, a computational physical model of voice production is used to generate time-domain signals pertaining to both the input glottal flow and the output speech pressure, for a wide range of vowels, fundamental frequencies, and voice qualities (involving co-variation of phonation type and loudness). Furthermore, for a fair comparison, the three key parameters of IAIF are selected by an exhaustive search to minimize the root-mean-square error between the estimated and reference glottal flow derivative in each analyzed frame and performance is assessed with two time-domain and two frequency-domain error measures. A conventional evaluation is also carried out with fixed parameter values determined by cross-validation. Results indicate that IOP tends to yield the lowest errors for nonback vowels (reducing errors by 31% on average compared with Standard), especially for not too high fundamental frequencies and not too pressed voice qualities; GFM becomes competitive for normal phonations when fixed parameter values are used; and in other cases, Standard IAIF is still recommended. In addition, the results suggest that not only the overall spectral tilt (as controlled by IOP and GFM) but also the balance between the levels of different spectral regions, can be important for accurate estimation of the glottal flow.
Parham Mokhtari, Brad H. Story, Paavo Alku, Hiroshi Ando
Speech Commun.2
2017 An acoustically-driven vocal tract model for stop consonant production
Brad H. Story, Kate Bunton
Speech Commun.1
2016 Formant measurement in children's speech based on spectral filtering
Brad H. Story, Kate Bunton
Speech Commun.1
2014 Automatic glottal inverse filtering with the Markov chain Monte Carlo method
Harri Auvinen, Tuomo Raitio, Manu Airaksinen, Samuli Siltanen, Brad H. Story, Paavo Alku
Comput. Speech Lang.5
2014 Quasi Closed Phase Glottal Inverse Filtering Analysis With Weighted Linear Prediction
abstract
This study presents a new glottal inverse filtering (GIF) technique based on closed phase analysis over multiple fundamental periods. The proposed quasi closed phase (QCP) analysis method utilizes weighted linear prediction (WLP) with a specific attenuated main excitation (AME) weight function that attenuates the contribution of the glottal source in the linear prediction model optimization. This enables the use of the autocorrelation criterion in linear prediction in contrast to the covariance criterion used in conventional closed phase analysis. The QCP method was compared to previously developed methods by using synthetic vowels produced with the conventional source-filter model as well as with a physical modeling approach. The obtained objective measures show that the QCP method improves the GIF performance in terms of errors in typical glottal source parametrizations for both low- and high-pitched vowels. Additionally, QCP was tested in a physiologically oriented vocoder, where the analysis/synthesis quality was evaluated with a subjective listening test indicating improved perceived quality for normal speaking style.
Manu Airaksinen, Tuomo Raitio, Brad H. Story, Paavo Alku
IEEE ACM Trans. Audio Speech Lang. Process.3
2013 Quasi closed phase analysis for glottal inverse filtering
Manu Airaksinen, Brad H. Story, Paavo Alku
INTERSPEECH2
2013 Phrase-level speech simulation with an airway modulation model of speech production
Brad H. Story
Comput. Speech Lang.1
2013 Joint Source-Filter Optimization for Accurate Vocal Tract Estimation Using Differential Evolution
abstract
In this work, we present a joint source-filter optimization approach for separating voiced speech into vocal tract (VT) and voice source components. The presented method is pitch-synchronous and thereby exhibits a high robustness against vocal jitter, shimmer and other glottal variations while covering various voice qualities. The voice source is modeled using the Liljencrants-Fant (LF) model, which is integrated into a time-varying auto-regressive speech production model with exogenous input (ARX). The non-convex optimization problem of finding the optimal model parameters is addressed by a heuristic, evolutionary optimization method called differential evolution. The optimization method is first validated in a series of experiments with synthetic speech. Estimated glottal source and VT parameters are the criteria used for comparison with the iterative adaptive inverse filter (IAIF) method and the linear prediction (LP) method under varying conditions such as jitter, fundamental frequency (f0) as well as environmental and glottal noise. The results show that the proposed method largely reduces the bias and standard deviation of estimated VT coefficients and glottal source parameters. Furthermore, the performance of the source-filter separation is evaluated in experiments using speech generated with a physical model of speech production. The proposed method reliably estimates glottal flow waveforms and lower formant frequencies. Results obtained for higher formant frequencies indicate that research on more accurate voice source models and their interaction with the VT is necessary to improve the source-filter separation. The proposed optimization approach promises to be a useful tool for future research addressing this topic.
Olaf Schleusing, Tomi Kinnunen, Brad H. Story, Jean-Marc Vesin
IEEE Trans. Speech Audio Process.3
2012 Improved formant frequency estimation from high-pitched vowels by downgrading the contribution of the glottal source with weighted linear prediction
abstract
Since performance of conventional linear prediction (LP) deteriorates in formant estimation of high-pitched voices, several all-pole modeling methods robust to F0 have been developed. This study compares five such previously known methods and proposes a new technique, Weighted Linear Prediction with Attenuated Main Excitation (WLP-AME). WLP-AME utilizes weighted linear prediction in which the square of the prediction error is multiplied with a weighting function that downgrades the contribution of the glottal source in the model optimization. Consequently, the resulting all-pole model is affected more by the vocal tract characteristics, which leads to more accurate formant estimates. By using synthetic vowels created with a physical modeling approach, the study shows that WLP-AME yields improved formant frequency estimates for high-pitched vowels in comparison to the previously known methods.
Paavo Alku, Jouni Pohjalainen, Martti Vainio, Anne-Maria Laukkanen, Brad H. Story
INTERSPEECH5
2006 Vocal-tract modeling: fractional elongation of segment lengths in a waveguide model with half-sample delays
abstract
Digital waveguide models are commonly used for simulating vocal-tract acoustics based on physiological data. In particular, waveguide models with half-sample delays are known to be well suited for speech production research. This paper presents enhancements to such a model, aimed at improved accuracy in mapping physiological vocal-tract data (shape and length of the airway) to waveguide parameters. The enhancements allow the length of the vocal tract to be continuously varied, thus enabling more realistic synthesis. This is achieved by smoothly varying the individual segment lengths of a piecewise-cylindrical representation of the airway, without altering the system sampling frequency. Fractional-delay filters are used for spatial interpolation of the digital waveguide model. The algorithms are validated by modeling the protrusion of lips, lowering of larynx and lengthening of intermediate segments for a static vowel shape
Siddharth Mathur, Brad H. Story
IEEE Trans. Speech Audio Process.2
2004 Evaluation of an inverse filtering technique using physical modeling of voice production
abstract
Glottal flows and sound pressure waveforms of four different fundamental frequencies were generated using a computational model of vocal fold vibration and acoustic wave propagation in order to evaluate the performance of an inverse filtering method. Four time-based parameters of the glottal flow were used in order to assess the accuracy of the inverse filtering technique. The results show that for most of the cases analyzed the relative error was less than 5 % when the time-based parameters extracted from the estimated glottal flows were compared to those obtained from the original flow waveforms produced by the physical model of the vocal fold vibration.
Paavo Alku, Matti Airas, Brad H. Story
INTERSPEECH3
1997 Considerations in voice transformation with physiologic scaling principles
Ingo R. Titze, Darrell Wong, Brad H. Story, Russell Long
Speech Commun.3