EDBT 2026 Demo / reviewers in the wild / expert
Brad H. Story
dblp:34/7156
· DBLP profile ↗
15ranked-venue papers
3as first author
0since 2021 · last 2020
0000-0002-6530-8781ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
3 papers |
Audio and music processing · 100% | |
| Artificial intelligence
1 paper |
Speech recognition and synthesis · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Audio and music processing › speech analysis
vocal tract estimation |
0.6 | 2 | 2020 | Time-Varying Quasi-Closed-Phase Analysis for Accurate Formant Tracking in Speech Signals · IEEE ACM Trans. Audio Speech Lang. Process. 2020 Joint Source-Filter Optimization for Accurate Vocal Tract Estimation Using Differential Evolution · IEEE Trans. Speech Audio Process. 2013 |
Audio and music processing › speech analysis
formant tracking |
0.4 | 1 | 2020 | Time-Varying Quasi-Closed-Phase Analysis for Accurate Formant Tracking in Speech Signals · IEEE ACM Trans. Audio Speech Lang. Process. 2020 |
Audio and music processing
speech processing |
0.4 | 1 | 2020 | Time-Varying Quasi-Closed-Phase Analysis for Accurate Formant Tracking in Speech Signals · IEEE ACM Trans. Audio Speech Lang. Process. 2020 |
Audio and music processing › speech analysis
glottal inverse filtering |
0.2 | 1 | 2014 | Quasi Closed Phase Glottal Inverse Filtering Analysis With Weighted Linear Prediction · IEEE ACM Trans. Audio Speech Lang. Process. 2014 |
Audio and music processing
speech production |
0.2 | 1 | 2013 | Joint Source-Filter Optimization for Accurate Vocal Tract Estimation Using Differential Evolution · IEEE Trans. Speech Audio Process. 2013 |
Natural language and speech › Speech recognition and synthesis › speech production
speech production modeling |
0.1 | 1 | 2006 | Vocal-tract modeling: fractional elongation of segment lengths in a waveguide model with half-sample delays · IEEE Trans. Speech Audio Process. 2006 |
Audio and music processing
speech synthesis |
0.1 | 1 | 2014 | Quasi Closed Phase Glottal Inverse Filtering Analysis With Weighted Linear Prediction · IEEE ACM Trans. Audio Speech Lang. Process. 2014 |
Audio and music processing › speech coding
vocoder |
0.1 | 1 | 2014 | Quasi Closed Phase Glottal Inverse Filtering Analysis With Weighted Linear Prediction · IEEE ACM Trans. Audio Speech Lang. Process. 2014 |
Audio and music processing
speech analysis |
0.0 | 1 | 2013 | Joint Source-Filter Optimization for Accurate Vocal Tract Estimation Using Differential Evolution · IEEE Trans. Speech Audio Process. 2013 |
Methods — techniques the papers use, named apart from their topics
time-varying linear prediction · 0.4quasi-closed-phase analysis · 0.4l1 optimization · 0.4weighted linear prediction · 0.2attenuated main excitation · 0.2liljencrants-fant model · 0.2differential evolution · 0.2auto-regressive model with exogenous input · 0.2spatial interpolation · 0.1fractional-delay filter · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Time-Varying Quasi-Closed-Phase Analysis for Accurate Formant Tracking in Speech SignalsabstractIn this paper, we propose a new method for the accurate estimation and tracking of formants in speech signals using time-varying quasi-closed-phase (TVQCP) analysis. Conventional formant tracking methods typically adopt a two-stage estimateand-track strategy wherein an initial set of formant candidates are estimated using short-time analysis (e.g., 10-50 ms), followed by a tracking stage based on dynamic programming or a linear state-space model. One of the main disadvantages of these approaches is that the tracking stage, however good it may be, cannot improve upon the formant estimation accuracy of the first stage. The proposed TVQCP method provides a single-stage formant tracking that combines the estimation and tracking stages into one. TVQCP analysis combines three approaches to improve formant estimation and tracking: (1) it uses temporally weighted quasi-closed-phase analysis to derive closed-phase estimates of the vocal tract with reduced interference from the excitation source, (2) it increases the residual sparsity by using the L1 optimization and (3) it uses time-varying linear prediction analysis overlong time windows (e.g., 100-200 ms) to impose a continuity constraint on the vocal tract model and hence on the formant trajectories. Formant tracking experiments with a wide variety of synthetic and natural speech signals show that the proposed TVQCP method performs better than conventional and popular formant tracking tools, such as Wavesurfer and Praat (based on dynamic programming), the KARMA algorithm (based on Kalman filtering), and DeepFormants (based on deep neural networks trained in a supervised manner). Matlab scripts for the proposed method can be found at: Dhananjaya Gowda, Sudarsana Reddy Kadiri, Brad H. Story, Paavo Alku |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | OPENGLOT - An open environment for the evaluation of glottal inverse filtering
Paavo Alku, Tiina Murtola, Jarmo Malinen, Juha Kuortti, Brad H. Story, Manu Airaksinen, Mika Salmi, Erkki Vilkman, Ahmed Geneid |
Speech Commun. | 5 |
| 2019 | Estimation of the glottal source from coded telephone speech using deep neural networks
N. P. Narendra, Manu Airaksinen, Brad H. Story, Paavo Alku |
Speech Commun. | 3 |
| 2018 | Estimation of the glottal flow from speech pressure signals: Evaluation of three variants of iterative adaptive inverse filtering using computational physical modelling of voice productionabstractThe aim of this study is to comparatively review and evaluate three variants of the glottal inverse filtering algorithm based on iterative adaptive inverse filtering (IAIF): the Standard algorithm, and two recently proposed variants that use iterative optimal preemphasis (IOP) and a glottal flow model (GFM), respectively. To enable an objective evaluation, a computational physical model of voice production is used to generate time-domain signals pertaining to both the input glottal flow and the output speech pressure, for a wide range of vowels, fundamental frequencies, and voice qualities (involving co-variation of phonation type and loudness). Furthermore, for a fair comparison, the three key parameters of IAIF are selected by an exhaustive search to minimize the root-mean-square error between the estimated and reference glottal flow derivative in each analyzed frame and performance is assessed with two time-domain and two frequency-domain error measures. A conventional evaluation is also carried out with fixed parameter values determined by cross-validation. Results indicate that IOP tends to yield the lowest errors for nonback vowels (reducing errors by 31% on average compared with Standard), especially for not too high fundamental frequencies and not too pressed voice qualities; GFM becomes competitive for normal phonations when fixed parameter values are used; and in other cases, Standard IAIF is still recommended. In addition, the results suggest that not only the overall spectral tilt (as controlled by IOP and GFM) but also the balance between the levels of different spectral regions, can be important for accurate estimation of the glottal flow. Parham Mokhtari, Brad H. Story, Paavo Alku, Hiroshi Ando |
Speech Commun. | 2 |
| 2017 | An acoustically-driven vocal tract model for stop consonant production
Brad H. Story, Kate Bunton |
Speech Commun. | 1 |
| 2016 | Formant measurement in children's speech based on spectral filtering
Brad H. Story, Kate Bunton |
Speech Commun. | 1 |
| 2014 | Automatic glottal inverse filtering with the Markov chain Monte Carlo method
Harri Auvinen, Tuomo Raitio, Manu Airaksinen, Samuli Siltanen, Brad H. Story, Paavo Alku |
Comput. Speech Lang. | 5 |
| 2014 | Quasi Closed Phase Glottal Inverse Filtering Analysis With Weighted Linear PredictionabstractThis study presents a new glottal inverse filtering (GIF) technique based on closed phase analysis over multiple fundamental periods. The proposed quasi closed phase (QCP) analysis method utilizes weighted linear prediction (WLP) with a specific attenuated main excitation (AME) weight function that attenuates the contribution of the glottal source in the linear prediction model optimization. This enables the use of the autocorrelation criterion in linear prediction in contrast to the covariance criterion used in conventional closed phase analysis. The QCP method was compared to previously developed methods by using synthetic vowels produced with the conventional source-filter model as well as with a physical modeling approach. The obtained objective measures show that the QCP method improves the GIF performance in terms of errors in typical glottal source parametrizations for both low- and high-pitched vowels. Additionally, QCP was tested in a physiologically oriented vocoder, where the analysis/synthesis quality was evaluated with a subjective listening test indicating improved perceived quality for normal speaking style. Manu Airaksinen, Tuomo Raitio, Brad H. Story, Paavo Alku |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2013 | Quasi closed phase analysis for glottal inverse filtering
Manu Airaksinen, Brad H. Story, Paavo Alku |
INTERSPEECH | 2 |
| 2013 | Phrase-level speech simulation with an airway modulation model of speech production
Brad H. Story |
Comput. Speech Lang. | 1 |
| 2013 | Joint Source-Filter Optimization for Accurate Vocal Tract Estimation Using Differential EvolutionabstractIn this work, we present a joint source-filter optimization approach for separating voiced speech into vocal tract (VT) and voice source components. The presented method is pitch-synchronous and thereby exhibits a high robustness against vocal jitter, shimmer and other glottal variations while covering various voice qualities. The voice source is modeled using the Liljencrants-Fant (LF) model, which is integrated into a time-varying auto-regressive speech production model with exogenous input (ARX). The non-convex optimization problem of finding the optimal model parameters is addressed by a heuristic, evolutionary optimization method called differential evolution. The optimization method is first validated in a series of experiments with synthetic speech. Estimated glottal source and VT parameters are the criteria used for comparison with the iterative adaptive inverse filter (IAIF) method and the linear prediction (LP) method under varying conditions such as jitter, fundamental frequency (f0) as well as environmental and glottal noise. The results show that the proposed method largely reduces the bias and standard deviation of estimated VT coefficients and glottal source parameters. Furthermore, the performance of the source-filter separation is evaluated in experiments using speech generated with a physical model of speech production. The proposed method reliably estimates glottal flow waveforms and lower formant frequencies. Results obtained for higher formant frequencies indicate that research on more accurate voice source models and their interaction with the VT is necessary to improve the source-filter separation. The proposed optimization approach promises to be a useful tool for future research addressing this topic. Olaf Schleusing, Tomi Kinnunen, Brad H. Story, Jean-Marc Vesin |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | Improved formant frequency estimation from high-pitched vowels by downgrading the contribution of the glottal source with weighted linear predictionabstractSince performance of conventional linear prediction (LP) deteriorates in formant estimation of high-pitched voices, several all-pole modeling methods robust to F0 have been developed. This study compares five such previously known methods and proposes a new technique, Weighted Linear Prediction with Attenuated Main Excitation (WLP-AME). WLP-AME utilizes weighted linear prediction in which the square of the prediction error is multiplied with a weighting function that downgrades the contribution of the glottal source in the model optimization. Consequently, the resulting all-pole model is affected more by the vocal tract characteristics, which leads to more accurate formant estimates. By using synthetic vowels created with a physical modeling approach, the study shows that WLP-AME yields improved formant frequency estimates for high-pitched vowels in comparison to the previously known methods. Paavo Alku, Jouni Pohjalainen, Martti Vainio, Anne-Maria Laukkanen, Brad H. Story |
INTERSPEECH | 5 |
| 2006 | Vocal-tract modeling: fractional elongation of segment lengths in a waveguide model with half-sample delaysabstractDigital waveguide models are commonly used for simulating vocal-tract acoustics based on physiological data. In particular, waveguide models with half-sample delays are known to be well suited for speech production research. This paper presents enhancements to such a model, aimed at improved accuracy in mapping physiological vocal-tract data (shape and length of the airway) to waveguide parameters. The enhancements allow the length of the vocal tract to be continuously varied, thus enabling more realistic synthesis. This is achieved by smoothly varying the individual segment lengths of a piecewise-cylindrical representation of the airway, without altering the system sampling frequency. Fractional-delay filters are used for spatial interpolation of the digital waveguide model. The algorithms are validated by modeling the protrusion of lips, lowering of larynx and lengthening of intermediate segments for a static vowel shape Siddharth Mathur, Brad H. Story |
IEEE Trans. Speech Audio Process. | 2 |
| 2004 | Evaluation of an inverse filtering technique using physical modeling of voice productionabstractGlottal flows and sound pressure waveforms of four different fundamental frequencies were generated using a computational model of vocal fold vibration and acoustic wave propagation in order to evaluate the performance of an inverse filtering method. Four time-based parameters of the glottal flow were used in order to assess the accuracy of the inverse filtering technique. The results show that for most of the cases analyzed the relative error was less than 5 % when the time-based parameters extracted from the estimated glottal flows were compared to those obtained from the original flow waveforms produced by the physical model of the vocal fold vibration. Paavo Alku, Matti Airas, Brad H. Story |
INTERSPEECH | 3 |
| 1997 | Considerations in voice transformation with physiologic scaling principles
Ingo R. Titze, Darrell Wong, Brad H. Story, Russell Long |
Speech Commun. | 3 |