Simon Stone

dblp:194/1224 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
8since 2021 · last 2024
0000-0002-1739-4953ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 5 since 2021
YearPublicationVenuePosition
2024 Monophthong vocal tract shapes are sufficient for articulatory synthesis of German primary diphthongs
abstract
German primary diphthongs are conventionally transcribed using the same symbols used for some monophthong vowels. However, if the corresponding vocal tract shapes are used for articulatory synthesis, the results often sound unnatural. Furthermore, there is no clear consensus in the literature if diphthongs have monopthong constituents and if so, which ones. This study therefore analyzed a set of audio recordings from the reference speaker of the state-of-the-art articulatory synthesizer VocalTractLab to identify likely candidates for the monophthong constituents of the German primary diphthongs. We then evaluated these candidates in a listening experiment with naive listeners to determine a naturalness ranking of these candidates and specialized diphthong shapes. The results showed that the German primary diphthongs can indeed be synthesized with no significant loss in naturalness by replacing the specialized diphthong shapes for the initial and final segments by shapes also used for monopthong vowels.
Simon Stone, Peter Birkholz
Speech Commun.1
2023 A Comparative Study of 3D and 1D Acoustic Simulations of the Higher Frequencies of Speech
abstract
Articulatory synthesis generates speech sounds by simulating the physical phenomena involved in speech production. The accuracy of the physical modelling is expected to affect the naturalness of the synthesis: the more realistic the description is, the greater the naturalness is expected to be. In this work, the accuracy of acoustic wave propagation in the vocal tract was evaluated with two perceptual experiments. Sustained vowels generated using a one-dimensional acoustic model, a three-dimensional acoustic model and an artificial bandwidth extension algorithm (without a physical basis) were compared. Since the difference between the acoustic methods tested affects mainly the frequencies above 4 kHz, we ensured that the low frequency part of the stimuli, up to 4 kHz, was similar. Thus, the participants' responses were based only on the differences at high frequency. The first experiment was a pair comparison, in which the participants had to select the more natural sounding stimuli. In the second experiment, the participants had to rate the naturalness of the stimuli on a linear scale. The results confirmed that a more accurate physical modeling leads to greater naturalness. However, this was limited to the phonemes /o/ and /u/, for which transverse resonances in the anterior vocal tract may play an important role that only a 3D acoustic simulation can accurately represent. It was also found that male stimuli were perceived as significantly more natural than female ones. However, voice quality did not affect naturalness.
Rémi Blandin, Simon Stone, Angélique Remacle, Vincent Didone, Peter Birkholz
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Carina - A Corpus of Aligned German Read Speech Including Annotations
abstract
This paper presents the semi-automatically created Corpus of Aligned Read Speech Including Annotations (CARInA), a speech corpus based on the German Spoken Wikipedia Corpus (GSWC). CARInA tokenizes, consolidates and organizes the vast, but rather unstructured material contained in GSWC. The contents are grouped by annotation completeness, and extended by canonic, morphosyntactic and prosodic annotations. The annotations are provided in BPF and TextGrid format. It contains 194 hours of speech material from 327 speakers, of which 124 hours are fully phonetically aligned and 30 hours are fully aligned at all annotation levels. CARInA is freely available1, designed to grow and improve over time, and suitable for large-scale speech analyses or machine learning tasks as illustrated by two examples shown in this paper.
Hannes Kath, Simon Stone, Stefan Rapp, Peter Birkholz
ICASSP2
2022 Relationship between the acoustic time intervals and tongue movements of German diphthongs
Arne-Lukas Fietkau, Simon Stone, Peter Birkholz
INTERSPEECH2
2022 Glottal inverse filtering based on articulatory synthesis and deep learning
Ingo Langheinrich, Simon Stone, Peter Birkholz
INTERSPEECH2
2022 PyRCN: A toolbox for exploration and application of Reservoir Computing Networks
Peter Steiner, Azarakhsh Jalalvand, Simon Stone, Peter Birkholz
Eng. Appl. Artif. Intell.3
2022 Articulatory Synthesis of Vocalized /r/ Allophones in German
abstract
Articulatory synthesis relies on precise, parametric vocal tract shapes to generate natural-sounding speech. In German, a particular challenge is the accurate synthesis of the vocalic /r/ allophones following vowels or in syllable coda position. Using established phonetic conventions, no satisfying results could be achieved so far, implying a possible shortcoming of these existing conventions. This study therefore analyzed a large number of natural recordings of the sounds in question from a single speaker to find the optimal number of target vocal tract shapes. Applying clustering techniques, the manifold of vocalic /r/ allophones could be reduced to two prototypical [ɐ] variants previously undescribed in the literature. As shown by a listening experiment, which of these two sounds was preferred in which context depended not only on the respective context vowel’s tenseness, but also on its openness and acoustic distance to other context vowels ending in the same [ɐ] variant. This indicates that the two different allophones might serve as a contrastive cue to help differentiate between otherwise similar context vowels.
Simon Stone, Yingming Gao, Peter Birkholz
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 Articulatory Data Recorder: A Framework for Real-Time Articulatory Data Recording
Alexander Wilbrandt, Simon Stone, Peter Birkholz
Interspeech2
2020 Cross-Speaker Silent-Speech Command Word Recognition Using Electro-Optical Stomatography
abstract
Speech recognition based on articulatory movements instead of the acoustic signal is of growing interest in the community. In this work, we present the results of a study using a novel measurement technology called Electro-Optical Stomatography to capture speech movements and use the acquired data to recognize a number of command words. The performance of the recognition system was evaluated using two vocabularies (one with 30 and one with 10 words) and four speakers. The speaker-dependent results were up to the state-of-the-art with average word accuracies of 97% to 99.5%, while the speaker-independent results exceeded it with average word accuracies of approx. 56% to 62%.
Simon Stone, Peter Birkholz
ICASSP1
2020 Prediction of Voicing and the F0 Contour from Electromagnetic Articulography Data for Articulation-to-Speech Synthesis
abstract
Articulation-to-speech synthesis based solely on supraglottal articulation requires some sort of intonation control. This paper examines to what extent the f0contour of an utterance can be predicted from such supraglottal articulation data. To that end, three groups of machine learning models (support vector machines, kernel ridge regression and neural networks) were trained and evaluated on the mngu0speech corpus containing synchronous articulatory and audio data. The best voiced/unvoiced/silence classification rates were achieved by a deep neural network with two hidden layers: 85.8 % with no look-ahead (important for on-line applications) and 86 % with a look-ahead of 50 ms. The best f0prediction model without look-ahead scored a root-mean-square error (RMSE) (when compared to the original f0contours) of 10.4 Hz using a neural network with one hidden layer, while the best prediction with a look-ahead of 50 ms was attained by kernel ridge regression and an RMSE of 10.3 Hz. The predicted f0contours were also subjectively evaluated in a listening test by manipulating the f0of the original speech files using PRAAT. The results are consistent with the objective evaluation.
Simon Stone, Peter Birkholz
ICASSP1
2020 Feature Engineering and Stacked Echo State Networks for Musical Onset Detection
abstract
In music analysis, one of the most fundamental tasks is note onset detection - detecting the beginning of new note events. As the target function of onset detection is related to other tasks, such as beat tracking or tempo estimation, onset detection is the basis for such related tasks. Furthermore, it can help to improve Automatic Music Transcription (AMT). Typically, different approaches for onset detection follow a similar outline: An audio signal is transformed into an Onset Detection Function (ODF), which should have rather low values (i.e. close to zero) for most of the time but with pronounced peaks at onset times, which can then be extracted by applying peak picking algorithms on the ODF. In the recent years, several kinds of neural networks were used successfully to compute the ODF from feature vectors. Currently, Convolutional Neural Networks (CNNs) define the state of the art. In this paper, we build up on an alternative approach to obtain a ODF by Echo State Networks (ESNs), which have achieved comparable results to CNNs in several tasks, such as speech and image recognition. In contrast to the typical iterative training procedures of deep learning architectures, such as CNNs or networks consisting of Long-Short-Term Memory Cells (LSTMs), in ESNs only a very small part of the weights is easily trained in one shot using linear regression. By comparing the performance of several feature extraction methods, pre-processing steps and introducing a new way to stack ESNs, we expand our previous approach to achieve results that fall between a bidirectional LSTM network and a CNN with relative improvements of 1.8 % and -1.4 %, respectively. For the evaluation, we used exactly the same 8-fold cross validation setup as for the reference results.
Peter Steiner, Azarakhsh Jalalvand, Simon Stone, Peter Birkholz
ICPR3
2019 Perceptual Optimization of an Enhanced Geometric Vocal Fold Model for Articulatory Speech Synthesis
Peter Birkholz, Susanne Drechsel, Simon Stone
INTERSPEECH3
2019 Articulatory Copy Synthesis Based on a Genetic Algorithm
Yingming Gao, Simon Stone, Peter Birkholz
INTERSPEECH2
2018 Non-Invasive Silent Phoneme Recognition Using Microwave Signals
abstract
Besides the recognition of audible speech, there is currently an increasing interest in the recognition of silent speech, which has a range of novel applications. A major obstacle for a wide spread of silent-speech technology is the lack of measurement methods for speech movements that are convenient, non-invasive, portable, and robust at the same time. Therefore, as an alternative to established methods, we examined to what extent different phonemes can be discriminated from the electromagnetic transmission and reflection properties of the vocal tract. To this end, we attached two Vivaldi antennas on the cheek and below the chin of two subjects. While the subjects produced 25 phonemes in multiple phonetic contexts each, we measured the electromagnetic transmission spectra from one antenna to the other, and the reflection spectra for each antenna (radar), in a frequency band from 2-12 GHz. Two classification methods (k-nearest neighbors and linear discriminant analysis) were trained to predict the phoneme identity from the spectral data. With linear discriminant analysis, cross-validated phoneme recognition rates of 93% and 85% were achieved for the two subjects. Although these results are speaker- and session-dependent, they suggest that electromagnetic transmission and reflection measurements of the vocal tract have great potential for future silent-speech interfaces.
Peter Birkholz, Simon Stone, Klaus Wolf, Dirk Plettemeier
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Construction and Evaluation of a Parametric One-Dimensional Vocal Tract Model
abstract
Articulatory speech synthesis based on aero-acoustic simulations of the vocal tract is computationally expensive and, therefore, requires simple yet precise models. Modeling the onedimensional vocal tract area function directly instead of a higher dimensional vocal tract model is an efficient way to minimize the computational overhead of the simulations. In this paper, we propose a new parametric vocal tract model that is controlled by six points and capable of modeling a large variety of vocal tract shapes. We geometrically and perceptually evaluated the model on a set of 22 reference area functions corresponding to German vowels and consonants. The model was able to geometrically approximate the reference area functions with a minimum root-mean-square error of 0.302 cm2, a maximum error of 1.142 cm2, and a median error of 0.891 cm2. After optimizations, a perceptual evaluation of the synthesis using our model in combination with a state-of-the-art aero-acoustic simulation achieved a vowel recognition rate of 90.7% and a consonant recognition rate of 73.2%.
Simon Stone, Michael Marxen, Peter Birkholz
IEEE ACM Trans. Audio Speech Lang. Process.1
2017 The Sound of Deception - What Makes a Speaker Credible?
Anne Schröder, Simon Stone, Peter Birkholz
INTERSPEECH2
2017 A Time-Warping Pitch Tracking Algorithm Considering Fast f0 Changes
Simon Stone, Peter Steiner, Peter Birkholz
INTERSPEECH1
2016 Silent-Speech Command Word Recognition Using Electro-Optical Stomatography
Simon Stone, Peter Birkholz
INTERSPEECH1