VLDB 2026 Research / reviewers in the wild / expert
Alice Turk
dblp:90/9329
· DBLP profile ↗
10ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0002-9627-3040ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Self-supervised Optimality-Guided Learning of Speech ArticulationabstractThis paper introduces a novel approach for modeling speech articulatory planning based on Optimal Control Theory. The presented approach uses an internal feed-forward controller model that learns to predict optimal articulatory commands minimizing a context-dependent objective function. This objective function combines conflicting tasks of minimizing articulatory effort and maximizing the recognition probability of a target vowel based on acoustic characteristics. We present a self-supervised optimality-guided architecture for training the feedforward internal model that directly uses the objective function as a training loss. Simulations involving isolated vowels of American-English show that online training of the internal model enables feedforward estimation of near-optimal articulatory parameters. Juraj Simko, Benjamin Elie, Alice Turk |
INTERSPEECH | 3 |
| 2024 | A data-driven model of acoustic speech intelligibility for optimization-based models of speech productionabstractThis paper presents a data-driven model of intelligibility which is intended to be used in an optimization-based model of speech production.The BiLSTM-based model is trained as a phoneme classifier and takes a sequence of real articulatory trajectories as input and returns the probability of phonemes over time.The optimization minimizes a cost function which is the weighted sum of the conflicting demands of being intelligible and least articulatory effort.The data-driven intelligibility model presented in this paper is used to compute the intelligibility score.Simulations support Lindblom's hypo-and hyper-articulation theory of speech, as the degree of hyper-articulation of speech can be modified and tuned along a continuum by balancing the importance given to both requirements of intelligibility and least articulatory effort. Benjamin Elie, Juraj Simko, Alice Turk |
INTERSPEECH | 3 |
| 2024 | Optimization-based planning of speech articulation using general Tau TheoryabstractThis paper presents a model of speech articulation planning and generation based on General Tau Theory and Optimal Control Theory. Because General Tau Theory assumes that articulatory targets are always reached, the model accounts for speech variation via context-dependent articulatory targets. Targets are chosen via the optimization of a composite objective function. This function models three different task requirements: maximal intelligibility, minimal articulatory effort and minimal utterance duration. The paper shows that systematic phonetic variability can be reproduced by adjusting the weights assigned to each task requirement. Weights can be adjusted globally to simulate different speech styles, and can be adjusted locally to simulate different levels of prosodic prominence. The solution of the optimization procedure contains Tau equation parameter values for each articulatory movement, namely position of the articulator at the movement offset, movement duration, and a parameter which relates to the shape of the movement’s velocity profile. The paper presents simulations which illustrate the ability of the model to predict or reproduce several well-known characteristics of speech. These phenomena include close-to-symmetric velocity profiles for articulatory movement, variation related to speech rate, centralization of unstressed vowels, lengthening of stressed vowels, lenition of unstressed lingual stop consonants, and coarticulation of stop consonants. Benjamin Elie, Juraj Simko, Alice Turk |
Speech Commun. | 3 |
| 2023 | Optimal control of speech with context-dependent articulatory targetsabstractThis paper presents a computational implementation of phonetic planning which consists of choosing the position of articulatory targets which satisfy conflicting linguistic and extra-linguistic requirements. We present a minimal model that considers intelligibility and least effort as task requirements. To achieve the context-dependent variability of targets, our model approximates intelligibility as a function of target phoneme recognition probability given a vector of articulatory parameters. Preliminary experiments show that our minimal computational model of phonetic planning is able to predict two types of hypoarticulation by adjusting the weight assigned to effort: vowel centralization and stop consonant lenition. Benjamin Elie, Juraj Simko, Alice Turk |
INTERSPEECH | 3 |
| 2023 | Estimating virtual targets for lingual stop consonants using general Tau theoryabstractThis paper investigates the existence and position of virtual targets during the production of stop consonants. Using the equations from general Tau theory to model the time-course of tongue constriction formation movements, targets were estimated by fitting these equations on observed tongue constriction variables extracted from real EMA data from 2 native speakers of English. Results suggest that targets are virtual for 50 to 60% of movements. For these movements, virtual targets of the tongue tip constriction are predicted to occur around 0.1 cm beyond the palate, and virtual targets for the tongue dorsum constriction are predicted to occur between 0.05 and 0.2 cm beyond the palate. Our results suggest that the time-course of movement is planned so that the onset of closure occurs with relatively high velocity: closure onset is generally located very close in time to the time of peak velocity. Benjamin Elie, Alice Turk |
INTERSPEECH | 2 |
| 2023 | Modeling trajectories of human speech articulators using general Tau theoryabstractThis paper presents an application of general Tau theory to the modeling and analysis of articulatory trajectories in speech. We evaluated the model using electromagnetic articulometry data from 12 native speakers of English reading a common text, where trajectories of the following sensors were fitted: lower and upper lips, jaw, and three tongue sensors. Additionally, we analyzed trajectories of the lip aperture signal. Our experiments show that the general Tau theory model gives a better fit than existing (i) methods based on critically damped oscillators, and (ii) a method based on sequential target approximation. These findings support the hypothesis of Tau-guided movements of articulators during speech production. In the second part of the paper, our Tau theory analysis shows that articulatory movements follow similar velocity profile distributions across speakers. In particular, the value of the shape parameter κ of the Tau theory equation is identically distributed across speakers, following a unimodal distribution. The statistical mode of the distribution corresponds to the value of κ that generates a symmetric velocity profile. The analysis of the statistical distribution of κ values also reveals that its variance decreases when greater articulatory effort is required, such that produced articulatory effort remains close to that predicted by the theoretical minimal cost function based on forces acting on the moving articulator. This provides new evidence that articulatory effort is optimized during speech production. Benjamin Elie, David N. Lee, Alice Turk |
Speech Commun. | 3 |
| 2013 | The edinburgh speech production facility doubletalk corpus
James M. Scobbie, Alice Turk, Christian Geng, Simon King 0001, Robin J. Lickley, Korin Richmond |
INTERSPEECH | 2 |
| 2011 | Accelerometer-Based Respiratory Measurement During SpeechabstractAccelerometer-based respiratory monitoring is a recent area of research based on the observation of small rotations at the chest wall due to breathing. Previous studies of this technique have begun to address some sources of interference e.g. subject movements, but have not investigated operation during speech production when breathing patterns are known to be substantially different to normal respiration. We demonstrate measurement of speech breathing with a wireless tri-axial accelerometer in a synchronously captured dataset, including annotated audio and electro-magnetic articulograph data. We find agreement between peaks in the accelerometer-derived rotation signal and manually annotated breath timings, and correlation between peak rotations and the duration of audible in breaths. In speech breathing the rotation rate signal does not appear to be a good proxy for airflow rate as previously suggested, and instead seems to better reflect the role of specific muscles around the accelerometer location. We conclude that the method can be usable during speech breathing, but that this difference should be considered. The method has some advantages for speech breathing research due to its unobtrusive nature. Andrew Bates, Martin J. Ling, Christian Geng, Alice Turk, D. K. Arvind 0001 |
BSN | 4 |
| 1998 | Vowel quality in spontaneous speech: what makes a good vowel?abstractClear speech is characterised by longer segmental durations and less target undershoot [9] which results in more extreme spectral features.This paper deals with the clarity of vowels produced in spontaneous speech in a large corpus of task-oriented dialogues.We present an automatic technique for measuring vowel clarity on the basis of a vowel's spectral characteristics.This technique was evaluated using a perceptual test.Subjects rated the 'goodness' of vowels with different spectral characteristics with controlled duration and amplitude and these results were compared with an automatic rating.Results indicated that although agreement between subjects and the automatic measurement was poor it was as poor as the agreement between subjects.On the basis of these results we address the following questions:1. Can subjects reliably judge the clarity of vowels excerpted from spontaneous speech without duration cues? 2. Can a statistical model [3] reliably predict the subjects' response to such vowels? Matthew P. Aylett, Alice Turk |
ICSLP | 2 |
| 1997 | The domain of accentual lengthening in Scottish EnglishabstractThis study describes speech production experiments designed to determine the domain of accentual lengthening in Scottish English. Results suggest that accentual lengthening affects not only the syllable which bears the pitch accent (phrasal stress), but extends rightwards beyond this syllable. Secondly, the amount of lengthening on a syllable adjacent to a pitch accent appears to depend upon its membership in a pitch accented unit. Several candidates for the accentual-lengthening unit are entertained. 1. INTRODUCTION The experiments presented in this paper were designed to determine the domain of the durational effects of accent (phrasal stress). Phrasal stress is phonologically associated with a particular vowel, or stress-bearing unit, but its durational correlates may extend beyond the unit with which it is associated. Knowledge about how far durational effects extend is crucial for modelling durational effects in automatic speech recognition and synthesis. Furthermore, evidence fo... Alice Turk, Laurence White |
EUROSPEECH | 1 |