Juraj Simko

dblp:77/9238 · DBLP profile ↗
← Back
24ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 22 · 4 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Enabling the replicability of speech synthesis perceptual evaluations
abstract
How speech synthesis is evaluated is nowadays questioned. Not only have conventional listening tests as a whole been proven a poor match for modern synthesis, but more fundamentally, important information (e.g., the question asked to the listener) is frequently missing in the report of the outcome of the evaluation despite the impact on the interpretation of the test results. This can lead to uncertainty about the validity of these evaluations. To address this issue, we propose standardising the structure of any evaluation report. To facilitate this standardisation, our contribution is twofold: an open-source subjective evaluation platform; and a set of reporting guidelines. The platform is designed to enable the development of easily shareable evaluation recipes. The set of guidelines complements the platform to support researchers in reporting their evaluation choices and analysis in more detail while relying on the recipe to describe the actual evaluation process.
Sébastien Le Maguer, Gwénolé Lecorvé, Damien Lolive, Naomi Harte, Juraj Simko
INTERSPEECH5
2025 Self-supervised Optimality-Guided Learning of Speech Articulation
abstract
This paper introduces a novel approach for modeling speech articulatory planning based on Optimal Control Theory. The presented approach uses an internal feed-forward controller model that learns to predict optimal articulatory commands minimizing a context-dependent objective function. This objective function combines conflicting tasks of minimizing articulatory effort and maximizing the recognition probability of a target vowel based on acoustic characteristics. We present a self-supervised optimality-guided architecture for training the feedforward internal model that directly uses the objective function as a training loss. Simulations involving isolated vowels of American-English show that online training of the internal model enables feedforward estimation of near-optimal articulatory parameters.
Juraj Simko, Benjamin Elie, Alice Turk
INTERSPEECH1
2024 A data-driven model of acoustic speech intelligibility for optimization-based models of speech production
abstract
This paper presents a data-driven model of intelligibility which is intended to be used in an optimization-based model of speech production.The BiLSTM-based model is trained as a phoneme classifier and takes a sequence of real articulatory trajectories as input and returns the probability of phonemes over time.The optimization minimizes a cost function which is the weighted sum of the conflicting demands of being intelligible and least articulatory effort.The data-driven intelligibility model presented in this paper is used to compute the intelligibility score.Simulations support Lindblom's hypo-and hyper-articulation theory of speech, as the degree of hyper-articulation of speech can be modified and tuned along a continuum by balancing the importance given to both requirements of intelligibility and least articulatory effort.
Benjamin Elie, Juraj Simko, Alice Turk
INTERSPEECH2
2024 Optimization-based planning of speech articulation using general Tau Theory
abstract
This paper presents a model of speech articulation planning and generation based on General Tau Theory and Optimal Control Theory. Because General Tau Theory assumes that articulatory targets are always reached, the model accounts for speech variation via context-dependent articulatory targets. Targets are chosen via the optimization of a composite objective function. This function models three different task requirements: maximal intelligibility, minimal articulatory effort and minimal utterance duration. The paper shows that systematic phonetic variability can be reproduced by adjusting the weights assigned to each task requirement. Weights can be adjusted globally to simulate different speech styles, and can be adjusted locally to simulate different levels of prosodic prominence. The solution of the optimization procedure contains Tau equation parameter values for each articulatory movement, namely position of the articulator at the movement offset, movement duration, and a parameter which relates to the shape of the movement’s velocity profile. The paper presents simulations which illustrate the ability of the model to predict or reproduce several well-known characteristics of speech. These phenomena include close-to-symmetric velocity profiles for articulatory movement, variation related to speech rate, centralization of unstressed vowels, lengthening of stressed vowels, lenition of unstressed lingual stop consonants, and coarticulation of stop consonants.
Benjamin Elie, Juraj Simko, Alice Turk
Speech Commun.2
2023 Optimal control of speech with context-dependent articulatory targets
abstract
This paper presents a computational implementation of phonetic planning which consists of choosing the position of articulatory targets which satisfy conflicting linguistic and extra-linguistic requirements. We present a minimal model that considers intelligibility and least effort as task requirements. To achieve the context-dependent variability of targets, our model approximates intelligibility as a function of target phoneme recognition probability given a vector of articulatory parameters. Preliminary experiments show that our minimal computational model of phonetic planning is able to predict two types of hypoarticulation by adjusting the weight assigned to effort: vowel centralization and stop consonant lenition.
Benjamin Elie, Juraj Simko, Alice Turk
INTERSPEECH2
2019 Prosodic Representations of Prominence Classification Neural Networks and Autoencoders Using Bottleneck Features
abstract
Prominence perception has been known to correlate with a complex interplay of the acoustic features of energy, fundamental frequency, spectral tilt, and duration. The contribution and importance of each of these features in distinguishing between prominent and non-prominent units in speech is not always easy to determine, and more so, the prosodic representations that humans and automatic classifiers learn have been difficult to interpret. This work focuses on examining the acoustic prosodic representations that binary prominence classification neural networks and autoencoders learn for prominence. We investigate the complex features learned at different layers of the network as well as the 10-dimensional bottleneck features (BNFs), for the standard acoustic prosodic correlates of prominence separately and in combination. We analyze and visualize the BNFs obtained from the prominence classification neural networks as well as their network activations. The experiments are conducted on a corpus of Dutch continuous speech with manually annotated prominence labels. Our results show that the prosodic representations obtained from the BNFs and higher-dimensional non-BNFs provide good separation of the two prominence categories, with, however, different partitioning of the BNF space for the distinct features, and the best overall separation obtained for F0.
Sofoklis Kakouros, Antti Suni, Juraj Simko, Martti Vainio
INTERSPEECH3
2019 Comparative Analysis of Prosodic Characteristics Using WaveNet Embeddings
abstract
Peer reviewed
Antti Suni, Marcin Wlodarczak, Martti Vainio, Juraj Simko
INTERSPEECH4
2018 Articulatory Consequences of Vocal Effort Elicitation Method
abstract
Articulatory features from two datasets, Slovak and Swedish, were compared to see whether different methods of eliciting loud speech (ambient noise vs visually presented loudness target) result in different articulatory behavior. The features studied were temporal and kinematic characteristics of lip separation within the closing and opening gestures of bilabial consonants, and of the tongue body movement from /i/ to /a/ through a bilabial consonant. The results indicate larger hyper- articulation in the speech elicited with visually presented target. While individual articulatory strategies are evident, the speaker groups agree on increasing the kinematic features equally within each gesture in response to the increased vocal effort. Another concerted strategy is keeping the tongue response at a minimum, presumably to preserve acoustic prerequisites necessary for the adequate vowel identity. While the method of visually presented loudness target elicits larger span of vocal effort, the two elicitation methods achieve comparable consistency per loudness conditions.
Elísabet Eir Cortes, Marcin Wlodarczak, Juraj Simko
INTERSPEECH3
2018 Prominence-based Evaluation of L2 Prosody
abstract
Peer reviewed
Heini Kallio, Antti Suni, Päivi Virkkunen, Juraj Simko
INTERSPEECH4
2017 Creak as a Feature of Lexical Stress in Estonian
abstract
Peer reviewed
Kätlin Aare, Pärtel Lippus, Juraj Simko
INTERSPEECH3
2017 Kinematic Signatures of Prosody in Lombard Speech
abstract
Peer reviewed
Stefan Benus, Juraj Simko, Mona Lehtinen
INTERSPEECH2
2017 Comparing Languages Using Hierarchical Prosodic Analysis
abstract
Peer reviewed
Juraj Simko, Antti Suni, Katri Hiovain, Martti Vainio
INTERSPEECH1
2017 Hierarchical representation and estimation of prosody using continuous wavelet transform
Antti Suni, Juraj Simko, Daniel Aalto, Martti Vainio
Comput. Speech Lang.2
2016 Congruency Effect Between Articulation and Grasping in Native English Speakers
abstract
Previous studies have shown congruency effects between specific speech articulations and manual grasping actions. For example, uttering the syllable [kα] facilitates power grip responses in terms of reaction time and response accuracy. A similar association of the syllable [ti] with precision grip has also been observed. As these congruency effects have been to date shown only for Finnish native speakers, this study explored whether the congruency effects generalize to native speakers of another language. The original experiments were therefore replicated with English participants (N=16). Several previous findings were reproduced, namely the association of syllables [kα] and [ke] with power grip and of [ti] and [te] with precision grip. However, the association of vowels [α] and [i] with power and precision grip, respectively, previously found for Finnish participants, was not significant for English speakers. This difference could be related to ambiguities of English orthography and pronunciation variations. It is possible that for English speakers seeing a certain written vowel activates several different phonological representations associated with that letter. If the congruency effects are based on interactions between specific phonological representations and grasp actions, this ambiguity might lead to weakening of the effects in the manner demonstrated here.
Mikko Tiainen, Fatima M. Felisberti, Kaisa Tiippana, Martti Vainio, Juraj Simko, Jirí Lukavský, Lari Vainio
INTERSPEECH5
2015 F0 discontinuity as a marker of prosodic boundary strength in lombard speech
abstract
Prosodic boundary strength (PBS) refers to the degree of disjuncture between two chunks of speech.It is affected by both linguistic and para-linguistic communicative intentions playing thus an important role in both speech generation and recognition tasks.Among several PBS signals, we focus in this paper on pitch-related discontinuities in boundaries conveying linguistically meaningful contrasts produced in increasing levels of ambient noise.We compare several measures of local and global pitch reset and use classifiers in an effort to better understand the relationship between the degree of ambient noise and F0 marking of PBS.Our results include a positive effect of some noise on boundary classification, better performance of local than global reset features, and more systematic behavior of F0 falls compared to rises.
Stefan Benus, Uwe D. Reichel, Juraj Simko
INTERSPEECH3
2015 Polysyllabic shortening and word-final lengthening in English
abstract
Windmann A, Simko J, Wagner P. Polysyllabic Shortening and Word-Final Lengthening in English. In: Proceedings of Interspeech 2015. 2015: 36-40.
Andreas Windmann, Juraj Simko, Petra Wagner
INTERSPEECH2
2015 Optimization-based modeling of speech timing
Andreas Windmann, Juraj Simko, Petra Wagner
Speech Commun.2
2014 A unified account of prominence effects in an optimization-based model of speech timing
abstract
Windmann A, Simko J, Wagner P. A Unified Account of Prominence Effects in an Optimization-Based Model of Speech Timing. In: Proceedings of Interspeech 2014. 2014: 159-163.
Andreas Windmann, Juraj Simko, Petra Wagner
INTERSPEECH2
2013 Language background affects the strength of the pitch bias in a duration discrimination task
abstract
The fundamental frequency of a complex sound modulates the perceived duration of a sound. Higher pitch sounds are perceived longer compared to lower pitch sounds as shown by several independent studies since 1973. In this paper, the effect of language background is studied: native speakers of Finnish and German participated in a two alternative forced choice duration discrimination experiment where the duration and frequency of two sounds are randomly varied. The overall duration discrimination sensitivity was similar to both groups but the speakers of Finnish were influenced more by the pitch in their judgements. In addition, the difference in the two sounds’ pitch period explained the response data better than the difference in pitch frequencies or the pitch interval. As the Finnish quantity system is known to employ both duration and pitch cues, the present results suggest that the speakers are shaped by the language environment even when the task is purely non-linguistic.
Daniel Aalto, Juraj Simko, Martti Vainio
INTERSPEECH2
2013 Modeling durational incompressibility
abstract
We show how incompressibility, a well-described property of some prosodic timing effects, can be accounted for in an optimization-based model of speech timing.Preliminary results of a corpus study are presented, replicating and generalizing previous findings on incompressibility as a function of increasing speaking rate.We then introduce the architecture of our model and present results of simulation experiments that reproduce the results of the corpus analysis.Results suggest that incompressibility can be interpreted as a consequence of tradeoffs between competing requirements of production efficiency and communicative efficacy.
Andreas Windmann, Juraj Simko, Britta Wrede, Petra Wagner
INTERSPEECH2
2013 Pitch and duration as a basis for entrainment of overlapped speech onsets
abstract
The present paper reports on the impact of pitch accents and duration on temporal organisation of overlapping speech onsets in spontaneous dialogue.We observe a non-random pattern of overlap initiations within intervals between consecutive pitch accents, thus extending our earlier reports of a similar effect within vowel-to-vowel intervals.The latter finding was interpreted as a tendency to start overlapped speech directly before perceptually prominent vocalic onsets.In an attempt to reconcile these results, we investigate whether the effect observed on vowel-to-vowel intervals is influenced by presence of pitch accents and lengthening, both of which are known to be correlated with perceptual prominence.We find a strong effect of duration, which, however, does not on its own account fully for the observed pattern, indicating that other correlates of prominence might be involved in guiding the timing of overlap onsets.
Marcin Wlodarczak, Juraj Simko, Petra Wagner
INTERSPEECH2
2012 Temporal entrainment in overlapped speech: Cross-linguistic study
abstract
Wlodarczak M, Simko J, Wagner P. Temporal entrainment in overlapped speech: Cross-linguistic study. In: 13th Annual Conference of the International Speech Communication Association 2012 (INTERSPEECH 2012). Vol. 1. Red Hook, NY: Curran; 2013: 614-617.
Marcin Wlodarczak, Juraj Simko, Petra Wagner
INTERSPEECH2
2011 Investigating the Stability of Intergestural Timing Relations
abstract
Simko J, Cummins F, Benus S. Investigating the stability of intergestural timing relations. Presented at the Interspeech, Florence, Italy.
Juraj Simko, Fred Cummins, Stefan Benus
INTERSPEECH1
2009 Sequencing of articulatory gestures using cost optimization
abstract
Within the framework of articulatory phonology (AP), gestures function as primitives, and their ordering in time is provided by a gestural score. Determining how they should be sequenced in time has been something of a challenge. We modify the task dynamic implementation of AP, by defining tasks to be the desired positions of physically embodied end effectors. This allows us to investigate the optimal sequencing of gestures based on a parametric cost function. Costs evaluated include precision of articulation, articulatory effort, and gesture duration. We find that a simple optimization using these costs results in stable gestural sequences that reproduce several known coarticulatory effects.
Juraj Simko, Fred Cummins
INTERSPEECH1