Yi Xu 0007

dblp:14/5580-7 · DBLP profile ↗
← Back
26ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0002-8541-2658ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 17 · 3 first-author · 7 since 2021
YearPublicationVenuePosition
2025 Learnability of English diphthongs: One dynamic target vs. two static targets
Anqi Xu 0002, Daniel R. van Niekerk, Branislav Gerazov, Paul Konstantin Krug, Santitham Prom-on, Peter Birkholz, Yi Xu 0007
Speech Commun.7
2024 Tone-syllable synchrony in Mandarin: New evidence and implications
abstract
Recent research has shown evidence based on a minimal contrast paradigm that consonants and vowels are articulatorily synchronized at the onset of the syllable. What remains less clear is the laryngeal dimension of the syllable, for which evidence of tone synchrony with the consonant-vowel syllable has been circumstantial. The present study assesses the precise tone-vowel alignment in Mandarin Chinese by applying the minimal contrast paradigm. The vowel onset is determined by detecting divergence points of F2 trajectories between a pair of disyllabic sequences with two contrasting vowels, and the onsets of tones are determined by detecting divergence points of f0 trajectories in contrasting disyllabic tone pairs, using generalized additive mixed models (GAMMs). The alignment of the divergence-determined vowel and tone onsets is then evaluated with linear mixed effect models (LMEMs) and their synchrony is validated with Bayes factors. The results indicate that tone and vowel onsets are fully synchronized. There is therefore evidence for strict alignment of consonant, vowel and tone as hypothesized in the synchronization model of the syllable. Also, with the newly established tone onset, the previously reported ‘anticipatory raising’ effect of tone now appears to occur within rather than before the articulatory syllable. Implications of these findings will be discussed.
Weiyi Kang, Yi Xu 0007
Speech Commun.2
2023 Self-Supervised Solution to the Control Problem of Articulatory Synthesis
abstract
Given an articulatory-to-acoustic forward model, it is a priori \nunknown how its motor control must be operated to achieve a \ndesired acoustic result. This control problem is a fundamental \nissue of articulatory speech synthesis and the cradle of acousticto-articulatory inversion, a discipline which attempts to address \nthe issue by the means of various methods. This work presents \nan end-to-end solution to the articulatory control problem, in \nwhich synthetic motor trajectories of Monte-Carlo-generated \nartificial speech are linked to input modalities (such as natural speech recordings or phoneme sequence input) via speakerindependent latent representations of a vector-quantized variational autoencoder. The proposed method is self-supervised and \nthus, in principle, synthesizer and speaker model independent.
Paul Konstantin Krug, Peter Birkholz, Branislav Gerazov, Daniel R. van Niekerk, Anqi Xu 0002, Yi Xu 0007
INTERSPEECH6
2023 Simulating vocal learning of spoken language: Beyond imitation
abstract
Computational approaches have an important role to play in understanding the complex process of speech acquisition, in general, and have recently been popular in studies of vocal learning in particular. In this article we suggest that two significant problems associated with imitative vocal learning of spoken language, the speaker normalisation and phonological correspondence problems, can be addressed by linguistically grounded auditory perception. In particular, we show how the articulation of consonant–vowel syllables may be learnt from auditory percepts that can represent either individual utterances by speakers with different vocal tract characteristics or ideal phonetic realisations. The result is an optimisation-based implementation of vocal exploration – incorporating semantic, auditory, and articulatory signals – that can serve as a basis for simulating vocal learning beyond imitation.
Daniel R. van Niekerk, Anqi Xu 0002, Branislav Gerazov, Paul Konstantin Krug, Peter Birkholz, Lorna F. Halliday, Santitham Prom-on, Yi Xu 0007
Speech Commun.8
2023 Artificial Vocal Learning Guided by Phoneme Recognition and Visual Information
abstract
This paper introduces a paradigm shift regarding vocal learning simulations, in which the communicative function of speech acquisition determines the learning process and intelligibility is considered the primary measure of learning success. Thereby, a novel approach for artificial vocal learning is presented that utilizes deep neural network-based phoneme recognition in order to calculate the speech acquisition objective function. This function guides a learning framework that involves the state-of-the-art articulatory speech synthesizer VocalTractLab as the motor-to-acoustic forward model. In this way, an extensive set of German phonemes, including most of the consonants and all stressed vowels, was produced successfully. The synthetic phonemes were rated as highly intelligible by human listeners. Furthermore, it is shown that visual speech information, such as lip and jaw movements, can be extracted from video recordings and be incorporated into the learning framework as an additional loss component during the optimization process. It was observed that this visual loss did not increase the overall intelligibility of phonemes. Instead, the visual loss acted as a regularization mechanism that facilitated the finding of more biologically plausible solutions in the articulatory domain.
Paul Konstantin Krug, Peter Birkholz, Branislav Gerazov, Daniel R. van Niekerk, Anqi Xu 0002, Yi Xu 0007
IEEE ACM Trans. Audio Speech Lang. Process.6
2022 Articulatory Synthesis for Data Augmentation in Phoneme Recognition
abstract
While numerous studies on automatic speech recognition have been published in recent years describing data augmentation strategies based on time or frequency domain signal processing, few works exist on the artificial extensions of training data sets using purely synthetic speech data.In this work, the German KIEL corpus was augmented with synthetic data generated with the state-of-the-art articulatory synthesizer VOCALTRACT-LAB.It is shown that the additional synthetic data can lead to a significantly better performance in single-phoneme recognition in certain cases, while at the same time, the performance can also decrease in other cases, depending on the degree of acoustic naturalness of the synthetic phonemes.As a result, this work can potentially guide future studies to improve the quality of articulatory synthesis via the link between synthetic speech production and automatic speech recognition.
Paul Konstantin Krug, Peter Birkholz, Branislav Gerazov, Daniel R. van Niekerk, Anqi Xu 0002, Yi Xu 0007
INTERSPEECH6
2022 Exploration strategies for articulatory synthesis of complex syllable onsets
abstract
High-quality articulatory speech synthesis has many potential applications in speech science and technology. However, developing appropriate mappings from linguistic specification to articulatory gestures is difficult and time consuming. In this paper we construct an optimisation-based framework as a first step towards learning these mappings without manual intervention. We demonstrate the production of CCV syllables and discuss the quality of the articulatory gestures with reference to coarticulation.
Daniel R. van Niekerk, Anqi Xu 0002, Branislav Gerazov, Paul Konstantin Krug, Peter Birkholz, Yi Xu 0007
INTERSPEECH6
2022 Evoc-Learn - High quality simulation of early vocal learning
Yi Xu 0007, Anqi Xu 0002, Daniel R. van Niekerk, Branislav Gerazov, Peter Birkholz, Paul Konstantin Krug, Santitham Prom-on, Lorna F. Halliday
INTERSPEECH1
2021 Segmental Alignment of English Syllables with Singleton and Cluster Onsets
abstract
Recent research has shown fresh evidence that consonant and vowel are synchronised at the syllable onset, as predicted by a number of theoretical models. The finding was made by using a minimal contrast paradigm to determine segment onset in Mandarin CV syllables, which differed from the conventional method of detecting gesture onset with a velocity threshold [1]. It has remained unclear, however, if CV co-onset also occurs between the nucleus vowel and a consonant cluster, as predicted by the articulatory syllable model [2]. This study applied the minimal contrast paradigm to British English in both CV and clusterV (CLV) syllables, and analysed the spectral patterns with signal chopping in conjunction with recurrent neural networks (RNN) with long short-term memory (LSTM) [3]. Results show that vowel onset is synchronised with the onset of the first consonant in a cluster, thus supporting the articulatory syllable model.
Zirui Liu 0003, Yi Xu 0007
Interspeech2
2021 Model-Based Exploration of Linking Between Vowel Articulatory Space and Acoustic Space
abstract
While the acoustic vowel space has been extensively studied in previous research, little is known about the high-dimensional articulatory space of vowels. The articulatory imaging techniques are limited to tracking only a few key articulators, leaving the rest of the articulators unmonitored. In the present study, we attempted to develop a detailed articulatory space obtained by training a 3D articulatory synthesizer to learn eleven British English vowels. An analysis-by-synthesis strategy was used to acoustically optimize vocal tract parameters that represent twenty articulatory dimensions. The results show that tongue height and retraction, larynx location and lip roundness are the most perceptually distinctive articulatory dimensions. Yet, even for these dimensions, there is a fair amount of articulatory overlap between vowels, unlike the fine-grained acoustic space. This method opens up the possibility of using modelling to investigate the link between speech production and perception.
Anqi Xu 0002, Daniel R. van Niekerk, Branislav Gerazov, Paul Konstantin Krug, Santitham Prom-on, Peter Birkholz, Yi Xu 0007
Interspeech7
2020 Coarticulation as Synchronised Sequential Target Approximation: An EMA Study
abstract
In this study we tested the hypothesis that consonant and vowel articulations start at the same time at syllable onset [1]. Articulatory data was collected for Mandarin Chinese using Electromagnetic Articulography (EMA), which tracks flesh-point movements in time and space. Unlike the traditional velocity threshold method [2], we used a triplet method based on the minimal pair paradigm [3] that detects divergence points between contrastive pairs of C or V respectively, before comparing their relative timing. Results show that articulatory onsets of consonant and vowel in CV syllables do not differ significantly from each other, which is consistent with the CV synchrony hypothesis. At the same time, the results also show some evidence that articulators that are shared by both C and V are engaged in sequential articulation, i.e., approaching the V target after approaching the C target.
Zirui Liu 0003, Yi Xu 0007, Feng-fan Hsieh
INTERSPEECH2
2020 Finding Intelligible Consonant-Vowel Sounds Using High-Quality Articulatory Synthesis
abstract
In this study, a state-of-the-art articulatory speech synthesiser was used as the basis for simulating the exploration of CV sounds imitating speech stimuli. By adopting a relevant kinematic model and systematically reducing the search space of consonant articulatory targets, intelligible CV sounds can be found. Derivative-free optimisation strategies were evaluated to speed up the process of exploring articulatory space and the possibility of using automatic speech recognition as a means of evaluating intelligibility was explored.
Daniel R. van Niekerk, Anqi Xu 0002, Branislav Gerazov, Paul Konstantin Krug, Peter Birkholz, Yi Xu 0007
INTERSPEECH6
2019 Prosodic encoding of focus in Hijazi Arabic
Muhammad Swaileh Alzaidi, Yi Xu 0007, Anqi Xu 0002
Speech Commun.2
2018 A Weighted Superposition of Functional Contours Model for Modelling Contextual Prominence of Elementary Prosodic Contours
abstract
International audience
Branislav Gerazov, Gérard Bailly, Yi Xu 0007
INTERSPEECH3
2017 Manipulation of the prosodic features of vocal tract length, nasality and articulatory precision using articulatory synthesis
Peter Birkholz, Lucia Martin, Yi Xu 0007, Stefan Scherbaum, Christiane Neuschaefer-Rube
Comput. Speech Lang.3
2014 Toward invariant functional representations of variable surface fundamental frequency contours: Synthesizing speech melody via model-based stochastic learning
Yi Xu 0007, Santitham Prom-on
Speech Commun.1
2013 Mora-based pre-low raising in Japanese pitch accent
abstract
This study is an attempt to understand the phonetic properties of pitch accent conditions in Japanese as related to the two observed versions of H tones. We tested the hypothesis that the higher version (accented H) results from pre-low raising (PLR) rather than being inherently higher. Correlation analysis reveals an inverse relation between accent peak and the following low tone, and that the strength of such correlations is affected by both peak-to-word-end distance (categorical effect) and within-mora time pressure (gradient), but the two effects work in opposite directions. We take this as evidence that the former effect is due to mora-level pre-planning while the latter is mechanical. These results suggest that in Japanese a low pitch target raises the preceding high target through anticipatory dissimilation. The findings of this study extend our previous understanding of the mechanisms of pitch production. Copyright © 2013 ISCA.
Yi Xu 0007, Santitham Prom-on
INTERSPEECH2
2013 Training an articulatory synthesizer with continuous acoustic data
abstract
This paper reports preliminary results of our effort to address the acoustic-to-articulatory inversion problem. We tested an approach that simulates speech production acquisition as a distal learning task, with acoustic signals of natural utterances in the form of MFCC as input, VocalTractLab — a 3D articulatory synthesizer controlled by target approximation models as the learner, and stochastic gradient descent as the training method. The approach was tested on a number of natural utterances, and the results were highly encouraging.
Santitham Prom-on, Peter Birkholz, Yi Xu 0007
INTERSPEECH3
2011 Simulating Post-L F0 Bouncing by Modeling Articulatory Dynamics
abstract
Post-L F 0 bouncing (post-L bouncing for short) is a prosodic phenomenon whereby F 0 is temporarily raised following a very low pitch. The phenomenon is quite robust, but is not widely known, and it has never been computationally modeled. This paper presents the results of our simulation of the phenomenon by modeling articulatory dynamics. Using the quantitative Target Approximation (qTA) model, we were able to simulate the F 0 rise after the Mandarin L tone by adding an acceleration adjustment to the initial state of the first post-L Neutral tone. Furthermore, a linear relationship was found between the added acceleration and the amount of F 0 lowering in the L tone. We interpreted the results as evidence that post-L bouncing is directly related to the articulatory mechanism of producing a very low pitch. Copyright © 2011 ISCA.
Santitham Prom-on, Yi Xu 0007, Fang Liu 0018
INTERSPEECH2
2010 Articulatory-functional modeling of speech prosody: a review
abstract
Natural prosody is produced by an articulatory system to convey communicative meanings. It is therefore desirable for prosody modeling to represent both articulatory mechanisms and communicative functions. There are doubts, however, as to whether such representation is necessary or beneficial if the aim of modeling is to just generate perceptually acceptable output. In this paper we briefly review models that have attempted to implement representations of either or both aspects of prosody. We show that, at least theoretically, it is beneficial to represent both articulatory mechanisms and communicative functions even if the goal is to just simulate surface prosody. Index Terms: speech prosody, modeling, PENTA, qTA 1.
Yi Xu 0007, Santitham Prom-on
INTERSPEECH1
2007 The neutral tone in question intonation in Mandarin
abstract
This study investigates how the neutral tone, when preceded by different full tones under different focus conditions, behaves in question intonation in Mandarin. Results indicate that 1) the preceding full/neutral tone largely determines the local F0 trajectory of the neutral tone, but the latter also gradually converges over the course of several neutral tone syllables, 2) post-focus lowering, which is caused by the effect of focus, occurs in both neutral-tone-ending and High-tone-ending sentences, with the interrogative intonation in questions realized as an upward shift starting from the focused word, and 3) sentence-final neutral tone has a falling contour even in questions, thus contrasting with sentence-final High tone, which has a rising contour in questions. Index Terms: neutral tone, focus, sentence type 1.
Fang Liu 0018, Yi Xu 0007
INTERSPEECH2
2006 Quantitative Target Approximation Model: Simulating Underlying Mechanisms of Tones and Intonations
abstract
This paper proposes a quantitative target approximation (qTA) model for simulating tone and intonation. Based on two theoretical models: the target approximation model (Y. Xu and Q.E. Wang, 2001) and the PENTA model (Y. Xu, 2005), the qTA model additionally incorporates several assumptions related to the underlying articulatory mechanisms, including (1) F0production can be represented by a second-order overdamped system, and (2) the system is controlled by a time-delayed feedback loop to sequentially approximate underlying pitch targets. We tested the model with the dataset from Y. Xu (1999). Two experiments were conducted to validate the model and to study the effect of tone, position, and focus. The results were satisfactory in term of the error rate and correlation
Santitham Prom-on, Yi Xu 0007, Bundit Thipakorn
ICASSP (1)2
2005 Speech melody as articulatorily implemented communicative functions
Yi Xu 0007
Speech Commun.1
2002 Segmentation of glides with tonal alignment as reference
abstract
This paper reports an attempt to determine the segmentation of glides using existing knowledge about tonal alignment as reference. It is found that the likely onset of a glide is much earlier than what is acoustically the most obvious, i.e., the point where the formants reach their extremes. It is further found that there is indication that the point of formant extremes may in fact be the point of glide release. 1. THE PROBLEM If we set aside the issue of overlapping gestures [2, 4], determining segmental boundaries can be quite straightforward in some cases, at least for practical purposes. In a sequence of two CV syllables such as /mama/, the boundary between the two syllables can be said to be at the onset of the second /m/. We may refer to this kind of boundary as the "de facto " syllable
Yi Xu 0007, Fang Liu 0018
INTERSPEECH1
2001 Pitch targets and their realization: Evidence from Mandarin Chinese
Yi Xu 0007, Emily Q. Wang
Speech Commun.1
1988 An acoustic-phonetic oriented system for synthesizing Chinese
Shun-an Yang, Yi Xu 0007
Speech Commun.2