VLDB 2026 Research / reviewers in the wild / expert
Kayoko Yanagisawa
dblp:91/9230
· DBLP profile ↗
18ranked-venue papers
4as first author
7since 2021 · last 2023
0000-0002-3444-7287ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 14 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Modelling Low-Resource Accents Without Accent-Specific TTS FrontendabstractThis work focuses on modelling a speaker’s accent that does not have a dedicated text-to-speech (TTS) frontend, including a grapheme-to-phoneme (G2P) module. Prior work on modelling accents assumes a phonetic transcription is available for the target accent, which might not be the case for low-resource, regional accents. In our work, we propose an approach whereby we first augment the target accent data to sound like the donor voice via voice conversion, then train a multi-speaker multi-accent TTS model on the combination of recordings and synthetic data, to generate the donor’s voice speaking in the target accent. Throughout the procedure, we use a TTS frontend developed for the same language but a different accent. We show qualitative and quantitative analysis where the proposed strategy achieves state-of-the-art results compared to other generative models. Our work demonstrates that low resource accents can be modelled with relatively little data and without developing an accent-specific TTS frontend. Audio samples of our model converting to multiple accents are available on our web page3. Georgi Tinchev, Marta Czarnowska, Kamil Deja, Kayoko Yanagisawa, Marius Cotescu |
ICASSP | 4 |
| 2023 | Cross-Lingual Knowledge Distillation via Flow-Based Voice Conversion for Robust Polyglot Text-to-Speech
Dariusz Piotrowski, Renard Korzeniowski, Alessio Falai, Sebastian Cygert, Kamil Pokora, Georgi Tinchev, Ziyao Zhang 0001, Kayoko Yanagisawa |
ICONIP (7) | 8 |
| 2023 | Comparing normalizing flows and diffusion models for prosody and acoustic modelling in text-to-speech
Guangyan Zhang, Thomas Merritt, Manuel Sam Ribeiro, Biel Tura Vecino, Kayoko Yanagisawa, Kamil Pokora, Abdelhamid Ezzerg, Sebastian Cygert, Ammar Abbas, Piotr Bilinski, Roberto Barra-Chicote, Daniel Korzekwa, Jaime Lorenzo-Trueba |
INTERSPEECH | 5 |
| 2022 | Creating New Voices using Normalizing FlowsabstractCreating realistic and natural-sounding synthetic speech remains a big challenge for voice identities unseen during training. As there is growing interest in synthesizing voices of new speakers, here we investigate the ability of normalizing flows in text-to-speech (TTS) and voice conversion (VC) modes to extrapolate from speakers observed during training to create unseen speaker identities. Firstly, we create an approach for TTS and VC, and then we comprehensively evaluate our methods and baselines in terms of intelligibility, naturalness, speaker similarity, and ability to create new voices. We use both objective and subjective metrics to benchmark our techniques on 2 evaluation tasks: zero-shot and new voice speech synthesis. The goal of the former task is to measure the precision of the conversion to an unseen voice. The goal of the latter is to measure the ability to create new voices. Extensive evaluations demonstrate that the proposed approach systematically allows to obtain state-of-the-art performance in zero-shot speech synthesis and creates various new voices, unobserved in the training set. We consider this work to be the first attempt to synthesize new voices based on mel-spectrograms and normalizing flows, along with a comprehensive analysis and comparison of the TTS and VC modes. Piotr Bilinski, Thomas Merritt, Abdelhamid Ezzerg, Kamil Pokora, Sebastian Cygert, Kayoko Yanagisawa, Roberto Barra-Chicote, Daniel Korzekwa |
INTERSPEECH | 6 |
| 2022 | Unify and Conquer: How Phonetic Feature Representation Affects Polyglot Text-To-Speech (TTS)abstractAn essential design decision for multilingual Neural Text-To-Speech (NTTS) systems is how to represent input linguistic features within the model.Looking at the wide variety of approaches in the literature, two main paradigms emerge, unified and separate representations.The former uses a shared set of phonetic tokens across languages, whereas the latter uses unique phonetic tokens for each language.In this paper, we conduct a comprehensive study comparing multilingual NTTS systems models trained with both representations.Our results reveal that the unified approach consistently achieves better crosslingual synthesis with respect to both naturalness and accent.Separate representations tend to have an order of magnitude more tokens than unified ones, which may affect model capacity.For this reason, we carry out an ablation study to understand the interaction of the representation type with the size of the token embedding.We find that the difference between the two paradigms only emerges above a certain threshold embedding size.This study provides strong evidence that unified representations should be the preferred paradigm when building multilingual NTTS systems. Ariadna Sánchez, Alessio Falai, Orazio Angelini, Kayoko Yanagisawa |
INTERSPEECH | 5 |
| 2022 | Mix and Match: An Empirical Study on Training Corpus Composition for Polyglot Text-To-Speech (TTS)abstractTraining multilingual Neural Text-To-Speech (NTTS) models using only monolingual corpora has emerged as a popular way for building voice cloning based Polyglot NTTS systems.In order to train these models, it is essential to understand how the composition of the training corpora affects the quality of multilingual speech synthesis.In this context, it is common to hear questions such as "Would including more Spanish data help my Italian synthesis, given the closeness of both languages?".Unfortunately, we found existing literature on the topic lacking in completeness in this regard.In the present work, we conduct an extensive ablation study aimed at understanding how various factors of the training corpora, such as language family affiliation, gender composition, and the number of speakers, contribute to the quality of Polyglot synthesis.Our findings include the observation that female speaker data are preferred in most scenarios, and that it is not always beneficial to have more speakers from the target language variant in the training corpus.The findings herein are informative for the process of data procurement and corpora building. Alessio Falai, Ariadna Sánchez, Orazio Angelini, Kayoko Yanagisawa |
INTERSPEECH | 5 |
| 2022 | Remap, Warp and Attend: Non-Parallel Many-to-Many Accent Conversion with Normalizing FlowsabstractRegional accents of the same language affect not only how words are pronounced (i.e., phonetic content), but also impact prosodic aspects of speech such as speaking rate and intonation. This paper investigates a novel flow-based approach to accent conversion using normalizing flows. The proposed approach revolves around three steps: remapping the phonetic conditioning, to better match the target accent, warping the duration of the converted speech, to better suit the target phonemes, and an attention mechanism that implicitly aligns source and target speech sequences. The proposed remap-warp-attend system enables adaptation of both phonetic and prosodic aspects of speech while allowing for source and converted speech signals to be of different lengths. Objective and subjective evaluations show that the proposed approach significantly outperforms a competitive CopyCat baseline model in terms of similarity to the target accent, naturalness and intelligibility. Abdelhamid Ezzerg, Thomas Merritt, Kayoko Yanagisawa, Piotr Bilinski, Magdalena Proszewska, Kamil Pokora, Renard Korzeniowski, Roberto Barra-Chicote, Daniel Korzekwa |
SLT | 3 |
| 2020 | Singing Synthesis: With a Little Help from my AttentionabstractWe present UTACO, a singing synthesis model based on an attention-based sequence-to-sequence mechanism and a vocoder based on dilated causal convolutions. These two classes of models have significantly affected the field of text-to-speech, but have never been thoroughly applied to the task of singing synthesis. UTACO demonstrates that attention can be successfully applied to the singing synthesis field and improves naturalness over the state of the art. The system requires considerably less explicit modelling of voice features such as F0 patterns, vibratos, and note and phoneme durations, than previous models in the literature. Despite this, it shows a strong improvement in naturalness with respect to previous neural singing synthesis models. The model does not require any durations or pitch patterns as inputs, and learns to insert vibrato autonomously according to the musical context. However, we observe that, by completely dispensing with any explicit duration modelling it becomes harder to obtain the fine control of timing needed to exactly match the tempo of a song. Orazio Angelini, Alexis Moinet, Kayoko Yanagisawa, Thomas Drugman |
INTERSPEECH | 3 |
| 2016 | Multi-stream spectral representation for statistical parametric speech synthesisabstractIn statistical parametric speech synthesis such as Hidden Markov Model (HMM) based synthesis, one of the problems is in the over-smoothing of parameters, which leads to a muffled sensation in the synthesised output. In this paper, we propose an approach in which the high frequency spectrum is modelled separately from the low frequency spectrum. The high frequency band, which does not carry much linguistic information, is clustered using a very large decision tree so as to generate parameters as close as possible to natural speech samples. The boundary frequency can be adjusted at synthesis time for each state. Subjective listening tests show that the proposed approach is significantly preferred over the conventional approach using a single spectrum stream. Samples synthesised using the proposed approach sound less muffled and more natural. Kayoko Yanagisawa, Ranniery Maia, Yannis Stylianou |
ICASSP | 1 |
| 2016 | Expressive visual text-to-speech as an assistive technology for individuals with autism spectrum conditionsabstractAdults with Autism Spectrum Conditions (ASC) experience marked difficulties in recognising the emotions of others and responding appropriately. The clinical characteristics of ASC mean that face to face or group interventions may not be appropriate for this clinical group. This article explores the potential of a new interactive technology, converting text to emotionally expressive speech, to improve emotion processing ability and attention to faces in adults with ASC. We demonstrate a method for generating a near-videorealistic avatar (XpressiveTalk), which can produce a video of a face uttering inputted text, in a large variety of emotional tones. We then demonstrate that general population adults can correctly recognize the emotions portrayed by XpressiveTalk. Adults with ASC are significantly less accurate than controls, but still above chance levels for inferring emotions from XpressiveTalk. Both groups are significantly more accurate when inferring sad emotions from XpressiveTalk compared to the original actress, and rate these expressions as significantly more preferred and realistic. The potential applications for XpressiveTalk as an assistive technology for adults with ASC is discussed. Sarah A. Cassidy, Björn Stenger, L. Van Dongen, Kayoko Yanagisawa, Vincent Wan, Simon Baron-Cohen, Roberto Cipolla |
Comput. Vis. Image Underst. | 4 |
| 2014 | Cluster adaptive training of average voice modelsabstractHidden Markov model based text-to-speech systems may be adapted so that the synthesised speech sounds like a particular person. The average voice model (AVM) approach uses linear transforms to achieve this while multiple decision tree cluster adaptive training (CAT) represents different speakers as points in a low dimensional space. This paper describes a novel combination of CAT and AVM for modelling speakers. CAT yields higher quality synthetic speech than AVMs but AVMs model the target speaker better. The resulting combination may be interpreted as a more powerful version of the AVM. Results show that the combination achieves better target speaker similarity when compared with both AVM and CAT while the speech quality is in-between AVM and CAT. Vincent Wan, Javier Latorre, Kayoko Yanagisawa, Mark J. F. Gales, Yannis Stylianou |
ICASSP | 3 |
| 2014 | Generating multiple-accent pronunciations for TTS using joint sequence model interpolationabstractStandard grapheme-to-phoneme (G2P) systems are trained using a homogeneous lexicon, for example one associated with a particular accent. In practice, a synthesis system may be required to handle multiple accents. Furthermore, a speaker rarely has a pure accent; accents vary continuously within and between regions of a country. Generating phonetic sequences for each accent is possible, but combining them to yield a single synthesis pronunciation is highly challenging. To address this problem, this paper considers a space of accents. The bases for these spaces are defined by statistical G2P models in the form of graphone models. A linear combination of these models define the accent space. By selecting a point in this continuous space, it is possible to specify the accent for an individual speaker. The performance of this approach is evaluated using an accent space defined by American, Scottish and British English. By moving around the accent space, it is shown that it is possible to synthesize speech from all these accents as well as a range of intermediate points. BalaKrishna Kolluru, Vincent Wan, Javier Latorre, Kayoko Yanagisawa, Mark J. F. Gales |
INTERSPEECH | 4 |
| 2014 | Voice expression conversion with factorised HMM-TTS modelsabstractThis paper proposes a method to modify the expression or emotion in a sample of speech without altering the speaker’s identity. The method exploits a statistical speech model that factorises the speaker identity from expressions using linear transforms. For this approach, the set of transforms that best fit the speaker and expression of the input speech sample are learned. They are then combined with the expression transforms of the desired expression taken from another speaker. Since the combined expression transform is factorised and contains information about expression only, it may be applied to the original speech sample to modify its expression to the desired one without altering the identity of the speaker. Notably, this method may be applied universally to any voice without the need for a parallel training corpus. Javier Latorre, Vincent Wan, Kayoko Yanagisawa |
INTERSPEECH | 3 |
| 2014 | Speech intonation for TTS: study on evaluation methodologyabstractThe standard evaluation of intonation models is by means of non-referenced subjective tests (pair or MOS) in which subjects rate the quality or compare different samples without any explicit reference. These tests are usually conducted on an isolated sentence basis. However, for a single sentence, with no contextual information, there are multiple valid intonations. A subject's preference over this range of intonation patterns may be highly personal. This paper investigates the degree to which this ambiguity in the appropriate intonation pattern impacts the assessments of prosody for speech synthesis systems. To examine this problem, the variance of the F0 pattern of several vocoded sentences was modified and subjects asked to compare multiple versions with different levels of modification in terms of preference/quality. Then, they were presented with the reference which defines the original intonation and asked about the similarity to that reference. The results show that subjects can identify the samples with no F0 variance modification when given a reference but they don't always prefer them. Thus, non-referenced tests with no context, though may help to analyse user acceptability, may not be appropriate to measure the performance of intonation models. Javier Latorre, Kayoko Yanagisawa, Vincent Wan, BalaKrishna Kolluru, Mark J. F. Gales |
INTERSPEECH | 2 |
| 2014 | Noise-robust TTS speaker adaptation with statistics smoothingabstractIn practical scenarios for speaker adaptation of speech synthesis systems, the quality of adaptation audio data may be poor. In these situations, it is necessary to make use of the available audio to capture the speaker attributes, whilst aiming to obtain a synthesis voice which does not have any of the lowquality attributes of the audio. One approach to achieving this is to define a sub-space of parametric synthesis parameters in which the adapted system must lie. Though this yields reasonable synthesis quality, target speaker similarity degrades. Quality is also affected in severe noise conditions. This paper describes a smoothing approach that addresses this problem. For a noisy target speaker, first a 'similar speaker' is selected from a database of speakers. Statistics from this speaker are then smoothed with those obtained from the target speaker. By appropriately combining the two sources of information, it is possible to balance similarity and quality. Results indicate that both the quality and similarity can be improved by smoothing, especially for severe noise conditions. The similarity performance, however, varies from speaker to speaker, indicating the importance of a reasonable automatic speaker selection method and the coverage of the candidate speaker pool. Kayoko Yanagisawa, Langzhou Chen, Mark J. F. Gales |
INTERSPEECH | 1 |
| 2013 | Photo-realistic expressive text to talking head synthesis
Vincent Wan, Art Blokland, Norbert Braunschweiler, Langzhou Chen, BalaKrishna Kolluru, Javier Latorre, Ranniery Maia, Björn Stenger, Kayoko Yanagisawa, Yannis Stylianou, Masami Akamine, Mark J. F. Gales, Roberto Cipolla |
INTERSPEECH | 10 |
| 2010 | A phonetic alternative to cross-language voice conversion in a text-dependent context: evaluation of speaker identityabstractSpoken language conversion (SLC) aims to generate utterances in the voice of a speaker but in a language unknown to them, using speech synthesis systems and speech processing techniques. Previous approaches to SLC have been based on cross-language voice conversion (VC), which has underlying assumptions that ignore phonetic and phonological differences between languages, leading to a reduction in intelligibility of the output. Accent morphing (AM) was proposed as an alternative approach, and its intelligibility performance was investigated in a previous study. AM attempts to preserve the voice characteristics of the target speaker whilst modifying their accent, using phonetic knowledge obtained from a native speaker of the target language. This paper examines AM and VC in terms of how similar the output sounds like the target speaker. AM achieved similarity ratings at least equivalent to VC, but the study highlighted various difficulties in evaluating speaker identity in a SLC context. © 2010 ISCA. Kayoko Yanagisawa, Mark A. Huckvale |
INTERSPEECH | 1 |
| 2008 | A phonetic assessment of cross-language voice conversionabstractCross-language voice conversion maps the speech of speaker S1 in language L1 to the voice of speaker S2 using knowledge only of how S2 speaks a different language L2. This mapping is usually performed using speech material from S1 and S2 that has been deemed "equivalent" in either acoustic or phonetic terms. This study investigates the issue of equivalence in more detail, and contrasts the performance of a voice conversion system operating in both mono-lingual and cross-lingual modes using Japanese and English. We show that voice conversion impacts the intelligibility of the converted speech, but to a significantly greater degree for cross-language conversion. A phonetic comparison of the monolingual and cross-language converted speech suggests that consonantal information is degraded in both conditions, but vowel information is degraded more in the cross-language condition. Copyright © 2008 ISCA. Kayoko Yanagisawa, Mark A. Huckvale |
INTERSPEECH | 1 |