EDBT 2026 Demo / reviewers in the wild / expert
Tomoki Koriyama
dblp:44/9232
· DBLP profile ↗
44ranked-venue papers
16as first author
14since 2021 · last 2026
0000-0002-8347-5604ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 41 · 15 first-author · 14 since 2021Artificial intelligence and machine learning · 28 · 9 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Speaker-conditioned phrase break prediction for text-to-speech with phoneme-level pre-trained language modelabstractThis paper advances phrase break prediction (also known as phrasing) in multi-speaker text-to-speech (TTS) systems. We integrate speaker-specific features by leveraging speaker embeddings to enhance the performance of the phrasing model. We further demonstrate that these speaker embeddings can capture speaker-related characteristics solely from the phrasing task. Besides, we explore the potential of pre-trained speaker embeddings for unseen speakers through a few-shot adaptation method. Furthermore, we pioneer the application of phoneme-level pre-trained language models to this TTS front-end task, which significantly boosts the accuracy of the phrasing model. Our methods are rigorously assessed through both objective and subjective evaluations, demonstrating their effectiveness. • Speaker-conditioned phrasing model improves accuracy in multi-speaker phrasing tasks. • We explore various speaker embeddings in phrasing models. • We apply phoneme-level pre-trained language models to enhance phrasing accuracy. • We propose a speaker adaptation method for few-shot phrasing tasks. • We verify that speaker embeddings learn human-aligned features via phrasing tasks. Yuki Saito 0001, Takaaki Saeki, Tomoki Koriyama, Wataru Nakata, Detai Xin, Hiroshi Saruwatari |
Speech Commun. | 4 |
| 2025 | Speaker-agnostic Emotion Vector for Cross-speaker Emotion Intensity Control
Masato Murata, Koichi Miyazaki, Tomoki Koriyama |
INTERSPEECH | 3 |
| 2025 | Eigenvoice Synthesis based on Model Editing for Speaker Generation
Masato Murata, Koichi Miyazaki, Tomoki Koriyama, Tomoki Toda |
INTERSPEECH | 3 |
| 2024 | VAE-based Phoneme Alignment Using Gradient Annealing and SSL Acoustic Features
Tomoki Koriyama |
INTERSPEECH | 1 |
| 2024 | An Attribute Interpolation Method in Speech Synthesis by Model Merging
Masato Murata, Koichi Miyazaki, Tomoki Koriyama |
INTERSPEECH | 3 |
| 2024 | Frame-Wise Breath Detection with Self-Training: An Exploration of Enhancing Breath Naturalness in Text-to-Speech
Tomoki Koriyama, Yuki Saito 0001 |
INTERSPEECH | 2 |
| 2023 | Structured State Space Decoder for Speech Recognition and SynthesisabstractAutomatic speech recognition (ASR) systems developed in recent years have shown promising results with self-attention models (e.g., Transformer and Conformer), which are replacing conventional recurrent neural networks. Meanwhile, a structured state space model (S4) has been recently proposed, producing promising results for various long-sequence modeling tasks, including raw speech classification. The S4 model can be trained in parallel, similar to the Transformer model. In this study, we applied S4 as a decoder for ASR and text-to-speech (TTS) tasks, respectively, by comparing it with the Transformer decoder. For the ASR task, our experimental results demonstrate that the proposed model achieves a competitive word error rate (WER) of 1.88%/4.25% on the LibriSpeech test-clean/test-other set and a character error rate (CER) of 3.80%/2.63%/2.98% on the CSJ eval1/eval2/eval3 set. Furthermore, the proposed model is more robust than the standard Transformer model, particularly for long-form speech on both the datasets. In the TTS task, the proposed method outperforms the Transformer baseline. Koichi Miyazaki, Masato Murata, Tomoki Koriyama |
ICASSP | 3 |
| 2023 | Duration-Aware Pause Insertion Using Pre-Trained Language Model for Multi-Speaker Text-To-SpeechabstractPause insertion, also known as phrase break prediction and phrasing, is an essential part of TTS systems because proper pauses with natural duration significantly enhance the rhythm and intelligibility of synthetic speech. However, conventional phrasing models ignore various speakers’ different styles of inserting silent pauses, which can degrade the performance of the model trained on a multi-speaker speech corpus. To this end, we propose more powerful pause insertion frameworks based on a pre-trained language model. Our approach uses bidirectional encoder representations from transformers (BERT) pre-trained on a large-scale text corpus, injecting speaker embeddings to capture various speaker characteristics. We also leverage duration-aware pause insertion for more natural multi-speaker TTS. We develop and evaluate two types of models. The first improves conventional phrasing models on the position prediction of respiratory pauses (RPs), i.e., silent pauses at word transitions without punctuation. It performs speaker-conditioned RP prediction considering contextual information and is used to demonstrate the effect of speaker information on the prediction. The second model is further designed for phoneme-based TTS models and performs duration-aware pause insertion, predicting both RPs and punctuation-indicated pauses (PIPs) that are categorized by duration. The evaluation results show that our models improve the precision and recall of pause insertion and the rhythm of synthetic speech. Tomoki Koriyama, Yuki Saito 0001, Takaaki Saeki, Detai Xin, Hiroshi Saruwatari |
ICASSP | 2 |
| 2022 | Predicting VQVAE-based Character Acting Style from Quotation-Annotated Text for Audiobook Speech Synthesis
Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, Yuki Saito 0001, Yusuke Ijima, Ryo Masumura, Hiroshi Saruwatari |
INTERSPEECH | 2 |
| 2022 | UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022abstractWe present the UTokyo-SaruLab mean opinion score (MOS) prediction system submitted to VoiceMOS Challenge 2022.The challenge is to predict the MOS values of speech samples collected from previous Blizzard Challenges and Voice Conversion Challenges for two tracks: a main track for in-domain prediction and an out-of-domain (OOD) track for which there is less labeled data from different listening tests.Our system is based on ensemble learning of strong and weak learners.Strong learners incorporate several improvements to the previous finetuning models of self-supervised learning (SSL) models, while weak learners use basic machine-learning methods to predict scores from SSL features.In the Challenge, our system had the highest score on several metrics for both the main and OOD tracks.In addition, we conducted ablation studies to investigate the effectiveness of our proposed methods. Takaaki Saeki, Detai Xin, Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, Hiroshi Saruwatari |
INTERSPEECH | 4 |
| 2021 | Harmonic WaveGAN: GAN-Based Speech Waveform Generation Model with Harmonic Structure Discriminator
Kazuki Mizuta, Tomoki Koriyama, Hiroshi Saruwatari |
Interspeech | 2 |
| 2021 | Sequence-to-Sequence Learning for Deep Gaussian Process Based Speech Synthesis Using Self-Attention GP Layer
Taiki Nakamura, Tomoki Koriyama, Hiroshi Saruwatari |
Interspeech | 2 |
| 2021 | Cross-Lingual Speaker Adaptation Using Domain Adaptation and Speaker Consistency Loss for Text-To-Speech Synthesis
Detai Xin, Yuki Saito 0001, Shinnosuke Takamichi, Tomoki Koriyama, Hiroshi Saruwatari |
Interspeech | 4 |
| 2021 | Deep Gaussian process based multi-speaker speech synthesis with latent speaker representationabstractThis paper proposes deep Gaussian process (DGP)-based frameworks for multi-speaker speech synthesis and speaker representation learning. A DGP has a deep architecture of Bayesian kernel regression, and it has been reported that DGP-based single speaker speech synthesis outperforms deep neural network (DNN)-based ones in the framework of statistical parametric speech synthesis. By extending this method to multiple speakers, it is expected that higher speech quality can be achieved with a smaller number of training utterances from each speaker. To apply DGPs to multi-speaker speech synthesis, we propose two methods: one using DGP with one-hot speaker codes, and the other using a deep Gaussian process latent variable model (DGPLVM). The DGP with one-hot speaker codes uses additional GP layers to transform speaker codes into latent speaker representations. The DGPLVM directly models the distribution of latent speaker representations and learns it jointly with acoustic model parameters. In this method, acoustic speaker similarity is expressed in terms of the similarity of the speaker representations, and thus, the voices of similar speakers are efficiently modeled. We experimentally evaluated the performance of the proposed methods in comparison with those of conventional DNN and variational autoencoder (VAE)-based frameworks, in terms of acoustic feature distortion and subjective speech quality. The experimental results demonstrate that (1) the proposed DGP-based and DGPLVM-based methods improve subjective speech quality compared with a feed-forward DNN-based method, and (2) even when the amount of training data for target speakers is limited, the DGPLVM-based method outperforms other methods, including the VAE-based one. Additionally, (3) by using a speaker representation randomly sampled from the learned speaker space, the DGPLVM-based method can generate voices of non-existent speakers. Kentaro Mitsui, Tomoki Koriyama, Hiroshi Saruwatari |
Speech Commun. | 2 |
| 2020 | Utterance-Level Sequential Modeling for Deep Gaussian Process Based Speech Synthesis Using Simple Recurrent UnitabstractThis paper presents a deep Gaussian process (DGP) model with a recurrent architecture for speech sequence modeling. DGP is a Bayesian deep model that can be trained effectively with the consideration of model complexity and is a kernel regression model that can have high expressibility. In the previous studies, it was shown that the DGP-based speech synthesis outperformed neural network-based one, in which both models used a feed-forward architecture. To improve the naturalness of synthetic speech, in this paper, we show that DGP can be applied to utterance-level modeling using recurrent architecture models. We adopt a simple recurrent unit (SRU) for the proposed model to achieve a recurrent architecture, in which we can execute fast speech parameter generation by using the high parallelization nature of SRU. The objective and subjective evaluation results show that the proposed SRU-DGP-based speech synthesis outperforms not only feed-forward DGP but also automatically tuned SRU- and long short-term memory (LSTM)-based neural networks. Tomoki Koriyama, Hiroshi Saruwatari |
ICASSP | 1 |
| 2020 | Multi-Speaker Text-to-Speech Synthesis Using Deep Gaussian ProcessesabstractMulti-speaker speech synthesis is a technique for modeling multiple speakers' voices with a single model.Although many approaches using deep neural networks (DNNs) have been proposed, DNNs are prone to overfitting when the amount of training data is limited.We propose a framework for multi-speaker speech synthesis using deep Gaussian processes (DGPs); a DGP is a deep architecture of Bayesian kernel regressions and thus robust to overfitting.In this framework, speaker information is fed to duration/acoustic models using speaker codes.We also examine the use of deep Gaussian process latent variable models (DGPLVMs).In this approach, the representation of each speaker is learned simultaneously with other model parameters, and therefore the similarity or dissimilarity of speakers is considered efficiently.We experimentally evaluated two situations to investigate the effectiveness of the proposed methods.In one situation, the amount of data from each speaker is balanced (speaker-balanced), and in the other, the data from certain speakers are limited (speaker-imbalanced). Subjective and objective evaluation results showed that both the DGP and DG-PLVM synthesize multi-speaker speech more effective than a DNN in the speaker-balanced situation.We also found that the DGPLVM outperforms the DGP significantly in the speakerimbalanced situation. Kentaro Mitsui, Tomoki Koriyama, Hiroshi Saruwatari |
INTERSPEECH | 2 |
| 2020 | Cross-Lingual Text-To-Speech Synthesis via Domain Adaptation and Perceptual Similarity Regression in Speaker Space
Detai Xin, Yuki Saito 0001, Shinnosuke Takamichi, Tomoki Koriyama, Hiroshi Saruwatari |
INTERSPEECH | 4 |
| 2020 | Investigating Effective Additional Contextual Factors in DNN-Based Spontaneous Speech Synthesis
Yuki Yamashita, Tomoki Koriyama, Yuki Saito 0001, Shinnosuke Takamichi, Yusuke Ijima, Ryo Masumura, Hiroshi Saruwatari |
INTERSPEECH | 2 |
| 2020 | DNN-based Speech Synthesis Using Abundant Tags of Spontaneous Speech CorpusabstractIn this paper, we investigate the effectiveness of using rich annotations in deep neural network (DNN)-based statistical speech synthesis. DNN-based frameworks typically use linguistic information as input features called context instead of directly using text. In such frameworks, we can synthesize not only reading-style speech but also speech with paralinguistic and nonlinguistic features by adding such information to the context. However, it is not clear what kind of information is crucial for reproducing paralinguistic and nonlinguistic features. Therefore, we investigate the effectiveness of rich tags in DNN-based speech synthesis according to the Corpus of Spontaneous Japanese (CSJ), which has a large amount of annotations on paralinguistic features such as prosody, disfluency, and morphological features. Experimental evaluation results shows that the reproducibility of paralinguistic features of synthetic speech was enhanced by adding such information as context. Yuki Yamashita, Tomoki Koriyama, Yuki Saito 0001, Shinnosuke Takamichi, Yusuke Ijima, Ryo Masumura, Hiroshi Saruwatari |
LREC | 2 |
| 2019 | A Training Method Using DNN-guided Layerwise Pretraining for Deep Gaussian ProcessesabstractThis paper proposes a novel framework of training method of deep Gaussian processes (DGPs). DGPs are deep architecture models based on stacked multiple GPs, which can overcome the limitation of single-layer GPs. Although a stochastic variational inference (SVI)- based method has been proposed for DGP training with an arbitrary amount of training data, it does not always update model parameters appropriately due to repeated Monte Carlo sampling of multiple GPs. To resolve this problem, we propose a pretraining method to determine initial parameters of DGPs. In the proposed method, layer-wise training of single-layer GPs is performed using the hidden-layer values obtained by a deep neural network (DNN) that has an analogous structure to a target DGP model. The proposed method utilizes the characteristics of single-layer GPs and deep neural network (DNN) whose training is easier than DGP's. Experimental results using two speech synthesis databases with approximately 600 K and 1.4 M training data points, respectively, which were composed of hundreds of input and output features, gave the effectiveness of the proposed method. Tomoki Koriyama, Takao Kobayashi |
ICASSP | 1 |
| 2019 | Generative Moment Matching Network-based Random Modulation Post-filter for DNN-based Singing Voice Synthesis and Neural Double-trackingabstractThis paper proposes a generative moment matching network (GMMN)-based post-filter that provides inter-utterance pitch variation for deep neural network (DNN)-based singing voice synthesis. The natural pitch variation of a human singing voice leads to a richer musical experience and is used in double-tracking, a recording method in which two performances of the same phrase are recorded and mixed to create a richer, layered sound. However, singing voices synthesized using conventional DNN-based methods never vary because the synthesis process is deterministic and only one waveform is synthesized from one musical score. To address this problem, we use a GMMN to model the variation of the modulation spectrum of the pitch contour of natural singing voices and add a randomized inter-utterance variation to the pitch contour generated by conventional DNN-based singing voice synthesis. Experimental evaluations suggest that 1) our approach can provide perceptible inter-utterance pitch variation while preserving speech quality. We extend our approach to double-tracking, and the evaluation demonstrates that 2) GMMN-based neural double-tracking is perceptually closer to natural double-tracking than conventional signal processing-based artificial double-tracking is. Hiroki Tamaru, Yuki Saito 0001, Shinnosuke Takamichi, Tomoki Koriyama, Hiroshi Saruwatari |
ICASSP | 4 |
| 2019 | Semi-Supervised Prosody Modeling Using Deep Gaussian Process Latent Variable Model
Tomoki Koriyama, Takao Kobayashi |
INTERSPEECH | 1 |
| 2019 | Statistical Parametric Speech Synthesis Using Deep Gaussian ProcessesabstractThis paper proposes a framework of speech synthesis based on deep Gaussian processes (DGPs), which is a deep architecture model composed of stacked Bayesian kernel regressions. In this method, we train a statistical model of transformation from contextual features to speech parameters in a similar manner to deep neural network (DNN)-based speech synthesis. To apply DGPs to a statistical parametric speech synthesis framework, our framework uses an approximation method, doubly stochastic variational inference, which is suitable for an arbitrary amount of data. Since the training of DGPs is based on the marginal likelihood that takes into account not only data fitting, but also model complexity, DGPs are less vulnerable to overfitting compared with DNNs. In experimental evaluations, we investigated a performance comparison of the proposed DGP-based framework with a feedforward DNN-based one. Subjective and objective evaluation results showed that our DGP framework yielded a higher mean opinion score and lower acoustic feature distortions than the conventional framework. Tomoki Koriyama, Takao Kobayashi |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2018 | GPR-based Thai speech synthesis using multi-level duration prediction
Decha Moungsri, Tomoki Koriyama, Takao Kobayashi |
Speech Commun. | 2 |
| 2017 | Duration prediction using multiple Gaussian process experts for GPR-based speech synthesisabstractThis paper proposes an alternative multi-level approach to duration prediction for improving prosody generation in statistical parametric speech synthesis using multiple Gaussian process experts. We use two duration models at different levels, specifically, syllable and phone. First, we individually train syllable- and phone-level duration models. Then, the predictive distributions of syllable and phone duration models are combined by product of Gaussians. The means of combined predictive distributions are used as predicted durations for synthetic speech. We show objective and subjective evaluation results for the proposed technique by comparing with the conventional ones when the techniques are applied to Gaussian process regression (GPR)-based speech synthesis. Decha Moungsri, Tomoki Koriyama, Takao Kobayashi |
ICASSP | 2 |
| 2017 | Sampling-Based Speech Parameter Generation Using Moment-Matching NetworksabstractThis paper presents sampling-based speech parameter generation using moment-matching networks for Deep Neural Network (DNN)-based speech synthesis.Although people never produce exactly the same speech even if we try to express the same linguistic and para-linguistic information, typical statistical speech synthesis produces completely the same speech, i.e., there is no inter-utterance variation in synthetic speech.To give synthetic speech natural inter-utterance variation, this paper builds DNN acoustic models that make it possible to randomly sample speech parameters.The DNNs are trained so that they make the moments of generated speech parameters close to those of natural speech parameters.Since the variation of speech parameters is compressed into a low-dimensional simple prior noise vector, our algorithm has lower computation cost than direct sampling of speech parameters.As the first step towards generating synthetic speech that has natural inter-utterance variation, this paper investigates whether or not the proposed sampling-based generation deteriorates synthetic speech quality.In evaluation, we compare speech quality of conventional maximum likelihood-based generation and proposed sampling-based generation.The result demonstrates the proposed generation causes no degradation in speech quality. Shinnosuke Takamichi, Tomoki Koriyama, Hiroshi Saruwatari |
INTERSPEECH | 2 |
| 2016 | A speaker adaptation technique for Gaussian process regression based speech synthesis using feature space transformabstractIn this paper, we propose a speaker adaptation technique for statistical parametric speech synthesis based on Gaussian process regression (GPR). Although it is reported that the GPR-based speech synthesis improves the naturalness of synthetic speech compared with the HMM-based speech synthesis, any speaker adaptation techniques for the GPR-based one have not been established. This is because GPR is a nonparametric model and hence it is impossible to directly apply linear transforms to model parameters. In the proposed technique, we introduce feature-space transform to achieve model adaptation in the framework of GPR-based speech synthesis. Experimental results of objective and subjective tests show that the proposed technique outperforms the conventional HMM-based speaker adaptation framework. Tomoki Koriyama, Syohei Oshio, Takao Kobayashi |
ICASSP | 1 |
| 2016 | Unsupervised Stress Information Labeling Using Gaussian Process Latent Variable Model for Statistical Speech Synthesis
Decha Moungsri, Tomoki Koriyama, Takao Kobayashi |
INTERSPEECH | 2 |
| 2015 | Prosody generation using frame-based Gaussian process regression and classification for statistical parametric speech synthesisabstractThis paper proposes novel models of F0 contours and phone durations using Gaussian process regression and classification (GPR and GPC) for statistical parametric speech synthesis. Although the use of frame-based GPR has shown the effectiveness of spectral feature modeling in previous studies, the application of GPR to prosodic features, i.e., F0 and phone duration, was not investigated sufficiently because the kernel function was designed for phonetic information only. In this paper, therefore, we propose a kernel function available for multiple units such as syllables, moras, and accent phrases. The proposed kernel function is based on temporal acoustic events like the beginning of accent phrase and the relative position between the target frame and the event is utilized for the kernel function. Experimental results of objective and subjective tests show that the GPR/GPC-based F0 and duration modeling improves the prediction accuracy of acoustic features compared with HMM-based speech synthesis. Tomoki Koriyama, Takao Kobayashi |
ICASSP | 1 |
| 2015 | A comparison of speech synthesis systems based on GPR, HMM, and DNN with a small amount of training data
Tomoki Koriyama, Takao Kobayashi |
INTERSPEECH | 1 |
| 2015 | Duration prediction using multi-level model for GPR-based speech synthesis
Decha Moungsri, Tomoki Koriyama, Takao Kobayashi |
INTERSPEECH | 2 |
| 2015 | HMM-based expressive singing voice synthesis with singing style control and robust pitch modeling
Takashi Nose, Misa Kanemoto, Tomoki Koriyama, Takao Kobayashi |
Comput. Speech Lang. | 3 |
| 2014 | Parametric speech synthesis based on Gaussian process regression using global variance and hyperparameter optimizationabstractThis paper examines two issues of a statistical speech synthesis approach based Gaussian process (GP) regression. Although GP-based speech synthesis can give higher performance in generating spectral parameters than the HMM-based one, a number of issues still remain. In this paper, we incorporate global variance (GV) feature to overcome over-smoothing problem into the parameter generation. Furthermore, in order to utilize an appropriate kernel function in accordance with actual data, we propose an EM-based kernel hyperparameter optimization technique. Objective and subjective evaluation results show that using GV and hyperparameter estimation enhanced the performance in spectral feature generation. Tomoki Koriyama, Takashi Nose, Takao Kobayashi |
ICASSP | 1 |
| 2014 | Accent type and phrase boundary estimation using acoustic and language models for automatic prosodic labelingabstractThis paper proposes an automatic prosodic labeling technique for constructing speech database used for speech synthesis.In the corpus-based Japanese speech synthesis, it is essential to use annotated speech data with prosodic information such as phrase boundaries and accent types.However, manual annotation is generally time-consuming and expensive.To overcome this problem, we propose an estimation technique of accent types and phrase boundaries from speech waveform and its transcribed text using both language and acoustic models.We use conditional random field (CRF) for the language model, and HMM for the acoustic model which has shown to be effective in prosody modeling in speech synthesis.By introducing HMM, continuously changing features of F0 contours are modeled well and this results in higher estimation accuracy than conventional techniques that use simple polygonal line approximation of F0 contours. Tomoki Koriyama, Hiroshi Suzuki, Takashi Nose, Takahiro Shinozaki, Takao Kobayashi |
INTERSPEECH | 1 |
| 2014 | Transform mapping using shared decision tree context clustering for HMM-based cross-lingual speech synthesis
Daiki Nagahama, Takashi Nose, Tomoki Koriyama, Takao Kobayashi |
INTERSPEECH | 3 |
| 2014 | Prosodic variation enhancement using unsupervised context labeling for HMM-based expressive speech synthesis
Yu Maeno, Takashi Nose, Takao Kobayashi, Tomoki Koriyama, Yusuke Ijima, Hideharu Nakajima, Hideyuki Mizuno, Osamu Yoshioka |
Speech Commun. | 4 |
| 2013 | Frame-level acoustic modeling based on Gaussian process regression for statistical nonparametric speech synthesisabstractThis paper proposes a new approach to text-to-speech based on Gaussian processes which are widely used to perform non-parametric Bayesian regression and classification. The Gaussian process regression model is designed for the prediction of frame-level acoustic features from the corresponding frame information. The frame information includes relative position in the phone and preceding and succeeding phoneme information obtained from linguistic information. In this paper, a frame context kernel is proposed as a similarity measure of respective frames. Experimental results using a small data set show the potential of the proposed approach without state-dependent dynamic features or decision-tree clustering used in a conventional HMM-based approach. Tomoki Koriyama, Takashi Nose, Takao Kobayashi |
ICASSP | 1 |
| 2013 | HMM-based expressive speech synthesis based on phrase-level F0 context labelingabstractThis paper proposes a technique for adding more prosodic variations to the synthetic speech in HMM-based expressive speech synthesis. We create novel phrase-level F0 context labels from the residual information of F0 features between original and synthetic speech for the training data. Specifically, we classify the difference of average log F0 values between the original and synthetic speech into three classes which have perceptual meanings, i.e., high, neutral, and low of relative pitch at the phrase level. We evaluate both ideal and practical cases using appealing and fairy tale speech recorded under a realistic condition. In the ideal case, we examine the potential of our technique to modify the F0 patterns under a condition where the original F0 contours of test sentences are known. In the practical case, we show how the users intuitively modify the pitch by changing the initial F0 context labels obtained from the input text. Yu Maeno, Takashi Nose, Takao Kobayashi, Tomoki Koriyama, Yusuke Ijima, Hideharu Nakajima, Hideyuki Mizuno, Osamu Yoshioka |
ICASSP | 4 |
| 2013 | Statistical nonparametric speech synthesis using sparse Gaussian processesabstractThis paper proposes a statistical nonparametric speech synthesis technique based on a sparse Gaussian process regression (GPR).In our previous study, we proposed GPR-based speech synthesis where each frame of synthesis units is modeled by a regression of Gaussian processes.Preliminary experiments of synthesizing several phones including both vowels and consonants showed a potential of the technique.In this paper, the previous work is extended to full-sentence speech synthesis using sparse GPs and context modification.Specifically, clusterbased sparse Gaussian processes such as local GPs and partially independent conditional (PIC) approximation are examined as a computationally feasible approach.Moreover, frame-level context is extended to include not only a position context from a current phone but also adjacent phones to generate smoothly changing speech parameters.Objective and subjective evaluation results show that the proposed technique outperforms the HMM-based speech synthesis with minimum generation error training. Tomoki Koriyama, Takashi Nose, Takao Kobayashi |
INTERSPEECH | 1 |
| 2013 | A style control technique for singing voice synthesis based on multiple-regression HSMM
Takashi Nose, Misa Kanemoto, Tomoki Koriyama, Takao Kobayashi |
INTERSPEECH | 3 |
| 2012 | An F0 modeling technique based on prosodic events for spontaneous speech synthesisabstractThis paper proposes a technique for effective modeling of F0 contours using prosodic-event-based HMM units for HMM-based spontaneous speech synthesis. The modeling unit corresponds to one of prosodic event segments such as pitch falling by accent and pitch rising by boundary pitch movement (BPM). Since the prosodic events of one phrase are generally less frequent than the changes of phonemes, the proposed unit is expected to reduce the number of model parameters of F0, which leads to robust parameter estimation. The objective and subjective experiments using spontaneous conversational speech data show that the proposed technique can significantly reduce the number of model parameters while keeping the naturalness of the synthetic speech. Tomoki Koriyama, Takashi Nose, Takao Kobayashi |
ICASSP | 1 |
| 2012 | Discontinuous Observation HMM for Prosodic-Event-Based F0 GenerationabstractThis paper examines F0 modeling and generation techniques for spontaneous speech synthesis.In the previous study, we proposed a prosodic-unit HMM where the synthesis unit is defined as a segment between two prosodic events represented by a ToBI label framework.To take the advantage of the prosodicunit HMM, continuous F0 sequences must be modeled from discontinuous F0 data including unvoiced regions.The conventional F0 models such as the MSD-HMM and the continuous F0 HMM are not always appropriate for such demand.To overcome this problem, we propose an alternative F0 model named discontinuous observation HMM (DO-HMM) where the unvoiced frames are regarded as missing data.We objectively evaluate the performance of the DO-HMM by comparing it with the conventional F0 modeling techniques and discuss the results. Tomoki Koriyama, Takashi Nose, Takao Kobayashi |
INTERSPEECH | 1 |
| 2011 | On the Use of Extended Context for HMM-Based Spontaneous Conversational Speech SynthesisabstractThis paper addresses an issue of prosodic variability of spontaneous speech in HMM-based spontaneous conversational speech synthesis.We propose an extended context set including peculiar information to spontaneous speech derived from the annotation data embedded in a large-scale of database of spontaneous Japanese.We show the effectiveness of the newly introduced contexts from the results of objective and subjective evaluation experiments.We also propose stopping criteria for decision-tree clustering to alleviate an over-fitting problem.Experimental results show that the restriction of the size of each leaf node can improve the quality of synthetic speech. Tomoki Koriyama, Takashi Nose, Takao Kobayashi |
INTERSPEECH | 1 |
| 2010 | Conversational spontaneous speech synthesis using average voice model
Tomoki Koriyama, Takashi Nose, Takao Kobayashi |
INTERSPEECH | 1 |