Yamato Ohtani

dblp:34/8763 · DBLP profile ↗
← Back
28ranked-venue papers
12as first author
10since 2021 · last 2025
0009-0009-2961-2821ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 28 · 12 first-author · 10 since 2021Artificial intelligence and machine learning · 22 · 10 first-author · 7 since 2021
YearPublicationVenuePosition
2025 Layer-wise Analysis for Quality of Multilingual Synthesized Speech
abstract
While supervised quality predictors for synthesized speech have demonstrated strong correlations with human ratings, their requirement for in-domain labeled training data hinders their generalization ability to new domains. Unsupervised approaches based on pretrained self-supervised learning (SSL) based models and automatic speech recognition (ASR) models are a promising alternative; however, little is known about how these models encode information about speech quality. Towards the goal of better understanding how different aspects of speech quality are encoded in a multilingual setting, we present a layer-wise analysis of multilingual pretrained speech models based on reference modeling. We find that features extracted from early SSL layers show correlations with human ratings of synthesized speech, and later layers of ASR models can predict quality of non-neural systems as well as intelligibility. We also demonstrate the importance of using well-matched reference data.
Erica Cooper, Takuma Okamoto, Yamato Ohtani, Tomoki Toda, Hisashi Kawai
ASRU3
2025 Voice Factor Control Using FIR-Based Fast Neural Vocoder for Speech Generation Applications
abstract
We have proposed a fast neural vocoder based on the source-filter model introducing finite impulse response (FIR) filters called FIRNet. FIRNet is highly compatible with digital signal processing (DSP) and can, therefore, generate waveforms from vocoder parameters and modified voice factors, such as tone, intonation, and timbre, using DSP. Although modern neural waveform generation systems, such as voice conversion and text-to-speech (TTS), have been able to generate human-like synthetic speech and imitate the reference speaker’s timbre, it is challenging for these systems to manually control arbitrary voice factors, unlike traditional TTS systems. By applying FIRNet to modern neural waveform generation systems, they can achieve arbitrary voice factor controllability. We will demonstrate two applications using FIRNet with DSP-based voice factor controls: one is analysis-synthesis, and the other is text-to-speech.
Yamato Ohtani, Takuma Okamoto, Tomoki Toda, Hisashi Kawai
ASRU1
2025 Mora-Level Prosody Prediction for Text-to-Speech Using Japanese BERT Without Accentual Labels
abstract
In practical text-to-speech (TTS) for pitch accent languages, such as Japanese, high-fidelity synthesis with correct prosody requires not only a phoneme sequence but also accentual information. Although accentual information can be obtained from accent dictionaries, words not included in the dictionaries and accent sandhi are sometimes synthesized with incorrect prosody, and manual registration of huge amounts of accent data is costly. Additionally, previous machine learning-based data-driven accent information estimation approaches for TTS also require huge quantities of handcrafted accentual labels during training. This paper proposes a data-driven prosody prediction method for Japanese TTS that uses Japanese BERT and does not require any accentual labels during training. A Japanese TTS acoustic model with mora-level (katakana sequence) input is first trained and mora-level fundamental frequency values (fo), which directly correspond to the prosody, are extracted for the training data using forced alignment. Then, a pre-trained Japanese BERT is finetuned for the mora-level foprediction task with word sequences including kanji and the corresponding katakana sequences as input and the mora-level foextracted using forced alignment as the prediction target. During TTS inference, the mora-level fosequence predicted by the finetuned Japanese BERT is input to the TTS acoustic model along with the katakana input, and correct prosodic synthesis can be realized thanks to this predicted fosequence. Experimental results demonstrate that the proposed method can realize the same synthesis quality and higher accent correctness compared with conventional neural TTS models with accentual labels.
Tadashi Ogura, Takuma Okamoto, Yamato Ohtani, Erica Cooper, Tomoki Toda, Hisashi Kawai
ICASSP3
2025 GST-BERT-TTS: Prosody Prediction Without Accentual Labels For Multi-Speaker TTS Using BERT With Global Style Tokens
Tadashi Ogura, Takuma Okamoto, Yamato Ohtani, Erica Cooper, Tomoki Toda, Hisashi Kawai
INTERSPEECH3
2024 FIRNet: Fundamental Frequency Controllable Fast Neural Vocoder With Trainable Finite Impulse Response Filter
abstract
Some neural vocoders with fundamental frequency (f0) control have succeeded in performing real-time inference on a single CPU while preserving the quality of the synthetic speech. However, compared with legacy vocoders based on signal processing, their inference speeds are still low. This paper proposes a neural vocoder based on the source-filter model with trainable time-variant finite impulse response (FIR) filters, to achieve a similar inference speed to legacy vocoders. In the proposed model, FIRNet, multiple FIR coefficients are predicted using the neural networks, and the speech waveform is then generated by convolving a mixed excitation signal with these FIR coefficients. Experimental results show that FIRNet can achieve an inference speed similar to legacy vocoders while maintaining f0controllability and natural speech quality.
Yamato Ohtani, Takuma Okamoto, Tomoki Toda, Hisashi Kawai
ICASSP1
2024 Convnext-TTS And Convnext-VC: Convnext-Based Fast End-To-End Sequence-To-Sequence Text-To-Speech And Voice Conversion
abstract
End-to-end (E2E) sequence-to-sequence (S2S) neural text-to-speech (TTS) models and E2E-S2S neural voice conversion (VC) models can achieve high-quality speech synthesis with a single neural network. To further improve the synthesis quality of E2E-S2S TTS and VC models and increase their inference speed, we propose a Transformer-free ConvNeXt-based encoder and decoder. Additionally, to further increase the inference speed, we propose ConvNeXt-TTS and ConvNeXt-VC, which include the WaveNeXt neural vocoder. This is also constructed from ConvNeXt blocks and can achieve much faster synthesis than HiFi-GAN. The results of experiments using the Hi-Fi-CAPTAIN corpus for the E2E-S2S-TTS and E2E-S2S-VC conditions demonstrate that the proposed ConvNeXt-based encoder and decoder can perform inference three times faster than a Transformer-based encoder and decoder while improving the synthesis quality. In particular, ConvNeXt-TTS and ConvNeXt-VC can achieve very fast E2E-S2S-TTS and E2E-S2S-VC with a real-time factor of 0.05 using a single-core CPU.
Takuma Okamoto, Yamato Ohtani, Tomoki Toda, Hisashi Kawai
ICASSP2
2024 Mobile PresenTra: NICT fast neural text-to-speech system on smartphones with incremental inference of MS-FC-HiFi-GAN for law-latency synthesis
Takuma Okamoto, Yamato Ohtani, Hisashi Kawai
INTERSPEECH2
2024 Challenge of Singing Voice Synthesis Using Only Text-To-Speech Corpus With FIRNet Source-Filter Neural Vocoder
Takuma Okamoto, Yamato Ohtani, Sota Shimizu, Tomoki Toda, Hisashi Kawai
INTERSPEECH2
2023 WaveNeXt: ConvNeXt-Based Fast Neural Vocoder Without ISTFT layer
abstract
A recently proposed neural vocoder, Vocos, can perform inference ten times faster than HiFi-GAN because of its use of ConvNeXt layers that can predict high-resolution short-time Fourier transform (STFT) spectra and an inverse STFT layer. To improve synthesis quality while preserving inference speed, this paper proposes an alternative ConvNeXt-based fast neural vocoder, WaveNeXt, in which the inverse STFT layer in Vocos is replaced with a trainable linear layer that can directly predict speech waveform samples without STFT spectra. Additionally, by integrating the JETS-based end-to-end text-to-speech (E2E TTS) framework, E2E TTS models can also be constructed with Vocos and WaveNeXt. Furthermore, full-band models with a sampling frequency of 48 kHz were investigated. The results of experiments for both the analysis-synthesis and E2E TTS conditions demonstrate that the proposed WaveNeXt can achieve higher quality synthesis than Vocos while preserving its inference speed.
Takuma Okamoto, Haruki Yamashita, Yamato Ohtani, Tomoki Toda, Hisashi Kawai
ASRU3
2022 Spoken-Text-Style Transfer with Conditional Variational Autoencoder and Content Word Storage
Daiki Yoshioka, Yusuke Yasuda, Noriyuki Matsunaga, Yamato Ohtani, Tomoki Toda
INTERSPEECH4
2020 A Cyclical Post-Filtering Approach to Mismatch Refinement of Neural Vocoder for Text-to-Speech Systems
abstract
Recently, the effectiveness of text-to-speech (TTS) systems combined with neural vocoders to generate high-fidelity speech has been shown.However, collecting the required training data and building these advanced systems from scratch are time and resource consuming.An economical approach is to develop a neural vocoder to enhance the speech generated by existing or low-cost TTS systems.Nonetheless, this approach usually suffers from two issues: 1) temporal mismatches between TTS and natural waveforms and 2) acoustic mismatches between training and testing data.To address these issues, we adopt a cyclic voice conversion (VC) model to generate temporally matched pseudo-VC data for training and acoustically matched enhanced data for testing the neural vocoders.Because of the generality, this framework can be applied to arbitrary TTS systems and neural vocoders.In this paper, we apply the proposed method with a state-of-the-art WaveNet vocoder for two different basic TTS systems, and both objective and subjective experimental results confirm the effectiveness of the proposed framework.
Yi-Chiao Wu, Patrick Lumban Tobing, Kazuki Yasuhara, Noriyuki Matsunaga, Yamato Ohtani, Tomoki Toda
INTERSPEECH5
2016 Voice Quality Control Using Perceptual Expressions for Statistical Parametric Speech Synthesis Based on Cluster Adaptive Training
Yamato Ohtani, Koichiro Mori, Masahiro Morita
INTERSPEECH1
2015 Emotional transplant in statistical speech synthesis based on emotion additive model
Yamato Ohtani, Yu Nasu, Masahiro Morita, Masami Akamine
INTERSPEECH1
2014 GMM-based bandwidth extension using sub-band basis spectrum model
Yamato Ohtani, Masatsune Tamura, Masahiro Morita, Masami Akamine
INTERSPEECH1
2012 Histogram-based spectral equalization for HMM-based speech synthesis using mel-LSP
Yamato Ohtani, Masatsune Tamura, Masahiro Morita, Takehiko Kagoshima, Masami Akamine
INTERSPEECH1
2012 HMM-based speech synthesis using sub-band basis spectrum model
Yamato Ohtani, Masatsune Tamura, Masahiro Morita, Takehiko Kagoshima, Masami Akamine
INTERSPEECH1
2011 Continuous F0 in the source-excitation generation for HMM-based TTS: Do we need voiced/unvoiced classification?
abstract
Most HMM-based TTS systems use a hard voiced/unvoiced classification to produce a discontinuous F0 signal which is used for the generation of the source-excitation. When a mixed source excitation is used, this decision can be based on two different sources of information: the state-specific MSD-prior of the F0 models, and/or the frame-specific features generated by the aperiodicity model. This paper examines the meaning of these variables in the synthesis process, their interaction, and how they affect the perceived quality of the generated speech The results of several perceptual experiments show that when using mixed excitation, subjects consistently prefer samples with very few or no false unvoiced errors, whereas a reduction in the rate of false voiced errors does not produce any perceptual improvement. This suggests that rather than using any form of hard voiced/unvoiced classification, e.g., the MSD-prior, it is better for synthesis to use a continuous F0 signal and rely on the frame-level soft voiced/unvoiced decision of the aperiodicity model.
Javier Latorre, Mark J. F. Gales, Sabine Buchholz, Kate M. Knill, Masatsune Tamura, Yamato Ohtani, Masami Akamine
ICASSP6
2010 Non-parallel training for many-to-many eigenvoice conversion
abstract
This paper presents a novel training method of an eigenvoice Gaussian mixture model (EV-GMM) effectively using non-parallel data sets for many-to-many eigenvoice conversion, which is a technique for converting an arbitrary source speaker's voice into an arbitrary target speaker's voice. In the proposed method, an initial EV-GMM is trained with the conventional method using parallel data sets consisting of a single reference speaker and multiple pre-stored speakers. Then, the initial EV-GMM is further refined using non-parallel data sets including a larger number of pre-stored speakers while considering the reference speaker's voices as hidden variables. The experimental results demonstrate that the proposed method yields significant quality improvements in converted speech by enabling us to use data of a larger number of pre-stored speakers.
Yamato Ohtani, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
ICASSP1
2010 Adaptive voice-quality control based on one-to-many eigenvoice conversion
abstract
This paper presents adaptive voice-quality control methods based on one-to-many eigenvoice conversion. To intuitively control the converted voice quality by manipulating a small number of control parameters, a multiple regression Gaussian mixture model (MR-GMM) has been proposed. The MR-GMM also allows us to estimate the optimum control parameters if target speech samples are available. However, its adaptation performance is limited because the number of control parameters is too small to widely model voice quality of various target speakers. To improve the adaptation performance while keeping capability of voice-quality control, this paper proposes an extended MR-GMM (EMR-GMM) with additional adaptive parameters to extend a subspace modeling target voice quality. Experimental results demonstrate that the EMR-GMM yields significant improvements of the adaptation performance while allowing us to intuitively control the converted voice quality.
Kumi Ohta, Tomoki Toda, Yamato Ohtani, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH3
2009 Cross-language voice conversion based on eigenvoices
abstract
INTERSPEECH2009: 10th Annual Conference of the International Speech Communication Association, September 6-10, 2009, Brighton, UK.
Malorie Charlier, Yamato Ohtani, Tomoki Toda, Alexis Moinet, Thierry Dutoit
INTERSPEECH2
2009 Many-to-many eigenvoice conversion with reference voice
abstract
In this paper, we propose many-to-many voice conversion (VC) techniques to convert an arbitrary source speaker's voice into an arbitrary target speaker's voice. We have proposed one-to-many eigenvoice conversion (EVC) and many-to-one EVC. In the EVC, an eigenvoice Gaussian mixture model (EV-GMM) is trained in advance using multiple parallel data sets of a reference speaker and many pre-stored speakers. The EV-GMM is flexibly adapted to an arbitrary speaker using a small amount of adaptation data without any linguistic constraints. In this paper, we achieve many-to-many VC by sequentially performing many-to-one EVC and one-to-many EVC through the reference speaker using the same EV-GMM. Experimental results demonstrate the effectiveness of the proposed many-to-many VC.
Yamato Ohtani, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH1
2008 Low-delay voice conversion based on maximum likelihood estimation of spectral parameter trajectory
abstract
As typical voice conversion methods, two spectral conversion processes have been proposed: 1) the frame-based conversion that converts spectral parameters frame by frame and 2) the trajectory-based conversion that converts all spectral parameters over an utterance simultaneously. The former process is capable of real-time conversion but it sometimes causes inappropriate spectral movements. On the other hand, the latter process provides the converted spectral parameters exhibiting proper dynamic characteristics but a batch process is inevitable. To achieve the real-time conversion process considering spectral dynamic characteristics, we propose a time-recursive conversion algorithm based on maximum likelihood estimation of spectral parameter trajectory. Experimental results show that the proposed method achieves the low-delay conversion process, e.g., only one frame delay, while keeping the conversion performance comparably high to that of the conventional trajectory-based conversion. Index Terms: speech synthesis, voice conversion, Gaussian mixture model, maximum likelihood estimation, time-recursive algorithm. 1.
Takashi Muramatsu, Yamato Ohtani, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH2
2008 An improved one-to-many eigenvoice conversion system
abstract
We have previously developed a one-to-many eigenvoice conversion (EVC) system enabling the conversion from a specific source speaker's voice into an arbitrary target speaker's voice. In this system, eigenvoice Gaussian mixture model (EV-GMM) is trained in advance with multiple parallel data sets composed of utterance pairs of the source and many pre-stored target speakers. The EV-GMM is effectively adapted to an arbitrary target speaker using a small amount of adaptation data. Although this system achieves the very flexible training of the conversion model, the quality of the converted speech is still not high enough. In order to alleviate this problem, we simultaneously apply the following promising techniques to the one-to-many EVC system: 1) STRAIGHT mixed excitation, 2) the conversion algorithm considering global variance, and 3) speaker adaptive training of the EV-GMM. Experimental results demonstrate that the proposed system causes remarkable improvements in the performance of EVC.
Yamato Ohtani, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH1
2008 Maximum a posteriori adaptation for many-to-one eigenvoice conversion
abstract
Many-to-one eigenvoice conversion (EVC) allows the conversion from an arbitrary speaker's voice into the pre-determined target speaker's voice. In this method, a canonical eigenvoice Gaussian mixture model is effectively adapted to any source speaker using only a few utterances as the adaptation data. In this paper, we propose a many-to-one EVC based on maximum a posteriori (MAP) adaptation for further improving the robustness of the adaptation process to the amount of adaptation data. Results of objective and subjective evaluations demonstrate that the proposed method is the most effective among the other conventional many-to-one VC methods when using any amount of adaptation data (e.g., from 300 ms to 16 utterances).
Daisuke Tani, Tomoki Toda, Yamato Ohtani, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH3
2007 One-to-Many and Many-to-One Voice Conversion Based on Eigenvoices
abstract
This paper describes two flexible frameworks of voice conversion (VC), i.e., one-to-many VC and many-to-one VC. One-to-many VC realizes the conversion from a user's voice as a source to arbitrary target speakers' ones and many-to-one VC realizes the conversion vice versa. We apply eigenvoice conversion (EVC) to both VC frameworks. Using multiple parallel data sets consisting of utterance-pairs of the user and multiple pre-stored speakers, an eigenvoice Gaussian mixture model (EV-GMM) is trained in advance. Unsupervised adaptation of the EV-GMM is available to construct the conversion model for arbitrary target speakers in one-to-many VC or arbitrary source speakers in many-to-one VC using only a small amount of their speech data. Results of various experimental evaluations demonstrate the effectiveness of the proposed VC frameworks.
Tomoki Toda, Yamato Ohtani, Kiyohiro Shikano
ICASSP (4)2
2007 Speaker adaptive training for one-to-many eigenvoice conversion based on Gaussian mixture model
abstract
One-to-many eigenvoice conversion (EVC) allows the conversion of a specific source speaker into arbitrary target speakers. Eigenvoice Gaussian mixture model (EV-GMM) is trained in advance with multiple parallel data sets consisting of the source speaker and many pre-stored target speakers. The EV-GMM is adapted for arbitrary target speakers using only a few utterances by estimating a small number of free parameters. Therefore, the initial EV-GMM directly affects the conversion performance of the adapted EV-GMM. In order to prepare a better initial model, this paper proposes Speaker Adaptive Training (SAT) of a canonical EV-GMM in one-to-many EVC. Results of objective and subjective evaluations demonstrate that SAT causes significant improvements in the performance of EVC.
Yamato Ohtani, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH1
2006 Maximum likelihood voice conversion based on GMM with STRAIGHT mixed excitation
abstract
The performance of voice conversion has been considerably improved through statistical modeling of spectral sequences. However, the converted speech still contains traces of artificial sounds. To alleviate this, it is necessary to statistically model a source sequence as well as a spectral sequence. In this paper, we introduce STRAIGHT mixed excitation to a framework of the voice conversion based on a Gaussian Mixture Model (GMM) on joint probability density of source and target features. We convert both spectral and source feature sequences based on Maximum Likelihood Estimation (MLE). Objective and subjective evaluation results demonstrate that the proposed source conversion produces strong improvements in both the converted speech quality and the conversion accuracy for speaker individuality.
Yamato Ohtani, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano
INTERSPEECH1
2006 Eigenvoice conversion based on Gaussian mixture model
abstract
This paper describes a novel framework of voice conversion (VC). We call it eigenvoice conversion (EVC). We apply EVC to the conversion from a source speaker's voice to arbitrary target speakers' voices. Using multiple parallel data sets consisting of utterance-pairs of the source and multiple pre-stored target speakers, a canonical eigenvoice GMM (EV-GMM) is trained in advance. That conversion model enables us to flexibly control the speaker individuality of the converted speech by manually setting weight parameters. In addition, the optimum weight set for a specific target speaker is estimated using only speech data of the target speaker without any linguistic restrictions. We evaluate the performance of EVC by a spectral distortion measure. Experimental results demonstrate that EVC works very well even if we use only a few utterances of the target speaker for the weight estimation.
Tomoki Toda, Yamato Ohtani, Kiyohiro Shikano
INTERSPEECH2