EDBT 2026 Demo / reviewers in the wild / expert
Kazuhiro Kobayashi
dblp:70/8239
· DBLP profile ↗
35ranked-venue papers
11as first author
9since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 11 first-author · 8 since 2021Artificial intelligence and machine learning · 20 · 5 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Electrolaryngeal Speech Intelligibility Enhancement through Robust Linguistic EncodersabstractWe propose a novel framework for electrolaryngeal speech intelligibility enhancement through the use of robust linguistic encoders. Pretraining and fine-tuning approaches have proven to work well in this task, but in most cases, various mismatches, such as the speech type mismatch (electrolaryngeal vs. typical) or a speaker mismatch between the datasets used in each stage, can deteriorate the conversion performance of this framework. To resolve this issue, we propose a linguistic encoder robust enough to project both EL and typical speech in the same latent space, while still being able to extract accurate linguistic information, creating a unified representation to reduce the speech type mismatch. Furthermore, we introduce HuBERT output features to the proposed framework for reducing the speaker mismatch, making it possible to effectively use a large-scale parallel dataset during pretraining. We show that compared to the conventional framework using mel-spectrogram input and output features, using the proposed framework enables the model to synthesize more intelligible and naturally sounding speech, as shown by a significant 16% improvement in character error rate and 0.83 improvement in naturalness score. Lester Phillip Violeta, Wen-Chin Huang, Ryuichi Yamamoto, Kazuhiro Kobayashi, Tomoki Toda |
ICASSP | 5 |
| 2023 | Low-Latency Electrolaryngeal Speech Enhancement Based on Fastspeech2-Based Voice Conversion and Self-Supervised Speech RepresentationabstractIn this paper, we propose a low-latency sequence-to-sequence speech enhancement technique for electrolaryngeal (EL) speech. A low-latency EL speech enhancement technique based on CLDNN was previously proposed to enable laryngectomees to produce relatively naturally sounding speech compared to the original EL speech. However, the naturalness of the enhanced speech is still far from that of the natural speech due to insufficient modeling accuracy of an acoustic feature sequence and speech waveform caused by frame-wise conversion processing. To solve this problem, in this paper, we propose a low-latency sequence-to-sequence EL speech enhancement technique based on CLDNN-FastSpeech2-based VC. Moreover, to improve various factors such as naturalness, intelligibility, and robustness, we also propose the following techniques: 1) utilizing a self-supervised speech representation (SSL) and 2) randomly sampling the predicted and ground-truth features to alleviate prediction errors in variance adaptor. The experimental results demonstrate that the proposed method yields better performance in both objective and subjective evaluations compared to the conventional frame-based EL speech enhancement technique. Kazuhiro Kobayashi, Tomoki Hayashi, Tomoki Toda |
ICASSP | 1 |
| 2022 | An Investigation of Streaming Non-Autoregressive sequence-to-sequence Voice ConversionabstractRecent advances in sequence-to-sequence (S2S) models have improved the quality of voice conversion (VC), but it requires the entire sequence to perform inference, which prevents using it in real-time applications. To address this issue, this paper extends the non-autoregressive (NAR) S2S-VC model to enable us to perform streaming VC. We introduce streamable architectures such as causal convolution and self-attention with causal masking for the FastSpeech2-based NAR-S2S-VC model. The streamable architecture also tries to convert durations, which are kept as is in conventional real-time VC methods. To further improve the performance of the streaming VC model, we utilize an instant knowledge distillation with a dual-mode architecture, which performs non-causal and causal inference by sharing the network parameters. Through the experimental evaluation with Japanese parallel corpus, we investigate the impact on performance caused by the streamable architecture. The experimental results reveal that the use of future context frames increases latency, but it improves the conversion quality and that the difference in the speaking rate affects the performance of streaming inference. Tomoki Hayashi, Kazuhiro Kobayashi, Tomoki Toda |
ICASSP | 2 |
| 2022 | Two-Stage Training Method for Japanese Electrolaryngeal Speech Enhancement Based on Sequence-to-Sequence Voice ConversionabstractSequence-to-sequence (seq2seq) voice conversion (VC) models have greater potential in converting electrolaryngeal (EL) speech to normal speech (EL2SP) compared to conventional VC models. However, EL2SP based on seq2seq VC requires a sufficiently large amount of parallel data for the model training and it suffers from significant performance degradation when the amount of training data is insufficient. To address this issue, we suggest a novel, two-stage strategy to optimize the performance on EL2SP based on seq2seq VC when a small amount of the parallel dataset is available. In contrast to utilizing high-quality data augmentations in previous studies, we first combine a large amount of imperfect synthetic parallel data of EL and normal speech, with the original dataset into VC training. Then, a second stage training is conducted with the original parallel dataset only. The results show that the proposed method progressively improves the performance of EL2SP based on seq2seq VC. Lester Phillip Violeta, Kazuhiro Kobayashi, Tomoki Toda |
SLT | 3 |
| 2021 | Mandarin Electrolaryngeal Speech Voice Conversion with Sequence-to-Sequence ModelingabstractThe electrolaryngeal speech (EL speech) is typically spoken with an electrolarynx device that generates excitation signals to substitute human vocal fold vibrations. Because the excitation signals cannot perfectly characterize sound sources generated by vocal folds, the naturalness and intelligibility of the EL speech are inevitably worse than that of the natural speech (NL speech). To improve speech naturalness, statistical models, such as Gaussian mixture models and deep-learning-based models, have been employed for EL speech voice conversion (ELVC). The ELVC task aims to convert EL speech into NL speech through an ELVC model. To implement a frame-wise ELVC system, accurate feature alignment is crucial for model training. However, the abnormal acoustic characteristics of the EL speech cause misalignments and accordingly limit the ELVC performance. To address this issue, we propose a novel ELVC system based on sequence-to-sequence (seq2seq) modeling with text-to-speech (TTS) pretraining. The seq2seq model involves an attention mechanism to concurrently perform representation learning and alignment. Meanwhile, TTS pretraining provides efficient training with limited data. Experimental results show that the proposed ELVC system yields notable improvements in terms of standardized evaluation metrics and subjective listening tests over a well-known frame-wise ELVC system. Ming-Chi Yen, Wen-Chin Huang, Kazuhiro Kobayashi, Yu-Huai Peng, Shu-Wei Tsai, Yu Tsao 0001, Tomoki Toda, Jyh-Shing Roger Jang, Hsin-Min Wang |
ASRU | 3 |
| 2021 | Non-Autoregressive Sequence-To-Sequence Voice ConversionabstractThis paper proposes a novel voice conversion (VC) method based on non-autoregressive sequence-to-sequence (NAR-S2S) models. Inspired by the great success of NAR-S2S models such as FastSpeech in text-to-speech (TTS), we extend the FastSpeech2 model for the VC problem. We introduce the convolution-augmented Transformer (Conformer) instead of the Transformer, making it possible to capture both local and global context information from the input sequence. Furthermore, we extend variance predictors to variance converters to explicitly convert the source speaker’s prosody components such as pitch and energy into the target speaker. The experimental evaluation with the Japanese speaker dataset, which consists of male and female speakers of 1,000 utterances, demonstrates that the proposed model enables us to perform more stable, faster, and better conversion than autoregressive S2S (AR-S2S) models such as Tacotron2 and Transformer. Tomoki Hayashi, Wen-Chin Huang, Kazuhiro Kobayashi, Tomoki Toda |
ICASSP | 3 |
| 2021 | Crank: An Open-Source Software for Nonparallel Voice Conversion Based on Vector-Quantized Variational AutoencoderabstractIn this paper, we present an open-source software for developing a nonparallel voice conversion (VC) system named crank. Although we have released an open-source VC software based on the Gaussian mixture model named sprocket in the last VC Challenge, it is not straightforward to apply any speech corpus because it is necessary to prepare parallel utterances of source and target speakers to model a statistical conversion function. To address this issue, in this study, we developed a new open-source VC software that enables users to model the conversion function by using only a nonparallel speech corpus. For implementing the VC software, we used a vector-quantized variational autoencoder (VQVAE). To rapidly examine the effectiveness of recent technologies developed in this research field, crank also supports several representative works for autoencoder-based VC methods such as the use of hierarchical architectures, cyclic architectures, generative adversarial networks, speaker adversarial training, and neural vocoders. Moreover, it is possible to automatically estimate objective measures such as mel-cepstrum distortion and pseudo mean opinion score based on MOSNet. In this paper, we describe representative functions developed in crank and make brief comparisons by objective evaluations. Kazuhiro Kobayashi, Wen-Chin Huang, Yi-Chiao Wu, Patrick Lumban Tobing, Tomoki Hayashi, Tomoki Toda |
ICASSP | 1 |
| 2021 | A Preliminary Study of a Two-Stage Paradigm for Preserving Speaker Identity in Dysarthric Voice ConversionabstractWe propose a new paradigm for maintaining speaker identity in dysarthric voice conversion (DVC). The poor quality of dysarthric speech can be greatly improved by statistical VC, but as the normal speech utterances of a dysarthria patient are nearly impossible to collect, previous work failed to recover the individuality of the patient. In light of this, we suggest a novel, two-stage approach for DVC, which is highly flexible in that no normal speech of the patient is required. First, a powerful parallel sequence-to-sequence model converts the input dysarthric speech into a normal speech of a reference speaker as an intermediate product, and a nonparallel, frame-wise VC model realized with a variational autoencoder then converts the speaker identity of the reference speech back to that of the patient while assumed to be capable of preserving the enhanced quality. We investigate several design options. Experimental evaluation results demonstrate the potential of our approach to improving the quality of the dysarthric speech while maintaining the speaker identity. Wen-Chin Huang, Kazuhiro Kobayashi, Yu-Huai Peng, Ching-Feng Liu, Yu Tsao 0001, Hsin-Min Wang, Tomoki Toda |
Interspeech | 2 |
| 2021 | Quasi-Periodic WaveNet: An Autoregressive Raw Waveform Generative Model With Pitch-Dependent Dilated Convolution Neural NetworkabstractIn this paper, a pitch-adaptive waveform generative model named Quasi-Periodic WaveNet (QPNet) is proposed to improve the limited pitch controllability of vanilla WaveNet (WN) using pitch-dependent dilated convolution neural networks (PDCNNs). Specifically, as a probabilistic autoregressive generation model with stacked dilated convolution layers, WN achieves high-fidelity audio waveform generation. However, the pure-data-driven nature and the lack of prior knowledge of audio signals degrade the pitch controllability of WN. For instance, it is difficult for WN to precisely generate the periodic components of audio signals when the given auxiliary fundamental frequency (F0) features are outside the F0range observed in the training data. To address this problem, QPNet with two novel designs is proposed. First, the PDCNN component is applied to dynamically change the network architecture of WN according to the given auxiliary F0features. Second, a cascaded network structure is utilized to simultaneously model the long- and short-term dependencies of quasi-periodic signals such as speech. The performances of single-tone sinusoid and speech generations are evaluated. The experimental results show the effectiveness of the PDCNNs for unseen auxiliary F0features and the effectiveness of the cascaded structure for speech generation. Yi-Chiao Wu, Tomoki Hayashi, Patrick Lumban Tobing, Kazuhiro Kobayashi, Tomoki Toda |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | Efficient Shallow Wavenet Vocoder Using Multiple Samples Output Based on Laplacian Distribution and Linear PredictionabstractThis paper presents a novel way for an efficient implementation scheme of shallow WaveNet vocoder with multiple samples (segment) output based on the use of Laplacian distribution and linear prediction. In our previous work, we have proposed a shallow architecture for WaveNet vocoder that utilizes only 9 dilated convolutional layers while capable of generating high-quality speech with the use of Laplacian distribution in speech samples modeling. However, there is still a lot of room for improvements to increase the computation efficiency, such as by the inference of segment output and the use of a more compact structure. In this work, we tackle this issue by proposing a simple implementation scheme of segment output modeling, that can be easily extended into other neural vocoders, where the Laplacian distribution parameters of multiple samples are estimated simultaneously. Further, to preserve the dependencies of the samples within the segment, we also propose utilizing linear prediction (LP) to compute the distribution parameters, where data- driven LP-coefficients are estimated by the WaveNet vocoder along with locations and scales. Finally, a shallower WaveNet vocoder with 6 layers is deployed. The experimental results demonstrate that the proposed LP-based Laplacian distribution can alleviate the quality degradation caused by segment generation. Patrick Lumban Tobing, Yi-Chiao Wu, Tomoki Hayashi, Kazuhiro Kobayashi, Tomoki Toda |
ICASSP | 4 |
| 2020 | Intelligibility Enhancement Based on Speech Waveform Modification Using Hearing Impairment
Shu Hikosaka, Shogo Seki, Tomoki Hayashi, Kazuhiro Kobayashi, Kazuya Takeda, Hideki Banno, Tomoki Toda |
INTERSPEECH | 4 |
| 2020 | Cyclic Spectral Modeling for Unsupervised Unit Discovery into Voice Conversion with Excitation and Waveform ModelingabstractWe present a novel approach of cyclic spectral modeling for unsupervised discovery of speech units into voice conversion with excitation network and waveform modeling.Specifically, we propose two spectral modeling techniques: 1) cyclic vectorquantized autoencoder (CycleVQVAE), and 2) cyclic variational autoencoder (CycleVAE).In CycleVQVAE, a discrete latent space is used for the speech units, whereas, in Cycle-VAE, a continuous latent space is used.The cyclic structure is developed using the reconstruction flow and the cyclic reconstruction flow of spectral features, where the latter is obtained by recycling the converted spectral features.This method is used to obtain a possible speaker-independent latent space because of marginalization on all possible speaker conversion pairs during training.On the other hand, speaker-dependent space is conditioned with a one-hot speaker-code.Excitation modeling is developed in a separate manner for CycleVQVAE, while it is in a joint manner for CycleVAE.To generate speech waveform, WaveNet-based waveform modeling is used.The proposed framework is entried for the ZeroSpeech Challenge 2020, and is capable of reaching a character error rate of 0.21, a speaker similarity score of 3.91, a mean opinion score of 3.84 for the naturalness of the converted speech in the 2019 voice conversion task. Patrick Lumban Tobing, Tomoki Hayashi, Yi-Chiao Wu, Kazuhiro Kobayashi, Tomoki Toda |
INTERSPEECH | 4 |
| 2019 | Development of a Real-time Bionic Voice Generation System based on Statistical Excitation PredictionabstractDespite the emergent progress in many fields of bionics, larynx amputees still lack a functional Bionic Voice source to overcome their voice disability. We have established the Pneumatic Bionic Voice (PBV) as a promising technology to generate a voice for larynx amputees. This summary provides an overview of our efforts to implement the PBV system in real-time together with a demonstration of its results. The PBV is an electronic adaptation of an old school mechanical voice source called the Pneumatic Artificial Larynx (PAL). The PAL is a non-invasive voice prosthesis, with an exceptionally high-quality voice, driven exclusively by respiration. This work proposes an approach to statistically model the PAL's voice generation mechanism from the respiration and to implement it in a real-time PBV system that larynx amputees can use to generate voice in continuous speech. Farzaneh Ahmadi, Kazuhiro Kobayashi, Tomoki Toda |
ASSETS | 2 |
| 2019 | Voice Conversion with Cyclic Recurrent Neural Network and Fine-tuned Wavenet VocoderabstractThis paper presents a novel framework for providing high-quality parallel voice conversion (VC) using a cyclic recurrent neural network (RNN) and a finely tuned WaveNet vocoder. Using the proposed system, we are tackling the quality degradation issue faced by WaveNet when it is fed with estimated (oversmoothed) speech features, such as mel-cepstrum parameters predicted from a statistical model. In VC, providing predicted features to fine-tune a pretrained WaveNet model is not straightforward owing to the difference in time-sequence alignment. To overcome this problem, we propose the use of a cyclic spectral conversion network that is capable of performing a conversion flow, i.e., source-to-target, and a cyclic flow, i.e., generate self-predicted target speaker features, and is trained by using both the conversion and cyclic losses. The experimental results demonstrate that, overall, the proposed system significantly improves the converted speech, resulting in a mean opinion score of 3.79 and a speaker similarity score of 73.86%. Patrick Lumban Tobing, Yi-Chiao Wu, Tomoki Hayashi, Kazuhiro Kobayashi, Tomoki Toda |
ICASSP | 4 |
| 2019 | Investigation of F0 Conditioning and Fully Convolutional Networks in Variational Autoencoder Based Voice ConversionabstractIn this work, we investigate the effectiveness of two techniques for improving variational autoencoder (VAE) based voice conversion (VC). First, we reconsider the relationship between vocoder features extracted using the high quality vocoders adopted in conventional VC systems, and hypothesize that the spectral features are in fact F0 dependent. Such hypothesis implies that during the conversion phase, the latent codes and the converted features in VAE based VC are in fact source F0 dependent. To this end, we propose to utilize the F0 as an additional input of the decoder. The model can learn to disentangle the latent code from the F0 and thus generates converted F0 dependent converted features. Second, to better capture temporal dependencies of the spectral features and the F0 pattern, we replace the frame wise conversion structure in the original VAE based VC framework with a fully convolutional network structure. Our experiments demonstrate that the degree of disentanglement as well as the naturalness of the converted speech are indeed improved. Wen-Chin Huang, Yi-Chiao Wu, Chen-Chou Lo, Patrick Lumban Tobing, Tomoki Hayashi, Kazuhiro Kobayashi, Tomoki Toda, Yu Tsao 0001, Hsin-Min Wang |
INTERSPEECH | 6 |
| 2019 | Robustness of Statistical Voice Conversion Based on Direct Waveform Modification Against Background Sounds
Yusuke Kurita, Kazuhiro Kobayashi, Kazuya Takeda, Tomoki Toda |
INTERSPEECH | 2 |
| 2019 | Non-Parallel Voice Conversion with Cyclic Variational AutoencoderabstractIn this paper, we present a novel technique for a non-parallel voice conversion (VC) with the use of cyclic variational autoencoder (CycleVAE)-based spectral modeling.In a variational autoencoder (VAE) framework, a latent space, usually with a Gaussian prior, is used to encode a set of input features.In a VAE-based VC, the encoded latent features are fed into a decoder, along with speaker-coding features, to generate estimated spectra with either the original speaker identity (reconstructed) or another speaker identity (converted).Due to the non-parallel modeling condition, the converted spectra can not be directly optimized, which heavily degrades the performance of a VAEbased VC.In this work, to overcome this problem, we propose to use CycleVAE-based spectral model that indirectly optimizes the conversion flow by recycling the converted features back into the system to obtain corresponding cyclic reconstructed spectra that can be directly optimized.The cyclic flow can be continued by using the cyclic reconstructed features as input for the next cycle.The experimental results demonstrate the effectiveness of the proposed CycleVAE-based VC, which yields higher accuracy of converted spectra, generates latent features with higher correlation degree, and significantly improves the quality and conversion accuracy of the converted speech. Patrick Lumban Tobing, Yi-Chiao Wu, Tomoki Hayashi, Kazuhiro Kobayashi, Tomoki Toda |
INTERSPEECH | 4 |
| 2019 | Quasi-Periodic WaveNet Vocoder: A Pitch Dependent Dilated Convolution Model for Parametric Speech GenerationabstractIn this paper, we propose a quasi-periodic neural network (QPNet) vocoder with a novel network architecture named pitch-dependent dilated convolution (PDCNN) to improve the pitch controllability of WaveNet (WN) vocoder. The effectiveness of the WN vocoder to generate high-fidelity speech samples from given acoustic features has been proved recently. However, because of the fixed dilated convolution and generic network architecture, the WN vocoder hardly generates speech with given F0 values which are outside the range observed in training data. Consequently, the WN vocoder lacks the pitch controllability which is one of the essential capabilities of conventional vocoders. To address this limitation, we propose the PDCNN component which has the time-variant adaptive dilation size related to the given F0 values and a cascade network structure of the QPNet vocoder to generate quasi-periodic signals such as speech. Both objective and subjective tests are conducted, and the experimental results demonstrate the better pitch controllability of the QPNet vocoder compared to the same and double sized WN vocoders while attaining comparable speech qualities. Index Terms: WaveNet, vocoder, quasi-periodic signal, pitch-dependent dilated convolution, pitch controllability Yi-Chiao Wu, Tomoki Hayashi, Patrick Lumban Tobing, Kazuhiro Kobayashi, Tomoki Toda |
INTERSPEECH | 4 |
| 2018 | Collapsed Speech Segment Detection and Suppression for WaveNet VocoderabstractIn this paper, we propose a technique to alleviate the quality degradation caused by collapsed speech segments sometimes generated by the WaveNet vocoder.The effectiveness of the WaveNet vocoder for generating natural speech from acoustic features has been proved in recent works.However, it sometimes generates very noisy speech with collapsed speech segments when only a limited amount of training data is available or significant acoustic mismatches exist between the training and testing data.Such a limitation on the corpus and limited ability of the model can easily occur in some speech generation applications, such as voice conversion and speech enhancement.To address this problem, we propose a technique to automatically detect collapsed speech segments.Moreover, to refine the detected segments, we also propose a waveform generation technique for WaveNet using a linear predictive coding constraint.Verification and subjective tests are conducted to investigate the effectiveness of the proposed techniques.The verification results indicate that the detection technique can detect most collapsed segments.The subjective evaluations of voice conversion demonstrate that the generation technique significantly improves the speech quality while maintaining the same speaker similarity. Yi-Chiao Wu, Kazuhiro Kobayashi, Tomoki Hayashi, Patrick Lumban Tobing, Tomoki Toda |
INTERSPEECH | 2 |
| 2018 | An Evaluation of Deep Spectral Mappings and WaveNet Vocoder for Voice ConversionabstractThis paper presents an evaluation of deep spectral mapping and WaveNet vocoder in voice conversion (VC). In our VC framework, spectral features of an input speaker are converted into those of a target speaker using the deep spectral mapping, and then together with the excitation features, the converted waveform is generated using WaveNet vocoder. In this work, we compare three different deep spectral mapping networks, i.e., a deep single density network (DSDN), a deep mixture density network (DMDN), and a long short-term memory recurrent neural network with an autoregressive output layer (LSTM-AR). Moreover, we also investigate several methods for reducing mismatches of spectral features used in WaveNet vocoder between training and conversion processes, such as some methods to alleviate oversmoothing effects of the converted spectral features, and another method to refine WaveNet using the converted spectral features. The experimental results demonstrate that the LSTM-AR yields nearly better spectral mapping accuracy than the others, and the proposed WaveNet refinement method significantly improves the naturalness of the converted waveform. Patrick Lumban Tobing, Tomoki Hayashi, Yi-Chiao Wu, Kazuhiro Kobayashi, Tomoki Toda |
SLT | 4 |
| 2018 | Intra-gender statistical singing voice conversion with direct waveform modification using log-spectral differentialabstractThis paper presents a novel intra-gender statistical singing voice conversion (SVC) technique with direct waveform modification based on the log-spectrum differential (DIFFSVC) that can convert the voice timbre of a source singer into that of a target singer without vocoder-based waveform generation of the converted singing voice. SVC makes it possible to convert the singing voice characteristics of an arbitrary source singer into those of an arbitrary target singer by converting some of its acoustic features, such as F0, aperiodicity, and spectral features based on a statistical conversion function. However, the sound quality of the converted singing voice is typically degraded compared with that of a natural singing voice, owing to various factors, such as analysis and modeling errors in the vocoding process and over-smoothing of the converted feature trajectory. To alleviate sound quality degradation, we propose a statistical conversion process that directly modifies the signal in the waveform domain by estimating the difference in the spectra of the source and target singers’ singing voices. Additionally, we propose the following several techniques for the DIFFSVC method: 1) derivation of a differential Gaussian mixture model (DIFFGMM) from a conventional Gaussian mixture model (GMM) and 2) a parameter generation algorithm considering the global variance (GV). The experimental results demonstrate that the proposed DIFFSVC methods enable significant improvements in the sound quality of the converted singing voice, while preserving the conversion accuracy of the singer’s identity compared with conventional SVC. Kazuhiro Kobayashi, Tomoki Toda, Satoshi Nakamura 0001 |
Speech Commun. | 1 |
| 2017 | An investigation of multi-speaker training for wavenet vocoderabstractIn this paper, we investigate the effectiveness of multi-speaker training for WaveNet vocoder. In our previous work, we have demonstrated that our proposed speaker-dependent (SD) WaveNet vocoder, which is trained with a single speaker's speech data, is capable of modeling temporal waveform structure, such as phase information, and makes it possible to generate more naturally sounding synthetic voices compared to conventional high-quality vocoder, STRAIGHT. However, it is still difficult to generate synthetic voices of various speakers using the SD-WaveNet due to its speaker-dependent property. Towards the development of speaker-independent WaveNet vocoder, we apply multi-speaker training techniques to the WaveNet vocoder and investigate its effectiveness. The experimental results demonstrate that 1) the multispeaker WaveNet vocoder still outperforms STRAIGHT in generating known speakers' voices but it is comparable to STRAIGHT in generating unknown speakers' voices, and 2) the multi-speaker training is effective for developing the WaveNet vocoder capable of speech modification. Tomoki Hayashi, Akira Tamamori, Kazuhiro Kobayashi, Kazuya Takeda, Tomoki Toda |
ASRU | 3 |
| 2017 | Statistical Voice Conversion with WaveNet-Based Waveform Generation
Kazuhiro Kobayashi, Tomoki Hayashi, Akira Tamamori, Tomoki Toda |
INTERSPEECH | 1 |
| 2017 | Speaker-Dependent WaveNet Vocoder
Akira Tamamori, Tomoki Hayashi, Kazuhiro Kobayashi, Kazuya Takeda, Tomoki Toda |
INTERSPEECH | 3 |
| 2017 | Articulatory Controllable Speech Modification Based on Statistical Inversion and Production MappingsabstractIn this paper, we present an innovative way of utilizing the natural relationship between speech sounds and articulatory movements by developing an articulatory controllable speech modification system. Specifically, we employ statistical acoustic-to-articulatory inversion mapping and articulatory-to-acoustic production mapping based on a Gaussian mixture model, allowing flexible modification of the model parameters and the independence of the text input features. Modification of an input speech signal through manipulation of the unobserved articulatory movements is achievable through a sequence of inversion and production mappings. To ensure the naturalness of articulatory movement trajectories, we introduce a method for manipulating articulatory parameters by considering their intercorrelation. Moreover, to generate high-quality modified speech sounds, we avoid the use of vocoder-based excitation generation by presenting several implementations of direct waveform modification capable of directly filtering an input speech signal using the differences in spectral parameters. The experimental results demonstrate that: 1) higher accuracy in the estimation of spectral parameters is achieved by using sequential inversion and production mappings than for conventional production mapping using measured articulatory parameters, 2) the method for manipulating articulatory parameters by considering their intercorrelation makes it possible to generate more natural trajectories of modified articulatory movements; 3) the implementations of the direct waveform modification method significantly improve the quality of modified speech sounds, even under varying speaking conditions; and 4) the controllability of the system is ensured by its capability of producing modified vowel sounds through the manipulation of appropriate articulatory configurations. Patrick Lumban Tobing, Kazuhiro Kobayashi, Tomoki Toda |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Implementation of F0 transformation for statistical singing voice conversion based on direct waveform modificationabstractThis paper presents a technique for transforming F0in a framework of statistical singing voice conversion with direct waveform modification based on spectrum differential (DIFFSVC). The DIFFSVC method converts voice timbre of singing voices of a source singer into that of a target singer without using vocoder-based waveform generation. Although this method achieves high sound quality of the converted singing voices, its use is limited to only intra-gender conversion without the need of F0transformation. To make it possible to also use the DIFFSVC method for cross-gender conversion, we propose a method to transform F0of an input singing voice for the DIFFSVC. The proposed method is also based on direct waveform modification using overlap-add process and filtering process. Results of subjective evaluations demonstrate that the proposed DIFFSVC method with F0transformation significantly improves sound quality of the converted singing voices while preserving the conversion accuracy of singer identity in the cross-gender conversion compared to the conventional SVC with vocoder. Kazuhiro Kobayashi, Tomoki Toda, Satoshi Nakamura 0001 |
ICASSP | 1 |
| 2016 | An estimation method of voice timbre evaluation values using feature extraction with Gaussian mixture model based on reference singerabstractThis paper presents an estimation method of voice timbre evaluation values for arbitrary singer's singing voices generated with a singing voice synthesis system towards the development of a singing voice retrieval system. The voice timbre evaluation values are numerical values corresponding to voice timbre expression words, such as "Age" and "Gender", and they usually need to be manually assigned to individual singers' singing voices through listening. To make it possible to automatically estimate them from given singer's singing voices, an acoustic feature to well capture only each singer's voice timbre is extracted with a Gaussian mixture model trained using parallel data between singing voices sung by many pre-stored target singers and same voices sung by a reference singer. Then, the voice timbre evaluation values are estimated from the extracted feature using regression models. The experimental results showed that the proposed method is capable of accurately estimating those values for some expression words, such as "Age" and "Gender", and nonlinear regression is effective for the expression words, "Powerfulness" and "Uniqueness." Soichi Yamane, Kazuhiro Kobayashi, Tomoki Toda, Tomoyasu Nakano, Masataka Goto, Satoshi Nakamura 0001 |
ICASSP | 2 |
| 2016 | The NU-NAIST Voice Conversion System for the Voice Conversion Challenge 2016
Kazuhiro Kobayashi, Shinnosuke Takamichi, Satoshi Nakamura 0001, Tomoki Toda |
INTERSPEECH | 1 |
| 2016 | F0 transformation techniques for statistical voice conversion with direct waveform modification with spectral differentialabstractThis paper presents several F0transformation techniques for statistical voice conversion (VC) with direct waveform modification with spectral differential (DIFFVC). Statistical VC is a technique to convert speaker identity of a source speaker's voice into that of a target speaker by converting several acoustic features, such as spectral and excitation features. This technique usually uses vocoder to generate converted speech waveforms from the converted acoustic features. However, the use of vocoder often causes speech quality degradation of the converted voice owing to insufficient parameterization accuracy. To avoid this issue, we have proposed a direct waveform modification technique based on spectral differential filtering and have successfully applied it to intra-gender singing VC (DIFFSVC) where excitation features are not necessary converted. Moreover, we have also applied it to cross-gender singing VC by implementing F0transformation with a constant rate such as one octave increase or decrease. On the other hand, it is not straightforward to apply the DIFFSVC framework to normal speech conversion because the F0transformation ratio widely varies depending on a combination of the source and target speakers. In this paper, we propose several F0transformation techniques for DIFFVC and compare their performance in terms of speech quality of the converted voice and conversion accuracy of speaker individuality. The experimental results demonstrate that the F0transformation technique based on waveform modification achieves the best performance among the proposed techniques. Kazuhiro Kobayashi, Tomoki Toda, Satoshi Nakamura 0001 |
SLT | 1 |
| 2015 | Statistical singing voice conversion based on direct waveform modification with global varianceabstractThis paper presents techniques to improve the quality of voices generated through statistical singing voice conversion with direct waveform modification based on spectrum differential (DIFFSVC). The DIFFSVC method makes it possible to convert singing voice characteristics of a source singer into those of a target singer without using vocoder-based waveform generation. However, quality of the converted singing voice still degrades compared to that of a natural singing voice due to various factors, such as the over-smoothing of the converted spectral parameter trajectory. To alleviate this over-smoothing, we propose a technique to restore the global variance of the converted spectral parameter trajectory within the framework of the DIFFSVC method. We also propose another technique to specifically avoid over-smoothing at unvoiced frames. Results of subjective and objective evaluations demonstrate that the proposed techniques significantly improve speech quality of the converted singing voice while preserving the conversion accuracy of singer identity compared to the conventional DIFFSVC. Kazuhiro Kobayashi, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 1 |
| 2015 | Articulatory controllable speech modification based on Gaussian mixture models with direct waveform modification using spectrum differentialabstractIn our previous work, we have developed a speech modification system capable of manipulating unobserved articulatory movements by sequentially performing speech-to-articulatory inversion mapping and articulatory-to-speech production mapping based on a Gaussian mixture model (GMM)-based statistical feature mapping technique. One of the biggest issues to be addressed in this system is quality degradation of the synthetic speech caused by modeling and conversion errors in a vocoderbased waveform generation framework. To address this issue, we propose several implementation methods of direct waveform modification. The proposed methods directly filter an input speech waveform with a time sequence of spectral differential parameters calculated between unmodified and modified spectral envelop parameters in order to avoid using vocoderbased excitation signal generation. The experimental results show that the proposed direct waveform modification methods yield significantly larger quality improvements in the synthetic speech while also keeping a capability of intuitively modifying phoneme sounds by manipulating the unobserved articulatory movements. Patrick Lumban Tobing, Kazuhiro Kobayashi, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2014 | Regression approaches to perceptual age control in singing voice conversionabstractThe perceptual age of a singing voice is the age of the singer as perceived by the listener, and is one of the notable characteristics that determines perceptions of a song. In this paper, we describe a novel voice timbre control technique based on the perceptual age for singing voice conversion (SVC). Singers can sing expressively by controlling prosody and voice timbre, but the varieties of voices that singers can produce are limited by physical constraints. Previous work has attempted to overcome the limitation through the use of statistical voice conversion. This technique makes it possible to convert singing voice timbre of an arbitrary source singer into that of an arbitrary target singer. However, it is still difficult to intuitively control singing voice characteristics by manipulating parameters corresponding to specific physical traits, such as gender and age. In this paper, we develop a technique for controlling the voice timbre based on perceptual age that maintains the singer's individuality. The experimental results show that the proposed voice timbre control method makes it possible to change the singer's perceptual age while not having an adverse effect on the perceived individuality. Kazuhiro Kobayashi, Tomoki Toda, Tomoyasu Nakano, Masataka Goto, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
ICASSP | 1 |
| 2014 | Statistical singing voice conversion with direct waveform modification based on the spectrum differentialabstractThis paper presents a novel statistical singing voice conversion (SVC) technique with direct waveform modification based on the spectrum differential that can convert voice timbre of a source singer into that of a target singer without using a vocoder to generate converted singing voice waveforms. SVC makes it possible to convert singing voice characteristics of an arbitrary source singer into those of an arbitrary target singer. However, speech quality of the converted singing voice is significantly degraded compared to that of a natural singing voice due to various factors, such as analysis and modeling errors in the vocoderbased framework. To alleviate this degradation, we propose a statistical conversion process that directly modifies the signal in the waveform domain by estimating the difference in the spectra of the source and target singers’ singing voices. The differential spectral feature is directly estimated using a differential Gaussian mixture model (GMM) that is analytically derived from the traditional GMM used as a conversion model in the conventional SVC. The experimental results demonstrate that the proposed method makes it possible to significantly improve speech quality in the converted singing voice while preserving the conversion accuracy of singer identity compared to the conventional SVC. Index Terms: singing voice, statistical voice conversion, vocoder, Gaussian mixture model, differential spectral compensation Kazuhiro Kobayashi, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 1 |
| 2013 | An investigation of acoustic features for singing voice conversion based on perceptual ageabstractIn this paper, we investigate the acoustic features that can be modified to control the perceptual age of a singing voice. Singers can sing expressively by controlling prosody and vocal timbre, but the varieties of voices that singers can produce are limited by physical constraints. Previous work has attempted to overcome this limitation through the use of statistical voice conversion. This technique makes it possible to convert singing voice characteristics of an arbitrary source singer into those of an arbitrary target singer. However, it is still difficult to intu-itively control singing voice characteristics by manipulating pa-rameters corresponding to specific physical traits, such as gen-der and age. In this paper, we focus on controlling the perceived age of the singer and, as a first step, perform an investigation of the factors that play a part in the listener’s perception of the singer’s age. The experimental results demonstrate that 1) the perceptual age of singing voices corresponds relatively well to the actual age of the singer, 2) speech analysis/synthesis pro-cessing and statistical voice conversion processing don’t cause adverse effects on the perceptual age of singing voices, and 3) prosodic features have a larger effect on the perceptual age than spectral features. Kazuhiro Kobayashi, Hironori Doi, Tomoki Toda, Tomoyasu Nakano, Masataka Goto, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 1 |
| 2002 | The Valence Patterns of Japanese Verbs Extracted From The EDR Corpus
Takano Ogino, Hitoshi Isahara, Kazuhiro Kobayashi |
LREC | 3 |