EDBT 2026 Demo / reviewers in the wild / expert
Kou Tanaka
dblp:140/2727
· DBLP profile ↗
37ranked-venue papers
12as first author
19since 2021 · last 2025
0009-0003-7107-607XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 11 first-author · 17 since 2021Artificial intelligence and machine learning · 21 · 5 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Rethinking Mean Opinion Scores in Speech Quality Assessment: Score Aggregation through Quantized Distribution FittingabstractThis study addresses the task of speech quality assessment (SQA), which aims to automatically predict the subjective quality of a given speech. Recent efforts have focused on training neural-based models to predict the mean opinion score (MOS) of speech samples produced by text-to-speech or voice conversion systems. We aim to enhance the performance of the models by a score aggregation method instead of MOS. The proposed method mitigates the effects of some issues arising from constraints imposed by limited options. Our method assumes annotators internally consider continuous scores and pick the nearest discrete rating. By modeling this process, we approximate the rating distribution by quantizing the latent continuous distribution. We then use the peak of the latent distribution, estimated through the loss between the distribution and actual ratings, as the new value instead of MOS. Experimental results demonstrate that substituting MOSNet’s target with this proposed value improves prediction performance. Yuto Kondo, Hirokazu Kameoka, Kou Tanaka, Takuhiro Kaneko |
ICASSP | 3 |
| 2025 | FasterVoiceGrad: Faster One-step Diffusion-Based Voice Conversion with Adversarial Diffusion Conversion Distillation
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Yuto Kondo |
INTERSPEECH | 3 |
| 2025 | Vocoder-Projected Feature Discriminator
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Yuto Kondo |
INTERSPEECH | 3 |
| 2025 | JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles
Yuto Kondo, Hirokazu Kameoka, Kou Tanaka, Takuhiro Kaneko |
INTERSPEECH | 3 |
| 2024 | Training Generative Adversarial Network-Based Vocoder with Limited Data Using Augmentation-Conditional DiscriminatorabstractA generative adversarial network (GAN)-based vocoder trained with an adversarial discriminator is commonly used for speech synthesis because of its fast, lightweight, and high-quality characteristics. However, this data-driven model requires a large amount of training data incurring high data-collection costs. This fact motivates us to train a GAN-based vocoder on limited data. A promising solution is to augment the training data to avoid overfitting. However, a standard discriminator is unconditional and insensitive to distributional changes caused by data augmentation. Thus, augmented speech (which can be extraordinary) may be considered real speech. To address this issue, we propose an augmentation-conditional discriminator (AugCondD) that receives the augmentation state as input in addition to speech, thereby assessing input speech according to augmentation state, without inhibiting the learning of the original non-augmented distribution. Experimental results indicate that AugCondD improves speech quality under limited data conditions while achieving comparable speech quality under sufficient data conditions.1 Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka |
ICASSP | 3 |
| 2024 | Selecting N-Lowest Scores for Training MOS Prediction ModelsabstractThe automatic speech quality assessment (SQA) has been extensively studied to predict the speech quality without time-consuming questionnaires. Recently, neural-based SQA models have been actively developed for speech samples produced by text-to-speech or voice conversion, with a primary focus on training mean opinion score (MOS) prediction models. The quality of each speech sample may not be consistent across the entire duration, and it remains unclear which segments of the speech receive the primary focus from humans when assigning subjective evaluation for MOS calculation. We hypothesize that when humans rate speech, they tend to assign more weight to low-quality speech segments, and the variance in ratings for each sample is mainly due to accidental assignment of higher scores when overlooking the poor quality speech segments. Motivated by the hypothesis, we analyze the VCC2018 and BVCC datasets. Based on the hypothesis, we propose the more reliable representative value Nlow-MOS, the mean of the N-lowest opinion scores. Our experiments show that LCC and SRCC improve compared to regular MOS when employing Nlow-MOS to MOSNet training. This result suggests that Nlow-MOS is a more intrinsic representative value of subjective speech quality and makes MOSNet a better comparator of VC models. Yuto Kondo, Hirokazu Kameoka, Kou Tanaka, Takuhiro Kaneko |
ICASSP | 3 |
| 2024 | FastVoiceGrad: One-step Diffusion-Based Voice Conversion with Adversarial Conditional Diffusion Distillation
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Yuto Kondo |
INTERSPEECH | 3 |
| 2024 | PRVAE-VC2: Non-Parallel Voice Conversion by Distillation of Speech Representations
Kou Tanaka, Hirokazu Kameoka, Takuhiro Kaneko, Yuto Kondo |
INTERSPEECH | 1 |
| 2024 | VoiceGrad: Non-Parallel Any-to-Many Voice Conversion With Annealed Langevin DynamicsabstractIn this paper, we propose a non-parallel any-to-many voice conversion (VC) method termedVoiceGrad. Inspired by WaveGrad, a recently introduced novel waveform generation method, VoiceGrad is based upon the concepts of score matching, Langevin dynamics, and diffusion models. The idea involves training a score approximator, a fully convolutional network with a U-Net structure, to predict the gradient of the log density of the speech feature sequences of multiple speakers. The trained score approximator can be used to perform VC by using annealed Langevin dynamics or reverse diffusion process to iteratively update an input feature sequence towards the nearest stationary point of the target distribution. Thanks to the nature of this concept, VoiceGrad enables any-to-many VC, a VC scenario in which the speaker of input speech can be arbitrary, and allows for non-parallel training, which requires no parallel utterances. Hirokazu Kameoka, Takuhiro Kaneko, Kou Tanaka, Nobukatsu Hojo, Shogo Seki |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Wave-U-Net Discriminator: Fast and Lightweight Discriminator for Generative Adversarial Network-Based Speech SynthesisabstractIn speech synthesis, a generative adversarial network (GAN), training a generator (speech synthesizer) and a discriminator in a min-max game, is widely used to improve speech quality. An ensemble of discriminators is commonly used in recent neural vocoders (e.g., HiFi-GAN) and end-to-end text-to-speech (TTS) systems (e.g., VITS) to scrutinize waveforms from multiple perspectives. Such discriminators allow synthesized speech to adequately approach real speech; however, they require an increase in the model size and computation time according to the increase in the number of discriminators. Alternatively, this study proposes a Wave-U-Net discriminator, which is a single but expressive discriminator with Wave-U-Net architecture. This discriminator is unique; it can assess a waveform in a sample-wise manner with the same resolution as the input signal, while extracting multilevel features via an encoder and decoder with skip connections. This architecture provides a generator with sufficiently rich information for the synthesized speech to be closely matched to the real speech. During the experiments, the proposed ideas were applied to a representative neural vocoder (HiFi-GAN) and an end-to-end TTS system (VITS). The results demonstrate that the proposed models can achieve comparable speech quality with a 2.31 times faster and 14.5 times more lightweight discriminator when used in HiFi-GAN and a 1.90 times faster and 9.62 times more lightweight discriminator when used in VITS.1 Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Shogo Seki |
ICASSP | 3 |
| 2023 | JSV-VC: Jointly Trained Speaker Verification and Voice Conversion ModelsabstractThis paper proposes a variational autoencoder (VAE)-based method for voice conversion (VC) on arbitrary source-target speaker pairs without parallel corpora, i.e., non-parallel any-to-any VC. One typical approach is to use speaker embeddings obtained from a speaker verification (SV) model as the condition for a VC model. However, converted speech is not guaranteed to reflect a target speaker’s characteristics in a naive combination of VC and SV models. Moreover, speaker embeddings are not designed for VC problems, leading to suboptimal conversion performance. To address these issues, the proposed method, JSV-VC, trains both VC and SV models jointly. The VC model is trained so that converted speech is verified as the target speaker in the SV model, while the SV model is trained in order to output consistent embeddings before and after the VC model. The experimental evaluation reveals that JSV-VC outperforms conventional any-to-any VC methods quantitatively and qualitatively. Shogo Seki, Hirokazu Kameoka, Kou Tanaka, Takuhiro Kaneko |
ICASSP | 3 |
| 2023 | iSTFTNet2: Faster and More Lightweight iSTFT-Based Neural Vocoder Using 1D-2D CNN
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Shogo Seki |
INTERSPEECH | 3 |
| 2023 | CFVC: Conditional Filtering for Controllable Voice Conversion
Kou Tanaka, Takuhiro Kaneko, Hirokazu Kameoka, Shogo Seki |
INTERSPEECH | 1 |
| 2022 | ISTFTNET: Fast and Lightweight Mel-Spectrogram Vocoder Incorporating Inverse Short-Time Fourier TransformabstractIn recent text-to-speech synthesis and voice conversion systems, a mel-spectrogram is commonly applied as an intermediate representation, and the necessity for a mel-spectrogram vocoder is increasing. A mel-spectrogram vocoder must solve three inverse problems: recovery of the original-scale magnitude spectrogram, phase reconstruction, and frequency-to-time conversion. A typical convolutional mel-spectrogram vocoder solves these problems jointly and implicitly using a convolutional neural network, including temporal upsampling layers, when directly calculating a raw waveform. Such an approach allows skipping redundant processes during waveform synthesis (e.g., the direct reconstruction of high-dimensional original-scale spectrograms). By contrast, the approach solves all problems in a black box and cannot effectively employ the time-frequency structures existing in a mel-spectrogram. We thus propose iSTFTNet, which replaces some output-side layers of the mel-spectrogram vocoder with the inverse short-time Fourier transform (iSTFT) after sufficiently reducing the frequency dimension using upsampling layers, reducing the computational cost from black-box modeling and avoiding redundant estimations of high-dimensional spectrograms. During our experiments, we applied our ideas to three HiFi-GAN variants and made the models faster and more lightweight with a reasonable speech quality.1 Takuhiro Kaneko, Kou Tanaka, Hirokazu Kameoka, Shogo Seki |
ICASSP | 2 |
| 2022 | CAUSE: Crossmodal Action Unit Sequence Estimation from Speech
Hirokazu Kameoka, Takuhiro Kaneko, Shogo Seki, Kou Tanaka |
INTERSPEECH | 4 |
| 2022 | MISRNet: Lightweight Neural Vocoder Using Multi-Input Single Shared Residual Blocks
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Shogo Seki |
INTERSPEECH | 3 |
| 2022 | Distilling Sequence-to-Sequence Voice Conversion Models for Streaming Conversion ApplicationsabstractThis paper describes a method for distilling a recurrent-based sequence-to-sequence (S2S) voice conversion (VC) model. Although the performance of recent VCs is becoming higher quality, streaming conversion is still a challenge when considering practical applications. To achieve streaming VC, the conversion model needs a streamable structure, a causal layer rather than a non-causal layer. Motivated by this constraint and recent advances in S2S learning, we apply the teacher-student framework to recurrent-based S2S- VC models. A major challenge is how to minimize degradation due to the use of causal layers which masks future input information. Experimental evaluations show that except for male-to-female speaker conversion, our approach is able to maintain the teacher model's performance in terms of subjective evaluations despite the streamable student model structure. Audio samples can be accessed on http://www.kecl.ntt.co.jp/people/tanaka.ko/projects/dists2svc. Kou Tanaka, Hirokazu Kameoka, Takuhiro Kaneko, Shogo Seki |
SLT | 1 |
| 2021 | Maskcyclegan-VC: Learning Non-Parallel Voice Conversion with Filling in FramesabstractNon-parallel voice conversion (VC) is a technique for training voice converters without a parallel corpus. Cycle-consistent adversarial network-based VCs (CycleGAN-VC and CycleGAN-VC2) are widely accepted as benchmark methods. However, owing to their insufficient ability to grasp time-frequency structures, their application is limited to mel-cepstrum conversion and not mel-spectrogram conversion despite recent advances in mel-spectrogram vocoders. To overcome this, CycleGAN-VC3, an improved variant of CycleGAN-VC2 that incorporates an additional module called time-frequency adaptive normalization (TFAN), has been proposed. However, an increase in the number of learned parameters is imposed. As an alternative, we propose MaskCycleGAN-VC, which is another extension of CycleGAN-VC2 and is trained using a novel auxiliary task called filling in frames (FIF). With FIF, we apply a temporal mask to the input mel-spectrogram and encourage the converter to fill in missing frames based on surrounding frames. This task allows the converter to learn time-frequency structures in a self-supervised manner and eliminates the need for an additional module such as TFAN. A subjective evaluation of the naturalness and speaker similarity showed that MaskCycleGAN-VC outperformed both CycleGAN-VC2 and CycleGAN-VC3 with a model size similar to that of CycleGAN-VC2.1 Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Nobukatsu Hojo |
ICASSP | 3 |
| 2021 | Many-to-Many Voice Transformer NetworkabstractThis paper proposes a voice conversion (VC) method based on a sequence-to-sequence (S2S) learning framework, which enables simultaneous conversion of the voice characteristics, pitch contour, and duration of input speech. We previously proposed an S2S-based VC method using a transformer network architecture called the voice transformer network (VTN). The original VTN was designed to learn only a mapping of speech feature sequences from one speaker to another. Here, the main idea we propose is an extension of the original VTN that can simultaneously learn mappings among multiple speakers. This extension, called the many-to-many VTN, enables us to fully use available training data collected from multiple speakers by capturing common latent features that can be shared across different speakers. It also allows us to introduce a training loss called the identity mapping loss to ensure that the input feature sequence will remain unchanged when the source and target speaker indices are the same. Using this particular loss for model training has been found to be extremely effective in improving the performance of the model at test time. We conducted speaker identity conversion experiments and found that our model obtained higher sound quality and speaker similarity than baseline methods. We also found that our model, with a slight modification to its architecture, can handle any-to-many conversion tasks reasonably well. Hirokazu Kameoka, Wen-Chin Huang, Kou Tanaka, Takuhiro Kaneko, Nobukatsu Hojo, Tomoki Toda |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | CycleGAN-VC3: Examining and Improving CycleGAN-VCs for Mel-Spectrogram ConversionabstractNon-parallel voice conversion (VC) is a technique for learning mappings between source and target speeches without using a parallel corpus.Recently, cycle-consistent adversarial network (CycleGAN)-VC and CycleGAN-VC2 have shown promising results regarding this problem and have been widely used as benchmark methods.However, owing to the ambiguity of the effectiveness of CycleGAN-VC/VC2 for mel-spectrogram conversion, they are typically used for mel-cepstrum conversion even when comparative methods employ mel-spectrogram as a conversion target.To address this, we examined the applicability of CycleGAN-VC/VC2 to mel-spectrogram conversion.Through initial experiments, we discovered that their direct applications compromised the time-frequency structure that should be preserved during conversion.To remedy this, we propose CycleGAN-VC3, an improvement of CycleGAN-VC2 that incorporates time-frequency adaptive normalization (TFAN).Using TFAN, we can adjust the scale and bias of the converted features while reflecting the time-frequency structure of the source mel-spectrogram.We evaluated CycleGAN-VC3 on inter-gender and intra-gender non-parallel VC.A subjective evaluation of naturalness and similarity showed that for every VC pair, CycleGAN-VC3 outperforms or is competitive with the two types of CycleGAN-VC2, one of which was applied to mel-cepstrum and the other to mel-spectrogram. 1 Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Nobukatsu Hojo |
INTERSPEECH | 3 |
| 2020 | Collection and Analysis of Dialogues Provided by Two Speakers Acting as OneabstractTsunehiro Arimoto, Ryuichiro Higashinaka, Kou Tanaka, Takahito Kawanishi, Hiroaki Sugiyama, Hiroshi Sawada, Hiroshi Ishiguro. Proceedings of the 21th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2020. Tsunehiro Arimoto, Ryuichiro Higashinaka, Kou Tanaka, Takahito Kawanishi, Hiroaki Sugiyama, Hiroshi Sawada, Hiroshi Ishiguro |
SIGdial | 3 |
| 2020 | Nonparallel Voice Conversion With Augmented Classifier Star Generative Adversarial NetworksabstractWe previously proposed a method that allows for nonparallel voice conversion (VC) by using a variant of generative adversarial networks (GANs) called StarGAN. The main features of our method, called StarGAN-VC, are as follows: First, it requires no parallel utterances, transcriptions, or time alignment procedures for speech generator training. Second, it can simultaneously learn mappings across multiple domains using a single generator network and thus fully exploit available training data collected from multiple domains to capture latent features that are common to all the domains. Third, it can generate converted speech signals quickly enough to allow real-time implementations and requires only several minutes of training examples to generate reasonably realistic-sounding speech. In this article, we describe three formulations of StarGAN, including a newly introduced novel StarGAN variant called “Augmented classifier StarGAN (A-StarGAN)”, and compare them in a nonparallel VC task. We also compare them with several baseline methods. Hirokazu Kameoka, Takuhiro Kaneko, Kou Tanaka, Nobukatsu Hojo |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | ConvS2S-VC: Fully Convolutional Sequence-to-Sequence Voice ConversionabstractThis article proposes a voice conversion (VC) method using sequence-to-sequence (seq2seq or S2S) learning, which flexibly converts not only the voice characteristics but also the pitch contour and duration of input speech. The proposed method, called ConvS2S-VC, has three key features. First, it uses a model with a fully convolutional architecture. This is particularly advantageous in that it is suitable for parallel computations using GPUs. It is also beneficial since it enables effective normalization techniques such as batch normalization to be used for all the hidden layers in the networks. Second, it achieves many-to-many conversion by simultaneously learning mappings among multiple speakers using only a single model instead of separately learning mappings between each speaker pair using a different model. This enables the model to fully utilize available training data collected from multiple speakers by capturing common latent features that can be shared across different speakers. Owing to this structure, our model works reasonably well even without source speaker information, thus making it able to handle any-to-many conversion tasks. Third, we introduce a mechanism, called the conditional batch normalization that switches batch normalization layers in accordance with the target speaker. This particular mechanism has been found to be extremely effective for our many-to-many conversion model. We conducted speaker identity conversion experiments and found that ConvS2S-VC obtained higher sound quality and speaker similarity than baseline methods. We also found from audio examples that it could perform well in various tasks including emotional expression conversion, electrolaryngeal speech enhancement, and English accent conversion. Hirokazu Kameoka, Kou Tanaka, Damian Kwasny, Takuhiro Kaneko, Nobukatsu Hojo |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Cyclegan-VC2: Improved Cyclegan-based Non-parallel Voice ConversionabstractNon-parallel voice conversion (VC) is a technique for learning the mapping from source to target speech without relying on parallel data. This is an important task, but it has been challenging due to the disadvantages of the training conditions. Recently, CycleGAN-VC has provided a breakthrough and performed comparably to a parallel VC method without relying on any extra data, modules, or time alignment procedures. However, there is still a large gap between the real target and converted speech, and bridging this gap remains a challenge. To reduce the gap, we propose CycleGAN-VC2, which is an improved version of CycleGAN-VC incorporating three new techniques: an improved objective (two-step adversarial losses), improved generator (2-1-2D CNN), and improved discriminator (PatchGAN). We evaluated our method on a non-parallel VC task and analyzed the effect of each technique in detail. An objective evaluation showed that these techniques help bring the converted feature sequence closer to the target in terms of both global and local structures, which we assess by using Mel-cepstral distortion and modulation spectra distance, respectively. A subjective evaluation showed that CycleGAN-VC2 outperforms CycleGAN-VC in terms of naturalness and similarity for every speaker pair, including intra-gender and inter-gender pairs. Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Nobukatsu Hojo |
ICASSP | 3 |
| 2019 | ATTS2S-VC: Sequence-to-sequence Voice Conversion with Attention and Context Preservation MechanismsabstractThis paper describes a method based on a sequence-to-sequence learning (Seq2Seq) with attention and context preservation mechanism for voice conversion (VC) tasks. Seq2Seq has been outstanding at numerous tasks involving sequence modeling such as speech synthesis and recognition, machine translation, and image captioning. In contrast to current VC techniques, our method 1) stabilizes and accelerates the training procedure by considering guided attention and proposed context preservation losses, 2) allows not only spectral envelopes but also fundamental frequency contours and durations of speech to be converted, 3) requires no context information such as phoneme labels, and 4) requires no time-aligned source and target speech data in advance. In our experiment, the proposed VC framework can be trained in only one day, using only one GPU of an NVIDIA Tesla K80, while the quality of the synthesized speech is higher than that of speech converted by Gaussian mixture model-based VC and is comparable to that of speech generated by recurrent neural network-based text-to-speech synthesis, which can be regarded as an upper limit on VC performance. Kou Tanaka, Hirokazu Kameoka, Takuhiro Kaneko, Nobukatsu Hojo |
ICASSP | 1 |
| 2019 | StarGAN-VC2: Rethinking Conditional Methods for StarGAN-Based Voice ConversionabstractNon-parallel multi-domain voice conversion (VC) is a technique for learning mappings among multiple domains without relying on parallel data.This is important but challenging owing to the requirement of learning multiple mappings and the nonavailability of explicit supervision.Recently, StarGAN-VC has garnered attention owing to its ability to solve this problem only using a single generator.However, there is still a gap between real and converted speech.To bridge this gap, we rethink conditional methods of StarGAN-VC, which are key components for achieving non-parallel multi-domain VC in a single model, and propose an improved variant called StarGAN-VC2.Particularly, we rethink conditional methods in two aspects: training objectives and network architectures.For the former, we propose a source-and-target conditional adversarial loss that allows all source domain data to be convertible to the target domain data.For the latter, we introduce a modulation-based conditional method that can transform the modulation of the acoustic feature in a domain-specific manner.We evaluated our methods on non-parallel multi-speaker VC.An objective evaluation demonstrates that our proposed methods improve speech quality in terms of both global and local structure measures.Furthermore, a subjective evaluation shows that StarGAN-VC2 outperforms StarGAN-VC in terms of naturalness and speaker similarity. Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Nobukatsu Hojo |
INTERSPEECH | 3 |
| 2019 | ACVAE-VC: Non-Parallel Voice Conversion With Auxiliary Classifier Variational AutoencoderabstractThis paper proposes a non-parallel voice conversion (VC) method using a variant of the conditional variational autoencoder (VAE) called an auxiliary classifier VAE. The proposed method has two key features. First, it adopts fully convolutional architectures to construct the encoder and decoder networks so that the networks can learn conversion rules that capture the time dependencies in the acoustic feature sequences of source and target speech. Second, it uses information-theoretic regularization for the model training to ensure that the information in the attribute class label will not be lost in the conversion process. With regular conditional VAEs, the encoder and decoder are free to ignore the attribute class label input. This can be problematic since in such a situation, the attribute class label will have little effect on controlling the voice characteristics of input speech at test time. Such situations can be avoided by introducing an auxiliary classifier and training the encoder and decoder so that the attribute classes of the decoder outputs are correctly predicted by the classifier. We also present several ways to convert the feature sequence of input speech using the trained encoder and decoder and compare them in terms of audio quality through objective and subjective evaluations. We confirmed experimentally that the proposed method outperformed baseline non-parallel VC systems and performed comparably to an open-source parallel VC system trained using a parallel corpus in a speaker identity conversion task. Hirokazu Kameoka, Takuhiro Kaneko, Kou Tanaka, Nobukatsu Hojo |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Vae-Space: Deep Generative Model of Voice Fundamental Frequency ContoursabstractModeling the speech generation process can provide flexible and interpretable ways to generate intended synthetic speech. In this paper, we present a deep generative model of fundamental frequency (F0) contours of normal speech and singing voices. The generative model we propose in this paper 1) is able to accurately decompose an F0contour into the sum of phrase and accent components of the Fujisaki model, a mathematical model describing the control mechanism of vocal fold vibration, without an iterative algorithm, and 2) can represent/generate F0contours of both normal speech and singing voices reasonably well. Kou Tanaka, Hirokazu Kameoka, Kazuho Morikawa |
ICASSP | 1 |
| 2018 | StarGAN-VC: non-parallel many-to-many Voice Conversion Using Star Generative Adversarial NetworksabstractThis paper proposes a method that allows non-parallel many-to-many voice conversion (VC) by using a variant of a generative adversarial network (GAN) called StarGAN. Our method, which we call StarGAN-VC, is noteworthy in that it (1) requires no parallel utterances, transcriptions, or time alignment procedures for speech generator training, (2) simultaneously learns many-to-many mappings across different attribute domains using a single generator network, (3) is able to generate converted speech signals quickly enough to allow real-time implementations and (4) requires only several minutes of training examples to generate reasonably realistic sounding speech. Subjective evaluation experiments on a non-parallel many-to-many speaker identity conversion task revealed that the proposed method obtained higher sound quality and speaker similarity than a state-of-the-art method based on variational autoencoding GANs. Hirokazu Kameoka, Takuhiro Kaneko, Kou Tanaka, Nobukatsu Hojo |
SLT | 3 |
| 2018 | Synthetic-to-Natural Speech Waveform Conversion Using Cycle-Consistent Adversarial NetworksabstractWe propose a learning-based filter that allows us to directly modify a synthetic speech waveform into a natural speech waveform. Speech-processing systems using a vocoder framework such as statistical parametric speech synthesis and voice conversion are convenient especially for a limited number of data because it is possible to represent and process interpretable acoustic features over a compact space, such as the fundamental frequency (F0) and mel-cepstrum. However, a well-known problem that leads to the quality degradation of generated speech is an over-smoothing effect that eliminates some detailed structure of generated/converted acoustic features. To address this issue, we propose a synthetic-to-natural speech waveform conversion technique that uses cycle-consistent adversarial networks and which does not require any explicit assumption about speech waveform in adversarial learning. In contrast to current techniques, since our modification is performed at the waveform level, we expect that the proposed method will also make it possible to generate “vocoder-less” sounding speech even if the input speech is synthesized using a vocoder framework. The experimental results demonstrate that our proposed method can 1) alleviate the over-smoothing effect of the acoustic features despite the direct modification method used for the waveform and 2) greatly improve the naturalness of the generated speech sounds. Kou Tanaka, Takuhiro Kaneko, Nobukatsu Hojo, Hirokazu Kameoka |
SLT | 1 |
| 2017 | Physically Constrained Statistical F0 Prediction for Electrolaryngeal Speech Enhancement
Kou Tanaka, Hirokazu Kameoka, Tomoki Toda, Satoshi Nakamura 0001 |
INTERSPEECH | 1 |
| 2016 | Statistical F0 prediction for electrolaryngeal speech enhancement considering generative process of F0 contours within product of experts frameworkabstractWe have previously proposed a statistical fundamental frequency (F0) prediction method that makes it possible to predict the underlying F0contour of electrolaryngeal (EL) speech from its spectral feature sequence. Although this method was shown to contribute to improving the naturalness of EL speech as a whole, the predicted F0contour was still unnatural compared with that in normal speech. One possible solution to improve the naturalness of the predicted F0contours would be to take account of the physical mechanism of vocal phonation. Recently a statistical model of voice F0contours was formulated by constructing a stochastic counterpart of the Fujisaki model, a well-founded mathematical model representing the control mechanism of vocal fold vibration. This paper proposes a Product-of-Experts model to incorporate this generative model of voice F0contours into the statistical F0prediction model. Based on the constructed model, we derive algorithms for parameter training and F0prediction. Experimental results revealed that the proposed method successfully outperformed our previously proposed method in terms of the naturalness of the predicted F0contours. Kou Tanaka, Hirokazu Kameoka, Tomoki Toda, Satoshi Nakamura 0001 |
ICASSP | 1 |
| 2015 | An Enhanced Electrolarynx with Automatic Fundamental Frequency Control based on Statistical PredictionabstractAn electrolarynx is a type of speaking aid device which is able to mechanically generate excitation sounds to help laryngectomees produce electrolaryngeal (EL) speech. Although EL speech is quite intelligible, its naturalness suffers from monotonous fundamental frequency patterns of the mechanical excitation sounds. To make it possible to generate more natural excitation sounds, we have proposed a method to automatically control the fundamental frequency of the sounds generated by the electrolarynx based on a statistical prediction model, which predicts the fundamental frequency patterns from the produced EL speech in real-time. In this paper, we develop a prototype system by implementing the proposed control method in an actual, physical electrolarynx and evaluate its performance. Kou Tanaka, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
ASSETS | 1 |
| 2015 | Non-audible murmur enhancement based on statistical conversion using air- and body-conductive microphones in noisy environmentsabstractNon-Audible Murmur (NAM) is an extremely soft whispered voice detected by a special body-conductive microphone called a NAM microphone. Although NAM is a promising medium for silent speech communication, its quality is significantly degraded by its faint volume and spectral changes caused by body-conductive recording. To improve the quality of NAM, several enhancement methods based on statistical voice conversion (VC) techniques have been proposed, and their effectiveness has been confirmed in quiet environments. However, it can be expected that NAM will be used not only in quiet, but also in noisy environments, and it is thus necessary to develop enhancement methods that will also work in these cases. In this paper, we propose a framework for NAM enhancement using not only the NAM microphone but also an air-conductive microphone. Airand body-conducted NAM signals are used as the input of VC to estimate a more naturally sounding speech signal. To clarify adverse effects of external noises on the performance of the proposed framework and investigate a possibility to alleviate them by revising VC models, we also implement noise-dependent VC models within the proposed framework. Experimental results demonstrate that the proposed framework yields significant improvements in the spectral conversion accuracy and listenability of enhanced speech under both quiet and noisy environments. Yusuke Tajiri, Kou Tanaka, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2014 | An evaluation of excitation feature prediction in a hybrid approach to electrolaryngeal speech enhancementabstractWe implement removing micro-prosody with low-pass filtering and avoiding Unvoiced/Voiced (U/V) prediction as part of a hybrid approach to improve statistical excitation prediction in electrolaryngeal (EL) speech enhancement. An electrolarynx is a device that artificially generates excitation sounds to enable laryngectomees to produce EL speech. Although proficient laryngectomees can produce quite intelligible EL speech, it sounds very unnatural due to the mechanical excitation produced by the device. Moreover, the excitation sounds produced by the device often leak outside, adding noise to EL speech. To address these issues, in our previous work, we proposed a hybrid method using a noise reduction method for enhancing spectral parameters and voice conversion method for predicting excitation parameters. In this paper, we evaluate the effect of removing micro-prosody with low-pass filtering and avoiding U/V prediction in the hybrid enhancement process. Kou Tanaka, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
ICASSP | 1 |
| 2014 | Direct F0 control of an electrolarynx based on statistical excitation feature prediction and its evaluation through simulationabstractAn electrolarynx is a device that artificially generates excitation sounds to enable laryngectomees to produce electrolaryngeal (EL) speech. Although proficient laryngectomees can produce quite intelligible EL speech, it sounds very unnatural due to the mechanical excitation produced by the device. To address this issue, we have proposed several EL speech enhancement methods using statistical voice conversion and showed that statistical prediction of excitation parameters, such as F0 patterns, was essential to significantly improve naturalness of EL speech. In these methods, the original EL speech is recorded with a microphone and the enhanced EL speech is presented from a loudspeaker in real time. This framework is effective for telecommunication but it is not suitable to face-to-face conversation because both the original EL speech and the enhanced EL speech are presented to listeners. In this paper, we propose direct F0 control of the electrolarynx based on statistical excitation prediction to develop an EL speech enhancement technique also effective for face-to-face conversation. F0 patterns of excitation signals produced by the electrolarynx are predicted in real time from the EL speech produced by the laryngectomee’s articulation of the excitation signals with previously predicted F0 values. A simulation experiment is conducted to evaluate the effectiveness of the proposed method. The experimental results demonstrate that the proposed method yields significant improvements in naturalness of EL speech while keeping its intelligibility high enough. Index Terms: laryngectomee, electrolarynx, electrolaryngeal speech, statistical excitation prediction, simulation evaluation Kou Tanaka, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 1 |
| 2013 | A hybrid approach to electrolaryngeal speech enhancement based on spectral subtraction and statistical voice conversionabstractWe present a hybrid approach to improving naturalness of electrolaryngeal (EL) speech while minimizing degradation in intelligibility. An electrolarynx is a device that artificially generates excitation sounds to enable laryngectomees to produce EL speech. Although proficient laryngectomees can produce quite intelligible EL speech, it sounds very unnatural due to the mechanical excitation produced by the device. Moreover, the excitation sounds produced by the device often leak outside, adding noise to EL speech. To address these issues, previous work has proposed methods for EL speech enhancement through either noise reduction or voice conversion. The former usually causes no degradation in intelligibility but yields only small improvements in naturalness as the mechanical excitation sounds remain essentially unchanged. On the other hand, the latter method significantly improves naturalness of EL speech using spectral and excitation parameters of natural voices converted from acoustic parameters of EL speech, but it usually causes degradation in intelligibility owing to errors in conversion. We propose a hybrid method using the noise reduction method for enhancing spectral parameters and voice conversion method for predicting excitation parameters. The experimental results demonstrate the proposed method yields significant improvements in naturalness compared with EL speech while keeping intelligibility high enough. Index Terms: speaking-aid, electrolaryngeal speech, spectral subtraction, voice conversion, hybrid approach Kou Tanaka, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 1 |