Hirokazu Kameoka

dblp:97/941 · DBLP profile ↗
← Back
121ranked-venue papers
24as first author
29since 2021 · last 2025
0000-0003-3102-0162ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 97 · 13 first-author · 22 since 2021Artificial intelligence and machine learning · 58 · 13 first-author · 16 since 2021Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Rethinking Mean Opinion Scores in Speech Quality Assessment: Score Aggregation through Quantized Distribution Fitting
abstract
This study addresses the task of speech quality assessment (SQA), which aims to automatically predict the subjective quality of a given speech. Recent efforts have focused on training neural-based models to predict the mean opinion score (MOS) of speech samples produced by text-to-speech or voice conversion systems. We aim to enhance the performance of the models by a score aggregation method instead of MOS. The proposed method mitigates the effects of some issues arising from constraints imposed by limited options. Our method assumes annotators internally consider continuous scores and pick the nearest discrete rating. By modeling this process, we approximate the rating distribution by quantizing the latent continuous distribution. We then use the peak of the latent distribution, estimated through the loss between the distribution and actual ratings, as the new value instead of MOS. Experimental results demonstrate that substituting MOSNet’s target with this proposed value improves prediction performance.
Yuto Kondo, Hirokazu Kameoka, Kou Tanaka, Takuhiro Kaneko
ICASSP2
2025 FasterVoiceGrad: Faster One-step Diffusion-Based Voice Conversion with Adversarial Diffusion Conversion Distillation
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Yuto Kondo
INTERSPEECH2
2025 Vocoder-Projected Feature Discriminator
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Yuto Kondo
INTERSPEECH2
2025 JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles
Yuto Kondo, Hirokazu Kameoka, Kou Tanaka, Takuhiro Kaneko
INTERSPEECH2
2024 Training Generative Adversarial Network-Based Vocoder with Limited Data Using Augmentation-Conditional Discriminator
abstract
A generative adversarial network (GAN)-based vocoder trained with an adversarial discriminator is commonly used for speech synthesis because of its fast, lightweight, and high-quality characteristics. However, this data-driven model requires a large amount of training data incurring high data-collection costs. This fact motivates us to train a GAN-based vocoder on limited data. A promising solution is to augment the training data to avoid overfitting. However, a standard discriminator is unconditional and insensitive to distributional changes caused by data augmentation. Thus, augmented speech (which can be extraordinary) may be considered real speech. To address this issue, we propose an augmentation-conditional discriminator (AugCondD) that receives the augmentation state as input in addition to speech, thereby assessing input speech according to augmentation state, without inhibiting the learning of the original non-augmented distribution. Experimental results indicate that AugCondD improves speech quality under limited data conditions while achieving comparable speech quality under sufficient data conditions.1
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka
ICASSP2
2024 Selecting N-Lowest Scores for Training MOS Prediction Models
abstract
The automatic speech quality assessment (SQA) has been extensively studied to predict the speech quality without time-consuming questionnaires. Recently, neural-based SQA models have been actively developed for speech samples produced by text-to-speech or voice conversion, with a primary focus on training mean opinion score (MOS) prediction models. The quality of each speech sample may not be consistent across the entire duration, and it remains unclear which segments of the speech receive the primary focus from humans when assigning subjective evaluation for MOS calculation. We hypothesize that when humans rate speech, they tend to assign more weight to low-quality speech segments, and the variance in ratings for each sample is mainly due to accidental assignment of higher scores when overlooking the poor quality speech segments. Motivated by the hypothesis, we analyze the VCC2018 and BVCC datasets. Based on the hypothesis, we propose the more reliable representative value Nlow-MOS, the mean of the N-lowest opinion scores. Our experiments show that LCC and SRCC improve compared to regular MOS when employing Nlow-MOS to MOSNet training. This result suggests that Nlow-MOS is a more intrinsic representative value of subjective speech quality and makes MOSNet a better comparator of VC models.
Yuto Kondo, Hirokazu Kameoka, Kou Tanaka, Takuhiro Kaneko
ICASSP2
2024 FastVoiceGrad: One-step Diffusion-Based Voice Conversion with Adversarial Conditional Diffusion Distillation
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Yuto Kondo
INTERSPEECH2
2024 PRVAE-VC2: Non-Parallel Voice Conversion by Distillation of Speech Representations
Kou Tanaka, Hirokazu Kameoka, Takuhiro Kaneko, Yuto Kondo
INTERSPEECH2
2024 VoiceGrad: Non-Parallel Any-to-Many Voice Conversion With Annealed Langevin Dynamics
abstract
In this paper, we propose a non-parallel any-to-many voice conversion (VC) method termedVoiceGrad. Inspired by WaveGrad, a recently introduced novel waveform generation method, VoiceGrad is based upon the concepts of score matching, Langevin dynamics, and diffusion models. The idea involves training a score approximator, a fully convolutional network with a U-Net structure, to predict the gradient of the log density of the speech feature sequences of multiple speakers. The trained score approximator can be used to perform VC by using annealed Langevin dynamics or reverse diffusion process to iteratively update an input feature sequence towards the nearest stationary point of the target distribution. Thanks to the nature of this concept, VoiceGrad enables any-to-many VC, a VC scenario in which the speaker of input speech can be arbitrary, and allows for non-parallel training, which requires no parallel utterances.
Hirokazu Kameoka, Takuhiro Kaneko, Kou Tanaka, Nobukatsu Hojo, Shogo Seki
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 Wave-U-Net Discriminator: Fast and Lightweight Discriminator for Generative Adversarial Network-Based Speech Synthesis
abstract
In speech synthesis, a generative adversarial network (GAN), training a generator (speech synthesizer) and a discriminator in a min-max game, is widely used to improve speech quality. An ensemble of discriminators is commonly used in recent neural vocoders (e.g., HiFi-GAN) and end-to-end text-to-speech (TTS) systems (e.g., VITS) to scrutinize waveforms from multiple perspectives. Such discriminators allow synthesized speech to adequately approach real speech; however, they require an increase in the model size and computation time according to the increase in the number of discriminators. Alternatively, this study proposes a Wave-U-Net discriminator, which is a single but expressive discriminator with Wave-U-Net architecture. This discriminator is unique; it can assess a waveform in a sample-wise manner with the same resolution as the input signal, while extracting multilevel features via an encoder and decoder with skip connections. This architecture provides a generator with sufficiently rich information for the synthesized speech to be closely matched to the real speech. During the experiments, the proposed ideas were applied to a representative neural vocoder (HiFi-GAN) and an end-to-end TTS system (VITS). The results demonstrate that the proposed models can achieve comparable speech quality with a 2.31 times faster and 14.5 times more lightweight discriminator when used in HiFi-GAN and a 1.90 times faster and 9.62 times more lightweight discriminator when used in VITS.1
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Shogo Seki
ICASSP2
2023 JSV-VC: Jointly Trained Speaker Verification and Voice Conversion Models
abstract
This paper proposes a variational autoencoder (VAE)-based method for voice conversion (VC) on arbitrary source-target speaker pairs without parallel corpora, i.e., non-parallel any-to-any VC. One typical approach is to use speaker embeddings obtained from a speaker verification (SV) model as the condition for a VC model. However, converted speech is not guaranteed to reflect a target speaker’s characteristics in a naive combination of VC and SV models. Moreover, speaker embeddings are not designed for VC problems, leading to suboptimal conversion performance. To address these issues, the proposed method, JSV-VC, trains both VC and SV models jointly. The VC model is trained so that converted speech is verified as the target speaker in the SV model, while the SV model is trained in order to output consistent embeddings before and after the VC model. The experimental evaluation reveals that JSV-VC outperforms conventional any-to-any VC methods quantitatively and qualitatively.
Shogo Seki, Hirokazu Kameoka, Kou Tanaka, Takuhiro Kaneko
ICASSP2
2023 iSTFTNet2: Faster and More Lightweight iSTFT-Based Neural Vocoder Using 1D-2D CNN
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Shogo Seki
INTERSPEECH2
2023 CFVC: Conditional Filtering for Controllable Voice Conversion
Kou Tanaka, Takuhiro Kaneko, Hirokazu Kameoka, Shogo Seki
INTERSPEECH3
2023 FastMVAE2: On Improving and Accelerating the Fast Variational Autoencoder-Based Source Separation Algorithm for Determined Mixtures
abstract
This article proposes a new source model and training scheme to improve the accuracy and speed of the multichannel variational autoencoder (MVAE) method. The MVAE method is a recently proposed powerful multichannel source separation method. It consists of pretraining a source model represented by a conditional VAE (CVAE) and then estimating separation matrices along with other unknown parameters so that the log-likelihood is non-decreasing given an observed mixture signal. Although the MVAE method has been shown to provide high source separation performance, one drawback is the computational cost of the backpropagation steps in the separation-matrix estimation algorithm. To overcome this drawback, a method called “FastMVAE” was subsequently proposed, which uses an auxiliary classifier VAE (ACVAE) to train the source model. By using the classifier and encoder trained in this way, the optimal parameters of the source model can be inferred efficiently, albeit approximately, in each step of the algorithm. However, the generalization capability of the trained ACVAE source model was not satisfactory, which led to poor performance in situations with unseen data. To improve the generalization capability, this article proposes a new model architecture (called the “ChimeraACVAE” model) and a training scheme based on knowledge distillation. The experimental results revealed that the proposed source model trained with the proposed loss function achieved better source separation performance with less computation time than FastMVAE. We also confirmed that our methods were able to separate 18 sources with a reasonably good accuracy.
Li Li 0063, Hirokazu Kameoka, Shoji Makino
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Attentionpit: Soft Permutation Invariant Training for Audio Source Separation with Attention Mechanism
abstract
Permutation invariant training (PIT) has recently attracted attention as a framework to achieve end-to-end time-domain audio source separation. Its goal is to train a separation network that takes a mixture signal as input and produces the J underlying source signals. Since the order of the output signals is arbitrary, the idea of PIT is to first find the best output-target assignment and then update the network parameters based on the error given by that assignment at each iteration. However, there are two known problems with PIT: One is that it has a time complexity of $\mathcal{O}\left( {J!} \right)$, which makes it infeasible as J increases, and the other is that it is prone to getting stuck in bad local optimal solutions due to the hard output-target assignment process. To overcome these problems simultaneously, in this paper, we propose AttentionPIT, which uses an attention mechanism to find soft output-target assignments for separation network training, and can be run in polynomial time in J, as with the recently proposed fast PIT variants such as SinkPIT and HungarianPIT. The training loss of AttentionPIT is fully differentiable, allowing us to simultaneously perform processes corresponding to soft output-target assignment and network parameter update through backpropagation. Experiments on the LibriMix corpus revealed that while AttentionPIT works reasonably well on its own, it works even better when combined with SinkPIT and HungarianPIT so that AttentionPIT is run only in the early stages of training.
Hirokazu Kameoka, Shogo Seki, Li Li 0063, Chihiro Watanabe
ICASSP1
2022 ISTFTNET: Fast and Lightweight Mel-Spectrogram Vocoder Incorporating Inverse Short-Time Fourier Transform
abstract
In recent text-to-speech synthesis and voice conversion systems, a mel-spectrogram is commonly applied as an intermediate representation, and the necessity for a mel-spectrogram vocoder is increasing. A mel-spectrogram vocoder must solve three inverse problems: recovery of the original-scale magnitude spectrogram, phase reconstruction, and frequency-to-time conversion. A typical convolutional mel-spectrogram vocoder solves these problems jointly and implicitly using a convolutional neural network, including temporal upsampling layers, when directly calculating a raw waveform. Such an approach allows skipping redundant processes during waveform synthesis (e.g., the direct reconstruction of high-dimensional original-scale spectrograms). By contrast, the approach solves all problems in a black box and cannot effectively employ the time-frequency structures existing in a mel-spectrogram. We thus propose iSTFTNet, which replaces some output-side layers of the mel-spectrogram vocoder with the inverse short-time Fourier transform (iSTFT) after sufficiently reducing the frequency dimension using upsampling layers, reducing the computational cost from black-box modeling and avoiding redundant estimations of high-dimensional spectrograms. During our experiments, we applied our ideas to three HiFi-GAN variants and made the models faster and more lightweight with a reasonable speech quality.1
Takuhiro Kaneko, Kou Tanaka, Hirokazu Kameoka, Shogo Seki
ICASSP3
2022 HBP: An Efficient Block Permutation Solver Using Hungarian Algorithm and Spectrogram Inpainting for Multichannel Audio Source Separation
abstract
This paper proposes a method called "Hungarian Block Permutation (HBP)" to solve the block permutation problem in frequency-domain multichannel audio source separation. Many methods for frequency-domain multichannel audio source separation are designed to simultaneously solve frequency-wise source separation and permutation alignment in determined cases. However, in practice, separation can fail due to permutation inconsistencies in different frequency blocks for various reasons, such as convergence to a locally optimal solution as a result of bad initialization. To correct permutation inconsistencies, the proposed HBP method first masks, for each separated signal, the frequency bands where the components from other sources are likely to be dominant, and then restores the components in those bands so that the restored spectrogram becomes closer to the original spectrogram of the corresponding source. The Hungarian algorithm is then used to perform permutation realignment in those bands in accordance with the restored spectrogram. The experimental results show that the proposed method can solve the permutation realignment and improve the separation performance even in the case of 18 speakers.
Li Li 0063, Hirokazu Kameoka, Shogo Seki
ICASSP2
2022 Investigation And Comparison of Optimization Methods for Variational Autoencoder-Based Underdetermined Multichannel Source Separation
abstract
In this paper, we investigate two algorithms for variational autoencoder (VAE)-based underdetermined multichannel source separation. We previously extended the multichannel VAE (MVAE) method for determined multichannel source separation and proposed the generalized MVAE (GMVAE) method for underdetermined multichannel source separation. The GMVAE method employs a conditional VAE (CVAE) as the source model representing the power spectrograms of the underlying sources present in a mixture. While we developed a convergence-guaranteed parameter estimation algorithm using a majorization-minimization/minorization-maximization (MM) algorithm, an expectation-maximization (EM) algorithm also allows us to design another algorithm with the same property. However, a comparison of the MM-based and EM-based algorithms has not yet been revealed. To elucidate this, we investigate the MM-based and EM-based algorithms for the GMVAE method, using an improved CVAE variant called auxiliary classifier VAE (ACVAE). The experimental results suggest that the EM-based algorithm takes less computational cost, achieving comparable separation performance with the MM-based algorithm.
Shogo Seki, Hirokazu Kameoka, Li Li 0063
ICASSP2
2022 CAUSE: Crossmodal Action Unit Sequence Estimation from Speech
Hirokazu Kameoka, Takuhiro Kaneko, Shogo Seki, Kou Tanaka
INTERSPEECH1
2022 MISRNet: Lightweight Neural Vocoder Using Multi-Input Single Shared Residual Blocks
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Shogo Seki
INTERSPEECH2
2022 Distilling Sequence-to-Sequence Voice Conversion Models for Streaming Conversion Applications
abstract
This paper describes a method for distilling a recurrent-based sequence-to-sequence (S2S) voice conversion (VC) model. Although the performance of recent VCs is becoming higher quality, streaming conversion is still a challenge when considering practical applications. To achieve streaming VC, the conversion model needs a streamable structure, a causal layer rather than a non-causal layer. Motivated by this constraint and recent advances in S2S learning, we apply the teacher-student framework to recurrent-based S2S- VC models. A major challenge is how to minimize degradation due to the use of causal layers which masks future input information. Experimental evaluations show that except for male-to-female speaker conversion, our approach is able to maintain the teacher model's performance in terms of subjective evaluations despite the streamable student model structure. Audio samples can be accessed on http://www.kecl.ntt.co.jp/people/tanaka.ko/projects/dists2svc.
Kou Tanaka, Hirokazu Kameoka, Takuhiro Kaneko, Shogo Seki
SLT2
2022 Application of non-negative matrix factorization in oncology: one approach for establishing precision medicine
abstract
The increase in the expectations of artificial intelligence (AI) technology has led to machine learning technology being actively used in the medical field. Non-negative matrix factorization (NMF) is a machine learning technique used for image analysis, speech recognition, and language processing; recently, it is being applied to medical research. Precision medicine, wherein important information is extracted from large-scale medical data to provide optimal medical care for every individual, is considered important in medical policies globally, and the application of machine learning techniques to this end is being handled in several ways. NMF is also introduced differently because of the characteristics of its algorithms. In this review, the importance of NMF in the field of medicine, with a focus on the field of oncology, is described by explaining the mathematical science of NMF and the characteristics of the algorithm, providing examples of how NMF can be used to establish precision medicine, and presenting the challenges of NMF. Finally, the direction regarding the effective use of NMF in the field of oncology is also discussed.
Ryuji Hamamoto, Ken Takasawa, Hidenori Machino, Kazuma Kobayashi, Satoshi Takahashi, Amina Bolatkan, Norio Shinkai, Akira Sakai, Rina Aoyama, Masayoshi Yamada, Ken Asada, Masaaki Komatsu, Koji Okamoto, Hirokazu Kameoka, Syuzo Kaneko
Briefings Bioinform.14
2021 SepNet: A Deep Separation Matrix Prediction Network for Multichannel Audio Source Separation
abstract
In this paper, we propose SepNet, a deep neural network (DNN) designed to predict separation matrices from multichannel observations. One well-known approach to blind source separation (BSS) involves independent component analysis (ICA). A recently developed method called independent low-rank matrix analysis (ILRMA) is one of its powerful variants. These methods allow the estimation of separation matrices based on deterministic iterative algorithms. Specifically, ILRMA is designed to update the separation matrix according to an update rule derived based on the majorization-minimization principle. Although ILRMA performs reasonably well under some conditions, there is still room for improvement in terms of both separation accuracy and computation time, especially for large-scale microphone arrays. The existence of a deterministic iterative algorithm that can find one of the stationary points of the BSS problem implies that a DNN can also play that role if designed and trained properly. Motivated by this, we propose introducing a DNN that learns to convert a predefined input (e.g., an identity matrix) into a true separation matrix in accordance with a multichannel observation. To enable it to find one of the multiple solutions corresponding to different permutations of the source indices, we further propose adopting a permutation invariant training strategy to train the network. By using a fully convolutional architecture, we can design the network so that the forward propagation can be computed efficiently. The experimental results revealed that SepNet was able to find separation matrices faster and with better separation accuracy than ILRMA for mixtures of two sources.
Shota Inoue, Hirokazu Kameoka, Li Li 0063, Shoji Makino
ICASSP2
2021 Maskcyclegan-VC: Learning Non-Parallel Voice Conversion with Filling in Frames
abstract
Non-parallel voice conversion (VC) is a technique for training voice converters without a parallel corpus. Cycle-consistent adversarial network-based VCs (CycleGAN-VC and CycleGAN-VC2) are widely accepted as benchmark methods. However, owing to their insufficient ability to grasp time-frequency structures, their application is limited to mel-cepstrum conversion and not mel-spectrogram conversion despite recent advances in mel-spectrogram vocoders. To overcome this, CycleGAN-VC3, an improved variant of CycleGAN-VC2 that incorporates an additional module called time-frequency adaptive normalization (TFAN), has been proposed. However, an increase in the number of learned parameters is imposed. As an alternative, we propose MaskCycleGAN-VC, which is another extension of CycleGAN-VC2 and is trained using a novel auxiliary task called filling in frames (FIF). With FIF, we apply a temporal mask to the input mel-spectrogram and encourage the converter to fill in missing frames based on surrounding frames. This task allows the converter to learn time-frequency structures in a self-supervised manner and eliminates the need for an additional module such as TFAN. A subjective evaluation of the naturalness and speaker similarity showed that MaskCycleGAN-VC outperformed both CycleGAN-VC2 and CycleGAN-VC3 with a model size similar to that of CycleGAN-VC2.1
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Nobukatsu Hojo
ICASSP2
2021 StarGAN-VC+ASR: StarGAN-Based Non-Parallel Voice Conversion Regularized by Automatic Speech Recognition
abstract
Preserving the linguistic content of input speech is essential during voice conversion (VC). The star generative adversarial network-based VC method (StarGAN-VC) is a recently developed method that allows non-parallel many-to-many VC. Although this method is powerful, it can fail to preserve the linguistic content of input speech when the number of available training samples is extremely small. To overcome this problem, we propose the use of automatic speech recognition to assist model training, to improve StarGAN-VC, especially in low-resource scenarios. Experimental results show that using our proposed method, StarGAN-VC can retain more linguistic information than vanilla StarGAN-VC.
Shoki Sakamoto, Akira Taniguchi, Tadahiro Taniguchi, Hirokazu Kameoka
Interspeech4
2021 X-DC: Explainable Deep Clustering Based on Learnable Spectrogram Templates
abstract
Deep neural networks (DNNs) have achieved substantial predictive performance in various speech processing tasks. Particularly, it has been shown that a monaural speech separation task can be successfully solved with a DNN-based method called deep clustering (DC), which uses a DNN to describe the process of assigning a continuous vector to each time-frequency (TF) bin and measure how likely each pair of TF bins is to be dominated by the same speaker. In DC, the DNN is trained so that the embedding vectors for the TF bins dominated by the same speaker are forced to get close to each other. One concern regarding DC is that the embedding process described by a DNN has a black-box structure, which is usually very hard to interpret. The potential weakness owing to the noninterpretable black box structure is that it lacks the flexibility of addressing the mismatch between training and test conditions (caused by reverberation, for instance). To overcome this limitation, in this letter, we propose the concept of explainable deep clustering (X-DC), whose network architecture can be interpreted as a process of fitting learnable spectrogram templates to an input spectrogram followed by Wiener filtering. During training, the elements of the spectrogram templates and their activations are constrained to be nonnegative, which facilitates the sparsity of their values and thus improves interpretability. The main advantage of this framework is that it naturally allows us to incorporate a model adaptation mechanism into the network thanks to its physically interpretable structure. We experimentally show that the proposed X-DC enables us to visualize and understand the clues for the model to determine the embedding vectors while achieving speech separation performance comparable to that of the original DC models.
Chihiro Watanabe, Hirokazu Kameoka
Neural Comput.2
2021 Pretraining Techniques for Sequence-to-Sequence Voice Conversion
abstract
Sequence-to-sequence (seq2seq) voice conversion (VC) models are attractive owing to their ability to convert prosody. Nonetheless, without sufficient data, seq2seq VC models can suffer from unstable training and mispronunciation problems in the converted speech, thus far from practical. To tackle these shortcomings, we propose to transfer knowledge from other speech processing tasks where large-scale corpora are easily available, typically text-to-speech (TTS) and automatic speech recognition (ASR). We argue that VC models initialized with such pretrained ASR or TTS model parameters can generate effective hidden representations for high-fidelity, highly intelligible converted speech. In this work, we examine our proposed method in a parallel, one-to-one setting. We employed recurrent neural network (RNN)-based and Transformer based models, and through systematical experiments, we demonstrate the effectiveness of the pretraining scheme and the superiority of Transformer based models over RNN-based models in terms of intelligibility, naturalness, and similarity.
Wen-Chin Huang, Tomoki Hayashi, Yi-Chiao Wu, Hirokazu Kameoka, Tomoki Toda
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 Many-to-Many Voice Transformer Network
abstract
This paper proposes a voice conversion (VC) method based on a sequence-to-sequence (S2S) learning framework, which enables simultaneous conversion of the voice characteristics, pitch contour, and duration of input speech. We previously proposed an S2S-based VC method using a transformer network architecture called the voice transformer network (VTN). The original VTN was designed to learn only a mapping of speech feature sequences from one speaker to another. Here, the main idea we propose is an extension of the original VTN that can simultaneously learn mappings among multiple speakers. This extension, called the many-to-many VTN, enables us to fully use available training data collected from multiple speakers by capturing common latent features that can be shared across different speakers. It also allows us to introduce a training loss called the identity mapping loss to ensure that the input feature sequence will remain unchanged when the source and target speaker indices are the same. Using this particular loss for model training has been found to be extremely effective in improving the performance of the model at test time. We conducted speaker identity conversion experiments and found that our model obtained higher sound quality and speaker similarity than baseline methods. We also found that our model, with a slight modification to its architecture, can handle any-to-many conversion tasks reasonably well.
Hirokazu Kameoka, Wen-Chin Huang, Kou Tanaka, Takuhiro Kaneko, Nobukatsu Hojo, Tomoki Toda
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 Harmonic-Temporal Factor Decomposition for Unsupervised Monaural Separation of Harmonic Sounds
abstract
We address the problem of separating a monaural mixture of harmonic sounds into the audio signals of individual semitones in an unsupervised manner. Unsupervised monaural audio source separation has thus far been mainly addressed by two approaches: one rooted in computational auditory scene analysis (CASA) and the other based on non-negative matrix factorization (NMF). These approaches focus on different clues for making source separation possible. A CASA-based method called harmonic-temporal clustering (HTC) focuses on a local time-frequency structure of individual sources, whereas NMF focuses on a global time-frequency structure of music spectrograms. These clues do not conflict with each other and can be used to achieve a more reliable audio source separation algorithm. Hence, we propose a monaural audio source separation framework, harmonic-temporal factor decomposition (HTFD), by developing a spectrogram model that encompasses the features of the models used in the NMF and HTC approaches. We further incorporate a source-filter model to build an extension of HTFD, source-filter HTFD (SF-HTFD). We derive efficient parameter estimation algorithms of HTFD and SF-HTFD based on the auxiliary function principle. We show, through music source separation experiments, the efficacy of HTFD and SF-HTFD compared with conventional methods. Furthermore, we demonstrate the effectiveness of HTFD and SF-HTFD for automatic musical key transposition.
Tomohiko Nakamura, Hirokazu Kameoka
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Voice Transformer Network: Sequence-to-Sequence Voice Conversion Using Transformer with Text-to-Speech Pretraining
abstract
We introduce a novel sequence-to-sequence (seq2seq) voice conversion (VC) model based on the Transformer architecture with text-to-speech (TTS) pretraining.Seq2seq VC models are attractive owing to their ability to convert prosody.While seq2seq models based on recurrent neural networks (RNNs) and convolutional neural networks (CNNs) have been successfully applied to VC, the use of the Transformer network, which has shown promising results in various speech processing tasks, has not yet been investigated.Nonetheless, their data-hungry property and the mispronunciation of converted speech make seq2seq models far from practical.To this end, we propose a simple yet effective pretraining technique to transfer knowledge from learned TTS models, which benefit from large-scale, easily accessible TTS corpora.VC models initialized with such pretrained model parameters are able to generate effective hidden representations for high-fidelity, highly intelligible converted speech.Experimental results show that such a pretraining scheme can facilitate data-efficient training and outperform an RNN-based seq2seq VC model in terms of intelligibility, naturalness, and similarity.
Wen-Chin Huang, Tomoki Hayashi, Yi-Chiao Wu, Hirokazu Kameoka, Tomoki Toda
INTERSPEECH4
2020 CycleGAN-VC3: Examining and Improving CycleGAN-VCs for Mel-Spectrogram Conversion
abstract
Non-parallel voice conversion (VC) is a technique for learning mappings between source and target speeches without using a parallel corpus.Recently, cycle-consistent adversarial network (CycleGAN)-VC and CycleGAN-VC2 have shown promising results regarding this problem and have been widely used as benchmark methods.However, owing to the ambiguity of the effectiveness of CycleGAN-VC/VC2 for mel-spectrogram conversion, they are typically used for mel-cepstrum conversion even when comparative methods employ mel-spectrogram as a conversion target.To address this, we examined the applicability of CycleGAN-VC/VC2 to mel-spectrogram conversion.Through initial experiments, we discovered that their direct applications compromised the time-frequency structure that should be preserved during conversion.To remedy this, we propose CycleGAN-VC3, an improvement of CycleGAN-VC2 that incorporates time-frequency adaptive normalization (TFAN).Using TFAN, we can adjust the scale and bias of the converted features while reflecting the time-frequency structure of the source mel-spectrogram.We evaluated CycleGAN-VC3 on inter-gender and intra-gender non-parallel VC.A subjective evaluation of naturalness and similarity showed that for every VC pair, CycleGAN-VC3 outperforms or is competitive with the two types of CycleGAN-VC2, one of which was applied to mel-cepstrum and the other to mel-spectrogram. 1
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Nobukatsu Hojo
INTERSPEECH2
2020 Nonparallel Voice Conversion With Augmented Classifier Star Generative Adversarial Networks
abstract
We previously proposed a method that allows for nonparallel voice conversion (VC) by using a variant of generative adversarial networks (GANs) called StarGAN. The main features of our method, called StarGAN-VC, are as follows: First, it requires no parallel utterances, transcriptions, or time alignment procedures for speech generator training. Second, it can simultaneously learn mappings across multiple domains using a single generator network and thus fully exploit available training data collected from multiple domains to capture latent features that are common to all the domains. Third, it can generate converted speech signals quickly enough to allow real-time implementations and requires only several minutes of training examples to generate reasonably realistic-sounding speech. In this article, we describe three formulations of StarGAN, including a newly introduced novel StarGAN variant called “Augmented classifier StarGAN (A-StarGAN)”, and compare them in a nonparallel VC task. We also compare them with several baseline methods.
Hirokazu Kameoka, Takuhiro Kaneko, Kou Tanaka, Nobukatsu Hojo
IEEE ACM Trans. Audio Speech Lang. Process.1
2020 ConvS2S-VC: Fully Convolutional Sequence-to-Sequence Voice Conversion
abstract
This article proposes a voice conversion (VC) method using sequence-to-sequence (seq2seq or S2S) learning, which flexibly converts not only the voice characteristics but also the pitch contour and duration of input speech. The proposed method, called ConvS2S-VC, has three key features. First, it uses a model with a fully convolutional architecture. This is particularly advantageous in that it is suitable for parallel computations using GPUs. It is also beneficial since it enables effective normalization techniques such as batch normalization to be used for all the hidden layers in the networks. Second, it achieves many-to-many conversion by simultaneously learning mappings among multiple speakers using only a single model instead of separately learning mappings between each speaker pair using a different model. This enables the model to fully utilize available training data collected from multiple speakers by capturing common latent features that can be shared across different speakers. Owing to this structure, our model works reasonably well even without source speaker information, thus making it able to handle any-to-many conversion tasks. Third, we introduce a mechanism, called the conditional batch normalization that switches batch normalization layers in accordance with the target speaker. This particular mechanism has been found to be extremely effective for our many-to-many conversion model. We conducted speaker identity conversion experiments and found that ConvS2S-VC obtained higher sound quality and speaker similarity than baseline methods. We also found from audio examples that it could perform well in various tasks including emotional expression conversion, electrolaryngeal speech enhancement, and English accent conversion.
Hirokazu Kameoka, Kou Tanaka, Damian Kwasny, Takuhiro Kaneko, Nobukatsu Hojo
IEEE ACM Trans. Audio Speech Lang. Process.1
2019 Joint Separation and Dereverberation of Reverberant Mixtures with Multichannel Variational Autoencoder
abstract
In this paper, we deal with a multichannel source separation problem under a highly reverberant condition. The multichannel variational autoencoder (MVAE) is a recently proposed source separation method that employs the decoder distribution of a conditional VAE (CVAE) as the generative model for the complex spectrograms of the underlying source signals. Although MVAE is notable in that it can significantly improve the source separation performance compared with conventional methods, its capability to separate highly reverberant mixtures is still limited since MVAE uses an instantaneous mixture model. To overcome this limitation, in this paper we propose extending MVAE to simultaneously solve source separation and dereverberation problems by formulating the separation system as a frequency-domain convolutive mixture model. A convergence-guaranteed algorithm based on the coordinate descent method is derived for the optimiza- tion. Experimental results revealed that the proposed method outperformed the conventional methods in terms of all the source separation criteria in highly reverberant environments.
Shota Inoue, Hirokazu Kameoka, Li Li 0063, Shogo Seki, Shoji Makino
ICASSP2
2019 Seeing through Sounds: Predicting Visual Semantic Segmentation Results from Multichannel Audio Signals
abstract
Sounds provide us with vast amounts of information about surrounding objects and can even remind us visual images of them. Is it possible to implement this noteworthy human ability on machines? In this paper, we study a new task that consists of predicting image recognition results in the form of semantic segmentation with given multichannel audio signals. Our approach uses a convolutional neural network that is designed to directly output semantic segmentation results by taking audio features as its inputs. A bilinear feature fusion scheme is incorporated that efficiently models underlying higher-order interactions between audio and visual sources. Experimental evaluations with both synthetic and real sound datasets show that our approach can recover the desired segmented images reasonably well.
Go Irie, Mirela Ostrek, Hirokazu Kameoka, Akisato Kimura, Takahito Kawanishi, Kunio Kashino
ICASSP4
2019 Cyclegan-VC2: Improved Cyclegan-based Non-parallel Voice Conversion
abstract
Non-parallel voice conversion (VC) is a technique for learning the mapping from source to target speech without relying on parallel data. This is an important task, but it has been challenging due to the disadvantages of the training conditions. Recently, CycleGAN-VC has provided a breakthrough and performed comparably to a parallel VC method without relying on any extra data, modules, or time alignment procedures. However, there is still a large gap between the real target and converted speech, and bridging this gap remains a challenge. To reduce the gap, we propose CycleGAN-VC2, which is an improved version of CycleGAN-VC incorporating three new techniques: an improved objective (two-step adversarial losses), improved generator (2-1-2D CNN), and improved discriminator (PatchGAN). We evaluated our method on a non-parallel VC task and analyzed the effect of each technique in detail. An objective evaluation showed that these techniques help bring the converted feature sequence closer to the target in terms of both global and local structures, which we assess by using Mel-cepstral distortion and modulation spectra distance, respectively. A subjective evaluation showed that CycleGAN-VC2 outperforms CycleGAN-VC in terms of naturalness and similarity for every speaker pair, including intra-gender and inter-gender pairs.
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Nobukatsu Hojo
ICASSP2
2019 Fast MVAE: Joint Separation and Classification of Mixed Sources Based on Multichannel Variational Autoencoder with Auxiliary Classifier
abstract
This paper proposes an alternative algorithm for the multi-channel variational autoencoder (MVAE), a recently proposed multichannel source separation approach. While MVAE is notable for its impressive source separation performance, its convergence-guaranteed optimization algorithm and the fact that it allows us to estimate source-class labels simultaneously with source separation, there are still two major drawbacks, namely, the high computational complexity and the unsatisfactory source classification accuracy. To overcome these drawbacks, the proposed method employs an auxiliary classifier VAE, which is an information-theoretic extension of the conditional VAE, for learning the generative model of the source spectrograms. Furthermore, with the trained auxiliary classifier, we introduce a novel algorithm for the optimization that can both reduce the computational time and improve the source classification performance. We call the proposed method "fast MVAE (fMVAE) ". Experimental evaluations revealed that fMVAE achieved source separation performance comparable to that of MVAE and a source classification accu-racy rate of about 80% while reducing computational time by about 93%.
Li Li 0063, Hirokazu Kameoka, Shoji Makino
ICASSP2
2019 ATTS2S-VC: Sequence-to-sequence Voice Conversion with Attention and Context Preservation Mechanisms
abstract
This paper describes a method based on a sequence-to-sequence learning (Seq2Seq) with attention and context preservation mechanism for voice conversion (VC) tasks. Seq2Seq has been outstanding at numerous tasks involving sequence modeling such as speech synthesis and recognition, machine translation, and image captioning. In contrast to current VC techniques, our method 1) stabilizes and accelerates the training procedure by considering guided attention and proposed context preservation losses, 2) allows not only spectral envelopes but also fundamental frequency contours and durations of speech to be converted, 3) requires no context information such as phoneme labels, and 4) requires no time-aligned source and target speech data in advance. In our experiment, the proposed VC framework can be trained in only one day, using only one GPU of an NVIDIA Tesla K80, while the quality of the synthesized speech is higher than that of speech converted by Gaussian mixture model-based VC and is comparable to that of speech generated by recurrent neural network-based text-to-speech synthesis, which can be regarded as an upper limit on VC performance.
Kou Tanaka, Hirokazu Kameoka, Takuhiro Kaneko, Nobukatsu Hojo
ICASSP2
2019 StarGAN-VC2: Rethinking Conditional Methods for StarGAN-Based Voice Conversion
abstract
Non-parallel multi-domain voice conversion (VC) is a technique for learning mappings among multiple domains without relying on parallel data.This is important but challenging owing to the requirement of learning multiple mappings and the nonavailability of explicit supervision.Recently, StarGAN-VC has garnered attention owing to its ability to solve this problem only using a single generator.However, there is still a gap between real and converted speech.To bridge this gap, we rethink conditional methods of StarGAN-VC, which are key components for achieving non-parallel multi-domain VC in a single model, and propose an improved variant called StarGAN-VC2.Particularly, we rethink conditional methods in two aspects: training objectives and network architectures.For the former, we propose a source-and-target conditional adversarial loss that allows all source domain data to be convertible to the target domain data.For the latter, we introduce a modulation-based conditional method that can transform the modulation of the acoustic feature in a domain-specific manner.We evaluated our methods on non-parallel multi-speaker VC.An objective evaluation demonstrates that our proposed methods improve speech quality in terms of both global and local structure measures.Furthermore, a subjective evaluation shows that StarGAN-VC2 outperforms StarGAN-VC in terms of naturalness and speaker similarity.
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Nobukatsu Hojo
INTERSPEECH2
2019 A Modified Algorithm for Multiple Input Spectrogram Inversion
abstract
We propose a new algorithm to estimate the phase of speech signal in the mixture of audio sources under the assumption that the magnitude spectrum of each source is given. The pre- vious method, multiple input spectrogram inversion algorithm (MISI), often performs poorly when the magnitude spectro- grams estimated are not accurate. This may be because it im- poses a strict constraint that the summation of source wave- forms should be exactly the same as the mixture waveform. Our proposing algorithm employs a new objective function in which this constraint is relaxed. In this objective function, the difference between the summation of source waveforms and the mixture waveform is the target to be minimized. The perfor- mance of our method, modified MISI is evaluated on two dif- ferent experimental settings. In both settings it improves the audio source separation performance compared to MISI.
Dongxiao Wang, Hirokazu Kameoka, Koichi Shinoda
INTERSPEECH2
2019 Supervised Determined Source Separation with Multichannel Variational Autoencoder
abstract
This letter proposes a multichannel source separation technique, the multichannel variational autoencoder (MVAE) method, which uses a conditional VAE (CVAE) to model and estimate the power spectrograms of the sources in a mixture. By training the CVAE using the spectrograms of training examples with source-class labels, we can use the trained decoder distribution as a universal generative model capable of generating spectrograms conditioned on a specified class index. By treating the latent space variables and the class index as the unknown parameters of this generative model, we can develop a convergence-guaranteed algorithm for supervised determined source separation that consists of iteratively estimating the power spectrograms of the underlying sources, as well as the separation matrices. In experimental evaluations, our MVAE produced better separation performance than a baseline method.
Hirokazu Kameoka, Li Li 0063, Shota Inoue, Shoji Makino
Neural Comput.1
2019 ACVAE-VC: Non-Parallel Voice Conversion With Auxiliary Classifier Variational Autoencoder
abstract
This paper proposes a non-parallel voice conversion (VC) method using a variant of the conditional variational autoencoder (VAE) called an auxiliary classifier VAE. The proposed method has two key features. First, it adopts fully convolutional architectures to construct the encoder and decoder networks so that the networks can learn conversion rules that capture the time dependencies in the acoustic feature sequences of source and target speech. Second, it uses information-theoretic regularization for the model training to ensure that the information in the attribute class label will not be lost in the conversion process. With regular conditional VAEs, the encoder and decoder are free to ignore the attribute class label input. This can be problematic since in such a situation, the attribute class label will have little effect on controlling the voice characteristics of input speech at test time. Such situations can be avoided by introducing an auxiliary classifier and training the encoder and decoder so that the attribute classes of the decoder outputs are correctly predicted by the classifier. We also present several ways to convert the feature sequence of input speech using the trained encoder and decoder and compare them in terms of audio quality through objective and subjective evaluations. We confirmed experimentally that the proposed method outperformed baseline non-parallel VC systems and performed comparably to an open-source parallel VC system trained using a parallel corpus in a speaker identity conversion task.
Hirokazu Kameoka, Takuhiro Kaneko, Kou Tanaka, Nobukatsu Hojo
IEEE ACM Trans. Audio Speech Lang. Process.1
2018 Speech Waveform Synthesis from MFCC Sequences with Generative Adversarial Networks
abstract
This paper proposes a method for generating speech from filterbank mel frequency cepstral coefficients (MFCC), which are widely used in speech applications, such as ASR, but are generally considered unusable for speech synthesis. First, we predict fundamental frequency and voicing information from MFCCs with an autoregressive recurrent neural net. Second, the spectral envelope information contained in MFCCs is converted to all-pole filters, and a pitch-synchronous excitation model matched to these filters is trained. Finally, we introduce a generative adversarial network -based noise model to add a realistic high-frequency stochastic component to the modeled excitation signal. The results show that high quality speech reconstruction can be obtained, given only MFCC information at test time.
Lauri Juvela, Bajibabu Bollepalli, Xin Wang 0037, Hirokazu Kameoka, Manu Airaksinen, Junichi Yamagishi, Paavo Alku
ICASSP4
2018 Joint Separation and Dereverberation of Reverberant Mixtures with Determined Multichannel Non-Negative Matrix Factorization
abstract
This paper proposes an extension of multichannel non-negative matrix factorization (MNMF) that simultaneously solves source separation and dereverberation. While MNMF was originally formulated under an underdetermined problem setting where sources can outnumber microphones, a determined counterpart of MNMF, which we call the determined MNMF (DMNMF), has recently been proposed with notable success. This approach is particularly notable in that the optimization process can be more than 30 times faster than the underdetermined version owing to the fact that it involves no matrix inversion computations. One drawback as regards all methods based on instantaneous mixture models, including MNMF, is that they are weak against long reverberation. To overcome this drawback, this paper proposes an extension of DMNMF using a frequency-domain convolutive mixture model. The optimization process of the proposed method consists of iteratively updating (i) the spectral parameters of each source using the majorization-minimization algorithm, (ii) the separation matrix using iterative projection, and (iii) the dereverberation filters using multichannel linear prediction. Experimental results showed that the proposed method yielded higher separation performance and dereverberation performance than the baseline method under highly reverberant environments.
Hideaki Kagami, Hirokazu Kameoka, Masahiro Yukawa
ICASSP2
2018 Deep Clustering with Gated Convolutional Networks
abstract
Deep clustering is a recently introduced deep learning-based method for speech separation. The idea is to model and train the mapping from each time-frequency (TF) region of a spectrogram to an embedding space so that the embedding features of the TF regions dominated by the same source are forced to get close to each other and those dominated by different sources are forced to get separated from each other. This allows us to construct binary masks by applying a regular clustering algorithm to the mapped embedding vectors of a test mixture signal. The original deep clustering uses a bidirectional long short-term memory (BLSTM) recurrent neural network (RNN) to model the embedding process. Although RNN-based architectures are indeed a natural choice for modeling long-term dependencies of time series data, recent work has shown that convolutional networks (CNNs) with gating mechanisms also have an excellent potential for capturing long-term structures. In addition, they are less prone to overfitting and are suitable for parallel computations. Motivated by these facts, this paper proposes adopting CNN-based architectures for deep clustering. Specifically, we use a gated CNN architecture, which was introduced to model word sequences for language modeling and was shown to outperform LSTM language models trained in a similar setting. We tested various CNN architectures on a monaural source separation task. The results revealed that the proposed architectures achieved better performance than the BLSTM-based architecture under the same training condition and comparable performance even with a smaller amount of training data.
Li Li 0063, Hirokazu Kameoka
ICASSP2
2018 Vae-Space: Deep Generative Model of Voice Fundamental Frequency Contours
abstract
Modeling the speech generation process can provide flexible and interpretable ways to generate intended synthetic speech. In this paper, we present a deep generative model of fundamental frequency (F0) contours of normal speech and singing voices. The generative model we propose in this paper 1) is able to accurately decompose an F0contour into the sum of phrase and accent components of the Fujisaki model, a mathematical model describing the control mechanism of vocal fold vibration, without an iterative algorithm, and 2) can represent/generate F0contours of both normal speech and singing voices reasonably well.
Kou Tanaka, Hirokazu Kameoka, Kazuho Morikawa
ICASSP2
2018 StarGAN-VC: non-parallel many-to-many Voice Conversion Using Star Generative Adversarial Networks
abstract
This paper proposes a method that allows non-parallel many-to-many voice conversion (VC) by using a variant of a generative adversarial network (GAN) called StarGAN. Our method, which we call StarGAN-VC, is noteworthy in that it (1) requires no parallel utterances, transcriptions, or time alignment procedures for speech generator training, (2) simultaneously learns many-to-many mappings across different attribute domains using a single generator network, (3) is able to generate converted speech signals quickly enough to allow real-time implementations and (4) requires only several minutes of training examples to generate reasonably realistic sounding speech. Subjective evaluation experiments on a non-parallel many-to-many speaker identity conversion task revealed that the proposed method obtained higher sound quality and speaker similarity than a state-of-the-art method based on variational autoencoding GANs.
Hirokazu Kameoka, Takuhiro Kaneko, Kou Tanaka, Nobukatsu Hojo
SLT1
2018 Synthetic-to-Natural Speech Waveform Conversion Using Cycle-Consistent Adversarial Networks
abstract
We propose a learning-based filter that allows us to directly modify a synthetic speech waveform into a natural speech waveform. Speech-processing systems using a vocoder framework such as statistical parametric speech synthesis and voice conversion are convenient especially for a limited number of data because it is possible to represent and process interpretable acoustic features over a compact space, such as the fundamental frequency (F0) and mel-cepstrum. However, a well-known problem that leads to the quality degradation of generated speech is an over-smoothing effect that eliminates some detailed structure of generated/converted acoustic features. To address this issue, we propose a synthetic-to-natural speech waveform conversion technique that uses cycle-consistent adversarial networks and which does not require any explicit assumption about speech waveform in adversarial learning. In contrast to current techniques, since our modification is performed at the waveform level, we expect that the proposed method will also make it possible to generate “vocoder-less” sounding speech even if the input speech is synthesized using a vocoder framework. The experimental results demonstrate that our proposed method can 1) alleviate the over-smoothing effect of the acoustic features despite the direct modification method used for the waveform and 2) greatly improve the naturalness of the generated speech sounds.
Kou Tanaka, Takuhiro Kaneko, Nobukatsu Hojo, Hirokazu Kameoka
SLT4
2018 Nonnegative Matrix Factorization With Basis Clustering Using Cepstral Distance Regularization
abstract
One successful approach for audio source separation involves applying nonnegative matrix factorization (NMF) to a magnitude spectrogram regarded as a nonnegative matrix. This can be interpreted as approximating the observed spectra at each time frame as the linear sum of the basis spectra scaled by time-varying amplitudes. This paper deals with the problem of the unsupervised instrument-wise source separation of polyphonic signals based on an extension of the NMF approach. We focus on the fact that each piece of music is typically played on a handful of musical instruments, which allows us to assume that the spectra of the underlying audio events in a polyphonic signal can be grouped into a reasonably small number of clusters in the mel-frequency cepstral coefficient (MFCC) domain. Based on this assumption, we propose formulating factorization of a magnitude spectrogram and clustering of the basis spectra in the MFCC domain as a joint optimization problem and derive a novel optimization algorithm based on the majorization–minimization principle. Experimental results revealed that our method was superior to a two-stage algorithm that consists of performing factorization followed by clustering the basis spectra, thus showing the advantage of the joint optimization approach.
Hirokazu Kameoka, Takuya Higuchi, Mikihiro Tanaka, Li Li 0063
IEEE ACM Trans. Audio Speech Lang. Process.1
2017 A majorization-minimization algorithm with projected gradient updates for time-domain spectrogram factorization
abstract
We previously introduced a framework called time-domain spectrogram factorization (TSF), which realizes nonnegative matrix factorization (NMF)-like source separation in the time domain. This framework is particularly noteworthy in that, while maintaining the ability of NMF to obtain a parts-based representation of magnitude spectra, it allows us to (i) circumvent the commonly made assumption with the NMF approach that the magnitude spectra of source components are additive and (ii) take account of the interdependence of the phase/amplitude components at different time-frequency points. In particular, the second factor has been overlooked despite its potential importance. Our previous study revealed that the conventional TSF algorithm was relatively slow due to large matrix inversions, and the early stopping of the algorithm often resulted in poor separation accuracy. To overcome this problem, this paper presents an iterative TSF solver using projected gradient updates. Simulation results show that the proposed TSF approach yields higher source separation performance than NMF and the other variants including the original TSF.
Hideaki Kagami, Hirokazu Kameoka, Masahiro Yukawa
ICASSP2
2017 Complex NMF with the generalized Kullback-Leibler divergence
abstract
We previously introduced a phase-aware variant of the non-negative matrix factorization (NMF) approach for audio source separation, which we call the “Complex NMF (CNMF).” This approach makes it possible to realize NMF-like signal decompositions in the complex time-frequency domain. One limitation of the CNMF framework is that the divergence measure is limited to only the Euclidean distance. Some previous studies have revealed that for source separation tasks with NMF, the generalized Kullback-Leibler (KL) divergence tends to yield higher accuracy than when using other divergence measures. This motivated us to believe that CNMF could achieve even greater source separation accuracy if we could derive an algorithm for a KL divergence counterpart of CNMF. In this paper, we start by defining the notion of the “dual” form of the CNMF formulation, derived from the original Euclidean CNMF, and show that a KL divergence counterpart of CNMF can be developed based on this dual formulation. We call this “KL-CNMF”. We further derive a convergence-guaranteed iterative algorithm for KL-CNMF based on a majorization-minimization scheme. The source separation experiments revealed that the proposed KL-CNMF yielded higher accuracy than the Euclidean CNMF and NMF with varying divergences.
Hirokazu Kameoka, Hideaki Kagami, Masahiro Yukawa
ICASSP1
2017 Generative adversarial network-based postfilter for statistical parametric speech synthesis
abstract
We propose a postfilter based on a generative adversarial network (GAN) to compensate for the differences between natural speech and speech synthesized by statistical parametric speech synthesis. In particular, we focus on the differences caused by over-smoothing, which makes the sounds muffled. Over-smoothing occurs in the time and frequency directions and is highly correlated in both directions, and conventional methods based on heuristics are too limited to cover all the factors (e.g., global variance was designed only to recover the dynamic range). To solve this problem, we focus on “spectral texture”, i.e., the details of the time-frequency representation, and propose a learning-based postfilter that captures the structures directly from the data. To estimate the true distribution, we utilize a GAN composed of a generator and a discriminator. This optimizes the generator to produce samples imitating the dataset according to the adversarial discriminator. This adversarial process encourages the generator to fit the true data distribution, i.e., to generate realistic spectral texture. Objective evaluation of experimental results shows that the GAN-based postfilter can compensate for detailed spectral structures including modulation spectrum, and subjective evaluation shows that its generated speech is comparable to natural speech.
Takuhiro Kaneko, Hirokazu Kameoka, Nobukatsu Hojo, Yusuke Ijima, Kaoru Hiramatsu, Kunio Kashino
ICASSP2
2017 Fast algorithm for statistical phrase/accent command estimation based on generative model incorporating spectral features
abstract
An important challenge in speech processing involves extracting non-linguistic information from a fundamental frequency (F0) contour of speech. We propose a fast algorithm for estimating the model parameters of the Fujisaki model, namely, the timings and magnitudes of the phrase and accent commands. Although a powerful parameter estimation framework based on a stochastic counterpart of the Fujisaki model has recently been proposed, it still had room for improvement in terms of both computational efficiency and parameter estimation accuracy. This paper describes our two contributions. First, we propose a hard expectation-maximization (EM) algorithm for parameter inference where the E step of the conventional EM algorithm is replaced with a point estimation procedure to accelerate the estimation process. Second, to improve the parameter estimation accuracy, we add a generative process of a spectral feature sequence to the generative model. This makes it possible to use linguistic or phonological information as an additional clue to estimate the timings of the accent commands. The experiments confirmed that the present algorithm was approximately 16 times faster and estimated parameters about 3% more accurately than the conventional algorithm.
Ryotaro Sato, Hirokazu Kameoka, Kunio Kashino
ICASSP2
2017 A noise suppression method for body-conducted soft speech based on non-negative tensor factorization of air- and body-conducted signals
abstract
This paper presents a novel noise suppression method to enhance soft speech recorded with a special body-conductive microphone called nonaudible murmur (NAM) microphone. NAM microphone is capable of detecting extremely soft speech, but the recorded soft speech easily suffers from external noise due to its faint volume. To effectively suppress noise on the body-conducted signals, an external noise monitoring framework using an air-conducive microphone has been proposed. In this study, we propose a noise suppression method for this framework based on a probabilistic observation model robust against phase variations. In the proposed method, noise suppression process is formulated as a special case of non-negative tensor factorization of the observed air- and body-conducted signals. Experimental results demonstrate that 1) the proposed method consistently outperforms the conventional method under real noisy environments and 2) the proposed method effectively deals with speech acoustic changes caused by the Lombard reflex.
Yusuke Tajiri, Hirokazu Kameoka, Tomoki Toda
ICASSP2
2017 DNN-SPACE: DNN-HMM-Based Generative Model of Voice F0 Contours for Statistical Phrase/Accent Command Estimation
Nobukatsu Hojo, Yasuhito Ohsugi, Yusuke Ijima, Hirokazu Kameoka
INTERSPEECH4
2017 Sequence-to-Sequence Voice Conversion with Similarity Metric Learned Using Generative Adversarial Networks
Takuhiro Kaneko, Hirokazu Kameoka, Kaoru Hiramatsu, Kunio Kashino
INTERSPEECH2
2017 Generative Adversarial Network-Based Postfilter for STFT Spectrograms
abstract
We propose a learning-based postfilter to reconstruct the high-fidelity spectral texture in short-term Fourier transform (STFT) spectrograms. In speech-processing systems, such as speech synthesis, voice conversion, and speech enhancement, the STFT spectrograms have been widely used as key acoustic representations. In these tasks, we normally need to precisely generate or predict the representations from inputs; however, generated spectra typically lack the fine structures close to the true data. To overcome these limitations and reconstruct spectra having finer structures, we propose a generative adversarial network (GAN)-based postfilter that is implicitly optimized to match the true feature distribution in adversarial learning. The challenge with this postfilter is that a GAN cannot be easily trained for very high-dimensional data such as the STFT. Therefore, we introduce a divide-and-concatenate strategy. We first divide the spectrograms into multiple frequency bands with overlap, train the GAN-based postfilter for the individual bands, and finally connect the bands with overlap. We applied our proposed postfilter to a DNN-based speech-synthesis task. The results show that our proposed postfilter can be used to reduce the gap between synthesized and target spectra, even in the highdimensional STFT domain.
Takuhiro Kaneko, Shinji Takaki, Hirokazu Kameoka, Junichi Yamagishi
INTERSPEECH3
2017 Speech Enhancement Using Non-Negative Spectrogram Models with Mel-Generalized Cepstral Regularization
Li Li 0063, Hirokazu Kameoka, Tomoki Toda, Shoji Makino
INTERSPEECH2
2017 Direct Modeling of Frequency Spectra and Waveform Generation Based on Phase Recovery for DNN-Based Speech Synthesis
abstract
In statistical parametric speech synthesis (SPSS) systems using the high-quality vocoder, acoustic features such as melcepstrum coefficients and F0 are predicted from linguistic features in order to utilize the vocoder to generate speech waveforms. However, the generated speech waveform generally suffers from quality deterioration such as buzziness caused by utilizing the vocoder. Although several attempts such as improving an excitation model have been investigated to alleviate the problem, it is difficult to completely avoid it if the SPSS system is based on the vocoder. To overcome this problem, there have recently been attempts to directly model waveform samples. Superior performance has been demonstrated, but computation time and latency are still issues. With the aim to construct another type of DNN-based speech synthesizer with neither the vocoder nor computational explosion, we investigated direct modeling of frequency spectra and waveform generation based on phase recovery. In this framework, STFT spectral amplitudes that include harmonics information derived from F0 are directly predicted through a DNN-based acoustic model and we use Griffin and Lim’s approach to recover phase and generate waveforms. The experimental results showed that the proposed system synthesized speech without buzziness and outperformed one generated from a conventional system using the vocoder.
Shinji Takaki, Hirokazu Kameoka, Junichi Yamagishi
INTERSPEECH2
2017 Physically Constrained Statistical F0 Prediction for Electrolaryngeal Speech Enhancement
Kou Tanaka, Hirokazu Kameoka, Tomoki Toda, Satoshi Nakamura 0001
INTERSPEECH2
2016 Sparse sound field decomposition with multichannel extension of complex NMF
abstract
A sparse sound field decomposition method using prior information on source signals in the time-frequency domain is proposed. Sparse sound field decomposition has been proved to be effective for various acoustic signal processing applications. Current methods for sparse decomposition are based only on the spatial sparsity of the source distribution. However, it can be assumed that possible source signals to be decomposed are approximately known in advance. To exploit this prior information, we incorporated the complex nonnegative factorization model into sparse sound field decomposition. Since the magnitude spectrum of the possible source signals can be trained in advance, accuracy of the sparse decomposition can be improved even when the source signals are highly correlated and the sources are in a highly noisy environment. In addition, the proposed decomposition algorithm is derived using the auxiliary function method. Numerical experiments indicated that the sparse decomposition performance was significantly improved using the proposed method.
Naoki Murata, Shoichi Koyama, Hirokazu Kameoka, Norihiro Takamune, Hiroshi Saruwatari
ICASSP3
2016 Shifted and convolutive source-filter non-negative matrix factorization for monaural audio source separation
abstract
This paper proposes an extension of non-negative matrix factorization (NMF), which combines the shifted NMF model with the source-filter model. Shifted NMF was proposed as a powerful approach for monaural source separation and multiple fundamental frequency (F0) estimation, which is particularly unique in that it takes account of the constant inter-harmonic spacings of a harmonic structure in log-frequency representations and uses a shifted copy of a spectrum template to represent the spectra of different F0s. However, for those sounds that follow the source-filter model, this assumption does not hold in reality, since the filter spectra are usually invariant under F0 changes. A more reasonable way to represent the spectrum of a different F0 is to use a shifted copy of a harmonic structure template as the excitation spectrum and keep the filter spectrum fixed. Thus, we can describe the spectrogram of a mixture signal as the sum of the products between the shifted copies of excitation spectrum templates and filter spectrum templates. Furthermore, the time course of filter spectra represents the dynamics of the timbre, which is important for characterizing the feature of an instrument sound. Thus, we further incorporate the non-negative matrix factor deconvolution (NMFD) model into the above model to describe the filter spectrogram. We derive a computationally efficient and convergence-guaranteed algorithm for estimating the unknown parameters of the constructed model based on the auxiliary function approach. Experimental results revealed that the proposed method outperformed shifted NMF in terms of the source separation accuracy.
Tomohiko Nakamura, Hirokazu Kameoka
ICASSP2
2016 Statistical F0 prediction for electrolaryngeal speech enhancement considering generative process of F0 contours within product of experts framework
abstract
We have previously proposed a statistical fundamental frequency (F0) prediction method that makes it possible to predict the underlying F0contour of electrolaryngeal (EL) speech from its spectral feature sequence. Although this method was shown to contribute to improving the naturalness of EL speech as a whole, the predicted F0contour was still unnatural compared with that in normal speech. One possible solution to improve the naturalness of the predicted F0contours would be to take account of the physical mechanism of vocal phonation. Recently a statistical model of voice F0contours was formulated by constructing a stochastic counterpart of the Fujisaki model, a well-founded mathematical model representing the control mechanism of vocal fold vibration. This paper proposes a Product-of-Experts model to incorporate this generative model of voice F0contours into the statistical F0prediction model. Based on the constructed model, we derive algorithms for parameter training and F0prediction. Experimental results revealed that the proposed method successfully outperformed our previously proposed method in terms of the naturalness of the predicted F0contours.
Kou Tanaka, Hirokazu Kameoka, Tomoki Toda, Satoshi Nakamura 0001
ICASSP2
2016 Majorisation-Minimisation Based Optimisation of the Composite Autoregressive System with Application to Glottal Inverse Filtering
abstract
The composite autoregressive system can be used to estimate a speech source-filter decomposition in a rigorous manner, thus having potential use in glottal inverse filtering. By introducing a suitable prior, spectral tilt can be introduced into the source component estimation to better correspond to human voice production. However, the current expectation-maximisation based composite autoregressive model optimisation leaves room for improvement in terms of speed. Inspired by majorisation-minimisation techniques used for nonnegative matrix factorisation, this work derives new update rules for the model, resulting in faster convergence compared to the original approach. Additionally, we present a new glottal inverse filtering method based on the composite autoregressive system and compare it with inverse filtering methods currently used in glottal excitation modelling for parametric speech synthesis. These initial results show that the proposed method performs comparatively well, sometimes outperforming the reference methods.
Lauri Juvela, Hirokazu Kameoka, Manu Airaksinen, Junichi Yamagishi, Paavo Alku
INTERSPEECH2
2016 Semi-Supervised Joint Enhancement of Spectral and Cepstral Sequences of Noisy Speech
Li Li 0063, Hirokazu Kameoka, Takuya Higuchi, Hiroshi Saruwatari
INTERSPEECH2
2016 Acoustic-to-Articulatory Inversion Mapping Based on Latent Trajectory Gaussian Mixture Model
Patrick Lumban Tobing, Tomoki Toda, Hirokazu Kameoka, Satoshi Nakamura 0001
INTERSPEECH3
2016 Determined Blind Source Separation Unifying Independent Vector Analysis and Nonnegative Matrix Factorization
abstract
This paper addresses the determined blind source separation problem and proposes a new effective method unifying independent vector analysis (IVA) and nonnegative matrix factorization (NMF). IVA is a state-of-the-art technique that utilizes the statistical independence between sources in a mixture signal, and an efficient optimization scheme has been proposed for IVA. However, since the source model in IVA is based on a spherical multivariate distribution, IVA cannot utilize specific spectral structures such as the harmonic structures of pitched instrumental sounds. To solve this problem, we introduce NMF decomposition as the source model in IVA to capture the spectral structures. The formulation of the proposed method is derived from conventional multichannel NMF (MNMF), which reveals the relationship between MNMF and IVA. The proposed method can be optimized by the update rules of IVA and single-channel NMF. Experimental results show the efficacy of the proposed method compared with IVA and MNMF in terms of separation accuracy and convergence speed.
Daichi Kitamura, Nobutaka Ono, Hiroshi Sawada, Hirokazu Kameoka, Hiroshi Saruwatari
IEEE ACM Trans. Audio Speech Lang. Process.4
2015 Multi-resolution signal decomposition with time-domain spectrogram factorization
abstract
This paper proposes a novel framework that makes it possible to realize non-negative matrix factorization (NMF)-like signal decompositions in the time-domain. This new formulation also allows for an extension to multi-resolution signal decomposition, which was not possible with the conventional NMF framework.
Hirokazu Kameoka
ICASSP1
2015 Efficient multichannel nonnegative matrix factorization exploiting rank-1 spatial model
abstract
This paper proposes a new efficient multichannel nonnegative matrix factorization (NMF) method. Recently, multichannel NMF (MNMF) has been proposed as a means of solving the blind source separation problem. This method estimates a mixing system of sources and attempts to separate them in a blind fashion. However, this method is strongly dependent on its initial values because there are no constraints in the spatial models. To solve this problem, we introduce a rank-1 spatial model into MNMF. The proposed method estimates a demixing matrix while representing sources using NMF bases and can be optimized by the update rules of independent vector analysis and conventional single-channel NMF. Experimental results show the efficacy of the proposed method in terms of robustness and convergence speed.
Daichi Kitamura, Nobutaka Ono, Hiroshi Sawada, Hirokazu Kameoka, Hiroshi Saruwatari
ICASSP4
2015 Lp-norm non-negative matrix factorization and its application to singing voice enhancement
abstract
Measures of sparsity are useful in many aspects of audio signal processing including speech enhancement, audio coding and singing voice enhancement, and the well-known method for these applications is non-negative matrix factorization (NMF), which decomposes a non-negative data matrix into two non-negative matrices. Although previous studies on NMF have focused on the sparsity of the two matrices, the sparsity of reconstruction errors between a data matrix and the two matrices is also important, since designing the sparsity is equivalent to assuming the nature of the errors. We propose a new NMF technique, which we called Lp-norm NMF, that minimizes the Lpnorm of the reconstruction errors, and derive a computationally efficient algorithm for Lp-norm NMF according to an auxiliary function principle. This algorithm can be generalized for the factorization of a real-valued matrix into the product of two real-valued matrices. We apply the algorithm to singing voice enhancement and show that adequately selecting p improves the enhancement.
Tomohiko Nakamura, Hirokazu Kameoka
ICASSP2
2015 Generative Modeling of Voice Fundamental Frequency Contours
abstract
This paper introduces a generative model of voice fundamental frequency (F0) contours that allows us to extract prosodic features from raw speech data. The present F0contour model is formulated by translating the Fujisaki model, a well-founded mathematical model representing the control mechanism of vocal fold vibration, into a probabilistic model described as a discrete-time stochastic process. There are two motivations behind this formulation. One is to derive a general parameter estimation framework for the Fujisaki model that allows the introduction of powerful statistical methods. The other is to construct an automatically trainable version of the Fujisaki model that we can incorporate into statistical-model-based text-to-speech synthesizers in such a way that the Fujisaki-model parameters can be learned from a speech corpus in a unified manner. It could also be useful for other speech applications such as emotion recognition, speaker identification, speech conversion and dialogue systems, in which prosodic information plays a significant role. We quantitatively evaluated the performance of the proposed Fujisaki model parameter extractor using real speech data. Experimental results revealed that our method was superior to a state-of-the-art Fujisaki model parameter extractor.
Hirokazu Kameoka, Kota Yoshizato, Tatsuma Ishihara, Kento Kadowaki, Yasunori Ohishi, Kunio Kashino
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 Multichannel Signal Separation Combining Directional Clustering and Nonnegative Matrix Factorization with Spectrogram Restoration
abstract
In this paper, to address problems in multichannel music signal separation, we propose a new hybrid method that combines directional clustering and advanced nonnegative matrix factorization (NMF). The aims of multichannel music signal separation technology is to extract a specific target signal from observed multichannel signals that contain multiple instrumental sounds. In previous studies, various methods using NMF have been proposed, but many problems remain including poor separation accuracy and lack of robustness. To solve these problems, we propose a new supervised NMF (SNMF) with spectrogram restoration and a hybrid method that concatenates the proposed SNMF after directional clustering. Via the extrapolation of supervised spectral bases, the proposed SNMF attempts both target signal separation and reconstruction of the lost target components, which are generated by preceding directional clustering. In addition, we experimentally reveal the trade-off between separation and extrapolation abilities and propose a new scheme for adaptive divergence, where the optimal divergence can be automatically changed in each time frame according to the local spatial conditions. The results of an evaluation experiment show that our proposed hybrid method outperforms the conventional music signal separation methods.
Daichi Kitamura, Hiroshi Saruwatari, Hirokazu Kameoka, Yu Takahashi, Kazunobu Kondo, Satoshi Nakamura 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 Resolution Warped Spectral Representation for Low-Delay and Low-Bit-Rate Audio Coder
abstract
We have devised a high-quality frequency-domain audio coder based on the state-of-the-art monaural wide-band coder aiming at its use in low-delay and low-bit-rate conditions. The coder efficiently represents frequency spectral envelopes of the target signals with low computational complexity using optimally prepared non-negative sparse matrices. The experimental results reveal that this representation has positive effects on the objective and subjective quality of the coder resulting in the comparable quality to the same bit rate of 3GPP Extended Adaptive Multi-Rate WideBand ( AMR-WB+), a coder which permits more than four times longer delay compared with the proposed coder. Consequently, this coder is suitable for applications in mobile communications, which require low delay and low complexity.
Ryosuke Sugiura, Yutaka Kamamoto, Noboru Harada, Hirokazu Kameoka, Takehiro Moriya
IEEE ACM Trans. Audio Speech Lang. Process.4
2015 Optimal Coding of Generalized-Gaussian-Distributed Frequency Spectra for Low-Delay Audio Coder With Powered All-Pole Spectrum Estimation
abstract
We present an optimal coding scheme that parameterizes the maximum-likelihood estimate of variance for frequency spectra belonging to the generalized Gaussian distribution, the distribution covering the Laplacian and the Gaussian. By slightly modifying the all-pole model of the conventional linear prediction (LP), we can estimate the variance with the same method as in LP, which has low computational costs. Experimental results show that incorporating the coding scheme in a state-of-the-art wide-band audio coder enhances its objective and subjective quality in a low-bit-rate and low-delay situation by increasing the compression efficiency. Thus, this coding scheme will be useful in applications like mobile communications, which requires highly efficient compression.
Ryosuke Sugiura, Yutaka Kamamoto, Noboru Harada, Hirokazu Kameoka, Takehiro Moriya
IEEE ACM Trans. Audio Speech Lang. Process.4
2014 Fast Signal Reconstruction from Magnitude Spectrogram of Continuous Wavelet Transform Based on Spectrogram Consistency
Tomohiko Nakamura, Hirokazu Kameoka
DAFx2
2014 Underdetermined blind separation and tracking of moving sources based ONDOA-HMM
abstract
This paper deals with the problem of the underdetermined blind separation and tracking of moving sources. In practical situations, sound sources such as human speakers can move freely and so blind separation algorithms must be designed to track the temporal changes of the impulse responses. We propose solving this problem through the posterior inference of the parameters in a generative model of an observed multichannel signal, formulated under the assumption of the sparsity of time-frequency components of speech and the continuity of speakers' movements. Specifically, we describe a generative model of mixture signals by incorporating a generative model of a time-varying frequency array response for each source, described using a path-restricted hidden Markov model (HMM). Each hidden state of the present HMM represents the direction of arrival (DOA) of each source, and so we call it a “DOA-HMM.” Through the posterior inference of the overall generative model, we can simultaneously track the DOAs of sources, separate source signals and perform permutation alignment. The experiment showed that the proposed algorithm provided a 6.20 dB improvement compared with the conventional method in terms of the signal-to-interference ratio.
Takuya Higuchi, Norihiro Takamune, Tomohiko Nakamura, Hirokazu Kameoka
ICASSP4
2014 Timbre replacement of harmonic and drum components for music audio signals
abstract
This paper presents a system that allows users to customize an audio signal of polyphonic music (input), without using musical scores, by replacing the frequency characteristics of harmonic sounds and the timbres of drum sounds with those of another audio signal of polyphonic music (reference). To develop the system, we first use a method that can separate the amplitude spectra of the input and reference signals into harmonic and percussive spectra. We characterize frequency characteristics of the harmonic spectra by two envelopes tracing spectral dips and peaks roughly, and the input harmonic spectra are modified such that their envelopes become similar to those of the reference harmonic spectra. The input and reference percussive spectrograms are further decomposed into those of individual drum instruments, and we replace the timbres of those drum instruments in the input piece with those in the reference piece. Through the subjective experiment, we show that our system can replace drum timbres and frequency characteristics adequately.
Tomohiko Nakamura, Hirokazu Kameoka, Kazuyoshi Yoshii, Masataka Goto
ICASSP2
2014 Mondrian hidden Markov model for music signal processing
abstract
This paper discusses a new extension of hidden Markov models that can capture clusters embedded in transitions between the hidden states. In our model, the state-transition matrices are viewed as representations of relational data reflecting a network structure between the hidden states. We specifically present a nonparametric Bayesian approach to the proposed state-space model whose network structure is represented by a Mondrian Process-based relational model. We show an application of the proposed model to music signal analysis through some experimental results.
Masahiro Nakano, Yasunori Ohishi, Hirokazu Kameoka, Ryo Mukai, Kunio Kashino
ICASSP3
2014 Mixture of Gaussian process experts for predicting sung melodic contour with expressive dynamic fluctuations
abstract
We present a generative model for predicting the sung melodic contour, i.e., F0contour, with expressive dynamic fluctuations, such as vibrato and portamento, for a given musical score. Although several studies have attempted to characterize such fluctuations, no systematic method has been developed for generating the F0contour with them in connection with musical notes. In our model, the relationship between a musical note sequence and F0contour is directly learned by a mixture of Gaussian process experts. This approach allows us to automatically characterize the fluctuations by utilizing the kernel function for each Gaussian process expert and predict the F0contour for an arbitrary musical note sequence. Experimental results show that our model can better predict the F0contour than a baseline method can. Additionally, we discuss the effective musical contexts and the amount of training data for the prediction.
Yasunori Ohishi, Daichi Mochihashi, Hirokazu Kameoka, Kunio Kashino
ICASSP3
2014 A unified approach for underdetermined blind signal separation and source activity detection by multichannel factorial hidden Markov models
abstract
This paper proposes to introduce a new model called “the multichannel factorial hidden Markov Model (MFHMM)” for underdetermined blind signal separation (BSS). For monaural source separation, one successful approach involves applying nonnegative matrix factorization (NMF) to the magnitude spectrogram of a mixture signal, interpreted as a non-negative matrix. Up to now, multichannel extensions of NMF, which allow for the use of spatial information as an additional clue for source separation, have been proposed by several authors and proven to be an effective approach for underdetermined BSS. This approach is based on the assumption that an observed signal is a mixture of a limited number of source signals each of which has a static power spectral density scaled by a time-varying amplitude. However, many source signals in real world are nonstationary in nature and the variations of the spectral densities are much richer in time. Moreover, many sources including speech tend to stay inactive for some while until they switch to an active mode, implying that the total power of a source may depend on its underlying state. To reasonably characterize such a non-stationary nature of source signals, this paper proposes to extend the multichannel NMF model by modeling the transition of the set consisting of the spectral densities and the total power of each source using a hidden Markov model (HMM). By letting each HMM contain states corresponding to active and inactive modes, we will show that voice activity detection and source separation can be solved simultaneously through parameter inference of the present model. The experiment showed that the proposed algorithm provided a 7.65 dB improvement compared with the conventional multichannel NMF in terms of the signal-to-distortion ratio. Index Terms: blind signal separation, source activity detection, a hidden Markov model, non-negative matrix factorization
Takuya Higuchi, Hirofumi Takeda, Tomohiko Nakamura, Hirokazu Kameoka
INTERSPEECH4
2014 Speech prosody generation for text-to-speech synthesis based on generative model of F0 contours
abstract
This paper deals with the problem of generating the fundamental frequency (F0) contour of speech from a text input for text-to-speech synthesis. We have previously introduced a statistical model describing the generating process of speech F0 contours, based on the discrete-time version of the Fujisaki model. One remarkable feature of this model is that it has allowed us to derive an efficient algorithm based on powerful statistical methods for estimating the Fujisaki-model parameters from raw F0 contours. To associate a sequence of the Fujisakimodel parameters with a text input based on statistical learning, this paper proposes extending this model to a context-dependent one. We further propose a parameter training algorithm for the present model based on a decision tree-based context clustering. Index Terms: Speech F0 contours, stochastic model, Fujisaki model, hidden Markov model, EM algorithm
Kento Kadowaki, Tatsuma Ishihara, Nobukatsu Hojo, Hirokazu Kameoka
INTERSPEECH4
2014 Harmonic/percussive sound separation based on anisotropic smoothness of spectrograms
abstract
This paper describes a method to separate a monaural music signal into harmonic components e.g., a guitar and percussive components, e.g., a snare drum. Separation of these two components is a useful preprocessing for many music information retrieval applications, and in addition, it can be used as a new kind of music equalizer in itself, which enables a music listener to adjust the ratio of the volume of the guitar and the drum freely by themselves. Because of these potential applications, there have been many attempts to develop such a technique, especially in the last decade. However, some of the state-of-the-art techniques have a drawback that they are based on costly operations, such as the multiplications of large-sized matrix, Monte Carlo method, etc., which may constitute barriers to the practical use on some small computers such as smart phones. In this paper, an efficient method that does not depend on these costly operations is described. In formulating the methods, the authors basically assumed only the “anisotropic smoothness” of music spectrogram, which can be one of the minimalistic model that reflects the natures of these instruments. To be specific, the authors just assumed that harmonic instruments are smooth in time, while the percussive instruments are smooth in frequency on a music spectrogram. In this paper, on the basis of the assumption, source separation methods are formulated as optimization problems that optimize the “anisotropic smoothness” under some conditions. Because of the simplicity of the model, the derived algorithms are quite simple. Experimental results show that the methods were effective compared to a state-of-the-art technique, and the computation time was much shorter than an existing method; specifically, it can process a three-minute song in around 4-20 seconds on a laptop PC.
Hideyuki Tachibana, Nobutaka Ono, Hirokazu Kameoka, Shigeki Sagayama
IEEE ACM Trans. Audio Speech Lang. Process.3
2013 Bayesian semi-supervised audio event transcription based on Markov indian buffet process
abstract
We present a novel generative model for audio event transcription that recognizes “events” on audio signals including multiple kinds of overlapping sounds. In the proposed model, firstly, the overlapping audio events are modeled based on nonnegative matrix factorization into which Bayesian nonparametric approaches: the Markov Indian buffet process and the Chinese restaurant process, are incorporated. This approach allows us to automatically transcribe the events while avoiding the model selection problem by assuming a countably infinite number of possible audio events in the input signal. Then, Bayesian logistic regression annotates the audio frames with the multiple event labels in a semi-supervised learning setup. Experimental results show that our model can better annotate an audio signal in comparison with a baseline method. Additionally, we verify that our infinite generative model is also able to detect unknown audio events that are not included in the training data.
Yasunori Ohishi, Daichi Mochihashi, Tomoko Matsui, Masahiro Nakano, Hirokazu Kameoka, Tomonori Izumitani, Kunio Kashino
ICASSP5
2013 Probabilistic speech F0 contour model incorporating statistical vocabulary model of phrase-accent command sequence
abstract
We have previously proposed a generative model of speech F0 contours, based on the discrete-time version of the Fujisaki model (a model of the mechanisim for controlling F0s through laryngeal muscles). One advantage of this model is that it allows us to apply statistical methods to estimate the Fujisakimodel parameters from speech F0 contours. This paper proposes a new generative model of speech F0 contours incorporating a vocabulary model of intonation patterns. A parameter inference algorithm for the present model is derived. We quantitatively evaluated the performance of our parameter inference algorithm.
Tatsuma Ishihara, Hirokazu Kameoka, Kota Yoshizato, Daisuke Saito, Shigeki Sagayama
INTERSPEECH2
2013 Generative modeling of speech F0 contours
abstract
This paper introduces our ongoing work on generative modeling of speech fundamental frequency (F0) contours for estimating prosodic features from raw speech data. The present F0 contour model is formulated by translating the Fujisaki model, a well-founded mathematical model representing the control mechanism of vocal fold vibration, into a probabilistic model described as a discrete-time stochastic process. The motivation behind this formulation is two fold. One is to derive a general parameter estimation framework for the Fujisaki model, allowing for the introduction of powerful statistical methods. The other is to construct an automatically trainable version of the Fujisaki model so that in future it can be used to develop a statistical speaking style conversion system or incorporated into existing text-to-speech synthesis systems to improve the naturalness and intelligibility of computer-generated speech. We also briefly introduce a generative model of F0 contours of singing voice developed under the same spirit. Index Terms: speech F0 contour, Fujisaki model, generative model, hidden Markov model, EM algorithm
Hirokazu Kameoka, Kota Yoshizato, Tatsuma Ishihara, Yasunori Ohishi, Kunio Kashino, Shigeki Sagayama
INTERSPEECH1
2013 Multichannel Extensions of Non-Negative Matrix Factorization With Complex-Valued Data
abstract
This paper presents new formulations and algorithms for multichannel extensions of non-negative matrix factorization (NMF). The formulations employ Hermitian positive semidefinite matrices to represent a multichannel version of non-negative elements. Multichannel Euclidean distance and multichannel Itakura-Saito (IS) divergence are defined based on appropriate statistical models utilizing multivariate complex Gaussian distributions. To minimize this distance/divergence, efficient optimization algorithms in the form of multiplicative updates are derived by using properly designed auxiliary functions. Two methods are proposed for clustering NMF bases according to the estimated spatial property. Convolutive blind source separation (BSS) is performed by the multichannel extensions of NMF with the clustering mechanism. Experimental results show that 1) the derived multiplicative update rules exhibited good convergence behavior, and 2) BSS tasks for several music sources with two microphones and three instrumental parts were evaluated successfully.
Hiroshi Sawada, Hirokazu Kameoka, Shoko Araki, Naonori Ueda
IEEE Trans. Speech Audio Process.2
2012 Constrained and regularized variants of non-negative matrix factorization incorporating music-specific constraints
abstract
Music spectrograms typically have many structural regularities that can be exploited to help solve the problem of decomposing a given spectrogram into distinct musically meaningful components. In this paper, we introduce new variants of the non-negative matrix factorization concept that incorporate music-specific constraints.
Hirokazu Kameoka, Masahiro Nakano, Kazuki Ochiai, Yutaka Imoto, Kunio Kashino, Shigeki Sagayama
ICASSP1
2012 Bayesian nonparametric music parser
abstract
This paper proposes a novel representation of music that can be used for similarity-based music information retrieval, and also presents a method that converts an input polyphonic audio signal to the proposed representation. The representation involves a 2-dimensional tree structure, where each node encodes the musical note and the dimensions correspond to the time and simultaneous multiple notes, respectively. Since the temporal structure and the synchrony of simultaneous events are both essential in music, our representation reflects them explicitly. In the conventional approaches to music representation from audio, note extraction is usually performed prior to structure analysis, but accurate note extraction has been a difficult task. In the proposed method, note extraction and structure estimation is performed simultaneously and thus the optimal solution is obtained with a unified inference procedure. That is, we propose an extended 2-dimensional infinite probabilistic context-free grammar and a sparse factor model for spectrogram analysis. An efficient inference algorithm, based on Markov chain Monte Carlo sampling and dynamic programming, is presented. The experimental results show the effectiveness of the proposed approach.
Masahiro Nakano, Yasunori Ohishi, Hirokazu Kameoka, Ryo Mukai, Kunio Kashino
ICASSP3
2012 Explicit beat structure modeling for non-negative matrix factorization-based multipitch analysis
abstract
This paper proposes model-based non-negative matrix factorization (NMF) for estimating basis spectra and activations, detecting note onsets and offsets, and determining beat locations, simultaneously. Multipitch analysis is a process of detecting the pitch and onset of each note from a musical signal. Conventional NMF-based approaches often lead to unsatisfactory results very possibly due to the lack of musically meaningful constraints. As music is highly structured in terms of the temporal regularity underlying the onset occurrences of notes, we use this rhythmic structure to constrain NMF by parametrically modeling each note activation with a Gaussian mixture and derive an algorithm for iteratively updating model parameters. It is experimentally shown that the proposed model outperforms the standard NMF algorithms as regards onset detection rate.
Kazuki Ochiai, Hirokazu Kameoka, Shigeki Sagayama
ICASSP2
2012 Efficient algorithms for multichannel extensions of Itakura-Saito nonnegative matrix factorization
abstract
This paper proposes new algorithms for multichannel extensions of nonnegative matrix factorization (NMF) with the Itakura-Saito (IS) divergence. We employ Hermitian positive definite matrices for modeling the covariance matrix of a multivariate complex Gaussian distribution. Such matrices are basically estimated for NMF bases, but a source separation task can be performed by introducing variables that relate NMF bases and sources. The new algorithms are derived by using a majorization scheme with properly designed auxiliary functions. The algorithms are in the form of multiplicative updates, and exhibit good convergence behavior. We have succeeded in separating a professionally produced music recording into its vocal and guitar components.
Hiroshi Sawada, Hirokazu Kameoka, Shoko Araki, Naonori Ueda
ICASSP2
2012 Comparative evaluations of various harmonic/percussive sound separation algorithms based on anisotropic continuity of spectrogram
abstract
In this paper, we explore several algorithms to find the best performing algorithm for harmonic and percussive sound separation (HPSS) based on anisotropic continuity of spectrogram through comparative evaluation of their experimental performance. Separating harmonic and percussive sounds is useful as a preprocessor for many music analysis purposes including chord estimation, rhythm analysis, and other music information retrieval tasks. We have introduced a method called “Harmonic/Percussive Sound Separation” (HPSS), that decomposes a music signal into two components by separating the spectrogram into horizontally-continuous and vertically-continuous components, which roughly correspond to harmonic and percussive sounds, respectively. Many possible ways exist to realize the HPSS algorithm based on this concept while it has been unknown which algorithm performs best. This paper describes the details of five different HPSS algorithms and compares their performances over real music signals.
Hideyuki Tachibana, Hirokazu Kameoka, Nobutaka Ono, Shigeki Sagayama
ICASSP2
2012 Designing various component analysis at will
Akisato Kimura, Hitoshi Sakano, Hirokazu Kameoka, Masashi Sugiyama
ICPR3
2012 A Stochastic Model of Singing Voice F0 Contours for Characterizing Expressive Dynamic Components
abstract
We present a novel stochastic model of singing voice fundamental frequency (F0) contours for characterizing expressive dynamic components, such as vibrato and portamento. Although dynamic components can be important features for any singing voice applications, modeling and extracting these components from a raw F0 contour have yet to be accomplished. Therefore, we describe a process for generating dynamic components explicitly and represent the process as a stochastic model. Then we develop an algorithm for estimating the model parameters based on statistical techniques. Experimental results show that our method successfully extracts the expressive components from raw F0 contours.
Yasunori Ohishi, Hirokazu Kameoka, Daichi Mochihashi, Kunio Kashino
INTERSPEECH2
2012 Hidden Markov Convolutive Mixture Model for Pitch Contour Analysis of Speech
abstract
This paper proposes a stochastic model of speech F0 contours, based on the stochastic formulation of the Fujisaki model. Our motivation for the stochastic formulation is twofold. Firstly, it allows us to derive a well-behaved algorithm for estimating the Fujisaki model parameters from a raw F0 contour. Secondly, it will open the door to incorporating the well-founded F0 contour model into various statistical speech processing problems. We quantitatively evaluated the performance of our method in terms of an Fujisaki-model parameter estimation accuracy using real speech data. Experimental results revealed that our method was superior to a state-of-the-art Fujisaki model parameter extractor. Index Terms: speech F0 contours, statistical model, Fujisaki model, hidden Markov model, EM algorithm
Kota Yoshizato, Hirokazu Kameoka, Daisuke Saito, Shigeki Sagayama
INTERSPEECH2
2011 Automatic video annotation via Hierarchical Topic Trajectory Model considering cross-modal correlations
abstract
We propose a new statistical model, named Hierarchical Topic Trajectory Model (HTTM), for acquiring a dynamically changing topic model that represents the relationship between video frames and associated text labels. Model parameter estimation, annotation and retrieval can be executed within a unified framework with a few computation. It is also easy to add new modals such as audio signal and geotags. Preliminary experiments on video annotation task with manually annotated video dataset indicate that our proposed method can improve the annotation accuracy.
Takuho Nakano, Akisato Kimura, Hirokazu Kameoka, Shigeki Miyabe, Shigeki Sagayama, Nobutaka Ono, Kunio Kashino, Takuya Nishimoto
ICASSP3
2011 Infinite-state spectrum model for music signal analysis
abstract
This paper presents a nonparametric Bayesian extension of non-negative matrix factorization (NMF) for music signal analysis. Instrument sounds often exhibit non-stationary spectral characteristics. We introduce infinite-state spectral bases into NMF to represent time-varying spectra in polyphonic music signals. We describe our extension of NMF with infinite-state spectral bases generated by the Dirichlet process in a statistical framework, derive an efficient optimization algorithm based on collapsed variational inference, and validate the framework on audio data.
Masahiro Nakano, Jonathan Le Roux, Hirokazu Kameoka, Nobutaka Ono, Shigeki Sagayama
ICASSP3
2011 Formulations and algorithms for multichannel complex NMF
abstract
This paper studies some formulations and algorithms for the multichannel extension of nonnegative matrix factorization (NMF). We model the inter-channel characteristics of each NMF basis, including both the amplitude ratios and the phase differences on a channel pair. The learned inter-channel characteristics provide useful information for binding each NMF basis to each source component in such a situation that multiple sources are mixed in a convolutive manner and observed at multiple microphones. Effective optimization algorithms based on majorization are derived by using properly designed auxiliary functions. Experimental results show that the algorithms converged favorably regardless of the initialization.
Hiroshi Sawada, Hirokazu Kameoka, Shoko Araki, Naonori Ueda
ICASSP2
2011 Automatic audio tag classification via semi-supervised canonical density estimation
abstract
We propose a novel semi-supervised method for building a statistical model that represents the relationship between sounds and text labels ("tags"). The proposed method, named semi-supervised canonical density estimation, makes use of unlabeled sound data in two ways: 1) a low-dimensional latent space representing topics of sounds is extracted by a semi-supervised variant of canonical correlation analysis, and 2) topic models are learned by multi-class extension of semi-supervised kernel density estimation in the topic space. Real-world audio tagging experiments indicate that our pro posed method improves the accuracy even when only a small number of labeled sounds are available.
Jun Takagi, Yasunori Ohishi, Akisato Kimura, Masashi Sugiyama, Makoto Yamada, Hirokazu Kameoka
ICASSP6
2011 I-Divergence-based dereverberation method with auxiliary function approach
abstract
This paper presents a dereverberation method based on I-divergence minimization, which is particularly suitable for music signals. Existing dereverberation methods, including one designed for music, sometimes distort instrument sounds and make staccato-like tones. The problems with the Itakura-Saito-divergence-based formulation of the existing methods are their tendency to excessive suppression of direct sound and the difficulty of incorporating and optimizing sophisticated source models suitable for music signals. The proposed I-divergence-based method can mitigate these problems. Employing the I-divergence measure enables us to avoid the direct sound suppression problem and to use powerful music spectrum models without complicating its optimization. We develop a convergence-guaranteed parameter estimation algorithm based on the auxiliary function approach. Experimental results reveal the effectiveness of the proposed dereverberation method.
Naoki Yasuraoka, Hirokazu Kameoka, Takuya Yoshioka, Hiroshi G. Okuno
ICASSP2
2011 Computational auditory induction as a missing-data model-fitting problem with Bregman divergence
Jonathan Le Roux, Hirokazu Kameoka, Nobutaka Ono, Alain de Cheveigné, Shigeki Sagayama
Speech Commun.2
2010 A sparse component model of source signals and its application to blind source separation
abstract
In this paper, we propose a new method of blind source separation (BSS) for music signals. Our method has the following characteristics: 1) the method is a combination of the sparseness-based model of source signals and the factorized basis model in nonnegative matrix factorization (NMF), 2) it is assumed that only one basis which structure source signals is active at each time-frequency bin of the observed signals, in order to degrade the degree of freedom, 3) parameter estimation algorithm is based on the EM algorithm regarding the index of the only one active basis as the hidden variable. We develop the formulation at a different point from NMF and show source separation performance in some simulation experiments.
Yu Kitano, Hirokazu Kameoka, Yosuke Izumi, Nobutaka Ono, Shigeki Sagayama
ICASSP2
2010 SemiCCA: Efficient Semi-supervised Learning of Canonical Correlations
abstract
Canonical correlation analysis (CCA) is a powerful tool for analyzing multi-dimensional paired data. However, CCA tends to perform poorly when the number of paired samples is limited, which is often the case in practice. To cope with this problem, we propose a semi-supervised variant of CCA named "Semi CCA" that allows us to incorporate additional unpaired samples for mitigating overfitting. The proposed method smoothly bridges the eigenvalue problems of CCA and principal component analysis (PCA), and thus its solution can be computed efficiently just by solving a single (generalized) eigenvalue problem as the original CCA. Preliminary experiments with artificially generated samples and PASCAL VOC data sets demonstrate the effectiveness of the proposed method.
Akisato Kimura, Hirokazu Kameoka, Masashi Sugiyama, Takuho Nakano, Eisaku Maeda, Hitoshi Sakano, Katsuhiko Ishiguro
ICPR2
2010 Statistical modeling of F0 dynamics in singing voices based on Gaussian processes with multiple oscillation bases
abstract
We present a novel statistical model for dynamics of various singing behaviors, such as vibrato and overshoot, in a fundamental frequency (F0) contour. These dynamics are the important cues for perceiving individuality of a singer, and can be a useful measure for various applications, such as singing skill evaluation and singing voice synthesis. While most previous studies have modeled the dynamics using a second-order linear system, the automatic and accurate estimation of model parameters has yet to be accomplished. In this paper, we first develop a complete stochastic representation of the second-order system with Gaussian processes from parametric discretization, and propose a complete, efficient scheme for parameter estimation using the Expectation-Maximization (EM) algorithm. Experimental results show that the proposed method can decompose an F0 contour into a musical component and a dynamics component. Finally, we discuss estimating singing styles from the model parameters for each singer.
Yasunori Ohishi, Hirokazu Kameoka, Daichi Mochihashi, Hidehisa Nagano, Kunio Kashino
INTERSPEECH2
2010 Speech Spectrum Modeling for Joint Estimation of Spectral Envelope and Fundamental Frequency
abstract
Although considerable effort has been devoted to both fundamental frequency (F0) and spectral envelope estimation in the field of speech processing, the problem of determiningF0and spectral envelopes has largely been tackled independently. IfF0were known in advance, then the spectral envelope could be estimated very reliably. On the other hand, if the spectral envelope were known in advance, then we could obtain a reliableF0estimate.F0and the spectral envelope, each of which is a prerequisite of the other, should thus be estimated jointly rather than independently in succession. On this basis, we develop a parametric speech spectrum model that allows us to estimate theF0and spectral envelope simultaneously. We confirmed experimentally the significant advantage of this joint estimation approach for bothF0estimation and spectral envelope estimation.
Hirokazu Kameoka, Nobutaka Ono, Shigeki Sagayama
IEEE Trans. Speech Audio Process.1
2009 Robust speech dereverberation based on non-negativity and sparse nature of speech spectrograms
abstract
This paper presents a blind dereverberation method designed to recover the subband envelope of an original speech signal from its reverberant version. The problem is formulated as a blind deconvolution problem with non-negative constraints, regularized by the sparse nature of speech spectrograms. We derive an iterative algorithm for its optimization, which can be seen as a special case of the non-negative matrix factor deconvolution. We confirmed through experiments that the algorithm is fast and robust to speaker movement.
Hirokazu Kameoka, Tomohiro Nakatani, Takuya Yoshioka
ICASSP1
2009 Complex NMF: A new sparse representation for acoustic signals
abstract
This paper presents a new sparse representation for acoustic signals which is based on a mixing model defined in the complex-spectrum domain (where additivity holds), and allows us to extract recurrent patterns of magnitude spectra that underlie observed complex spectra and the phase estimates of constituent signals. An efficient iterative algorithm is derived, which reduces to the multiplicative update algorithm for non-negative matrix factorization developed by Lee under a particular condition.
Hirokazu Kameoka, Nobutaka Ono, Kunio Kashino, Shigeki Sagayama
ICASSP1
2009 Composite Autoregressive System for Sparse Source-filter Representation of speech
abstract
This paper presents a new generative model for speech signals called a ldquocomposite autoregressive systemrdquo. This model consists of a composite dictionary incorporating a set of the power spectral densities (PSDs) of excitation sources and a set of all-pole filters where the gain of each pair of excitation and filter elements is allowed to vary over time. We use this model to develop a computationally efficient scheme for generating a sparse mixture representation of speech based on the Expectation-Maximization algorithm. The algorithm iteratively updates the excitation PSDs and the gains through the update formulae, which reduce under a particular condition to the multiplicative update rule for non-negative matrix factorization with the Itakura-Saito distance criterion, and the all-pole parameters using the Levinson-Durbin algorithm.
Hirokazu Kameoka, Kunio Kashino
ISCAS1
2008 Auxiliary function approach to parameter estimation of constrained sinusoidal model for monaural speech separation
abstract
We introduce in this paper an auxiliary function approach to parameter estimation of the constrained sinusoidal model, which enables us to derive a complex-spectrum-domain EM-like multiple F0estimation algorithm. Through simulations, we evaluated the performance of the presented method in the ability to avoid locally optimal solutions. We implemented a monaural speech separation system based on the presented method and confirmed its performance on compound signals of real speech.
Hirokazu Kameoka, Nobutaka Ono, Shigeki Sagayama
ICASSP1
2008 Harmonic-Temporal-Timbral Clustering (HTTC) for the analysis of multi-instrument polyphonic music signals
abstract
In this paper, we discuss a new approach named Harmonic-Temporal-Timbral Clustering (HTTC) for the analysis of single-channel audio signal of multi-instrument polyphonic music to estimate the pitch, onset timing, power and duration of all the acoustic events and to classify them into timbre categories simultaneously. Each acoustic event is modeled by a harmonic structure and a smooth envelope both represented by Gaussian mixtures. Based on the similarity between these spectro-temporal structures, timbres are clustered to form timbre categories. The entire process is mathematically formulated as a minimization problem for the I-divergence between the HTTC parametric model and the observed spectrogram of the music audio signal to simultaneously update harmonic, temporal and timbral model parameters through the EM algorithm. Some experimental results are presented to discuss the performance of the algorithm.
Kenichi Miyamoto, Hirokazu Kameoka, Takuya Nishimoto, Nobutaka Ono, Shigeki Sagayama
ICASSP2
2008 Modulation analysis of speech through orthogonal FIR filterbank optimization
abstract
Newborns must learn to structure incoming acoustic information into segments, words, phrases, etc., before they can start to learn language. This process is thought to rely on modulation structure of the speech waveform induced by segmental or prosodic regularities within the speech heard by the infant. Here, we investigate the process by which the initial acoustic processing required by modulation analysis can itself be tuned by exposure to the regularities of speech. Starting from the classic definition of modulation, as applied within channels of the peripheral filter, we formulate a mathematical framework in which the structure of initial spectral filtering is adapted for modulation analysis. Our working hypothesis is that the human ear and brain are adapted to the analysis of modulation, via a data-driven learning process on the scale of development (or possibly evolution). Simulation results are presented and a comparison with filterbanks classically used in signal processing is done.
Jonathan Le Roux, Hirokazu Kameoka, Nobutaka Ono, Shigeki Sagayama, Alain de Cheveigné
ICASSP2
2008 Parameter estimation method of F0 control model for singing voices
abstract
In this paper, we propose a novel representation of F0 contours that provides a computationally efficient algorithm for automatically estimating the parameters of a F0 control model for singing voices. Although the best known F0 control model, based on a second-order system with a piece-wise constant function as its input, can generate F0 contours of natural singing voices, this model has no means of learning the model parameters from observed F0 contours automatically. Therefore, by modeling the piece-wise constant function by Hidden Markov Models (HMM) and approximating the second order differential equation by the difference equation, we estimate model parameters optimally based on iteration of Viterbi training and an LPC-like solver. Our representation is a generative model and can identify both the target musical note sequence and the dynamics of singing behaviors included in the F0 contours. Our experimental results show that the proposed method can separate the dynamics from the target musical note sequence and generate the F0 contours using estimated model parameters.
Yasunori Ohishi, Hirokazu Kameoka, Kunio Kashino, Kazuya Takeda
INTERSPEECH2
2008 Specmurt Analysis of Polyphonic Music Signals
abstract
This paper introduces a new music signal processing method to extract multiple fundamental frequencies, which we call specmurt analysis. In contrast with cepstrum which is the inverse Fourier transform of log-scaled power spectrum with linear frequency, specmurt is defined as the inverse Fourier transform of linear power spectrum with log-scaled frequency. Assuming that all tones in a polyphonic sound have a common harmonic pattern, the sound spectrum can be regarded as a sum of linearly stretched common harmonic structures along frequency. In the log-frequency domain, it is formulated as the convolution of a common harmonic structure and the distribution density of the fundamental frequencies of multiple tones. The fundamental frequency distribution can be found by deconvolving the observed spectrum with the assumed common harmonic structure, where the common harmonic structure is given heuristically or quasi-optimized with an iterative algorithm. The efficiency of specmurt analysis is experimentally demonstrated through generation of a piano-roll-like display from a polyphonic music signal and automatic sound-to-MIDI conversion. Multipitch estimation accuracy is evaluated over several polyphonic music signals and compared with manually annotated MIDI data.
Shoichiro Saito, Hirokazu Kameoka, Keigo Takahashi, Takuya Nishimoto, Shigeki Sagayama
IEEE Trans. Speech Audio Process.2
2007 Probabilistic Approach to Automatic Music Transcription from Audio Signals
abstract
We discuss automatic music transcription from audio input to music score by integrating our probabilistic approaches to multipitch spectral analysis, rhythm recognition and tempo estimation. In spectral analysis, acoustic energies in spectrogram are clustered into acoustic objects (i.e., music notes) with our method called harmonic-temporal-structured clustering (HTC) utilizing EM algorithm over a structured Gaussian mixture with constraints of harmonic structure and temporal smoothness. After onset and offset timings are found from separated energies of music notes through note power envelope modeling to obtain the piano-roll representation, the rhythm and tempo are simultaneously recognized and estimated in terms of maximum posterior probability given a probabilistic note duration models with HMM (hidden Markov model) and probabilistic "rhythm vocabulary." Variable tempo is also modeled by a smooth analytic curve. Rhythm recognition and tempo estimation is alternately performed to iteratively maximize the joint posterior probability. Experimental results are also shown.
Kenichi Miyamoto, Hirokazu Kameoka, Haruto Takeda, Takuya Nishimoto, Shigeki Sagayama
ICASSP (2)2
2007 Harmonic-Temporal Clustering of Speech for Single and Multiple F0 Contour Estimation in Noisy Environments
abstract
We present in this paper a novel F0contour estimation method based on a parametric description of the wavelet power spectrum of speech that accounts for its structure simultaneously in time and frequency directions. We model the speech spectrum as a sequence of spectral clusters governed by a smooth common F0contour expressed as a spline curve. The harmonic and temporal structure of these clusters and their common F0contour are estimated simultaneously. Through experimental comparisons with existing methods, we show that our algorithm is competitive on clean single-speaker speech, and that it outperforms existing methods both in the presence of noise and for the estimation of multiple F0contours of cochannel concurrent speech.
Jonathan Le Roux, Hirokazu Kameoka, Nobutaka Ono, Alain de Cheveigné, Shigeki Sagayama
ICASSP (4)2
2007 Automatic Decision of Piano Fingering Based on a Hidden Markov Models
Yuichiro Yonebayashi, Hirokazu Kameoka, Shigeki Sagayama
IJCAI2
2007 A Multipitch Analyzer Based on Harmonic Temporal Structured Clustering
abstract
This paper proposes a multipitch analyzer called the harmonic temporal structured clustering (HTC) method, that jointly estimates pitch, intensity, onset, duration, etc., of each underlying source in a multipitch audio signal. HTC decomposes the energy patterns diffused in time-frequency space, i.e., the power spectrum time series, into distinct clusters such that each has originated from a single source. The problem is equivalent to approximating the observed power spectrum time series by superimposed HTC source models, whose parameters are associated with the acoustic features that we wish to extract. The update equations of the HTC are explicitly derived by formulating the HTC source model with a Gaussian kernel representation. We verified through experiments the potential of the HTC method
Hirokazu Kameoka, Takuya Nishimoto, Shigeki Sagayama
IEEE Trans. Speech Audio Process.1
2007 Single and Multiple F0 Contour Estimation Through Parametric Spectrogram Modeling of Speech in Noisy Environments
abstract
This paper proposes a novel $F_{0}$ contour estimation algorithm based on a precise parametric description of the voiced parts of speech derived from the power spectrum. The algorithm is able to perform in a wide variety of noisy environments as well as to estimate the $F_{0}$ s of cochannel concurrent speech. The speech spectrum is modeled as a sequence of spectral clusters governed by a common $F_{0}$ contour expressed as a spline curve. These clusters are obtained by an unsupervised 2-D time-frequency clustering of the power density using a new formulation of the EM algorithm, and their common $F_{0}$ contour is estimated at the same time. A smooth $F_{0}$ contour is extracted for the whole utterance, linking together its voiced parts. A noise model is used to cope with nonharmonic background noise, which would otherwise interfere with the clustering of the harmonic portions of speech. We evaluate our algorithm in comparison with existing methods on several tasks, and show 1) that it is competitive on clean single-speaker speech, 2) that it outperforms existing methods in the presence of noise, and 3) that it outperforms existing methods for the estimation of multiple $F_{0}$ contours of cochannel concurrent speech.
Jonathan Le Roux, Hirokazu Kameoka, Nobutaka Ono, Alain de Cheveigné, Shigeki Sagayama
IEEE Trans. Speech Audio Process.2
2006 Speech analyzer using a joint estimation model of spectral envelope and fine structure
abstract
We have been working on a new speech analyzer based on a parametric representation of speech governed by the F0 parameter, towards practical human-machine interfaces. As a precise estimation of the frequency response of the vocal tract from a real speech signal requires the power of each component of the harmonic structure to be accurately estimated, one hopes to have a high-precision estimation of F0. At the same time, under the empirical constraint that speech spectral envelopes are usually smooth in the power domain, half pitch errors can be significantly avoided. Therefore, F0 and the envelope should be estimated jointly rather than separately through an optimal estimation of the spectral envelope and the spectral fine structure. In this article, we introduce a new speech analysis method using a spectral model with a composite function of envelope and fine structure models. Index Terms: parametric speech analyzer, speech synthesis, pitch estimation, spectral envelope estimation.
Hirokazu Kameoka, Jonathan Le Roux, Nobutaka Ono, Shigeki Sagayama
INTERSPEECH1
2005 Audio stream segregation of multi-pitch music signal based on time-space clustering using Gaussian kernel 2-dimensional model
abstract
The paper describes a novel approach for audio stream segregation of a multi-pitch music signal. We propose a parameter-constrained time-frequency spectrum model expressing both a harmonic spectral structure and a temporal curve of the power envelope with Gaussian kernels. MAP estimation of the model parameters using the EM algorithm provides fundamental frequency, onset and offset time, spectral envelope and power envelope of every underlying audio stream. Our proposed method showed high accuracy in a pitch name estimation task of several pieces of real music performance data.
Hirokazu Kameoka, Takuya Nishimoto, Shigeki Sagayama
ICASSP (3)1
2004 Separation of harmonic structures based on tied Gaussian mixture model and information criterion for concurrent sounds
abstract
A method for the separation of harmonic structures of cochannel input concurrent sounds is described. A model for multiple harmonic structures is constructed with a mixture of tied Gaussian mixtures, from which a single harmonic structure is modeled. Our algorithm enables estimation of both the number and the shape of the underlying harmonic structures, based on a maximum likelihood estimation of the model parameters using the EM algorithm and an information criterion. It operates without restriction on the number of mixed sounds and varieties of sound sources, and extracts accurate fundamental frequencies continuously with simple procedures in the spectral domain. Experiments showed high performance of the algorithm for both simultaneous speech and polyphonic music.
Hirokazu Kameoka, Takuya Nishimoto, Shigeki Sagayama
ICASSP (4)1
2004 Multi-pitch trajectory estimation of concurrent speech based on harmonic GMM and nonlinear kalman filtering
abstract
Abstract This paper describes a multi-pitch tracking algorithmof 1-channel simultaneous multiple speech. The algo-rithm selectively carries out the two alternative processesat each frame: frame-independent-process and frame-dependent-process. The former is the one we have previ-ously proposed[6], that gives good estimates of the num-ber of speakers and F 0 s with a single-frame-processing.The latter corresponds to the topic mainly described inthis paper, that recursively tracks F 0 s using nonlinearKalman filtering. We tested our algorithm on simulta-neous speech signal data and showed higher performancethan when the frame-independent-process was only used. 1. Introduction 1-channel multi-pitch estimation technique may con-tribute to various applications, such as spontaneous dia-logue speech recognition, that allows competitive speech,noise robust speech recognition, especially where noiseis a harmonic signal(e.g., telephone ring, back groundmusic, etc.), and also many music applications. How-ever, multi-pitch estimation of non-stationary signalsis hardly simple due to the complex factors such asspectral overlap, poor frequency resolution and spec-tral widening in short-time analysis, etc. Various ap-proaches concerning to this problem have convention-ally been attempted [1, 2, 3], while two important taskshave been left unsolved. Firstly, there has been no ro-bust way of estimating the number of speakers, andmost of the methods were obliged to assume for sim-plicity that the number is known a priori. Secondly, thedouble/half(harmonics/subharmonics) pitch error has stillbeen one of the most critical problem where convinc-ing solutions are not yet proposed. One may say bothproblems share the same difficulty of defining physicallyor mathematically proper criteria. Until now, we haveproposed a GMM(Gaussian mixture model)-based multi-pitch estimation algorithm that works as a single-frame-processing and gives solutions to the two tasks statedabove according to the information criterion[6]. This al-gorithm has not yet taken into account any time depen-dency property, that would ensure the improvements in
Takuya Nishimoto, Shigeki Sagayama, Hirokazu Kameoka
INTERSPEECH3