VLDB 2026 Research / reviewers in the wild / expert
Takuma Okamoto
dblp:132/9091
· DBLP profile ↗
33ranked-venue papers
23as first author
20since 2021 · last 2025
0000-0001-9913-4647ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 23 first-author · 18 since 2021Artificial intelligence and machine learning · 16 · 10 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Layer-wise Analysis for Quality of Multilingual Synthesized SpeechabstractWhile supervised quality predictors for synthesized speech have demonstrated strong correlations with human ratings, their requirement for in-domain labeled training data hinders their generalization ability to new domains. Unsupervised approaches based on pretrained self-supervised learning (SSL) based models and automatic speech recognition (ASR) models are a promising alternative; however, little is known about how these models encode information about speech quality. Towards the goal of better understanding how different aspects of speech quality are encoded in a multilingual setting, we present a layer-wise analysis of multilingual pretrained speech models based on reference modeling. We find that features extracted from early SSL layers show correlations with human ratings of synthesized speech, and later layers of ASR models can predict quality of non-neural systems as well as intelligibility. We also demonstrate the importance of using well-matched reference data. Erica Cooper, Takuma Okamoto, Yamato Ohtani, Tomoki Toda, Hisashi Kawai |
ASRU | 2 |
| 2025 | Voice Factor Control Using FIR-Based Fast Neural Vocoder for Speech Generation ApplicationsabstractWe have proposed a fast neural vocoder based on the source-filter model introducing finite impulse response (FIR) filters called FIRNet. FIRNet is highly compatible with digital signal processing (DSP) and can, therefore, generate waveforms from vocoder parameters and modified voice factors, such as tone, intonation, and timbre, using DSP. Although modern neural waveform generation systems, such as voice conversion and text-to-speech (TTS), have been able to generate human-like synthetic speech and imitate the reference speaker’s timbre, it is challenging for these systems to manually control arbitrary voice factors, unlike traditional TTS systems. By applying FIRNet to modern neural waveform generation systems, they can achieve arbitrary voice factor controllability. We will demonstrate two applications using FIRNet with DSP-based voice factor controls: one is analysis-synthesis, and the other is text-to-speech. Yamato Ohtani, Takuma Okamoto, Tomoki Toda, Hisashi Kawai |
ASRU | 2 |
| 2025 | Speech Masking System Based on Spatially Separated Multiple TTS Maskers With A Compact Circular Loudspeaker ArrayabstractSpeech masking system with spatially separated maskers is proposed by utilizing multiple sound spot synthesis and multi-speaker neural text-to-speech (TTS) technologies. Although previous systems introduce time-reversed signals of target speech for maskers, these meaningless sounds are unpleasant to listeners. Conversely, the proposed method introduces maskers with the same voice quality as a target speech generated by multi-speaker neural TTS model with global style tokens. Additionally, these TTS-based maskers are synthesized by multiple sound spot synthesis in multiple directions. Then, the target speech can only be heard at the target direction while multiple TTS-based meaningful maskers can also be easily heard at the other directions without discomfort. We implement a speech masking demo system with a compact circular array of 16 loudspeakers carried out with a backpack. In the demonstration, the proposed speech masking system with spatially separated multiple TTS-based maskers is demonstrated. Takuma Okamoto |
ASRU | 1 |
| 2025 | Mora-Level Prosody Prediction for Text-to-Speech Using Japanese BERT Without Accentual LabelsabstractIn practical text-to-speech (TTS) for pitch accent languages, such as Japanese, high-fidelity synthesis with correct prosody requires not only a phoneme sequence but also accentual information. Although accentual information can be obtained from accent dictionaries, words not included in the dictionaries and accent sandhi are sometimes synthesized with incorrect prosody, and manual registration of huge amounts of accent data is costly. Additionally, previous machine learning-based data-driven accent information estimation approaches for TTS also require huge quantities of handcrafted accentual labels during training. This paper proposes a data-driven prosody prediction method for Japanese TTS that uses Japanese BERT and does not require any accentual labels during training. A Japanese TTS acoustic model with mora-level (katakana sequence) input is first trained and mora-level fundamental frequency values (fo), which directly correspond to the prosody, are extracted for the training data using forced alignment. Then, a pre-trained Japanese BERT is finetuned for the mora-level foprediction task with word sequences including kanji and the corresponding katakana sequences as input and the mora-level foextracted using forced alignment as the prediction target. During TTS inference, the mora-level fosequence predicted by the finetuned Japanese BERT is input to the TTS acoustic model along with the katakana input, and correct prosodic synthesis can be realized thanks to this predicted fosequence. Experimental results demonstrate that the proposed method can realize the same synthesis quality and higher accent correctness compared with conventional neural TTS models with accentual labels. Tadashi Ogura, Takuma Okamoto, Yamato Ohtani, Erica Cooper, Tomoki Toda, Hisashi Kawai |
ICASSP | 2 |
| 2025 | GST-BERT-TTS: Prosody Prediction Without Accentual Labels For Multi-Speaker TTS Using BERT With Global Style Tokens
Tadashi Ogura, Takuma Okamoto, Yamato Ohtani, Erica Cooper, Tomoki Toda, Hisashi Kawai |
INTERSPEECH | 2 |
| 2025 | Simultaneous Speech Translation Integrated Compact Multiple Sound Spot Synthesis System On A Laptop Carried Out With A Backpack
Takuma Okamoto, Michiyo Kono |
INTERSPEECH | 1 |
| 2024 | FIRNet: Fundamental Frequency Controllable Fast Neural Vocoder With Trainable Finite Impulse Response FilterabstractSome neural vocoders with fundamental frequency (f0) control have succeeded in performing real-time inference on a single CPU while preserving the quality of the synthetic speech. However, compared with legacy vocoders based on signal processing, their inference speeds are still low. This paper proposes a neural vocoder based on the source-filter model with trainable time-variant finite impulse response (FIR) filters, to achieve a similar inference speed to legacy vocoders. In the proposed model, FIRNet, multiple FIR coefficients are predicted using the neural networks, and the speech waveform is then generated by convolving a mixed excitation signal with these FIR coefficients. Experimental results show that FIRNet can achieve an inference speed similar to legacy vocoders while maintaining f0controllability and natural speech quality. Yamato Ohtani, Takuma Okamoto, Tomoki Toda, Hisashi Kawai |
ICASSP | 2 |
| 2024 | Convnext-TTS And Convnext-VC: Convnext-Based Fast End-To-End Sequence-To-Sequence Text-To-Speech And Voice ConversionabstractEnd-to-end (E2E) sequence-to-sequence (S2S) neural text-to-speech (TTS) models and E2E-S2S neural voice conversion (VC) models can achieve high-quality speech synthesis with a single neural network. To further improve the synthesis quality of E2E-S2S TTS and VC models and increase their inference speed, we propose a Transformer-free ConvNeXt-based encoder and decoder. Additionally, to further increase the inference speed, we propose ConvNeXt-TTS and ConvNeXt-VC, which include the WaveNeXt neural vocoder. This is also constructed from ConvNeXt blocks and can achieve much faster synthesis than HiFi-GAN. The results of experiments using the Hi-Fi-CAPTAIN corpus for the E2E-S2S-TTS and E2E-S2S-VC conditions demonstrate that the proposed ConvNeXt-based encoder and decoder can perform inference three times faster than a Transformer-based encoder and decoder while improving the synthesis quality. In particular, ConvNeXt-TTS and ConvNeXt-VC can achieve very fast E2E-S2S-TTS and E2E-S2S-VC with a real-time factor of 0.05 using a single-core CPU. Takuma Okamoto, Yamato Ohtani, Tomoki Toda, Hisashi Kawai |
ICASSP | 1 |
| 2024 | Mobile PresenTra: NICT fast neural text-to-speech system on smartphones with incremental inference of MS-FC-HiFi-GAN for law-latency synthesis
Takuma Okamoto, Yamato Ohtani, Hisashi Kawai |
INTERSPEECH | 1 |
| 2024 | Challenge of Singing Voice Synthesis Using Only Text-To-Speech Corpus With FIRNet Source-Filter Neural Vocoder
Takuma Okamoto, Yamato Ohtani, Sota Shimizu, Tomoki Toda, Hisashi Kawai |
INTERSPEECH | 1 |
| 2023 | WaveNeXt: ConvNeXt-Based Fast Neural Vocoder Without ISTFT layerabstractA recently proposed neural vocoder, Vocos, can perform inference ten times faster than HiFi-GAN because of its use of ConvNeXt layers that can predict high-resolution short-time Fourier transform (STFT) spectra and an inverse STFT layer. To improve synthesis quality while preserving inference speed, this paper proposes an alternative ConvNeXt-based fast neural vocoder, WaveNeXt, in which the inverse STFT layer in Vocos is replaced with a trainable linear layer that can directly predict speech waveform samples without STFT spectra. Additionally, by integrating the JETS-based end-to-end text-to-speech (E2E TTS) framework, E2E TTS models can also be constructed with Vocos and WaveNeXt. Furthermore, full-band models with a sampling frequency of 48 kHz were investigated. The results of experiments for both the analysis-synthesis and E2E TTS conditions demonstrate that the proposed WaveNeXt can achieve higher quality synthesis than Vocos while preserving its inference speed. Takuma Okamoto, Haruki Yamashita, Yamato Ohtani, Tomoki Toda, Hisashi Kawai |
ASRU | 1 |
| 2023 | Continuous Action Space-Based Spoken Language Acquisition Agent Using Residual Sentence Embedding and Transformer DecoderabstractStudies on spoken language acquisition agents aim to understand the mechanism of human language learning and to realize it on computers. Existing open vocabulary agents first perform unsupervised word learning from speech signals to construct a word dictionary as a discrete action space and then conduct reinforcement learning to understand the use of the words in the dictionary through interaction with dialogue partners. A limitation is that they have difficulty pronouncing multi-word utterances. This study proposes an agent that generates multi-word waveform utterances using a continuous action space. The conventional agent uses a vision-focusing mechanism to accelerate dialogue-based learning by guiding the agent’s attention to those concepts in its eyesight. In contrast, the proposed agent replaces it with residual sentence embedding combined with vision features used as the action space. The agent consists of speech and image input front-ends, a transformer language model of pseudo-action space. Experimental results show that the agent learns multi-word utterances assisted by unsupervised learning algorithms using unlabeled speech and image data sets. Ryota Komatsu, Yusuke Kimura, Takuma Okamoto, Takahiro Shinozaki |
ICASSP | 3 |
| 2023 | E2E-S2S-VC: End-To-End Sequence-To-Sequence Voice Conversion
Takuma Okamoto, Tomoki Toda, Hisashi Kawai |
INTERSPEECH | 1 |
| 2023 | Harmonic-Net: Fundamental Frequency and Speech Rate Controllable Fast Neural VocoderabstractThere is a need to improve the synthesis quality of HiFi-GAN-based real-time neural speech waveform generative models on CPUs while preserving the controllability of fundamental frequency ($f_{\mathrm{o}}$) and speech rate (SR). For this purpose, we propose Harmonic-Net and Harmonic-Net+, which introduce two extended functions into the HiFi-GAN generator. The first extension is a downsampling network, named the excitation signal network, that hierarchically receives multi-channel excitation signals corresponding to$f_{\mathrm{o}}$. The second extension is the layerwise pitch-dependent dilated convolutional network (LW-PDCNN), which can flexibly change its receptive fields depending on the input$f_{\mathrm{o}}$to handle large fluctuations in$f_{\mathrm{o}}$for the upsampling-based HiFi-GAN generator. The proposed explicit input of excitation signals and LW-PDCNNs corresponding to$f_{\mathrm{o}}$are expected to realize high-quality synthesis for the normal and$f_{\mathrm{o}}$-conversion conditions and for the SR-conversion condition. The results of experiments for unseen speaker synthesis, full-band singing voice synthesis, and text-to-speech synthesis show that the proposed method with harmonic waves corresponding to$f_{\mathrm{o}}$can achieve higher synthesis quality than conventional methods in all (i.e., normal,$f_{\mathrm{o}}$-conversion, and SR-conversion) conditions. Keisuke Matsubara, Takuma Okamoto, Ryoichi Takashima, Tetsuya Takiguchi, Tomoki Toda, Hisashi Kawai |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Neural speech-rate conversion with multispeaker WaveNet vocoderabstractSpeech-rate conversion technology, which can expand or compress speech waveforms while preserving the pitch of the sound, is traditionally realized by signal-processing-based approaches. To improve the synthesis quality, this paper proposes a machine-learning-based approach using neural vocoders, to perform neural speech-rate conversion. The proposed approach introduces a multispeaker WaveNet vocoder trained with a multispeaker corpus. Speech-rate conversion for many and unspecified speakers, not included in the training data, is realized by resampling acoustic features or hidden features along the time direction in inference. In experiments, the multispeaker WaveNet vocoder was trained using the JVS corpus and two types of resampling methods were compared. Conventional WSOLA and STRAIGHT were also compared as signal-processing-based baselines. The test sets included Japanese speaker corpora for the monolingual condition, and an English multispeaker corpus (CMU ARCTIC) for the cross-lingual condition. The results of the experiments demonstrate that the proposed approach with resampling of hidden features can achieve higher quality speech-rate conversion than the conventional methods, in both monolingual and cross-lingual conditions, except for speakers with low fundamental frequency in conversion of fast speech. Takuma Okamoto, Keisuke Matsubara, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
Speech Commun. | 1 |
| 2021 | Multi-Stream HiFi-GAN with Data-Driven Waveform DecompositionabstractAlthough a HiFi-GAN vocoder can synthesize high-fidelity speech waveforms in real time on CPUs, there is a tradeoff between synthesis quality and inference speed. To increase inference speed while maintaining synthesis quality, a multi-band structure is introduced to HiFi-GAN. However, it cannot be trained well because of the strong constraint imposed by the fixed multi-band structure. As an alternative approach, Multi-stream MelGAN and HiFi-GAN are proposed, in which the fixed synthesis filter in Multi-band MelGAN is replaced by a trainable convolutional layer with the same structure. In contrast to Multi-band MelGAN, the proposed methods use the trainable synthesis filter to decompose speech waveforms in a data-driven manner. To evaluate the proposed Multi-stream HiFi-GAN as an entire real-time neural text-to-speech system on CPUs, a fast acoustic model, based on Parallel Tacotron 2 with forced alignment and accentual label input, was implemented. The results of experiments-using Japanese male, female, and multi-speaker corpora-indicate that Multi-stream HiFi-GAN can increase synthesis speed while improving or maintaining synthesis quality in analysis-synthesis and text-to-speech conditions for single-speaker models and unseen speaker synthesis for multi-speaker models, compared with the original HiFi-GAN. Takuma Okamoto, Tomoki Toda, Hisashi Kawai |
ASRU | 1 |
| 2021 | High-Intelligibility Speech Synthesis for Dysarthric Speakers with LPCNet-Based TTS and CycleVAE-Based VCabstractThis paper presents a high-intelligibility speech synthesis method for persons with dysarthria caused by athetoid cerebral palsy. The muscular control of such speakers is unstable because of their athetoid symptoms, and their pronunciation is unclear, which makes it difficult for them to communicate. In this paper, we present a method for generating highly intelligible speech that preserves the individuality of dysarthric speakers by combining Transformer-TTS, CycleVAE-VC, and a LPCNet vocoder. Rather than repairing prosody from the dysarthric speech, this method transfers the dysarthric speaker’s individuality to the speech of a healthy person generated by TTS synthesis. This task is both important and challenging. From the results of our evaluation experiments, we confirmed that the proposed method can partially transfer the individuality of the target dysarthric speaker while maintaining the intelligibility of the source speech. Keisuke Matsubara, Takuma Okamoto, Ryoichi Takashima, Tetsuya Takiguchi, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
ICASSP | 2 |
| 2021 | Close-Talking Recording with Planarly Distributed MicrophonesabstractThis paper provides a close-talking recording method with microphones in arbitrary planar distributions based on sound pressure interpolation. In this method, sound pressures recorded by microphones located near the center of multiple co-centered circular microphone arrays are interpolated using sound pressures recorded by these arrays via multiple 2D cylindrical harmonic analyses. If the sound sources are outside the spherical boundary formed by the maximum radius of the arrays, the interpolation is completed. However, if the sound sources are inside the spherical boundary, the interpolation is not completed in principle. Using this principle, close-talking recording is realized as a residual response between recorded and interpolated sound pressures. The recording radial sensitivity can be controlled by varying the maximum array radius. By introducing a kernel ridge regression-based sound field interpolation method, the approach can be extended to the microphones in arbitrary planar distributions. Simulations confirm the effectiveness of these close- talking recording methods. Takuma Okamoto |
ICASSP | 1 |
| 2021 | Noise Level Limited Sub-Modeling for Diffusion Probabilistic VocodersabstractAlthough diffusion probabilistic vocoders WaveGrad and DiffWave can realize real-time high-fidelity speech synthesis with a simple loss function in training, all noise components with over the full range of noise levels are predicted by one model in all iterations. This paper proposes a simple but effective noise level-limited sub-modeling framework for diffusion probabilistic vocoders Sub-WaveGrad and Sub-DiffWave. In the proposed method, DiffWave conditioned on a continuous noise level like WaveGrad, and spectral enhancement post-filtering are also provided. The proposed Sub-WaveGrad and Sub-DiffWave models are realized using 10 sub-models. These models are separately trained with different noise level limits, and only necessary sub-models are used according to the noise schedule during inference. The results of experiments using a Japanese female speech corpus indicate that both the proposed Sub-WaveGrad and Sub-DiffWave outperform vanilla WaveGrad and DiffWave in terms of the model accuracy and synthesis quality while retaining the inference speed. Takuma Okamoto, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
ICASSP | 1 |
| 2021 | Quasi-Periodic Parallel WaveGAN: A Non-Autoregressive Raw Waveform Generative Model With Pitch-Dependent Dilated Convolution Neural NetworkabstractIn this paper, we propose a quasi-periodic parallel WaveGAN (QPPWG) waveform generative model, which applies a quasi-periodic (QP) structure to a parallel WaveGAN (PWG) model using pitch-dependent dilated convolution networks (PDCNNs). PWG is a small-footprint GAN-based raw waveform generative model, whose generation time is much faster than real time because of its compact model and non-autoregressive (non-AR) and non-causal mechanisms. Although PWG achieves high-fidelity speech generation, the generic and simple network architecture lacks pitch controllability for an unseen auxiliary fundamental frequency (F0) feature such as a scaled F0. To improve the pitch controllability and speech modeling capability, we apply a QP structure with PDCNNs to PWG, which introduces pitch information to the network by dynamically changing the network architecture corresponding to the auxiliary F0feature. Both objective and subjective experimental results show that QPPWG outperforms PWG when the auxiliary F0feature is scaled. Moreover, analyses of the intermediate outputs of QPPWG also show better tractability and interpretability of QPPWG, which respectively models spectral and excitation-like signals using the cascaded fixed and adaptive blocks of the QP structure. Yi-Chiao Wu, Tomoki Hayashi, Takuma Okamoto, Hisashi Kawai, Tomoki Toda |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Transformer-Based Text-to-Speech with Weighted Forced AttentionabstractThis paper investigates state-of-the-art Transformer- and FastSpeech-based high-fidelity neural text-to-speech (TTS) with full-context label input for pitch accent languages. The aim is to realize faster training than conventional Tacotron-based models. Introducing phoneme durations into Tacotron-based TTS models improves both synthesis quality and stability. Therefore, a Transformer-based acoustic model with weighted forced attention obtained from phoneme durations is proposed to improve synthesis accuracy and stability, where both encoder-decoder attention and forced attention are used with a weighting factor. Furthermore, FastSpeech without a duration predictor, in which the phoneme durations are predicted by another conventional model, is also investigated. The results of experiments using a Japanese female corpus and the WaveGlow vocoder indicate that the proposed Transformer using forced attention with a weighting factor of 0.5 outperforms other models, and removing the duration predictor from FastSpeech improves synthesis quality, although the proposed weighted forced attention does not improve synthesis stability. Takuma Okamoto, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
ICASSP | 1 |
| 2020 | Quasi-Periodic Parallel WaveGAN Vocoder: A Non-Autoregressive Pitch-Dependent Dilated Convolution Model for Parametric Speech GenerationabstractIn this paper, we propose a parallel WaveGAN (PWG)-like neural vocoder with a quasi-periodic (QP) architecture to improve the pitch controllability of PWG. PWG is a compact non-autoregressive (non-AR) speech generation model, whose generative speed is much faster than real time. While utilizing PWG as a vocoder to generate speech on the basis of acoustic features such as spectral and prosodic features, PWG generates high-fidelity speech. However, when the input acoustic features include unseen pitches, the pitch accuracy of PWG-generated speech degrades because of the fixed and generic network of PWG without prior knowledge of speech periodicity. The proposed QPPWG adopts a pitch-dependent dilated convolution network (PDCNN) module, which introduces the pitch information into PWG via the dynamically changed network architecture, to improve the pitch controllability and speech modeling capability of vanilla PWG. Both objective and subjective evaluation results show the higher pitch accuracy and comparable speech quality of QPPWG-generated speech when the QPPWG model size is only 70 % of that of vanilla PWG. Yi-Chiao Wu, Tomoki Hayashi, Takuma Okamoto, Hisashi Kawai, Tomoki Toda |
INTERSPEECH | 3 |
| 2019 | Tacotron-Based Acoustic Model Using Phoneme Alignment for Practical Neural Text-to-Speech SystemsabstractAlthough sequence-to-sequence (seq2seq) models with attention mechanism in neural text-to-speech (TTS) systems, such as Tacotron 2, can jointly optimize duration and acoustic models, and realize high-fidelity synthesis compared with conventional duration-acoustic pipeline models, these involve a risk that speech samples cannot be sometimes successfully synthesized due to the attention prediction errors. Therefore, these seq2seq models cannot be directly introduced in practical TTS systems. On the other hand, the conventional pipeline models are broadly used in practical TTS systems since there are few crucial prediction errors in the duration model. For realizing high-quality practical TTS systems without attention prediction errors, this paper investigates Tacotron-based acoustic models with phoneme alignment instead of attention. The phoneme durations are first obtained from HMM-based forced alignment and the duration model is a simple bidirectional LSTM-based network. Then, a seq2seq model with forced alignment instead of attention is investigated and an alternative model with Tacotron decoder and phoneme duration is proposed. The results of experiments with full-context label input using WaveGlow vocoder indicate that the proposed model can realize a high-fidelity TTS system for Japanese with a real-time factor of 0.13 using a GPU without attention prediction errors compared with the seq2seq models. Takuma Okamoto, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
ASRU | 1 |
| 2019 | Horizontal 3D Sound Field Recording and 2.5D Synthesis with Omni-directional Circular ArraysabstractAlthough 2.5D sound field synthesis with a circular loudspeaker array can be used in a 3D sound field, a 2D sound field, instead of a 3D sound field, is assumed for a sound field recording with a circular microphone array. This paper presents a horizontal 3D sound field recording and 2.5D synthesis method used in 3D sound fields with multiple co-centered omni-directional circular microphone arrays and a circular loudspeaker array without vertical derivative measurements. The spherical harmonic spectrums used for 2.5D synthesis are extracted from the recorded horizontal sound pressures and the driving function of a circular sound source is analytically derived based on 2.5D higher-order Ambisonics. The results of computer simulations indicate the effectiveness of the proposed method, in comparison with the conventional least squares-based pressure matching, and 2D cylindrical harmonic analysis-based approaches. Takuma Okamoto |
ICASSP | 1 |
| 2019 | Investigations of Real-time Gaussian Fftnet and Parallel Wavenet Neural Vocoders with Simple Acoustic FeaturesabstractThis paper examines four approaches to improving real-time neural vocoders with simple acoustic features (SAF) constructed from fundamental frequency and mel-cepstra rather than mel-spectrograms. The investigations are as follows: 1) the effectiveness of single Gaussian (SG) autoregressive (AR) WaveNet and FFTNet vocoders with SAF, 2) the possibility of SG parallel WaveNet vocoder training and synthesis with SAF, 3) the impact of noise shaping on SG AR neural vocoders, and 4) the efficacy of bandwidth extension to synthesize speech waveforms at a sampling frequency of 24 kHz by SG AR neural vocoders from SAF for that of 16 kHz. The results of experiments indicate that SG AR WaveNet and real-time SG AR FFTNet vocoders with noise shaping using SAF can realize sufficient synthesis quality with bandwidth extension effect. Moreover, a real-time SG parallel WaveNet vocoder can also be trained using SAF. Takuma Okamoto, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
ICASSP | 1 |
| 2019 | Real-Time Neural Text-to-Speech with Sequence-to-Sequence Acoustic Model and WaveGlow or Single Gaussian WaveRNN Vocoders
Takuma Okamoto, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
INTERSPEECH | 1 |
| 2018 | An Investigation of Subband Wavenet Vocoder Covering Entire Audible Frequency Range with Limited Acoustic FeaturesabstractAlthough a WaveNet vocoder can synthesize more natural-sounding speech waveforms than conventional vocoders with sampling frequencies of 16 and 24 kHz, it is difficult to directly extend the sampling frequency to 48 kHz to cover the entire human audible frequency range for higher-quality synthesis because the model size becomes too large to train with a consumer GPU. For a WaveNet vocoder with a sampling frequency of 48 kHz with a consumer GPU, this paper introduces a subband WaveNet architecture to a speaker-dependent WaveNet vocoder and proposes a subband WaveNet vocoder. In experiments, each conditional subband WaveNet with a sampling frequency of 8 kHz was well trained using a consumer GPU. The results of subjective evaluations with a Japanese male speech corpus indicate that the proposed subband WaveNet vocoder with 36-dimensional simple acoustic features significantly outperformed the conventional source-filter model-based vocoders including STRAIGHT with 86-dimensional features. Takuma Okamoto, Kentaro Tachibana, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
ICASSP | 1 |
| 2018 | Improving FFTNet Vocoder with Noise Shaping and Subband ApproachesabstractAlthough FFTNet neural vocoders can synthesize speech waveforms in real time, the synthesized speech quality is worse than that of WaveNet vocoders. To improve the synthesized speech quality of FFTNet while ensuring real-time synthesis, residual connections are introduced to enhance the prediction accuracy. Additionally, time-invariant noise shaping and subband approaches, which significantly improve the synthesized speech quality of WaveNet vocoders, are applied. A subband FFTNet vocoder with multiband input is also proposed to directly compensate the phase shift between subbands. The proposed approaches are evaluated through experiments using a Japanese male corpus with a sampling frequency of 16 kHz. The results are compared with those synthesized by the STRAIGHT vocoder without mel-cepstral compression and those from conventional FFTNet and WaveNet vocoders. The proposed approaches are shown to successfully improve the synthesized speech quality of the FFTNet vocoder. In particular, the use of noise shaping enables FFTNet to significantly outperform the STRAIGHT vocoder. Takuma Okamoto, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
SLT | 1 |
| 2017 | Subband wavenet with overlapped single-sideband filterbanksabstractCompared with conventional vocoders, deep neural network-based raw audio generative models, such as WaveNet and SampleRNN, can more naturally synthesize speech signals, although the synthesis speed is a problem, especially with high sampling frequency. This paper provides subband WaveNet based on multirate signal processing for high-speed and high-quality synthesis with raw audio generative models. In the training stage, speech waveforms are decomposed and decimated into subband short waveforms with a low sampling rate, and each subband WaveNet network is trained using each subband stream. In the synthesis stage, each generated signal is up-sampled and integrated into a fullband speech signal. The results of objective and subjective experiments for unconditional WaveNet with a sampling frequency of 32 kHz indicate that the proposed subband WaveNet with a square-root Hann window-based overlapped 9-channel single-sideband filterbank can realize about four times the synthesis speed and improve the synthesized speech quality more than the conventional fullband WaveNet. Takuma Okamoto, Kentaro Tachibana, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
ASRU | 1 |
| 2017 | Analytical approach to 2.5D sound field control using a circular double-layer array of fixed-directivity loudspeakersabstractThis paper provides a regularization-free analytical approach to 2.5D interior and exterior sound field control using a circular double-layer array of fixed-directivity loudspeakers not only to provide a desired sound field inside the array but also to reduce the sound energy outside of it in the horizontal plane. The proposed method analytically derives the driving functions of the inner and outer circular arrays based on 2.5D spherical harmonic expansion to simultaneously control both sound fields inside and outside the array by a mode-matching framework. The results of computer simulations show that the proposed method is certainly effective compared with the conventional least squares approach that requires regularization schemes in terms of synthesis accuracy inside the listening zone over a wideband frequency range and the acoustic contrast between the quiet zone and the synthesis center at low frequencies. Takuma Okamoto |
ICASSP | 1 |
| 2016 | 2.5D higher order ambisonics for a sound field described by angular spectrum coefficientsabstractThis paper derives an analytical solution to convert sound field representation from the angular spectrum to the circular harmonics expansion. A sound field is decomposed to plane waves by the spatial Fourier transform and represented by the angular spectrum. A plane wave is also described by the circular harmonics expansion. In the proposed formulation, these two representations are integrated and a sound field can be represented by the circular harmonics expansion with the angular spectrum coefficients. For actual implementations, the driving function of a circular sound source for 2.5D higher order Ambisonics is analytically derived from a sound field described by the angular spectrum. The results of the computer simulations show that the proposed method with a circular loudspeaker array can reproduce a sound field recorded by a linear microphone array with appropriate accuracy around the center of the circular array. Takuma Okamoto |
ICASSP | 1 |
| 2015 | Near-field sound propagation based on a circular and linear array combinationabstractThis paper proposes a new method for realizing 3D near-field sound propagation. Its main concept is that the sound pressure radiated from a circular sound source outside the circle is completely canceled out using another linear sound source. The appropriate driving functions and the reproduced sound pressures of both sound sources are analytically derived based on the two-dimensional spatial Fourier transform. The radiation property reproduced by a circular source outside the circle is different from that inside the circle. As a result, three-dimensional near-field sound propagation can be realized within the radius of a circular source because the total radiated sound pressure outside the circle can only be completely canceled out. Compared with previous methods, the propagation distance can be controlled by changing the circle's radius. The results of computer simulations suggest that the proposed method can realize effective three-dimensional near-field sound propagation. Takuma Okamoto |
ICASSP | 1 |
| 2014 | Generation of multiple sound zones by spatial filtering in wavenumber domain using a linear array of loudspeakersabstractNovel signal processing is proposed for generating acoustically bright and dark zones at arbitrary horizontal positions using a linear array of loudspeakers. Most conventional methods are based on the numerical calculation of the inverse of the spatial correlation matrix between control points and the positions of the loudspeakers. However, such methods are quite unstable because the acoustic inverse problem is very ill-conditioned. On the other hand, in the proposed method, spatial filters in the wavenumber domain are analytically derived by introducing the spectral division method and modeling sound pressures at the control line as a rectangular window corresponding to bright and dark zones. Computer simulation results show that the proposed method can generate acoustically bright and dark zones more accurately than the conventional acoustic energy difference maximization method. Takuma Okamoto |
ICASSP | 1 |