Hiroki Kanagawa

dblp:136/5329 · DBLP profile ↗
← Back
11ranked-venue papers
9as first author
8since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 9 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 5 first-author · 4 since 2021
YearPublicationVenuePosition
2024 Noise-Robust Zero-Shot Text-to-Speech Synthesis Conditioned on Self-Supervised Speech-Representation Model with Adapters
abstract
The zero-shot text-to-speech (TTS) method, based on speaker embeddings extracted from reference speech using self-supervised learning (SSL) speech representations, can reproduce speaker characteristics very accurately. However, this approach suffers from degradation in speech synthesis quality when the reference speech contains noise. In this paper, we propose a noise-robust zero-shot TTS method. We incorporated adapters into the SSL model, which we fine-tuned with the TTS model using noisy reference speech. In addition, to further improve performance, we adopted a speech enhancement (SE) front-end. With these improvements, our proposed SSL-based zero-shot TTS achieved high-quality speech synthesis with noisy reference speech. Through the objective and subjective evaluations, we confirmed that the proposed method is highly robust to noise in reference speech, and effectively works in combination with SE.
Kenichi Fujita, Hiroshi Sato 0002, Takanori Ashihara, Hiroki Kanagawa, Marc Delcroix, Takafumi Moriya, Yusuke Ijima
ICASSP4
2024 Knowledge Distillation from Self-Supervised Representation Learning Model with Discrete Speech Units for Any-to-Any Streaming Voice Conversion
Hiroki Kanagawa, Yusuke Ijima
INTERSPEECH1
2024 Pre-training Neural Transducer-based Streaming Voice Conversion for Faster Convergence and Alignment-free Training
Hiroki Kanagawa, Takafumi Moriya, Yusuke Ijima
INTERSPEECH1
2023 Enhancement of Text-Predicting Style Token With Generative Adversarial Network for Expressive Speech Synthesis
abstract
This work proposes an advanced text-predicting style embedding for expressive speech synthesis. Text-predicting global style token (TPGST) predicts style embedding from text instead of reference speech and uses it to condition a text-to-speech synthesis (TTS) model, resulting in style TTS without reference speech. Although this minimizes the style embedding’s L1-loss between that extracted from reference speech and that predicted during training, predicted embedding tends to be over-smoothed. To overcome this issue, the proposed method uses the generative adversarial network (GAN) in training style predictors. This not only improves style reproduction, but also aims to reduce style conditioning mismatch during TTS model training. We also utilize TTS text embeddings as in other related work, as well as word information via BERT in order to find better style distributions in GAN. An evaluation of subjective style reproduction demonstrates that 1) the proposed method outperforms conventional TPGST, and 2) the use of words yielded by BERT provides even better performance. Our style predictor is also effective in attaining unseen style TTS for "seen" and "unseen" speakers.
Hiroki Kanagawa, Yusuke Ijima
ICASSP1
2023 VC-T: Streaming Voice Conversion Based on Neural Transducer
Hiroki Kanagawa, Takafumi Moriya, Yusuke Ijima
INTERSPEECH1
2022 Multi-Sample Subband Wavernn Via Multivariate Gaussian
abstract
This paper proposes a high-speed neural vocoder for CPU implementation. Two approaches for speeding up autoregressive neural vocoders have been proposed, 1) simultaneous multiple sample generation and 2) subband signal-based vocoder; so far they have been employed independently. Our neural vocoder is extremely fast as it generates multiple samples of subband signals simultaneously. Although there is an association between each subband signal, the conventional subband-based vocoder can degrade quality because each subband signal is generated from an independent probability distribution. To overcome this problem, we also introduce waveform generation that takes account of the association of each subband by employing multivariate Gaussian. Experiments show that 1) our proposed method is 1.81 times as fast as the conventional subband WaveRNN on a single-threaded CPU; 2) it outperformed the conventional method in a subjective evaluation in terms of naturalness, and achieved a mean opinion score (MOS) of 4.08 on text-to-speech task.
Hiroki Kanagawa, Yusuke Ijima
ICASSP1
2022 Joint Modeling of Multi-Sample and Subband Signals for Fast Neural Vocoding on CPU
Hiroki Kanagawa, Yusuke Ijima, Hiroyuki Toda
INTERSPEECH1
2022 SIMD-Size Aware Weight Regularization for Fast Neural Vocoding on CPU
abstract
This paper proposes weight regularization for a faster neural vocoder. Pruning time-consuming DNN modules is a promising way to realize a real-time vocoder on a CPU (e.g. WaveRNN, LPCNet). Regularization that encourages sparsity is also effective in avoiding the quality degradation created by pruning. However, the orders of weight matrices must be contiguous in SIMD size for fast vocoding. To ensure this order, we propose explicit SIMD size aware regularization. Our proposed method reshapes a weight matrix into a tensor so that the weights are aligned by group size in advance, and then computes the group Lasso-like regularization loss. Experiments on 70% sparse subband WaveRNN show that pruning in conventional Lasso and column-wise group Lasso degrades the synthetic speech's naturalness. The vocoder with proposed regularization 1) achieves comparable naturalness to that without pruning and 2) performs meaningfully faster than other conventional vocoders using regularization.
Hiroki Kanagawa, Yusuke Ijima
SLT1
2020 Lightweight LPCNet-Based Neural Vocoder with Tensor Decomposition
Hiroki Kanagawa, Yusuke Ijima
INTERSPEECH1
2018 Efficient Building Strategy with Knowledge Distillation for Small-Footprint Acoustic Models
abstract
In this paper, we propose a novel training strategy for deep neural network (DNN) based small-footprint acoustic models. The accuracy of DNN-based automatic speech recognition (ASR) systems can be greatly improved by leveraging large amounts of data to improve the level of expression. DNNs use many parameters to enhance recognition performance. Unfortunately, resource-constrained local devices are unable to run complex DNN-based ASR systems. For building compact acoustic models, the knowledge distillation (KD) approach is often used. KD uses a large, well-trained model that outputs target labels to train a compact model. However, the standard KD cannot fully utilize the large model outputs to train compact models because the soft logits provide only rough information. We assume that the large model must give more useful hints to the compact model. We propose an advanced KD that uses mean squared error to minimize the discrepancies between the final hidden layer outputs. We evaluate our proposal on recorded speech data sets assuming car-and home-use scenarios, and show that our models achieve lower character error rates than the conventional KD approach or from-scratch training on computation resource-constrained devices.
Takafumi Moriya, Hiroki Kanagawa, Kiyoaki Matsui, Takaaki Fukutomi, Yusuke Shinohara, Yoshikazu Yamaguchi, Manabu Okamoto, Yushi Aono
SLT2
2013 Speaker-independent style conversion for HMM-based expressive speech synthesis
abstract
This paper proposes a technique for creating target speaker's expressive-style model from the target speaker's neutral style speech in HMM-based speech synthesis. The technique is based on the style adaptation using linear transforms where speaker-independent transformation matrices are estimated in advance using pairs of neutral and target-style speech data of multiple speakers. By applying the obtained transformation matrices to a new speaker's neutral-style model, we can convert the style expressivity of the acoustic model to the target style without preparing any target-style speech of the speaker. In addition, we introduce a speaker adaptive training (SAT) framework into the transform estimation to reduce the acoustic difference among speakers. We subjectively evaluate the performance of the style conversion in terms of the naturalness, speaker similarity, and style reproducibility.
Hiroki Kanagawa, Takashi Nose, Takao Kobayashi
ICASSP1