Yusuke Ijima

dblp:67/8052 · DBLP profile ↗
← Back
46ranked-venue papers
8as first author
24since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 45 · 8 first-author · 24 since 2021Artificial intelligence and machine learning · 28 · 6 first-author · 14 since 2021
YearPublicationVenuePosition
2026 Effect of individual characteristics on impressions of one's own recorded voice
abstract
This study aims to identify individual characteristics such as age, gender, personality traits, and values that influence the perception of one’s own recorded voice. While previous studies have shown that the perception of one’s own recorded voice is different from that of others, and that these differences are influenced by individual characteristics, only a limited number of individual characteristics were examined in past research. In our study, we conducted a large-scale subjective experiment with 141 Japanese participants using multiple individual characteristics. Participants evaluated impressions of their own recorded voices and the voices of others, and we analyzed the relationship between each of the individual characteristics and the voice impressions. Our findings showed that individual characteristics such as the frequency of listening to one’s own recorded voice (which had not been examined in the previous studies) influenced the perception of one’s own recorded voice. We further analyzed the use of combinations of multiple individual characteristics, including those that influenced impressions in a single use, to predict impressions of one’s own recorded voice and found that they were better predicted by the combination of multiple individual characteristics than by the use of a single individual characteristic. • We show that impressions, such as familiarity, of one’s own recorded voice are different from those of others. • We show that multiple individual characteristics, such as the frequency of listening to one’s own recorded voice, influence the impression of one’s own recorded voice. • We show that combinations of multiple individual characteristics predicts the impression of one’s own recorded voice better than the use of a single individual characteristic.
Hikaru Yanagida, Yusuke Ijima, Naohiro Tawara
Speech Commun.2
2025 Voice Impression Control in Zero-Shot TTS
abstract
Para-/non-linguistic information in speech is pivotal in shaping the listeners' impression. Although zero-shot text-to-speech (TTS) has achieved high speaker fidelity, modulating subtle para-/non-linguistic information to control perceived voice characteristics, i.e., impressions, remains challenging. We have therefore developed a voice impression control method in zero-shot TTS that utilizes a low-dimensional vector to represent the intensities of various voice impression pairs (e.g., dark-bright). The results of both objective and subjective evaluations have demonstrated our method's effectiveness in impression control. Furthermore, generating this vector via a large language model enables target-impression generation from a natural language description of the desired impression, thus eliminating the need for manual optimization. Audio examples are available on our demo page (https://ntt-hilab-gensp.github.io/is2025voiceimpression/).
Kenichi Fujita, Shota Horiguchi, Yusuke Ijima
INTERSPEECH3
2024 What Do Self-Supervised Speech and Speaker Models Learn? New Findings from a Cross Model Layer-Wise Analysis
abstract
Self-supervised learning (SSL) has attracted increased attention for learning meaningful speech representations. Speech SSL models, such as WavLM, employ masked prediction training to encode general-purpose representations. In contrast, speaker SSL models, exemplified by DINO-based models, adopt utterance-level training objectives primarily for speaker representation. Understanding how these models represent information is essential for refining model efficiency and effectiveness. Unlike the various analyses of speech SSL, there has been limited investigation into what information speaker SSL captures and how its representation differs from speech SSL or other fully-supervised speaker models. This paper addresses these fundamental questions. We explore the capacity to capture various speech properties by applying SUPERB evaluation probing tasks to speech and speaker SSL models. We also examine which layers are predominantly utilized for each task to identify differences in how speech is represented. Furthermore, we conduct direct comparisons to measure the similarities between layers within and across models. Our analysis unveils that 1) the capacity to represent content information is somewhat unrelated to enhanced speaker representation, 2) specific layers of speech SSL models would be partly specialized in capturing linguistic information, and 3) speaker SSL models tend to disregard linguistic information but exhibit more sophisticated speaker representation.
Takanori Ashihara, Marc Delcroix, Takafumi Moriya, Kohei Matsuura, Taichi Asami, Yusuke Ijima
ICASSP6
2024 Noise-Robust Zero-Shot Text-to-Speech Synthesis Conditioned on Self-Supervised Speech-Representation Model with Adapters
abstract
The zero-shot text-to-speech (TTS) method, based on speaker embeddings extracted from reference speech using self-supervised learning (SSL) speech representations, can reproduce speaker characteristics very accurately. However, this approach suffers from degradation in speech synthesis quality when the reference speech contains noise. In this paper, we propose a noise-robust zero-shot TTS method. We incorporated adapters into the SSL model, which we fine-tuned with the TTS model using noisy reference speech. In addition, to further improve performance, we adopted a speech enhancement (SE) front-end. With these improvements, our proposed SSL-based zero-shot TTS achieved high-quality speech synthesis with noisy reference speech. Through the objective and subjective evaluations, we confirmed that the proposed method is highly robust to noise in reference speech, and effectively works in combination with SE.
Kenichi Fujita, Hiroshi Sato 0002, Takanori Ashihara, Hiroki Kanagawa, Marc Delcroix, Takafumi Moriya, Yusuke Ijima
ICASSP7
2024 STYLECAP: Automatic Speaking-Style Captioning from Speech Based on Speech and Language Self-Supervised Learning Models
abstract
We propose StyleCap, a method to generate natural language descriptions of speaking styles appearing in speech. Although most of conventional techniques for para-/non-linguistic information recognition focus on the category classification or the intensity estimation of pre-defined labels, they cannot provide the reasoning of the recognition result in an interpretable manner. StyleCap is a first step towards an end-to-end method for generating speaking-style prompts from speech, i.e., automatic speaking-style captioning. StyleCap is trained with paired data of speech and natural language descriptions. We train neural networks that convert a speech representation vector into prefix vectors that are fed into a large language model (LLM)-based text decoder. We explore an appropriate text decoder and speech feature representation suitable for this new task. The experimental results demonstrate that our StyleCap leveraging richer LLMs for the text decoder, speech self-supervised learning (SSL) features, and sentence rephrasing augmentation improves the accuracy and diversity of generated speaking-style captions. Samples of speaking-style captions generated by our StyleCap are publicly available1.
Kazuki Yamauchi, Yusuke Ijima, Yuki Saito 0001
ICASSP2
2024 Lightweight Zero-shot Text-to-Speech with Mixture of Adapters
Kenichi Fujita, Takanori Ashihara, Marc Delcroix, Yusuke Ijima
INTERSPEECH4
2024 Knowledge Distillation from Self-Supervised Representation Learning Model with Discrete Speech Units for Any-to-Any Streaming Voice Conversion
Hiroki Kanagawa, Yusuke Ijima
INTERSPEECH2
2024 Pre-training Neural Transducer-based Streaming Voice Conversion for Faster Convergence and Alignment-free Training
Hiroki Kanagawa, Takafumi Moriya, Yusuke Ijima
INTERSPEECH3
2023 Enhancement of Text-Predicting Style Token With Generative Adversarial Network for Expressive Speech Synthesis
abstract
This work proposes an advanced text-predicting style embedding for expressive speech synthesis. Text-predicting global style token (TPGST) predicts style embedding from text instead of reference speech and uses it to condition a text-to-speech synthesis (TTS) model, resulting in style TTS without reference speech. Although this minimizes the style embedding’s L1-loss between that extracted from reference speech and that predicted during training, predicted embedding tends to be over-smoothed. To overcome this issue, the proposed method uses the generative adversarial network (GAN) in training style predictors. This not only improves style reproduction, but also aims to reduce style conditioning mismatch during TTS model training. We also utilize TTS text embeddings as in other related work, as well as word information via BERT in order to find better style distributions in GAN. An evaluation of subjective style reproduction demonstrates that 1) the proposed method outperforms conventional TPGST, and 2) the use of words yielded by BERT provides even better performance. Our style predictor is also effective in attaining unseen style TTS for "seen" and "unseen" speakers.
Hiroki Kanagawa, Yusuke Ijima
ICASSP2
2023 SpeechGLUE: How Well Can Self-Supervised Speech Models Capture Linguistic Knowledge?
Takanori Ashihara, Takafumi Moriya, Kohei Matsuura, Tomohiro Tanaka, Yusuke Ijima, Taichi Asami, Marc Delcroix, Yukinori Honma
INTERSPEECH5
2023 VC-T: Streaming Voice Conversion Based on Neural Transducer
Hiroki Kanagawa, Takafumi Moriya, Yusuke Ijima
INTERSPEECH3
2023 A stimulus-organism-response model of willingness to buy from advertising speech using voice quality
Mizuki Nagano, Yusuke Ijima, Sadao Hiroya
INTERSPEECH2
2023 Influence of Personal Traits on Impressions of One's Own Voice
Hikaru Yanagida, Yusuke Ijima, Naohiro Tawara
INTERSPEECH2
2022 Multi-Sample Subband Wavernn Via Multivariate Gaussian
abstract
This paper proposes a high-speed neural vocoder for CPU implementation. Two approaches for speeding up autoregressive neural vocoders have been proposed, 1) simultaneous multiple sample generation and 2) subband signal-based vocoder; so far they have been employed independently. Our neural vocoder is extremely fast as it generates multiple samples of subband signals simultaneously. Although there is an association between each subband signal, the conventional subband-based vocoder can degrade quality because each subband signal is generated from an independent probability distribution. To overcome this problem, we also introduce waveform generation that takes account of the association of each subband by employing multivariate Gaussian. Experiments show that 1) our proposed method is 1.81 times as fast as the conventional subband WaveRNN on a single-threaded CPU; 2) it outperformed the conventional method in a subjective evaluation in terms of naturalness, and achieved a mean opinion score (MOS) of 4.08 on text-to-speech task.
Hiroki Kanagawa, Yusuke Ijima
ICASSP2
2022 Joint Modeling of Multi-Sample and Subband Signals for Fast Neural Vocoding on CPU
Hiroki Kanagawa, Yusuke Ijima, Hiroyuki Toda
INTERSPEECH2
2022 Predicting VQVAE-based Character Acting Style from Quotation-Annotated Text for Audiobook Speech Synthesis
Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, Yuki Saito 0001, Yusuke Ijima, Ryo Masumura, Hiroshi Saruwatari
INTERSPEECH5
2022 SIMD-Size Aware Weight Regularization for Fast Neural Vocoding on CPU
abstract
This paper proposes weight regularization for a faster neural vocoder. Pruning time-consuming DNN modules is a promising way to realize a real-time vocoder on a CPU (e.g. WaveRNN, LPCNet). Regularization that encourages sparsity is also effective in avoiding the quality degradation created by pruning. However, the orders of weight matrices must be contiguous in SIMD size for fast vocoding. To ensure this order, we propose explicit SIMD size aware regularization. Our proposed method reshapes a weight matrix into a tensor so that the weights are aligned by group size in advance, and then computes the group Lasso-like regularization loss. Experiments on 70% sparse subband WaveRNN show that pruning in conventional Lasso and column-wise group Lasso degrades the synthetic speech's naturalness. The vocoder with proposed regularization 1) achieves comparable naturalness to that without pruning and 2) performs meaningfully faster than other conventional vocoders using regularization.
Hiroki Kanagawa, Yusuke Ijima
SLT2
2021 Robust Speech-Age Estimation Using Local Maximum Mean Discrepancy Under Mismatched Recording Conditions
abstract
A recently proposed time-delay neural network (TDNN)-based age estimation system has yielded state-of-the-art performance in speech-age estimation tasks. However, the performance of this TDNN-based system can seriously degrade when the recording conditions of each utterance are different in the training and testing phases. To tackle this problem, we examine the efficiencies of a series of unsupervised domain adaptation (UDA) methods to obtain the model invariance against the difference of these conditions. In particular, we propose using local maximum mean discrepancy (LMMD) with soft-target labels to consider an ordinal relationship between age labels. In most UDA methods, the model is trained to obtain domain invariant representations by minimizing the statistical difference of the distributions between labeled source and unlabeled target data without considering their age class labels. In contrast, our LMMD-based approach locally minimizes the differences in their distributions on each age class while considering adjacent age classes using soft-target labels. We conducted speech-age estimation experiments on in-house datasets under mismatched conditions including different background noise, reverberation, and microphones. The experimental comparison demonstrated that the LMMD-based method contributed to efficiently reducing the effect of mismatches of input data, yielding significant improvements over other UDA methods, such as MMD and reverse gradients.
Naohiro Tawara, Atsunori Ogawa, Yuki Kitagishi, Hosana Kamiyama, Yusuke Ijima
ASRU5
2021 Speech Emotion Recognition Based on Listener Adaptive Models
abstract
This paper presents a novel speech emotion recognition scheme that can deal with the individuality of emotion perception. Most conventional methods directly model the majority decision of multiple listener’s perceived emotions. However, emotion perception varies with the listener, which means the conventional methods can mismatch the recognition results to human perception. In order to mitigate this problem, we propose a Listener Adaptive (LA) model that reflects emotion recognition criteria of each listener. One-hot listener codes with several adaptation layers are employed in the LA model. The LA model yields the posterior probabilities of the listener-specific perceived emotions. Majority-voted emotion can be also estimated by averaging, in the LA model, the posterior probabilities for all listeners. Experiments on two emotional speech datasets demonstrate that the proposed approach offers improved listener-wise perceived emotion recognition performance in natural speech.
Atsushi Ando, Ryo Masumura, Hiroshi Sato 0002, Takafumi Moriya, Takanori Ashihara, Yusuke Ijima, Tomoki Toda
ICASSP6
2021 Simpleflat: A Simple Whole-Network Pre-Training Approach for RNN Transducer-Based End-to-End Speech Recognition
abstract
Recurrent neural network-transducer (RNN-T) is promising for building time-synchronous end-to-end automatic speech recognition (ASR) systems, in part because it does not need frame-wise alignment between input features and target labels in the training step. Although training without alignment is beneficial, it makes it difficult to discern the relation between input features and output token sequences. This, in effect, degrades RNN-T performance. Our solution is SimpleFlat (SF), a novel and simple whole-network pretraining approach for RNN-T. SF extracts frame-wise alignments on-the-fly from the training dataset, and does not require any external resources. We distribute equal numbers of target tokens to each frame following RNN-T encoder output lengths by repeating each token. The frame-wise tokens so created are shifted, and also used as the prediction network inputs. Therefore, SF can be implemented by cross entropy loss computation as in autoregressive model training. Experiments on Japanese and English ASR tasks demonstrate that SF can effectively improve various RNN-T architectures.
Takafumi Moriya, Takanori Ashihara, Tomohiro Tanaka, Tsubasa Ochiai, Hiroshi Sato 0002, Atsushi Ando, Yusuke Ijima, Ryo Masumura, Yusuke Shinohara
ICASSP7
2021 Phoneme Duration Modeling Using Speech Rhythm-Based Speaker Embeddings for Multi-Speaker Speech Synthesis
Kenichi Fujita, Atsushi Ando, Yusuke Ijima
Interspeech3
2021 Phonetic and Prosodic Information Estimation from Texts for Genuine Japanese End-to-End Text-to-Speech
Naoto Kakegawa, Sunao Hara, Masanobu Abe, Yusuke Ijima
Interspeech4
2021 Impact of Emotional State on Estimation of Willingness to Buy from Advertising Speech
Mizuki Nagano, Yusuke Ijima, Sadao Hiroya
Interspeech2
2021 Model architectures to extrapolate emotional expressions in DNN-based text-to-speech
Katsuki Inoue, Sunao Hara, Masanobu Abe, Nobukatsu Hojo, Yusuke Ijima
Speech Commun.5
2020 Lightweight LPCNet-Based Neural Vocoder with Tensor Decomposition
Hiroki Kanagawa, Yusuke Ijima
INTERSPEECH2
2020 Investigating Effective Additional Contextual Factors in DNN-Based Spontaneous Speech Synthesis
Yuki Yamashita, Tomoki Koriyama, Yuki Saito 0001, Shinnosuke Takamichi, Yusuke Ijima, Ryo Masumura, Hiroshi Saruwatari
INTERSPEECH5
2020 DNN-based Speech Synthesis Using Abundant Tags of Spontaneous Speech Corpus
abstract
In this paper, we investigate the effectiveness of using rich annotations in deep neural network (DNN)-based statistical speech synthesis. DNN-based frameworks typically use linguistic information as input features called context instead of directly using text. In such frameworks, we can synthesize not only reading-style speech but also speech with paralinguistic and nonlinguistic features by adding such information to the context. However, it is not clear what kind of information is crucial for reproducing paralinguistic and nonlinguistic features. Therefore, we investigate the effectiveness of rich tags in DNN-based speech synthesis according to the Corpus of Spontaneous Japanese (CSJ), which has a large amount of annotations on paralinguistic features such as prosody, disfluency, and morphological features. Experimental evaluation results shows that the reproducibility of paralinguistic features of synthetic speech was enhanced by adding such information as context.
Yuki Yamashita, Tomoki Koriyama, Yuki Saito 0001, Shinnosuke Takamichi, Yusuke Ijima, Ryo Masumura, Hiroshi Saruwatari
LREC5
2019 End-to-End Automatic Speech Recognition with a Reconstruction Criterion Using Speech-to-Text and Text-to-Speech Encoder-Decoders
Ryo Masumura, Hiroshi Sato 0002, Tomohiro Tanaka, Takafumi Moriya, Yusuke Ijima, Takanobu Oba
INTERSPEECH5
2018 Soft-Target Training with Ambiguous Emotional Utterances for DNN-Based Speech Emotion Classification
abstract
This paper presents a novel emotion classification method for natural speech. One of the problems in the state-of-the-art method based on Deep Neural Network (DNN) is the paucity of the training data compared to model complexity. To solve this problem, this paper utilizes the ambiguous emotional utterances, utterances that have no dominant target emotion label. While previous work ignored ambiguous emotional utterances for training, the proposed method leverages all annotated labels via soft-target training. In addition, this paper modifies the soft-target training in order to effectively handle both clear and ambiguous emotional utterances. Experiments show that the proposed method yields performance improvements in terms of both weighted and unweighted accuracies.
Atsushi Ando, Satoshi Kobashikawa, Hosana Kamiyama, Ryo Masumura, Yusuke Ijima, Yushi Aono
ICASSP5
2018 Neural Confnet Classification: Fully Neural Network Based Spoken Utterance Classification Using Word Confusion Networks
abstract
This paper describes neural ConfNet classification, a novel fully neural network based spoken utterance classification method that uses word confusion networks (ConfNets). Our motivation is to establish a spoken utterance classification method that can precisely understand natural language and robustly handle automatic speech recognition (ASR) errors. Remarkable progress has been made in neural networks for accurate modeling, however, most previous methods could not handle ASR errors since they were developed for reference transcriptions. Therefore, in our work we utilized ConfNets, which are compact and efficient graph representations of ASR hypotheses. Our idea is to regard the ConfNet as a sequence of bag-of-weighted-arcs and introduce a mechanism that converts the bag-of-weighted-arcs into a continuous representation called a modified weighted sum representation. This enables us to flexibly connect ConfNets to arbitrary model structures developed for reference transcriptions. We demonstrate the effectiveness of the neural ConfNet classification in dialogue act, extended named entity, and question type classification tasks.
Ryo Masumura, Yusuke Ijima, Taichi Asami, Hirokazu Masataki, Ryuichiro Higashinaka
ICASSP2
2018 Non-Parallel Voice Conversion Using Variational Autoencoders Conditioned by Phonetic Posteriorgrams and D-Vectors
abstract
This paper proposes novel frameworks for non-parallel voice conversion (VC) using variational autoencoders (VAEs). Although conventional VAE-based VC models can be trained using non-parallel speech corpora with given speaker representations, phonetic contents of the converted speech tend to vanish because of an over-regularization issue often observed in latent variables of the VAEs. To overcome the issue, this paper proposes a VAE-based non-parallel VC conditioned by not only the speaker representations but also phonetic contents of speech represented as phonetic posteriorgrams (PPGs). Since the phonetic contents are given during the training, we can expect that the VC models effectively learn speaker-independent latent features of speech. Focusing on the point, this paper also extends the conventional VAE-based non-parallel VC to many-to-many VC that can convert arbitrary speakers' characteristics into another arbitrary speakers' ones. We investigate two methods to estimate speaker representations for speakers not included in speech corpora used for training VC models: 1) adapting conventional speaker codes, and 2) using d-vectors for the speaker representations. Experimental results demonstrate that 1) PPGs successfully improve both naturalness and speaker similarity of the converted speech, and 2) both speaker codes and d-vectors can be adopted to the VAE-based many-to-many non-parallel VC.
Yuki Saito 0001, Yusuke Ijima, Kyosuke Nishida, Shinnosuke Takamichi
ICASSP2
2017 Generative adversarial network-based postfilter for statistical parametric speech synthesis
abstract
We propose a postfilter based on a generative adversarial network (GAN) to compensate for the differences between natural speech and speech synthesized by statistical parametric speech synthesis. In particular, we focus on the differences caused by over-smoothing, which makes the sounds muffled. Over-smoothing occurs in the time and frequency directions and is highly correlated in both directions, and conventional methods based on heuristics are too limited to cover all the factors (e.g., global variance was designed only to recover the dynamic range). To solve this problem, we focus on “spectral texture”, i.e., the details of the time-frequency representation, and propose a learning-based postfilter that captures the structures directly from the data. To estimate the true distribution, we utilize a GAN composed of a generator and a discriminator. This optimizes the generator to produce samples imitating the dataset according to the adversarial discriminator. This adversarial process encourages the generator to fit the true data distribution, i.e., to generate realistic spectral texture. Objective evaluation of experimental results shows that the GAN-based postfilter can compensate for detailed spectral structures including modulation spectrum, and subjective evaluation shows that its generated speech is comparable to natural speech.
Takuhiro Kaneko, Hirokazu Kameoka, Nobukatsu Hojo, Yusuke Ijima, Kaoru Hiramatsu, Kunio Kashino
ICASSP4
2017 DNN-SPACE: DNN-HMM-Based Generative Model of Voice F0 Contours for Statistical Phrase/Accent Command Estimation
Nobukatsu Hojo, Yasuhito Ohsugi, Yusuke Ijima, Hirokazu Kameoka
INTERSPEECH3
2017 Prosody Aware Word-Level Encoder Based on BLSTM-RNNs for DNN-Based Speech Synthesis
Yusuke Ijima, Nobukatsu Hojo, Ryo Masumura, Taichi Asami
INTERSPEECH1
2016 An Investigation of DNN-Based Speech Synthesis Using Speaker Codes
Nobukatsu Hojo, Yusuke Ijima, Hideyuki Mizuno
INTERSPEECH2
2016 Objective Evaluation Using Association Between Dimensions Within Spectral Features for Statistical Parametric Speech Synthesis
Yusuke Ijima, Taichi Asami, Hideyuki Mizuno
INTERSPEECH1
2015 Sub-band text-to-speech combining sample-based spectrum with statistically generated spectrum
Tadashi Inai, Sunao Hara, Masanobu Abe, Yusuke Ijima, Noboru Miyazaki, Hideyuki Mizuno
INTERSPEECH4
2015 Statistical model training technique based on speaker clustering approach for HMM-based speech synthesis
Yusuke Ijima, Noboru Miyazaki, Hideyuki Mizuno, Sumitaka Sakauchi
Speech Commun.1
2014 Prosodic variation enhancement using unsupervised context labeling for HMM-based expressive speech synthesis
Yu Maeno, Takashi Nose, Takao Kobayashi, Tomoki Koriyama, Yusuke Ijima, Hideharu Nakajima, Hideyuki Mizuno, Osamu Yoshioka
Speech Commun.5
2013 HMM-based expressive speech synthesis based on phrase-level F0 context labeling
abstract
This paper proposes a technique for adding more prosodic variations to the synthetic speech in HMM-based expressive speech synthesis. We create novel phrase-level F0 context labels from the residual information of F0 features between original and synthetic speech for the training data. Specifically, we classify the difference of average log F0 values between the original and synthetic speech into three classes which have perceptual meanings, i.e., high, neutral, and low of relative pitch at the phrase level. We evaluate both ideal and practical cases using appealing and fairy tale speech recorded under a realistic condition. In the ideal case, we examine the potential of our technique to modify the F0 patterns under a condition where the original F0 contours of test sentences are known. In the practical case, we show how the users intuitively modify the pitch by changing the initial F0 context labels obtained from the input text.
Yu Maeno, Takashi Nose, Takao Kobayashi, Tomoki Koriyama, Yusuke Ijima, Hideharu Nakajima, Hideyuki Mizuno, Osamu Yoshioka
ICASSP5
2012 Similar Speaker Selection Technique Based on Distance Metric Learning with Perceptual Voice Quality Similarity
Yusuke Ijima, Mitsuaki Isogai, Hideyuki Mizuno
INTERSPEECH1
2011 Correlation Analysis of Acoustic Features with Perceptual Voice Quality Similarity for Similar Speaker Selection
Yusuke Ijima, Mitsuaki Isogai, Hideyuki Mizuno
INTERSPEECH1
2011 HMM-Based Emphatic Speech Synthesis Using Unsupervised Context Labeling
abstract
This paper describes an approach to HMM-based expressive speech synthesis which does not require any supervised labeling process for emphasis context. We use appealing-style speech whose sentences were taken from real domains. To reduce the cost for labeling speech data with an emphasis context for the model training, we propose an unsupervised labeling technique of the emphasis context based on the difference between original and generated F0 patterns of training sentences. Although the criterion for the emphasis labeling is quite simple, subjective evaluation results reveal that the unsupervised labeling is comparable to the labeling conducted carefully by a human in terms of speech naturalness and emphasis reproducibility. Index Terms: HMM-based speech synthesis, expressive speech, emphasis expression, unsupervised labeling, F0 generation
Yu Maeno, Takashi Nose, Takao Kobayashi, Yusuke Ijima, Hideharu Nakajima, Hideyuki Mizuno, Osamu Yoshioka
INTERSPEECH4
2009 Emotional speech recognition based on style estimation and adaptation with multiple-regression HMM
abstract
This paper proposes a technique for emotional speech recognition which enables us to extract paralinguistic information as well as linguistic information contained in speech signal. The technique is based on style estimation and style adaptation using multiple-regression HMM. Recognition process consists of two stages. In the first stage, a style vector that represents the emotional expression category and intensity of its variation of input speech is estimated on a sentence-by-sentence basis. Then the acoustic models are adapted using the estimated style vector and standard HMM-based speech recognition is performed in the second stage. We assess the performance of the proposed technique on the recognition of acted emotional speech uttered by both professional narrators and non-professional speakers and show the effectiveness of the technique.
Yusuke Ijima, Makoto Tachibana, Takashi Nose, Takao Kobayashi
ICASSP1
2009 Speaking style adaptation for spontaneous speech recognition using multiple-regression HMM
Yusuke Ijima, Takeshi Matsubara, Takashi Nose, Takao Kobayashi
INTERSPEECH1
2008 An on-line adaptation technique for emotional speech recognition using style estimation with multiple-regression HMM
Yusuke Ijima, Makoto Tachibana, Takashi Nose, Takao Kobayashi
INTERSPEECH1