Ryoichi Takashima

dblp:22/8758 · DBLP profile ↗
← Back
37ranked-venue papers
13as first author
12since 2021 · last 2025
0000-0002-9808-0250ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 11 first-author · 6 since 2021Artificial intelligence and machine learning · 16 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Speaker-dependent Continuous Speech Recognition for Individuals with Cerebral Palsy Using Weighted Finite-State Transducer and Text-to-Speech Synthesis
abstract
Despite remarkable advances in automatic speech recognition (ASR) technology, existing systems have not achieved sufficient recognition accuracy for speech recognition of individuals with cerebral palsy.Speech recognition for individuals with speech disorders faces acoustic challenges because the speech characteristics of these individuals differ from those of individuals without speech disorders.Additionally, Japanese ASR faces unique linguistic challenges due to the mixed character set including kanji (Chinese characters), hiragana, and katakana (Japanese phonetic syllabary).Adapting end-to-end ASR models to this task requires large amounts of training data.However, collecting sufficient amounts of speech data for training is difficult because recording speech from individuals with cerebral palsy is a large burden on them.This paper revisit a weighted finite-state transducer based hybrid speaker-dependent ASR system for individuals with cerebral palsy, which decomposes the system into a speaker-dependent acoustic model, a pronunciation dictionary, and a language model.This approach is effective when speech data is limited, as it uses speech data only for training the acoustic model while other components are learned only from text data.Furthermore, to enhance the speaker-dependent acoustic model, we introduce data augmentation using text-to-speech synthesis and multi-step model adaptation using synthetic speech.Experimental validation using speech samples from individuals with cerebral palsy demonstrates that the proposed methodology achieves superior performance compared to the state-of-the-art end-to-end Whisper (ASR system).
Takeru Otani, Ryo Aihara, Ryoichi Takashima, Tetsuya Takiguchi, Tatsuhiko Saito
ASSETS3
2025 Highly Intelligible Text-to-Speech System Based on Weighted Averaging of Parameters for Individuals with Spinal Muscular Atrophy
abstract
To support communication for individuals with dysarthria who have difficulty producing intelligible speech, text-to-speech (TTS) systems are gaining attention.However, conventional TTS systems synthesize speech using voices that differ from those of the users themselves, which can create a sense of psychological distance between the user and their communication partner.Deep neural network-based TTS models can accurately reproduce trained speech and can generate speech resembling the user's own voice by training on the user's speech data.However, the models trained on speech with dysarthria also replicate the unintelligibility of the original speech, making them unsuitable for communication support.This paper focuses on dysarthria caused by spinal muscular atrophy (SMA), and proposes a method to construct a TTS model that synthesizes intelligible speech while preserving the voice characteristics of a speaker with dysarthria.The proposed approach involves computing a weighted average of the parameters of a TTS model trained on speech from an SMA speaker and a TTS model trained on speech from a speaker without dysarthria.The experimental results confirm that the synthesized speech generated using the proposed method maintains the voice quality of the SMA speaker while being more intelligible than that produced by conventional methods. CCS Concepts• Social and professional topics → Assistive technologies.
Yusuke Yagi, Ryoichi Takashima, Chiho Sasaki, Tetsuya Takiguchi
ASSETS2
2025 Revisiting WFST-based Hybrid Japanese Speech Recognition System for Individuals with Organic Speech Disorders
Naoki Hojo, Ryoichi Takashima, Chihiro Sugiyama, Nobukazu Tanaka, Kanji Nohara, Kazunori Nozaki, Tetsuya Takiguchi
INTERSPEECH2
2025 Zero-Shot Learning for Acoustic Event Classification Using an Attribute Vector and Conditional GAN
Kohei Uehara, Ryoichi Takashima, Tetsuya Takiguchi
INTERSPEECH2
2025 Operatic Singing Voice Synthesis From Inexperienced Voice Considering Tempo and Vowel Change
Aoto Sugahara, Soma Kishimoto, Yuji Adachi, Kiyoto Tai, Ryoichi Takashima, Tetsuya Takiguchi
MMM (3)5
2024 Individuality-Preserving Speech Synthesis for Spinal Muscular Atrophy with a Tracheotomy
abstract
Aphasia and dysarthria are the two main language disorders that cause difficulty in speech. This study focuses on articulation disorders, particularly among individuals with spinal muscular atrophy (SMA) whose speech is challenging to comprehend. Specifically, it addresses communication support through text-to-speech synthesis technology that maintains the speaker’s individuality. Previous research on individuals with SMA who have undergone tracheotomy surgery has predominantly centered on postoperative care environments unrelated to speech communication, with few precedents in the study of communication support using speech synthesis technology. Therefore, this study aims to develop a speech synthesis system that preserves the speaker’s individuality while producing clearer speech. This is performed by fine-tuning a pre-trained speech synthesis model, initially trained on a large corpus of speech by those with no speech impediment, using a small amount of speech of the target person with SMA. Subjective evaluations using both actual and synthesized speech demonstrated that the system could adequately learn the speaker’s individuality and produce synthesized speech with slightly improved clarity.
Minori Iwata, Ryoichi Takashima, Chiho Sasaki, Tetsuya Takiguchi
ASSETS2
2024 Self-supervised learning using unlabeled speech with multiple types of speech disorder for disordered speech recognition
abstract
This paper investigates a training method of an automatic speech recognition (ASR) model for people with speech disorders. Because the characteristics of their speech differ significantly from those of the typical speech, in order to recognize the speech of a user with a disorder, the system needs to be trained with the user’s speech in advance. However, recording speech from people with disorders is a large burden for them, and therefore, it is difficult to collect a sufficient amount of speech for training. To address this issue, this study investigates the use of two types of speech as training data. The first type is unlabeled speech, which can be easily collected but lacks text labels (e.g., spontaneous speech in daily life). To utilize the unlabeled speech for training an ASR model, a self-supervised learning approach is employed. The second type involves utilizing speech data from individuals with different types of speech disorders. In our system, besides the user’s speech, the speech of individuals with the same type of disorder and even different types of disorders is also incorporated. Experimental results demonstrated that using unlabeled speech and speech from multiple types of disorders led to reduced recognition error rates.
Ryoichi Takashima, Takeru Otani, Ryo Aihara, Tetsuya Takiguchi, Shinya Taguchi
ASSETS1
2023 Zero-Shot Sound Event Classification Using a Sound Attribute Vector with Global and Local Feature Learning
abstract
This paper introduces a zero-shot sound event classification (ZS-SEC) method to identify sound events that have never occurred in training data. In our previous work, we proposed a ZS-SEC method using sound attribute vectors (SAVs), where a deep neural network model infers attribute information that describes the sound of an event class instead of inferring its class label directly. Our previous method showed that it could classify unseen events to some extent; however, the accuracy for unseen events was far inferior to that for seen events. In this paper, we propose a new ZS-SEC method that can learn discriminative global features and local features simultaneously to enhance SAV-based ZS-SEC. In the proposed method, while the global features are learned in order to discriminate the event classes in the training data, the spectro-temporal local features are learned in order to regress the attribute information using attribute prototypes. The experimental results show that our proposed method can improve the accuracy of SAV-based ZS-SEC and can visualize the region in the spectrogram related to each attribute.
Xunquan Chen, Ryoichi Takashima, Tetsuya Takiguchi
ICASSP3
2023 Harmonic-Net: Fundamental Frequency and Speech Rate Controllable Fast Neural Vocoder
abstract
There is a need to improve the synthesis quality of HiFi-GAN-based real-time neural speech waveform generative models on CPUs while preserving the controllability of fundamental frequency ($f_{\mathrm{o}}$) and speech rate (SR). For this purpose, we propose Harmonic-Net and Harmonic-Net+, which introduce two extended functions into the HiFi-GAN generator. The first extension is a downsampling network, named the excitation signal network, that hierarchically receives multi-channel excitation signals corresponding to$f_{\mathrm{o}}$. The second extension is the layerwise pitch-dependent dilated convolutional network (LW-PDCNN), which can flexibly change its receptive fields depending on the input$f_{\mathrm{o}}$to handle large fluctuations in$f_{\mathrm{o}}$for the upsampling-based HiFi-GAN generator. The proposed explicit input of excitation signals and LW-PDCNNs corresponding to$f_{\mathrm{o}}$are expected to realize high-quality synthesis for the normal and$f_{\mathrm{o}}$-conversion conditions and for the SR-conversion condition. The results of experiments for unseen speaker synthesis, full-band singing voice synthesis, and text-to-speech synthesis show that the proposed method with harmonic waves corresponding to$f_{\mathrm{o}}$can achieve higher synthesis quality than conventional methods in all (i.e., normal,$f_{\mathrm{o}}$-conversion, and SR-conversion) conditions.
Keisuke Matsubara, Takuma Okamoto, Ryoichi Takashima, Tetsuya Takiguchi, Tomoki Toda, Hisashi Kawai
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Speaker-Targeted Audio-Visual Speech Recognition Using a Hybrid CTC/Attention Model with Interference Loss
abstract
Audio-visual (AV)-automatic speech recognition (ASR) can improve speech recognition accuracy by using lip images, especially in noisy environments. The recently proposed AV Align system integrates speech and image features based on a cross-modal attention mechanism, where attention weights for visual features are estimated by using acoustic features as queries. Although AV Align shows an improvement in recognition accuracy in background noise environments, we have observed that the recognition accuracy degrades significantly in interference speaker environments, where a target speech and an interfering speech overlap each other. In order to improve the speech recognition accuracy of the target speaker in such situations, we propose a method that combines the auxiliary loss function that maximizes the recognition accuracy of the interference speaker and the CTC loss function for training the AV-ASR model. The experimental results using the TCD-TIMIT dataset show that the use of these auxiliary loss functions improves the performance of target-speaker speech recognition in interference speaker environments.
Ryota Tsunoda, Ryo Aihara, Ryoichi Takashima, Tetsuya Takiguchi, Yoshie Imai
ICASSP3
2021 High-Intelligibility Speech Synthesis for Dysarthric Speakers with LPCNet-Based TTS and CycleVAE-Based VC
abstract
This paper presents a high-intelligibility speech synthesis method for persons with dysarthria caused by athetoid cerebral palsy. The muscular control of such speakers is unstable because of their athetoid symptoms, and their pronunciation is unclear, which makes it difficult for them to communicate. In this paper, we present a method for generating highly intelligible speech that preserves the individuality of dysarthric speakers by combining Transformer-TTS, CycleVAE-VC, and a LPCNet vocoder. Rather than repairing prosody from the dysarthric speech, this method transfers the dysarthric speaker’s individuality to the speech of a healthy person generated by TTS synthesis. This task is both important and challenging. From the results of our evaluation experiments, we confirmed that the proposed method can partially transfer the individuality of the target dysarthric speaker while maintaining the intelligibility of the source speech.
Keisuke Matsubara, Takuma Okamoto, Ryoichi Takashima, Tetsuya Takiguchi, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai
ICASSP3
2021 Multimodal fusion for indoor sound source localization
Ryoichi Takashima, Xingchen Guo, Zhihong Zhang 0001, Xuexin Xu, Tetsuya Takiguchi, Edwin R. Hancock
Pattern Recognit.2
2020 FasterRCNN Monitoring of Road Damages: Competition and Deployment
abstract
Maintaining aging infrastructure is a challenge currently faced by local and national administrators all around the world. An important prerequisite for efficient infrastructure maintenance is to continuously monitor (i.e., quantify the level of safety and reliability) the state of very large structures. Meanwhile, computer vision has made impressive strides in recent years, mainly due to successful applications of deep learning models. These novel progresses are allowing the automation of vision tasks, which were previously impossible to automate, offering promising possibilities to assist administrators in optimizing their infrastructure maintenance operations. In this context, the IEEE 2020 global Road Damage Detection (RDD) Challenge is giving an opportunity for deep learning and computer vision researchers to get involved and help accurately track pavement damages on road networks. This paper proposes two contributions to that topic: In a first part, we detail our solution to the RDD Challenge. In a second part, we present our efforts in deploying our model on a local road network, explaining the proposed methodology and encountered challenges.
Tristan Hascoet, Andreas Persch, Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki
IEEE BigData4
2020 Two-Step Acoustic Model Adaptation for Dysarthric Speech Recognition
abstract
This paper introduces a model adaptation approach for a speaker-dependent dysarthric speech recognition system. The dysarthria we focus on in this paper is caused by athetoid cerebral palsy, which causes involuntary muscle movements in those with the disease. For this reason, the dysarthric people's speech is often unstable and difficult for conventional automatic speech recognition (ASR) systems to recognize. A model-adaptation approach, which adapts an ASR model to dysarthric speech, is one possible solution. However, because the difference in speaking styles between dysarthric and non-dysarthric people is so significant, the conventional adaptation method is not able to sufficiently adapt the model to the dysarthric speech. In our proposed two-step model-adaptation approach, an ASR model is first adapted to the general speaking style of multiple dysarthric speakers, and then the adapted model is further adapted for the target speaker. From our experiments on an ASR task, our two-step adaptation approach showed better performance than a conventional one-step adaptation approach.
Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki
ICASSP1
2020 Dysarthric Speech Recognition Based on Deep Metric Learning
Yuki Takashima, Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki
INTERSPEECH2
2019 Investigation of Sequence-level Knowledge Distillation Methods for CTC Acoustic Models
abstract
This paper presents knowledge distillation (KD) methods for training connectionist temporal classification (CTC) acoustic models. In a previous study, we proposed a KD method based on the sequence-level cross-entropy, and showed that the conventional KD method based on the frame-level cross-entropy did not work effectively for CTC acoustic models, whereas the proposed method improved the performance of the models. In this paper, we investigate the implementation of sequence-level KD for CTC models and propose a lattice-based sequence-level KD method. Experiments investigating model compression and the training of a noise-robust model using the Wall Street Journal (WSJ) and CHiME4 datasets demonstrate that the sequence-level KD methods improve the performance of CTC acoustic models on both two tasks, and show that the lattice-based method can compute the sequence-level KD more efficiently than the N-best-based method proposed in our previous work.
Ryoichi Takashima, Sheng Li 0010, Hisashi Kawai
ICASSP1
2019 Auxiliary Interference Speaker Loss for Target-Speaker Speech Recognition
abstract
In this paper, we propose a novel auxiliary loss function for target-speaker automatic speech recognition (ASR). Our method automatically extracts and transcribes target speaker's utterances from a monaural mixture of multiple speakers speech given a short sample of the target speaker. The proposed auxiliary loss function attempts to additionally maximize interference speaker ASR accuracy during training. This will regularize the network to achieve a better representation for speaker separation, thus achieving better accuracy on the target-speaker ASR. We evaluated our proposed method using two-speaker-mixed speech in various signal-to-interference-ratio conditions. We first built a strong target-speaker ASR baseline based on the state-of-the-art lattice-free maximum mutual information. This baseline achieved a word error rate (WER) of 18.06% on the test set while a normal ASR trained with clean data produced a completely corrupted result (WER of 84.71%). Then, our proposed loss further reduced the WER by 6.6% relative to this strong baseline, achieving a WER of 16.87%. In addition to the accuracy improvement, we also showed that the auxiliary output branch for the proposed loss can even be used for a secondary ASR for interference speakers' speech.
Naoyuki Kanda, Shota Horiguchi, Ryoichi Takashima, Yusuke Fujita, Kenji Nagamatsu, Shinji Watanabe 0001
INTERSPEECH3
2018 An Investigation of a Knowledge Distillation Method for CTC Acoustic Models
abstract
End-to-end acoustic models, such as connectionist temporal classification (CTC) and the attention model, have been studied, and their speech recognition accuracies come close to those of conventional deep neural network (DNN)-hidden Markov models. However, most high-performance end-to-end models are not suitable for real-time (streaming) speech recognition because they are based on bidirectional recurrent neural networks (RNNs). In this study, to improve the performance of unidirectional RNN-based CTC, which is suitable for real-time processing, we investigate the knowledge distillation (KD)-based model compression method for training a CTC acoustic model. we evaluate a frame-level KD method and a sequence-level KD method for CTC model. The speech recognition experiments on Wall Street Journal tasks demonstrate that, the frame-level KD worsens the WERs ofunidirectional CTC model, whereas sequence-level KD can improve the WERs of the model.
Ryoichi Takashima, Sheng Li 0010, Hisashi Kawai
ICASSP1
2018 CTC Loss Function with a Unit-Level Ambiguity Penalty
abstract
This paper presents a modified loss function for training connectionist temporal classification (CTC)-based acoustic models. CTC-based acoustic models have been studied as alternatives to conventional hidden Markov models (HMMs), but have often shown worse performance than conventional deep neural network (DNN)-HMM hybrid models. In this paper, we attempt to identify the primary factor preventing CTC-based models from achieving their full potential, and hypothesize this constraint lies in the ambiguity in the identification boundaries among unit-level labels (phonemes or characters). In accordance with this hypothesis, we propose a modified CTC loss function using an ambiguity penalty. This penalty is defined by the conditional entropy and works to increase the separation metrics among unit-level labels. We evaluate the proposed method on the WSJ and CHiME4 tasks, and demonstrate that our modification improves the word error rate compared with that of the conventional CTC-based model when the training dataset is small.
Ryoichi Takashima, Sheng Li 0010, Hisashi Kawai
ICASSP1
2018 Improving CTC-based Acoustic Model with Very Deep Residual Time-delay Neural Networks
Sheng Li 0010, Xugang Lu, Ryoichi Takashima, Tatsuya Kawahara, Hisashi Kawai
INTERSPEECH3
2018 Improving Very Deep Time-Delay Neural Network With Vertical-Attention For Effectively Training CTC-Based ASR Systems
abstract
The very deep neural network has recently been proposed for speech recognition and achieves significant performance. It has excellent potential for integration with end-to-end (E2E) training. Connectionist temporal classification (CTC) has shown great potential in E2E acoustic modeling. In this study, we investigate deep architectures and techniques which are suitable for CTC-based acoustic modeling. We propose a very deep residual time-delay CTC neural network (VResTD-CTC). How to select a suitable deep architecture optimized with the CTC objective function is crucial for obtaining the state of the art performance. Excellent performances can be obtained by selecting deep architecture for non-E2E ASR systems modeling with tied-triphone states. However, these optimized structures do not guarantee to achieve better or comparable performances on E2E (e.g., CTC-based) systems modeling with dynamic acoustic units. For solving this problem and further leveraging the system performance, we introduce the vertical-attention mechanism to reweight the residual blocks at each time step. Speech recognition experiments show our proposed model significantly outperforms the DNN and LSTM-based (both bidirectional and unidirectional) CTC baseline models.
Sheng Li 0010, Xugang Lu, Ryoichi Takashima, Tatsuya Kawahara, Hisashi Kawai
SLT3
2017 Incremental training and constructing the very deep convolutional residual network acoustic models
abstract
Inspired by the successful applications in image recognition, the very deep convolutional residual network (ResNet) based model has been applied in automatic speech recognition (ASR). However, the computational load is heavy for training the ResNet with a large quantity of data. In this paper, we propose an incremental model training framework to accelerate the training process of the ResNet. The incremental model training framework is based on the unequal importance of each layer and connection in the ResNet. The modules with important layers and connections are regarded as a skeleton model, while those left are regarded as an auxiliary model. The total depth of the skeleton model is quite shallow compared to the very deep full network. In our incremental training, the skeleton model is first trained with the full training data set. Other layers and connections belonging to the auxiliary model are gradually attached to the skeleton model and tuned. Our experiments showed that the proposed incremental training obtained comparable performances and faster training speed compared with the model training as a whole without consideration of the different importance of each layer.
Sheng Li 0010, Xugang Lu, Ryoichi Takashima, Tatsuya Kawahara, Hisashi Kawai
ASRU4
2016 Data Augmentation Using Multi-Input Multi-Output Source Separation for Deep Neural Network Based Acoustic Modeling
Yusuke Fujita, Ryoichi Takashima, Takeshi Homma, Masahito Togami
INTERSPEECH2
2015 Unified ASR system using LGM-based source separation, noise-robust feature extraction, and word hypothesis selection
abstract
In this paper, we propose a unified system that incorporates speech source separation and automatic speech recognition for various noise environments. There are three features in the proposed system. The first feature of the proposed method is the LGM (local Gaussian modeling) based source separation with the efficient permutation alignment method that integrates a power spectrum correlation based method and a direction-of-arrival (DOA) based method. Evaluation results show that using the separated speech with the baseline acoustic modeling method reduces the word error rate (WER) significantly. The second feature of the proposed method is multi-condition training with per-utterance normalized features and noise-aware features in the acoustic modeling step. In this paper, we show that the proposed training method is effective even when an input signal has been distorted through the source separation step. The third feature is the word hypothesis selection method for integrating multiple recognition results. The proposed selection method estimates correct words based on a recognizer's confidence and co-occurrence characteristics. The evaluation results show that the proposed selection method outperforms the conventional recognizer output voting error reduction (ROVER) method. The proposed system is evaluated using the third CHiME challenge dataset. Evaluation results show that the proposed system resulted in an improvement of 66.1% over the baseline system.
Yusuke Fujita, Ryoichi Takashima, Takeshi Homma, Rintaro Ikeshita, Yohei Kawaguchi, Takashi Sumiyoshi, Takashi Endo, Masahito Togami
ASRU2
2014 Frequency domain acoustic echo reduction based on Kalman smoother with time-varying noise covariance matrix
abstract
In this paper, we propose a novel acoustic-echo-reduction technique at a time-frequency domain, which is optimally combined with speech enhancement. Unlike conventional echo reduction techniques which minimizes only residual power of the far-end acoustic echo signal, the proposed method minimizes summation of the residual echo signal and distortion of the near-end speech signal from a minimum mean square error (MMSE) perspective. The proposed method performs echo reduction with speech enhancement and parameter optimization in an iterative manner based on the expectation-maximization (EM) algorithm. The E step is corresponding with the echo reduction and speech enhancement based on the Kalman smoother with a time-varying covariance matrix for the observation noise term, which reflects the time-varying characteristics of speech sources. By using the time-varying covariance matrix, we can enhance speech sources effectively with acoustic echo reduction. Associated with the time-varying covariance matrix, a new optimization scheme of parameters for the M step is derived in this paper. Experimental results with impulse responses which was recorded under a real meeting room show that the proposed method can effectively enhance a near-end speech signal when there are a near-end speech signal and a far-end acoustic echo signal.
Masahito Togami, Yohei Kawaguchi, Ryoichi Takashima
ICASSP3
2013 Individuality-preserving voice conversion for articulation disorders based on non-negative matrix factorization
abstract
We present in this paper a voice conversion (VC) method for a person with an articulation disorder resulting from athetoid cerebral palsy. The movement of such speakers is limited by their athetoid symptoms, and their consonants are often unstable or unclear, which makes it difficult for them to communicate. In this paper, exemplar-based spectral conversion using Non-negative Matrix Factorization (NMF) is applied to a voice with an articulation disorder. To preserve the speaker's individuality, we used a combined dictionary that is constructed from the source speaker's vowels and target speaker's consonants. Experimental results indicate that the performance of NMF-based VC is considerably better than conventional GMM-based VC.
Ryo Aihara, Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki
ICASSP2
2013 Prediction of unlearned position based on local regression for single-channel talker localization using acoustic transfer function
abstract
This paper presents a sound-source (talker) localization method using only a single microphone. In our previous work, we discussed the single-channel sound-source localization method based on the discrimination of the acoustic transfer function. However, that method requires the training of the acoustic transfer function for each possible position in advance, and it is difficult to estimate the position that has not been pre-trained. In order to estimate such unlearned positions, in this paper, we discuss a single-channel talker localization method based on a regression model, which predicts the position from the acoustic transfer function. For training the regression model, we use the local regression approach, which trains the regression model from only training samples that are similar to the evaluation data. Considering both the linear and non-linear regression models, the effectiveness of this method has been confirmed by sound-source localization experiments performed in different room environments.
Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki
ICASSP1
2013 Exemplar-based individuality-preserving voice conversion for articulation disorders in noisy environments
abstract
We present in this paper a noise robust voice conversion (VC) method for a person with an articulation disorder resulting from athetoid cerebral palsy. The movements of such speakers are limited by their athetoid symptoms, and their consonants are often unstable or unclear, which makes it difficult for them to communicate. In this paper, exemplar-based spectral conversion using Non-negative Matrix Factorization (NMF) is applied to a voice with an articulation disorder in real noisy environments. In this paper, in order to deal with background noise, an input noisy source signal is decomposed into the clean source exemplars and noise exemplars by NMF. Also, to preserve the speaker’s individuality, we use a combined dictionary that was constructed from the source speaker’s vowels and target speaker’s consonants. The effectiveness of this method was confirmed by comparing its effectiveness with that of a conventional Gaussian Mixture Model (GMM)-based method. Index Terms: Voice Conversion, NMF, Articulation Disorders, Noise Robustness, Assistive Technologies
Ryo Aihara, Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki
INTERSPEECH2
2013 Voice conversion in high-order eigen space using deep belief nets
abstract
This paper presents a voice conversion technique using Deep Belief Nets (DBNs) to build high-order eigen spaces of the source/target speakers, where it is easier to convert the source speech to the target speech than in the traditional cepstrum space. DBNs have a deep architecture that automatically discovers abstractions to maximally express the original input features. If we train the DBNs using only the speech of an individual speaker, it can be considered that there is less phonological information and relatively more speaker individuality in the output features at the highest layer. Training the DBNs for a source speaker and a target speaker, we can then connect and convert the speaker individuality abstractions using Neural Networks (NNs). The converted abstraction of the source speaker is then brought back to the cepstrum space using an inverse process of the DBNs of the target speaker. We conducted speakervoice conversion experiments and confirmed the efficacy of our method with respect to subjective and objective criteria, comparing it with the conventional Gaussian Mixture Model-based method.
Toru Nakashika, Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki
INTERSPEECH2
2012 A new multiple-kernel-learning weighting method for localizing human brain magnetic activity
abstract
This paper shows that pattern classification based on machine learning is a powerful tool to analyze human brain activity data obtained by magnetoencephalography (MEG). We propose a new weighting method using a multiple kernel learning (MKL) algorithm to localize the brain area contributing to the accurate vowel discrimination. Our MKL simultaneously estimates both the classification boundary and the weight of each MEG sensor; MEG amplitude obtained from each pair of sensors is an element of the feature vector. The estimated weight indicates how the corresponding sensor is useful for classifying the MEG response patterns. Our results show both the large-weight MEG sensors mainly in a language area of the brain and the high classification accuracy (73.0%) in the 100 ~ 200 ms latency range.
Tetsuya Takiguchi, Toshiaki Imada, Ryoichi Takashima, Yasuo Ariki, Jo-Fu Lotus Lin, Patricia K. Kuhl, Masaki Kawakatsu, Makoto Kotani
ICASSP3
2012 Estimation of Talker's Head Orientation Based on Discrimination of the Shape of Cross-power Spectrum Phase Coefficients
abstract
This paper presents a talker’s head orientation estimation method using 2-channel microphones. In recent research, some approaches based on a network of microphone arrays have been proposed in order to estimate the talker’s head orientation. In those methods, the talker’s head orientation is estimated using the sound amplitude or peak value of CSP (Cross-power Spectrum Phase) coefficients obtained from each microphone array. However, microphone array network systems need many microphone arrays to be set along the walls of a given room so that sub-microphone arrays surround the user. In this paper, we focus on the shape of the CSP coefficients affected by the reverberation, which depends on the talker’s position and the head orientation. In our proposed method, we use not only the peak value but also the other values of the CSP coefficients as feature vectors, and the talker’s position and the head orientation are estimated by discriminating the CSP vector. The effectiveness of this method has been confirmed by talker localization and head orientation estimation experiments performed in a real environment.
Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki
INTERSPEECH1
2012 Exemplar-based voice conversion in noisy environment
abstract
This paper presents a voice conversion (VC) technique for noisy environments, where parallel exemplars are introduced to encode the source speech signal and synthesize the target speech signal. The parallel exemplars (dictionary) consist of the source exemplars and target exemplars, having the same texts uttered by the source and target speakers. The input source signal is decomposed into the source exemplars, noise exemplars obtained from the input signal, and their weights (activities). Then, by using the weights of the source exemplars, the converted signal is constructed from the target exemplars. We carried out speaker conversion tasks using clean speech data and noise-added speech data. The effectiveness of this method was confirmed by comparing its effectiveness with that of a conventional Gaussian Mixture Model (GMM)-based method.
Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki
SLT1
2011 Feature selection based on Multiple Kernel Learning for single-channel sound source localization using the acoustic transfer function
abstract
This paper presents a sound source (talker) localization method using only a single microphone. In our previous work [1], we discussed the single-channel sound source localization method, where the acoustic transfer function from a user's position is estimated by using a Hidden Markov Model (HMM) of clean speech in the cepstral domain. In this paper, each cepstral dimension of the acoustic transfer function is newly selected in order to select the cepstral dimensions having information that is useful for classifying the user's position. Then, we propose a feature selection method for the cepstral parameter using Multiple Kernel Learning (MKL) to define the base kernels for each cepstral dimension (scalar) of the acoustic transfer function. The user's position is trained and classified by Support Vector Machine (SVM). The effectiveness of this method has been confirmed by sound source (talker) localization experiments performed in a room environment.
Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki
ICASSP1
2011 Agglomerative Hierarchical Clustering of Emotions in Speech Based on Subjective Relative Similarity
Ryoichi Takashima, Tohru Nagano, Ryuki Tachibana, Masafumi Nishimura
INTERSPEECH1
2011 Single-Channel Head Orientation Estimation Based on Discrimination of Acoustic Transfer Function
abstract
This paper presents a talker’s head orientation estimation method using only a single microphone, where phoneme HMMs (Hidden Markov Models) of clean speech are introduced to separate the acoustic transfer function at the user’s position and head orientation. The frame sequence of the acoustic transfer function is estimated by maximizing the likelihood of training data uttered from a given position with a given head orientation. Using the separated frame sequence data, the user’s position and the head orientation are trained by Support Vector Machine (SVM) in advance. Then, for each test utterance, the frame sequence of the acoustic transfer function is separated based on the maximum likelihood estimation using the label sequence obtained from the phoneme recognition, and the user’s position and head orientation are estimated by discriminating the separated acoustic transfer function using SVM. The effectiveness of this method has been confirmed by talker localization and head orientation estimation experiments performed in a real environment. Index Terms: single channel, talker localization, head orientation, acoustic transfer function
Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki
INTERSPEECH1
2010 HMM-based separation of acoustic transfer function for single-channel sound source localization
abstract
This paper presents a sound source (talker) localization method using only a single microphone, where a HMM (Hidden Markov Model) of clean speech is introduced to estimate the acoustic transfer function from a user's position. The new method is able to carry out this estimation without measuring impulse responses. The frame sequence of the acoustic transfer function is estimated by maximizing the likelihood of training data uttered from a given position, where the cepstral parameters are used to effectively represent useful clean speech. Using the estimated frame sequence data, the GMM (Gaussian Mixture Model) of the acoustic transfer function is created to deal with the influence of a room impulse response. Then, for each test data set, we find a maximum-likelihood GMM from among the estimated GMMs corresponding to each position. The effectiveness of this method has been confirmed by talker localization experiments performed in a room environment.
Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki
ICASSP1
2009 Monaural sound-source-direction estimation using the acoustic transfer function of an active microphone
Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki
FUSION1