Tetsuya Takiguchi

dblp:79/4485 · DBLP profile ↗
← Back
115ranked-venue papers
9as first author
20since 2021 · last 2025
0000-0001-5005-7679ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 91 · 8 first-author · 9 since 2021Artificial intelligence and machine learning · 52 · 2 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 10 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 since 2021Databases, data management, data science and information retrieval · 3
YearPublicationVenuePosition
2025 Speaker-dependent Continuous Speech Recognition for Individuals with Cerebral Palsy Using Weighted Finite-State Transducer and Text-to-Speech Synthesis
abstract
Despite remarkable advances in automatic speech recognition (ASR) technology, existing systems have not achieved sufficient recognition accuracy for speech recognition of individuals with cerebral palsy.Speech recognition for individuals with speech disorders faces acoustic challenges because the speech characteristics of these individuals differ from those of individuals without speech disorders.Additionally, Japanese ASR faces unique linguistic challenges due to the mixed character set including kanji (Chinese characters), hiragana, and katakana (Japanese phonetic syllabary).Adapting end-to-end ASR models to this task requires large amounts of training data.However, collecting sufficient amounts of speech data for training is difficult because recording speech from individuals with cerebral palsy is a large burden on them.This paper revisit a weighted finite-state transducer based hybrid speaker-dependent ASR system for individuals with cerebral palsy, which decomposes the system into a speaker-dependent acoustic model, a pronunciation dictionary, and a language model.This approach is effective when speech data is limited, as it uses speech data only for training the acoustic model while other components are learned only from text data.Furthermore, to enhance the speaker-dependent acoustic model, we introduce data augmentation using text-to-speech synthesis and multi-step model adaptation using synthetic speech.Experimental validation using speech samples from individuals with cerebral palsy demonstrates that the proposed methodology achieves superior performance compared to the state-of-the-art end-to-end Whisper (ASR system).
Takeru Otani, Ryo Aihara, Ryoichi Takashima, Tetsuya Takiguchi, Tatsuhiko Saito
ASSETS4
2025 Highly Intelligible Text-to-Speech System Based on Weighted Averaging of Parameters for Individuals with Spinal Muscular Atrophy
abstract
To support communication for individuals with dysarthria who have difficulty producing intelligible speech, text-to-speech (TTS) systems are gaining attention.However, conventional TTS systems synthesize speech using voices that differ from those of the users themselves, which can create a sense of psychological distance between the user and their communication partner.Deep neural network-based TTS models can accurately reproduce trained speech and can generate speech resembling the user's own voice by training on the user's speech data.However, the models trained on speech with dysarthria also replicate the unintelligibility of the original speech, making them unsuitable for communication support.This paper focuses on dysarthria caused by spinal muscular atrophy (SMA), and proposes a method to construct a TTS model that synthesizes intelligible speech while preserving the voice characteristics of a speaker with dysarthria.The proposed approach involves computing a weighted average of the parameters of a TTS model trained on speech from an SMA speaker and a TTS model trained on speech from a speaker without dysarthria.The experimental results confirm that the synthesized speech generated using the proposed method maintains the voice quality of the SMA speaker while being more intelligible than that produced by conventional methods. CCS Concepts• Social and professional topics → Assistive technologies.
Yusuke Yagi, Ryoichi Takashima, Chiho Sasaki, Tetsuya Takiguchi
ASSETS4
2025 Revisiting WFST-based Hybrid Japanese Speech Recognition System for Individuals with Organic Speech Disorders
Naoki Hojo, Ryoichi Takashima, Chihiro Sugiyama, Nobukazu Tanaka, Kanji Nohara, Kazunori Nozaki, Tetsuya Takiguchi
INTERSPEECH7
2025 Zero-Shot Learning for Acoustic Event Classification Using an Attribute Vector and Conditional GAN
Kohei Uehara, Ryoichi Takashima, Tetsuya Takiguchi
INTERSPEECH3
2025 Operatic Singing Voice Synthesis From Inexperienced Voice Considering Tempo and Vowel Change
Aoto Sugahara, Soma Kishimoto, Yuji Adachi, Kiyoto Tai, Ryoichi Takashima, Tetsuya Takiguchi
MMM (3)6
2025 A Robot that Supports Collaborative Art Appreciation through Visual Thinking Strategies
abstract
Social robots are increasingly being used as interactive partners to facilitate people’s understanding and appreciation of art. Considering people's stages of aesthetic development is essential to richer understanding of artworks, but such a viewpoint is less of a focus in human-robot interaction contexts. Therefore, we developed a robot system for collaborative art appreciation using a Visual Thinking Strategies (VTS) method to enhance people’s engagement with art by considering their stages of aesthetic experience. We conducted an experiment to investigate the effectiveness of our system in supporting art appreciation of participants in a laboratory setting. We also investigated the effects of embodiment, i.e., the physical body of a robot, on art-appreciation support. The experiment results indicate that our system significantly increased the intention to use it, which is related to social acceptance, and embodiment significantly influenced likeability, perceived intelligence, and perceived enjoyment.
Minori Iwata, Masahiro Shiomi, Tetsuya Takiguchi
RO-MAN3
2024 Individuality-Preserving Speech Synthesis for Spinal Muscular Atrophy with a Tracheotomy
abstract
Aphasia and dysarthria are the two main language disorders that cause difficulty in speech. This study focuses on articulation disorders, particularly among individuals with spinal muscular atrophy (SMA) whose speech is challenging to comprehend. Specifically, it addresses communication support through text-to-speech synthesis technology that maintains the speaker’s individuality. Previous research on individuals with SMA who have undergone tracheotomy surgery has predominantly centered on postoperative care environments unrelated to speech communication, with few precedents in the study of communication support using speech synthesis technology. Therefore, this study aims to develop a speech synthesis system that preserves the speaker’s individuality while producing clearer speech. This is performed by fine-tuning a pre-trained speech synthesis model, initially trained on a large corpus of speech by those with no speech impediment, using a small amount of speech of the target person with SMA. Subjective evaluations using both actual and synthesized speech demonstrated that the system could adequately learn the speaker’s individuality and produce synthesized speech with slightly improved clarity.
Minori Iwata, Ryoichi Takashima, Chiho Sasaki, Tetsuya Takiguchi
ASSETS4
2024 Self-supervised learning using unlabeled speech with multiple types of speech disorder for disordered speech recognition
abstract
This paper investigates a training method of an automatic speech recognition (ASR) model for people with speech disorders. Because the characteristics of their speech differ significantly from those of the typical speech, in order to recognize the speech of a user with a disorder, the system needs to be trained with the user’s speech in advance. However, recording speech from people with disorders is a large burden for them, and therefore, it is difficult to collect a sufficient amount of speech for training. To address this issue, this study investigates the use of two types of speech as training data. The first type is unlabeled speech, which can be easily collected but lacks text labels (e.g., spontaneous speech in daily life). To utilize the unlabeled speech for training an ASR model, a self-supervised learning approach is employed. The second type involves utilizing speech data from individuals with different types of speech disorders. In our system, besides the user’s speech, the speech of individuals with the same type of disorder and even different types of disorders is also incorporated. Experimental results demonstrated that using unlabeled speech and speech from multiple types of disorders led to reduced recognition error rates.
Ryoichi Takashima, Takeru Otani, Ryo Aihara, Tetsuya Takiguchi, Shinya Taguchi
ASSETS4
2024 Attempts on detecting Alzheimer's disease by fine-tuning pre-trained model with Gaze Data
abstract
This poster presents a study on detecting Alzheimer’s disease (AD) using deep learning from gaze data. In this study, we modify an existing pre-trained deep neural network model, gazeNet, for transfer learning. The results suggest the possibility of applying this method to mild cognitive impairment screening tests.
Junichi Nagasawa, Yuichi Nakata, Mamoru Hiroe, Yutaka Kawaguchi, Yuji Maegawa, Naoki Hojo, Tetsuya Takiguchi, Minoru Nakayama, Maki Uchimura, Yuma Sonoda, Hisatomo Kowa, Takashi Nagamatsu
ETRA8
2024 Attempts on detecting Alzheimer's disease by fine-tuning pre-trained model with Gaze Data
abstract
Early detection of Alzheimer’s disease (AD) is important but difficult. Screening for AD using neuropsychological tests such as mini-mental state examination (MMSE) is time-consuming and burdensome for patients. Recently, several methods have been reported for detecting AD based on eye movements. However, analyzing eye movements requires considerable effort. Although machine learning from eye movement data is a strong candidate for labor-saving, it requires large datasets. In this study, we modify an existing pre-trained deep neural network model, gazeNet, for transfer learning. For evaluation, we exclusively used data from one participant and fine-tuned the model using data from all the remaining participants. We repeated this procedure separately for each of the 14 participants. The results of eye movement during the antisaccade task were not satisfactory for the discrimination of AD, and detailed analysis suggested that the data might potentially have a correlation with MMSE scores in the mild cognitive impairment range.
Junichi Nagasawa, Yuichi Nakata, Mamoru Hiroe, Yutaka Kawaguchi, Yuji Maegawa, Naoki Hojo, Tetsuya Takiguchi, Minoru Nakayama, Maki Uchimura, Yuma Sonoda, Hisatomo Kowa, Takashi Nagamatsu
ETRA8
2024 Effects of Listening Behaviors of a Social Robot on Adult's Motivation and Performance in Piano Practice
abstract
The landscape of education with social robots is evolving, especially within the realm of music education. However, past studies have focused on children as music learners and such verbal behaviors as praise. Therefore, it remains unknown whether existing studies are effective for adult learners in musical education as well as how to effectively design the non-verbal behaviors of robots. This study investigates the effective behaviors of social robots by comparing three kinds of listening behaviors: none, nodding (simple listening), and enjoying (affective listening). We developed a music education support system that consists of a social robot and a MIDI keyboard and conducted an experiment with adult participants. Our experimental results described the advantages of affective-listening behavior during music education over simple listening and non-listening based on gender.
Ryuto Matsusaka, Masahiro Shiomi, Tetsuya Takiguchi
RO-MAN3
2023 Zero-Shot Sound Event Classification Using a Sound Attribute Vector with Global and Local Feature Learning
abstract
This paper introduces a zero-shot sound event classification (ZS-SEC) method to identify sound events that have never occurred in training data. In our previous work, we proposed a ZS-SEC method using sound attribute vectors (SAVs), where a deep neural network model infers attribute information that describes the sound of an event class instead of inferring its class label directly. Our previous method showed that it could classify unseen events to some extent; however, the accuracy for unseen events was far inferior to that for seen events. In this paper, we propose a new ZS-SEC method that can learn discriminative global features and local features simultaneously to enhance SAV-based ZS-SEC. In the proposed method, while the global features are learned in order to discriminate the event classes in the training data, the spectro-temporal local features are learned in order to regress the attribute information using attribute prototypes. The experimental results show that our proposed method can improve the accuracy of SAV-based ZS-SEC and can visualize the region in the spectrogram related to each attribute.
Xunquan Chen, Ryoichi Takashima, Tetsuya Takiguchi
ICASSP4
2023 Harmonic-Net: Fundamental Frequency and Speech Rate Controllable Fast Neural Vocoder
abstract
There is a need to improve the synthesis quality of HiFi-GAN-based real-time neural speech waveform generative models on CPUs while preserving the controllability of fundamental frequency ($f_{\mathrm{o}}$) and speech rate (SR). For this purpose, we propose Harmonic-Net and Harmonic-Net+, which introduce two extended functions into the HiFi-GAN generator. The first extension is a downsampling network, named the excitation signal network, that hierarchically receives multi-channel excitation signals corresponding to$f_{\mathrm{o}}$. The second extension is the layerwise pitch-dependent dilated convolutional network (LW-PDCNN), which can flexibly change its receptive fields depending on the input$f_{\mathrm{o}}$to handle large fluctuations in$f_{\mathrm{o}}$for the upsampling-based HiFi-GAN generator. The proposed explicit input of excitation signals and LW-PDCNNs corresponding to$f_{\mathrm{o}}$are expected to realize high-quality synthesis for the normal and$f_{\mathrm{o}}$-conversion conditions and for the SR-conversion condition. The results of experiments for unseen speaker synthesis, full-band singing voice synthesis, and text-to-speech synthesis show that the proposed method with harmonic waves corresponding to$f_{\mathrm{o}}$can achieve higher synthesis quality than conventional methods in all (i.e., normal,$f_{\mathrm{o}}$-conversion, and SR-conversion) conditions.
Keisuke Matsubara, Takuma Okamoto, Ryoichi Takashima, Tetsuya Takiguchi, Tomoki Toda, Hisashi Kawai
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Speaker-Independent Emotional Voice Conversion via Disentangled Representations
abstract
Emotional Voice Conversion (EVC) technology aims to transfer emotional state in speech while keeping the linguistic information and speaker identity unchanged. Prior studies on EVC have been limited to perform the conversion for a specific speaker or a predefined set of multiple speakers seen in the training stage. When encountering arbitrary speakers that may be unseen during training (outside the set of speakers used in training), existing EVC methods have limited conversion capabilities. However, converting the emotion of arbitrary speakers, even those unseen during the training procedure, in one model is much more challenging and much more attractive in real-world scenarios. To address this problem, in this study, we propose SIEVC, a novel speaker-independent emotional voice conversion framework for arbitrary speakers via disentangled representation learning. The proposed method employs the autoencoder framework to disentangle the emotion information and emotion-independent information of each input speech into separated representation spaces. To achieve better disentanglement, we incorporate mutual information minimization into the training process. In addition, adversarial training is applied to enhance the quality of the generated audio signals. Finally, speaker-independent EVC for arbitrary speakers could be achieved by only replacing the emotion representations of source speech with the target ones. The experimental results demonstrate that the proposed EVC model outperforms the baseline models in terms of objective and subjective evaluation for both seen and unseen speakers.
Xunquan Chen, Xuexin Xu, Zhihong Zhang 0001, Tetsuya Takiguchi, Edwin R. Hancock
IEEE Trans. Multim.5
2022 Speaker-Targeted Audio-Visual Speech Recognition Using a Hybrid CTC/Attention Model with Interference Loss
abstract
Audio-visual (AV)-automatic speech recognition (ASR) can improve speech recognition accuracy by using lip images, especially in noisy environments. The recently proposed AV Align system integrates speech and image features based on a cross-modal attention mechanism, where attention weights for visual features are estimated by using acoustic features as queries. Although AV Align shows an improvement in recognition accuracy in background noise environments, we have observed that the recognition accuracy degrades significantly in interference speaker environments, where a target speech and an interfering speech overlap each other. In order to improve the speech recognition accuracy of the target speaker in such situations, we propose a method that combines the auxiliary loss function that maximizes the recognition accuracy of the interference speaker and the CTC loss function for training the AV-ASR model. The experimental results using the TCD-TIMIT dataset show that the use of these auxiliary loss functions improves the performance of target-speaker speech recognition in interference speaker environments.
Ryota Tsunoda, Ryo Aihara, Ryoichi Takashima, Tetsuya Takiguchi, Yoshie Imai
ICASSP4
2022 Where Do Humans Build Levees? A Case Study on the Contiguous United States
abstract
Understanding where and why human build levees offers several values: From a hydrological perspective, integration of levees to global flood models has been shown to improve their accuracy. From an Economic intelligence perspective, levee locations provide precious insights into past and future urban developments. However, very little data exists on the location of levees at a global scale, which hinders our ability to reach a global understanding of this question. One rare exception is the National Levee Database (NLD) dataset provided by the U.S. Army Corps of Engineers (USACE). In this study, we hypothesize that levees are built at locations where human activity and flood risk coexist, and develop predictive models that output the probability of levee existence at the hydrological catchment level. Quantitative analysis of these models using the NLD dataset allows to validate our hypothesis, with several important nuances, which we discuss at length.
M. Ikegawa, Tristan Hascoet, Victor Pellet, Tetsuya Takiguchi, D. Yamazaki
IGARSS7
2022 Building a Knowledge-Based Dialogue System with Text Infilling
abstract
In recent years, generation-based dialogue systems using state-of-the-art (SoTA) transformerbased models have demonstrated impressive performance in simulating human-like conversations.To improve the coherence and knowledge utilization capabilities of dialogue systems, knowledge-based dialogue systems integrate retrieved graph knowledge into transformer-based models.However, knowledge-based dialog systems sometimes generate responses without using the retrieved knowledge.In this work, we propose a method in which the knowledge-based dialogue system can constantly utilize the retrieved knowledge using text infilling.Text infilling is the task of predicting missing spans of a sentence or paragraph.We utilize this text infilling to enable dialog systems to fill incomplete responses with the retrieved knowledge.Our proposed dialogue system has been proven to generate significantly more correct responses than baseline dialogue systems.
Tetsuya Takiguchi, Yasuo Ariki
SIGDIAL2
2022 Direction of arrival estimation for indoor environments based on acoustic composition model with a single microphone
Xingchen Guo, Xuexin Xu, Xunquan Chen, Rong Jia, Zhihong Zhang 0001, Tetsuya Takiguchi, Edwin R. Hancock
Pattern Recognit.7
2021 High-Intelligibility Speech Synthesis for Dysarthric Speakers with LPCNet-Based TTS and CycleVAE-Based VC
abstract
This paper presents a high-intelligibility speech synthesis method for persons with dysarthria caused by athetoid cerebral palsy. The muscular control of such speakers is unstable because of their athetoid symptoms, and their pronunciation is unclear, which makes it difficult for them to communicate. In this paper, we present a method for generating highly intelligible speech that preserves the individuality of dysarthric speakers by combining Transformer-TTS, CycleVAE-VC, and a LPCNet vocoder. Rather than repairing prosody from the dysarthric speech, this method transfers the dysarthric speaker’s individuality to the speech of a healthy person generated by TTS synthesis. This task is both important and challenging. From the results of our evaluation experiments, we confirmed that the proposed method can partially transfer the individuality of the target dysarthric speaker while maintaining the intelligibility of the source speech.
Keisuke Matsubara, Takuma Okamoto, Ryoichi Takashima, Tetsuya Takiguchi, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai
ICASSP4
2021 Multimodal fusion for indoor sound source localization
Ryoichi Takashima, Xingchen Guo, Zhihong Zhang 0001, Xuexin Xu, Tetsuya Takiguchi, Edwin R. Hancock
Pattern Recognit.6
2020 FasterRCNN Monitoring of Road Damages: Competition and Deployment
abstract
Maintaining aging infrastructure is a challenge currently faced by local and national administrators all around the world. An important prerequisite for efficient infrastructure maintenance is to continuously monitor (i.e., quantify the level of safety and reliability) the state of very large structures. Meanwhile, computer vision has made impressive strides in recent years, mainly due to successful applications of deep learning models. These novel progresses are allowing the automation of vision tasks, which were previously impossible to automate, offering promising possibilities to assist administrators in optimizing their infrastructure maintenance operations. In this context, the IEEE 2020 global Road Damage Detection (RDD) Challenge is giving an opportunity for deep learning and computer vision researchers to get involved and help accurately track pavement damages on road networks. This paper proposes two contributions to that topic: In a first part, we detail our solution to the RDD Challenge. In a second part, we present our efforts in deploying our model on a local road network, explaining the proposed methodology and encountered challenges.
Tristan Hascoet, Andreas Persch, Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki
IEEE BigData5
2020 Two-Step Acoustic Model Adaptation for Dysarthric Speech Recognition
abstract
This paper introduces a model adaptation approach for a speaker-dependent dysarthric speech recognition system. The dysarthria we focus on in this paper is caused by athetoid cerebral palsy, which causes involuntary muscle movements in those with the disease. For this reason, the dysarthric people's speech is often unstable and difficult for conventional automatic speech recognition (ASR) systems to recognize. A model-adaptation approach, which adapts an ASR model to dysarthric speech, is one possible solution. However, because the difference in speaking styles between dysarthric and non-dysarthric people is so significant, the conventional adaptation method is not able to sufficiently adapt the model to the dysarthric speech. In our proposed two-step model-adaptation approach, an ASR model is first adapted to the general speaking style of multiple dysarthric speakers, and then the adapted model is further adapted for the target speaker. From our experiments on an ASR task, our two-step adaptation approach showed better performance than a conventional one-step adaptation approach.
Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki
ICASSP2
2020 Dysarthric Speech Recognition Based on Deep Metric Learning
Yuki Takashima, Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki
INTERSPEECH3
2019 On Zero-Shot Recognition of Generic Objects
abstract
Many recent advances in computer vision are the results of a healthy competition among researchers on high quality, task-specific, benchmarks. After a decade of active research, zero-shot learning (ZSL) models accuracy on the Imagenet benchmark remains far too low to be considered for practical object recognition applications. In this paper, we argue that the main reason behind this apparent lack of progress is the poor quality of this benchmark. We highlight major structural flaws of the current benchmark and analyze different factors impacting the accuracy of ZSL models. We show that the actual classification accuracy of existing ZSL models is significantly higher than was previously thought as we account for these flaws. We then introduce the notion of structural bias specific to ZSL datasets. We discuss how the presence of this new form of bias allows for a trivial solution to the standard benchmark and conclude on the need for a new benchmark. We then detail the semi-automated construction of a new benchmark to address these flaws.
Tristan Hascoet, Yasuo Ariki, Tetsuya Takiguchi
CVPR3
2019 End-to-end Dysarthric Speech Recognition Using Multiple Databases
abstract
We present in this paper an end-to-end automatic speech recognition (ASR) system for a person with an articulation disorder resulting from athetoid cerebral palsy. In the case of a person with this type of articulation disorder, the speech style is quite different from that of a physically unimpaired person, and the amount of their speech data available to train the model is limited because their burden is large due to strain on the speech muscles. Therefore, the performance of ASR systems for people with an articulation disorder degrades significantly. In this paper, we propose an end-to-end ASR framework trained by not only the speech data of a Japanese person with an articulation disorder but also the speech data of a physically unimpaired Japanese person and a non-Japanese person with an articulation disorder to relieve the lack of training data of a target speaker. An end-to-end ASR model encapsulates an acoustic and language model jointly. In our proposed model, an acoustic model portion is shared between persons with dysarthria, and a language model portion is assigned to each language regardless of dysarthria. Experimental results show the merit of our proposed approach of using multiple databases for speech recognition.
Yuki Takashima, Tetsuya Takiguchi, Yasuo Ariki
ICASSP2
2019 Emotional Voice Conversion Using Dual Supervised Adversarial Networks With Continuous Wavelet Transform F0 Features
abstract
In emotional voice conversion (VC) tasks, it is difficult to deal with a simple representation of fundamental frequency (F0), which is the most important feature in emotional voice representation. In order to address this issue, we propose the adaptive scales continuous wavelet transform (ADS-CWT) method to systematically capture F0 features of different temporal levels, which can represent different prosodic aspects, ranging from micro-prosody to sentences. Moreover, in an emotional VC task, each dataset is paired with the labeled emotional voice and neutral voice, which can be regarded as a dual task. Owing to, first, dual supervised learning's ability to improve the training performances by using the leveraging probabilistic connection between the dual tasks to enhance the learning from labeled data and, second, generative adversarial networks' (GANs') ability to mitigate the over-smoothing problem caused in the low-level data space when converting the acoustic features, we further present a novel training framework for emotional VC using GANs combined with dual supervised learning, named as dual supervised adversarial networks. In emotional VC experiments, we confirmed the high similarity performance of our method when using limited labeled data for emotional VC. Our method achieves good and consistent performance, in both objective and subjective evaluations.
Zhaojie Luo, Tetsuya Takiguchi, Yasuo Ariki
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 Polar Transformation on Image Features for Orientation-Invariant Representations
abstract
The choice of image feature representation plays a crucial role in the analysis of visual information. Although vast numbers of alternative robust feature representation models have been proposed to improve the performance of different visual tasks, most existing feature representations [e.g., handcrafted features or convolutional neural networks (CNNs)] have a relatively limited capacity to capture the highly orientation-invariant (rotation/reversal) features. The net consequence is suboptimal visual performance. To address these problems, this study adopts a novel transformational approach, which investigates the potential of using polar feature representations. Our low level consists of a histogram of oriented gradient, which is then binned using annular spatial bin-type cells applied to the polar gradient. This gives gradient binning invariance for feature extraction. In this way, the descriptors have significantly enhanced orientation-invariant capabilities. The proposed feature representation, calledorientation-invariant histograms of oriented gradients, is capable of accurately processing visual tasks (e.g., facial expression recognition). In the context of the CNN architecture, we propose two polar convolution operations, referred to as full polar convolution and local polar convolution, and use these to develop polar architectures for the CNN orientation-invariant representation. Experimental results show that the proposed orientation-invariant image representation, based on polar models for both handcrafted features and deep learning features, is both competitive with state-of-the-art methods and maintains compact representation on a set of challenging benchmark image datasets.
Zhaojie Luo, Zhihong Zhang 0001, Faliang Huang, Zhiling Ye, Tetsuya Takiguchi, Edwin R. Hancock
IEEE Trans. Multim.6
2018 Parallel-Data-Free Dictionary Learning for Voice Conversion Using Non-Negative Tucker Decomposition
abstract
Voice conversion (VC) is a technique where only speaker-specific information in source speech is converted while preserving the associated phonological information. Nonnegative Matrix Factorization (NMF)-based VC has been researched because of the natural-sounding voice it produces compared with conventional Gaussian Mixture Model-based VC. In conventional NMF- VC, parallel data are used to train the models; therefore, unnatural pre-processing of speech data to make parallel data is needed. NMF-VC also tends to be a large model because this method has many parallel exemplars for the dictionary matrix; therefore, the computational cost is high. In this paper, we propose a novel parallel dictionary learning method using non-negative Tucker decomposition (NTD) which uses tensor decomposition and decomposes an input observation into a set of mode matrices and one core tensor. Our proposed NTD-based dictionary learning method estimates the dictionary matrix for NMF- VC without using parallel data. Experimental results show that our proposed method outperforms conventional non-parallel VC methods.
Yuki Takashima, Hajime Yano, Toru Nakashika, Tetsuya Takiguchi, Yasuo Ariki
ICASSP4
2018 Oil Price Forecasting Using Supervised GANs with Continuous Wavelet Transform Features
abstract
This paper proposes a novel approach based on a supervised Generative Adversarial Networks (GANs) model that forecasts the crude oil prices with Adaptive Scales Continuous Wavelet Transform (AS-CWT). In our study, we first confirmed that the possibility of using Continuous Wavelet Transform (CWT) to decompose an oil price series into various components, such as the sequence of days, weeks, months and years, so that the decomposed new time series can be used as inputs for a deep-learning (DL) training model. Second, we find that applying the proposed adaptive scales in the CWT method can strengthen the dependence of inputs and provide more useful information, which can improve the forecasting performance. Finally, we use the supervised GANs model as a training model, which can provide more accurate forecasts than those of the naive forecast (NF) model and other nonlinear models, such as Neural Networks (NNs), and Deep Belief Networks (DBNs) when dealing with a limited amount of oil prices data.
Zhaojie Luo, Xiao Jing Cai, Katsuyuki Tanaka, Tetsuya Takiguchi, Takuji Kinkyo, Shigeyuki Hamori
ICPR5
2018 Sound Recovery Considering the Vibration Direction of an Object in a Video
abstract
When a sound hits an object, it causes the surface of that object to vibrate. Some research has been carried out on the recovering of sounds by extracting the vibrations that have been recorded on high-speed videos. This research is expected to be applied in the field of surveillance and security because sounds can be recorded from far away. The vibration of objects due to sound is so fast and minute that it is invisible, but it is possible to observe the changes in objects as the movement of each pixel by using the high-speed video. In this paper, we propose a sound-recovery method focusing on the vibration direction of the object. It is considered that a better sound can be obtained by trying to recover the sound based on the direction of the largest vibration. First, the sound is recovered in a certain direction (initial direction). Next, the sound is recovered in the vibration direction that has the largest correlation with the pre-first recovered sound. Repeating this process, the sound can be recovered in the direction of the largest vibration. We recovered sounds from several objects in videos and ascertained the effectiveness of the method.
Yohei Fuse, Yusuke Yasumi, Tetsuya Takiguchi
ISM3
2018 Spectrum Enhancement of Singing Voice Using Deep Learning
abstract
In this paper, we propose a novel singing-voice enhancement system that makes the singing voice of amateurs similar to that of professional opera singers, where the singing voice of amateurs is emphasized by using a singing voice of a professional opera singer on a frequency band that represents the remarkable characteristic of the professional singer. Moreover, our proposed singing-voice enhancement based on highway networks is able to convert any song (that a professional opera singer does not sing). As a result of our experiments, the singing voice of the amateur singer at the middle-high frequency range which contains a lot of frequency components that affect glossiness was emphasized while maintaining speaker characteristics.
Ryuka Nanzaka, Tsuyoshi Kitamura, Tetsuya Takiguchi, Yuji Adachi, Kiyoto Tai
ISM3
2017 A Bayesian nonparametric multimodal data modeling framework for video emotion recognition
abstract
Video emotion recognition as an emerging research field has been attracting more and more focus in recent years. However, such work is quite challenging, since human emotions are hard to differentiate precisely due to its complexity and diversity, moreover, the expressions of sentiment in a content-rich video are sparse. Previous studies presented a number of approaches to try to learn human emotions on video level by exploiting various video features. However, most of works just used simple low-level video features such as hand-crafted image features, and they also did not consider the further latent connections among different multimodal data within a video. To tackle these problems, we develop a novel Bayesian non-parametric multimodal data modeling framework to learn the emotions from video, where the adopted image data are deep features extracted from key frames of video via convolutional neural networks (CNNs), and the adopted audio data are Mel-frequency cepstral coefficient (MFCC) features. In this framework, we then use a symmetric correspondence hierarchical Dirichlet processes (Sym-cHDP) model to mine their latent emotional events (topics) between image features and audio features. Finally, the effectiveness of our framework is demonstrated via comprehensive experimentations.
Zhaojie Luo, Koji Eguchi, Tetsuya Takiguchi, Tsukasa Omoto
ICME4
2017 Phoneme-Discriminative Features for Dysarthric Speech Conversion
Ryo Aihara, Tetsuya Takiguchi, Yasuo Ariki
INTERSPEECH2
2017 Emotional Voice Conversion with Adaptive Scales F0 Based on Wavelet Transform Using Limited Amount of Emotional Data
Zhaojie Luo, Tetsuya Takiguchi, Yasuo Ariki
INTERSPEECH3
2016 Selection of an optimum random matrix using a genetic algorithm for acoustic feature extraction
abstract
This paper describes a selection technique of an optimum random matrix using a genetic algorithm for speech recognition based on random projections. Random projections have been suggested as a means of dimensionality reduction, where the original data are projected onto a subspace using a random matrix. Moreover, as we are able to produce various random matrices, it may be possible to find a transform matrix that is superior to conventional transformation matrices among random matrices. In this paper, a genetic algorithm is introduced to find an optimum random matrix. Its effectiveness is confirmed by word recognition experiments.
Yuichiro Kataoka, Toru Nakashika, Ryo Aihara, Tetsuya Takiguchi, Yasuo Ariki
ICIS4
2016 Lip reading using a dynamic feature of lip images and convolutional neural networks
abstract
In this paper, a lip-reading method using a novel dynamic feature of lip images is proposed. The dynamic feature of lip images is calculated as the first-order regression coefficients using a few neighboring frames (images). It constiutes a better representation of the time derivatives to the basic static image. The dynamic feature is processed by using convolution neural networks (CNNs), which are able to reduce the negative influence caused by shaking of the subject and face alignment blurring at the feature-extraction level. Its effectiveness has been confirmed by word-recognition experiments comparing the proposed method with the conventional static (original) image.
Yuki Takashima, Tetsuya Takiguchi, Yasuo Ariki
ICIS3
2016 Emotional voice conversion using deep neural networks with MCC and F0 features
abstract
An artificial neural network is one of the most important models for training features in a voice conversion task. Typically, Neural Networks (NNs) are not effective in processing low-dimensional F0 features, thus this causes that the performance of those methods based on neural networks for training Mel Cepstral Coefficients (MCC) are not outstanding. However, F0 can robustly represent various prosody signals (e.g., emotional prosody). In this study, we propose an effective method based on the NNs to train the normalized-segment-F0 features (NSF0) for emotional prosody conversion. Meanwhile, the proposed method adopts deep belief networks (DBNs) to train spectrum features for voice conversion. By using these approaches, the proposed method can change the spectrum and the prosody for the emotional voice at the same time. Moreover, the experimental results show that the proposed method outperforms other state-of-the-art methods for voice emotional conversion.
Zhaojie Luo, Tetsuya Takiguchi, Yasuo Ariki
ICIS2
2016 Semi-non-negative matrix factorization using alternating direction method of multipliers for voice conversion
abstract
Voice conversion (VC) is being widely researched in the field of speech processing because of increased interest in using such processing in applications such as personalized Text-To-Speech systems. A VC method using Non-negative Matrix Factorization (NMF) has been researched because of its natural sounding voice, however, huge memory usage and high computational times have been reported as problems. We present in this paper a new VC method using Semi-Non-negative Matrix Factorization (Semi-NMF) using the Alternating Direction Method of Multipliers (ADMM) in order to tackle the problems associated with NMF-based VC. Dictionary learning using Semi-NMF can create a compact dictionary, and ADMM enables faster convergence than conventional Semi-NMF. Experimental results show that our proposed method is 76 times faster than conventional NMF, and its conversion quality is almost the same as that of the conventional method.
Ryo Aihara, Tetsuya Takiguchi, Yasuo Ariki
ICASSP2
2016 Modeling deep bidirectional relationships for image classification and generation
abstract
This paper presents a novel probabilistic model that represents a joint probability of two visible variables with a deep architecture, called a deep relational model (DRM). The model stacks several layers from one visible layer on to another visible layer, sandwiching hidden layers between them. As with restricted Boltzmann machines (RBMs) and deep Boltzmann machines (DBMs), all connections (weights) between two adjacent layers are undirected. During the maximum-likelihood (ML)-based training, the network attempts to capture latent complex relationships between two visible variables (e.g., an image showing a certain number and its corresponding label) thanks to its deep architecture. Unlike deep neural networks, 1) the proposed DRM is a totally generative model, and 2) the weights can be optimized in a probabilistic manner. This paper presents and discusses the experiments conduced to evaluate our DRM's performance in recognition and generation tasks.
Toru Nakashika, Tetsuya Takiguchi, Yasuo Ariki
ICASSP2
2016 Parallel Dictionary Learning for Voice Conversion Using Discriminative Graph-Embedded Non-Negative Matrix Factorization
Ryo Aihara, Tetsuya Takiguchi, Yasuo Ariki
INTERSPEECH2
2016 Audio-Visual Speech Recognition Using Bimodal-Trained Bottleneck Features for a Person with Severe Hearing Loss
Yuki Takashima, Ryo Aihara, Tetsuya Takiguchi, Yasuo Ariki, Nobuyuki Mitani, Kiyohiro Omori, Kaoru Nakazono
INTERSPEECH3
2016 Multiple Non-Negative Matrix Factorization for Many-to-Many Voice Conversion
abstract
A novel voice conversion (VC) method for arbitrary speakers is proposed. Non-negative matrix factorization (NMF) has recently been applied to exemplar-based VC. It offers noise robustness and naturalness of the converted voice, compared with widely used Gaussian mixture model-based VC. However, because NMF-based VC requires parallel training data from source and target speakers, the voice of arbitrary speakers cannot be converted in this framework. In this study, we propose the multiple non-negative matrix factorization (Multi-NMF) to allow the implementation of many-to-many, exemplar-based VC. Our experimental results demonstrate that the conversion quality of the proposed method is close to that of conventional one-to-one VC, even though the proposed method requires neither the source speakers' spectra, nor the target speakers' spectra, to be included in the training set.
Ryo Aihara, Tetsuya Takiguchi, Yasuo Ariki
IEEE ACM Trans. Audio Speech Lang. Process.2
2016 Non-Parallel Training in Voice Conversion Using an Adaptive Restricted Boltzmann Machine
abstract
In this paper, we present a voice conversion (VC) method that does not use any parallel data while training the model. VC is a technique where only speaker-specific information in source speech is converted while keeping the phonological information unchanged. Most of the existing VC methods rely on parallel data-pairs of speech data from the source and target speakers uttering the same sentences. However, the use of parallel data in training causes several problems: 1) the data used for the training are limited to the predefined sentences, 2) the trained model is only applied to the speaker pair used in the training, and 3) mismatches in alignment may occur. Although it is, thus, fairly preferable in VC not to use parallel data, a nonparallel approach is considered difficult to learn. In our approach, we achieve nonparallel training based on a speaker adaptation technique and capturing latent phonological information. This approach assumes that speech signals are produced from a restricted Boltzmann machine-based probabilistic model, where phonological information and speaker-related information are defined explicitly. Speaker-independent and speaker-dependent parameters are simultaneously trained under speaker adaptive training. In the conversion stage, a given speech signal is decomposed into phonological and speaker-related information, the speaker-related information is replaced with that of the desired speaker, and then voice-converted speech is obtained by mixing the two. Our experimental results showed that our approach outperformed another nonparallel approach, and produced results similar to those of the popular conventional Gaussian mixture models-based method that used parallel data in subjective and objective criteria.
Toru Nakashika, Tetsuya Takiguchi, Yasuhiro Minami
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 Facial expression recognition with multithreaded cascade of rotation-invariant HOG
abstract
We propose a novel and general framework, named the multithreading cascade of rotation-invariant histograms of oriented gradients (McRiHOG) for facial expression recognition (FER). In this paper, we attempt to solve two problems about high-quality local feature descriptors and robust classifying algorithm for FER. The first solution is that we adopt annular spatial bins type HOG (Histograms of Oriented Gradients) descriptors to describe local patches. In this way, it significantly enhances the descriptors in regard to rotation-invariant ability and feature description accuracy; The second one is that we use a novel multithreading cascade to simultaneously learn multiclass data. Multithreading cascade is implemented through non-interfering boosting channels, which are respectively built to train weak classifiers for each expression. The superiority of McRiHOG over current state-of-the-art methods is clearly demonstrated by evaluation experiments based on three popular public databases, CK+, MMI, and AFEW.
Tetsuya Takiguchi, Yasuo Ariki
ACII2
2015 Activity-mapping non-negative matrix factorization for exemplar-based voice conversion
abstract
Voice conversion (VC) is being widely researched in the field of speech processing because of increased interest in using such processing in applications such as personalized Text-To-Speech systems. We present in this paper an exemplar-based VC method us- ing Non-negative Matrix Factorization (NMF), which is different from conventional statistical VC. In our previous exemplar-based VC method, input speech is represented by the source dictionary and its sparse coefficients. The source and the target dictionaries are fully coupled and the converted voice is constructed from the source coefficients and the target dictionary. In this paper, we propose an Activity-mapping NMF approach and introduce mapping matrices between source and target sparse coefficients. The effectiveness of this method was confirmed by comparing its effectiveness with that of a conventional Gaussian Mixture Model (GMM)-based method and a conventional NMF-based method.
Ryo Aihara, Tetsuya Takiguchi, Yasuo Ariki
ICASSP2
2015 Multithreading AdaBoost framework for object recognition
abstract
Our research focuses on the study of effective feature description and robust classifier technique, proposing a novel learning framework, which is capable of processing multiclass objects recognition simultaneously and accurately. The framework adopts rotation-invariant histograms of oriented gradients (Ri-HOG) as feature descriptors. Most of the existing HOG techniques are computed on a dense grid of uniformly-spaced cells and use overlapping local contrast of rectangular blocks for normalization. However, we adopt annular spatial bins type cells and apply the radial gradient to attain gradient binning invariance for feature extraction. In this way, it significantly enhances HOG in regard to rotation-invariant ability and feature description accuracy; The classifier is derived from AdaBoost algorithm, but it is ameliorated and implemented through non-interfering boosting channels, which are respectively built to train weak classifiers for each object category. In this way, the boosting cascade can allow the weak classifier to be trained to fit complex distributions. The proposed method is valid on PASCAL VOC 2007 database and it achieves the state-of-the-arts performance.
Tetsuya Takiguchi, Yasuo Ariki
ICIP2
2015 Sparse nonlinear representation for voice conversion
abstract
In voice conversion, sparse-representation-based methods have recently been garnering attention because they are, relatively speaking, not affected by over-fitting or over-smoothing problems. In these approaches, voice conversion is achieved by estimating a sparse vector that determines which dictionaries of the target speaker should be used, calculated from the matching of the input vector and dictionaries of the source speaker. The sparse-representation-based voice conversion methods can be broadly divided into two approaches: 1) an approach that uses raw acoustic features in the training data as parallel dictionaries, and 2) an approach that trains parallel dictionaries from the training data. In our approach, we follow the latter approach and systematically estimate the parallel dictionaries using a joint-density restricted Boltzmann machine with sparse constraints. Through voice-conversion experiments, we confirmed the high-performance of our method, comparing it with the conventional Gaussian mixture model (GMM)-based approach, and a non-negative matrix factorization (NMF)-based approach, which is based on sparse representation.
Toru Nakashika, Tetsuya Takiguchi, Yasuo Ariki
ICME2
2015 Individuality-Preserving Voice Reconstruction for Articulation Disorders Using Text-to-Speech Synthesis
abstract
This paper presents a speech synthesis method for people with articulation disorders. Because the movements of such speakers are limited by their athetoid symptoms, their prosody is often unstable and their speech rate differs from that of a physically unimpaired person, which causes their speech to be less intelligible and, consequently, makes communication with physically unimpaired persons difficult. In order to deal with these problems, this paper describes a Hidden Markov Model(HMM)-based text-to-speech synthesis approach that preserves the individuality of a person with an articulation disorder and aids them in their communication. In our method, a duration model of a physically unimpaired person is used for the HMM synthesis system and an F0 model in the system is trained using the F0 patterns of the physically unimpaired person, with the average F0 being converted to the target F0 in advance. In order to preserve the target speaker's individuality, a spectral model is built from target spectra. Through experimental evaluations, we have confirmed that the proposed method successfully synthesizes intelligible speech while maintaining the target speaker's individuality.
Reina Ueda, Tetsuya Takiguchi, Yasuo Ariki
ICMI2
2015 Word-Error Correction of Continuous Speech Recognition Based on Normalized Relevance Distance
Yohei Fusayasu, Katsuyuki Tanaka, Tetsuya Takiguchi, Yasuo Ariki
IJCAI3
2015 Many-to-many voice conversion based on multiple non-negative matrix factorization
abstract
We present in this paper an exemplar-based Voice Conversion (VC) method using Non-negative Matrix Factorization (NMF), which is different from conventional statistical VC. NMF-based VC has advantages of noise robustness and naturalness of converted voice compared to Gaussian Mixture Model (GMM)based VC. However, because NMF-based VC is based on parallel training data of source and target speakers, we cannot convert the voice of arbitrary speakers in this framework. In this paper, we propose a many-to-many VC method that makes use of Multiple Non-negative Matrix Factorization (Multi-NMF). By using Multi-NMF, an arbitrary speaker’s voice is converted to another arbitrary speaker’s voice without the need for any input or output speaker training data. We assume that this method is flexible because we can adopt it to voice quality control or noise robust VC. Index Terms: voice conversion, speech synthesis, many-tomany, exemplar-based, NMF
Ryo Aihara, Tetsuya Takiguchi, Yasuo Ariki
INTERSPEECH2
2015 Content-based Image Retrieval Using Rotation-invariant Histograms of Oriented Gradients
abstract
Our research focuses on the question of feature descriptors for robust effective computing, proposing a novel feature representation method, namely, rotation-invariant histograms of oriented gradients (Ri-HOG) for image retrieval. Most of the existing HOG techniques are computed on a dense grid of uniformly-spaced cells and use overlapping local contrast of rectangular blocks for normalization. However, we adopt annular spatial bins type cells and apply radial gradient to attain gradient binning invariance for feature extraction. In this way, it significantly enhances HOG in regard to rotation-invariant ability and feature descripting accuracy. In experiments, the proposed method is evaluated on Corel-5k and Corel-10k datasets. The experimental results demonstrate that the proposed method is much more effective than many existing image feature descriptors for content-based image retrieval.
Toru Nakashika, Tetsuya Takiguchi, Yasuo Ariki
ICMR3
2015 Voice Conversion Using RNN Pre-Trained by Recurrent Temporal Restricted Boltzmann Machines
abstract
This paper presents a voice conversion (VC) method that utilizes the recently proposed probabilistic models called recurrent temporal restricted Boltzmann machines (RTRBMs). One RTRBM is used for each speaker, with the goal of capturing high-order temporal dependencies in an acoustic sequence. Our algorithm starts from the separate training of one RTRBM for a source speaker and another for a target speaker using speaker-dependent training data. Because each RTRBM attempts to discover abstractions to maximally express the training data at each time step, as well as the temporal dependencies in the training data, we expect that the models represent the linguistic-related latent features in high-order spaces. In our approach, we convert (match) features of emphasis for the source speaker to those of the target speaker using a neural network (NN), so that the entire network (consisting of the two RTRBMs and the NN) acts as a deep recurrent NN and can be fine-tuned. Using VC experiments, we confirm the high performance of our method, especially in terms of objective criteria, relative to conventional VC methods such as approaches based on Gaussian mixture models and on NNs.
Toru Nakashika, Tetsuya Takiguchi, Yasuo Ariki
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 Voice conversion based on Non-negative matrix factorization using phoneme-categorized dictionary
abstract
We present in this paper an exemplar-based voice conversion (VC) method using a phoneme-categorized dictionary. Sparse representation-based VC using Non-negative matrix factorization (NMF) is employed for spectral conversion between different speakers. In our previous NMF-based VC method, source exemplars and target exemplars are extracted from parallel training data, having the same texts uttered by the source and target speakers. The input source signal is represented using the source exemplars and their weights. Then, the converted speech is constructed from the target exemplars and the weights related to the source exemplars. However, this exemplar-based approach needs to hold all the training exemplars (frames), and it may cause mismatching of phonemes between input signals and selected exemplars. In this paper, in order to reduce the mismatching of phoneme alignment, we propose a phoneme-categorized sub-dictionary and a dictionary selection method using NMF. By using the sub-dictionary, the performance of VC is improved compared to a conventional NMF-based VC. The effectiveness of this method was confirmed by comparing its effectiveness with that of a conventional Gaussian Mixture Model (GMM)-based method and a conventional NMF-based method.
Ryo Aihara, Toru Nakashika, Tetsuya Takiguchi, Yasuo Ariki
ICASSP3
2014 Multimodal voice conversion using non-negative matrix factorization in noisy environments
abstract
This paper presents a multimodal voice conversion (VC) method for noisy environments. In our previous NMF-based VC method, source exemplars and target exemplars are extracted from parallel training data, in which the same texts are uttered by the source and target speakers. The input source signal is then decomposed into source exemplars, noise exemplars obtained from the input signal, and their weights. Then, the converted speech is constructed from the target exemplars and the weights related to the source exemplars. In this paper, we propose a multimodal VC that improves the noise robustness in our NMF-based VC method. By using the joint audio-visual features as source features, the performance of VC is improved compared to a previous audio-input NMF-based VC method. The effectiveness of this method was confirmed by comparing its effectiveness with that of a conventional Gaussian Mixture Model (GMM)-based method.
Kenta Masaka, Ryo Aihara, Tetsuya Takiguchi, Yasuo Ariki
ICASSP3
2014 Voice conversion in time-invariant speaker-independent space
abstract
In this paper, we present a voice conversion (VC) method that utilizes conditional restricted Boltzmann machines (CRBMs) for each speaker to obtain time-invariant speaker-independent spaces where voice features are converted more easily than those in an original acoustic feature space. First, we train two CRBMs for a source and target speaker independently using speaker-dependent training data (without the need to parallelize the training data). Then, a small number of parallel data are fed into each CRBM and the high-order features produced by the CRBMs are used to train a concatenating neural network (NN) between the two CRBMs. Finally, the entire network (the two CRBMs and the NN) is fine-tuned using the acoustic parallel data. Through voice-conversion experiments, we confirmed the high performance of our method in terms of objective and subjective evaluations, comparing it with conventional GMM, NN, and speaker-dependent DBN approaches.
Toru Nakashika, Tetsuya Takiguchi, Yasuo Ariki
ICASSP2
2014 3D-Object Recognition Based on LLC Using Depth Spatial Pyramid
abstract
Recently introduced high-accuracy RGB-D cameras are capable of providing high quality three-dimension information (color and depth information) easily. The overall shape of the object can be understood by acquiring depth information. However, conventional methods adopted this camera use depth information only to extract the local feature. To improve the object recognition accuracy, in our approach, the overall object shape is expressed by the depth spatial pyramid based on depth information. In more detail, multiple features within each sub-region of the depth spatial pyramid are pooled. As a result, the feature representation including the depth topological information is constructed. We use histogram of oriented normal vectors (HONV) designed to capture local geometric characteristics as 3D local features and locality-constrained linear coding (LLC) to project each descriptor into its local-coordinate system. As a result of image recognition, the proposed method has improved the recognition rate compared with conventional methods.
Toru Nakashika, Takafumi Hori, Tetsuya Takiguchi, Yasuo Ariki
ICPR3
2014 Error correction of automatic speech recognition based on normalized web distance
abstract
In this paper, we focus on the problems associated with error correction of automatic speech recognition (ASR) based on confusion networks. The problems discussed are the availability of corpus in terms of calculating the semantic score and performance degradation for error correction using N -gram due to the null transitions in the confusion networks. In attempt to solve these problems, first, we employ Normalized Web Distance as a measure for semantic similarity between words that are located far from each other. The advantage of Normalized Web Distance is that it may use the Internet and so on for learning semantic similarity, which might solve the problem of corpus availability. Secondly, an error correction model without null nodes in confusion networks is trained using conditional random fields in order to improve the performance of error correction using N -grams. Index Terms: confusion network, conditional random fields, word-error correction, normalized web distance
E. Byambakhishig, Katsuyuki Tanaka, Ryo Aihara, Toru Nakashika, Tetsuya Takiguchi, Yasuo Ariki
INTERSPEECH5
2014 Multimodal exemplar-based voice conversion using lip features in noisy environments
abstract
This paper presents a multimodal voice conversion (VC) method for noisy environments. In our previous exemplarbased VC method, source exemplars and target exemplars are extracted from parallel training data, in which the same texts are uttered by the source and target speakers. The input source signal is then decomposed into source exemplars, noise exemplars obtained from the input signal, and their weights. Then, the converted speech is constructed from the target exemplars and the weights related to the source exemplars. In this paper, we propose a multimodal VC method that improves the noise robustness of our previous exemplar-based VC method. As visual features, we use not only conventional DCT but also the features extracted from Active Appearance Model (AAM) applied to the lip area of a face image. Furthermore, we introduce the combination weight between audio and visual features and formulate a new cost function in order to estimate the audiovisual exemplars. By using the joint audio-visual features as source features, the VC performance is improved compared to a previous audio-input exemplar-based VC method. The effectiveness of this method was confirmed by comparing its effectiveness with that of a conventional Gaussian Mixture Model (GMM)-based method. Index Terms: voice conversion, multimodal, image features, non-negative matrix factorization, noise robustness
Kenta Masaka, Ryo Aihara, Tetsuya Takiguchi, Yasuo Ariki
INTERSPEECH3
2014 High-order sequence modeling using speaker-dependent recurrent temporal restricted boltzmann machines for voice conversion
abstract
This paper presents a voice conversion (VC) method that utilizes recently proposed recurrent temporal restricted Boltzmann machines (RTRBMs) for each speaker, with the goal of capturing high-order temporal dependencies in an acoustic sequence. Our algorithm starts from the separate training of two RTRBMs for a source and target speaker using speaker-dependent training data. Since each RTRBM attempts to discover abstractions at each time step, as well as the temporal dependencies in the training data, we expect that the models represent the speaker-specific latent features in the high-order spaces. In our approach, we run conversion from such speaker-specificemphasized features of the source speaker to those of the target speaker using a neural network (NN), so that the entire network (the two RTRBMs ant the NN) forms a deep recurrent neural network and can be fine-tuned. Through VC experiments, we confirmed the high performance of our method especially in terms of objective criteria in comparison to conventional VC methods such as Gaussian mixture model (GMM)-based approaches.
Toru Nakashika, Tetsuya Takiguchi, Yasuo Ariki
INTERSPEECH2
2013 Individuality-preserving voice conversion for articulation disorders based on non-negative matrix factorization
abstract
We present in this paper a voice conversion (VC) method for a person with an articulation disorder resulting from athetoid cerebral palsy. The movement of such speakers is limited by their athetoid symptoms, and their consonants are often unstable or unclear, which makes it difficult for them to communicate. In this paper, exemplar-based spectral conversion using Non-negative Matrix Factorization (NMF) is applied to a voice with an articulation disorder. To preserve the speaker's individuality, we used a combined dictionary that is constructed from the source speaker's vowels and target speaker's consonants. Experimental results indicate that the performance of NMF-based VC is considerably better than conventional GMM-based VC.
Ryo Aihara, Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki
ICASSP3
2013 Sparse representation for outliers suppression in semi-supervised image annotation
abstract
Recently, generic object recognition (automatic image annotation) that achieves human-like vision using a computer has being looked to for use in robot vision, automatic categorization of images, and retrieval of images. For the annotation, semi-supervised learning, which incorporates a large amount of unsupervised training data (unlabeled data) along with a small amount of supervised data (labeled data), is expected to be an effective tool as it reduces the burden of manual annotation. However, some unlabeled data in semi-supervised models contains outliers that negatively affect the parameter estimation on the training stage. Such outliers often cause the over-fitting problem especially when a small amount of training data is used. In this paper, we propose a practical method to prevent the over-fitting in semi-supervised learning, suppressing existing outliers by sparse representation. In our experiments we got 4 points improvement comparing conventional semi-supervised methods, SemiNB and TSVM.
Toru Nakashika, Takeshi Okumura, Tetsuya Takiguchi, Yasuo Ariki
ICASSP3
2013 Prediction of unlearned position based on local regression for single-channel talker localization using acoustic transfer function
abstract
This paper presents a sound-source (talker) localization method using only a single microphone. In our previous work, we discussed the single-channel sound-source localization method based on the discrimination of the acoustic transfer function. However, that method requires the training of the acoustic transfer function for each possible position in advance, and it is difficult to estimate the position that has not been pre-trained. In order to estimate such unlearned positions, in this paper, we discuss a single-channel talker localization method based on a regression model, which predicts the position from the acoustic transfer function. For training the regression model, we use the local regression approach, which trains the regression model from only training samples that are similar to the evaluation data. Considering both the linear and non-linear regression models, the effectiveness of this method has been confirmed by sound-source localization experiments performed in different room environments.
Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki
ICASSP2
2013 Exemplar-based individuality-preserving voice conversion for articulation disorders in noisy environments
abstract
We present in this paper a noise robust voice conversion (VC) method for a person with an articulation disorder resulting from athetoid cerebral palsy. The movements of such speakers are limited by their athetoid symptoms, and their consonants are often unstable or unclear, which makes it difficult for them to communicate. In this paper, exemplar-based spectral conversion using Non-negative Matrix Factorization (NMF) is applied to a voice with an articulation disorder in real noisy environments. In this paper, in order to deal with background noise, an input noisy source signal is decomposed into the clean source exemplars and noise exemplars by NMF. Also, to preserve the speaker’s individuality, we use a combined dictionary that was constructed from the source speaker’s vowels and target speaker’s consonants. The effectiveness of this method was confirmed by comparing its effectiveness with that of a conventional Gaussian Mixture Model (GMM)-based method. Index Terms: Voice Conversion, NMF, Articulation Disorders, Noise Robustness, Assistive Technologies
Ryo Aihara, Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki
INTERSPEECH3
2013 Voice conversion in high-order eigen space using deep belief nets
abstract
This paper presents a voice conversion technique using Deep Belief Nets (DBNs) to build high-order eigen spaces of the source/target speakers, where it is easier to convert the source speech to the target speech than in the traditional cepstrum space. DBNs have a deep architecture that automatically discovers abstractions to maximally express the original input features. If we train the DBNs using only the speech of an individual speaker, it can be considered that there is less phonological information and relatively more speaker individuality in the output features at the highest layer. Training the DBNs for a source speaker and a target speaker, we can then connect and convert the speaker individuality abstractions using Neural Networks (NNs). The converted abstraction of the source speaker is then brought back to the cepstrum space using an inverse process of the DBNs of the target speaker. We conducted speakervoice conversion experiments and confirmed the efficacy of our method with respect to subjective and objective criteria, comparing it with the conventional Gaussian Mixture Model-based method.
Toru Nakashika, Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki
INTERSPEECH3
2013 Two-step correction of speech recognition errors based on n-gram and long contextual information
abstract
This paper presents a fully automatic word error correction on a confusion network that makes use of long contextual information. However, a problem with long contextual information is that improvement of the recognition accuracy is minimal because of the word errors surrounding words. In this paper, recognition errors are first reduced by error correction using Ngram features. After that, the long-distance context scores are applied to the correction of the residual recognition errors. Index Terms: confusion network, conditional random fields, word-error correction, long contextual information
Ryohei Nakatani, Tetsuya Takiguchi, Yasuo Ariki
INTERSPEECH2
2013 Robust facial expressions recognition using 3D average face and ameliorated adaboost
abstract
One of the most crucial techniques associated with Computer Vision is technology that deals with facial recognition, especially, the automatic estimation of facial expressions. However, in real-time facial expression recognition, when a face turns sideways, the expressional feature extraction becomes difficult as the view of camera changes and recognition accuracy degrades significantly. Therefore, quite many conventional methods are proposed, which are based on static images or limited to situations in which the face is viewed from the front. In this paper, a method that uses Look-Up-Table (LUT) AdaBoost combining with the three-dimensional average face is proposed to solve the problem mentioned above. In order to evaluate the proposed method, the experiment compared with the conventional method was executed. These approaches show promising results and very good success rates. This paper covers several methods that can improve results by making the system more robust.
Yasuo Ariki, Tetsuya Takiguchi
ACM Multimedia3
2012 Generic object recognition by graph structural expression
abstract
This paper describes a method for generic object recognition using graph structural expression. In recent years, generic object recognition by computer is finding extensive use in a variety of fields, including robotic vision and image retrieval. Conventional methods use a bag-of-features (BoF) approach, which expresses the image as an appearance frequency histogram of visual words by quantizing SIFT (Scale-Invariant Feature Transform) features. However, there is a problem associated with this approach, namely that the location information and the relationship between keypoints (both of which are important as structural information) are lost. To deal with this problem, in the proposed method, the graph is constructed by connecting SIFT keypoints with lines. As a result, the keypoints maintain their relationship, and then structural representation with location information is achieved. Since graph representation is not suitable for statistical work, the graph is embedded into a vector space according to the graph edit distance. The experiment results on an image dataset of 10 classes showed that, the proposed method improved the recognition rate by 14.08%.
Takahiro Hori, Tetsuya Takiguchi, Yasuo Ariki
ICASSP2
2012 Super-resolution by GMM based conversion using self-reduction image
abstract
In recent years, super-resolution techniques in the field of computer vision have been studied actively owing to the potential applicability in various fields. In this paper, we propose a single-image, super-resolution approach using GMM (Gaussian Mixture Model)-based conversion. The conversion function is constructed by GMM using the input image and its self-reduction image. The high-resolution image is obtained by applying the conversion function to the enlarged input image without any outside database. We confirmed the effectiveness of this proposed method through the experiments.
Yuki Ogawa, Yasuo Ariki, Tetsuya Takiguchi
ICASSP3
2012 A new multiple-kernel-learning weighting method for localizing human brain magnetic activity
abstract
This paper shows that pattern classification based on machine learning is a powerful tool to analyze human brain activity data obtained by magnetoencephalography (MEG). We propose a new weighting method using a multiple kernel learning (MKL) algorithm to localize the brain area contributing to the accurate vowel discrimination. Our MKL simultaneously estimates both the classification boundary and the weight of each MEG sensor; MEG amplitude obtained from each pair of sensors is an element of the feature vector. The estimated weight indicates how the corresponding sensor is useful for classifying the MEG response patterns. Our results show both the large-weight MEG sensors mainly in a language area of the brain and the high classification accuracy (73.0%) in the 100 ~ 200 ms latency range.
Tetsuya Takiguchi, Toshiaki Imada, Ryoichi Takashima, Yasuo Ariki, Jo-Fu Lotus Lin, Patricia K. Kuhl, Masaki Kawakatsu, Makoto Kotani
ICASSP1
2012 Acoustic model transformations based on random projections
abstract
This paper proposes a novel acoustic model transformation method for speech recognition based on random projections. Random projections have been suggested as a means of dimensionality reduction, where the original data are projected onto a subspace using a random matrix. Moreover, as we are able to produce various random matrices, it may be possible to find a transform matrix that is superior to conventional transformation matrices among random matrices. In our previous work, a random-projection-based feature combination technique has been proposed but had a high computational cost. In order to deal with this cost, in this paper, we introduce random projections on the acoustic model domain, where linear transformations are applied to an acoustic model using random matrices. Its effectiveness is confirmed by word recognition experiments on noisy speech.
Tetsuya Takiguchi, Mariko Yoshii, Yasuo Ariki, Jeff A. Bilmes
ICASSP1
2012 3D tracking of soccer players using time-situation graph in monocular image sequence
Hiroki Itoh, Tetsuya Takiguchi, Yasuo Ariki
ICPR2
2012 Local-feature-map Integration Using Convolutional Neural Networks for Music Genre Classification
abstract
International audience
Toru Nakashika, Christophe Garcia, Tetsuya Takiguchi
INTERSPEECH3
2012 Estimation of Talker's Head Orientation Based on Discrimination of the Shape of Cross-power Spectrum Phase Coefficients
abstract
This paper presents a talker’s head orientation estimation method using 2-channel microphones. In recent research, some approaches based on a network of microphone arrays have been proposed in order to estimate the talker’s head orientation. In those methods, the talker’s head orientation is estimated using the sound amplitude or peak value of CSP (Cross-power Spectrum Phase) coefficients obtained from each microphone array. However, microphone array network systems need many microphone arrays to be set along the walls of a given room so that sub-microphone arrays surround the user. In this paper, we focus on the shape of the CSP coefficients affected by the reverberation, which depends on the talker’s position and the head orientation. In our proposed method, we use not only the peak value but also the other values of the CSP coefficients as feature vectors, and the talker’s position and the head orientation are estimated by discriminating the CSP vector. The effectiveness of this method has been confirmed by talker localization and head orientation estimation experiments performed in a real environment.
Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki
INTERSPEECH2
2012 Super-resolution Using GMM and PLS Regression
abstract
In recent years, super-resolution techniques in the field of computer vision have been studied in earnest owing to the potential applicability of such technology in a variety of fields. In this paper, we propose a single-image, super-resolution approach using a Gaussian Mixture Model (GMM) and Partial Least Squares (PLS) regression. A GMM-based super-resolution technique is shown to be more efficient than previously known techniques, such as sparse-coding-based techniques. But the GMM-based conversion may result in over fitting. In this paper, an effective technique for preventing over fitting, which combines PLS regression with a GMM, is proposed. The conversion function is constructed using the input image and its self-reduction image. The high-resolution image is obtained by applying the conversion function to the enlarged input image without any outside database. We confirmed the effectiveness of this proposed method through our experiments.
Yuki Ogawa, Takahiro Hori, Tetsuya Takiguchi, Yasuo Ariki
ISM3
2012 Robust AAM-based audio-visual speech recognition against face direction changes
abstract
As one of the techniques for robust speech recognition under noisy environments, audio-visual speech recognition (AVSR) using lip dynamic scene information together with audio information is attracting attention, and the research has advanced in recent years. However, in visual speech recognition (VSR), when a face turns sideways, the shape of the lip as viewed from the camera changes and the recognition accuracy degrades significantly. Therefore, many of the conventional VSR methods are limited to situations in which the face is viewed from the front. This paper proposes a VSR method to convert faces viewed from various directions into faces that are viewed from the front using Active Appearance Models (AAM). In the experiment, even when the face direction changes about 30 degrees relative to a frontal view, the recognition accuracy improved significantly.
Yuto Komai, Tetsuya Takiguchi, Yasuo Ariki
ACM Multimedia3
2012 Exemplar-based voice conversion in noisy environment
abstract
This paper presents a voice conversion (VC) technique for noisy environments, where parallel exemplars are introduced to encode the source speech signal and synthesize the target speech signal. The parallel exemplars (dictionary) consist of the source exemplars and target exemplars, having the same texts uttered by the source and target speakers. The input source signal is decomposed into the source exemplars, noise exemplars obtained from the input signal, and their weights (activities). Then, by using the weights of the source exemplars, the converted signal is constructed from the target exemplars. We carried out speaker conversion tasks using clean speech data and noise-added speech data. The effectiveness of this method was confirmed by comparing its effectiveness with that of a conventional Gaussian Mixture Model (GMM)-based method.
Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki
SLT2
2011 Generic object recognition using automatic region extraction and dimensional feature integration utilizing multiple kernel learning
abstract
Recently, in generic object recognition research, a classification technique based on integration of image features is garnering much attention. However, with a classifying technique using feature integration, there are some features that may cause incorrect recognition of objects and a large amount of noise that causes a degradation in the recognition accuracy of image data. In this paper, we propose feature selection in an object area that is restricted by removing its back ground region, and multiple kernel learning (MKL) to weight each dimension, as well as the features themselves. This enables accurate and effective weighting since the weight is computed for each dimension using the selected feature. Experimental results indicate the validity of automatic feature selection. Classification performance is improved by using a background removing technique that utilizes saliency maps and graph cuts, and each dimensional weighting method using MKL.
Toru Nakashika, Akira Suga, Tetsuya Takiguchi, Yasuo Ariki
ICASSP3
2011 Feature selection based on Multiple Kernel Learning for single-channel sound source localization using the acoustic transfer function
abstract
This paper presents a sound source (talker) localization method using only a single microphone. In our previous work [1], we discussed the single-channel sound source localization method, where the acoustic transfer function from a user's position is estimated by using a Hidden Markov Model (HMM) of clean speech in the cepstral domain. In this paper, each cepstral dimension of the acoustic transfer function is newly selected in order to select the cepstral dimensions having information that is useful for classifying the user's position. Then, we propose a feature selection method for the cepstral parameter using Multiple Kernel Learning (MKL) to define the base kernels for each cepstral dimension (scalar) of the acoustic transfer function. The user's position is trained and classified by Support Vector Machine (SVM). The effectiveness of this method has been confirmed by sound source (talker) localization experiments performed in a room environment.
Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki
ICASSP2
2011 Probabilistic Spectrum Envelope: Categorized Audio-Features Representation for NMF-Based Sound Decomposition
abstract
NMF (Non-negative Matrix Factorization) has been one of the most useful techniques for audio signal analysis in recent years. In particular, supervised NMF, in which a large number of samples is used for analyzing a signal, is garnering much attention in sound source separation or noise reduction research. However, because such methods require all the possible samples for the analysis, it is hard to build a practical system based on this method. In this paper, we propose a novel method of signal analysis that combines the NMF and probabilistic approaches. In this approach, it is assumed that each audio-source category (such as phonemes or musical instruments) has an environment-invariant feature, called a probabilistic spectrum envelope (PSE). At the start, the PSE of each category is learned using a technique based on Gaussian Process Regression. Then, the observed spectrum is analyzed using a combination of supervised NMF and Genetic Algorithm with pre-trained PSEs. Index Terms: signal analysis, source separation, non-negative matrix factorization, probabilistic spectrum envelope, Gaussian process, genetic algorithm
Toru Nakashika, Tetsuya Takiguchi, Yasuo Ariki
INTERSPEECH2
2011 Single-Channel Head Orientation Estimation Based on Discrimination of Acoustic Transfer Function
abstract
This paper presents a talker’s head orientation estimation method using only a single microphone, where phoneme HMMs (Hidden Markov Models) of clean speech are introduced to separate the acoustic transfer function at the user’s position and head orientation. The frame sequence of the acoustic transfer function is estimated by maximizing the likelihood of training data uttered from a given position with a given head orientation. Using the separated frame sequence data, the user’s position and the head orientation are trained by Support Vector Machine (SVM) in advance. Then, for each test utterance, the frame sequence of the acoustic transfer function is separated based on the maximum likelihood estimation using the label sequence obtained from the phoneme recognition, and the user’s position and head orientation are estimated by discriminating the separated acoustic transfer function using SVM. The effectiveness of this method has been confirmed by talker localization and head orientation estimation experiments performed in a real environment. Index Terms: single channel, talker localization, head orientation, acoustic transfer function
Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki
INTERSPEECH2
2011 Image Annotation with Concept Level Feature Using PLSA+CCA
Tetsuya Takiguchi, Yasuo Ariki
MMM (2)2
2011 Audio-Visual Speech Recognition Based on AAM Parameter and Phoneme Analysis of Visual Feature
Yuto Komai, Yasuo Ariki, Tetsuya Takiguchi
PSIVT (1)3
2010 HMM-based separation of acoustic transfer function for single-channel sound source localization
abstract
This paper presents a sound source (talker) localization method using only a single microphone, where a HMM (Hidden Markov Model) of clean speech is introduced to estimate the acoustic transfer function from a user's position. The new method is able to carry out this estimation without measuring impulse responses. The frame sequence of the acoustic transfer function is estimated by maximizing the likelihood of training data uttered from a given position, where the cepstral parameters are used to effectively represent useful clean speech. Using the estimated frame sequence data, the GMM (Gaussian Mixture Model) of the acoustic transfer function is created to deal with the influence of a room impulse response. Then, for each test data set, we find a maximum-likelihood GMM from among the estimated GMMs corresponding to each position. The effectiveness of this method has been confirmed by talker localization experiments performed in a room environment.
Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki
ICASSP2
2010 Evaluation of random-projection-based feature combination on speech recognition
abstract
Random projection has been suggested as a means of dimensionality reduction, where the original data are projected onto a subspace using a random matrix. It represents a computationally simple method that approximately preserves the Euclidean distance of any two points through the projection. Moreover, as we are able to produce various random matrices, there may be some possibility of finding a random matrix that gives a better speech recognition accuracy among these random matrices. In this paper, we investigate the feasibility of random projection for speech feature extraction. To obtain an optimal result from among many (infinite) random matrices, a vote-based random-projection combination is introduced in this paper, where ROVER combination is applied to random-projection-based features. Its effectiveness is confirmed by word recognition experiments.
Tetsuya Takiguchi, Jeff A. Bilmes, Mariko Yoshii, Yasuo Ariki
ICASSP1
2010 Structuring a gene network using a multiresolution independence test
abstract
In order to structure a gene network, a score-based approach is often used. A score-based approach, however, is problematic because by assuming a probability distribution, one is prevented from finding other dependent relationships with other genes. In this research, we structured a gene network from observed gene expression data using a multiresolution independence test and a conditional independence test, which is the non-parametric method proposed by Margaritis for learning the structure of Bayesian networks without making any probability distribution assumptions. The experimental results achieved an improvement in sensitivity of 0.05, and an improvement in specificity of 0.01.
Takayuki Yamamoto, Tetsuya Takiguchi, Yasuo Ariki
ICASSP2
2010 Generic Object Recognition by Tree Conditional Random Field Based on Hierarchical Segmentation
abstract
Generic object recognition by a computer is strongly required in various fields like robot vision and image retrieval in recent years. Conventional methods use Conditional Random Field (CRF) that recognizes the class of each region using the features extracted from the local regions and the class co-occurrence between the adjoining regions. However, there is a problem that the discriminative ability of the features extracted from local regions is insufficient, and these methods is not robust to the scale variance. To solve this problem, we propose a method that integrates the recognition results in multi-scales by tree conditional random field based on hierarchical segmentation. As a result of the image dataset of 7 classes, the proposed method has improved the recognition rate by 2.2%.
Takeshi Okumura, Tetsuya Takiguchi, Yasuo Ariki
ICPR2
2010 Speech synthesis by modeling harmonics structure with multiple function
abstract
In this paper, we present a new approach for the speech synthesis, in which speech utterances are synthesized using the parameters of spectro-modeling function (Multiple function). With this approach, only harmonic-parts are extracted from the phoneme spectrum, and the time-varying spectrum corresponding to the harmonics or sinusoidal components is modeled using the Multiple function. We introduce two types of the functions, and present the method to estimate the parameters of each function using the observed phoneme spectrum. In the synthesis stage, speech signals are generated from the parameters of the Multiple function. The advantage of this method is that it only requires a few speech synthesis parameters. We discuss the effectiveness of our proposed method through experimental results.
Toru Nakashika, Ryuki Tachibana, Masafumi Nishimura, Tetsuya Takiguchi, Yasuo Ariki
INTERSPEECH4
2010 Multimodal speech recognition of a person with articulation disorders using AAM and MAF
abstract
We investigated the speech recognition of a person with articulation disorders resulting from athetoid cerebral palsy. The articulation of speech tends to become unstable due to strain on speech-related muscles, and that causes degradation of speech recognition. Therefore, we use multiple acoustic frames (MAF) as an acoustic feature to solve this problem. Further, in a real environment, current speech recognition systems do not have sufficient performance due to noise influence. In addition to acoustic features, visual features are used to increase noise robustness in a real environment. However, there are recognition problems resulting from the tendency of those suffering from cerebral palsy to move their head erratically. We investigate a pose-robust audio-visual speech recognition method using an Active Appearance Model (AAM) to solve this problem for people with articulation disorders resulting from athetoid cerebral palsy. AAMs are used for face tracking to extract pose-robust facial feature points. Its effectiveness is confirmed by word recognition experiments on noisy speech of a person with articulation disorders.
Chikoto Miyamoto, Yuto Komai, Tetsuya Takiguchi, Yasuo Ariki, Ichao Li
MMSP3
2009 Human Action Recognition Using HDP by Integrating Motion and Location Information
Yasuo Ariki, Takuya Tonaru, Tetsuya Takiguchi
ACCV (2)3
2009 Monaural sound-source-direction estimation using the acoustic transfer function of an active microphone
Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki
FUSION2
2009 System request detection in human conversation based on multi-resolution Gabor wavelet features
abstract
For a hands-free speech interface, it is important to detect commands in spontaneous utterances. Usual voice activity detection systems can only distinguish speech frames from nonspeech frames, but they cannot discriminate whether the detected speech section is a command for a system or not. In this paper, in order to analyze the difference between system requests and spontaneous utterances, we focus on fluctuations in a long period, such as prosodic articulation, and fluctuations in a short period, such as phoneme articulation. The use of multi-resolution analysis using Gabor wavelet on a Log-scale Mel-frequency Filter-bank clarifies the different characteristics of system commands and spontaneous utterances. Experiments using our robot dialog corpus show that the accuracy of the proposed method is 92.6% in F-measure, while the conventional power and prosody-based method is just 66.7%. Index Terms: dialog system, voice activity detection, system request detection
Tomoyuki Yamagata, Tetsuya Takiguchi, Yasuo Ariki
INTERSPEECH2
2008 Digital camera work for soccer video production with event recognition and accurate ball tracking by switching search method
abstract
In this paper, we propose a method of digital zooming by automatically recognizing the soccer game events such as penalty kick and free kick based on player and ball tracking. We also propose an efficient and stable ball tracking method by switching search methods between global search and local search. In the frames where the ball is lost as well as in the first frame, the global search with normalized cross-correlation is performed. Then the local search with particle filter continues the tracking. Once the system has recognized that the local search had failed by receiving a sequence of low probabilities of the ball tracking results, the search method is switched to the global search. We carried out three ball tracking experiments; global search, local search and the proposed switching search method. As a result, the experiments showed the effectiveness of the proposed method.
Yasuo Ariki, Tetsuya Takiguchi, Kazuki Yano
ICME2
2008 Graph cuts by using local texture features of wavelet coefficient for image segmentation
abstract
This paper proposes an approach to image segmentation using iterated graph cuts based on local texture features of wavelet coefficient. Using multiresolution analysis based on Haar wavelet, low-frequency range (smoothed image) is used for n-link and high-frequency range (local texture features) is used for t-link along with color histogram. The proposed method can segment the object region with noisy edges and colors similar to the background, but heavy texture change. Experimental results illustrate the validity of our method.
Keita Fukuda, Tetsuya Takiguchi, Yasuo Ariki
ICME2
2008 3D human posture estimation using the HOG features from monocular image
abstract
In this paper, we propose a method to estimate the 3D human posture from monocular image without using the markers. A 3D human body is expressed by a multi-joint model, and a set of the joint angles describes a posture. The proposed method estimates the posture using histograms of oriented gradients(HOG) feature vectors that can express the shape of the object in the input image obtained from monocular camera. In addition, the feature dimension of the background region is reduced for reliability by principal component analysis (PCA) computed at every block of HOG. The joint angles in human multi-joint model are estimated by linear regression analysis applied to its feature vector extracted from the input image. As a result of comparison experiment with the shape contexts features, the RMS error was reduced by about 5.35 degrees.
Katsunori Onishi, Tetsuya Takiguchi, Yasuo Ariki
ICPR2
2008 Object recognition and segmentation using SIFT and Graph Cuts
abstract
In this paper, we propose a method of object recognition and segmentation using Scale-Invariant Feature Transform (SIFT) and Graph Cuts. SIFT feature is invariant for rotations, scale changes, and illumination changes and it is often used for object recognition. However, in previous object recognition work using SIFT, the object region is simply presumed by the affine-transformation and the accurate object region was not segmented. On the other hand, Graph Cuts is proposed as a segmentation method of a detail object region. But it was necessary to give seeds manually. By combing SIFT and Graph Cuts, in our method, the existence of objects is recognized first by vote processing of SIFT keypoints. After that, the object region is cut out by Graph Cuts using SIFT keypoints as seeds. Thanks to this combination, both recognition and segmentation are performed automatically under cluttered backgrounds including occlusion.
Akira Suga, Keita Fukuda, Tetsuya Takiguchi, Yasuo Ariki
ICPR3
2008 Integration of metamodel and acoustic model for speech recognition
abstract
We investigated the speech recognition of a person with artic-ulation disorders resulting from athetoid cerebral palsy. The articulation of the first speech tends to become unstable due to strain on speech-related muscles, and that causes degradation of speech recognition. Therefore, we proposed a robust feature extraction method based on PCA (Principal Component Analy-sis) instead of MFCC [1]. In this paper, we discuss our effort to integrate a Metamodel [2] and Acoustic model approach. Meta-model has a technique for incorporating a model of a speaker’s confusion matrix into the ASR process in such a way as to increase recognition accuracy. Its effectiveness has been con-firmed by word recognition experiments. Index Terms: articulation disorders, PCA, feature extraction, model integration, dysarthric speech
Hironori Matsumasa, Tetsuya Takiguchi, Yasuo Ariki, Ichao Li, Toshitaka Nakabayashi
INTERSPEECH2
2008 Sudden noise reduction based on GMM with noise power estimation
abstract
This paper describes a method for reducing sudden noise using noise detection and classification methods, and noise power estimation. Sudden noise detection and classification have been dealt with in our previous study. In this paper, GMM-based noise reduction is performed using the detection and classification results. As a result of classification, we can determine the kind of noise we are dealing with, but the power is unknown. In this paper, this problem is solved by combining an estimation of noise power with the noise reduction method. In our experiments, the proposed method achieved good performance for recognition of utterances overlapped by sudden noises.
Nobuyuki Miyake, Tetsuya Takiguchi, Yasuo Ariki
INTERSPEECH2
2008 CENSREC-4: development of evaluation framework for distant-talking speech recognition under reverberant environments
abstract
In this paper, we newly introduce a collection of databases and evaluation tools called CENSREC-4, which is an evaluation framework for distant-talking speech under hands-free conditions. Distant-talking speech recognition is crucial for a handsfree speech interface. Therefore, we measured room impulse responses to investigate reverberant speech recognition in various environments. The data contained in CENSREC-4 are connected digit utterances, as in CENSREC-1. Two subsets are included in the data: basic data sets and extra data sets. The basic data sets are used for the evaluation environment for the room impulse response-convolved speech data. The extra data sets consist of simulated and recorded data. An evaluation framework is only provided for the basic data sets as evaluation tools. The results of evaluation experiments proved that CENSREC-4 is an effective database for evaluating the new dereverberation method because the traditional dereverberation process had difficulty sufficiently improving the recognition performance. Index Terms: Various environments, Impulse response, Convolution, Real recorded data, Evaluation framework
Masato Nakayama, Takanobu Nishiura, Yuki Denda, Norihide Kitaoka, Kazumasa Yamamoto, Takeshi Yamada, Satoru Tsuge, Chiyomi Miyajima, Masakiyo Fujimoto, Tetsuya Takiguchi, Satoshi Tamura, Tetsuji Ogawa, Shigeki Matsuda, Shingo Kuroiwa, Kazuya Takeda, Satoshi Nakamura 0001
INTERSPEECH10
2008 Evaluation Framework for Distant-talking Speech Recognition under Reverberant Environments: newest Part of the CENSREC Series -
Takanobu Nishiura, Masato Nakayama, Yuki Denda, Norihide Kitaoka, Kazumasa Yamamoto, Takeshi Yamada, Satoru Tsuge, Chiyomi Miyajima, Masakiyo Fujimoto, Tetsuya Takiguchi, Satoshi Tamura, Shingo Kuroiwa, Kazuya Takeda, Satoshi Nakamura 0001
LREC10
2008 Tagging Video Contents with Positive/Negative Interest Based on User's Facial Expression
Masanori Miyahara, Masaki Aoki, Tetsuya Takiguchi, Yasuo Ariki
MMM3
2007 Development of VAD evaluation framework CENSREC-1-C and investigation of relationship between VAD and speech recognition performance
abstract
Voice activity detection (VAD) plays an important role in speech processing including speech recognition, speech enhancement, and speech coding in noisy environments. We developed an evaluation framework for VAD in such environments, called corpus and environment for noisy speech recognition 1 concatenated (CENSREC-1-C). This framework consists of noisy continuous digit utterances and evaluation tools for VAD results. By adoptiong two evaluation measures, one for frame-level detection performance and the other for utterance-level detection performance, we provide the evaluation results of a power-based VAD method as a baseline. When using VAD in speech recognizer, the detected speech segments are extended to avoid the loss of speech frames and the pause segments are then absorbed by a pause model. We investigate the balance of an explicit segmentation by VAD and an implicit segmentation by a pause model using an experimental simulation of segment extension and show that a small extension improves speech recognition.
Norihide Kitaoka, Kazumasa Yamamoto, Tomohiro Kusamizu, Seiichi Nakagawa, Takeshi Yamada, Satoru Tsuge, Chiyomi Miyajima, Takanobu Nishiura, Masato Nakayama, Yuki Denda, Masakiyo Fujimoto, Tetsuya Takiguchi, Satoshi Tamura, Shingo Kuroiwa, Kazuya Takeda, Satoshi Nakamura 0001
ASRU12
2007 PCA-based feature extraction for fluctuation in speaking style of articulation disorders
abstract
We investigated the speech recognition of a person with articulation disorders resulting from athetoid cerebral palsy. Recently, the accuracy of speaker-independent speech recognition has been remarkably improved by the use of stochastic modeling of speech. However, the use of those acoustic models causes degradation of speech recognition for a person with different speech styles (e.g., articulation disorders). In this paper, we discuss our efforts to build an acoustic model for a person with articulation disorders. The articulation of the first speech tends to become unstable due to strain on muscles and that causes degradation of speech recognition. Therefore, we propose a robust feature extraction method based on PCA (Principal Component Analysis) instead of MFCC. Its effectiveness is confirmed by word recognition experiments. Index Terms: articulation disorders, PCA, feature extraction
Hironori Matsumasa, Tetsuya Takiguchi, Yasuo Ariki, Ichao Li, Toshitaka Nakabayashi
INTERSPEECH2
2007 Language modeling using PLSA-based topic HMM
abstract
In this paper, we propose a PLSA-based language model for sports-related live speech. This model is implemented using a unigram rescaling technique that combines a topic model and an n-gram. In the conventional method, unigram rescaling is performed with a topic distribution estimated from a recognized transcription history. This method can improve the performance, but it cannot express topic transition. By incorporating the concept of topic transition, it is expected that the recognition performance will be improved. Thus, the proposed method employs a “Topic HMM” instead of a history to estimate the topic distribution. The Topic HMM is an Ergodic HMM that expresses typical topic distributions as well as topic transition probabilities. Word accuracy results from our experiments confirmed the superiority of the proposed method over a trigram and a PLSA-based conventional method that uses a recognized history.
Atsushi Sako, Tetsuya Takiguchi, Yasuo Ariki
INTERSPEECH2
2007 System request detection in conversation based on acoustic and speaker alternation features
abstract
For a hands-free speech interface, it is important to detect com-mands in spontaneous utterances. To discriminate commands from human-human conversations by acoustic features, it is ef-ficient to consider the head and the tail of an utterance. The dif-ferent characteristics of system requests and spontaneous utter-ances appear on these parts of an utterance. Experiment shows that by separating the head and the tail of an utterance, the ac-curacy of detection was improved. And also, considering the al-ternation of speakers using two channel microphones improved the performance. Although detecting system requests using lin-guistic features shows high accuracy, combining acoustic and turn-taking features lift up the performance. Index Terms: system request detection, utterance verification, SVM, speech recognition, turn-taking
Tomoyuki Yamagata, Atsushi Sako, Tetsuya Takiguchi, Yasuo Ariki
INTERSPEECH3
2007 Voice activity detection by lip shape tracking using EBGM
abstract
We propose a voice activity detection of a target speaker (driver) in a car by integrating lip movement and acoustic processing. To prevent the wrong detection caused by nontarget speakers using only acoustic processing, the proposed system extracts the lip movement of the target speaker by measuring the lip aspect ratio. An infrared camera is used to cope with the change of lighting environment. In order to extract the lip from gray scale images, Elastic Bunch Graph Matching is employed. Experimental results showed the proposed system improved the precision rate in the voice activity detection by approximately 40% compared to the method using only acoustic processing in a car.
Masaki Aoki, Ken Masuda, Hiroyoshi Matsuda, Tetsuya Takiguchi, Yasuo Ariki
ACM Multimedia4
2006 Robust Feature Extraction using Kernel PCA
abstract
We investigate a robust speech feature extraction method using kernel PCA (principal component analysis). Kernel PCA has been suggested for various image processing tasks requiring an image model such as, e.g., denoising, where a noise-free image is constructed from a noisy input image. Much research for robust speech feature extraction has been done, but it is difficult to completely remove the non-stationary noise or reverberation. The most commonly used noise-removal techniques are based on the spectral-domain operation, and then for the speech recognition, MFCC (mel frequency cepstral coefficient) is computed, where DCT (discrete cosine transform) is applied to the mel-scale filter bank output. In this paper, we propose robust feature extraction based on kernel PCA instead of DCT, where the main speech element is projected onto low-order features, while noise or reverberant element is projected onto high-order ones. Its effectiveness is confirmed by word recognition experiments on reverberant speech
Tetsuya Takiguchi, Yasuo Ariki
ICASSP (1)1
2006 Phoneme recognition based on fisher weight map to higher-order local auto-correlation
abstract
In this paper, we propose a new feature extraction method based on higher-order local auto-correlation (HLAC) and Fisher weight map (FWM). Widely used MFCC features lack temporal dynamics. To solve this problem, 35 types of local auto-correlation features are computed within two-dimensional local regions. These local features are accumulated over more global regions by weighting high scores on the discriminative areas where the typical features among all phonemes are well expressed. This score map is called Fisher weight map. We verified the effectiveness of the HLAC and FWM through vowel recognition and total phoneme recognition.
Yasuo Ariki, Shunsuke Kato 0001, Tetsuya Takiguchi
INTERSPEECH3
2005 Situation based speech recognition for structuring baseball live games
abstract
It is a difficult problem to recognize baseball live speech because the speech is rather fast, noisy, emotional and disfluent due to rephrasing, repetition, mistake and grammatical deviation caused by spontaneous speaking style. To solve these problems, we have been studied the speech recognition method incorporating the baseball game task-dependent knowledge as well as an announcer’s emotion in commentary speech [1]. In addition, in this paper, we propose the situation prediction model based on word co-occurrence. Owing to these proposed models, speech recognition errors are effectively prevented. This method is formalized in the framework of probability theory and implemented in the conventional speech decoding (Viterbi) algorithm. The experimental results showed that the proposed approach improved the structuring and segmentation accuracy as well as keywords accuracy.
Atsushi Sako, Tetsuya Takiguchi, Yasuo Ariki
INTERSPEECH2
2005 Recognition of hands-free speech and hand pointing action for conversational TV
abstract
In this paper, we propose a structure and components of a conversational television set(TV) to which we can ask anything on the broadcasted contents and receive the interesting information from the TV. The conversational TV is composed of two types of processing; back end processing and front end processing. In the back end processing, broadcasted contents are analyzed using speech and video recognition techniques and both of the meta data and the structure are extracted. In the front end processing, human speech and hand action are recognized to understand the user intention. We show some applications, being developed in this conversational TV with multi-modal interactions, such as word explanation, human information retrieval, event retrieval in soccer and baseball video games with contextual awareness.
Yasuo Ariki, Tetsuya Takiguchi, Atsushi Sako
ACM Multimedia2
2004 Acoustic model adaptation using first order prediction for reverberant speech
abstract
The paper describes a hands-free speech recognition technique based on acoustic model adaptation to reverberant speech. In hands-free speech recognition, the recognition accuracy is degraded by reverberation, since each segment of speech is affected by the reflection energy of the preceding segment. To compensate for the reflection signal, we introduce a frame-by-frame adaptation method, adding the reflection signal to the means of the acoustic model. The reflection signal is approximated by a first-order linear prediction from the preceding frame, and the linear prediction coefficient is estimated by a maximum likelihood method by using the EM algorithm, which maximizes the likelihood of the adaptation data. Its effectiveness is confirmed by word recognition experiments on reverberant speech.
Tetsuya Takiguchi, Masafumi Nishimura
ICASSP (1)1
2001 HMM-separation-based speech recognition for a distant moving speaker
abstract
This paper presents a hands-free speech recognition method based on HMM composition and separation for speech contaminated not only by additive noise but also by an acoustic transfer function. The method realizes an improved user interface such that a user is not encumbered by microphone equipment in noisy and reverberant environments. The use of HMM composition has already been proposed for countering additive noise. In this paper, the same approach is extended to handle convolutional acoustic distortion in a reverberant room, by using an HMM to model the acoustic transfer function. The states of this HMM correspond to different positions of the sound source. It can represent the positions of the sound sources, even if the speaker moves. This paper also proposes a new method, HMM separation, for estimating the HMM parameters of the acoustic transfer function on the basis of a maximum likelihood manner. The proposed method is obtained through the reverse of the process of HMM composition, where the model parameters are estimated by maximizing the likelihood of adaptation data uttered from an unknown position. Therefore, measurement of impulse responses is not required. The paper also describes the performance of the proposed methods for recognizing real distant-talking speech. The results of experiments clarify the effectiveness of the proposed method.
Tetsuya Takiguchi, Satoshi Nakamura 0001, Kiyohiro Shikano
IEEE Trans. Speech Audio Process.1
2000 Speech recognition for a distant moving speaker based on HMM composition and separation
abstract
This paper describes a hands-free speech recognition method based on HMM composition and separation for speech contaminated not only by additive noise but also by an acoustic transfer function. The method realizes an improved user interface such that a user is not encumbered by microphone equipment in noisy and reverberant environments. In this approach, an attempt is made to model acoustic transfer functions by means of an ergodic HMM. The states of this HMM correspond to different positions of the sound source. It can represent the positions of the sound sources, even if the speaker moves. The HMM parameters of the acoustic transfer function are estimated by HMM separation. The method is obtained through the reverse of the process of HMM composition, where the model parameters are estimated by maximizing the likelihood of adaptation data uttered from an unknown position. Therefore, measurement of impulse responses is not required. In this paper, we record the speech of a distant moving speaker in real environments. The results of experiments for the speech of a distant moving speaker clarified the effectiveness of HMM composition and separation.
Tetsuya Takiguchi, Satoshi Nakamura 0001, Kiyohiro Shikano
ICASSP1
1998 Evaluation of model adaptation by HMM decomposition on telephone speech recognition
abstract
In this paper, we evaluate performance of model adaptation by the previously proposed HMM decomposition method on telephone speech recognition. The HMM decomposition method separates a composed HMM into a known phoneme HMM and an unknown noise and channel HMM by maximum likelihood (ML) estimation of the HMM parameters. A transfer function (telephone channel) HMM is estimated using adaptation speech data by applying the HMM decomposition twice in the linear spectral domain for noise and in the cepstral domain for channel. The telephone speech data for evaluation are recorded through 10 kinds of ordinary analog telephone handsets and cordless telephone handsets. The test results show that the average phrase accuracy with the clean speech HMMs is 60.9% for the ordinary analog telephone handsets, and 19.6% for the cordless telephone handsets. By the HMM decomposition method, the average phrase accuracy is improved to 78.1% for the ordinary analog telephone handsets, and 50.5% for the cordless telephone handsets.
Tetsuya Takiguchi, Satoshi Nakamura 0001, Kiyohiro Shikano, Masatoshi Morishima, Toshihiro Isobe
ICSLP1
1997 Model adaptation based on HMM decomposition for reverberant speech recognition
abstract
The performance of a speech recognizer is degraded drastically in reverberant environments. The authors propose a novel algorithm which can model an observation signal by composition of HMMs of clean speech, noise and an acoustic transfer function. However, estimating HMM parameters of the acoustic transfer function is still a serious problem. In their previous paper, they measured real impulse responses of training positions in an experiment room. It is inconvenient and unrealistic to measure impulse responses for every possible new experiment room. The paper presents a new method for estimating HMM parameters of the acoustic transfer function from some adaptation data by using an HMM decomposition algorithm which is an inverse process of the HMM composition. Its effectiveness is confirmed by a series of speaker dependent and independent word recognition experiments on simulated distant-talking speech data.
Tetsuya Takiguchi, Satoshi Nakamura 0001, Kiyohiro Shikano
ICASSP1
1996 Noise and room acoustics distorted speech recognition by HMM composition
abstract
This paper presents a robust speech recognition method based on the HMM composition for the noisy room acoustics distorted speech. The method realizes an improved user interface such as the user is not encumbered by microphone equipment. The proposed HMM composition is obtained by naturally extending the HMM composition method of an additive noise to that of the convolutional room acoustics distortion. The HMM composition is conducted by 2 steps: (1) composition of HMMs of a speech and acoustical transfer function in the cepstrum domain, and (2) composition of distorted speech and noise HMMs in the linear spectral domain. The speaker dependent/independent word recognition experiments are carried out using the speech database contaminated by the additive noise and convolutional room acoustics distortion. The evaluation experiments are also conducted for unknown testing sound source positions. These results clarified the effectiveness of the proposed method.
Satoshi Nakamura 0001, Tetsuya Takiguchi, Kiyohiro Shikano
ICASSP2