VLDB 2026 Research / reviewers in the wild / expert
Tetsuya Takiguchi
dblp:79/4485
· DBLP profile ↗
115ranked-venue papers
9as first author
20since 2021 · last 2025
0000-0001-5005-7679ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 91 · 8 first-author · 9 since 2021Artificial intelligence and machine learning · 52 · 2 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 10 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 since 2021Databases, data management, data science and information retrieval · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Speaker-dependent Continuous Speech Recognition for Individuals with Cerebral Palsy Using Weighted Finite-State Transducer and Text-to-Speech SynthesisabstractDespite remarkable advances in automatic speech recognition (ASR) technology, existing systems have not achieved sufficient recognition accuracy for speech recognition of individuals with cerebral palsy.Speech recognition for individuals with speech disorders faces acoustic challenges because the speech characteristics of these individuals differ from those of individuals without speech disorders.Additionally, Japanese ASR faces unique linguistic challenges due to the mixed character set including kanji (Chinese characters), hiragana, and katakana (Japanese phonetic syllabary).Adapting end-to-end ASR models to this task requires large amounts of training data.However, collecting sufficient amounts of speech data for training is difficult because recording speech from individuals with cerebral palsy is a large burden on them.This paper revisit a weighted finite-state transducer based hybrid speaker-dependent ASR system for individuals with cerebral palsy, which decomposes the system into a speaker-dependent acoustic model, a pronunciation dictionary, and a language model.This approach is effective when speech data is limited, as it uses speech data only for training the acoustic model while other components are learned only from text data.Furthermore, to enhance the speaker-dependent acoustic model, we introduce data augmentation using text-to-speech synthesis and multi-step model adaptation using synthetic speech.Experimental validation using speech samples from individuals with cerebral palsy demonstrates that the proposed methodology achieves superior performance compared to the state-of-the-art end-to-end Whisper (ASR system). Takeru Otani, Ryo Aihara, Ryoichi Takashima, Tetsuya Takiguchi, Tatsuhiko Saito |
ASSETS | 4 |
| 2025 | Highly Intelligible Text-to-Speech System Based on Weighted Averaging of Parameters for Individuals with Spinal Muscular AtrophyabstractTo support communication for individuals with dysarthria who have difficulty producing intelligible speech, text-to-speech (TTS) systems are gaining attention.However, conventional TTS systems synthesize speech using voices that differ from those of the users themselves, which can create a sense of psychological distance between the user and their communication partner.Deep neural network-based TTS models can accurately reproduce trained speech and can generate speech resembling the user's own voice by training on the user's speech data.However, the models trained on speech with dysarthria also replicate the unintelligibility of the original speech, making them unsuitable for communication support.This paper focuses on dysarthria caused by spinal muscular atrophy (SMA), and proposes a method to construct a TTS model that synthesizes intelligible speech while preserving the voice characteristics of a speaker with dysarthria.The proposed approach involves computing a weighted average of the parameters of a TTS model trained on speech from an SMA speaker and a TTS model trained on speech from a speaker without dysarthria.The experimental results confirm that the synthesized speech generated using the proposed method maintains the voice quality of the SMA speaker while being more intelligible than that produced by conventional methods. CCS Concepts• Social and professional topics → Assistive technologies. Yusuke Yagi, Ryoichi Takashima, Chiho Sasaki, Tetsuya Takiguchi |
ASSETS | 4 |
| 2025 | Revisiting WFST-based Hybrid Japanese Speech Recognition System for Individuals with Organic Speech Disorders
Naoki Hojo, Ryoichi Takashima, Chihiro Sugiyama, Nobukazu Tanaka, Kanji Nohara, Kazunori Nozaki, Tetsuya Takiguchi |
INTERSPEECH | 7 |
| 2025 | Zero-Shot Learning for Acoustic Event Classification Using an Attribute Vector and Conditional GAN
Kohei Uehara, Ryoichi Takashima, Tetsuya Takiguchi |
INTERSPEECH | 3 |
| 2025 | Operatic Singing Voice Synthesis From Inexperienced Voice Considering Tempo and Vowel Change
Aoto Sugahara, Soma Kishimoto, Yuji Adachi, Kiyoto Tai, Ryoichi Takashima, Tetsuya Takiguchi |
MMM (3) | 6 |
| 2025 | A Robot that Supports Collaborative Art Appreciation through Visual Thinking StrategiesabstractSocial robots are increasingly being used as interactive partners to facilitate people’s understanding and appreciation of art. Considering people's stages of aesthetic development is essential to richer understanding of artworks, but such a viewpoint is less of a focus in human-robot interaction contexts. Therefore, we developed a robot system for collaborative art appreciation using a Visual Thinking Strategies (VTS) method to enhance people’s engagement with art by considering their stages of aesthetic experience. We conducted an experiment to investigate the effectiveness of our system in supporting art appreciation of participants in a laboratory setting. We also investigated the effects of embodiment, i.e., the physical body of a robot, on art-appreciation support. The experiment results indicate that our system significantly increased the intention to use it, which is related to social acceptance, and embodiment significantly influenced likeability, perceived intelligence, and perceived enjoyment. Minori Iwata, Masahiro Shiomi, Tetsuya Takiguchi |
RO-MAN | 3 |
| 2024 | Individuality-Preserving Speech Synthesis for Spinal Muscular Atrophy with a TracheotomyabstractAphasia and dysarthria are the two main language disorders that cause difficulty in speech. This study focuses on articulation disorders, particularly among individuals with spinal muscular atrophy (SMA) whose speech is challenging to comprehend. Specifically, it addresses communication support through text-to-speech synthesis technology that maintains the speaker’s individuality. Previous research on individuals with SMA who have undergone tracheotomy surgery has predominantly centered on postoperative care environments unrelated to speech communication, with few precedents in the study of communication support using speech synthesis technology. Therefore, this study aims to develop a speech synthesis system that preserves the speaker’s individuality while producing clearer speech. This is performed by fine-tuning a pre-trained speech synthesis model, initially trained on a large corpus of speech by those with no speech impediment, using a small amount of speech of the target person with SMA. Subjective evaluations using both actual and synthesized speech demonstrated that the system could adequately learn the speaker’s individuality and produce synthesized speech with slightly improved clarity. Minori Iwata, Ryoichi Takashima, Chiho Sasaki, Tetsuya Takiguchi |
ASSETS | 4 |
| 2024 | Self-supervised learning using unlabeled speech with multiple types of speech disorder for disordered speech recognitionabstractThis paper investigates a training method of an automatic speech recognition (ASR) model for people with speech disorders. Because the characteristics of their speech differ significantly from those of the typical speech, in order to recognize the speech of a user with a disorder, the system needs to be trained with the user’s speech in advance. However, recording speech from people with disorders is a large burden for them, and therefore, it is difficult to collect a sufficient amount of speech for training. To address this issue, this study investigates the use of two types of speech as training data. The first type is unlabeled speech, which can be easily collected but lacks text labels (e.g., spontaneous speech in daily life). To utilize the unlabeled speech for training an ASR model, a self-supervised learning approach is employed. The second type involves utilizing speech data from individuals with different types of speech disorders. In our system, besides the user’s speech, the speech of individuals with the same type of disorder and even different types of disorders is also incorporated. Experimental results demonstrated that using unlabeled speech and speech from multiple types of disorders led to reduced recognition error rates. Ryoichi Takashima, Takeru Otani, Ryo Aihara, Tetsuya Takiguchi, Shinya Taguchi |
ASSETS | 4 |
| 2024 | Attempts on detecting Alzheimer's disease by fine-tuning pre-trained model with Gaze DataabstractThis poster presents a study on detecting Alzheimer’s disease (AD) using deep learning from gaze data. In this study, we modify an existing pre-trained deep neural network model, gazeNet, for transfer learning. The results suggest the possibility of applying this method to mild cognitive impairment screening tests. Junichi Nagasawa, Yuichi Nakata, Mamoru Hiroe, Yutaka Kawaguchi, Yuji Maegawa, Naoki Hojo, Tetsuya Takiguchi, Minoru Nakayama, Maki Uchimura, Yuma Sonoda, Hisatomo Kowa, Takashi Nagamatsu |
ETRA | 8 |
| 2024 | Attempts on detecting Alzheimer's disease by fine-tuning pre-trained model with Gaze DataabstractEarly detection of Alzheimer’s disease (AD) is important but difficult. Screening for AD using neuropsychological tests such as mini-mental state examination (MMSE) is time-consuming and burdensome for patients. Recently, several methods have been reported for detecting AD based on eye movements. However, analyzing eye movements requires considerable effort. Although machine learning from eye movement data is a strong candidate for labor-saving, it requires large datasets. In this study, we modify an existing pre-trained deep neural network model, gazeNet, for transfer learning. For evaluation, we exclusively used data from one participant and fine-tuned the model using data from all the remaining participants. We repeated this procedure separately for each of the 14 participants. The results of eye movement during the antisaccade task were not satisfactory for the discrimination of AD, and detailed analysis suggested that the data might potentially have a correlation with MMSE scores in the mild cognitive impairment range. Junichi Nagasawa, Yuichi Nakata, Mamoru Hiroe, Yutaka Kawaguchi, Yuji Maegawa, Naoki Hojo, Tetsuya Takiguchi, Minoru Nakayama, Maki Uchimura, Yuma Sonoda, Hisatomo Kowa, Takashi Nagamatsu |
ETRA | 8 |
| 2024 | Effects of Listening Behaviors of a Social Robot on Adult's Motivation and Performance in Piano PracticeabstractThe landscape of education with social robots is evolving, especially within the realm of music education. However, past studies have focused on children as music learners and such verbal behaviors as praise. Therefore, it remains unknown whether existing studies are effective for adult learners in musical education as well as how to effectively design the non-verbal behaviors of robots. This study investigates the effective behaviors of social robots by comparing three kinds of listening behaviors: none, nodding (simple listening), and enjoying (affective listening). We developed a music education support system that consists of a social robot and a MIDI keyboard and conducted an experiment with adult participants. Our experimental results described the advantages of affective-listening behavior during music education over simple listening and non-listening based on gender. Ryuto Matsusaka, Masahiro Shiomi, Tetsuya Takiguchi |
RO-MAN | 3 |
| 2023 | Zero-Shot Sound Event Classification Using a Sound Attribute Vector with Global and Local Feature LearningabstractThis paper introduces a zero-shot sound event classification (ZS-SEC) method to identify sound events that have never occurred in training data. In our previous work, we proposed a ZS-SEC method using sound attribute vectors (SAVs), where a deep neural network model infers attribute information that describes the sound of an event class instead of inferring its class label directly. Our previous method showed that it could classify unseen events to some extent; however, the accuracy for unseen events was far inferior to that for seen events. In this paper, we propose a new ZS-SEC method that can learn discriminative global features and local features simultaneously to enhance SAV-based ZS-SEC. In the proposed method, while the global features are learned in order to discriminate the event classes in the training data, the spectro-temporal local features are learned in order to regress the attribute information using attribute prototypes. The experimental results show that our proposed method can improve the accuracy of SAV-based ZS-SEC and can visualize the region in the spectrogram related to each attribute. Xunquan Chen, Ryoichi Takashima, Tetsuya Takiguchi |
ICASSP | 4 |
| 2023 | Harmonic-Net: Fundamental Frequency and Speech Rate Controllable Fast Neural VocoderabstractThere is a need to improve the synthesis quality of HiFi-GAN-based real-time neural speech waveform generative models on CPUs while preserving the controllability of fundamental frequency ($f_{\mathrm{o}}$) and speech rate (SR). For this purpose, we propose Harmonic-Net and Harmonic-Net+, which introduce two extended functions into the HiFi-GAN generator. The first extension is a downsampling network, named the excitation signal network, that hierarchically receives multi-channel excitation signals corresponding to$f_{\mathrm{o}}$. The second extension is the layerwise pitch-dependent dilated convolutional network (LW-PDCNN), which can flexibly change its receptive fields depending on the input$f_{\mathrm{o}}$to handle large fluctuations in$f_{\mathrm{o}}$for the upsampling-based HiFi-GAN generator. The proposed explicit input of excitation signals and LW-PDCNNs corresponding to$f_{\mathrm{o}}$are expected to realize high-quality synthesis for the normal and$f_{\mathrm{o}}$-conversion conditions and for the SR-conversion condition. The results of experiments for unseen speaker synthesis, full-band singing voice synthesis, and text-to-speech synthesis show that the proposed method with harmonic waves corresponding to$f_{\mathrm{o}}$can achieve higher synthesis quality than conventional methods in all (i.e., normal,$f_{\mathrm{o}}$-conversion, and SR-conversion) conditions. Keisuke Matsubara, Takuma Okamoto, Ryoichi Takashima, Tetsuya Takiguchi, Tomoki Toda, Hisashi Kawai |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Speaker-Independent Emotional Voice Conversion via Disentangled RepresentationsabstractEmotional Voice Conversion (EVC) technology aims to transfer emotional state in speech while keeping the linguistic information and speaker identity unchanged. Prior studies on EVC have been limited to perform the conversion for a specific speaker or a predefined set of multiple speakers seen in the training stage. When encountering arbitrary speakers that may be unseen during training (outside the set of speakers used in training), existing EVC methods have limited conversion capabilities. However, converting the emotion of arbitrary speakers, even those unseen during the training procedure, in one model is much more challenging and much more attractive in real-world scenarios. To address this problem, in this study, we propose SIEVC, a novel speaker-independent emotional voice conversion framework for arbitrary speakers via disentangled representation learning. The proposed method employs the autoencoder framework to disentangle the emotion information and emotion-independent information of each input speech into separated representation spaces. To achieve better disentanglement, we incorporate mutual information minimization into the training process. In addition, adversarial training is applied to enhance the quality of the generated audio signals. Finally, speaker-independent EVC for arbitrary speakers could be achieved by only replacing the emotion representations of source speech with the target ones. The experimental results demonstrate that the proposed EVC model outperforms the baseline models in terms of objective and subjective evaluation for both seen and unseen speakers. Xunquan Chen, Xuexin Xu, Zhihong Zhang 0001, Tetsuya Takiguchi, Edwin R. Hancock |
IEEE Trans. Multim. | 5 |
| 2022 | Speaker-Targeted Audio-Visual Speech Recognition Using a Hybrid CTC/Attention Model with Interference LossabstractAudio-visual (AV)-automatic speech recognition (ASR) can improve speech recognition accuracy by using lip images, especially in noisy environments. The recently proposed AV Align system integrates speech and image features based on a cross-modal attention mechanism, where attention weights for visual features are estimated by using acoustic features as queries. Although AV Align shows an improvement in recognition accuracy in background noise environments, we have observed that the recognition accuracy degrades significantly in interference speaker environments, where a target speech and an interfering speech overlap each other. In order to improve the speech recognition accuracy of the target speaker in such situations, we propose a method that combines the auxiliary loss function that maximizes the recognition accuracy of the interference speaker and the CTC loss function for training the AV-ASR model. The experimental results using the TCD-TIMIT dataset show that the use of these auxiliary loss functions improves the performance of target-speaker speech recognition in interference speaker environments. Ryota Tsunoda, Ryo Aihara, Ryoichi Takashima, Tetsuya Takiguchi, Yoshie Imai |
ICASSP | 4 |
| 2022 | Where Do Humans Build Levees? A Case Study on the Contiguous United StatesabstractUnderstanding where and why human build levees offers several values: From a hydrological perspective, integration of levees to global flood models has been shown to improve their accuracy. From an Economic intelligence perspective, levee locations provide precious insights into past and future urban developments. However, very little data exists on the location of levees at a global scale, which hinders our ability to reach a global understanding of this question. One rare exception is the National Levee Database (NLD) dataset provided by the U.S. Army Corps of Engineers (USACE). In this study, we hypothesize that levees are built at locations where human activity and flood risk coexist, and develop predictive models that output the probability of levee existence at the hydrological catchment level. Quantitative analysis of these models using the NLD dataset allows to validate our hypothesis, with several important nuances, which we discuss at length. M. Ikegawa, Tristan Hascoet, Victor Pellet, Tetsuya Takiguchi, D. Yamazaki |
IGARSS | 7 |
| 2022 | Building a Knowledge-Based Dialogue System with Text InfillingabstractIn recent years, generation-based dialogue systems using state-of-the-art (SoTA) transformerbased models have demonstrated impressive performance in simulating human-like conversations.To improve the coherence and knowledge utilization capabilities of dialogue systems, knowledge-based dialogue systems integrate retrieved graph knowledge into transformer-based models.However, knowledge-based dialog systems sometimes generate responses without using the retrieved knowledge.In this work, we propose a method in which the knowledge-based dialogue system can constantly utilize the retrieved knowledge using text infilling.Text infilling is the task of predicting missing spans of a sentence or paragraph.We utilize this text infilling to enable dialog systems to fill incomplete responses with the retrieved knowledge.Our proposed dialogue system has been proven to generate significantly more correct responses than baseline dialogue systems. Tetsuya Takiguchi, Yasuo Ariki |
SIGDIAL | 2 |
| 2022 | Direction of arrival estimation for indoor environments based on acoustic composition model with a single microphone
Xingchen Guo, Xuexin Xu, Xunquan Chen, Rong Jia, Zhihong Zhang 0001, Tetsuya Takiguchi, Edwin R. Hancock |
Pattern Recognit. | 7 |
| 2021 | High-Intelligibility Speech Synthesis for Dysarthric Speakers with LPCNet-Based TTS and CycleVAE-Based VCabstractThis paper presents a high-intelligibility speech synthesis method for persons with dysarthria caused by athetoid cerebral palsy. The muscular control of such speakers is unstable because of their athetoid symptoms, and their pronunciation is unclear, which makes it difficult for them to communicate. In this paper, we present a method for generating highly intelligible speech that preserves the individuality of dysarthric speakers by combining Transformer-TTS, CycleVAE-VC, and a LPCNet vocoder. Rather than repairing prosody from the dysarthric speech, this method transfers the dysarthric speaker’s individuality to the speech of a healthy person generated by TTS synthesis. This task is both important and challenging. From the results of our evaluation experiments, we confirmed that the proposed method can partially transfer the individuality of the target dysarthric speaker while maintaining the intelligibility of the source speech. Keisuke Matsubara, Takuma Okamoto, Ryoichi Takashima, Tetsuya Takiguchi, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai |
ICASSP | 4 |
| 2021 | Multimodal fusion for indoor sound source localization
Ryoichi Takashima, Xingchen Guo, Zhihong Zhang 0001, Xuexin Xu, Tetsuya Takiguchi, Edwin R. Hancock |
Pattern Recognit. | 6 |
| 2020 | FasterRCNN Monitoring of Road Damages: Competition and DeploymentabstractMaintaining aging infrastructure is a challenge currently faced by local and national administrators all around the world. An important prerequisite for efficient infrastructure maintenance is to continuously monitor (i.e., quantify the level of safety and reliability) the state of very large structures. Meanwhile, computer vision has made impressive strides in recent years, mainly due to successful applications of deep learning models. These novel progresses are allowing the automation of vision tasks, which were previously impossible to automate, offering promising possibilities to assist administrators in optimizing their infrastructure maintenance operations. In this context, the IEEE 2020 global Road Damage Detection (RDD) Challenge is giving an opportunity for deep learning and computer vision researchers to get involved and help accurately track pavement damages on road networks. This paper proposes two contributions to that topic: In a first part, we detail our solution to the RDD Challenge. In a second part, we present our efforts in deploying our model on a local road network, explaining the proposed methodology and encountered challenges. Tristan Hascoet, Andreas Persch, Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki |
IEEE BigData | 5 |
| 2020 | Two-Step Acoustic Model Adaptation for Dysarthric Speech RecognitionabstractThis paper introduces a model adaptation approach for a speaker-dependent dysarthric speech recognition system. The dysarthria we focus on in this paper is caused by athetoid cerebral palsy, which causes involuntary muscle movements in those with the disease. For this reason, the dysarthric people's speech is often unstable and difficult for conventional automatic speech recognition (ASR) systems to recognize. A model-adaptation approach, which adapts an ASR model to dysarthric speech, is one possible solution. However, because the difference in speaking styles between dysarthric and non-dysarthric people is so significant, the conventional adaptation method is not able to sufficiently adapt the model to the dysarthric speech. In our proposed two-step model-adaptation approach, an ASR model is first adapted to the general speaking style of multiple dysarthric speakers, and then the adapted model is further adapted for the target speaker. From our experiments on an ASR task, our two-step adaptation approach showed better performance than a conventional one-step adaptation approach. Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki |
ICASSP | 2 |
| 2020 | Dysarthric Speech Recognition Based on Deep Metric Learning
Yuki Takashima, Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki |
INTERSPEECH | 3 |
| 2019 | On Zero-Shot Recognition of Generic ObjectsabstractMany recent advances in computer vision are the results of a healthy competition among researchers on high quality, task-specific, benchmarks. After a decade of active research, zero-shot learning (ZSL) models accuracy on the Imagenet benchmark remains far too low to be considered for practical object recognition applications. In this paper, we argue that the main reason behind this apparent lack of progress is the poor quality of this benchmark. We highlight major structural flaws of the current benchmark and analyze different factors impacting the accuracy of ZSL models. We show that the actual classification accuracy of existing ZSL models is significantly higher than was previously thought as we account for these flaws. We then introduce the notion of structural bias specific to ZSL datasets. We discuss how the presence of this new form of bias allows for a trivial solution to the standard benchmark and conclude on the need for a new benchmark. We then detail the semi-automated construction of a new benchmark to address these flaws. Tristan Hascoet, Yasuo Ariki, Tetsuya Takiguchi |
CVPR | 3 |
| 2019 | End-to-end Dysarthric Speech Recognition Using Multiple DatabasesabstractWe present in this paper an end-to-end automatic speech recognition (ASR) system for a person with an articulation disorder resulting from athetoid cerebral palsy. In the case of a person with this type of articulation disorder, the speech style is quite different from that of a physically unimpaired person, and the amount of their speech data available to train the model is limited because their burden is large due to strain on the speech muscles. Therefore, the performance of ASR systems for people with an articulation disorder degrades significantly. In this paper, we propose an end-to-end ASR framework trained by not only the speech data of a Japanese person with an articulation disorder but also the speech data of a physically unimpaired Japanese person and a non-Japanese person with an articulation disorder to relieve the lack of training data of a target speaker. An end-to-end ASR model encapsulates an acoustic and language model jointly. In our proposed model, an acoustic model portion is shared between persons with dysarthria, and a language model portion is assigned to each language regardless of dysarthria. Experimental results show the merit of our proposed approach of using multiple databases for speech recognition. Yuki Takashima, Tetsuya Takiguchi, Yasuo Ariki |
ICASSP | 2 |
| 2019 | Emotional Voice Conversion Using Dual Supervised Adversarial Networks With Continuous Wavelet Transform F0 FeaturesabstractIn emotional voice conversion (VC) tasks, it is difficult to deal with a simple representation of fundamental frequency (F0), which is the most important feature in emotional voice representation. In order to address this issue, we propose the adaptive scales continuous wavelet transform (ADS-CWT) method to systematically capture F0 features of different temporal levels, which can represent different prosodic aspects, ranging from micro-prosody to sentences. Moreover, in an emotional VC task, each dataset is paired with the labeled emotional voice and neutral voice, which can be regarded as a dual task. Owing to, first, dual supervised learning's ability to improve the training performances by using the leveraging probabilistic connection between the dual tasks to enhance the learning from labeled data and, second, generative adversarial networks' (GANs') ability to mitigate the over-smoothing problem caused in the low-level data space when converting the acoustic features, we further present a novel training framework for emotional VC using GANs combined with dual supervised learning, named as dual supervised adversarial networks. In emotional VC experiments, we confirmed the high similarity performance of our method when using limited labeled data for emotional VC. Our method achieves good and consistent performance, in both objective and subjective evaluations. Zhaojie Luo, Tetsuya Takiguchi, Yasuo Ariki |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Polar Transformation on Image Features for Orientation-Invariant RepresentationsabstractThe choice of image feature representation plays a crucial role in the analysis of visual information. Although vast numbers of alternative robust feature representation models have been proposed to improve the performance of different visual tasks, most existing feature representations [e.g., handcrafted features or convolutional neural networks (CNNs)] have a relatively limited capacity to capture the highly orientation-invariant (rotation/reversal) features. The net consequence is suboptimal visual performance. To address these problems, this study adopts a novel transformational approach, which investigates the potential of using polar feature representations. Our low level consists of a histogram of oriented gradient, which is then binned using annular spatial bin-type cells applied to the polar gradient. This gives gradient binning invariance for feature extraction. In this way, the descriptors have significantly enhanced orientation-invariant capabilities. The proposed feature representation, calledorientation-invariant histograms of oriented gradients, is capable of accurately processing visual tasks (e.g., facial expression recognition). In the context of the CNN architecture, we propose two polar convolution operations, referred to as full polar convolution and local polar convolution, and use these to develop polar architectures for the CNN orientation-invariant representation. Experimental results show that the proposed orientation-invariant image representation, based on polar models for both handcrafted features and deep learning features, is both competitive with state-of-the-art methods and maintains compact representation on a set of challenging benchmark image datasets. Zhaojie Luo, Zhihong Zhang 0001, Faliang Huang, Zhiling Ye, Tetsuya Takiguchi, Edwin R. Hancock |
IEEE Trans. Multim. | 6 |
| 2018 | Parallel-Data-Free Dictionary Learning for Voice Conversion Using Non-Negative Tucker DecompositionabstractVoice conversion (VC) is a technique where only speaker-specific information in source speech is converted while preserving the associated phonological information. Nonnegative Matrix Factorization (NMF)-based VC has been researched because of the natural-sounding voice it produces compared with conventional Gaussian Mixture Model-based VC. In conventional NMF- VC, parallel data are used to train the models; therefore, unnatural pre-processing of speech data to make parallel data is needed. NMF-VC also tends to be a large model because this method has many parallel exemplars for the dictionary matrix; therefore, the computational cost is high. In this paper, we propose a novel parallel dictionary learning method using non-negative Tucker decomposition (NTD) which uses tensor decomposition and decomposes an input observation into a set of mode matrices and one core tensor. Our proposed NTD-based dictionary learning method estimates the dictionary matrix for NMF- VC without using parallel data. Experimental results show that our proposed method outperforms conventional non-parallel VC methods. Yuki Takashima, Hajime Yano, Toru Nakashika, Tetsuya Takiguchi, Yasuo Ariki |
ICASSP | 4 |
| 2018 | Oil Price Forecasting Using Supervised GANs with Continuous Wavelet Transform FeaturesabstractThis paper proposes a novel approach based on a supervised Generative Adversarial Networks (GANs) model that forecasts the crude oil prices with Adaptive Scales Continuous Wavelet Transform (AS-CWT). In our study, we first confirmed that the possibility of using Continuous Wavelet Transform (CWT) to decompose an oil price series into various components, such as the sequence of days, weeks, months and years, so that the decomposed new time series can be used as inputs for a deep-learning (DL) training model. Second, we find that applying the proposed adaptive scales in the CWT method can strengthen the dependence of inputs and provide more useful information, which can improve the forecasting performance. Finally, we use the supervised GANs model as a training model, which can provide more accurate forecasts than those of the naive forecast (NF) model and other nonlinear models, such as Neural Networks (NNs), and Deep Belief Networks (DBNs) when dealing with a limited amount of oil prices data. Zhaojie Luo, Xiao Jing Cai, Katsuyuki Tanaka, Tetsuya Takiguchi, Takuji Kinkyo, Shigeyuki Hamori |
ICPR | 5 |
| 2018 | Sound Recovery Considering the Vibration Direction of an Object in a VideoabstractWhen a sound hits an object, it causes the surface of that object to vibrate. Some research has been carried out on the recovering of sounds by extracting the vibrations that have been recorded on high-speed videos. This research is expected to be applied in the field of surveillance and security because sounds can be recorded from far away. The vibration of objects due to sound is so fast and minute that it is invisible, but it is possible to observe the changes in objects as the movement of each pixel by using the high-speed video. In this paper, we propose a sound-recovery method focusing on the vibration direction of the object. It is considered that a better sound can be obtained by trying to recover the sound based on the direction of the largest vibration. First, the sound is recovered in a certain direction (initial direction). Next, the sound is recovered in the vibration direction that has the largest correlation with the pre-first recovered sound. Repeating this process, the sound can be recovered in the direction of the largest vibration. We recovered sounds from several objects in videos and ascertained the effectiveness of the method. Yohei Fuse, Yusuke Yasumi, Tetsuya Takiguchi |
ISM | 3 |
| 2018 | Spectrum Enhancement of Singing Voice Using Deep LearningabstractIn this paper, we propose a novel singing-voice enhancement system that makes the singing voice of amateurs similar to that of professional opera singers, where the singing voice of amateurs is emphasized by using a singing voice of a professional opera singer on a frequency band that represents the remarkable characteristic of the professional singer. Moreover, our proposed singing-voice enhancement based on highway networks is able to convert any song (that a professional opera singer does not sing). As a result of our experiments, the singing voice of the amateur singer at the middle-high frequency range which contains a lot of frequency components that affect glossiness was emphasized while maintaining speaker characteristics. Ryuka Nanzaka, Tsuyoshi Kitamura, Tetsuya Takiguchi, Yuji Adachi, Kiyoto Tai |
ISM | 3 |
| 2017 | A Bayesian nonparametric multimodal data modeling framework for video emotion recognitionabstractVideo emotion recognition as an emerging research field has been attracting more and more focus in recent years. However, such work is quite challenging, since human emotions are hard to differentiate precisely due to its complexity and diversity, moreover, the expressions of sentiment in a content-rich video are sparse. Previous studies presented a number of approaches to try to learn human emotions on video level by exploiting various video features. However, most of works just used simple low-level video features such as hand-crafted image features, and they also did not consider the further latent connections among different multimodal data within a video. To tackle these problems, we develop a novel Bayesian non-parametric multimodal data modeling framework to learn the emotions from video, where the adopted image data are deep features extracted from key frames of video via convolutional neural networks (CNNs), and the adopted audio data are Mel-frequency cepstral coefficient (MFCC) features. In this framework, we then use a symmetric correspondence hierarchical Dirichlet processes (Sym-cHDP) model to mine their latent emotional events (topics) between image features and audio features. Finally, the effectiveness of our framework is demonstrated via comprehensive experimentations. Zhaojie Luo, Koji Eguchi, Tetsuya Takiguchi, Tsukasa Omoto |
ICME | 4 |
| 2017 | Phoneme-Discriminative Features for Dysarthric Speech Conversion
Ryo Aihara, Tetsuya Takiguchi, Yasuo Ariki |
INTERSPEECH | 2 |
| 2017 | Emotional Voice Conversion with Adaptive Scales F0 Based on Wavelet Transform Using Limited Amount of Emotional Data
Zhaojie Luo, Tetsuya Takiguchi, Yasuo Ariki |
INTERSPEECH | 3 |
| 2016 | Selection of an optimum random matrix using a genetic algorithm for acoustic feature extractionabstractThis paper describes a selection technique of an optimum random matrix using a genetic algorithm for speech recognition based on random projections. Random projections have been suggested as a means of dimensionality reduction, where the original data are projected onto a subspace using a random matrix. Moreover, as we are able to produce various random matrices, it may be possible to find a transform matrix that is superior to conventional transformation matrices among random matrices. In this paper, a genetic algorithm is introduced to find an optimum random matrix. Its effectiveness is confirmed by word recognition experiments. Yuichiro Kataoka, Toru Nakashika, Ryo Aihara, Tetsuya Takiguchi, Yasuo Ariki |
ICIS | 4 |
| 2016 | Lip reading using a dynamic feature of lip images and convolutional neural networksabstractIn this paper, a lip-reading method using a novel dynamic feature of lip images is proposed. The dynamic feature of lip images is calculated as the first-order regression coefficients using a few neighboring frames (images). It constiutes a better representation of the time derivatives to the basic static image. The dynamic feature is processed by using convolution neural networks (CNNs), which are able to reduce the negative influence caused by shaking of the subject and face alignment blurring at the feature-extraction level. Its effectiveness has been confirmed by word-recognition experiments comparing the proposed method with the conventional static (original) image. Yuki Takashima, Tetsuya Takiguchi, Yasuo Ariki |
ICIS | 3 |
| 2016 | Emotional voice conversion using deep neural networks with MCC and F0 featuresabstractAn artificial neural network is one of the most important models for training features in a voice conversion task. Typically, Neural Networks (NNs) are not effective in processing low-dimensional F0 features, thus this causes that the performance of those methods based on neural networks for training Mel Cepstral Coefficients (MCC) are not outstanding. However, F0 can robustly represent various prosody signals (e.g., emotional prosody). In this study, we propose an effective method based on the NNs to train the normalized-segment-F0 features (NSF0) for emotional prosody conversion. Meanwhile, the proposed method adopts deep belief networks (DBNs) to train spectrum features for voice conversion. By using these approaches, the proposed method can change the spectrum and the prosody for the emotional voice at the same time. Moreover, the experimental results show that the proposed method outperforms other state-of-the-art methods for voice emotional conversion. Zhaojie Luo, Tetsuya Takiguchi, Yasuo Ariki |
ICIS | 2 |
| 2016 | Semi-non-negative matrix factorization using alternating direction method of multipliers for voice conversionabstractVoice conversion (VC) is being widely researched in the field of speech processing because of increased interest in using such processing in applications such as personalized Text-To-Speech systems. A VC method using Non-negative Matrix Factorization (NMF) has been researched because of its natural sounding voice, however, huge memory usage and high computational times have been reported as problems. We present in this paper a new VC method using Semi-Non-negative Matrix Factorization (Semi-NMF) using the Alternating Direction Method of Multipliers (ADMM) in order to tackle the problems associated with NMF-based VC. Dictionary learning using Semi-NMF can create a compact dictionary, and ADMM enables faster convergence than conventional Semi-NMF. Experimental results show that our proposed method is 76 times faster than conventional NMF, and its conversion quality is almost the same as that of the conventional method. Ryo Aihara, Tetsuya Takiguchi, Yasuo Ariki |
ICASSP | 2 |
| 2016 | Modeling deep bidirectional relationships for image classification and generationabstractThis paper presents a novel probabilistic model that represents a joint probability of two visible variables with a deep architecture, called a deep relational model (DRM). The model stacks several layers from one visible layer on to another visible layer, sandwiching hidden layers between them. As with restricted Boltzmann machines (RBMs) and deep Boltzmann machines (DBMs), all connections (weights) between two adjacent layers are undirected. During the maximum-likelihood (ML)-based training, the network attempts to capture latent complex relationships between two visible variables (e.g., an image showing a certain number and its corresponding label) thanks to its deep architecture. Unlike deep neural networks, 1) the proposed DRM is a totally generative model, and 2) the weights can be optimized in a probabilistic manner. This paper presents and discusses the experiments conduced to evaluate our DRM's performance in recognition and generation tasks. Toru Nakashika, Tetsuya Takiguchi, Yasuo Ariki |
ICASSP | 2 |
| 2016 | Parallel Dictionary Learning for Voice Conversion Using Discriminative Graph-Embedded Non-Negative Matrix Factorization
Ryo Aihara, Tetsuya Takiguchi, Yasuo Ariki |
INTERSPEECH | 2 |
| 2016 | Audio-Visual Speech Recognition Using Bimodal-Trained Bottleneck Features for a Person with Severe Hearing Loss
Yuki Takashima, Ryo Aihara, Tetsuya Takiguchi, Yasuo Ariki, Nobuyuki Mitani, Kiyohiro Omori, Kaoru Nakazono |
INTERSPEECH | 3 |
| 2016 | Multiple Non-Negative Matrix Factorization for Many-to-Many Voice ConversionabstractA novel voice conversion (VC) method for arbitrary speakers is proposed. Non-negative matrix factorization (NMF) has recently been applied to exemplar-based VC. It offers noise robustness and naturalness of the converted voice, compared with widely used Gaussian mixture model-based VC. However, because NMF-based VC requires parallel training data from source and target speakers, the voice of arbitrary speakers cannot be converted in this framework. In this study, we propose the multiple non-negative matrix factorization (Multi-NMF) to allow the implementation of many-to-many, exemplar-based VC. Our experimental results demonstrate that the conversion quality of the proposed method is close to that of conventional one-to-one VC, even though the proposed method requires neither the source speakers' spectra, nor the target speakers' spectra, to be included in the training set. Ryo Aihara, Tetsuya Takiguchi, Yasuo Ariki |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Non-Parallel Training in Voice Conversion Using an Adaptive Restricted Boltzmann MachineabstractIn this paper, we present a voice conversion (VC) method that does not use any parallel data while training the model. VC is a technique where only speaker-specific information in source speech is converted while keeping the phonological information unchanged. Most of the existing VC methods rely on parallel data-pairs of speech data from the source and target speakers uttering the same sentences. However, the use of parallel data in training causes several problems: 1) the data used for the training are limited to the predefined sentences, 2) the trained model is only applied to the speaker pair used in the training, and 3) mismatches in alignment may occur. Although it is, thus, fairly preferable in VC not to use parallel data, a nonparallel approach is considered difficult to learn. In our approach, we achieve nonparallel training based on a speaker adaptation technique and capturing latent phonological information. This approach assumes that speech signals are produced from a restricted Boltzmann machine-based probabilistic model, where phonological information and speaker-related information are defined explicitly. Speaker-independent and speaker-dependent parameters are simultaneously trained under speaker adaptive training. In the conversion stage, a given speech signal is decomposed into phonological and speaker-related information, the speaker-related information is replaced with that of the desired speaker, and then voice-converted speech is obtained by mixing the two. Our experimental results showed that our approach outperformed another nonparallel approach, and produced results similar to those of the popular conventional Gaussian mixture models-based method that used parallel data in subjective and objective criteria. Toru Nakashika, Tetsuya Takiguchi, Yasuhiro Minami |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Facial expression recognition with multithreaded cascade of rotation-invariant HOGabstractWe propose a novel and general framework, named the multithreading cascade of rotation-invariant histograms of oriented gradients (McRiHOG) for facial expression recognition (FER). In this paper, we attempt to solve two problems about high-quality local feature descriptors and robust classifying algorithm for FER. The first solution is that we adopt annular spatial bins type HOG (Histograms of Oriented Gradients) descriptors to describe local patches. In this way, it significantly enhances the descriptors in regard to rotation-invariant ability and feature description accuracy; The second one is that we use a novel multithreading cascade to simultaneously learn multiclass data. Multithreading cascade is implemented through non-interfering boosting channels, which are respectively built to train weak classifiers for each expression. The superiority of McRiHOG over current state-of-the-art methods is clearly demonstrated by evaluation experiments based on three popular public databases, CK+, MMI, and AFEW. Tetsuya Takiguchi, Yasuo Ariki |
ACII | 2 |
| 2015 | Activity-mapping non-negative matrix factorization for exemplar-based voice conversionabstractVoice conversion (VC) is being widely researched in the field of speech processing because of increased interest in using such processing in applications such as personalized Text-To-Speech systems. We present in this paper an exemplar-based VC method us- ing Non-negative Matrix Factorization (NMF), which is different from conventional statistical VC. In our previous exemplar-based VC method, input speech is represented by the source dictionary and its sparse coefficients. The source and the target dictionaries are fully coupled and the converted voice is constructed from the source coefficients and the target dictionary. In this paper, we propose an Activity-mapping NMF approach and introduce mapping matrices between source and target sparse coefficients. The effectiveness of this method was confirmed by comparing its effectiveness with that of a conventional Gaussian Mixture Model (GMM)-based method and a conventional NMF-based method. Ryo Aihara, Tetsuya Takiguchi, Yasuo Ariki |
ICASSP | 2 |
| 2015 | Multithreading AdaBoost framework for object recognitionabstractOur research focuses on the study of effective feature description and robust classifier technique, proposing a novel learning framework, which is capable of processing multiclass objects recognition simultaneously and accurately. The framework adopts rotation-invariant histograms of oriented gradients (Ri-HOG) as feature descriptors. Most of the existing HOG techniques are computed on a dense grid of uniformly-spaced cells and use overlapping local contrast of rectangular blocks for normalization. However, we adopt annular spatial bins type cells and apply the radial gradient to attain gradient binning invariance for feature extraction. In this way, it significantly enhances HOG in regard to rotation-invariant ability and feature description accuracy; The classifier is derived from AdaBoost algorithm, but it is ameliorated and implemented through non-interfering boosting channels, which are respectively built to train weak classifiers for each object category. In this way, the boosting cascade can allow the weak classifier to be trained to fit complex distributions. The proposed method is valid on PASCAL VOC 2007 database and it achieves the state-of-the-arts performance. Tetsuya Takiguchi, Yasuo Ariki |
ICIP | 2 |
| 2015 | Sparse nonlinear representation for voice conversionabstractIn voice conversion, sparse-representation-based methods have recently been garnering attention because they are, relatively speaking, not affected by over-fitting or over-smoothing problems. In these approaches, voice conversion is achieved by estimating a sparse vector that determines which dictionaries of the target speaker should be used, calculated from the matching of the input vector and dictionaries of the source speaker. The sparse-representation-based voice conversion methods can be broadly divided into two approaches: 1) an approach that uses raw acoustic features in the training data as parallel dictionaries, and 2) an approach that trains parallel dictionaries from the training data. In our approach, we follow the latter approach and systematically estimate the parallel dictionaries using a joint-density restricted Boltzmann machine with sparse constraints. Through voice-conversion experiments, we confirmed the high-performance of our method, comparing it with the conventional Gaussian mixture model (GMM)-based approach, and a non-negative matrix factorization (NMF)-based approach, which is based on sparse representation. Toru Nakashika, Tetsuya Takiguchi, Yasuo Ariki |
ICME | 2 |
| 2015 | Individuality-Preserving Voice Reconstruction for Articulation Disorders Using Text-to-Speech SynthesisabstractThis paper presents a speech synthesis method for people with articulation disorders. Because the movements of such speakers are limited by their athetoid symptoms, their prosody is often unstable and their speech rate differs from that of a physically unimpaired person, which causes their speech to be less intelligible and, consequently, makes communication with physically unimpaired persons difficult. In order to deal with these problems, this paper describes a Hidden Markov Model(HMM)-based text-to-speech synthesis approach that preserves the individuality of a person with an articulation disorder and aids them in their communication. In our method, a duration model of a physically unimpaired person is used for the HMM synthesis system and an F0 model in the system is trained using the F0 patterns of the physically unimpaired person, with the average F0 being converted to the target F0 in advance. In order to preserve the target speaker's individuality, a spectral model is built from target spectra. Through experimental evaluations, we have confirmed that the proposed method successfully synthesizes intelligible speech while maintaining the target speaker's individuality. Reina Ueda, Tetsuya Takiguchi, Yasuo Ariki |
ICMI | 2 |
| 2015 | Word-Error Correction of Continuous Speech Recognition Based on Normalized Relevance Distance
Yohei Fusayasu, Katsuyuki Tanaka, Tetsuya Takiguchi, Yasuo Ariki |
IJCAI | 3 |
| 2015 | Many-to-many voice conversion based on multiple non-negative matrix factorizationabstractWe present in this paper an exemplar-based Voice Conversion (VC) method using Non-negative Matrix Factorization (NMF), which is different from conventional statistical VC. NMF-based VC has advantages of noise robustness and naturalness of converted voice compared to Gaussian Mixture Model (GMM)based VC. However, because NMF-based VC is based on parallel training data of source and target speakers, we cannot convert the voice of arbitrary speakers in this framework. In this paper, we propose a many-to-many VC method that makes use of Multiple Non-negative Matrix Factorization (Multi-NMF). By using Multi-NMF, an arbitrary speaker’s voice is converted to another arbitrary speaker’s voice without the need for any input or output speaker training data. We assume that this method is flexible because we can adopt it to voice quality control or noise robust VC. Index Terms: voice conversion, speech synthesis, many-tomany, exemplar-based, NMF Ryo Aihara, Tetsuya Takiguchi, Yasuo Ariki |
INTERSPEECH | 2 |
| 2015 | Content-based Image Retrieval Using Rotation-invariant Histograms of Oriented GradientsabstractOur research focuses on the question of feature descriptors for robust effective computing, proposing a novel feature representation method, namely, rotation-invariant histograms of oriented gradients (Ri-HOG) for image retrieval. Most of the existing HOG techniques are computed on a dense grid of uniformly-spaced cells and use overlapping local contrast of rectangular blocks for normalization. However, we adopt annular spatial bins type cells and apply radial gradient to attain gradient binning invariance for feature extraction. In this way, it significantly enhances HOG in regard to rotation-invariant ability and feature descripting accuracy. In experiments, the proposed method is evaluated on Corel-5k and Corel-10k datasets. The experimental results demonstrate that the proposed method is much more effective than many existing image feature descriptors for content-based image retrieval. Toru Nakashika, Tetsuya Takiguchi, Yasuo Ariki |
ICMR | 3 |
| 2015 | Voice Conversion Using RNN Pre-Trained by Recurrent Temporal Restricted Boltzmann MachinesabstractThis paper presents a voice conversion (VC) method that utilizes the recently proposed probabilistic models called recurrent temporal restricted Boltzmann machines (RTRBMs). One RTRBM is used for each speaker, with the goal of capturing high-order temporal dependencies in an acoustic sequence. Our algorithm starts from the separate training of one RTRBM for a source speaker and another for a target speaker using speaker-dependent training data. Because each RTRBM attempts to discover abstractions to maximally express the training data at each time step, as well as the temporal dependencies in the training data, we expect that the models represent the linguistic-related latent features in high-order spaces. In our approach, we convert (match) features of emphasis for the source speaker to those of the target speaker using a neural network (NN), so that the entire network (consisting of the two RTRBMs and the NN) acts as a deep recurrent NN and can be fine-tuned. Using VC experiments, we confirm the high performance of our method, especially in terms of objective criteria, relative to conventional VC methods such as approaches based on Gaussian mixture models and on NNs. Toru Nakashika, Tetsuya Takiguchi, Yasuo Ariki |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Voice conversion based on Non-negative matrix factorization using phoneme-categorized dictionaryabstractWe present in this paper an exemplar-based voice conversion (VC) method using a phoneme-categorized dictionary. Sparse representation-based VC using Non-negative matrix factorization (NMF) is employed for spectral conversion between different speakers. In our previous NMF-based VC method, source exemplars and target exemplars are extracted from parallel training data, having the same texts uttered by the source and target speakers. The input source signal is represented using the source exemplars and their weights. Then, the converted speech is constructed from the target exemplars and the weights related to the source exemplars. However, this exemplar-based approach needs to hold all the training exemplars (frames), and it may cause mismatching of phonemes between input signals and selected exemplars. In this paper, in order to reduce the mismatching of phoneme alignment, we propose a phoneme-categorized sub-dictionary and a dictionary selection method using NMF. By using the sub-dictionary, the performance of VC is improved compared to a conventional NMF-based VC. The effectiveness of this method was confirmed by comparing its effectiveness with that of a conventional Gaussian Mixture Model (GMM)-based method and a conventional NMF-based method. Ryo Aihara, Toru Nakashika, Tetsuya Takiguchi, Yasuo Ariki |
ICASSP | 3 |
| 2014 | Multimodal voice conversion using non-negative matrix factorization in noisy environmentsabstractThis paper presents a multimodal voice conversion (VC) method for noisy environments. In our previous NMF-based VC method, source exemplars and target exemplars are extracted from parallel training data, in which the same texts are uttered by the source and target speakers. The input source signal is then decomposed into source exemplars, noise exemplars obtained from the input signal, and their weights. Then, the converted speech is constructed from the target exemplars and the weights related to the source exemplars. In this paper, we propose a multimodal VC that improves the noise robustness in our NMF-based VC method. By using the joint audio-visual features as source features, the performance of VC is improved compared to a previous audio-input NMF-based VC method. The effectiveness of this method was confirmed by comparing its effectiveness with that of a conventional Gaussian Mixture Model (GMM)-based method. Kenta Masaka, Ryo Aihara, Tetsuya Takiguchi, Yasuo Ariki |
ICASSP | 3 |
| 2014 | Voice conversion in time-invariant speaker-independent spaceabstractIn this paper, we present a voice conversion (VC) method that utilizes conditional restricted Boltzmann machines (CRBMs) for each speaker to obtain time-invariant speaker-independent spaces where voice features are converted more easily than those in an original acoustic feature space. First, we train two CRBMs for a source and target speaker independently using speaker-dependent training data (without the need to parallelize the training data). Then, a small number of parallel data are fed into each CRBM and the high-order features produced by the CRBMs are used to train a concatenating neural network (NN) between the two CRBMs. Finally, the entire network (the two CRBMs and the NN) is fine-tuned using the acoustic parallel data. Through voice-conversion experiments, we confirmed the high performance of our method in terms of objective and subjective evaluations, comparing it with conventional GMM, NN, and speaker-dependent DBN approaches. Toru Nakashika, Tetsuya Takiguchi, Yasuo Ariki |
ICASSP | 2 |
| 2014 | 3D-Object Recognition Based on LLC Using Depth Spatial PyramidabstractRecently introduced high-accuracy RGB-D cameras are capable of providing high quality three-dimension information (color and depth information) easily. The overall shape of the object can be understood by acquiring depth information. However, conventional methods adopted this camera use depth information only to extract the local feature. To improve the object recognition accuracy, in our approach, the overall object shape is expressed by the depth spatial pyramid based on depth information. In more detail, multiple features within each sub-region of the depth spatial pyramid are pooled. As a result, the feature representation including the depth topological information is constructed. We use histogram of oriented normal vectors (HONV) designed to capture local geometric characteristics as 3D local features and locality-constrained linear coding (LLC) to project each descriptor into its local-coordinate system. As a result of image recognition, the proposed method has improved the recognition rate compared with conventional methods. Toru Nakashika, Takafumi Hori, Tetsuya Takiguchi, Yasuo Ariki |
ICPR | 3 |
| 2014 | Error correction of automatic speech recognition based on normalized web distanceabstractIn this paper, we focus on the problems associated with error correction of automatic speech recognition (ASR) based on confusion networks. The problems discussed are the availability of corpus in terms of calculating the semantic score and performance degradation for error correction using N -gram due to the null transitions in the confusion networks. In attempt to solve these problems, first, we employ Normalized Web Distance as a measure for semantic similarity between words that are located far from each other. The advantage of Normalized Web Distance is that it may use the Internet and so on for learning semantic similarity, which might solve the problem of corpus availability. Secondly, an error correction model without null nodes in confusion networks is trained using conditional random fields in order to improve the performance of error correction using N -grams. Index Terms: confusion network, conditional random fields, word-error correction, normalized web distance E. Byambakhishig, Katsuyuki Tanaka, Ryo Aihara, Toru Nakashika, Tetsuya Takiguchi, Yasuo Ariki |
INTERSPEECH | 5 |
| 2014 | Multimodal exemplar-based voice conversion using lip features in noisy environmentsabstractThis paper presents a multimodal voice conversion (VC) method for noisy environments. In our previous exemplarbased VC method, source exemplars and target exemplars are extracted from parallel training data, in which the same texts are uttered by the source and target speakers. The input source signal is then decomposed into source exemplars, noise exemplars obtained from the input signal, and their weights. Then, the converted speech is constructed from the target exemplars and the weights related to the source exemplars. In this paper, we propose a multimodal VC method that improves the noise robustness of our previous exemplar-based VC method. As visual features, we use not only conventional DCT but also the features extracted from Active Appearance Model (AAM) applied to the lip area of a face image. Furthermore, we introduce the combination weight between audio and visual features and formulate a new cost function in order to estimate the audiovisual exemplars. By using the joint audio-visual features as source features, the VC performance is improved compared to a previous audio-input exemplar-based VC method. The effectiveness of this method was confirmed by comparing its effectiveness with that of a conventional Gaussian Mixture Model (GMM)-based method. Index Terms: voice conversion, multimodal, image features, non-negative matrix factorization, noise robustness Kenta Masaka, Ryo Aihara, Tetsuya Takiguchi, Yasuo Ariki |
INTERSPEECH | 3 |
| 2014 | High-order sequence modeling using speaker-dependent recurrent temporal restricted boltzmann machines for voice conversionabstractThis paper presents a voice conversion (VC) method that utilizes recently proposed recurrent temporal restricted Boltzmann machines (RTRBMs) for each speaker, with the goal of capturing high-order temporal dependencies in an acoustic sequence. Our algorithm starts from the separate training of two RTRBMs for a source and target speaker using speaker-dependent training data. Since each RTRBM attempts to discover abstractions at each time step, as well as the temporal dependencies in the training data, we expect that the models represent the speaker-specific latent features in the high-order spaces. In our approach, we run conversion from such speaker-specificemphasized features of the source speaker to those of the target speaker using a neural network (NN), so that the entire network (the two RTRBMs ant the NN) forms a deep recurrent neural network and can be fine-tuned. Through VC experiments, we confirmed the high performance of our method especially in terms of objective criteria in comparison to conventional VC methods such as Gaussian mixture model (GMM)-based approaches. Toru Nakashika, Tetsuya Takiguchi, Yasuo Ariki |
INTERSPEECH | 2 |
| 2013 | Individuality-preserving voice conversion for articulation disorders based on non-negative matrix factorizationabstractWe present in this paper a voice conversion (VC) method for a person with an articulation disorder resulting from athetoid cerebral palsy. The movement of such speakers is limited by their athetoid symptoms, and their consonants are often unstable or unclear, which makes it difficult for them to communicate. In this paper, exemplar-based spectral conversion using Non-negative Matrix Factorization (NMF) is applied to a voice with an articulation disorder. To preserve the speaker's individuality, we used a combined dictionary that is constructed from the source speaker's vowels and target speaker's consonants. Experimental results indicate that the performance of NMF-based VC is considerably better than conventional GMM-based VC. Ryo Aihara, Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki |
ICASSP | 3 |
| 2013 | Sparse representation for outliers suppression in semi-supervised image annotationabstractRecently, generic object recognition (automatic image annotation) that achieves human-like vision using a computer has being looked to for use in robot vision, automatic categorization of images, and retrieval of images. For the annotation, semi-supervised learning, which incorporates a large amount of unsupervised training data (unlabeled data) along with a small amount of supervised data (labeled data), is expected to be an effective tool as it reduces the burden of manual annotation. However, some unlabeled data in semi-supervised models contains outliers that negatively affect the parameter estimation on the training stage. Such outliers often cause the over-fitting problem especially when a small amount of training data is used. In this paper, we propose a practical method to prevent the over-fitting in semi-supervised learning, suppressing existing outliers by sparse representation. In our experiments we got 4 points improvement comparing conventional semi-supervised methods, SemiNB and TSVM. Toru Nakashika, Takeshi Okumura, Tetsuya Takiguchi, Yasuo Ariki |
ICASSP | 3 |
| 2013 | Prediction of unlearned position based on local regression for single-channel talker localization using acoustic transfer functionabstractThis paper presents a sound-source (talker) localization method using only a single microphone. In our previous work, we discussed the single-channel sound-source localization method based on the discrimination of the acoustic transfer function. However, that method requires the training of the acoustic transfer function for each possible position in advance, and it is difficult to estimate the position that has not been pre-trained. In order to estimate such unlearned positions, in this paper, we discuss a single-channel talker localization method based on a regression model, which predicts the position from the acoustic transfer function. For training the regression model, we use the local regression approach, which trains the regression model from only training samples that are similar to the evaluation data. Considering both the linear and non-linear regression models, the effectiveness of this method has been confirmed by sound-source localization experiments performed in different room environments. Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki |
ICASSP | 2 |
| 2013 | Exemplar-based individuality-preserving voice conversion for articulation disorders in noisy environmentsabstractWe present in this paper a noise robust voice conversion (VC) method for a person with an articulation disorder resulting from athetoid cerebral palsy. The movements of such speakers are limited by their athetoid symptoms, and their consonants are often unstable or unclear, which makes it difficult for them to communicate. In this paper, exemplar-based spectral conversion using Non-negative Matrix Factorization (NMF) is applied to a voice with an articulation disorder in real noisy environments. In this paper, in order to deal with background noise, an input noisy source signal is decomposed into the clean source exemplars and noise exemplars by NMF. Also, to preserve the speaker’s individuality, we use a combined dictionary that was constructed from the source speaker’s vowels and target speaker’s consonants. The effectiveness of this method was confirmed by comparing its effectiveness with that of a conventional Gaussian Mixture Model (GMM)-based method. Index Terms: Voice Conversion, NMF, Articulation Disorders, Noise Robustness, Assistive Technologies Ryo Aihara, Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki |
INTERSPEECH | 3 |
| 2013 | Voice conversion in high-order eigen space using deep belief netsabstractThis paper presents a voice conversion technique using Deep Belief Nets (DBNs) to build high-order eigen spaces of the source/target speakers, where it is easier to convert the source speech to the target speech than in the traditional cepstrum space. DBNs have a deep architecture that automatically discovers abstractions to maximally express the original input features. If we train the DBNs using only the speech of an individual speaker, it can be considered that there is less phonological information and relatively more speaker individuality in the output features at the highest layer. Training the DBNs for a source speaker and a target speaker, we can then connect and convert the speaker individuality abstractions using Neural Networks (NNs). The converted abstraction of the source speaker is then brought back to the cepstrum space using an inverse process of the DBNs of the target speaker. We conducted speakervoice conversion experiments and confirmed the efficacy of our method with respect to subjective and objective criteria, comparing it with the conventional Gaussian Mixture Model-based method. Toru Nakashika, Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki |
INTERSPEECH | 3 |
| 2013 | Two-step correction of speech recognition errors based on n-gram and long contextual informationabstractThis paper presents a fully automatic word error correction on a confusion network that makes use of long contextual information. However, a problem with long contextual information is that improvement of the recognition accuracy is minimal because of the word errors surrounding words. In this paper, recognition errors are first reduced by error correction using Ngram features. After that, the long-distance context scores are applied to the correction of the residual recognition errors. Index Terms: confusion network, conditional random fields, word-error correction, long contextual information Ryohei Nakatani, Tetsuya Takiguchi, Yasuo Ariki |
INTERSPEECH | 2 |
| 2013 | Robust facial expressions recognition using 3D average face and ameliorated adaboostabstractOne of the most crucial techniques associated with Computer Vision is technology that deals with facial recognition, especially, the automatic estimation of facial expressions. However, in real-time facial expression recognition, when a face turns sideways, the expressional feature extraction becomes difficult as the view of camera changes and recognition accuracy degrades significantly. Therefore, quite many conventional methods are proposed, which are based on static images or limited to situations in which the face is viewed from the front. In this paper, a method that uses Look-Up-Table (LUT) AdaBoost combining with the three-dimensional average face is proposed to solve the problem mentioned above. In order to evaluate the proposed method, the experiment compared with the conventional method was executed. These approaches show promising results and very good success rates. This paper covers several methods that can improve results by making the system more robust. Yasuo Ariki, Tetsuya Takiguchi |
ACM Multimedia | 3 |
| 2012 | Generic object recognition by graph structural expressionabstractThis paper describes a method for generic object recognition using graph structural expression. In recent years, generic object recognition by computer is finding extensive use in a variety of fields, including robotic vision and image retrieval. Conventional methods use a bag-of-features (BoF) approach, which expresses the image as an appearance frequency histogram of visual words by quantizing SIFT (Scale-Invariant Feature Transform) features. However, there is a problem associated with this approach, namely that the location information and the relationship between keypoints (both of which are important as structural information) are lost. To deal with this problem, in the proposed method, the graph is constructed by connecting SIFT keypoints with lines. As a result, the keypoints maintain their relationship, and then structural representation with location information is achieved. Since graph representation is not suitable for statistical work, the graph is embedded into a vector space according to the graph edit distance. The experiment results on an image dataset of 10 classes showed that, the proposed method improved the recognition rate by 14.08%. Takahiro Hori, Tetsuya Takiguchi, Yasuo Ariki |
ICASSP | 2 |
| 2012 | Super-resolution by GMM based conversion using self-reduction imageabstractIn recent years, super-resolution techniques in the field of computer vision have been studied actively owing to the potential applicability in various fields. In this paper, we propose a single-image, super-resolution approach using GMM (Gaussian Mixture Model)-based conversion. The conversion function is constructed by GMM using the input image and its self-reduction image. The high-resolution image is obtained by applying the conversion function to the enlarged input image without any outside database. We confirmed the effectiveness of this proposed method through the experiments. Yuki Ogawa, Yasuo Ariki, Tetsuya Takiguchi |
ICASSP | 3 |
| 2012 | A new multiple-kernel-learning weighting method for localizing human brain magnetic activityabstractThis paper shows that pattern classification based on machine learning is a powerful tool to analyze human brain activity data obtained by magnetoencephalography (MEG). We propose a new weighting method using a multiple kernel learning (MKL) algorithm to localize the brain area contributing to the accurate vowel discrimination. Our MKL simultaneously estimates both the classification boundary and the weight of each MEG sensor; MEG amplitude obtained from each pair of sensors is an element of the feature vector. The estimated weight indicates how the corresponding sensor is useful for classifying the MEG response patterns. Our results show both the large-weight MEG sensors mainly in a language area of the brain and the high classification accuracy (73.0%) in the 100 ~ 200 ms latency range. Tetsuya Takiguchi, Toshiaki Imada, Ryoichi Takashima, Yasuo Ariki, Jo-Fu Lotus Lin, Patricia K. Kuhl, Masaki Kawakatsu, Makoto Kotani |
ICASSP | 1 |
| 2012 | Acoustic model transformations based on random projectionsabstractThis paper proposes a novel acoustic model transformation method for speech recognition based on random projections. Random projections have been suggested as a means of dimensionality reduction, where the original data are projected onto a subspace using a random matrix. Moreover, as we are able to produce various random matrices, it may be possible to find a transform matrix that is superior to conventional transformation matrices among random matrices. In our previous work, a random-projection-based feature combination technique has been proposed but had a high computational cost. In order to deal with this cost, in this paper, we introduce random projections on the acoustic model domain, where linear transformations are applied to an acoustic model using random matrices. Its effectiveness is confirmed by word recognition experiments on noisy speech. Tetsuya Takiguchi, Mariko Yoshii, Yasuo Ariki, Jeff A. Bilmes |
ICASSP | 1 |
| 2012 | 3D tracking of soccer players using time-situation graph in monocular image sequence
Hiroki Itoh, Tetsuya Takiguchi, Yasuo Ariki |
ICPR | 2 |
| 2012 | Local-feature-map Integration Using Convolutional Neural Networks for Music Genre ClassificationabstractInternational audience Toru Nakashika, Christophe Garcia, Tetsuya Takiguchi |
INTERSPEECH | 3 |
| 2012 | Estimation of Talker's Head Orientation Based on Discrimination of the Shape of Cross-power Spectrum Phase CoefficientsabstractThis paper presents a talker’s head orientation estimation method using 2-channel microphones. In recent research, some approaches based on a network of microphone arrays have been proposed in order to estimate the talker’s head orientation. In those methods, the talker’s head orientation is estimated using the sound amplitude or peak value of CSP (Cross-power Spectrum Phase) coefficients obtained from each microphone array. However, microphone array network systems need many microphone arrays to be set along the walls of a given room so that sub-microphone arrays surround the user. In this paper, we focus on the shape of the CSP coefficients affected by the reverberation, which depends on the talker’s position and the head orientation. In our proposed method, we use not only the peak value but also the other values of the CSP coefficients as feature vectors, and the talker’s position and the head orientation are estimated by discriminating the CSP vector. The effectiveness of this method has been confirmed by talker localization and head orientation estimation experiments performed in a real environment. Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki |
INTERSPEECH | 2 |
| 2012 | Super-resolution Using GMM and PLS RegressionabstractIn recent years, super-resolution techniques in the field of computer vision have been studied in earnest owing to the potential applicability of such technology in a variety of fields. In this paper, we propose a single-image, super-resolution approach using a Gaussian Mixture Model (GMM) and Partial Least Squares (PLS) regression. A GMM-based super-resolution technique is shown to be more efficient than previously known techniques, such as sparse-coding-based techniques. But the GMM-based conversion may result in over fitting. In this paper, an effective technique for preventing over fitting, which combines PLS regression with a GMM, is proposed. The conversion function is constructed using the input image and its self-reduction image. The high-resolution image is obtained by applying the conversion function to the enlarged input image without any outside database. We confirmed the effectiveness of this proposed method through our experiments. Yuki Ogawa, Takahiro Hori, Tetsuya Takiguchi, Yasuo Ariki |
ISM | 3 |
| 2012 | Robust AAM-based audio-visual speech recognition against face direction changesabstractAs one of the techniques for robust speech recognition under noisy environments, audio-visual speech recognition (AVSR) using lip dynamic scene information together with audio information is attracting attention, and the research has advanced in recent years. However, in visual speech recognition (VSR), when a face turns sideways, the shape of the lip as viewed from the camera changes and the recognition accuracy degrades significantly. Therefore, many of the conventional VSR methods are limited to situations in which the face is viewed from the front. This paper proposes a VSR method to convert faces viewed from various directions into faces that are viewed from the front using Active Appearance Models (AAM). In the experiment, even when the face direction changes about 30 degrees relative to a frontal view, the recognition accuracy improved significantly. Yuto Komai, Tetsuya Takiguchi, Yasuo Ariki |
ACM Multimedia | 3 |
| 2012 | Exemplar-based voice conversion in noisy environmentabstractThis paper presents a voice conversion (VC) technique for noisy environments, where parallel exemplars are introduced to encode the source speech signal and synthesize the target speech signal. The parallel exemplars (dictionary) consist of the source exemplars and target exemplars, having the same texts uttered by the source and target speakers. The input source signal is decomposed into the source exemplars, noise exemplars obtained from the input signal, and their weights (activities). Then, by using the weights of the source exemplars, the converted signal is constructed from the target exemplars. We carried out speaker conversion tasks using clean speech data and noise-added speech data. The effectiveness of this method was confirmed by comparing its effectiveness with that of a conventional Gaussian Mixture Model (GMM)-based method. Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki |
SLT | 2 |
| 2011 | Generic object recognition using automatic region extraction and dimensional feature integration utilizing multiple kernel learningabstractRecently, in generic object recognition research, a classification technique based on integration of image features is garnering much attention. However, with a classifying technique using feature integration, there are some features that may cause incorrect recognition of objects and a large amount of noise that causes a degradation in the recognition accuracy of image data. In this paper, we propose feature selection in an object area that is restricted by removing its back ground region, and multiple kernel learning (MKL) to weight each dimension, as well as the features themselves. This enables accurate and effective weighting since the weight is computed for each dimension using the selected feature. Experimental results indicate the validity of automatic feature selection. Classification performance is improved by using a background removing technique that utilizes saliency maps and graph cuts, and each dimensional weighting method using MKL. Toru Nakashika, Akira Suga, Tetsuya Takiguchi, Yasuo Ariki |
ICASSP | 3 |
| 2011 | Feature selection based on Multiple Kernel Learning for single-channel sound source localization using the acoustic transfer functionabstractThis paper presents a sound source (talker) localization method using only a single microphone. In our previous work [1], we discussed the single-channel sound source localization method, where the acoustic transfer function from a user's position is estimated by using a Hidden Markov Model (HMM) of clean speech in the cepstral domain. In this paper, each cepstral dimension of the acoustic transfer function is newly selected in order to select the cepstral dimensions having information that is useful for classifying the user's position. Then, we propose a feature selection method for the cepstral parameter using Multiple Kernel Learning (MKL) to define the base kernels for each cepstral dimension (scalar) of the acoustic transfer function. The user's position is trained and classified by Support Vector Machine (SVM). The effectiveness of this method has been confirmed by sound source (talker) localization experiments performed in a room environment. Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki |
ICASSP | 2 |
| 2011 | Probabilistic Spectrum Envelope: Categorized Audio-Features Representation for NMF-Based Sound DecompositionabstractNMF (Non-negative Matrix Factorization) has been one of the most useful techniques for audio signal analysis in recent years. In particular, supervised NMF, in which a large number of samples is used for analyzing a signal, is garnering much attention in sound source separation or noise reduction research. However, because such methods require all the possible samples for the analysis, it is hard to build a practical system based on this method. In this paper, we propose a novel method of signal analysis that combines the NMF and probabilistic approaches. In this approach, it is assumed that each audio-source category (such as phonemes or musical instruments) has an environment-invariant feature, called a probabilistic spectrum envelope (PSE). At the start, the PSE of each category is learned using a technique based on Gaussian Process Regression. Then, the observed spectrum is analyzed using a combination of supervised NMF and Genetic Algorithm with pre-trained PSEs. Index Terms: signal analysis, source separation, non-negative matrix factorization, probabilistic spectrum envelope, Gaussian process, genetic algorithm Toru Nakashika, Tetsuya Takiguchi, Yasuo Ariki |
INTERSPEECH | 2 |
| 2011 | Single-Channel Head Orientation Estimation Based on Discrimination of Acoustic Transfer FunctionabstractThis paper presents a talker’s head orientation estimation method using only a single microphone, where phoneme HMMs (Hidden Markov Models) of clean speech are introduced to separate the acoustic transfer function at the user’s position and head orientation. The frame sequence of the acoustic transfer function is estimated by maximizing the likelihood of training data uttered from a given position with a given head orientation. Using the separated frame sequence data, the user’s position and the head orientation are trained by Support Vector Machine (SVM) in advance. Then, for each test utterance, the frame sequence of the acoustic transfer function is separated based on the maximum likelihood estimation using the label sequence obtained from the phoneme recognition, and the user’s position and head orientation are estimated by discriminating the separated acoustic transfer function using SVM. The effectiveness of this method has been confirmed by talker localization and head orientation estimation experiments performed in a real environment. Index Terms: single channel, talker localization, head orientation, acoustic transfer function Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki |
INTERSPEECH | 2 |
| 2011 | Image Annotation with Concept Level Feature Using PLSA+CCA
Tetsuya Takiguchi, Yasuo Ariki |
MMM (2) | 2 |
| 2011 | Audio-Visual Speech Recognition Based on AAM Parameter and Phoneme Analysis of Visual Feature
Yuto Komai, Yasuo Ariki, Tetsuya Takiguchi |
PSIVT (1) | 3 |
| 2010 | HMM-based separation of acoustic transfer function for single-channel sound source localizationabstractThis paper presents a sound source (talker) localization method using only a single microphone, where a HMM (Hidden Markov Model) of clean speech is introduced to estimate the acoustic transfer function from a user's position. The new method is able to carry out this estimation without measuring impulse responses. The frame sequence of the acoustic transfer function is estimated by maximizing the likelihood of training data uttered from a given position, where the cepstral parameters are used to effectively represent useful clean speech. Using the estimated frame sequence data, the GMM (Gaussian Mixture Model) of the acoustic transfer function is created to deal with the influence of a room impulse response. Then, for each test data set, we find a maximum-likelihood GMM from among the estimated GMMs corresponding to each position. The effectiveness of this method has been confirmed by talker localization experiments performed in a room environment. Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki |
ICASSP | 2 |
| 2010 | Evaluation of random-projection-based feature combination on speech recognitionabstractRandom projection has been suggested as a means of dimensionality reduction, where the original data are projected onto a subspace using a random matrix. It represents a computationally simple method that approximately preserves the Euclidean distance of any two points through the projection. Moreover, as we are able to produce various random matrices, there may be some possibility of finding a random matrix that gives a better speech recognition accuracy among these random matrices. In this paper, we investigate the feasibility of random projection for speech feature extraction. To obtain an optimal result from among many (infinite) random matrices, a vote-based random-projection combination is introduced in this paper, where ROVER combination is applied to random-projection-based features. Its effectiveness is confirmed by word recognition experiments. Tetsuya Takiguchi, Jeff A. Bilmes, Mariko Yoshii, Yasuo Ariki |
ICASSP | 1 |
| 2010 | Structuring a gene network using a multiresolution independence testabstractIn order to structure a gene network, a score-based approach is often used. A score-based approach, however, is problematic because by assuming a probability distribution, one is prevented from finding other dependent relationships with other genes. In this research, we structured a gene network from observed gene expression data using a multiresolution independence test and a conditional independence test, which is the non-parametric method proposed by Margaritis for learning the structure of Bayesian networks without making any probability distribution assumptions. The experimental results achieved an improvement in sensitivity of 0.05, and an improvement in specificity of 0.01. Takayuki Yamamoto, Tetsuya Takiguchi, Yasuo Ariki |
ICASSP | 2 |
| 2010 | Generic Object Recognition by Tree Conditional Random Field Based on Hierarchical SegmentationabstractGeneric object recognition by a computer is strongly required in various fields like robot vision and image retrieval in recent years. Conventional methods use Conditional Random Field (CRF) that recognizes the class of each region using the features extracted from the local regions and the class co-occurrence between the adjoining regions. However, there is a problem that the discriminative ability of the features extracted from local regions is insufficient, and these methods is not robust to the scale variance. To solve this problem, we propose a method that integrates the recognition results in multi-scales by tree conditional random field based on hierarchical segmentation. As a result of the image dataset of 7 classes, the proposed method has improved the recognition rate by 2.2%. Takeshi Okumura, Tetsuya Takiguchi, Yasuo Ariki |
ICPR | 2 |
| 2010 | Speech synthesis by modeling harmonics structure with multiple functionabstractIn this paper, we present a new approach for the speech synthesis, in which speech utterances are synthesized using the parameters of spectro-modeling function (Multiple function). With this approach, only harmonic-parts are extracted from the phoneme spectrum, and the time-varying spectrum corresponding to the harmonics or sinusoidal components is modeled using the Multiple function. We introduce two types of the functions, and present the method to estimate the parameters of each function using the observed phoneme spectrum. In the synthesis stage, speech signals are generated from the parameters of the Multiple function. The advantage of this method is that it only requires a few speech synthesis parameters. We discuss the effectiveness of our proposed method through experimental results. Toru Nakashika, Ryuki Tachibana, Masafumi Nishimura, Tetsuya Takiguchi, Yasuo Ariki |
INTERSPEECH | 4 |
| 2010 | Multimodal speech recognition of a person with articulation disorders using AAM and MAFabstractWe investigated the speech recognition of a person with articulation disorders resulting from athetoid cerebral palsy. The articulation of speech tends to become unstable due to strain on speech-related muscles, and that causes degradation of speech recognition. Therefore, we use multiple acoustic frames (MAF) as an acoustic feature to solve this problem. Further, in a real environment, current speech recognition systems do not have sufficient performance due to noise influence. In addition to acoustic features, visual features are used to increase noise robustness in a real environment. However, there are recognition problems resulting from the tendency of those suffering from cerebral palsy to move their head erratically. We investigate a pose-robust audio-visual speech recognition method using an Active Appearance Model (AAM) to solve this problem for people with articulation disorders resulting from athetoid cerebral palsy. AAMs are used for face tracking to extract pose-robust facial feature points. Its effectiveness is confirmed by word recognition experiments on noisy speech of a person with articulation disorders. Chikoto Miyamoto, Yuto Komai, Tetsuya Takiguchi, Yasuo Ariki, Ichao Li |
MMSP | 3 |
| 2009 | Human Action Recognition Using HDP by Integrating Motion and Location Information
Yasuo Ariki, Takuya Tonaru, Tetsuya Takiguchi |
ACCV (2) | 3 |
| 2009 | Monaural sound-source-direction estimation using the acoustic transfer function of an active microphone
Ryoichi Takashima, Tetsuya Takiguchi, Yasuo Ariki |
FUSION | 2 |
| 2009 | System request detection in human conversation based on multi-resolution Gabor wavelet featuresabstractFor a hands-free speech interface, it is important to detect commands in spontaneous utterances. Usual voice activity detection systems can only distinguish speech frames from nonspeech frames, but they cannot discriminate whether the detected speech section is a command for a system or not. In this paper, in order to analyze the difference between system requests and spontaneous utterances, we focus on fluctuations in a long period, such as prosodic articulation, and fluctuations in a short period, such as phoneme articulation. The use of multi-resolution analysis using Gabor wavelet on a Log-scale Mel-frequency Filter-bank clarifies the different characteristics of system commands and spontaneous utterances. Experiments using our robot dialog corpus show that the accuracy of the proposed method is 92.6% in F-measure, while the conventional power and prosody-based method is just 66.7%. Index Terms: dialog system, voice activity detection, system request detection Tomoyuki Yamagata, Tetsuya Takiguchi, Yasuo Ariki |
INTERSPEECH | 2 |
| 2008 | Digital camera work for soccer video production with event recognition and accurate ball tracking by switching search methodabstractIn this paper, we propose a method of digital zooming by automatically recognizing the soccer game events such as penalty kick and free kick based on player and ball tracking. We also propose an efficient and stable ball tracking method by switching search methods between global search and local search. In the frames where the ball is lost as well as in the first frame, the global search with normalized cross-correlation is performed. Then the local search with particle filter continues the tracking. Once the system has recognized that the local search had failed by receiving a sequence of low probabilities of the ball tracking results, the search method is switched to the global search. We carried out three ball tracking experiments; global search, local search and the proposed switching search method. As a result, the experiments showed the effectiveness of the proposed method. Yasuo Ariki, Tetsuya Takiguchi, Kazuki Yano |
ICME | 2 |
| 2008 | Graph cuts by using local texture features of wavelet coefficient for image segmentationabstractThis paper proposes an approach to image segmentation using iterated graph cuts based on local texture features of wavelet coefficient. Using multiresolution analysis based on Haar wavelet, low-frequency range (smoothed image) is used for n-link and high-frequency range (local texture features) is used for t-link along with color histogram. The proposed method can segment the object region with noisy edges and colors similar to the background, but heavy texture change. Experimental results illustrate the validity of our method. Keita Fukuda, Tetsuya Takiguchi, Yasuo Ariki |
ICME | 2 |
| 2008 | 3D human posture estimation using the HOG features from monocular imageabstractIn this paper, we propose a method to estimate the 3D human posture from monocular image without using the markers. A 3D human body is expressed by a multi-joint model, and a set of the joint angles describes a posture. The proposed method estimates the posture using histograms of oriented gradients(HOG) feature vectors that can express the shape of the object in the input image obtained from monocular camera. In addition, the feature dimension of the background region is reduced for reliability by principal component analysis (PCA) computed at every block of HOG. The joint angles in human multi-joint model are estimated by linear regression analysis applied to its feature vector extracted from the input image. As a result of comparison experiment with the shape contexts features, the RMS error was reduced by about 5.35 degrees. Katsunori Onishi, Tetsuya Takiguchi, Yasuo Ariki |
ICPR | 2 |
| 2008 | Object recognition and segmentation using SIFT and Graph CutsabstractIn this paper, we propose a method of object recognition and segmentation using Scale-Invariant Feature Transform (SIFT) and Graph Cuts. SIFT feature is invariant for rotations, scale changes, and illumination changes and it is often used for object recognition. However, in previous object recognition work using SIFT, the object region is simply presumed by the affine-transformation and the accurate object region was not segmented. On the other hand, Graph Cuts is proposed as a segmentation method of a detail object region. But it was necessary to give seeds manually. By combing SIFT and Graph Cuts, in our method, the existence of objects is recognized first by vote processing of SIFT keypoints. After that, the object region is cut out by Graph Cuts using SIFT keypoints as seeds. Thanks to this combination, both recognition and segmentation are performed automatically under cluttered backgrounds including occlusion. Akira Suga, Keita Fukuda, Tetsuya Takiguchi, Yasuo Ariki |
ICPR | 3 |
| 2008 | Integration of metamodel and acoustic model for speech recognitionabstractWe investigated the speech recognition of a person with artic-ulation disorders resulting from athetoid cerebral palsy. The articulation of the first speech tends to become unstable due to strain on speech-related muscles, and that causes degradation of speech recognition. Therefore, we proposed a robust feature extraction method based on PCA (Principal Component Analy-sis) instead of MFCC [1]. In this paper, we discuss our effort to integrate a Metamodel [2] and Acoustic model approach. Meta-model has a technique for incorporating a model of a speaker’s confusion matrix into the ASR process in such a way as to increase recognition accuracy. Its effectiveness has been con-firmed by word recognition experiments. Index Terms: articulation disorders, PCA, feature extraction, model integration, dysarthric speech Hironori Matsumasa, Tetsuya Takiguchi, Yasuo Ariki, Ichao Li, Toshitaka Nakabayashi |
INTERSPEECH | 2 |
| 2008 | Sudden noise reduction based on GMM with noise power estimationabstractThis paper describes a method for reducing sudden noise using noise detection and classification methods, and noise power estimation. Sudden noise detection and classification have been dealt with in our previous study. In this paper, GMM-based noise reduction is performed using the detection and classification results. As a result of classification, we can determine the kind of noise we are dealing with, but the power is unknown. In this paper, this problem is solved by combining an estimation of noise power with the noise reduction method. In our experiments, the proposed method achieved good performance for recognition of utterances overlapped by sudden noises. Nobuyuki Miyake, Tetsuya Takiguchi, Yasuo Ariki |
INTERSPEECH | 2 |
| 2008 | CENSREC-4: development of evaluation framework for distant-talking speech recognition under reverberant environmentsabstractIn this paper, we newly introduce a collection of databases and evaluation tools called CENSREC-4, which is an evaluation framework for distant-talking speech under hands-free conditions. Distant-talking speech recognition is crucial for a handsfree speech interface. Therefore, we measured room impulse responses to investigate reverberant speech recognition in various environments. The data contained in CENSREC-4 are connected digit utterances, as in CENSREC-1. Two subsets are included in the data: basic data sets and extra data sets. The basic data sets are used for the evaluation environment for the room impulse response-convolved speech data. The extra data sets consist of simulated and recorded data. An evaluation framework is only provided for the basic data sets as evaluation tools. The results of evaluation experiments proved that CENSREC-4 is an effective database for evaluating the new dereverberation method because the traditional dereverberation process had difficulty sufficiently improving the recognition performance. Index Terms: Various environments, Impulse response, Convolution, Real recorded data, Evaluation framework Masato Nakayama, Takanobu Nishiura, Yuki Denda, Norihide Kitaoka, Kazumasa Yamamoto, Takeshi Yamada, Satoru Tsuge, Chiyomi Miyajima, Masakiyo Fujimoto, Tetsuya Takiguchi, Satoshi Tamura, Tetsuji Ogawa, Shigeki Matsuda, Shingo Kuroiwa, Kazuya Takeda, Satoshi Nakamura 0001 |
INTERSPEECH | 10 |
| 2008 | Evaluation Framework for Distant-talking Speech Recognition under Reverberant Environments: newest Part of the CENSREC Series -
Takanobu Nishiura, Masato Nakayama, Yuki Denda, Norihide Kitaoka, Kazumasa Yamamoto, Takeshi Yamada, Satoru Tsuge, Chiyomi Miyajima, Masakiyo Fujimoto, Tetsuya Takiguchi, Satoshi Tamura, Shingo Kuroiwa, Kazuya Takeda, Satoshi Nakamura 0001 |
LREC | 10 |
| 2008 | Tagging Video Contents with Positive/Negative Interest Based on User's Facial Expression
Masanori Miyahara, Masaki Aoki, Tetsuya Takiguchi, Yasuo Ariki |
MMM | 3 |
| 2007 | Development of VAD evaluation framework CENSREC-1-C and investigation of relationship between VAD and speech recognition performanceabstractVoice activity detection (VAD) plays an important role in speech processing including speech recognition, speech enhancement, and speech coding in noisy environments. We developed an evaluation framework for VAD in such environments, called corpus and environment for noisy speech recognition 1 concatenated (CENSREC-1-C). This framework consists of noisy continuous digit utterances and evaluation tools for VAD results. By adoptiong two evaluation measures, one for frame-level detection performance and the other for utterance-level detection performance, we provide the evaluation results of a power-based VAD method as a baseline. When using VAD in speech recognizer, the detected speech segments are extended to avoid the loss of speech frames and the pause segments are then absorbed by a pause model. We investigate the balance of an explicit segmentation by VAD and an implicit segmentation by a pause model using an experimental simulation of segment extension and show that a small extension improves speech recognition. Norihide Kitaoka, Kazumasa Yamamoto, Tomohiro Kusamizu, Seiichi Nakagawa, Takeshi Yamada, Satoru Tsuge, Chiyomi Miyajima, Takanobu Nishiura, Masato Nakayama, Yuki Denda, Masakiyo Fujimoto, Tetsuya Takiguchi, Satoshi Tamura, Shingo Kuroiwa, Kazuya Takeda, Satoshi Nakamura 0001 |
ASRU | 12 |
| 2007 | PCA-based feature extraction for fluctuation in speaking style of articulation disordersabstractWe investigated the speech recognition of a person with articulation disorders resulting from athetoid cerebral palsy. Recently, the accuracy of speaker-independent speech recognition has been remarkably improved by the use of stochastic modeling of speech. However, the use of those acoustic models causes degradation of speech recognition for a person with different speech styles (e.g., articulation disorders). In this paper, we discuss our efforts to build an acoustic model for a person with articulation disorders. The articulation of the first speech tends to become unstable due to strain on muscles and that causes degradation of speech recognition. Therefore, we propose a robust feature extraction method based on PCA (Principal Component Analysis) instead of MFCC. Its effectiveness is confirmed by word recognition experiments. Index Terms: articulation disorders, PCA, feature extraction Hironori Matsumasa, Tetsuya Takiguchi, Yasuo Ariki, Ichao Li, Toshitaka Nakabayashi |
INTERSPEECH | 2 |
| 2007 | Language modeling using PLSA-based topic HMMabstractIn this paper, we propose a PLSA-based language model for sports-related live speech. This model is implemented using a unigram rescaling technique that combines a topic model and an n-gram. In the conventional method, unigram rescaling is performed with a topic distribution estimated from a recognized transcription history. This method can improve the performance, but it cannot express topic transition. By incorporating the concept of topic transition, it is expected that the recognition performance will be improved. Thus, the proposed method employs a “Topic HMM” instead of a history to estimate the topic distribution. The Topic HMM is an Ergodic HMM that expresses typical topic distributions as well as topic transition probabilities. Word accuracy results from our experiments confirmed the superiority of the proposed method over a trigram and a PLSA-based conventional method that uses a recognized history. Atsushi Sako, Tetsuya Takiguchi, Yasuo Ariki |
INTERSPEECH | 2 |
| 2007 | System request detection in conversation based on acoustic and speaker alternation featuresabstractFor a hands-free speech interface, it is important to detect com-mands in spontaneous utterances. To discriminate commands from human-human conversations by acoustic features, it is ef-ficient to consider the head and the tail of an utterance. The dif-ferent characteristics of system requests and spontaneous utter-ances appear on these parts of an utterance. Experiment shows that by separating the head and the tail of an utterance, the ac-curacy of detection was improved. And also, considering the al-ternation of speakers using two channel microphones improved the performance. Although detecting system requests using lin-guistic features shows high accuracy, combining acoustic and turn-taking features lift up the performance. Index Terms: system request detection, utterance verification, SVM, speech recognition, turn-taking Tomoyuki Yamagata, Atsushi Sako, Tetsuya Takiguchi, Yasuo Ariki |
INTERSPEECH | 3 |
| 2007 | Voice activity detection by lip shape tracking using EBGMabstractWe propose a voice activity detection of a target speaker (driver) in a car by integrating lip movement and acoustic processing. To prevent the wrong detection caused by nontarget speakers using only acoustic processing, the proposed system extracts the lip movement of the target speaker by measuring the lip aspect ratio. An infrared camera is used to cope with the change of lighting environment. In order to extract the lip from gray scale images, Elastic Bunch Graph Matching is employed. Experimental results showed the proposed system improved the precision rate in the voice activity detection by approximately 40% compared to the method using only acoustic processing in a car. Masaki Aoki, Ken Masuda, Hiroyoshi Matsuda, Tetsuya Takiguchi, Yasuo Ariki |
ACM Multimedia | 4 |
| 2006 | Robust Feature Extraction using Kernel PCAabstractWe investigate a robust speech feature extraction method using kernel PCA (principal component analysis). Kernel PCA has been suggested for various image processing tasks requiring an image model such as, e.g., denoising, where a noise-free image is constructed from a noisy input image. Much research for robust speech feature extraction has been done, but it is difficult to completely remove the non-stationary noise or reverberation. The most commonly used noise-removal techniques are based on the spectral-domain operation, and then for the speech recognition, MFCC (mel frequency cepstral coefficient) is computed, where DCT (discrete cosine transform) is applied to the mel-scale filter bank output. In this paper, we propose robust feature extraction based on kernel PCA instead of DCT, where the main speech element is projected onto low-order features, while noise or reverberant element is projected onto high-order ones. Its effectiveness is confirmed by word recognition experiments on reverberant speech Tetsuya Takiguchi, Yasuo Ariki |
ICASSP (1) | 1 |
| 2006 | Phoneme recognition based on fisher weight map to higher-order local auto-correlationabstractIn this paper, we propose a new feature extraction method based on higher-order local auto-correlation (HLAC) and Fisher weight map (FWM). Widely used MFCC features lack temporal dynamics. To solve this problem, 35 types of local auto-correlation features are computed within two-dimensional local regions. These local features are accumulated over more global regions by weighting high scores on the discriminative areas where the typical features among all phonemes are well expressed. This score map is called Fisher weight map. We verified the effectiveness of the HLAC and FWM through vowel recognition and total phoneme recognition. Yasuo Ariki, Shunsuke Kato 0001, Tetsuya Takiguchi |
INTERSPEECH | 3 |
| 2005 | Situation based speech recognition for structuring baseball live gamesabstractIt is a difficult problem to recognize baseball live speech because the speech is rather fast, noisy, emotional and disfluent due to rephrasing, repetition, mistake and grammatical deviation caused by spontaneous speaking style. To solve these problems, we have been studied the speech recognition method incorporating the baseball game task-dependent knowledge as well as an announcer’s emotion in commentary speech [1]. In addition, in this paper, we propose the situation prediction model based on word co-occurrence. Owing to these proposed models, speech recognition errors are effectively prevented. This method is formalized in the framework of probability theory and implemented in the conventional speech decoding (Viterbi) algorithm. The experimental results showed that the proposed approach improved the structuring and segmentation accuracy as well as keywords accuracy. Atsushi Sako, Tetsuya Takiguchi, Yasuo Ariki |
INTERSPEECH | 2 |
| 2005 | Recognition of hands-free speech and hand pointing action for conversational TVabstractIn this paper, we propose a structure and components of a conversational television set(TV) to which we can ask anything on the broadcasted contents and receive the interesting information from the TV. The conversational TV is composed of two types of processing; back end processing and front end processing. In the back end processing, broadcasted contents are analyzed using speech and video recognition techniques and both of the meta data and the structure are extracted. In the front end processing, human speech and hand action are recognized to understand the user intention. We show some applications, being developed in this conversational TV with multi-modal interactions, such as word explanation, human information retrieval, event retrieval in soccer and baseball video games with contextual awareness. Yasuo Ariki, Tetsuya Takiguchi, Atsushi Sako |
ACM Multimedia | 2 |
| 2004 | Acoustic model adaptation using first order prediction for reverberant speechabstractThe paper describes a hands-free speech recognition technique based on acoustic model adaptation to reverberant speech. In hands-free speech recognition, the recognition accuracy is degraded by reverberation, since each segment of speech is affected by the reflection energy of the preceding segment. To compensate for the reflection signal, we introduce a frame-by-frame adaptation method, adding the reflection signal to the means of the acoustic model. The reflection signal is approximated by a first-order linear prediction from the preceding frame, and the linear prediction coefficient is estimated by a maximum likelihood method by using the EM algorithm, which maximizes the likelihood of the adaptation data. Its effectiveness is confirmed by word recognition experiments on reverberant speech. Tetsuya Takiguchi, Masafumi Nishimura |
ICASSP (1) | 1 |
| 2001 | HMM-separation-based speech recognition for a distant moving speakerabstractThis paper presents a hands-free speech recognition method based on HMM composition and separation for speech contaminated not only by additive noise but also by an acoustic transfer function. The method realizes an improved user interface such that a user is not encumbered by microphone equipment in noisy and reverberant environments. The use of HMM composition has already been proposed for countering additive noise. In this paper, the same approach is extended to handle convolutional acoustic distortion in a reverberant room, by using an HMM to model the acoustic transfer function. The states of this HMM correspond to different positions of the sound source. It can represent the positions of the sound sources, even if the speaker moves. This paper also proposes a new method, HMM separation, for estimating the HMM parameters of the acoustic transfer function on the basis of a maximum likelihood manner. The proposed method is obtained through the reverse of the process of HMM composition, where the model parameters are estimated by maximizing the likelihood of adaptation data uttered from an unknown position. Therefore, measurement of impulse responses is not required. The paper also describes the performance of the proposed methods for recognizing real distant-talking speech. The results of experiments clarify the effectiveness of the proposed method. Tetsuya Takiguchi, Satoshi Nakamura 0001, Kiyohiro Shikano |
IEEE Trans. Speech Audio Process. | 1 |
| 2000 | Speech recognition for a distant moving speaker based on HMM composition and separationabstractThis paper describes a hands-free speech recognition method based on HMM composition and separation for speech contaminated not only by additive noise but also by an acoustic transfer function. The method realizes an improved user interface such that a user is not encumbered by microphone equipment in noisy and reverberant environments. In this approach, an attempt is made to model acoustic transfer functions by means of an ergodic HMM. The states of this HMM correspond to different positions of the sound source. It can represent the positions of the sound sources, even if the speaker moves. The HMM parameters of the acoustic transfer function are estimated by HMM separation. The method is obtained through the reverse of the process of HMM composition, where the model parameters are estimated by maximizing the likelihood of adaptation data uttered from an unknown position. Therefore, measurement of impulse responses is not required. In this paper, we record the speech of a distant moving speaker in real environments. The results of experiments for the speech of a distant moving speaker clarified the effectiveness of HMM composition and separation. Tetsuya Takiguchi, Satoshi Nakamura 0001, Kiyohiro Shikano |
ICASSP | 1 |
| 1998 | Evaluation of model adaptation by HMM decomposition on telephone speech recognitionabstractIn this paper, we evaluate performance of model adaptation by the previously proposed HMM decomposition method on telephone speech recognition. The HMM decomposition method separates a composed HMM into a known phoneme HMM and an unknown noise and channel HMM by maximum likelihood (ML) estimation of the HMM parameters. A transfer function (telephone channel) HMM is estimated using adaptation speech data by applying the HMM decomposition twice in the linear spectral domain for noise and in the cepstral domain for channel. The telephone speech data for evaluation are recorded through 10 kinds of ordinary analog telephone handsets and cordless telephone handsets. The test results show that the average phrase accuracy with the clean speech HMMs is 60.9% for the ordinary analog telephone handsets, and 19.6% for the cordless telephone handsets. By the HMM decomposition method, the average phrase accuracy is improved to 78.1% for the ordinary analog telephone handsets, and 50.5% for the cordless telephone handsets. Tetsuya Takiguchi, Satoshi Nakamura 0001, Kiyohiro Shikano, Masatoshi Morishima, Toshihiro Isobe |
ICSLP | 1 |
| 1997 | Model adaptation based on HMM decomposition for reverberant speech recognitionabstractThe performance of a speech recognizer is degraded drastically in reverberant environments. The authors propose a novel algorithm which can model an observation signal by composition of HMMs of clean speech, noise and an acoustic transfer function. However, estimating HMM parameters of the acoustic transfer function is still a serious problem. In their previous paper, they measured real impulse responses of training positions in an experiment room. It is inconvenient and unrealistic to measure impulse responses for every possible new experiment room. The paper presents a new method for estimating HMM parameters of the acoustic transfer function from some adaptation data by using an HMM decomposition algorithm which is an inverse process of the HMM composition. Its effectiveness is confirmed by a series of speaker dependent and independent word recognition experiments on simulated distant-talking speech data. Tetsuya Takiguchi, Satoshi Nakamura 0001, Kiyohiro Shikano |
ICASSP | 1 |
| 1996 | Noise and room acoustics distorted speech recognition by HMM compositionabstractThis paper presents a robust speech recognition method based on the HMM composition for the noisy room acoustics distorted speech. The method realizes an improved user interface such as the user is not encumbered by microphone equipment. The proposed HMM composition is obtained by naturally extending the HMM composition method of an additive noise to that of the convolutional room acoustics distortion. The HMM composition is conducted by 2 steps: (1) composition of HMMs of a speech and acoustical transfer function in the cepstrum domain, and (2) composition of distorted speech and noise HMMs in the linear spectral domain. The speaker dependent/independent word recognition experiments are carried out using the speech database contaminated by the additive noise and convolutional room acoustics distortion. The evaluation experiments are also conducted for unknown testing sound source positions. These results clarified the effectiveness of the proposed method. Satoshi Nakamura 0001, Tetsuya Takiguchi, Kiyohiro Shikano |
ICASSP | 2 |