EDBT 2026 Demo / reviewers in the wild / expert
Satoshi Nakamura 0001
dblp:57/1548-1
· DBLP profile ↗
401ranked-venue papers
19as first author
37since 2021 · last 2025
0000-0001-6956-3803ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 285 · 14 first-author · 18 since 2021Artificial intelligence and machine learning · 267 · 10 first-author · 28 since 2021Human-computer interaction and ubiquitous computing · 16 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 10 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Systems, architecture and hardware · 2Software engineering, systems software and programming languages · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning Marmoset Vocal Patterns with a Masked Autoencoder for Robust Call Segmentation, Classification, and Caller IdentificationabstractThe marmoset, a highly vocal primate, is a key model for studying social-communicative behavior. Unlike human speech, marmoset vocalizations are less structured, highly variable, and recorded in noisy, low-resource conditions. Learning marmoset communication requires joint call segmentation, classification, and caller identification-challenging domain tasks. Previous CNNs handle local patterns but struggle with long-range temporal structure. We applied Transformers using self-attention for global dependencies. However, Transformers show overfitting and instability on small, noisy annotated datasets. To address this, we pretrain Transformers with MAE-a self-supervised method reconstructing masked segments from hundreds of hours of unannotated marmoset recordings. The pretraining improved stability and generalization. Results show MAE-pretrained Transformers outperform CNNs, demonstrating modern self-supervised architectures effectively model low-resource non-human vocal communication. Shinnosuke Takamichi, Sakriani Sakti, Satoshi Nakamura 0001 |
ASRU | 4 |
| 2024 | NAIST-SIC-Aligned: An Aligned English-Japanese Simultaneous Interpretation CorpusabstractIt remains a question that how simultaneous interpretation (SI) data affects simultaneous machine translation (SiMT). Research has been limited due to the lack of a large-scale training corpus. In this work, we aim to fill in the gap by introducing NAIST-SIC-Aligned, which is an automatically-aligned parallel English-Japanese SI dataset. Starting with a non-aligned corpus NAIST-SIC, we propose a two-stage alignment approach to make the corpus parallel and thus suitable for model training. The first stage is coarse alignment where we perform a many-to-many mapping between source and target sentences, and the second stage is fine-grained alignment where we perform intra- and inter-sentence filtering to improve the quality of aligned pairs. To ensure the quality of the corpus, each step has been validated either quantitatively or qualitatively. This is the first open-sourced large-scale parallel SI dataset in the literature. We also manually curated a small test set for evaluation purposes. Our results show that models trained with SI data lead to significant improvement in translation quality and latency over baselines. We hope our work advances research on SI corpora construction and SiMT. Our data will be released upon the paper’s acceptance. Jinming Zhao, Katsuhito Sudoh, Satoshi Nakamura 0001, Yuka Ko, Kosuke Doi, Ryo Fukuda |
LREC/COLING | 3 |
| 2024 | LLMs Are Zero-Shot Context-Aware Simultaneous TranslatorsabstractThe advent of transformers has fueled progress in machine translation.More recently large language models (LLMs) have come to the spotlight thanks to their generality and strong performance in a wide range of language tasks, including translation.Here we show that open-source LLMs perform on par with or better than some state-of-the-art baselines in simultaneous machine translation (SiMT) tasks, zero-shot.We also demonstrate that injection of minimal background information, which is easy with an LLM, brings further performance gains, especially on challenging technical subject-matter.This highlights LLMs' potential for building next generation of massively multilingual, context-aware and terminologically accurate SiMT systems that require no resource-intensive training or fine-tuning.The code is available at https://github.com/RomanKoshkin/toLLMatch. Roman Koshkin, Katsuhito Sudoh, Satoshi Nakamura 0001 |
EMNLP | 3 |
| 2024 | Subspace Representations for Soft Set Operations and Sentence SimilaritiesabstractYoichi Ishibashi, Sho Yokoi, Katsuhito Sudoh, Satoshi Nakamura. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Yoichi Ishibashi, Sho Yokoi, Katsuhito Sudoh, Satoshi Nakamura 0001 |
NAACL-HLT | 4 |
| 2024 | Continual few-shot patch-based learning for anime-style colorizationabstractThe automatic colorization of anime line drawings is a challenging problem in production pipelines. Recent advances in deep neural networks have addressed this problem; however, collectingmany images of colorization targets in novel anime work before the colorization process starts leads to chicken-and-egg problems and has become an obstacle to using them in production pipelines. To overcome this obstacle, we propose a new patch-based learning method for few-shot anime-style colorization. The learning method adopts an efficient patch sampling technique with position embedding according to the characteristics of anime line drawings. We also present a continuous learning strategy that continuously updates our colorization model using new samples colorized by human artists. The advantage of our method is that it can learn our colorization model from scratch or pre-trained weights using only a few pre- and post-colorized line drawings that are created by artists in their usual colorization work. Therefore, our method can be easily incorporated within existing production pipelines. We quantitatively demonstrate that our colorizationmethod outperforms state-of-the-art methods. Akinobu Maejima, Seitaro Shinagawa, Hiroyuki Kubo, Takuya Funatomi, Tatsuo Yotsukura, Satoshi Nakamura 0001, Yasuhiro Mukaigawa |
Comput. Vis. Media | 6 |
| 2024 | Adaptive virtual agent: Design and evaluation for real-time human-agent interactionabstractWhen we converse, we adapt our behaviors to our interlocutors. The adaptation can serve to indicate our engagement which can also elicit enhancement of the involvement of others. Virtual agents (or socially interactive virtual agents) that play the role of interaction partners can improve the human users’ interaction experience by displaying continuous and adaptive behaviors in real time. Virtual agents have been used in multiple domains to improve user interaction and performance. The promising results of the endowment of adaptation to agents in increasing the agents’ perception and user experience were shown in previous studies. In this paper, we develop an adaptive virtual agent that renders real-time adaptive behaviors based on the behaviors shown by its human interlocutor. The ASAP model rendering reciprocally adaptive agent behavior was employed to realize the system. The system consists of four main parts: perception of social signals, agent adaptive behavior generation, agent visualization (i.e. rendering of the agent’s verbal and nonverbal behavior), and communication of signals. To showcase the usefulness of our adaptive agent, as a proof-of-concept we choose the e-health application of cognitive behavior therapy (CBT), which identifies and rectifies biased and irrational thoughts (or automatic thoughts). Through this study, we show the importance of giving the agent reciprocal adaptation capability notably in enhancing the user experience and the effectiveness of the CBT session. We validate the importance of endowing such adaptation capability by studying the difference between agents that are reciprocal adaptive, solely expressive (with mismatched behavior), and inexpressive (in a still posture) via questionnaires and measures related to the agent perception (naturalness, human-likeliness, synchrony, and engagement) for user experience and the CBT effectiveness (mood, anxiety, stress, and cognitive change). These results highlight the value of making virtual agents adapt in real time. This could lead to agents being capable of providing more personalized and interactive experiences for a wide range of applications. Also, we have collected a new human-agent interaction (HAI) database, HAI-CBT database, which is publicly available to the research community. Jieyeon Woo, Kazuhiro Shidara, Catherine Achard, Hiroki Tanaka, Satoshi Nakamura 0001, Catherine Pelachaud |
Int. J. Hum. Comput. Stud. | 5 |
| 2024 | Improving Speech Translation Accuracy and Time Efficiency With Fine-Tuned wav2vec 2.0-Based Speech SegmentationabstractSpeech translation (ST) automatically converts utterances in a source language into text in another language. Splitting continuous speech into shorter segments, known as speech segmentation, plays an important role in ST. Recent segmentation methods trained to mimic the segmentation of ST corpora have surpassed traditional approaches. Tsiamas et al. [1] proposed a segmentation frame classifier (SFC) based on a pre-trained speech encoder called wav2vec 2.0. Their method, named SHAS, retains 95-98% of the BLEU score for ST corpus segmentation. However, the segments generated by SHAS are very different from ST corpus segmentation and tend to be longer with multiple combined utterances. This is due to SHAS's reliance on length heuristics, i.e., it splits speech into segments of easily translatable length without fully considering the potential for ST improvement by splitting them into even shorter segments. Longer segments often degrade translation quality and ST's time efficiency. In this study, we extended SHAS to improve ST translation accuracy and efficiency by splitting speech into shorter segments that correspond to sentences. We introduced a simple segmentation avlgorithm using the moving average of SFC predictions without relying on length heuristics and explored wav2vec 2.0 fine-tuning for improved speech segmentation prediction. Our experimental results reveal that our speech segmentation method significantly improved the quality and the time efficiency of speech translation compared to SHAS. Ryo Fukuda, Katsuhito Sudoh, Satoshi Nakamura 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Evaluating the Robustness of Discrete PromptsabstractDiscrete prompts have been used for finetuning Pre-trained Language Models for diverse NLP tasks.In particular, automatic methods that generate discrete prompts from a small set of training instances have reported superior performance.However, a closer look at the learnt prompts reveals that they contain noisy and counter-intuitive lexical constructs that would not be encountered in manuallywritten prompts.This raises an important yet understudied question regarding the robustness of automatically learnt discrete prompts when used in downstream tasks.To address this question, we conduct a systematic study of the robustness of discrete prompts by applying carefully designed perturbations into an application using AutoPrompt and then measure their performance in two Natural Language Inference (NLI) datasets.Our experimental results show that although the discrete prompt-based method remains relatively robust against perturbations to NLI inputs, they are highly sensitive to other types of perturbations such as shuffling and deletion of prompt tokens.Moreover, they generalize poorly across different NLI datasets.We hope our findings will inspire future work on robust discrete prompt learning.1 Yoichi Ishibashi, Danushka Bollegala, Katsuhito Sudoh, Satoshi Nakamura 0001 |
EACL | 4 |
| 2023 | Acceptability and Trustworthiness of Virtual Agents by Effects of Theory of Mind and Social Skills TrainingabstractWe constructed a social skills training system using virtual agents and developed a new training module for four basic tasks: declining, requesting, praising, and listening. Previous work demonstrated that a virtual agent's theory of mind influences the building of trust between agents and users. The purpose of this study is to explore the effect of trustworthiness, acceptability, familiarity, and likeability on the agents' theory of mind and the social skills training contents. In our experiment, 29 participants rated the trustworthiness and acceptability of the virtual agent after watching a video that featured levels of theory of mind and social skills training. Their ratings were obtained using self-evaluation measures at each stage. We confirmed that our users' trust and acceptability of the virtual agent were significantly changed depending on the level of the virtual agent's theory of mind. We also confirmed that the users' trust and acceptability in the trainer tended to improve after the social skills training. Hiroki Tanaka, Takeshi Saga, Kota Iwauchi, Satoshi Nakamura 0001 |
FG | 4 |
| 2023 | Multimodal Voice Activity Prediction: Turn-taking Events Detection in Expert-Novice ConversationabstractPredicting the timing of utterances in dyadic conversations is essential for achieving natural interactions between humans and virtual agents. Since the former often use non-verbal cues to adjust the order of their speech, this study proposes a multimodal model incorporating non-verbal features using a Transformer-based voice activity prediction model. First, in line with previous research, we reproduced a baseline model that utilized audio features (audio waveform, voice activity frame, and voice activity history) as inputs. To this baseline model, we added non-verbal features: gaze direction, action units, head pose, and articular points. We compared our multimodal model with the baseline model to investigate the impact of non-verbal cues on voice activity prediction. We utilized a dyadic expert-novice conversation dataset and evaluated the average outcomes across ten model trainings. Results revealed that our proposed models with all the features improved the accuracy of the next speaker prediction by 2.3% and back-channeled prediction by 1.8% (p-value < 0.025). In particular, action units may contribute significantly to the turn-shift and back-channeled predictions. This study demonstrates that including non-verbal features in Transformer-based turn-taking models enhances the efficacy of models for predicting voice activity in dyadic conversations. Kazuyo Onishi, Hiroki Tanaka, Satoshi Nakamura 0001 |
HAI | 3 |
| 2023 | Self-Adaptive Incremental Machine Speech Chain for Lombard TTS with High-Granularity ASR Feedback in Dynamic Noise ConditionabstractA common approach for text-to-speech (TTS) in noisy conditions is offline fine-tuning, which is generally utilized on static noises and predefined conditions. We recently proposed a self-adaptive TTS in machine speech chain inference that enables TTS to control its voices in statically and dynamically noisy environments based on auditory feedback from automatic speech recognition (ASR) and speech-to-noise ratio (SNR) recognition. However, that study only investigated the system on synthetic Lombard speech data. Furthermore, the ASR feedback was at a lower granularity based only on the loss of the positive character class. In this paper, we improve the self-adaptive TTS using character-vocabulary level ASR feedback at higher granularity, considering the losses in the positive and negative classes. We focus on a self-adaptive incremental TTS (Adapt-ITTS) with a short-term feedback mechanism that aims for low latency adaptation for dynamically noisy situations. In experiments, our proposed Adapt- ITTS successfully improved intelligibility in noisy conditions based on synthetic and natural Lombard speech data on the Wall Street Journal and Hurricane datasets, respectively. Sashi Novitasari, Sakriani Sakti, Satoshi Nakamura 0001 |
ICASSP | 3 |
| 2023 | Average Token Delay: A Latency Metric for Simultaneous Translation
Yasumasa Kano, Katsuhito Sudoh, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2023 | Inter-connection: Effective Connection between Pre-trained Encoder and Decoder for Speech Translation
Yuta Nishikawa, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2023 | Reflective action selection based on positive-unlabeled learning and causality detection modelabstractTask-oriented dialogue systems need to take appropriate actions not only for clear user requests but also for ambiguous and vague ones. In this study, “ambiguous” denotes that although users have potential requests, they failed to clearly define and verbalize their content and conditions which can be associated with system actions. For such ambiguous requests, taking reflective actions is one plausible choice for such systems. In our study, “reflective” denotes taking actions that satisfy user requests before the users themselves clarify their demands. We constructed such a reflective dialogue agent by collecting a corpus that includes pairs of ambiguous user requests and corresponding reflective system actions on sightseeing navigation with a smartphone. Since annotating every possible combination of user requests and system actions is impossible, this study built a corpus where one reflective action is annotated to one ambiguous user request. To train an action selection model on such incomplete training data in which only one action is associated with a request, we applied the positive/unlabeled (PU) learning method, which assumes that only part of the data is labeled with positive examples. In addition, we enhanced the action selection by extracting and distilling knowledge that corresponds to causality from the training data using a causality detection model. The experimental results show that both the PU learning method and the causality detection model improved the performances of the reflective action selection compared to the conventional positive/negative (PN) learning method. Shohei Tanaka, Koichiro Yoshino, Katsuhito Sudoh, Satoshi Nakamura 0001 |
Comput. Speech Lang. | 4 |
| 2022 | 3rd Workshop on Social Affective Multimodal Interaction for Health (SAMIH)abstractThis workshop discusses how interactive, multimodal technology such as virtual agents can be used in social skills training for measuring and training social-affective interactions. Sensing technology now enables analyzing user’s behaviors and physiological signals. Various signal processing and machine learning methods can be used for such prediction tasks. Such social signal processing and tools can be applied to measure and reduce social stress in everyday situations, including public speaking at schools and workplaces. Hiroki Tanaka, Satoshi Nakamura 0001, Kazuhiro Shidara, Jean-Claude Martin, Catherine Pelachaud |
ICMI | 2 |
| 2022 | Speech Segmentation Optimization using Segmented Bilingual Speech Corpus for End-to-end Speech TranslationabstractSpeech segmentation, which splits long speech into short segments, is essential for speech translation (ST).Popular VAD tools like WebRTC VAD 1 have generally relied on pause-based segmentation.Unfortunately, pauses in speech do not necessarily match sentence boundaries, and sentences can be connected by a very short pause that is difficult to detect by VAD.In this study, we propose a speech segmentation method using a binary classification model trained using a segmented bilingual speech corpus.We also propose a hybrid method that combines VAD and the above speech segmentation method.Experimental results reveal that the proposed method is more suitable for cascade and end-to-end ST systems than conventional segmentation methods.The hybrid approach further improves the translation performance. Ryo Fukuda, Katsuhito Sudoh, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2022 | Applying Syntax-Prosody Mapping Hypothesis and Prosodic Well-Formedness Constraints to Neural Sequence-to-Sequence Speech Synthesis
Kei Furukawa, Takeshi Kishiyama, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2022 | Multimodal Persuasive Dialogue Corpus using Teleoperated Android
Seiya Kawano, Muteki Arioka, Akishige Yuguchi, Kenta Yamamoto, Koji Inoue, Tatsuya Kawahara, Satoshi Nakamura 0001, Koichiro Yoshino |
INTERSPEECH | 7 |
| 2022 | Improved Consistency Training for Semi-Supervised Sequence-to-Sequence ASR via Speech Chain Reconstruction and Self-TranscribingabstractConsistency regularization has recently been applied to semi-supervised sequence-to-sequence (S2S) automatic speech recognition (ASR).This principle encourages an ASR model to output similar predictions for the same input speech with different perturbations.The existing paradigm of semi-supervised S2S ASR utilizes SpecAugment as data augmentation and requires a static teacher model to produce pseudo transcripts for untranscribed speech.However, this paradigm fails to take full advantage of consistency regularization.First, the masking operations of SpecAugment may damage the linguistic contents of the speech, thus influencing the quality of pseudo labels.Second, S2S ASR requires both input speech and prefix tokens to make the next prediction.The static prefix tokens made by the offline teacher model cannot match dynamic pseudo labels during consistency training.In this work, we propose an improved consistency training paradigm of semi-supervised S2S ASR.We utilize speech chain reconstruction as the weak augmentation to generate high-quality pseudo labels.Moreover, we demonstrate that dynamic pseudo transcripts produced by the student ASR model benefit the consistency training.Experiments on LJSpeech and LibriSpeech corpora show that compared to supervised baselines, our improved paradigm achieves a 12.2% CER improvement in the single-speaker setting and 38.6% in the multi-speaker setting. Heli Qi, Sashi Novitasari, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2022 | USB: A Unified Semi-supervised Learning Benchmark for ClassificationabstractSemi-supervised learning (SSL) improves model generalization by leveraging massive unlabeled data to augment limited labeled samples. However, currently, popular SSL evaluation protocols are often constrained to computer vision (CV) tasks. In addition, previous work typically trains deep neural networks from scratch, which is time-consuming and environmentally unfriendly. To address the above issues, we construct a Unified SSL Benchmark (USB) for classification by selecting 15 diverse, challenging, and comprehensive tasks from CV, natural language processing (NLP), and audio processing (Audio), on which we systematically evaluate the dominant SSL methods, and also open-source a modular and extensible codebase for fair evaluation of these SSL methods. We further provide the pre-trained versions of the state-of-the-art neural models for CV tasks to make the cost affordable for further tuning. USB enables the evaluation of a single SSL algorithm on more tasks from multiple domains but with less cost. Specifically, on a single NVIDIA V100, only 39 GPU days are required to evaluate FixMatch on 15 tasks in USB while 335 GPU days (279 GPU days on 4 CV datasets except for ImageNet) are needed on 5 CV tasks with TorchSSL. Yidong Wang 0003, Hao Chen 0102, Wang Sun, Ran Tao 0013, Wenxin Hou, Linyi Yang, Zhi Zhou 0007, Lan-Zhe Guo, Heli Qi, Zhen Wu 0002, Yufeng Li 0008, Satoshi Nakamura 0001, Wei Ye 0004, Marios Savvides, Bhiksha Raj, Takahiro Shinozaki, Bernt Schiele, Jindong Wang 0001, Xing Xie 0001, Yue Zhang 0004 |
NeurIPS | 14 |
| 2022 | Tackling multiple object tracking with complicated motions - Re-designing the integration of motion and appearance
Fan Yang 0032, Zheng Wang 0007, Yang Wu 0001, Sakriani Sakti, Satoshi Nakamura 0001 |
Image Vis. Comput. | 5 |
| 2022 | A Machine Speech Chain Approach for Dynamically Adaptive Lombard TTS in Static and Dynamic Noise EnvironmentsabstractRecent end-to-end text-to-speech synthesis (TTS) systems have successfully synthesized high-quality speech. However, TTS speech intelligibility degrades in noisy environments because most of these systems were not designed to handle noisy environments. Several works attempted to address this problem by using offline fine-tuning to adapt their TTS to noisy conditions. Unlike machines, humans never perform offline fine-tuning. Instead, they speak with the Lombard effect in noisy places, where they dynamically adjust their vocal effort to improve the audibility of their speech. This ability is supported by the speech chain mechanism, which involves auditory feedback passing from speech perception to speech production. This paper proposes an alternative approach to TTS in noisy environments that is closer to the human Lombard effect. Specifically, we implement Lombard TTS in a machine speech chain framework to synthesize speech with dynamic adaptation. Our TTS performs adaptation by generating speech utterances based on the auditory feedback that consists of the automatic speech recognition (ASR) loss as the speech intelligibility measure and the speech-to-noise ratio (SNR) prediction as power measurement. Two versions of TTS are investigated: non-incremental TTS with utterance-level feedback and incremental TTS (ITTS) with short-term feedback to reduce the delay without significant performance loss. Furthermore, we evaluate the TTS systems in both static and dynamic noise conditions. Our experimental results show that auditory feedback enhanced the TTS speech intelligibility in noise. Sashi Novitasari, Sakriani Sakti, Satoshi Nakamura 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Modeling Unsupervised Empirical Adaptation by DPGMM and DPGMM-RNN Hybrid Model to Extract Perceptual Features for Low-Resource ASRabstractSpeech feature extraction is critical for ASR systems. Such successful features as MFCC and PLP use filterbank techniques to model log-scaled speech perception but fail to model the adaptation of human speech perception by hearing experiences. Infant perception that is adapted by hearing speech without text may cause permanent brain state modifications (engrams) that serve as a physical fundamental basis for lifetime speech perception formation. This realization motivates us to propose to model such an unsupervised adaptation process, where adaptation denotes perception that is affected or changed by the history of experiences, with the Dirichlet Process Gaussian Mixture Model (DPGMM) and the DPGMM-RNN hybrid model to extract perceptual features to improve ASR. Our proposed features extend MFCC features with posteriorgrams extracted from the DPGMM algorithm or the DPGMM-RNN hybrid model. Our analysis shows that the DPGMM and DPGMM-RNN model perplexities agree with infant auditory perplexity to support that the proposed features are perceptual. Our ASR results verify the effectiveness of the proposed unsupervised features in such tasks as LVCSR on WSJ and ASR on noisy low-resource telephone conversations, compared with the supervised bottleneck features from Kaldi in ASR performance. Sakriani Sakti, Jinsong Zhang 0001, Satoshi Nakamura 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Clustering of Human Movement Trajectories based on Distributional Representations Derived from Bi-directional LSTM Network with Geographical CoordinatesabstractAs the ubiquity of such wearable devices as smart-phones continues to deepen its presence in modern societies, it has become possible to analyze and visualize people who are moving as part of a trajectory of big data. In this study, we cluster human movement trajectories using time-series distributional representations. For the clustering, we calculated the distance of the representation vectors derived from neural network models. Previous work leveraged the Long short-term memory (LSTM) network to train the next mesh prediction. In this study, we propose using the Bi-directional LSTM (Bi-LSTM) network and the integrated additional geographical coordinates (latitude and longitude information) in models to accurately predict the next mesh and construct user clusters. As a result, we improved the accuracy of the next mesh prediction and obtained and visualized clusters of human movement trajectories. Hiroki Tanaka, Takeshi Saga, Satoshi Nakamura 0001 |
IEEE BigData | 3 |
| 2021 | 2nd Workshop on Social Affective Multimodal Interaction for Health (SAMIH)abstractThis workshop discusses how interactive, multimodal technology such as virtual agents can be used in social skills training for measuring and training social-affective interactions. Sensing technology now enables analyzing user’s behaviors and physiological signals. Various signal processing and machine learning methods can be used for such prediction tasks. Such social signal processing and tools can be applied to measure and reduce social stress in everyday situations, including public speaking at schools and workplaces. Hiroki Tanaka, Satoshi Nakamura 0001, Jean-Claude Martin, Catherine Pelachaud |
ICMI | 2 |
| 2021 | Weakly-Supervised Speech-to-Text Mapping with Visually Connected Non-Parallel Speech-Text Data Using Cyclic Partially-Aligned Transformer
Johanes Effendi, Sakriani Sakti, Satoshi Nakamura 0001 |
Interspeech | 3 |
| 2021 | ASR Posterior-Based Loss for Multi-Task End-to-End Speech Translation
Yuka Ko, Katsuhito Sudoh, Sakriani Sakti, Satoshi Nakamura 0001 |
Interspeech | 4 |
| 2021 | Dynamically Adaptive Machine Speech Chain Inference for TTS in Noisy Environment: Listen and Speak Louder
Sashi Novitasari, Sakriani Sakti, Satoshi Nakamura 0001 |
Interspeech | 3 |
| 2021 | Unsupervised Neural-Based Graph Clustering for Variable-Length Speech Representation Discovery of Zero-Resource Languages
Shun Takahashi, Sakriani Sakti, Satoshi Nakamura 0001 |
Interspeech | 3 |
| 2021 | Transcribing Paralinguistic Acoustic Cues to Target Language Text in Transformer-Based Speech-to-Text Translation
Hirotaka Tokuyama, Sakriani Sakti, Katsuhito Sudoh, Satoshi Nakamura 0001 |
Interspeech | 4 |
| 2021 | ARTA: Collection and Classification of Ambiguous Requests and Thoughtful ActionsabstractHuman-assisting systems such as dialogue systems must take thoughtful, appropriate actions not only for clear and unambiguous user requests, but also for ambiguous user requests, even if the users themselves are not aware of their potential requirements.To construct such a dialogue agent, we collected a corpus and developed a model that classifies ambiguous user requests into corresponding system actions.In order to collect a high-quality corpus, we asked workers to input antecedent user requests whose pre-defined actions could be regarded as thoughtful.Although multiple actions could be identified as thoughtful for a single user request, annotating all combinations of user requests and system actions is impractical.For this reason, we fully annotated only the test data and left the annotation of the training data incomplete.In order to train the classification model on such training data, we applied the positive/unlabeled (PU) learning method, which assumes that only a part of the data is labeled with positive examples.The experimental results show that the PU learning method achieved better performance than the general positive/negative (PN) learning method to classify thoughtful actions given an ambiguous user request. Shohei Tanaka, Koichiro Yoshino, Katsuhito Sudoh, Satoshi Nakamura 0001 |
SIGDIAL | 4 |
| 2021 | Transformer-Based Direct Speech-To-Speech Translation with TranscoderabstractTraditional speech translation systems use a cascade manner that concatenates speech recognition (ASR), machine translation (MT), and text-to-speech (TTS) synthesis to translate speech from one language to another language in a step-by-step manner. Unfortunately, since those components are trained separately, MT often struggles to handle ASR errors, resulting in unnatural translation results. Recently, one work attempted to construct direct speech translation in a single model. The model used a multi-task scheme that learns to predict not only the target speech spectrograms directly but also the source and target phoneme transcription as auxiliary tasks. However, that work was only evaluated Spanish-English language pairs with similar syntax and word order. With syntactically distant language pairs, speech translation requires distant word order, and thus direct speech frame-to-frame alignments become difficult. Another direction was to construct a single deep-learning framework while keeping the step-by-step translation process. However, such studies focused only on speech-to-text translation. Furthermore, all of these works were based on a recurrent neural net-work (RNN) model. In this work, we propose a step-by-step scheme to a complete end-to-end speech-to-speech translation and propose a Transformer-based speech translation using Transcoder. We compare our proposed and multi-task model using syntactically similar and distant language pairs. Takatomo Kano, Sakriani Sakti, Satoshi Nakamura 0001 |
SLT | 3 |
| 2021 | Incorporating Discriminative DPGMM Posteriorgrams for Low-Resource ASRabstractThe first step in building an ASR system is to extract proper speech features. The ideal speech features for ASR must also have high discriminabilities between linguistic units and be robust to such non-linguistic factors as gender, age, emotions, or noise. The discriminabilities of various features have been compared in several Zerospeech challenges to discover linguistic units without any transcriptions, in which the posteriorgrams of DPGMM clustering show strong discriminability and get several top results of ABX discrimination scores between phonemes. This paper appends DPGMM posteriorgrams to increase the discriminability of acoustic features to enhance ASR systems. To the best of our knowledge, DPGMM features, which are usually applied to such tasks as spoken term detection and zero resources tasks, have not been applied to large vocabulary continuous speech recognition (LVCSR) before. DPGMM clustering can dynamically change the number of Gaussians until each one fits one segmental pattern of the whole speech corpus with the highest probability such that the linguistic units of different segmental patterns are clearly discriminated. Our experimental results on the WSJ corpora show our proposal stably improves ASR systems and provides even more improvement for smaller datasets with fewer resources. Sakriani Sakti, Satoshi Nakamura 0001 |
SLT | 3 |
| 2021 | ReMOT: A model-agnostic refinement for multiple object tracking
Fan Yang 0032, Sakriani Sakti, Yang Wu 0001, Satoshi Nakamura 0001 |
Image Vis. Comput. | 5 |
| 2021 | Towards Tokenization and Part-of-Speech Tagging for Khmer: Data and DiscussionabstractAs a highly analytic language, Khmer has considerable ambiguities in tokenization and part-of-speech (POS) tagging processing. This topic is investigated in this study. Specifically, a 20,000-sentence Khmer corpus with manual tokenization and POS-tagging annotation is released after a series of work over the last 4 years. This is the largest morphologically annotated Khmer dataset as of 2020, when this article was prepared. Based on the annotated data, experiments were conducted to establish a comprehensive benchmark on the automatic processing of tokenization and POS-tagging for Khmer. Specifically, a support vector machine, a conditional random field (CRF) , a long short-term memory (LSTM) -based recurrent neural network, and an integrated LSTM-CRF model have been investigated and discussed. As a primary conclusion, processing at morpheme-level is satisfactory for the provided data. However, it is intrinsically difficult to identify further grammatical constituents of compounds or phrases because of the complex analytic features of the language. Syntactic annotation and automatic parsing for Khmer will be scheduled in the near future. Hour Kaing, Chenchen Ding, Masao Utiyama, Eiichiro Sumita, Sam Sethserey, Sopheap Seng, Katsuhito Sudoh, Satoshi Nakamura 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 8 |
| 2021 | Tackling Perception Bias in Unsupervised Phoneme Discovery Using DPGMM-RNN Hybrid Model and Functional LoadabstractThe human perception of phonemes is biased against speech sounds. The lack of correspondence between perceptual phonemes and acoustic signals forms a big challenge in designing unsupervised algorithms to distinguish phonemes from sound. We propose the DPGMM-RNN hybrid model that improves phoneme categorization by relieving the fragmentation problem. We also merge segments with low functional load, which is the work done by segment contrasts to differentiate between utterances, just like humans who convert unambiguous segments into phonemes as units for immediate perception. Our results show that the DPGMM-RNN hybrid model relieves the fragmentation problem and improves phoneme discriminability. The minimal functional load merge compresses a segment system, preserves information and keeps phoneme discriminability. Sakriani Sakti, Jinsong Zhang 0001, Satoshi Nakamura 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Instance-Level Heterogeneous Domain Adaptation for Limited-Labeled Sketch-to-Photo RetrievalabstractAlthough sketch-to-photo retrieval has a wide range of applications, it is costly to obtain paired and rich-labeled ground truth. Differently, photo retrieval data is easier to acquire. Therefore, previous works pre-train their models on rich-labeled photo retrieval data (i.e., source domain) and then fine-tune them on the limited-labeled sketch-to-photo retrieval data (i.e., target domain). However, without co-training source and target data, source domain knowledge might be forgotten during the fine-tuning process, while simply co-training them may cause negative transfer due to domain gaps. Moreover, identity label spaces of source data and target data are generally disjoint and therefore conventional category-level Domain Adaptation (DA) is not directly applicable. To address these issues, we propose an Instance-level Heterogeneous Domain Adaptation (IHDA) framework. We apply the fine-tuning strategy for identity label learning, aiming to transfer the instance-level knowledge in an inductive transfer manner. Meanwhile, labeled attributes from the source data are selected to form a shared label space for source and target domains. Guided by shared attributes, DA is utilized to bridge cross-dataset domain gaps and heterogeneous domain gaps, which transfers instance-level knowledge in a transductive transfer manner. Experiments show that our method has set a new state of the art on three sketch-to-photo image retrieval benchmarks without extra annotations, which opens the door to train more effective models on limited-labeled heterogeneous image retrieval tasks. Fan Yang 0032, Yang Wu 0001, Zheng Wang 0007, Sakriani Sakti, Satoshi Nakamura 0001 |
IEEE Trans. Multim. | 6 |
| 2020 | Automatic Machine Translation Evaluation using Source Language Inputs and Cross-lingual Language ModelabstractWe propose an automatic evaluation method of machine translation that uses source language sentences regarded as additional pseudo references.The proposed method evaluates a translation hypothesis in a regression model.The model takes the paired source, reference, and hypothesis sentence all together as an input.A pretrained large scale cross-lingual language model encodes the input to sentence-pair vectors, and the model predicts a human evaluation score with those vectors.Our experiments show that our proposed method using Crosslingual Language Model (XLM) trained with a translation language modeling (TLM) objective achieves a higher correlation with human judgments than a baseline method that uses only hypothesis and reference sentences.Additionally, using source sentences in our proposed method is confirmed to improve the evaluation performance. Kosuke Takahashi, Katsuhito Sudoh, Satoshi Nakamura 0001 |
ACL | 3 |
| 2020 | Incorporating Noisy Length Constraints into Transformer with Length-aware Positional EncodingsabstractNeural Machine Translation often suffers from an under-translation problem due to its limited modeling of output sequence lengths.In this work, we propose a novel approach to training a Transformer model using length constraints based on length-aware positional encoding (PE).Since length constraints with exact target sentence lengths degrade translation performance, we add random noise within a certain window size to the length constraints in the PE during the training.In the inference step, we predict the output lengths using input sequences and a BERTbased length prediction model.Experimental results in an ASPEC English-to-Japanese translation showed the proposed method produced translations with lengths close to the reference ones and outperformed a vanilla Transformer by 3.22 points in BLEU on short sentences within ten subwords.The average translation results using our length prediction model were also better than another baseline method using input lengths for the length constraints.The proposed noise injection improved robustness for length prediction errors, especially within the window size. Yui Oka, Katsuki Chousa, Katsuhito Sudoh, Satoshi Nakamura 0001 |
COLING | 4 |
| 2020 | Improving Spoken Language Understanding by Wisdom of CrowdsabstractSpoken language understanding (SLU), which converts user requests in natural language to machine-interpretable expressions, is becoming an essential task.The lack of training data is an important problem, especially for new system tasks, because existing SLU systems are based on statistical approaches.In this paper, we proposed to use two sources of the "wisdom of crowds," crowdsourcing and knowledge community website, for improving the SLU system.We firstly collected paraphrasing variations for new system tasks through crowdsourcing as seed data, and then augmented them using similar questions from a knowledge community website.We investigated the effects of the proposed data augmentation method in SLU task, even with small seed data.In particular, the proposed architecture augmented more than 120,000 samples to improve SLU accuracies. Koichiro Yoshino, Kana Ikeuchi, Katsuhito Sudoh, Satoshi Nakamura 0001 |
COLING | 4 |
| 2020 | Using Panoramic Videos for Multi-Person Localization and Tracking In A 3D Panoramic Coordinateabstract3D panoramic multi-person localization and tracking are prominent in many applications, however, conventional methods using LiDAR equipment could be economically expensive and also computationally inefficient due to the processing of point cloud data. In this work, we propose an effective and efficient approach at a low cost. First, we obtain panoramic videos with four normal cameras. Then, we transform human locations from a 2D panoramic image coordinate to a 3D panoramic camera coordinate using camera geometry and human bio-metric property (i.e., height). Finally, we generate 3D tracklets by associating human appearance and 3D trajectory. We verify the effectiveness of our method on three datasets including a new one built by us, in terms of 3D single-view multi-person localization, 3D single-view multi-person tracking, and 3D panoramic multi-person localization and tracking. Our code and dataset are available at https://github.com/fandulu/MPLT. Fan Yang 0032, Feiran Li, Yang Wu 0001, Sakriani Sakti, Satoshi Nakamura 0001 |
ICASSP | 5 |
| 2020 | DEJA-VU: Double Feature Presentation and Iterated Loss in Deep Transformer NetworksabstractDeep acoustic models typically receive features in the first layer of the network, and process increasingly abstract representations in the subsequent layers. Here, we propose to feed the input features at multiple depths in the acoustic model. As our motivation is to allow acoustic models to re-examine their input features in light of partial hypotheses we introduce intermediate model heads and loss function. We study this architecture in the context of deep Transformer networks, and we use an attention mechanism over both the previous layer activations and the input features. To train this model's intermediate output hypothesis, we apply the objective function at each layer right before feature re-use. We find that the use of such iterated loss significantly improves performance by itself, as well as enabling input feature re-use. We present results on both Librispeech, and a large scale video dataset, with relative improvements of 10 - 20% for Librispeech and 3.2 - 13% for videos. Andros Tjandra, Chunxi Liu, Frank Zhang 0001, Xiaohui Zhang 0007, Yongqiang Wang 0005, Gabriel Synnaeve, Satoshi Nakamura 0001, Geoffrey Zweig |
ICASSP | 7 |
| 2020 | Social Affective Multimodal Interaction for HealthabstractThis workshop discusses how interactive, multimodal technology such as virtual agents can be used in social skills training for measuring and training social-affective interactions. Sensing technology now enables analyzing user's behaviors and physiological signals. Various signal processing and machine learning methods can be used for such prediction tasks. Such social signal processing and tools can be applied to measure and reduce social stress in everyday situations, including public speaking at schools and workplaces. Hiroki Tanaka, Satoshi Nakamura 0001, Jean-Claude Martin, Catherine Pelachaud |
ICMI | 2 |
| 2020 | Augmenting Images for ASR and TTS Through Single-Loop and Dual-Loop Multimodal Chain FrameworkabstractPrevious research has proposed a machine speech chain to enable automatic speech recognition (ASR) and text-to-speech synthesis (TTS) to assist each other in semi-supervised learning and to avoid the need for a large amount of paired speech and text data.However, that framework still requires a large amount of unpaired (speech or text) data.A prototype multimodal machine chain was then explored to further reduce the need for a large amount of unpaired data, which could improve ASR or TTS even when no more speech or text data were available.Unfortunately, this framework relied on the image retrieval (IR) model, and thus it was limited to handling only those images that were already known during training.Furthermore, the performance of this framework was only investigated with single-speaker artificial speech data.In this study, we revamp the multimodal machine chain framework with image generation (IG) and investigate the possibility of augmenting image data for ASR and TTS using single-loop and dual-loop architectures on multispeaker natural speech data.Experimental results revealed that both single-loop and dual-loop multimodal chain frameworks enabled ASR and TTS to improve their performance using an image-only dataset. Johanes Effendi, Andros Tjandra, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2020 | Incremental Machine Speech Chain Towards Enabling Listening While Speaking in Real-TimeabstractInspired by a human speech chain mechanism, a machine speech chain framework based on deep learning was recently proposed for the semi-supervised development of automatic speech recognition (ASR) and text-to-speech synthesis (TTS) systems.However, the mechanism to listen while speaking can be done only after receiving entire input sequences.Thus, there is a significant delay when encountering long utterances.By contrast, humans can listen to what they speak in real-time, and if there is a delay in hearing, they won't be able to continue speaking.In this work, we propose an incremental machine speech chain towards enabling machine to listen while speaking in real-time.Specifically, we construct incremental ASR (ISR) and incremental TTS (ITTS) by letting both systems improve together through a short-term loop.Our experimental results reveal that our proposed framework is able to reduce delays due to long utterances while keeping a comparable performance to the non-incremental basic machine speech chain. Sashi Novitasari, Andros Tjandra, Tomoya Yanagita, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2020 | Combining Audio and Brain Activity for Predicting Speech Quality
Ivan Halim Parmonangan, Hiroki Tanaka, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2020 | Transformer VQ-VAE for Unsupervised Unit Discovery and Speech Synthesis: ZeroSpeech 2020 ChallengeabstractIn this paper, we report our submitted system for the ZeroSpeech 2020 challenge on Track 2019.The main theme in this challenge is to build a speech synthesizer without any textual information or phonetic labels.In order to tackle those challenges, we build a system that must address two major components such as 1) given speech audio, extract subword units in an unsupervised way and 2) resynthesize the audio from novel speakers.The system also needs to balance the codebook performance between the ABX error rate and the bitrate compression rate.Our main contribution here is we proposed Transformer-based VQ-VAE for unsupervised unit discovery and Transformerbased inverter for the speech synthesis given the extracted codebook.Additionally, we also explored several regularization methods to improve performance even further. Andros Tjandra, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2020 | Neural Speech Completion
Kazuki Tsunematsu, Johanes Effendi, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2020 | Emotional Speech Corpus for Persuasive Dialogue SystemabstractExpressing emotion is known as an efficient way to persuade one’s dialogue partner to accept one’s claim or proposal. Emotional expression in speech can express the speaker’s emotion more directly than using only emotion expression in the text, which will lead to a more persuasive dialogue. In this paper, we built a speech dialogue corpus in a persuasive scenario that uses emotional expressions to build a persuasive dialogue system with emotional expressions. We extended an existing text dialogue corpus by adding variations of emotional responses to cover different combinations of broad dialogue context and a variety of emotional states by crowd-sourcing. Then, we recorded emotional speech consisting of of collected emotional expressions spoken by a voice actor. The experimental results indicate that the collected emotional expressions with their speeches have higher emotional expressiveness for expressing the system’s emotion to users. Sara Asai, Koichiro Yoshino, Seitaro Shinagawa, Sakriani Sakti, Satoshi Nakamura 0001 |
LREC | 5 |
| 2020 | Improving neural machine translation through phrase-based soft forced decoding
Jingyi Zhang 0001, Masao Utiyama, Eiichiro Sumita, Graham Neubig, Satoshi Nakamura 0001 |
Mach. Transl. | 5 |
| 2020 | End-to-End Speech Translation With Transcoding by Multi-Task Learning for Distant Language PairsabstractDirectly translating spoken utterances from a source language to a target language is challenging because it requires a fundamental transformation in both linguistic and para/non-linguistic features. Traditional speech-to-speech translation approaches concatenate automatic speech recognition (ASR), text-to-text machine translation (MT), and text-to-speech synthesizer (TTS) by text information. The current state-of-the-art models for ASR, MT, and TTS have mainly been built using deep neural networks, in particular, an attention-based encoder-decoder neural network with an attention mechanism. Recently, several works have constructed end-to-end direct speech-to-text translation by combining ASR and MT into a single model. However, the usefulness of these models has only been investigated on language pairs of similar syntax and word order (e.g., English-French or English-Spanish). For syntactically distant language pairs (e.g., English-Japanese), speech translation requires distant word reordering. Furthermore, parallel texts with corresponding speech utterances that are suitable for training end-to-end speech translation are generally unavailable. Collecting such corpora is usually time-consuming and expensive. This article proposes the first attempt to build an end-to-end direct speech-to-text translation system on syntactically distant language pairs that suffer from long-distance reordering. We train the model on English (subject-verb-object (SVO) word order) and Japanese (SOV word order) language pairs. To guide the attention-based encoder-decoder model on this difficult problem, we construct end-to-end speech translation with transcoding and utilize curriculum learning (CL) strategies that gradually train the network for end-to-end speech translation tasks by adapting the decoder or encoder parts. We use TTS for data augmentation to generate corresponding speech utterances from the existing parallel text data. Our experiment results show that the proposed approach provides significant improvements compared with conventional cascade models and the direct speech translation approach that uses a single model without transcoding and CL strategies. Takatomo Kano, Sakriani Sakti, Satoshi Nakamura 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Multi-Source Neural Machine Translation With Missing DataabstractMachine translation is rife with ambiguities in word ordering and word choice, and even with the advent of machine-learning methods that learn to resolve this ambiguity based on statistics from large corpora, mistakes are frequent. Multi-source translation is an approach that attempts to resolve these ambiguities by exploiting multiple inputs (e.g. sentences in three different languages) to increase translation accuracy. These methods are trained on multilingual corpora, which include the multiple source languages and the target language, and then at test time uses information from both source languages while generating the target. While there are many of these multilingual corpora, such as multilingual translations of TED talks or European parliament proceedings, in practice, many multilingual corpora are not complete due to the difficulty to provide translations in all of the relevant languages. Existing studies on multi-source translation did not explicitly handle such situations, and thus are only applicable to complete corpora that have all of the languages of interest, severely limiting their practical applicability. In this article, we examine approaches for multi-source neural machine translation (NMT) that can learn from and translate such incomplete corpora. Specifically, we propose methods to deal with incomplete corpora at both training time and test time. For training time, we examine two methods: (1) a simple method that simply replaces missing source translations with a special NULL symbol, and (2) a data augmentation approach that fills in incomplete parts with source translations created from multi-source NMT. For test-time, we examine methods that use multi-source translation even when only a single source is provided by first translating into an additional auxiliary language using standard NMT, then using multi-source translation on the original source and this generated auxiliary language sentence. Extensive experiments demonstrate that the proposed training-time and test-time methods both significantly improve translation performance. Yuta Nishimura, Katsuhito Sudoh, Graham Neubig, Satoshi Nakamura 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | Machine Speech ChainabstractDespite the close relationship between speech perception and production, research in automatic speech recognition (ASR) and text-to-speech synthesis (TTS) has progressed more or less independently without exerting much mutual influence. In human communication, on the other hand, a closed-loop speech chain mechanism with auditory feedback from the speaker's mouth to her ear is crucial. In this paper, we take a step further and develop a closed-loop machine speech chain model based on deep learning. The sequence-to-sequence model in closed-loop architecture allows us to train our model on the concatenation of both labeled and unlabeled data. While ASR transcribes the unlabeled speech features, TTS attempts to reconstruct the original speech waveform based on the text from ASR. In the opposite direction, ASR also attempts to reconstruct the original text transcription given the synthesized speech. To the best of our knowledge, this is the first deep learning framework that integrates human speech perception and production behaviors. Our experimental results show that the proposed approach significantly improved performance over that from separate systems that were only trained with labeled data. Andros Tjandra, Sakriani Sakti, Satoshi Nakamura 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Corrections to "Machine Speech Chain"abstractPresents corrections to author information for the above named paper. Andros Tjandra, Sakriani Sakti, Satoshi Nakamura 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Listening While Speaking and Visualizing: Improving ASR Through Multimodal ChainabstractPreviously, a machine speech chain, which is based on sequence-to-sequence deep learning, was proposed to mimic speech perception and production behavior. Such chains separately processed listening and speaking by automatic speech recognition (ASR) and text-to-speech synthesis (TTS) and simultaneously enabled them to teach each other in semi-supervised learning when they received unpaired data. Unfortunately, this speech chain study is limited to speech and textual modalities. In fact, natural communication is actually multimodal and involves both auditory and visual sensory systems. Although the said speech chain reduces the requirement of having a full amount of paired data, in this case we still need a large amount of unpaired data. In this research, we take a further step and construct a multimodal chain and design a closely knit chain architecture that combines ASR, TTS, image captioning, and image production models into a single framework. The framework allows the training of each component without requiring a large number of parallel multimodal data. Our experimental results also show that an ASR can be further trained without speech and text data and cross-modal data augmentation remains possible through our proposed chain, which improves the ASR performance. Johanes Effendi, Andros Tjandra, Sakriani Sakti, Satoshi Nakamura 0001 |
ASRU | 4 |
| 2019 | Neural Machine Translation with Acoustic EmbeddingabstractNeural machine translation (NMT) has successfully redefined the state of the art in machine translation on several language pairs. One popular framework models the translation process end-to-end using attentional encoder-decoder architecture and treats each word in the vectors of intermediate representation. These embedding vectors are sensitive to the meaning of words and allow semantically similar words to be near each other in the vector spaces and share their statistical power. Unfortunately, the model often maps such similar words too closely, which complicates distinguishing them. Consequently, NMT systems often mistranslate words that seem natural in the context but do not reflect the content of the source sentence. Incorporating auxiliary information usually enhances the discriminability. In this research, we integrate acoustic information within NMT by multi-task learning. Here, our model learns how to embed and translate word sequences based on their acoustic and semantic differences by helping it choose the correct output word based on its meaning and pronunciation. Our experiment results show that our proposed approach provides more significant improvement than the standard text-based transformer NMT model in BLEU score evaluation. Takatomo Kano, Sakriani Sakti, Satoshi Nakamura 0001 |
ASRU | 3 |
| 2019 | Zero-Shot Code-Switching ASR and TTS with Multilingual Machine Speech ChainabstractConstructing automatic speech recognition (ASR) and text-to-speech (TTS) for code-switching in a supervised fashion poses a challenge since a large amount of code-switching speech and the corresponding transcription are usually unavailable. The machine speech chain mechanism can be utilized to achieve semi-supervised learning. The framework enables ASR and TTS to assist each other when they receive unpaired data since it allows them to infer the missing pair and optimize the models with reconstruction loss. In this study, we handle multiple language pairs of code-switching by integrating language embeddings into the machine speech chain and investigate whether the model can perform with code-switching language pairs that are never explicitly seen during training. Experimental results reveal that the proposed approach improves the performance of the multilingual code-switching language pairs with which the model was trained and can also perform with unknown code-switching language pairs without directly learning on it. Sahoko Nakayama, Andros Tjandra, Sakriani Sakti, Satoshi Nakamura 0001 |
ASRU | 4 |
| 2019 | Speech-to-Speech Translation Between Untranscribed Unknown LanguagesabstractIn this paper, we explore a method for training speech-to-speech translation tasks without any transcription or linguistic supervision. Our proposed method consists of two steps: First, we train and generate discrete representation with unsupervised term discovery with a discrete quantized autoencoder. Second, we train a sequence-to-sequence model that directly maps the source language speech to the target languages discrete representation. Our proposed method can directly generate target speech without any auxiliary or pre-training steps with a source or target transcription. To the best of our knowledge, this is the first work that performed pure speech-to-speech translation between untranscribed unknown languages. Andros Tjandra, Sakriani Sakti, Satoshi Nakamura 0001 |
ASRU | 3 |
| 2019 | Speech Artifact Removal from Eeg Recordings of Spoken Word Production with Tensor DecompositionabstractResearch about brain activities involving spoken word production is considerably underdeveloped because of the undiscovered characteristics of speech artifacts, which contaminate electroencephalogram (EEG) signals and prevent the inspection of the underlying cognitive processes. To fuel further EEG research with speech production, a method using three-mode tensor decomposition (time x space x frequency) is proposed to perform speech artifact removal. Tensor decomposition enables simultaneous inspection of multiple modes, which suits the multi-way nature of EEG data. In a picture-naming task, we collected raw data with speech artifacts by placing two electrodes near the mouth to record lip EMG. Based on our evaluation, which calculated the correlation values between grand-averaged speech artifacts and the lip EMG, tensor decomposition outperformed the former methods that were based on independent component analysis (ICA) and blind source separation (BSS), both in detecting speech artifact (0.985) and producing clean data (0.101). Our proposed method correctly preserved the components unrelated to speech, which was validated by computing the correlation value between the grand-averaged raw data without EOG and cleaned data before the speech onset (0.92-0.94). Holy Lovenia, Hiroki Tanaka, Sakriani Sakti, Ayu Purwarianti, Satoshi Nakamura 0001 |
ICASSP | 5 |
| 2019 | End-to-end Feedback Loss in Speech Chain Framework via Straight-through EstimatorabstractThe speech chain mechanism integrates automatic speech recognition (ASR) and text-to-speech synthesis (TTS) modules into a single cycle during training. In our previous work, we applied a speech chain mechanism as a semi-supervised learning. It provides the ability for ASR and TTS to assist each other when they receive unpaired data and let them infer the missing pair and optimize the model with reconstruction loss. If we only have speech without transcription, ASR generates the most likely transcription from the speech data, and then TTS uses the generated transcription to reconstruct the original speech features. However, in previous papers, we just limited our back-propagation to the closest module, which is the TTS part. One reason is that back-propagating the error through the ASR is challenging due to the output of the ASR being discrete tokens, creating non-differentiability between the TTS and ASR. In this paper, we address this problem and describe how to thoroughly train a speech chain end-to-end for reconstruction loss using a straight-through estimator (ST). Experimental results revealed that, with sampling from ST-Gumbel-Softmax, we were able to update ASR parameters and improve the ASR performances by 11% relative CER reduction compared to the baseline. Andros Tjandra, Sakriani Sakti, Satoshi Nakamura 0001 |
ICASSP | 3 |
| 2019 | Cross-lingual Speech-based Tobi Label Generation Using Bidirectional LstmabstractIn this paper we investigate the automatic generation of ToBI-style prosody labels. The work is motivated by the idea of using prosodic information to facilitate the automatic lexicon discovery for unseen and under-resourced languages for which sufficient training data is not available. Specifically, the prosodic boundaries are meant to serve as additional top-down information in the word segmentation step. To this end we attempt to apply the trained Japanese models cross-lingually on a language not seen in training (English). We generate break index labels, using only the speech signal as input, with no additional information given at test time in the form of transcripts or prior word segmentations. The labels are generated using bidirectional LSTMs trained on spontaneous Japanese speech. We evaluate the quality of these labels using established metrics, with an F1 score of 0.55 for cross-lingual prosodic break detection (given a tolerance of 80 ms). Marco Vetter, Sakriani Sakti, Satoshi Nakamura 0001 |
ICASSP | 3 |
| 2019 | Neural Conversation Model Controllable by Given Dialogue Act Based on Adversarial Learning and Label-aware ObjectiveabstractBuilding a controllable neural conversation model (NCM) is an important task.In this paper, we focus on controlling the responses of NCMs by using dialogue act labels of responses as conditions.We introduce an adversarial learning framework for the task of generating conditional responses with a new objective to a discriminator, which explicitly distinguishes sentences by using labels.This change strongly encourages the generation of label-conditioned sentences.We compared the proposed method with some existing methods for generating conditional responses.The experimental results show that our proposed method has higher controllability for dialogue acts even though it has higher or comparable naturalness to existing methods. Seiya Kawano, Koichiro Yoshino, Satoshi Nakamura 0001 |
INLG | 3 |
| 2019 | An Incremental Turn-Taking Model for Task-Oriented Dialog SystemsabstractIn a human-machine dialog scenario, deciding the appropriate time for the machine to take the turn is an open research problem. In contrast, humans engaged in conversations are able to timely decide when to interrupt the speaker for competitive or non-competitive reasons. In state-of-the-art turn-by-turn dialog systems the decision on the next dialog action is taken at the end of the utterance. In this paper, we propose a token-by-token prediction of the dialog state from incremental transcriptions of the user utterance. To identify the point of maximal understanding in an ongoing utterance, we a) implement an incremental Dialog State Tracker which is updated on a token basis (iDST) b) re-label the Dialog State Tracking Challenge 2 (DSTC2) dataset and c) adapt it to the incremental turn-taking experimental scenario. The re-labeling consists of assigning a binary value to each token in the user utterance that allows to identify the appropriate point for taking the turn. Finally, we implement an incremental Turn Taking Decider (iTTD) that is trained on these new labels for the turn-taking decision. We show that the proposed model can achieve a better performance compared to a deterministic handcrafted turn-taking algorithm. Andrei Catalin, Koichiro Yoshino, Yukitoshi Murase, Satoshi Nakamura 0001, Giuseppe Riccardi |
INTERSPEECH | 4 |
| 2019 | Sequence-to-Sequence Learning via Attention Transfer for Incremental Speech RecognitionabstractAttention-based sequence-to-sequence automatic speech recognition (ASR) requires a significant delay to recognize long utterances because the output is generated after receiving entire input sequences. Although several studies recently proposed sequence mechanisms for incremental speech recognition (ISR), using different frameworks and learning algorithms is more complicated than the standard ASR model. One main reason is because the model needs to decide the incremental steps and learn the transcription that aligns with the current short speech segment. In this work, we investigate whether it is possible to employ the original architecture of attention-based ASR for ISR tasks by treating a full-utterance ASR as the teacher model and the ISR as the student model. We design an alternative student network that, instead of using a thinner or a shallower model, keeps the original architecture of the teacher model but with shorter sequences (few encoder and decoder states). Using attention transfer, the student network learns to mimic the same alignment between the current input short speech segments and the transcription. Our experiments show that by delaying the starting time of recognition process with about 1.7 sec, we can achieve comparable performance to one that needs to wait until the end. Sashi Novitasari, Andros Tjandra, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2019 | Speech Quality Evaluation of Synthesized Japanese Speech Using EEG
Ivan Halim Parmonangan, Hiroki Tanaka, Sakriani Sakti, Shinnosuke Takamichi, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2019 | VQVAE Unsupervised Unit Discovery and Multi-Scale Code2Spec Inverter for Zerospeech Challenge 2019abstractWe describe our submitted system for the ZeroSpeech Challenge 2019.The current challenge theme addresses the difficulty of constructing a speech synthesizer without any text or phonetic labels and requires a system that can (1) discover subword units in an unsupervised way, and (2) synthesize the speech with a target speaker's voice.Moreover, the system should also balance the discrimination score ABX, the bit-rate compression rate, and the naturalness and the intelligibility of the constructed voice.To tackle these problems and achieve the best tradeoff, we utilize a vector quantized variational autoencoder (VQ-VAE) and a multi-scale codebook-tospectrogram (Code2Spec) inverter trained by mean square error and adversarial loss.The VQ-VAE extracts the speech to a latent space, forces itself to map it into the nearest codebook and produces compressed representation.Next, the inverter generates a magnitude spectrogram to the target voice, given the codebook vectors from VQ-VAE.In our experiments, we also investigated several other clustering algorithms, including K-Means and GMM, and compared them with the VQ-VAE result on ABX scores and bit rates.Our proposed approach significantly improved the intelligibility (in CER), the MOS, and discrimination ABX scores compared to the official ZeroSpeech 2019 baseline or even the topline. Andros Tjandra, Berrak Sisman, Mingyang Zhang 0003, Sakriani Sakti, Haizhou Li 0001, Satoshi Nakamura 0001 |
INTERSPEECH | 6 |
| 2019 | Make Skeleton-based Action Recognition Model Smaller, Faster and BetterabstractAlthough skeleton-based action recognition has achieved great success in recent years, most of the existing methods may suffer from a large model size and slow execution speed. To alleviate this issue, we analyze skeleton sequence properties to propose a Double-feature Double-motion Network (DD-Net) for skeleton-based action recognition. By using a lightweight network structure (i.e., 0.15 million parameters), DD-Net can reach a super fast speed, as 3,500 FPS on an ordinary GPU (e.g., GTX 1080Ti), or, 2,000 FPS on an ordinary CPU (e.g., Intel E5-2620). By employing robust features, DD-Net achieves state-of-the-art performance on our experiment datasets: SHREC (i.e., hand actions) and JHMDB (i.e., body actions). Our code is on https://github.com/fandulu/DD-Net. Fan Yang 0032, Yang Wu 0001, Sakriani Sakti, Satoshi Nakamura 0001 |
MMAsia | 4 |
| 2019 | Classification of alkaloids according to the starting substances of their biosynthetic pathways using graph convolutional neural networksabstractBACKGROUND: Alkaloids, a class of organic compounds that contain nitrogen bases, are mainly synthesized as secondary metabolites in plants and fungi, and they have a wide range of bioactivities. Although there are thousands of compounds in this class, few of their biosynthesis pathways are fully identified. In this study, we constructed a model to predict their precursors based on a novel kind of neural network called the molecular graph convolutional neural network. Molecular similarity is a crucial metric in the analysis of qualitative structure-activity relationships. However, it is sometimes difficult for current fingerprint representations to emphasize specific features for the target problems efficiently. It is advantageous to allow the model to select the appropriate features according to data-driven decisions for extracting more useful information, which influences a classification or regression problem substantially. RESULTS: In this study, we applied a neural network architecture for undirected graph representation of molecules. By encoding a molecule as an abstract graph and applying "convolution" on the graph and training the weight of the neural network framework, the neural network can optimize feature selection for the training problem. By incorporating the effects from adjacent atoms recursively, graph convolutional neural networks can extract the features of latent atoms that represent chemical features of a molecule efficiently. In order to investigate alkaloid biosynthesis, we trained the network to distinguish the precursors of 566 alkaloids, which are almost all of the alkaloids whose biosynthesis pathways are known, and showed that the model could predict starting substances with an averaged accuracy of 97.5%. CONCLUSION: We have showed that our model can predict more accurately compared to the random forest and general neural network when the variables and fingerprints are not selected, while the performance is comparable when we carefully select 507 variables from 18000 dimensions of descriptors. The prediction of pathways contributes to understanding of alkaloid synthesis mechanisms and the application of graph based neural network models to similar problems in bioinformatics would therefore be beneficial. We applied our model to evaluate the precursors of biosynthesis of 12000 alkaloids found in various organisms and found power-low-like distribution. Ryohei Eguchi, Naoaki Ono, Aki Hirai, Tetsuo Katsuragi, Satoshi Nakamura 0001, Ming Huang 0002, Md. Altaf-Ul-Amin, Shigehiko Kanaya |
BMC Bioinform. | 5 |
| 2019 | Associative knowledge feature vector inferred on external knowledge base for dialog state tracking
Yukitoshi Murase, Koichiro Yoshino, Satoshi Nakamura 0001 |
Comput. Speech Lang. | 3 |
| 2019 | Positive Emotion Elicitation in Chat-Based Dialogue SystemsabstractWe aim to draw on an important overlooked potential of affective dialogue systems-their application to promote positive emotional states, similar to that of emotional support between humans. This can be achieved by eliciting a more positive emotional valence throughout a dialogue system interaction, i.e., positive emotion elicitation. Existing works on emotion elicitation have not yet paid attention to the emotional benefit for the users. Moreover, a positive emotion elicitation corpus does not yet exist despite the growing number of emotion-rich corpora. Towards this goal, first, we propose a response retrieval approach for positive emotion elicitation by utilizing examples of emotion appraisal from a dialogue corpus. Second, we efficiently construct a corpus using the proposed retrieval method, by replacing responses in a dialogue with those that elicit a more positive emotion. We validate the corpus through crowdsourcing to ensure its quality. Finally, we propose a novel neural network architecture for an emotion-sensitive neural chat-based dialogue system, optimized on the constructed corpus to elicit positive emotion. Objective and subjective evaluations show that the proposed methods result in dialogue responses that are more natural and elicit a more positive emotional response. Further analyses of the results are discussed in this paper. Nurul Lubis, Sakriani Sakti, Koichiro Yoshino, Satoshi Nakamura 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2018 | Eliciting Positive Emotion through Affect-Sensitive Dialogue Response Generation: A Neural Network ApproachabstractAn emotionally-competent computer agent could be a valuable assistive technology in performing various affective tasks. For example caring for the elderly, low-cost ubiquitous chat therapy, and providing emotional support in general, by promoting a more positive emotional state through dialogue system interaction. However, despite the increase of interest in this task, existing works face a number of shortcomings: system scalability, restrictive modeling, and weak emphasis on maximizing user emotional experience. In this paper, we build a fully data driven chat-oriented dialogue system that can dynamically mimic affective human interactions by utilizing a neural network architecture. In particular, we propose a sequence-to-sequence response generator that considers the emotional context of the dialogue. An emotion encoder is trained jointly with the entire network to encode and maintain the emotional context throughout the dialogue. The encoded emotion information is then incorporated in the response generation process. We train the network with a dialogue corpus that contains positive-emotion eliciting responses, collected through crowd-sourcing. Objective evaluation shows that incorporation of emotion into the training process helps reduce the perplexity of the generated responses, even when a small dataset is used. Subsequent subjective evaluation shows that the proposed method produces responses that are more natural and likely to elicit a more positive emotion. Nurul Lubis, Sakriani Sakti, Koichiro Yoshino, Satoshi Nakamura 0001 |
AAAI | 4 |
| 2018 | TRANS-AM: Discovery Method of Optimal Input Vectors Corresponding to Objective Variables
Yu Suzuki 0001, Koichiro Yoshino, Satoshi Nakamura 0001 |
DaWaK | 4 |
| 2018 | Information Filtering Method for Twitter Streaming Data Using Human-in-the-Loop Machine Learning
Yu Suzuki 0001, Satoshi Nakamura 0001 |
DEXA (2) | 2 |
| 2018 | Graph Regularized Tensor Factorization for Single-Trial EEG AnalysisabstractThis study proposes a tensor factorization algorithm for electroencephalographies (EEGs) that incorporates the geometric structure of the electrode location. The purpose is removing noise caused by EEG activities which are irrelevant to stimuli presented to a subject from single-trial event-related potential (ERP) data. Canonical polyadic decomposition (CPD) is extended by adding a regularization term that controls the spatial smoothness of the decomposed components on a scalp. An initialization method using geometrical information is also proposed. The geometric structure of an EEG signal is expressed as an undirected graph where the similarities between electrodes are defined by their relative distances on a scalp. The effectiveness is demonstrated in a noise-removing experiment using pseudo-ERP, where the proposed method achieved better performance than the conventional CPD. Hayato Maki, Hiroki Tanaka, Sakriani Sakti, Satoshi Nakamura 0001 |
ICASSP | 4 |
| 2018 | Sequence-to-Sequence Asr Optimization Via Reinforcement LearningabstractDespite the success of sequence-to-sequence approaches in automatic speech recognition (ASR) systems, the models still suffer from several problems, mainly due to the mismatch between the training and inference conditions. In the sequence-to-sequence architecture, the model is trained to predict the grapheme of the current time-step given the input of speech signal and the ground-truth grapheme history of the previous time-steps. However, it remains unclear how well the model approximates real-world speech during inference. Thus, generating the whole transcription from scratch based on previous predictions is complicated and errors can propagate over time. Furthermore, the model is optimized to maximize the likelihood of training data instead of error rate evaluation metrics that actually quantify recognition quality. This paper presents an alternative strategy for training sequence-to-sequence ASR models by adopting the idea of reinforcement learning (RL). Unlike the standard training scheme with maximum likelihood estimation, our proposed approach utilizes the policy gradient algorithm. We can (1) sample the whole transcription based on the model's prediction in the training process and (2) directly optimize the model with negative Levenshtein distance as the reward. Experimental results demonstrate that we significantly improved the performance compared to a model trained onlv with maximum likelihood estimation. Andros Tjandra, Sakriani Sakti, Satoshi Nakamura 0001 |
ICASSP | 3 |
| 2018 | Listening Skills Assessment through Computer AgentsabstractSocial skills training, performed by human trainers, is a well-established method for obtaining appropriate skills in social interaction. Previous work automated the process of social skills training by developing a dialogue system that teaches social skills through interaction with a computer agent. Even though previous work that simulated social skills training considered speaking skills, human social skills trainers take into account other skills such as listening. In this paper, we propose assessment of user listening skills during conversation with computer agents toward automated social skills training. We recorded data of 27 Japanese graduate students interacting with a female agent. The agent spoke to the participants about a recent memorable story and how to make a telephone call, and the participants listened. Two expert external raters assessed the participants' listening skills. We manually extracted features relating to eye fixation and behavioral cues of the participants, and confirmed that a simple linear regression with selected features can correctly predict a user's listening skills with above 0.45 correlation coefficient. Hiroki Tanaka, Hideki Negoro, Hidemi Iwasaka, Satoshi Nakamura 0001 |
ICMI | 4 |
| 2018 | Tensor Decomposition for Compressing Recurrent Neural NetworkabstractIn the machine learning fields, Recurrent Neural Network (RNN) has become a popular architecture for sequential data modeling. However, behind the impressive performance, RNNs require a large number of parameters for both training and inference. In this paper, we are trying to reduce the number of parameters and maintain the expressive power from RNN simultaneously. We utilize several tensor decompositions method including CANDECOMP/PARAFAC (CP), Tucker decomposition and Tensor Train (TT) to re-parameterize the Gated Recurrent Unit (GRU) RNN. We evaluate all tensor-based RNNs performance on sequence modeling tasks with a various number of parameters. Based on our experiment results, TT-GRU achieved the best results in a various number of parameters compared to other decomposition methods. Andros Tjandra, Sakriani Sakti, Satoshi Nakamura 0001 |
IJCNN | 3 |
| 2018 | Compressing End-to-end ASR Networks by Tensor-Train Decomposition
Takuma Mori, Andros Tjandra, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2018 | Machine Speech Chain with One-shot Speaker AdaptationabstractIn previous work, we developed a closed-loop speech chain model based on deep learning, in which the architecture enabled the automatic speech recognition (ASR) and text-to-speech synthesis (TTS) components to mutually improve their performance.This was accomplished by the two parts teaching each other using both labeled and unlabeled data.This approach could significantly improve model performance within a single-speaker speech dataset, but only a slight increase could be gained in multi-speaker tasks.Furthermore, the model is still unable to handle unseen speakers.In this paper, we present a new speech chain mechanism by integrating a speaker recognition model inside the loop.We also propose extending the capability of TTS to handle unseen speakers by implementing one-shot speaker adaptation.This enables TTS to mimic voice characteristics from one speaker to another with only a one-shot speaker sample, even from a text without any speaker information.In the speech chain loop mechanism, ASR also benefits from the ability to further learn an arbitrary speakers characteristics from the generated speech waveform, resulting in a significant improvement in the recognition rate. Andros Tjandra, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2018 | Detection of Dementia from Responses to Atypical Questions Asked by Embodied Conversational Agents
Tsuyoki Ujiro, Hiroki Tanaka, Hiroyoshi Adachi, Hiroaki Kazui, Manabu Ikeda, Takashi Kudo, Satoshi Nakamura 0001 |
INTERSPEECH | 7 |
| 2018 | Incremental TTS for Japanese Language
Tomoya Yanagita, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2018 | Construction of English-French Multimodal Affective Conversational Corpus from TV Dramas
Sashi Novitasari, Quoc Truong Do, Sakriani Sakti, Dessi Puji Lestari, Satoshi Nakamura 0001 |
LREC | 5 |
| 2018 | Dialogue Scenario Collection of Persuasive Dialogue with Emotional Expressions via Crowdsourcing
Koichiro Yoshino, Yoko Ishikawa, Masahiro Mizukami, Yu Suzuki 0001, Sakriani Sakti, Satoshi Nakamura 0001 |
LREC | 6 |
| 2018 | Japanese Dialogue Corpus of Information Navigation and Attentive Listening Annotated with Extended ISO-24617-2 Dialogue Act Tags
Koichiro Yoshino, Hiroki Tanaka, Kyoshiro Sugiyama, Makoto Kondo, Satoshi Nakamura 0001 |
LREC | 5 |
| 2018 | Guiding Neural Machine Translation with Retrieved Translation PiecesabstractJingyi Zhang, Masao Utiyama, Eiichro Sumita, Graham Neubig, Satoshi Nakamura. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Jingyi Zhang 0001, Masao Utiyama, Eiichiro Sumita, Graham Neubig, Satoshi Nakamura 0001 |
NAACL-HLT | 5 |
| 2018 | Unsupervised Counselor Dialogue Clustering for Positive Emotion Elicitation in Neural Dialogue SystemabstractPositive emotion elicitation seeks to improve user's emotional state through dialogue system interaction, where a chatbased scenario is layered with an implicit goal to address user's emotional needs.Standard neural dialogue system approaches still fall short in this situation as they tend to generate only short, generic responses.Learning from expert actions is critical, as these potentially differ from standard dialogue acts.In this paper, we propose using a hierarchical neural network for response generation that is conditioned on 1) expert's action, 2) dialogue context, and 3) user emotion, encoded from user input.We construct a corpus of interactions between a counselor and 30 participants following a negative emotional exposure to learn expert actions and responses in a positive emotion elicitation scenario.Instead of relying on the expensive, labor intensive, and often ambiguous human annotations, we unsupervisedly cluster the expert's responses and use the resulting labels to train the network.Our experiments and evaluation show that the proposed approach yields lower perplexity and generates a larger variety of responses. Nurul Lubis, Sakriani Sakti, Koichiro Yoshino, Satoshi Nakamura 0001 |
SIGDIAL Conference | 4 |
| 2018 | Toward Multi-Features Emphasis Speech Translation: Assessment of Human Emphasis Production and Perception with Speech and Text CluesabstractEmphasis is an important factor of human speech that helps convey emotion and the focused information of utterances. Recently, studies have been conducted on speech-to-speech translation to preserve the emphasis information from the source language to the target language. However, since different cultures have various ways of expressing emphasis, just considering the acoustic-to-acoustic feature emphasis translation may not always reflect the experiences of users. On the other hand, emphasis can be expressed at various levels in both text and speech. However, it remains unclear how we communicate emphasis in a different form (acoustic/linguistic) with different levels and whether we can perceive the difference between different levels of emphasis or observe the similarity of the same emphasis levels in both text and speech forms. In this paper, we conducted analyses on human perception of emphasis with both speech and text clues through crowd-sourced evaluations. The results indicate that although participants can distinguish among emphasis levels and perceive the same emphasis level between speech and text, many ambiguities still exist at certain emphasis levels. Thus, our result provides insight into what needs to be handled during the emphasis translation process. Quoc Truong Do, Sakriani Sakti, Satoshi Nakamura 0001 |
SLT | 3 |
| 2018 | Optimizing Neural Response Generator with Emotional Impact InformationabstractThe potential of dialogue systems to address user's emotional need has steadily grown. In particular, we focus on dialogue systems application to promote positive emotional states, similar to that of emotional support between humans. Positive emotion elicitation takes form as chat-based dialogue interactions that is layered with an implicit goal to improve user's emotional state. To this date, existing approaches have only relied on mimicking the target responses without considering their emotional impact, i.e. the change of emotional state they cause on the listener, in the model itself. In this paper, we propose explicitly utilizing emotional impact information to optimize neural dialogue system towards generating responses that elicit positive emotion. We examine two emotion-rich corpora with different data collection scenarios: Wizard-of-Oz and spontaneous. Evaluation shows that the proposed method yields lower perplexity, as well as produces responses that are perceived as more natural and likely to elicit a more positive emotion. Nurul Lubis, Sakriani Sakti, Koichiro Yoshino, Satoshi Nakamura 0001 |
SLT | 4 |
| 2018 | Speech Chain for Semi-Supervised Learning of Japanese-English Code-Switching ASR and TTSabstractCode-switching (CS) speech, in which speakers alternate between two or more languages in the same utterance, often occurs in multilingual communities. Such a phenomenon poses challenges for spoken language technologies: automatic speech recognition (ASR) and text-to-speech synthesis (TTS), since the systems need to be able to handle the input in a multilingual setting. We may find code-switching text or code-switching speech in social media, but parallel speech and the transcriptions of code-switching data, which are suitable for training ASR and TTS, are generally unavailable. In this paper, we utilize a speech chain framework based on deep learning to enable ASR and TTS to learn code-switching in a semi-supervised fashion. We base our system on Japanese-English conversational speech. We first separately train the ASR and TTS systems with parallel speech-text of monolingual data (supervised learning) and perform a speech chain with only code-switching text or code-switching speech (unsupervised learning). Experimental results reveal that such closed-loop architecture allows ASR and TTS to learn from each other and improve the performance even without any parallel code-switching data. Sahoko Nakayama, Andros Tjandra, Sakriani Sakti, Satoshi Nakamura 0001 |
SLT | 4 |
| 2018 | Adaptive Wavenet Vocoder for Residual Compensation in GAN-Based Voice ConversionabstractIn this paper, we propose to use generative adversarial networks (GAN) together with a WaveNet vocoder to address the over-smoothing problem arising from the deep learning approaches to voice conversion, and to improve the vocoding quality over the traditional vocoders. As GAN aims to minimize the divergence between the natural and converted speech parameters, it effectively alleviates the over-smoothing problem in the converted speech. On the other hand, WaveNet vocoder allows us to leverage from the human speech of a large speaker population, thus improving the naturalness of the synthetic voice. Furthermore, for the first time, we study how to use WaveNet vocoder for residual compensation to improve the voice conversion performance. The experiments show that the proposed voice conversion framework consistently outperforms the baselines. Berrak Sisman, Mingyang Zhang 0003, Sakriani Sakti, Haizhou Li 0001, Satoshi Nakamura 0001 |
SLT | 5 |
| 2018 | Multi-Scale Alignment and Contextual History for Attention Mechanism in Sequence-to-Sequence ModelabstractA sequence-to-sequence model is a neural network module for mapping two sequences of different lengths. The sequence-to-sequence model has three core modules: encoder, decoder, and attention. Attention is the bridge that connects the encoder and decoder modules and improves model performance in many tasks. In this paper, we propose two ideas to improve sequence-to-sequence model performance by enhancing the attention module. First, we maintain the history of the location and the expected context from several previous time-steps. Second, we apply multiscale convolution from several previous attention vectors to the current decoder state. We utilized our proposed framework for sequence-to-sequence speech recognition and text-to-speech systems. The results reveal that our proposed extension can improve performance significantly compared to a standard attention baseline. Andros Tjandra, Sakriani Sakti, Satoshi Nakamura 0001 |
SLT | 3 |
| 2018 | An end-to-end model for cross-lingual transformation of paralinguistic information
Takatomo Kano, Shinnosuke Takamichi, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura 0001 |
Mach. Transl. | 6 |
| 2018 | Intra-gender statistical singing voice conversion with direct waveform modification using log-spectral differentialabstractThis paper presents a novel intra-gender statistical singing voice conversion (SVC) technique with direct waveform modification based on the log-spectrum differential (DIFFSVC) that can convert the voice timbre of a source singer into that of a target singer without vocoder-based waveform generation of the converted singing voice. SVC makes it possible to convert the singing voice characteristics of an arbitrary source singer into those of an arbitrary target singer by converting some of its acoustic features, such as F0, aperiodicity, and spectral features based on a statistical conversion function. However, the sound quality of the converted singing voice is typically degraded compared with that of a natural singing voice, owing to various factors, such as analysis and modeling errors in the vocoding process and over-smoothing of the converted feature trajectory. To alleviate sound quality degradation, we propose a statistical conversion process that directly modifies the signal in the waveform domain by estimating the difference in the spectra of the source and target singers’ singing voices. Additionally, we propose the following several techniques for the DIFFSVC method: 1) derivation of a differential Gaussian mixture model (DIFFGMM) from a conventional Gaussian mixture model (GMM) and 2) a parameter generation algorithm considering the global variance (GV). The experimental results demonstrate that the proposed DIFFSVC methods enable significant improvements in the sound quality of the converted singing voice, while preserving the conversion accuracy of the singer’s identity compared with conventional SVC. Kazuhiro Kobayashi, Tomoki Toda, Satoshi Nakamura 0001 |
Speech Commun. | 3 |
| 2018 | Sequence-to-Sequence Models for Emphasis Speech TranslationabstractSpeech-to-speech translation (S2ST) systems are capable of breaking language barriers in cross-lingual communication by translating speech across languages. Recent studies have introduced many improvements that allow existing S2ST systems to handle not only linguistic meaning but also paralinguistic information such as emphasis by proposing additional emphasis estimation and translation components. However, the approach used for emphasis translation is not optimal for sequence translation tasks and fails to easily handle the long-term dependencies of words and emphasis levels. It also requires the quantization of emphasis levels and treats them as discrete labels instead of continuous values. Moreover, the whole translation pipeline is fairly complex and slow because all components are trained separately without joint optimization. In this paper, we make two contributions: 1) we propose an approach that can handle continuous emphasis levels based on sequence-to-sequence models, and 2) we combine machine and emphasis translation into a single model, which greatly simplifies the translation pipeline and make it easier to perform joint optimization. Our results on an emphasis translation task indicate that our translation models outperform previous models by a large margin in both objective and subjective tests. Experiments on a joint translation model also show that our models can perform joint translation of words and emphasis with one-word delays instead of full-sentence delays while preserving the translation performance of both tasks. Quoc Truong Do, Sakriani Sakti, Satoshi Nakamura 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Dirichlet Process Mixture of Mixtures Model for Unsupervised Subword ModelingabstractWe develop a parallelizable Markov chain Monte Carlo sampler for a Dirichlet process mixture of mixtures model. Our sampler jointly infers a codebook and clusters. The codebook is a global collection of components. Clusters are mixtures, defined over the codebook. We combine a nonergodic Gibbs sampler with two layers of split and merge samplers on codebook and mixture level to form a valid ergodic chain. We design an additional switch sampler for components that supports convergence in our experimental results. In the use case of unsupervised subword modeling, we show that our method infers complex classes from real speech feature vectors that consistently show higher quality on several evaluation metrics. At the same time, we infer fewer classes that represent subword units more consistently and show longer durations, compared to a standard Dirichlet process mixture model sampler. Michael Heck, Sakriani Sakti, Satoshi Nakamura 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | Processing negative emotions through social communication: Multimodal database construction and analysisabstractEmotion-rich data is pre-requisite in the efforts of transferring emotional aspects of human communication into Human-Computer Interaction (HCI). An important facet of human social-affective interaction is its ability to facilitate social sharing of emotion, a fundamental part of the emotional processes. When conducted properly, such an interaction can give a positive effect to emotion-related problems. However, there is still a lack of resources that are: 1) explicitly designed for studying the emotional problems commonly encountered in everyday life, and 2) involving a professional as an expert in the conversation. In this paper, we present recordings of dyadic social-affective interactions between a professional counselor as an expert and 30 participants, summing up to 23 hours and 41 minutes of material. In each interaction, a negative emotion inducer is shown to the dyad, and the goal of the expert is to aid emotion processing and elicit a positive emotional change through the interaction. Specifically, we aim to observe how an external party can guide and facilitate emotion processing, especially after a negative emotional response in a commonly encountered social situation. The construction, development, and analysis of the database is detailed in this paper. Nurul Lubis, Michael Heck, Sakriani Sakti, Koichiro Yoshino, Satoshi Nakamura 0001 |
ACII | 5 |
| 2017 | Neural Machine Translation via Binary Code PredictionabstractIn this paper, we propose a new method for calculating the output layer in neural machine translation systems.The method is based on predicting a binary code for each word and can reduce computation time/memory requirements of the output layer to be logarithmic in vocabulary size in the best case.In addition, we also introduce two advanced approaches to improve the robustness of the proposed model: using error-correcting codes and combining softmax and binary codes.Experiments on two English ↔ Japanese bidirectional translation tasks show proposed models achieve BLEU scores that approach the softmax, while reducing memory usage to the order of less than 1/10 and improving decoding speed on CPUs by x5 to x10. Yusuke Oda, Philip Arthur, Graham Neubig, Koichiro Yoshino, Satoshi Nakamura 0001 |
ACL (1) | 5 |
| 2017 | Feature optimized DPGMM clustering for unsupervised subword modeling: A contribution to zerospeech 2017abstractThis paper describes our unsupervised subword modeling pipeline for the zero resource speech challenge (ZeroSpeech) 2017. Our approach is built around the Dirichlet process Gaussian mixture model (DPGMM) that we use to cluster speech feature vectors into a dynamically sized set of classes. By considering each class an acoustic unit, speech can be represented as sequence of class posteriorgrams. We enhance this method by automatically optimizing the DPGMM sampler's input features in a multi-stage clustering framework, where we unsupervisedly learn transformations using LDA, MLLT and (basis) fMLLR to reduce variance in the features. We show that this optimization considerably boosts the subword modeling quality, according to the performance on the ABX phone discriminability task. For the first time, we apply inferred subword models to previously unseen data from a new set of speakers. We demonstrate our method's good generalization and the effectiveness of its blind speaker adaptation in extensive experiments on a multitude of datasets. Our pipeline has very little need for hyper-parameter adjustment and is entirely unsupervised, i.e., it only takes raw audio recordings as input, without requiring any pre-defined segmentation, explicit speaker IDs or other meta data. Michael Heck, Sakriani Sakti, Satoshi Nakamura 0001 |
ASRU | 3 |
| 2017 | Listening while speaking: Speech chain by deep learningabstractDespite the close relationship between speech perception and production, research in automatic speech recognition (ASR) and text-to-speech synthesis (TTS) has progressed more or less independently without exerting much mutual influence on each other. In human communication, on the other hand, a closed-loop speech chain mechanism with auditory feedback from the speaker's mouth to her ear is crucial. In this paper, we take a step further and develop a closed-loop speech chain model based on deep learning. The sequence-to-sequence model in close-loop architecture allows us to train our model on the concatenation of both labeled and unlabeled data. While ASR transcribes the unlabeled speech features, TTS attempts to reconstruct the original speech waveform based on the text from ASR. In the opposite direction, ASR also attempts to reconstruct the original text transcription given the synthesized speech. To the best of our knowledge, this is the first deep learning model that integrates human speech perception and production behaviors. Our experimental results show that the proposed approach significantly improved the performance more than separate systems that were only trained with labeled data. Andros Tjandra, Sakriani Sakti, Satoshi Nakamura 0001 |
ASRU | 3 |
| 2017 | Attention-based Wav2Text with feature transfer learningabstractConventional automatic speech recognition (ASR) typically performs multi-level pattern recognition tasks that map the acoustic speech waveform into a hierarchy of speech units. But, it is widely known that information loss in the earlier stage can propagate through the later stages. After the resurgence of deep learning, interest has emerged in the possibility of developing a purely end-to-end ASR system from the raw waveform to the transcription without any predefined alignments and hand-engineered models. However, the successful attempts in end-to-end architecture still used spectral-based features, while the successful attempts in using raw waveform were still based on the hybrid deep neural network - Hidden Markov model (DNN-HMM) framework. In this paper, we construct the first end-to-end attention-based encoder-decoder model to process directly from raw speech waveform to the text transcription. We called the model as Attention-based Wav2Text. To assist the training process of the end-to-end model, we propose to utilize a feature transfer learning. Experimental results also reveal that the proposed Attention-based Wav2Text model directly with raw waveform could achieve a better result in comparison with the attentional encoder-decoder model trained on standard front-end filterbank features. Andros Tjandra, Sakriani Sakti, Satoshi Nakamura 0001 |
ASRU | 3 |
| 2017 | A trade-off between estimation accuracy of worker quality and task complexityabstractIn crowdsourcing, many people are less capable of producing quality work, and there are those who work inadequately. We can improve the quality of work, and also we can decrease time and wages if we eliminate poor workers and give extra instruction to their workers. Therefore, estimating work quality is essential for uncovering poor workers. In existing studies, the response behavior of workers was used to estimate their quality. However, in these studies, the authors only apply to complicated tasks that have many types of response behavior. In this paper, we propose a method for estimating the quality of workers by their response behavior by intentionally complicating a simple task. By doing so, we can get more accurate and detailed response behavior. By using accurate and detailed response behavior even in simple tasks that have few types of response behavior, the estimation accuracy of low-quality workers improved. However, workers had to work for slightly longer. Yoshitaka Matsuda, Yu Suzuki 0001, Satoshi Nakamura 0001 |
IEEE BigData | 3 |
| 2017 | Tracking liking state in brain activity while watching multiple moviesabstractEmotion is a valuable information in various applications ranging from human-computer interaction to automated multimedia content delivery. Conventional methods to recognize emotion were based on speech prosody cues, facial expression, and body language. However, this information may not appear when people watch a movie. In recent years, some studies have started to use electroencephalogram (EEG) signals in recognizing emotion. But, the EEG data were entirely analyzed in each scene of movies for emotion classification. Thus, the detailed information of emotional state changes cannot be extracted. In this study, we utilize EEG to track affective state during watching multiple movies. Experiments were done by measuring continuous liking state during watching three types of movies, and then constructing subject dependent emotional state tracking model. We used support vector machine (SVM) as a classifier, and support vector regression (SVR) for regression. As a result, the best classification accuracy was 77.6%, and the best regression model achieved 0.645 of correlation coefficient between actual liking state and predicted liking state. These results demonstrate that continuous emotional state can be predicted by our EEG-based method. Naoto Terasawa, Hiroki Tanaka, Sakriani Sakti, Satoshi Nakamura 0001 |
ICMI | 4 |
| 2017 | Acquisition and Assessment of Semantic Content for the Generation of Elaborateness and Indirectness in Spoken Dialogue SystemsabstractIn a dialogue system, the dialogue manager selects one of several system actions and thereby determines the system’s behaviour. Defining all possible system actions in a dialogue system by hand is a tedious work. While efforts have been made to automatically generate such system actions, those approaches are mostly focused on providing functional system behaviour. Adapting the system behaviour to the user becomes a difficult task due to the limited amount of system actions available. We aim to increase the adaptability of a dialogue system by automatically generating variants of system actions. In this work, we introduce an approach to automatically generate action variants for elaborateness and indirectness. Our proposed algorithm extracts RDF triplets from a knowledge base and rates their relevance to the original system action to find suitable content. We show that the results of our algorithm are mostly perceived similarly to human generated elaborateness and indirectness and can be used to adapt a conversation to the current user and situation. We also discuss where the results of our algorithm are still lacking and how this could be improved: Taking into account the conversation topic as well as the culture of the user is likely to have beneficial effect on the user’s perception. Louisa Pragst, Koichiro Yoshino, Wolfgang Minker, Satoshi Nakamura 0001, Stefan Ultes |
IJCNLP(1) | 4 |
| 2017 | Local Monotonic Attention Mechanism for End-to-End Speech And Language ProcessingabstractRecently, encoder-decoder neural networks have shown impressive performance on many sequence-related tasks. The architecture commonly uses an attentional mechanism which allows the model to learn alignments between the source and the target sequence. Most attentional mechanisms used today is based on a global attention property which requires a computation of a weighted summarization of the whole input sequence generated by encoder states. However, it is computationally expensive and often produces misalignment on the longer input sequence. Furthermore, it does not fit with monotonous or left-to-right nature in several tasks, such as automatic speech recognition (ASR), grapheme-to-phoneme (G2P), etc. In this paper, we propose a novel attention mechanism that has local and monotonic properties. Various ways to control those properties are also explored. Experimental results on ASR, G2P and machine translation between two languages with similar sentence structures, demonstrate that the proposed encoder-decoder model with local monotonic attention could achieve significant performance improvements and reduce the computational complexity in comparison with the one that used the standard global attention architecture. Andros Tjandra, Sakriani Sakti, Satoshi Nakamura 0001 |
IJCNLP(1) | 3 |
| 2017 | Improving Neural Machine Translation through Phrase-based Forced DecodingabstractCompared to traditional statistical machine translation (SMT), neural machine translation (NMT) often sacrifices adequacy for the sake of fluency. We propose a method to combine the advantages of traditional SMT and NMT by exploiting an existing phrase-based SMT model to compute the phrase-based decoding cost for an NMT output and then using the phrase-based decoding cost to rerank the n-best NMT outputs. The main challenge in implementing this approach is that NMT outputs may not be in the search space of the standard phrase-based decoding algorithm, because the search space of phrase-based SMT is limited by the phrase-based translation rule table. We propose a soft forced decoding algorithm, which can always successfully find a decoding path for any NMT output. We show that using the forced decoding cost to rerank the NMT outputs can successfully improve translation quality on four different language pairs. Jingyi Zhang 0001, Masao Utiyama, Eiichiro Sumita, Graham Neubig, Satoshi Nakamura 0001 |
IJCNLP(1) | 5 |
| 2017 | Compressing recurrent neural network with tensor trainabstractRecurrent Neural Network (RNN) are a popular choice for modeling temporal and sequential tasks and achieve many state-of-the-art performance on various complex problems. However, most of the state-of-the-art RNNs have millions of parameters and require many computational resources for training and predicting new data. This paper proposes an alternative RNN model to reduce the number of parameters significantly by representing the weight parameters based on Tensor Train (TT) format. In this paper, we implement the TT-format representation for several RNN architectures such as simple RNN and Gated Recurrent Unit (GRU). We compare and evaluate our proposed RNN model with uncompressed RNN model on sequence classification and sequence prediction tasks. Our proposed RNNs with TT-format are able to preserve the performance while reducing the number of RNN parameters significantly up to 40 times smaller. Andros Tjandra, Sakriani Sakti, Satoshi Nakamura 0001 |
IJCNN | 3 |
| 2017 | Toward Expressive Speech Translation: A Unified Sequence-to-Sequence LSTMs Approach for Translating Words and Emphasis
Quoc Truong Do, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2017 | Ensembles of Multi-Scale VGG Acoustic Models
Michael Heck, Masayuki Suzuki, Takashi Fukuda, Gakuto Kurata, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2017 | Structured-Based Curriculum Learning for End-to-End English-Japanese Speech TranslationabstractSequence-to-sequence attentional-based neural network architectures have been shown to provide a powerful model for machine translation and speech recognition.Recently, several works have attempted to extend the models for end-to-end speech translation task.However, the usefulness of these models were only investigated on language pairs with similar syntax and word order (e.g., English-French or English-Spanish).In this work, we focus on end-to-end speech translation tasks on syntactically distant language pairs (e.g., English-Japanese) that require distant word reordering.To guide the encoder-decoder attentional model to learn this difficult problem, we propose a structured-based curriculum learning strategy.Unlike conventional curriculum learning that gradually emphasizes difficult data examples, we formalize learning strategies from easier network structures to more difficult network structures.Here, we start the training with end-to-end encoder-decoder for speech recognition or text-based machine translation task then gradually move to end-to-end speech translation task.The experiment results show that the proposed approach could provide significant improvements in comparison with the one without curriculum learning. Takatomo Kano, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2017 | Physically Constrained Statistical F0 Prediction for Electrolaryngeal Speech Enhancement
Kou Tanaka, Hirokazu Kameoka, Tomoki Toda, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2017 | Subject-Independent Classification of Japanese Spoken Sentences by Multiple Frequency Bands Phase Pattern of EEG Response During Speech Perception
Hiroki Tanaka, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2017 | Information Navigation System with Discovering User InterestsabstractWe demonstrate an information navigation system for sightseeing domains that has a dialogue interface for discovering user interests for tourist activities.The system discovers interests of a user with focus detection on user utterances, and proactively presents related information to the discovered user interest.A partially observable Markov decision process (POMDP)-based dialogue manager, which is extended with user focus states, controls the behavior of the system to provide information with several dialogue acts for providing information.We transferred the belief-update function and the policy of the manager from other system trained on a di↵erent domain to show the generality of defined dialogue acts for our information navigation system. Koichiro Yoshino, Yu Suzuki 0001, Satoshi Nakamura 0001 |
SIGDIAL Conference | 3 |
| 2017 | Semantically readable distributed representation learning for social media miningabstractThe problem with distributed representations generated by neural networks is that the meaning of the features is difficult to understand. We propose a new method that gives a specific meaning to each node of a hidden layer by introducing a manually created word semantic vector dictionary into the initial weights and by using paragraph vector models. Our experimental results demonstrated that weights obtained based on learning and weights based on the dictionary are more strongly correlated in a closed test and more weakly correlated in an open test, compared with the results of a control test. Additionally, we found that the learned vector are better than the performance of the existing paragraph vector in the evaluation of the sentiment analysis task. Finally, we determined the readability of document embedding in a user test. The definition of readability in this paper is that people can understand the meaning of large weighted features of distributed representations. A total of 52.4% of the top five weighted hidden nodes were related to tweets where one of the paragraph vector models learned the document embedding. Because each hidden node maintains a specific meaning, the proposed method succeeds in improving readability. Ikuo Keshi, Yu Suzuki 0001, Koichiro Yoshino, Satoshi Nakamura 0001 |
WI | 4 |
| 2017 | Transcribing against time
Matthias Sperber, Graham Neubig, Jan Niehues, Satoshi Nakamura 0001, Alex Waibel |
Speech Commun. | 4 |
| 2017 | Preserving Word-Level Emphasis in Speech-to-Speech TranslationabstractSpeech-to-speech translation (S2ST) is a technology that translates speech across languages, which can remove barriers in cross-lingual communication. In the conventional S2ST systems, the linguistic meaning of speech was translated, but paralinguistic information conveying other features of the speech such as emotion or emphasis were ignored. In this paper, we propose a method to translate paralinguistic information, specifically focusing on emphasis. The method consists of a series of components that can accurately translate emphasis using all acoustic features of speech. First, linear-regression hidden semi-Markov models (LRHSMMs) are used to estimate a real-numbered emphasis value for every word in an utterance, resulting in a sequence of values for the utterance. After that the emphasis translation module translates the estimated emphasis sequence into a target language emphasis sequence using a conditional random field model considering the features of emphasis levels, words, and part-of-speech tags. Finally, the speech synthesis module synthesizes emphasized speech with LR-HSMMs, taking into account the translated emphasis sequence and transcription. The results indicate that our translation model can translate emphasis information, correctly emphasizing words in the target language with 91.6% F-measure by objective evaluation. A listening test with human subjects further showed that they could identify the emphasized words with 87.8% F-measure, and that the naturalness of the audio was preserved. Quoc Truong Do, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2016 | A Continuous Space Rule Selection Model for Syntax-based Statistical Machine Translation
Jingyi Zhang 0001, Masao Utiyama, Eiichiro Sumita, Graham Neubig, Satoshi Nakamura 0001 |
ACL (1) | 5 |
| 2016 | Learning a Lexicon and Translation Model from Phoneme LatticesabstractLanguage documentation begins by gathering speech.Manual or automatic transcription at the word level is typically not possible because of the absence of an orthography or prior lexicon, and though manual phonemic transcription is possible, it is prohibitively slow.On the other hand, translations of the minority language into a major language are more easily acquired.We propose a method to harness such translations to improve automatic phoneme recognition.The method assumes no prior lexicon or translation model, instead learning them from phoneme lattices and translations of the speech being transcribed.Experiments demonstrate phoneme error rate improvements against two baselines and the model's ability to learn useful bilingual lexical entries. Oliver Adams, Graham Neubig, Trevor Cohn, Steven Bird, Quoc Truong Do, Satoshi Nakamura 0001 |
EMNLP | 6 |
| 2016 | Incorporating Discrete Translation Lexicons into Neural Machine TranslationabstractNeural machine translation (NMT) often makes mistakes in translating low-frequency content words that are essential to understanding the meaning of the sentence.We propose a method to alleviate this problem by augmenting NMT systems with discrete translation lexicons that efficiently encode translations of these low-frequency words.We describe a method to calculate the lexicon probability of the next word in the translation candidate by using the attention vector of the NMT model to select which source word lexical probabilities the model should focus on.We test two methods to combine this probability with the standard NMT probability: (1) using it as a bias, and (2) linear interpolation.Experiments on two corpora show an improvement of 2.0-2.3BLEU and 0.13-0.44NIST score, and faster convergence time. 1 Philip Arthur, Graham Neubig, Satoshi Nakamura 0001 |
EMNLP | 3 |
| 2016 | Implementation of F0 transformation for statistical singing voice conversion based on direct waveform modificationabstractThis paper presents a technique for transforming F0in a framework of statistical singing voice conversion with direct waveform modification based on spectrum differential (DIFFSVC). The DIFFSVC method converts voice timbre of singing voices of a source singer into that of a target singer without using vocoder-based waveform generation. Although this method achieves high sound quality of the converted singing voices, its use is limited to only intra-gender conversion without the need of F0transformation. To make it possible to also use the DIFFSVC method for cross-gender conversion, we propose a method to transform F0of an input singing voice for the DIFFSVC. The proposed method is also based on direct waveform modification using overlap-add process and filtering process. Results of subjective evaluations demonstrate that the proposed DIFFSVC method with F0transformation significantly improves sound quality of the converted singing voices while preserving the conversion accuracy of singer identity in the cross-gender conversion compared to the conventional SVC with vocoder. Kazuhiro Kobayashi, Tomoki Toda, Satoshi Nakamura 0001 |
ICASSP | 3 |
| 2016 | Noise suppression method for body-conducted soft speech enhancement based on external noise monitoringabstractThis paper presents a novel approach to suppressing adverse effects of external noise on body-conducted soft speech for silent speech communication in noisy environments. Nonaudible murmur (NAM) microphone as one of the body-conductive microphones is capable of detecting very soft speech. However, body-conducted soft speech easily suffers from external noise owing to its faint volume. To address this issue, the proposed method additionally uses an air-conductive microphone to detect only an external noise signal and uses the detected external noise signal to suppress its effect on the body-conducted soft speech. A semi-blind source separation technique is appüed to the proposed method for estimating a linear filter to suppress the noise components without voice activity detection. Experimental results demonstrate that the proposed method yields 10 dB SNR improvements in 80 dBA noisy conditions and also yields significant improvements in sound quality of body-conducted soft speech. Yusuke Tajiri, Tomoki Toda, Satoshi Nakamura 0001 |
ICASSP | 3 |
| 2016 | Statistical F0 prediction for electrolaryngeal speech enhancement considering generative process of F0 contours within product of experts frameworkabstractWe have previously proposed a statistical fundamental frequency (F0) prediction method that makes it possible to predict the underlying F0contour of electrolaryngeal (EL) speech from its spectral feature sequence. Although this method was shown to contribute to improving the naturalness of EL speech as a whole, the predicted F0contour was still unnatural compared with that in normal speech. One possible solution to improve the naturalness of the predicted F0contours would be to take account of the physical mechanism of vocal phonation. Recently a statistical model of voice F0contours was formulated by constructing a stochastic counterpart of the Fujisaki model, a well-founded mathematical model representing the control mechanism of vocal fold vibration. This paper proposes a Product-of-Experts model to incorporate this generative model of voice F0contours into the statistical F0prediction model. Based on the constructed model, we derive algorithms for parameter training and F0prediction. Experimental results revealed that the proposed method successfully outperformed our previously proposed method in terms of the naturalness of the predicted F0contours. Kou Tanaka, Hirokazu Kameoka, Tomoki Toda, Satoshi Nakamura 0001 |
ICASSP | 4 |
| 2016 | An estimation method of voice timbre evaluation values using feature extraction with Gaussian mixture model based on reference singerabstractThis paper presents an estimation method of voice timbre evaluation values for arbitrary singer's singing voices generated with a singing voice synthesis system towards the development of a singing voice retrieval system. The voice timbre evaluation values are numerical values corresponding to voice timbre expression words, such as "Age" and "Gender", and they usually need to be manually assigned to individual singers' singing voices through listening. To make it possible to automatically estimate them from given singer's singing voices, an acoustic feature to well capture only each singer's voice timbre is extracted with a Gaussian mixture model trained using parallel data between singing voices sung by many pre-stored target singers and same voices sung by a reference singer. Then, the voice timbre evaluation values are estimated from the extracted feature using regression models. The experimental results showed that the proposed method is capable of accurately estimating those values for some expression words, such as "Age" and "Gender", and nonlinear regression is effective for the expression words, "Powerfulness" and "Uniqueness." Soichi Yamane, Kazuhiro Kobayashi, Tomoki Toda, Tomoyasu Nakano, Masataka Goto, Satoshi Nakamura 0001 |
ICASSP | 6 |
| 2016 | Personalized unknown word detection in non-native language reading using eye gazeabstractThis paper proposes a method to detect unknown words during natural reading of non-native language text by using eye-tracking features. A previous approach utilizes gaze duration and word rarity features to perform this detection. However, while this system can be used by trained users, its performance is not sufficient during natural reading by untrained users. In this paper, we 1) apply support vector machines (SVM) with novel eye movement features that were not considered in the previous work and 2) examine the effect of personalization. The experimental results demonstrate that learning using SVMs and proposed eye movement features improves detection performance as measured by F-measure and that personalization further improves results. Rui Hiraoka, Hiroki Tanaka, Sakriani Sakti, Graham Neubig, Satoshi Nakamura 0001 |
ICMI | 5 |
| 2016 | Automatic detection of very early stage of dementia through multimodal interaction with computer avatarsabstractThis paper proposes a new approach to detecting very early stage of dementia automatically. We develop a computer avatar with spoken dialog functionalities that produces natural spoken queries referring to Mini Mental State Examination, Wechsler Memory Scale-Revised and other related questions. Multimodal interactive data of spoken dialogues from 18 participants (9 dementias and 9 healthy controls) are recorded, and audiovisual features are extracted. We confirm that the support vector machines can classify into two groups with 0.94 detection performance as measured by areas under ROC curve. It is found that our system has possibilities to detect very early stage of dementia through spoken dialog with our computer avatars. Hiroki Tanaka, Hiroyoshi Adachi, Norimichi Ukita, Takashi Kudo, Satoshi Nakamura 0001 |
ICMI | 5 |
| 2016 | Fast text anonymization using k-anonyminityabstractIn this paper, we propose a method for anonymizing unstructured texts using a quasi-identifier list. In our method, the system redacts from some parts of quasi-identifiers in the texts to the alternate characters such as "*", in order to prevent re-identification of information which should be kept in secrecy. However, this method has a room for an improvement for keeping the information on the original text as is. If the system anonymizes the texts and keeps the original texts as much as possible, the accuracy of the outputs by data mining techniques for the anonymized texts should be useful. Our method anonymizes quasi-identifiers to remain substrings which do not contribute to re-identification, in order to keep the information on the original texts as is. Wakana Maeda, Yu Suzuki 0001, Satoshi Nakamura 0001 |
iiWAS | 3 |
| 2016 | Gated Recurrent Neural Tensor NetworkabstractRecurrent Neural Networks (RNNs), which are a powerful scheme for modeling temporal and sequential data need to capture long-term dependencies on datasets and represent them in hidden layers with a powerful model to capture more information from inputs. For modeling long-term dependencies in a dataset, the gating mechanism concept can help RNNs remember and forget previous information. Representing the hidden layers of an RNN with more expressive operations (i.e., tensor products) helps it learn a more complex relationship between the current input and the previous hidden layer information. These ideas can generally improve RNN performances. In this paper, we proposed a novel RNN architecture that combine the concepts of gating mechanism and the tensor product into a single model. By combining these two concepts into a single RNN, our proposed models learn long-term dependencies by modeling with gating units and obtain more expressive and direct interaction between input and hidden layers using a tensor product on 3-dimensional array (tensor) weight parameters. We use Long Short Term Memory (LSTM) RNN and Gated Recurrent Unit (GRU) RNN and combine them with a tensor product inside their formulations. Our proposed RNNs, which are called a Long-Short Term Memory Recurrent Neural Tensor Network (LSTMRNTN) and Gated Recurrent Unit Recurrent Neural Tensor Network (GRURNTN), are made by combining the LSTM and GRU RNN models with the tensor product. We conducted experiments with our proposed models on word-level and character-level language modeling tasks and revealed that our proposed models significantly improved their performance compared to our baseline models. Andros Tjandra, Sakriani Sakti, Ruli Manurung, Mirna Adriani, Satoshi Nakamura 0001 |
IJCNN | 5 |
| 2016 | Transferring Emphasis in Speech Translation Using Hard-Attentional Neural Network Models
Quoc Truong Do, Sakriani Sakti, Graham Neubig, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2016 | A Hybrid System for Continuous Word-Level Emphasis Modeling Based on HMM State Clustering and Adaptive Training
Quoc Truong Do, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2016 | Supervised Learning of Acoustic Models in a Zero Resource Setting to Improve DPGMM Clustering
Michael Heck, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2016 | The NU-NAIST Voice Conversion System for the Voice Conversion Challenge 2016
Kazuhiro Kobayashi, Shinnosuke Takamichi, Satoshi Nakamura 0001, Tomoki Toda |
INTERSPEECH | 3 |
| 2016 | Acoustic-to-Articulatory Inversion Mapping Based on Latent Trajectory Gaussian Mixture Model
Patrick Lumban Tobing, Tomoki Toda, Hirokazu Kameoka, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2016 | Unsupervised Joint Estimation of Grapheme-to-Phoneme Conversion Systems and Acoustic Model Adaptation for Non-Native Speech Recognition
Satoshi Tsujioka, Sakriani Sakti, Koichiro Yoshino, Graham Neubig, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2016 | Unsupervised Phoneme Segmentation of Previously Unseen Languages
Marco Vetter, Markus Müller 0001, Fatima Hamlaoui, Graham Neubig, Satoshi Nakamura 0001, Sebastian Stüker, Alex Waibel |
INTERSPEECH | 5 |
| 2016 | Construction of Japanese Audio-Visual Emotion Database and Its Application in Emotion Recognition
Nurul Lubis, Randy Gomez, Sakriani Sakti, Keisuke Nakamura, Koichiro Yoshino, Satoshi Nakamura 0001, Kazuhiro Nakadai |
LREC | 6 |
| 2016 | Optimizing Computer-Assisted Transcription Quality with Iterative User Interfaces
Matthias Sperber, Graham Neubig, Satoshi Nakamura 0001, Alex Waibel |
LREC | 3 |
| 2016 | Selecting Syntactic, Non-redundant Segments in Active Learning for Machine TranslationabstractAkiva Miura, Graham Neubig, Michael Paul, Satoshi Nakamura. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Akiva Miura, Graham Neubig, Michael Paul, Satoshi Nakamura 0001 |
HLT-NAACL | 4 |
| 2016 | Cultural Communication Idiosyncrasies in Human-Computer InteractionabstractIn this work, we investigate whether the cultural idiosyncrasies found in humanhuman interaction may be transferred to human-computer interaction.With the aim of designing a culture-sensitive dialogue system, we designed a user study creating a dialogue in a domain that has the potential capacity to reveal cultural differences.The dialogue contains different options for the system output according to cultural differences.We conducted a survey among Germans and Japanese to investigate whether the supposed differences may be applied in human-computer interaction.Our results show that there are indeed differences, but not all results are consistent with the cultural models. Juliana Miehle, Koichiro Yoshino, Louisa Pragst, Stefan Ultes, Satoshi Nakamura 0001, Wolfgang Minker |
SIGDIAL Conference | 5 |
| 2016 | Analyzing the Effect of Entrainment on Dialogue ActsabstractEntrainment is a factor in dialogue that affects not only human-human but also human-machine interaction. While entrainment on the lexical level is well documented, less is known about how entrainment affects dialogue on a more abstract, structural level. In this paper, we investigate the effect of entrainment on dialogue acts and on lexical choice given dialogue acts, as well as how entrainment changes during a dialogue. We also define a novel measure of entrainment to measure these various types of entrainment. These results may serve as guidelines for dialogue systems that would like to entrain with users in a similar manner. Masahiro Mizukami, Koichiro Yoshino, Graham Neubig, David R. Traum, Satoshi Nakamura 0001 |
SIGDIAL Conference | 5 |
| 2016 | Iterative training of a DPGMM-HMM acoustic unit recognizer in a zero resource scenarioabstractIn this paper we propose a framework for building a full-fledged acoustic unit recognizer in a zero resource setting, i.e., without any provided labels. For that, we combine an iterative Dirichlet process Gaussian mixture model (DPGMM) clustering framework with a standard pipeline for supervised GMM-HMM acoustic model (AM) and n-gram language model (LM) training, enhanced by a scheme for iterative model re-training. We use the DPGMM to cluster feature vectors into a dynamically sized set of acoustic units. The frame based class labels serve as transcriptions of the audio data and are used as input to the AM and LM training pipeline. We show that iterative unsupervised model re-training of this DPGMM-HMM acoustic unit recognizer improves performance according to an ABX sound class discriminability task based evaluation. Our results show that the learned models generalize well and that sound class discriminability benefits from contextual information introduced by the language model. Our systems are competitive with supervisedly trained phone recognizers, and can beat the baseline set by DPGMM clustering. Michael Heck, Sakriani Sakti, Satoshi Nakamura 0001 |
SLT | 3 |
| 2016 | F0 transformation techniques for statistical voice conversion with direct waveform modification with spectral differentialabstractThis paper presents several F0transformation techniques for statistical voice conversion (VC) with direct waveform modification with spectral differential (DIFFVC). Statistical VC is a technique to convert speaker identity of a source speaker's voice into that of a target speaker by converting several acoustic features, such as spectral and excitation features. This technique usually uses vocoder to generate converted speech waveforms from the converted acoustic features. However, the use of vocoder often causes speech quality degradation of the converted voice owing to insufficient parameterization accuracy. To avoid this issue, we have proposed a direct waveform modification technique based on spectral differential filtering and have successfully applied it to intra-gender singing VC (DIFFSVC) where excitation features are not necessary converted. Moreover, we have also applied it to cross-gender singing VC by implementing F0transformation with a constant rate such as one octave increase or decrease. On the other hand, it is not straightforward to apply the DIFFSVC framework to normal speech conversion because the F0transformation ratio widely varies depending on a combination of the source and target speakers. In this paper, we propose several F0transformation techniques for DIFFVC and compare their performance in terms of speech quality of the converted voice and conversion accuracy of speaker individuality. The experimental results demonstrate that the F0transformation technique based on waveform modification achieves the best performance among the proposed techniques. Kazuhiro Kobayashi, Tomoki Toda, Satoshi Nakamura 0001 |
SLT | 3 |
| 2016 | Deep bottleneck features and sound-dependent i-vectors for simultaneous recognition of speech and environmental soundsabstractIn speech interfaces, it is often necessary to understand the overall auditory environment, not only recognizing what is being said, but also being aware of the location or actions surrounding the utterance. However, automatic speech recognition (ASR) becomes difficult when recognizing speech with environmental sounds. Standard solutions treat environmental sounds as noise, and remove them to improve ASR performance. On the other hand, most studies on environmental sounds construct classifiers for environmental sounds only, without interference of spoken utterances. But, in reality, such separate situations almost never exist. This study attempts to address the problem of simultaneous recognition of speech and environmental sounds. Particularly, we examine the possibility of using deep neural network (DNN) techniques to recognize speech and environmental sounds simultaneously, and improve the accuracy of both tasks under respective noisy conditions. First, we investigate DNN architectures including two parallel single-task DNNs, and a single multi-task DNN. However, we found direct multi-task learning of simultaneous speech and environmental recognition to be difficult. Therefore, we further propose a method that combines bottleneck features and sound-dependent i-vectors within this framework. Experimental evaluation results reveal that the utilizing bottleneck features and i-vectors as the input of DNNs can help to improve accuracy of each recognition task. Sakriani Sakti, Seiji Kawanishi, Graham Neubig, Koichiro Yoshino, Satoshi Nakamura 0001 |
SLT | 5 |
| 2016 | Learning local word reorderings for hierarchical phrase-based statistical machine translation
Jingyi Zhang 0001, Masao Utiyama, Eiichiro Sumita, Hai Zhao 0001, Graham Neubig, Satoshi Nakamura 0001 |
Mach. Transl. | 6 |
| 2016 | Learning cooperative persuasive dialogue policies using framing
Takuya Hiraoka, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
Speech Commun. | 5 |
| 2016 | Postfilters to Modify the Modulation Spectrum for Statistical Parametric Speech SynthesisabstractThis paper presents novel approaches based on modulation spectrum (MS) for high-quality statistical parametric speech synthesis, including text-to-speech (TTS) and voice conversion (VC). Although statistical parametric speech synthesis offers various advantages over concatenative speech synthesis, the synthetic speech quality is still not as good as that of concatenative speech synthesis or the quality of natural speech. One of the biggest issues causing the quality degradation is the over-smoothing effect often observed in the generated speech parameter trajectories. Global variance (GV) is known as a feature well correlated with the over-smoothing effect, and the effectiveness of keeping the GV of the generated speech parameter trajectories similar to those of natural speech has been confirmed. However, the quality gap between natural speech and synthetic speech is still large. In this paper, we propose using the MS of the generated speech parameter trajectories as a new feature to effectively quantify the over-smoothing effect. Moreover, we propose postfilters to modify the MS utterance by utterance or segment by segment to make the MS of synthetic speech close to that of natural speech. The proposed postfilters are applicable to various synthesizers based on statistical parametric speech synthesis. We first perform an evaluation of the proposed method in the framework of hidden Markov model (HMM)-based TTS, examining its properties from different perspectives. Furthermore, effectiveness of the proposed postfilters are also evaluated in Gaussian mixture model (GMM)-based VC and classification and regression trees (CART)-based TTS (a.k.a., CLUSTERGEN). The experimental results demonstrate that 1) the proposed utterance-level postfilter achieves quality comparable to the conventional generation algorithm considering the GV, and yields significant improvements by applying to the GV-based generation algorithm in HMM-based TTS, 2) the proposed segment-level postfilter capable of achieving low-delay synthesis also yields significant improvements in synthetic speech quality, and 3) the proposed postfilters are also effective in not only HMM-based TTS but also GMM-based VC and CLUSTERGEN. Shinnosuke Takamichi, Tomoki Toda, Alan W. Black, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2016 | Teaching Social Communication Skills Through Human-Agent InteractionabstractThere are a large number of computer-based systems that aim to train and improve social skills. However, most of these do not resemble the training regimens used by human instructors. In this article, we propose a computer-based training system that follows the procedure of social skills training (SST), a well-established method to decrease human anxiety and discomfort in social interaction, and acquire social skills. We attempt to automate the process of SST by developing a dialogue system named the automated social skills trainer , which teaches social communication skills through human-agent interaction. The system includes a virtual avatar that recognizes user speech and language information and gives feedback to users. Its design is based on conventional SST performed by human participants, including defining target skills, modeling, role-play, feedback, reinforcement, and homework. We performed a series of three experiments investigating (1) the advantages of using computer-based training systems compared to human-human interaction (HHI) by subjectively evaluating nervousness, ease of talking, and ability to talk well; (2) the relationship between speech language features and human social skills; and (3) the effect of computer-based training using our proposed system. Results of our first experiment show that interaction with an avatar decreases nervousness and increases the user's subjective impression of his or her ability to talk well compared to interaction with an unfamiliar person. The experimental evaluation measuring the relationship between social skill and speech and language features shows that these features have a relationship with social skills. Finally, experiments measuring the effect of performing SST with the proposed application show that participants significantly improve their skill, as assessed by separate evaluators, by using the system for 50 minutes. A user survey also shows that the users thought our system is useful and easy to use, and that interaction with the avatar felt similar to HHI. Hiroki Tanaka, Sakriani Sakti, Graham Neubig, Tomoki Toda, Hideki Negoro, Hidemi Iwasaka, Satoshi Nakamura 0001 |
ACM Trans. Interact. Intell. Syst. | 7 |
| 2015 | Syntax-based Simultaneous Translation through Prediction of Unseen Syntactic ConstituentsabstractYusuke Oda, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Yusuke Oda, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
ACL (1) | 5 |
| 2015 | The NAIST ASR system for the 2015 Multi-Genre Broadcast challenge: On combination of deep learning systems using a rank-score functionabstractThe Multi-Genre Broadcast challenge is an official challenge of the IEEE Automatic Speech Recognition and Understanding Workshop. This paper presents NAISTs contribution to the premiere of this challenge. The presented speech-to-text system for English makes use of various front-ends (e.g., MFCC, i-vector and FBANK), DNN acoustic models and several language models for decoding and rescoring (N-gram, RNNLM). Subsets of the training data with varying sizes were evaluated with respect to the overall training quality. Two speech segmentation systems were developed for the challenge, based on DNNs and GMM-HMMs. Recognition was performed in three stages: Decoding, lattice rescoring and system combination. This paper focuses on the system combination experiments and presents a rank-score based system weighting approach, which gave better performance compared to a normal system combination strategy. The DNN based ASR system trained on MFCC + i-vector features with the sMBR training criterion gives the best performance of 27.8% WER, and thus significantly outperforms the baseline DNN-HMM sMBR yielding 33.7% WER. Quoc Truong Do, Michael Heck, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura 0001 |
ASRU | 6 |
| 2015 | A study of social-affective communication: Automatic prediction of emotion triggers and responses in television talk showsabstractAdvancements in spoken language technologies have allowed users to interact with computers in an increasingly natural manner. However, most conversational agents or dialogue systems are yet to consider emotional awareness in interaction. To consider emotion in these situations, social-affective knowledge in conversational agents is essential. In this paper, we present a study of the social-affective process in natural conversation from television talk shows. We analyze occurrences of emotion (emotional responses), and the events that elicit them (emotional triggers). We then utilize our analysis for prediction to model the ability of a dialogue system to decide an action and response in an affective interaction. This knowledge has great potential to incorporate emotion into human-computer interaction. Experiments in two languages, English and Indonesian, show that automatic prediction performance surpasses random guessing accuracy. Nurul Lubis, Sakriani Sakti, Graham Neubig, Koichiro Yoshino, Tomoki Toda, Satoshi Nakamura 0001 |
ASRU | 6 |
| 2015 | Adaptive selection from multiple response candidates in example-based dialogueabstractIn spoken dialogue systems, dialogue modeling is one of the most important factors for contributing to user satisfaction improvement. Especially in Example-Based Dialogue Modeling (EBDM), effective methods to build dialogue example databases and to select response utterances from examples are the keys for improving dialogue quality. In dialogue corpora, it often have plural appropriate responses for one utterance. However, the system merges these plural appropriate responses into the one system response, thus, it does not try to use plural responses properly by user preference. In fact, responses that each user thinks to be preferable are different. In this paper, we propose a framework that select an appropriate response from plural appropriate response candidates satisfies users. It has a multi-response example database, and selects an appropriate response based on collaborative filtering. Experimental results showed that the proposed framework were successfully choosing appropriate responses, considering multi-response candidates improves user satisfaction to 4.1 from 3.7 of single response, and the adaptive response selection method increased user satisfaction from 3.7 to 4.3. Masahiro Mizukami, Hideaki Kizuki, Toshio Nomura, Graham Neubig, Koichiro Yoshino, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
ASRU | 8 |
| 2015 | Incremental sentence compression using LSTM recurrent networksabstractMany of the current sentence compression techniques attempt to produce a shortened form of a sentence by relying on syntactic structure such as dependency tree representations. While the performance of sentence compression has been improving, these approaches require a full parse of the sentence before performing sentence compression, making it difficult to perform compression in real time. In this paper, we examine the possibilities of performing incremental sentence compression using long short-term memory (LSTM) recurrent neural networks (RNN). The decision of whether to remove a word is done at each time step, without waiting for the end of the sentence. Various RNN parameters are investigated, including the number of layers and network connections. Furthermore, we also propose using a pretraining method in which the network is pretrained as an autoencoder. Experimental results reveal that our method obtains compression rates similar to human references and a better accuracy than the state-of-the-art tree transduction models. Sakriani Sakti, Faiz Ilham, Graham Neubig, Tomoki Toda, Ayu Purwarianti, Satoshi Nakamura 0001 |
ASRU | 6 |
| 2015 | Stochastic Gradient Variational Bayes for deep learning-based ASRabstractMany successful methods for training deep neural networks (DNN) rely on an unsupervised pretraining algorithm. It is particularly effective when the number of labeled training samples is not large enough, because pretraining method helps to initialize the parameter values in the appropriate range near a local good minimum, for further discriminative finetuning. However, while the improvement is impressive, training DNN is difficult because the objective function of DNN is highly non-convex function of the parameters. To avoid placing the parameter that generalizes poorly, a robust generative modelling is necessary. This paper explore an alternative of generative modelling for pretraining DNN-based acoustic modelling using Stochastic Gradient Variational Bayes (SGVB) within autoencoder framework called Variational Bayes Autoencoder (VBAE). It performs an efficient approximate inference and learning with directed probabilistic graphical models. During fine-tuning, probabilistic encoder parameters with latent variable components are then used in discriminative training for acoustic model. Here, we investigate the performances of DNN-based acoustic model using the proposed pretrained VBAE in comparison with widely used pretraining algorithms like Restricted Boltzmann Machine (RBM) and Stacked Denoising Autoencoder (SDAE). The results reveal that VBAE pretraining with Gaussian latent variables gave the best performance. Andros Tjandra, Sakriani Sakti, Satoshi Nakamura 0001, Mirna Adriani |
ASRU | 3 |
| 2015 | An Enhanced Electrolarynx with Automatic Fundamental Frequency Control based on Statistical PredictionabstractAn electrolarynx is a type of speaking aid device which is able to mechanically generate excitation sounds to help laryngectomees produce electrolaryngeal (EL) speech. Although EL speech is quite intelligible, its naturalness suffers from monotonous fundamental frequency patterns of the mechanical excitation sounds. To make it possible to generate more natural excitation sounds, we have proposed a method to automatically control the fundamental frequency of the sounds generated by the electrolarynx based on a statistical prediction model, which predicts the fundamental frequency patterns from the produced EL speech in real-time. In this paper, we develop a prototype system by implementing the proposed control method in an actual, physical electrolarynx and evaluate its performance. Kou Tanaka, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
ASSETS | 5 |
| 2015 | A Binarized Neural Network Joint Model for Machine TranslationabstractThe neural network joint model (NNJM), which augments the neural network language model (NNLM) with an m-word source context window, has achieved large gains in machine translation accuracy, but also has problems with high normalization cost when using large vocabularies.Training the NNJM with noise-contrastive estimation (NCE), instead of standard maximum likelihood estimation (MLE), can reduce computation cost.In this paper, we propose an alternative to NCE, the binarized NNJM (BNNJM), which learns a binary classifier that takes both the context and target words as input, and can be efficiently trained using MLE.We compare the BNNJM and NNJM trained by NCE on various translation tasks. Jingyi Zhang 0001, Masao Utiyama, Eiichiro Sumita, Graham Neubig, Satoshi Nakamura 0001 |
EMNLP | 5 |
| 2015 | WFST-based structural classification integrating dnn acoustic features and RNN language features for speech recognitionabstractThis paper proposes a method to train Weighted Finite State Transducer (WFST) based structural classifiers using deep neural network (DNN) acoustic features and recurrent neural network (RNN) language features for speech recognition. Structural classification is an effective approach to achieve highly accurate recognition of structured data in which the classifier is optimized to maximize the discriminative performance using different kinds of features. A WFST-based classifier, which can integrate acoustic, pronunciation, and language features embedded in a composed WFST, was recently extended to incorporate DNN bottleneck (DNNBN) features. In this paper, we further investigate the integration of a RNN language model (RNNLM) with the WFST classifier. To this end, we introduce a lattice rescoring method using a RNNLM for efficient classifier training. In a lecture transcription task, we reduced the word error rate from 19.2% to 18.6% by optimizing the WFST parameters for the DNNBN acoustic and RNNLM language features. Quoc Truong Do, Satoshi Nakamura 0001, Marc Delcroix, Takaaki Hori |
ICASSP | 2 |
| 2015 | EEG signal enhancement using multi-channel wiener filter with a spatial correlation priorabstractEvent-related potentials (ERPs) of electroencephalogram (EEG) are often used as features for brain machine interfaces or for analysis of brain activities. However, as EEG signals easily suffer from various artifacts, ERPs are often collapsed and hard to observe. There are several attempts at using multi-channel EEG signals to enhance EEG signals of interest and make ERPs more clearly observed. For example, a previous work has proposed a blind EEG signal separation method using a multi-channel Wiener filter designed with a probabilistic generative model of observed EEG signals. This method copes with the under-determination of EEG signal separation by assuming sparseness of each EEG component in the time-frequency domain. Although this method blindly separates EEG signals into individual EEG components using time-varying scaled spatial correlation matrices, target EEG components, such as P300 of ERP, are often known in advance in some applications. In this paper, inspired by this previous work, we propose a probabilistic EEG signal enhancement method using a multi-channel Wiener filter, newly incorporating prior information of the spatial correlation matrices related to the target EEG component in the probabilistic generative model to improve performance of EEG signal enhancement. An experimental evaluation for P300 enhancement shows that the proposed method significantly reduces artifacts. Hayato Maki, Tomoki Toda, Sakriani Sakti, Graham Neubig, Satoshi Nakamura 0001 |
ICASSP | 5 |
| 2015 | Statistical modeling of binaural signal and its application to binaural source separationabstractThis paper addresses a new statistical model of binaural signals and its application to efficient binaural source separation. Binaural source separation is always required to retain a spatial cue of the separated sound, such as a head-related transfer function (HRTF). However, the direct use of an HRTF is not realistic because this information is normally not known in advance. To cope with this problem, first, we focus on the difference between signal probability density functions at both ears, which can be blindly estimated by using our previous work on higher-order statistics. Next, we derive a sound-localization-preserved generalized minimum mean-square error short-time spectral amplitude estimator. Objective and subjective experiments show the efficacy of the proposed method in terms of spatial quality. Yuki Murota, Daichi Kitamura, Shoichi Koyama, Hiroshi Saruwatari, Satoshi Nakamura 0001 |
ICASSP | 5 |
| 2015 | Parameter generation algorithm considering Modulation Spectrum for HMM-based speech synthesisabstractThis paper proposes a novel parameter generation algorithm for high-quality speech generation in Hidden Markov Model (HMM)-based speech synthesis. One of the biggest issues causing significant quality degradation is the over-smoothing effect often observed in generated parameter trajectories. Global Variance (GV) is known as a feature well correlated with the over-smoothing effect and a metric on the GV of the generated parameters is effectively used as a penalty term in the conventional parameter generation. However, the quality of the synthetic speech is far from that of the natural speech. Recently, we have found that a Modulation Spectrum (MS) of the generated parameters, which is also regarded as an extension of the GV, is more sensitively correlated with the over-smoothing effect than the GV. This paper incorporates a metric on the MS as a new penalty term in the proposed parameter generation algorithm. The experimental results demonstrate that the proposed parameter generation algorithm considering the MS yields significant improvements in synthetic speech quality compared to the conventional parameter generation algorithm considering the GV. Shinnosuke Takamichi, Tomoki Toda, Alan W. Black, Satoshi Nakamura 0001 |
ICASSP | 4 |
| 2015 | Modulation spectrum-constrained trajectory training algorithm for GMM-based Voice ConversionabstractThis paper presents a novel training algorithm for Gaussian Mixture Model (GMM)-based Voice Conversion (VC). One of the advantages of GMM-based VC is computationally efficient conversion processing enabling to achieve real-time VC applications. On the other hand, the quality of the converted speech is still significantly worse than that of natural speech. In order to address this problem while preserving the computationally efficient conversion processing, the proposed training method enables 1) to use a consistent optimization criterion between training and conversion and 2) to compensate a Modulation Spectrum (MS) of the converted parameter trajectory as a feature sensitively correlated with over-smoothing effects causing quality degradation of the converted speech. The experimental results demonstrate that the proposed algorithm yields significant improvements in term of both the converted speech quality and the conversion accuracy for speaker individuality compared to the basic training algorithm. Shinnosuke Takamichi, Tomoki Toda, Alan W. Black, Satoshi Nakamura 0001 |
ICASSP | 4 |
| 2015 | Combination of two-dimensional cochleogram and spectrogram features for deep learning-based ASRabstractThis paper explores the use of auditory features based on cochleograms; two dimensional speech features derived from gammatone filters within the convolutional neural network (CNN) framework. Furthermore, we also propose various possibilities to combine cochleogram features with log-mel filter banks or spectrogram features. In particular, we combine within low and high levels of CNN framework which we refer to as low-level and high-level feature combination. As comparison, we also construct the similar configuration with deep neural network (DNN). Performance was evaluated in the framework of hybrid neural network - hidden Markov model (NN-HMM) system on TIMIT phoneme sequence recognition task. The results reveal that cochleogram-spectrogram feature combination provides significant advantages. The best accuracy was obtained by high-level combination of two dimensional cochleogram-spectrogram features using CNN, achieved up to 8.2% relative phoneme error rate (PER) reduction from CNN single features or 19.7% relative PER reduction from DNN single features. Andros Tjandra, Sakriani Sakti, Graham Neubig, Tomoki Toda, Mirna Adriani, Satoshi Nakamura 0001 |
ICASSP | 6 |
| 2015 | Preserving word-level emphasis in speech-to-speech translation using linear regression HSMMsabstractIn speech, emphasis is an important type of paralinguistic information that helps convey the focus of an utterance, new information, and emotion. If emphasis can be incorporated into a speech-to-speech (S2S) translation system, it will be possible to convey this information across the language barrier. However, previous related work focuses only on the translation of particular prosodic features, such as F0, or works with emphasis but focuses on extremely small vocabularies, such as the 10 digits. In this paper, we describe a new S2S method that is able to translate the emphasis across languages and consider multiple features of emphasis such as power, F0, and duration over larger vocabularies. We do so by introducing two new components: word-level emphasis estimation using linear regression hidden semi-Markov models, and emphasis translation that translates the word-level emphasis to the target language with conditional random fields. The text-to-speech synthesis system is also modified to be able to synthesize emphasized speech. The result shows that our system can translate the emphasis correctly with 91.6% F -measure for objective test, and 87.8% for subjective test. Quoc Truong Do, Shinnosuke Takamichi, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura 0001 |
INTERSPEECH | 6 |
| 2015 | Statistical singing voice conversion based on direct waveform modification with global varianceabstractThis paper presents techniques to improve the quality of voices generated through statistical singing voice conversion with direct waveform modification based on spectrum differential (DIFFSVC). The DIFFSVC method makes it possible to convert singing voice characteristics of a source singer into those of a target singer without using vocoder-based waveform generation. However, quality of the converted singing voice still degrades compared to that of a natural singing voice due to various factors, such as the over-smoothing of the converted spectral parameter trajectory. To alleviate this over-smoothing, we propose a technique to restore the global variance of the converted spectral parameter trajectory within the framework of the DIFFSVC method. We also propose another technique to specifically avoid over-smoothing at unvoiced frames. Results of subjective and objective evaluations demonstrate that the proposed techniques significantly improve speech quality of the converted singing voice while preserving the conversion accuracy of singer identity compared to the conventional DIFFSVC. Kazuhiro Kobayashi, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2015 | Speed or accuracy? a study in evaluation of simultaneous speech translationabstractSimultaneous speech translation is a technology that attempts to reduce the delay inherent in speech translation by beginning translation before the end of explicit sentence boundaries. Despite best efforts, there is still often a trade-off between speed and accuracy in these systems, with systems with less delay also achieving lower accuracy. However, somewhat surprisingly, there is no previous work examining the relative importance of speed and accuracy, and thus given two systems with various speeds and accuracies, it is difficult to say with certainty which is better. In this paper, we make the first steps towards evaluation of simultaneous speech translation systems in consideration of both speed and accuracy. We collect user evaluations of speech translation results with different levels of accuracy and delay, and using this data to learn the parameters of an evaluation measure that can judge the trade-off between these two factors. Based on these results, we find that considering both accuracy and delay in the evaluation of speech translation results helps improve correlations with human judgements, and that users placed higher relative importance on reducing delay when results were presented through text, rather than speech. Takashi Mieno, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2015 | A latent variable model for joint pause prediction and dependency parsingabstractThe prosody of speech is closely related to syntactic structure of the spoken sentence, and thus analysis models that jointly consider these two types of information are promising. However, manual annotation of syntactic information and prosodic information such as pauses is laborious, and thus it can be difficult to obtain sufficient data to train such joint models. In this paper, we tackle this problem by introducing a joint pause prediction and dependency parsing model that treats pauses between consecutive words as latent variables. Using this model, it is possible to learn from not only data labeled with both syntax and pause information, but also data labeled with only syntactic information, which can be obtained in larger quantities. Experiments find that a joint pause prediction and dependency parsing model obtains better pause prediction F-measure than a decision-tree-based baseline trained on the same data, and that the addition of more data using the proposed latent variable model leads for further gains of up to 11.6 points in F-measure. The Tung Nguyen, Graham Neubig, Hiroyuki Shindo, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
INTERSPEECH | 6 |
| 2015 | Non-native speech synthesis preserving speaker individuality based on partial correction of prosodic and phonetic characteristicsabstractThis paper presents a novel non-native speech synthesis technique that preserves the individuality of a non-native speaker. Cross-lingual speech synthesis based on voice conversion or HMM-based speech synthesis, which synthesizes foreign language speech of a specific non-native speaker reflecting the speaker-dependent acoustic characteristics extracted from the speaker’s natural speech in his/her mother tongue, tends to cause a degradation of speaker individuality in synthetic speech compared to intra-lingual speech synthesis. This paper proposes a new approach to cross-lingual speech synthesis that preserves speaker individuality by explicitly using non-native speech spoken by the target speaker. Although the use of nonnative speech makes it possible to preserve the speaker individuality in the synthesized target speech, naturalness is significantly degraded as the speech is directly affected by unnatural prosody and pronunciation often caused by differences in the linguistic systems of the source and target languages. To improve naturalness while preserving speaker individuality, we propose (1) a prosodic correction method based on model adaptation, and (2) a phonetic correction method based on spectrum replacement for unvoiced consonants. The experimental results demonstrate that these proposed methods are capable of significantly improving naturalness while preserving the speaker individuality in synthetic speech. Yuji Oshima, Shinnosuke Takamichi, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 6 |
| 2015 | Non-audible murmur enhancement based on statistical conversion using air- and body-conductive microphones in noisy environmentsabstractNon-Audible Murmur (NAM) is an extremely soft whispered voice detected by a special body-conductive microphone called a NAM microphone. Although NAM is a promising medium for silent speech communication, its quality is significantly degraded by its faint volume and spectral changes caused by body-conductive recording. To improve the quality of NAM, several enhancement methods based on statistical voice conversion (VC) techniques have been proposed, and their effectiveness has been confirmed in quiet environments. However, it can be expected that NAM will be used not only in quiet, but also in noisy environments, and it is thus necessary to develop enhancement methods that will also work in these cases. In this paper, we propose a framework for NAM enhancement using not only the NAM microphone but also an air-conductive microphone. Airand body-conducted NAM signals are used as the input of VC to estimate a more naturally sounding speech signal. To clarify adverse effects of external noises on the performance of the proposed framework and investigate a possibility to alleviate them by revising VC models, we also implement noise-dependent VC models within the proposed framework. Experimental results demonstrate that the proposed framework yields significant improvements in the spectral conversion accuracy and listenability of enhanced speech under both quiet and noisy environments. Yusuke Tajiri, Kou Tanaka, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 6 |
| 2015 | Modulation spectrum-constrained trajectory training algorithm for HMM-based speech synthesis
Shinnosuke Takamichi, Tomoki Toda, Alan W. Black, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2015 | Articulatory controllable speech modification based on Gaussian mixture models with direct waveform modification using spectrum differentialabstractIn our previous work, we have developed a speech modification system capable of manipulating unobserved articulatory movements by sequentially performing speech-to-articulatory inversion mapping and articulatory-to-speech production mapping based on a Gaussian mixture model (GMM)-based statistical feature mapping technique. One of the biggest issues to be addressed in this system is quality degradation of the synthetic speech caused by modeling and conversion errors in a vocoderbased waveform generation framework. To address this issue, we propose several implementation methods of direct waveform modification. The proposed methods directly filter an input speech waveform with a time sequence of spectral differential parameters calculated between unmodified and modified spectral envelop parameters in order to avoid using vocoderbased excitation signal generation. The experimental results show that the proposed direct waveform modification methods yield significantly larger quality improvements in the synthetic speech while also keeping a capability of intuitively modifying phoneme sounds by manipulating the unobserved articulatory movements. Patrick Lumban Tobing, Kazuhiro Kobayashi, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 6 |
| 2015 | Automated Social Skills TrainerabstractSocial skills training is a well-established method to decrease human anxiety and discomfort in social interaction, and acquire social skills. In this paper, we attempt to automate the process of social skills training by developing a dialogue system named "automated social skills trainer," which provides social skills training through human-computer interaction. The system includes a virtual avatar that recognizes user speech and language information and gives feedback to users to improve their social skills. Its design is based on conventional social skills training performed by human participants, including defining target skills, modeling, role-play, feedback, reinforcement, and homework. An experimental evaluation measuring the relationship between social skill and speech and language features shows that these features have a relationship with autistic traits. Additional experiments measuring the effect of performing social skills training with the proposed application show that most participants improve their skill by using the system for 50 minutes. Hiroki Tanaka, Sakriani Sakti, Graham Neubig, Tomoki Toda, Hideki Negoro, Hidemi Iwasaka, Satoshi Nakamura 0001 |
IUI | 7 |
| 2015 | Pseudogen: A Tool to Automatically Generate Pseudo-Code from Source CodeabstractUnderstanding the behavior of source code written in an unfamiliar programming language is difficult. One way to aid understanding of difficult code is to add corresponding pseudo-code, which describes in detail the workings of the code in a natural language such as English. In spite of its usefulness, most source code does not have corresponding pseudo-code because it is tedious to create. This paper demonstrates a tool Pseudogen that makes it possible to automatically generate pseudo-code from source code using statistical machine translation (SMT). Pseudogen currently supports generation of English or Japanese pseudo-code from Python source code, and the SMT framework makes it easy for users to create new generators for their preferred source code/pseudo-code pairs. Hiroyuki Fudaba, Yusuke Oda, Koichi Akabe, Graham Neubig, Hideaki Hata, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
ASE | 8 |
| 2015 | Learning to Generate Pseudo-Code from Source Code Using Statistical Machine Translation (T)abstractPseudo-code written in natural language can aid the comprehension of source code in unfamiliar programming languages. However, the great majority of source code has no corresponding pseudo-code, because pseudo-code is redundant and laborious to create. If pseudo-code could be generated automatically and instantly from given source code, we could allow for on-demand production of pseudo-code without human effort. In this paper, we propose a method to automatically generate pseudo-code from source code, specifically adopting the statistical machine translation (SMT) framework. SMT, which was originally designed to translate between two natural languages, allows us to automatically learn the relationship between source code/pseudo-code pairs, making it possible to create a pseudo-code generator with less human effort. In experiments, we generated English or Japanese pseudo-code from Python statements using SMT, and find that the generated pseudo-code is largely accurate, and aids code understanding. Yusuke Oda, Hiroyuki Fudaba, Graham Neubig, Hideaki Hata, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
ASE | 7 |
| 2015 | Ckylark: A More Robust PCFG-LA ParserabstractYusuke Oda, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations. 2015. Yusuke Oda, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
HLT-NAACL | 5 |
| 2015 | Reinforcement Learning in Multi-Party Trading DialogabstractIn this paper, we apply reinforcement learning (RL) to a multi-party trading scenario where the dialog system (learner) trades with one, two, or three other agents.We experiment with different RL algorithms and reward functions.The negotiation strategy of the learner is learned through simulated dialog with trader simulators.In our experiments, we evaluate how the performance of the learner varies depending on the RL algorithm used and the number of traders.Our results show that (1) even in simple multi-party trading dialog tasks, learning an effective negotiation policy is a very hard problem; and (2) the use of neural fitted Q iteration combined with an incremental reward function produces negotiation policies as effective or even better than the policies of two strong hand-crafted baselines. Takuya Hiraoka, Kallirroi Georgila, Elnaz Nouri, David R. Traum, Satoshi Nakamura 0001 |
SIGDIAL Conference | 5 |
| 2015 | Semantic Parsing of Ambiguous Input through Paraphrasing and VerificationabstractWe propose a new method for semantic parsing of ambiguous and ungrammatical input, such as search queries. We do so by building on an existing semantic parsing framework that uses synchronous context free grammars (SCFG) to jointly model the input sentence and output meaning representation. We generalize this SCFG framework to allow not one, but multiple outputs. Using this formalism, we construct a grammar that takes an ambiguous input string and jointly maps it into both a meaning representation and a natural language paraphrase that is less ambiguous than the original input. This paraphrase can be used to disambiguate the meaning representation via verification using a language model that calculates the probability of each paraphrase. Philip Arthur, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
Trans. Assoc. Comput. Linguistics | 5 |
| 2015 | Multichannel Signal Separation Combining Directional Clustering and Nonnegative Matrix Factorization with Spectrogram RestorationabstractIn this paper, to address problems in multichannel music signal separation, we propose a new hybrid method that combines directional clustering and advanced nonnegative matrix factorization (NMF). The aims of multichannel music signal separation technology is to extract a specific target signal from observed multichannel signals that contain multiple instrumental sounds. In previous studies, various methods using NMF have been proposed, but many problems remain including poor separation accuracy and lack of robustness. To solve these problems, we propose a new supervised NMF (SNMF) with spectrogram restoration and a hybrid method that concatenates the proposed SNMF after directional clustering. Via the extrapolation of supervised spectral bases, the proposed SNMF attempts both target signal separation and reconstruction of the lost target components, which are generated by preceding directional clustering. In addition, we experimentally reveal the trade-off between separation and extrapolation abilities and propose a new scheme for adaptive divergence, where the optimal divergence can be automatically changed in each time frame according to the local spatial conditions. The results of an evaluation experiment show that our proposed hybrid method outperforms the conventional music signal separation methods. Daichi Kitamura, Hiroshi Saruwatari, Hirokazu Kameoka, Yu Takahashi, Kazunobu Kondo, Satoshi Nakamura 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2014 | Discriminative Language Models as a Tool for Machine Translation Error Analysis
Koichi Akabe, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
COLING | 5 |
| 2014 | Reinforcement Learning of Cooperative Persuasive Dialogue Policies using Framing
Takuya Hiraoka, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
COLING | 5 |
| 2014 | Acquiring a Dictionary of Emotion-Provoking EventsabstractHoa Trong Vu, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura. Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics, volume 2: Short Papers. 2014. Hoa Trong Vu, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
EACL | 5 |
| 2014 | Regression approaches to perceptual age control in singing voice conversionabstractThe perceptual age of a singing voice is the age of the singer as perceived by the listener, and is one of the notable characteristics that determines perceptions of a song. In this paper, we describe a novel voice timbre control technique based on the perceptual age for singing voice conversion (SVC). Singers can sing expressively by controlling prosody and voice timbre, but the varieties of voices that singers can produce are limited by physical constraints. Previous work has attempted to overcome the limitation through the use of statistical voice conversion. This technique makes it possible to convert singing voice timbre of an arbitrary source singer into that of an arbitrary target singer. However, it is still difficult to intuitively control singing voice characteristics by manipulating parameters corresponding to specific physical traits, such as gender and age. In this paper, we develop a technique for controlling the voice timbre based on perceptual age that maintains the singer's individuality. The experimental results show that the proposed voice timbre control method makes it possible to change the singer's perceptual age while not having an adverse effect on the perceived individuality. Kazuhiro Kobayashi, Tomoki Toda, Tomoyasu Nakano, Masataka Goto, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
ICASSP | 7 |
| 2014 | Narrow Adaptive Regularization of weights for grapheme-to-phoneme conversionabstractAs the speech recognition field proceeds to open domain and multilingual tasks, the need for robust g2p conversion has been increasing. Towards this objective, we propose a new g2p conversion training method based on the Narrow Adaptive Regularization of Weights (NAROW) online learning algorithm. NAROW improves over its predecessor AROW by automatically adjusting hyperparameters to reduce mistake bounds, and ensuring that the learning rate is not updated when features for the input data have already been updated enough. The contribution of this paper is first to extend NAROW to structured learning, and show the inequality to bound the maximum number of errors in structured NAROW. In experiments, our proposed approach significantly improved over MIRA with consistent phoneme error rate reductions of 1.3-3.8% on a variety of dictionaries. Keigo Kubo, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura 0001 |
ICASSP | 5 |
| 2014 | Music signal separation based on Bayesian spectral amplitude estimator with automatic target prior adaptationabstractIn this paper, we propose a new approach for addressing music signal separation based on the generalized Bayesian estimator with automatic prior adaptation. This method consists of three parts, namely, the generalized MMSE-STSA estimator with a flexible target signal prior, the NMF-based dynamic interference spectrogram estimator, and closed-form parameter estimation for the statistical model of the target signal based on higher-order statistics. The statistical model parameter of the hidden target signal can be detected automatically for optimal Bayesian estimation with online target-signal prior adaptation. Our experimental evaluation can show the efficacy of the proposed method. Yuki Murota, Daichi Kitamura, Shunsuke Nakai, Hiroshi Saruwatari, Satoshi Nakamura 0001, Yu Takahashi, Kazunobu Kondo |
ICASSP | 5 |
| 2014 | A postfilter to modify the modulation spectrum in HMM-based speech synthesisabstractIn this paper, we propose a postfilter to compensate modulation spectrum in HMM-based speech synthesis. In order to alleviate over-smoothing effects which is a main cause of quality degradation in HMM-based speech synthesis, it is necessary to consider features that can capture over-smoothing. Global Variance (GV) is one well-known example of such a feature, and the effectiveness of parameter generation algorithm considering GV have been confirmed. However, the quality gap between natural speech and synthetic speech is still large. In this paper, we introduce the Modulation Spectrum (MS) of speech parameter trajectory as a new feature to effectively capture the over-smoothing effect, and we propose a postfilter based on the MS. The MS is represented as a power spectrum of the parameter trajectory. The generated speech parameter sequence is filtered to ensure that its MS has a pattern similar to natural speech. Experimental results show quality improvements when the proposed methods are applied to spectral and F0components, compared with conventional methods considering GV. Shinnosuke Takamichi, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
ICASSP | 5 |
| 2014 | An evaluation of excitation feature prediction in a hybrid approach to electrolaryngeal speech enhancementabstractWe implement removing micro-prosody with low-pass filtering and avoiding Unvoiced/Voiced (U/V) prediction as part of a hybrid approach to improve statistical excitation prediction in electrolaryngeal (EL) speech enhancement. An electrolarynx is a device that artificially generates excitation sounds to enable laryngectomees to produce EL speech. Although proficient laryngectomees can produce quite intelligible EL speech, it sounds very unnatural due to the mechanical excitation produced by the device. Moreover, the excitation sounds produced by the device often leak outside, adding noise to EL speech. To address these issues, in our previous work, we proposed a hybrid method using a noise reduction method for enhancing spectral parameters and voice conversion method for predicting excitation parameters. In this paper, we evaluate the effect of removing micro-prosody with low-pass filtering and avoiding U/V prediction in the hybrid enhancement process. Kou Tanaka, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
ICASSP | 5 |
| 2014 | A hearing impairment simulation method using audiogram-based approximation of auditory charatecteristicsabstractHearing impairment simulation is an effective technique to educate normal-hearing people about auditory perception of the hearing-impaired. Because auditory characteristics of the hearing impaired vary greatly from person-to-person, personalization of the hearing impairment simulation systems is essential to accurately simulate these individual differences. However, measurement of auditory characteristics of individuals is time-consuming work. In this paper, we propose a hearing impairment simulation method that is easily applied to individual hearing-impaired persons. Auditory filter characteristics and gain characteristics are estimated from easily measurable audiograms of each individual. We also implement a method for manually adjusting the hearing impairment level to improve accuracy of the proposed hearing impairment simulation. An experimental evaluation is conducted to compare intelligibility between hearing-impaired and normal-hearing persons with the proposed hearing impairment simulation. The experimental results show that the proposed method effectively makes the word correct rate and phoneme confusion tendency of the normal hearing persons similar to those of the hearing impaired persons. Index Terms: hearing-impairment simulation, personalization, auditory filter characteristics, gain characteristics, audiogram, Nozomi Jinbo, Shinnosuke Takamichi, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 6 |
| 2014 | Statistical singing voice conversion with direct waveform modification based on the spectrum differentialabstractThis paper presents a novel statistical singing voice conversion (SVC) technique with direct waveform modification based on the spectrum differential that can convert voice timbre of a source singer into that of a target singer without using a vocoder to generate converted singing voice waveforms. SVC makes it possible to convert singing voice characteristics of an arbitrary source singer into those of an arbitrary target singer. However, speech quality of the converted singing voice is significantly degraded compared to that of a natural singing voice due to various factors, such as analysis and modeling errors in the vocoderbased framework. To alleviate this degradation, we propose a statistical conversion process that directly modifies the signal in the waveform domain by estimating the difference in the spectra of the source and target singers’ singing voices. The differential spectral feature is directly estimated using a differential Gaussian mixture model (GMM) that is analytically derived from the traditional GMM used as a conversion model in the conventional SVC. The experimental results demonstrate that the proposed method makes it possible to significantly improve speech quality in the converted singing voice while preserving the conversion accuracy of singer identity compared to the conventional SVC. Index Terms: singing voice, statistical voice conversion, vocoder, Gaussian mixture model, differential spectral compensation Kazuhiro Kobayashi, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2014 | Structured soft margin confidence weighted learning for grapheme-to-phoneme conversionabstractIn recent years, structured online discriminative learning methods using second order statistics have been shown to outperform conventional generative and discriminative models in the grapheme-to-phoneme (g2p) conversion task. However, these methods update the parameters by sequentially using N-best hypotheses predicted with the current parameters. Thus, the parameters appearing in early hypotheses are overfitted compared with those in later hypotheses. In this paper, we propose a novel method called structured soft margin confidence weighted learning, which extends multi-class confidence weighted learning to structured learning. The proposed method extends multiclass CW in two ways, allowing for improved robustness to overfitting: (1) regularization inspired by soft margin support vector machines, allowing for margin error, and (2) update using N-best hypotheses simultaneously and interdependently. In an evaluation experiment on the g2p conversion task, the proposed method improved over all other approaches in terms of phoneme error rate with a significant difference. Index Terms: g2p conversion, out-of-vocabulary word, online discriminative training, structured learning, confidence weighted algorithm Keigo Kubo, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2014 | Data-driven generation of text balloons based on linguistic and acoustic features of a comics-anime corpusabstractMost automatic speech recognition systems existing today are still limited to recognizing what is being said, without being concerned with how it is being said. On the other hand, research on emotion recognition from speech has recently gained considerable interest, but how those emotions could be expressed in text-based communication has not been widely investigated. Our long-term goal is to construct expressive speech-to-text systems that conveys all information from acoustic speech, including verbal message, emotional state, speaker condition, and background noise, into unified text-based communication. In this preliminary study, we start with developing a system that can convey emotional speech into text-based communication by way of text balloons. As there exist many possible ways to generate the text balloons, we propose to utilize linguistic and acoustic features based on comic books and anime films. Experimental results reveal that expressive text is more preferable than static text, and the system is able to estimate the shape of text balloons with 87.01% accuracy. Index Terms: data-driven approaches, expressive text generation, linguistic and acoustic features Sho Matsumiya, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2014 | Direct F0 control of an electrolarynx based on statistical excitation feature prediction and its evaluation through simulationabstractAn electrolarynx is a device that artificially generates excitation sounds to enable laryngectomees to produce electrolaryngeal (EL) speech. Although proficient laryngectomees can produce quite intelligible EL speech, it sounds very unnatural due to the mechanical excitation produced by the device. To address this issue, we have proposed several EL speech enhancement methods using statistical voice conversion and showed that statistical prediction of excitation parameters, such as F0 patterns, was essential to significantly improve naturalness of EL speech. In these methods, the original EL speech is recorded with a microphone and the enhanced EL speech is presented from a loudspeaker in real time. This framework is effective for telecommunication but it is not suitable to face-to-face conversation because both the original EL speech and the enhanced EL speech are presented to listeners. In this paper, we propose direct F0 control of the electrolarynx based on statistical excitation prediction to develop an EL speech enhancement technique also effective for face-to-face conversation. F0 patterns of excitation signals produced by the electrolarynx are predicted in real time from the EL speech produced by the laryngectomee’s articulation of the excitation signals with previously predicted F0 values. A simulation experiment is conducted to evaluate the effectiveness of the proposed method. The experimental results demonstrate that the proposed method yields significant improvements in naturalness of EL speech while keeping its intelligibility high enough. Index Terms: laryngectomee, electrolarynx, electrolaryngeal speech, statistical excitation prediction, simulation evaluation Kou Tanaka, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2014 | Articulatory controllable speech modification based on statistical feature mapping with Gaussian mixture modelsabstractThis paper presents a novel speech modification method capable of controlling unobservable articulatory parameters based on a statistical feature mapping technique with Gaussian Mixture Models (GMMs). In previous work [1], the GMM-based statistical feature mapping was successfully applied to acousticto-articulatory inversion mapping and articulatory-to-acoustic production mapping separately. In this paper, these two mapping frameworks are integrated to a unified framework to develop a novel speech modification system. The proposed system sequentially performs the inversion and the production mapping, making it possible to modify phonemic sounds of an input speech signal by intuitively manipulating articulatory parameters estimated from the input speech signal. We also propose a manipulation method to automatically compensate for unmodified articulatory movements considering inter-dimensional correlation of the articulatory parameters. The proposed system is implemented for a single English speaker and its effectiveness is evaluated experimentally. The experimental results demonstrate that the proposed system is capable of modifying phonemic sounds by manipulating the estimated articulatory movements and higher speech quality is achieved by considering the inter-dimensional correlation in the manipulation. Index Terms: speech modification, acoustic-to-articulatory inversion mapping, articulatory-to-acoustic production mapping, Gaussian mixture model, inter-dimensional correlation Patrick Lumban Tobing, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001, Ayu Purwarianti |
INTERSPEECH | 5 |
| 2014 | Towards Multilingual Conversations in the Medical Domain: Development of Multilingual Medical Data and A Network-based ASR System
Sakriani Sakti, Keigo Kubo, Sho Matsumiya, Graham Neubig, Tomoki Toda, Satoshi Nakamura 0001, Fumihiro Adachi, Ryosuke Isotani |
LREC | 6 |
| 2014 | Collection of a Simultaneous Translation Corpus for Comparative Analysis
Hiroaki Shimizu, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
LREC | 5 |
| 2014 | Emotion recognition on Indonesian television talk showsabstractAs interaction between human and computer continues to develop to the most natural form possible, it becomes more and more urgent to incorporate emotion in the equation. The field continues to develop, yet exploration of the subject in Indonesian is still very lacking. This paper presents the first study of emotion recognition in Indonesian, including the construction of the first emotionally colored speech corpus in the language, and the building of an emotion classifier through an optimized machine learning process. We construct our corpus using television talk show recordings in various topics of discussion, yielding colorful emotional utterances. In our machine learning experiment, we employ the support vector machine (SVM) algorithm with feature selection and parameter optimization to ensure the best resulting model possible. Evaluation of the experiment result shows recognition accuracy of 68.31% at best. Nurul Lubis, Dessi Puji Lestari, Ayu Purwarianti, Sakriani Sakti, Satoshi Nakamura 0001 |
SLT | 5 |
| 2014 | Improving the robustness of example-based dialog retrieval using recursive neural network paraphrase identificationabstractPrevious work on example-based chat-oriented dialog systems utilizing real human-to-human conversation has shown promising results. However, most previous methods use relatively simple retrieval techniques, resulting in weakness to out of vocabulary (OOV) database queries and inadequate handling of interactions between words in the sentence. To overcome this problem, in this paper we propose a method to utilize recursive neural network paraphrase identification to improve the accuracy and robustness of example-based dialog response retrieval. We model our dialog-pair database and user input query with distributed word representations, and employ recursive autoencoders and dynamic pooling to determine whether two sentences with arbitrary length have the same meaning. The distributed representations have the potential to improve handling of OOV cases, and the recursive structure can reduce confusion in example matching. We evaluate the system performance based on objective and subjective metrics. Lasguido Nio, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura 0001 |
SLT | 5 |
| 2014 | On-the-fly user modeling for cost-sensitive correction of speech transcriptsabstractWe propose an on-the-fly updating framework for cost-sensitive manual correction of automatically recognized speech transcripts. This framework trains cost-models during the transcription process, and does not require the transcriber enrollment necessary in previous work. We use a baseline method that optimizes a segmentation into segments to supervise or not to supervise in a cost-sensitive fashion that minimizes human effort, and introduce a much faster algorithm for computing such a segmentation that can be used for on-the-fly updates. Besides removing the need to carry out enrollments, experiments show that our updating framework results in 28% higher human supervision efficiency than previous cost-sensitive approaches. Matthias Sperber, Graham Neubig, Satoshi Nakamura 0001, Alex Waibel |
SLT | 3 |
| 2014 | Musical-noise-free blind speech extraction integrating microphone array and iterative spectral subtractionabstractIn this paper, we propose a musical-noise-free blind speech extraction method using a microphone array for application to nonstationary noise. In our previous study, it was found that optimized iterative spectral subtraction (SS) results in speech enhancement with almost no musical noise generation, but this method is valid only for stationary noise. The proposed method consists of iterative blind dynamic noise estimation by, e.g., independent component analysis (ICA) or multichannel Wiener filtering, and musical-noise-free speech extraction by modified iterative SS, where multiple iterative SS is applied to each channel while maintaining the multichannel property reused for the dynamic noise estimators. Also, in relation to the proposed method, we discuss the justification of applying ICA to signals nonlinearly distorted by SS. From objective and subjective evaluations simulating a real-world hands-free speech communication system, we reveal that the proposed method outperforms the conventional methods. Ryoichi Miyazaki, Hiroshi Saruwatari, Satoshi Nakamura 0001, Kiyohiro Shikano, Kazunobu Kondo, Jonathan Blanchette, Martin Bouchard 0001 |
Signal Process. | 3 |
| 2014 | Segmentation for Efficient Supervised Language Annotation with an Explicit Cost-Utility TradeoffabstractIn this paper, we study the problem of manually correcting automatic annotations of natural language in as efficient a manner as possible. We introduce a method for automatically segmenting a corpus into chunks such that many uncertain labels are grouped into the same chunk, while human supervision can be omitted altogether for other segments. A tradeoff must be found for segment sizes. Choosing short segments allows us to reduce the number of highly confident labels that are supervised by the annotator, which is useful because these labels are often already correct and supervising correct labels is a waste of effort. In contrast, long segments reduce the cognitive effort due to context switches. Our method helps find the segmentation that optimizes supervision efficiency by defining user models to predict the cost and utility of supervising each segment and solving a constrained optimization problem balancing these contradictory objectives. A user study demonstrates noticeable gains over pre-segmented, confidence-ordered baselines on two natural language processing tasks: speech transcription and word segmentation. Matthias Sperber, Mirjam Simantzik, Graham Neubig, Satoshi Nakamura 0001, Alex Waibel |
Trans. Assoc. Comput. Linguistics | 4 |
| 2013 | Dialogue management for leading the conversation in persuasive dialogue systemsabstractIn this research, we propose a probabilistic dialogue modeling method for persuasive dialogue systems that interact with the user based on a specific goal, and lead the user to take actions that the system intends from candidate actions satisfying the user's needs. As a baseline system, we develop a dialogue model assuming the user makes decisions based on preference. Then we improve the model by introducing methods to guide the user from topic to topic. We evaluate the system knowledge and dialogue manager in a task that tests the system's persuasive power, and find that the proposed method is effective in this respect. Takuya Hiraoka, Yuki Yamauchi, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
ASRU | 6 |
| 2013 | Evaluation of a singing voice conversion method based on many-to-many eigenvoice conversionabstractIn this paper, we evaluate our proposed singing voice conver-sion method from various perspectives. To enable singers to freely control their voice timbre of singing voice, we have pro-posed a singing voice conversion method based on many-to-many eigenvoice conversion (EVC) that enables to convert the voice timbre of an arbitrary source singer into that of another arbitrary target singer using a probabilistic model. Further-more, to easily develop training data consisting of multiple par-allel data sets between a single reference singer and many other singers, a technique for efficiently and effectively generating the parallel data sets from nonparallel singing voice data sets of many singers using a singing-to-singing synthesis system have been proposed. However, we have never conducted sufficient investigations into the effectiveness of these proposed methods. In this paper, we conduct both objective and subjective eval-uations to carefully investigate the effectiveness of proposed methods. Moreover, the differences between singing voice con-version and speaking voice conversion are also analyzed. Ex-perimental results show that our proposed method succeeds in enabling people to control their own voice timbre by using only an extremely small amount of the target singing voice. Index Terms: singing voice, voice conversion, eigenvoice con-version, singing-to-singing synthesis, performance evaluation Hironori Doi, Tomoki Toda, Tomoyasu Nakano, Masataka Goto, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2013 | Simple, lexicalized choice of translation timing for simultaneous speech translation
Tomoki Fujita, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2013 | Generalizing continuous-space translation of paralinguistic informationabstractIn previous work, we proposed a model for speech-to-speech translation that is sensitive to paralinguistic information such as duration and power of spoken words [1]. This model uses linear regression to map source acoustic features to target acoustic features directly and in continuous space. However, while the model is effective, it faces scalability issues as a single model must be trained for every word, which makes it difficult to generalize to words for which we do not have parallel speech. In this work we first demonstrate that simply training a linear regression model on all words is not sufficient to express paralinguistic translation. We next describe a neural network model that has sufficient expressive power to perform paralinguistic translation with a single model. We evaluate the proposed method on a digit translation task and show that we achieve similar results with a single neural network-based model as previous work did using word-dependent models. Index Terms: speech translation, paralinguistic information, linear regression, neural network Takatomo Kano, Shinnosuke Takamichi, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura 0001 |
INTERSPEECH | 6 |
| 2013 | An investigation of acoustic features for singing voice conversion based on perceptual ageabstractIn this paper, we investigate the acoustic features that can be modified to control the perceptual age of a singing voice. Singers can sing expressively by controlling prosody and vocal timbre, but the varieties of voices that singers can produce are limited by physical constraints. Previous work has attempted to overcome this limitation through the use of statistical voice conversion. This technique makes it possible to convert singing voice characteristics of an arbitrary source singer into those of an arbitrary target singer. However, it is still difficult to intu-itively control singing voice characteristics by manipulating pa-rameters corresponding to specific physical traits, such as gen-der and age. In this paper, we focus on controlling the perceived age of the singer and, as a first step, perform an investigation of the factors that play a part in the listener’s perception of the singer’s age. The experimental results demonstrate that 1) the perceptual age of singing voices corresponds relatively well to the actual age of the singer, 2) speech analysis/synthesis pro-cessing and statistical voice conversion processing don’t cause adverse effects on the perceptual age of singing voices, and 3) prosodic features have a larger effect on the perceptual age than spectral features. Kazuhiro Kobayashi, Hironori Doi, Tomoki Toda, Tomoyasu Nakano, Masataka Goto, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 8 |
| 2013 | Grapheme-to-phoneme conversion based on adaptive regularization of weight vectorsabstractThe current state-of-the-art approach in grapheme-to-phoneme (g2p) conversion is structured learning based on the Margin Infused Relaxed Algorithm (MIRA), which is an online discriminative training method for multiclass classification. However, it is known that the aggressive weight update method of MIRA is prone to overfitting, even if the current example is an outlier or noisy. Adaptive Regularization of Weight Vectors (AROW) has been proposed to resolve this problem for binary classification. In addition, AROW’s update rule is simpler and more efficient than that of MIRA, allowing for more efficient training. Although AROW has these advantages, it has not been applied to g2p conversion yet. In this paper, we first apply AROW to g2p conversion which is structured learning problem. In an evaluation that employed a dataset including noisy data our proposed approach achieves a 5.3% error reduction rate compared to MIRA implemented in DirecTL+ in terms of phoneme error rate while requiring only 78% the training time. Index Terms:g2p conversion, out-of-vocabulary word, online discriminative training, structured learning, AROW Keigo Kubo, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2013 | A digital signal processor implementation of silent/electrolaryngeal speech enhancement based on real-time statistical voice conversionabstractIn this paper, we present a digital signal processor (DSP) implementation of real-time statistical voice conversion (VC) for silent speech enhancement and electrolaryngeal speech enhancement. As a silent speech interface, we focus on nonaudible murmur (NAM), which can be used in situations where audible speech is not acceptable. Electrolaryngeal speech is one of the typical types of alaryngeal speech produced by an alternative speaking method for laryngectomees. However, the sound quality of NAM and electrolaryngeal speech suffers from lack of naturalness. VC has proven to be one of the promising approaches to address this problem, and it has been successfully implemented on devices with sufficient computational resources. An implementation on devices that are highly portable but have limited computational resources would greatly contribute to its practical use. In this paper we further implement real-time VC on a DSP. To implement the two speech enhancement systems based on real-time VC, one from NAM to a whispered voice and the other from electrolaryngeal speech to a natural voice, we propose several methods for reducing computational cost while preserving conversion accuracy. We conduct experimental evaluations and show that real-time VC is capable of running on a DSP with little degradation. Index Terms: statistical voice conversion, real-time processing, reduction of computational cost, DSP, non-audible murmur, electrolaryngeal speech Takuto Moriguchi, Tomoki Toda, Motoaki Sano, Hiroshi Sato 0002, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 7 |
| 2013 | An empirical comparison of joint optimization techniques for speech translationabstractSpeech translation (ST) systems consist of three major components: automatic speech recognition (ASR), machine translation (MT), and speech synthesis (SS). In general the ASR system is tuned independently to minimize word error rate (WER), but previous research has shown that ASR and MT can be jointly optimized to improve translation quality [1]. Independently, many techniques have recently been proposed for the optimization of MT, such as empirical comparison of joint optimization using minimum error rate training (MERT) [2], pairwise ranking optimization (PRO) [3] and the batch margin infused relaxed algorithm (MIRA) [4]. The first contribution of this paper is an empirical comparison of these techniques in the context of joint optimization. As the last two methods are able to use sparse features, we also introduce lexicalized features using the frequencies of recognized words. In addition, motivated by initial results, we propose a hybrid optimization method that changes the translation evaluation measure depending on the features to be optimized. Experimental results for the best combination of algorithm and features show a gain of 1.3 BLEU points at 27% of the computational cost of previous joint optimization methods. Index Terms: speech translation, machine translation, automatic speech recognition, joint optimization Masaya Ohgushi, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2013 | Efficient speech transcription through respeakingabstractWe propose a method for efficient off-line speech transcription through respeaking. Speech is segmented into smaller utterances using an initial automatic transcript. Respeaking is performed segment by segment, while confidence filtering helps save supervision effort. We conduct detailed experiments comparing speaking vs. typing, sequential vs. confidence-ordered supervision, and examine the effect of the respeaking word error rate on correction efficiency. Our results demonstrate that the proposed method can not only outperform typing in terms of correction efficiency, but is also much less demanding for the respeakers than traditional respeaking methods, consequently helping to keep costs down. Matthias Sperber, Graham Neubig, Christian Fügen, Satoshi Nakamura 0001, Alex Waibel |
INTERSPEECH | 4 |
| 2013 | Improvements to HMM-based speech synthesis based on parameter generation with rich context modelsabstractIn this paper, we improve parameter generation with rich context models by modifying an initialization method and further apply it to both spectral and F0 components in HMM-based speech synthesis. To alleviate over-smoothing effects caused by the traditional parameter generation methods, we have previously proposed an iterative parameter generation method with rich context models. It has been reported that this method yields quality improvements in synthetic speech but there are still limitations. This is because 1) this generation method still suffers from the over-smoothing effect, as it uses the parameters generated by the traditional method as an initial parameters, which strongly affect on the finally generated parameters and 2) it is applied to only the spectral component. To address these issues, we propose 1) an initialization method to generate less smoothed but more discontinuous initial parameters that tend to yield better generated parameters, and 2) a parameter generation method with rich context models for the F0 component. Experimental results show that the proposed methods yield significant improvements in quality of synthetic speech. Index Terms: HMM-based speech synthesis, rich context models, GMM, context clustering, over-smoothing, MSD-HMM Shinnosuke Takamichi, Tomoki Toda, Yoshinori Shiga, Sakriani Sakti, Graham Neubig, Satoshi Nakamura 0001 |
INTERSPEECH | 6 |
| 2013 | A hybrid approach to electrolaryngeal speech enhancement based on spectral subtraction and statistical voice conversionabstractWe present a hybrid approach to improving naturalness of electrolaryngeal (EL) speech while minimizing degradation in intelligibility. An electrolarynx is a device that artificially generates excitation sounds to enable laryngectomees to produce EL speech. Although proficient laryngectomees can produce quite intelligible EL speech, it sounds very unnatural due to the mechanical excitation produced by the device. Moreover, the excitation sounds produced by the device often leak outside, adding noise to EL speech. To address these issues, previous work has proposed methods for EL speech enhancement through either noise reduction or voice conversion. The former usually causes no degradation in intelligibility but yields only small improvements in naturalness as the mechanical excitation sounds remain essentially unchanged. On the other hand, the latter method significantly improves naturalness of EL speech using spectral and excitation parameters of natural voices converted from acoustic parameters of EL speech, but it usually causes degradation in intelligibility owing to errors in conversion. We propose a hybrid method using the noise reduction method for enhancing spectral parameters and voice conversion method for predicting excitation parameters. The experimental results demonstrate the proposed method yields significant improvements in naturalness compared with EL speech while keeping intelligibility high enough. Index Terms: speaking-aid, electrolaryngeal speech, spectral subtraction, voice conversion, hybrid approach Kou Tanaka, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2013 | Multilingual Speech-to-Speech Translation System: VoiceTraabstractThis study presents an overview of VoiceTra, which was developed by NICT and released as the world's first network-based multilingual speech-to-speech translation system for smartphones, and describes in detail its multilingual speech recognition, its multilingual translation, and its multilingual speech synthesis in regards to field experiments. We show the effects of system updates using the data collected from field experiments to improve our acoustic and language models. Shigeki Matsuda, Xinhui Hu, Yoshinori Shiga, Hideki Kashioka, Chiori Hori, Keiji Yasuda, Hideo Okuma, Masao Uchiyama, Eiichiro Sumita, Hisashi Kawai, Satoshi Nakamura 0001 |
MDM (2) | 11 |
| 2013 | A-STAR: Toward translating Asian spoken languages
Sakriani Sakti, Michael Paul, Andrew M. Finch, Shinsuke Sakai, Thang Tat Vu, Noriyuki Kimura, Chiori Hori, Eiichiro Sumita, Satoshi Nakamura 0001, Jun Park, Chai Wutiwiwatchai, Bo Xu 0002, Hammam Riza, Karunesh Arora, Haizhou Li 0001 |
Comput. Speech Lang. | 9 |
| 2012 | A bootstrapping approach for SLU portability to a new language by inducting unannotated user queriesabstractThis paper proposes a bootstrapping method of constructing a new spoken language understanding (SLU) system in a target language by utilizing statistical machine translation given an SLU module in some source language. The main challenge in this work is to induct unannotated automatic speech recognition results of user queries in the source language collected through a spoken dialog system, which is under public test. In order to select candidate expressions from among erroneous translation results stemming from problems with speech recognition and machine translation, we use back-translation results to check whether the translation result maintains the semantic meaning of the original sentence. We demonstrate that the proposed scheme can effectively prefer suitable sentences for inclusion in the training data as well as help improve the SLU module for the target language. Teruhisa Misu, Etsuo Mizukami, Hideki Kashioka, Satoshi Nakamura 0001, Haizhou Li 0001 |
ICASSP | 4 |
| 2012 | An Evaluation of Parameter Generation Methods with Rich Context Models in HMM-Based Speech Synthesis
Shinnosuke Takamichi, Tomoki Toda, Yoshinori Shiga, Hisashi Kawai, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 6 |
| 2011 | Blind noise suppression for Non-Audible Murmur recognition with stereo signal processingabstractIn this paper, we propose a blind noise suppression method for Non-Audible Murmur (NAM) recognition. NAM is a very soft whispered voice detected with NAM microphone, which is one of the body-conductive microphones. Due to its recording mechanism, the detected signal suffers from noise caused by speaker's movements. In the proposed method using a stereo signal detected with two NAM microphones, the noise is estimated with blind source separation, and then, spectral subtraction is performed in each channel to reduce the noise. Moreover, channel selection is performed frame by frame to generate less distorted monaural NAM signal. Experimental results show that 1) word accuracy in large vocabulary continuous NAM recognition is degraded from 69.2% to 53.6% by the noise and 2) it is significantly recovered to 63.3% in a simulated situation and 58.6% in a real situation with the proposed method. Shunta Ishii, Tomoki Toda, Hiroshi Saruwatari, Sakriani Sakti, Satoshi Nakamura 0001 |
ASRU | 5 |
| 2011 | Unsupervised determination of efficient Korean LVCSR units using a Bayesian Dirichlet process modelabstractKorean is an agglutinative language that does not have explicit word boundaries. It is also a highly inflective language that exhibits severe coarticulation effects. These characteristics pose a challenge in developing large-vocabulary continuous speech recognition (LVCSR) systems. Many existing Korean LVCSR systems attempt to overcome these difficulties by defining a set of "word" units using morphological analysis (rule-based) or statistical methods. These approaches usually require a great deal of linguistic knowledge or at least some explicit information about the statistical distribution of the units. However, exceptions or uncommon words (e.g., foreign proper nouns) still exist that cannot be covered by rules alone. In this paper, we investigate the use of an unsupervised, nonparametric Bayesian approach to automatically determining efficient units for a Korean LVCSR system. Specifically, we utilize a Dirichlet process model trained using Bayesian inference through block Gibbs sampling. Our approach provides a principled way of learning units without explicit linguistic knowledge or any static parameters. Experiments were conducted on a travel domain corpus, which includes many foreign words and proper nouns. In our experiments we compared our method to a set of state-of-the-art baseline systems that relied on either morphological analysis or segmentation heuristics. Our system was able to produce a considerably more compact set of "word" units than the best baseline system (the lexical dictionary was approximately half the size), with a recognition accuracy 5.89% higher in terms of the relative word error rate than the best baseline system. Sakriani Sakti, Andrew M. Finch, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura 0001 |
ICASSP | 5 |
| 2011 | Increasing discriminative capability on MAP-based mapping function estimation for acoustic model adaptationabstractIn this study, we propose increasing discriminative power on the maximum a posteriori (MAP)-based mapping function estimation for acoustic model adaptation. Based on the effective and stable learning advantages of MAP-based estimation, we incorporate a discriminative term and derive a new objective function. By applying the new function for online mapping function estimation, we developed discriminative maximum a posteriori (DMAP) linear regression (DMAPLR) and DMAP-based ensemble speaker and speaking environment modeling (DMAP-based ESSEM). We evaluate the DMAPLR and DMAP-based ESSEM on the Aurora-2 task in a supervised adaptation mode. The experimental results show that both DMAPLR and DMAP-based ESSEM consistently provide improvements over their ML-based and MAP-based counterparts irrespective of using one, two, or three adaptation utterances. From the improvements, we confirm the strong effect of increasing discriminative capability on the MAP-based mapping function estimation. Moreover, we verify that including multiple knowledge sources in the objective function can efficiently enhance model adaptation performance. When compared with the baseline result DMAP-ESSEM achieves a 15.96% (9.21% to 7.74%) average word error rate (WER) reduction using only one adaptation utterance. Yu Tsao 0001, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura 0001 |
ICASSP | 4 |
| 2011 | A sampling-based environment population projection approach for rapid acoustic model adaptationabstractWe propose an environment population projection (EPP) approach for rapid acoustic model adaptation to reduce environment mismatches with limited amounts of adaptation data. This approach consists of two stages: population construction and projection. In the population construction stage, we apply a sampling scheme on the adaptation data to construct an environment population based on acoustic models prepared in the training phase. With this sampling procedure, the environment samples in the population characterize diverse acoustic information embedded in the adaptation data. Next, the projection stage estimates a function to map the environment population into one set of acoustic models that matches the testing condition. With a well constructed environment population, a simple projection function can enable the EPP approach to accurately characterize the testing environment even with a small amount of adaptation data. To examine the rapid adaptation ability of EPP, we used only one adaptation utterance and tested performance in both supervised and unsupervised adaptation modes on Aurora-2 and Aurora-2J tasks. It is found that EPP achieves satisfactory performance under both modes for both tasks. On the Aurora-2J task for example, EPP gives a clear improvement of a 13.87% (8.58% to 7.39%) word error rate (WER) reduction over our baseline in the unsupervised adaptation mode. Yu Tsao 0001, Shigeki Matsuda, Shinsuke Sakai, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura 0001 |
ICASSP | 6 |
| 2011 | Adaptive Regularization Framework for Robust Voice Activity Detection
Xugang Lu, Masashi Unoki, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2011 | User Study of Spoken Decision Support SystemabstractThis paper presents the results of the user evaluation of spo- ken decision support dialogue systems, which help users select from a set of alternatives. Thus far, we have modeled this deci- sion support dialogue as a partially observable Markov decision process (POMDP), and optimized its dialogue strategy to maxi- mize the value of the user’s decision. In this paper, we present a comparative evaluation of the optimized dialogue strategy with several baseline strategies, and demonstrate that the optimized dialogue strategy that was effective in user simulation experi- ments works well in an evaluation by real users. Teruhisa Misu, Kiyonori Ohtake, Chiori Hori, Hisashi Kawai, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2011 | Toward Construction of Spoken Dialogue System that Evokes Users' Spontaneous Backchannels
Teruhisa Misu, Etsuo Mizukami, Yoshinori Shiga, Shinichi Kawamoto, Hisashi Kawai, Satoshi Nakamura 0001 |
SIGDIAL Conference | 6 |
| 2011 | Sub-band temporal modulation envelopes and their normalization for automatic speech recognition in reverberant environments
Xugang Lu, Masashi Unoki, Satoshi Nakamura 0001 |
Comput. Speech Lang. | 3 |
| 2011 | Temporal modulation normalization for robust speech feature extraction and recognition
Xugang Lu, Shigeki Matsuda, Masashi Unoki, Satoshi Nakamura 0001 |
Multim. Tools Appl. | 4 |
| 2010 | Brazilian portuguese acoustic model training based on data borrowing from other language
Kazuhiko Abe, Sakriani Sakti, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2010 | Cluster-based language model for spoken document retrieval using NMF-based document clustering
Xinhui Hu, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2010 | Construction and evaluations of an annotated Chinese conversational corpus in travel domain for the language model of speech recognition
Xinhui Hu, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2010 | Voice activity detection in a reguarized reproducing kernel hilbert space
Xugang Lu, Masashi Unoki, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2010 | Utilizing a noisy-channel approach for Korean LVCSR
Sakriani Sakti, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2010 | Active learning of confidence measure function in robot language acquisition frameworkabstractIn an object manipulation dialogue, a robot may misunderstand an ambiguous command from a user, such as “Place the cup down (on the table),” potentially resulting in an accident. Although making confirmation questions before all motion will decrease the risk of this failure, the user will find it more convenient if confirmation questions are not made under trivial situations. This paper proposes a method for estimating ambiguity in the commands by introducing an active learning framework with Bayesian logistic regression to human-robot spoken dialogue. We conducted physical experiments in which a user and a manipulator-based robot communicated in spoken language to manipulate toys. Komei Sugiura, Naoto Iwahashi, Hideki Kashioka, Satoshi Nakamura 0001 |
IROS | 4 |
| 2010 | Dialogue Acts Annotation for NICT Kyoto Tour Dialogue Corpus to Construct Statistical Dialogue Systems
Kiyonori Ohtake, Teruhisa Misu, Chiori Hori, Hideki Kashioka, Satoshi Nakamura 0001 |
LREC | 5 |
| 2010 | Modeling Spoken Decision Making Dialogue and Optimization of its Dialogue Strategy
Teruhisa Misu, Komei Sugiura, Kiyonori Ohtake, Chiori Hori, Hideki Kashioka, Hisashi Kawai, Satoshi Nakamura 0001 |
SIGDIAL Conference | 7 |
| 2010 | Integrating lip-synch into game production workflow: "Sengoku BASARA 3" (Copyright restrictions prevent ACM from providing the full text for this article)abstractNo abstract available. Shinichi Kawamoto, Tatsuo Yotsukura, Satoshi Nakamura 0001, Junya Yamamoto, Tsunenori Shirahama, Hakuei Yamamoto |
SIGGRAPH ASIA (Sketches) | 3 |
| 2010 | Dialogue strategy optimization to assist user's decision for spoken consulting dialogue systemsabstractThis paper addresses a user model and dialogue state definition in spoken consulting dialogue systems that help users in making decision. When selecting from a set of alternatives, users have various decision criteria for making decision. Users often do not have a definite goal or criteria for selection, and thus they may find not only what kind of information the system can provide but their own preference or factors that they should emphasize. In this paper, we model such consulting dialogue as partially observable Markov decision process (POMDP). We then present an optimization of dialogue strategy to help users make better decisions. Teruhisa Misu, Komei Sugiura, Kiyonori Ohtake, Chiori Hori, Hideki Kashioka, Hisashi Kawai, Satoshi Nakamura 0001 |
SLT | 7 |
| 2010 | Temporal contrast normalization and edge-preserved smoothing of temporal modulation structures of speech for robust speech recognition
Xugang Lu, Shigeki Matsuda, Masashi Unoki, Satoshi Nakamura 0001 |
Speech Commun. | 4 |
| 2009 | Weighted finite state transducer based statistical dialog managementabstractWe proposed a dialog system using a weighted finite-state transducer (WFST) in which user concept and system action tags are input and output of the transducer, respectively. The WFST-based platform for dialog management enables us to combine various statistical models for dialog management (DM), user input understanding and system action generation, and then search the best system action in response to user inputs among multiple hypotheses. To test the potential of the WFST-based DM platform using statistical models, we constructed a dialog system using a human-to-human spoken dialog corpus for hotel reservation, which is annotated with Interchange Format (IF). A scenario WFST and a spoken language understanding (SLU) WFST were obtained from the corpus and then composed together and optimized. We evaluated the detection accuracy of the system next action tags using Mean Reciprocal Ranking (MRR). Finally, we constructed a full WFST-based dialog system by composing SLU, scenario and sentence generation (SG) WFSTs. Humans read the system responses in natural language and judged the quality of the responses. We confirmed that the WFST-based DM platform was capable of handling various spoken language and scenarios when the user concept and system action tags are consistent and distinguishable. Chiori Hori, Kiyonori Ohtake, Teruhisa Misu, Hideki Kashioka, Satoshi Nakamura 0001 |
ASRU | 5 |
| 2009 | The Asian network-based speech-to-speech translation systemabstractThis paper outlines the first Asian network-based speech-to-speech translation system developed by the Asian Speech Translation Advanced Research (A-STAR) consortium. The system was designed to translate common spoken utterances of travel conversations from a certain source language into multiple target languages in order to facilitate multiparty travel conversations between people speaking different Asian languages. Each A-STAR member contributes one or more of the following spoken language technologies: automatic speech recognition, machine translation, and text-to-speech through Web servers. Currently, the system has successfully covered 9 languages-namely, 8 Asian languages (Hindi, Indonesian, Japanese, Korean, Malay, Thai, Vietnamese, Chinese) and additionally, the English language. The system's domain covers about 20,000 travel expressions, including proper nouns that are names of famous places or attractions in Asian countries. In this paper, we discuss the difficulties involved in connecting various different spoken language translation systems through Web servers. We also present speech-translation results on the first A-STAR demo experiments carried out in July 2009. Sakriani Sakti, Noriyuki Kimura, Michael Paul, Chiori Hori, Eiichiro Sumita, Satoshi Nakamura 0001, Jun Park, Chai Wutiwiwatchai, Bo Xu 0002, Hammam Riza, Karunesh Arora, Haizhou Li 0001 |
ASRU | 6 |
| 2009 | MAP estimation of online mapping parameters in ensemble speaker and speaking environment modelingabstractRecently, an ensemble speaker and speaking environment modeling (ESSEM) framework was proposed to enhance automatic speech recognition performance under adverse conditions. In the online phase of ESSEM, the prepared environment structure in the offline stage is transformed to a set of acoustic models for the target testing environment by using a mapping function. In the original ESSEM framework, the mapping function parameters are estimated based on a maximum likelihood (ML) criterion. In this study, we propose to use a maximum a posteriori (MAP) criterion to calculate the mapping function to avoid a possible over-fitting problem that can degrade the accuracy of environment characterization. For the MAP estimation, we also study two types of prior densities, namely, clustered prior and hierarchical prior, in this paper. On the Aurora-2 task using either type of prior densities, MAP-based ESSEM can achieve better performance than ML-based ESSEM, especially under low SNR conditions. When comparing to our best baseline results, the MAP-based ESSEM achieves a 14.97% (5.41% to 4.60%) word error rate reduction in average at a signal to noise ratio of 0 dB to 20 dB over the three testing sets. Yu Tsao 0001, Shigeki Matsuda, Satoshi Nakamura 0001, Chin-Hui Lee 0001 |
ASRU | 3 |
| 2009 | Automatic voice assignment tool for Instant Casting movie SystemabstractIn Instant Casting movie System, a personal CG character is automatically generated. The character resembles a participant in a face geometry and texture. However, the voice of a character was an alternative voice determined by the gender of the participant. Therefore sometimes it's not enough to identify the personality of a CG character. In this paper, an automatic pre-scored voice assignment tool for a personal CG character is presented. Voice is essential to identify a personal character as well as a face feature. Our proposed system selects the most similar voice to the participants from voice database, and assigns it as a voice of CG character. Voice similarity criterion is presented by combination of eight acoustic features. After assigning voice data to a personal character, the voice track is played back in synchronization with the movement of the CG character. 60 voice variations have been prepared to our voice database. Validity of the assigned voice has been evaluated by MOS value. The proposed method has achieved 68% of the theoretical figure that is calculated by preliminary experiments. Yoshihiro Adachi, Shinichi Kawamoto, Tatsuo Yotsukura, Shigeo Morishima, Satoshi Nakamura 0001 |
ICASSP | 5 |
| 2009 | Statistical dialog management applied to WFST-based dialog systemsabstractWe have proposed an expandable dialog scenario description and platform to manage dialog systems using a weighted finite-state transducer (WFST) in which user concept and system action tags are input and output of the transducer, respectively. In this paper, we apply this framework to statistical dialog management in which a dialog strategy is acquired from a corpus of human-to-human conversation for hotel reservation. A scenario WFST for dialog management was automatically created from an N-gram model of a tag sequence that was annotated in the corpus with Interchange Format (IF). Additionally, a word-to-concept WFST for spoken language understanding (SLU) was obtained from the same corpus. The acquired scenario WFST and SLU WFST were composed together and then optimized. We evaluated the proposed WFST-based statistic dialog management in terms of correctness to detect the next system actions and have confirmed the automatically acquired dialog scenario from a corpus can manage dialog reasonably on the WFST-based dialog management platform. Chiori Hori, Kiyonori Ohtake, Teruhisa Misu, Hideki Kashioka, Satoshi Nakamura 0001 |
ICASSP | 5 |
| 2009 | Temporal contrast normalization and edge-preserved smoothing on temporal modulation structure for robust speech recognitionabstractIn this paper, we propose a two-step processing algorithm which adaptively normalizes the temporal modulation of speech to extract robust speech feature for automatic speech recognition systems. The first step processing is to normalize the temporal modulation contrast (TMC) of the cepstral time series for both clean and noisy speech. The second step processing is to smooth the normalized temporal modulation structure to reduce the artifacts due to noise while preserving the speech modulation events (edges). We tested our algorithm on speech recognition experiments in additive noise condition (AURORA-2J data corpus), reverberant noise condition (convolution of clean speech utterances from AURORA-2J with a smart room impulse response), and noisy condition with both reverberant and additive noise (air conditioner noise in a smart room). For comparison, the ETSI advanced front-end (AFE) algorithm was used. Our results showed that the algorithm provided: (1) for additive noise condition, 57.26% relative word error reduction (RWER) rate for clean conditional training (59.37% for AFE), and 33.52% RWER rate for multi-conditional training (35.77% for AFE), (2) for reverberant condition, 51.28% RWER rate (10.17% for AFE) and (3) for noisy condition with both reverberant and additive noise, 71.74% RWER rate (48.86% for AFE). Xugang Lu, Shigeki Matsuda, Masashi Unoki, Tohru Shimizu, Satoshi Nakamura 0001 |
ICASSP | 5 |
| 2009 | CART-based modeling of Chinese tonal patterns with a functional model tracing the fundamental frequency trajectoriesabstractWe propose an approach to modeling Chinese tonal patterns, focusing on the basic fundamental frequency (F0) patterns characterized by the contextual linguistic features that can be directly extracted from text. We analyze tonal patterns as sparse target points (tonal F0peaks and valleys) and represent them in parametric form within the framework of a functional F0model. The relationships between the target points and underlying linguistic features are trained using classification and regression tree analysis (CARTs), and this functional model is used to trace the F0trajectories when training the CARTs and to synthesize a tonal pattern from the target points predicted by the CARTs. Our experiments indicate that the proposed method has low F0prediction errors. Utilization of the F0ranges measured from training samples could significantly reduce the influences of differences in voice ranges on training a speaker-independent model. Furthermore, the most important roles in characterizing tonal patterns were played by a few linguistic features such as lexical tone context and the distinction between voiced from unvoiced initials. Jinfu Ni, Shinsuke Sakai, Tohru Shimizu, Satoshi Nakamura 0001 |
ICASSP | 4 |
| 2009 | Optimal learning of P-Layer additive F0 models with cross-validationabstractIn this paper, we present the derivation of the backfitting training algorithms for generic p-layer additive F0models for arbitrary positive integer p. We have presented the special cases of the algorithms with p = 2 and p = 3 that have been successfully applied to the modelings of Japanese and English F0contours, whereas the derivation of the algorithm was presented only for the two-layer case. The additive F0model have smoothing parameters that establish a trade-off between the fit to the training data and the smoothness of the fitted curves, which have been all set to unity in the previous works. In this paper, we also present an optimal approach to set the values of these parameters using cross validation. We performed the training using the Boston University Radio News Corpus and confirmed the effectiveness of the proposed method. Shinsuke Sakai, Tatsuya Kawahara, Tohru Shimizu, Satoshi Nakamura 0001 |
ICASSP | 4 |
| 2009 | Recent advances in WFST-based dialog system
Chiori Hori, Kiyonori Ohtake, Teruhisa Misu, Hideki Kashioka, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2009 | Subband temporal modulation spectrum normalization for automatic speech recognition in reverberant environments
Xugang Lu, Masashi Unoki, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2009 | A decision tree-based clustering approach to state definition in an excitation modeling framework for HMM-based speech synthesisabstractThis paper presents a decision tree-based algorithm to cluster residual segments assuming an excitation model based on statedependent filtering of pulse train and white noise. The decision tree construction principle is the same as the one applied to speech recognition. Here parent nodes are split using the residual maximum likelihood criterion. Once these excitation decision trees are constructed for residual signals segmented by full context models, using questions related to the full context of the training sentences, they can be utilized for excitation modeling in speech synthesis based on hidden Markov models (HMM). Experimental results have shown that the algorithm in question is very effective in terms of clustering residual signals given segmentation, pitch marks and full context questions, resulting in filters with good residual modeling properties. Ranniery Maia, Tomoki Toda, Keiichi Tokuda, Shinsuke Sakai, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2009 | A study on soft margin estimation of linear regression parameters for speaker adaptationabstractWe formulate a framework for soft margin estimation-based linear regression (SMELR) and apply it to supervised speaker adaptation. Enhanced separation capability and increased discriminative ability are two key properties in margin-based discriminative training. For the adaptation process to be able to flexibly utilize any amount of data, we also propose a novel interpolation scheme to linearly combine the speaker independent (SI) and speaker adaptive SMELR (SMELR/SA) models. The two proposed SMELR algorithms were evaluated on a Japanese large vocabulary continuous speech recognition task. Both the SMELR and interpolated SI+SMELR/SA techniques showed improved speech adaptation performance in comparison with the well-known maximum likelihood linear regression (MLLR) method. We also found that the interpolation framework works even more effectively than SMELR when the amount of adaptation data is relatively small. Shigeki Matsuda, Yu Tsao 0001, Jinyu Li 0001, Satoshi Nakamura 0001, Chin-Hui Lee 0001 |
INTERSPEECH | 4 |
| 2009 | Annotating communicative function and semantic content in dialogue act for construction of consulting dialogue systemsabstractOur goal in this study is to train a dialogue manager that can handle consulting dialogues through spontaneous interactions from a tagged dialogue corpus. We have collected 130 hours of consulting dialogues in sightseeing guidance domain. This paper provides our taxonomy of dialogue act (DA) annotation that can describe two aspects of utterances. One is a communicative function (speech act), and the other is a semantic content of an utterance. We provide an overview of the Kyoto tour guide dialogue corpus and a preliminary analysis using the dialogue act tags. Teruhisa Misu, Kiyonori Ohtake, Chiori Hori, Hideki Kashioka, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2009 | A close look into the probabilistic concatenation model for corpus-based speech synthesis
Shinsuke Sakai, Ranniery Maia, Hisashi Kawai, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2009 | Bayesian learning of confidence measure function for generation of utterances and motions in object manipulation dialogue taskabstractThis paper proposes a method that generates motions and utterances in an object manipulation dialogue task. The proposed method integrates belief modules for speech, vision, and motions into a probabilistic framework so that a user’s utterances can be understood based on multimodal information. Responses to the utterances are optimized based on an integrated confidence measure function for the integrated belief modules. Bayesian logistic regression is used for the learning of the confidence measure function. The experimental results revealed that the proposed method reduced the failure rate from 12% down to 2.6% while the rejection rate was less than 24%. Index Terms: multimodal spoken dialogue system, robot language acquisition, confidence, Bayesian logistic regression Komei Sugiura, Naoto Iwahashi, Hideki Kashioka, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2009 | Automatic pronunciation scoring of words and sentences independent from the non-native's first language
Tobias Cincarek, Rainer Gruhn, Christian Hacker, Elmar Nöth, Satoshi Nakamura 0001 |
Comput. Speech Lang. | 5 |
| 2008 | Development of Indonesian Large Vocabulary Continuous Speech Recognition System within A-STAR Project
Sakriani Sakti, Eka Kelana, Hammam Riza, Shinsuke Sakai, Konstantin Markov, Satoshi Nakamura 0001 |
IJCNLP | 6 |
| 2008 | Dialog management using weighted finite-state transducers
Chiori Hori, Kiyonori Ohtake, Teruhisa Misu, Hideki Kashioka, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2008 | Improved novelty detection for online GMM based speaker diarizationabstractDetection of speakers which have not been seen before is an essential part of every online speaker diarization system. New speaker detection accuracy has direct impact on the overall diarization performance. In our previous system, for novelty detection we used global GMM likelihood ratio (LR) threshold. However, as the system analysis showed, the optimal threshold depends on the speaker gender as well as on the number of registered speakers. In this paper, we present the results of this analysis and the approach we have taken to solve this problem. First, we use different thresholds for male and female speakers, and second, for each gender before the thresholding we apply likelihood ratio mean and variance normalization. This greatly reduced the threshold dependency on the number of speakers and allowed to use fixed threshold for each gender. The LR distribution statistics are collected online and updated every time new likelihood ratio is calculated. Experiments on the TCSTAR database showed that compared with the previous global threshold method, the new novelty detection approach reduces the speaker diarization error rate up to 35%. Konstantin Markov, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2008 | CENSREC-4: development of evaluation framework for distant-talking speech recognition under reverberant environmentsabstractIn this paper, we newly introduce a collection of databases and evaluation tools called CENSREC-4, which is an evaluation framework for distant-talking speech under hands-free conditions. Distant-talking speech recognition is crucial for a handsfree speech interface. Therefore, we measured room impulse responses to investigate reverberant speech recognition in various environments. The data contained in CENSREC-4 are connected digit utterances, as in CENSREC-1. Two subsets are included in the data: basic data sets and extra data sets. The basic data sets are used for the evaluation environment for the room impulse response-convolved speech data. The extra data sets consist of simulated and recorded data. An evaluation framework is only provided for the basic data sets as evaluation tools. The results of evaluation experiments proved that CENSREC-4 is an effective database for evaluating the new dereverberation method because the traditional dereverberation process had difficulty sufficiently improving the recognition performance. Index Terms: Various environments, Impulse response, Convolution, Real recorded data, Evaluation framework Masato Nakayama, Takanobu Nishiura, Yuki Denda, Norihide Kitaoka, Kazumasa Yamamoto, Takeshi Yamada, Satoru Tsuge, Chiyomi Miyajima, Masakiyo Fujimoto, Tetsuya Takiguchi, Satoshi Tamura, Tetsuji Ogawa, Shigeki Matsuda, Shingo Kuroiwa, Kazuya Takeda, Satoshi Nakamura 0001 |
INTERSPEECH | 16 |
| 2008 | Evaluation Framework for Distant-talking Speech Recognition under Reverberant Environments: newest Part of the CENSREC Series -
Takanobu Nishiura, Masato Nakayama, Yuki Denda, Norihide Kitaoka, Kazumasa Yamamoto, Takeshi Yamada, Satoru Tsuge, Chiyomi Miyajima, Masakiyo Fujimoto, Tetsuya Takiguchi, Satoshi Tamura, Shingo Kuroiwa, Kazuya Takeda, Satoshi Nakamura 0001 |
LREC | 14 |
| 2008 | Spoken Dialog System for Next Generation Knowledge AccessabstractThis paper described our development dialog system on Kyoto tourist information assistance. Dialog part of our system helped user to make an appropriate query. Information analysis part would be assisted for user to select the retrieved information. Nowadays we can get most information through the Internet. However, we have a trouble to pick up expected information from the huge results with conventional search engines. Especially in mobile terminal, we are confronted with great difficulties for two factors. One is that most of users cannot make an appropriate query because their request is vague with theirselves. The other is that the retrieved information has huge variation and mobile terminal has small area for displaying them. Therefore, we aim to develop technologies for the users to input their requests by familiar way and clarify what they want to know with displaying the retrieved information with suitable method. Hideki Kashioka, Susumu Akamine, Takafumi Nakanishi, Hisashi Miyamori, Koji Zettsu, Yutaka Kidawara, Satoshi Nakamura 0001 |
MDM | 7 |
| 2008 | Post-recording tool for instant casting movie systemabstractThis paper proposes a universal user-friendly post-recording tool for an Instant Casting Movie System (ICS) that enables anyone to be a movie star using his or her own voice and faces. A personal CG character is automatically generated by scanning one's face geometry and image in ICS. Voice is as essential to identify a person as face. However, a character's voice is only based on gender in ICS. We proposed a novel voice recording tool for participants of all ages in a short time. Post-recording tasks are very difficult because speakers should speak in synchronization with the mouth movements of the CG characters. Therefore this task is generally recorded by professional voice actors. Our proposed tool has the following four features: 1) various supporting information for synchronization with voice and mouth movement timing for users; 2) automatic post-processing of recorded voices for compositing mixed audio; 3) intuitively displays operation for people of all ages; and 4) handles multiple users in parallel for quick recording. We developed a prototype speech synchronization system using a post-recording tool and conducted subjective evaluation experiments of it. Over 60% of the subjects responded that the tool's interface can be operated easily. Shinichi Kawamoto, Tatsuo Yotsukura, Shigeo Morishima, Satoshi Nakamura 0001 |
ACM Multimedia | 4 |
| 2008 | Efficient lip-synch tool for 3D cartoon animationabstractAbstract We propose a set of algorithms to efficiently make speech animation for 3D cartoon characters. Our prototype system is based on blendshapes, a linear interpolation technique, which is widely used in facial animation practice. In our system, a few base target shapes of the character, prerecorded voice, and its transcription are required as input. We describe a simple technique that amplifies the target shapes from few inputs using a generic database of viseme mouth shapes. We also introduce additional lip‐synch editing parameters that allow designers to quickly tune the lip movements. Based on these, we implement our prototype system as a Maya plug‐in. The demonstration movies created with this system illustrate well the practicality of our approach. Copyright © 2008 John Wiley & Sons, Ltd. Shinichi Kawamoto, Tatsuo Yotsukura, Ken Anjyo, Satoshi Nakamura 0001 |
Comput. Animat. Virtual Worlds | 4 |
| 2008 | A Robust Speech Recognition System for Communication Robots in Noisy EnvironmentsabstractThe application range of communication robots could be widely expanded by the use of automatic speech recognition (ASR) systems with improved robustness for noise and for speakers of different ages. In past researches, several modules have been proposed and evaluated for improving the robustness of ASR systems in noisy environments. However, this performance might be degraded when applied to robots, due to problems caused by distant speech and the robot's own noise. In this paper, we implemented the individual modules in a humanoid robot, and evaluated the ASR performance in a real-world noisy environment for adults' and children's speech. The performance of each module was verified by adding different levels of real environment noise recorded in a cafeteria. Experimental results indicated that our ASR system could achieve over 80% word accuracy in 70-dBA noise. Further evaluation of adult speech recorded in a real noisy environment resulted in 73% word accuracy. Carlos Toshinori Ishi, Shigeki Matsuda, Takayuki Kanda 0001, Takatoshi Jitsuhiro, Hiroshi Ishiguro, Satoshi Nakamura 0001, Norihiro Hagita |
IEEE Trans. Robotics | 6 |
| 2007 | NICT-ATR Speech-to-Speech Translation System
Eiichiro Sumita, Tohru Shimizu, Satoshi Nakamura 0001 |
ACL | 3 |
| 2007 | Development of VAD evaluation framework CENSREC-1-C and investigation of relationship between VAD and speech recognition performanceabstractVoice activity detection (VAD) plays an important role in speech processing including speech recognition, speech enhancement, and speech coding in noisy environments. We developed an evaluation framework for VAD in such environments, called corpus and environment for noisy speech recognition 1 concatenated (CENSREC-1-C). This framework consists of noisy continuous digit utterances and evaluation tools for VAD results. By adoptiong two evaluation measures, one for frame-level detection performance and the other for utterance-level detection performance, we provide the evaluation results of a power-based VAD method as a baseline. When using VAD in speech recognizer, the detected speech segments are extended to avoid the loss of speech frames and the pause segments are then absorbed by a pause model. We investigate the balance of an explicit segmentation by VAD and an implicit segmentation by a pause model using an experimental simulation of segment extension and show that a small extension improves speech recognition. Norihide Kitaoka, Kazumasa Yamamoto, Tomohiro Kusamizu, Seiichi Nakagawa, Takeshi Yamada, Satoru Tsuge, Chiyomi Miyajima, Takanobu Nishiura, Masato Nakayama, Yuki Denda, Masakiyo Fujimoto, Tetsuya Takiguchi, Satoshi Tamura, Shingo Kuroiwa, Kazuya Takeda, Satoshi Nakamura 0001 |
ASRU | 16 |
| 2007 | Never-ending learning system for on-line speaker diarizationabstractIn this paper, we describe newhigh-performanceon-line speaker diarization system which works faster than real-time and has very low latency. It consists of several modules including voice activity detection, novel speaker detection, speaker gender and speaker identity classification. Allmodules share a set of Gaussian mixturemodels (GMM) representing pause, male and female speakers, and each individual speaker. Initially, there are only three GMMs for pause and two speaker genders, trained in advance from some data. During the speaker diarization process, for each speech segment it is decidedwhether it comes from a new speaker or from already known speaker. In case of a new speaker, his/her gender is identified, and then, from the corresponding gender GMM, a new GMM is spawned by copying its parameters. This GMM is learned on-line using the speech segment data and from this point it is used to represent the new speaker. All individual speaker models are produced in this way. In the case of an old speaker, s/he is identified and the correspondingGMMis again learned on-line. In order to prevent an unlimited grow of the speaker model number, those models that have not been selected as winners for a long period of time are deleted from the system. This allows the system to be able to perform its task indefinitely in addition to being capable of self-organization, i.e. unsupervised adaptive learning, and preservation of the learned knowledge, i.e. speakers. Such functionalities are attributed to the so called Never-Ending Learning systems. For evaluation, we used part of the TC-STAR database consisting of European Parliament Plenary speeches. The results show that this system achieves a speaker diarization error rate of 4.6% with latency of at most 3 seconds. Konstantin Markov, Satoshi Nakamura 0001 |
ASRU | 2 |
| 2007 | Use of Poisson Processes to Generate Fundamental Frequency ContoursabstractThe prosodic contributions to voice fundamental frequency (F0) contours can be analyzed into a series of sparser tonal targets (F0peaks and valleys). The transitions through these targets are interpolated by spline or filtering functions to predict the shape of F0contours. A functional model was proposed in the previous work for this purpose. This paper presents an enhanced version of this model achieved by replacing its decay filter with a Poisson-process-induced filter. It is enhanced because the former is a special case of the latter. The new filter manages to delay the decaying process while interpolations are being uttered. A target point can thus act as target levels, if necessary. The algorithms for estimating parameters, which were implemented on computers, are also presented. Experiments conducted on thousands of observed F0contours, including Mandarin, Japanese, and English, indicate that the enhanced version significantly facilitates their automatic parameterization. Jinfu Ni, Satoshi Nakamura 0001 |
ICASSP (4) | 2 |
| 2007 | Never-ending learning with dynamic hidden Markov networkabstractCurrent automatic speech recognition systems have two distinctive modes of operation: training and recognition. After the training, system parameters are fixed, and if a mismatch between training and testing conditions occurs, an adaptation procedure is commonly applied. However, the adaptation methods change the system parameters in such a way that previously learned knowledge is irrecoverably destroyed. In searching for a solution to this problem and motivated by the results of recent neuro-biological studies, we have developed a network of hidden Markov states that is capable of unsupervised on-line adaptive learning while preserving the previously acquired knowledge. Speech patterns are represented by state sequences or paths through the network. The network can detect previously unseen patterns, and if such a new pattern is encountered, it is learned by adding new states and transitions to the network. Paths and states corresponding to spurious events or ”noises” and, therefore, rarely visited, are gradually removed. Thus, the network can grow and shrink when needed, i.e. it dynamically changes its structure. The learning process continues as long as the network lasts, i.e. theoretically forever, so it is called neverending learning. The output of the network is the best state sequence and the decoding is done concurrently with the learning. Thus the network always operates in a single learning/decoding mode. Initial experiments with a small database of isolated spelled letters showed that the Dynamic Hidden Markov network is indeed capable of never-ending learning and can perfectly recognize previously learned speech patterns. Index Terms: never-ending learning, life-long learning, dynamic hidden markov network, self-organization, topology representation. Konstantin Markov, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2007 | An HMM acoustic model incorporating various additional knowledge sourcesabstractWe introduce a method of incorporating additional knowledge sources into an HMM-based statistical acoustic model. The probabilistic relationship between information sources is first learned through a Bayesian network to easily integrate any ad-ditional knowledge sources that might come from any domain and then the global joint probability density function (PDF) of the model is formulated. Where the model becomes too com-plex and direct BN inference is intractable, we utilize a junction tree algorithm to decompose the global joint PDF into a linked set of local conditional PDFs. This way, a simplified form of the model can be constructed and reliably estimated using a limited amount of training data. Here, we apply this framework to in-corporate accents, gender, and wide-phonetic knowledge infor-mation at the HMM phonetic model level. The performance of the proposed method was evaluated on an LVCSR task using two different types of accented English speech data. Experi-mental results revealed that our method improves word accu-racy with respect to standard HMM. Index Terms: acoustic model, knowledge incorporation, Bayesian network, junction tree algorithm, wide-phonetic knowledge. 1. Sakriani Sakti, Konstantin Markov, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2007 | Multichannel Bin-Wise Robust Frequency-Domain Adaptive Filtering and Its Application to Adaptive BeamformingabstractLeast-squares error (LSE) or mean-squared error (MSE) optimization criteria lead to adaptive filters that are highly sensitive to impulsive noise. The sensitivity to noise bursts increases with the convergence speed of the adaptation algorithm and limits the performance of signal processing algorithms, especially when fast convergence is required, as for example, in adaptive beamforming for speech and audio signal acquisition or acoustic echo cancellation. In these applications, noise bursts are frequently due to undetected double-talk. In this paper, we present impulsive noise robust multichannel frequency-domain adaptive filters (MC-FDAFs) based on outlier-robust M-estimation using a Newton algorithm and a discrete Newton algorithm, which are especially designed for frequency bin-wise adaptation control. Bin-wise adaptation and control in the frequency-domain enables the application of the outlier-robust MC-FDAFs to a generalized sidelobe canceler (GSC) using an adaptive blocking matrix for speech and audio signal acquisition. It is shown that the improved robustness leads to faster convergence and to higher interference suppression relative to nonrobust adaptation algorithms, especially during periods of strong interference Wolfgang Herbordt, Herbert Buchner, Satoshi Nakamura 0001, Walter Kellermann |
IEEE Trans. Speech Audio Process. | 3 |
| 2007 | Out-of-Domain Utterance Detection Using Classification Confidences of Multiple TopicsabstractOne significant problem for spoken language systems is how to cope with users' out-of-domain (OOD) utterances which cannot be handled by the back-end application system. In this paper, we propose a novel OOD detection framework, which makes use of the classification confidence scores of multiple topics and applies a linear discriminant model to perform in-domain verification. The verification model is trained using a combination of deleted interpolation of the in-domain data and minimum-classification-error training, and does not require actual OOD data during the training process, thus realizing high portability. When applied to the "phrasebook" system, a single utterance read-style speech task, the proposed approach achieves an absolute reduction in OOD detection errors of up to 8.1 points (40% relative) compared to a baseline method based on the maximum topic classification score. Furthermore, the proposed approach realizes comparable performance to an equivalent system trained on both in-domain and OOD data, while requiring no OOD data during training. We also apply this framework to the "machine-aided-dialogue" corpus, a spontaneous dialogue speech task, and extend the framework in two manners. First, we introduce topic clustering which enables reliable topic confidence scores to be generated even for indistinct utterances, and second, we implement methods to effectively incorporate dialogue context. Integration of these two methods into the proposed framework significantly improves OOD detection performance, achieving a further reduction in equal error rate (EER) of 7.9 points Ian Lane, Tatsuya Kawahara, Tomoko Matsui, Satoshi Nakamura 0001 |
IEEE Trans. Speech Audio Process. | 4 |
| 2007 | Incorporating Knowledge Sources Into a Statistical Acoustic Model for Spoken Language Communication SystemsabstractThis paper introduces a general framework for incorporating additional sources of knowledge into an HMM-based statistical acoustic model. Since the knowledge sources are often derived from different domains, it may be difficult to formulate a probabilistic function of the model without learning the causal dependencies between the sources. We utilized a Bayesian network framework to solve this problem. The advantages of this graphical model framework are 1) it allows the probabilistic relationship between information sources to be learned and 2) it facilitates the decomposition of the joint probability density function (PDF) into a linked set of local conditional PDFs. This way, a simplified form of the model can be constructed and reliably estimated using a limited amount of training data. We applied this framework to the problem of incorporating wide-phonetic knowledge information, which often suffers from a sparsity of data and memory constraints. We evaluated how well the proposed method performed on an large-vocabulary continuous speech recognition (LVCSR) task using English speech data that contained two different types of accents. The experimental results revealed that it improved the word accuracy with respect to standard HMM, with or without additional sources of knowledge. Sakriani Sakti, Konstantin Markov, Satoshi Nakamura 0001 |
IEEE Trans. Computers | 3 |
| 2006 | Sequential Non-Stationary Noise Tracking Using Particle Filtering with Switching Dynamical SystemabstractThis paper addresses a speech recognition problem in non-stationary noise environments: the estimation of noise sequences. To solve this problem, we present a particle filter-based sequential noise estimation method for the front-end processing of speech recognition. In the proposed method, the particle filter is defined by a dynamical system based on Polyak averaging and feedback. We also introduce a switching dynamical system into the particle filter to cope with the state transition characteristics of non-stationary noise. In the evaluation results, we observed that the proposed method improves speech recognition accuracy in the results of non-stationary noise environments by a noise compensation method with stationary noise assumptions Masakiyo Fujimoto, Satoshi Nakamura 0001 |
ICASSP (1) | 2 |
| 2006 | Incorporation of Pentaphone-Context Dependency Based on Hybrid Hmm/Bn Acoustic Modeling FrameworkabstractThis paper presents a new method of modeling pentaphone-context units using the hybrid HMM/BN acoustic modeling. Rather than modeling pentaphones explicitly, in this approach we extend the modeled phonetic context within the triphone framework, since the probabilistic dependencies between the triphone context unit and the second preceding/following contexts are incorporated into the triphone state output distributions by means of the BN. Another advantage is that we can use a standard decoding system by assuming the next preceding/following context variables hidden during recognition. In this study, the performance of pentaphone HMM/BN model was evaluated with our LVCSR system by phoneme recognition and by large-vocabulary continuous word recognition tasks. In both cases, we observed consistently improved performance over the standard HMM based triphone model with the same number of parameters Sakriani Sakti, Konstantin Markov, Satoshi Nakamura 0001 |
ICASSP (1) | 3 |
| 2006 | Automatic Derivation of a Phoneme Set with Tone Information for Chinese Speech Recognition Based on Mutual Information CriterionabstractAn appropriate approach to model tone information is helpful for building Chinese large vocabulary continuous speech recognition system. We propose to derive an efficient phoneme set of tone-dependent sub-word units to build a recognition system, by iteratively merging a pair of tone-dependent units according to the principle of minimal loss of the mutual information. The mutual information is measured between the word tokens and their phoneme transcriptions in a training text corpus, based on the system lexical and language model. The approach has the capability to keep discriminative tonal (and phoneme) contrasts that are most helpful for disambiguating homophone words due to lack of tones, and merge those tonal (and phoneme) contrasts that are not important for word disambiguation for the recognition task. This enable a flexible selection of phoneme set according to a balance between the MI information amount and the number of phonemes. We applied the method to traditional phoneme set of Initial/Finals, and derived several phoneme sets with different number of units. Speech recognition experiments using the derived sets showed their effectiveness. Jinsong Zhang 0001, Xinhui Hu, Satoshi Nakamura 0001 |
ICASSP (1) | 3 |
| 2006 | Forward-backwards training of hybrid HMM/BN acoustic modelsabstractIn this paper, we describe an application of the Forward-Backwards (F-B) algorithm for maximum likelihood training of hybrid HMM/Bayesian Network (BN) acoustic models. Previously, HMM/BN parameter estimation was based on a Viterbi training algorithm that requires two passes over the training data: one for BN learning and one for updating HMM transition probabilities. In this work, we first analyze the F-B training for a conventional HMM and show that the state PDF parameter estimation is analogous to weighted-data classifier training. The gamma variable of the Forward-Backwards algorithm plays the role of the data weight. From this perspective, it is straightforward to apply FB-based training to the HMM/BN models since the BN learning algorithm allows training with weighted data. Experiments on accented speech (American, British and Australian English) show that F-B training outperforms the previous Viterbi learning approach and that the HMM/BN model achieved better performance than the conventional HMM. Index Terms: forward-backwards algorithm, HMM/BN, weighted data training, accent modeling. Konstantin Markov, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2006 | CENSREC2: corpus and evaluation environments for in car continuous digit speech recognition
Satoshi Nakamura 0001, Masakiyo Fujimoto, Kazuya Takeda |
INTERSPEECH | 1 |
| 2006 | The use of Bayesian network for incorporating accent, gender and wide-context dependency informationabstractWe propose a new method of incorporating the additional knowledge of accent, gender, and wide-context dependency information into ASR systems by utilizing the advantages of Bayesian networks. First, we only incorporate pentaphone-context dependency information. After that, accent and gender information are also integrated. In this method, we can easily extend conventional triphone HMMs to cover various sources of knowledge. The probabilistic dependencies between a triphone context unit and additional knowledge are learned through a BN. Another advantage is that during recognition, additional knowledge variables are assumed to be hidden, so that the existing standard triphone-based decoding system can be used without modification. The performance of the proposed model was evaluated on an LVCSR task using two different types of accented English speech data. Experimental results show that this proposed method improves word accuracy with respect to standard triphone models. Index Terms: acoustic modeling, bayesian network, knowledge incorporation, wide-context dependency. Sakriani Sakti, Konstantin Markov, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2006 | Speech recognition of foreign out-of-vocabulary words using a hierarchical language model
Hirofumi Yamamoto, Gen-ichiro Kikui, Satoshi Nakamura 0001, Yoshinori Sagisaka |
INTERSPEECH | 3 |
| 2006 | Oriental COCOSDA: Past, Present and Future
Shuichi Itahashi, Chiu-yu Tseng, Satoshi Nakamura 0001 |
LREC | 3 |
| 2006 | Developing Client-Server Speech Translation PlatformabstractThis paper describes a client-server speech translation platform designed for use at mobile terminals. Because terminals and servers are connected via a 3G public mobile phone networks, speech translation services are available at various places with thin client. This platform realizes hands-free communication and robustness for real use of speech translation in noisy environments. A microphone array and new noise suppression technique improves speech recognition performance, and a corpus-based approach enables wide coverage, robustness and portability to new languages and domains. The experimental result for evaluating the communicability of speakers of different languages shows that task completion rates using the speech translation system of 85% and 75% are achieved for Japanese- English and Japanese-Chinese, respectively. The system also has the ability to convey approximately one item of information per 2 utterances (one turn) on average for both Japanese-English and Japanese-Chinese in a task-oriented dialogue. Tohru Shimizu, Yutaka Ashikari, Toshiyuki Takezawa, Masahide Mizushima, Gen-ichiro Kikui, Yutaka Sasaki, Satoshi Nakamura 0001 |
MDM | 7 |
| 2006 | Integration of articulatory and spectrum features based on the hybrid HMM/BN modeling framework
Konstantin Markov, Satoshi Nakamura 0001 |
Speech Commun. | 3 |
| 2006 | HMM-based noise-robust feature compensation
Akira Sasou, Futoshi Asano, Satoshi Nakamura 0001, Kazuyo Tanaka |
Speech Commun. | 3 |
| 2006 | The ATR multilingual speech-to-speech translation systemabstractIn this paper, we describe the ATR multilingual speech-to-speech translation (S2ST) system, which is mainly focused on translation between English and Asian languages (Japanese and Chinese). There are three main modules of our S2ST system: large-vocabulary continuous speech recognition, machine text-to-text (T2T) translation, and text-to-speech synthesis. All of them are multilingual and are designed using state-of-the-art technologies developed at ATR. A corpus-based statistical machine learning framework forms the basis of our system design. We use a parallel multilingual database consisting of over 600 000 sentences that cover a broad range of travel-related conversations. Recent evaluation of the overall system showed that speech-to-speech translation quality is high, being at the level of a person having a Test of English for International Communication (TOEIC) score of 750 out of the perfect score of 990. Satoshi Nakamura 0001, Konstantin Markov, Hiromi Nakaiwa, Gen-ichiro Kikui, Hisashi Kawai, Takatoshi Jitsuhiro, Jinsong Zhang 0001, Hirofumi Yamamoto, Eiichiro Sumita, Seiichi Yamamoto |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | Particle Filter Based Non-Stationary Noise Tracking for Robust Speech RecognitionabstractThis paper addresses the main speech recognition problem in nonstationary noise environments: the estimation of noise sequences. To solve this problem, we present a particle filter-based sequential noise estimation method for front-end processing of speech recognition in noise. In the proposed method, a noise sequence is estimated through a sequential importance sampling step, then a residual resampling step, and finally a Markov chain Monte Carlo step with Metropolis-Hastings sampling. The estimated noise sequence is applied to MMSE-based clean speech estimation method. The evaluations were conducted on speech recognition in highly nonstationary noise environments. In the evaluation results, we observed that the proposed method improves speech recognition accuracy in non-stationary noise environments over noise compensation with stationary noise assumptions. Masakiyo Fujimoto, Satoshi Nakamura 0001 |
ICASSP (1) | 2 |
| 2005 | Joint optimization of LCMV beamforming and acoustic echo cancellation for automatic speech recognitionabstractFor full-duplex hands-free acoustic human/machine interfaces, a combination of acoustic echo cancellation and speech enhancement is often required in order to suppress acoustic echoes, local interference and noise. In order to exploit positive synergies between acoustic echo cancellation and speech enhancement optimally, we previously presented a combined least-squares (LS) optimization criterion for the integration of acoustic echo cancellation and adaptive linearly-constrained minimum variance (LCMV) beamforming (Herbordt, W. et al., Proc. EURASIP European Sig. Process. Conf., 2004). By means of speech recognition experiments, we now illustrate the efficiency of the proposed solution in situations with high levels of background noise and with time-varying echo paths and frequent double-talk. Wolfgang Herbordt, Satoshi Nakamura 0001, Walter Kellermann |
ICASSP (3) | 2 |
| 2005 | Modeling Successive Frame Dependencies with Hybrid HMM/BN Acoustic ModelabstractMost current state-of-the-art speech recognition systems use the hidden Markov model (HMM) for modeling the acoustical characteristics of a speech signal. In the first-order HMM, speech data are assumed to be independently and identically distributed (iid), meaning that there is no dependency between neighboring feature vectors. Another assumption is that the current vector depends only on the current HMM state. In practice, however, these assumptions are not true. We describe a hybrid HMM/BN (Bayesian network) acoustic model, where the dependency of the current speech vector on the previous vector and on the previous state is also learned and used in speech recognition. This is possible because the state probability distribution is modeled by a BN. Previous instances of the state and speech feature vector are represented by additional variables of the BN and the probabilistic dependencies between them, and their current instances are learned during training. During recognition, the likelihood of the current feature vector is inferred from the BN where the previous state and previous feature vector are treated as hidden. We have evaluated this hybrid HMM/BN model with our LVCSR system by phoneme recognition and by large-vocabulary continuous word recognition tasks. In both cases, we observed improved performance over the conventional Gaussian mixture HMM. Konstantin Markov, Satoshi Nakamura 0001 |
ICASSP (1) | 2 |
| 2005 | Online cepstral filtering using a sequential EM approach with Polyak averaging and feedbackabstractWe propose an online filtering algorithm that aims to alleviate the decrease we see in ASR performance when the speech is corrupted by additive noise. Using an initial estimate of the noise distribution, the algorithm updates the noise model on a frame synchronous basis. Using Polyak averaging we obtain a sequence of robust, frame-synchronous noise model estimates, and a minimum mean square error (MMSE) filter is used to denoise the cepstral coefficients. The algorithm is compared to a batch version which uses several iterations of the EM-algorithm over the complete utterance to estimate the noise model, and it is shown that the performance obtained using the averaging of the noise model is comparable to the batch performance. Tor André Myrvoll, Satoshi Nakamura 0001 |
ICASSP (1) | 2 |
| 2005 | Spoken dialog system and its evaluation of geographic information system for elderly persons' mobility support
Takatoshi Jitsuhiro, Shigeki Matsuda, Yutaka Ashikari, Satoshi Nakamura 0001, Ikuko Eguchi Yairi, Seiji Igi |
INTERSPEECH | 4 |
| 2005 | Outlier detection for acoustic model training using robust statisticsabstractIn this paper, we propose an acoustic model training technique which is robust against outliers such as clipping, unexpected noise, poorly pronounced word segments, or mistranscriptions, which deteriorate the quality of the acoustic models and in turn decrease speech recognition performance. The outlier-robust acoustic model training technique is based on a maximum likelihood (ML) criterion and automatically detects and removes outliers from the training data. Experiments with artificially contaminated mis-transcribed training data show that nearly the same word error rate can be obtained for contaminated data using the proposed technique as for uncontaminated data. Application to a dialogue speech database with unknown outliers reduces the errors by 4.03 %. Shigeki Matsuda, Wolfgang Herbordt, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2005 | Incorporating a Bayesian wide phonetic context model for acoustic rescoringabstractThis paper presents a method for improving acoustic model precision by incorporating wide phonetic context units in speech recognition. The wide phonetic context model is constructed from several narrower context-dependent models based on the Bayesian framework. Such a composition is performed in order to avoid the crucial problem of a limited availability of training data and to reduce the model complexity. To enhance the model reliability due to unseen contexts and limited training data, flooring and deleted interpolation techniques are used. Experimental results show that this method gives improvement of the word accuracy with respect to the standard triphone model. Sakriani Sakti, Satoshi Nakamura 0001, Konstantin Markov |
INTERSPEECH | 2 |
| 2005 | Tone nucleus-based multi-level robust acoustic tonal modeling of sentential F0 variations for Chinese continuous speech tone recognition
Jinsong Zhang 0001, Satoshi Nakamura 0001, Keikichi Hirose |
Speech Commun. | 2 |
| 2005 | Maximum likelihood sub-band adaptation for robust speech recognition
Donglai Zhu, Satoshi Nakamura 0001, Kuldip K. Paliwal, Renhua Wang |
Speech Commun. | 2 |
| 2004 | Automatic generation of non-uniform HMM structures based on variational Bayesian approachabstractWe propose using the variational Bayesian (VB) approach for automatically creating nonuniform, context-dependent HMM topologies in speech recognition. The maximum likelihood (ML) criterion is generally used to create HMM topologies. However, it has an over-fitting problem. Information criteria have been used to overcome this problem, but theoretically they cannot be applied to complicated models like HMM. Recently, to avoid these problems, the VB approach has been developed in the machine-learning field. We introduce the VB approach to the successive state splitting (SSS) algorithm, which can create both contextual and temporal variations for HMM. We define the prior and posterior probability densities and free energy with latent variables as split and stop criteria. Experimental results show that the proposed method can automatically create a more efficient model and obtain better performance, especially for vowels, than the original method. Takatoshi Jitsuhiro, Satoshi Nakamura 0001 |
ICASSP (1) | 2 |
| 2004 | Out-of-domain detection based on confidence measures from multiple topic classificationabstractOne significant problem for spoken language systems is how to cope with users' OOD (out-of-domain) utterances which cannot be handled by the back-end system. In this paper, we propose a novel OOD detection framework, which makes use of classification confidence scores of multiple topics and trains a linear discriminant in-domain verifier using gradient probabilistic descent (GPD). Training is based on deleted interpolation of the in-domain data, and thus does not require actual OOD data, providing high portability. Three topic classification schemes of word N-gram models, latent semantic analysis (LSA), and support vector machines (SVM) are evaluated, and SVM is shown to have the greatest discriminative ability. In an OOD detection task, the proposed approach achieves an absolute reduction in equal error rate (EER) of 6.5% compared to a baseline method based on a simple combination of multiple-topic classifications. Furthermore, comparison with a system trained using OOD data demonstrates that the proposed training scheme realizes comparable performance while requiring no knowledge of the OOD data set. Ian Lane, Tatsuya Kawahara, Tomoko Matsui, Satoshi Nakamura 0001 |
ICASSP (1) | 4 |
| 2004 | Minimum mean square error filtering of noisy cepstral coefficients with applications to ASRabstractIn our previous work (2003), we investigated a new approach to robust speech recognition. An exact procedure was developed to filter noisy cepstral coefficients in the mean-square-error sense, and it was shown that this method outperformed the well known vector Taylor series (VTS) approach, which in turn is based on linear approximations to the non-linear filtering problem. Unfortunately. the procedure presented involved several integral equations with no known closed form solution. Numerical integration techniques were needed, which in turn led to slow performance, and in some cases, numerical problems. In this work we address this problem by using piecewise approximations to the integrands, which in turn yield closed form solutions. The revised procedure is tested on a subset of the Aurora 2 database, and the results are compared with the original numerical integration based approach, as well as VTS. Tor André Myrvoll, Satoshi Nakamura 0001 |
ICASSP (1) | 2 |
| 2004 | Speech recognition for multiple non-native accent groups with speaker-group-dependent acoustic modelsabstractIn this paper, the recognition performance for non-native English speech with two different kinds of speaker-groupdependent acoustic models is investigated. The approaches for creating speaker groups include knowledge-based grouping of non-native speakers by their first language, and the automatic clustering of speakers. Clustering is based on speakerdependent acoustic models in speaker Eigenspace. The acoustic model for each speaker group is obtained by bootstrapping with pre-segmented speech data or adaptation of a speakerindependent native baseline model. For the decoding of a nonnative speaker’s utterance not seen during the training or adaptation phase, the selection of a model suitable to cope with the accent characteristics of that speaker is necessary. Here, ideal selection via an oracle and parallel decoding are examined. Evaluation is conducted in a hotel reservation task for five major accent groups, including German, French, Indonesian, Chinese and Japanese speakers. Recognition results with speakerdependent and an accent-independent non-native model will also be reported. Tobias Cincarek, Rainer Gruhn, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2004 | A statistical lexicon for non-native speech recognition
Rainer Gruhn, Konstantin Markov, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2004 | Increasing the mixture components of non-uniform HMM structures based on a variational Bayesian approachabstractWe propose using the Variational Bayesian (VB) approach for automatically creating non-uniform, context-dependent HMM topologies. Although the Maximum Likelihood (ML) criterion is generally used to create HMM topologies, it has an overfitting problem. Recently, to avoid this problem, the VB approach has been applied to create acoustic models for speech recognition. We introduce the VB approach to the Successive State Splitting (SSS) algorithm, which can create both contextual and temporal variations for HMMs. Experimental results show that the proposed method can automatically create a more efficient model than the original method. Furthermore, we evaluated a method to increase the number of mixture components by using the VB approach and considering temporal structures. The VB approach obtained the best performance with a smaller number of mixture components in comparison with that obtained by using ML based methods. Takatoshi Jitsuhiro, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2004 | Topic classification and verification modeling for out-of-domain utterance detection
Tatsuya Kawahara, Ian Lane, Tomoko Matsui, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2004 | Robust verification of recognized words in noiseabstractIn this paper we investigate robust word verification in noise using the generalized word posterior probability (GWPP). In computing GWPP, reduced search space, relaxed time registrations of hypothesized words in the word graph, and optimal acoustic and language model weights are employed. The sensitivity of word verification errors with respect to the parameters of GWPP was tested under different SNR conditions. We found that around the optimal parameter settings, there exists a relatively stable region where the total number of word verification errors is fairly insensitive (robust) to the exact choice of the optimal values. Cross-SNR condition tests using a large vocabulary, speaker independent, continuous Japanese speech database (Basic Travel Expression Corpus) confirms the robustness of the GWPP based word verification in different SNR’s. 1. Wai Kit Lo, Frank K. Soong, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2004 | Integration of articulatory dynamic parameters in HMM/BN based speech recognition systemabstractIn this paper, we describe several approaches to integra-tion of the articulatory dynamic parameters along with ar-ticulatory position data into a HMM/BN model based au-tomatic speech recognition system. This work is a contin-uation of our previous study, where we have successfully combined speech acoustic features in form of MFCC with articulatory position observations. Articulatory dynamic parameters are represented by velocity and acceleration coefficients calculated as first and second derivatives of the articulatory position data. All these features are in-tegrated using the HMM/BN acoustic model where each feature corresponds to different Bayesian Network vari-able. By changing the BN topology we can change the way articulatory and acoustic parameters are combined. The evaluation experiments showed that the effect of the articulatory dynamic features greatly depends on the BN structure and that careful data analysis is essential in gain-ing knowledge about the underlying dependencies be-tween different information sources. In comparison with conventional HMM system trained on acoustic data only, the HMM/BN system achieved significant improvement of the recognition performance. 1. Konstantin Markov, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2004 | Speech recognition system robust to noise and speaking stylesabstractIt is difficult to recognize speech distorted by various factors, especially when an ASR system contains only a single acoustic model. One solution is to use multiple acoustic models, one model for each different condition. In this paper, we discuss a parallel decoding-based ASR system that is robust to the noise type, SNR, speaker gender and speaking style. Our system consists of two recognition channels based on MFCC and Differential MFCC (DMFCC) features. Each channel has several acoustic models depending on SNR, speaker gender and speaking style, and each acoustic model is adapted by fast noise adaptation. From each channel, one hypothesis is selected based on its likelihood. The final recognition result is obtained by combining hypotheses from the two channels. We evaluate the performance of our system by normal and hyperarticulated test speech data contaminated by various types of noise at different SNR levels. Experiments demonstrate that the system could achieve recognition accuracy in excess of 80% for the normal speaking style data at a SNR of 0 dB. For hyper-articulated speech data, the recognition accuracy improved from about 10% to over 45% compared to a system without acoustic models for hyperarticulated speech. Shigeki Matsuda, Takatoshi Jitsuhiro, Konstantin Markov, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2004 | Online minimum mean square error filtering of noisy cepstral coefficients using a sequential EM algorithm
Tor André Myrvoll, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2004 | Indonesian speech recognition for hearing and speaking impaired peopleabstractThis paper outlines our efforts in developing Indonesian speech recognition for hearing and speaking impaired people. The lack of speech-enabling technology and research, as well as a shortage of data on the Indonesian language presents a major challenge for us to deal with. Difficulties arise in developing an Indonesian speech corpus since Indonesian is actually most people’s second language after their own ethnic native language. Collecting all of the possible languages and dialects of the tribes recognized in Indonesia is still the biggest problem we face. In speech recognition, segmented utterances according to labels are usually used as a starting point for training speech models. This segmentation strategy is also one of the main issues. Initialization training utterances with flat segmentation would not give sufficient performance. Here, we used an English speech recognizer to set initial segmentation of Indonesian utterances. This method produced a significant improvement of up to 40% in performance. Sakriani Sakti, Arry Akhmad Arman, Satoshi Nakamura 0001, Paulus Hutagaol |
INTERSPEECH | 3 |
| 2004 | HMM-based feature compensation method: an evaluation using the AURORA2abstractIn this paper, we describe an HMM-based featurecompensation method. The proposed method compensates for noise-corrupted features in the MFCC domain using the output probability density functions (pdf) of the Hidden Markov Models (HMM). In compensating the features, the output pdfs are adaptively weighted according to forward path probabilities. Because of this, the proposed method can minimize degradation of featurecompensation accuracy due to temporally changing noise environment. We evaluated the proposed method based on the AURORA2 database. All the experiments were conducted in a clean condition. The experiment results indicate that the proposed method, combined with cepstral mean subtraction, can achieve a word accuracy of 87.64%. We also show that the proposed method is useful in a transient pulse noise environment. Akira Sasou, Kazuyo Tanaka, Satoshi Nakamura 0001, Futoshi Asano |
INTERSPEECH | 3 |
| 2004 | Optimal acoustic and language model weights for minimizing word verification errorsabstractGeneralized word posterior probability (GWPP), a confidence measure for verifying recognized words, needs to equalize and weight acoustic and language model likelihood contributions to minimize verification errors. In this study, we investigate the word verification error surface and use it to optimize these weights and the corresponding verification threshold in a development set. We test three different search algorithms for finding the optimal parameters, including: a full grid search, a gradient-based steepest descent search, and a downhill simplex search. The three search methods yield very similar solutions. Proper acoustic and language model weights, especially the ratio between them, changes with the relative importance (reliability) between the two knowledge sources. For a narrow beam width, the role of the acoustic model is less critical than language model in GWPP-based word verification, which is due to the noisy acoustic information maintained in a narrow beam. Using a large vocabulary continuous Japanese speech database (Basic Travel Expression Corpus), the largest relative improvement obtained is 33.2 % for confidence error rate and 38.7 % for a modified word accuracy. 1. Frank K. Soong, Wai Kit Lo, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2004 | Efficient tone classification of speaker independent continuous Chinese speech using anchoring based discriminating featuresabstractAnchoring based discriminating features were proposed efficient for tone discrimination of Chinese continuous speech, and have been successfully applied before to tone classification of speaker dependent experiment. This paper presents its application to speaker independent tone classification experiments. Furthermore, we made detailed comparison experiments on the efficiencies of three groups of features: the left context dependent, the right context dependent anchoring F0 features, and the conventional F0 features. Experimental results showed that a combination of all three groups achieved a significant improvement of absolute 6.4% from 82.6% by the baseline system to 89.0%. When the three groups of features are used individually, both groups of the anchoring features led to better results than the conventional features, and the left context dependent anchoring features led to the highest performance. Jinsong Zhang 0001, Satoshi Nakamura 0001, Keikichi Hirose |
INTERSPEECH | 2 |
| 2004 | Noise adaptive speech recognition based on sequential noise parameter estimation
Kaisheng Yao, Kuldip K. Paliwal, Satoshi Nakamura 0001 |
Speech Commun. | 3 |
| 2004 | Introduction to the Special Issue on Spontaneous Speech Processing
Sadaoki Furui, Mary E. Beckman, Julia Hirschberg, Shuichi Itahashi, Tatsuya Kawahara, Satoshi Nakamura 0001, Shri Narayanan |
IEEE Trans. Speech Audio Process. | 6 |
| 2003 | Hybrid HMM/BN LVCSR system integrating multiple acoustic featuresabstractIn current HMM based speech recognition systems, it is difficult to supplement acoustic spectrum features with additional information such as pitch, gender, articulator positions, etc. On the other hand, dynamic Bayesian networks (DBN) allow for easy combination of different features and make use of conditional dependencies between them. However, lack of efficient algorithms has prevented their application in large vocabulary continuous speech recognition. The hybrid HMM/BN acoustic model, where HMM are used for modeling of temporal speech characteristics and state probability model is represented by BN, provides a trade off solution to the problem. In this paper we describe the HMM/BN acoustic model and LVCSR system built upon this model. In the HMM/BN model, in addition to speech observation variable, state BN has two more discrete variables representing speaker gender and pitch frequency. Evaluation results on WSJ database showed lower word error rate with respect to the same complexity conventional HMM acoustic model when there is enough training data to estimate reliable HMM/BN parameters. Konstantin Markov, Satoshi Nakamura 0001 |
ICASSP (1) | 2 |
| 2003 | An evaluation of adaptive beamformer based on average speech spectrum for noisy speech recognitionabstractDistant-talking speech recognition in noisy environments is indispensable for self-moving robots or teleconference systems. However, background noise and room reverberations seriously degrade the sound-capture quality in real acoustic environments. A microphone array is an ideal candidate as an effective method for capturing distant-talking speech. AMNOR (Adaptive Microphone-array for NOise Reduction) was proposed as an adaptive beamformer for capturing the desired distant signals in noisy environments by Y. Kaneda and J. Ohga (see IEEE Trans. Acoust. Speech Sig. Process., vol.ASSP-34, no.6, p.1391-1400, 1986). Although the AMNOR has been proven effective, it can be further improved if we know the spectrum characteristics of the desired distant signals in advance. Regarding speech as a desired distant signal, we have designed an AMNOR based on the average speech spectrum. We particularly focus on the performance of the proposed AMNOR for distant-talking speech capture and recognition. Evaluation experiments in real acoustic environments confirm that ASR (automatic speech recognition) performance in noisy environments is improved by 5-10% using our AMNOR. In addition, the proposed AMNOR provides better noise reduction performance than that of the conventional AMNOR. Takanobu Nishiura, Masato Nakayama, Satoshi Nakamura 0001 |
ICASSP (1) | 3 |
| 2003 | A multilevel framework to model the inherently confounding nature of sentential F0sentential F0 contours contours for recognizing Chinese lexical tonesabstractThis paper presents a multilevel framework to cope with the complex variations in Chinese sentential F0 contours in order to recognize lexical tones. Tone nucleus model is to get rid of the influence of intrinsic F0 transition loci at sub-syllable level. The pitch anchoring concept is used to normalize tonal F0 contours at syllable level. The hypo- and hyper-intonation model is used to account for the interplay of tone coarticulation and higher level prosodic effects. The whole approach achieved significant higher performance than the conventional method. Jinsong Zhang 0001, Keikichi Hirose, Satoshi Nakamura 0001 |
ICASSP (1) | 3 |
| 2003 | An evaluation of adaptive beamformer based on average speech spectrum for noisy speech recognitionabstractDistant-talking speech recognition in noisy environments is indispensable for self-moving robots or tele-conference systems. However, background noise and room reverberations seriously degrade the sound-capture quality in real acoustic environments. A microphone array is an ideal candidate as an effective method for capturing distant-talking speech. AMNOR (adaptive microphone-array for noise reduction) was proposed as an adaptive beamformer for capturing the desired distant signals in noisy environments by Kaneda et al. Although the AMNOR has been proven effective, it can be further improved if we know spectrum characteristics of the desired distant signals in advance. Therefore, we regarded speech as a desired distant signal and designed an AMNOR based on the average speech spectrum. In this paper, we particularly focused on the performance of AMNOR based on the average speech spectrum for distant-talking speech capture and recognition. As a result of evaluation experiments in real acoustic environments, we confirmed that the ASR (automatic speech recognition) performance was improved 5-10% by using AMNOR based on the average speech spectrum in noisy environments. In addition, the proposed AMNOR provides better noise reduction performance than that of conventional AMNOR. Takanobu Nishiura, Masato Nakayama, Satoshi Nakamura 0001 |
ICME | 3 |
| 2003 | Detection and separation of speech segment using audio and video information fusion
Futoshi Asano, Yoichi Motomura, Hideki Asoh, Takashi Yoshimura, Naoyuki Ichimura, Kiyoshi Yamamoto, Nobuhiko Kitawaki, Satoshi Nakamura 0001 |
INTERSPEECH | 8 |
| 2003 | Missing feature theory applied to robust speech recognition over IP network
Toshiki Endo, Shingo Kuroiwa, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2003 | A semi-blind source separation method for hands-free speech recognition of multiple talkersabstractIn this paper, we present a beamforming based semi-blind source separation technique, which can be applied efficiently for hands-free speech recognition of multiple talkers (including moving talkers, too). The main difference from the conventional blind source separation techniques lies in the fact that the proposed method does not attempt to separate explicitly the unknown signals in a pre-processing pass before speech recognition. In fact, localization of multiple talkers, separation of the signals, and speech recognition are integrated in a single pass. Each time frame, beams formed by a delay-and-sum beamformer are steered to every direction, and speech information is extracted. A modified Viterbi formula provides n-best hypotheses for each direction and word hypotheses. At the final frame, all hypotheses are clustered based on their direction information. The clusters, which correspond to the talkers include information about the recognized speech of the multiple talkers and about their direction. Experiments for recognition of two and three talkers showed very promising results. In the case of two talkers, and using simulated clean data we achieved for `top 5' hypotheses a recognition rate of series 95.02% on average, which is very promising result. Panikos Heracleous, Satoshi Nakamura 0001, Kiyohiro Shikano |
INTERSPEECH | 2 |
| 2003 | Automatic generation of non-uniform context-dependent HMM topologies based on the MDL criterion
Takatoshi Jitsuhiro, Tomoko Matsui, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2003 | Hierarchical topic classification for dialog speech recognition based on language model switchingabstractA speech recognition architecture combining topic detection and topic-dependent language modeling is proposed. In this architecture, a hierarchical back-off mechanism is introduced to improve system robustness. Detailed topic models are applied when topic detection is confident, and wider models that cover multiple topics are applied in cases of uncertainty. In this paper, two topic detection methods are evaluated for the architecture: unigram likelihood and SVM (Support Vector Machine). On the ATR Basic Travel Expression corpus, both topic detection methods provide a comparable reduction in WER of 10.0% and 11.1 % respectively over a single language model system. Finally the proposed re-decoding approach is compared with an equivalent system based on re-scoring. It is shown that redecoding is vital to provide optimal recognition performance. 1. Ian Lane, Tatsuya Kawahara, Tomoko Matsui, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2003 | Hybrid HMM/BN ASR system integrating spectrum and articulatory featuresabstractIn this paper, we describe automatic speech recognition system where features extracted from human speech production system in form of articulatory movements data are effectively integrated in the acoustic model for improved recognition performance. The system is based on the hybrid HMM/BN model, which allows for easy integration of different speech features by modeling probabilistic dependencies between them. In addition, features like articulatory movements, which are difficult or impossible to obtain during recognition, can be left hidden, in fact eliminating the need of their extraction. The system was evaluated in phoneme recognition task on small database consisting of three speakers’ data in speaker dependent and multi-speaker modes. In both cases, we obtained higher recognition rates compared to conventional, spectrum based HMM system with the same number of parameters. Konstantin Markov, Yosuke Iizuka, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2003 | Noise reduction using paired-microphones on non-equally-spaced microphone arrangement
Mitsunori Mizumachi, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2003 | Environmental sound source identification based on hidden Markov model for robust speech recognitionabstractIn real acoustic environments, humans communicate with each other through speech by focusing on the target speech among environmental sounds. We can easily identify the target sound from other environmental sounds. For hands-free speech recognition, the identification of the target speech from environmental sounds is imperative. This mechanism may also be important for a self-moving robot to sense the acoustic environments and communicate with humans. Therefore, this paper first proposes Hidden Markov Model (HMM)-based environmental sound source identification. Environmental sounds are modeled by three states of HMMs and evaluated using 92 kinds of environmental sounds. The identification accuracy was 95.4%. This paper also proposes a new HMM composition method that composes speech HMMs and an HMM of categorized environmental sounds for robust environmental sound-added speech recognition. As a result of the evaluation experiments, we confirmed that the proposed HMM composition outperforms the conventional HMM composition with speech HMMs and a noise (environmental sound) HMM trained using noise periods prior to the target speech in a captured signal. Takanobu Nishiura, Satoshi Nakamura 0001, Kazuhiro Miki, Kiyohiro Shikano |
INTERSPEECH | 2 |
| 2003 | Adaptation of acoustic model using the gain-adapted HMM decomposition methodabstractIn a real environment, it is essential to adapt an acoustic model to variations in background noises in order to realize robust speech recognition. In this paper, we construct an extended acoustic model by combining a mismatch model with a clean acoustic model trained using only clean speech. We assume the mismatch model conforms to a Gaussian distribution with timevarying population parameters. The proposed method adapts the extended acoustic model to the noises by estimating the population parameters using a Gaussian Mixture Model (GMM) and Gain-Adapted Hidden Markov Model (GA-HMM) decomposition method. We performed recognition experiments under noisy conditions using the AURORA2 database in order to confirm the effectiveness of the proposed method. Akira Sasou, Futoshi Asano, Kazuyo Tanaka, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2003 | Integration of noise reduction algorithms for Aurora2 taskabstractTo achieve high recognition performance for a wide variety of noise and for a wide range of signal-to-noise ratios, this paper presents the integration of four noise reduction algorithms: spectral subtraction with smoothing of time direction, temporal domain SVD-based speech enhancement, GMM-based speech estimation and KLT-based comb-filtering. Recognition results on the Aurora2 task show that the effectiveness of these algorithms and their combinations strongly depends on noise conditions, and excessive noise reduction tends to degrade recognition performance in multicondition training. Takeshi Yamada, Jiro Okada, Kazuya Takeda, Norihide Kitaoka, Masakiyo Fujimoto, Shingo Kuroiwa, Kazumasa Yamamoto, Takanobu Nishiura, Mitsunori Mizumachi, Satoshi Nakamura 0001 |
INTERSPEECH | 10 |
| 2003 | Model based noisy speech recognition with environment parameters estimated by noise adaptive speech recognition with priorabstractWe have proposed earlier a noise adaptive speech recognition approach for recognizing speech corrupted by nonstationary noise and channel distortion. In this paper, we extend this approach. Instead of maximum likelihood estimation of environment parameters (as done in our previous work), the present method estimates environment parameters within the Bayesian framework that is capable of incorporating prior knowledge of the environment. Experiments are conducted on a database that contains digit utterances contaminated by channel distortion and nonstationary noise. Results show that this method performs better than the previous methods. Kaisheng Yao, Kuldip K. Paliwal, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2003 | Maximum likelihood sub-band weighting for robust speech recognitionabstractSub-band speech recognition approaches have been proposed for robust speech recognition, where full-band power spectra are divided into several sub-bands and then likelihoods or cepstral vectors of the sub-bands are merged depending on their reliability. In conventional sub-band approaches, correlations across the sub-bands are not modeled and the merging weights can only be set experientially or estimated during training procedures, which may not match observed data. The methods further degrade performance for clean speech. We proposed a novel sub-band approach, where frequency sub-bands are multiplied with weighting factors and merged, which considers sub-band dependence and proves to be more robust than both full-band and conventional sub-band approaches. And further the weighting factors can be obtained by using the maximum-likelihood estimation approaches in order to minimize the mismatch between the trained models and the observed features. Finally we evaluated our methods on both the Aurora2 task and the Resource Management task and showed improvement of performance on the two tasks consistently. 1. Donglai Zhu, Satoshi Nakamura 0001, Kuldip K. Paliwal, Renhua Wang |
INTERSPEECH | 2 |
| 2003 | Model-based talking face synthesis for anthropomorphic spoken dialog agent systemabstractTowards natural human-machine communication, interface technologies by way of speech and image information have been intensively developed. An anthropomorphic dialog agent is an ideal system, which integrates spoken dialog and natural facial expressions. This paper reports on our project aiming to create a general-purpose toolkit for building an easily customizable anthropomorphic agent. There have been almost no tools so far such as intuitive, easy to understand, fully interactive, and open source. Our anthropomorphic agent is designed to fulfill these requirements. This toolkit consists four modules, multi modal dialog integration, speech recognition, speech synthesis, and face image synthesis. These modules are highly modularized and interlinked by a simple communication protocols.In this paper, we focus on the construction of an agent's face image synthesis. For this part lip movement control synchronous to the speech signal and facial emotion expression are the most important parts. We developed the face image synthesis module (FSM) that only requires one frontal face image, and can be used by any skill level of users. A user's original agent can be generated by easy adjustment of the frontal face image and the generic wire-frame model. The paper describes overall system diagram and specifically the agent's face image synthesis part. Tatsuo Yotsukura, Shigeo Morishima, Satoshi Nakamura 0001 |
ACM Multimedia | 3 |
| 2003 | Cepstrum derived from differentiated power spectrum for robust speech recognition
Jingdong Chen, Kuldip K. Paliwal, Satoshi Nakamura 0001 |
Speech Commun. | 3 |
| 2002 | Audio-visual speech translation with automatic lip syncqronization and face tracking based on 3-D head modelabstractSpeech-to-speech translation has been studied to realize natural human communication beyond language barriers. Toward further multi-modal natural communication, visual information such as face and lip movements will be necessary. In this paper, we introduce a multi-modal English-to-Japanese and Japanese-to-English translation system that also translates the speaker's speech motion while synchronizing it to the translated speech. To retain the speaker's facial expression, we substitute only the speech organ's image with the synthesized one, which is made by a three-dimensional wire-frame model that is adaptable to any speaker. Our approach enables image synthesis and translation with an extremely small database. We conduct subjective evaluation by connected digit discrimination using data with and without audiovisual lip-synchronicity. The results confirm the sufficient quality of the proposed audio-visual translation system. Shigeo Morishima, Shin Ogata, Kazumasa Murai, Satoshi Nakamura 0001 |
ICASSP | 4 |
| 2002 | Robust bi-modal speech recognition based on state synchronous modeling and stream weight optimizationabstractThere have been higher demands recently for Automatic Speech Recognition (ASR) systems able to operate robustly in acoustically noisy environments. This paper proposes a method to effectively integrate audio and visual information in audio-visual (bi-modal) ASR systems. Such integration inevitably necessitates modeling of the synchronization of the audio and visual information. To address the time lag and correlation problems in individual features between speech and lip movements, we introduce a type of integrated HMM modeling of audio-visual information based on a family of HMM composition. The proposed model can represent state synchronicity not only within a phoneme but also between phonemes. Furthermore, we also propose a rapid stream weight optimization based on GPD algorithm for noisy bi-modal speech recognition. Evaluation experiments show that the proposed method improves the recognition accuracy for noisy speech. In SNR=0dB our proposed method attained 16% higher performance compared to a product HMMs without the synchronicity re-estimation. Satoshi Nakamura 0001, Ken'ichi Kumatani, Satoshi Tamura |
ICASSP | 1 |
| 2002 | Talker localization in a real acoustic environment based on DOA estimation and statistical sound source identificationabstractFor a hands-free speech interface, it is very important to capture distant talking speech with high quality. A microphone array is an ideal candidate for this purpose. However, this approach requires localizing the target talker. Conventional talker localization algorithms in multiple sound source environments not only have difficulty localizing the multiple sound sources accurately, but also have difficulty localizing the target talker among known multiple sound source positions. To cope with these problems, we propose a new talker localization algorithm consisting of two algorithms. One is DOA (Direction Of Arrival) estimation algorithm for multiple sound source localization based on CSP (Cross-power Spectrum Phase) coefficient addition method. The other is statistical sound source identification algorithm based on GMM (Gaussian Mixture Model) for localizing the target talker position among localized multiple sound sources. In this paper, we particularly focus on the talker localization performance based on the combination of these two algorithms with a microphone array. Takanobu Nishiura, Satoshi Nakamura 0001, Kiyohiro Shikano |
ICASSP | 2 |
| 2002 | Noise adaptive speech recognition in time-varying noise based on sequential kullback proximal algorithmabstractWe present a noise adaptive speech recognition approach, where time-varying noise parameter estimation and Viterbi process are combined together. The Viterbi process provides approximated joint likelihood of active partial paths and observation sequence given the noise parameter sequence estimated till previous frame. The joint likelihood after normalization provides approximation to the posterior probabilities of state sequences for an EM-type recursive process based on sequential Kullback proximal algorithm to estimate the current noise parameter. The combined process can easily be applied to perform continuous speech recognition in presence of non-stationary noise. Experiments were conducted in simulated and real non-stationary noises. Results showed that the noise adaptive system provides significant improvements in word accuracy as compared to the baseline system (without noise compensation) and the normal noise compensation system (which assumes the noise to be stationary). Kaisheng Yao, Kuldip K. Paliwal, Satoshi Nakamura 0001 |
ICASSP | 3 |
| 2002 | Real time face detection for multimodal speech recognitionabstractWe propose a real time system to detect the speaker's frontal face for multimodal speech recognition. It is widely acknowledged that automatic speech recognizers, as well as humans, can improve recognition performance by adding visual modality, i.e., the speaker's facial image to audio modality. Visual modality also provides inaudible information, such as the speaker's facial orientation, and the location of the mouth. To acquire this information, we have to localize the speaker's face in real time. Our system is a combination of skin color detection and spatial feature detection. The color-based detection is fast but depends on the skin and the background color, while the special feature detection requires more computation. We applied color-based pruning to reduce the search space for the spatial feature detection. By detecting the facial orientation, the proposed method functions as a "face to talk" switch in place of the "push to talk" switch. In our experiment, pruning based on color reduced 53-97% of the search space, and 98.9% of the frontal face was detected correctly by the subsequent spatial detector. Kazumasa Murai, Satoshi Nakamura 0001 |
ICME (2) | 2 |
| 2002 | Design and collection of acoustic sound data for hands-free speech recognition and sound scene understandingabstractThe sound data for open evaluation is necessary for studies such as sound source localization, sound retrieval, sound recognition and hands-free speech recognition in real acoustic environments. This paper reports on our project for acoustic data collection. There are many kinds of sound scenes in real environments. The sound scene is specified by sound sources and room acoustics. The number of combinations of the sound sources, source positions and rooms is huge in real acoustic environments. We assumed that the sound in the environments can be simulated by convolution of the isolated sound sources and impulse responses. As an isolated sound source, hundred kinds of environment sounds and speech sounds are collected. The impulse responses are collected in various acoustic environments. Additionally we collected sounds from a moving source. In this paper, progress of our sound scene database collection project and application to environment sound recognition and hands-free speech recognition are described. Satoshi Nakamura 0001, Kazuo Hiyane, Futoshi Asano, Yutaka Kaneda, Takeshi Yamada, Takanobu Nishiura, Tetsunori Kobayashi, Shiro Ise, Hiroshi Saruwatari |
ICME (2) | 1 |
| 2002 | An evaluation of sound source identification with RWCP sound scene database in real acoustic environmentsabstractIt is very important for a hands-free speech interface to capture distant speech with high quality. A microphone array is an ideal candidate for this purpose. However, this approach requires localizing the target talker. Conventional talker localization methods in multiple sound source environments not only have difficulty localizing the multiple sound sources accurately, but also have difficulty localizing the target talker among known multiple sound source positions. To cope with these problems, we propose a new talker localization method consisting of two algorithms. One algorithm is for multiple sound source localization based on CSP (cross-power spectrum phase) analysis. The other algorithm is for sound source identification among localized multiple sound sources towards talker localization. We particularly focus on the latter statistical sound source identification among localized multiple sound sources with statistical speech and environmental sound models based on GMMs (Gaussian mixture models) and a microphone array towards talker localization. We especially evaluate the performance of the proposed algorithms with the RWCP sound scene database in real acoustic environments (RWCP-DB). Takanobu Nishiura, Satoshi Nakamura 0001 |
ICME (2) | 2 |
| 2002 | Multi-Modal Translation System and Its EvaluationabstractSpeech-to-speech translation has been studied to realize natural human communication beyond language barriers. Toward further multi-modal natural communication, visual information such as face and lip movements will be necessary. We introduce a multi-modal English-to-Japanese and Japanese-to-English translation system that also translates the speaker's speech motion while synchronizing it to the translated speech. To retain the speaker's facial expression, we substitute only the speech organ's image with the synthesized one, which is made by a three-dimensional wire-frame model that is adaptable to any speaker. Our approach enables image synthesis and translation with an extremely small database. We conduct subjective evaluation tests using the connected digit discrimination test using data with and without audio-visual lip-synchronization. The results confirm the significant quality of the proposed audio-visual translation system and the importance of lip-synchronization. Shigeo Morishima, Satoshi Nakamura 0001 |
ICMI | 2 |
| 2002 | 3-D N-Best Search for Simultaneous Recognition of Distant-Talking Speech of Multiple TalkersabstractA microphone array is a promising solution for realizing hands-free speech recognition in real environments. Accurate talker localization is very important for speech recognition using the microphone array. However, localization of a moving talker is difficult in noisy reverberant environments. Talker localization errors degrade the performance of speech recognition. To solve the problem, we proposed a new speech recognition algorithm which considers multiple talker direction hypotheses simultaneously. The proposed algorithm performs Viterbi search in 3-dimensional trellis space composed of talker directions, input frames, and HMM states. In this paper we describe a new simultaneous recognition algorithm for distant-talking speech of multiple talkers using the extended 3D N-best search algorithm. The algorithm exploits path distance-based clustering and a likelihood normalization technique appeared to be necessary in order to build an efficient system for our purpose. We evaluated the proposed method using reverberated data, which are those simulated by the image method and recorded in a real room. The image method was used to find the accuracy-reverberation time relationship, and real data was used to evaluate the real performance of our algorithm. The Top 3 result of simultaneous word accuracy was 73.02% under 162 ms reverberation time using the image method. Satoshi Nakamura 0001, Panikos Heracleous |
ICMI | 1 |
| 2002 | Multi-Modal Temporal Asynchronicity Modeling by Product HMMs for RobustabstractThe demand for audio-visual speech recognition (AVSR) has increased in order to make speech recognition systems robust to acoustic noise. There are two kinds of research issue in audio-visual speech recognition, such as integration modeling considering asynchronicity between modalities and adaptive information weighting according information reliability. This paper proposes a method to effectively integrate audio and visual information. Such integration, inevitably, necessitates modeling the synchronization and asynchronization of audio and visual information. To address the time lag and correlation problems in individual features between speech and lip movements, we introduce a type of integrated HMM modeling of audio-visual information based on a family of a product HMM. The proposed model can represent state synchronicity not only within a phoneme, but also between phonemes. Furthermore, we also propose a rapid stream weight optimization based on the GPD algorithm for noisy, bimodal speech recognition. Evaluation experiments show that the proposed method improves the recognition accuracy for noisy speech. When SNR=0 dB our proposed method attained 16% higher performance compared to a product HMM without synchronicity re-estimation. Satoshi Nakamura 0001, Ken'ichi Kumatani, Satoshi Tamura |
ICMI | 1 |
| 2002 | Weighted graph based decision tree optimization for high accuracy acoustic modeling
Jinsong Zhang 0001, Satoshi Nakamura 0001, Chin-Hui Lee 0001, Tat-Seng Chua |
INTERSPEECH | 3 |
| 2002 | HMM COmposition-based rapid model adaptation using a priori noise GMM adaptation evaluation on Aurora2 corpus
Masaki Ida, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2002 | Modeling HMM state distributions with Bayesian networks
Konstantin Markov, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2002 | The 2ch hybrid subtractive beamformer applied to line sound sources
Mitsunori Mizumachi, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2002 | Suitable design of adaptive beamformer based on average speech spectrum for noisy speech recognitionabstractRecognition of distant-talking speech is indispensable for self-moving robots or teleconference systems. However, background noise and room reverberations seriously degrade the sound capture quality in real acoustic environments. A microphone array is an ideal candidate as an effective method for capturing distant-talking speech. AMNOR (Adaptive Microphone-array for NOise Reduction) was proposed an adaptive beamformer for capturing the desired distant signals in noisy environments by Kaneda et al. Although AMNOR has proven itself effective, it could be further improved if we knew the spectrum characteristics of desired distant signals in advance. Therefore, in this paper we regard speech as a desired distant signal and design AMNOR based on the average speech spectrum for distant-talking speech capture and recognition. As a result of evaluation experiments in real acoustic environments, we could confirm that the ASR (Automatic Speech Recognition) performance was improved 5 ~ 10% by AMNOR based on average speech spectrum in noisy environments. Takanobu Nishiura, Satoshi Nakamura 0001, Yuka Okada, Takeshi Yamada, Kiyohiro Shikano |
INTERSPEECH | 2 |
| 2002 | Speaking rate compensation based on likelihood criterion in acoustic model training and decoding
Kozo Okuda, Tatsuya Kawahara, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2002 | Noise adaptive speech recognition with acoustic models trained from noisy speech evaluated on Aurora-2 databaseabstractIn this paper, we apply the noise adaptive speech recognition for noisy speech recognition in non-stationary noise to the situation that acoustic models are trained from noisy speech. We justify it by that the noise adaptive speech recognition includes iterative processes between a noise parameter estimation step and a model adaptation step, which can possibly do non-linear mapping between the original training space and that for recognition. Experiments were performed onAurora-2 task with multi-conditional training set which includes noisy utterances. Through experiments, we observed that the noise adaptive speech recognition can have better performance than the baseline system trained from multiconditional training set without noise adaptive speech recognition. 1. Kaisheng Yao, Kuldip K. Paliwal, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2002 | Evaluation of a noise adaptive speech recognition system on the Aurora 3 database
Kaisheng Yao, Donglai Zhu, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2002 | Modeling varying pauses to develop robust acoustic models for recognizing noisy conversational speech
Jinsong Zhang 0001, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2002 | The Present Status of Speech Database in Japan: Development, Management, and Application to Speech Research
Hisao Kuwabara, Shuichi Itahashi, Mikio Yamamoto, Toshiyuki Takezawa, Satoshi Nakamura 0001, Kazuya Takeda |
LREC | 5 |
| 2002 | Distant-talking speech recognition based on a 3-D Viterbi search using a microphone arrayabstractThis paper focuses on microphone arrays to realize distant-talking speech recognition in real environments. In distant-talking situations, users can speak at arbitrary positions while moving. Therefore, it,is very important for high quality speech acquisition using microphone arrays to localize a talker accurately. However, it is very difficult to localize a moving talker in noisy and reverberant environments. The talker localization errors result in performance degradation of speech recognition. One way to solve this problem is to integrate the speech recognition process and the talker localization into a unified framework. This paper proposes a new speech recognition algorithm based on a three-dimensional (3-D) Viterbi search. The 3-D Viterbi method extracts a direction-time sequence of parameter vectors by steering a beam to every direction in every frame, then finds the most likely path in a 3-D trellis space composed of talker directions, input frames and HMM states. This means that speech recognition and talker localization are performed simultaneously within a statistical framework. To evaluate the performance of the 3-D Viterbi method, recognition experiments for real environment data were carried out. The results confirmed that the 3-D Viterbi method drastically improves the recognition performance for the moving talker case as well as for the fixed-position talker case. Takeshi Yamada, Satoshi Nakamura 0001, Kiyohiro Shikano |
IEEE Trans. Speech Audio Process. | 2 |
| 2002 | Statistical multimodal integration for audio-visual speech processingabstractSensory information is indispensable for living things. It is also important for living things to integrate multiple types of senses to understand their surroundings. In human communications, human beings must further integrate the multimodal senses of audition and vision to understand intention. In this paper, we describe speech related modalities since speech is the most important media to transmit human intention. To date, there have been a lot of studies concerning technologies in speech communications, but performance levels still have room for improvement. For instance, although speech recognition has achieved remarkable progress, the speech recognition performance still seriously degrades in acoustically adverse environments. On the other hand, perceptual research has proved the existence of the complementary integration of audio speech and visual face movements in human perception mechanisms. Such research has stimulated attempts to apply visual face information to speech recognition and synthesis. This paper introduces works on audio-visual speech recognition, speech to lip movement mapping for audio-visual speech synthesis, and audio-visual speech translation. Satoshi Nakamura 0001 |
IEEE Trans. Neural Networks | 1 |
| 2001 | A microphone array-based 3-D N-best search algorithm for the simultaneous recognition of multiple sound sources in real environmentsabstractDeals with the recognition of distant talking speech and, particularly, with the simultaneous recognition of multiple sound sources. A problem that must be solved in the recognition of distant talking speech is talker localization. In some approaches, the talker is localized by using short- and long-term power. The 3-D Viterbi search based method proposed by Yamada et al.(1998), integrates talker localization and speech recognition. This method provides high recognition rates but its application is restricted to the presence of one talker. In order to deal with multiple talkers, we extended the 3-D Viterbi search method to a 3-D N-best search method enabling the recognition of multiple sound sources. The paper describes our baseline 3-D N-best search-based system and two additional techniques, namely, a likelihood normalization technique and a path distance-based clustering technique. The paper also describes experiments carried out in order to evaluate the performance of the system. Panikos Heracleous, Satoshi Nakamura 0001, Kiyohiro Shikano |
ICASSP | 2 |
| 2001 | Discriminative training of HMM using maximum normalized likelihood algorithmabstractWe present the maximum normalized likelihood estimation (MNLE) algorithm and its application for discriminative training of hidden Markov models (HMMs) for continuous speech recognition. The objective of this algorithm is to maximize the normalized frame likelihood of training data. Instead of gradient descent techniques usually applied for objective function optimization in other discriminative algorithms such as the minimum classification error (MCE) and maximum mutual information (MMI), we used a modified expectation-maximization (EM) algorithm which greatly simplifies and speeds up the training procedure. Evaluation experiments showed better recognition rates compared, to both the maximum likelihood (ML) training method and MCE/GPD discriminative method. In addition, the MNLE algorithm showed better generalization abilities and was faster than MCE/GPD. Konstantin Markov, Seiichi Nakagawa, Satoshi Nakamura 0001 |
ICASSP | 3 |
| 2001 | An Adaptive Integration Based On Product Hmm For Audio-Visual Speech RecognitionabstractThere have been higher demands recently for Automatic Speech Recognition (ASR) systems able to operate robustly in acoustically noisy environments. This paper proposes a method to effectively integrate audio and visual information in audiovisual (bi-modal) ASR systems. For such integration, the following issues are important: (1) The synchronization of the audio and visual information, and (2) The optimization of a system in its environment. In (1), the individual feature of the speech and lip movements has the time lag, and has the correlation. To address this problem, we introduce an integration method using HMM composition. In (2), we examine whether the GPD algorithm can adaptively estimate the stream weights. Evaluation experiments show that the proposed method improves the recognition accuracy for noisy speech. Ken'ichi Kumatani, Satoshi Nakamura 0001, Kiyohiro Shikano |
ICME | 2 |
| 2001 | Automatic Face Tracking And Model Match-Move In Video Sequence Using 3d Face Model
Takafumi Misawa, Kazumasa Murai, Satoshi Nakamura 0001, Shigeo Morishima |
ICME | 3 |
| 2001 | Trends of Learning Technology Standard
Shigeo Morishima, Shin Ogata, Satoshi Nakamura 0001 |
ICME | 3 |
| 2001 | Speech Detection By Facial Image For Multimodal Speech Recognition
Kazumasa Murai, Ken'ichi Kumatani, Satoshi Nakamura 0001 |
ICME | 3 |
| 2001 | Automatic Steering Of Microphone Array And Video Camera Toward Multi-Lingual Tele-Conference Through Speech-To-Speech Translation
Takanobu Nishiura, Rainer Gruhn, Satoshi Nakamura 0001 |
ICME | 3 |
| 2001 | Model-Based Lip Synchronization With Automatically Translated Systhetic Voice Toward A Multi-Modal Translation SystemabstractIn this paper, we introduce a multi-modal English-to-Japanese and Japanese-to-English translation system that also translates the speaker's speech motion while synchronizing it to the translated speech. To retain the speaker's facial expression, we substitute only the speech organ's image with the synthesized one, which is made by a three-dimensional wire-frame model that is adaptable to any speaker. Our approach enables image synthesis and translation with an extremely small database. Shin Ogata, Kazumasa Murai, Satoshi Nakamura 0001, Shigeo Morishima |
ICME | 3 |
| 2001 | Sub-band based additive noise removal for robust speech recognitionabstractGriffith Sciences, Griffith School of Engineering Jingdong Chen, Kuldip K. Paliwal, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2001 | Noise reduction using paired-microphones for both far-field and near-field sound sources
Mitsunori Mizumachi, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2001 | Statistical sound source identification in a real acoustic environment for robust speech recognition using a microphone arrayabstractIt is very important for a hands-free speech interface to capture distant talking speech with high quality. A microphone array is an ideal candidate for this purpose. However, this approach requires localizing the target talker. To cope with this problem, we propose a new talker localization method consisting of two algorithms. One algorithm is for multiple sound source localization based on CSP (Cross-power Spectrum Phase) analysis. The other algorithm is for sound source identification among localized multiple sound sources towards talker localization. In this paper, we particularly focus on the latter statistical sound source identification among localized multiple sound sources with statistical speech and environmental sound models based on GMMs (Gaussian Mixture Models) and a microphone array towards talker localization. Takanobu Nishiura, Satoshi Nakamura 0001, Kiyohiro Shikano |
INTERSPEECH | 2 |
| 2001 | Towards the creation of acoustic models for stressed Japanese speech
Kozo Okuda, Tomoko Matsui, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2001 | Feature extraction and model-based noise compensation for noisy speech recognition evaluated on AURORA 2 taskabstractNo Full Text Kaisheng Yao, Jingdong Chen, Kuldip K. Paliwal, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2001 | Sequential noise compensation by a sequential kullback proximal algorithmabstractNo Full Text Kaisheng Yao, Kuldip K. Paliwal, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2001 | A hybrid approach to enhance task portability of acoustic models in Chinese speech recognitionabstractThis paper presents our approach to enhance the portability of acoustic models by mitigating the phonetic mismatch arising from a new testing task which is rather different from the training data. The approach is a hybrid one which combines knowledge-based context categorization to generate a context rich set of subword units, and data-driven-based acoustic model clustering on the level of context category. Compared with the conventional approach of only phonetic decision tree based model clustering and unseen model generation, the new approach improved greatly the desired subword coverage for the new testing domain, and achieved an error rate reduction by 10.8% for Chinese character accuracy in the recognition experiments. Together with the effect of the newly adopted basic units of 9 glottal stops, we achieved a total 23.5% error rate reduction in the testing compared to the baseline system. Jinsong Zhang 0001, Shuwu Zhang, Yoshinori Sagisaka, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2001 | HMM-separation-based speech recognition for a distant moving speakerabstractThis paper presents a hands-free speech recognition method based on HMM composition and separation for speech contaminated not only by additive noise but also by an acoustic transfer function. The method realizes an improved user interface such that a user is not encumbered by microphone equipment in noisy and reverberant environments. The use of HMM composition has already been proposed for countering additive noise. In this paper, the same approach is extended to handle convolutional acoustic distortion in a reverberant room, by using an HMM to model the acoustic transfer function. The states of this HMM correspond to different positions of the sound source. It can represent the positions of the sound sources, even if the speaker moves. This paper also proposes a new method, HMM separation, for estimating the HMM parameters of the acoustic transfer function on the basis of a maximum likelihood manner. The proposed method is obtained through the reverse of the process of HMM composition, where the model parameters are estimated by maximizing the likelihood of adaptation data uttered from an unknown position. Therefore, measurement of impulse responses is not required. The paper also describes the performance of the proposed methods for recognizing real distant-talking speech. The results of experiments clarify the effectiveness of the proposed method. Tetsuya Takiguchi, Satoshi Nakamura 0001, Kiyohiro Shikano |
IEEE Trans. Speech Audio Process. | 2 |
| 2000 | Localization of multiple sound sources based on a CSP analysis with a microphone arrayabstractAccurate localization of multiple sound sources is indispensable for the microphone array-based high quality sound capture. For single sound source localization, the CSP (cross-power spectrum phase analysis) method has been proposed. The CSP method localizes a sound source as a crossing point of sound directions estimated using different microphone pairs. However, when localizing multiple sound sources, the CSP method has a problem that the localization accuracy is degraded due to cross-correlation among different sound sources. To solve this problem, this paper proposes a new method which suppresses the undesired cross-correlation by synchronous addition of CSP coefficients derived from multiple microphone pairs. Experiment results in a real room showed that the proposed method improves the localization accuracy when increasing the number of the synchronous addition. Takanobu Nishiura, Takeshi Yamada, Satoshi Nakamura 0001, Kiyohiro Shikano |
ICASSP | 3 |
| 2000 | Speech recognition for a distant moving speaker based on HMM composition and separationabstractThis paper describes a hands-free speech recognition method based on HMM composition and separation for speech contaminated not only by additive noise but also by an acoustic transfer function. The method realizes an improved user interface such that a user is not encumbered by microphone equipment in noisy and reverberant environments. In this approach, an attempt is made to model acoustic transfer functions by means of an ergodic HMM. The states of this HMM correspond to different positions of the sound source. It can represent the positions of the sound sources, even if the speaker moves. The HMM parameters of the acoustic transfer function are estimated by HMM separation. The method is obtained through the reverse of the process of HMM composition, where the model parameters are estimated by maximizing the likelihood of adaptation data uttered from an unknown position. Therefore, measurement of impulse responses is not required. In this paper, we record the speech of a distant moving speaker in real environments. The results of experiments for the speech of a distant moving speaker clarified the effectiveness of HMM composition and separation. Tetsuya Takiguchi, Satoshi Nakamura 0001, Kiyohiro Shikano |
ICASSP | 2 |
| 2000 | Robust fundamental frequency estimation using instantaneous frequencies of harmonic componentsabstractThis paper proposes a noise-tolerant method for fundamental frequency (F0) extraction. This method includes several new ideas, including the estimation of the instantaneous frequencies of the higher harmonic components, and the design of an adaptive weighting function based on a bandwidth equation that combines the F0 information in the harmonic components. To evaluate the proposed method, we constructed a relatively large database of simultaneous recordings of speech waveforms and EGG (Electro Glotto Graphy). The database consists of 30 sentences pronounced by 14 male and 14 female normal subjects, i.e., 840 sentences in total. The duration of the sound is about 35 minutes including about 20 minutes of voicing. The experiments were performed with additive noise for four pitch extraction methods, i.e., the proposed method, the original TEMPO, an improved cepstrum method, and a common F0 extraction program in ESPS. The results were as follows: 1) the proposed method is always better than any of the other methods when the SNR is greater than about 2 dB; 2) for high SNR values (> 15 dB), the correct rates of the proposed method and the original TEMPO are about 95% and much better than the improved cepstrum method (92%) and the ESPS function (89%); and 3) all of the methods degrade to less than 62% when the SNR is 0 dB. As a result, the proposed method improves the performance for low SNR values and also maintains high accuracy inherent from the original TEMPO for high SNR values. Yoshinori Atake, Toshio Irino, Hideki Kawahara, Jinlin Lu, Satoshi Nakamura 0001, Kiyohiro Shikano |
INTERSPEECH | 5 |
| 2000 | A block cosine transform and its application in speech recognition
Jingdong Chen, Kuldip K. Paliwal, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2000 | Cellular-phone based speech-to-speech translation system ATR-MATRIX
Rainer Gruhn, Harald Singer, Hajime Tsukada, Masaki Naito, Atsushi Nishino, Atsushi Nakamura, Yoshinori Sagisaka, Satoshi Nakamura 0001 |
INTERSPEECH | 8 |
| 2000 | Frame level likelihood transformations for ASR and utterance verificationabstractIn most of the current speech recognition systems based on HMM, existing decoding and utterance veri cation methods make use of state output likelihood as a measure of the acoustic match between the input data and the acoustic models. In this paper, we present a new and more generalized approach to the formation of the acoustic match score. The essence of this approach is to transform the likelihood of each acoustic vector with respect to any particular HMM state according to some non-linear function. We have investigated two types of such transformation functions. The rst one, performs likelihood normalization, and the second one transforms likelihoods into exponentially ordered weights. The transformed likelihoods, as new acoustic scores, are used further for decoding, recognition and veri cation instead of the conventional likelihoods. In our evaluation experiments we used TIMIT database for phoneme recognition and veri cation and a database of 710 speakers and a total of 4252 distinct words, for isolated word recognition and veri cation. The results we achieved show that the transformed likelihood scores, in average, increase slightly the recognition accuracy and reduce the veri cation error rates up to 30%. Konstantin Markov, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2000 | Analysis of acoustic models trained on a large-scale Japanese speech database
Tomoko Matsui, Masaki Naito, Yoshinori Sagisaka, Kozo Okuda, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2000 | Design of robust subtractive beamformer for noisy speech recognition
Mitsunori Mizumachi, Masato Akagi, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2000 | Stream weight optimization of speech and lip image sequence for audio-visual speech recognitionabstractBimodal speech recognition systems, with the use of visual information to supplement acoustic information, have been shown to yield better recognition performance than purely acoustic systems, especially when background noise is present. The early integration strategy for HMM-based audio-visual speech recognition is one promising approach, where the output probability is obtaned by product of output probabilites of audio and visual streams. This paper addresses a novel method which optimizes stream weights so as to maximize recognition performance. The proposed method estimates the stream weights based on a normalized log likelihood which is derived by ratio of likelihood of a correct word and highest likelihood of incorrect words. The isolated word recognition experiment results show that the audio-visual speech recognition by proposed method attains 56.2% (10 dB), 55.2% (0dB) and 15.2% (20dB) better performance compared to that only using audio information. The results also show the proposed method can reduce a number of adaptation words. Satoshi Nakamura 0001, Hidetoshi Ito, Kiyohiro Shikano |
INTERSPEECH | 1 |
| 2000 | Multimodal corpora for human-machine interaction research
Satoshi Nakamura 0001, Keiko Watanuki, Toshiyuki Takezawa, Satoru Hayamizu |
INTERSPEECH | 1 |
| 2000 | Residual noise compensation by a sequential EM algorithm for robust speech recognition in nonstationary noiseabstractWe model noise as a stationary component plus a time varying residual. The stationary part is estimated off-line and compensated using Log-Add noise compensation. The time varying residual is estimated and compensated using a sequential EM algorithm. The residual noise compensation proceeds in parallel with the recognition process. Experimental results demonstrate that the proposed algorithm improves the recognition performance not only in highly nonstationary noise but also in slow-varying noise, compared with Log-Add noise compensation alone. Kaisheng Yao, Bertram E. Shi, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2000 | Discriminating Chinese lexical tones by anchoring F0 features
Jinsong Zhang 0001, Satoshi Nakamura 0001, Keikichi Hirose |
INTERSPEECH | 2 |
| 2000 | Acoustical Sound Database in Real Environments for Sound Scene Understanding and Hands-Free Speech Recognition
Satoshi Nakamura 0001, Kazuo Hiyane, Futoshi Asano, Takanobu Nishiura, Takeshi Yamada |
LREC | 1 |
| 2000 | Speech enhancement based on the subspace methodabstractA method of speech enhancement using microphone-array signal processing based on the subspace method is proposed and evaluated. The method consists of the following two stages corresponding to the different types of noise. In the first stage, less-directional ambient noise is reduced by eliminating the noise-dominant subspace. It is realized by weighting the eigenvalues of the spatial correlation matrix. This is based on the fact that the energy of less-directional noise spreads over all eigenvalues while that of directional components is concentrated on a few dominant eigenvalues. In the second stage, the spectrum of the target source is extracted from the mixture of spectra of the multiple directional components remaining in the modified spatial correlation matrix by using a minimum variance beamformer. Finally, the proposed method is evaluated in both a simulated model environment and a real environment. Futoshi Asano, Satoru Hayamizu, Takeshi Yamada, Satoshi Nakamura 0001 |
IEEE Trans. Speech Audio Process. | 4 |
| 1999 | Simultaneous recognition of multiple sound sources based on 3-d n-best search using microphone arrayabstractThe recognition of distant talking speech in a noisy and reverberant environments is key issue in any speech recognition system. A so-called hands-free speech recognition system plays an important role in the natural and friendly human-machine interface. Considering the practical use of a speech recognition system, we realize that such a system has to deal, also, with the case of the presence of multiple sound sources, including multiple talkers, as well as other noise sources. This paper proposes a novel method which recognizes multiple talkers simultaneously in real environments by extending the 3-D Viterbi search to a 3-D N-best search algorithm. While the 3-D Viterbi method finds the most likely path in the 3-D trellis space, the proposed method considers multiple hypotheses for each direction in every frame. Combinations of the direction sequence and the phoneme sequence of multiple sources are included in the N-best list. The paper investigates the performance of the proposed method through experiments using real utterances of multiple talkers. Panikos Heracleous, Takeshi Yamada, Satoshi Nakamura 0001, Kiyohiro Shikano |
EUROSPEECH | 3 |
| 1999 | Data collection in real acoustical environments for sound scene understanding and hands-free speech recognitionabstractThis paper describes a sound scene database necessary for studies such as sound source localization, sound retrieval, sound recognition and hands-free speech recognition in real acoustical environments.This paper reports on a project for collection of the sound scene data supported by Real World Computing Partnership(RWCP).There are many kinds of sound scenes in real environments.The sound scene is denoted by sound sources and room acoustics.The numb e r o f c o m bination of the sound sources, source positions and rooms is huge in real acoustical environments.Two approaches are taken to build the sound scene database in the early stage of the project.The rst approach is to collect isolated sound sources of many kinds of non-speech sounds and speech sounds.The second approach is to collect impulse responses in various acoustical environments.The sound in the environments can be simulated by convolution of the isolated sound sources and impulse responses.In a later stage, the sound scene data in real acoustical environments is planned to be collected using a three dimensional microphone array.In this paper, the plan and progress of our sound scene database project are described.1. Satoshi Nakamura 0001, Kazuo Hiyane, Futoshi Asano, Takeshi Yamada, Takashi Endo |
EUROSPEECH | 1 |
| 1998 | Lip Movement Synthesis from Speech Based on Hidden Markov Models
Eli Yamamoto, Satoshi Nakamura 0001, Kiyohiro Shikano |
FG | 2 |
| 1998 | Efficient representation of short-time phase based on group delayabstractAn efficient representation of short-time phase characteristics of speech sounds is proposed, based on findings which suggest the perceptual importance of phase characteristics. Subjective tests indicated that the synthesized speech sounds by the proposed method are indistinguishable from the original speech sounds with a moderate data compression. The proposed representation uses lower-order coefficients of the inverse Fourier transform of the group delay of speech. It also alleviates the voiced/unvoiced decision, which is an indispensable part in conventional speech coding algorithms. These features make our method potentially very useful in many applications like speech morphing. Hideki Banno, Jinlin Lu, Satoshi Nakamura 0001, Kiyohiro Shikano, Hideki Kawahara |
ICASSP | 3 |
| 1998 | Robust speech recognition in car environmentsabstractA user-friendly speech interface in a car cabin is highly needed for safety reasons. This paper describes a robust speech recognition method that can cope with additive noise and multiplicative distortions. A known additive noise, a source signal of which is available, might be canceled by NLMS-VAD (normalized least mean squares with frame-wise voice activity detection). On the other hand, an unknown additive noise, a source signal of which is not available, is suppressed with CSS (continuous spectral subtraction). Furthermore, various multiplicative distortions are simultaneously compensated with E-CMN (exact cepstrum mean normalization) which is speaker dependent/environment-dependent CMN for speech/non-speech. Evaluation results of the proposed method for car cabin environments are finally described. Makoto Shozakai, Satoshi Nakamura 0001, Kiyohiro Shikano |
ICASSP | 2 |
| 1998 | Hands-free speech recognition based on 3-D Viterbi search using a microphone arrayabstractA microphone array is a promising solution for realizing hands-free speech recognition in real environments. Accurate talker localization is very important for speech recognition using a microphone array. However localization of a moving talker is difficult in noisy reverberant environments. Talker localization errors degrade the performance of speech recognition. To solve the problem, this paper proposes a new speech recognition algorithm which considers multiple talker direction hypotheses simultaneously. The proposed algorithm performs a Viterbi search in 3-dimensional trellis space composed of talker directions, input frames, and HMM states. As a result, a locus of the talker and a phoneme sequence of the speech are obtained by finding an optimal path with the highest likelihood. To evaluate the performance of the proposed algorithm, speech recognition experiments are carried out on simulated data and real environment data. These results show that the proposed algorithm works well even if the talker moves. Takeshi Yamada, Satoshi Nakamura 0001, Kiyohiro Shikano |
ICASSP | 2 |
| 1998 | Creating speaker independent HMM models for restricted database using STRAIGHT-TEMPO morphingabstractICSLP1998: the 5th International Conference on Spoken Language Processing, November 30 - December 4, 1998, Sydney, Australia. Alexandre Girardi, Kiyohiro Shikano, Satoshi Nakamura 0001 |
ICSLP | 3 |
| 1998 | Evaluation of model adaptation by HMM decomposition on telephone speech recognitionabstractIn this paper, we evaluate performance of model adaptation by the previously proposed HMM decomposition method on telephone speech recognition. The HMM decomposition method separates a composed HMM into a known phoneme HMM and an unknown noise and channel HMM by maximum likelihood (ML) estimation of the HMM parameters. A transfer function (telephone channel) HMM is estimated using adaptation speech data by applying the HMM decomposition twice in the linear spectral domain for noise and in the cepstral domain for channel. The telephone speech data for evaluation are recorded through 10 kinds of ordinary analog telephone handsets and cordless telephone handsets. The test results show that the average phrase accuracy with the clean speech HMMs is 60.9% for the ordinary analog telephone handsets, and 19.6% for the cordless telephone handsets. By the HMM decomposition method, the average phrase accuracy is improved to 78.1% for the ordinary analog telephone handsets, and 50.5% for the cordless telephone handsets. Tetsuya Takiguchi, Satoshi Nakamura 0001, Kiyohiro Shikano, Masatoshi Morishima, Toshihiro Isobe |
ICSLP | 2 |
| 1998 | An effect of adaptive beamforming on hands-free speech recognition based on 3-d viterbi searchabstractICSLP1998: the 5th International Conference on Spoken Language Processing, November 30 - December 4, 1998, Sydney, Australia. Takeshi Yamada, Satoshi Nakamura 0001, Kiyohiro Shikano |
ICSLP | 2 |
| 1998 | Speech-to-lip movement synthesis based on the EM algorithm using audio-visual HMMsabstractICSLP1998: the 5th International Conference on Spoken Language Processing, November 30 - December 4, 1998, Sydney, Australia. Eli Yamamoto, Satoshi Nakamura 0001, Kiyohiro Shikano |
ICSLP | 2 |
| 1998 | Compression algorithm of trigram language models based on maximum likelihood estimationabstractICSLP1998: the 5th International Conference on Spoken Language Processing, November 30 - December 4, 1998, Sydney, Australia. Norimichi Yodo, Kiyohiro Shikano, Satoshi Nakamura 0001 |
ICSLP | 3 |
| 1998 | Speech-to-lip movement synthesis maximizing audio-visual joint probability based on EM algorithmabstractWe investigate methods using the hidden Markov model (HMM) to drive a lip movement sequence with input speech. We have already investigated a mapping method based on the Viterbi decoding algorithm which converts an input speech to a lip movement sequence through the most likely HMM state sequence conducted by audio HMMs. However, the method contains a substantial problem of producing errors along incorrectly decoded HMM states. This paper newly proposes a method to re-estimate the visual parameters using the HMMs of the audio-visual joint probability under the expectation-maximization (EM) algorithm. In experiments, the proposed mapping method using the EM algorithm shows an error reduction of 26% compared to a method using the Viterbi algorithm at incorrectly decoded bi-labial consonants. Satoshi Nakamura 0001, Eli Yamamoto, Kiyohiro Shikano |
MMSP | 1 |
| 1998 | Lip movement synthesis from speech based on Hidden Markov Models
Eli Yamamoto, Satoshi Nakamura 0001, Kiyohiro Shikano |
Speech Commun. | 2 |
| 1997 | Model adaptation based on HMM decomposition for reverberant speech recognitionabstractThe performance of a speech recognizer is degraded drastically in reverberant environments. The authors propose a novel algorithm which can model an observation signal by composition of HMMs of clean speech, noise and an acoustic transfer function. However, estimating HMM parameters of the acoustic transfer function is still a serious problem. In their previous paper, they measured real impulse responses of training positions in an experiment room. It is inconvenient and unrealistic to measure impulse responses for every possible new experiment room. The paper presents a new method for estimating HMM parameters of the acoustic transfer function from some adaptation data by using an HMM decomposition algorithm which is an inverse process of the HMM composition. Its effectiveness is confirmed by a series of speaker dependent and independent word recognition experiments on simulated distant-talking speech data. Tetsuya Takiguchi, Satoshi Nakamura 0001, Kiyohiro Shikano |
ICASSP | 2 |
| 1997 | Maximum likelihood successive state splitting algorithm for tied-mixture HMNETabstractThis paper describes a new approach to ML-SSS (Maximum Likelihood Successive State Splitting) algorithm that uses tied- mixture representation of the output probability density function instead of a single Gaussian during the splitting phase of the ML-SSS algorithm. The tied-mixture representation results in a better state split gain, because it is able to measure diferences in the phoneme environment space that ML-SSS can not. With this more informative gain the new algorithm can choose a better split state and corresponding data. Phoneme clustering experiments were conducted which lead up to 38% of error reduction if compared to the ML-SSS algorithm. Alexandre Girardi, Harald Singer, Kiyohiro Shikano, Satoshi Nakamura 0001 |
EUROSPEECH | 4 |
| 1997 | Microphone array design measures for hands-free speech recognitionabstractEUROSPEECH1997: the 5th European Conference on Speech Communication and Technology , September 22-25, 1997, Rhodes, Greece. Masaaki Inoue, Satoshi Nakamura 0001, Takeshi Yamada, Kiyohiro Shikano |
EUROSPEECH | 2 |
| 1997 | Improved bimodal speech recognition using tied-mixture HMMs and 5000 word audio-visual synchronous databaseabstractEUROSPEECH1997: the 5th European Conference on Speech Communication and Technology , September 22-25, 1997, Rhodes, Greece. Satoshi Nakamura 0001, Ron Nagai, Kiyohiro Shikano |
EUROSPEECH | 1 |
| 1997 | Room acoustics and reverberation: impact on hands-free recognitionabstractHands-free speech recognition is a very important issue for a natural human machine interface. The distant talking speech in real environments is distorted by noise and reverberation of the room. This paper introduces characteristics of the room acoustical distortion and their influences on speech recognition accuracy. Then the paper tries to give a prospect of the solution based on previous studies and our research efforts. Especially a microphone array based-method and a model adaptation method are discussed. The microphone array can reduce the influences of the acoustical distortion by beam-forming. On the other hand, the model adaptation method can estimate the acoustical transfer function and adapt the speech models against the distorted observation signals. Furthermore, this paper also addresses hands-free speech recognition by incorporating automatic lip reading. Satoshi Nakamura 0001, Kiyohiro Shikano |
EUROSPEECH | 1 |
| 1997 | A non-iterative model-adaptive e-CMN/PMC approach for speech recognition in car environmentsabstractThis paper investigates the Cepstrum Mean Normalization(CMN) which has been widely acknowledged useful for compensation of multiplicative distortions. However, the performance of usual CMN is limited because the normalization by a single cepstrum mean vector is not enough to compensate many factors of multiplicative distortion in real environments. To solve this problem, a new method E-CMN is proposed. The method estimates two cepstrum mean vectors, one for speech and the other for non-speech for each speaker and subtracts them from an input cepstrum This method is capable of compensating various kinds of multiplicative distortion collectively to normalize input spectra. Furthermore, a new model-adaptive approach E-CMN/PMC, based on E- CMN and HMM composition, is proposed for environments with additive noise and multiplicative distortions. This method is simplified in a sense that it is possible to add speech models and an additive noise model without any iterative operations. Matching gains for all frequency bands of speech models to the noise model are uniquely estimated as a cepstrum mean vector for speech. The performance of E-CMN/PMC in adverse car environsnents is finally evaluated. Makoto Shozakai, Satoshi Nakamura 0001, Kiyohiro Shikano |
EUROSPEECH | 2 |
| 1996 | Noise and room acoustics distorted speech recognition by HMM compositionabstractThis paper presents a robust speech recognition method based on the HMM composition for the noisy room acoustics distorted speech. The method realizes an improved user interface such as the user is not encumbered by microphone equipment. The proposed HMM composition is obtained by naturally extending the HMM composition method of an additive noise to that of the convolutional room acoustics distortion. The HMM composition is conducted by 2 steps: (1) composition of HMMs of a speech and acoustical transfer function in the cepstrum domain, and (2) composition of distorted speech and noise HMMs in the linear spectral domain. The speaker dependent/independent word recognition experiments are carried out using the speech database contaminated by the additive noise and convolutional room acoustics distortion. The evaluation experiments are also conducted for unknown testing sound source positions. These results clarified the effectiveness of the proposed method. Satoshi Nakamura 0001, Tetsuya Takiguchi, Kiyohiro Shikano |
ICASSP | 1 |
| 1996 | Robust speech recognition with speaker localization by a microphone arrayabstractThis paper proposes robust speech recognition with Speaker Localization by a Arrayed Microphone (SLAM) to realize hands-free speech interface in noisy environments.In order to localize a speaker direction accurately in low SNR conditions, a speaker localization algorithm based on extracting a pitch harmonics is introduced.To evaluate the performance of the proposed system, speech recognition experiments are carried out both in computer simulation and real environments.These results show that the proposed system attains the much higher speech recognition performance than that of a single microphone not only in computer simulation but also in real environments. Takeshi Yamada, Satoshi Nakamura 0001, Kiyohiro Shikano |
ICSLP | 2 |
| 1993 | Robust word spotting in adverse car environmentsabstractThis paper presents a novel word recognition technique which allows hand-free dialing under adverse car environments. The algorithm is based on speaker dependent word spotting. The paper compared and investigated the methods of spectral subtraction, short time modified coherence, multi-microphone, dynamic and accelerated features, weighted distance measures and multi-templates by word recognition experiments. The experiments are carried out using real speech database uttered in adverse car environments. The experiments show that the method using sinusoidal weighted lifter and the method using multi-templates are robust against the adverse environments. Keywords: word spotting, distance measure, adverse environment 1. INTRODUCTION Recent technology realizes a small size mobile cellular telephone. This enables easy communication to anybody at any time on the way. However, it has a problem when we use the telephone while driving. Dialing and a telephone call while driving are very dang... Satoshi Nakamura 0001, Toshio Akabane, Seiji Hamaguchi |
EUROSPEECH | 1 |
| 1991 | A neural speaker model for speaker clusteringabstractA speaker model using a neural network is proposed for reference speaker clustering on speaker independent speech recognition. Speaker individuality is embedded in not only a static short time spectrum and a pitch frequency, but also a dynamic spectral pattern and pitch pattern. In conventional modeling, speaker individuality is based on the former static features. The authors try to capture the latter dynamic features, of speaker by a neural speaker model. Two methods, neural prediction modeling by multilayer perceptron and learning matrix vector-quantization, are considered for the speaker modeling. Using the measures of speaker modeling, speaker clustering of the reference patterns based on mutual information is carried out for speaker independent speech recognition.> Satoshi Nakamura 0001, Toshio Akabane |
ICASSP | 1 |
| 1990 | ATR HMM-LR continuous speech recognition systemabstractAn improvement of the hidden Markov model (HMM) LR continuous-speech recognizer using multiple codebooks, HMM state duration control and fuzzy vector quantization is described. The system recognizes Japanese phrases (Bunsetsu) according to a context-free grammar including 1035 words. In speaker-dependent conditions, a phrase recognition rate of 88.4% (99.0% for the top five candidates) was attained. The system was tested with speaker-adaptation based on a codebook mapping algorithm. An average speaker-adapted phrase recognition rate of 81.6% (98.8% for the top five candidates) was attained.> Toshiyuki Hanazawa, Kenji Kita, Satoshi Nakamura 0001, Takeshi Kawabata, Kiyohiro Shikano |
ICASSP | 3 |
| 1990 | Supplementation of HMM for articulatory variation in speaker adaptationabstractA method of dealing with articulatory speaker variations in hidden Markov models (HMMs) for speaker adaptation is proposed. Speech data from many speakers are spectrally mapped onto a standard speaker. These data are used to teach the HMM the interspeaker articulatory variations that subsist across the spectral mapping. The proposed method is compared to other adaptation methods through the /b,d,g/ recognition task. The results show 82.5% recognition accuracy, which is better than the rates of other methods. Evaluation experiments on a Japanese all phoneme recognition task and a continuous-speech recognition task are reported. Average recognition rates for Japanese all phonemes are 71.3% and 93.2%, for the best candidate and the top-three candidates, respectively. These are 0.7% and 1.5% higher than the rates of the basic spectrum mapping method. In the continuous-speech recognition experiment, average phrase recognition rates are 74.9% and 96.2%, for the best candidate and the top-five candidates, respectively.> Hiroaki Hattori, Satoshi Nakamura 0001, Kiyohiro Shikano |
ICASSP | 2 |
| 1990 | A comparative study of spectral mapping for speaker adaptationabstractA comparative study of speaker adaptation using neural network spectral mapping and fuzzy vector quantization (VQ)-based spectral mapping is described. The speaker adaptation experiments were carried out using a database of 216 phonetically balanced words uttered by three speakers. The accuracy of spectral mapping is measured and evaluated by interspeaker spectral distortion. The results show the fuzzy VQ-based spectral mapping algorithm to be about 6% better in spectral distortion than nonlinear spectral mapping by a feedforward neural network. An investigation of the actual spectrogram shows that the fuzzy VQ-based spectral mapping works better than the neural network mapping.> Satoshi Nakamura 0001, Kiyohiro Shikano |
ICASSP | 1 |
| 1990 | Speaker weighted training of HMM using multiple reference speakers
Hiroaki Hattori, Satoshi Nakamura 0001, Kiyohiro Shikano, Shigeki Sagayama |
ICSLP | 2 |
| 1989 | Speaker adaptation applied to HMM and neural networksabstractThe authors propose a speaker adaptation algorithm which does not depend on speech recognition algorithms. The proposed spectral mapping algorithm is based on three ideas: (1) accurate representation of the input vector by separate vector quantization and fuzzy vector quantization, (2) continuous spectral mapping from one speaker to another by fuzzy mapping, and (3) accurate establishment of spectral correspondence based on the fuzzy relationship of the membership function obtained from supervised training. The spectrum dynamic features are also utilized. The algorithm is applied to hidden Markov models (HMMs) and neural networks and evaluated using a database of 216 phonetically balanced words and 5240 important Japanese words uttered by three speakers. The HMM speaker adapted recognition rate for /b,d,g/ is 79.5%. The average recognition rate for the top-three choices is about 91%. The algorithm was applied to neural networks and resulted in almost the same performance. The algorithm was also applied to voice conversion, and a preference score of 65.6% was obtained.> Satoshi Nakamura 0001, Kiyohiro Shikano |
ICASSP | 1 |