EDBT 2026 Demo / reviewers in the wild / expert
Tatsuya Kawahara
dblp:92/635
· DBLP profile ↗
302ranked-venue papers
36as first author
64since 2021 · last 2026
0000-0002-2686-2296ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 214 · 25 first-author · 45 since 2021Graphics, computer vision, multimedia, augmented reality and games · 211 · 32 first-author · 35 since 2021Human-computer interaction and ubiquitous computing · 17 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding EvaluationabstractWeihua Zheng, Zhengyuan Liu, Tanmoy Chakraborty, Weiwen Xu, Xiaoxue Gao, Bryan Chen Zhengyu Tan, Bowei Zou, Chang Liu, Yujia Hu, Xing Xie, Xiaoyuan Yi, Jing Yao, Chaojun Wang, Long Li, Rui Liu, Huiyao Liu, Koji Inoue, Ryuichi Sumida, Tatsuya Kawahara, Fan Xu, Lingyu Ye, Wei Tian, Dongjun Kim, Jimin Jung, Jaehyung Seo, Nadya Yuki Wangsajaya, Pham Minh Duc, Ojasva Saxena, Palash Nandi, Xiyan Tao, Wiwik Karlina, Tuan Luong, Keertana Arun Vasan, Roy Ka-Wei Lee, Nancy F. Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhengyuan Liu, Tanmoy Chakraborty 0002, Weiwen Xu, Xiaoxue Gao, Bryan Chen Zhengyu Tan, Bowei Zou, Chang Liu 0071, Xing Xie 0001, Xiaoyuan Yi, Jing Yao 0003, Chaojun Wang, Rui Liu 0019, Huiyao Liu, Koji Inoue, Ryuichi Sumida, Tatsuya Kawahara, Lingyu Ye, Jimin Jung, Jaehyung Seo, Nadya Yuki Wangsajaya, Pham Minh Duc, Ojasva Saxena, Palash Nandi, Xiyan Tao, Wiwik Karlina, Tuan Luong, Keertana Arun Vasan, Roy Ka-Wei Lee, Nancy F. Chen |
ACL (1) | 19 |
| 2026 | Adapting Pretrained Models to Endangered Languages in Japan: A Comparative Study on Ryukyuan and Ainu Speech Recognition
Kohei Matsuura, Takanori Ashihara, Tatsuya Kawahara |
LREC | 3 |
| 2026 | On the Structure of Address in Multi-Party Dialogue: From Discrete Labels to Continuous LevelsabstractIn multi-party dialogues between a dialogue system and multiple users, identifying to whom an utterance is addressed is a key challenge. Prior work has typically treated addressee detection as a multi-class classification task, selecting a single label representing an individual participant or the group. This formulation assumes that address is inherently discrete and has primarily been used for predicting turn-taking. In this paper, we revisit this assumption by analyzing address as a continuous phenomenon. Using a multi-party human dialogue corpus annotated by multiple annotators, we construct both binary address labels derived from majority-vote addressee labels and continuous address levels inferred from annotator judgments using a latent-variable model. We then examine how these representations relate to turn-taking as well as listener behaviors, including gaze and backchannels. Our results show that, in addition to turn-taking, both gaze and backchannels are associated with address. Furthermore, models using continuous address levels achieve better predictive fit than those using discrete labels, suggesting that address may exhibit graded structure. Finally, we discuss the future directions of addressee detection research based on the findings of this study. Taiga Mori, Koji Inoue, Divesh Lala, Tatsuya Kawahara |
SIGDIAL | 4 |
| 2026 | I Understand How You Feel: Enhancing Deeper Emotional Support Through Multilingual Emotional Validation in Dialogue SystemabstractEmotional validation - explicitly acknowledging that a user’s feelings make sense - has proven therapeutic value but has received little computational attention. The emotional validation in dialogue systems can be decomposed into (i) validating response identification, (ii) validation timing detection, and (iii) validating response generation. To support research on all three subtasks, we release M-EDESConv, a 120k English–Japanese multilingual corpus created through hybrid manual–automatic annotation, and M-TESC, a multilingual spoken-dialogue test set. For timing detection, we propose MEGUMI, a Multilingual Emotion-aware Gated Unit for Mutual Integration, that fuses frozen XLM-RoBERTa semantics with language-specific emotion encoders via cross-modal attention and gated fusion. MEGUMI shows superior performance on both the M-EDESConv and M-TESC datasets, both objectively and subjectively. Finally, our EmoValidBench benchmarks of GPT-4.1 Nano and Llama-3.1 8B indicate that current LLMs generate contextually similar, diverse validating responses, but emotional understanding remains a major area for improvement. Zi Haur Pang, Yahui Fu 0001, Koji Inoue, Tatsuya Kawahara |
SIGDIAL | 4 |
| 2026 | Bridging Speech Emotion Recognition and Personality: Dataset and Temporal Interaction Condition NetworkabstractThis study investigates the interaction between personality traits and emotion expression, exploring how personality information can improve speech emotion recognition (SER). We collect the personality annotation for the IEMOCAP dataset, making it the first speech dataset that contains both emotion and personality annotations (PA-IEMOCAP), and enabling direct integration of personality traits into SER. Statistical analysis on this dataset identified significant correlations between per sonality traits and emotional expressions. To extract finegrained personality features, we propose a temporal interaction condition network (TICN), in which personality features are integrated with HuBERT-based acoustic features for SER. Experiments show that incorporating ground-truth personality traits significantly enhances valence recognition, improving the concordance correlation coefficient (CCC) from 0.698 to 0.785 compared to the baseline without personality information. For practical applications in dialogue systems where personality information about the user is unavailable, we develop a front-end module of automatic personality recognition. Using these automatically predicted traits as inputs to our proposed TICN model, we achieve a CCC of 0.776 for valence recognition, representing an 11.17% relative improvement over the baseline. These findings confirm the effectiveness of personality-aware SER and provide a solid foundation for further exploration in personality-aware speech processing applications. Yuan Gao 0040, Yahui Fu 0001, Chenhui Chu, Tatsuya Kawahara |
IEEE Trans. Affect. Comput. | 5 |
| 2026 | Robot-Mediated Multi-Party Conversation Aimed at Affect Improvement for Psychiatric PatientsabstractThis paper describes a multi-party attentive listening system that interacts with two persons to familiarize each other through conversation. We are mainly targeting social implementation in hospitals to contribute to the rehabilitation of people with psychiatric disorders to promote affect improvement in terms of pleasure and arousal. We conducted an experiment in a psychiatric outpatient-daycare program. Twenty daycare attendees participated in a three-party conversation session between a pair of two humans and a humanoid robot. One of the paired participants talked about his/her favorite topic and was attentively listened to by the other and the robot. In a subsequent session, the human pairs switched each other's roles. The subjective evaluations showed that both the pleasure and arousal of the participants were significantly improved after the conversation. The participants rated the impression of the robot as easier to talk with than strangers. They also rated that they could understand and feel familiar with significantly more their human talk partner after the conversational session. The multiple linear regression analysis showed that participants became more pleasant and verbal when stimulated by backchannels and questions from both the human listener and the robot. It suggests that those who expressed themself using more words received positive impressions. Keiko Ochi, Divesh Lala, Koji Inoue, Tatsuya Kawahara, Hirokazu Kumazaki |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | KyotoMOS2: MOS Prediction for Speech Across Multiple Sampling RatesabstractWe propose KyotoMOS2, an automatic MOS prediction system capable of evaluating speech across varying sampling rates. We design a sampling-rate-aware SSL-MOS subsystem and evaluate 32 variants on both original and resampled audio. Based on system-level SRCC and MSE, 13 subsystems are selected and fused for the final prediction. A system-level SRCC-based early stopping strategy is used to train the fusion model. Our system (T13) ranked 2nd on SRCC in the AudioMOS Track 3. Wangjin Zhou, Keisuke Imoto, Tatsuya Kawahara |
ASRU | 4 |
| 2025 | Can LLMs be Surprised? Evaluation and Analysis of Surprise Expression of LLMsabstractWe investigate methods for virtual agents or humanoid robots to express surprise in dialogue by using LLMs. We created a newly annotated dialogue dataset focusing on surprise. Our findings indicate that accurately expressing surprise in dialogue is a challenging task. They also suggest directions for improvement–such as more appropriate modeling of commonness–and identify the feature of surprise itself. Motoori Takeuchi, Koji Inoue, Keiko Ochi, Tatsuya Kawahara |
HAI | 4 |
| 2025 | Leveraging IPA and Articulatory Features as Effective Inductive Biases for Multilingual ASR TrainingabstractIn recent advancements in end-to-end ASR, large-scale self-supervised or weakly supervised models have achieved a significant milestone. However, it remains challenging to train consistently high-performing multilingual models, transferable to languages without much resource. In this study, we propose embedding universal phonological knowledge to multilingual ASR by predicting international phonetic alphabet (IPA) targets and universal articulatory features alongside primary grapheme targets. These additions are expected to provide effective inductive bias or regularization for predicting grapheme targets across various languages. In the experiments, which involve fine-tuning a pre-trained XLS-R model using 10,400 hours of data across 120 languages from the Common Voice corpus, our proposed method achieved a 6.81% relative reduction in character error rate. Masato Mimura, Tatsuya Kawahara |
ICASSP | 3 |
| 2025 | Extending Whisper for Emotion Prediction Using Word-level Pseudo LabelsabstractThis paper extends Whisper’s automatic speech recognition (ASR) capabilities to perform speech-based emotion recognition (SER) by incorporating word-level emotion classification alongside ASR output. We generate four emotion pseudo-labels (neutral, happy, sad, angry) for each word using a pretrained frame-level SER model, and Whisper is fine-tuned for joint ASR and emotion classification at the word level. Sentence-level emotion labels are masked during training to encourage the transformer to use the ASR output for word-level emotion prediction. During inference, word-level predictions are combined with sentence-level predictions through majority voting to generate the final sentence-level label. When evaluated on the IEMOCAP dataset, our method maintains Whisper’s ASR word error rate while improving the SER weighted accuracy from 74.4% to 76.4% and the unweighted average recall from 77.1% to 79.0%. Kwok Chin Yuen, Sheng Li 0010, Jia Qi Yip, Chenhui Chu, Tatsuya Kawahara, Chng Eng Siong |
ICASSP | 5 |
| 2025 | InvoxSVC: Any-to-any Zero-shot Singing Voice Conversion with In-Context Learning in Latent Flow MatchingabstractRecent advancements in singing voice conversion (SVC) have focused on achieving zero-shot, any-to-any voice transformation capabilities. Many approaches attempt to modify voice characteristics by incorporating global timbre variables into acoustic models. However, these methods often depend heavily on the capabilities of timbre extractors and lack an understanding of temporal local information. This limitation poses challenges, particularly in replicating specific voice qualities such as those of children. To address this issue, we introduce InvoxSVC, a latent flow matching model (LFM) designed for rapid and precise singing voice conversion with a particular emphasis on capturing temporal local features. While reducing the residual timbral information in the source singing encoding through singer-guidance, InvoxSVC enhances the model’s ability to capture temporal nuances by integrating in-context learning during inference. Additionally, the model employs a pre-trained high-fidelity variational autoencoder (VAE) to improve waveform generation. In comparative evaluations, InvoxSVC outperforms the open-source project So-VITS-SVC in both objective and subjective assessments. Wangjin Zhou, Tianjiao Du, Wenhao Guan, Chenglin Xu, Yi Zhao 0006, Tatsuya Kawahara |
ICME | 7 |
| 2025 | CCMI 2025: Cross-Cultural Multimodal Interaction
Koji Inoue, Shogo Okada, Divesh Lala, Sahba Zojaji, Nancy F. Chen, Tatsuya Kawahara |
ICMI | 6 |
| 2025 | Real-time Generation of Various Types of Nodding for Avatar Attentive Listening System
Kazushi Kato, Koji Inoue, Divesh Lala, Keiko Ochi, Tatsuya Kawahara |
ICMI | 5 |
| 2025 | Triadic Multi-party Voice Activity Projection for Turn-taking in Spoken Dialogue SystemsabstractTurn-taking is a fundamental component of spoken dialogue, however conventional studies mostly involve dyadic settings. This work focuses on applying voice activity projection (VAP) to predict upcoming turn-taking in triadic multi-party scenarios. The goal of VAP models is to predict the future voice activity for each speaker utilizing only acoustic data. This is the first study to extend VAP into triadic conversation. We trained multiple models on a Japanese triadic dataset where participants discussed a variety of topics. We found that the VAP trained on triadic conversation outperformed the baseline for all models but that the type of conversation affected the accuracy. This study establishes that VAP can be used for turn-taking in triadic dialogue scenarios. Future work will incorporate this triadic VAP turn-taking model into spoken dialogue systems. Mikey Elmers, Koji Inoue, Divesh Lala, Tatsuya Kawahara |
INTERSPEECH | 4 |
| 2025 | Multi-lingual and Zero-Shot Speech Recognition by Incorporating Classification of Language-Independent Articulatory Features
Ryo Magoshi, Shinsuke Sakai, Tatsuya Kawahara |
INTERSPEECH | 4 |
| 2025 | Switch Conformer with Universal Phonetic Experts for Multilingual ASR
Masato Mimura, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2025 | Simple and Effective Content Encoder for Singing Voice Conversion via SSL-Embedding Dimension Reduction
Wangjin Zhou, Tianjiao Du, Chenglin Xu, Sheng Li 0010, Yi Zhao 0006, Tatsuya Kawahara |
INTERSPEECH | 6 |
| 2025 | A Noise-Robust Turn-Taking System for Real-World Dialogue Robots: A Field ExperimentabstractTurn-taking is a crucial aspect of human-robot interaction, directly influencing conversational fluidity and user engagement. While previous research has explored turn-taking models in controlled environments, their robustness in real-world settings remains underexplored. In this study, we propose a noise-robust voice activity projection (VAP) model, based on a Transformer architecture, to enhance real-time turn-taking in dialogue robots. To evaluate the effectiveness of the proposed system, we conducted a field experiment in a shopping mall, comparing the VAP system with a conventional cloud-based speech recognition system. Our analysis covered both subjective user evaluations and objective behavioral analysis. The results showed that the proposed system significantly reduced response latency, leading to a more natural conversation where both the robot and users responded faster. The subjective evaluations suggested that faster responses contribute to a better interaction experience. Koji Inoue, Yuki Okafuji, Jun Baba, Yoshiki Ohira, Katsuya Hyodo, Tatsuya Kawahara |
IROS | 6 |
| 2025 | Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity ProjectionabstractKoji Inoue, Divesh Lala, Gabriel Skantze, Tatsuya Kawahara. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Koji Inoue, Divesh Lala, Gabriel Skantze, Tatsuya Kawahara |
NAACL (Long Papers) | 4 |
| 2025 | Prompt-Guided Turn-Taking PredictionabstractTurn-taking prediction models are essential components in spoken dialogue systems and conversational robots. Recent approaches leverage transformer-based architectures to predict speech activity continuously and in real-time. In this study, we propose a novel model that enables turn-taking prediction to be dynamically controlled via textual prompts. This approach allows intuitive and explicit control through instructions such as “faster” or “calmer,” adapting dynamically to conversational partners and contexts. The proposed model builds upon a transformer-based voice activity projection (VAP) model, incorporating textual prompt embeddings into both channel-wise transformers and a cross-channel transformer. We evaluated the feasibility of our approach using over 950 hours of human-human spoken dialogue data. Since textual prompt data for the proposed approach was not available in existing datasets, we utilized a large language model (LLM) to generate synthetic prompt sentences. Experimental results demonstrated that the proposed model improved prediction accuracy and effectively varied turn-taking timing behaviors according to the textual prompts. Koji Inoue, Mikey Elmers, Yahui Fu 0001, Zi Haur Pang, Divesh Lala, Keiko Ochi, Tatsuya Kawahara |
SIGDIAL | 7 |
| 2024 | Multilingual Turn-taking Prediction Using Voice Activity ProjectionabstractThis paper investigates the application of voice activity projection (VAP), a predictive turn-taking model for spoken dialogue, on multilingual data, encompassing English, Mandarin, and Japanese. The VAP model continuously predicts the upcoming voice activities of participants in dyadic dialogue, leveraging a cross-attention Transformer to capture the dynamic interplay between participants. The results show that a monolingual VAP model trained on one language does not make good predictions when applied to other languages. However, a multilingual model, trained on all three languages, demonstrates predictive performance on par with monolingual models across all languages. Further analyses show that the multilingual model has learned to discern the language of the input signal. We also analyze the sensitivity to pitch, a prosodic cue that is thought to be important for turn-taking. Finally, we compare two different audio encoders, contrastive predictive coding (CPC) pre-trained on English, with a recent model based on multilingual wav2vec 2.0 (MMS). Koji Inoue, Bing'er Jiang, Erik Ekstedt, Tatsuya Kawahara, Gabriel Skantze |
LREC/COLING | 4 |
| 2024 | Enhancing Two-Stage Finetuning for Speech Emotion Recognition Using AdaptersabstractThis study investigates the effective finetuning of a pretrained model using adapters for speech emotion recognition (SER). Since emotion is related with linguistic and prosodic information and also other attributes such as gender and speaking style, a framework of multi-task learning (MTL) has been shown to be effective for SER. However, the learning targets of automatic speech recognition (ASR) and other attribute recognition are apparently in conflict. Therefore, we propose to employ different adaptation methods for different tasks in multiple finetuning stages. Since ASR is the most challenging and also influential for SER, in the first stage, we finetune all parameters of the pretrained model for ASR and SER. In the second stage, we incorporate adapters to finetune the model for gender and style recognition in addition to SER by freezing the parameters of the main Transformer model tuned for ASR. Experimental evaluations which extensively compare different adaptation methods using the IEMOCAP dataset demonstrate that the proposed approach achieves a significant improvement from the simple MTL. Yuan Gao 0040, Chenhui Chu, Tatsuya Kawahara |
ICASSP | 4 |
| 2024 | Diffusion-Based Speech Enhancement with Joint Generative and Predictive DecodersabstractDiffusion-based generative speech enhancement (SE) has recently received attention, but reverse diffusion remains time-consuming. One solution is to initialize the reverse diffusion process with enhanced features estimated by a predictive SE system. However, the pipeline structure currently does not consider for a combined use of generative and predictive decoders. The predictive decoder allows us to use the further complementarity between predictive and diffusion-based generative SE. In this paper, we propose a unified system that use jointly generative and predictive decoders across two levels. The encoder encodes both generative and predictive information at the shared encoding level. At the decoded feature level, we fuse the two decoded features by generative and predictive decoders. Specifically, the two SE modules are fused in the initial and final diffusion steps: the initial fusion initializes the diffusion process with the predictive SE to improve convergence, and the final fusion combines the two complementary SE outputs to enhance SE performance. Experiments conducted on the Voice-Bank dataset demonstrate that incorporating predictive information leads to faster decoding and higher PESQ scores compared with other score-based diffusion SE (StoRM and SGMSE+). Kazuki Shimada, Masato Hirano, Takashi Shibuya 0001, Yuichiro Koyama, Shusuke Takahashi, Tatsuya Kawahara, Yuki Mitsufuji |
ICASSP | 8 |
| 2024 | Zero- and Few-Shot Sound Event Localization and DetectionabstractSound event localization and detection (SELD) systems estimate direction-of-arrival (DOA) and temporal activation for sets of target classes. Neural network (NN)-based SELD systems have performed well in various sets of target classes, but they only output the DOA and temporal activation of preset classes trained before inference. To customize target classes after training, we tackle zero- and few-shot SELD tasks, in which we set new classes with a text sample or a few audio samples. While zero-shot sound classification tasks are achievable by embedding from contrastive language-audio pretraining (CLAP), zero-shot SELD tasks require assigning an activity and a DOA to each embedding, especially in overlapping cases. To tackle the assignment problem in overlapping cases, we propose an embed-ACCDOA model, which is trained to output track-wise CLAP embedding and corresponding activity-coupled Cartesian direction-of-arrival (ACCDOA). In our experimental evaluations on zero- and few-shot SELD tasks, the embed-ACCDOA model showed better location-dependent scores than a straightforward combination of the CLAP audio encoder and a DOA estimation model. Moreover, the proposed combination of the embed-ACCDOA model and CLAP audio encoder with zero-or few-shot samples performed comparably to an official baseline system trained with complete train data in an evaluation dataset. Kazuki Shimada, Kengo Uchida, Yuichiro Koyama, Takashi Shibuya 0001, Shusuke Takahashi, Yuki Mitsufuji, Tatsuya Kawahara |
ICASSP | 7 |
| 2024 | MOS-FAD: Improving Fake Audio Detection Via Automatic Mean Opinion Score PredictionabstractIEEE Automatic Mean Opinion Score (MOS) prediction is employed to evaluate the quality of synthetic speech. This study extends the application of predicted MOS to the task of Fake Audio Detection (FAD) as we expect that MOS can be used to assess how close synthesized speech is to the natural human voice. We propose MOS-FAD, where MOS can be leveraged at two key points in FAD: training data selection and model fusion. In training data selection, we demonstrate that MOS enables effective filtering of samples from unbalanced datasets. In the model fusion, our results demonstrate that incorporating MOS as a gating mechanism in FAD model fusion enhances overall performance. Wangjin Zhou, Zhengdong Yang, Chenhui Chu, Sheng Li 0010, Raj Dabre, Yi Zhao 0006, Tatsuya Kawahara |
ICASSP | 7 |
| 2024 | Speech Emotion Recognition with Multi-level Acoustic and Semantic Information Extraction and Interaction
Yuan Gao 0040, Chenhui Chu, Tatsuya Kawahara |
INTERSPEECH | 4 |
| 2024 | Efficient and Robust Long-Form Speech Recognition with Hybrid H3-ConformerabstractRecently, Conformer has achieved state-of-the-art performance in many speech recognition tasks.However, the Transformer-based models show significant deterioration for long-form speech, such as lectures, because the self-attention mechanism becomes unreliable with the computation of the square order of the input length.To solve the problem, we incorporate a kind of state-space model, Hungry Hungry Hippos (H3), to replace or complement the multi-head self-attention (MHSA).H3 allows for efficient modeling of long-form sequences with a linear-order computation.In experiments using two datasets of CSJ and LibriSpeech, our proposed H3-Conformer model performs efficient and robust recognition of long-form speech.Moreover, we propose a hybrid of H3 and MHSA and show that using H3 in higher layers and MHSA in lower layers provides significant improvement in online recognition.We also investigate a parallel use of H3 and MHSA in all layers, resulting in the best performance. Tomoki Honda, Shinsuke Sakai, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2024 | Entrainment Analysis and Prosody Prediction of Subsequent Interlocutor's Backchannels in Dialogue
Keiko Ochi, Koji Inoue, Divesh Lala, Tatsuya Kawahara |
INTERSPEECH | 4 |
| 2024 | Dual-path Adaptation of Pretrained Feature Extraction Module for Robust Automatic Speech Recognition
Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2024 | StyEmp: Stylizing Empathetic Response Generation via Multi-Grained Prefix Encoder and Personality ReinforcementabstractRecent approaches for empathetic response generation mainly focus on emotional resonance and user understanding, without considering the system's personality.Consistent personality is evident in real human expression and is important for creating trustworthy systems.To address this problem, we propose StyEmp, which aims to stylize the empathetic response generation with a consistent personality.Specifically, it incorporates a multi-grained prefix mechanism designed to capture the intricate relationship between a system's personality and its empathetic expressions.Furthermore, we introduce a personality reinforcement module that leverages contrastive learning to calibrate the generation model, ensuring that responses are both empathetic and reflective of a distinct personality.Automatic and human evaluations on the EMPATHETICDIALOGUES benchmark show that StyEmp outperforms competitive baselines in terms of both empathy and personality expressions.Our code is available at https://github.com/fuyahuii/StyEmp. Yahui Fu 0001, Chenhui Chu, Tatsuya Kawahara |
SIGDIAL | 3 |
| 2024 | Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech RecognitionabstractSerialized output training (SOT) attracts increasing attention due to its convenience and flexibility for multi-speaker automatic speech recognition (ASR). However, it is not easy to train with attention loss only. In this paper, we propose the overlapped encoding separation (EncSep) to fully utilize the benefits of the connectionist temporal classification (CTC) and attention (CTC-Attention) hybrid loss. This additional separator is inserted after the encoder to extract the multi-speaker information with CTC losses. Furthermore, we propose the serialized speech information guidance SOT (GEncSep) to further utilize the separated encodings. The separated streams are concatenated to provide single-speaker information to guide attention during decoding. The experimental results on Libri2Mix and Libri3Mix show that the single-speaker encoding can be separated from the overlapped encoding. The CTC loss helps to improve the encoder representation under complex scenarios (three-speaker and noisy conditions), which makes the EncSep have a relative improvement of more than 8% and 6% on the noisy Libri2Mix and Libri3Mix evaluation sets, respectively. GEncSep further improved performance, which was more than 12% and 9% relative improvement for the noisy Libri2Mix and Libri3Mix evaluation sets. Yuan Gao 0040, Zhaoheng Ni, Tatsuya Kawahara |
SLT | 4 |
| 2024 | A large-scale television advertising dataset for detailed impression analysisabstractAbstract Creating impressive video content such as movies and advertisements is a very important yet challenging task in business that requires both a sense of creativity and a lot of experience. Even professionals cannot necessarily invoke the impressions and emotions that they have aimed at. Many video advertisements are created and then disappear without giving a large impact on viewers. This paper presents a large-scale dataset of television (TV) advertisements that consists of 14,490 videos. The impressions of each video such as the recognition rate and interestingness rate are from the results of questionnaires answered by 620 participants. We also present a baseline for predicting the impression effects of TV advertisements by using visual and audio information, metadata such as broadcasting pattern, business category, the popularity of the casts, and text information including texts appearing on videos and narrations in audios. We predict four impressions of the viewers: 1) how much participants remember the video afterward, 2) how much they feel like buying the product/service, 3) how much they become interested in the product/service, and 4) how much they like the content of the advertisement itself. By combining images, audio, metadata, cast data, and text data, our baseline method is able to predict such impressions with a correlation of 0.69-0.82, much better than using a single-modal feature such as visual data or audio data only. This paper also gives some possible applications such as estimating the importance scores of each key frame, which gives us informative insights about how to make the advertisement content more impressive. Shunsuke Nakamura, Tatsuya Kawahara, Gen Tamura, Toshihiko Yamasaki |
Multim. Tools Appl. | 4 |
| 2024 | Waveform-Domain Speech Enhancement Using Spectrogram Encoding for Robust Speech RecognitionabstractWhile waveform-domain speech enhancement (SE) has been extensively investigated in recent years and achieves state-of-the-art performance in many datasets, spectrogram-based SE tends to show robust and stable enhancement behavior. In this paper, we propose a waveform-spectrogram hybrid method (WaveSpecEnc) to improve the robustness of waveform-domain SE. WaveSpecEnc refines the corresponding temporal feature map by spectrogram encoding in each encoder layer. Incorporating spectral information provides robust human hearing experience performance. However, it has a minor automatic speech recognition (ASR) improvement. Thus, we improve it for robust ASR by further utilizing spectrogram encoding information (WaveSpecEnc+) to both the SE front-end and ASR back-end. Experimental results using the CHiME-4 dataset show that ASR performance in real evaluation sets is consistently improved with the proposed method, which outperformed others, including DEMUCS and Conv-Tasnet. Refining in the shallow encoder layers is very effective, and the effect is confirmed even with a strong ASR baseline using WavLM. Masato Mimura, Tatsuya Kawahara |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2024 | Refining Synthesized Speech Using Speaker Information and Phone Masking for Data Augmentation of Speech RecognitionabstractWhile end-to-end automatic speech recognition (ASR) has shown impressive performance, it requires a huge amount of speech and transcription data. The conversion of domain-matched text to speech (TTS) has been investigated as one approach to data augmentation. The quality and diversity of the synthesized speech are critical in this approach. To ensure quality, a neural vocoder is widely used to generate speech waveforms in conventional studies, but it requires a huge amount of computation and another conversion to spectral-domain features such as the log-Mel filterbank (lmfb) output typically used for ASR. In this study, we explore the direct refinement of these features. Unlike conventional speech enhancement, we can use information on the ground-truth phone sequences of the speech and designated speaker to improve the quality and diversity. This process is realized as a Mel-to-Mel network, which can be placed after a text-to-Mel synthesis system such as FastSpeech 2. These two networks can be trained jointly. Moreover, semantic masking is applied to the lmfb features for robust training. Experimental evaluations demonstrate the effect of phone information, speaker information, and semantic masking. For speaker information, x-vector performs better than the simple speaker embedding. The proposed method achieves even better ASR performance with a much shorter computation time than the conventional method using a vocoder. Sei Ueno, Akinobu Lee, Tatsuya Kawahara |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Domain and Language Adaptation Using Heterogeneous Datasets for Wav2vec2.0-Based Speech Recognition of Low-Resource LanguageabstractWe address the effective finetuning of a large-scale pretrained model for automatic speech recognition (ASR) of lowresource languages with only a one-hour matched dataset. The finetuning is composed of domain adaptation and language adaptation, and they are conducted by using heterogeneous datasets, which are matched with either domain or language. For effective adaptation, we incorporate auxiliary tasks of domain identification and language identification with multi-task learning. Moreover, the embedding result of the auxiliary tasks is fused to the encoder output of the pretrained model for ASR. Experimental evaluations on the Khmer ASR using the corpus of ECCC (the Extraordinary Chambers in the Courts of Cambodia) demonstrate that first conducting domain adaption and then language adaption is effective. In addition, multi-tasking with domain identification and fusing the domain ID embedding gives the best performance, which is a CER improvement of 6.47% absolute from the baseline finetuning method. Soky Kak, Sheng Li 0010, Chenhui Chu, Tatsuya Kawahara |
ICASSP | 4 |
| 2023 | Time-Domain Speech Enhancement Assisted by Multi-Resolution Frequency Encoder and DecoderabstractTime-domain speech enhancement (SE) has recently been intensively investigated. Among recent works, DEMUCS [1] introduces multi-resolution STFT loss to enhance performance. However, some resolutions used for STFT contain non-stationary signals, and it is challenging to learn multi-resolution frequency losses simultaneously with only one output. For better use of multi-resolution frequency information, we supplement multiple spectrograms in different frame lengths into the time-domain encoders. They extract stationary frequency information in both narrowband and wideband. We also adopt multiple decoder outputs, each of which computes its corresponding resolution frequency loss. Experimental results show that (1) it is more effective to fuse stationary frequency features than non-stationary features in the encoder, and (2) the multiple outputs consistent with the frequency loss improve performance. Experiments on the Voice-Bank dataset show that the proposed method obtained a 0.14 PESQ improvement. Masato Mimura, Longbiao Wang, Jianwu Dang 0001, Tatsuya Kawahara |
ICASSP | 5 |
| 2023 | Two-stage Finetuning of Wav2vec 2.0 for Speech Emotion Recognition with ASR and Gender Pretraining
Yuan Gao 0040, Chenhui Chu, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2023 | Embedding Articulatory Constraints for Low-resource Speech Recognition Based on Large Pre-trained Model
Masato Mimura, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2023 | RealPersonaChat: A Realistic Persona Chat Corpus with Interlocutors' Own Personalities
Sanae Yamashita, Koji Inoue, Shota Mochizuki, Tatsuya Kawahara, Ryuichiro Higashinaka |
PACLIC | 5 |
| 2023 | Robotic Backchanneling in Online Conversation Facilitation: A Cross-Generational StudyabstractJapan faces many challenges related to its aging society, including increasing rates of cognitive decline in the population and a shortage of caregivers. Efforts have begun to explore solutions using artificial intelligence (AI), especially socially embodied intelligent agents and robots that can communicate with people. Yet, there has been little research on the compatibility of these agents with older adults in various everyday situations. To this end, we conducted a user study to evaluate a robot that functions as a facilitator for a group conversation protocol designed to prevent cognitive decline. We modified the robot to use backchannelling, a natural human way of speaking, to increase receptiveness of the robot and enjoyment of the group conversation experience. We conducted a cross-generational study with young adults and older adults. Qualitative analyses indicated that younger adults perceived the backchannelling version of the robot as kinder, more trustworthy, and more acceptable than the non-backchannelling robot. Finally, we found that the robot’s backchannelling elicited nonverbal backchanneling in older participants. Sota Kobuki, Katie Seaborn, Seiki Tokunaga, Kosuke Fukumori, Shun Hidaka, Kazuhiro Tamura, Koji Inoue, Tatsuya Kawahara, Mihoko Otake |
RO-MAN | 8 |
| 2023 | Reasoning before Responding: Integrating Commonsense-based Causality Explanation for Empathetic Response GenerationabstractRecent approaches to empathetic response generation try to incorporate commonsense knowledge or reasoning about the causes of emotions to better understand the user's experiences and feelings.However, these approaches mainly focus on understanding the causalities of context from the user's perspective, ignoring the system's perspective.In this paper, we propose a commonsense-based causality explanation approach for diverse empathetic response generation that considers both the user's perspective (user's desires and reactions) and the system's perspective (system's intentions and reactions).We enhance ChatGPT's ability to reason for the system's perspective by integrating in-context learning with commonsense knowledge.Then, we integrate the commonsense-based causality explanation with both ChatGPT and a T5-based model.Experimental evaluations demonstrate that our method outperforms other comparable methods on both automatic and human evaluations. Yahui Fu 0001, Koji Inoue, Chenhui Chu, Tatsuya Kawahara |
SIGDIAL | 4 |
| 2023 | Character expression for spoken dialogue systems with semi-supervised learning using Variational Auto-EncoderabstractCharacter of spoken dialogue systems is important not only for giving a positive impression of the system but also for gaining rapport from users. We have proposed a character expression model for spoken dialogue systems. The model expresses three character traits (extroversion, emotional instability, and politeness) of spoken dialogue systems by controlling spoken dialogue behaviors: utterance amount, backchannel, filler, and switching pause length. One major problem in training this model is that it is costly and time-consuming to collect many pair data of character traits and behaviors. To address this problem, semi-supervised learning is proposed based on a variational auto-encoder that exploits both the limited amount of labeled pair data and unlabeled corpus data. It was confirmed that the proposed model can express given characters more accurately than a baseline model with only supervised learning. We also implemented the character expression model in a spoken dialogue system for an autonomous android robot, and then conducted a subjective experiment with 75 university students to confirm the effectiveness of the character expression for specific dialogue scenarios. The results showed that expressing a character in accordance with the dialogue task by the proposed model improves the user’s impression of the appropriateness in formal dialogue such as job interview. Kenta Yamamoto, Koji Inoue, Tatsuya Kawahara |
Comput. Speech Lang. | 3 |
| 2023 | Alignment Knowledge Distillation for Online Streaming Attention-Based Speech RecognitionabstractThis article describes an efficient training method for online streaming attention-based encoder-decoder (AED) automatic speech recognition (ASR) systems. AED models have achieved competitive performance in offline scenarios by jointly optimizing all components. They have recently been extended to an online streaming framework via models such as monotonie chunkwise attention (MoChA). However, the elaborate attention calculation process is not robust against long-form speech utterances. Moreover, the sequence-level training objective and time-restricted streaming encoder cause a nonnegligible delay in token emission during inference. To address these problems, we propose CTC synchronous training (CTC-ST), in which CTC alignments are leveraged as a reference for token boundaries to enable a MoChA model to learn optimal monotonie input-output alignments. We formulate a purely end-to-end training objective to synchronize the boundaries of MoChA to those of CTC. The CTC model shares an encoder with the MoChA model to enhance the encoder representation. Moreover, the proposed method provides alignment information learned in the CTC branch to the attention-based decoder. Therefore, CTC-ST can be regarded as self-distillation of alignment knowledge from CTC to MoChA. Experimental evaluations on a variety of benchmark datasets show that the proposed method significantly reduces recognition errors and emission latency simultaneously. The robustness to long-form and noisy speech is also demonstrated. We compare CTC-ST with several methods that distill alignment knowledge from a hybrid ASR system and show that the CTC-ST can achieve a comparable tradeoff of accuracy and latency without relying on external alignment information. Hirofumi Inaguma, Tatsuya Kawahara |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Backchannel Generation Model for a Third Party Listener AgentabstractIn this work we propose a listening agent which can be used in a conversation between two humans. We firstly conduct a corpus analysis to identify three different categories of backchannel which the agent can use - responsive interjections, expressive interjections and shared laughs. From this data we train and evaluate a continuous backchannel generation model consisting of separate timing and form prediction models. We then conduct a subjective experiment to compare our model to random, dyadic, and ground truth models. We find that our model outperforms a random baseline and is comparable to the dyadic model despite the low evaluation of expressive interjections. We suggest that the perception of expressive interjections contribute significantly to the perception of the agent’s empathy and understanding of the conversation. The results also show the need for a more robust model to generate expressive interjections, perhaps aided by the use of linguistic features. Divesh Lala, Koji Inoue, Tatsuya Kawahara, Kei Sawada |
HAI | 3 |
| 2022 | Alzheimer's Dementia Detection through Spontaneous Dialogue with Proactive Robotic ListenersabstractAs the aging of society continues to accelerate, Alzheimer's Disease (AD) has received more and more attention from not only medical but also other fields, such as computer science, over the past decade. Since speech is considered one of the effective ways to diagnose cognitive decline, AD detection from speech has emerged as a hot topic. Nevertheless, such approaches fail to tackle several key issues: 1) AD is a complex neurocognitive disorder which means it is inappropriate to conduct AD detection using utterance information alone while ignoring dialogue infor-mation; 2) Utterances of AD patients contain many disfluencies that affect speech recognition yet are helpful to diagnosis; 3) AD patients tend to speak less, causing dialogue breakdown as the disease progresses. This fact leads to a small number of utterances, which may cause detection bias. Therefore, in this paper, we propose a novel AD detection architecture consisting of two major modules: an ensemble AD detector and a proactive listener. This architecture can be embedded in the dialogue system of conversational robots for healthcare. Yuanchao Li, Catherine Lai, Divesh Lala, Koji Inoue, Tatsuya Kawahara |
HRI | 5 |
| 2022 | Phone-Informed Refinement of Synthesized Mel Spectrogram for Data Augmentation in Speech RecognitionabstractWhile recent end-to-end automatic speech recognition (ASR) models achieve high performance, we need to prepare an abundant amount of training data, which is a barrier to apply them to a specific domain. To mitigate the lack of training data, text-to-speech (TTS) systems have been utilized to leverage text-only data to efficiently generate paired data for training the ASR model. The widely-used procedure first generates a Mel spectrogram from text data, then converts it into a waveform, and converts it again to a Mel spectrogram. The vocoder is often used to alleviate the difference between real and synthesized speech, but it requires a huge amount of run-time. In this work, we propose a phone-informed post-processing network that refines Mel spectrograms without using the vocoder. The proposed network consumes not only Mel spectrograms but also text information to use phone sequence information for refinement. Experimental evaluations demonstrate that the proposed network achieves better WERs than the vocoder network in an English domain adaptation task (LibriSpeech to TED-LIUM 2; read speech to spontaneous speech) in a much smaller amount of data generation time. It is also shown the use of phone information is critical for the improvement. We also confirm the effect of the proposed model in a Japanese domain adaptation task (CSJ-SPS to CSJ-APS; everyday topic to academic topic). Sei Ueno, Tatsuya Kawahara |
ICASSP | 2 |
| 2022 | Selective Multi-Task Learning For Speech Emotion Recognition Using Corpora Of Different StylesabstractWhile speech emotion recognition (SER) has been actively studied, the amount and variations of training data are limited compared with speech recognition and speaker recognition tasks. Therefore, it is promising to combine multiple corpora to train a generalized SER model. However, the manner of emotion expression is different according to the settings, task domains, and languages. In particular, there is a mismatch between acted datasets and spontaneous datasets since the former includes much more rich and explicit emotion expressions than the latter. In this paper, we investigate effective combination methods based on multi-task learning (MTL) considering the style attribute. We also hypothesize the neutral expression, which has the largest number of samples, is not affected by the style, and thus propose a selective MTL method that applies MTL to emotion categories except for the neutral category. Experimental evaluations using the IEMOCAP database and a call center dataset confirm the effect of the combination of the two corpora, MTL, and the proposed selective MTL. Heran Zhang, Masato Mimura, Tatsuya Kawahara, Kenkichi Ishizuka |
ICASSP | 3 |
| 2022 | Non-autoregressive Error Correction for CTC-based ASR with Phone-conditioned Masked LM
Hayato Futami, Hirofumi Inaguma, Sei Ueno, Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
INTERSPEECH | 6 |
| 2022 | Leveraging Simultaneous Translation for Enhancing Transcription of Low-resource Language via Cross Attention Mechanism
Soky Kak, Sheng Li 0010, Masato Mimura, Chenhui Chu, Tatsuya Kawahara |
INTERSPEECH | 5 |
| 2022 | Multimodal Persuasive Dialogue Corpus using Teleoperated Android
Seiya Kawano, Muteki Arioka, Akishige Yuguchi, Kenta Yamamoto, Koji Inoue, Tatsuya Kawahara, Satoshi Nakamura 0001, Koichiro Yoshino |
INTERSPEECH | 6 |
| 2022 | End-to-end Speech-to-Punctuated-Text RecognitionabstractConventional automatic speech recognition systems do not produce punctuation marks which are important for the readability of the speech recognition results.They are also needed for subsequent natural language processing tasks such as machine translation.There have been a lot of works on punctuation prediction models that insert punctuation marks into speech recognition results as post-processing.However, these studies do not utilize acoustic information for punctuation prediction and are directly affected by speech recognition errors.In this study, we propose an end-to-end model that takes speech as input and outputs punctuated texts.This model is expected to predict punctuation robustly against speech recognition errors while using acoustic information.We also propose to incorporate an auxiliary loss to train the model using the output of the intermediate layer and unpunctuated texts.Through experiments, we compare the performance of the proposed model to that of a cascaded system.The proposed model achieves higher punctuation prediction accuracy than the cascaded system without sacrificing the speech recognition error rate.It is also demonstrated that the multi-task learning using the intermediate output against the unpunctuated text is effective.Moreover, the proposed model has only about 1/7th of the parameters compared to the cascaded system. Jumon Nozaki, Tatsuya Kawahara, Kenkichi Ishizuka, Taiichi Hashimoto |
INTERSPEECH | 2 |
| 2022 | Monaural Speech Enhancement Based on Spectrogram Decomposition for Convolutional Neural Network-sensitive Feature Extraction
Longbiao Wang, Sheng Li 0010, Jianwu Dang 0001, Tatsuya Kawahara |
INTERSPEECH | 5 |
| 2022 | Simultaneous Job Interview System Using Multiple Semi-autonomous AgentsabstractIn recent years, spoken dialogue systems have been used in job interviews where an applicant talks to a system that asks pre-defined questions, called on-demand and self-paced job interviews.We propose a simultaneous job interview system, where one interviewer can conduct one-on-one interviews with multiple applicants simultaneously by cooperating with multiple autonomous interview dialogue systems.However, it is challenging for interviewers to monitor and understand all parallel interviews done by the autonomous system simultaneously.To address this issue, we implement two automatic dialogue understanding functions: (1) response evaluation of each applicant's responses and (2) keyword extraction for a summary of the responses.In this system, interviewers can intervene in a dialogue session when needed and smoothly ask a proper question that elaborates the interview.We have conducted a pilot experiment where an interviewer conducted simultaneous job interviews with three candidates. Haruki Kawai, Yusuke Muraki, Kenta Yamamoto, Divesh Lala, Koji Inoue, Tatsuya Kawahara |
SIGDIAL | 6 |
| 2022 | Computationally-Efficient Overdetermined Blind Source Separation Based on Iterative Source SteeringabstractThis paper describesa computationally-efficient optimization algorithm for the blind source separation (BSS) of overdetermined mixtures. In the determined case, a matrix-inversion-free iterative source steering (ISS) algorithm has been proposed for estimating a square demixing matrix as a computationally-efficient alternative to the popular iterative projection (IP) algorithm. The IP algorithm is based on source-wise (i.e., row-wise) updates of the demixing matrix, and lends itself naturally to an extension to overdetermined independent vector analysis (IVA) called OverIVA. In contrast, the ISS algorithm changes the whole demixing matrix at every update, making its extension to the overdetermined case non-trivial. In this paper, we propose a modified ISS algorithm for OverIVA fully exploiting the computational savings of ISS. We also derive an overdetermined extension of independent low-rank matrix analysis (OverILRMA) with the modified ISS algorithm. Experimental results showed that the proposed ISS-based OverIVA and OverILRMA were comparable or superior to the conventional IP-based counterparts in speech separation performance while achieving lower computational cost. Yicheng Du, Robin Scheibler, Masahito Togami, Kazuyoshi Yoshii, Tatsuya Kawahara |
IEEE Signal Process. Lett. | 5 |
| 2022 | Autoregressive Moving Average Jointly-Diagonalizable Spatial Covariance Analysis for Joint Source Separation and DereverberationabstractThis paper describes a computationally-efficient statistical approach to joint (semi-)blind source separation and dereverberation for multichannel noisy reverberant mixture signals. A standard approach to source separation is to formulate a generative model of a multichannel mixture spectrogram that consists of source and spatial models representing the time-frequency power spectral densities (PSDs) and spatial covariance matrices (SCMs) of source images, respectively, and find the maximum-likelihood estimates of these parameters. A state-of-the-art blind source separation method in this thread of research is fast multichannel nonnegative matrix factorization (FastMNMF) based on the low-rank PSDs and jointly-diagonalizable full-rank SCMs. To perform mutually-dependent separation and dereverberation jointly, in this paper we integrate both moving average (MA) and autoregressive (AR) models that represent the early reflections and late reverberations of sources, respectively, into the FastMNMF formalism. Using a pretrained deep generative model of speech PSDs as a source model, we realize semi-blind joint speech separation and dereverberation. We derive an iterative optimization algorithm based on iterative projection or iterative source steering for jointly and efficiently updating the AR parameters and the SCMs. Our experimental results showed the superiority of the proposed ARMA extension over its AR- or MA-ablated version in a speech separation and/or dereverberation task. Kouhei Sekiguchi, Yoshiaki Bando, Aditya Arie Nugraha, Mathieu Fontaine 0002, Kazuyoshi Yoshii, Tatsuya Kawahara |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2021 | ASR Rescoring and Confidence Estimation with ElectraabstractIn automatic speech recognition (ASR) rescoring, the hypothesis with the fewest errors should be selected from the$n$-best list using a language model (LM). However, LMs are usually trained to maximize the likelihood of correct word sequences, not to detect ASR errors. We propose an ASR rescoring method for directly detecting errors with ELECTRA, which is originally a pre-training method for NLP tasks. ELECTRA is pre-trained to predict whether each word is replaced by BERT or not, which can simulate ASR error detection on large text corpora. To make this pre-training closer to ASR error detection, we further propose an extended version of ELECTRA called phone-attentive ELECTRA (P-ELECTRA). In the pre-training of P-ELECTRA, each word is replaced by a phone-to-word conversion model, which leverages phone information to generate acoustically similar words. Since our rescoring method is optimized for detecting errors, it can also be used for word-level confidence estimation. Experimental evaluations on the Librispeech and TED-LIUM2 corpora show that our rescoring method with ELECTRA is competitive with conventional rescoring methods with faster inference. ELECTRA also performs better in confidence estimation than BERT because it can learn to detect inappropriate words not only in fine-tuning but also in pre-training. Hayato Futami, Hirofumi Inaguma, Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
ASRU | 5 |
| 2021 | Data Augmentation for ASR Using TTS Via a Discrete RepresentationabstractWhile end-to-end automatic speech recognition (ASR) has achieved high performance, it requires a huge amount of paired speech and transcription data for training. Recently, data augmentation methods have actively been investigated. One method is to use a text-to-speech (TTS) system to gen-erate speech data from text-only data and use the generated speech for data augmentation, but it has been found that the synthesized log Mel-scale filterbank (lmfb) features could have a serious mismatch with the real speech features. In this study, we propose a data augmentation method via a discrete speech representation. The TTS model predicts discrete ID sequences instead of lmfb features, and the ASR also uses the ID sequences as training data. We expect that the use of a discrete representation based on vq-wav2vec not only makes TTS training easier but also mitigates the mismatch with real data. Experimental evaluations show that the pro-posed method outperforms the data augmentation method using the conventional TTS. We found that it reduces speaker dependency, and the generated features are distributed more closely to the real ones. Sei Ueno, Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
ASRU | 4 |
| 2021 | ORTHROS: non-autoregressive end-to-end speech translation With dual-decoderabstractFast inference speed is an important goal towards real-world deployment of speech translation (ST) systems. End-to-end (E2E) models based on the encoder-decoder architecture are more suitable for this goal than traditional cascaded systems, but their effectiveness regarding decoding speed has not been explored so far. Inspired by recent progress in non-autoregressive (NAR) methods in text-based translation, which generates target tokens in parallel by eliminating conditional dependencies, we study the problem of NAR decoding for E2E-ST. We propose a novel NAR E2E-ST framework, Orthros, in which both NAR and autoregressive (AR) decoders are jointly trained on the shared speech encoder. The latter is used for selecting better translation among various length candidates generated from the former, which dramatically improves the effectiveness of a large length beam with negligible overhead. We further investigate effective length prediction methods from speech inputs and the impact of vocabulary sizes. Experiments on four benchmarks show the effectiveness of the proposed method in improving inference speed while maintaining competitive translation quality compared to state-of-the-art AR E2E-ST systems. Hirofumi Inaguma, Yosuke Higuchi, Kevin Duh, Tatsuya Kawahara, Shinji Watanabe 0001 |
ICASSP | 4 |
| 2021 | StableEmit: Selection Probability Discount for Reducing Emission Latency of Streaming Monotonic Attention ASRabstractWhile attention-based encoder-decoder (AED) models have been successfully extended to the online variants for streaming automatic speech recognition (ASR), such as monotonic chunkwise attention (MoChA), the models still have a large label emission latency because of the unconstrained end-to-end training objective. Previous works tackled this problem by leveraging alignment information to control the timing to emit tokens during training. In this work, we propose a simple alignment-free regularization method, StableEmit, to encourage MoChA to emit tokens earlier. StableEmit discounts the selection probabilities in hard monotonic attention for token boundary detection by a constant factor and regularizes them to recover the total attention mass during training. As a result, the scale of the selection probabilities is increased, and the values can reach a threshold for token emission earlier, leading to a reduction of emission latency and deletion errors. Moreover, StableEmit can be combined with methods that constraint alignments to further improve the accuracy and latency. Experimental evaluations with LSTM and Conformer encoders demonstrate that StableEmit significantly reduces the recognition errors and the emission latency simultaneously. We also show that the use of alignment information is complementary in both metrics. Hirofumi Inaguma, Tatsuya Kawahara |
Interspeech | 2 |
| 2021 | VAD-Free Streaming Hybrid CTC/Attention ASR for Unsegmented RecordingabstractIn this work, we propose novel decoding algorithms to enable streaming automatic speech recognition (ASR) on unsegmented long-form recordings without voice activity detection (VAD), based on monotonic chunkwise attention (MoChA) with an auxiliary connectionist temporal classification (CTC) objective. We propose a block-synchronous beam search decoding to take advantage of efficient batched output-synchronous and low-latency input-synchronous searches. We also propose a VAD-free inference algorithm that leverages CTC probabilities to determine a suitable timing to reset the model states to tackle the vulnerability to long-form data. Experimental evaluations demonstrate that the block-synchronous decoding achieves comparable accuracy to the label-synchronous one. Moreover, the VAD-free inference can recognize long-form speech robustly for up to a few hours. Hirofumi Inaguma, Tatsuya Kawahara |
Interspeech | 2 |
| 2021 | Source and Target Bidirectional Knowledge Distillation for End-to-end Speech TranslationabstractHirofumi Inaguma, Tatsuya Kawahara, Shinji Watanabe. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Hirofumi Inaguma, Tatsuya Kawahara, Shinji Watanabe 0001 |
NAACL-HLT | 2 |
| 2021 | A multi-party attentive listening robot which stimulates involvement from side participantsabstractWe demonstrate the moderating abilities of a multi-party attentive listening robot system when multiple people are speaking in turns.Our conventional one-on-one attentive listening system generates listener responses such as backchannels, repeats, elaborating questions, and assessments.In this paper, additional robot responses that stimulate a listening user (side participant) to become more involved in the dialogue are proposed.The additional responses elicit assessments and questions from the side participant, making the dialogue more empathetic and lively. Koji Inoue, Hiromi Sakamoto, Kenta Yamamoto, Divesh Lala, Tatsuya Kawahara |
SIGDIAL | 5 |
| 2021 | ERICA: An Empathetic Android Companion for Covid-19 QuarantineabstractOver the past year, research in various domains, including Natural Language Processing (NLP), has been accelerated to fight against the COVID-19 pandemic, yet such research has just started on dialogue systems.In this paper, we introduce an end-to-end dialogue system which aims to ease the isolation of people under self-quarantine.We conduct a control simulation experiment to assess the effects of the user interface, a web-based virtual agent called Nora vs. the android ERICA via a video call.The experimental results show that the android offers a more valuable user experience by giving the impression of being more empathetic and engaging in the conversation due to its nonverbal information, such as facial expressions and body gestures.Demo video available at https://youtu.be/PLPEBXLeKJI. Etsuko Ishii, Genta Indra Winata, Samuel Cahyawijaya, Divesh Lala, Tatsuya Kawahara, Pascale Fung |
SIGDIAL | 5 |
| 2021 | Multi-Referenced Training for Dialogue Response GenerationabstractIn open-domain dialogue response generation, a dialogue context can be continued with diverse responses, and the dialogue models should capture such one-to-many relations.In this work, we first analyze the training objective of dialogue models from the view of Kullback-Leibler divergence (KLD) and show that the gap between the real world probability distribution and the single-referenced data's probability distribution prevents the model from learning the one-to-many relations efficiently.Then we explore approaches to multireferenced training in two aspects.Data-wise, we generate diverse pseudo references from a powerful pretrained model to build multireferenced data that provides a better approximation of the real-world distribution.Modelwise, we propose to equip variational models with an expressive prior, named linear Gaussian model (LGM).Experimental results of automated evaluation and human evaluation show that the methods yield significant improvements over baselines.1 Tianyu Zhao 0001, Tatsuya Kawahara |
SIGDIAL | 2 |
| 2020 | Designing Precise and Robust Dialogue Response EvaluatorsabstractAutomatic dialogue response evaluator has been proposed as an alternative to automated metrics and human evaluation.However, existing automatic evaluators achieve only moderate correlation with human judgement and they are not robust.In this work, we propose to build a reference-free evaluator and exploit the power of semi-supervised training and pretrained (masked) language models.Experimental results demonstrate that the proposed evaluator achieves a strong correlation (> 0.6) with human judgement and generalizes robustly to diverse responses and corpora.We open-source the code and data in https://github.com/ ZHAOTING/dialog-processing. Tianyu Zhao 0001, Divesh Lala, Tatsuya Kawahara |
ACL | 3 |
| 2020 | Topic-relevant Response Generation using Optimal Transport for an Open-domain Dialog SystemabstractConventional neural generative models tend to generate safe and generic responses which have little connection with previous utterances semantically and would disengage users in a dialog system.To generate relevant responses, we propose a method that employs two types of constraints -topical constraint and semantic constraint.Under the hypothesis that a response and its context have higher relevance when they share the same topics, the topical constraint encourages the topics of a response to match its context by conditioning response decoding on topic words' embeddings.The semantic constraint, which encourages a response to be semantically related to its context by regularizing the decoding objective function with semantic distance, is proposed.Optimal transport is applied to compute a weighted semantic distance between the representation of a response and the context.Generated responses are evaluated by automatic metrics, as well as human judgment, showing that the proposed method can generate more topic-relevant and content-rich responses than conventional models. Shuying Zhang, Tianyu Zhao 0001, Tatsuya Kawahara |
COLING | 3 |
| 2020 | Job Interviewer Android with Elaborate Follow-up Question GenerationabstractA job interview is a domain that takes advantage of an android robot's human-like appearance and behaviors. In this work, our goal is to implement a system in which an android plays the role of an interviewer so that users may practice for a real job interview. Our proposed system generates elaborate follow-up questions based on responses from the interviewee. We conducted an interactive experiment to compare the proposed system against a baseline system that asked only fixed-form questions. We found that this system was significantly better than the baseline system with respect to the impression of the interview and the quality of the questions, and that the presence of the android interviewer was enhanced by the follow-up questions. We also found a similar result when using a virtual agent interviewer, except that presence was not enhanced. Koji Inoue, Kohei Hara, Divesh Lala, Kenta Yamamoto, Shizuka Nakamura, Katsuya Takanashi, Tatsuya Kawahara |
ICMI | 7 |
| 2020 | End-to-End Speech-to-Dialog-Act RecognitionabstractSpoken language understanding, which extracts intents and/or semantic concepts in utterances, is conventionally formulated as a post-processing of automatic speech recognition.It is usually trained with oracle transcripts, but needs to deal with errors by ASR.Moreover, there are acoustic features which are related with intents but not represented with the transcripts.In this paper, we present an end-to-end model that directly converts speech into dialog acts without the deterministic transcription process.In the proposed model, the dialog act recognition network is conjunct with an acoustic-to-word ASR model at its latent layer before the softmax layer, which provides a distributed representation of word-level ASR decoding information.Then, the entire network is fine-tuned in an end-to-end manner.This allows for stable training as well as robustness against ASR errors.The model is further extended to conduct DA segmentation jointly.Evaluations with the Switchboard corpus demonstrate that the proposed method significantly improves dialog act recognition accuracy from the conventional pipeline framework. Viet-Trung Dang, Tianyu Zhao 0001, Sei Ueno, Hirofumi Inaguma, Tatsuya Kawahara |
INTERSPEECH | 5 |
| 2020 | End-to-End Speech Emotion Recognition Combined with Acoustic-to-Word ASR Model
Sei Ueno, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2020 | Distilling the Knowledge of BERT for Sequence-to-Sequence ASRabstractAttention-based sequence-to-sequence (seq2seq) models have achieved promising results in automatic speech recognition (ASR).However, as these models decode in a left-to-right way, they do not have access to context on the right.We leverage both left and right context by applying BERT as an external language model to seq2seq ASR through knowledge distillation.In our proposed method, BERT generates soft labels to guide the training of seq2seq ASR.Furthermore, we leverage context beyond the current utterance as input to BERT.Experimental evaluations show that our method significantly improves the ASR performance from the seq2seq baseline on the Corpus of Spontaneous Japanese (CSJ).Knowledge distillation from BERT outperforms that from a transformer LM that only looks at left context.We also show the effectiveness of leveraging context beyond the current utterance.Our method outperforms other LM application approaches such as n-best rescoring and shallow fusion, while it does not require extra inference cost. Hayato Futami, Hirofumi Inaguma, Sei Ueno, Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
INTERSPEECH | 6 |
| 2020 | CTC-Synchronous Training for Monotonic Attention ModelabstractMonotonic chunkwise attention (MoChA) has been studied for the online streaming automatic speech recognition (ASR) based on a sequence-to-sequence framework. In contrast to connectionist temporal classification (CTC), backward probabilities cannot be leveraged in the alignment marginalization process during training due to left-to-right dependency in the decoder. This results in the error propagation of alignments to subsequent token generation. To address this problem, we propose CTC-synchronous training (CTC-ST), in which MoChA uses CTC alignments to learn optimal monotonic alignments. Reference CTC alignments are extracted from a CTC branch sharing the same encoder with the decoder. The entire model is jointly optimized so that the expected boundaries from MoChA are synchronized with the alignments. Experimental evaluations of the TEDLIUM release-2 and Librispeech corpora show that the proposed method significantly improves recognition, especially for long utterances. We also show that CTC-ST can bring out the full potential of SpecAugment for MoChA. Hirofumi Inaguma, Masato Mimura, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2020 | Enhancing Monotonic Multihead Attention for Streaming ASRabstractWe investigate a monotonic multihead attention (MMA) by extending hard monotonic attention to Transformer-based automatic speech recognition (ASR) for online streaming applications. For streaming inference, all monotonic attention (MA) heads should learn proper alignments because the next token is not generated until all heads detect the corresponding token boundaries. However, we found not all MA heads learn alignments with a naïve implementation. To encourage every head to learn alignments properly, we propose HeadDrop regularization by masking out a part of heads stochastically during training. Furthermore, we propose to prune redundant heads to improve consensus among heads for boundary detection and prevent delayed token generation caused by such heads. Chunkwise attention on each MA head is extended to the multihead counterpart. Finally, we propose head-synchronous beam search decoding to guarantee stable streaming inference. Hirofumi Inaguma, Masato Mimura, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2020 | Generative Adversarial Training Data Adaptation for Very Low-Resource Automatic Speech RecognitionabstractIt is important to transcribe and archive speech data of endangered languages for preserving heritages of verbal culture and automatic speech recognition (ASR) is a powerful tool to facilitate this process. However, since endangered languages do not generally have large corpora with many speakers, the performance of ASR models trained on them are considerably poor in general. Nevertheless, we are often left with a lot of recordings of spontaneous speech data that have to be transcribed. In this work, for mitigating this speaker sparsity problem, we propose to convert the whole training speech data and make it sound like the test speaker in order to develop a highly accurate ASR system for this speaker. For this purpose, we utilize a CycleGAN-based non-parallel voice conversion technology to forge a labeled training data that is close to the test speaker's speech. We evaluated this speaker adaptation approach on two low-resource corpora, namely, Ainu and Mboshi. We obtained 35-60% relative improvement in phone error rate on the Ainu corpus, and 40% relative improvement was attained on the Mboshi corpus. This approach outperformed two conventional methods namely unsupervised adaptation and multilingual training with these two corpora. Kohei Matsuura, Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
INTERSPEECH | 4 |
| 2020 | Semi-Supervised Learning for Character Expression of Spoken Dialogue Systems
Kenta Yamamoto, Koji Inoue, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2020 | Speech Corpus of Ainu Folklore and End-to-end Speech Recognition for Ainu LanguageabstractAinu is an unwritten language that has been spoken by Ainu people who are one of the ethnic groups in Japan. It is recognized as critically endangered by UNESCO and archiving and documentation of its language heritage is of paramount importance. Although a considerable amount of voice recordings of Ainu folklore has been produced and accumulated to save their culture, only a quite limited parts of them are transcribed so far. Thus, we started a project of automatic speech recognition (ASR) for the Ainu language in order to contribute to the development of annotated language archives. In this paper, we report speech corpus development and the structure and performance of end-to-end ASR for Ainu. We investigated four modeling units (phone, syllable, word piece, and word) and found that the syllable-based model performed best in terms of both word and phone recognition accuracy, which were about 60% and over 85% respectively in speaker-open condition. Furthermore, word and phone accuracy of 80% and 90% has been achieved in a speaker-closed setting. We also found out that a multilingual ASR training with additional speech corpora of English and Japanese further improves the speaker-open test accuracy. Kohei Matsuura, Sei Ueno, Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
LREC | 5 |
| 2020 | An Attentive Listening System with Android ERICA: Comparison of Autonomous and WOZ InteractionsabstractWe describe an attentive listening system for the autonomous android robot ERICA.The proposed system generates several types of listener responses: backchannels, repeats, elaborating questions, assessments, generic sentimental responses, and generic responses.In this paper, we report a subjective experiment with 20 elderly people.First, we evaluated each system utterance excluding backchannels and generic responses, in an offline manner.It was found that most of the system utterances were linguistically appropriate, and they elicited positive reactions from the subjects.Furthermore, 58.2% of the responses were acknowledged as being appropriate listener responses.We also compared the proposed system with a WOZ system where a human operator was operating the robot.From the subjective evaluation, the proposed system achieved comparable scores in basic skills of attentive listening such as encouragement to talk, focused on the talk, and actively listening.It was also found that there is still a gap between the system and the WOZ for more sophisticated skills such as dialogue understanding, showing interest, and empathy towards the user. Koji Inoue, Divesh Lala, Kenta Yamamoto, Shizuka Nakamura, Katsuya Takanashi, Tatsuya Kawahara |
SIGdial | 6 |
| 2020 | Cross-Lingual Transfer Learning of Non-Native Acoustic Modeling for Pronunciation Error Detection and DiagnosisabstractIn computer-assisted pronunciation training (CAPT), the scarcity of large-scale non-native corpora and human expert annotations are two fundamental challenges to non-native acoustic modeling. Most existing approaches of acoustic modeling in CAPT are based on non-native corpora while there are so many living languages in the world. It is impractical to collect and annotate every non-native speech corpus considering different language pairs. In this work, we address non-native acoustic modeling (both on phonetic and articulatory level) based on transfer learning. In order to effectively train acoustic models of non-native speech without using such data, we propose to exploit two large native speech corpora of learner's native language (L1) and target language (L2) to model cross-lingual phenomena. This kind of transfer learning can provide a better feature representation of non-native speech. Experimental evaluations are carried out for Japanese speakers learning English. We first demonstrate the proposed acoustic-phone model achieves a lower word error rate in non-native speech recognition. It also improves the pronunciation error detection based on goodness of pronunciation (GOP) score. For diagnosis of pronunciation errors, the proposed acoustic-articulatory modeling method is effective for providing detailed feedback at the articulation level. Richeng Duan, Tatsuya Kawahara, Masatake Dantsuji, Hiroaki Nanjo |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Fast Multichannel Nonnegative Matrix Factorization With Directivity-Aware Jointly-Diagonalizable Spatial Covariance Matrices for Blind Source SeparationabstractThis article describes a computationally-efficient blind source separation (BSS) method based on the independence, low-rankness, and directivity of the sources. A typical approach to BSS is unsupervised learning of a probabilistic model that consists of a source model representing the time-frequency structure of source images and a spatial model representing their inter-channel covariance structure. Building upon the low-rank source model based on nonnegative matrix factorization (NMF), which has been considered to be effective for inter-frequency source alignment, multichannel NMF (MNMF) assumes source images to follow multivariate complex Gaussian distributions with unconstrained full-rank spatial covariance matrices (SCMs). An effective way of reducing the computational cost and initialization sensitivity of MNMF is to restrict the degree of freedom of SCMs. While a variant of MNMF called independent low-rank matrix analysis (ILRMA) severely restricts SCMs to rank-1 matrices under an idealized condition that only directional and less-echoic sources exist, we restrict SCMs to jointly-diagonalizable yet full-rank matrices in a frequency-wise manner, resulting in FastMNMF1. To help inter-frequency source alignment, we then propose FastMNMF2 that shares the directional feature of each source over all frequency bins. To explicitly consider the directivity or diffuseness of each source, we also propose rank-constrained FastMNMF that enables us to individually specify the ranks of SCMs. Our experiments showed the superiority of FastMNMF over MNMF and ILRMA in speech separation and the effectiveness of the rank constraint in speech enhancement. Kouhei Sekiguchi, Yoshiaki Bando, Aditya Arie Nugraha, Kazuyoshi Yoshii, Tatsuya Kawahara |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2019 | Multilingual End-to-End Speech TranslationabstractIn this paper, we propose a simple yet effective framework for multilingual end-to-end speech translation (ST), in which speech utterances in source languages are directly translated to the desired target languages with a universal sequence-to-sequence architecture. While multilingual models have shown to be useful for automatic speech recognition (ASR) and machine translation (MT), this is the first time they are applied to the end-to-end ST problem. We show the effectiveness of multilingual end-to-end ST in two scenarios: one-to-many and many-to-many translations with publicly available data. We experimentally confirm that multilingual end-to-end ST models significantly outperform bilingual ones in both scenarios. The generalization of multilingual training is also evaluated in a transfer learning scenario to a very low-resource language pair. All of our codes and the database are publicly available to encourage further research in this emergent multilingual ST topic11Available at https://github.com/espnet/espnet.. Hirofumi Inaguma, Kevin Duh, Tatsuya Kawahara, Shinji Watanabe 0001 |
ASRU | 3 |
| 2019 | Transfer Learning of Language-independent End-to-end ASR with Language Model FusionabstractThis work explores better adaptation methods to low-resource languages using an external language model (LM) under the framework of transfer learning. We first build a language-independent ASR system in a unified sequence-to-sequence (S2S) architecture with a shared vocabulary among all languages. During adaptation, we perform LM fusion transfer, where an external LM is integrated into the decoder network of the attention-based S2S model in the whole adaptation stage, to effectively incorporate linguistic context of the target language. We also investigate various seed models for transfer learning. Experimental evaluations using the IARPA BABEL data set show that LM fusion transfer improves performances on all target five languages compared with simple transfer learning when the external text data is available. Our final system drastically reduces the performance gap from the hybrid systems. Hirofumi Inaguma, Murali Karthick Baskar, Tatsuya Kawahara, Shinji Watanabe 0001 |
ICASSP | 4 |
| 2019 | Multi-speaker Sequence-to-sequence Speech Synthesis for Data Augmentation in Acoustic-to-word Speech RecognitionabstractThe acoustic-to-word (A2W) automatic speech recognition (ASR) realizes very fast decoding with a simple architecture and achieves state-of-the-art performance. However, the A2W model suffers from the out-of-vocabulary (OOV) word problem and cannot use text-only data to improve the language modeling capability. Meanwhile, sequence-to-sequence neural speech synthesis has also been developed and achieved naturalness comparable to human speech. We investigate leveraging sequence-to-sequence neural speech synthesis to augment training data for the ASR system in a target domain. While speech synthesis model is usually trained with single speaker data, ASR needs to cover a variety of speakers. In this work, we extend the speech synthesizer so that it can output speech of many speakers. The multi-speaker speech synthesizer is trained with a large corpus in the source domain, then used to generate acoustic features from texts of the target domain. These synthesized speech features are combined with real speech features of the source domain to train an attention-based A2W model. Experimental results show that the A2W model trained with the multi-speaker model achieved a significant improvement over the baseline and the single speaker model. Sei Ueno, Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
ICASSP | 4 |
| 2019 | Smooth Turn-taking by a Robot Using an Online Continuous Model to Generate Turn-taking CuesabstractTurn-taking in human-robot interaction is a crucial part of spoken dialogue systems, but current models do not allow for human-like turn-taking speed seen in natural conversation. In this work we propose combining two independent prediction models. A continuous model predicts the upcoming end of the turn in order to generate gaze aversion and fillers as turn-taking cues. This prediction is done while the user is speaking, so turn-taking can be done with little silence between turns, or even overlap. Once a speech recognition result has been received at a later time, a second model uses the lexical information to decide if or when the turn should actually be taken. We constructed the continuous model using the speaker’s prosodic features as inputs and evaluated its online performance. We then conducted a subjective experiment in which we implemented our model in an android robot and asked participants to compare it to one without turn-taking cues, which produces a response when a speech recognition result is received. We found that using both gaze aversion and a filler was preferred when the continuous model correctly predicted the upcoming end of turn, while using only gaze aversion was better if the prediction was wrong. Divesh Lala, Koji Inoue, Tatsuya Kawahara |
ICMI | 3 |
| 2019 | ERICA and WikiTalkabstractThe demo shows ERICA, a highly realistic female android robot, and WikiTalk, an application that helps robots to talk about thousands of topics using information from Wikipedia. The combination of ERICA and WikiTalk results in more natural and engaging human-robot conversations. Divesh Lala, Graham Wilcock, Kristiina Jokinen, Tatsuya Kawahara |
IJCAI | 4 |
| 2019 | End-to-End Articulatory Attribute Modeling for Low-Resource Multilingual Speech Recognition
Sheng Li 0010, Chenchen Ding, Xugang Lu, Tatsuya Kawahara, Hisashi Kawai |
INTERSPEECH | 5 |
| 2019 | Turn-Taking Prediction Based on Detection of Transition Relevance Place
Kohei Hara, Koji Inoue, Katsuya Takanashi, Tatsuya Kawahara |
INTERSPEECH | 4 |
| 2019 | Analysis of Effect and Timing of Fillers in Natural Turn-Taking
Divesh Lala, Shizuka Nakamura, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2019 | Investigating Radical-Based End-to-End Speech Recognition Systems for Chinese Dialects and Japanese
Sheng Li 0010, Xugang Lu, Chenchen Ding, Tatsuya Kawahara, Hisashi Kawai |
INTERSPEECH | 5 |
| 2019 | Improving Transformer-Based Speech Recognition Systems with Compressed Structure and Speech Attributes Augmentation
Sheng Li 0010, Raj Dabre, Xugang Lu, Tatsuya Kawahara, Hisashi Kawai |
INTERSPEECH | 5 |
| 2019 | Improved End-to-End Speech Emotion Recognition Using Self Attention Mechanism and Multitask Learning
Yuanchao Li, Tianyu Zhao 0001, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2019 | Joint dialog act segmentation and recognition in human conversations using attention to dialog context
Tianyu Zhao 0001, Tatsuya Kawahara |
Comput. Speech Lang. | 2 |
| 2019 | Semi-Supervised Multichannel Speech Enhancement With a Deep Speech PriorabstractThis paper describes a semi-supervised multichannel speech enhancement method that uses clean speech data for prior training. Although multichannel nonnegative matrix factorization (MNMF) and its constrained variant called independent low-rank matrix analysis (ILRMA) have successfully been used for unsupervised speech enhancement, the low-rank assumption on the power spectral densities (PSDs) of all sources (speech and noise) does not hold in reality. To solve this problem, we replace a low-rank speech model with a deep generative speech model, i.e., formulate a probabilistic model of noisy speech by integrating a deep speech model, a low-rank noise model, and a full-rank or rank-1 model of spatial characteristics of speech and noise. The deep speech model is trained from clean speech data in an unsupervised auto-encoding variational Bayesian manner. Given multichannel noisy speech spectra, the full-rank or rank-1 spatial covariance matrices and PSDs of speech and noise are estimated in an unsupervised maximum-likelihood manner. Experimental results showed that the full-rank version of the proposed method was significantly better than MNMF, ILRMA, and the rank-1 version. We confirmed that the initialization-sensitivity and local-optimum problems of MNMF with many spatial parameters can be solved by incorporating the precise speech model. Kouhei Sekiguchi, Yoshiaki Bando, Aditya Arie Nugraha, Kazuyoshi Yoshii, Tatsuya Kawahara |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2019 | Unsupervised Speech Enhancement Based on Multichannel NMF-Informed Beamforming for Noise-Robust Automatic Speech RecognitionabstractThis paper describes multichannel speech enhancement for improving automatic speech recognition (ASR) in noisy environments. Recently, the minimum variance distortionless response (MVDR) beamforming has widely been used because it works well if the steering vector of speech and the spatial covariance matrix (SCM) of noise are given. To estimating such spatial information, conventional studies take a supervised approach that classifies each time-frequency (TF) bin into noise or speech by training a deep neural network (DNN). The performance of ASR, however, is degraded in an unknown noisy environment. To solve this problem, we take an unsupervised approach that decomposes each TF bin into the sum of speech and noise by using multichannel nonnegative matrix factorization (MNMF). This enables us to accurately estimate the SCMs of speech and noise not from observed noisy mixtures but from separated speech and noise components. In this paper, we propose online MVDR beamforming by effectively initializing and incrementally updating the parameters of MNMF. Another main contribution is to comprehensively investigate the performances of ASR obtained by various types of spatial filters, i.e., time-invariant and variant versions of MVDR beamformers and those of rank-1 and full-rank multichannel Wiener filters, in combination with MNMF. The experimental results showed that the proposed method outperformed the state-of-the-art DNN-based beamforming method in unknown environments that did not match training data. Kazuki Shimada, Yoshiaki Bando, Masato Mimura, Katsutoshi Itoyama, Kazuyoshi Yoshii, Tatsuya Kawahara |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2018 | Statistical Speech Enhancement Based on Probabilistic Integration of Variational Autoencoder and Non-Negative Matrix FactorizationabstractThis paper presents a statistical method of single-channel speech enhancement that uses a variational autoencoder (VAE) as a prior distribution on clean speech. A standard approach to speech enhancement is to train a deep neural network (DNN) to take noisy speech as input and output clean speech. Although this supervised approach requires a very large amount of pair data for training, it is not robust against unknown environments. Another approach is to use non-negative matrix factorization (NMF) based on basis spectra trained on clean speech in advance and those adapted to noise on the fly. This semi-supervised approach, however, causes considerable signal distortion in enhanced speech due to the unrealistic assumption that speech spectrograms are linear combinations of the basis spectra. Replacing the poor linear generative model of clean speech in NMF with a VAE-a powerful nonlinear deep generative model-trained on clean speech, we formulate a unified probabilistic generative model of noisy speech. Given noisy speech as observed data, we can sample clean speech from its posterior distribution. The proposed method outperformed the conventional DNN-based method in unseen noisy environments. Yoshiaki Bando, Masato Mimura, Katsutoshi Itoyama, Kazuyoshi Yoshii, Tatsuya Kawahara |
ICASSP | 5 |
| 2018 | Efficient Learning of Articulatory Models Based on Multi-Label Training and Label Correction for Pronunciation LearningabstractArticulatory feedback is effective for computer-assisted pronunciation training (CAPT) systems. This paper investigates efficient model learning methods for providing articulatory information to language learners. We first propose an articulatory attribute modeling method based on a multi-label learning scheme. Then, the models are further enhanced with a simple and effective training label correction method. These proposed methods are evaluated in three tasks: native attribute recognition, pronunciation error detection of non-native speech, and non-native speech recognition. Experimental results show that proposed methods significantly improve the conventional deep neural network (DNN) based articulatory models. Richeng Duan, Tatsuya Kawahara, Masatake Dantsuji, Hiroaki Nanjo |
ICASSP | 2 |
| 2018 | An End-to-End Approach to Joint Social Signal Detection and Automatic Speech RecognitionabstractSocial signals such as laughter and fillers are often observed in natural conversation, and they play various roles in human-to-human communication. Detecting these events is useful for transcription systems to generate rich transcription and for dialogue systems to behave as we do such as synchronized laughing or attentive listening. We have studied an end-to-end approach to directly detect social signals from speech by using connectionist temporal classification (CTC), which is one of the end-to-end sequence labelling models. In this work, we propose a unified framework that integrates social signal detection (SSD) and automatic speech recognition (ASR). We investigate several reference labelling methods regarding social signals. Experimental evaluations demonstrate that our end-to-end framework significantly outperforms the conventional DNN-HMM system with regard to SSD performance as well as the character error rate (CER). Hirofumi Inaguma, Masato Mimura, Koji Inoue, Kazuyoshi Yoshii, Tatsuya Kawahara |
ICASSP | 5 |
| 2018 | Audio-Visual Conversation Analysis by Smart Posterboard and Humanoid RobotabstractThis paper addresses audio-visual signal processing for conversation analysis, which involves multi-modal behavior detection and mental-state recognition. We have investigated prediction of turn-taking by the audience in a poster session from their multi-modal behaviors, and found out that the eye-gaze provides an important cue compared with head nodding and verbal backchannels. This finding has been applied to audio-visual speaker diarization by combining eye-gaze information. We are now investigating engagement recognition in human-robot interaction based on the same scheme. Robust and realtime detection of laughing, backchannels and nodding is realized based on LSTM-CTC. We introduce a latent “character” model to cope with the subjectivity and variations of engagement annotations. Experimental evaluations demonstrate that (1) the latent character model is effective, (2) automatic behavior detection is robust and does not degrade the engagement recognition accuracy, and (3) the eye-gaze is the most important feature among others. Tatsuya Kawahara, Koji Inoue, Divesh Lala, Katsuya Takanashi |
ICASSP | 1 |
| 2018 | Unsupervised Beamforming Based on Multichannel Nonnegative Matrix Factorization for Noisy Speech RecognitionabstractThis paper presents unsupervised multichannel speech enhancement for noisy speech recognition. Time-frequency (TF) mask estimation has actively been studied for estimating the steering vectors and spatial covariance matrices of speech and noise used for beamforming. The state-of-the-art approach to mask estimation is to use deep neural networks (DNN s) for classifying the TF bins of observed signals into speech and noise. Such a supervised approach, however, does not work well in an unknown environment. To accurately estimate the spatial covariance matrices in an unsupervised manner, we perform blind source separation (BSS) based on multichannel nonnegative matrix factorization (MNMF) for decomposing each TF bin into the components of speech and the other sources (noise). To clarify a suitable type of beamforming for MNMF, we tested both time- invariant and time-varying versions of the minimum variance distortionless response (MVDR) beamforming in addition to standard multichannel Wiener filtering (MWF). The experimental results showed that our MNMF-based beamforming approach outperformed the state-of-the-art DNN-based beamforming method in unknown environments that do not match the training data. Kazuki Shimada, Yoshiaki Bando, Masato Mimura, Katsutoshi Itoyama, Kazuyoshi Yoshii, Tatsuya Kawahara |
ICASSP | 6 |
| 2018 | Acoustic-to-Word Attention-Based Model Complemented with Character-Level CTC-Based ModelabstractThis paper addresses end-to-end speech recognition which directly maps acoustic features to a word sequence. The acoustic-to-word model is attractive since it does not require an external language model and an elaborate decoder, resulting in extremely simple and fast decoding. The apparent drawback of this modeling is sparseness of training data, particularly for less frequent words. In this paper, we propose a framework complemented with a character-level model. Joint training of the word-level model with the character-level model enhances the generality of deep learning of feature extraction and classification processes, preventing it from overfitting. Moreover, the character-level model is used to decode out-of-vocabulary (OOV) words that are not covered by the word-level model. Since there are choices of connectionist temporal classification (CTC) and attention-based models in the end-to-end recognition, we also explore optimal combination for the hybrid system. Evaluations on the Corpus of Spontaneous Japanese (CSJ) show that (1) the acoustic-to-word attention-based model outperforms CTC, (2) multitask learning (MTL) with character-level CTC model is effective, and (3) the hybrid system achieves comparable or even better accuracy than the standard DNN-HMM system with a decoding speed faster by a factor of 25. Sei Ueno, Hirofumi Inaguma, Masato Mimura, Tatsuya Kawahara |
ICASSP | 4 |
| 2018 | Evaluation of Real-time Deep Learning Turn-taking Models for Multiple Dialogue ScenariosabstractThe task of identifying when to take a conversational turn is an important function of spoken dialogue systems. The turn-taking system should also ideally be able to handle many types of dialogue, from structured conversation to spontaneous and unstructured discourse. Our goal is to determine how much a generalized model trained on many types of dialogue scenarios would improve on a model trained only for a specific scenario. To achieve this goal we created a large corpus of Wizard-of-Oz conversation data which consisted of several different types of dialogue sessions, and then compared a generalized model with scenario-specific models. For our evaluation we go further than simply reporting conventional metrics, which we show are not informative enough to evaluate turn-taking in a real-time system. Instead, we process results using a performance curve of latency and false cut-in rate, and further improve our model's real-time performance using a finite-state turn-taking machine. Our results show that the generalized model greatly outperformed the individual model for attentive listening scenarios but was worse in job interview scenarios. This implies that a model based on a large corpus is better suited to conversation which is more user-initiated and unstructured. We also propose that our method of evaluation leads to more informative performance metrics in a real-time system. Divesh Lala, Koji Inoue, Tatsuya Kawahara |
ICMI | 3 |
| 2018 | Prediction of Turn-taking Using Multitask Learning with Prediction of Backchannels and Fillers
Kohei Hara, Koji Inoue, Katsuya Takanashi, Tatsuya Kawahara |
INTERSPEECH | 4 |
| 2018 | Engagement Recognition in Spoken Dialogue via Neural Network by Aggregating Different Annotators' Models
Koji Inoue, Divesh Lala, Katsuya Takanashi, Tatsuya Kawahara |
INTERSPEECH | 4 |
| 2018 | Improving CTC-based Acoustic Model with Very Deep Residual Time-delay Neural Networks
Sheng Li 0010, Xugang Lu, Ryoichi Takashima, Tatsuya Kawahara, Hisashi Kawai |
INTERSPEECH | 5 |
| 2018 | Forward-Backward Attention Decoder
Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2018 | Encoder Transfer for Attention-based Acoustic-to-word Speech Recognition
Sei Ueno, Takafumi Moriya, Masato Mimura, Shinsuke Sakai, Yusuke Shinohara, Yoshikazu Yamaguchi, Yushi Aono, Tatsuya Kawahara |
INTERSPEECH | 8 |
| 2018 | Voice Input Tutoring System for Older Adults using Input Stumble DetectionabstractMany older adults are interested in smartphones but encounter difficulties in self-instruction and need support, especially text input. Voice input is a useful option for text input, but also presents some difficulties for older adults.In this paper, we propose a tutoring system for voice input that detects input stumbles using a statistical approach and provides instructions to overcome them. We construct the tutoring system based on the data from a user study with novice older adults. In an evaluation experiment, the number of input stumble and the sentence completion time of the participants using the tutoring system were significantly smaller than those without it. The results showed that the tutoring system resulted in the improvement of the efficiency of voice input for novice older adults. Toshiyuki Hagiya, Keiichiro Hoashi, Tatsuya Kawahara |
IUI | 3 |
| 2018 | A Unified Neural Architecture for Joint Dialog Act Segmentation and Recognition in Spoken Dialog SystemabstractIn spoken dialog systems (SDSs), dialog act (DA) segmentation and recognition provide essential information for response generation.A majority of previous works assumed ground-truth segmentation of DA units, which is not available from automatic speech recognition (ASR) in SDS.We propose a unified architecture based on neural networks, which consists of a sequence tagger for segmentation and a classifier for recognition.The DA recognition model is based on hierarchical neural networks to incorporate the context of preceding sentences.We investigate sharing some layers of the two components so that they can be trained jointly and learn generalized features from both tasks.An evaluation on the Switchboard Dialog Act (SwDA) corpus shows that the jointly-trained models outperform independently-trained models, single-step models, and other reported results in DA segmentation, recognition, and joint tasks. Tianyu Zhao 0001, Tatsuya Kawahara |
SIGDIAL Conference | 2 |
| 2018 | Improving Very Deep Time-Delay Neural Network With Vertical-Attention For Effectively Training CTC-Based ASR SystemsabstractThe very deep neural network has recently been proposed for speech recognition and achieves significant performance. It has excellent potential for integration with end-to-end (E2E) training. Connectionist temporal classification (CTC) has shown great potential in E2E acoustic modeling. In this study, we investigate deep architectures and techniques which are suitable for CTC-based acoustic modeling. We propose a very deep residual time-delay CTC neural network (VResTD-CTC). How to select a suitable deep architecture optimized with the CTC objective function is crucial for obtaining the state of the art performance. Excellent performances can be obtained by selecting deep architecture for non-E2E ASR systems modeling with tied-triphone states. However, these optimized structures do not guarantee to achieve better or comparable performances on E2E (e.g., CTC-based) systems modeling with dynamic acoustic units. For solving this problem and further leveraging the system performance, we introduce the vertical-attention mechanism to reweight the residual blocks at each time step. Speech recognition experiments show our proposed model significantly outperforms the DNN and LSTM-based (both bidirectional and unidirectional) CTC baseline models. Sheng Li 0010, Xugang Lu, Ryoichi Takashima, Tatsuya Kawahara, Hisashi Kawai |
SLT | 5 |
| 2018 | Improving OOV Detection and Resolution with External Language Models in Acoustic-to-Word ASRabstractAcoustic-to-word (A2W) end-to-end automatic speech recognition (ASR) systems have attracted attention because of an extremely simplified architecture and fast decoding. To alleviate data sparseness issues due to infrequent words, the combination with an acoustic-to-character (A2C) model is investigated. Moreover, the A2C model can be used to recover-of-vocabulary (OOV) words that are not covered by the A2W model, but this requires accurate detection of OOV words. A2W models learn contexts with both acoustic and transcripts; therefore they tend to falsely recognize OOV words as words in the vocabulary. In this paper, we tackle this problem by using external language models (LM), which are trained only with transcriptions and have better linguistic information to detect OOV words. The A2C model is used to resolve these OOV words. Experimental evaluations show that external LMs have the effects of not only reducing errors but also increasing the number of detected OOV words, and the proposed method significantly improves performances in English conversational and Japanese lecture corpora, especially for-of-domain scenario. We also investigate the impact of the vocabulary size of A2W models and the data size for training LMs. Moreover, our approach can reduce the vocabulary size several times with marginal performance degradation. Hirofumi Inaguma, Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
SLT | 4 |
| 2018 | Leveraging Sequence-to-Sequence Speech Synthesis for Enhancing Acoustic-to-Word Speech RecognitionabstractEncoder-decoder models for acoustic-to-word (A2W) automatic speech recognition (ASR) are attractive for their simplicity of architecture and run-time latency while achieving state-of-the-art performances. However, word-based models commonly suffer from the-of-vocabulary (OOV) word problem. They also cannot leverage text data to improve their language modeling capability. Recently, sequence-to-sequence neural speech synthesis models trainable from corpora have been developed and shown to achieve naturalness com- parable to recorded human speech. In this paper, we explore how we can leverage the current speech synthesis technology to tailor the ASR system for a target domain by preparing only a relevant text corpus. From a set of target domain texts, we generate speech features using a sequence-to-sequence speech synthesizer. These artificial speech features together with real speech features from conventional speech corpora are used to train an attention-based A2W model. Experimental results show that the proposed approach improves the word accuracy significantly compared to the baseline trained only with the real speech, although synthetic part of the training data comes only from a single female speaker voice. Masato Mimura, Sei Ueno, Hirofumi Inaguma, Shinsuke Sakai, Tatsuya Kawahara |
SLT | 5 |
| 2018 | Exploiting automatic speech recognition errors to enhance partial and synchronized caption for facilitating second language listening
Maryam Sadat Mirzaei, Kourosh Meshgi, Tatsuya Kawahara |
Comput. Speech Lang. | 3 |
| 2018 | Speech Enhancement Based on Bayesian Low-Rank and Sparse Decomposition of Multichannel Magnitude SpectrogramsabstractThis paper presents a blind multichannel speech enhancement method that can deal with the time-varying layout of microphones and sound sources. Since nonnegative tensor factorization (NTF) separates a multichannel magnitude (or power) spectrogram into source spectrograms without phase information, it is robust against the time-varying mixing system. This method, however, requires prior information such as the spectral bases (templates) of each source spectrogram in advance. To solve this problem, we develop a Bayesian model called robust NTF (Bayesian RNTF) that decomposes a multichannel magnitude spectrogram into target speech and noise spectrograms based on their sparseness and low rankness. Bayesian RNTF is applied to the challenging task of speech enhancement for a microphone array distributed on a hose-shaped rescue robot. When the robot searches for victims under collapsed buildings, the layout of the microphones changes over time and some of them often fail to capture target speech. Our method robustly works under such situations, thanks to its characteristic of time-varying mixing system. Experiments using a 3-m hose-shaped rescue robot with eight microphones show that the proposed method outperforms conventional blind methods in enhancement performance by the signal-to-noise ratio of 1.03 dB. Yoshiaki Bando, Katsutoshi Itoyama, Masashi Konyo, Satoshi Tadokoro, Kazuhiro Nakadai, Kazuyoshi Yoshii, Tatsuya Kawahara, Hiroshi G. Okuno |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2018 | Bayesian Multichannel Audio Source Separation Based on Integrated Source and Spatial ModelsabstractThis paper presents new statistical methods of multichannel audio source separation based on unified source and spatial models that, respectively, represent the generative process of latent source spectrograms and that of observed mixture spectrograms. One possibility of the source model is a factor model based on nonnegative matrix factorization that represents each time-frequency (TF) bin as the weighted sum of basis spectra. Another possibility is a mixture model inspired by latent Dirichlet allocation that exclusively classifies each TF bin into one of basis spectra. Similarly, the spatial model can either be a factor model that represents each TF bin as the weighted sum of source spectra or a mixture model that classifies each bin into one of those spectra. To unify these models in a principled manner and incorporate prior knowledge of a microphone array, we propose hierarchical Bayesian models of all the source-spatial combinations (factor-factor, mixture-factor, factor-mixture, and mixture-mixture models) and derive efficient Gibbs sampling algorithms for posterior inference. Experimental results showed that the proposed unified models outperformed the state-of-the-art method using only the spatial mixture model. Among the four unified models, the spatial factor model tended to work better than the spatial mixture model in exchange for larger computational cost, and the choice of source models had a little impact on the performance and computational cost. Kousuke Itakura, Yoshiaki Bando, Eita Nakamura, Katsutoshi Itoyama, Kazuyoshi Yoshii, Tatsuya Kawahara |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2017 | Incremental training and constructing the very deep convolutional residual network acoustic modelsabstractInspired by the successful applications in image recognition, the very deep convolutional residual network (ResNet) based model has been applied in automatic speech recognition (ASR). However, the computational load is heavy for training the ResNet with a large quantity of data. In this paper, we propose an incremental model training framework to accelerate the training process of the ResNet. The incremental model training framework is based on the unequal importance of each layer and connection in the ResNet. The modules with important layers and connections are regarded as a skeleton model, while those left are regarded as an auxiliary model. The total depth of the skeleton model is quite shallow compared to the very deep full network. In our incremental training, the skeleton model is first trained with the full training data set. Other layers and connections belonging to the auxiliary model are gradually attached to the skeleton model and tuned. Our experiments showed that the proposed incremental training obtained comparable performances and faster training speed compared with the model training as a whole without consideration of the different importance of each layer. Sheng Li 0010, Xugang Lu, Ryoichi Takashima, Tatsuya Kawahara, Hisashi Kawai |
ASRU | 5 |
| 2017 | Cross-domain speech recognition using nonparallel corpora with cycle-consistent adversarial networksabstractAutomatic speech recognition (ASR) systems often does not perform well when it is used in a different acoustic domain from the training time, such as utterances spoken in noisy environments or in different speaking styles. We propose a novel approach to cross-domain speech recognition based on acoustic feature mappings provided by a deep neural network, which is trained using nonparallel speech corpora from two different domains and using no phone labels. For training a target domain acoustic model, we generate “fake” target speech features from the labeleld source domain features using a mapping Gf. We can also generate “fake” source features for testing from the target features using the backward mapping Gbwhich has been learned simultaneously with G f. The mappings G f and Gbare trained as adversarial networks using a conventional adversarial loss and a cycle-consistency loss criterion that encourages the backward mapping to bring the translated feature back to the original as much as possible such that Gb(Gf (x)) ≈ x. In a highly challenging task of model adaptation only using domain speech features, our method achieved up to 16 % relative improvements in WER in the evaluation using the CHiME3 real test data. The backward mapping was also confirmed to be effective with a speaking style adaptation task. Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
ASRU | 3 |
| 2017 | Utterance Behavior of Users While Playing Basketball with a Virtual Teammate
Divesh Lala, Yuanchao Li, Tatsuya Kawahara |
ICAART (1) | 3 |
| 2017 | Effective articulatory modeling for pronunciation error detection of L2 learner without non-native training dataabstractFor effective articulatory feedback in computer-assisted pronunciation training (CAPT) systems, we address effective articulatory models of second language (L2) learners' speech without using such data, which is difficult to collect and annotate in a large scale. Context-dependent articulatory attributes (placement and manner of articulation) are modeled based on deep neural network (DNN). In order to efficiently train the non-native articulatory models, we exploit large speech corpora of native and target language to model inter-language phenomena. This multi-lingual learning is then combined with multi-task learning, which uses phone-classification as a sub-task. These methods are applied to Mandarin Chinese pronunciation learning by Japanese native speakers. Effects are confirmed in the native attribute classification and pronunciation error detection of non-native speech. Richeng Duan, Tatsuya Kawahara, Masatake Dantsuji, Jinsong Zhang 0001 |
ICASSP | 2 |
| 2017 | Bayesian multichannel nonnegative matrix factorization for audio source separation and localizationabstractThis paper presents a Bayesian extension of multichannel nonnegative matrix factorization (MNMF) that decomposes the complex spectrograms of mixture signals recorded by a microphone array into basis spectra, their temporal activations, and the spatial correlation matrices of sources (directions) in the time-frequency-channel domain. Although the original MNMF can be used in a blind setting, prior knowledge of a microphone array is useful for improving source separation. The impulse response (spatial correlation matrix) of each direction can be measured in an anechoic room, however, it differs from that in a real environment where the microphone array is used. To solve this, we propose a unified Bayesian model of source separation and localization by introducing a prior distribution determined by an anechoic spatial correlation matrix on a real spatial correlation matrix with respect to each direction. This enables us to adaptively estimate a real spatial correlation matrix and the direction of each source. Experimental results showed that our method outperformed the original MNMF and the state-of-the-art methods with prior knowledge in terms of signal-to-distortion ratio (SDR) even when the method was used in an unknown environment with acoustic characteristics different from those of the anechoic room. Kousuke Itakura, Yoshiaki Bando, Eita Nakamura, Katsutoshi Itoyama, Kazuyoshi Yoshii, Tatsuya Kawahara |
ICASSP | 6 |
| 2017 | Semi-supervised ensemble DNN acoustic model trainingabstractIt is very important to exploit abundant unlabeled speech for improving the acoustic model training in automatic speech recognition (ASR). Semi-supervised training methods incorporate unlabeled data in addition to labeled data to enhance the model training, but it encounters the error-prone label problem. The ensemble training scheme trains a set of models and combines them to make the model more general and robust, but it has not been applied to the unlabeled data. In this work, we propose an effective semi-supervised training of deep neural network (DNN) acoustic models by incorporating the diversity among the ensemble of models. The resultant model improved the performance in the lecture transcription task. Moreover, the proposed method has also shown a potential for DNN adaptation. Sheng Li 0010, Xugang Lu, Shinsuke Sakai, Masato Mimura, Tatsuya Kawahara |
ICASSP | 5 |
| 2017 | Joint Learning of Dialog Act Segmentation and Recognition in Spoken Dialog Using Neural NetworksabstractDialog act segmentation and recognition are basic natural language understanding tasks in spoken dialog systems. This paper investigates a unified architecture for these two tasks, which aims to improve the model’s performance on both of the tasks. Compared with past joint models, the proposed architecture can (1) incorporate contextual information in dialog act recognition, and (2) integrate models for tasks of different levels as a whole, i.e. dialog act segmentation on the word level and dialog act recognition on the segment level. Experimental results show that the joint training system outperforms the simple cascading system and the joint coding system on both dialog act segmentation and recognition tasks. Tianyu Zhao 0001, Tatsuya Kawahara |
IJCNLP(1) | 2 |
| 2017 | Social Signal Detection in Spontaneous Dialogue Using Bidirectional LSTM-CTC
Hirofumi Inaguma, Koji Inoue, Masato Mimura, Tatsuya Kawahara |
INTERSPEECH | 4 |
| 2017 | Combined Multi-Channel NMF-Based Robust Beamforming for Noisy Speech Recognition
Masato Mimura, Yoshiaki Bando, Kazuki Shimada, Shinsuke Sakai, Kazuyoshi Yoshii, Tatsuya Kawahara |
INTERSPEECH | 6 |
| 2017 | Analysis of the Relationship Between Prosodic Features of Fillers and its Forms or Occurrence Positions
Shizuka Nakamura, Ryosuke Nakanishi, Katsuya Takanashi, Tatsuya Kawahara |
INTERSPEECH | 4 |
| 2017 | Attentive listening system with backchanneling, response generation and flexible turn-takingabstractAttentive listening systems are designed to let people, especially senior people, keep talking to maintain communication ability and mental health.This paper addresses key components of an attentive listening system which encourages users to talk smoothly.First, we introduce continuous prediction of end-of-utterances and generation of backchannels, rather than generating backchannels after end-point detection of utterances.This improves subjective evaluations of backchannels.Second, we propose an effective statement response mechanism which detects focus words and responds in the form of a question or partial repeat.This can be applied to any statement.Moreover, a flexible turn-taking mechanism is designed which uses backchannels or fillers when the turnswitch is ambiguous.These techniques are integrated into a humanoid robot to conduct attentive listening.We test the feasibility of the system in a pilot experiment and show that it can produce coherent dialogues during conversation. Divesh Lala, Pierrick Milhorat, Koji Inoue, Masanari Ishida, Katsuya Takanashi, Tatsuya Kawahara |
SIGDIAL Conference | 6 |
| 2016 | Data selection from multiple ASR systems' hypotheses for unsupervised acoustic model trainingabstractThis paper addresses unsupervised training of DNN acoustic model, by exploiting a large amount of unlabeled data with CRF-based classifiers. In the proposed scheme, we obtain ASR hypotheses by complementary GMM and DNN based ASR systems. Then, a set of dedicated classifiers are designed and trained to select the better hypothesis and verify the selected data. It is demonstrated that the classifiers can effectively filter usable data from unlabeled data for acoustic model training. The proposed method achieved significant improvement in the ASR accuracy from the baseline system, and it outperformed the models trained from the data selected based on the confidence measure scores (CMS) and also from the simple ROVER-based system combination. Sheng Li 0010, Yuya Akita, Tatsuya Kawahara |
ICASSP | 3 |
| 2016 | Multimodal interaction with the autonomous Android ERICAabstractWe demonstrate an interactive conversation with an android named ERICA. In this demonstration the user can converse with ERICA on a number of topics. We demonstrate both the dialog management system and the eye gaze behavior of ERICA used for indicating attention and turn taking. Divesh Lala, Pierrick Milhorat, Koji Inoue, Tianyu Zhao 0001, Tatsuya Kawahara |
ICMI | 5 |
| 2016 | Prediction and Generation of Backchannel Form for Attentive Listening Systems
Tatsuya Kawahara, Takashi Yamaguchi, Koji Inoue, Katsuya Takanashi, Nigel G. Ward |
INTERSPEECH | 1 |
| 2016 | Joint Optimization of Denoising Autoencoder and DNN Acoustic Model Based on Multi-Target Learning for Noisy Speech Recognition
Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2016 | Managing Dialog and Joint Actions for Virtual Basketball Teammates
Divesh Lala, Tatsuya Kawahara |
IVA | 2 |
| 2016 | ERICA: The ERATO Intelligent Conversational AndroidabstractThe development of an android with convincingly lifelike appearance and behavior has been a long-standing goal in robotics, and recent years have seen great progress in many of the technologies needed to create such androids. However, it is necessary to actually integrate these technologies into a robot system in order to assess the progress that has been made towards this goal and to identify important areas for future work. To this end, we are developing ERICA, an autonomous android system capable of conversational interaction, featuring advanced sensing and speech synthesis technologies, and arguably the most humanlike android built to date. Although the project is ongoing, initial development of the basic android platform has been completed. In this paper we present an overview of the requirements and design of the platform, describe the development process of an interactive application, report on ERICA's first autonomous public demonstration, and discuss the main technical challenges that remain to be addressed in order to create humanlike, autonomous androids. Dylan F. Glas, Takashi Minato, Carlos Toshinori Ishi, Tatsuya Kawahara, Hiroshi Ishiguro |
RO-MAN | 4 |
| 2016 | Talking with ERICA, an autonomous androidabstractWe demonstrate dialogues with an autonomous android ERICA, who has an appearance like a human being.Currently, ERICA plays two social roles: a laboratory guide and a counselor.It is designed to follow the protocols of human dialogue to make the user comfortable: (1) having a chat before the main talk, (2) proactively asking questions, and (3) conveying proper feedbacks.The combination of the human-like appearance and the appropriate behaviors according to her social roles allows for symbiotic human-robot interaction. Koji Inoue, Pierrick Milhorat, Divesh Lala, Tianyu Zhao 0001, Tatsuya Kawahara |
SIGDIAL Conference | 5 |
| 2016 | Semi-Supervised Acoustic Model Training by Discriminative Data Selection From Multiple ASR Systems' HypothesesabstractWhile the performance of ASR systems depends on the size of the training data, it is very costly to prepare accurate and faithful transcripts. In this paper, we investigate a semisupervised training scheme, which takes the advantage of huge quantities of unlabeled video lecture archive, particularly for the deep neural network (DNN) acoustic model. In the proposed method, we obtain ASR hypotheses by complementary GMM- and DNN-based ASR systems. Then, a set of CRF-based classifiers is trained to select the correct hypotheses and verify the selected data. The proposed hypothesis combination shows higher quality compared with the conventional system combination method (ROVER). Moreover, compared with the conventional data selection based on confidence measure score, our method is demonstrated more effective for filtering usable data. Significant improvement in the ASR accuracy is achieved over the baseline system and in comparison with the models trained with the conventional system combination and data selection methods. Sheng Li 0010, Yuya Akita, Tatsuya Kawahara |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Language model adaptation for academic lectures using character recognition result of presentation slidesabstractFor automatic speech recognition (ASR) of lectures, texts of presentation slides are expected to be useful for adapting a language model, while slide texts are not always available in a machine-readable form. In this paper, we propose a language model adaptation framework that uses character recognition results of slide images in a lecture video. Since character recognition results contain many errors, we introduce a filtering method based on morphological and topic information. Then we perform linear interpolation of the baseline language model with the filtered results and also relevant texts which are selected automatically from a text database using the filtered results. We further conduct a cache-based adaptation method on the resulting language model, in which keywords in the filtered results are cached and used to boost the word probability. In an experimental evaluation over real lectures, we obtained a significant improvement of ASR performance by this adaptation framework. Yuya Akita, Yizheng Tong, Tatsuya Kawahara |
ICASSP | 3 |
| 2015 | Deep autoencoders augmented with phone-class feature for reverberant speech recognitionabstractThis paper addresses reverberant speech recognition based on front-end processing using DAE (Deep AutoEncoder) coupled with DNN (Deep Neural Network) acoustic model. DAE can effectively and flexibly learn mapping from corrupted speech to the original clean speech based on the deep learning scheme. While this mapping is conventionally conducted only with the acoustic information, we presume the mapping is also dependent on the phone information. Therefore, we propose a new scheme (pDAE), which augments a phone-class feature to the standard acoustic features as input. Two types of the phone-class feature are investigated. One is the hard recognition result of monophones, and the other is a soft representation derived from the posterior outputs of monophone DNN. In the evaluation on the Reverb Challenge 2014 task, the augmented feature in either type results in a significant improvement (7-8% relative) from the standard DAE. It is also shown that using the soft representation in the training phase is critical. Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
ICASSP | 3 |
| 2015 | Enhanced speaker diarization with detection of backchannels using eye-gaze information in poster conversationsabstractWe propose multi-modal speaker diarization using acoustic and eye-gaze information in poster conversations. Eye-gaze information plays an important role in turn-taking, thus it is useful for predicting speech activity. In this paper, a variety of eyegaze features are elaborated and combined with the acoustic information by the multi-modal integration model. Moreover, we introduce another model to detect backchannels, which involve different eye-gaze behaviors. This enhances the diarization result by filtering meaningful utterances such as questions and comments. Experimental evaluations in real poster sessions demonstrate that eye-gaze information contributes to improvement of diarization accuracy under noisy environments, and its weight is automatically determined according to the Signal-toNoise Ratio (SNR). Index Terms: speaker diarization, backchannel, multi-modal, eye-gaze, poster conversation Koji Inoue, Yukoh Wakabayashi, Hiromasa Yoshimoto, Katsuya Takanashi, Tatsuya Kawahara |
INTERSPEECH | 5 |
| 2015 | Discriminative data selection for lightly supervised training of acoustic model using closed caption textsabstractWe present a novel data selection method for lightly supervised training of acoustic model, which exploits a large amount of data with closed caption texts but not faithful transcripts. In the proposed scheme, a sequence of the closed caption text and that of the ASR hypothesis by the baseline system are aligned. Then, a set of dedicated classifiers is designed and trained to select the correct one among them or reject both. It is demonstrated that the classifiers can effectively filter the usable data for acoustic model training without tuning any threshold parameters. A significant improvement in the ASR accuracy is achieved from the baseline system and also in comparison with the conventional method of lightly supervised training based on simple matching and confidence measure scores. Sheng Li 0010, Yuya Akita, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2015 | Ensemble speaker modeling using speaker adaptive training deep neural network for speaker adaptationabstractIn this paper, we introduce an ensemble speaker modeling using a speaker adaptive training (SAT) deep neural network (SAT-DNN). We first train a speaker-independent DNN (SIDNN) acoustic model as a universal speaker model (USM). Based on the USM, a SAT-DNN is used to obtain a set of speaker-dependent models by assuming that all other layers except one speaker-dependent (SD) layer are shared among speakers. The speaker ensemble matrix is created by concatenating all of the SD neural weight matrices. With matrix factorization technique, an ensemble speaker subspace is extracted. When testing, an initial model for each target speaker is selected in this ensemble speaker subspace. Then, adaptation is carried out to obtain the final acoustic model for testing. In order to reduce the number of adaptation parameters, low-rank speaker subspace is further explored. We test our algorithm on lecture transcription task. Experimental results showed that our proposed method is effective for unsupervised speaker adaptation. Index Terms: speaker adaptation, deep neural networks, ensemble modeling, lecture transcription Sheng Li 0010, Xugang Lu, Yuya Akita, Tatsuya Kawahara |
INTERSPEECH | 4 |
| 2015 | Speech dereverberation using long short-term memoryabstractRecently, neural networks have been used for not only phone recognition but also denoising and dereverberation. However, the conventional denoising deep autoencoder (DAE) based on the feed-forward structure is not capable of handling very long speech frames of reverberation. LSTM can be effectively trained to reduce the average error between the enhanced signal and the original clean signal by considering the effect of the long past time frames. In this paper, we demonstrate that considering as long as the maximum reverberation time of the database is effective. Since the effect of reverberation varies depending on the phone-class of the whole speech context, we augment the input of the autoencoder with the phone-class information of the past frames as well as the current frame and call this version of the LSTM autoencoder pLSTM. In the speech recognition experiment using the data set of Reverb Challenge 2014, the LSTM front-end reduced the WER of the multicondition DNN-HMM by 14.5%, and the use of the phone class feature yielded in pLSTM further improvement of 7.5%. The performance with the pLSTM is comparable to that of pDAE, while the number of parameters is only 1/25-1/8. Index Terms: Speech Dereverberation, Long Short-Term Memory (LSTM), Deep Autoencoder (DAE) Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2015 | Conversational system for information navigation based on POMDP with user focus tracking
Koichiro Yoshino, Tatsuya Kawahara |
Comput. Speech Lang. | 2 |
| 2014 | Partial and Synchronized Caption Generation to Develop Second Language Listening Skill
Maryam Sadat Mirzaei, Yuya Akita, Tatsuya Kawahara |
ICCE | 3 |
| 2014 | Speaker diarization using eye-gaze information in multi-party conversationsabstractWe present a novel speaker diarization method by using eye-gaze information in multi-party conversations. In real environ-ments, speaker diarization or speech activity detection of each participant of the conversation is challenging because of distant talking and ambient noise. In contrast, eye-gaze information is robust against acoustic degradation, and it is presumed that eye-gaze behavior plays an important role in turn-taking and thus in predicting utterances. The proposed method stochastically integrates eye-gaze information with acoustic information for speaker diarization. Specifically, three models are investigated for multi-modal integration in this paper. Experimental eval-uations in real poster sessions demonstrate that the proposed method improves accuracy of speaker diarization from the base-line acoustic method. Index Terms: speaker diarization, multi-modal interaction, eye-gaze Koji Inoue, Yukoh Wakabayashi, Hiromasa Yoshimoto, Tatsuya Kawahara |
INTERSPEECH | 4 |
| 2014 | Information Navigation System Based on POMDP that Tracks User FocusabstractWe present a spoken dialogue system for navigating information (such as news ar-ticles), and which can engage in small talk. At the core is a partially observ-able Markov decision process (POMDP), which tracks user’s state and focus of at-tention. The input to the POMDP is pro-vided by a spoken language understanding (SLU) component implemented with lo-gistic regression (LR) and conditional ran-dom fields (CRFs). The POMDP selects one of six action classes; each action class is implemented with its own module. 1 Koichiro Yoshino, Tatsuya Kawahara |
SIGDIAL Conference | 2 |
| 2014 | Lexicon optimization based on discriminative learning for automatic speech recognition of agglutinative language
Mijit Ablimit, Tatsuya Kawahara, Askar Hamdulla |
Speech Commun. | 2 |
| 2014 | Multiparty Interaction Understanding Using Smart Multimodal Digital SignageabstractThis paper presents a novel multimodal system designed for multi-party human-human interaction analysis. The design of human-machine interfaces for multiple users is challenging because simultaneous processing of actions and reactions have to be consistent. The proposed system consists of a large display equipped with multiple sensing devices: microphone array, HD video cameras, and depth sensors. Multiple users positioned in front of the panel freely interact using voice or gesture while looking at the displayed content, without wearing any particular devices (such as motion capture sensors or head mounted devices). Acoustic and visual information is captured and processed jointly using established and state-of-the-art techniques to obtain individual speech and gaze direction. Furthermore, a new framework is proposed to model A/V multimodal interaction between verbal and nonverbal communication events. Dynamics of audio signals obtained from speaker diarization and head poses extracted from video images are modeled using hybrid dynamical systems (HDS). We show that HDS temporal structure characteristics can be used for multimodal interaction level estimation, which is useful feedback that can help to improve multi-party communication experience. Experimental results using synthetic and real-world datasets of group communication such as poster presentations show the feasibility of the proposed multimodal system. Tony Tung, Randy Gomez, Tatsuya Kawahara, Takashi Matsuyama |
IEEE Trans. Hum. Mach. Syst. | 3 |
| 2013 | Incorporating semantic information to selection of web texts for language model of spoken dialogue systemabstractA novel text selection approach for training a language model (LM) with Web texts is proposed for automatic speech recognition (ASR) of spoken dialogue systems. Compared to the conventional approach based on perplexity criterion, the proposed approach introduces a semantic-level relevance measure with the back-end knowledge base used in the dialogue system. We focus on the predicate-argument (P-A) structure characteristic to the domain in order to filter semantically relevant sentences in the domain. Several choices of statistical models and combination methods with the perplexity measure are investigated in this paper. Experimental evaluations in two different domains demonstrate the effectiveness and generality of the proposed approach. The combination method realizes significant improvement not only in ASR accuracy but also in semantic and dialogue-level accuracy. Koichiro Yoshino, Shinsuke Mori, Tatsuya Kawahara |
ICASSP | 3 |
| 2013 | Hands-free human-robot communication robust to speaker's radial positionabstractIn this paper we present a method in room transfer function (RTF) estimation, employed specifically for dereverberation in hands-free human-robot communication.We introduce a radial distance compensation scheme which significantly improved the RTF estimate robust to the speech power variation due to changes in speaker's radial position. The proposed method is implemented in two levels; first, waveform-level compensation is executed to reflect the change in power caused by the change of radial position to the RTF. We generated possible RTF estimates within a close neighbourhood based on curve fitting. Then, we select among these estimates the optimal RTF based on acoustic model likelihood criterion, the same criterion employed in automatic speech recognition (ASR) systems. The latter is referred to as acoustic model-level compensation, which links the generated RTF to the ASR. We note that in ASR application, both waveform and acoustic models play an important role in achieving optimal performance. Thus, the synergistic effect of the two processes guarantee ASR performance improvement when used in conjunction with our ASR-based dereverberation scheme. Experimental evaluation show robustness in recognition performance when used in hands-free human-robot communication environment. Randy Gomez, Keisuke Nakamura, Kazuhiro Nakadai, Ui-Hyun Kim, Hiroshi G. Okuno, Tatsuya Kawahara |
ICRA | 6 |
| 2013 | Predicate Argument Structure Analysis using Partially Annotated Corpora
Koichiro Yoshino, Shinsuke Mori, Tatsuya Kawahara |
IJCNLP | 3 |
| 2013 | Estimation of interest and comprehension level of audience through multi-modal behaviors in poster conversationsabstractWe address the estimation of the interest and comprehension level of an audience in poster sessions. Compared to lecture presentations, the audience’s behaviors such as gazing and backchannels are more observable in poster presentations. These multi-modal behaviors are presumably related with their interest and comprehension level. We also assume that the interest and comprehension level can be judged by particular speech acts of the audience such as questions and reactive tokens. First, we make a preliminary analysis on their correlation. Next, we investigate the relationship between the audience’s behaviors and the question type. Then, we conduct prediction of questions and their type based on the multi-modal behaviors during the relevant topic segment. Experimental results show that verbal backchannels and eye-gaze patterns are good predictors to this task, and also the combination of the multi-modal features is effective. Index Terms: multi-modal interaction, behavioral analysis, eye-gaze, backchannel Tatsuya Kawahara, Soichiro Hayashi, Katsuya Takanashi |
INTERSPEECH | 1 |
| 2013 | Substring-based machine translation
Graham Neubig, Taro Watanabe, Shinsuke Mori, Tatsuya Kawahara |
Mach. Transl. | 4 |
| 2012 | Machine Translation without Words through Substring Alignment
Graham Neubig, Taro Watanabe, Shinsuke Mori, Tatsuya Kawahara |
ACL (1) | 4 |
| 2012 | Language Modeling for Spoken Dialogue System based on Filtering using Predicate-Argument Structures
Koichiro Yoshino, Shinsuke Mori, Tatsuya Kawahara |
COLING | 3 |
| 2012 | Multi-party human-robot interaction with distant-talking speech recognitionabstractSpeech is one of the most natural medium for human communication, which makes it vital to human-robot interaction. In real environments where robots are deployed, distant-talking speech recognition is difficult to realize due to the effects of reverberation. This leads to the degradation of speech recognition and understanding, and hinders a seamless human-robot interaction. To minimize this problem, traditional speech enhancement techniques optimized for human perception are adopted to achieve robustness in human-robot interaction. However, human and machine perceive speech differently: an improvement in speech recognition performance may not automatically translate to an improvement in human-robot interaction experience (as perceived by the users). In this paper, we propose a method in optimizing speech enhancement techniques specifically to improve automatic speech recognition (ASR) with emphasis on the human-robot interaction experience. Experimental results using real reverberant data in a multi-party conversation, show that the proposed method improved human-robot interaction experience in severe reverberant conditions compared to the traditional techniques. Randy Gomez, Tatsuya Kawahara, Keisuke Nakamura, Kazuhiro Nakadai |
HRI | 2 |
| 2012 | Transcription System Using Automatic Speech Recognition for the Japanese Parliament (Diet)abstractThis article describes a new automatic transcription system in the Japanese Parliament which deploys our automatic speech recognition (ASR) technology. To achieve high recognition performance in spontaneous meeting speech, we have investigated an efficient training scheme with minimal supervision which can exploit a huge amount of real data. Specifically, we have proposed a lightly-supervised training scheme based on statistical language model transformation, which fills the gap between faithful transcripts of spoken utterances and final texts for documentation. Once this mapping is trained, we no longer need faithful transcripts for training both acoustic and language models. Instead, we can fully exploit the speech and text data available in Parliament as they are. This scheme also realizes a sustainable ASR system which evolves, i.e. update/re-train the models, only with speech and text generated during the system operation. The ASR system has been deployed in the Japanese Parliament since 2010, and consistently achieved character accuracy of nearly 90%, which is useful for streamlining the transcription process. Tatsuya Kawahara |
IAAI | 1 |
| 2012 | Discriminative approach to lexical entry selection for automatic speech recognition of agglutinative languageabstractIn agglutinative languages, selection of lexical unit is not obvious. Morpheme unit is usually adopted to ensure the sufficient coverage, but many morphemes are short, resulting in weak constraints and possible confusions. In this paper, we propose a discriminative approach to select lexical entries which will directly contribute to ASR error reduction. We define an evaluation function for each word by a set of features and their weights, and the measure for optimization by the difference of WERs by the morpheme-based model and by the word-based model. Then, the weights of the features are learned by a perceptron algorithm. Finally, word (or sub-word) entries with higher evaluation scores are selected to be added to the lexicon. This method is successfully applied to an Uyghur large-vocabulary continuous speech recognition system, resulting in a significant reduction of WER and the lexicon size. Further improvement is achieved by combining with a statistical method based on mutual information criterion. Mijit Ablimit, Tatsuya Kawahara, Askar Hamdulla |
ICASSP | 2 |
| 2012 | Automatic Transcription of Lecture Speech using Language Model Based on Speaking-Style Transformation of Proceeding TextsabstractFor language modeling of spontaneous speech recognition, we propose a style transformation approach, which transforms written texts to a spoken-style language model. Since these two styles are largely different and thus direct transformation is difficult, we cascade two transformation methods; rule-based transformation to rewrite written-style texts to intermediate “verbatim” texts, and statistical transformation of language model from the verbatim style to the spoken style which is suitable for ASR. In an experimental evaluation on real lecture speech, the proposed transformation approach achieved higher performance than the conventional linear interpolation method. Index Terms: automatic speech recognition, lecture speech, language model, style transformation Yuya Akita, Makoto Watanabe, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2012 | Dereverberation based on Wavelet Packet Filtering for Robust Automatic Speech Recognition
Tatsuya Kawahara, Randy Gomez |
INTERSPEECH | 1 |
| 2012 | Prediction of Turn-Taking by Combining Prosodic and Eye-Gaze Information in Poster ConversationsabstractWe investigate turn-taking behaviors in conversations in poster sessions. While the poster presenter holds most of the turns during sessions, the audience’s utterances are more important and should not be missed. In this paper, therefore, prediction of turn-taking by the audience is addressed. It is classified into two sub-tasks: prediction of speaker change and prediction of the next speaker. We made analysis on eye-gaze information and its relationship with turn-taking, introducing joint eye-gaze events by the presenter and audience. We also parameterize backchannel patterns of the audience. As a result of machine learning with these features, it is found that combination of prosodic features of the presenter and the joint eye-gaze features is effective for predicting speaker change, while eyegaze duration and backchannels preceding the speaker change are useful for predicting the next speaker among the audience. Index Terms: multi-party interaction, turn-taking, prosody, eye-gaze Tatsuya Kawahara, Takuma Iwatate, Katsuya Takanashi |
INTERSPEECH | 1 |
| 2012 | Comparative Analysis of Intensity between Native Speakers and Japanese Speakers of EnglishabstractIntensity has been reported as a reliable acoustical correlate of stress accent for English language, but not of pitch accent for Japanese language. This difference between English and Japanese languages is presumed to shape the characteristics of intensity in English spoken by Japanese (Japanese English, henceforth). Based on this presumption, the intensity of words in sentence utterances for Japanese English is compared to that for native speakers’ English (native English, henceforth). Statistical analysis shows that nouns for Japanese English are produced with less intensity, whereas most function words are with more intensity than those for native English. A correlation is recognized between the above results and the proficiency in Japanese English. Index Terms: power patterns, amplitude, prominence, Tomoko Nariai, Kazuyo Tanaka, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2012 | Designing an Evaluation Framework for Spoken Term Detection and Spoken Document Retrieval at the NTCIR-9 SpokenDoc Task
Tomoyosi Akiba, Hiromitsu Nishizaki, Kiyoaki Aikawa, Tatsuya Kawahara, Tomoko Matsui |
LREC | 4 |
| 2012 | Multi-modal Sensing and Analysis of Poster Conversations: Toward Smart Posterboard
Tatsuya Kawahara |
SIGDIAL Conference | 1 |
| 2012 | A monotonic statistical machine translation approach to speaking style transformation
Graham Neubig, Yuya Akita, Shinsuke Mori, Tatsuya Kawahara |
Comput. Speech Lang. | 4 |
| 2011 | An Unsupervised Model for Joint Phrase Alignment and Extraction
Graham Neubig, Taro Watanabe, Eiichiro Sumita, Shinsuke Mori, Tatsuya Kawahara |
ACL | 5 |
| 2011 | Automatic Comma Insertion of Lecture Transcripts Based on Multiple AnnotationsabstractTo enhance readability and usability of speech recognition results, automatic punctuation is an essential process. In this paper, we address automatic comma prediction based on conditional random fields (CRF) using lexical, syntactic and pause information. Since there is large disagreement in comma insertion between humans, we model individual tendencies of punctuation using annotations given by multiple annotators, and combine these models by voting and interpolation frameworks. Experimental evaluations on real lecture speech demonstrated that the combination of individual punctuation models achieves higher prediction accuracy for commas agreed by all annotators and those given by individual annotators. Yuya Akita, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2011 | Denoising Using Optimized Wavelet Filtering for Automatic Speech RecognitionabstractWe present an improved denoising method based on filtering of the noisy wavelet coefficients using a Wiener gain for automatic speech recognition (ASR). We optimize the wavelet parameters for speech and different noise profiles to achieve a better estimate of the Wiener gain for effective filtering. Moreover, we introduce a scaling parameter in the Wiener gain to minimize mismatch caused by distortion during the denoising process. Experimental results in large vocabulary continuous speech recognition (LVCSR) show that the proposed method is effective and robust to different noise conditions. IndexTerms: Speech recognition, Robustness, Denoising and Wavelet Randy Gomez, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2011 | Spoken Dialogue System based on Information Extraction using Similarity of Predicate Argument Structures
Koichiro Yoshino, Shinsuke Mori, Tatsuya Kawahara |
SIGDIAL Conference | 3 |
| 2010 | Using online model comparison in the Variational Bayes framework for online unsupervised Voice Activity DetectionabstractThis paper presents the use of online Variational Bayes method for online Voice Activity Detection (VAD) in an unsupervised context. In conventional VAD, the final step often relies on state machines whose parameters are heuristically tuned. The goal of this study is to propose a solid statistical scheme for VAD using online model comparison which is provided from the Variational Bayes framework. In this scheme, two models are estimated online in parallel: one for the noise-only situation, and the other for the noise-plus-signal situation The VAD decision is done automatically depending on the selected model. An experimental evaluation on the CENSREC-1-C database shows a significant improvement by the proposed method compared to conventional statistical VAD methods. David Cournapeau, Shinji Watanabe 0001, Atsushi Nakamura, Tatsuya Kawahara |
ICASSP | 4 |
| 2010 | Optimizing spectral subtraction and wiener filtering for robust speech recognition in reverberant and noisy conditionsabstractSpeech enhancement is a common approach to address the effects of degradation due to noise and channel contamination. This approach is intended to suppress unwanted signal and recover the clean speech. In this paper, we focus on two simple and low-computational methods: Wiener filtering (WF) and spectral subtraction (SS). Conventionally, these are formulated with no relation with automatic speech recognition (ASR). We propose to optimize the conventional speech enhancement technique in relation with likelihood of the acoustic model. We also exploit these simple speech enhancement techniques that are originally designed for denoising, to address reverberation as well. In the experiment with real noisy and reverberant environments, we have achieved significant improvement in recognition performance using the proposed approach. Randy Gomez, Tatsuya Kawahara |
ICASSP | 2 |
| 2010 | Improved statistical models for SMT-based speaking style transformationabstractAutomatic speech recognition (ASR) results contain not only ASR errors, but also disfluencies and colloquial expressions that must be corrected to create readable transcripts. We take the approach of statistical machine translation (SMT) to “translate” from ASR results into transcript-style text. We introduce two novel modeling techniques in this framework: a context-dependent translation model, which allows for usage of context to accurately model translation probabilities, and log-linear interpolation of conditional and joint probabilities, which allows for frequently observed translation patterns to be given higher priority. The system is implemented using weighted finite state transducers (WFST). On an evaluation using ASR results and manual transcripts of meetings of the Japanese Diet (national congress), the proposed methods showed a significant increase in accuracy over traditional modeling techniques. Graham Neubig, Yuya Akita, Shinsuke Mori, Tatsuya Kawahara |
ICASSP | 4 |
| 2010 | Semi-automated update of automatic transcription system for the Japanese national congressabstractUpdate of acoustic and language models is vital to maintain performance of automatic speech recognition (ASR) systems. To alleviate efforts for updating models, we propose a “semi-automated ” framework for the ASR system of the Japanese National Congress. The framework consists of our speaking-style transformation (SST) and lightly-supervised training (LSV) approaches, which can automatically generate spoken-style training texts and labels from documents like meeting minutes. An experimental evaluation demonstrated that this update framework improved the ASR performance for the latest meeting data. We also address an estimation method of the ASR accuracy based on SST, which uses minutes as reference texts and does not require verbatim transcripts. Index Terms: Spontaneous speech recognition, congressional speech, lightly-supervised training, speaking-style transformation 1. Yuya Akita, Masato Mimura, Graham Neubig, Tatsuya Kawahara |
INTERSPEECH | 4 |
| 2010 | An improved wavelet-based dereverberation for robust automatic speech recognitionabstractThis paper presents an improved wavelet-based dereverberation method for automatic speech recognition (ASR). Dereverberation is based on filtering reverberant wavelet coefficients with the Wiener gains to suppress the effect of the late reflections. Optimization of the wavelet parameters using acoustic model enables the system to estimate the clean speech and late reflections effectively. This results to a better estimate of the Wiener gains for dereverberation in the ASR application. Additional tuning of the parameters of the Wiener gain in relation with the acoustic model further improves the dereverberation process for ASR. In the experiment with real reverberant data, we have achieved a significant improvement in ASR accuracy. Randy Gomez, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2010 | Constructing Japanese test collections for spoken term detectionabstractSpoken Document Retrieval (SDR) and Spoken Term Detection (STD) have been two of the most intensively investigated topics in spoken document processing research according to the establishment of the SDR and STD test collections by the Text REtrieval Conference (TREC) and NIST. Because Japanese spoken document processing researchers also requires such test collections for SDR and STD, we have established a working group to develop these collections in Special Interest Group-Spoken Language Processing (SIG-SLP) of the Information Processing Society of Japan. The working group has constructed and made available a test collection for SDR, and is now constructing new test collections for STD that will be open to researchers. The present paper introduces the policies, outline, and schedule of the new test collections. Then, the new test collections are compared with the NIST STD test collections. Index Terms: spoken term detection, test collection 1. Yoshiaki Itoh 0001, Hiromitsu Nishizaki, Xinhui Hu, Hiroaki Nanjo, Tomoyosi Akiba, Tatsuya Kawahara, Seiichi Nakagawa, Tomoko Matsui, Yoichi Yamashita, Kiyoaki Aikawa |
INTERSPEECH | 6 |
| 2010 | Classroom note-taking system for hearing impaired students using automatic speech recognition adapted to lecturesabstractWe are developing a real-time lecture transcription system for hearing impaired students in university classrooms. The automatic speech recognition (ASR) system is adapted to individual lecture courses and lecturers, to enhance the recognition accuracy. The ASR results are selectively corrected by a human editor, through a dedicated interface, before presenting to the students. An efficient adaptation scheme of the ASR modules has been investigated in this work. The system was tested for a hearing-impaired student in a lecture course on civil engineering. Compared with the current manual note-taking scheme offered by two volunteers, the proposed system generated almost double amount of texts with one human editor. Index Terms: hearing impaired, note-taking, automatic speech recognition, lectures, adaptation Tatsuya Kawahara, Norihiro Katsumaru, Yuya Akita, Shinsuke Mori |
INTERSPEECH | 1 |
| 2010 | Detection of hot spots in poster conversations based on reactive tokens of audienceabstractWe present a novel scheme for indexing “hot spots ” in conversations, such as poster sessions, based on the reaction of the audience. Specifically, we focus on laughters and non-lexical reactive tokens, which are presumably related with funny spots and interesting spots, respectively. A robust detection method of these acoustic events is realized by combining BIC-based segmentation and GMM-based classification, with additional verifiers for reactive tokens. Subjective evaluations suggest that hot spots associated with reactive tokens are consistently useful while those with laughters are not so reliable. Furthermore, we investigate prosodic patterns of those reactive tokens which are closely related with the interest level. Index Terms: audio indexing, acoustic event detection, hot spots, reactive token, prosody Tatsuya Kawahara, Kouhei Sumi, Zhi-Qiang Chang, Katsuya Takanashi |
INTERSPEECH | 1 |
| 2010 | Learning a language model from continuous speechabstractThis paper presents a new approach to language model construction, learning a language model not from text, but directly from continuous speech. A phoneme lattice is created using acoustic model scores, and Bayesian techniques are used to robustly learn a language model from this noisy input. A novel sampling technique is devised that allows for the integrated learning of word boundaries and an n-gram language model with no prior linguistic knowledge. The proposed techniques were used to learn a language model directly from continuous, potentially large-vocabulary speech. This language model was able to significantly reduce the ASR phoneme error rate over a separate set of test data, and the proposed lattice processing and lexical acquisition techniques were found to be important factors in this improvement. Index Terms: language acquisition, word segmentation, Pitman-Yor language model, Bayesian learning Graham Neubig, Masato Mimura, Shinsuke Mori, Tatsuya Kawahara |
INTERSPEECH | 4 |
| 2010 | Bayes risk-based dialogue management for document retrieval system with speech interface
Teruhisa Misu, Tatsuya Kawahara |
Speech Commun. | 2 |
| 2010 | Statistical Transformation of Language and Pronunciation Models for Spontaneous Speech RecognitionabstractWe propose a novel approach based on a statistical transformation framework for language and pronunciation modeling of spontaneous speech. Since it is not practical to train a spoken-style model using numerous spoken transcripts, the proposed approach generates a spoken-style model by transforming an orthographic model trained with document archives such as the minutes of meetings and the proceedings of lectures. The transformation is based on a statistical model estimated using a small amount of a parallel corpus, which consists of faithful transcripts aligned with their orthographic documents. Patterns of transformation, such as substitution, deletion, and insertion of words, are extracted with their word and part-of-speech (POS) contexts, and transformation probabilities are estimated based on occurrence statistics in a parallel aligned corpus. For pronunciation modeling, subword-based mapping between baseforms and surface forms is extracted with their occurrence counts, then a set of rewrite rules with their probabilities are derived as a transformation model. Spoken-style language and pronunciation (surface forms) models can be predicted by applying these transformation patterns to a document-style language model and baseforms in a lexicon, respectively. The transformed models significantly reduced perplexity and word error rates (WERs) in a task of transcribing congressional meetings, even though the domains and topics were different from the parallel corpus. This result demonstrates the generality and portability of the proposed framework. Yuya Akita, Tatsuya Kawahara |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Robust Speech Recognition Based on Dereverberation Parameter Optimization Using Acoustic Model LikelihoodabstractAutomatic speech recognition (ASR) in reverberant environments is a challenging task. Most dereverberation techniques address this problem through signal processing and enhances the reverberant waveform independent from the speech recognizer. In this paper, we propose a novel scheme to perform dereverberation in relation with the likelihood of the back-end ASR system. Our proposed approach effectively selects the dereverberation parameters, in the form of multiband scale factors, so that they improve the likelihood of the acoustic model. Then, the acoustic model is retrained using the optimal parameters. During the recognition phase, we implement additional optimization of the parameters. By using Gaussian mixture model (GMM), the process for selecting the scale factors become efficient. Moreover, we remove the dependency of the adopted dereverberation technique on the room impulse response (RIR) measurement, by using an artificial RIR generator and selecting based on the acoustic likelihood. Experimental results show significant improvement in recognition performance with the proposed method over the conventional approach. Randy Gomez, Tatsuya Kawahara |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Speech Activity Detection for Multi-Party Conversation Analyses Based on Likelihood Ratio Test on Spatial MagnitudeabstractThis paper proposes a microphone array-based speech activity detection (SAD) method for analyzing multi-party conversations recorded in the presence of noise. In particular, the proposed method considers conversations where the number of speakers and speaker locations cannot be restricted, such as when standing and talking, and at poster sessions. When we observe such conversations, there are directional noise sources and diffuse noise that affect the direction of arrival estimations of the target speech signals. To detect speech activity without a priori knowledge about the speakers and noise environments, a likelihood ratio test (LRT)-based SAD method is applied to spatial magnitude, which are estimated by using the time-frequency masking of the observed spectra. The proposed method can exploit the enhanced signals obtained from time-frequency masking, and works even in the presence of environmental noise. Experiments with recorded simulated poster sessions confirmed that the proposed method could outperform conventional methods based on the LRT for a single channel, magnitude coherence, or crosspower spectrum phase. Kentaro Ishizuka, Shoko Araki, Tatsuya Kawahara |
IEEE Trans. Speech Audio Process. | 3 |
| 2009 | New perspectives on spoken language understanding: Does machine need to fully understand speech?abstractSpoken language understanding (SLU) has been traditionally formulated to extract meanings or concepts of user utterances in the context of human-machine dialogue. With the broadened coverage of spoken language processing, the tasks and methodologies of SLU have been changed accordingly. The back-end of spoken dialogue systems now consist of not only relational databases (RDB) but also general documents, incorporating information retrieval (IR) and question-answering (QA) techniques. This paradigm shift and the author's approaches are reviewed. SLU is also being designed to cover human-human dialogues and multi-party conversations. Major approaches to ¿understand¿ human-human speech communication and a new approach based on the lister's reactions are reviewed. As a whole, these trends are apparently not oriented for full understanding of spoken language, but for robust extraction of clue information. Tatsuya Kawahara |
ASRU | 1 |
| 2009 | Language model transformation applied to lightly supervised training of acoustic model for congress meetingsabstractFor effective training of acoustic and language models for spontaneous speech such as meetings, it is significant to exploit the texts available in a large scale, which may not be faithful transcripts of the utterances. We have proposed a language model transformation scheme to cope with the differences between verbatim transcripts of spontaneous utterances and human-made transcripts such as those in proceedings. In this paper, we investigate its application to lightly supervised training of the acoustic model. By transforming the corresponding text in the proceedings, we can generate a very constrained model to predict the actual utterances. The experimental evaluation with the transcription system for the Japanese Congress meetings demonstrated that the proposed scheme can generate accurate labels for acoustic model training and thus realizes the comparable ASR (Automatic Speech Recognition) performance to the case using manual transcripts. Tatsuya Kawahara, Masato Mimura, Yuya Akita |
ICASSP | 1 |
| 2009 | Optimal learning of P-Layer additive F0 models with cross-validationabstractIn this paper, we present the derivation of the backfitting training algorithms for generic p-layer additive F0models for arbitrary positive integer p. We have presented the special cases of the algorithms with p = 2 and p = 3 that have been successfully applied to the modelings of Japanese and English F0contours, whereas the derivation of the algorithm was presented only for the two-layer case. The additive F0model have smoothing parameters that establish a trade-off between the fit to the training data and the smoothness of the fitted curves, which have been all set to unity in the previous works. In this paper, we also present an optimal approach to set the values of these parameters using cross validation. We performed the training using the Boston University Radio News Corpus and confirmed the effectiveness of the proposed method. Shinsuke Sakai, Tatsuya Kawahara, Tohru Shimizu, Satoshi Nakamura 0001 |
ICASSP | 2 |
| 2009 | Automatic transcription system for meetings of the Japanese national congressabstractThis paper presents an automatic speech recognition (ASR) system for assisting meeting record creation of the National Congress of Japan. The system is designed to cope with spontaneous characteristics of meeting speech, as well as a variety of topics and speakers. For acoustic model, minimum phone error (MPE) training is applied with several normalization techniques. For language model, we have proposed statistical style transformation to generate spoken-style N-grams and their statistics. We also introduce statistical modeling of pronunciation variation in spontaneous speech. The ASR system was evaluated on real congressional meetings, and achieved word accuracy of 84%. It is also suggested that the ASR-based transcripts with this accuracy level is usable for editing meeting records. Index Terms: Spontaneous speech recognition, congressional speech, minimum phone error training, statistical style transformation Yuya Akita, Masato Mimura, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2009 | Optimization of dereverberation parameters based on likelihood of speech recognizerabstractSpeech recognition under reverberant condition is a difficult task. Most dereverberation techniques used to address this problem enhance the reverberant waveform independent from that of the speech recognizer. In this paper, we improve the conventional Spectral Subtraction-based (SS) dereverberation technique. In our proposed approach, the dereverberation parameters are optimized to improve the likelihood of the acoustic model. The system is capable of adaptively fine-tuning these parameters jointly with acoustic model training. Additional optimization is also implemented during decoding of the test utterances. We have evaluated using real reverberant data and experimental results show that the proposed method significantly improves the recognition performance over the conventional approach. Randy Gomez, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2009 | A WFST-based log-linear framework for speaking-style transformationabstractWhen attempting to make transcripts from automatic speech recognition results, disfluency deletion, transformation of colloquial expressions, and insertion of dropped words must be performed to ensure that the final product is clean transcriptstyle text. This paper introduces a system for the automatic transformation of the spoken word to transcript-style language that enables not only deletion of disfluencies, but also substitutions of colloquial expressions and insertion of dropped words. A number of potentially useful features are combined in a loglinear probabilistic framework, and the utility of each is examined. The system is implemented using weighted finite state transducers (WFSTs) to allow for easy combination of features and integration with other WFST-based systems. On evaluation, the best system achieved a 5.37 % word error rate, a 5.49% absolute gain over a rule-based baseline and a 1.54 % absolute gain over a simple noisy-channel model. Index Terms: speaking style transformation, disfluency detection, weighted finite state transducers, log-linear model Graham Neubig, Shinsuke Mori, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2009 | Acoustic event detection for spotting "hot spots" in podcastsabstractThis paper presents a method to detect acoustic events that can be used to find “hot spots ” in podcast programs. We focus on meaningful non-verbal audible reactions which suggest hot spots such as laughter and reactive tokens. In order to detect this kind of short events and segment the counterpart utterances, we need accurate audio segmentation and classification, dealing with various recording environments and background music. Thus, we propose a method for automatically estimating and switching penalty weights for the BIC-based segmentation depending on background environments. Experimental results show significant improvement in detection accuracy by proposed method compared to when using a constant penalty weight. Index Terms: acoustic event detection, laughter detection, podcast, Bayesian Information Criterion Kouhei Sumi, Tatsuya Kawahara, Jun Ogata, Masataka Goto |
INTERSPEECH | 2 |
| 2009 | A Model of Temporally Changing User Behaviors in a Deployed Spoken Dialogue System
Kazunori Komatani, Tatsuya Kawahara, Hiroshi G. Okuno |
UMAP | 2 |
| 2009 | Computer Assisted Language Learning system based on dynamic question generation and error prediction for automatic speech recognition
Hongcui Wang, Christopher J. Waple, Tatsuya Kawahara |
Speech Commun. | 3 |
| 2008 | Using variational bayes free energy for unsupervised voice activity detectionabstractThis paper addresses the problem of Voice Active Detection (VAD) in noisy environments. We introduce Variational Bayes approach to EM for classification to replace the heuristic state machines. The Variational Bayes approach provides an explicit approximation of the evidence called Free Energy. Free Energy is used to assess the reliability of the classification model, and can be periodically updated with a small number of samples. We apply this scheme to the detection of invalid classification caused in noise-only portions for more reliable VAD, avoiding some of the heuristics conventionally used in many VAD algorithms. An experimental evaluation is conducted on the CENSREC-1-C database for VAD evaluation, and the proposed method gives a significant improvement. David Cournapeau, Tatsuya Kawahara |
ICASSP | 2 |
| 2008 | Automatic lecture transcription by exploiting presentation slide information for language model adaptationabstractThe paper addresses language model adaptation for automatic lecture transcription by fully exploiting presentation slide information used in the lecture. As the text in the presentation slides is small in its size and fragmentary in its content, a robust adaptation scheme is addressed by focusing on the keyword and topic information. Several methods are investigated and combined; first, global topic adaptation is conducted based on PLSA (Probabilistic Latent Semantic Analysis) using keywords appearing in all slides. Web text is also retrieved to enhance the relevant text. Then, local preference of the keywords are reflected with a cache model by referring to the slide used during each utterance. Experimental evaluations on real lectures show that the proposed method combining the global and local slide information achieves a significant improvement of recognition accuracy, especially in the detection rate of content keywords. Tatsuya Kawahara, Yusuke Nemoto, Yuya Akita |
ICASSP | 1 |
| 2008 | Admissible stopping in viterbi beam search for unit selection in concatenative speech synthesisabstractCorpus-based concatenative speech synthesis is very popular these days due to its highly natural speech quality. The amount of computation required in the run time, however, is often quite large and various approaches have been proposed for reducing this runtime computation. In this paper, we propose early stopping schemes for Viterbi beam search in the unit selection, with which we can stop early in the local Viterbi maximization for each unit as well as in the exploration of candidate units for a given target. It takes advantage of the fact that the space of the acoustic parameters of the database units is closed and certain upper bounds of the concatenation scores can be precomputed. The proposed method for early stopping is admissible in that it does not change the result of the Viterbi beam search if the upper bounds are properly computed. Experiments show that the proposed methods of admissible stopping effectively reduce the amount of computation required in the Viterbi beam search while keeping its result unchanged. Shinichi Sakai, Tatsuya Kawahara, Shun Nakamura |
ICASSP | 2 |
| 2008 | GMM and HMM training by aggregated EM algorithm with increased ensemble sizes for robust parameter estimationabstractIn order to compensate for the weaknesses of the expectation maximization (EM) algorithm to over-training and to improve model performance for new data, we have recently proposed aggregated EM (Ag-EM) algorithm that introduces bagging-like approach in the framework of the EM algorithm and have shown that it gives similar improvements as cross-validation EM (CV EM) over conventional EM. However, a limitation with the experiments was that the number of multiple models used in the aggregation operation or the ensemble size was fixed to a small value. Here, we investigate the relationship between the ensemble size and the performance as well as giving a theoretical discussion with the order of the computational cost. The algorithm is first analyzed using simulated data and then applied to large vocabulary speech recognition on oral presentations. Both of these experiments show that Ag-EM outperforms CV-EM by using larger ensemble sizes. Takahiro Shinozaki, Tatsuya Kawahara |
ICASSP | 2 |
| 2008 | Effective error prediction using decision tree for ASR grammar network in call systemabstractCALL (computer assisted language learning) systems using ASR (automatic speech recognition) for second language learning have received increasing interest recently. However, it still remains a challenge to achieve high speech recognition performance, including accurate detection of erroneous utterances by non-native speakers. Conventionally, possible error patterns, based on linguistic knowledge, are added to the ASR grammar network. However, this approach easily falls in the trade-off of coverage of errors and the increase of perplexity. To solve the problem, we propose a method based on a decision tree to learn effective prediction of errors made by non-native speakers. An experimental evaluation with a number of foreign students in our university shows that the proposed method can effectively generate an ASR grammar network, given a target sentence, to achieve both better coverage of errors and smaller perplexity, resulting in significant improvement in ASR accuracy. Hongcui Wang, Tatsuya Kawahara |
ICASSP | 2 |
| 2008 | Statistical speech activity detection based on spatial power distribution for analyses of poster presentationsabstractThis paper proposes a microphone array based statistical speech activity detection (SAD) method for analyses of poster presentations recorded in the presence of noise. Such poster presentations are a kind of multi-party conversation, where the number of speakers and speaker location are unrestricted, and directional noise sources affect the direction of arrival of the target speech signals. To detect speech activity in such cases without a priori knowledge about the speakers and noise environments, we applied a likelihood ratio test based SAD method to spatial power distributions. The proposed method can exploit the enhanced signals obtained from timefrequency masking, and work even in the presence of environmental noise by utilizing the a priori signal-to-noise ratios of the spatial power distributions. Experiments with recorded poster presentations confirmed that the proposed method significantly improves the SAD accuracies compared with those obtained with a frequency spectrum based statistical SAD method. Index Terms: speech activity detection, microphone arrays, multi-party conversations, spatial power distribution Kentaro Ishizuka, Shoko Araki, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2008 | Multi-modal recording, analysis and indexing of poster sessionsabstractA new project on multi-modal analysis of poster sessions is introduced. We have designed an environment dedicated to recording of poster conversations using multiple sensors, and collected a number of sessions, to which a variety of multi-modal information is annotated, including utterance units for individual speakers, backchannels, nodding, gazing, and pointing. Automatic speaker diarization, that is a combination of speech activity detection and speaker identification, is conducted using a set of distant microphones, and a reasonable performance is obtained. Then, we investigate automatic classification of conversation segments into two modes: presentation mode and question-answer mode. Preliminary experiments show that multi-modal features on non-verbal behaviors play a significant role in the indexing of this kind of conversations. Index Terms: multi-modal corpus, poster conversation, speaker diarization, non-verbal information Tatsuya Kawahara, Hisao Setoguchi, Katsuya Takanashi, Kentaro Ishizuka, Shoko Araki |
INTERSPEECH | 1 |
| 2008 | Detection of feeling through back-channels in spoken dialogue
Tatsuya Kawahara, Masayoshi Toyokura, Teruhisa Misu, Chiori Hori |
INTERSPEECH | 1 |
| 2008 | Predicting ASR errors by exploiting barge-in rate of individual users for spoken dialogue systemsabstractWe exploit the barge-in rate of individual users to predict automatic speech recognition (ASR) errors. A barge-in is a situation in which a user starts speaking during a system prompt, and it can be detected even when ASR results are not reliable. Such features not using ASR results can be a clue for managing a situation in which user utterances cannot be successfully recognized. Since individual users in our system can be identified by their phone numbers, we accumulate how often each user barges in and use this rate as a user profile for determining whether a current “barge-in ” utterance should be accepted or not. We furthermore set a window that reflects the temporal transition of the user’s behavior as they get accustomed to the system. Experimental results show that setting the window improves the prediction accuracy of whether the utterance should be accepted or not. The experiments also clarify the minimum window width for improving accuracy. Index Terms: spoken dialogue system, user modeling, barge-in 1. Kazunori Komatani, Tatsuya Kawahara, Hiroshi G. Okuno |
INTERSPEECH | 2 |
| 2008 | Extracting word-pronunciation pairs from comparable set of text and speechabstractOne of the problems in text-to-speech (TTS) systems and speech-to-text (STT) systems is pronunciation estimation of unknown words. In this paper, we propose a method for extracting unknown words and their pronunciations from similar sets of Japanese text data and speech data. Out-of-vocabulary words are extracted from text with a stochastic model and pronunciations hypotheses are generated. These entries are verified by conducting automatic speech recognition on audio data. In this work, we use news articles and broadcast TV news covering similar topics. Most extracted pairs turned out to be correct according to a human judges. We also tested the TTS frontend enhanced with these entries on other web news articles, and observed an improvement in the pronunciation estimation accuracy of 9.2 % (relative). The proposed method can be used to realize a spoken language processing system that acquires and updates its lexicon automatically. 1. Tetsuro Sasada, Shinsuke Mori, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2008 | Aggregated cross-validation and its efficient application to Gaussian mixture optimizationabstractWe have previously proposed a cross-validation (CV) based Gaussian mixture optimization method that efficiently optimizes the model structure based on CV likelihood.In this study, we propose aggregated cross-validation (AgCV) that introduces a bagging-like approach in the CV framework to reinforce the model selection ability.While a single model is used in CV to evaluate a held-out subset, AgCV uses multiple models to reduce the variance in the score estimation.By integrating AgCV instead of CV in the Gaussian mixture optimization algorithm, an AgCV likelihood based Gaussian mixture optimization algorithm is obtained.The algorithm works efficiently by using sufficient statistics and can be applied to large models such as Gaussian mixture HMM.The proposed algorithm is evaluated by speech recognition experiments on oral presentations and it is shown that lower word error rates are obtained by the AgCV optimization method when compared to CV and MDL based methods. Takahiro Shinozaki, Sadaoki Furui, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2008 | A Japanese CALL system based on dynamic question generation and error prediction for ASRabstractWe have developed a new CALL system to aid students learning Japanese as a second language. The system offers students the chance to practice the Japanese grammar and vocabulary, by creating their own sentences based on visual prompts, before receiving feedback on their mistakes. Questions are dynamically generated along with sentence patterns of the lesson point, to realize variety and flexibility of the lesson. Students can give their answers with either text input or speech input. To enhance speech recognition performance, a decision tree-based method is incorporated to predict possible errors made by non-native speakers for each generated sentence on the fly. Trials of the system have been conducted with a number of foreign students in our university, and positive feedbacks were obtained. Hongcui Wang, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2008 | Test Collections for Spoken Document Retrieval from Lecture Audio Data
Tomoyosi Akiba, Kiyoaki Aikawa, Yoshiaki Itoh 0001, Tatsuya Kawahara, Hiroaki Nanjo, Hiromitsu Nishizaki, Norihito Yasuda, Yoichi Yamashita, Katunobu Itou |
LREC | 4 |
| 2007 | HMM training based on CV-EM and CV Gaussian mixture optimizationabstractA combination of the cross-validation EM (CV-EM) algorithm and the cross-validation (CV) Gaussian mixture optimization method is explored. CV-EM and CV Gaussian mixture optimization are our previously proposed training algorithms that use CV likelihood instead of the conventional training set likelihood for robust model estimation. Since CV-EM is a parameter optimization method and CV Gaussian mixture optimization is a structure optimization algorithm, these methods can be combined. Large vocabulary speech recognition experiments are performed on oral presentations. It is shown that both CV-EM and CV Gaussian mixture optimization give lower word error rates than the conventional EM, and their combination is effective to further reduce the word error rate. Takahiro Shinozaki, Tatsuya Kawahara |
ASRU | 2 |
| 2007 | Topic-Independent Speaking-Style Transformation of Language Model for Spontaneous Speech RecognitionabstractFor language modeling of spontaneous speech, we propose a novel approach, based on the statistical machine translation framework, which transforms a document-style model to the spoken style. For better coverage and more reliable estimation, incorporation of POS (part-of-speech) information is explored in addition to lexical information. In this paper, we investigate several methods that combine POS-based model or integrate POS information in the ME (maximum entropy) scheme. They achieve significant reduction in perplexity and WER in a meeting transcription task. Moreover, the model is applied to different domains or committee meetings of different topics. As a result, even larger perplexity reduction is achieved compared with the case tested in the same domain. The result demonstrates the generality and portability of the model. Yuya Akita, Tatsuya Kawahara |
ICASSP (4) | 2 |
| 2007 | Automatic Detection of Sentence and Clause Units using Local Syntactic DependencyabstractFor robust detection of sentence and clause units in spontaneous speech such as lectures and meetings, we propose a novel cascaded chunking strategy which incorporates syntactic and semantic information. Application of general syntactic parsing is difficult for spontaneous speech having ill-formed sentences and disfluencies, especially for erroneous transcripts generated by ASR systems. Therefore, we focus on the local syntactic dependency of adjacent words and phrases, and train binary classifiers based on SVM (support vector machines) for this purpose. An experimental evaluation using spontaneous talks of the CSJ (Corpus of Spontaneous Japanese) demonstrates that the proposed dependency analysis can be robustly performed and is effective for clause/sentence unit detection in ASR outputs. Tatsuya Kawahara, Masahiro Saikou, Katsuya Takanashi |
ICASSP (4) | 1 |
| 2007 | Speech-Based Interactive Information Guidance System using Question-Answering TechniqueabstractThis paper addresses an interactive framework for information navigation based on document knowledge base. In conventional audio guidance systems, such as those deployed in museums, the information flow is one-way and the content is fixed. In order to make an interactive guidance system, we propose the application of question-answering (QA) techniques. Since users tend to use anaphoric expressions in successive questions, we investigate appropriate handling of contextual information based on topic detection, together with the effect of using N-best information in ASR output. Moreover, we apply the QA technique to generation of system-initiative information recommendation. A navigation system on Kyoto city information was implemented. Effectiveness of the proposed techniques was confirmed through a field trial by a number of real novice users. Teruhisa Misu, Tatsuya Kawahara |
ICASSP (4) | 2 |
| 2007 | An Interactive Framework for Document Retrieval and Presentation with Question-Answering Function in Restricted Domain
Teruhisa Misu, Tatsuya Kawahara |
IEA/AIE | 2 |
| 2007 | PLSA-based topic detection in meetings for adaptation of lexicon and language modelabstractA topic detection approach based on a probabilistic framework is proposed to realize topic adaptation of speech recognition systems for long speech archives such as meetings. Since topics in such speech are not clearly defined unlike news stories, we adopt a probabilistic representation of topics based on probabilistic latent semantic analysis (PLSA). A topical sub-space is constructed by PLSA, and speech segments are projected to the subspace, then each segment is represented by a vector which consists of topic probabilities obtained by the projection. Topic detection is performed by clustering these vectors, and topic adaptation is done by collecting relevant texts based on the similarity in this probabilistic representation. In experimental evaluations, the proposed approach demonstrated significant reduction of perplexity and outof-vocabulary rates as well as robustness against ASR errors. Yuya Akita, Yusuke Nemoto, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2007 | Evaluation of real-time voice activity detection based on high order statisticsabstractWe have proposed a method for real-time, unsupervised voice activity detection (VAD). In this paper, problems of feature selection and classification scheme are addressed. The feature is based on High Order Statistics (HOS) to discriminate close and far-field talk, enhanced by a feature derived from the normalized autocorrelation. Comparative effectiveness on several HOS is shown. The classification is done in real-time with a recursive, online EM algorithm. The algorithm is evaluated on the CENSREC-1-C database, which is used for VAD evaluation for automatic speech recognition (ASR) [1], and the proposed method is confirmed to significantly outperform the baseline energy-based method. Index Terms: Voice activity detection, online EM, high order statistics David Cournapeau, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2007 | Analyzing temporal transition of real user's behaviors in a spoken dialogue systemabstractManaging various behaviors of real users is indispensable for spoken dialogue systems to operate adequately in real environments. We have analyzed various users ’ behaviors using data collected over 34 months from the Kyoto City Bus Information System. We focused on “barge-in ” and added barge-in rates to our analysis. Temporal transitions of users ’ behaviors, such as automatic speech recognition (ASR) accuracy, task success rates and barge-in rates, were initially investigated. We then examined the relationship between ASR accuracy and barge-in rates. Analysis revealed that the ASR accuracy of utterances inputted with barge-ins was lower because many novices, who were not accustomed to the timing when to utter, used the system. We also observed that the ASR accuracy of utterances with barge-ins differed based on the barge-in rates of individual users. The results indicate that the barge-in rate can be used as a novel user profile for detecting ASR errors. Index Terms: spoken dialogue system, real user behavior, barge-in Kazunori Komatani, Tatsuya Kawahara, Hiroshi G. Okuno |
INTERSPEECH | 2 |
| 2007 | Bayes risk-based optimization of dialogue management for document retrieval system with speech interfaceabstractAbstract We propose an efficient dialogue management for an informa-tionnavigationsystembasedonadocument knowledge base. Itis expected that incorporation of appropriate N-best candidatesofASRandcontextualinformationwillimprovethesystemper-formance. The system also has several choices in generatingresponses or confirmations. In this paper, this selection is opti-mizedasminimizationofBayesriskbasedonrewardforcorrectinformation presentation and penalty for redundant turns. Wehave evaluated this strategy with our spoken dialogue system“Dialogue Navigator for Kyoto City”, which also has question-answering capability. Effectiveness of the proposed frameworkwas confirmed in the success rate of retrieval and the averagenumber of turns for information access. Index Terms : spoken dialogue system, dialogue management,Bayes risk 1. Introduction The target of spoken dialogue systems is being extended fromsimple databases such as flight information to general docu-ments including manuals and newspaper articles[1]. In suchsystems, the automatic speech recognition (ASR) result of theuser utterance is matched against a set of target documents us-ing the vector space model, and documents with high matchingscores are presented to the user. These kinds of applicationsare expected to be useful especially when retrieving informa-tion with keyboardless devices such as a PDA, a tablet PC anda car navigation system. We have developed “Dialogue Navi-gator for Kyoto City”, which can make interactive guidance byincorporating the question-answering (QA) function as well asgeneral document retrieval.In these systems, making confirmation is needed to elim-inate misunderstandings caused by ASR errors, but so manyredundant confirmations are bothersome to users. There havebeenmanystudiesthatdealwithefficientdialoguemanagementto make confirmation[2, 3, 4]. However, most of them are de-signed for relational database (RDB) query tasks, which havea definite set of keywords, and they are not directly applica-ble to document retrieval tasks, in which every word is used inmatching. In document retrieval tasks, therefore, it is more rea-sonable to make confirmation considering the confidence scoreof retrieval.In addition, our system classifies user utterances to queriesand questions, and generates appropriate responses for respec-tive inputs. Unlike conventional QA tasks, such as TREC QATrack[5], it is not obvious whether the utterance is a query orU1: Please explain the Golden Pavilion.S1: The Golden Pavilion is one of the buildings in theRokuon-ji in Kyoto, and is the main attraction of thetemple sites. The entire pavilion except the basementfloor is covered with pure gold leaf.U2: When was it built?S2: The Golden Pavilion was originally built in 1397to serve as a retirement villa for Shogun AshikagaYoshimitsu.U3: Then, please tell me its history.......Figure 1: Example dialogue of Dialogue Navigator for KyotoCitya question. In addition, an exact answer for a question doesnot necessarily exist in the document sets. Therefore, it is notalways optimal to respond the question with its answer alone.In this paper, we address the extension of conventional op-timization methods of dialogue management, to be applicableto general document retrieval tasks with QA function. Specif-ically, we propose a dialogue management that optimizes thechoicesinresponsegenerationbyminimizingBayesrisk,basedon reward for correct information presentation and penalty forredundant turns, which are defined by the score of documentretrieval and answer extraction. Teruhisa Misu, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2007 | Gaussian mixture optimization for HMM based on efficient cross-validationabstractA Gaussian mixture optimization method is explored using cross-validation likelihood as an objective function instead of the conventional training set likelihood. The optimization is based on reducing the number of mixture components by selecting and merging a pair of Gaussians step by step base on the objective function so as to remove redundant components and improve the generality of the model. Cross-validation likelihood is more appropriate for avoiding over-fitting than the conventional likelihood and can be efficiently computed using sufficient statistics. It results in a better Gaussian pair selection and provides a termination criterion that does not rely on empirical thresholds. Large-vocabulary speech recognition experiments on oral presentations show that the cross-validation method gives a smaller word error rate with an automatically determined model size than a baseline training procedure that does not perform the optimization. Index Terms: speech recognition, HMM, Gaussian mixture, cross-validation, sufficient statistics Takahiro Shinozaki, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2007 | Evaluating and optimizing Japanese tutor system featuring dynamic question generation and interactive guidanceabstractWe are developing a new CALL system to aid students learning Japanese as a second language. This system is designed to allow students to create their own sentences based on visual prompts, receiving feedback based on their mistakes. The questions are dynamically generated, resulting in a large variety of challenges. The students may choose to receive guidance in order to complete each task, selecting the level of help that best suits their needs. A scoring system is also incorporated, which awards a grade to students based on the errors made and hints used. The trial of the system has been conducted with twenty one students, providing the statistics of actual errors and hint usages. With these data, we have trained the weights of the scoring system by taking into account the impact of each issue on the proficiency of the students. The validity of the estimated score is generally confirmed by predicting the proficiency of the students. Index Terms: CALL, second language learning, Japanese Christopher J. Waple, Hongcui Wang, Tatsuya Kawahara, Yasushi Tsubota, Masatake Dantsuji |
INTERSPEECH | 3 |
| 2007 | Real-Time Continuous Speech Recognition System on SH-4A MicroprocessorabstractTo expand CSR (continuous speech recognition) software to the mobile environmental use, we have developed embedded version of Julius (embedded Julius). Julius is open source CSR software, and has been used by many researchers and developers in Japan as a standard decoder on PCs. In this paper, we describe an implementation of the embedded Julius on a SH-4A microprocessor. SH-4A is a high-end 32-bit MPU (720 MIPS) with on-chip FPU. However, further computational reduction is necessary for the embedded Julius to operate realtime. Applying some optimizations, the embedded Julius achieves real-time processing on the SH-4A. The experimental results show 0.89 times RT(real-time), resulting 4.0 times faster than baseline CSR. We also evaluated the embedded Julius on large vocabulary (20,000 words). It shows almost real-time processing (1.25 times RT). Hiroaki Kokubo, Nobuo Hataoka, Akinobu Lee, Tatsuya Kawahara, Kiyohiro Shikano |
MMSP | 4 |
| 2007 | Out-of-Domain Utterance Detection Using Classification Confidences of Multiple TopicsabstractOne significant problem for spoken language systems is how to cope with users' out-of-domain (OOD) utterances which cannot be handled by the back-end application system. In this paper, we propose a novel OOD detection framework, which makes use of the classification confidence scores of multiple topics and applies a linear discriminant model to perform in-domain verification. The verification model is trained using a combination of deleted interpolation of the in-domain data and minimum-classification-error training, and does not require actual OOD data during the training process, thus realizing high portability. When applied to the "phrasebook" system, a single utterance read-style speech task, the proposed approach achieves an absolute reduction in OOD detection errors of up to 8.1 points (40% relative) compared to a baseline method based on the maximum topic classification score. Furthermore, the proposed approach realizes comparable performance to an equivalent system trained on both in-domain and OOD data, while requiring no OOD data during training. We also apply this framework to the "machine-aided-dialogue" corpus, a spontaneous dialogue speech task, and extend the framework in two manners. First, we introduce topic clustering which enables reliable topic confidence scores to be generated even for indistinct utterances, and second, we implement methods to effectively incorporate dialogue context. Integration of these two methods into the proposed framework significantly improves OOD detection performance, achieving a further reduction in equal error rate (EER) of 7.9 points Ian Lane, Tatsuya Kawahara, Tomoko Matsui, Satoshi Nakamura 0001 |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Detection of Quotations and Inserted Clauses and Its Application to Dependency Structure Analysis in Spontaneous Japanese
Ryoji Hamabe, Kiyotaka Uchimoto, Tatsuya Kawahara, Hitoshi Isahara |
ACL | 3 |
| 2006 | Efficient Estimation of Language Model Statistics of Spontaneous Speech Via Statistical Transformation ModelabstractOne of the most significant problems in language modeling of spontaneous speech such as meetings and lectures is that only limited amount of matched training data, i.e. faithful transcript for the relevant task domain, is available. In this paper, we propose a novel transformation approach to estimate language model statistics of spontaneous speech from a document-style text database, which is often available with a large scale. The proposed statistical transformation model is designed for modeling characteristic linguistic phenomena in spontaneous speech and estimating their occurrence probabilities. These contextual patterns and probabilities are derived from a small amount of parallel aligned corpus of the faithful transcripts and their document-style texts. To realize wide coverage and reliable estimation, a model based on part-of-speech (POS) is also prepared to provide a back-off scheme from a word-based model. The approach has been successfully applied to estimation of the language model for National Congress meetings from their minute archives, and significant reduction of test-set perplexity is achieved Yuya Akita, Tatsuya Kawahara |
ICASSP (1) | 2 |
| 2006 | Sentence boundary detection of spontaneous Japanese using statistical language model and support vector machinesabstractAbstract This paper presents two different approaches utilizing sta-tisticallanguagemodel(SLM)andsupportvectormachines(SVM) for sentence boundary detection of spontaneousJapanese. In the SLM-based approach, linguistic likeli-hoods and occurrence of pause are used to determine sen-tence boundaries. To suppress false alarms, heuristic pat-terns of end-of-sentence expressions are also incorporated.On the other hand, SVM is adopted to realize robust classi-fication against a wide variety of expressions and speechrecognition errors. Detection is performed by an SVM-based text chunker using lexical and pause information asfeatures. We evaluated these approaches on manual and au-tomatic transcription of spontaneous lectures and speeches,and achieved F-measures of 0.85 and 0.78, respectively. Index Terms : sentence boundary detection, spontaneousspeech, statistical language model, support vector ma-chines. 1. Introduction Recent advance of automatic speech recognition (ASR)technology, especially for spontaneous speech, enables var-ious applications such as spoken document archiving andretrieval, speech summarization and speech translation. Toorganizeaspokendocumentinastructuredformandtogiveuseful indices, transcriptions should be segmented into ap-propriate units like sentences. Moreover, these applicationsare usually built by combining an ASR system with naturallanguage processing (NLP) systems such as a parser and amachine translator, which often assume that input text is asentence. However,sentencesinspontaneousspeechareill-formed, and sentence boundaries are indistinct. Output textby ASR systems is just a sequence of words and has no ex-plicitsentenceboundaries, sothefurtherstepofsegmentingthe ASR output is required for these applications.Automatic boundary detection of spoken sentences hasbeen explored mainly on broadcast news (BN)tasks[1,2,3]and conversational telephone speech (CTS) tasks[3, 4, 5] inEnglish. As features for detection, pause, prosodic and lin-guistic information is often used. Most popular approach isa combination of prosodic and linguistic information[2, 3],which realizes high performance on BN and CTS tasks.Prosody-based approaches[1, 5] have also been investi-gated. Meanwhile, linguistic information is not used by it-self, since most of these works were performed on Englishdata, where cue words or expressions of sentence bound-aries are not easily defined.In Japanese, cue expressions are typically observed atthe end of sentences and expected to be useful for bound-ary detection. However, variety of such expressions is solarge in spoken Japanese, that it is hard to collect sufficientamountofdatafortrainingstatisticalmodelssuchasamax-imum entropy (ME) model. Moreover, many of cue expres-sions consist of particles, which are apparently difficult tobedetectedinASR.Thus,robustnessforASRerrorsshouldalso be investigated.In this paper, we address two approaches of sen-tence boundary detection for spontaneous Japanese. Asframeworks of detection, we adopt and compare statisti-cal language model (SLM) and support vector machines(SVM). The proposed approaches are evaluated with reallectures and speeches included in the Corpus of Sponta-neous Japanese (CSJ). Yuya Akita, Masahiro Saikou, Hiroaki Nanjo, Tatsuya Kawahara |
INTERSPEECH | 4 |
| 2006 | Voice activity detector based on enhanced cumulant of LPC residual and on-line EM algorithmabstractThis paper addresses the problem of segmenting audio data recorded with embedded devices for the purpose of intelligent sensing in the context of multi-modal interactions. We propose a real-time method for robust speech detection in natural, noisy environments. It is based on a fusion of high order statistics of the LPC residual and autocorrelation, and adopts an on-line version of Expectation Maximization algorithm for the classification. Experimental evaluations show that the proposed method provides better detection performance under different types of natural noises, working robustly against other voices in the context of multi-speaker interactive situations. As the proposed method is based on features which have a low computational cost, and has a small latency, it is suitable for real-time tracking applications. David Cournapeau, Tatsuya Kawahara, Kenji Mase, Tomoji Toriyama |
INTERSPEECH | 2 |
| 2006 | Detection of quotations and inserted clauses and its application to dependency structure analysis in spontaneous JapaneseabstractJapanese dependency structure is usually represented by relationships between phrasal units called bunsetsus. One of the biggest problems with dependency structure analysis in spontaneous speech is that clause boundaries are ambiguous. This paper describes a method for detecting the boundaries of quotations and inserted clauses and that for improving the dependency accuracy by applying the detected boundaries to dependency structure analysis. The quotations and inserted clauses are determined by using an SVM-based text chunking method that considers information on morphemes, pauses, fillers, etc. The information on automatically analyzed dependency structure is also used to detect the beginning of the clauses. Our evaluation experiment using Corpus of Spontaneous Japanese (CSJ) showed that the automatically estimated boundaries of quotations and inserted clauses helped to improve the accuracy of dependency structure analysis. Ryoji Hamabe, Kiyotaka Uchimoto, Tatsuya Kawahara, Hitoshi Isahara |
INTERSPEECH | 3 |
| 2006 | Evaluation of voice activity detection by combining multiple features with weight adaptationabstractFor noise-robust automatic speech recognition (ASR), we propose a novel voice activity detection (VAD) method based on a combination of multiple features. The scheme uses a weighted combination of four conventionalVAD features: amplitude level, zero crossing rate, spectral information, and Gaussian mixture model (GMM) likelihood. The weights for combination are adaptively updated using minimum classification error (MCE) training. In this paper, we first investigate the effect of adaptation of the combination weights and GMM parameters, and demonstrate that the weights can be effectively adapted with a single utterance. Then, we present application of the method to ASR. It is confirmed that the proposed method significantly outperforms conventional methods in various noise conditions. Index Terms: speech recognition, voice activity detection, MCE training, noise adaptation Yusuke Kida, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2006 | A bootstrapping approach for developing language model of new spoken dialogue systems by selecting web textsabstractThis paper proposes a bootstrapping method of constructing statistical language models for new spoken dialogue systems by collecting and selecting sentences from the World Wide Web (WWW). To make effective search queries that cover the target domain in full detail, we exploit the document set described about the target domain as seeding data. An important issue is how to filter the retrieved Web pages, since all of the retrieved Web texts are not necessarily suitable as training data. We induct an existing dialogue corpus of different domain to prefer the texts of spoken style. The proposed method was evaluated on two different tasks of software support and sightseeing guidance, and significant reduction of the word error rate was achieved. We show that it is vital to incorporate the dialogue corpus, though not relevant to the target domain, in the text selection phase. Index Terms: speech recognition, language model, spoken dialogue system, web text selection. Teruhisa Misu, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2006 | Decision tree-based training of probabilistic concatenation models for corpus-based speech synthesisabstractThe measure of the goodness, or cost, of concatenating synthesis units plays an important role in concatenative speech synthesis. In this paper, we present a probabilistic approach to concatenation modeling in which the goodness of concatenation is represented as the conditional probability of observing the spectral shape of a unit given the previous unit and the current phonetic context. This conditional probability is modeled by a conditional Gaussian density whose mean vector has a form of linear transform of the past spectral shape. A phonetic decision-tree based parameter tying is performed to achieve a robust training that balances between model complexity and the amount of training data available. The concatenation models are implemented in a corpus-based speech synthesizer trained with a CMU Arctic database and the effectiveness of the proposed method was confirmed by a subjective listening test. Index Terms: speech synthesis, unit selection, join costs. Shinsuke Sakai, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2006 | Prototyping a call system for students of Japanese using dynamic diagram generation and interactive hintsabstractIn this paper we discuss the concept and design of a new CALL (Computer Assisted Language Learning) system being developed to aid students learning Japanese as a second language. The system is being designed to allow students to create and speak their own sentences based on visual prompts, before receiving feedback on their mistakes. The students may choose to receive guidance in order to complete each task, selecting the level of help that best suits their needs. Having described the concept of the system and its design, we discuss some tests recently carried out using a prototype of the system, summarize the results obtained, and give some thought as to the significance of these results on the future development of the system. Index Terms: CALL, second language learning, Japanese, sentence generation, interactive help. Christopher J. Waple, Yasushi Tsubota, Masatake Dantsuji, Tatsuya Kawahara |
INTERSPEECH | 4 |
| 2006 | Dependency-structure Annotation to Corpus of Spontaneous Japanese
Kiyotaka Uchimoto, Ryoji Hamabe, Takehiko Maruyama, Katsuya Takanashi, Tatsuya Kawahara, Hitoshi Isahara |
LREC | 5 |
| 2006 | Embedded Julius: Continuous Speech Recognition Software for MicroprocessorabstractTo expand CSR (continuous speech recognition) software to the mobile environmental use, we have developed embedded version of "Julius". Julius is open source CSR software, and has been used by many researchers and developers in Japan as a standard decoder on PCs. Julius works as a real time decoder on a PC. However further computational reduction is necessary to use Julius on a microprocessor. Further cost reduction is needed. For reducing cost of calculating pdfs (probability density function), Julius adopts a GMS (Gaussian mixture selection) method. In this paper, we modify the GMS method to realize a continuous speech recognizer on microprocessors. This approach does not change the structure of acoustic models in consistency with that used by conventional Julius, and enables developers to use acoustic models developed by popular modeling tools. On simulation, the proposed method has archived 20% reduction of computational costs compared to conventional GMS, 40% reduction compared to no GMS. Finally, the embedded version of Julius was tested on a developmental hardware platform named "T-engine". The proposed method showed 2.23 of RTF (real time factor) resulting 79% of that of no GMS without any degradation of recognition performance Hiroaki Kokubo, Hiroaki Hataoka, Akinobu Lee, Tatsuya Kawahara, Kiyohiro Shikano |
MMSP | 4 |
| 2006 | Dialogue strategy to clarify user's queries for document retrieval system with speech interface
Teruhisa Misu, Tatsuya Kawahara |
Speech Commun. | 2 |
| 2005 | Generalized Statistical Modeling of Pronunciation Variations using Variable-length Phone ContextabstractPronunciation variation modeling is one of the major issues in automatic transcription of spontaneous speech. We present statistical modeling of subword-based mapping between baseforms and surface forms using a large-scale spontaneous speech corpus (CSJ). Variation patterns of phone sequences are automatically extracted together with their contexts of up to two preceding and following phones, which are decided by their occurrence statistics. We then derive a set of rewrite rules with their probabilities and variable-length phone contexts. The model effectively predicts pronunciation variations depending on the phone context using a back-off scheme. Since it is based on phone sequences, the model is applicable to any lexicon to generate appropriate surface forms. The proposed method was evaluated on two transcription tasks whose domains are different from the training corpus (CSJ), and significant reduction of word error rates was achieved. Yuya Akita, Tatsuya Kawahara |
ICASSP (1) | 2 |
| 2005 | Incorporating Dialogue Context and Topic Clustering in Out-of-Domain DetectionabstractThe detection and handling of OOD (out-of-domain) user utterances are significant problems for spoken language systems. We have proposed a novel OOD detection framework, which makes use of classification confidence scores of multiple topics. In this paper, we extend this framework in order to handle natural language dialogue. Specifically, two issues are addressed. First, to effectively incorporate dialogue context, we investigate methods to combine multiple utterances at various stages of the OOD detection process. Second, to improve robustness on spontaneous speech, we introduce a topic clustering scheme which provides reliable topic classification confidence even for indistinct utterances. The system is evaluated on natural dialogue via the ATR speech-to-speech translation system, and a significant improvement in OOD detection accuracy was achieved by incorporating the two proposed techniques. Ian Lane, Tatsuya Kawahara |
ICASSP (1) | 2 |
| 2005 | A New ASR Evaluation Measure and Minimum Bayes-Risk Decoding for Open-domain Speech UnderstandingabstractA new evaluation measure of speech recognition and a decoding strategy for keyword-based open-domain speech understanding are presented. Conventionally, WER (word error rate) has been widely used as an evaluation measure of speech recognition, which treats all words in a uniform manner. We define a weighted keyword error rate (WKER) which gives a weight on errors from a viewpoint of information retrieval. We first demonstrate that this measure is more appropriate for predicting the performance of key sentence indexing of oral presentations. Then, we formulate a decoding method to minimize WKER based on a minimum Bayes-risk (MBR) framework, and show that the decoding method works reasonably for improving WKER and key sentence indexing. Hiroaki Nanjo, Tatsuya Kawahara |
ICASSP (1) | 2 |
| 2005 | Voice activity detection based on optimally weighted combination of multiple features
Yusuke Kida, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2005 | Utterance verification incorporating in-domain confidence and discourse coherence measuresabstractConventional confidence measures for assessing the reliability of ASR output are typically derived from “low-level ” information which is obtained during speech recognition decoding. In contract to these approaches, we propose a novel utterance verification scheme which incorporates confidence measures derived from “high-level ” knowledge sources. Specifically, we investigate two measures: in-domain confidence, the degree of match between the input utterance and the application domain of the back-end system, and discourse coherence, the consistency between consecutive utterances in a dialogue session. A joint verification confidence is generated by combining these two measures with an orthodox measure based on GPP (generalized posterior probability). The proposed verification scheme was evaluated on spontaneous dialogue via the ATR speech-tospeech translation system. The two proposed measures were effective in improving verification accuracy. 1. Ian Lane, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2005 | Dialogue strategy to clarify user's queries for document retrieval system with speech interfaceabstractAbstract This paper proposes a dialogue strategy for clarifying and constraining queries to document retrieval systems with speech input interfaces. It is indispensable for spoken dialogue systems to interpret user’s intention robustly in the presence of speech recognition errors and extraneous expressions characteristic of spontaneous speech. In speech input, moreover, users’ queries tend to be vague, and they may need to be clarified through dialogue in order to extract sufficient information to get meaningful retrieval results. In conventional database query tasks, it is easy to cope with these problems by extracting and confirming keywords based on semantic slots. However, it is not straightforward to apply such a methodology to general document retrieval tasks. In this paper, we first introduce two statistical measures for identifying critical portions to be confirmed. The relevance score (RS) represents the matching degree with the document set. The significance score (SS) detects portions that affect retrieval results. With these measures, the system can generate confirmations to handle speech recognition errors, prior to and after the retrieval, respectively. Then, we propose a dialogue strategy for generating clarifications to narrow down the retrieved items, especially when many documents are matched because of a vague input query. The optimal clarification question is dynamically selected based on information gain (IG) – the reduction in the number of matched items. A set of possible clarification questions is prepared using various knowledge sources. As a bottom-up knowledge source, we extract a list of words that can take a number of objects and potentially causes ambiguity, using a dependency structure analysis of the document texts. This is complemented by top-down knowledge sources of metadata and hand-crafted questions. Our dialogue strategy is implemented and evaluated against a software support knowledge base of 40 K entries. We demonstrate that our strategy significantly improves the success rate of retrieval. Teruhisa Misu, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2005 | Minimum Bayes-risk decoding considering word significance for information retrieval systemabstractThe paper addresses a new evaluation measure of automatic speech recognition (ASR) and a decoding strategy oriented for speech-based information retrieval (IR). Although word error rate (WER), which treats all words in a uniform manner, has been widely used as an evaluation measure of ASR, significance of words are different in speech understanding or IR. In this paper, we define a new ASR evaluation measure, namely, weighted word error rate (WWER) that gives a weight on errors from a viewpoint of IR. Then, we formulate a decoding method to minimize WWER based on Minimum BayesRisk (MBR) framework, and show that the decoding method improves WWER and IR accuracy. Hiroaki Nanjo, Teruhisa Misu, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2005 | Trigger-based language model adaptation for automatic meeting transcriptionabstractWe present a novel trigger-based language model adaptation method oriented to the transcription of meetings. In meetings, the topic is focused and consistent throughout the whole session, therefore keywords can be correlated over long distances. The trigger-based language model is designed to capture such long-distance dependencies, but it is typically constructed from a large corpus, which is usually too general to derive task-dependent trigger pairs. In the proposed method, we make use of the initial speech recognition results to extract task-dependent trigger pairs and to estimate their statistics. Moreover, we introduce a back-off scheme that also exploits the statistics estimated from a large corpus. The proposed model reduced the testset perplexity twice as much as the typical trigger-based language model constructed from a large corpus, and achieved a remarkable perplexity reduction of 41% over the baseline when combined with an adapted trigram language model. Carlos Troncoso, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2005 | Speaker model selection based on the Bayesian information criterion applied to unsupervised speaker indexingabstractIn conventional speaker recognition tasks, the amount of training data is almost the same for each speaker, and the speaker model structure is uniform and specified manually according to the nature of the task and the available size of the training data. In real-world speech data such as telephone conversations and meetings, however, serious problems arise in applying a uniform model because variations in the utterance durations of speakers are large, with numerous short utterances. We therefore propose a flexible framework in which an optimal speaker model (GMM or VQ) is automatically selected based on the Bayesian Information Criterion (BIC) according to the amount of training data available. The framework makes it possible to use a discrete model when the data is sparse, and to seamlessly switch to a continuous model after a large amount of data is obtained. The proposed framework was implemented in unsupervised speaker indexing of a discussion audio. For a real discussion archive with a total duration of 10 hours, we demonstrate that the proposed method has higher indexing performance than that of conventional methods. The speaker index is also used to adapt a speaker-independent acoustic model to each participant for automatic transcription of the discussion. We demonstrate that speaker indexing with our method is sufficiently accurate for adaptation of the acoustic model. Masafumi Nishida, Tatsuya Kawahara |
IEEE Trans. Speech Audio Process. | 2 |
| 2005 | User Modeling in Spoken Dialogue Systems to Generate Flexible Guidance
Kazunori Komatani, Shinichi Ueno, Tatsuya Kawahara, Hiroshi G. Okuno |
User Model. User Adapt. Interact. | 3 |
| 2004 | Efficient Confirmation Strategy for Large-scale Text Retrieval Systems with Spoken Dialogue Interface
Kazunori Komatani, Teruhisa Misu, Tatsuya Kawahara, Hiroshi G. Okuno |
COLING | 3 |
| 2004 | Dependency Structure Analysis and Sentence Boundary Detection in Spontaneous Japanese
Kazuya Shitaoka, Kiyotaka Uchimoto, Tatsuya Kawahara, Hitoshi Isahara |
COLING | 3 |
| 2004 | Out-of-domain detection based on confidence measures from multiple topic classificationabstractOne significant problem for spoken language systems is how to cope with users' OOD (out-of-domain) utterances which cannot be handled by the back-end system. In this paper, we propose a novel OOD detection framework, which makes use of classification confidence scores of multiple topics and trains a linear discriminant in-domain verifier using gradient probabilistic descent (GPD). Training is based on deleted interpolation of the in-domain data, and thus does not require actual OOD data, providing high portability. Three topic classification schemes of word N-gram models, latent semantic analysis (LSA), and support vector machines (SVM) are evaluated, and SVM is shown to have the greatest discriminative ability. In an OOD detection task, the proposed approach achieves an absolute reduction in equal error rate (EER) of 6.5% compared to a baseline method based on a simple combination of multiple-topic classifications. Furthermore, comparison with a system trained using OOD data demonstrates that the proposed training scheme realizes comparable performance while requiring no knowledge of the OOD data set. Ian Lane, Tatsuya Kawahara, Tomoko Matsui, Satoshi Nakamura 0001 |
ICASSP (1) | 2 |
| 2004 | Real-time word confidence scoring using local posterior probabilities on tree trellis searchabstractConfidence scoring based on word posterior probability is usually performed as a post process of speech recognition decoding, and also needs a large number of word hypotheses to get enough confidence quality. We propose a simple way of computing the word confidence using estimated posterior probability while decoding. At the word expansion of stack decoding search, the local sentence likelihoods that contain heuristic scores of unreached segment are directly used to compute the posterior probabilities. Experimental results showed that, although the likelihoods are not optimal, we can provide slightly better confidence measures compared with N-best lists, while the computation is faster than the 100-best method because no N-best decoding is required. Akinobu Lee, Kiyohiro Shikano, Tatsuya Kawahara |
ICASSP (1) | 3 |
| 2004 | Automatic indexing of key sentences for lecture archives using statistics of presumed discourse markersabstractAutomatic extraction of key sentences from lecture audio archives is addressed. The method makes use of the characteristic expressions used in initial utterances of sections, which are defined as discourse markers and derived in a totally unsupervised manner based on word statistics. The statistics of the presumed discourse markers are then used to define the importance of the sentences. It is also combined with the conventional tf-idf measure of content words. Experimental results using a large corpus of lectures confirm the effectiveness of the method based on the discourse markers and its combination with the keyword-based method. It is also shown that the method is robust against ASR errors and sentence segmentation accuracy is more vital. Thus, we also enhance segmentation by incorporating prosodic information. Hiroaki Nanjo, Tasuku Kitade, Tatsuya Kawahara |
ICASSP (1) | 3 |
| 2004 | Speaker indexing and adaptation using speaker clustering based on statistical model selectionabstractThe paper addresses unsupervised speaker indexing and automatic speech recognition of discussions. In speaker indexing, there are two cases, where the number of speakers is unknown beforehand and where the number is known. When the specified number is unknown, it is difficult to apply to various data because it needs to determine several parameters like threshold. In addition, serious problems arise in applying a uniform model because variations in the utterance durations of speakers are large. We thus propose a method which can robustly perform speaker indexing for the two cases using a flexible framework in which an optimal speaker model (GMM or VQ) is selected based on the BIC (Bayesian information criterion). Moreover, we propose a combination method of speaker adaptation based on speaker selection and the indexing method. For real discussion archives, we demonstrated that indexing performance is higher than that of conventional methods for the two cases and speech recognition performance was improved by the combination method. Masafumi Nishida, Tatsuya Kawahara |
ICASSP (1) | 2 |
| 2004 | Automatic audio archiving system for panel discussionsabstractWe present an automatic audio archiving system suitable for panel discussions. In our archive framework, audio data, speech transcription, speaker and content based indices are integrated in order to realize efficient archive browsing. Speaker indexing is performed in a totally unsupervised manner. The speaker information is also used for enhancing the automatic speech recognition system. These results are aligned with audio segments. Moreover we also introduce a novel indexing of utterances based on discourse tags that represent intentions and importance of utterances. A discourse tagger combining rule based and statistical methods is developed to automatically generate high-level indices. Finally, these results are combined and encoded using an MPEG-7 framework, resulting in highly portable archives. Yuya Akita, Masahiro Hasegawa, Tatsuya Kawahara |
ICME | 3 |
| 2004 | Recognition of Emotional States in Spoken Dialogue with a Robot
Kazunori Komatani, Ryosuke Ito, Tatsuya Kawahara, Hiroshi G. Okuno |
IEA/AIE | 3 |
| 2004 | Language model adaptation based on PLSA of topics and speakersabstractWe address an adaptation method of statistical language models to topics and speaker characteristics for automatic transcription of meetings and discussions. A baseline language model is a mixture of two models, which are trained with different corpora covering various topics and speakers, respectively. Then, probabilistic latent semantic analysis (PLSA) is performed on the same respective corpora and the initial ASR result to provide unigram probabilities conditioned on input speech. Finally, the baseline model is adapted by scaling N-gram probabilities with these unigram probabilities. For speaker adaptation purpose, we make use of spontaneous speech corpus (CSJ) in which a large number of speakers gave talks for given topics. Experimental evaluation with real discussions showed that both topic and speaker adaptation improved test-set perplexity and word accuracy. Yuya Akita, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2004 | Practical use of English pronunciation system for Japanese students in the CALL classroomabstractWe have developed an English pronunciation learning system which estimates the intelligibility of students’ speech and ranks their errors from the viewpoint of improving their intelligibility. We have begun using this system in a CALL class at Kyoto University. We have evaluated system performance through the use of questionnaires and analysis of speech data logged in the server, and will present our findings in this paper. Tatsuya Kawahara, Masatake Dantsuji, Yasushi Tsubota |
INTERSPEECH | 1 |
| 2004 | Topic classification and verification modeling for out-of-domain utterance detection
Tatsuya Kawahara, Ian Lane, Tomoko Matsui, Satoshi Nakamura 0001 |
INTERSPEECH | 1 |
| 2004 | Recent progress of open-source LVCSR engine julius and Japanese model repositoryabstractContinuous Speech Recognition Consortium (CSRC) was founded for further enhancement of Japanese Dictation Toolkit that had been developed by the support of a Japanese agency. Overview of its product software is reported in this paper. The open-source LVCSR (large vocabulary continuous speech recognition) engine Julius has been improved both in performance and functionality, and it is also ported to Microsoft Windows in compliance with SAPI (Speech API). The software is now used for not a few languages and plenty of applications. For plug-and-play speech recognition in various applications, we have also compiled a repository of acoustic and language models for Japanese. Especially, the acoustic model set realizes wider coverage of user generations and speech-input environments. Tatsuya Kawahara, Akinobu Lee, Kazuya Takeda, Katunobu Itou, Kiyohiro Shikano |
INTERSPEECH | 1 |
| 2004 | Automatic transformation of lecture transcription into document style using statistical frameworkabstractThis paper addresses automatic transformation from spoken style texts to written style texts. Exact transcriptions and speech recognition results of live lectures include many spoken language expressions, and thus, are not suitable for documents and need to be edited. In this paper, we present a method of applying of the statistical approach used in machine translation to this postprocessing task. Specifically, we implement the correction of colloquial expressions, the deletion of fillers, the insertion of periods, and the insertion of particles in an integrated manner. A preliminary evaluation confirms that the statistical transformation framework works well and we achieved high recall and precision rate of period and particle insertion. Tatsuya Kawahara, Kazuya Shitaoka, Hiroaki Nanjo |
INTERSPEECH | 1 |
| 2004 | Dependency structure analysis and sentence boundary detection in spontaneous Japanese
Tatsuya Kawahara, Kiyotaka Uchimoto, Hitoshi Isahara, Kazuya Shitaoka |
INTERSPEECH | 1 |
| 2004 | Automatic extraction of key sentences from oral presentations using statistical measure based on discourse markersabstractAutomatic extraction of key sentences from academic presentation speeches is addressed. The method makes use of the characteristic expressions used in initial utterances of sections, which are defined as discourse markers and derived in a totally unsupervised manner based on word statistics. The statistics of the discourse markers are then used to define the importance of the sentences. It is also combined with the conventional tf-idf measure of content words. Comprehensive evaluation using the Corpus of Spontaneous Japanese and a variety of experimental setups is presented in this paper. We carefully designed the evaluation scheme to be compared to human performance. The proposed method using the discourse markers shows consistent effectiveness in the key sentence extraction. Based on the indexing, we realize efficient browsing of lecture audio archives. Tasuku Kitade, Tatsuya Kawahara, Hiroaki Nanjo |
INTERSPEECH | 2 |
| 2004 | Example-based training of dialogue planning incorporating user and situation modelsabstractTo provide a high level of usability, spoken dialogue systems must generate cooperative responses for a wide variety of users and situations. We introduce a dialogue planning scheme incorporating user and situation models making such dialogue adaptation possible. Manually developing a set of dialogue rules to account for all possible model combinations, would be very difficult and obstruct system portability. To overcome this problem, we propose a novel example-based training scheme for dialogue planning, where example dialogues from a role-playing simulation are collected and a machine learning approach is used to train the dialogue planner. The proposed scheme is evaluated on the Kyoto city voice portal, a multi-domain spoken dialogue system. Subjects participated in a role-playing simulation where they selected appropriate system responses at each dialogue turn based on a given scenario. Experimental results show that the system successfully trains the dialogue planner and provides reasonable system performance. Ian Lane, Tatsuya Kawahara, Shinichi Ueno |
INTERSPEECH | 2 |
| 2004 | Confirmation strategy for document retrieval systems with spoken dialog interfaceabstractAdequate confirmation is indispensable in spoken dialog systems to eliminate misunderstandings caused by speech recognition errors. Spoken language also inherently includes redundant expressions such as disfluency and out-of-domain phrases, which do not contribute to task achievement. It is easy to define a set of keywords to be confirmed for conventional database query tasks, but not straightforward in general document retrieval tasks. In this paper, we propose two statistical measures for identifying portions to be confirmed. A relevance score (RS) represents matching degree with the document set. A significance score (SS) detects portions that consequently affect the retrieval results. With these measures, the system can generate confirmation prior to and posterior to the retrieval, respectively. The strategy is implemented and evaluated with retrieval from software support knowledge base of 40K entries. It is shown that the proposed strategy using the two measures is more efficient than using the conventional confidence measure. Teruhisa Misu, Tatsuya Kawahara, Kazunori Komatani |
INTERSPEECH | 2 |
| 2004 | Introduction to the Special Issue on Spontaneous Speech Processing
Sadaoki Furui, Mary E. Beckman, Julia Hirschberg, Shuichi Itahashi, Tatsuya Kawahara, Satoshi Nakamura 0001, Shri Narayanan |
IEEE Trans. Speech Audio Process. | 5 |
| 2004 | Automatic indexing of lecture presentations using unsupervised learning of presumed discourse markersabstractA new method for automatic detection of section boundaries and extraction of key sentences from lecture audio archives is proposed. The method makes use of 'discourse markers' (DMs), which are characteristic expressions used in initial utterances of sections, together with pause and language model information. The DMs are derived in a totally unsupervised manner based on word statistics. An experimental evaluation using the Corpus of Spontaneous Japanese (CSJ) demonstrates that the proposed method provides better indexing of section boundaries compared with a simple baseline method using pause information only, and that it is robust against speech recognition errors. The method is also applied to extraction of key sentences that can index the section topics. The statistics of the presumed DMs are used to define the importance of sentences, which favors potentially section-initial ones. The measure is also combined with the conventional tf-idf measure based on content words. Experimental results confirm the effectiveness of using the DMs in combination with the keyword-based method. The paper also describes a statistical framework for transforming raw speech transcriptions into the document style for defining appropriate sentence units and improving readability. Tatsuya Kawahara, Masahiro Hasegawa, Kazuya Shitaoka, Tasuku Kitade, Hiroaki Nanjo |
IEEE Trans. Speech Audio Process. | 1 |
| 2004 | Language model and speaking rate adaptation for spontaneous presentation speech recognitionabstractThe paper addresses adaptation methods to language model and speaking rate (SR) of individual speakers which are two major problems in automatic transcription of spontaneous presentation speech. To cope with a large variation in expression and pronunciation of words depending on the speaker, firstly, we investigate the effect of statistical and context-dependent pronunciation modeling. Secondly, we present unsupervised methods of language model adaptation to a specific speaker and a topic by 1) selecting similar texts based on the word perplexity and TF-IDF measure and 2) making direct use of the initial recognition result for generating an enhanced model. We confirm that all proposed adaptation methods and their combinations reduce the perplexity and word error rate. We also present a decoding strategy adapted to the SR. In spontaneous speech, SR is generally fast and may vary a lot. We also observe different error tendencies for portions of presentations where speech is fast or slow. Therefore, we propose a SR-dependent decoding strategy that applies the most appropriate acoustic analysis, phone models, and decoding parameters according to the SR. Several methods are investigated and their selective application leads to improved accuracy. The combined effect of the two proposed adaptation methods is also confirmed in transcription of real academic presentation. Hiroaki Nanjo, Tatsuya Kawahara |
IEEE Trans. Speech Audio Process. | 2 |
| 2003 | Flexible Guidance Generation Using User Model in Spoken Dialogue SystemsabstractWe address appropriate user modeling in order to generate cooperative responses to each user in spoken dialogue systems. Unlike previous studies that focus on user's knowledge or typical kinds of users, the user model we propose is more comprehensive. Specifically, we set up three dimensions of user models: skill level to the system, knowledge level on the target domain and the degree of hastiness. Moreover, the models are automatically derived by decision tree learning using real dialogue data collected by the system. We obtained reasonable classification accuracy for all dimensions. Dialogue strategies based on the user modeling are implemented in Kyoto city bus information system that has been developed at our laboratory. Experimental evaluation shows that the cooperative responses adaptive to individual users serve as good guidance for novice users without increasing the dialogue duration for skilled users. Kazunori Komatani, Shinichi Ueno, Tatsuya Kawahara, Hiroshi G. Okuno |
ACL | 3 |
| 2003 | Language model switching based on topic detection for dialog speech recognitionabstractAn efficient, scalable speech recognition architecture is proposed for multidomain dialog systems by combining topic detection and topic-dependent language modeling. The inferred domain is automatically detected from the user's utterance, and speech recognition is then performed with an appropriate domain-dependent language model. The architecture improves accuracy and efficiency over current approaches and is scaleable to a large number of domains. In this paper, a novel framework using a multilayer hierarchy of language models is introduced in order to improve robustness against topic detection errors. The proposed system provides a relative reduction in WER of 10.5% over a single language model system. Furthermore it achieves an accuracy that is comparable to using multiple language models in parallel while using only a fraction of the computational cost. Ian Lane, Tatsuya Kawahara, Tomoko Matsui |
ICASSP (1) | 2 |
| 2003 | Unsupervised speaker indexing using speaker model selection based on Bayesian information criterionabstractThe paper addresses unsupervised speaker indexing for discussion audio archives. In discussions, the speaker changes frequently, thus the duration of utterances is very short and its variation is large, which causes significant problems in applying conventional methods such as model adaptation and variance-BIC (Bayesian information criterion) methods. We propose a flexible framework that selects an optimal speaker model (GMM or VQ) based on the BIC according to the duration of utterances. When the speech segment is short, the simple and robust VQ-based method is expected to be chosen, while GMM can be reliably trained for long segments. For a discussion archive having a total duration of 10 hours, it is demonstrated that the proposed method achieves higher indexing performance than that of conventional methods. Masafumi Nishida, Tatsuya Kawahara |
ICASSP (1) | 2 |
| 2003 | Unsupervised speaker indexing using anchor models and automatic transcription of discussionsabstractWe present unsupervised speaker indexing combined with automatic speech recognition (ASR) for speech archives such as discussions. Our proposed indexing method is based on anchor models, by which we define a feature vector based on the similarity with speakers of a large scale speech database. Several techniques are introduced to improve discriminant ability. ASR is performed using the results of this indexing. No discussion corpus is available to train acoustic and language models. So we applied the speaker adaptation technique to the baseline acoustic model based on the indexing. We also constructed a language model by merging two models that cover different linguistic features. We achieved the speaker indexing accuracy of 93% and the significant improvement of ASR for real discussion data. Yuya Akita, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2003 | Spoken dialogue system for queries on appliance manuals using hierarchical confirmation strategyabstractWe address a dialogue framework for queries on manuals of electric appliances with a speech interface. Users can makequeriesbyunconstrainedspeech, fromwhichkeywords are extracted and matched to the items in the manual. As a result, so many items are usually obtained. Thus, we introduce an effective dialogue strategy which narrows down the items using a tree structure extracted from the manual. Three cost functions are presented and compared to minimize the number of dialogue turns. We have evaluated the system performance on VTR manual query task. The numberof averagedialogueturnsis reduced to 71% using our strategy compared with a conventional method that makes confirmation in turn according to the matching likelihood. Thus, the proposed system helps users find their intended items more efficiently. Tatsuya Kawahara, Ryosuke Ito, Kazunori Komatani |
INTERSPEECH | 1 |
| 2003 | User modeling in spoken dialogue systems for flexible guidance generationabstractWe address appropriate user modeling in order to generate cooperative responses to each user in spoken dialogue systems. Unlike previous studies that focus on users’ knowledge or typical kinds of users, the proposed user model is more comprehensive. Specifically, we set up three dimensions of user models: skill level to the system, knowledge level on the target domain and degree of hastiness. Moreover, the models are automatically derived by decision tree learning using real dialogue data. We obtained reasonable classification accuracy for all dimensions. Dialogue strategies based on the user modeling are implemented in Kyoto city bus information system that has been developed at our laboratory. Experimental evaluation shows that the cooperative responses adaptive to individual users serve as good guidance for novice users without increasing the dialogue duration for skilled users. Kazunori Komatani, Shinichi Ueno, Tatsuya Kawahara, Hiroshi G. Okuno |
INTERSPEECH | 3 |
| 2003 | Hierarchical topic classification for dialog speech recognition based on language model switchingabstractA speech recognition architecture combining topic detection and topic-dependent language modeling is proposed. In this architecture, a hierarchical back-off mechanism is introduced to improve system robustness. Detailed topic models are applied when topic detection is confident, and wider models that cover multiple topics are applied in cases of uncertainty. In this paper, two topic detection methods are evaluated for the architecture: unigram likelihood and SVM (Support Vector Machine). On the ATR Basic Travel Expression corpus, both topic detection methods provide a comparable reduction in WER of 10.0% and 11.1 % respectively over a single language model system. Finally the proposed re-decoding approach is compared with an equivalent system based on re-scoring. It is shown that redecoding is vital to provide optimal recognition performance. 1. Ian Lane, Tatsuya Kawahara, Tomoko Matsui, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2003 | Speaker model selection using Bayesian information criterion for speaker indexing and speaker adaptationabstractThis paper addresses unsupervised speaker indexing for discussion audio archives. We propose a flexible framework that selects an optimal speaker model (GMM or VQ) based on the Bayesian Information Criterion (BIC) according to input utterances. The framework makes it possible to use a discrete model whenthedataissparse, andtoseamlesslyswitchtoacontinuous model after a large cluster is obtained. The speaker indexing is also applied and evaluated at automatic speech recognition of discussions by adapting a speaker-independent acoustic model to each participant. It is demonstrated that indexing with our method is sufficiently accurate for the speaker adaptation. Masafumi Nishida, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2002 | Efficient Dialogue Strategy to Find Users' Intended Items from Information Query Results
Kazunori Komatani, Tatsuya Kawahara, Ryosuke Ito, Hiroshi G. Okuno |
COLING | 2 |
| 2002 | Automatic indexing of lecture speech by extracting topic-independent discourse markersabstractAutomatic detection of section (sub-topic) boundaries in lecture speech is addressed. The method makes use of the characteristic expressions used in initial utterances of sections defined as discourse makers, as well as pause and language model information. The discourse markers are derived in a totally unsupervised manner based on word statistics used in the information retrieval technique. The statistics is used to select candidates picked up by other information. Experimental results show that the proposed method realizes better indexing performance (better precision at high recall rates) than the simple baseline method using pause information only. Moreover, it is shown to be robust against speech recognition errors. Tatsuya Kawahara, Masahiro Hasegawa |
ICASSP | 1 |
| 2002 | Speaking-rate dependent decoding and adaptation for spontaneous lecture speech recognitionabstractThis paper addresses the problem of speaking rate in large vocabulary spontaneous speech recognition. In spontaneous lecture speech, the speaking rate is generally fast and may vary a lot within a talk. We also observed different error tendencies for fast and slow speech segments. Therefore, we first present a speaking-rate dependent decoding strategy that applies the most adequate acoustic analysis, phone models and decoding parameters according to the speaking rate. Several methods are investigated and their selective application leads to accuracy improvement. We also propose to make use of speaking-rate information in speaker adaptation, in which the different adapted models are set up for fast and slow utterances. It is confirmed that the method is more effective than normal adaptation. Hiroaki Nanjo, Tatsuya Kawahara |
ICASSP | 2 |
| 2002 | Modeling and automatic detection of English sentence stress for computer-assisted English prosody learning system
Kazunori Imoto, Yasushi Tsubota, Antoine Raux, Tatsuya Kawahara, Masatake Dantsuji |
INTERSPEECH | 4 |
| 2002 | Speaking rate compensation based on likelihood criterion in acoustic model training and decoding
Kozo Okuda, Tatsuya Kawahara, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2002 | Automatic intelligibility assessment and diagnosis of critical pronunciation errors for computer-assisted pronunciation learningabstractWe introduce a novel method to diagnose pronunciation errors that are most critical to the intelligibility of L2 learners. A preliminary study showed that error rates computed by a speech recognition-based system can be used to characterize intelligibility. We deduce a probabilistic algorithm to derive intelligibility from error rates. We also define an error priority function that indicates which errors are most critical to intelligibility. Experimental results proved the validity of the approach. Antoine Raux, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2002 | Recognition and verification of English by Japanese students for computer-assisted language learning system
Yasushi Tsubota, Tatsuya Kawahara, Masatake Dantsuji |
INTERSPEECH | 2 |
| 2002 | Belief network based disambiguation of object reference in spoken dialogue system for robotabstractWe are studying joint activity in which a remote robot finds an object by communicating with the user over a voice-only channel. We focus on how the robot disambiguates the reference of the uttered word or phrase to the target object. For example, by “cup”, one may refer to a “teacup”, a “coffee cup”, or even a “glass” under some situations. This reference (hereafter, “object reference”) is user-dependent. We confirm that a user model of object references is significant by conducting a survey of 12 subjects. In addition to ambiguity of object reference, actual systems should cope with two other sources of uncertainty in speech and image recognition. We present a Belief Network based probabilistic reasoning system to determine the object reference. The resulting system demonstrates that the number of interactions needed to find a common reference is reduced as the user model is refined. Yoko Yamakata, Tatsuya Kawahara, Hiroshi G. Okuno |
INTERSPEECH | 2 |
| 2002 | Continuous Speech Recognition Consortium an Open Repository for CSR Tools and Models
Akinobu Lee, Tatsuya Kawahara, Kazuya Takeda, Masato Mimura, Atsushi Yamada, Akinori Ito, Katunobu Itou, Kiyohiro Shikano |
LREC | 2 |
| 2001 | Gaussian mixture selection using context-independent HMMabstractWe address a method to efficiently select Gaussian mixtures for fast acoustic likelihood computation. It makes use of context-independent models for selection and back-off of corresponding triphone models. Specifically, for the k-best phone models by the preliminary evaluation, triphone models of higher resolution are applied, and others are assigned likelihoods with the monophone models. This selection scheme assigns more reliable back-off likelihoods to the un-selected states than the conventional Gaussian selection based on a VQ codebook. It can also incorporate efficient Gaussian pruning at the preliminary evaluation, which offsets the increased size of the pre-selection model. Experimental results show that the proposed method achieves comparable performance as the standard Gaussian selection, and performs much better under aggressive pruning condition. Together with the phonetic tied-mixture modeling, acoustic matching cost is reduced to almost 14% with little loss of accuracy. Akinobu Lee, Tatsuya Kawahara, Kiyohiro Shikano |
ICASSP | 2 |
| 2001 | Domain-independent spoken dialogue platform using key-phrase spotting based on combined language modelabstractWe present a portable platform for spoken dialogue systems and its experimental evaluation. Conventional development of speech interfaces involves much labor cost in either describing a task grammar or collecting a task corpus. Our platform automatically generates a lexicon and a language model of keyphrases based on task description and structure of the domain database. By spotting key-phrases using both the generated grammar and word 2-gram model trained with dialogue corpora of similar domains, we realize flexible speech understanding on a variety of utterances. Furthermore, adopting a GUI that explicitly displays acceptable utterance patterns is effective in guiding user utterances within the system’s capability. We evaluate the generated spoken dialogue system using 24 novice users. The number of unacceptable utterances are significantly reduced with the simple phrase grammar and GUI. And the phrase spotter using the combined language model improves the semantic accuracy by 15.5% compared with the conventional method decoding the whole sentence with a fixed grammar. Kazunori Komatani, Katsuaki Tanaka, Hiroaki Kashima, Tatsuya Kawahara |
INTERSPEECH | 4 |
| 2001 | Julius - an open source real-time large vocabulary recognition engineabstractJulius is a high-performance, two-pass LVCSR decoder for researchers and developers. Based on word 3-gram and context-dependent HMM, it can perform almost real-time decoding on most current PCs in 20k word dictation task. Major search techniques are fully incorporated such as tree lexicon, N-gram factoring, cross-word context dependency handling, enveloped beam search, Gaussian pruning, Gaussian selection, etc. Besides search efficiency, it is also modularized carefully to be independent from model structures, and various HMM types are supported such as shared-state triphones and tied-mixture models, with any number of mixtures, states, or phones. Standard formats are adopted to cope with other free modeling toolkit. The main platform is Linux and other Unix workstations, and partially works on Windows. Julius is distributed with open license together with source codes, and has been used by many researchers and developers in Japan. Akinobu Lee, Tatsuya Kawahara, Kiyohiro Shikano |
INTERSPEECH | 2 |
| 2001 | Speaking rate dependent acoustic modeling for spontaneous lecture speech recognitionabstractThe paper addresses large vocabulary spontaneous speech recognition focusing on acoustic modeling that considers the speaking rate. Using the real lecture speech corpus collected under the priority research project in Japan, we have made baseline acoustic model, and evaluated on the automatic transcription of oral presentations by experienced speakers and obtained word accuracy of 58.2%. Compared with read speech, we have observed significant difference in the speaking rate. To handle fast and poorly articulated phone segments, several extensions of the modeling are explored. Specifically, we introduce stateskipping modeling, speech rate-dependent model, and syllable sub-word modeling. As a result, we reduced the word error rate by absolute 0.8%-2.0%. We also address a language modeling especially on effective use of various large text corpora. Hiroaki Nanjo, Kazuomi Kato, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2000 | Flexible Mixed-Initiative Dialogue Management using Concept-Level Confidence Measures of Speech Recognizer Output
Kazunori Komatani, Tatsuya Kawahara |
COLING | 2 |
| 2000 | A new phonetic tied-mixture model for efficient decodingabstractA phonetic tied-mixture (PTM) model for efficient large vocabulary continuous speech recognition is presented. It is synthesized from context-independent phone models with 64 mixture components per state by assigning different mixture weights according to the shared states of triphones. Mixtures are then re-estimated for optimization. The model achieves a word error rate of 7.0% with a 20000-word dictation of newspaper corpus, which is comparable to the best figure by the triphone of much higher resolutions. Compared with conventional PTMs that share Gaussians by all states, the proposed model is easily trained and reliably estimated. Furthermore, the model enables the decoder to perform efficient Gaussian pruning. It is found out that computing only two out of 64 components does not cause any loss of accuracy. Several methods for the pruning are proposed and compared, and the best one reduced the computation to about 20%. Akinobu Lee, Tatsuya Kawahara, Kazuya Takeda, Kiyohiro Shikano |
ICASSP | 2 |
| 2000 | Overview of an intelligent system for information retrieval based on human-machine dialogue through spoken language
Hiroya Fujisaki, Katsuhiko Shirai, Shuji Doshita, Seiichi Nakagawa, Keikichi Hirose, Shuichi Itahashi, Tatsuya Kawahara, Sumio Ohno, Hideaki Kikuchi, Kenji Abe, Shinya Kiriyama |
INTERSPEECH | 7 |
| 2000 | Modelling of the perception of English sentence stress for computer-assisted language learningabstractABSTRACTFor learning foreign language pronunciation, prosodic fea-tures are important as much as, or more than segmen-tal features. For Japanese speakers, one of difficulties tolearn English pronunciation is rhythm because of the dif-ferences between two languages: mora-timing rhythm andstress-timing rhythm. In order to correct errors in rhythm,the method of evaluating sentence stress that constitutesrhythm is very significant. In this paper, we present amethod of automatic detecting sentence stress syllables forthe evaluation criterion. Using a linear discriminant func-tion of pitch, intensity and vowel duration, about 90% ofthe syllables were correctly detected as to sentence stress.Also we analyzed the different and common characteristicsamong different English native speakers. The results re-vealed that the perception of the sentence stress among 11native speakers had general agreement with respect to howto integrate three features.1. INTRODUCTIONProsodic features play an important role in human commu-nications as to make clear focal point of topics, and alsoto emphasize or express one’s intention. However, the dif-ferences between English and Japanese in the expressionsof stress and rhythm cause serious difficulties for learnersmastering prosodic patterns. Japanese speakers tend to ut-ter English pronunciation in monotonous rhythm, whichoccasionally misleads their communication. So effectiveevaluationand instructionofrhythmare essentialtoacquirecorrect English pronunciation.Recent advancement in speech technology enables usto develop a computer-assisted language learning (CALL)system. In the previous studies[1][2][3], effective evalua-tion criteria for segmental features or intonation were pro-posed. Also, the method of evaluating and instructing En-glish word accent for Japanese was proposed[4]. However,the evaluation criterion of sentence stress, which is one ofthe most essential factors of English rhythm, is not estab-lished. So we propose a method of automatic detectingsentence stress syllables for a base of an effective CALLsystem.It is known that English syllable stresses consist ofpitch, intensity and duration. These features correspondto fundamental frequency, power and vowel duration, re-spectively. In [6], the rhythm instruction was realized byjudging the stress with duration and vowel quality, not withpitch or intensity. According to [8], pitch and power alsobecome key factors of sentence stress. So we utilized threefeatures to detect sentence stresses. However, it is not clearwhich acoustic feature is the most important or how thesefeatures are integrated when English native speakers per-ceive sentence stresses. We investigate these issues and tryto establish a universal evaluation criterion of English sen-tence stresses.In order to estimate the appearance of sentence stresssyllables from contours of fundamental frequency andpower, each acoustic feature is to be normalized and quan-tized by syllable units. Then we adopt a linear discrim-inant function that integrates these acoustic features withweights. The weights, which reflect the importance of eachacoustic feature in perceiving English sentence stresses,are estimated by discriminant analysis using the TIMITdatabase.From the degree of agreement between natives’ per-ception and our method by a linear discriminant function,we verify the validity of our model. Also by comparisonwith the weights that are estimated from different natives’labels, we see the common and different characteristicsamong different natives’ perception.2. SPEECH MATERIALAs materials for the study, we picked up 310 sentencesproduced by English native speakers with New Englanddialect from the TIMIT database. Those speech sampleswere labeled by eleven English native speakers. From theirperception, we eliminated inadequate speech samples thathad neutral declarative rhythm to improve the reliability ofspeech materials and those labels. Kazunori Imoto, Masatake Dantsuji, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2000 | Automatic transcription of lecture speech using topic-independent language modelingabstractWe approach lecture speech recognition with a topicindependent language model and its adaptation. As lecture speech has its characteristic style that is different from newspapers and conversations, dedicated language modeling is needed. The problem is that, although lectures have many keywords specific to the topic and fields, available corpus of each domain is limited in size. Thus, we introduce topic-independent modeling with a vocabulary selection mechanism based on a mutual information criterion. It realizes better coverage and accuracy with small complexity than the conventional word frequency-based method. This baseline model is adapted to specific lectures using preprint texts. We have tried automatic transcription of oral presentations and achieved a word error rate of 23.6% on the average. Kazuomi Kato, Hiroaki Nanjo, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2000 | Free software toolkit for Japanese large vocabulary continuous speech recognitionabstractA sharable software repository for Japanese LVCSR (Large Vocabulary Continuous Speech Recognition) is introduced. It is designed as a baseline platform for research and developed by researchers of different academic institutes under a governmental support. The repository consists of a recognition engine (Julius), Japanese acoustic models and statistical language models as well as Japanese morphological analysis tools. These modules can be easily integrated and replaced under a plug-and-play framework, which makes it possible to fairly evaluate components and to develop specific application systems. Assessment of these modules and systems in a 20000-word dictation task is reported. The software repository is freely available to the public. Tatsuya Kawahara, Akinobu Lee, Tetsunori Kobayashi, Kazuya Takeda, Nobuaki Minematsu, Shigeki Sagayama, Katunobu Itou, Akinori Ito, Mikio Yamamoto, Atsushi Yamada, Takehito Utsuro, Kiyohiro Shikano |
INTERSPEECH | 1 |
| 2000 | Generating effective confirmation and guidance using two-level confidence measures for dialogue systemsabstractWe present a method to generate effective confirmation and guidance using concept-level confidence measures (CM) derived from speech recognizer output in order to handle speech recognition errors. We define two conceptlevel CM, which are on content-words and on semanticattributes, using 10-best outputs of the speech recognizer and parsing with phrase-level grammars. Content-word CM is useful for selecting plausible interpretations. Less confident interpretations are given to confirmation process, and non-confident ones are rejected. The strategy improved the interpretation accuracy by 11.5%. Moreover, the semantic-attribute CM is used to estimate user’s intention and generates system-initiative guidances even when successful interpretation is not obtained. Kazunori Komatani, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2000 | Automatic diagnosis of recognition errors in large vocabulary continuous speech recognition systemsabstractAutomatic diagnosis of recognition errors in large vocabulary continuous speech recognition (LVCSR) systems is addressed. It consists of two steps. The first step is to identify the module that causes recognition errors for every erroneous segment. This statistics points out which modules to be revised. The second step is to analyze the causes of the errors in detail. Specifically, the triphone and N-gram entries related to the errors are listed. The diagnostic information provides directions for improvement. This diagnosis has been applied to three LVCSR systems: read speech dictation system, lecture speech transcription system and dialogue speech recognition system. We have observed different and interesting diagnosis results. In the dictation system, the diagnosis is useful for improving our decoder Julius. In the lecture and dialogue speech recognition systems, problems in acoustic and language modeling are made clear. Hiroaki Nanjo, Akinobu Lee, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2000 | Computer-assisted English vowel learning system for Japanese speakers using cross language formant structuresabstractWe present a novel Computer-Assisted Language Learning (CALL) system for Japanese students who learn English as a second language. We regard formant structure of Japanese vowels pronounced by Japanese learners of English as their own formant structure. This structure is transformed to learner’s ideal English formant structure based on the relationship between English and Japanese articulation charts both of which are corresponded with formant structure. When the learners’ English pronunciation is input, it is compared with the estimated ideal one, and articulatory instructions are given. We verified that the mapping and estimation of English vowel parameters are correct with bilingual speakers’ speech and observed the learning effect with five students who tried the system. Yasushi Tsubota, Masatake Dantsuji, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2000 | IPA Japanese Dictation Free Software Project
Katunobu Itou, Kiyohiro Shikano, Tatsuya Kawahara, Kazuya Takeda, Atsushi Yamada, Akinori Ito, Takehito Utsuro, Tetsunori Kobayashi, Nobuaki Minematsu, Mikio Yamamoto, Shigeki Sagayama, Akinobu Lee |
LREC | 3 |
| 1999 | Topic independent language model for key-phrase detection and verificationabstractA topic independent lexical and language modeling for robust key-phrase detection and verification is presented. Instead of assuming a domain specific lexicon and language model, our model is designed to characterize filler phrases depending on the speaking-style, thus can be trained with large corpora of different topics but the same style. Mutual information criterion is used to select topic independent filler words and their N-gram model is used for verification of key-phrase hypotheses. A dialogue-style dependent filler model improves the key-phrase detection in different dialogue applications. A lecture-style dependent model is trained with transcriptions of various oral presentations by filtering out topic specific words. It performs much better verification of key-phrases uttered during lectures of different topics compared with the conventional syllable-based model and large vocabulary model. Tatsuya Kawahara, Shuji Doshita |
ICASSP | 1 |
| 1998 | Automatic pronunciation error detection and guidance for foreign language learningabstractWe propose an e ective application of speech recognition to foreign language pronunciation learning. The objective of our system is to detect pronunciation errors and provide diagnostic feedback through speech processing and recognition methods. Automatic pronunciation error detection is used for two kinds of mispronunciation, that is mistake and linguistical inheritance. The correlation between automatic detection and human judgement shows its reliability. For the feedback guidance to an erroneous phone, we set up classi ers for the well-recognized articulatory features, the place of articulation and the manner of articulation, in order to identify the cause of incorrect articulation. It provides feedback guidance on how to correct mispronunciation. Chul-Ho Jo, Tatsuya Kawahara, Shuji Doshita, Masatake Dantsuji |
ICSLP | 2 |
| 1998 | Speaking-style dependent lexicalized filler model for key-phrase detection and verification
Tatsuya Kawahara, Kentaro Ishizuka, Shuji Doshita |
ICSLP | 1 |
| 1998 | Sharable software repository for Japanese large vocabulary continuous speech recognitionabstractThe project of Japanese LVCSR (Large Vocabulary Continuous Speech Recognition) platform is introduced. It is a collaboration of researchers of different academic institutes and intended to develop a sharable software repository of not only databases but also models and programs. The platform consists of a standard recognition engine, Japanese phone models and Japanese statistical language models. A set of Japanese phone HMMs are trained with ASJ (Acoustic Society of Japan) databases of 20K sentence utterances per each gender. Japanese word N-gram (2-gram and 3-gram) models are constructed with a corpus of Mainichi newspaper of four years. The recognition engine JULIUS is developed for assessment of both acoustic and language models. The modules are integrated as a Japanese LVCSR system and evaluated on 5000-word dictation task. The software repository is available to the public. Tatsuya Kawahara, Tetsunori Kobayashi, Kazuya Takeda, Nobuaki Minematsu, Katunobu Itou, Mikio Yamamoto, Atsushi Yamada, Takehito Utsuro, Kiyohiro Shikano |
ICSLP | 1 |
| 1998 | An efficient two-pass search algorithm using word trellis indexabstractWe propose an e cient two-pass search algorithm for LVCSR. Instead of conventional word graph, the rst preliminary pass generates \word trellis index, keeping track of all survived word hypotheses within the beam every time-frame. As it represents all found word boundaries non-deterministically, we can (1) obtain accurate sentence-dependent hypotheses on the second search, and (2) avoid expensive word-pair approximation on the rst pass. The second pass performs an e cient stack decoding search, where the index is referred to as predicted word list and heuristics. Experimental results on 5,000-word Japanese dictation task show that, compared with the word-graph method, this trellis-based method runs with less than 1/10 memory cost while keeping high accuracy. Finally, by handling inter-word context dependency, we achieved the word error rate of 5.6%. Akinobu Lee, Tatsuya Kawahara, Shuji Doshita |
ICSLP | 2 |
| 1998 | Prosodic analysis of fillers and self-repair in Japanese speech
Felix C. M. Quimbo, Tatsuya Kawahara, Shuji Doshita |
ICSLP | 2 |
| 1998 | Flexible speech understanding based on combined key-phrase detection and verificationabstractWe propose a novel speech understanding strategy based on combined detection and verification of semantically tagged key-phrases in spontaneous spoken utterances. Key-phrases are defined in a top-down manner so as to constitute semantic slots. Their detection directly leads to robust understanding. A phrase network realizes both a wide coverage and a reasonable constraint for detection. A subword-based verifier is then incorporated to reduce false alarms in detection and attach confidence measures of the detected phrases. This set of phrase confidence measures, when incorporated in a spoken dialogue system, forms a basis for designing intelligent speech interfaces that accept only verified key-phrases and reprompt users to clarify unspecified or unrecognized portions. Several forms of confidence measures based on subword-level tests are investigated. The proposed approach was tested on field data collected from real-world trial applications. The combined detection and verification strategy drastically improves the accuracy in handling out-of-grammar utterances over the conventional decoding approaches while maintaining the performance for in-grammar utterances. Tatsuya Kawahara, Biing-Hwang Juang |
IEEE Trans. Speech Audio Process. | 1 |
| 1997 | Combining key-phrase detection and subword-based verification for flexible speech understandingabstractA flexible speech understanding framework combining key-phrase detection and verification is presented. Detection of semantically-tagged key-phrases directly leads to robust understanding. In order to select reliable detection and eliminate false alarms, utterance verification technique is incorporated. A phrase verifier combines subword-based likelihood ratios of correct models and anti-subword alternate models. A confidence measure that focuses on mis-matched subwords is proposed and demonstrated as the most effective. The combined strategy drastically improves the semantic accuracy for out-of-grammar utterances, while maintaining the performance for in-grammar samples. We also found that utterance verification applied after grammar-based decoding is not so effective as the proposed detection and verification strategy. Tatsuya Kawahara, Biing-Hwang Juang |
ICASSP | 1 |
| 1997 | Task adaptation using MAP estimation in N-gram language modelingabstractDescribes a method of task adaptation in N-gram language modeling for accurately estimating the N-gram statistics from the small amount of data of the target task. Assuming a task-independent N-gram to be a-priori knowledge, the N-gram is adapted to a target task by MAP (maximum a-posteriori probability) estimation. Experimental results showed that the perplexities of the task-adapted models were 15% (trigram) and 24% (bigram) lower than those of the task-independent model, and that the perplexity reduction of the adaptation went up to a maximum of 39% when the amount of text data in the adapted task was very small. Hirokazu Masataki, Yoshinori Sagisaka, Kazuya Hisaki, Tatsuya Kawahara |
ICASSP | 4 |
| 1996 | Concept-based phrase spotting approach for spontaneous speech understandingabstractIn order to realize robust speech understanding, we present a phrase spotting approach. The use of concept-based phrases as the core unit for understanding is advantageous because the phrase-level constraint realizes wider coverage and stable matching and they are directly mapped to semantic cases. The phrase spotting and the sentence-level parsing are formulated as a progressive search. It realizes an optimal search to spot phrase candidates with Viterbi scoring and A* search to combine the phrase candidates into optimal sentence hypotheses. The approach achieved higher detection rates and robust interpretation of ill-formed utterances. We also examined the effect of the background language model. It is shown that lexical knowledge in the background is vital for spotting and the use of the acoustic score of the filler model is significant for parsing. Tatsuya Kawahara, Norihide Kitaoka, Shuji Doshita |
ICASSP | 1 |
| 1996 | Key-phrase detection and verification for flexible speech understanding
Tatsuya Kawahara, Biing-Hwang Juang |
ICSLP | 1 |
| 1994 | Heuristic search integrating syntactic, semantic and dialog-level constraintsabstractWe present a new model of speech understanding, based on the cooperation of the speech recognizer and language analyzer, which interacts with the knowledge sources while keeping its modularity. The semantic analyzer is realized with a semantic network that represents the possible concepts in a task. The speech recognizer based on an LR parser interacts with the semantic analyzer to eliminate invalid hypotheses at an early stage. The coupling of a loose grammar and interactive semantic analysis accepts ill-formed sentences while filtering out non-sense ones, thus realizes robust understanding. Dialog-level knowledge is also incorporated to constrain both the syntactic and the semantic knowledge sources. The key to guide the search efficiently is powerful heuristics. The relationship between the heuristic power and search efficiency is examined experimentally. The stochastic word bigram is derived from the probabilistic LR grammar as A*-admissible heuristics.> Tatsuya Kawahara, Masahiro Araki, Shuji Doshita |
ICASSP (2) | 1 |
| 1994 | Keyword and phrase spotting with heuristic language model
Tatsuya Kawahara, Toshihiko Munetsugu, Norihide Kitaoka, Shuji Doshita |
ICSLP | 1 |
| 1992 | HMM based on pair-wise Bayes classifiersabstractA novel hidden Markov model (HMM) architecture which realizes both high discriminating ability and stochastic scoring is presented. In modifying continuous HMM so that the states of the models are best separated, distinctive features vary for different states or different models, and different discriminant functions should be made for different competing states. For every pair of the states, a Bayes classifier which performs a vector transformation based on discriminant analysis is constructed. Each classifier ranks the two states and computes a relative value of the probabilities. Output probabilities of the HMM states are obtained by combining and normalizing the results of the pair-wise classifications. Training of the classifiers and HMMs is done interactively and iteratively so that they are optimized totally. Experimental results show that the method, called the pair-wise Bayes classifier-HMM (PWBC-HMM) is more effective than the conventional HMM. It realizes robust recognition by modifying pattern space to fully separate confusing classes, while retaining analog outputs by statistical Bayes classifiers.> Tatsuya Kawahara, Shuji Doshita |
ICASSP | 1 |
| 1991 | Phoneme recognition by combining discriminant analysis and HMMabstractThe authors present a novel phoneme recognition method which combines two stochastic methods; discriminant analysis and the hidden Markov model (HMM) method. The HMM is powerful in time-warping and in capturing the global dynamic features, but its discriminating ability is not sufficient. The approach used is to construct the HMM with a phonetic element classifier front-end. Each phonetic element belongs to one phoneme and represents a local pattern of the phoneme. The classifier is a modified version of discriminant analysis, that is, a combination of Bayes classifiers. It extracts optimal features to separate the phonetic elements and consequently contributes to separate HMMs of the phonemes. Furthermore, the score of the classifier is combined with that of the HMM. Since the classifier is based on a statistical method, combination of the scores is straightforward both in theory and in practice. Experimental results showed that the combined method is more effective than the conventional VQ (vector quantization) HMM and that utilizing the score of the classifier on the local features is significant.> Tatsuya Kawahara, Shuji Doshita |
ICASSP | 1 |
| 1991 | Unsupervised speaker normalization by speaker Markov model converter for speaker-independent speech recognition
Pascale Fung, Tatsuya Kawahara, Shuji Doshita |
EUROSPEECH | 2 |
| 1990 | Phoneme recognition by combining Bayesian linear discriminations of selected pairs of classes
Tatsuya Kawahara, Toru Ogawa, Shigeyoshi Kitazawa, Shuji Doshita |
ICSLP | 1 |