Takahiro Shinozaki

dblp:06/6505 · DBLP profile ↗
← Back
64ranked-venue papers
19as first author
21since 2021 · last 2025
0000-0001-8114-8450ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 53 · 18 first-author · 13 since 2021Artificial intelligence and machine learning · 42 · 12 first-author · 14 since 2021
YearPublicationVenuePosition
2025 Deep Generic Representations for Domain-Generalized Anomalous Sound Detection
abstract
Developing a reliable anomalous sound detection (ASD) system requires robustness to noise, adaptation to domain shifts, and effective performance with limited training data. Current leading methods rely on extensive labeled data for each target machine type to train feature extractors using Outlier-Exposure (OE) techniques, yet their performance on the target domain remains sub-optimal. In this paper, we present Gen-Rep, which utilizes generic feature representations from a robust, large-scale pre-trained feature extractor combined with kNN for domain-generalized ASD, without the need for fine-tuning. GenRep incorporates MemMixup, a simple approach for augmenting the target memory bank using nearest source samples, paired with a domain normalization technique to address the imbalance between source and target domains. GenRep outperforms the best OE-based approach without a need for labeled data with an Official Score of 73.79% on the DCASE2023T2 Eval set and demonstrates robustness under limited data scenarios. The code is available open-source1.
Phurich Saengthong, Takahiro Shinozaki
ICASSP2
2024 Self-Supervised Speaker Verification with Adaptive Threshold and Hierarchical Training
abstract
In self-supervised speaker verification, the quality of generated pseudo labels becomes a bottleneck for the performance. This work introduces a dynamic threshold within the iterative DIstillation with NO labels (DINO) framework. We employ a Gaussian Mixture Model (GMM) to model the loss distribution of the training data. The GMM has two components: one represents samples with reliable labels, and the other with un-reliable ones. These components help us determine a thresh-old for retaining samples with reliable labels. Furthermore, to take advantage of the different sensitivity of network layers to label noise, we further introduce hierarchical training to reduce the negative impact of unreliable labels. Compared to the baseline with a fixed threshold, our two strategies result in an 8.9% relative improvement on the Vox-O trial of the Voxceleb1 evaluation dataset.
Zehua Zhou, Haoyuan Yang, Takahiro Shinozaki
ICASSP3
2024 Self-Supervised Syllable Discovery Based on Speaker-Disentangled Hubert
abstract
Self-supervised speech representation learning has become essential for extracting meaningful features from untranscribed audio. Recent advances highlight the potential of deriving discrete symbols from the features correlated with linguistic units, which enables text-less training across diverse tasks. In particular, sentence-level Self-Distillation of the pretrained HuBERT (SD-HuBERT) induces syllabic structures within latent speech frame representations extracted from an intermediate Transformer layer. In SD-HuBERT, sentence-level representation is accumulated from speech frame features through self-attention layers using a special CLS token. However, we observe that the information aggregated in the CLS token correlates more with speaker identity than with linguistic content. To address this, we propose a speech-only self-supervised fine-tuning approach that separates syllabic units from speaker information. Our method introduces speaker perturbation as data augmentation and adopts a frame-level training objective to prevent the CLS token from aggregating paralinguistic information. Experimental results show that our approach surpasses the current state-of-the-art method in most syllable segmentation and syllabic unit quality metrics on Librispeech, underscoring its effectiveness in promoting syllabic organization within speech-only models1.1Codes and models:https://github.com/ryota-komatsu/speaker_disentangled_hubert
Ryota Komatsu, Takahiro Shinozaki
SLT2
2023 Continuous Action Space-Based Spoken Language Acquisition Agent Using Residual Sentence Embedding and Transformer Decoder
abstract
Studies on spoken language acquisition agents aim to understand the mechanism of human language learning and to realize it on computers. Existing open vocabulary agents first perform unsupervised word learning from speech signals to construct a word dictionary as a discrete action space and then conduct reinforcement learning to understand the use of the words in the dictionary through interaction with dialogue partners. A limitation is that they have difficulty pronouncing multi-word utterances. This study proposes an agent that generates multi-word waveform utterances using a continuous action space. The conventional agent uses a vision-focusing mechanism to accelerate dialogue-based learning by guiding the agent’s attention to those concepts in its eyesight. In contrast, the proposed agent replaces it with residual sentence embedding combined with vision features used as the action space. The agent consists of speech and image input front-ends, a transformer language model of pseudo-action space. Experimental results show that the agent learns multi-word utterances assisted by unsupervised learning algorithms using unlabeled speech and image data sets.
Ryota Komatsu, Yusuke Kimura, Takuma Okamoto, Takahiro Shinozaki
ICASSP4
2023 FreeMatch: Self-adaptive Thresholding for Semi-supervised Learning
Yidong Wang 0003, Hao Chen 0102, Qiang Heng, Wenxin Hou, Zhen Wu 0002, Jindong Wang 0001, Marios Savvides, Takahiro Shinozaki, Bhiksha Raj, Bernt Schiele, Xing Xie 0001
ICLR9
2023 Memory Network-Based End-To-End Neural ES-KMeans for Improved Word Segmentation
Yu Iwamoto, Takahiro Shinozaki
INTERSPEECH2
2022 Margin Calibration for Long-Tailed Visual Recognition
Yidong Wang 0003, Wenxin Hou, Zhen Wu 0002, Jindong Wang 0001, Takahiro Shinozaki
ACML6
2022 Exploiting Unlabeled Data for Target-Oriented Opinion Words Extraction
abstract
Target-oriented Opinion Words Extraction (TOWE) is a fine-grained sentiment analysis task that aims to extract the corresponding opinion words of a given opinion target from the sentence. Recently, deep learning approaches have made remarkable progress on this task. Nevertheless, the TOWE task still suffers from the scarcity of training data due to the expensive data annotation process. Limited labeled data increase the risk of distribution shift between test data and training data. In this paper, we propose exploiting massive unlabeled data to reduce the risk by increasing the exposure of the model to varying distribution shifts. Specifically, we propose a novel Multi-Grained Consistency Regularization (MGCR) method to make use of unlabeled data and design two filters specifically for TOWE to filter noisy data at different granularity. Extensive experimental results on four TOWE benchmark datasets indicate the superiority of MGCR compared with current state-of-the-art methods. The in-depth analysis also demonstrates the effectiveness of the different-granularity filters.
Yidong Wang 0003, Hao Wu 0059, Ao Liu 0008, Wenxin Hou, Zhen Wu 0002, Jindong Wang 0001, Takahiro Shinozaki, Manabu Okumura, Yue Zhang 0004
COLING7
2022 Hybrid RNN-T/Attention-Based Streaming ASR with Triggered Chunkwise Attention and Dual Internal Language Model Integration
abstract
In this paper we propose improvements to our recently proposed hybrid RNN-T/Attention architecture that includes a shared encoder followed by recurrent neural network-transducer (RNN-T) and triggered attention-based decoders (TAD). The use of triggered attention enables the attention-based decoder (AD) to operate in a streaming manner. When a trigger point is detected by RNN-T, TAD uses the context from the start-of-speech up to that trigger point to compute the attention weights. Consequently, the computation costs and the memory consumptions are quadratically increased with the duration of the utterances because all input features must be stored and used to re-compute the attention weights. In this paper, we use a short context from a few frames prior to each trigger point for attention weight computation resulting in reduced computation and memory costs. We call the proposed framework triggered chunkwise AD (TCAD). We also investigate the effectiveness of internal language model (ILM) estimation approach using both ILMs of RNN-T and TCAD heads for improving RNN-T performance. We confirm in experiments with public and private datasets covering various scenarios that TCAD achieves superior recognition performance while reducing computation costs compared to TAD.
Takafumi Moriya, Takanori Ashihara, Atsushi Ando, Hiroshi Sato 0002, Tomohiro Tanaka, Kohei Matsuura, Ryo Masumura, Marc Delcroix, Takahiro Shinozaki
ICASSP9
2022 Streaming Target-Speaker ASR with Neural Transducer
Takafumi Moriya, Hiroshi Sato 0002, Tsubasa Ochiai, Marc Delcroix, Takahiro Shinozaki
INTERSPEECH5
2022 Self-Supervised Learning with Multi-Target Contrastive Coding for Non-Native Acoustic Modeling of Mispronunciation Verification
Longfei Yang, Jinsong Zhang 0001, Takahiro Shinozaki
INTERSPEECH3
2022 Augmented Adversarial Self-Supervised Learning for Early-Stage Alzheimer's Speech Detection
Longfei Yang, Wenqing Wei, Sheng Li 0010, Jiyi Li, Takahiro Shinozaki
INTERSPEECH5
2022 Censer: Curriculum Semi-supervised Learning for Speech Recognition Based on Self-supervised Pre-training
abstract
Recent studies have shown that the benefits provided by selfsupervised pre-training and self-training (pseudo-labeling) are complementary.Semi-supervised fine-tuning strategies under the pre-training framework, however, remain insufficiently studied.Besides, modern semi-supervised speech recognition algorithms either treat unlabeled data indiscriminately or filter out noisy samples with a confidence threshold.The dissimilarities among different unlabeled data are often ignored.In this paper, we propose Censer, a semi-supervised speech recognition algorithm based on self-supervised pre-training to maximize the utilization of unlabeled data.The pre-training stage of Censer adopts wav2vec2.0and the fine-tuning stage employs an improved semisupervised learning algorithm from slimIPL, which leverages unlabeled data progressively according to their pseudo labels' qualities.We also incorporate a temporal pseudo label pool and an exponential moving average to control the pseudo labels' update frequency and to avoid model divergence.Experimental results on Libri-Light and LibriSpeech datasets manifest our proposed method achieves better performance compared to existing approaches while being more unified.
Songjun Cao, Takahiro Shinozaki
INTERSPEECH6
2022 USB: A Unified Semi-supervised Learning Benchmark for Classification
abstract
Semi-supervised learning (SSL) improves model generalization by leveraging massive unlabeled data to augment limited labeled samples. However, currently, popular SSL evaluation protocols are often constrained to computer vision (CV) tasks. In addition, previous work typically trains deep neural networks from scratch, which is time-consuming and environmentally unfriendly. To address the above issues, we construct a Unified SSL Benchmark (USB) for classification by selecting 15 diverse, challenging, and comprehensive tasks from CV, natural language processing (NLP), and audio processing (Audio), on which we systematically evaluate the dominant SSL methods, and also open-source a modular and extensible codebase for fair evaluation of these SSL methods. We further provide the pre-trained versions of the state-of-the-art neural models for CV tasks to make the cost affordable for further tuning. USB enables the evaluation of a single SSL algorithm on more tasks from multiple domains but with less cost. Specifically, on a single NVIDIA V100, only 39 GPU days are required to evaluate FixMatch on 15 tasks in USB while 335 GPU days (279 GPU days on 4 CV datasets except for ImageNet) are needed on 5 CV tasks with TorchSSL.
Yidong Wang 0003, Hao Chen 0102, Wang Sun, Ran Tao 0013, Wenxin Hou, Linyi Yang, Zhi Zhou 0007, Lan-Zhe Guo, Heli Qi, Zhen Wu 0002, Yufeng Li 0008, Satoshi Nakamura 0001, Wei Ye 0004, Marios Savvides, Bhiksha Raj, Takahiro Shinozaki, Bernt Schiele, Jindong Wang 0001, Xing Xie 0001, Yue Zhang 0004
NeurIPS18
2022 Multi-Domain Dialogue State Tracking with Top-K Slot Self Attention
abstract
As an important component of task-oriented dialogue systems, dialogue state tracking is designed to track the dialogue state through the conversations between users and systems.Multi-domain dialogue state tracking is a challenging task, in which the correlation among different domains and slots needs to consider.Recently, slot self-attention is proposed to provide a data-driven manner to handle it.However, a full-support slot self-attention may involve redundant information interchange.In this paper, we propose a top-k attention-based slot self-attention for multi-domain dialogue state tracking.In the slot self-attention layers, we force each slot to involve information from the other k prominent slots and mask the rest out.The experimental results on two mainstream multi-domain task-oriented dialogue datasets, MultiWOZ 2.0 and MultiWOZ 2.4, present that our proposed approach is effective to improve the performance of multi-domain dialogue state tracking.We also find that the best result is obtained when each slot interchanges information with only a few slots.
Longfei Yang, Jiyi Li, Sheng Li 0010, Takahiro Shinozaki
SIGDIAL4
2022 Exploiting Adapters for Cross-Lingual Low-Resource Speech Recognition
abstract
Cross-lingual speech adaptation aims to solve the problem of leveraging multiple rich-resource languages to build models for a low-resource target language. Since the low-resource language has limited training data, speech recognition models can easily overfit. Adapter is a versatile module that can be plugged into Transformer for parameter-efficient learning. In this paper, we propose to use adapters for parameter-efficient cross-lingual speech adaptation. Based on our previous MetaAdapter that implicitly leverages adapters, we propose a novel algorithm called SimAdapter for explicitly learning knowledge from adapters. Our algorithms can be easily integrated into the Transformer structure. MetaAdapter leverages meta-learning to transfer the general knowledge from training data to the test language. SimAdapter aims to learn the similarities between the source and target languages during fine-tuning using the adapters. We conduct extensive experiments on five-low-resource languages in the Common Voice dataset. Results demonstrate that MetaAdapter and SimAdapter can reduce WER by 2.98% and 2.55% with only 2.5% and 15.5% of trainable parameters compared to the strong full-model fine-tuning baseline. Moreover, we show that these two novel algorithms can be integrated for better performance with up to 3.55% relative WER reduction.
Wenxin Hou, Han Zhu 0004, Yidong Wang 0003, Jindong Wang 0001, Tao Qin 0001, Renjun Xu, Takahiro Shinozaki
IEEE ACM Trans. Audio Speech Lang. Process.7
2021 Meta-Adapter: Efficient Cross-Lingual Adaptation With Meta-Learning
abstract
Transfer learning from a multilingual model has shown favorable results on low-resource automatic speech recognition (ASR). However, full-model fine-tuning generates a separate model for every target language and is not suitable for deploying and maintaining in production. The key challenge lies in how to efficiently extend the pre-trained model with fewer parameters. In this paper, we propose to combine the adapter module with meta-learning algorithms to achieve high recognition performance under low-resource settings and improve the parameter-efficiency of the model. Extensive experiments show that our methods can achieve comparable or even superior recognition rates than the state-of-the-art baselines on low-resource languages, especially under very-low-resource conditions, with a significantly smaller model profile.
Wenxin Hou, Yidong Wang 0003, Shengzhou Gao, Takahiro Shinozaki
ICASSP4
2021 Cross-Domain Speech Recognition with Unsupervised Character-Level Distribution Matching
abstract
End-to-end automatic speech recognition (ASR) can achieve promising performance with large-scale training data. However, it is known that domain mismatch between training and testing data often leads to a degradation of recognition accuracy. In this work, we focus on the unsupervised domain adaptation for ASR and propose CMatch, a Character-level distribution matching method to perform fine-grained adaptation between each character in two domains. First, to obtain labels for the features belonging to each character, we achieve frame-level label assignment using the Connectionist Temporal Classification (CTC) pseudo labels. Then, we match the character-level distributions using Maximum Mean Discrepancy. We train our algorithm using the self-training technique. Experiments on the Libri-Adapt dataset show that our proposed approach achieves 14.39% and 16.50% relative Word Error Rate (WER) reduction on both cross-device and cross-environment ASR. We also comprehensively analyze the different strategies for frame-level label assignment and Transformer adaptations.
Wenxin Hou, Jindong Wang 0001, Xu Tan 0003, Tao Qin 0001, Takahiro Shinozaki
Interspeech5
2021 FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo Labeling
abstract
The recently proposed FixMatch achieved state-of-the-art results on most semi-supervised learning (SSL) benchmarks. However, like other modern SSL algorithms, FixMatch uses a pre-defined constant threshold for all classes to select unlabeled data that contribute to the training, thus failing to consider different learning status and learning difficulties of different classes. To address this issue, we propose Curriculum Pseudo Labeling (CPL), a curriculum learning approach to leverage unlabeled data according to the model's learning status. The core of CPL is to flexibly adjust thresholds for different classes at each time step to let pass informative unlabeled data and their pseudo labels. CPL does not introduce additional parameters or computations (forward or backward propagation). We apply CPL to FixMatch and call our improved algorithm FlexMatch. FlexMatch achieves state-of-the-art performance on a variety of SSL benchmarks, with especially strong performances when the labeled data are extremely limited or when the task is challenging. For example, FlexMatch achieves 13.96% and 18.96% error rate reduction over FixMatch on CIFAR-100 and STL-10 datasets respectively, when there are only 4 labels per class. CPL also significantly boosts the convergence speed, e.g., FlexMatch can use only 1/5 training time of FixMatch to achieve even better performance. Furthermore, we show that CPL can be easily adapted to other SSL algorithms and remarkably improve their performances. We open-source our code at https://github.com/TorchSSL/TorchSSL.
Yidong Wang 0003, Wenxin Hou, Hao Wu 0059, Jindong Wang 0001, Manabu Okumura, Takahiro Shinozaki
NeurIPS7
2021 Unsupervised Acoustic-to-Articulatory Inversion Neural Network Learning Based on Deterministic Policy Gradient
abstract
This paper presents an unsupervised learning method of deep neural networks that perform acoustic-to-articulatory inversion for arbitrary utterances. Conventional unsupervised acoustic-to-articulatory inversion methods are based on the analysis-by-synthesis approach and non-linear optimization algorithms. One limitation is that they require time-consuming iterative optimizations to obtain articulatory parameters for a given target speech segment. Neural networks, after learning their relationship, can obtain these articulatory parameters without an iterative optimization. However, conventional methods need supervised learning and paired acoustic and articulatory samples. We propose a hybrid auto-encoder based unsupervised learning framework for the acoustic-to-articulatory inversion neural networks that can capture context information. The essential point of the framework is making the training effective. We investigate several reinforcement learning algorithms and show the usefulness of the deterministic policy gradient. Experimental results demonstrate that the proposed method can infer articulatory parameters not only for training set segments but also for unseen utterances. Averaged reconstruction errors achieved for open test samples are similar to or even lower than the conventional method that directly optimizes the articulatory parameters in a closed condition.
Hayato Shibata, Mingxin Zhang 0008, Takahiro Shinozaki
SLT3
2021 Non-native acoustic modeling for mispronunciation verification based on language adversarial representation learning
Longfei Yang, Kaiqi Fu, Jinsong Zhang 0001, Takahiro Shinozaki
Neural Networks4
2020 Dual Inheritance Evolution Strategy for Deep Neural Network Optimization
abstract
Deep neural networks (DNNs) need intensive tuning of their configurations such as network structures and learning conditions. The tuning is a type of black-box optimization problem where evolutionary algorithms are applicable. A distinctive property in evolutionary optimization of DNN configurations is that there is a double structure in the optimization; the evolutionary algorithm optimizes a chromosome representing the DNN configuration while an individual DNN with the configuration learns from training data typically by back-propagation. With an aim to obtain better-optimized DNNs by evolutionary algorithms, we propose a dual inheritance evolution strategy based on an analogy to human brain evolution where gene and culture co-evolves. The proposed method is an extension of a conventional evolution strategy by introducing an additional pass to directly propagate culture or knowledge from ancestor DNNs to descendant DNNs by integrating teacher-student learning. We apply the proposed method to the automatic tuning of an end-to-end neural network-based speech recognition system. Experimental results show that the proposed method produces a smaller model with higher recognition performance than a baseline optimization based on the Covariance Matrix Adaptation Evolution Strategy (CMA-ES).
Kent Hino, Yusuke Kimura, Takahiro Shinozaki
CEC4
2020 Spoken Language Acquisition Based on Reinforcement Learning and Word Unit Segmentation
abstract
The process of spoken-language acquisition has been one of the topics of greatest interest to linguists for decades. By uti-lizing modern machine learning techniques, we simulated this process on computers, which helps to understand it and develop new possibilities of applying this concept on intelligent robots, among other things. This paper proposes a new framework for simulating spoken-language acquisition by combining reinforcement learning and unsupervised learning methods. Our experiments also show that a spoken language can be acquired considerably faster by identifying potential word segments from collected ambient sounds in an unsupervised manner.
Shengzhou Gao, Wenxin Hou, Tomohiro Tanaka, Takahiro Shinozaki
ICASSP4
2020 Unsupervised Sound Source Localization From Audio-Image Pairs Using Input Gradient Map
abstract
Humans easily and routinely identify an image region that corresponds to an observed sound in their daily lives. The task is formulated as an unsupervised sound source localization without using tagged data. Recently, several methods have been proposed that utilize the activation of hidden or output layers of neural networks, such as an attention layer or feature maps in a convolutional neural network (CNN). We propose another strategy that obtains a localization map at the input side, applying the widely used input gradient method. It is computationally efficient and can be easily applied to any existing techniques because it is free from the network structure. Taking advantage of it, we propose a combination method with existing methods for higher sound localization performance. Experiments are performed using the Flickr-SoundNet data set. When a pre-trained image front-end was used, the proposed method gives better results than the attention-based method. For a completely unsupervised condition, the gradient method provides comparable performance as the conventional methods; the best results are obtained by this combination method.
Tomohiro Tanaka, Takahiro Shinozaki
ICPR2
2020 Large-Scale End-to-End Multilingual Speech Recognition and Language Identification with Multi-Task Learning
Wenxin Hou, Bairong Zhuang, Longfei Yang, Jiatong Shi, Takahiro Shinozaki
INTERSPEECH6
2020 Pronunciation Erroneous Tendency Detection with Language Adversarial Represent Learning
Longfei Yang, Kaiqi Fu, Jinsong Zhang 0001, Takahiro Shinozaki
INTERSPEECH4
2020 Sound-Image Grounding Based Focusing Mechanism for Efficient Automatic Spoken Language Acquisition
Mingxin Zhang 0008, Tomohiro Tanaka, Wenxin Hou, Shengzhou Gao, Takahiro Shinozaki
INTERSPEECH5
2020 Time-Domain Target-Speaker Speech Separation with Waveform-Based Speaker Embedding
Jianshu Zhao, Shengzhou Gao, Takahiro Shinozaki
INTERSPEECH3
2019 Efficient Free Keyword Detection Based on CNN and End-to-End Continuous DP-Matching
abstract
For continuous keyword detection, the advantage of dynamic programming (DP) matching is that it can detect any keyword without re-training the system. In previous research, higher detection accuracy was reported using 2D-RNN based DP matching than using conventional DP and embedding methods. However, 2D-RNN based DP matching has a high computational cost. In order to address this problem, we combine a convolutional neural network (CNN) and 2D-RNN based DP matching into a unified framework which, based on the kernel size and the number of CNN layers, has a polynomial order effect on reducing the computational cost. Experimental results, using Google Speech Commands Dataset and the CHiME-3 challenge's noise data, demonstrate that our proposed model improves open keyword detection performance, compared to the embedding-based baseline system, while it is nine times faster than previous 2D-RNN DP matching.
Tomohiro Tanaka, Takahiro Shinozaki
ASRU2
2019 Effective and Stable Neuron Model Optimization Based on Aggregated CMA-ES
abstract
Computer simulations have facilitated our understanding of the dynamic behavior of the brain and the effect of the medical treatment such as deep brain stimulation. For improving the simulation model, it is essential to develop a method for optimizing parameters of a neuron model from available experimental data. In this paper, we apply Covariance Matrix Adaptation Evolutionary Strategy (CMA-ES) to the parameter optimization problem, and compare it with widely used conventional approaches including genetic algorithm (GA) and the Nelder-Mead method. A problem we have observed with CMA-ES is that the performance highly depends on the initial condition. To overcome the problem, we extend CMA-ES by making an aggregation of evolution. We analyze a public dataset recorded from a rat neocortical neuron, which shows that the proposed approach achieves higher performance than the conventional methods.
Takahiro Shinozaki, Ryota Kobayashi
ICASSP2
2019 Evolution-Strategy-Based Automation of System Development for High-Performance Speech Recognition
abstract
The state-of-the-art large vocabulary speech recognition systems consist of several components including hidden Markov model and deep neural network. To realize the highest recognition performance, numerous meta-parameters specifying the designs and training setups of these components must be optimized. A prominent obstacle in system development is the laborious effort required by human experts in tuning these meta-parameters. To automate the process, we propose to tune the meta-parameters of a whole large vocabulary speech recognition system using the evolution strategy with a multi-objective Pareto optimization. As the result of the evolution, the system is optimized for both low word error rate and compact model size. Since the approach requires repeated training and evaluation of the recognition systems that require large computation, we make use of parallel computation on cloud computers. Experimental results show the effectiveness of the proposed approach by discovering appropriate configuration for large vocabulary speech recognition systems automatically.
Takafumi Moriya, Tomohiro Tanaka, Takahiro Shinozaki, Shinji Watanabe 0001, Kevin Duh
IEEE ACM Trans. Audio Speech Lang. Process.3
2018 Reinforcement Learning of Speech Recognition System Based on Policy Gradient and Hypothesis Selection
abstract
Automatic speech recognition (ASR) systems have achieved high recognition performance for several tasks. However, the performance of such systems is dependent on the tremendously costly development work of preparing vast amounts of task-matched transcribed speech data for supervised training. The key problem here is the cost of transcribing speech data. The cost is repeatedly required to support new languages and new tasks. Assuming broad network services for transcribing speech data for many users, a system would become more self-sufficient and more useful if it possessed the ability to learn from very light feedback from the users without annoying them. In this paper, we propose a general reinforcement learning framework for ASR systems based on the policy gradient method. As a particular instance of the framework, we also propose a hypothesis selection-based reinforcement learning method. The proposed framework provides a new view for several existing training and adaptation methods. The experimental results show that the proposed method improves the recognition performance compared to unsupervised adaptation.
Taku Kato, Takahiro Shinozaki
ICASSP2
2017 Composite embedding systems for ZeroSpeech2017 Track1
abstract
This paper investigates novel composite embedding systems for language-independent high-performance feature extraction using triphone-based DNN-HMM and character-based end-to-end speech recognition systems. The DNN-HMM is trained with phoneme transcripts based on a large-scale Japanese ASR recipe included in the Kaldi toolkit from the Corpus of Spontaneous Japanese (CSJ) with some modifications. The end-to-end ASR system is based on a hybrid architecture consisting of an attention-based encoder-decoder and connectionist temporal classification. This model is trained with multi-language speech data using character transcripts in a pure end-to-end fashion without requiring phonemic representation. Posterior features, PCA-transformed features, and bottleneck features are extracted from the two systems; then, various combinations of features are explored. Additionally, a bypassed autoencoder (bypassed AE) is proposed to normalize speaker characteristics in an unsupervised manner. An evaluation using the ABX test showed that the DNN-HMM-based CSJ bottleneck features resulted in a good performance regardless of the input language. The pre-activation vectors extracted from the multilingual end-to-end system with PCA provided a somewhat better performance than did the CSJ bottleneck features. The bypassed AE yielded an improved performance over a baseline AE. The lowest error rates were obtained by composite features that concatenated the end-to-end features with the CSJ bottleneck features.
Hayato Shibata, Taku Kato, Takahiro Shinozaki, Shinji Watanabe 0001
ASRU3
2017 Semi-Supervised Learning of a Pronunciation Dictionary from Disjoint Phonemic Transcripts and Text
Takahiro Shinozaki, Shinji Watanabe 0001, Daichi Mochihashi, Graham Neubig
INTERSPEECH1
2016 Automated structure discovery and parameter tuning of neural network language model based on evolution strategy
abstract
Long short-term memory (LSTM) recurrent neural network based language models are known to improve speech recognition performance. However, significant effort is required to optimize network structures and training configurations. In this study, we automate the development process using evolutionary algorithms. In particular, we apply the covariance matrix adaptation-evolution strategy (CMA-ES), which has demonstrated robustness in other black box hyper-parameter optimization problems. By flexibly allowing optimization of various meta-parameters including layer wise unit types, our method automatically finds a configuration that gives improved recognition performance. Further, by using a Pareto based multi-objective CMA-ES, both WER and computational time were reduced jointly: after 10 generations, relative WER and computational time reductions for decoding were 4.1% and 22.7% respectively, compared to an initial baseline system whose WER was 8.7%.
Tomohiro Tanaka, Takafumi Moriya, Takahiro Shinozaki, Shinji Watanabe 0001, Takaaki Hori, Kevin Duh
SLT3
2015 Automation of system building for state-of-the-art large vocabulary speech recognition using evolution strategy
abstract
When building a state-of-the-art speech recognition system, the laborious effort required by human experts in tuning numerous parameters remains a prominent obstacle. The goal of this paper is to automate the process. We propose to tune DNN-HMM based large vocabulary speech recognition systems using the covariance matrix adaptation evolution strategy (CMA-ES) with a multi-objective Pareto optimization. This optimizes systems to achieve both high-accuracy and compact model size. An additional advantage of our approach is that it is efficiently parallelizable and easily adapted to cloud computing services. We performed experiments on the Corpus of Spontaneous Japanese (CSJ) using the TSUBAME 2.5 supercomputer. Compared with a strong manually tuned configuration borrowed from a similar system, our approach automatically discovered systems with lower WER by 0.48%, and systems with 59% smaller model size while keeping WER constant. The optimized training script is released in the Kaldi speech recognition toolkit as the first publicly available recipe for Japanese large vocabulary speech recognition.
Takafumi Moriya, Tomohiro Tanaka, Takahiro Shinozaki, Shinji Watanabe 0001, Kevin Duh
ASRU3
2015 Structure discovery of deep neural network based on evolutionary algorithms
abstract
Deep neural networks (DNNs) are constructed by considering highly complicated configurations including network structure and several tuning parameters (number of hidden states and learning rate in each layer), which greatly affect the performance of speech processing applications. To reach optimal performance in such systems, deep understanding and expertise in DNNs is necessary, which limits the development of DNN systems to skilled experts. To overcome the problem, this paper proposes an efficient optimization strategy for DNN structure and parameters using evolutionary algorithms. The proposed approach parametrizes the DNN structure by a directed acyclic graph, and the DNN structure is represented by a simple binary vector. Genetic algorithm and covariance matrix adaptation evolution strategy efficiently optimize the performance jointly with respect to the above binary vector and the other tuning parameters. Experiments on phoneme recognition and spoken digit detection tasks show the effectiveness of the proposed approach by discovering the appropriate DNN structure automatically.
Takahiro Shinozaki, Shinji Watanabe 0001
ICASSP1
2014 Accent type and phrase boundary estimation using acoustic and language models for automatic prosodic labeling
abstract
This paper proposes an automatic prosodic labeling technique for constructing speech database used for speech synthesis.In the corpus-based Japanese speech synthesis, it is essential to use annotated speech data with prosodic information such as phrase boundaries and accent types.However, manual annotation is generally time-consuming and expensive.To overcome this problem, we propose an estimation technique of accent types and phrase boundaries from speech waveform and its transcribed text using both language and acoustic models.We use conditional random field (CRF) for the language model, and HMM for the acoustic model which has shown to be effective in prosody modeling in speech synthesis.By introducing HMM, continuously changing features of F0 contours are modeled well and this results in higher estimation accuracy than conventional techniques that use simple polygonal line approximation of F0 contours.
Tomoki Koriyama, Hiroshi Suzuki, Takashi Nose, Takahiro Shinozaki, Takao Kobayashi
INTERSPEECH4
2013 Reverberant speech recognition based on denoising autoencoder
abstract
Denoising autoencoder is applied to reverberant speech recognition as a noise robust front-end to reconstruct clean speech spectrum from noisy input. In order to capture context effects of speech sounds, a window of multiple short-windowed spectral frames are concatenated to form a single input vector. Additionally, a combination of short and long-term spectra is investigated to properly handle long impulse response of reverberation while keeping necessary time resolution for speech recognition. Experiments are performed using the CENSREC-4dataset that is designed as an evaluation framework for distant-talking speech recognition. Experimental results show that the proposed denoising autoencoder based front-end using the shortwindowed spectra gives better results than conventional methods. By combining the long-term spectra, further improvement is obtained. The recognition accuracy by the proposed method using the short and long-term spectra is 97.0% for the open condition test set of the dataset, whereas it is 87.8% when a multicondition training based baseline is used. As a supplemental experiment, large vocabulary speech recognition is also performed and the effectiveness of the proposed method has been confirmed. Index Terms: Denoising autoencoder, reverberant speech recognition, restricted Boltzmann machine, distant-talking speech recognition, CENSREC-4
Takaaki Ishii, Hiroki Komiyama, Takahiro Shinozaki, Yasuo Horiuchi, Shingo Kuroiwa
INTERSPEECH3
2012 Unsupervised CV language model adaptation based on direct likelihood maximization sentence selection
abstract
Direct likelihood maximization selection (DLMS) selects a subset of language model training data so that likelihood of in-domain development data is maximized. By using recognition hypothesis instead of the in-domain development data, it can be used for unsupervised adaptation. We apply DLMS to iterative unsupervised adaptation for presentation speech recognition. A problem of the iterative unsupervised adaptation is that adapted models are estimated including recognition errors and it limits the adaptation performance. To solve the problem, we introduce the framework of unsupervised cross-validation (CV) adaptation that has originally been proposed for acoustic model adaptation. Large vocabulary speech recognition experiments show that the CV approach is effective for DLMS based adaptation reducing 19.3% of error rate by an initial model to 18.0%.
Takahiro Shinozaki, Yasuo Horiuchi, Shingo Kuroiwa
ICASSP1
2012 HMM Based Continuous EOG Recognition for Eye-input Speech Interface
abstract
To provide an efficient means of communication for those who cannot move muscles of the whole body except eyes due to amyotrophic lateral sclerosis (ALS), we are developing a speech synthesis interface that is based on electrooculogram (EOG) input.EOG is an electrical signal that is observed through electrodes attached on the skin around eyes and reflects eye position.A key component of the system is a continuous recognizer for the EOG signal.In this paper, we propose and investigate a hidden Markov model (HMM) based EOG recognizer applying continuous speech recognition techniques.In the experiments, we evaluate the recognition system both in user dependent and independent conditions.It is shown that 96.1% of recognition accuracy is obtained for five classes of eye actions by a user dependent system using six channels.While it is difficult to obtain good performance by a user independent system, it is shown that maximum likelihood linear regression (MLLR) adaptation helps for EOG recognition.
Fuming Fang, Takahiro Shinozaki, Yasuo Horiuchi, Shingo Kuroiwa, Sadaoki Furui, Toshimitsu Musha
INTERSPEECH2
2011 Sentence Selection by Direct Likelihood Maximization for Language Model Adaptation
abstract
A general framework of language model task adaptation is to select documents in a large training set based on a language model estimated on a development data.However, this strategy has a deficiency that the selected documents are biased to the most frequent patterns in the development data.To address this problem, a new task adaptation method is proposed that selects documents in the training set so as to directly reduce the perplexity on the development set.Moreover, a weighting method to modify the perplexity objective function is proposed to improve the generalization to unseen data.The proposed adaptation methods are evaluated by large vocabulary speech recognition experiments.It is shown that the proposed adaptation with the weighting term produces a compact-size model that gives consistently lower word error rates for different tasks.
Takahiro Shinozaki, Yu Kubota, Sadaoki Furui, Eiji Utsunomiya, Yasutaka Shindoh
INTERSPEECH1
2010 Investigations on ensemble based unsupervised adaptation methods
abstract
We have previously proposed unsupervised cross-validation (CV) adaptation that introduces CV into an iterative unsupervised batch mode adaptation framework to suppress the influence of errors in an internally generated recognition hypothesis and have shown that it improves recognition performance. However, a limitation was that the experiments were performed using only a clean speech recognition task with a ML trained initial acoustic model. Another limitation was that only the CV method was investigated while there was a possibility of using other ensemble methods. In this study, we evaluate the CV method using a discriminatively trained baseline and a noisy speech recognition task. As an alternative to CV adaptation, unsupervised aggregated (Ag) adaptation is proposed and investigated that introduces a bagging like idea instead of CV. Experimental results show that CV and Ag adaptations consistently give larger improvements than the conventional batch adaptation but the former is more advantageous in terms of computational cost.
Yu Kubota, Takahiro Shinozaki, Sadaoki Furui
ICASSP2
2009 Unsupervisec cross-validation adaptation algorithms for improved adaptation performance
abstract
An unsupervised cross-validation adaptation algorithm and its variation are proposed that introduce the idea of cross-validation in the unsupervised batch-mode adaptation framework to improve the adaptation performance. The first algorithm is constructed on a general adaptation technique such as MLLR and can be used in combination with any adaptation method. The second algorithm is a modified version of the first algorithm and works with lower computational cost by assuming MLLR. These algorithms are extensions of our previously proposed CV training methods and are useful to suppress the negative effect of the conventional unsupervised batch-mode adaptation process that reinforces the errors included in automatic transcriptions. The proposed algorithms were evaluated in domain adaptation, speaker adaptation, and in their combination for large vocabulary spontaneous speech recognition. When the domain and speaker adaptations were combined using a read speech initial model, the relative word error rate reduction by the proposed method was 29% whereas the reduction by the conventional approach was 23%.
Takahiro Shinozaki, Yu Kubota, Sadaoki Furui
ICASSP1
2009 Target speech GMM-based spectral compensation for noise robust speech recognition
abstract
Abstract To improve speech recognition performance in adverse condi-tions, a noise compensation method is proposed that applies atransformation in the spectral domain whose parameters are op-timized based on likelihood of speech GMM modeled on thefeature domain. The idea is that additive and convolutionalnoises have mathematically simple expression in the spectraldomain while speech characteristics are better modeled in thefeature domain such as MFCC. The proposed method worksas a feature extraction front-end that is independent from de-coding engine, and has ability to compensate for non-stationaryadditive and convolutional noises with a short time delay. Itincludes spectral subtraction as a special case when no param-eter optimization is performed. Experiments were performedusing the AURORA-2J database. It has been shown that signif-icantly higher recognition performance is obtained by the pro-posed method than spectral subtraction. Index Terms : noisy speech recognition, spectrum, Gaussianmixture model
Takahiro Shinozaki, Sadaoki Furui
INTERSPEECH1
2008 GMM and HMM training by aggregated EM algorithm with increased ensemble sizes for robust parameter estimation
abstract
In order to compensate for the weaknesses of the expectation maximization (EM) algorithm to over-training and to improve model performance for new data, we have recently proposed aggregated EM (Ag-EM) algorithm that introduces bagging-like approach in the framework of the EM algorithm and have shown that it gives similar improvements as cross-validation EM (CV EM) over conventional EM. However, a limitation with the experiments was that the number of multiple models used in the aggregation operation or the ensemble size was fixed to a small value. Here, we investigate the relationship between the ensemble size and the performance as well as giving a theoretical discussion with the order of the computational cost. The algorithm is first analyzed using simulated data and then applied to large vocabulary speech recognition on oral presentations. Both of these experiments show that Ag-EM outperforms CV-EM by using larger ensemble sizes.
Takahiro Shinozaki, Tatsuya Kawahara
ICASSP1
2008 Aggregated cross-validation and its efficient application to Gaussian mixture optimization
abstract
We have previously proposed a cross-validation (CV) based Gaussian mixture optimization method that efficiently optimizes the model structure based on CV likelihood.In this study, we propose aggregated cross-validation (AgCV) that introduces a bagging-like approach in the CV framework to reinforce the model selection ability.While a single model is used in CV to evaluate a held-out subset, AgCV uses multiple models to reduce the variance in the score estimation.By integrating AgCV instead of CV in the Gaussian mixture optimization algorithm, an AgCV likelihood based Gaussian mixture optimization algorithm is obtained.The algorithm works efficiently by using sufficient statistics and can be applied to large models such as Gaussian mixture HMM.The proposed algorithm is evaluated by speech recognition experiments on oral presentations and it is shown that lower word error rates are obtained by the AgCV optimization method when compared to CV and MDL based methods.
Takahiro Shinozaki, Sadaoki Furui, Tatsuya Kawahara
INTERSPEECH1
2008 Cross-validation and aggregated EM training for robust parameter estimation
Takahiro Shinozaki, Mari Ostendorf
Comput. Speech Lang.1
2007 HMM training based on CV-EM and CV Gaussian mixture optimization
abstract
A combination of the cross-validation EM (CV-EM) algorithm and the cross-validation (CV) Gaussian mixture optimization method is explored. CV-EM and CV Gaussian mixture optimization are our previously proposed training algorithms that use CV likelihood instead of the conventional training set likelihood for robust model estimation. Since CV-EM is a parameter optimization method and CV Gaussian mixture optimization is a structure optimization algorithm, these methods can be combined. Large vocabulary speech recognition experiments are performed on oral presentations. It is shown that both CV-EM and CV Gaussian mixture optimization give lower word error rates than the conventional EM, and their combination is effective to further reduce the word error rate.
Takahiro Shinozaki, Tatsuya Kawahara
ASRU1
2007 Model Complexity Selection and Cross-Validation EM Training for Robust Speaker Diarization
abstract
Accurate modeling of speaker clusters is important in the task of speaker diarization. Creating accurate models involves both selection of the model complexity and optimum training given the data. Using models with fixed complexity and trained using the standard EM algorithm poses a risk of overfitting, which can lead to a reduction in diarization performance. In this paper a technique proposed by the author to estimate the complexity of a model is combined with a novel training algorithm called "cross-validation EM" to control the number of training iterations. This combination leads to more robust speaker modeling and results in an increase in speaker diarization performance. Tests on the NIST RT (MDM) datasets for meetings show a relative improvement of 10.6% relative on the test set.
Xavier Anguera Miró, Takahiro Shinozaki, Chuck Wooters, Javier Hernando
ICASSP (4)2
2007 Cross-Validation EM Training for Robust Parameter Estimation
abstract
A new maximum likelihood training algorithm is proposed that compensates for weaknesses of the EM algorithm by using cross-validation likelihood in the expectation step to avoid overtraining. By using a set of sufficient statistics associated with a partitioning of the training data, as in parallel EM, the algorithm has the same order of computational requirements as the original EM algorithm. Analyses using a GMM with artificial data show the proposed algorithm is more robust for overtraining than the conventional EM algorithm. Large vocabulary recognition experiments on Mandarin broadcast news data show that the method makes better use of more parameters and gives lower recognition error rates than EM training.
Takahiro Shinozaki, Mari Ostendorf
ICASSP (4)1
2007 Gaussian mixture optimization for HMM based on efficient cross-validation
abstract
A Gaussian mixture optimization method is explored using cross-validation likelihood as an objective function instead of the conventional training set likelihood. The optimization is based on reducing the number of mixture components by selecting and merging a pair of Gaussians step by step base on the objective function so as to remove redundant components and improve the generality of the model. Cross-validation likelihood is more appropriate for avoiding over-fitting than the conventional likelihood and can be efficiently computed using sufficient statistics. It results in a better Gaussian pair selection and provides a termination criterion that does not rely on empirical thresholds. Large-vocabulary speech recognition experiments on oral presentations show that the cross-validation method gives a smaller word error rate with an automatically determined model size than a baseline training procedure that does not perform the optimization. Index Terms: speech recognition, HMM, Gaussian mixture, cross-validation, sufficient statistics
Takahiro Shinozaki, Tatsuya Kawahara
INTERSPEECH1
2006 Hmm State Clustering Based on Efficient Cross-Validation
abstract
Decision tree state clustering is explored using a cross validation likelihood criterion. Cross-validation likelihood is more reliable than conventional likelihood and can be efficiently computed using sufficient statistics. It results in a better tying structure and provides a termination criterion that does not rely on empirical thresholds. Large vocabulary recognition experiments on conversational telephone speech show that, for large numbers of tied states, the cross-validation method gives more robust results
Takahiro Shinozaki
ICASSP (1)1
2006 Investigation on Mandarin broadcast news speech recognition
abstract
This paper describes the authors' efforts in developing a competitive Mandarin broadcast news speech recognizer. They have successfully incorporated the most popular speech technologies into their system. More importantly, they present two novel algorithms for smoothing pitch features and segmenting Chinese characters into word units. In addition, they propose to borrow the principle of point-wise mutual information for creating a Chinese word lexicon automatically. Their final system achieved a 6.0% character error rate (CER) on dev04 and a 16.0% CER on eval04 with simpler acoustic models, less training data, and simpler decoding architecture compared with other state-of-the-art systems. This system is equally competitive.
Mei-Yuh Hwang, Wen Wang 0001, Takahiro Shinozaki
INTERSPEECH4
2005 Cluster-based modeling for ubiquitous speech recognition
abstract
In order to realize speech recognition systems that can achieve high recognition accuracy for ubiquitous speech, it is crucial to make the systems flexible enough to cope with a large variability of spontaneous speech. This paper investigates two speech recognition methods that can adapt to speech variation using a large number of models trained based on clustering techniques; one automatically builds a model adapted to input speech using recognition hypotheses and clustered models, and the other directly uses clustered models in parallel. Both methods have been confirmed to be effective by evaluation experiments using presentation speech. Although the latter method needs a large amount of computation, it has an advantage in that it can be applied to online recognition, since it does not need recognition hypotheses. The former method can also be applied to online recognition, if the text of proceedings for the presentation can be used in place of recognition hypotheses. 1.
Sadaoki Furui, Tomohisa Ichiba, Takahiro Shinozaki, Edward W. D. Whittaker, Koji Iwano
INTERSPEECH3
2005 Data sampling for improved speech recognizer training
Takahiro Shinozaki, Mari Ostendorf, Les E. Atlas
INTERSPEECH1
2004 Spontaneous speech recognition using a massively parallel decoder
abstract
Since spontaneous utterances include many variations, speakerand task-independent general models do not work well.This paper proposes combining cluster-based language and acoustic models based on the framework of Massively Parallel Decoder (MPD).The MPD is a parallel decoder that has a large number of decoding units, in which each unit is assigned to each combination of element models.It runs efficiently on a parallel computer, and thus the turnaround time is comparable to conventional decoders using a single model and a processor.In the experiments conducted using lecture speeches from the Corpus of Spontaneous Japanese, two types of cluster models have been investigated: lecture-based cluster models and utterancebased cluster models.It has been confirmed that utterancebased cluster models give significantly lower recognition error rate than lecture-based cluster models in both language and acoustic modeling.It has also been shown that roughly 100 decoding units are enough in terms of recognition rate, and in the best setting, 12% reduction in word error rate was obtained in comparison with the conventional decoder.
Takahiro Shinozaki, Sadaoki Furui
INTERSPEECH1
2003 Unsupervised class-based language model adaptation for spontaneous speech recognition
abstract
This paper proposes an unsupervised, batch-type, class-based language model adaptation method for spontaneous speech recognition. The word classes are automatically determined by maximizing the average mutual information between the classes using a training set. A class-based language model is built based on recognition hypotheses obtained using a general word-based language model, and linearly interpolated with the general language model. All the input utterances are re-recognized using the adapted language model. The proposed method was applied to the recognition of spontaneous presentations and was found to be effective in improving the recognition accuracy for all the presentations. The best condition was found to be using 100 word classes, and in this condition 2.3% of the absolute value improvement in the word accuracy averaged over all the speakers was achieved.
Tadasuke Yokoyama, Takahiro Shinozaki, Koji Iwano, Sadaoki Furui
ICASSP (1)2
2003 Time adjustable mixture weights for speaking rate fluctuation
abstract
Abstract Oneof themostseriousproblemsin spontaneousspeechrecog-nition is the degradation of recognition accuracy due to thespeaking rate fluctuation in an utterance. This paper proposesa method for adjusting mixture weights of an HMM frame byframe depending on the local speaking rate. The proposedmethodisimplementedusingtheBayesiannetworkframework.A hidden variable representing the variation of the “mode” ofthespeakingrateisintroducedanditsvaluecontrolsthemixtureweights of Gaussian mixtures. Model training and maximumprobability assignment of the variablesare conducted usingtheEM/GEMandinferencealgorithmsforBayesiannetworks. TheBayesian network is used to rescore the acoustic likelihood ofthe hypotheses in N-best lists. Experimental results show thatthe proposed method improves word accuracy by 1.6% for theabsolute value on meeting speech given the speaking rate in-formation, whereas improvement by a regression HMM is lesssignificant. 1. Introduction Conventional HMM-based recognizers suffer a lower recogni-tion rate for spontaneous speech. One of the significant factorsreducing the recognition rate is the variable nature of sponta-neousutterances[1]. For example,thespeakingratecanchangeeven within one utterance. A possible strategy to manage thisproblem is first estimating the speaking rate and then adjustingarecognizer basedonthe speakingrate. Thispaperinvestigateshow to use the speaking rate information to control probabilitydensity functions in the HMM.Since estimating the speaking rate is itself a difficult prob-lem and assuming an accurate detection is unrealistic, themethod of adjusting the decoder accordingto the speaking rateshould be probabilistic. This paper proposes a method of ad-justing mixture weights of the HMM at every frame based onthe speaking rate. A hidden variable that represents a “mode”ofthespeakingratecontrolsthemixtureweights. Theproposedmethod can model complex changes of acoustic characteristicswith only asmall increaseof the numberof parameters.The proposed method is implemented using the Bayesiannetwork framework to realize detailed control of HMM param-eters[2]. WeuseGMTK[3]for modelparametertraining usingthe EM/GEM algorithms and for decoding. The network ac-ceptsthespeakingrateinadditiontotheusualacousticfeatures.The proposed method is used to rescore the acoustic likelihoodof N-best hypotheses.This paper is organized as follows. In Section 2, the pro-posed method is formulated as a Bayesian network. In Sec-tion 3, an experimental set up is described. Training processesof the Bayesian network are described in Section 4, and ob-tained parameters related to the speaking rate are analyzed inSection 5. The proposed method is applied to meeting speechrecognition in Section 6 and the results are discussed in Sec-tion 7. It is shown that the proposed method is more effectivein improving the recognition ratethan aregressionHMM givenspeaking rate information. Finally, in Section 8, the paper isconcluded.
Takahiro Shinozaki, Sadaoki Furui
INTERSPEECH1
2002 Analysis on individual differences in automatic transcription of spontaneous presentations
abstract
This paper reports an analysis of individual differences in spontaneous presentation speech recognition performances. Ten minutes from each presentation given by 50 male speakers, for a total of 500 minutes, has been automatically recognized for the analysis. Correlation and regression analyses were applied to the word recognition accuracy and various speaker attributes. A restricted set of the speaker attributes comprising the speaking rate, the out of vocabulary rate and the repair rate was found to be most significant to yield individual differences in the word accuracy. Unsupervised MLLR speaker adaptation worked well for improving the word accuracy but did not change the structure of the individual differences. Approximately half of the variance in the word accuracy was explained by a regression model using the limited set of three attributes.
Takahiro Shinozaki, Sadaoki Furui
ICASSP1
2002 A new lexicon optimization method for LVCSR based on linguistic and acoustic characteristics of words
abstract
This paper proposes a new lexicon optimization method to improve recognition rate of large scale spontaneous speech recognition.Occurrence count and length of a word has strong correlation with difficulty of recognizing the word.First, we investigate the relation and make a word correctness probability model.The proposed method optimizes the lexicon by making compound words or phrases step by step based on the word correctness probability model so as to improve the estimated recognition rate of the system.The optimization method is applied to a large scale Japanese spontaneous speech corpus.Experimental results show that the language model using the optimized lexicon improves the recognition rate.
Takahiro Shinozaki, Sadaoki Furui
INTERSPEECH1
2001 Ubiquitous speech processing
abstract
In the ubiquitous (pervasive) computing era, it is expected that everybody will access information services anytime anywhere, and these services are expected to augment various human intelligent activities. Speech recognition technology can play an important role in this era by providing: (a) conversational systems for accessing information services and (b) systems for transcribing, understanding and summarizing ubiquitous speech documents such as meetings, lectures, presentations and voicemails. In the former systems, robust conversation using wireless handheld/hands-free devices in the real mobile computing environment will be crucial as will multimodal speech recognition technology. To create the latter systems, the ability to understand and summarize speech documents is one of the key requirements. The paper presents technological perspectives and introduces several research activities being conducted from these standpoints in our research group.
Sadaoki Furui, Koji Iwano, Chiori Hori, Takahiro Shinozaki, Yohei Saito, Satoshi Tamura
ICASSP4
2001 Towards automatic transcription of spontaneous presentations
abstract
This paper reports various investigations on recognizing spontaneous presentation speech in connection with the "Spontaneous Speech" national project started in 1999.Presentation speech uttered by 10 male speakers of approximately 4.5 hours duration has been recognized.Experimental results show that acoustic and language modeling based on an actual spontaneous speech corpus is far more effective than conventional modeling based on read speech.The recognition accuracy has a wide speaker-tospeaker variability according to the speaking rate, the number of fillers, the number of repairs, etc.It was confirmed that unsupervised speaker adaptation of acoustic models was effective to improve the recognition accuracy.The recognition accuracy for spontaneous speech is, however, still rather low, and there remains a large number of research issues.
Takahiro Shinozaki, Chiori Hori, Sadaoki Furui
INTERSPEECH1
2000 Toward the realization of spontaneous speech recognition - introduction of a Japanese priority program and preliminary results -
abstract
Although high recognition accuracy can be obtained for speech in the form of reading a written text or similar by using state-of-the art speech recognition technology, the accuracy is quite poor for freely spoken spontaneous speech. From this perspective, a five-year national project for raising the technological level of speech recognition and understanding commenced in Japan in 1999. The project focuses on building a large-scale spontaneous speech corpus and acoustic and linguistic modeling for spontaneous speech recognition and summarization. This paper reports some results of preliminary experiments which have been conducted at Tokyo Institute of Technology. Experimental results show that acoustic and language modeling based on the actual spontaneous speech corpus is far more effective than modeling based on read speech. It was also shown that our proposed automatic speech summarization method could effectively extract relatively important information and remove redundant and irrelevant information. 1.
Sadaoki Furui, Kikuo Maekawa, Hitoshi Isahara, Takahiro Shinozaki, Takashi Ohdaira
INTERSPEECH4