Takashi Fukuda

dblp:98/4675 · DBLP profile ↗
← Back
47ranked-venue papers
26as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 46 · 26 first-author · 10 since 2021Artificial intelligence and machine learning · 30 · 16 first-author · 7 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities
abstract
Granite-speech LLMs are compact and efficient speech language models specifically designed for English ASR1and automatic speech translation (AST). The models were trained by modality aligning granite-3.3-instruct to speech on publicly available open-source corpora. Comprehensive benchmarking on English ASR shows that they outperform several competitors’ models that were trained on orders of magnitude more proprietary data, and they keep pace on English-to-X AST for major European languages, Japanese, and Mandarin. The speech-specific components are: a conformer acoustic encoder using block attention and self-conditioning trained with connectionist temporal classification, a windowed query-transformer speech modality adapter used to do temporal downsampling of the acoustic embeddings and map them to the LLM text embedding space, and LoRA adapters to further fine-tune the text LLM. The models are freely available on HuggingFace2under a permissive Apache 2.0 license.1The latest models (revision 3.3.2) support multilingual ASR in English, French, German, Spanish and Portuguese and bidirectional speech translation to and from English. This paper covers the initial English-only release.2https://huggingface.co/ibm-granite/granite-speech-3.3-2b (and…-8b).
George Saon, Avihu Dekel, Alexi Brooks, Tohru Nagano, Abraham Daniels, Aharon Satt, Ashish R. Mittal, Brian Kingsbury, David Haws, Edmilson da Silva Morais, Gakuto Kurata, Hagai Aronowitz, Ibrahim Ibrahim, Hong-Kwang Jeff Kuo, Kate Soule, Luis A. Lastras, Masayuki Suzuki, Ron Hoory, Samuel Thomas 0001, Sashi Novitasari, Takashi Fukuda, Vishal Sunder, Zvi Kons
ASRU21
2025 Knowledge Distillation Based Training of Unified Conformer CTC Models for Multi-form ASR
abstract
There is an on-going body of research on training separate dedicated models for either short-form or long-form utterances. Multi-form acoustic models that are simply trained on combined data from long-form and short-form utterances often suffer from various negative impacts due to the diversity of a speaking style, an accent, and a recording condition. In addition, a linguistic mismatch that comes from an utterance length is also another factor of the degradation. In this paper we investigate novel techniques for training unified Conformer-based models on multi-form speech data obtained from diverse domains and sources to serve multiple downstream applications with a single model. Our approach incorporates chunk-wise short-term discriminative knowledge distillation with an encoder embedding masking and mitigates the aforementioned problems that appear for single unified models. We show the benefit of our proposed technique on long and short-form ASR test sets by comparing our models against several variants trained by mixing utterances with various audio lengths. The proposed technique provides a significant improvement of up to 8.5% relative WER reduction over baseline systems that operate at a similar decoding cost.
Takashi Fukuda, Gakuto Kurata, George Saon
ICASSP1
2025 Improving End-to-end Mixed-case ASR with Knowledge Distillation and Integration of Voice Activity Cues
Sashi Novitasari, Takashi Fukuda, Gakuto Kurata
INTERSPEECH2
2025 Voice Activity-based Text Segmentation for ASR Text Denormalization
Sashi Novitasari, Takashi Fukuda, Gakuto Kurata
INTERSPEECH2
2023 Effective Training of RNN Transducer Models on Diverse Sources of Speech and Text Data
abstract
This paper proposes a novel modeling framework for effective training of end-to-end automatic speech recognition (ASR) models on various sources of data from diverse domains: speech paired with clean ground truth transcripts, speech with noisy pseudo transcripts from semi-supervised decodes and unpaired text-only data. In our proposed approach, we build a recurrent neural network transducer (RNN-T) model with a shared multimodal encoder, multi-branch prediction networks and a shared common joint network. To train on unpaired text-only data sets along with transcribed speech data, the shared encoder is trained to process both speech and text modalities. Differences in data from multiple domains are effectively handled by training a multi-branch prediction network on various different data sets before an interpolation step combines the multi-branch prediction networks back into a computationally-efficient single branch. We show the benefit of our proposed technique on several ASR test sets by comparing our models to those trained by simple data mixing. The technique provides a significant relative improvement of up to 6% over baseline systems operating at a similar decoding cost.
Takashi Fukuda, Samuel Thomas 0001
ICASSP1
2022 Improving Generalization of Deep Neural Network Acoustic Models with Length Perturbation and N-best Based Label Smoothing
abstract
We introduce two techniques, length perturbation and n-best based label smoothing, to improve generalization of deep neural network (DNN) acoustic models for automatic speech recognition (ASR).Length perturbation is a data augmentation algorithm that randomly drops and inserts frames of an utterance to alter the length of the speech feature sequence.N-best based label smoothing randomly injects noise to ground truth labels during training in order to avoid overfitting, where the noisy labels are generated from n-best hypotheses.We evaluate these two techniques extensively on the 300-hour Switchboard (SWB300) dataset and an in-house 500-hour Japanese (JPN500) dataset using recurrent neural network transducer (RNNT) acoustic models for ASR.We show that both techniques improve the generalization of RNNT models individually and they can also be complementary.In particular, they yield good improvements over a strong SWB300 baseline and give state-of-art performance on SWB300 using RNNT models.
George Saon, Tohru Nagano, Masayuki Suzuki, Takashi Fukuda, Brian Kingsbury, Gakuto Kurata
INTERSPEECH5
2022 Global RNN Transducer Models For Multi-dialect Speech Recognition
Takashi Fukuda, Samuel Thomas 0001, Masayuki Suzuki, Gakuto Kurata, George Saon, Brian Kingsbury
INTERSPEECH1
2022 Improving ASR Robustness in Noisy Condition Through VAD Integration
Sashi Novitasari, Takashi Fukuda, Gakuto Kurata
INTERSPEECH2
2021 Generalized Knowledge Distillation from an Ensemble of Specialized Teachers Leveraging Unsupervised Neural Clustering
abstract
This paper proposes an improved generalized knowledge distillation framework with multiple dissimilar teacher networks, each of which is specialized for a specific domain, to make a deployable student network more robust to challenging acoustic environments. In this paper, we first address a method to partition the training data for constructing ensembles of the teachers from unsupervised neural clustering with features based on context-dependent phonemes representing each acoustic domain. Second, we illustrate how a single student network designed from partitioned data is effectively trained with multiple specialized teachers. During the training step, the weights of the student network are updated using a composite two-part cross entropy loss obtained from a pair consisting of a specialized teacher corresponding to input speech and a generalized teacher trained with a balanced data set. Unlike system combination methods, we aim to incorporate the benefits from multiple models into a single student network via knowledge distillation that does not increase any computational costs during the decoding time. The improvement of the proposed technique is shown on acoustically diverse signals contaminated by challenging practical noises.
Takashi Fukuda, Gakuto Kurata
ICASSP1
2021 Knowledge Distillation Based Training of Universal ASR Source Models for Cross-Lingual Transfer
Takashi Fukuda, Samuel Thomas 0001
Interspeech1
2020 Implicit Transfer of Privileged Acoustic Information in a Generalized Knowledge Distillation Framework
Takashi Fukuda, Samuel Thomas 0001
INTERSPEECH1
2019 Mixed Bandwidth Acoustic Modeling Leveraging Knowledge Distillation
abstract
Training of mixed bandwidth acoustic models have recently been realized by incorporating special Mel filterbanks. To fit information into every filterbank bin available across both narrowband and wideband data, these filterbanks pad zeros at high frequency ranges of narrowband data. Although these methods succeed in decreasing word error rates (WER) on broadband data, they fail to improve on narrowband signals. In this paper, we propose methods to mitigate these effects with generalized knowledge distillation. In our method, specialized teacher networks are first trained on lossless acoustic features with full scale Mel filterbanks. While training student networks, privileged knowledge from these teacher networks is then used to compensate for missing information at high frequencies introduced by the special Mel filterbanks. We show the benefit of the proposed technique for both narrowband (10% relative WER improvement) and wideband data (7.5% relative WER improvement) on the Aurora 4 task over traditional methods.
Takashi Fukuda, Samuel Thomas 0001
ASRU1
2019 Data Augmentation Based on Vowel Stretch for Improving Children's Speech Recognition
abstract
Prolongation is a speech disfluency that lengthens some portions of speech utterances. It is frequently observed in children's spontaneous speech, while it is rare in read speech. To make acoustic models more robust to children's spontaneous speech, collecting a large amount of children's speech data containing prolongation is usually required, which is very impractical in many cases. To tackle this problem, we propose a novel data augmentation method that virtually generates additional data by simulating prolongation. The method inserts pseudo frames into specific positions of speech utterances to simulate prolongation. The acoustic features of the inserted frames are calculated from the original frames on both sides. This is based on our analysis that many of vowels are actually stretched in children's spontaneous speech. Our proposed procedure can generate partially stretched utterances with low computational costs, unlike a conventional speed or tempo perturbation method that extends and shrinks entire utterances at a uniform rate. The effectiveness of the proposed method were confirmed with the experiments of acoustic model adaptations, in which our proposed method focusing on vowel stretch showed consistent improvement compared with conventional speed and tempo perturbation approach.
Tohru Nagano, Takashi Fukuda, Masayuki Suzuki, Gakuto Kurata
ASRU2
2019 Automated Testing of Basic Recognition Capability for Speech Recognition Systems
abstract
Automatic speech recognition systems transform speech audio data into text data, i.e., word sequences, as the recognition results. These word sequences are generally defined by the language model of the speech recognition system. Therefore, the capability of the speech recognition system to translate audio data obtained by typically pronouncing word sequences that are accepted by the language model into word sequences that are equivalent to the original ones can be regarded as a basic capability of the speech recognition systems. This work describes a testing method that checks whether speech recognition systems have this basic recognition capability. The method can verify the basic capability by performing the testing separately from recognition robustness testing. It can also be fully automated. We constructed a test automation system and evaluated though several experiments whether it could detect defects in speech recognition systems. The results demonstrate that the test automation system can effectively detect basic defects at an early phase of speech recognition development or refinement.
Futoshi Iwama, Takashi Fukuda
ICST2
2019 Direct Neuron-Wise Fusion of Cognate Neural Networks
Takashi Fukuda, Masayuki Suzuki, Gakuto Kurata
INTERSPEECH1
2018 Data Augmentation Improves Recognition of Foreign Accented Speech
Takashi Fukuda, Raul Fernandez, Andrew Rosenberg, Samuel Thomas 0001, Bhuvana Ramabhadran, Alexander Sorin, Gakuto Kurata
INTERSPEECH1
2018 Detecting breathing sounds in realistic Japanese telephone conversations and its application to automatic speech recognition
Takashi Fukuda, Osamu Ichikawa, Masafumi Nishimura
Speech Commun.1
2017 Effective joint training of denoising feature space transforms and Neural Network based acoustic models
abstract
Neural Network (NN) based acoustic frontends, such as denoising autoencoders, are actively being investigated to improve the robustness of NN based acoustic models to various noise conditions. In recent work the joint training of such frontends with backend NNs has been shown to significantly improve speech recognition performance. In this paper, we propose an effective algorithm to jointly train such a denoising feature space transform and a NN based acoustic model with various kinds of data. Our proposed method first pretrains a Convolutional Neural Network (CNN) based denoising frontend and then jointly trains this frontend with a NN backend acoustic model. In the unsupervised pretraining stage, the frontend is designed to estimate clean log Mel-filterbank features from noisy log-power spectral input features. A subsequent multi-stage training of the proposed frontend, with the dropout technique applied only at the joint layer between the frontend and backend NNs, leads to significant improvements in the overall performance. On the Aurora-4 task, our proposed system achieves an average WER of 9.98%. This is a 9.0% relative improvement over one of the best reported speaker independent baseline system's performance. A final semi-supervised adaptation of the frontend NN, similar to feature space adaptation, reduces the average WER to 7.39%, a further relative WER improvement of 25%.
Takashi Fukuda, Osamu Ichikawa, Gakuto Kurata, Ryuki Tachibana, Samuel Thomas 0001, Bhuvana Ramabhadran
ICASSP1
2017 Harmonic feature fusion for robust neural network-based acoustic modeling
abstract
Acoustic modeling with deep learning has drastically improved the performance of automatic speech recognition (ASR) where the main stream of the acoustic feature is still log-Mel filtered one. While the log-Mel filtered features lose harmonic-structure information, they still include useful information for ASR. Several attempts have been made to integrate higher-resolution information into the network. In order to improve the ASR accuracy in noisy conditions, we propose new features integrated into acoustic modeling to represent which parts in the time-frequency domain have a distinct harmonic structure, since it is partially observed in noisy environments. The new features are combined with the standard acoustic features, and the network is trained with them using various noisy data. Through these operations, it learns the acoustic features with a kind of quality tag describing which parts are clean or degraded. Our model reduced the word error rate in an Aurora-4 task by 10.3% in DNN compared with the strong baseline while retaining the high accuracy in clean test cases.
Osamu Ichikawa, Takashi Fukuda, Masayuki Suzuki, Gakuto Kurata, Bhuvana Ramabhadran
ICASSP2
2017 Efficient Knowledge Distillation from an Ensemble of Teachers
Takashi Fukuda, Masayuki Suzuki, Gakuto Kurata, Samuel Thomas 0001, Jia Cui, Bhuvana Ramabhadran
INTERSPEECH1
2017 Ensembles of Multi-Scale VGG Acoustic Models
Michael Heck, Masayuki Suzuki, Takashi Fukuda, Gakuto Kurata, Satoshi Nakamura 0001
INTERSPEECH3
2017 Factorial Modeling for Effective Suppression of Directional Noise
Osamu Ichikawa, Takashi Fukuda, Gakuto Kurata, Steven J. Rennie
INTERSPEECH2
2016 Convolutional neural network pre-trained with projection matrices on linear discriminant analysis
abstract
Recently, the hybrid architecture of a neural network (NN) and a hidden Markov model (HMM) has shown significant improvement on automatic speech recognition (ASR) over the conventional Gaussian mixture model (GMM)-based system. The convolutional neural network (CNN), a successful NN-based system, can represent local spectral variations spanning the time-frequency space. Meanwhile, spectro-temporal features have been widely studied to make ASR more robust. Typically, the spectro-temporal features are extracted from acoustic spectral patterns using a 2D filtering process. Convolutional layers in CNN that have various local windows can also be regarded as an efficient feature extractor to capture 2D spectral variations. In a standard procedure, the local windows in CNN are initialized randomly before the pre-training and are iteratively updated with a back propagation algorithm in the pre-training and fine-tuning steps. In this paper, we explore using projection matrices composed of eigenvectors estimated by linear discriminant analysis (LDA) objective function as initial weights for the first convolutional layer in CNN. From analysis of the local windows trained by the proposed method, we can see the eigenvectors of LDA has desirable properties as initial weights of CNN. The proposed method yielded a 8.1% relative improvement compared to CNN with local weights initialized randomly.
Takashi Fukuda, Osamu Ichikawa, Ryuki Tachibana
ICASSP1
2014 Regularized feature-space discriminative adaptation for robust ASR
Takashi Fukuda, Osamu Ichikawa, Masafumi Nishimura, Steven J. Rennie, Vaibhava Goel
INTERSPEECH1
2013 Channel-mapping for speech corpus recycling
abstract
The performance of automatic speech recognition (ASR) is heavily dependent on the acoustic environment in the target domain. Large investments have focused on ways to record speech data in specific environments. In contrast, recent Internet services using hand-held devices such as smartphones have created opportunities to acquire huge amounts of “live” speech data at low cost. There are practical demands to reuse this abundant data in different acoustic environments. To transform such source data for a target domain, developers can use channel mapping and noise addition. However, channel mapping of the data is difficult without stereo mapping data or impulse response data. We tested GMM-based channel mapping with a vector Taylor series (VTS) formulation on a per-utterance basis. We found this type of channel mapping effectively simulated our target domain data.
Osamu Ichikawa, Steven J. Rennie, Takashi Fukuda, Masafumi Nishimura
ICASSP3
2012 Constructing ensembles of dissimilar acoustic models using hidden attributes of training data
abstract
One of the objectives in acoustic modeling is to realize robust statistical models against the wide variety of acoustic conditions that are present in real world environments. As large amounts of training data become available, modeling subsets of the data with similar acoustic qualities can be done accurately and multiple acoustic models are jointly used as a form of system combination or model selection. In this paper, we propose a method to partition the training data for constructing ensembles of acoustic models using metadata attributes such as SNR, speaking rate, and duration via a binary tree. The metadata attribute used at each binary split in the decision tree is obtained using a metric proposed in this paper that is cosine-similarity based. The resulting multiple models are combined using voting techniques such as n-best ROVER. The proposed method improved the recognition accuracy by up to 4% relative over the state-of-the-art system on a large vocabulary continuous speech recognition voice search task.
Takashi Fukuda, Ryuki Tachibana, Upendra V. Chaudhari, Bhuvana Ramabhadran, Puming Zhan
ICASSP1
2012 Model-based noise reduction leveraging frequency-wise confidence metric for in-car speech recognition
abstract
Model-based approaches for noise reduction effectively improve the performance of automatic speech recognition in noisy environments. Most of them use the Minimum Mean Square Estimate (MMSE) criterion for de-noised speech estimates. In general, an observation has speech-dominant bands and noise-dominant bands in the Mel spectral domain. This paper introduces a method to add weight to speech-dominated bands when evaluating the posterior probability of each speech state, as these bands are generally more reliable. To leverage high-resolution information in the Mel domain, we use Local Peak Weight (LPW) as the confidence metric for the degree of speech dominance. This information is also used to regulate the amount of compensation that is applied to each frequency band during feature reconstruction under an integrated probabilistic model. The method produced relative word error rate improvements of up to 33.8% over the baseline MMSE method on an isolated word task with car noise.
Osamu Ichikawa, Steven J. Rennie, Takashi Fukuda, Masafumi Nishimura
ICASSP3
2011 Frame-level AnyBoost for LVCSR with the MMI Criterion
abstract
This paper propose a variant of AnyBoost for a large vocabulary continuous speech recognition (LVCSR) task. AnyBoost is an efficient algorithm to train an ensemble of weak learners by gradient descent for an objective function.We present a novel training procedure that trains acoustic models via the MMI criterion using data that is weighted proportional to the summation of the posterior functions of previous round of weak learners. Optimized for system combination by n-best ROVER at runtime, data weights for a new weak learner are computed as a weighted summation of posteriors of previous weak learners. We compare a frame-based version and a sentence-based version of our proposed algorithm with a frame-based AdaBoost algorithm. We will present results on a voice search task trained with different amounts of data with gains of 5.1% to 7.5% relative in WER can be obtained by three rounds of boosting.
Ryuki Tachibana, Takashi Fukuda, Upendra V. Chaudhari, Bhuvana Ramabhadran, Puming Zhan
ASRU2
2011 Combining Feature Space Discriminative Training with Long-Term Spectro-Temporal Features for Noise-Robust Speech Recognition
Takashi Fukuda, Osamu Ichikawa, Masafumi Nishimura
INTERSPEECH1
2011 Breath-Detection-Based Telephony Speech Phrasing
Takashi Fukuda, Osamu Ichikawa, Masafumi Nishimura
INTERSPEECH1
2010 Improved voice activity detection using static harmonic features
abstract
Accurate voice activity detection (VAD) is important for robust automatic speech recognition (ASR) systems. We have proposed a statistical-model-based VAD using the long-term temporal information in speech, which shows good robustness against noise in an automobile environment. For further improvement, this paper describes a new method to exploit harmonic structure information with statistical models. In our approach, local peaks considered to be harmonic structures are extracted, without explicit pitch detection and voiced-unvoiced classification. The proposed method including both long-term temporal and static harmonic features led to considerable improvements under low SNR conditions in our VAD testing. In addition, the word error rate was reduced by 29.1% in a test that included a full ASR system.
Takashi Fukuda, Osamu Ichikawa, Masafumi Nishimura
ICASSP1
2009 Dynamic features in the linear domain for robust automatic speech recognition in a reverberant environment
Osamu Ichikawa, Takashi Fukuda, Ryuki Tachibana, Masafumi Nishimura
INTERSPEECH2
2008 Local peak enhancement combined with noise reduction algorithms for robust automatic speech recognition in automobiles
abstract
The accuracy of automatic speech recognition in automobiles is significantly degraded in very low SNR (Signal to Noise Ratio) situations such as “Fan high” or “Window open”. In such cases, speech signals are often buried in broadband noise. In this paper, we propose a novel approach for such situations that utilizes harmonic structures in the human voice. It pursues two objectives. (1) Unlike comb filtering, it should not rely on F0 detection or voiced/unvoiced detection, since they are not accurate enough in noisy environments. (2) It should work with existing noise reduction algorithms. In our new approach, an observed power spectrum is directly converted into a filter for speech enhancement by retaining only the local peaks considered to be harmonic structures. In our experiments, we reduced the word error rate significantly in realistic automobile environments, and our approach showed further improvements when used with existing noise reduction algorithms.
Osamu Ichikawa, Takashi Fukuda, Masafumi Nishimura
ICASSP2
2008 Phone-duration-dependent long-term dynamic features for a stochastic model-based voice activity detection
Takashi Fukuda, Osamu Ichikawa, Masafumi Nishimura
INTERSPEECH1
2008 Short- and long-term dynamic features for robust speech recognition
Takashi Fukuda, Osamu Ichikawa, Masafumi Nishimura
INTERSPEECH1
2005 Pitch-Synchronous ZCPA (PS-ZCPA)-Based Feature Extraction with Auditory Masking
abstract
A pitch-synchronous (PS) auditory feature extraction method, based on ZCPA (zero-crossings peak-amplitudes), has been proposed (Ghulam, M. et al., Proc. ICSLP04, 2004) and was shown to be more robust than the conventional ZCPA (Kim, D.S. et al., IEEE Trans. Speech Audio Process., vol.7, no.1, p.55-69, 1999). We examine the effect of auditory masking, both simultaneous and temporal, in the PS-ZCPA method. We also observe the effect of varying the number of histogram bins on the way to find out the optimum parameters of the proposed method. Experimental results demonstrate the improved performance of the PS-ZCPA method achieved by embedding auditory masking into it; for example, with both the masking methods embedded, the performance increases to 73.71% from the 69.92% obtained without masking for PS-ZCPA, while it showed little improvement with an increased number of histogram bins.
Muhammad Ghulam, Takashi Fukuda, Junsei Horikawa, Tsuneo Nitta
ICASSP (1)2
2005 Designing multiple distinctive phonetic feature extractors for canonicalization by using clustering technique
abstract
Acoustic models of an HMM-based classifier include various types of hidden factors such as speaker-specific characteristics and acoustic environments. If there exist a canonicalization process that represses the decrease of differences in acoustic-likelihood among categories resulted from hidden factors, a robust ASR system can be realized. We have previously proposed the canonicalization process of featureparameters composed of three distinctive phonetic feature (DPF) extractors focused on a gender factor. This paper describes an attempt to design multiple DPF extractors corresponding to unspecific hidden factors, as well as to introduce a noise suppressor that is targeted for the canonicalization of a noise factor. In an experiment on Japanese version AURORA2 database (AURORA2-J), the proposed system achieved significant improvements when combining the canonicalization process with the noise reduction technique based on a two-stage Wiener filter.
Takashi Fukuda, Muhammad Ghulam, Tsuneo Nitta
INTERSPEECH1
2004 Canonicalization of feature parameters for automatic speech recognition
abstract
Acoustic models (AMs) of an HMM-based classifier include various types of hidden variables such as gender type, speaking rate, and acoustic environment. If there exists a canonicalization process that reduces the influence of the hidden variables from the AMs, a robust automatic speech recognition (ASR) system can be realized. In this paper, we describe the configuration of a canonicalization process targeting gender type as a hidden variable. The proposed canonicalization process is composed of multiple distinctive phonetic feature (DPF) extractors corresponding to the hidden variable and a DPF selector in which the distance between input DPF and AMs is compared. In a DPF extraction stage, an input sequence of acoustic feature vectors is mapped onto three DPF spaces corresponding to male, female, and neutral voice by using three multilayer neural networks (MLNs). Experiments are carried out by comparing (A) the combination of the canonicalized DPF and a single HMM classifier, and (B) the combination of a single acoustic feature (MFCC) and multiple HMM classifiers. The result shows that the proposed canonicalization method outperforms both of the conventional ASR with MFCC and a single HMM and the ASR with multiple HMMs in spite of less memories and computation time.
Takashi Fukuda, Tsuneo Nitta
INTERSPEECH1
2004 A noise-robust feature extraction method based on pitch-synchronous ZCPA for ASR
abstract
In this paper, we propose a novel feature extraction method based on an auditory nervous system for robust automatic speech recognition (ASR). In the proposed method, a pitchsynchronous mechanism is embedded in ZCPA (ZeroCrossings Peak-Amplitudes), which has previously been shown to outperform the conventional features in the presence of noise. A noise-robust non-delayed pitch determination algorithm (PDA) is also developed. In the experiment, the proposed pitch-synchronous ZCPA (PS-ZCPA) was proved more robust than the original ZCPA method. Moreover, a simple noise subtraction (NS) method is also integrated in the proposed method and the performance was evaluated using the Aurora-2J database. The experimental results showed the superiority of the proposed PS-ZCPA method with NS over the PS-ZCPA method without NS.
Muhammad Ghulam, Takashi Fukuda, Junsei Horikawa, Tsuneo Nitta
INTERSPEECH2
2003 Distinctive phonetic feature extraction for robust speech recognition
abstract
The paper describes an attempt to extract distinctive phonetic features (DPFs) that represent articulatory gestures in linguistic theory by using a multilayer neural network (MLN) and to apply the DPFs to noise-robust speech recognition. In the DPF extraction stage, after converting a speech signal to acoustic features composed of local features (LFs), an MLN with 33 output units, corresponding to context-dependent DPFs of 11 DPFs, 11 preceding context DPFs, and 11 following context DPFs, maps the LFs to DPFs. The proposed DPF parameters without MFCC (Mel-frequency cepstral coefficients) were firstly evaluated in comparison with a standard parameter set of MFCC and dynamic features on a word recognition task using clean speech; the result showed the same performance as that of the standard set. Noise robustness of these parameters was then tested with four types of additive noise and the proposed DPF parameters outperformed the standard set except for one additive noise type.
Takashi Fukuda, Wataru Yamamoto, Tsuneo Nitta
ICASSP (2)1
2003 Noise-robust ASR by using distinctive phonetic features approximated with logarithmic normal distribution of HMM
Takashi Fukuda, Tsuneo Nitta
INTERSPEECH1
2003 Noise-robust automatic speech recognition using orthogonalized distinctive phonetic feature vectors
abstract
Abstract With the aim of using an automatic speech recognition (ASR) system in practical environments, various approaches focused on noise-robustness such as noise adaptation and reduction techniques have been investigated. We have previously proposed a distinctive phonetic feature (DPF) parameter set for a noise-robust ASR system, which reduced the effect of high-level additive noise[1]. This paper describes an attempt to apply an orthogonalized DPF parameter set as an input of HMMs. In our proposed method, orthogonal bases are calculated using conventional DPF vectors that represent 38 Japanese phonemes, then the Karhunen-Loeve transform (KLT) is used to orthogonalize the DPFs, output from a multi-layer neural network (MLN), by using the orthogonal bases. In experiments, orthogonalized DPF parameters were firstly compared with original DPF parameters on an isolated spoken-word recognition task with clean speech. Noise robustness was then tested with four types of additive noise. The proposed orthogonalized DPFs can reduce the error rate in an isolated spoken-word recognition task both with clean speech and with speech contaminated by additive noise. Furthermore, we achieved significant improvements over a baseline system with MFCC and dynamic feature-set when combining the orthogonalized DPFs with conventional static MFCCs and ∆P.
Takashi Fukuda, Tsuneo Nitta
INTERSPEECH1
2003 Voice quality normalization in an utterance for robust ASR
abstract
In this paper, we propose a novel method of normalizing the voice quality in an utterance for both clean speech and speech contaminated by noise. The normalization method is applied to the N-best hypotheses from an HMM-based classifier, then an SM (Sub-space Method)-based verifier tests the hypotheses after normalizing the monophone scores together with the HMMbased likelihood score. The HMM-SM-based speech recognition system was proposed previously [1, 2] and successfully implemented on a speaker-independent word recognition task and an OOV word rejection task. We extend the proposed system to a connected digit string recognition task by exploring the effect of the voice quality normalization in an utterance for robust ASR and compare it with the HMM-based recognition systems with utterance-level normalization, word-level normalization, monophone-level normalization, and state-level normalization. Experimental results performed on connected 4digit strings showed that the word accuracy was significantly improved from 95.7% obtained by the typical HMM-based system with utterance-level normalization to 98.2% obtained by the HMM-SM-based system for clean speech, from 88.1% to 91.5% for noise-added speech with SNR=10dB, and from 72.4% to 76.4% for noise-added speech with SNR=5dB, while the other HMM-based systems also showed lower performances.
Muhammad Ghulam, Takashi Fukuda, Tsuneo Nitta
INTERSPEECH2
2002 Confidence scoring for accurate HMM-based word recognition by using SM-based monophone score normalization
abstract
In this paper, we propose a novel confidence scoring method that is applied to N-best hypotheses output from an HMM-based classifier. In the first pass of the proposed method, the HMM-based classifier with monophone models outputs N-best hypotheses and boundaries of all the monophones in the hypotheses. In the second pass, an SM(sub-space method)-based verifier tests the hypotheses by comparing confidence scores. We discuss how to convert a monophone similarity score of SM into a likelihood score, how to normalize the variations of acoustic quality in an utterance, and how to combine an HMM-based likelihood of word level and an SM-based likelihood of monophone level. In the experiments performed on speaker-independent word recognition, the proposed confidence scoring method significantly improves correct word recognition rate from 95.3% obtained by the standard HMM classifier to 98.0%.
Takaharu Sato, Muhammad Ghulam, Takashi Fukuda, Tsuneo Nitta
ICASSP3
2002 Improving performance of an HMM-based ASR system by using monophone-level normalized confidence measure
abstract
In this paper, we propose a novel confidence scoring method that is applied to N-best hypotheses output from an HMM-based classifier. In the first pass of the proposed method, the HMM-based classifier with monophone models outputs N-best hypotheses (word candidates) and boundaries of all the monophones in the hypotheses. In the second pass, an SM (Sub-space Method)-based verifier tests the hypotheses by comparing confidence scores. We discuss how to convert a monophone similarity score of SM into a likelihood score, how to normalize the variations of acoustic quality in an utterance, how to combine an HMM-based likelihood of word level and an SM-based likelihood of monophone level, and also how to accept the correct words and reject OOV words. In the experiments performed on speaker-independent word recognition, the proposed confidence scoring method significantly reduced word error rate from 4.7% obtained by the standard HMM classifier to 2.0%, and it also reduced the equal error rate from 9.0% to 6.5% in an unknown word rejection task.
Muhammad Ghulam, Takashi Fukuda, Takaharu Sato, Tsuneo Nitta
INTERSPEECH2
2001 Peripheral features for HMM-based speech recognition
abstract
This paper describes an attempt to extract peripheral features of a point c(t/sub i/,q/sub j/) on a time-quefrency (TQ) pattern by observing n/spl times/n neighborhoods of the point, and then to incorporate these peripheral features into the MFCC-based feature extractor of a speech recognition system as a replacement to dynamic features. In the design of the feature extractor, firstly, the orthogonal bases extracted directly from speech data by using the Karhunen-Loeve transform (KLT) of 7/spl times/3 blocks on a TQ pattern are adopted as the peripheral features, then, the upper two primal bases are selected and simplified in the form of /spl utri//sub t/-operator and /spl utri//sub q/-operator. The proposed feature-set of MFCC and peripheral features shows significant improvements in comparison with the standard feature-set of MFCC and dynamic features in experiments with an HMM-based automatic speech recognition (ASR) system. The reason for the increased performance is discussed in terms of minimal-pair tests.
Takashi Fukuda, Masashi Takigawa, Tsuneo Nitta
ICASSP1
2000 A novel feature extraction using multiple acoustic feature planes for HMM-based speech recognition
Tsuneo Nitta, Masashi Takigawa, Takashi Fukuda
INTERSPEECH3