EDBT 2026 Demo / reviewers in the wild / expert
Norihide Kitaoka
dblp:07/6964
· DBLP profile ↗
70ranked-venue papers
10as first author
13since 2021 · last 2026
0000-0003-2028-8585ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 54 · 9 first-author · 8 since 2021Artificial intelligence and machine learning · 52 · 9 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Corpus for Personalized Dialogue Breakdown Repair in Japanese Open-Domain Conversations
Kazuya Tsubokura, Yurie Iribe, Norihide Kitaoka |
LREC | 3 |
| 2026 | Generation of Dialogue Breakdown Repair Utterance in Non-task Oriented ConversationabstractRecent advances in large language models (LLMs) have significantly improved the response quality of dialogue systems. However, the issue of dialogue breakdowns where systems produce utterances that confuse users, such as those containing misinformation or lacking common sense still persists. Since such breakdowns can negatively affect users’ impressions of dialogue systems, it is essential to appropriately repair the flow of conversation. Nevertheless, no existing method can robustly handle a wide range of breakdown types. To address this issue, this study aims to develop a dialogue breakdown repair generation system that can robustly handle various types of dialogue breakdowns. Specifically, we train an LLM using the Dialogue Breakdown Repair Corpus, which contains repair utterances corresponding to diverse breakdown scenarios. As a result, we construct a model specialized in generating repair utterances for various breakdown types and demonstrate that it achieves higher accuracy than existing models. Kazuya Tsubokura, Yurie Iribe, Norihide Kitaoka |
SIGDIAL | 3 |
| 2025 | Backchannel prediction for natural spoken dialog systems using general speaker and listener information
Yoshinori Fukunaga, Ryota Nishimura, Kengo Ohta, Norihide Kitaoka |
INTERSPEECH | 4 |
| 2025 | Fine-tuning Parakeet-TDT for Dysarthric Speech Recognition in the Speech Accessibility Project Challenge
Kaito Takahashi, Keigo Hojo, Toshimitsu Sakai, Yukoh Wakabayashi, Norihide Kitaoka |
INTERSPEECH | 5 |
| 2025 | Domain adaptation using non-parallel target domain corpus for self-supervised learning-based automatic speech recognitionabstractThe recognition accuracy of conventional automatic speech recognition (ASR) systems depends heavily on the amount of speech and associated transcription data available in the target domain for model training. However, preparing parallel speech and text data each time a model is trained for a new domain is costly and time-consuming. To solve this problem, we propose a method of domain adaptation that does not require the use of a large amount of parallel target domain training data, as most of the data used for model training is not from the target domain. Instead, only target domain speech is used for model training, along with non-target domain speech and its parallel text data, i.e., the domains and contents of the two types of training data do not correspond to one another. Collecting this type of training data is relatively inexpensive. Domain adaptation is performed in two steps: (1) A pre-trained wav2vec 2.0 model is further pre-trained using a large amount of target domain speech data and is then fine-tuned using a large amount of non-target domain speech and its transcriptions. (2) The density ratio approach (DRA) is applied during inference to a language model (LM) trained using target domain text unrelated to, and independently from, the wav2vec 2.0 training. Experimental evaluation illustrated that the proposed domain adaptation obtained character error rate (CER) 10.4 pts lower than baseline with wav2vec 2.0 and 3.9 pts with XLS-R under the situation that the parallel target domain data is unavailable against the target domain test set, achieving 34.4% and 16.2% reductions in relative CER. Takahiro Kinouchi, Atsunori Ogawa, Yukoh Wakabayashi, Kengo Ohta, Norihide Kitaoka |
Speech Commun. | 5 |
| 2024 | Boosting CTC-based ASR using inter-layer attention-based CTC lossabstractThis paper addresses improving the performance of CTC-based models, which leverage the intermediate outputs of all encoder layers with an attention mechanism.Several previous studies have used the intermediate outputs of the encoder layer to modify CTC-based models.Here, we focus on the role of the Transformer encoder layer, and each encoder layer is computed for two CTC losses by weighting the intermediate outputs of its lower and upper layers using an attention mechanism.By dividing the layer into two groups, it is expected to be possible to calculate the loss, taking into account both acoustic and linguistic features.Experimental results showed that the proposed method improved the baseline recognition performance of TEDLIUM2 speech data, achieving a WER of 9.9% on the dev set and 11.8% on the test set.Our method outperformed the conventional methods for WER with only slightly increased inference speed measured by RTF. Keigo Hojo, Yukoh Wakabayashi, Kengo Ohta, Atsunori Ogawa, Norihide Kitaoka |
INTERSPEECH | 5 |
| 2024 | Text-only Domain Adaptation for CTC-based Speech Recognition through Substitution of Implicit Linguistic Information in the Search Space
Tatsunari Takagi, Yukoh Wakabayashi, Atsunori Ogawa, Norihide Kitaoka |
INTERSPEECH | 4 |
| 2023 | Relationships Between Gender, Personality Traits and Features of Multi-Modal Data to Responses to Spoken Dialog Systems Breakdown
Kazuya Tsubokura, Yurie Iribe, Norihide Kitaoka |
INTERSPEECH | 3 |
| 2023 | A new speech corpus of super-elderly Japanese for acoustic modelingabstractThe development of accessible speech recognition technology will allow the elderly to more easily access electronically stored information. However, the necessary level of recognition accuracy for elderly speech has not yet been achieved using conventional speech recognition systems, due to the unique features of the speech of elderly people. To address this problem, we have created a new speech corpus named EARS (Elderly Adults Read Speech), consisting of the recorded read speech of 123 super-elderly Japanese people (average age: 83.1), as a resource for training automated speech recognition models for the elderly. In this study, we investigated the acoustic features of super-elderly Japanese speech using our new speech corpus. In comparison to the speech of less elderly Japanese speakers, we observed a slower speech rate and extended vowel duration for both genders, a slight increase in fundamental frequency for males, and a slight decrease in fundamental frequency for females. To demonstrate the efficacy of our corpus, we also conducted speech recognition experiments using two different acoustic models (DNN-HMM and transformer-based), trained with a combination of data from our corpus and speech data from three conventional Japanese speech corpora. When using the DNN-HMM trained with EARS and speech data from existing corpora, the character error rate (CER) was reduced by 7.8% (to just over 9%), compared to a CER of 16.9% when using only the baseline training corpora. We also investigated the effect of training the models with various amounts of EARS data, using a simple data expansion method. The acoustic models were also trained for various numbers of epochs without any modifications. When using the Transformer-based end-to-end speech recognizer, the character error rate was reduced by 3.0% (to 11.4%) by using a doubled EARS corpus with the baseline data for training, compared to a CER of 13.4% when only data from the baseline training corpora were used. Meiko Fukuda, Ryota Nishimura, Hiromitsu Nishizaki, Koharu Horii, Yurie Iribe, Kazumasa Yamamoto, Norihide Kitaoka |
Comput. Speech Lang. | 7 |
| 2022 | End-to-End Spontaneous Speech Recognition Using Disfluency Labeling
Koharu Horii, Meiko Fukuda, Kengo Ohta, Ryota Nishimura, Atsunori Ogawa, Norihide Kitaoka |
INTERSPEECH | 6 |
| 2022 | Elderly Conversational Speech Corpus with Cognitive Impairment Test and Pilot Dementia Detection Experiment Using Acoustic Characteristics of Speech in Japanese DialectsabstractThere is a need for a simple method of detecting early signs of dementia which is not burdensome to patients, since early diagnosis and treatment can often slow the advance of the disease. Several studies have explored using only the acoustic and linguistic information of conversational speech as diagnostic material, with some success. To accelerate this research, we recorded natural conversations between 128 elderly people living in four different regions of Japan and interviewers, who also administered the Hasegawa’s Dementia Scale-Revised (HDS-R), a cognitive impairment test. Using our elderly speech corpus and dementia test results, we propose an SVM-based screening method which can detect dementia using the acoustic features of conversational speech even when regional dialects are present. We accomplish this by omitting some acoustic features, to limit the negative effect of differences between dialects. When using our proposed method, a dementia detection accuracy rate of about 91% was achieved for speakers from two regions. When speech from four regions was used in a second experiment, the discrimination rate fell to 76.6%, but this may have been due to using only sentence-level acoustic features in the second experiment, instead of sentence and phoneme-level features as in the previous experiment. This is an on-going research project, and additional investigation is needed to understand differences in the acoustic characteristics of phoneme units in the conversational speech collected from these four regions, to determine whether the removal of formants and other features can improve the dementia detection rate. Meiko Fukuda, Ryota Nishimura, Maina Umezawa, Kazumasa Yamamoto, Yurie Iribe, Norihide Kitaoka |
LREC | 6 |
| 2021 | Response type selection for chat-like spoken dialog systems based on LSTM and multi-task learningabstractWe propose a method of automatically selecting appropriate responses in conversational spoken dialog systems by explicitly determining the correct response type that is needed first, based on a comparison of the user’s input utterance with many other utterances. Response utterances are then generated based on this response type designation (back channel, changing the topic, expanding the topic, etc.). This allows the generation of more appropriate responses than conventional end-to-end approaches, which only use the user’s input to directly generate response utterances. As a response type selector, we propose an LSTM-based encoder–decoder framework utilizing acoustic and linguistic features extracted from input utterances. In order to extract these features more accurately, we utilize not only input utterances but also response utterances in the training corpus. To do so, multi-task learning using multiple decoders is also investigated. To evaluate our proposed method, we conducted experiments using a corpus of dialogs between elderly people and an interviewer. Our proposed method outperformed conventional methods using either a point-wise classifier based on Support Vector Machines, or a single-task learning LSTM. The best performance was achieved when our two response type selectors (one trained using acoustic features, and the other trained using linguistic features) were combined, and multi-task learning was also performed. Kengo Ohta, Ryota Nishimura, Norihide Kitaoka |
Speech Commun. | 3 |
| 2021 | Normalization of Transliterated Mongolian Words Using Seq2Seq Model with Limited DataabstractThe huge increase in social media use in recent years has resulted in new forms of social interaction, changing our daily lives. Due to increasing contact between people from different cultures as a result of globalization, there has also been an increase in the use of the Latin alphabet, and as a result a large amount of transliterated text is being used on social media. In this study, we propose a variety of character level sequence-to-sequence (seq2seq) models for normalizing noisy, transliterated text written in Latin script into Mongolian Cyrillic script, for scenarios in which there is a limited amount of training data available. We applied performance enhancement methods, which included various beam search strategies, N-gram-based context adoption, edit distance-based correction and dictionary-based checking, in novel ways to two basic seq2seq models. We experimentally evaluated these two basic models as well as fourteen enhanced seq2seq models, and compared their noisy text normalization performance with that of a transliteration model and a conventional statistical machine translation (SMT) model. The proposed seq2seq models improved the robustness of the basic seq2seq models for normalizing out-of-vocabulary (OOV) words, and most of our models achieved higher normalization performance than the conventional method. When using test data during our text normalization experiment, our proposed method which included checking each hypothesis during the inference period achieved the lowest word error rate (WER = 13.41%), which was 4.51% fewer errors than when using the conventional SMT method. Zolzaya Byambadorj, Ryota Nishimura, Altangerel Ayush, Norihide Kitaoka |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2020 | Improving Speech Recognition for the Elderly: A New Corpus of Elderly Japanese Speech and Investigation of Acoustic Modeling for Speech RecognitionabstractIn an aging society like Japan, a highly accurate speech recognition system is needed for use in electronic devices for the elderly, but this level of accuracy cannot be obtained using conventional speech recognition systems due to the unique features of the speech of elderly people. S-JNAS, a corpus of elderly Japanese speech, is widely used for acoustic modeling in Japan, but the average age of its speakers is 67.6 years old. Since average life expectancy in Japan is now 84.2 years, we are constructing a new speech corpus, which currently consists of the utterances of 221 speakers with an average age of 79.2, collected from four regions of Japan. In addition, we expand on our previous study (Fukuda, 2019) by further investigating the construction of acoustic models suitable for elderly speech. We create new acoustic models and train them using a combination of existing Japanese speech corpora (JNAS, S-JNAS, CSJ), with and without our ‘super-elderly’ speech data, and conduct speech recognition experiments. Our new acoustic models achieve word error rates (WER) as low as 13.38%, exceeding the results of our previous study in which we used the CSJ acoustic model adapted for elderly speech (17.4% WER). Meiko Fukuda, Hiromitsu Nishizaki, Yurie Iribe, Ryota Nishimura, Norihide Kitaoka |
LREC | 5 |
| 2019 | Small-Footprint Magic Word Detection Method Using Convolutional LSTM Neural Network
Taiki Yamamoto, Ryota Nishimura, Masayuki Misaki, Norihide Kitaoka |
INTERSPEECH | 4 |
| 2016 | Speech Corpus Spoken by Young-old, Old-old and Oldest-old Japanese
Yurie Iribe, Norihide Kitaoka, Shuhei Segawa |
LREC | 2 |
| 2015 | Integration of deep bottleneck features for audio-visual speech recognition
Hiroshi Ninomiya, Norihide Kitaoka, Satoshi Tamura, Yurie Iribe, Kazuya Takeda |
INTERSPEECH | 2 |
| 2015 | Modeling of Physical Characteristics of Speech under StressabstractThis letter presents a method to perform the classification of speech under stress based on physical characteristics. A physical model is proposed to model airflow patterns in the physiological system in order to represent the process of speech production under psychological stress, and physical parameters characterizing airflow variations in the vocal folds, the vocal tract, and laryngeal ventricle are explored. Experimental evaluations show that the physical parameters are effective for the classification of stressed speech. Takatoshi Jitsuhiro, Chiyomi Miyajima, Norihide Kitaoka, Kazuya Takeda |
IEEE Signal Process. Lett. | 4 |
| 2014 | Effect of acoustic and linguistic contexts on human and machine speech recognition
Norihide Kitaoka, Daisuke Enami, Seiichi Nakagawa |
Comput. Speech Lang. | 1 |
| 2013 | Analysis and modeling of entrainment in chorus singingabstractThe dynamics of the contour of the fundamental frequency (F0) of singing voices in a chorus is analyzed from the view point of `entrainment' in singing behavior. The One-Mass-Two-Spring (OMTS) coupled system is used as the mathematical model of the contour of the F0of singing voices that are concurrently singing the same melody. Using this model, the characteristics of the F0dynamics of a voice singing in a chorus are parameterized by the mass, the coefficients of friction, and the spring factors of an OMTS system. It is experimentally confirmed that a steepest decent method can estimate the four model parameters, so that the model can generate the F0contour of a voice singing in a chorus with a less than 44.4 cents of RMS error. Preliminary experiments also show that experienced and novice singers can be correctly identified using the parameters of our model, because their entrainment behaviors are significantly different. Motonari Kawagishi, Shota Kawabuchi, Chiyomi Miyajima, Norihide Kitaoka, Kazuya Takeda |
ICASSP | 4 |
| 2013 | Estimation of vocal tract parameters for the classification of speech under stressabstractIn this work, we propose a method for the classification of speech under stress that is based on a physical model. Using this method, the characteristics of the vocal folds and the vocal tract are taken into consideration, based on the process of speech production. In addition to vocal fold parameters, we estimate parameters of the vocal tract representing cross-sectional areas and vocal tract length, by fitting a two-mass model to real speech. Results show that calculation of vocal tract length for each speaker can improve the accuracy of the estimation of other physical parameters. Analysis is performed under vowel-dependent and vowel-independent conditions, showing that the proposed physical features are effective for the classification of neutral and stressed speech. Takatoshi Jitsuhiro, Chiyomi Miyajima, Norihide Kitaoka, Kazuya Takeda |
ICASSP | 4 |
| 2013 | Classification of speech under stress by modeling the aerodynamics of the laryngeal ventricle
Takatoshi Jitsuhiro, Chiyomi Miyajima, Norihide Kitaoka, Kazuya Takeda |
INTERSPEECH | 4 |
| 2012 | Physical characteristics of vocal folds during speech under stressabstractWe focus on variations in the glottal source of speech production, which is essential for understanding the generation of speech under psychological stress. In this paper, a two-mass vocal fold model is fitted to estimate the stiffness parameters of vocal folds during speech, and the stiffness parameters are then analyzed in order to classify recorded samples into neutral and stressed speech. Mechanisms of vocal folds under stress are derived from the experimental results. We propose using a Muscle Tension Ratio (MTR) to identify speech under stress. Our results show that MTR is more effective than a conventional method of stress measurement. Takatoshi Jitsuhiro, Chiyomi Miyajima, Norihide Kitaoka, Kazuya Takeda |
ICASSP | 4 |
| 2012 | Classification of Stressed Speech Using Physical Parameters Derived from Two-Mass Model
Takatoshi Jitsuhiro, Chiyomi Miyajima, Norihide Kitaoka, Kazuya Takeda |
INTERSPEECH | 4 |
| 2012 | Causal analysis of task completion errors in spoken music retrieval interactions
Sunao Hara, Norihide Kitaoka, Kazuya Takeda |
LREC | 2 |
| 2011 | Robust seed model training for speaker adaptation using pseudo-speaker features generated by inverse CMLLR transformationabstractIn this paper, we propose a novel acoustic model training method which is suitable for speaker adaptation in speech recognition. Our method is based on feature generation from a small amount of speakers' data. For decades, speaker adaptation methods have been widely used. Such adaptation methods need some amount of adaptation data and if the data is not sufficient, speech recognition performance degrade significantly. If the seed models to be adapted to a specific speaker can widely cover more speakers, speaker adaptation can perform robustly. To make such robust seed models, we adopt inverse maximum likelihood linear regression (MLLR) transformation-based feature generation, and then train our seed models using these features. First we obtain MLLR transformation matrices from a limited number of existing speakers. Then we extract the bases of the MLLR transformation matrices using PCA. The distribution of the weight parameters to express the MLLR transformation matrices for the existing speakers is estimated. Next we generate pseudo-speaker MLLR transformations by sampling the weight parameters from the distribution, and apply the inverse of the transformation to the normalized existing speaker features to generate the pseudo-speakers' features. Finally, using these features, we train the acoustic seed models. Using this seed models, we obtained better speaker adaptation results than using simply environmentally adapted models. Arata Itoh, Sunao Hara, Norihide Kitaoka, Kazuya Takeda |
ASRU | 3 |
| 2011 | Driver risk evaluation based on acceleration, deceleration, and steering behaviorabstractWe propose a driver risk evaluation method based on the analysis of driving data captured with drive recorders. To evaluate the acceleration behavior of each driver we plot the maximum acceleration per minute to velocity on a two dimensional plane and approximate the distribution by linear regression. We assume that the higher the y-intercept of the line, the quicker the driver accelerates from a stop, and the higher the x-intercept, the higher the preferred speed of travel. To evaluate deceleration behavior, brake pedal operation patterns are classified into four types, based on how the brake is depressed and released. We evaluate deceleration risk levels based on these four braking pattern categories. Steering behavior is evaluated based on the relationship between the radius of road curvature and road design speed as defined in the road construction ordinance. Some correlation is observed between our evaluation results and those manually scored by risk consultants. Chiyomi Miyajima, Hiroki Ukai, Atsumi Naito, Hideomi Amata, Norihide Kitaoka, Kazuya Takeda |
ICASSP | 5 |
| 2011 | Detection of Task-Incomplete Dialogs Based on Utterance-and-Behavior Tag N-Gram for Spoken Dialog SystemsabstractWe propose a method of detecting “task incomplete” dialogs in spoken dialog systems using N-gram-based dialog models. We used a database created during a field test in which inexperienced users used a client-server music retrieval system with a spoken dialog interface on their own PCs. In this study, the dialog for a music retrieval task consisted of a sequence of user and system tags that related their utterances and behaviors. The dialogs were manually classified into two classes: the dialog either completed the music retrieval task or it didn’t. We then detected dialogs that did not complete the task, using N-gram probability models or a Support Vector Machine with N-gram feature vectors trained using manually classified dialogs. Off-line and on-line detection experiments were conducted on a large amount of real data, and the results show that our proposed method achieved good classification performance. Sunao Hara, Norihide Kitaoka, Kazuya Takeda |
INTERSPEECH | 2 |
| 2011 | Alternative Frequency Scale Cepstral Coefficient for Robust Sound Event Recognition
Yi Ren Leng, Tran Huy Dat, Norihide Kitaoka, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2011 | Analysis of Real-World Driver's FrustrationabstractThis paper investigates a method for estimating a driver's spontaneous frustration in the real world. In line with a specific definition of emotion, the proposed method integrates information about the environment, the driver's emotional state, and the driver's responses in a single model. Driving data are recorded using an instrumented vehicle on which multiple sensors are mounted. While driving, drivers also interact with an automatic speech recognition (ASR) system to retrieve and play music. Using a Bayesian network, we combine knowledge on the driving environment assessed through data annotation, speech recognition errors, the driver's emotional state (frustration), and the driver's responses measured through facial expressions, physiological condition, and gas- and brake-pedal actuation. Experiments are performed with data from 20 drivers. We discuss the relevance of the proposed model and features of frustration estimation. When all of the available information is used, the overall estimation achieves a true positive rate of 80% and a false positive rate of 9% (i.e., the system correctly estimates 80% of the frustration and, when drivers are not frustrated, makes mistakes 9% of the time). Lucas Malta, Chiyomi Miyajima, Norihide Kitaoka, Kazuya Takeda |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2010 | Automatic detection of task-incompleted dialog for spoken dialog system based on dialog act n-gramabstractIn this paper, we propose a method of detecting task- incompleted users for a spoken dialog system using an N-gram- based dialog history model. We collected a large amount of spoken dialog data accompanied by usability evaluation scores by users in real environments. The database was made by a field test in which naive users used a client-server music retrieval system with a spoken dialog interface on their own PCs. An N-gram model was trained from sequences that consist of user dialog acts and/or system dialog acts for two dialog classes, that is, the dialog completed the music retrieval task or the dialog incompleted the task. Then the system detects unknown dialogs that is not completed the task based on the N-gram likelihood. Experiments were conducted on large real data, and the results show that our proposed method achieved good classification performance. When the classifier correctly detected all of the task-incompleted dialogs, our proposed method achieved a false detection rate of 6%. Sunao Hara, Norihide Kitaoka, Kazuya Takeda |
INTERSPEECH | 2 |
| 2010 | Selective gammatone filterbank feature for robust sound event recognition
Yi Ren Leng, Tran Huy Dat, Norihide Kitaoka, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2010 | A browsing and retrieval system for driving dataabstractWith the increased presence and recent advances of drive recorders, rich driving data that include video, vehicle acceleration signals, driver speech, GPS data, and several sensor signals can be continuously recorded and stored. These advances enable researchers to study driving behavior more extensively for traffic safety. However, increasing the variety and the amount of driving data complicates the simultaneous browsing of various data and finding desired data from large databases. In this study, we develop a browsing and retrieval system for driving data that provides a multi-modal data browser, query- and similarity-based retrieval functions, and a fast browsing function that skips redundant scenes. For sharing data with several users, this system can be used via networks from PCs or smartphones, This system uses a time-series active search, which has been successfully used for fast search of audio and video data, as its retrieval function algorithm. In a few seconds, this system can retrieve driving scenes that are similar to an input scene from 80,000 scenes. Retrieval performance was compared in various retrieval conditions by changing the codebook size of the vector quantization for the histogram features and a combination of driving signals. Experimental results showed that more than 97% retrieval performance was achieved for driving behaviors of left/right turns and curves using a combination of such complementary information as steering angles and lateral acceleration. We also compared the proposed method to a conventional image-based retrieval method using subjective similarity scores of driving scenes. Our proposed system retrieved similar scenes with about a 75% retrieval performance that was five points higher than a conventional image-based retrieval method. It is because image-based method is sensitive to changes of image in the area except in the region of interest for driving data retrieval. The fast browsing function also skipped scenes that could not be skipped by an image-based method. Masashi Naito, Chiyomi Miyajima, Takanori Nishino, Norihide Kitaoka, Kazuya Takeda |
Intelligent Vehicles Symposium | 4 |
| 2010 | Estimation Method of User Satisfaction Using N-gram-based Dialog History Model for Spoken Dialog System
Sunao Hara, Norihide Kitaoka, Kazuya Takeda |
LREC | 2 |
| 2009 | Spoken dialog strategy based on understanding graph searchabstractWe regarded information retrieval as a graph search problem and proposed several novel dialog strategies that can recover from misrecognition through a spoken dialog that traverses the graph. To recover from misrecognition without seeking confirmation, our system kept multiple understanding hypotheses at each turn and searched for a globally optimal hypothesis in the graph whose nodes express understanding states across user utterances in a whole dialog. As for a dialog strategy, we introduced a new criterion based on efficiency in information retrieval and consistency with understanding hypotheses to select an appropriate system response. Using such criterion, the system removes the ambiguity so that users do not feel that a response that conflicts with the actual user intent is unnatural. We developed a spoken dialog system using these techniques and showed dialog examples in which misrecognition was naturally corrected. We also showed that our strategy was efficient in terms of the number of turns. Yuji Kinoshita, Chiyomi Miyajima, Norihide Kitaoka, Kazuya Takeda |
ICASSP | 3 |
| 2009 | Feature transformation based on discriminant analysis preserving local structure for speech recognitionabstractTo improve speech recognition performance, a feature transformation based on discriminant analysis has been widely used to reduce redundant dimensions of features. Linear discriminant analysis (LDA) and heteroscedastic discriminant analysis (HDA) are often used for this purpose, and a generalization method for LDA and HDA called power LDA (PLDA) has been proposed. However, these methods may result in unexpected dimensionality reduction for multimodal data. It is important to preserve the local structure of the data in reducing the dimensionality of multimodal data. In this paper we introduce two methods, locality preserving HDA and locality preserving PLDA. We also give an efficient calculation scheme to obtain an optimal projection. Makoto Sakai, Norihide Kitaoka, Kazuya Takeda |
ICASSP | 2 |
| 2009 | Subjective experiments on influence of response timing in spoken dialoguesabstractTo verify the validity of analysis results relating to dialogue rhythm from earlier studies, we produced spoken dialogues based on analysis results relating to response timing and the other spoken dialogues, and performed subjective experiments to investigateparameters suchas the naturalness of the dialogue, the incongruity of the synthesized speech, and the ease of comprehension of the utterances. We used very short task-oriented four-turn dialogues using synthesized speech in Experiment 1, and approx. one-minute free-conversation dialogues in Experiment 2 using natural human speech and synthesized speech. As a result, we were able to show that a natural response timing exists for utterances, and that response timings that conform to the utterance contents are felt to be more natural, thus demonstrating the validity of the analysis results relating to dialogue rhythm. Index Terms: Speech analysis, Speech communication, Interactive systems, User interface human factors, User interfaces Toshihiko Itoh, Norihide Kitaoka, Ryota Nishimura |
INTERSPEECH | 2 |
| 2009 | A multimedia corpus of driving behaviorsabstractIn this paper we present our multimedia corpus of real-world driving data (NUDrive), built with the primary objective of firming foundations for applying digital signal processing technologies in the vehicular environment. NUDrive is a content rich corpus composed of driving, speech, video, and physiological signals. So far, we have collected data from 250 drivers, who drove an instrumented vehicle under very similar conditions. In order to provide a more meaningful description of the situations drivers experience, a comprehensive data annotation protocol is proposed. We also briefly present a multimedia processing system, which uses information from various sources in NUDrive to implement a context-dependent estimation of a driver's spontaneous frustration. Results are encouraging and stress the relevance of content rich driving corpora to driver behavior modeling. Lucas Malta, Akira Ozaki, Chiyomi Miyajima, Norihide Kitaoka, Kazuya Takeda |
MMSP | 4 |
| 2008 | An integrative recognition method for speech and gesturesabstractWe propose an integrative recognition method of speech accompanied with gestures such as pointing. Simultaneously generated speech and pointing complementarily help the recognition of both, and thus the integration of these multiple modalities may improve recognition performance. As an example of such multimodal speech, we selected the explanation of a geometry problem. While the problem was being solved, speech and fingertip movements were recorded with a close-talking microphone and a 3D position sensor. To find the correspondence between utterance and gestures, we propose probability distribution of the time gap between the starting times of an utterance and gestures. We also propose an integrative recognition method using this distribution. We obtained approximately 3-point improvement for both speech and fingertip movement recognition performance with this method. Madoka Miki, Chiyomi Miyajima, Takanori Nishino, Norihide Kitaoka, Kazuya Takeda |
ICMI | 4 |
| 2008 | Class lecture summarization taking into account consecutiveness of important sentencesabstractThis paper presents a novel sentence extraction framework that takes into account the consecutiveness of important sentences using a Support Vector Machine (SVM). Generally, most ex-tractive summarizers do not take context information into ac-count, but do take into account the redundancy over the entire summarization. However, there must exist relationships among the extracted sentences. Actually, we can observe these rela-tionships as consecutiveness among the sentences. We deal with this consecutiveness by using dynamic and difference features to decide if a sentence needs to be extracted or not. Since impor-tant sentences tend to be extracted consecutively, we just used the decision made for the previous sentence as the dynamic fea-ture. We used the differences between the current and previ-ous feature values for the difference feature, since adjacent sen- Yasuhisa Fujii, Kazumasa Yamamoto, Norihide Kitaoka, Seiichi Nakagawa |
INTERSPEECH | 3 |
| 2008 | CENSREC-4: development of evaluation framework for distant-talking speech recognition under reverberant environmentsabstractIn this paper, we newly introduce a collection of databases and evaluation tools called CENSREC-4, which is an evaluation framework for distant-talking speech under hands-free conditions. Distant-talking speech recognition is crucial for a handsfree speech interface. Therefore, we measured room impulse responses to investigate reverberant speech recognition in various environments. The data contained in CENSREC-4 are connected digit utterances, as in CENSREC-1. Two subsets are included in the data: basic data sets and extra data sets. The basic data sets are used for the evaluation environment for the room impulse response-convolved speech data. The extra data sets consist of simulated and recorded data. An evaluation framework is only provided for the basic data sets as evaluation tools. The results of evaluation experiments proved that CENSREC-4 is an effective database for evaluating the new dereverberation method because the traditional dereverberation process had difficulty sufficiently improving the recognition performance. Index Terms: Various environments, Impulse response, Convolution, Real recorded data, Evaluation framework Masato Nakayama, Takanobu Nishiura, Yuki Denda, Norihide Kitaoka, Kazumasa Yamamoto, Takeshi Yamada, Satoru Tsuge, Chiyomi Miyajima, Masakiyo Fujimoto, Tetsuya Takiguchi, Satoshi Tamura, Tetsuji Ogawa, Shigeki Matsuda, Shingo Kuroiwa, Kazuya Takeda, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2008 | Analysis of relationship between impression of human-to-human conversations and prosodic change and its modelingabstractIf a dialog system could respond to a user as naturally as a human, the interaction would be smoother. Imitating human prosodic characteristics of utterances is important in computer-to-human natural interaction. To develop a cooperative/friendly spoken dialog system, we analyzed the correlation between the fundamental frequency’s synchrony tendency, or overlap fre-quency, and subjective measures of “liveliness”, “familiarity”, and “informality ” in human-to-human dialogs. We also mod-eled the properties of these features to realize chat-like conver-sations in our spoken dialog system. Index Terms: spoken dialog system, prosody control, response timing, chat-like conversation Ryota Nishimura, Norihide Kitaoka, Seiichi Nakagawa |
INTERSPEECH | 2 |
| 2008 | Building and combining document and music spaces for music query-by-webpage systemabstractBuilding and combining document and music spaces of songs are discussed for a new music recommendation applica-tion, which uses commonly read texts such as Web log as query input. The most important application of this flexible recom-mendation system is its music query-by-Webpage, from which a song that appropriately matches Webpage is automatically played. The key idea of the proposed system is to train a lin-ear transformation between document and music spaces so that query documents can be mapped onto a music space in which similarities based on acoustic characteristics is represented. The basic system has been trained using 2,650 pairs of song and review texts. Through experimental evaluations, we show the effectiveness of the system, which is three times better than the previous system. Web text as a training corpus and a bigram representation for the document vector are also investigated for the purpose of improving the system, and their effectiveness is also confirmed. Index Terms: Music, information retrieval, music similarity, latent semantic analysis, multimedia databases Ryoei Takahashi, Yasunori Ohishi, Norihide Kitaoka, Kazuya Takeda |
INTERSPEECH | 3 |
| 2008 | Blind dereverberation based on CMN and spectral subtraction by multi-channel LMS algorithmabstractWe proposed a blind dereverberation method based on spectral subtraction by Multi-Channel Least Mean Square (MCLMS) al-gorithm for distant-talking speech recognition in our previous study [1]. In this paper, we discuss the problems of the pro-posed method and present some solutions. In a distant-talking environment, the length of channel impulse response is longer than the short-term spectral analysis window. By treating the late reverberation as additive noise, a noise reduction technique based on spectral subtraction was proposed to estimate power spectrum of the clean speech using power spectra of the dis-torted speech and the unknown impulse responses. To estimate the power spectra of the impulse responses, a Variable Step-Size Unconstrained MCLMS (VSS-UMCLMS) algorithm for iden-tifying the impulse responses in a time domain was extended to a frequency domain. To reduce the effect of the estimation error of channel impulse response, we normalize the early re-verberation by CMN instead of the spectral subtraction used by the estimated impulse response in this paper. Furthermore, our proposed method is combined with a conventional delay-and-sum beamforming. We conducted the experiments on distorted speech signal simulated by convolving multi-channel impulse responses with clean speech. The modified proposed method achieved a relative error reduction rate of 22.7 % from conven-tional CMN and 12.0 % from the original proposed method, re-spectively. By combining the modified proposed method with the beamforming, a furthermore improvement (relative error re-duction rate of 23.3%) was achieved. Index Terms: distant-talking speech recognition, blind dere-verberation, Multi-channel LMS, spectral subtraction, CMN. 1. Longbiao Wang, Seiichi Nakagawa, Norihide Kitaoka |
INTERSPEECH | 3 |
| 2008 | Evaluation Framework for Distant-talking Speech Recognition under Reverberant Environments: newest Part of the CENSREC Series -
Takanobu Nishiura, Masato Nakayama, Yuki Denda, Norihide Kitaoka, Kazumasa Yamamoto, Takeshi Yamada, Satoru Tsuge, Chiyomi Miyajima, Masakiyo Fujimoto, Tetsuya Takiguchi, Satoshi Tamura, Shingo Kuroiwa, Kazuya Takeda, Satoshi Nakamura 0001 |
LREC | 4 |
| 2008 | In-car Speech Data Collection along with Various Multimodal Signals
Akira Ozaki, Sunao Hara, Takashi Kusakawa, Chiyomi Miyajima, Takanori Nishino, Norihide Kitaoka, Katunobu Itou, Kazuya Takeda |
LREC | 6 |
| 2007 | Development of VAD evaluation framework CENSREC-1-C and investigation of relationship between VAD and speech recognition performanceabstractVoice activity detection (VAD) plays an important role in speech processing including speech recognition, speech enhancement, and speech coding in noisy environments. We developed an evaluation framework for VAD in such environments, called corpus and environment for noisy speech recognition 1 concatenated (CENSREC-1-C). This framework consists of noisy continuous digit utterances and evaluation tools for VAD results. By adoptiong two evaluation measures, one for frame-level detection performance and the other for utterance-level detection performance, we provide the evaluation results of a power-based VAD method as a baseline. When using VAD in speech recognizer, the detected speech segments are extended to avoid the loss of speech frames and the pause segments are then absorbed by a pause model. We investigate the balance of an explicit segmentation by VAD and an implicit segmentation by a pause model using an experimental simulation of segment extension and show that a small extension improves speech recognition. Norihide Kitaoka, Kazumasa Yamamoto, Tomohiro Kusamizu, Seiichi Nakagawa, Takeshi Yamada, Satoru Tsuge, Chiyomi Miyajima, Takanobu Nishiura, Masato Nakayama, Yuki Denda, Masakiyo Fujimoto, Tetsuya Takiguchi, Satoshi Tamura, Shingo Kuroiwa, Kazuya Takeda, Satoshi Nakamura 0001 |
ASRU | 1 |
| 2007 | Generalization of Linear Discriminant Analysis used in Segmental Unit Input HMM for Speech RecognitionabstractTo precisely model the time dependency of features is one of the important issues for speech recognition. Segmental unit input HMM with a dimensionality reduction method is widely used to address this issue. Linear discriminant analysis (LDA) and heteroscedastic discriminant analysis (HDA) are classical and popular approaches to reduce dimensionality. However, it is difficult to find one particular criterion suitable for any kind of data set in carrying out dimensionality reduction while preserving discriminative information. In this paper, we propose a new framework which we call power linear discriminant analysis (PLDA). PLDA can describe various criteria including LDA and HDA with one parameter. Experimental results show that the PLDA is more effective than PCA, LDA, and HDA for various data sets. Makoto Sakai, Norihide Kitaoka, Seiichi Nakagawa |
ICASSP (4) | 2 |
| 2007 | Robust Distant Speech Recognition by Combining Position-Dependent CMN with Conventional CMNabstractWe proposed an environmentally robust speech recognition method based on position-dependent cepstral mean normalization (PD-CMN) to compensate for channel distortion depending on speaker position. PDCMN can efficiently compensate for the channel transmission characteristics while it cannot normalize speaker variation because position-dependent cepstral mean does not contain speaker characteristics. Conventional CMN can compensate for the speaker variation while it cannot obtain good recognition performance for short utterances. In this paper, we propose a robust distant speech recognition by combining position-dependent CMN with the conventional CMN to address the above problems. The position-dependent cepstral mean is linearly combined with conventional cepstral mean with following two types of processing. The first method is to use a fixed weighting coefficient over whole test data to obtain the combinational CMN, which is called fixed-weight combinational CMN. The second method is to calculate the output probability of multiple features compensated by a variable weighting coefficient at each frame, and a single decoder using these output probabilities is used to perform speech recognition, which is called variable-weight combinational CMN. We conducted the experiments of our proposed method using small vocabulary (100 words) distant isolated word recognition in a real environment. The proposed variable-weight combinational CMN method achieved a relative error reduction rate of 56.3% from conventional CMN and 22.2% from PDCMN, respectively. Longbiao Wang, Norihide Kitaoka, Seiichi Nakagawa |
ICASSP (4) | 2 |
| 2007 | Statistical segmentation and recognition of fingertip trajectories for a gesture interfaceabstractThis paper presents a virtual push button interface created by drawing a shape or line in the air with a fingertip. As an example of such a gesture-based interface, we developed a four-button interface for entering multi-digit numbers by pushing gestures within an invisible 2x2 button matrix inside a square drawn by the user. Trajectories of fingertip movements entering randomly chosen multi-digit numbers are captured with a 3D position sensor mounted on the the forefinger's tip. We propose a statistical segmentation method for the trajectory of movements and a normalization method that is associated with the direction and size of gestures. The performance of the proposed method is evaluated in HMM-based gesture recognition. The recognition rate of 60.0% was improved to 91.3% after applying the normalization method. Kazuhiro Morimoto, Chiyomi Miyajima, Norihide Kitaoka, Katunobu Itou, Kazuya Takeda |
ICMI | 3 |
| 2007 | Automatic extraction of cue phrases for important sentences in lecture speech and automatic lecture speech summarizationabstractWe automatically extract the summaries of spoken class lectures. This paper presents a novel method for sentence extraction-based automatic speech summarization. We propose a technique that extracts “cue phrases for im-portant sentences (CPs) ” that often appear in important sen-tences. We formulate CP extraction as a labeling problem of word sequences and use Conditional Random Fields (CRF) [1] for labeling. Automatic summarization using CP extraction re-sults as features yields precisions of 0.603 and 0.556 when us-ing manual transcriptions and Automatic Speech Recognition (ASR) results, respectively. Combining the features derived from the CPs and tradi-tional features (including repeated words, words repeated in a slide text, and term frequency (tf), which are surface linguistic information, and speech power and duration, which are prosodic features) [2, 3], we obtained better summarization performance with a κ-value of 0.380, a F-measure of 0.539, and a Rouge-4 of 0.709. Index Terms: automatic speech summarization, sentence ex-traction, speech synthesis Yasuhisa Fujii, Norihide Kitaoka, Seiichi Nakagawa |
INTERSPEECH | 2 |
| 2007 | Prosody change and response timing analysis in spontaneously spoken dialogs and their modeling in a spoken dialog systemabstractIf a dialog system were to respond to a user as naturally as a human, interaction would be smoother. Imitating the human prosodic behavior of utterances is important in computer-human natural conversations. In this paper, to develop a coopera-tive/friendly spoken dialog system, we analyzed the correlations between F0 synchrony tendency or overlap frequency and sub-jective measures: “liveliness, ” “familiarity, ” and “informality” in human-human dialogs. We also modeled the properties of these features and implemented the model on our dialog system that generated the response timing of aizuchi (back-channel), turn-taking based on a decision tree in real time, and dynamical F0 changes to realize chat-like conversations. Index Terms: spoken dialog system, prosody control, response timing Ryota Nishimura, Norihide Kitaoka, Seiichi Nakagawa |
INTERSPEECH | 2 |
| 2007 | Selection of optimal dimensionality reduction method using chernoff bound for segmental unit input HMMabstractTo precisely model the time dependency of features, segmental unit input HMM with a dimensionality reduction method has been widely used for speech recognition. Linear discriminant analysis (LDA) and heteroscedastic discriminant analysis (HDA) are popular approaches to reduce the dimen-sionality. We have proposed another dimensionality reduction method called power linear discriminant analysis (PLDA) to se-lect the best dimensionality reduction method that yields the highest recognition performance. This selection process on the basis of trial and error requires much time to train HMMs and to test the recognition performance for each dimensionality re-duction method. In this paper we propose a performance comparison method without training or testing. We show that the proposed method using the Chernoff bound can rapidly and accurately evaluate the relative recognition performance. Index Terms: speech recognitoin, feature extraction, multidi-mensional signal processing Makoto Sakai, Norihide Kitaoka, Seiichi Nakagawa |
INTERSPEECH | 2 |
| 2007 | Robust distant speaker recognition based on position-dependent CMN by combining speaker-specific GMM with speaker-adapted HMM
Longbiao Wang, Norihide Kitaoka, Seiichi Nakagawa |
Speech Commun. | 2 |
| 2006 | Noisy speech recognition based on selection of multiple noise suppression methods using noise GMMsabstractTo achieve high recognition performance for a wide variety of noise and for a wide range of signal-to-noise ratio, this paper presents integration methods of four noise reduction algorithms: spectral subtraction with smoothing of time direction, temporal domain SVD-based speech enhancement, GMM-based speech estimation and KLT-based comb-filtering. In this paper, we proposed two types of combination methods of noise suppression algorithms: selection of front-end processor and combination of results from multiple recognition processes. Recognition results on the AURORA-2J task showed the effectiveness of our proposed methods. Intex Terms: Noisy speech recongition, noise suppression method selection, AURORA-2J Norihide Kitaoka, Souta Hamaguchi, Seiichi Nakagawa |
INTERSPEECH | 1 |
| 2006 | A spoken Dialog System with Automatic Recovery Mechanism from misrecognitionabstractWe proposed a novel dialog strategy which can recover from mis- recognition through a spoken dialog. To recover from the misrecog- nition without confirmation, our system kept multiple understanding hypotheses at each turn and 'searched' for a globally optimal hypothesis across user's utterances in a whole dialog. As for a dialog strategy, we introduced a new criterion based on 'efficiency for convergence' and 'consistency with understanding hypotheses' to select an appropriate system response. Using such criterion, the system removes the ambiguity without making the user feel unnatural in relation to the response conflicting with actual user intent. We also proposed to adopt the repetition utterance detection to update the understanding hypotheses. We developed a spoken dialog system using these techniques and showed some dialog examples in which misrecognition was naturally corrected. We also showed that our strategy was efficient in terms of the number of turns. Norihide Kitaoka, Hirotoshi Yano, Seiichi Nakagawa |
SLT | 1 |
| 2005 | Multimodal interface for organization name input based on combination of isolated word recognition and continuous base-word recognitionabstractWe investigate a multimodal interface for organization name input to forms in WWW. The user first utters an organization name in an open vocabulary domain to the system. The system recognizes it with a combination method of isolated word recognition and continuous “base-word ” recognition. Word candidates and a base-word lattice obtained by this recognition procedure are displayed on a touch panel. Then the user chooses a sequence or base-words from the candidates by pen touch to construct the desired name and the name is sent to the client. This recognition method performs well in the organization name recognition task, in which a very large vocabulary size and its successive update is needed when using isolated word recognition. Our interface design also elevates users ’ input ability of organization names. 1. Norihide Kitaoka, Hironori Oshikawa, Seiichi Nakagawa |
INTERSPEECH | 1 |
| 2005 | Robust distant speaker recognition based on position dependent cepstral mean normalizationabstractIn a distant environment, channel distortion may drastically degrade speaker recognition performance. In this paper, we propose a robust speaker recognition method based on posi-tion dependent Cepstral Mean Normalization (CMN) to com-pensate the channel distortion depending on the speaker posi-tion. It is shown in [1] that the position dependent CMN is robust for speech recognition in a distant environment. We ex-tend this method to the speaker recognition and show that this method is much effective to speaker recognition. In the train-ing stage, the system measures the transmission characteristics according to the speaker positions from some grid points to the microphone in the room and estimated the compensation pa-rameters a priori. In the recognition stage, the system esti-mates the speaker position and adopts the estimated compen-sation parameters corresponding to the estimated position, and then the system applies the CMN to the speech and performs speaker recognition. In our past study, we proposed a new text-independent speaker recognition method by combining speaker-specific Gaussian Mixture Models (GMMs) with syllable-based HMMs adapted to the speakers by MAP [2]. The robustness of this speaker recognition method for the change of the speaking style in close-talking environment was evaluated in [2]. We in-tegrated this method to the proposed position dependent CMN for distant speaker recognition. Our experiments showed that the proposed method improved the speaker recognition perfor-mance remarkably in a distant environment. 1. Longbiao Wang, Norihide Kitaoka, Seiichi Nakagawa |
INTERSPEECH | 2 |
| 2005 | Robust distant speech recognition based on position dependent CMN using a novel multiple microphone processing techniqueabstractIn a distant environment, channel distortion may drastically degrade speech recognition performances. In this paper, we propose a robust multiple microphone speech processing approach based on position dependent Cepstral Mean Normalization (CMN). In the training stage, the system measures the transmission characteristics according to the speaker positions from some grid points in the room and estimated the compensation parameters a priori. In the recognition stage, the system estimates the speaker position and adopts the estimated compensation parameters corresponding to the estimated position, and then the system applies the CMN to the speech and performs speech recognition for each microphone. Finally, the maximum vote or the maximum summation likelihood of whole channels (that is, multiple microphones) is used to obtain the final result. In our proposed method, we use utterances emitted from a loudspeaker located at various positions to estimate compensation parameters for a convenient sake, and we also compensate the mismatch between the cepstral means of utterances spoken by human and those emitted from the loudspeaker. Our experiments showed that the proposed method improved the performances of speech recognition system in a distant environment efficiently and it could also compensate the mismatch between voices from human and loudspeaker well. Longbiao Wang, Norihide Kitaoka, Seiichi Nakagawa |
INTERSPEECH | 2 |
| 2004 | Robust distant speech recognition based on position dependent CMN
Norihide Kitaoka, Longbiao Wang, Seiichi Nakagawa |
INTERSPEECH | 1 |
| 2004 | Speech interface for name input based on combination of recognition methods using syllable-based n-gram and word dictionary
Hironori Oshikawa, Norihide Kitaoka, Seiichi Nakagawa |
INTERSPEECH | 2 |
| 2003 | Detection and recognition of correction utterance in spontaneously spoken dialogabstractRecently, the performance of speech recognition was drastically improved, and the products with the interface based on speech recognition have been realized. However, when we communicate with computers through a speech interface, misrecognition is inevitable, and it is difficult to recover from it because of the immaturity of the interface. Users try to recover from misrecognition by a repetition of the same content. So, the detection of user’s repetition is helpful for a system to detect its misunderstanding, and to recover from the misrecognition. In this paper, we assume the utterance which includes repetitions a correction and propose a method to detect correction utterances in spontaneously spoken dialog using a word spotting based on DTW (dynamic time warping) and N -best hypotheses overlapping measure. As a result, we achieved recall rate of 92.7% and precision of 89.1%. Moreover, we tried to improve recognition accuracy using the detection. Using the choice of vocabulary and grammar setup based on the detection, we achieved improvement in recognition performance from 42.7% to 50.0% for correction utterance and from 70.5% to 77.9% for non-correction utterance. Norihide Kitaoka, Naoko Kakutani, Seiichi Nakagawa |
INTERSPEECH | 1 |
| 2003 | Comparison of effects of acoustic and language knowledge on spontaneous speech perception/recognition between human and automatic speech recognizer
Norihide Kitaoka, Masahisa Shingu, Seiichi Nakagawa |
INTERSPEECH | 1 |
| 2003 | Generation of natural response timing using decision tree based on prosodic and linguistic information
Masashi Takeuchi, Norihide Kitaoka, Seiichi Nakagawa |
INTERSPEECH | 2 |
| 2003 | Integration of noise reduction algorithms for Aurora2 taskabstractTo achieve high recognition performance for a wide variety of noise and for a wide range of signal-to-noise ratios, this paper presents the integration of four noise reduction algorithms: spectral subtraction with smoothing of time direction, temporal domain SVD-based speech enhancement, GMM-based speech estimation and KLT-based comb-filtering. Recognition results on the Aurora2 task show that the effectiveness of these algorithms and their combinations strongly depends on noise conditions, and excessive noise reduction tends to degrade recognition performance in multicondition training. Takeshi Yamada, Jiro Okada, Kazuya Takeda, Norihide Kitaoka, Masakiyo Fujimoto, Shingo Kuroiwa, Kazumasa Yamamoto, Takanobu Nishiura, Mitsunori Mizumachi, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2002 | Detection and recognition of repaired speech on misrecognized utterances for speech input of car navigation system
Naoko Kakutani, Norihide Kitaoka, Seiichi Nakagawa |
INTERSPEECH | 2 |
| 2002 | Evaluation of spectral subtraction with smoothing of time direction on the Aurora 2 task
Norihide Kitaoka, Seiichi Nakagawa |
INTERSPEECH | 1 |
| 2002 | Speaker independent speech recognition using features based on glottal sound source
Norihide Kitaoka, Daisuke Yamada, Seiichi Nakagawa |
INTERSPEECH | 1 |
| 1996 | Concept-based phrase spotting approach for spontaneous speech understandingabstractIn order to realize robust speech understanding, we present a phrase spotting approach. The use of concept-based phrases as the core unit for understanding is advantageous because the phrase-level constraint realizes wider coverage and stable matching and they are directly mapped to semantic cases. The phrase spotting and the sentence-level parsing are formulated as a progressive search. It realizes an optimal search to spot phrase candidates with Viterbi scoring and A* search to combine the phrase candidates into optimal sentence hypotheses. The approach achieved higher detection rates and robust interpretation of ill-formed utterances. We also examined the effect of the background language model. It is shown that lexical knowledge in the background is vital for spotting and the use of the acoustic score of the filler model is significant for parsing. Tatsuya Kawahara, Norihide Kitaoka, Shuji Doshita |
ICASSP | 2 |
| 1994 | Keyword and phrase spotting with heuristic language model
Tatsuya Kawahara, Toshihiko Munetsugu, Norihide Kitaoka, Shuji Doshita |
ICSLP | 3 |