VLDB 2026 Research / reviewers in the wild / expert
Nobuaki Minematsu
dblp:33/2121
· DBLP profile ↗
199ranked-venue papers
30as first author
23since 2021 · last 2025
0000-0002-8778-9555ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 189 · 29 first-author · 21 since 2021Artificial intelligence and machine learning · 150 · 24 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Benchmarking Prosody Encoding in Discrete Speech TokensabstractRecently, discrete tokens derived from self-supervised learning (SSL) models via k-means clustering have been actively studied as pseudo-text in speech language models and as efficient intermediate representations for various tasks. However, these discrete tokens are typically learned in advance, separately from the training of language models or downstream tasks. As a result, choices related to discretization, such as the SSL model used or the number of clusters, must be made heuristically. In particular, speech language models are expected to understand and generate responses that reflect not only the semantic content but also prosodic features. Yet, there has been limited research on the ability of discrete tokens to capture prosodic information. To address this gap, this study conducts a comprehensive analysis focusing on prosodic encoding based on their sensitivity to the artificially modified prosody, aiming to provide practical guidelines for designing discrete tokens. Kentaro Onda, Satoru Fukayama, Daisuke Saito, Nobuaki Minematsu |
ASRU | 4 |
| 2025 | LangInLab: Augmenting Engineering Lab Instruction with Vision- and Voice-Enabled AI Agents for Language LearningabstractLangInLab integrates vision- and voice-enabled GPT agents into engineering labs to support situated English learning. Structured role-play lets students practice technical English without altering curricula. A preliminary study showed high usability and perceived learning benefits, including improved confidence and scientific English use, indicating the system’s potential for scalable EMI (English-Medium Instruction) support in STEM education. Masanori Shigi, Zackary Rackauckas, Yuka Akiyama, Nobuaki Minematsu |
HAI | 4 |
| 2025 | A Perception-Based L2 Speech Intelligibility Indicator: Leveraging a Rater's Shadowing and Sequence-to-sequence Voice Conversion
Haopeng Geng, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2025 | Discrete Tokens Exhibit Interlanguage Speech Intelligibility Benefit: an Analytical Study Towards Accent-robust ASR Only with Native Speech Data
Kentaro Onda, Keisuke Imoto, Satoru Fukayama, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 5 |
| 2025 | Prosodically Enhanced Foreign Accent Simulation by Discrete Token-based Resynthesis Only with Native Speech Corpora
Kentaro Onda, Keisuke Imoto, Satoru Fukayama, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 5 |
| 2024 | A ChatGPT-based oral Q&A practice system for first-time student participants in international conferences
Mayuko Aiba, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2024 | Exploring Pre-trained Speech Model for Articulatory Feature Extraction in Dysarthric Speech Using ASR
Yuqin Lin, Longbiao Wang, Jianwu Dang 0001, Nobuaki Minematsu |
INTERSPEECH | 4 |
| 2024 | A Pilot Study of GSLM-based Simulation of Foreign Accentuation Only Using Native Speech Corpora
Kentaro Onda, Joonyong Park, Nobuaki Minematsu, Daisuke Saito |
INTERSPEECH | 3 |
| 2024 | Acceleration of Posteriorgram-based DTW by Distilling the Class-to-class Distances Encoded in the Classifier Used to Calculate Posteriors
Haitong Sun, Jaehyun Choi, Nobuaki Minematsu, Daisuke Saito |
INTERSPEECH | 3 |
| 2024 | Analysis and Visualization of Directional Diversity in Listening Fluency of World Englishes Speakers in the Framework of Mutual Shadowing
Yu Tomita, Yingxiang Gao, Nobuaki Minematsu, Noriko Nakanishi, Daisuke Saito |
INTERSPEECH | 3 |
| 2023 | Multiple Acoustic Features Speech Emotion Recognition Using Cross-Attention TransformerabstractSpeech emotion recognition (SER) is a challenging task whose performance heavily relies on suitable affect-salient representations. Recently, transformer has exhibited outstanding qualities in learning relevant representations associated with this task. However, a normal transformer is only able to process the uni-source input, and there is often only one kind of input feature in a transformer-based SER system, which may cause limited knowledge. In this paper, we attempt to use the cross-attention transformer (CAT) to handle bi-source input. We propose a novel SER system to better fuse three types of acoustic features – raw waveform data, spectrogram, and MFCC using CAT. Experiments conducted on the IEMOCAP benchmark dataset have shown that our proposed system can achieve a 73.80% weighted accuracy (WA) and 74.25% unweighted accuracy (UA), which outperforms existing state-of-the-art approaches. Yurun He, Nobuaki Minematsu, Daisuke Saito |
ICASSP | 2 |
| 2023 | Hierarchical Softmax for End-To-End Low-Resource Multilingual Speech RecognitionabstractLow-resource speech recognition has been long-suffering from insufficient training data. In this paper, we propose an approach that leverages neighboring languages to improve low-resource scenario performance, founded on the hypothesis that similar linguistic units in neighboring languages exhibit comparable term frequency distributions, which enables us to construct a Huffman tree for performing multilingual hierarchical Softmax decoding. This hierarchical structure enables cross-lingual knowledge sharing among similar tokens, thereby enhancing low-resource training outcomes. Empirical analyses demonstrate that our method is effective in improving the accuracy and efficiency of low-resource speech recognition. Qianying Liu, Zhuo Gong, Zhengdong Yang, Sheng Li 0010, Chenchen Ding, Nobuaki Minematsu, Hao Huang 0009, Fei Cheng 0002, Chenhui Chu, Sadao Kurohashi |
ICASSP | 7 |
| 2023 | Automatic Prediction of Language Learners' Listenability Using Speech and Text Features Extracted from Listening Drills
Yingxiang Gao, Jaehyun Choi, Nobuaki Minematsu, Noriko Nakanishi, Daisuke Saito |
INTERSPEECH | 3 |
| 2023 | A Unified Framework to Improve Learners' Skills of Perception and Production Based on Speech Shadowing and Overlapping
Nobuaki Minematsu, Noriko Nakanishi, Yingxiang Gao, Haitong Sun |
INTERSPEECH | 1 |
| 2022 | Quantifying Discriminability between NMF BasesabstractDiscriminative nonnegative matrix factorization (DNMF) has been investigated as a promising basis-learning method for monaural source separation. To the best of our knowledge, however, no good and sound discussion has been made on quantitative definition of discriminability and it is difficult to evaluate how discriminative DNMF is actually. This paper introduces a quantitative measure to calculate how discriminative two NMF bases are. From the viewpoint of our measure, we compare three basis-learning methods of plain NMF, DNMF, and minimum-volume (min-vol) NMF. Experimental results of monaural speech separation reveal that min-vol NMF actually learns as discriminative bases as DNMF and achieves the best separation performance. This is probably because min-vol NMF can learn the most compact basis possible that can cover training data. Eisuke Konno, Daisuke Saito, Nobuaki Minematsu |
ICASSP | 3 |
| 2022 | Text-to-speech synthesis using spectral modeling based on non-negative autoencoder
Takeru Gorai, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2022 | Gradual Improvements Observed in Learners' Perception and Production of L2 Sounds Through Continuing Shadowing Practices on a Daily Basis
Takuya Kunihara, Chuanbo Zhu 0001, Nobuaki Minematsu, Noriko Nakanishi |
INTERSPEECH | 3 |
| 2022 | Detection of Learners' Listening Breakdown with Oral Dictation and Its Use to Model Listening Skill Improvement Exclusively Through Shadowing
Takuya Kunihara, Chuanbo Zhu 0001, Daisuke Saito, Nobuaki Minematsu, Noriko Nakanishi |
INTERSPEECH | 4 |
| 2022 | Automatic Prediction of Intelligibility of Words and Phonemes Produced Orally by Japanese Learners of EnglishabstractThe practical goal for language learning is smooth communication with others, and many teachers have a strong focus on measurement of not accentedness but intelligibility, often regarded as correctness of actual understanding. However, automatic prediction of intelligibility has not been well developed especially for smaller units such as words and phonemes. This is mainly because of difficulty of measuring while-listening behaviors of listeners, and thus it was difficult to build an L2 speech corpus of a sufficient size with intelligibility annotation to train a network-based predictor. In this paper, we annotate intelligibility using oral dictation with a small delay, i.e., shadowing, to collect a large enough corpus from two raters with different language backgrounds. Since perceived intelligibility depends on their language background, inter-rater difference should be taken into account. Therefore with this corpus, a multi-rater neural model is built to predict each rater's intelligibility of the individual words and phonemes in L2 speech. Two tasks are examined, i.e., regression of intelligibility scores and classification of a given segment to be intelligible or not. Results show that our model has higher F1 scores than intra-rater agreements, indicating that our model can simulate the two raters accurately well although they have different language background. Chuanbo Zhu 0001, Takuya Kunihara, Daisuke Saito, Nobuaki Minematsu, Noriko Nakanishi |
SLT | 4 |
| 2022 | Voice Conversion Based on Deep Neural Networks for Time-Variant Linear TransformationsabstractThis paper describes a novel framework of voice conversion to improve the conversion performance against the amount of training data. In voice conversion, deep neural networks are used as conversion models that map source to target features. In this framework, it generally needs a larger amount of training data and bigger models to build more accurate conversion models. This condition, however, will reduce the usability of voice conversion. In this paper, in order to improve the conversion performance versus the amount of training data, a top-down knowledge is introduced into models as prior. We expect that we can take advantage of top-down knowledge we have instead of preparing a large amount of data. In the proposed method, the conversion process of features is restricted to time-variant linear transformation on cepstral space. It explicitly utilizes an attribute of voice conversion i.e. homo-domain mapping, which is not common in automatic speech recognition or text-to-speech synthesis. In other words, in VC, the input and output are on the same feature domain. In addition, it also makes it possible to explicitly consider the physical difference between speakers such as the difference of vocal tract length. The assumption of the homo-domain mapping is related to conversion methods based on spectral differentials, and then the relation is discussed in the paper. Experiments demonstrate the effectiveness of our proposal and the way that the constraint of linear transformation works is investigated. Gaku Kotani, Daisuke Saito, Nobuaki Minematsu |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Multi-Granularity Annotation of Instantaneous Intelligibility of Learners' Utterances Based on Shadowing TechniquesabstractThe practical goal of pronunciation training is to acquire an intelligible enough pronunciation, not a native-like pronunciation. In our studies [1], [2], we proposed a method that can annotate instantaneous intelligibility of a given L2 English utterance by monitoring listeners' listening behaviors. The listeners were asked to shadow the L2 utterance and the degree of being inarticulate in shadowing was automatically quantified to be used as scores of the instantaneous intelligibility, which were shown to be highly correlated with subjective intelligibility scores. In the present paper, we make an objective assessment of the proposed method, where the automatic scores are compared with those calculated objectively based on manual transcripts of the shadowings. Experiments show that the former scores have such a high correlation as 0.935 with the latter scores, which is higher than correlation obtained using another kind of automatic scores calculated by transcribing the shadowings with ASR. Further, since intelligibility is sometimes discussed in pronunciation training in such smaller units as phonemes, our method is experimentally applied to intelligibility annotation with multiple granularity. Experiments show a high validity of our method to calculate instantaneous intelligibility in units of word, syllable, phoneme, and frame. Chuanbo Zhu 0001, Ryo Hakoda, Daisuke Saito, Nobuaki Minematsu, Noriko Nakanishi, Tazuko Nishimura |
ASRU | 4 |
| 2021 | Lexical Density Analysis of Word Productions in Japanese English Using Acoustic Word Embeddings
Shintaro Ando, Nobuaki Minematsu, Daisuke Saito |
Interspeech | 2 |
| 2021 | Optimized Prediction of Fluency of L2 English Based on Interpretable Network Using Quantity of Phonation and Quality of PronunciationabstractThis paper presents results of a joint project between an engineering team of a university and an educational team of another to develop an online fluency assessment system for Japanese learners of English. A picture description corpus of English spoken by 90 learners and 10 native speakers was used, where fluency was rated by other 10 native raters for each speaker manually. The assessment system was built to predict the averaged manual scores. For system development, a special focus was put on two separate purposes. The assessment system was trained in such an analytical way that teachers can know and discuss which speech features contribute more to fluency prediction, and in such a technical way that teachers' knowledge can be involved for training the system, which can be further optimized using an interpretable network. Experiments showed that quality-of-pronunciation features are much more helpful than quantity-of-phonation features, and the optimized system reached an extremely high correlation of 0.956 with the averaged manual scores, which is higher than the maximum of inter-rater correlations (0.910). Ayano Yasukagawa, Daisuke Saito, Nobuaki Minematsu, Kazuya Saito |
SLT | 4 |
| 2020 | Converting Written Language to Spoken Language with Neural Machine Translation for Language ModelingabstractWhen building a language model (LM) for spontaneous speech, the ideal situation is to have a large amount of spoken, in-domain training data. Having such abundant data, however, is not realistic. We address this problem by generating texts in spoken language from those in written language by using a neural machine translation (NMT) model. We collected faithful transcripts of fully spontaneous speech and corresponding written versions and used them as a parallel corpus to train the NMT model. We used top-k random sampling, which generates a large variety of texts of higher quality as compared to other generation methods for NMT. We indicate that the NMT model is capable of converting written texts in a certain domain to spoken texts, and that the converted texts are effective for training LMs. Our experimental results show significant improvement of speech recognition accuracy with the LMs. Shintaro Ando, Masayuki Suzuki, Nobuyasu Itoh, Gakuto Kurata, Nobuaki Minematsu |
ICASSP | 5 |
| 2020 | Shadowability Annotation with Fine Granularity on L2 Utterances and its Improvement with Native Listeners' Script-Shadowing
Zhenchao Lin, Ryo Takashima, Daisuke Saito, Nobuaki Minematsu, Noriko Nakanishi |
INTERSPEECH | 4 |
| 2020 | Discriminative Method to Extract Coarse Prosodic Structure and its Application for Statistical Phrase/Accent Command Estimation
Yuma Shirahata, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2019 | Analysis of Native Listeners' Facial Microexpressions While Shadowing Non-Native Speech - Potential of Shadowers' Facial Expressions for Comprehensibility Prediction
Tasavat Trisitichoke, Shintaro Ando, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 4 |
| 2019 | Many-to-Many and Completely Parallel-Data-Free Voice Conversion Based on Eigenspace DNNabstractMedia conversion of image, text, speech, etc., generally requires a large amount of parallel data for training a conversion model. Recently, methods for training the model using no or a small amount of parallel data draw researchers' attention. In many-to-many voice conversion, since it is often hard to collect parallel data from every pair of speakers, the conversion models requiring no parallel data are desired. Conventional many-to-many voice conversion models required a large amount of prestored parallel data to acquire prior knowledge of the entire speaker space. Then, a specific model from an arbitrary speaker to another can be realized by adapting a few model parameters. Although these conversion models certainly do not use parallel data in an adaptation step, they still use parallel data for prior training. In this study, we aim at realizing completely parallel-data-free and many-to-many voice conversion. The proposed method uses both Eigenvoice Gaussian mixture models (EVGMM) and Deep neural network (DNN). EVGMM is a many-to-many conversion model that constructs the entire speaker space (called eigenspace) by analyzing mean vectors of Gaussian mixture models and it is used in our method to decompose training speakers' features into their eigenspace components. By using the speaker features and the obtained components as pseudo parallel data, multiple DNNs are trained to realize conversion between them. With these DNNs, features of any target speaker can be represented by a weighted sum of the components. It should be noted that all the processes of our proposal do not require any parallel data. A key technique is to estimate covariance terms of EVGMM with no parallel data. Experiments indicate that individuality scores of the proposed method using no parallel data are comparable enough to those of a baseline system trained with parallel data. Tetsuya Hashimoto, Daisuke Saito, Nobuaki Minematsu |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | A Study of Objective Measurement of Comprehensibility through Native Speakers' Shadowing of Learners' Utterances
Yusuke Inoue 0004, Suguru Kabashima, Daisuke Saito, Nobuaki Minematsu, Kumi Kanamura, Yutaka Yamauchi |
INTERSPEECH | 4 |
| 2018 | A Comparative Study of Statistical Conversion of Face to Voice Based on Their Subjective Impressions
Yasuhito Ohsugi, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2018 | DNN-Based Scoring of Language Learners' Proficiency Using Learners' Shadowings and Native Listeners' Responsive ShadowingsabstractThis paper investigates DNN-based scoring techniques when they are applied to two tasks related to foreign language education. One is a conventional task, which attempts to predict a language learner's overall proficiency of oral communication. For this aim, learners' shadowing utterances are assessed automatically. The other is a very new and novel task, which attempts to predict intelligibility or comprehensibility of a learner's pronunciation. In this task, native listeners' responsive shadowings are assessed. For both the tasks, similar technical frameworks are tested, where DNN-based phoneme posteriors, posteriogram-based DTW scores, ASR-based accuracies, shadowing latencies, etc are used to train regression models, which aim to predict manually rated scores. Experiments show that, in both the tasks, the correlation between the DNN-based predicted scores and the averaged human scores is higher than or at least comparable to the averaged correlation between the scores of human raters. This fact clearly indicates that our proposed automatic rating module can be introduced to language education as another human rater. Suguru Kabashima, Yusuke Inoue 0004, Daisuke Saito, Nobuaki Minematsu |
SLT | 4 |
| 2017 | Parallel-Data-Free Many-to-Many Voice Conversion Based on DNN Integrated with Eigenspace Using a Non-Parallel Speech Corpus
Tetsuya Hashimoto, Hidetsugu Uchida, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 4 |
| 2017 | Use of Global and Acoustic Features Associated with Contextual Factors to Adapt Language Models for Spontaneous Speech Recognition
Shohei Toyama, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2017 | Acoustic-to-Articulatory Mapping Based on Mixture of Probabilistic Canonical Correlation Analysis
Hidetsugu Uchida, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2017 | Automatic Scoring of Shadowing Speech Based on DNN Posteriors and Their DTW
Junwei Yue, Fumiya Shiozawa, Shohei Toyama, Yutaka Yamauchi, Kayoko Ito, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 7 |
| 2016 | Divergence estimation based on deep neural networks and its use for language identificationabstractIn this paper, we propose a method to estimate statistical divergence between probability distributions by a DNN-based discriminative approach and its use for language identification tasks. Since statistical divergence is generally defined as a functional of two probability density functions, these density functions are usually represented in a parametric form. Then, if a mismatch exists between the assumed distribution and its true one, the obtained divergence becomes erroneous. In our proposed method, by using Bayes' theorem, the statistical divergence is estimated by using DNN as discriminative estimation model. In our method, the divergence between two distributions is able to be estimated without assuming a specific form for these distributions. When the amount of data available for estimation is small, however, it becomes intractable to calculate the integral of the divergence function over all the feature space and to train neural networks. To mitigate this problem, two solutions are introduced; a model adaptation method for DNN and a sampling approach for integration. We apply this approach to language identification tasks, where the obtained divergences are used to extract a speech structure. Experimental results show that our approach can improve the performance of language identification by 10.85% relative compared to the conventional approach based on i-vector. Yosuke Kashiwagi, Congying Zhang, Daisuke Saito, Nobuaki Minematsu |
ICASSP | 4 |
| 2016 | Automatic Assessment and Error Detection of Shadowing Speech: Case of English Spoken by Japanese Learners
Shuju Shi, Yosuke Kashiwagi, Shohei Toyama, Junwei Yue, Yutaka Yamauchi, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 7 |
| 2016 | Prediction of the Articulatory Movements of Unseen Phonemes of a Speaker Using the Speech Structure of Another Speaker
Hidetsugu Uchida, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2016 | Voice Conversion Based on Matrix Variate Gaussian Mixture Model Using Multiple Frame Features
Hidetsugu Uchida, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 4 |
| 2016 | Speaker Representations for Speaker Adaptation in Multiple Speakers' BLSTM-RNN-Based Speech Synthesis
Yi Zhao 0006, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2016 | Improved prediction of the accent gap between speakers of English for individual-based clustering of World EnglishesabstractThe term of “World Englishes” describes the current state of English and one of their main characteristics is a large diversity of pronunciation, called accents. In our previous studies, we developed several techniques to realize effective clustering and visualization of the diversity. For this aim, the accent gap between two speakers has to be quantified independently of extra-linguistic factors such as age and gender. To realize this, a unique representation of speech, called speech structure, which is theoretically invariant against these factors, was applied to represent pronunciation. In the current study, by controlling the degree of invariance, we attempt to improve accent gap prediction. Two techniques are tested: DNN-based model-free estimation of divergence and multi-stream speech structures. In the former, instead of estimating separability between two speech events based on some model assumptions, DNN-based class posteriors are utilized for estimation. In the latter, by deriving one speech structure for each sub-space of acoustic features, constrained invariance is realized. Our proposals are tested in terms of the correlation between reference accent gaps and the predicted and quantified gaps. Experiments show that the correlation is improved from 0.718 to 0.730. Fumiya Shiozawa, Daisuke Saito, Nobuaki Minematsu |
SLT | 3 |
| 2016 | Phonetisaurus: Exploring grapheme-to-phoneme conversion with joint n-gram models in the WFST frameworkabstractAbstract This paper provides an analysis of several practical issues related to the theory and implementation of Grapheme-to-Phoneme (G2P) conversion systems utilizing the Weighted Finite-State Transducer paradigm. The paper addresses issues related to system accuracy, training time and practical implementation. The focus is on joint n-gram models which have proven to provide an excellent trade-off between system accuracy and training complexity. The paper argues in favor of simple, productive approaches to G2P, which favor a balance between training time, accuracy and model complexity. The paper also introduces the first instance of using joint sequence RnnLMs directly for G2P conversion, and achieves new state-of-the-art performance via ensemble methods combining RnnLMs and n-gram based models. In addition to detailed descriptions of the approach, minor yet novel implementation solutions, and experimental results, the paper introducesPhonetisaurus, a fully-functional, flexible, open-source, BSD-licensed G2P conversion toolkit, which leverages the OpenFst library. The work is intended to be accessible to a broad range of readers. Josef R. Novak, Nobuaki Minematsu, Keikichi Hirose |
Nat. Lang. Eng. | 2 |
| 2015 | Statistical acoustic-to-articulatory mapping unified with speaker normalization based on voice conversionabstractThis paper proposes a model of speaker-normalized acoustic-toarticulatory mapping using statistical voice conversion. A mapping function from acoustic parameters to articulatory parameters is usually developed with a single speaker’s parallel data. Hence the constructed mapping model can work appropriately only for this specific speaker, and applying this model to other speakers degrades the performance of acoustic-to-articulatory mapping. In this paper, two models of speaker conversion and acoustic-to-articulatory mapping are implemented using Gaussian Mixture Models (GMM), and by integrating these two models, we propose two methods of speaker-normalized acoustic-to-articulatory mapping. One is concatenating these models sequentially, and the other integrates the two models into a unified model, where acoustic parameters of a speaker can be converted directly to articulatory parameters of another speaker. Experiments show that both methods can improve the mapping accuracy and that the latter method works better than the former method. Especially in the case of velar stop consonants, the mapping accuracy is higher by 0.6 mm. Index Terms: acoustic-to-articulatory mapping, Gaussian mixture model, voice conversion, speaker normalization Hidetsugu Uchida, Daisuke Saito, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 3 |
| 2015 | Automatic recognition of Japanese vowel length accounting for speaking rate and motivated by perception analysis
Greg Short, Keikichi Hirose, Mariko Kondo, Nobuaki Minematsu |
Speech Commun. | 4 |
| 2015 | Discriminative re-ranking for automatic speech recognition by leveraging invariant structures
Masayuki Suzuki, Gakuto Kurata, Masafumi Nishimura, Nobuaki Minematsu |
Speech Commun. | 4 |
| 2014 | Improved and robust prediction of pronunciation distance for individual-basis clustering of World Englishes pronunciationabstractEnglish is the only language available for global communication and is used by approximately 1.5 billions of speakers. It is also known to have a large diversity of pronunciation due to the influence of speakers' mother tongue, called accents. Our project aims at creating a global and individual-basis map of English pronunciations to be used in teaching and learning World Englishes (WE) as well as research studies of WE [1, 2]. Creating the map mathematically requires a distance matrix in terms of pronunciation differences among all the speakers considered, and technically requires a method of predicting the pronunciation distance between any pair of the speakers only by using their speech samples. In our previous study [3], we combined invariant pronunciation structure analysis [4, 5, 6, 7] and Support Vector Regression (SVR) to predict the inter-speaker pronunciation distances. In this paper, several techniques are introduced and examined whether they can increase accuracy and robustness of prediction. Experiments show that the correlation between IPA-based reference distances and the predicted distances is increased from 0.805 to 0.903, which is over the correlation of 0.829 that is obtained by using the phoneme-based ground truth distances. Shun Kasahara, S. Kitahara, Nobuaki Minematsu, Han-Ping Shen, Takehiko Makino, Daisuke Saito, K. Hiorse |
ICASSP | 3 |
| 2014 | Semi-supervised noise dictionary adaptation for exemplar-based noise robust speech recognitionabstractThe exemplar-based approaches, which model signals as a sparse linear combination of exemplars of signals, are proved to have state-of-the-art performance in noise robust ASR, especially on low SNRs. However, since both the speech exemplars and noise exemplars are built from training data and are fixed throughout the process of enhancing speech features, the conventional approach is especially weak for unknown types of noise. Therefore, in this paper, we propose a semi-supervised approach which automatically adapt noise exemplars to the target noise, while keeping the speech exemplars fixed. Continuous digits recognition experiments show that this approach is much more robust for unknown noise. The recognition errors are reduced by 36.2%. Yi Luan, Daisuke Saito, Yosuke Kashiwagi, Nobuaki Minematsu, Keikichi Hirose |
ICASSP | 4 |
| 2014 | Application of matrix variate Gaussian mixture model to statistical voice conversionabstractThis paper describes a novel approach to construct a mapping function between a given speaker pair using probability density functions (PDF) of matrix variate. In voice conversion studies, two important functions should be realized: 1) precise modeling of both the source and target feature spaces, and 2) construction of a proper transform function between these spaces. Voice conversion based on Gaussian mixture model (GMM) is the de facto standard because of their flexibility and easiness in handling. In GMM-based approaches, a joint vector space of the source and target is first constructed, and the joint PDF of the two vectors is modeled as GMM in the joint vector space. The joint vector approach mainly focuses on precise modeling of the ‘joint’ feature space, and does not always construct a proper transform between two feature spaces. In contrast, the proposed method constructs the joint PDF as GMM in a matrix variate space whose row and column respectively correspond to the two functions, and it has potential to precisely model both the characteristics of the feature spaces and the relation between the source and target spaces. Daisuke Saito, Hidenobu Doi, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 3 |
| 2013 | Discriminative piecewise linear transformation based on deep learning for noise robust automatic speech recognitionabstractIn this paper, we propose the use of deep neural networks to expand conventional methods of statistical feature enhancement based on piecewise linear transformation. Stereo-based piecewise linear compensation for environments (SPLICE), which is a powerful statistical approach for feature enhancement, models the probabilistic distribution of input noisy features as a mixture of Gaussians. However, soft assignment of an input vector to divided regions is sometimes done inadequately and the vector comes to go through inadequate conversion. Especially when conversion has to be linear, the conversion performance will be easily degraded. Feature enhancement using neural networks is another powerful approach which can directly model a non-linear relationship between noisy and clean feature spaces. In this case, however, it tends to suffer from over-fitting problems. In this paper, we attempt to mitigate this problem by reducing the number of model parameters to estimate. Our neural network is trained whose output layer is associated with the states in the clean feature space, not in the noisy feature space. This strategy makes the size of the output layer independent of the kind of a given noisy environment. Firstly, we characterize the distribution of clean features as a Gaussian mixture model and then, by using deep neural networks, estimate discriminatively the state in the clean space that an input noisy feature corresponds to. Experimental evaluations using the Aurora 2 dataset demonstrate that our proposed method has the best performance compared to conventional methods. Yosuke Kashiwagi, Daisuke Saito, Nobuaki Minematsu, Keikichi Hirose |
ASRU | 3 |
| 2013 | Automatic pronunciation clustering using a World English archive and pronunciation structure analysisabstractEnglish is the only language available for global communication. Due to the influence of speakers' mother tongue, however, those from different regions inevitably have different accents in their pronunciation of English. The ultimate goal of our project is creating a global pronunciation map of World Englishes on an individual basis, for speakers to use to locate similar English pronunciations. If the speaker is a learner, he can also know how his pronunciation compares to other varieties. Creating the map mathematically requires a matrix of pronunciation distances among all the speakers considered. This paper investigates invariant pronunciation structure analysis and Support Vector Regression (SVR) to predict the inter-speaker pronunciation distances. In experiments, the Speech Accent Archive (SAA), which contains speech data of worldwide accented English, is used as training and testing samples. IPA narrow transcriptions in the archive are used to prepare reference pronunciation distances, which are then predicted based on structural analysis and SVR, not with IPA transcriptions. Correlation between the reference distances and the predicted distances is calculated. Experimental results show very promising results and our proposed method outperforms by far a baseline system developed using an HMM-based phoneme recognizer. Han-Ping Shen, Nobuaki Minematsu, Takehiko Makino, Steven H. Weinberger, Teeraphon Pongkittiphan, Chung-Hsien Wu 0001 |
ASRU | 2 |
| 2013 | Improved estimation of femininity using GMM supervectors and SVR for voice therapy of Gender Identity Disorder ClientsabstractThis paper proposes a new method of estimating perceptual femininity (PF) of an input utterance using Gaussian Mixture Model (GMM) supervectors and support vector regression (SVR). The method is used to develop a femininity estimation tool, which is introduced to voice therapy of Gender Identity Disorder (GID) clients, especially MtF (Male to Female) transsexuals. In our previous study [1], we developed a PF estimator, where a male GMM and a female GMM of spectral features and those of pitch features were built and their likelihood scores of an input utterance were combined by linear regression to estimate PF. In this work, inspired by recent speaker recognition models [2], we replace the four likelihood scores from the four GMMs with supervectors composed by a spectral GMM and a pitch GMM estimated from an input utterance. Further, instead of simple linear regression, we introduce SVR, which is discriminative linear regression. Experiments using an MtF speech corpus show that the proposed method improves correlation between human and machine scores of PF and also reduces squared prediction error. Chengshuo Wang, Masayuki Suzuki, Nobuaki Minematsu, Kyoko Sakuraba, Keikichi Hirose |
ICASSP | 3 |
| 2013 | Artificial bandwidth extension based on regularized piecewise linear mapping with discriminative region weighting and long-Span featuresabstractArtificial Bandwidth Extension (ABE) has been introduced to improve perceived speech quality and intelligibility of narrowband telephone speech. Most of the existing algorithms divided ABE into 2 sub-problems, namely extension of the excitation signal and that of the spectral envelope. In this paper, we propose a new method for spectral envelope extension based on REgularized piecewise linear mapping with DIscriminative region weighting And Long-span features (REDIAL). REDIAL is a revised version of SPLICE, a well-known method for speech enhancement. In REDIAL, however, discriminative model is introduced for space division step of the original SPLICE. The proposed REDIAL-based method approximates non-linear transformation from narrowband features to their wideband counterpart by a summation of piecewise linear transformations. The proposed method was compared with the widely used GMM-based method, through objective and subjective evaluations in both speaker-dependent and speaker-independent conditions. Both evaluations showed that the proposed method significantly outperforms the conventional GMM-based method. Nguyen Duc Duy, Masayuki Suzuki, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 3 |
| 2013 | A free online accent and intonation dictionary for teachers and learners of Japanese
Hiroko Hirano, Ibuki Nakamura, Nobuaki Minematsu, Masayuki Suzuki, Chieko Nakagawa, Noriko Nakamura, Yukinori Tagawa, Keikichi Hirose, Hiroya Hashimoto |
INTERSPEECH | 3 |
| 2013 | Generation of fundamental frequency contours for Thai speech synthesis using tone nucleus modelabstractUTokyo Repositoryは本学で生産されたさまざまな学術成果を電子的形態で集中的に蓄積・保存し、世界に発信することを目的としたインターネット上の発信拠点です。 The UTokyo Repository is the system to store and provide digital resources created by members of the University of Tokyo. Its main purpose is to develop digital collections, make them available online, and preserve them for long-term access. Oraphan Krityakien, Keikichi Hirose, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2013 | Development of a web framework for teaching and learning Japanese prosody: OJAD (online Japanese accent dictionary)abstractThis paper introduces the first online and free framework for teaching and learning Japanese prosody including word accent and phrase intonation. This framework is called OJAD (Online Japanese Accent Dictionary) [1] and it provides three functions. 1) Visual, auditory, systematic, and comprehensive illustration of patterns of accent change (accent sandhi) of verbs and adjectives. Here only the changes caused by twelve kinds of fundamental conjugation are focused upon. 2) Visual illustration of the accent pattern of a given verbal expression, which is a combination of a verb and its postpositional auxiliary words. 3) Visual illustration of the pitch pattern of an any given sentence and the expected positions of accent nuclei in the sentence. The third function is implemented by using an accent change prediction module that we developed for Japanese text-to-speech (TTS) synthesizers [2, 3]. Experiments show that accent nucleus assignment to given texts by the proposed framework is much more accurate than that by native speakers. Subjective assessment and objective assessment by teachers and learners show very high pedagogical effectiveness of the framework. Index Terms: language education, Japanese prosody, accent sandhi, OJAD, speech synthesis, assessment experiments Ibuki Nakamura, Nobuaki Minematsu, Masayuki Suzuki, Hiroko Hirano, Chieko Nakagawa, Noriko Nakamura, Yukinori Tagawa, Keikichi Hirose, Hiroya Hashimoto |
INTERSPEECH | 2 |
| 2013 | Failure transitions for joint n-gram models and G2p conversionabstractThis work investigates two related issues in the area of WFST-based G2P conversion. The first is the impact that the approach utilized to convert a target word to an equivalent finite-state ma-chine has on downstream decoding efficiency. The second issue considered is the impact that the approach utilized to represent the joint n-gram model via the WFST framework has on the speed and accuracy of the system. In the latter case two novel algorithms are proposed, which extend the work from [1] to enable the use of failure-transitions with joint n-gram models. All solutions presented in this work are available as part of the open-source, BSD-licensed Phonetisaurus G2P toolkit [2]. Index Terms: G2P, WFST, model conversion Josef R. Novak, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 2 |
| 2013 | Unsupervised optimal phoneme segmentation: theory and experimental evaluationabstractAutomatic phoneme segmentation of a speech sequence is a basic problem in speech engineering. This study investigates unsupervised phoneme segmentation without using prior information on linguistic contents and acoustic models of an input sequence. The authors formulate the unsupervised segmentation as an optimal problem by means of maximum likelihood, and show that the optimal segmentation corresponds to minimising the coding length of the input sequence. Under different assumptions, five different objective functions are developed, namely log determinant, rate distortion (RD), Bayesian log determinant, Mahalanobis distance and Euclidean distance objectives. The authors prove that the optimal segmentations have the transformation‐invariant properties, introduce a time‐constrained agglomerative clustering algorithm to find the optimal segmentations, and propose an efficient implementation of the algorithm by using integration functions. The experiments are carried out on the TIMIT database to compare the above five objective functions. The results show that RD achieves the best performance, and the proposed method outperforms the previous unsupervised segmentation methods. Yu Qiao 0001, Dean Luo, Nobuaki Minematsu |
IET Signal Process. | 3 |
| 2013 | Japanese lexical accent recognition for a CALL system by deriving classification equations with perceptual experiments
Greg Short, Keikichi Hirose, Nobuaki Minematsu |
Speech Commun. | 3 |
| 2013 | Feature Enhancement With Joint Use of Consecutive Corrupted and Noise Feature Vectors With Discriminative Region WeightingabstractThis paper proposes a feature enhancement method that can achieve high speech recognition performance in a variety of noise environments with feasible computational cost. As the well-known Stereo-based Piecewise Linear Compensation for Environments (SPLICE) algorithm, the proposed method learns piecewise linear transformation to map corrupted feature vectors to the corresponding clean features, which enables efficient operation. To make the feature enhancement process adaptive to changes in noise, the piecewise linear transformation is performed by using a subspace of the joint space of corrupted and noise feature vectors, where the subspace is chosen such that classes (i.e., Gaussian mixture components) of underlying clean feature vectors can be best predicted. In addition, we propose utilizing temporally adjacent frames of corrupted and noise features in order to leverage dynamic characteristics of feature vectors. To prevent overfitting caused by the high dimensionality of the extended feature vectors covering the neighboring frames, we introduce regularized weighted minimum mean square error criterion. The proposed method achieved relative improvements of 34.2% and 22.2% over SPLICE under the clean and multi-style conditions, respectively, on the Aurora 2 task. Masayuki Suzuki, Takuya Yoshioka, Shinji Watanabe 0001, Nobuaki Minematsu, Keikichi Hirose |
IEEE Trans. Speech Audio Process. | 4 |
| 2012 | Unseen noise robust speech recognition using adaptive piecewise linear transformationabstractSPLICE is one of the speech enhancement methods based on feature conversion, which shows a high performance with a relatively small amount of calculation. After modeling noisy speech features as GMM, conversion functions are obtained for individual GMM components. The original SPLICE estimates clean feature vectors as a weighted summation of the converted versions of input vectors. Since the conversion functions are determined and fixed only by using training data, the effectiveness of the original SPLICE will be lower in the case of unseen noisy environments. In this paper, we propose a novel method to adapt the conversion functions to work well in unseen environments. First, to realize adaptive conversion functions, we characterize those functions using their super vectors. Then, we conduct PCA on the super vectors to reduce the number of parameters to be adapted. By representing the super vectors through their PCA-based base functions and weights, we implement an efficient adaptation method of conversion functions, which we call Eigen-SPLICE here after. Evaluation experiments show that Eigen-SPLICE has reduced word error rate by 21.0% relative to the conventional SPLICE, and by 24.1% relative to EMS SPLICE in the test set B of the AURORA-2 task. Keigo Chijiiwa, Masayuki Suzuki, Nobuaki Minematsu, Keikichi Hirose |
ICASSP | 3 |
| 2012 | MFCC enhancement using joint corrupted and noise feature space for highly non-stationary noise environmentsabstractOne of the most effective approaches to noise robust speech recognition is to remove the noise effect directly from corrupted MFCC vectors. However, VTS enhancement, which is a typical method for performing MFCC enhancement, provides limited improvement when the noise is highly non-stationary. This is because the VTS enhancement method cannot use a time-varying noise model to keep the computational cost at an acceptable level. This paper proposes a method that can enhance MFCC vectors and their dynamic parameters by using noise estimates that change on a frame-by-frame basis at a practical computational cost. The proposed method employs stereo data-based feature mapping like the well known SPLICE algorithm. The novelty of the proposed method lies in that it uses the joint space spanned by a concatenated vector of corrupted and noise features. It is also proposed to use linear discriminant analysis to effectively reduce the dimensionality of the joint space. The proposed method achieves 19.1% and 8.3% relative error reduction from the SPLICE and noise-mean normalized SPLICE algorithms, respectively. Masayuki Suzuki, Takuya Yoshioka, Shinji Watanabe 0001, Nobuaki Minematsu, Keikichi Hirose |
ICASSP | 4 |
| 2012 | Improved Automatic Extraction of Generation Process Model Commands and Its use for Generating Fundamental Frequency Contours for Training HMM-based Speech Synthesis
Hiroya Hashimoto, Keikichi Hirose, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2012 | Improved Prediction of Japanese Word Accent Sandhi Using CRFabstractIn Japanese, every content word has its own mora-based H/L pitch pattern when it is uttered in isolation, called accent type. When reading out a written sentence, however, this lexical H/L pattern is often changed according to the context, known as word accent sandhi. In our previous work, an accent sandhi predictor was developed using CRF [1], and in this paper, the predictor is improved through feature engineering especially fo-cusing on phrases including numerals and those including loan-words. This is because our previous work showed that the pre-diction performance was relatively low for those phrases. To optimize the features used for CRF, it is critical to take into ac-count the mechanism of word accent sandhi. We review linguis-tic and technical literature that attempted to characterize accent sandhi in the phrases including numerals and loanwords and, by reflecting these characteristics, the features are re-designed. Experiments show that the proposed predictor improved the per-formance relatively by 37 % and 41%, respectively. Index Terms: word accent sandhi, accent nucleus, text-to-speech, Japanese education, rule-based, corpus-based, CRF Nobuaki Minematsu, Shumpei Kobayashi, Shinya Shimizu, Keikichi Hirose |
INTERSPEECH | 1 |
| 2012 | Dynamic Grammars with Lookahead Composition for WFST-based Speech RecognitionabstractAutomatic Speech Recognition (ASR) applications often em-ploy a mixture of static and dynamic grammar components, and can thus benefit from the ability to efficiently modify the sys-tem vocabulary and other parameters in an on-line mode. This paper presents a novel, generic approach to dynamic grammar handling in the context of the Weighted Finite-State Transducer (WFST) paradigm. The method relies on a straightforward ex-tension of the lexicon and underlying grammar components, and leverages the ideas of on-the-fly composition and delayed construction to efficiently generate the recognition search space on-the-fly. The alternative partitioning of component models that this approach implies can also result in significant stor-age savings. In contrast to previous works in this area, the proposed method relies only on generic WFST operations and the context-dependency, lexicon and grammar components that form the basis of standard ASR cascades. Josef R. Novak, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 2 |
| 2012 | Improving WFST-based G2P Conversion with Alignment Constraints and RNNLM N-best RescoringabstractThis work introduces a modified WFST-based mul-tiple to multiple EM-driven alignment algorithm for Grapheme-to-Phoneme (G2P) conversion, and pre-liminary experimental results applying a Recurrent Neural Network Language Model (RNNLM) as an N-best rescoring mechanism for G2P conversion. The alignment algorithm leverages the WFST framework and introduces several simple structural constraints which yield a small but consistent improvement in Word Accuracy (WA) on a selection of standard base-lines. The RNNLM rescoring further extends these gains and achieves state-of-the-art performance on four standard G2P datasets. The system is also shown to be significantly faster than existing solu-tions. Finally, the complete WFST-based G2P frame-work is provided as an open-source toolkit. Josef R. Novak, Nobuaki Minematsu, Keikichi Hirose, Chiori Hori, Hideki Kashioka, Paul R. Dixon |
INTERSPEECH | 2 |
| 2012 | Effects of Speaker Adaptive Training on Tensor-based Arbitrary Speaker ConversionabstractThis paper introduces speaker adaptive training techniques to tensor-based arbitrary speaker conversion. In voice conversion studies, realization of conversion from/to an arbitrary speaker’s voice is one of the important objectives. For this purpose, eigen-voice conversion (EVC), which is based on an eigenvoice Gaus-sian mixture model (EV-GMM), was proposed. Although the EVC can effectively construct the conversion model for arbi-trary target speakers using only a few utterances, increase of the utterances used to construct the conversion model does not always improve the conversion performance. This is because the EV-GMMmethod has an inherent problem in representation of GMM supervectors. We previously proposed tensor-based speaker space as a solution for this problem, and realized more flexible control of speaker characteristics. In this paper, to aim larger improvement of the performance of VC, speaker adaptive training and tensor-based speaker representation are integrated. The proposed method can construct the flexible and precise con-version model, and experimental results of one-to-many voice conversion demonstrate the effectiveness of the proposed ap-proach. Index Terms: voice conversion, Gaussian mixture model, eigenvoice, Tucker decomposition, speaker adaptive training Daisuke Saito, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 2 |
| 2012 | Discriminative Reranking for LVCSR Leveraging Invariant StructureabstractAn invariant structure is one of the long-span acoustic represen-tations, where acoustic variations caused by non-linguistic fac-tors are effectively removed from speech. We present in this pa-per a new method to leverage the invariant structures as features of discriminative reranking for Large Vocabulary Continuous Speech Recognition (LVCSR). First we use a traditional HMM-based LVCSR system to get a list of N-best candidates with phone alignments and construct an invariant structure for each candidate using its phone alignment. Here, the invariant struc-ture is composed of lengths between every two phonemes in the candidate. Then we estimate a score of each phoneme-pair in the invariant structure, and rerank the N-best candidates using a weighted sum of the phoneme-pair scores, where the weights are trained discriminatively by averaged perceptron. Experi-mental results show a relative CER improvement of 6.69 % over the baseline HMM-based LVCSR system. Index Terms: Invariant Structure, LVCSR, Discriminative reranking 1. Masayuki Suzuki, Gakuto Kurata, Masafumi Nishimura, Nobuaki Minematsu |
INTERSPEECH | 4 |
| 2012 | Audio-visual feature integration based on piecewise linear transformation for noise robust automatic speech recognitionabstractMultimodal speech recognition is a promising approach to realize noise robust automatic speech recognition (ASR), and is currently gathering the attention of many researchers. Multimodal ASR utilizes not only audio features, which are sensitive to background noises, but also non-audio features such as lip shapes to achieve noise robustness. Although various methods have been proposed to integrate audio-visual features, there are still continuing discussions on how the vest integration of audio and visual features is realized. Weights of audio and visual features should be decided according to the noise features and levels: in general, larger weights to visual features when the noise level is low and vice versa, but how it can be controlled? In this paper, we propose a method based on piecewise linear transformation in feature integration. In contrast to other feature integration methods, our proposed method can appropriately change the weight depending on a state of an observed noisy feature, which has information both on uttered phonemes and environmental noise. Experiments on noisy speech recognition are conducted following to CENSREC-1-AV, and word error reduction rate around 24% is realized in average as compared to a decision fusion method. Yosuke Kashiwagi, Masayuki Suzuki, Nobuaki Minematsu, Keikichi Hirose |
SLT | 3 |
| 2012 | Performance improvement of automatic pronunciation assessment in a noisy classroomabstractIn recent years Computer-Assisted Language Learning (CALL) systems have been widely used in foreign language education. Some systems use automatic speech recognition (ASR) technologies to detect pronunciation errors and estimate the proficiency level of individual students. When speech recording is done in a CALL classroom, however, utterances of a student are always recorded with those of the others in the same class. The latter utterances are just background noise, and the performance of automatic pronunciation assessment is degraded especially when a student is surrounded with very active students. To solve this problem, we apply a noise reduction technique, Stereo-based Piecewise Linear Compensation for Environments (SPLICE), and the compensated feature sequences are input to a Goodness Of Pronunciation (GOP) assessment system. Results show that SPLICE-based noise reduction works very well as a means to improve the assessment performance in a noisy classroom. Yi Luan, Masayuki Suzuki, Yutaka Yamauchi, Nobuaki Minematsu, Shuhei Kato, Keikichi Hirose |
SLT | 4 |
| 2012 | Automatic Chinese pronunciation error detection using SVM trained with structural featuresabstractPronunciation errors are often made by learners of a foreign language. To build a Computer-Assisted Language Learning (CALL) system to support them, automatic error detection is essential. In this study, Japanese learners of Chinese are focused on. We investigated in automatic detection of their typical and frequent phoneme production errors. For this aim, four databases are newly created and we propose a detection method using Support Vector Machine (SVM) with structural features. The proposed method is compared to two baseline methods of Goodness Of Pronunciation (GOP) and Likelihood Ratio (LR) under the task of phoneme error detection. Experiments show that the proposed method performs much better than both of the two baseline methods. For example, the false rejection rate is reduced by as much as 82%. However, the results also indicate some drawbacks of using SVM with structural features. In this paper, we discuss merits and demerits of the proposed method and in what kind of real applications it works effectively. Tongmu Zhao, Akemi Hoshino, Masayuki Suzuki, Nobuaki Minematsu, Keikichi Hirose |
SLT | 4 |
| 2012 | A method for generation of Mandarin F0 contours based on tone nucleus model and superpositional model
Qinghua Sun, Keikichi Hirose, Nobuaki Minematsu |
Speech Commun. | 3 |
| 2012 | Statistical Voice Conversion Based on Noisy Channel ModelabstractThis paper describes a novel framework of voice conversion effectively using both a joint density model and a speaker model. In voice conversion studies, approaches based on the Gaussian mixture model (GMM) with probabilistic densities of joint vectors of a source and a target speakers are widely used to estimate a transform function between both the speakers. However, to achieve sufficient quality, these approaches require a parallel corpus which contains plenty of utterances with the same linguistic content spoken by both the speakers. In addition, the joint density GMM methods often suffer from overtraining effects when the amount of training data is small. To compensate for these problems, we propose a voice conversion framework, which integrates the speaker GMM of the target with the joint density model using a noisy channel model. The proposed method trains the joint density model with a few parallel utterances, and the speaker model with nonparallel data of the target, independently. It can ease the burden on the source speaker. Experiments demonstrate the effectiveness of the proposed method, especially when the amount of the parallel corpus is small. Daisuke Saito, Shinji Watanabe 0001, Atsushi Nakamura, Nobuaki Minematsu |
IEEE Trans. Speech Audio Process. | 4 |
| 2011 | Decision of response timing for incremental speech recognition with reinforcement learningabstractIn spoken dialog systems, it is important to reduce the delay in generating a response to a user's utterance. We investigate the use of incremental recognition results which can be obtained from a speech recognition engine before the input utterance ends. To enable the system to respond correctly before the end of the utterance, it is desired to utilize the incremental results effectively, although they are not reliable enough. We formulate this problem as a decision making task, in which the system makes choices iteratively either to answer based on previous observations, or to wait until the next observation. The reinforcement learning can be applied to the problem. As the results of experiments, the users highly evaluate the proposed method which estimate completion time of a user's utterance by using the results of speech recognition based on mora units. Takuya Nishimoto, Nobuaki Minematsu |
ASRU | 3 |
| 2011 | Improved F0 modeling and generation in voice conversionabstractF0 is an acoustic feature that varies largely from one speaker to an other. F0 is characterized by a discontinuity in the transition between voiced and unvoiced sounds that presents an obstacle to GMM modeling for use in voice conversion. A Multi-Space Distribution (MSD) [5] can be used to model unvoiced and voiced F0 regions in a linearly weighted mixture. However, the use of two incompatible probabilistic spaces, for example a continuous probability density for voiced observations, and a discrete probability for unvoiced observations, may result in an imprecise voiced/unvoiced (v/u) conversion in a maximum likelihood (ML) sense. In this paper we propose to use voicing strength, characterized by the normalized correlation coefficient magnitude, as calculated from F0 feature extraction, as an additional feature for improving F0 modeling and the v/u decision in the context of voice conversion. The proposed method was evaluated on male-to-female voice conversion tasks in both Mandarin and English. Objective tests showed that the approach is effective in reducing the Root Mean Square Error, while the results for subjective metrics including AB preference and ABX speaker similarity tests also showed gains. Aki Kunikoshi, Yao Qian, Frank K. Soong, Nobuaki Minematsu |
ICASSP | 4 |
| 2011 | High accurate model-integration-based voice conversion using dynamic features and model structure optimizationabstractThis paper combines a parameter generation algorithm and a model optimization approach with the model-integration-based voice con version (MIVC). We have proposed probabilistic integration of a joint density model and a speaker model to mitigate a requirement of the parallel corpus in voice conversion (VC) based on Gaussian Mixture Model (GMM). As well as the other VC methods, MIVC also suffers from the problems; the degradation of the perceptual quality caused by the discontinuity through the parameter trajectory, and the difficulty to optimize the model structure. To solve the problems, this paper proposes a parameter generation algorithm constrained by dynamic features for the first problem and an information criterion including mutual influences between the joint density model and the speaker model for the second problem. Experimental results show that the first approach improved the performance of VC and the second approach appropriately predicted the optimal number of mixtures of the speaker model for our MIVC. Daisuke Saito, Shinji Watanabe 0001, Atsushi Nakamura, Nobuaki Minematsu |
ICASSP | 4 |
| 2011 | Adaptation of Prosody in Speech Synthesis by Changing Command Values of the Generation Process Model of Fundamental Frequency
Keikichi Hirose, Keiko Ochi, Ryusuke Mihara, Hiroya Hashimoto, Daisuke Saito, Nobuaki Minematsu |
INTERSPEECH | 6 |
| 2011 | Gesture Design of Hand-to-Speech Converter Derived from Speech-to-Hand Converter Based on Probabilistic Integration ModelabstractWhen dysarthrics, individuals with speaking disabilities, try to communicate using speech, they often have no choice but to use speech synthesizers which require them to type word sym-bols or sound symbols. Input by this method often makes real-time communication troublesome and dysarthric users struggle to have smooth flowing conversations. In this study, we are developing a novel speech synthesizer where speech is gener-ated through hand motions rather than symbol input. In re-cent years, statistical voice conversion techniques have been proposed based on space mapping between given parallel ut-terances. By applying these methods, a hand space was mapped to a vowel space and a converter from hand motions to vowel transitions was developed. It reported that the proposed method is effective enough to generate the five Japanese vowels. In this paper, we discuss the expansion of this system to conso-nant generation. In order to create the gestures for consonants, a Speech-to-Hand conversion system is firstly developed using parallel data for vowels, in which consonants are not included. Then, we are able to automatically search for candidates for consonant gestures for a Hand-to-Speech system. Index Terms: Dysarthria, speech production, hand motions, media conversion, arrangement of gestures and vowels Aki Kunikoshi, Yu Qiao 0001, Daisuke Saito, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 4 |
| 2011 | Measurement of Objective Intelligibility of Japanese Accented English Using ERJ (English Read by Japanese) DatabaseabstractIn many schools, English is taught as international communication tool and the goal of English pronunciation training is generally to acquire intelligible enough pronunciation, which is not always native-sounding pronunciation. However, the definition of the intelligible pronunciation is not easy because it depends on the speaking skill of a speaker, the predictability of a content, and the language background of a listener. One kind of accented pronunciation, which is intelligible enough for some listeners, is often less intelligible for others. This paper focuses on objective intelligibility of Japanese English through the ears of American English speakers with little exposure to Japanese English. A large listening test was conducted using ERJ (English Read by Japanese) database. A balanced subset of this database were presented over a telephone line to the American listeners who were asked to repeat what they heard. Totally, 17,416 repetitive responses were collected and they were transcribed manually. This paper describes the design of this experiment and some results of analyzing the results of transcription. Index Terms: English pronunciation training, foreign accent, intelligibility, listening test, ERJ database Nobuaki Minematsu, Koji Okabe, Keisuke Ogaki, Keikichi Hirose |
INTERSPEECH | 1 |
| 2011 | Painless WFST Cascade Construction for LVCSR - TransducersaurusabstractThis paper introduces the Transducersaurus toolkit which provides a set of classes for generating each of the fundamental components of a typical WFST ASR cascade, including a Context-dependency transducer, a Lexicon, a stochastic language model and an optional silence class model. The toolkit further implements a simple scripting language in order to facilitate the construction of cascades with a variety of popular combination and optimization methods and provides integrated support for the T 3 and Juicer WFST decoders, and both Sphinx and HTK format acoustic models. New results for two standard WSJ tasks are also provided, comparing a variety of cascade construction and optimization algorithms. These results illustrate the flexibility of the toolkit as well as the tradeoffs inherent in various build algorithms. Index Terms: Speech Recognition, WFST, LVCSR Josef R. Novak, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 2 |
| 2011 | A Study on Bag of Gaussian Model with Application to Voice Conversion
Yu Qiao 0001, Tong Tong 0001, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2011 | One-to-Many Voice Conversion Based on Tensor Representation of Speaker SpaceabstractThis paper describes a novel approach to flexible control of speaker characteristics using tensor representation of speaker space. In voice conversion studies, realization of conversion from/to an arbitrary speaker’s voice is one of the important objectives. For this purpose, eigenvoice conversion (EVC) based on an eigenvoice Gaussian mixture model (EV-GMM) was proposed. In the EVC, similarly to speaker recognition approaches, a speaker space is constructed based on GMM supervectors which are high-dimensional vectors derived by concatenating the mean vectors of each of the speaker GMMs. In the speaker space, each speaker is represented by a small number of weight parameters of eigen-supervectors. In this paper, we revisit construction of the speaker space by introducing the tensor analysis of training data set. In our approach, each speaker is represented as a matrix of which the row and the column respectively correspond to the Gaussian component and the dimension of the mean vector, and the speaker space is derived by the tensor analysis of the set of the matrices. Our approach can solve an inherent problem of supervector representation, and it improves the performance of voice conversion. Experimental results of oneto-many voice conversion demonstrate the effectiveness of the proposed approach. Index Terms: voice conversion, Gaussian mixture model, eigenvoice, tensor analysis, Tucker decomposition Daisuke Saito, Keisuke Yamamoto, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 3 |
| 2011 | Continuous Digits Recognition Leveraging Invariant StructureabstractRecently, an invariant structure of speech was proposed, where the inevitable acoustic variations caused by non-linguistic fac-tors are effectively removed from speech. The invariant struc-ture was applied to isolated word recognition and the experi-mental results showed good performance. However, the pre-vious method can’t apply to continuous speech recognition di-rectly because there was no efficient decoding algorithm. In this paper, we propose a method to leverage the invariant structure in continuous digits recognition. We use a traditional HMM-based Automatic Speech Recognition (ASR) system to get N-best lists with phone alignments. Then we construct invariant structures using these phone alignments and re-rank the N-best lists by investigating which hypothesis is structurally more valid. Experimental results show a relative WER improvement of 17.4 % over the baseline HMM-based ASR system. Masayuki Suzuki, Gakuto Kurata, Masafumi Nishimura, Nobuaki Minematsu |
INTERSPEECH | 4 |
| 2011 | Prosody Conversion for Emotional Mandarin Speech Synthesis Using the Tone Nucleus Model
Miaomiao Wen, Keikichi Hirose, Nobuaki Minematsu |
INTERSPEECH | 4 |
| 2010 | HMM-based sequence-to-frame mapping for voice conversionabstractVoice conversion can be reduced to a problem to find a transformation function between the corresponding speech sequences of two speakers. Perhaps the most voice conversions methods are GMM-based statistical mapping methods. However, the classical GMM-based mapping is frame-to-frame, and cannot take account of the contextual information existing over a speech sequence. It is well known that HMM yields an efficient method to model the density of a whole speech sequence and has found great successes in speech recognition and synthesis. Inspired by this fact, this paper studies how to use HMM for voice conversion. We derive an HMM-based sequence-to-frame mapping function with statistical analysis. Different from previous HMM-based voice conversion methods that used forced alignment for segmentation and transform frames aligned to a state with its associated linear transformation, our method has a soft mapping function as a weighted summation of linear transformations. The weights are calculated as the HMM posterior probabilities of frames. We also propose and compare two methods to learn the parameters of our mapping functions, namely least square error estimation and maximum likelihood estimation. We carried out experiments to examine the proposed HMM-based method for voice conversion. Yu Qiao 0001, Daisuke Saito, Nobuaki Minematsu |
ICASSP | 3 |
| 2010 | Regularized-MLLR speaker adaptation for computer-assisted language learning systemabstractIn this paper, we propose a novel speaker adaptation technique, regularized-MLLR, for Computer Assisted Language Learn-ing (CALL) systems. This method uses a linear combination of a group of teachers ’ transformation matrices to represent each target learner’s transformation matrix, thus avoids the over-adaptation problem that erroneous pronunciations come to be judged as good pronunciations after conventional MLLR speaker adaptation, which uses learners ’ “imperfect ” speech as target utterances of adaptation. Experiments of automatic scor-ing and error detection on public databases show that the pro-posed method outperforms conventional MLLR adaption in pronunciation evaluation and can avoid the problem of over adaptation. Index Terms: Computer Assisted Language Learning (CALL), speaker adaption, pronunciation evaluation, goodness of pro-nunciation (GOP), maximum likelihood linear regression (MLLR) 1. Dean Luo, Yu Qiao 0001, Nobuaki Minematsu, Yutaka Yamauchi, Keikichi Hirose |
INTERSPEECH | 3 |
| 2010 | Probabilistic integration of joint density model and speaker model for voice conversionabstractThis paper describes a novel approach to voice conversion using both a joint density model and a speaker model. In voice con-version studies, approaches based on Gaussian Mixture Model (GMM) with probabilistic densities of joint vectors of a source and a target speakers are widely used to estimate a transfor-mation. However, for sufficient quality, they require a parallel corpus which contains plenty of utterances with the same lin-guistic content spoken by both the speakers. In addition, the joint density GMM methods often suffer from over-training ef-fects when the amount of training data is small. To compensate for these problems, we propose a novel approach to integrate the speaker GMM of the target with the joint density model using probabilistic formulation. The proposed method trains the joint density model with a few parallel utterances, and the speaker model with non-parallel data of the target, independently. It eases the burden on the source speaker. Experiments demon-strate the effectiveness of the proposed method, especially when the amount of the parallel corpus is small. Index Terms: voice conversion, joint density model, speaker model, probabilistic unification Daisuke Saito, Shinji Watanabe 0001, Atsushi Nakamura, Nobuaki Minematsu |
INTERSPEECH | 4 |
| 2010 | Integration of multilayer regression analysis with structure-based pronunciation assessmentabstractAutomatic pronunciation assessment has several difficulties. Adequacy in controlling the vocal organs is often estimated from the spectral envelopes of input utterances but the envelope patterns are also affected by other factors such as speaker iden-tity. Recently, a new method of speech representation was pro-posed where these non-linguistic variations are effectively re-moved through modeling only the contrastive aspects of speech features. This speech representation is called speech struc-ture. However, the often excessively high dimensionality of the speech structure can degrade the performance of structure-based pronunciation assessment. To deal with this problem, we integrate multilayer regression analysis with the structure-based assessment. The results show higher correlation between hu-man and machine scores and also show much higher robustness to speaker differences compared to widely used GOP-based analysis. Index Terms: CALL, speech structure, regression, GOP 1. Masayuki Suzuki, Yu Qiao 0001, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 3 |
| 2010 | Improved generation of fundamental frequency in HMM-based speech synthesis using generation process model
Miaomiao Wen, Keikichi Hirose, Nobuaki Minematsu |
INTERSPEECH | 4 |
| 2010 | Improving Mandarin segmental duration prediction with automatically extracted syntax features
Miaomiao Wen, Keikichi Hirose, Nobuaki Minematsu |
INTERSPEECH | 4 |
| 2009 | A study on Hidden Structural Model and its application to labeling sequencesabstractThis paper proposes hidden structure model (HSM) for statistical modeling of sequence data. The HSM generalizes our previous proposal on structural representation by introducing hidden states and probabilistic models. Compared with the previous structural representation, HSM not only can solve the problem of misalignment of events, but also can conduct structure-based decoding, which allows us to apply HSM to general speech recognition tasks. Different from HMM, HSM accounts for the probability of both locally absolute and globally contrastive features. This paper focuses on the fundamental formulation and theories of HSM. We also develop methods for the problems of state inference, probability calculation and parameter estimation of HSM. Especially, we show that the state inference of HSM can be reduced to a quadratic programming problem. We carry out two experiments to examine the performance of HSM on labeling sequences. The first experiment tests HSM by using artificially transformed sequences, and the second experiment is based on a Japanese corpus of connected vowel utterances. The experimental results demonstrate the effectiveness of HSM. Yu Qiao 0001, Masayuki Suzuki, Nobuaki Minematsu |
ASRU | 3 |
| 2009 | Sub-structure-based estimation of pronunciation proficiency and classification of learnersabstractAutomatic estimation of pronunciation proficiency has its specific difficulty. Adequacy in controlling the vocal organs can be estimated from spectral envelopes of input utterances but the envelope patterns are also affected easily by different speakers. To develop a pedagogically sound method for automatic estimation, the envelope changes caused by linguistic factors and those by extra-linguistic factors should be properly separated. For this aim, in our previous study [1], we proposed a mathematically-guaranteed and linguistically-valid speaker-invariant representation of pronunciation, called speech structure. After the proposal, we have examined that representation also for ASR [2], [3], [4] and, through these works, we have learned better how to apply speech structures to various tasks. In this paper, we focus on a proficiency estimation experiment done in [1] and, based on our recently proposed techniques for the structures, we carry out that experiment again but under new and different conditions. Here, we use smaller units of structural analysis, speaker-invariant substructures, and relative structural distances between a learner and a teacher. Results show that correlations between human and machine rating are improved and also show extremely higher robustness to speaker differences compared to widely used GOP scores. Further, we also demonstrate that the proposed representation can classify learners purely based on their pronunciation proficiency, not affected by their age and gender. Masayuki Suzuki, Nobuaki Minematsu, Dean Luo, Keikichi Hirose |
ASRU | 2 |
| 2009 | Control of prosodic focus in corpus-based generation of fundamental frequency contours of Japanese based on the generation process modelabstractA total corpus-based process of generating prosodic features from text is developed. The process first predicts pauses and phone durations, and then generates F0contours. Since F0contour generation is based on the generation process model, it is rather easy to manipulate the generated F0contours in command level. A method was developed for generating sentence F0contours, when a focus is placed in one of the ldquobunsetsurdquo of an utterance. The method is to predict differences in the F0model commands between with and without focus utterances, and apply them to the F0model commands predicted beforehand by the baseline method. The validity of the method was proved by the experiment on F0contour generation and speech synthesis. Keiko Ochi, Keikichi Hirose, Nobuaki Minematsu |
ICASSP | 3 |
| 2009 | Mixture of Probabilistic Linear Regressions: A unified view of GMM-based mapping techiquesabstractThis paper introduces a model of mixture of probabilistic linear regressions (MPLR) to learn a mapping function between two feature spaces. The MPLR consists of weighted combination of several probabilistic linear regressions, whose parameters are estimated by using matrix calculation. The mixture nature of MPLR allows it to model nonlinear transformation. The formulation of MPLR is general and independent of the types of the density models used. Two well-known GMM-based mapping methods for voice conversion [1, 2] can be regarded as special cases of MPLR. This unified view not only provides insights to the GMM-based mapping techniques, but also indicates methods to improve them. Compared to [1], our formulation of MPLR avoids solving complex linear equations and yields a faster estimation of the transform parameters. As for [2], the MPLR estimation provides a modified mapping function which overcomes an implicit problem in [2]-s mapping function. We carried out experiments to compare the MPLR-based methods with the traditional GMM-based methods [1, 2] on a voice conversion task. The experimental results show that the MPLR-based methods always have better performance in various parameter setups. Yu Qiao 0001, Nobuaki Minematsu |
ICASSP | 2 |
| 2009 | Affine invariant features and their application to speech recognitionabstractThis paper proposes a set of affine invariant features (AIFs) for sequence data. The proposed AIFs can be calculated directly from the sequence data, and their invariance to affine transformation is proved mathematically through algebraic calculation. We apply the AIFs to speech recognition. Since the vocal tract length (VTL) difference causes to frequency warping which can be approximated well by affine transform on cepstral features, the AIFs of cepstral sequence provide robust features for VTL variations. We experimentally examine the invariance of AIFs of speech signals, and apply AIFs for Japanese isolated word recognition. The experimental results show that the combination of AIFs with MFCC or MFCC+Delta can lead to higher recognition rates than MFCC or MFCC+Delta only. Especially in the mismatched experiments, the combination with AIFs can reduce the error rates about 30% when compared to MFCC or MFCC+Delta only. The AIFs are expected to have other applications than speech recognition, since their invariance is general. Yu Qiao 0001, Masayuki Suzuki, Nobuaki Minematsu |
ICASSP | 3 |
| 2009 | Speech generation from hand gestures based on space mappingabstractIndividuals with speaking disabilities, particularly people suf-fering from dysarthria, often use a TTS synthesizer for speech communication. Since users always have to type sound symbols and the synthesizer reads them out in a monotonous style, the use of the current synthesizers usually renders real-time opera-tion and lively communication difficult. This is why dysarthric users often fail to control the flow of conversation. In this pa-per, we propose a novel speech generation framework which makes use of hand gestures as input. People usually use tongue gesture transitions for speech generation but we develop a spe-cial glove, by wearing which, speech sounds are generated from hand gesture transitions. For development, GMM-based voice conversion techniques (mapping techniques) are applied to esti-mate a mapping function between a space of hand gestures and another space of speech sounds. In this paper, as an initial trial, a mapping between hand gestures and Japanese vowel sounds is estimated so that topological features of the selected gestures in a feature space and those of the five Japanese vowels in a cepstrum space are equalized. Experiments show that the spe-cial glove can generate good Japanese vowel transitions with voluntary control of duration and articulation. Index Terms: Dysarthria, speech production, hand motions, media conversion, arrangement of gestures and vowels Aki Kunikoshi, Yu Qiao 0001, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 3 |
| 2009 | Analysis and utilization of MLLR speaker adaptation technique for learners' pronunciation evaluationabstractIn this paper, we investigate the effects and problems of MLLR speaker adaptation when applied to pronunciation evaluation. Automatic scoring and error detection experiments are conducted on two publicly available databases of Japanese learners’ English pronunciation. As we expected, overadaptation causes misjudge of pronunciation accuracy. Following these experiments, two novel methods, Forced-aligned GOP scoring and Regularized-MLLR adaptation, are proposed to solve the adverse effects of MLLR adaption. Experimental results show that the proposed methods can better utilize MLLR adaptation and avoid over-adaptation. Index Terms: Computer Assisted Language Learning (CALL), speaker adaption, pronunciation evaluation, goodness of pronunciation (GOP), maximum likelihood linear regression (MLLR) Dean Luo, Yu Qiao 0001, Nobuaki Minematsu, Yutaka Yamauchi, Keikichi Hirose |
INTERSPEECH | 3 |
| 2009 | Structural analysis of dialects, sub-dialects and sub-sub-dialects of ChineseabstractIn China, there are hundred kinds of dialects. By traditional dialectology, they are classified into seven big dialect regions and most of them also have many sub-dialects and sub-subdialects. As they are different in various linguistic aspects, people from different dialect regions often cannot communicate orally. But for the sub-dialects of one dialect region, although they are sometimes still mutually unintelligible, more common features are shared. In this paper, a dialect pronunciation structure, which has been used successfully in dialectbased speaker classification in our previous work [1], is examined for the task of speaker classification and distance measurement among cities based on sub-dialects of Mandarin. Using the finals of the dialectal utterances of a specific list of written characters, a dialect pronunciation structure is built for every speaker in a data set and these speakers are classified based on the distances among their structures. Then, the results of classifying 16 Mandarin speakers based on their sub-dialects show that they are linguistically classified with little influence of their age and gender. Finally, distances among sub-sub-dialects are similarly calculated and evaluated. All the results show high validity and accordance to linguistic studies. Index Terms: Sub-dialects of Mandarin, pronunciation structure, speaker classification Xuebin Ma, Akira Nemoto, Nobuaki Minematsu, Yu Qiao 0001, Keikichi Hirose |
INTERSPEECH | 3 |
| 2009 | On invariant structural representation for speech recognition: theoretical validation and experimental improvementabstractOne of the most challenging problems in speech recognition is to deal with inevitable acoustic variations caused by nonlinguistic factors. Recently, an invariant structural representation of speech was proposed [1], where the non-linguistic variations are effectively removed though modeling the dynamic and contrastive aspects of speech signals. This paper describes our recent progresses on this problem. Theoretically, we prove that the maximum likelihood based decomposition can lead to the same structural representations for a sequence and its transformed version. Practically, we introduce a method of discriminant analysis of eigen-structure to deal with two limitations of structural representations, namely, high dimensionality and too strong invariance. In the 1st experiment, we evaluate the proposed method through recognizing connected Japanese vowels. The proposed method achieves a recognition rate 99.0%, which is higher than those of the previous structure based recognition methods [2, 3, 4] and word HMM. In the 2nd experiment, we examine the recognition performance of structural representations to vocal tract length (VTL) differences. The experimental results indicate that structural representations have much more robustness to VTL changes than HMM. w Moreover, the proposed method is about 60 times faster than the previous ones. Index Terms: Speech recognition, invariant structure, PCA, discriminative analysis Yu Qiao 0001, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 2 |
| 2009 | How to improve TTS systems for emotional expressivityabstractSeveral experiments have been carried out that revealed weaknesses of the current Text-To-Speech (TTS) systems in their emotional expressivity. Although some TTS systems allow XML-based representations of prosodic and/or phonetic variables, few publications considered, as a pre-processing stage, the use of intelligent text processing to detect affective information that can be used to tailor the parameters needed for emotional expressivity. This paper describes a technique for an automatic prosodic parameterization based on affective clues. This technique recognizes the affective information conveyed in a text and, accordingly to its emotional connotation, assigns appropriate pitch accents and other prosodic parameters by XML-tagging. This pre-processing assists the TTS system to generate synthesized speech that contains emotional clues. The experimental results are encouraging and suggest the possibility of suitable emotional expressivity in speech synthesis. Antonio Rui Ferreira Rebordão, Shaikh Mostafa Al Masum, Keikichi Hirose, Nobuaki Minematsu |
INTERSPEECH | 4 |
| 2009 | Optimal event search using a structural cost function - improvement of structure to speech conversionabstractThis paper describes a new and improved method for the frame-work of structure to speech conversion we previously proposed. Most of the speech synthesizers take a phoneme sequence as input and generate speech by converting each of the phonemes into its corresponding sound. In other words, they simulate a human process of reading text out. However, infants usually acquire speech communication ability without text or phoneme sequences. Since their phonemic awareness is very immature, they can hardly decompose an utterance into a sequence of phones or phonemes. As developmental psychology claims, in-fants acquire the holistic sound patterns of words from the utter-ances of their parents, called word Gestalt, and they reproduce them with their vocal tubes. This behavior is called vocal im-itation. In our previous studies, the word Gestalt was defined physically and a method of extracting it from a word utterance was proposed. We already applied the word Gestalt to ASR, CALL, and also speech generation, which we call structure to speech conversion. Unlike reading machines, our framework simulates infants ’ vocal imitation. In this paper, a method for improving our speech generation framework based on a struc-tural cost function is proposed and evaluated. Index Terms: speech synthesis, the structural representation, vocal imitation, a structural cost function Daisuke Saito, Yu Qiao 0001, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 3 |
| 2009 | A Theory of Phase Singularities for Image Representation and its Applications to Object Tracking and Image MatchingabstractThis paper studies phase singularities (PSs) for image representation. We show that PSs calculated with Laguerre-Gauss filters contain important information and provide a useful tool for image analysis. PSs are invariant to image translation and rotation. We introduce several invariant features to characterize the core structures around PSs and analyze the stability of PSs to noise addition and scale change. We also study the characteristics of PSs in a scale space, which lead to a method to select key scales along phase singularity curves. We demonstrate two applications of PSs: object tracking and image matching. In object tracking, we use the iterative closest point algorithm to determine the correspondences of PSs between two adjacent frames. The use of PSs allows us to precisely determine the motions of tracked objects. In image matching, we combine PSs and scale-invariant feature transform (SIFT) descriptor to deal with the variations between two images and examine the proposed method on a benchmark database. The results indicate that our method can find more correct matching pairs with higher repeatability rates than some well-known methods. Yu Qiao 0001, Wei Wang 0333, Nobuaki Minematsu, Jianzhuang Liu, Mitsou Takeda, Xiaoou Tang |
IEEE Trans. Image Process. | 3 |
| 2008 | Multi-stream parameterization for structural speech recognitionabstractRecently, a novel and structural representation of speech was proposed [1, 2], where the inevitable acoustic variations caused by nonlinguistic factors are effectively removed from speech. This structural representation captures only microphone- and speaker-invariant speech contrasts or dynamics and uses no absolute or static acoustic properties directly such as spectrums. In our previous study, the new representation was applied to recognizing a sequence of isolated vowels [3]. The structural models trained with a single speaker outperformed the conventional HMMs trained with more than four thousand speakers even in the case of noisy speech. We also applied the new models to recognizing utterances of connected vowels [4]. In the current paper, a multiple stream structuralization method is proposed to improve the performance of the structural recognition framework. The proposed method only with 8 training speakers shows the very comparable performance to that of the conventional 4,130-speaker triphone-based HMMs. Satoshi Asakawa, Nobuaki Minematsu, Keikichi Hirose |
ICASSP | 2 |
| 2008 | Unsupervised optimal phoneme segmentation: Objectives, algorithm and comparisonsabstractPhoneme segmentation is a fundamental problem in many speech recognition and synthesis studies. Unsupervised phoneme segmentation assumes no knowledge on linguistic contents and acoustic models, and thus poses a challenging problem. The essential question here is what is the optimal segmentation. This paper formulates the optimal segmentation problem into a probabilistic framework. Using statistics and information theory analysis, we develop three different objective functions, namely, summation of square error (SSE), log determinant (LD) and rate distortion (RD). Specially, RD function is derived from information rate distortion theory and can be related to human signal perception mechanism. We introduce a time-constrained agglomerative clustering algorithm to find the optimal segmentations. We also propose an efficient method to implement the algorithm by using integration functions. We carry out experiments on TIMIT database to compare the above three objective functions. The results show that rate distortion achieves the best performance and indicate that our method outperforms the recently published unsupervised segmentation methods. Yu Qiao 0001, Naoya Shimomura, Nobuaki Minematsu |
ICASSP | 3 |
| 2008 | Phase singularities for image representation and matchingabstractPhase features are widely used in image processing and representation due to their stability to deformation and noise. However, phase singularities,where the signals vanish, are generally regarded as harmful and unreliable facts. In this paper, on the contrary, we will show that phase singularities calculated by Laguerre-Gauss filter contain important information of input image and can provide a reliable representation for image matching. We show that the positions of phase singularities are invariant to translation and rotation. Usually, it is possible to recover the input image up to a constant scaling only from the positions of phase singularities. We study phase singularities in scale space, which allows us to determine the "intrinsic scales" of key phase singularities. We introduce three physical measures of the local structures of phase singularities and combine these measures with SIFT descriptor for image matching. We execute experiments on benchmark database to examine the proposed methods. The results indicate that the proposed method can achieve comparable performance with certain well-known methods. Yu Qiao 0001, Wei Wang 0333, Nobuaki Minematsu, Jianzhuang Liu, Xiaoou Tang |
ICASSP | 3 |
| 2008 | Directional dependency of cepstrum on vocal tract lengthabstractIN this paper, we prove that the direction of cepstrum vectors strongly depends on vocal tract length and that this dependency is represented as rotation in the n dimensional cepstrum space. In speech recognition studies, vocal tract length normalization (VTLN) techniques are widely used to cancel age- and gender-differences. In VTLN, a frequency warping is often carried out and it can be implemented as a linear transformation in a cepstrum space; c = Ac. However, the geometric properties of this transformation matrix A have not been well discussed. In this study, its properties are made clear using n dimensional geometry and it is shown that the matrix rotates any cepstrum vector similarly and apparently. Experimental results using resynthesized speech demonstrate that cepstrum vectors extracted from a speaker of 180 [cm] in height and those from another speaker of 120 [cm] in height are reasonably orthogonal. This result makes clear one of the reasons why children's speech is very difficult for conventional speech recognizers to deal with adequately. Daisuke Saito, Ryo Matsuura, Satoshi Asakawa, Nobuaki Minematsu, Keikichi Hirose |
ICASSP | 4 |
| 2008 | Automatic pronunciation evaluation of language learners' utterances generated through shadowingabstractIn foreign language learning, shadowing has been used as a method for improving speaking and listening ability. In this method, learners are required to repeat a presented native utterance as closely and quickly as possible. Since learners have to follow the speaking rate of the presented utterance, their pronunciation often becomes very inarticulate and unintelligible. These features of shadowing make it very difficult to build a reliable scoring system for shadowing productions. In this paper, two techniques are proposed and investigated for automatic scoring of shadowing productions. Experiments show that good correlations are found between automatic scores and TOEIC overall proficiency scores. Index Terms: shadowing, automatic scoring, articulatory effort, goodness of pronunciation, bottom-up clustering Dean Luo, Naoya Shimomura, Nobuaki Minematsu, Yutaka Yamauchi, Keikichi Hirose |
INTERSPEECH | 3 |
| 2008 | Robust voiced/unvoiced speech classification using empirical mode decomposition and periodic correlation modelabstractAbstract This paper presents a method of voiced/unvoiced (V/Uv) classification of noisy speech signals. Empirical mode decomposition (EMD), a newly developed tool to analyze nonlinear and non-stationary signals is used to filter the additive noise with the speech signal. The normalized autocorrelation of the filtered speech signal is computed to enhance the periodicity if any. It is considered that the voiced speech signal is periodically correlated and the unvoiced signal is not. A statistical model of determining periodic correlation is used to differentiate voiced and unvoiced speech with low SNR. The experimental results show that the use of EMD improves the classification performance and the overall efficiency is noticeable as compared to other existing algorithms. Index Terms : empirical mode decomposition, normalized autocorrelation, periodic correlation, voiced/unvoiced speech 1. Introduction Reliable classification of short time speech signal into voiced and unvoiced is a crucial preprocessing step in many speech processing applications and is essential in most analysis and synthesis system. For example: different strategy could be adopted for voiced and unvoiced parts in speech enhancement using spectral subtraction. The essence of classification is to determine whether the speech production system involves the vibration of the vocal cords [1]. The discrimination problem is an important one and has been worked on extensively during the last three decades [2]. The discrimination can effectively be performed using a single feature or parameter which is closely associated with the voicing and non-voicing activities of speech signal. Many algorithms have been reported for solving the detection problem [3] – [7]. In [3], Gaussian mixture model with cepstrum coefficients features is proposed for robust V/Uv classification. A higher order statistics (HOS) based method is proposed in [4] for V/Uv detection and pitch estimation simultaneously. The matching pursuit algorithm is used in [5] with Gabor decomposition. The wavelet transform is proposed in pitch and V/Uv detection in [6]. A statistical model applied in autocorrelation domain is also reported in [7]. In most of the existing algorithms are not so much noise robust and also the intensive threshold and training data are required for classification. Such requirements are troublesome for the use in application domain. The proposed method is noise robust and based on the statistical model for periodicity detection in speech signal without any training requirement. To reduce the effect of noise on speech signal, a data adaptive time domain filtering is proposed using newly developed empirical mode decomposition method [8]. Although speech signal is non-stationary in nature, Fourier based frequency domain filtering assumes that it is piecewise stationary. The speech decomposition is performed by fitting some predefined bases without satisfying its non-stationary nature. Whereas, EMD based approach decompose the speech signal as non-stationary time series and hence better performance in noise filtering. A method for determining whether an observed time series contains a periodically correlated sequence is employed here. It is based on the statistical tests for the coherence between spectral components for the presence of a periodically correlated covariance structure in a time series [9]. The autocorrelation function (ACF) makes the periodicity more prominent if any. The proposed periodic correlation model is applied in the autocorrelation domain rather than original time domain of the speech signal. The periodicity detection method is implemented in spectral domain to classify the speech segment into voiced or unvoiced one based on that it contains periodic correlated sequence or not respectively. Md. Khademul Islam Molla, Keikichi Hirose, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2008 | Control of prosodic focus in corpus-based generation of fundamental frequency based on the generation process model
Keiko Ochi, Keikichi Hirose, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2008 | Metric learning for unsupervised phoneme segmentationabstractUnsupervised phoneme segmentation aims at dividing a speech stream into phonemes without using any prior knowledge of lin-guistic contents and acoustic models. In [1], we formulated this problem into an optimization framework, and developed an ob-jective function, summation of squared error (SSE) based on the Euclidean distance of cepstral features. However, it is un-known whether or not Euclidean distance yields the best metric to estimate the goodness of segmentations. In this paper, we study how to learn a good metric to improve the performance of segmentation. We propose two criteria for learning met-ric: Minimum of Summation Variance (MSV) and Maximum of Discrimination Variance (MDV). The experimental results on TIMIT database indicate that the use of learning metric can achieve better segmentation performances. The best recall rate of this paper is 81.8 % (20ms windows), compared to 77.5% of [1]. We also introduce an iterative algorithm to learn met-ric without using labeled data, which achieves similar results as those with labeled data. Index Terms: Unsupervised phoneme segmentation, optimiza-tion, Mahalanobis distance, metric learning Yu Qiao 0001, Nobuaki Minematsu |
INTERSPEECH | 2 |
| 2008 | f-divergence is a generalized invariant measure between distributionsabstractFinding measures (or features) invariant to inevitable variations caused by non-linguistical factors (transformations) is a funda-mental yet important problem in speech recognition. Recently, Minematsu [1, 2] proved that Bhattacharyya distance (BD) be-tween two distributions is invariant to invertible transforms on feature space, and develop an invariant structural representation of speech based on it. There is a question: which kind of mea-sures can be invariant? In this paper, we prove that f-divergence yields a generalized family of invariant measures, and show that all the invariant measures have to be written in the forms of f-divergence. Many famous distances and divergences in in-formation and statistics, such as Bhattacharyya distance (BD), KL-divergence, Hellinger distance, can be written into forms of f-divergence. As an application, we carried out experiments on recognizing the utterances of connected Japanese vowels. The experimental results show that BD and KL have the best perfor-mance among the measures compared. Index Terms: f-divergence, invariant measure, invertible transformation, speech recognition Yu Qiao 0001, Nobuaki Minematsu |
INTERSPEECH | 2 |
| 2008 | Structure to speech conversion - speech generation based on infant-like vocal imitationabstractThis paper proposes a new framework of speech generation by imitating “infants ’ vocal imitation”. Most of the speech synthe-sizers take a phoneme sequence as input and generate speech by converting each of the phonemes into a sound sequentially. In other words, they simulate a human process of reading text out. However, infants usually acquire speech generation abil-ity without text or phoneme sequences. Since their phonemic awareness is very immature, they can hardly decompose a word utterance into a sequence of phones. In this situation, as devel-opmental psychology states, infants acquire the holistic sound pattern of words from the utterances of their parents, called word Gestalt, and they reproduce it with their vocal tubes. This behavior is called vocal imitation. In our previous studies, the word Gestalt was defined physically and a method of extract-ing it from an utterance was proposed and used successfully for ASR and CALL. In this paper, a method of converting the word Gestalt back to speech is proposed and evaluated. Unlike a read-ing machine, our proposal simulates infants ’ vocal imitation. Index Terms: speech synthesis, vocal imitation, word Gestalt, invariant structure, Bhattacharyya distance, searching problem Daisuke Saito, Satoshi Asakawa, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 3 |
| 2008 | Decomposition of rotational distortion caused by VTL difference using eigenvalues of its transformation matrixabstractIn speech recognition studies, vocal tract length normalization (VTLN) techniques are widely used to cancel age- and gender-difference. In VTLN, the distortion is often modeled as a lin-ear transform in a cepstrum space; ĉ=Ac. In our previous study, the geometrical properties ofA were discussed and it was shown that the matrix can be approximated as rotation matrix. In this study, a new method of better approximating A is pro-posed. Using eigenvalues ofA, its quasi-rotational distortion is factorized into multiple rotation operations and multiple magni-fication operations. Using this method, the intrinsic ambiguity of the rotation angle used in our previous study is resolved. In-stead, multiple rotation angles are introduced to understand bet-ter what kind of geometrical distortionsA induces to cepstrum vectors. Experiments show the validity of the new method and a new speech feature is also derived by the new method. Index Terms: frequency warping, rotation matrix, vocal tract length, eigenvalue, rotational plane Daisuke Saito, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 2 |
| 2008 | Corpus-based synthesis of Mandarin speech with F0 contours generated by superposing tone components on rule-generated phrase componentsabstractMandarin speech synthesis was conducted by generating prosodic features by the proposed method and segmental features by HMM-based method. The proposed method generates sentence fundamental frequency (F0) contours by representing them as a superposition of tone components on phrase components. The tone components are realized by concatenating their fragments at tone nuclei predicted by a corpus-based method, while the phrase components are generated by rules under the generation process model (F0model) framework. The method includes prediction of phoneme/pause durations in a statistical method as the first step. Through a listening test on the quality of synthetic speech, it was shown that a better quality was obtainable by the method as compared to that by the full HMM-based method. It was also shown that a better quality is obtainable as compared to the case of generating F0contours without super-positional scheme. Keikichi Hirose, Qinghua Sun, Nobuaki Minematsu |
SLT | 3 |
| 2008 | Filled pauses as cues to the complexity of upcoming phrases for native and non-native listeners
Michiko Watanabe, Keikichi Hirose, Yasuharu Den, Nobuaki Minematsu |
Speech Commun. | 4 |
| 2007 | Random discriminant structure analysis for automatic recognition of connected vowelsabstractThe universal structure of speech [1, 2], proves to be invariant to transformations in feature space, and thus provides a robust representation for speech recognition. One of the difficulties of using structure representation is due to its high dimensionality. This not only increases computational cost but also easily suffers from the curse of dimensionality [3, 4]. In this paper, we introduce random discriminant structure analysis (RDSA) to deal with this problem. Based on the observation that structural features are highly correlated and include large redundancy, the RDSA combines random feature selection and discriminative analysis to calculate several low dimensional and discriminative representations from an input structure. Then an individual classifier is trained for each representation and the outputs of each classifier are integrated for the final classification decision. Experimental results on connected Japanese vowel utterances show that our approach achieves a recognition rate of 98.3% based on the training data of 8 speakers, which is higher than that (97.4%) of HMMs trained with the utterances of 4,130 speakers. Yu Qiao 0001, Satoshi Asakawa, Nobuaki Minematsu |
ASRU | 3 |
| 2007 | Development of a Femininity Estimator using Speaker Recognition Techniques for Voice Therapy of Gender Identity Disorder ClientsabstractThis paper describes the development of an estimator of perceptual femininity (PF) of an input utterance using speaker recognition techniques. The estimator was designed for its clinical use and the target speakers are gender identity disorder (GID) clients, especially MtF (male to female) transsexuals. The voice therapy for MtFs is composed of three kinds of training; 1) raising the baseline F0range, 2) changing the baseline voice quality, and 3) enhancing fo dynamics to produce an exaggerated intonation pattern. The first two focus on static acoustic properties of speech and the voice quality is mainly controlled by size and shape of the articulators, which can be acoustically characterized by the spectral envelope. Gaussian mixture models (GMM) of fo values and spectrums were built separately for biologically male speakers and female ones. Using the four models, PF was estimated automatically for each of 142 utterances of 111 MtFs. The estimated values were compared with the PF values obtained through listening tests. Results showed very high correlation (R=0.86), which is comparable to the intra-rater correlation. Nobuaki Minematsu, Kazutaka Maruyama, Kyoko Sakuraba, Keikichi Hirose, Niro Tayama, Satoshi Imaizumi, Toshio Yamauchi |
ICASSP (4) | 1 |
| 2007 | Automatic recognition of connected vowels only using speaker-invariant representation of speech dynamicsabstractSpeech acoustics vary due to differences in gender, age, microphone, room, lines, and a variety of factors. In speech recognition research, to deal with these inevitable non-linguistic variations, thousands of speakers in different acoustic conditions were prepared to train acoustic models of individual phonemes. Recently, a novel representation of speech dynamics was proposed [1, 2], where the above non-linguistic factors are effectively removed from speech as if pitch information is removed from spectrum by its smoothing. This representation captures only speaker- and microphone-invariant speech dynamics and no absolute or static acoustic properties such as spectrums are used. With them, speaker identity has to remain in speech representation. In our previous study, the new representation was applied to recognizing a sequence of isolated vowels [3]. The Satoshi Asakawa, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 2 |
| 2007 | EMD based soft-thresholding for speech enhancementabstractThis paper introduces a novel speech enhancement method based on Empirical Mode Decomposition (EMD) and softthresholding algorithms. A modified soft thresholding strategy is adapted to the intrinsic mode functions (IMF) of the noisy speech. Due to the characteristics of EMD, each obtained IMF of the noisy signal will have different noise and speech energy distribution, thus will have a different noise variance. Based on this specific noise variance, by applying the proposed thresholding algorithm to each IMF separately, it is possible to effectively extract the existing noise components. The experimental results suggest that the proposed method is significantly more effective in removing the noise components from the noisy speech signal compared to recently reported techniques. The significantly better SNR improvement and the speech quality prove the superiority of the proposed algorithm. Index Terms: speech enhancement, empirical mode decomposition, soft-thresholding Erhan Deger, Md. Khademul Islam Molla, Keikichi Hirose, Nobuaki Minematsu, Md. Kamrul Hasan 0001 |
INTERSPEECH | 4 |
| 2007 | F0 models show Chinese speakers of Japanese insert intonational boundaries and drop pitchabstractWe used a command-response additive F0 model to analyze F0 patterns of Japanese spoken by native speakers of Mandarin Chinese. Compared to native speakers of Japanese, we found that Chinese speakers exhibit the following characteristics: (a) higher pitch, (b) more phrases, (c) bunsetsu decomposition, and (d) utterance-final plunging. These characteristics physically manifest themselves as: (a) higher baseline F0, (b) more phrase commands, (c) more accent commands, and (d) negative commands. These characteristics may be subjectively perceived as: (a) tinnier speech (possible L1 marker but does not degrade communication), (b) disjoint phrases (requires mental consolidation), (c) choppy prosodic words (requires reconstruction), and (d) abrupt utterance termination (possibly misconstrued as emphatic or rude). We believe these difficulties arose from tonal and syllable-timed interference, which can be overcome by prosodic control and planning. Index Terms: F0 contour, command-response model, L2 learning, Japanese, accent, phrase. Hiroko Hirano, Keikichi Hirose, Goh Kawai, Wentao Gu, Nobuaki Minematsu |
INTERSPEECH | 5 |
| 2007 | Corpus-based generation of prosodic features from text based on generation process modelabstractA total scheme of generating prosodic features from a text input was constructed. The method consists of corpus-based prediction of pauses, phone durations and fundamental frequencies (F0's), in this order, and information predicted in an earlier process is utilized in the following processes. Since prediction of F0's is done on the command values of F0 contour generation process model instead of direct F0 values, a stable and flexible control of F0 contours is possible. By adding constraints on the accent command timings as a post processing, a better quality was realized when speech was synthesized using prosodic features generated by the method. Validity of the developed method was confirmed through the listening test of the synthetic speech. Index Terms: speech synthesis, prosodic features, F0 contour 1. Keikichi Hirose, Keiko Ochi, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2007 | Structural assessment of language learners' pronunciationabstractSpeaker-invariant structural representation of speech was proposed [1], where only the phonic contrasts between speech sounds were extracted to form their external structure. The acoustic substances were completely discarded. Considering a mapping function between speaker A’s acoustic space and B’s space, the speech dynamics was mathematically proven to be invariant between the two irrespective of the form of the function [2]. This structural and dynamic representation was applied to describe the pronunciation of learners [3]. Since the nonlinguistic factors were removed effectively, the representation could highlighted the non-nativeness in the individual pronunciations. For vowel learning, it was automatically estimated for each of the learners which vowels to correct by priority [4]. Unlike the conventional approach, the estimation was done without the direct use of sound substances such as spectrums. In this paper, using the vowel charts of the learners plotted by an expert phonetician, the validity of this contrastive or relative approach is examined by comparing it with the conventional absolute approach. Results show the high validity of the proposed method. Index Terms: phonic contrasts, pronunciation structure, CALL Nobuaki Minematsu, K. Kamata, Satoshi Asakawa, Takehiko Makino, Tazuko Nishimura, Keikichi Hirose |
INTERSPEECH | 1 |
| 2007 | Pitch estimation of noisy speech signals using empirical mode decompositionabstractThis paper presents a pitch estimation method of noisy speech signal using empirical mode decomposition (EMD). The normalized autocorrelation function (NACF) of the noisy speech signal is decomposed into a finite set of band-limited signals termed as intrinsic mode functions (IMFs) using EMD. The periodicity of one IMF is supposed to be equal to the accurate pitch period. A conventional autocorrelation based pitch period detection method is used to select the IMF with pitch period. The accurate pitch period is obtained from the selected IMF. The pitch estimation performance in term of gross pitch error (GPE) of the proposed algorithm is compared with recently proposed methods. The experimental results show that the EMD based algorithm performs better in pitch estimation of noisy speech. Index Terms: empirical mode decomposition, pitch estimation, normalized autocorrelation. Md. Khademul Islam Molla, Keikichi Hirose, Nobuaki Minematsu, Md. Kamrul Hasan 0001 |
INTERSPEECH | 3 |
| 2007 | A framework of reply speech generation for concept-to-speech conversion in spoken dialogue systemsabstractDue to recent advancements in speech technologies, a large number of spoken dialogue systems have been constructed. However, since most of them adopt existing text-to-speech syn-thesizers, it is rather difficult to reflect the linguistic informa-tion obtained during the reply sentence generation well in out-put speech. A framework is necessary for correctly reflecting higher-level linguistic information, such as syntactic structure and discourse information. We have constructed a spoken dia-logue system on road guidance and realized concept-to-speech conversion, where output speech is generated in a unified pro-cess. Tag LISP forms keep the syntactic structures throughout the process in order to reflect the linguistic information in the prosody of output speech. Furthermore, by making it possible to insert not only words but also phrase templates in tags, various sentences were generated with a minor increase of templates. Validity of the methods is shown through experiments. Index Terms: spoken dialogue system, concept-to-speech con-version 1. Seiya Takada, Yuji Yagi, Keikichi Hirose, Nobuaki Minematsu |
INTERSPEECH | 4 |
| 2007 | Features of pauses and conjunctions at syntactic and discourse boundaries in Japanese monologuesabstractSyntactic and discourse boundaries are signalled by prosodic cues as well as linguistic cues in speech. We investigated whether there is a correspondence between prosodic or linguistic cues and the boundary strengths. We measured the rates of filled pauses (FPs) and conjunctions, and the durations of silent and filled pauses and conjunctions at four types of boundaries in casual presentations in Japanese. The results showed that the rates of FPs and conjunctions and the durations of silent pauses correspond to the boundary strengths. However, no significant correspondence was found between the duration of FPs or conjunctions and the boundary strengths. The results suggest that how long the speaker pauses and whether he or she utters a FP or a conjunction is relevant to the boundary strengths. However, the durations of FPs and conjunctions are likely to be affected by the other factors such as planning difficulties of the following parts of speech. Index Terms: syntactic boundary, discourse boundary, silent pause, filled pause, conjunction Michiko Watanabe, Yasuharu Den, Keikichi Hirose, Shusaku Miwa, Nobuaki Minematsu |
INTERSPEECH | 5 |
| 2006 | Para-Linguistic Information Represented as Distortion of the Acoustic Universal Structure In SpeechabstractSpeech acoustics varies from speaker to speaker, microphone to microphone, etc. Recently, a novel method was proposed to separate these static non-linguistic features from speech as spectral smoothing can separate pitch information from speech (N. Minematsu, 2005). Absolute properties of speech events, such as formants and spectrums, are completely discarded and only the phonic differences or contrasts between the events are extracted to form their external structure. This structure is called the acoustic universal structure and regarded as physical implementation of structural phonology because the structure is considered to represent only the linguistic and para-linguistic information. In this paper, the structural size is focused on and its correlation with the para-linguistic information is examined. Results showed that the size can be interpreted as magnitude of articulatory efforts made in speech production Nobuaki Minematsu, Satoshi Asakawa, Keikichi Hirose |
ICASSP (1) | 1 |
| 2006 | Localization Based Separation of Mixed Audio Signals with Binary Masking of Hilbert SpectrumabstractThis paper presents a method of audio signal separation from stereo mixtures using binary masking in time-frequency (TF) domain based on the spatial location of the audio sources. The TF representation of audio signal is obtained by Hubert spectrum (HS). The Hubert transformation together with empirical mode decomposition (EMD) produces HS which is a fine-resolution TF representation of any nonlinear and non-stationary signal. The sources are localized in the space of time and intensity differences between two microphones' signals. The separation is performed by masking the target signal in TF domain considering that the sources are disjoint orthogonal. The experimental results of the proposed method show a noticeable improvement of separation efficiency Md. Khademul Islam Molla, Keikichi Hirose, Nobuaki Minematsu |
ICASSP (5) | 3 |
| 2006 | Unfilled pauses in Japanese sentences read aloud by non-native learnersabstractPerception experiments suggest that natives judge non-native unfilled pauses as indiscriminate and indecisive. Multiple regression analyses of unfilled pauses indicate a connection between syntactic structure and pause location and duration. Native speakers uniformly pause at large syntactic breaks with marked duration, whereas non-natives' unfilled pauses are spread over various locations, possibly reflecting limited syntactic planning. Our method might be used to synthesize appropriate unfilled pauses in text-to-speech systems, and to train pausing behavior in automated pronunciation learning systems for non-native learners. Index Terms: unfilled pauses, second language learning, syntactic structure, perception experiments, multiple regression analyses Hiroko Hirano, Goh Kawai, Keikichi Hirose, Nobuaki Minematsu |
INTERSPEECH | 4 |
| 2006 | Corpus-based generation of fundamental frequency contours using generation process model and considering emotional focusesabstractWe formerly conducted emotional speech synthesis using our corpus-based method of generating fundamental frequency (F0) contours from text. The method predicts command values of F0 contour generation process model instead of directly predicting F0 value of each time frame. A better control of F0 contours was realized by taking the emotional level of each bunsetsu into account: adding information on which bunsetsu(s) the emotion is especially placed to the command predictor inputs. In the case of anger, F0 contours closer to the target contours are obtained by adding emotional levels. Speech synthesis was conducted by generating F0 contours in two ways: using commands predicted by taking emotional levels into account and those not. The result of perceptual experiment indicated that emotion was conveyed well by adding emotional levels. Index Terms: speech synthesis, emotion, F0 contour 1. Keikichi Hirose, Yasufumi Asano, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2006 | Tone recognition of continuous speech of standard Chinese using neural network and tone nucleus modelabstractA method is developed for recognizing lexical tone types of Standard Chinese syllables in continuous speech. Neural network (four-layered perceptron) is adopted as classifier. The method includes two steps; first recognizing tone types using prosodic features of voiced part, and then re-recognizing by viewing only on tone nucleus, which is a portion of the syllable showing rather stable fundamental frequency (F0) contour regardless of tone types of the preceding and following syllables. The voiced part (or tone nucleus) is divided into 20 segments, and F 0, delta-F 0, F 0 slope and short-term energy of each segment are served as inputs to the neural network. In order to cope with tone coarticulation, prosodic feature parameters for the last 5 segments of the preceding syllable and the initial 5 segments of the following syllable are included in the neural network inputs. Information on syllable length is also added to the inputs. Tone recognition experiment was conducted for a female speaker's utterances included in HKU96 corpus. The average recognition rate was 86.5 % including neutral tone syllables, when the tone nucleus model was not used. It increased to 86.9 %, when the model was used. The obtained rate is higher by more than 3 points as compared to that obtained by the hidden-Markov-model-based tone recognizer developed by the authors formerly. Index Terms: tone recognition, tone nucleus model, neural network, Standard Chinese Keikichi Hirose, Nobuaki Minematsu |
INTERSPEECH | 4 |
| 2006 | Development of a program for self assessment of Japanese pronunciation by English learnersabstractA program for self assessment of Japanese pronunciation by Englishspeaking learners was developed using a language model built with input from a language teacher in collaboration with speech engineers. This collaboration enhanced the program's capacity for accurate assessment and provides practical support to users by linking evaluation with feedback, and an editorial function of error patterns. The program drew positive responses from participants in a trial run. This paper discusses the development of our language model, the function and evaluation of this self assessment program. Chiharu Tsurutani, Yutaka Yamauchi, Nobuaki Minematsu, Dean Luo, Kazutaka Maruyama, Keikichi Hirose |
INTERSPEECH | 3 |
| 2006 | Factors affecting speakers² choice of fillers in Japanese presentationsabstractDisfluencies are intrinsic in spontaneous speech. Although it is known that there is a wide range of frequencies and types of disfluencies among speech, little is known about factors affecting speakers∂ choice of disfluency types. We first conducted a correspondence analysis using ratios of seven types of fillers, and other disfluencies, in 174 presentations and quantified the data. We conducted a cluster analysis using 4 dimension scores from the correspondence analysis, and extracted five filler-type groups. We then examined frequent types of presentations (formal or casual) and speaker attributes (gender and age) in each group. The results indicate that speakers∂choice of filler types is affected by speech levels, speakers∂gender and age, and that relevant factors differ depending on the type of fillers. Index Terms: disfluency type, fillers, sociolinguistic factors, speaker variation, Japanese Michiko Watanabe, Yasuharu Den, Keikichi Hirose, Shusaku Miwa, Nobuaki Minematsu |
INTERSPEECH | 5 |
| 2006 | Localization based audio source separation by sub-band beamformingabstractIn this paper, a localization based approach of audio signal separation from binary mixtures is carried out. The audio sources are localized in the spatial domain (azimuth plane) using the delay and amplitude variation cues between two microphones' signals. A coherence based technique is introduced here to localize the audio sources in adverse acoustical environment. The mixture signals are decomposed into a desired number of sub-bands with empirical mode decomposition (EMD) which is a data adaptive filtering scheme suitable for nonlinear and non-stationary signals. Data independent minimum variance beamforming is employed to separate the component sources in underdetermined condition (more sources than sensors). The experimental results of the proposed algorithm show noticeable separation efficiency. It is also found that the sub-band implementation improves the performance compared with and full-band approach Md. Khademul Islam Molla, Keikichi Hirose, Nobuaki Minematsu |
ISCAS | 3 |
| 2006 | Structural Representation of the pronunciation and its Use for CallabstractThis paper applies the structural representation of the pronunciation for computer-aided language learning (CALL). This representation was proposed to remove non-linguistic features such as age, gender, speaker, etc from speech acoustics (N. Minematsu et al., 2005). The removal was performed by extracting only the interrelations of speech events and discarding their absolute properties such as formants and spectrum envelopes. All the extracted interrelations mathematically form the external phonological structure of the events. Using this representation, in S. Asakawa et al., (2005), the vowel structure of a language learner was extracted and it was shown that the structural development via training can be traced and visualized adequately. This structural visualization can be regarded as pronunciation portfolio (N. Minematsu et al., 2004). This paper shows that the new representation can classify the language learners adequately and indicate which vowels should be corrected by priority. Nobuaki Minematsu, Satoshi Asakawa, Keikichi Hirose |
SLT | 1 |
| 2005 | Improved concept-to-speech generation in a dialogue system on road guidanceabstractAlthough in most spoken dialogue systems, text-to-speech conversion devices are used for reply speech generation. However, use of such devices makes it difficult to well reflect higher-level linguistic (and para-/non- linguistic) information obtainable during sentence generation process on reply speech. This situation degrades the reply speech quality mainly from the aspect of prosodic features. A method is necessary to directly converting content of reply into speech. This method, known as concept-to-speech conversion, was realized for the reply speech generation in our spoken dialogue system on road guidance. It is an improved version of our formerly developed one for an agent dialogue system. Reply sentence generation was conducted by pasting words and/or phrases at tag positions of a sentence frame, which was prepared in a tag-LISP form. In order to realize the concept-to-speech conversion, syntactic structure of phrases in user's inputs is kept and is utilized for the sentence generation. Several improvements, such as prosodic phrase boundary positioning using probability of word sequences, are also added to prosodic control in speech synthesis. In the spoken dialogue system, a user was guided to reach a place marked on a map through conversation. Several schemes on dialogue management were implemented to solve the problems caused due to the imperfect information on the roads given to the user and the system. A trial use of the system showed that a smooth conversation between the user and the system was possible. The result clearly indicated a better prosodic control for the newly developed method as compared to the original method. Yuji Yagi, Keikichi Hirose, Seiya Takada, Nobuaki Minematsu |
CW | 4 |
| 2005 | Mathematical Evidence of the Acoustic Universal Structure in SpeechabstractThe paper shows mathematically that there exists an acoustic universal structure in speech, which can be interpreted as a physical implementation of structural phonology. The structure has completely no dimensions of multiplicative and linear transformational distortions, which are inevitably involved in speech communication as differences of vocal tract shape, gender, age, microphone, room, line, hearing characteristics, and so on. A speech event, such as a phone, is probabilistically modeled as a distribution of parameters calculated by a linear transformation of a log spectrum, e.g., cepstrum. A set of events, such as a word, is relatively captured as structure composed of the distributions. An n-point structure is uniquely determined by fixing the lengths of its /sub n/C/sub 2/ diagonal lines, namely, the distance matrix among the n points. The distance between two distributions is calculated as a Bhattacharyya distance. The resulting structure has very interesting characteristics. Multiplicative and linear transformational distortions are geometrically interpreted as shift and rotation of the structure, respectively. This fact implies that there always exists a distortion-free communication channel between a speaker and a listener. Nobuaki Minematsu |
ICASSP (1) | 1 |
| 2005 | Structural representation of the non-native pronunciationsabstractAcoustic representation of speech provided by phonetics, spectrogram, is noisy representation in that it shows every acoustic aspect of speech. Age, gender, shape, microphone, room, line, etc. are completely irrelevant to the pronunciation assessment. However, the spectrogram is affected inevitably by these factors. Recently, a novel acoustic representation of speech was proposed, where dimensions of these non-linguistic factors can hardly be seen[1, 2]. Every acoustic substance of speech is discarded and only their interrelations are extracted to represent the pronunciation structurally. Using this method, individual learners were described as distorted phonemic structures[3] and automatic scoring of the pronunciation was investigated[3, 4]. This paper describes two new analyses using the proposed method. The first analysis is done to examine whether the method can trace the development of a student’s pronunciation appropriately using only a limited amount of speech. The second one focuses on the prosodic aspect of the pronunciation, especially stressed and unstressed vowels. The former indicates that the proposed method can show history of the student’s development adequately and the latter clarifies that size of the pronunciation structure is highly correlated with the pronunciation proficiency. 1. Satoshi Asakawa, Nobuaki Minematsu, Toshiko Isei-Jaakkola, Keikichi Hirose |
INTERSPEECH | 2 |
| 2005 | Corpus-based extraction of F0 contour generation process model parametersabstractA corpus-based method was developed for automatic extraction of the F0 contour generation process model parameters (phrase and accent commands). The method first smoothes the observed F0 contour by a piecewise 3rd order polynomial function and finds points of inflection. Then several parameters related to the points are used as input parameters for the predictor of the model commands. Finally the predicted commands are tuned to the observed contour by the analysis-by-synthesis. An experiment was conducted using ATR 504 sentence speech corpus, and the performance close to a rule-based method, also developed by the authors, was obtained. An experiment was further conducted by adding linguistic information of the content of the utterance (such as accent type, depth of bunsetsu boundary) to input parameters. The performance was largely improved; the extraction rates reached around 90 % for phrase commands and 84 % for accent commands. Keikichi Hirose, Yusuke Furuyama, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2005 | Multi-band approach of audio source discrimination with empirical mode decompositionabstractAbstract This paper presents a content-based approach of audio source indexing without any prior knowledge about the sources. The empirical mode decomposition ( EMD ) scheme, capable of decomposing nonlinear and non-stationary signals into some bases, is employed to implement the sub-band approach of the audio discrimination technique. The feature vectors are derived from each of selected sub-bands of the target frame. Linear predictive cepstrum coefficient ( LPCC ) is used as the main feature vector and Kullback-Leibler divergence ( KLd ) is performed as the scoring function to measure the similarity of the feature vectors. The higher order statistics ( HOS ) is employed to compute the LPCC . The use of HOS makes LPCC less affected by Gaussian noise. The experimental results show that the sub-band approach produces better discrimination efficiency than that of the full-band technique. This discrimination method is also suitable to solve the source permutation ambiguity in separation of multiple and concurrent moving sources from the mixture(s). Md. Khademul Islam Molla, Keikichi Hirose, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2005 | Japanese vowel recognition based on structural representation of speechabstractSpeech acoustics varies from speaker to speaker, microphone to microphone, room to room, line to line, etc. Physically speaking, every speech sample is distorted. Socially speaking, however, speech is the easiest communication media for humans. In order to cope with the inevitable distortions, speech engineers have built HMMs with speech data of hundreds or thousands of speakers and the models are called speaker-independent models. But they often need to be adapted to the input speaker or environment and this fact claims that the speaker-independent models are not really speaker-independent. Recently, a novel acoustic representation of speech was proposed, where dimensions of the above distortions can hardly be seen. It discards every acoustic substance of speech and captures only their interrelations to represent speech acoustics structurally. The new representation can be interpreted linguistically as physical implementation of structural phonology and also psychologically as speech Gestalt. In this paper, the first recognition experiment was carried out to investigate the performance of the new representation. The results showed that the new models trained from a single speaker with no normalization can outperform the conventional models trained from 4,130 speakers with CMN. Takao Murakami, Kazutaka Maruyama, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 3 |
| 2005 | Generation of fundamental frequency contours for Mandarin speech synthesis based on tone nucleus modelabstractA new method for generating sentence F 0 contours of Mandarin speech is proposed.The method assumes the F 0 contour generation process model, but generates the tone and phrase components in different ways and sums them to produce a sentence F 0 contour.The tone component is generated concatenating F 0 patterns of tone nuclei, which are predicted by a corpus-based scheme (binary decision trees).Experiments of F 0 contour generation were conducted by using 100 news utterances by a female speaker.The results showed that the method could generate F0 contours close to those of target speech.A perceptual evaluation was also conducted on the synthetic speech using the F0 contours generated by the method.An average score of 4.5 in a 5-point scale indicates the high naturalness, verifying the validity of the method. Qinghua Sun, Keikichi Hirose, Wentao Gu, Nobuaki Minematsu |
INTERSPEECH | 4 |
| 2005 | Filled pauses as cues to the complexity of following phrasesabstractCorpus based studies of spontaneous speech showed that filled pauses tended to precede relatively long and complex constituents. We examined whether listeners made use of such a tendency in speech processing. We tested the hypothesis that when listeners heard filled pauses they tended to expect a relatively long and complex phrase to follow. In the experiment participants listened to sentences referring to both simple and compound shapes presented on a computer screen. Their task was to press a button as soon as they had identified the shape that they heard. The sentences involved two factors: complexity and fluency. As the complexity factor, a half of the sentences described compound � shapes with long and complex phrases and�the other half described simple shapes with short and simple phrases.�As the fluency factor phrases describing a shape had a preceding filled pause, a preceding silent pause of the same length�as the filled pause, or no preceding pause. The results showed that response times for the complex phrases were significantly shorter after�filled or silent pauses than when there was no pause. In contrast, there was no significant difference between the three conditions for the simple phrases. The results support the hypothesis and indicate that it is the duration of filled pauses that give listeners cues to the complexity of upcoming phrases. 1. Michiko Watanabe, Keikichi Hirose, Yasuharu Den, Nobuaki Minematsu |
INTERSPEECH | 4 |
| 2005 | Synthesis of F0 contours using generation process model parameters predicted from unlabeled corpora: application to emotional speech synthesis
Keikichi Hirose, Kentaro Sato, Yasufumi Asano, Nobuaki Minematsu |
Speech Commun. | 4 |
| 2004 | Yet another acoustic representation of speech soundsabstractThis paper proposes yet another representation of speech sounds. The proposed speech modeling can remove both multiplicative and linear transformational distortion from speech theoretically. It means that speech sounds are represented without being affected by any static distortion inevitably involved in production, encoding, transmission, decoding, and hearing processes, such as differences in vocal tract length, gender, age, microphone, room, line, auditory characteristics, and so on. The method acoustically models not individual phones but their entire system, where only acoustic interrelation embedded in all the kinds of phones is focused. Since the method provides us with no absolute acoustic properties of phones, it cannot recognize or synthesize even a single phone. On the contrary, the proposed method is shown to be able to be applied to pronunciation assessment effectively and reliably, where the proficiency of pronunciation is estimated without using acoustic models of the individual phones directly in the matching. Nobuaki Minematsu |
ICASSP (1) | 1 |
| 2004 | N-gram language modeling of Japanese using bunsetsu boundariesabstractA new scheme of N-gram language modeling was pro-posed for Japanese, where word N-grams were calculated separately for the two cases: crossing and not crossing bunsetsu boundaries. Here, bunsetsu is a basic gram-matical (and pronunciation) unit of Japanese. A similar scheme using accent phrase boundaries instead of bun-setsu boundaries has already been proposed by the au-thors with a certain success, but it suffered from the train-ing data shortage, because assignment of accent phrase boundaries requires a speech corpus. In contrast, bun-setsu boundaries can be detected automatically from a written text with a rather high accuracy using a parser. It was shown from the experiment that a perplexity re-duction was possible by estimating bunsetsu boundaries from the history longer than N-1 words in the case of N-gram modeling and by selecting one from two types of models (crossing and not crossing bunsetsu bound-aries) according to the estimation. When 1 or 3 years of Mainichi Newspaper corpus was used for the training of tri-grams, the proposed scheme could reduce the perplex-ity by around 8 % from the baseline modeling (without separation). The proposed language modeling was ap-plied to a continuous speech recognition, and it showed that an improvement in word recognition rate was possi-ble especially when the training corpus was small (1 year of newspaper). 1. Sungyup Chung, Keikichi Hirose, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2004 | Use of prosodic features for speech recognitionabstractProsody is known to play an important role in human speech perception process. Therefore, there is an increasing need to use prosodic features for the advancement of speech recognition technology. However, prosody is related to various levels of information, from linguistic, para-linguistic, to non-linguistic, and, therefore, its acoustic manifestation is rather complicated with large variations. This fact prevents prosody to be incorporated in speech recognition process. In the current paper, discussions are given on how we can utilize prosodic features, showing our research works as examples. First, an idea of including word likelihood viewed from the accent type into the recognition process is shown. Second, a scheme of using prosody to control the pruning size in the decoding process is given. Prosodic features should be modeled rather differently form segmental features. Lastly, a new language model constructed by including prosodic events is explained. 1. Keikichi Hirose, Nobuaki Minematsu |
INTERSPEECH | 2 |
| 2004 | Pronunciation assessment based upon the compatibility between a learner's pronunciation structure and the target language's lexical structureabstractNative-sounding vs. intelligible. This has been a controver-sial issue for a long time in language learning and many teach-ers claim that the intelligible pronunciation should be the goal. What is the physical definition of the intelligibility? The cur-rent work shows a very good candidate answer to this ques-tion. The author proposed a new paradigm of observing speech acoustics based upon structural phonology, where all the kinds of speech events are viewed as an entire structure and this struc-ture was shown to be mathematically invariant with any static non-linguistic features such as age, gender, size, shape, micro-phone, room, line, and so on. This acoustic structure is purely linguistic and the phoneme-level structure is regarded as the pronunciation structure of individual students. This structure is matched with another linguistic structure, the lexical structure of the target language, and degree of compatibility between the two different levels of structures is calculated, which is defined as the intelligibility in this work. To increase the intelligibility, different instructions should be prepared for different students because no two students are the same. The proposed method can show the order of phonemes to learn, which is appropriate to a student and different from that of the others. 1. Nobuaki Minematsu |
INTERSPEECH | 1 |
| 2004 | Pronunciation assessment based upon the phonological distortions observed in language learners' utterancesabstractSpeech representation provided by acoustic phonetics, spectro-gram, is very noisy representation in that it shows every acoustic aspect of speech. Age, gender, size, shape, microphone, room and line are completely irrelevant to speech recognition, pro-nunciation assessment, and so on. But the spectrogram is af-fected easily by these factors. This is the very essential reason why speech systems are sometimes unreliable and the author supposes that the education should not endure this inevitable characteristics. The author proposed a novel method of acous-tic representation of speech where no dimensions of the above factors exist. The method was derived by implementing struc-tural phonology on physics. This paper examines whether the new representation of speech can provide a good tool of pro-nunciation assessment. Results of the experiments with good and intentionally-bad pronunciations of a single speaker showed that all the students are acoustically located between the two pronunciations, indicating that all the students are judged to be acoustically closer to the speaker than the speaker himself is. This result shows that the proposed method can delete the irrel-evant factors and is extremely reliable and effective in CALL. 1. Nobuaki Minematsu |
INTERSPEECH | 1 |
| 2004 | Audio source separation from the mixture using empirical mode decomposition with independent subspace analysisabstractIn this paper we decompose the Hilbert Spectrum of an audio mixture into a number of subspaces to segregate the sources. Empirical mode decomposition (EMD) together with Hilbert transform produces Hilbert spectrum (HS), which is a fine-resolution time-frequency representation of a non-stationary signal. EMD decomposes the mixture signal into some intrinsic oscillatory modes called intrinsic mode function (IMF). HS is constructed from the instantaneous frequency responses of IMFs. Some frequency independent basis vectors are derived using independent component analysis (ICA). Kulback-Laibler divergence based k-means clustering algorithm is proposed to group the basis vectors into number of desired sources. Then projecting HS on to the grouped basis vectors derives the independent source subspaces. The time domain source signals are assembled by applying some post processing on the subspaces. We have also produced some experimental results using our proposed separation algorithm. 1. Md. Khademul Islam Molla, Keikichi Hirose, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2004 | Clause types and filed pauses in Japanese spontaneous monologues
Michiko Watanabe, Yasuharu Den, Keikichi Hirose, Nobuaki Minematsu |
INTERSPEECH | 4 |
| 2003 | Use of linguistic information for automatic extraction of f_0 contour generation process model parametersabstractA method was developed to utilize linguistic information (lexical accent types and syntactic boundaries) to improve the performance of the automatic extraction of the F0 contour generation process model commands. The extraction scheme is first to smooth the observed F0 contour by a piecewise 3rd order polynomial function and to locate accent command positions by taking the derivative of the function. If the results of automatic extraction differ from those estimated from the linguistic information, they are modified according to the several rules. The results showed that some errors could be corrected by the use of linguistic information, especially when the initial word of an accent phrase is type 0 (flat) accent. As a whole, the correct extraction rate (recall rate) was increased from 79.8 % to 82.3 % for phrase commands and from 81.6 % to 85.9 % for accent commands. Keikichi Hirose, Yusuke Furuyama, Shuichi Narusawa, Nobuaki Minematsu, Hiroya Fujisaki |
INTERSPEECH | 4 |
| 2003 | A pronunciation training system for Japanese lexical accents with corrective feedback in learner's voiceabstractA system was developed for teaching non-Japanese learners pronunciation of Japanese lexical accents. The system first identifies word accent types in a learner's utterance using F0 change between two adjacent morae as the feature parameter. As for the representative F0 value for a mora, we defined one with a good match to the perceived pitch. The system notices the user if his/her pronunciation is good or not, and, then, generates audio and visual corrective feedbacks. Using TDPSOLA technique, the learner's utterance is modified in its prosodic features by referring to teacher's features, and offered to the learner as the audio corrective feedback. The visual feedback is also offered to enhance the modifications that occurred. Accent type pronunciation training experiments were conducted for 8 non-Japanese speakers, and the results showed that the training process could be facilitated by the feedbacks especially when they were asked to pronounce sentences. Keikichi Hirose, Frédéric Gendrin, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2003 | Corpus-based synthesis of fundamental frequency contours of Japanese using automatically-generated prosodic corpus and generation process modelabstractA corpus-based method of generating fundamental frequency (F0) contours of various speaking styles from text was developed. Instead of directly predicting F0 values, the method predicts command values of the F0 contour generation process model. Because of the model constraint, the resulting F0 contour keeps certain naturalness even when the prediction is done incorrectly. The method includes a scheme of automatic extraction of the model commands, which is necessary to prepare the training corpuses for various speaking styles. By introducing constraints on phrase command locations, a better extraction was realized, led to a better performance of the method. Speech synthesis was conducted using HMM speech synthesizer for calm speech and three types of emotional speech. The perceptual experiment showed the designated emotions could be well conveyed with the F0 contours generated by the developed method. Keikichi Hirose, Takayuki Ono, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2003 | Speech generation from concept for realizing conversation with an agent in a virtual roomabstractA concept to speech generation was realized in an agent dialogue system, where an agent (a stuffed animal) walked around in a small room constructed on a computer display to complete some jobs with instructions from a user. The communication between the user and the agent was done through speech. If the agent could not complete the job because of some difficulties, it tried to solve the problems through conversations with the user. Different from other spoken dialogue systems, the speech output from the agent was generated directly from the concept, and was synthesized using higher linguistic information. This scheme could largely improve the prosodic quality of speech output. In order to realize the concept to speech conversion, the linguistic information was handled as a tree structure in the whole dialogue process. Keikichi Hirose, Junji Tago, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2003 | CART-based factor analysis of intelligibility reduction in Japanese English
Nobuaki Minematsu, Changchen Guo, Keikichi Hirose |
INTERSPEECH | 1 |
| 2003 | Prosodic analysis and modeling of the NAGAUTA singing to synthesize its prosodic patterns from the standard notationabstractNAGAUTA (長唄) is a classical style of the Japanese singing. It has very original and unique prosodic patterns in its singing, where an abrupt and sharp change of F0 is always observed at a transition from a note to another. This F0 change is often found even where the transition is not accompanied by a change of tone. In this paper, we propose a model to synthesize this unique F0 pattern from the standard notation. Further, this paper shows an interesting phenomenon about power movements at the F0 changes. Acoustic analysis of NAGAUTA singing samples reveals that sharp increases of F0 and sharp decreases of power are observed synchronously. Although no discussion on physical mechanisms of this phenomenon is done in this paper, another model to generate this unique power pattern is also proposed. Evaluation experiments are done through listening and their results indicate high validity of the two proposed models. Nobuaki Minematsu, Bungo Matsuoka, Keikichi Hirose |
INTERSPEECH | 1 |
| 2003 | Improvement of non-native speech recognition by effectively modeling frequently observed pronunciation habitsabstractIn this paper, two techniques are proposed to enhance the nonnative (Japanese English) speech recognition performance. The first technique effectively integrates orthographic representation of a phoneme as an additional context in state clustering in training tied-state triphones. Non-native speakers often learned the target language not through their ears but through their eyes and it is easily assumed that their pronunciation of a phoneme may depend upon its grapheme. Here, correspondence between a vowel and its grapheme is automatically extracted and used as an additional context in the state clustering. The second technique elaborately couples a Japanese English acoustic model using triphones, mapping between the two models should be carefully trained because phoneme sets of both the models are different. Here, several phoneme recognition experiments are done to induce the mapping, and based upon the mapping, a tentative method of the coupling is examined. Results of LVCSR experiments show high validity of both the proposed methods. Nobuaki Minematsu, Koichi Osaki, Keikichi Hirose |
INTERSPEECH | 1 |
| 2003 | Automatic estimation of perceptual age using speaker modeling techniquesabstractThis paper proposes a technique to estimate speakers’ perceptual age automatically only with acoustic information of their utterances. Firstly, we experimentally collected data of how old individual speakers in databasessound to listeners. Speech samples of approximately 500 male speakers with a very wide range of the real age were presented to listeners, who were asked to estimate the age only by hearing. Using the results, the perceptual age of the individual speakers was defined in two ways as label (averaged age over the listeners) and distribution. Then, each of the speakers was acoustically modeled by GMMs. Finally, the perceptual age of an input speaker was estimated as weighted sum of the perceptual age of all the other speakers in the databases, where the weight for speaker i was calculated as a function of likelihood score of the input speaker as speaker i. Experiments showed that correlation was about 0.9 between the perceptual age estimated by the listening test and that estimated by the proposed method. This paper also introduces some techniques to realize robust estimation of the perceptual age. Nobuaki Minematsu, Keita Yamauchi, Keikichi Hirose |
INTERSPEECH | 1 |
| 2003 | Considerations on vowel durations for Japanese CALL systemabstractDue to various difficulties in pronunciation, utterances by nonnative speakers may be lacking in fluency. The Japanese pronunciation is said to have mora-synchronism, and, therefore, we assume that the disfluency may cause larger variations in vowel durations. Analyses of vowel (and CV) durations were conducted for Japanese sentence utterances by 2 non-Japanese speakers and one Japanese speaker (all female speakers). Larger variations were clearly observed in non-Japanese utterances. Then, 10 Japanese speakers were asked to rate the non-Japanese utterances. Strong negative correlations were observed between durational variations and pronunciation ratings. Based on the result, a method was developed for automatic evaluation of nonJapanese utterances. The ratings by the method were shown to be close to those by native speakers. Also, in order to offer a corrective feedback in learner’s voice, non-Japanese utterances were modified in their vowel durations by referring to native Japanese utterances. The modification was done using TD–PSOLA scheme. The result of listening test indicated some improvements in nativeness. Taro Mouri, Keikichi Hirose, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2003 | Estimation of resonant characteristics based on AR-HMM modeling and spectral envelope conversion of vowel soundsabstractA new method was developed for accurately separating source and articulation filter characteristics of speech. This method is based on the AR-HMM modeling, where the residual waveform is expressed as the output sequence from an HMM. To realize an accurate analysis, a scheme of dividing HMM state was newly introduced. Using the AR-filter parameter values obtained through the analysis, we can construct a vocoder-type formant synthesizer, where the residual waveform is used as the excitation source. Through the listening test on the vowel sounds synthesized using AR-filter from a vowel and excitation waveform from another vowel, it was shown that a “flexible” synthesis with a high controllability on the acoustic parameters were possible by our formant synthesis configuration. Nobuyuki Nishizawa, Keikichi Hirose, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2003 | Mora F0 representation for accent type identification in continuous speech and considerations on its relation with perceived pitch values
Carlos Toshinori Ishi, Keikichi Hirose, Nobuaki Minematsu |
Speech Commun. | 3 |
| 2003 | Data-driven generation of F0 contours using a superpositional model
Atsuhiro Sakurai, Keikichi Hirose, Nobuaki Minematsu |
Speech Commun. | 3 |
| 2002 | Automatic estimation of one's age with his/her speech based upon acoustic modeling techniques of speakersabstractThis paper proposes a technique which automatically estimates speakers' age only with acoustic, not linguistic, information of their utterances. This method is based upon speaker recognition techniques. In the current work, we firstly divided speakers of two databases, JNAS and S(senior)-JNAS, into two groups by listening tests. One group has only the speakers whose speech sounds so aged that one should take special care when he/she talks to them. The other group has the remaining speakers of the two databases. After that, each speaker group was modeled with GMM. Experiments of automatic identification of elderly speakers showed the correct identification rate of 91 %. To improve the performance, two prosodic features were considered, i.e, speech rate and local perturbation of power. Using these features, the identification rate was improved to 95%. Finally, using scores calculated by integrating GMMs with prosodic features, experiments were carried out to automatically estimate speakers' age. The results showed high correlation between speakers' age estimated subjectively by humans and automatically calculated score of ‘agedness’. Nobuaki Minematsu, Mariko Sekiguchi, Keikichi Hirose |
ICASSP | 1 |
| 2002 | A method for automatic extraction of model parameters from fundamental frequency contours of speechabstractThe process of generating the F0contour of speech has been modeled quite accurately in mathematical tenns by Fujisaki and his coworkers, but the extraction of parameters of the underlying commands from an observed F0contour is an inverse problem that can be solved only by successive approximation. In order to guarantee an efficient and accurate search for the solution, one needs to start with a set of initial values that are close enough to the optimum. This paper presents a method for pre-processing a measured F0contour to obtain its approximation consisting of third-order polynomial segments that are continuous and differentiable everywhere. It is shown that the proposed method allows one to obtain first-order approximations to the parameters of accent commands for about 90% of all the accent commands, and of phrase commands for about 84% of all the phrase commands. Shuichi Narusawa, Nobuaki Minematsu, Keikichi Hirose, Hiroya Fujisaki |
ICASSP | 2 |
| 2002 | Improved corpus-based synthesis of fundamental frequency contours using generation process modelabstractA corpus-based method of generating fundamental frequency (F 0) contours of various speaking styles from text was developed. Instead of directly predicting F 0 values, the method predicts command values of the F 0 contour generation process model. Because of the model constraint, the resulting F 0 contour keeps certain naturalness even when the prediction is done incorrectly. The method includes a scheme of automatic extraction of the model commands, which is necessary to prepare the training corpuses for various speaking styles. By introducing constraints on phrase command locations, a better extraction was realized, led to a better performance of the method. Speech synthesis was conducted using HMM speech synthesizer for calm speech and three types of emotional speech. The perceptual experiment showed the designated emotions could be well conveyed with the F 0 contours generated by the developed method. 1. Keikichi Hirose, Masaya Eto, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2002 | Statistical language modeling with prosodic boundaries and its use for continuous speech recognition
Keikichi Hirose, Nobuaki Minematsu, Makoto Terao |
INTERSPEECH | 2 |
| 2002 | Robust speech recognition using inter-speaker and intra-speaker adaptation
Baojie Li, Keikichi Hirose, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2002 | Integration of MLLR adaptation with pronunciation proficiency adaptation for non-native speech recognitionabstractTo recognize non-native speech, larger acoustic/linguistic distortions must be handled adequately in acoustic modeling, language modeling, lexical modeling, and/or decoding strategy. In this paper, a novel method to enhance MLLR adaptation of acoustic models for non-native speech recognition is proposed. In the case of native speech recognition, MLLR speaker adaptation was successfully introduced because it enables efficient adaptation with a small number of adaptation data by using a regression tree of Gaussian mixtures of HMMs. However, as for non-native speech, most of the cases, the regression tree built from the baseline HMMs does not match with pronunciation proficiency of a speaker. This paper provides a solution for this problem, where the speaker’s proficiency is automatically estimated and the tree suited for the proficiency is built, which can be viewed as proficiency adaptation. Recognition experiments show that MLLR with the new tree raises the averaged error reduction rate up to about 30 % from the baseline MLLR performance of approximately 20 %. Nobuaki Minematsu, Gakuto Kurata, Keikichi Hirose |
INTERSPEECH | 1 |
| 2002 | Corpus-based analysis of English spoken by Japanese students in view of the entire phonemic system of English
Nobuaki Minematsu, Gakuto Kurata, Keikichi Hirose |
INTERSPEECH | 1 |
| 2002 | Acoustic modeling of sentence stress using differential features between syllables for English rhythm learning system development
Nobuaki Minematsu, Satoshi Kobashikawa, Keikichi Hirose, Donna Erickson |
INTERSPEECH | 1 |
| 2002 | Automatic extraction of model parameters from fundamental frequency contours of English utterancesabstractThe generation process model of the fundamental frequency contours (F0 contours) of speech is known to be capable of generating F0 contours quite close to observed ones. The extraction of model parameters from an observed contour, however, requires an iterative process starting from a set of initial parameter values. In order to guarantee a rapid convergence to an optimum solution, the values should be appropriate ones. We already have developed a method of automatically extracting these from given F0 contours, and applied it to Japanese sentences with good results. The method is based on approximating an observed contour by a continuous curve differentiable everywhere. In the present paper, it was applied to English utterances. Experiments were conducted for 4 native speakers’ utterances with 14.5% and 17.5% of average miss and false alarm rates for the accent commands, and 35.7% and 15.5% for the phrase commands. Shuichi Narusawa, Nobuaki Minematsu, Keikichi Hirose, Hiroya Fujisaki |
INTERSPEECH | 2 |
| 2002 | Separation of voiced source characteristics and vocal tract transfer function characteristics for speech sounds by iterative analysis based on AR-HMM modelabstractA new method was developed for the separation of source and transfer function characteristics of speech sounds, with an aim of utilizing it to “flexible ” speech synthesis. The method is based on representing source waveform by an HMM, and transfer function by the AR process (AR-HMM model). As compared to methods based on ARX model, where a parametric representation is assumed for source waveform, a better separation is possible. By introducing a process of recursively deleting real poles of AR filters, which represent source waveform features, and including them into HMM source waveform, the resulting AR filters may correctly represent transfer function features. Experiments were conducted for Japanese vowel sounds in continuous speech, and the results were compared with those by conventional LP analysis and AR-HMM model analysis without recursive process. After representing obtained source and transfer function features respectively as DFT cepstrum and LPC cepstrum, variations of cepstrum parameters for each vowel sound were compared for the three analysis methods. The smallest variations were obtained by the proposed method, indicating that the proposed method can separate source and transfer function features well, and, thus, has potential ability of generating good quality of speech when applied to “flexible” speech synthesis. 1. Nobuyuki Nishizawa, Keikichi Hirose, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2002 | English Speech Database Read by Japanese Learners for CALL System Development
Nobuaki Minematsu, Yoshihiro Tomiyama, Kei Yoshimoto, Katsumasa Shimizu, Seiichi Nakagawa, Masatake Dantsuji, Shozo Makino |
LREC | 1 |
| 2001 | Generation of F0 contours using a model-constrained data-driven methodabstractIntroduces a model-constrained, data-driven method for generating fundamental frequency contours in Japanese text-to-speech synthesis. In the training phase, the parameters of a command-response F/sub 0/ contour generation model are learned by a prediction module, which can be a neural network or a set of binary regression trees. The input features consist of linguistic information related to accentual phrases that can be automatically derived from text, such as the position of the accentual phrase in the utterance, number of morae, accent type, and parts-of-speech. In the synthesis phase, the prediction module is used to generate appropriate values of model parameters. The use of the parametric model restricts the degrees of freedom of the problem, facilitating data-driven learning. Experimental results show that the method makes it possible to generate quite natural F/sub 0/ contours with a relatively small training database. Atsuhiro Sakurai, Keikichi Hirose, Nobuaki Minematsu |
ICASSP | 3 |
| 2001 | Corpus-based synthesis of fundamental frequency contours based on a generation process modelabstractA mode-constrained corpus-based synthesis strategy was developed for fundamental frequency (F 0 ) contours of Japanese sentences.In the training phase, the relationship between linguistic factors and the command values (amplitudes and locations) of F 0 contour generation process model was learned for a prediction module; a neural network in the current paper.Input parameters consist of linguistic information related to accentual phrases that can be automatically driven from text, such as the position of the accentual phrase in the utterance, number of morae, accent type, and morphological information.In the synthesis phase, the prediction module is used to generate command values of the model.The synthesis method was also realized based on multiple linear regression analysis to examine how each input parameter contributes to the F 0 contour generation.The use of the parametric model restricts the degrees of freedom of the mapping between linguistic and prosodic features, and thus enables to generate appropriate values even with limited training data.Experimental results showed that the method could generate F 0 contours quite close to those by the rulebased method. Keikichi Hirose, Masaya Eto, Nobuaki Minematsu, Atsuhiro Sakurai |
INTERSPEECH | 3 |
| 2001 | Identification of accent and intonation in sentences for CALL systemsabstractIn order to construct a CALL (Computer Aided Language Learning) system that can teach learners accent and intonation of Japanese, it's necessary to automatically identify accent types and intonation types in sentence utterances.For this purpose, several acoustic (prosodic) features of speech were investigated taking their effects on human perception into account.For the accent type identification method, the use of average values of F0 in mora and target values of F0 in mora final was evaluated in CV and VC units.Average values of VC units and target values of CV units showed better performance in the identification task.As for the intonation identification, several acoustic features were investigated to represent 6 types of sentence final tones, each conveying different information of intention and perceptual impression.The proposed acoustic features for relative duration and sentence final pitch change showed good correspondence to perceptual features. Carlos Toshinori Ishi, Nobuaki Minematsu, Ryuji Nishide, Keikichi Hirose |
INTERSPEECH | 2 |
| 2001 | Use of topic knowledge in spoken dialogue information retrieval system for academic documentsabstractAn efficient search function based on topic estimation was integrated to our spoken dialogue system for academic document information retrieval. The following two points were mainly studied: 1) to properly categorize documents (to be retrieved) into related topics, and 2) to facilitate retrieval process using topic knowledge. For the first point, a method was developed to calculate recursively the relevance scores of retrieval words and documents for topics. Effects of the recursive process were proved through experimental results; better classification of retrieval words and documents into topics was realized. As for the second point, retrieval range was limited into topics estimated from retrieval words. It was shown through experiments of retrieval task solving that necessary number of dialogue turns (therefore, period of dialogue) could be largely reduced by the range limitation; a smooth retrieval process was proved to be realized using topic knowledge. Shinya Kiriyama, Keikichi Hirose, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2001 | Instantaneous estimation of accentuation habits for Japanese students to learn English pronunciationabstractMore and more eorts have been recently made to apply speech technologies to language learning[1]{ [4].The authors have been especially focusing on Japanese manners of generating English word stress.This is because accentuation habits inevitable to Japanese learners can be easily found in their stress generation.In our previous studies, a stressed syllable detector and an accentuation habit estimator were developed[5]{ [7], where the estimated habits of individual learners accorded well with their English pronunciation prociency rated by English teachers.However, the estimation methods in our previous studies required several dozens of word utterances or a relatively large amount of computation even when a single word utterance enabled the estimation.In this paper, we i n v estigated a method which required only a single word utterance with a small computation cost.Results showed that similar tendencies can be found between the habits estimated in our previous study and those in the current one. Naoki Nakamura, Nobuaki Minematsu, Seiichi Nakagawa |
INTERSPEECH | 2 |
| 2000 | Analytical and perceptual study on the role of acoustic features in realizing emotional speechabstractInvestigation was conducted on how prosodic features of emotional speech changed depending on emotion levels. The analysis results on fundamental frequency (F0) contours and speech rates implied that humans have several ways to express emotions and use them rather randomly. Investigation was also conducted on what acoustic features were important to express emotions. Perceptual experiments using synthetic speech with copied acoustic features of target speech indicated importance of the segmental features other than the prosodic features. Especially, a high importance was observed in the case of happiness. 1. Keikichi Hirose, Nobuaki Minematsu, Hiromichi Kawanami |
INTERSPEECH | 2 |
| 2000 | Identification of Japanese double-mora phonemes considering speaking rate for the use in CALL systemsabstractFig.1. Two utterances “Sorewa oQtodesu ” (fast speech) and “Sorewa otodesu ” (slow speech) of a same speaker. The duration of the long phone /Qt / in the upper utterance is shorter than the short phone /t / in the lower utterance, showing the danger of using absolute values. Carlos Toshinori Ishi, Keikichi Hirose, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2000 | Free software toolkit for Japanese large vocabulary continuous speech recognitionabstractA sharable software repository for Japanese LVCSR (Large Vocabulary Continuous Speech Recognition) is introduced. It is designed as a baseline platform for research and developed by researchers of different academic institutes under a governmental support. The repository consists of a recognition engine (Julius), Japanese acoustic models and statistical language models as well as Japanese morphological analysis tools. These modules can be easily integrated and replaced under a plug-and-play framework, which makes it possible to fairly evaluate components and to develop specific application systems. Assessment of these modules and systems in a 20000-word dictation task is reported. The software repository is freely available to the public. Tatsuya Kawahara, Akinobu Lee, Tetsunori Kobayashi, Kazuya Takeda, Nobuaki Minematsu, Shigeki Sagayama, Katunobu Itou, Akinori Ito, Mikio Yamamoto, Atsushi Yamada, Takehito Utsuro, Kiyohiro Shikano |
INTERSPEECH | 5 |
| 2000 | Efficient search strategy in large vocabulary continuous speech recognition using prosodic boundary information
Shi-wook Lee, Keikichi Hirose, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2000 | Modeling phone correlation for speaker adaptive speech recognitionabstractInformation of phone relationships is regarded as acting an important role in speech recognition. It has been successfully exploited in many speaker adaptation approaches. In this paper, we propose a new approach, named Phone Pair Model (PPM) re-scoring, to utilize phone relationships for speaker-adaptive speech recognition. PPM re-scoring approach does not really adapt model parameters to a new speaker. It just uses some pre-registered phones ' samples from the speaker being recognized, to re-calculate the likelihood of phones that has been calculated on conventional phone HMMs, resulting in a more correct recognition result. Additionally, it can deal with not only inter-speaker acoustic variations but also intra-speaker acoustic variations adequately. Results of two recognition experiments, one using phone HMMs only and the other incorporating phone HMMs with the PPMs, showed that even by using only a few vowel samples as the pre-registered phones, PP-M re-scoring approach brought an increase in recognition rate. 1. Baojie Li, Keikichi Hirose, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2000 | Performance comparison among HMM, DTW, and human abilities in terms of identifying stress patterns of word utterancesabstractWe have been focusing on applying speech technologies to pronunciation learning. In our previous study[1], a stressed syllable detector was implemented by using stressed syllable HMMs and unstressed ones. And using the detector internally, several systems were implemented[2]. However, their development did not necessarily require the use of HMMs as an acoustic modeling method. In this paper, an HMM-based method, a DTW-based method, and a human strategy only with visual inspection were compared in terms of their performance in judging whether two utterances of a word have the same stress pattern, e.g. récord and recórd. Here, one utterance was given by a Japanese learner and the other one was done by a native speaker. Experiments showed that HMMs gave us the higher performance Nobuaki Minematsu, Yukiko Fujisawa, Seiichi Nakagawa |
INTERSPEECH | 1 |
| 2000 | Quality improvement of PSOLA analysis-synthesis using partial zero-phase conversionabstractThis paper discusses two issues of the quality improvement of F0 modified speech based upon PSOLA analysissynthesis. Previous studies[1][2] pointed out that the location of a window of PSOLA influences the quality of synthesized speech and one of them claimed that the center of a window should be located at a pitch pulse in source waveforms. However, pitch pulse detection sometimes fails due to undesired acoustic events. In this paper, several methods are experimentally examined to reduce pitch pulse detection errors. Even when the detection is done correctly, F0 modified re-synthesized speech sometimes causes “echoes ” in the re-arranged waveforms. This is mainly caused by a pitch pulse with small sharpness or by that with two relatively high pulses, not pitch pulses, before and after it. To suppress the echoes with little loss of naturalness, partial zero/π-phase conversion is proposed here. Experiments show the high validity of the proposed methods in improving the quality of re-synthesized speech. 1. Nobuaki Minematsu, Seiichi Nakagawa |
INTERSPEECH | 1 |
| 2000 | Instantaneous estimation of prosodic pronunciation habits for Japanese students to learn English pronunciationabstractMore and more efforts have been recently made to apply speech technologies to language learning and develop CALL systems[1]–[4]. The authors have been focusing on Japanese manners of generating English word stress. This is because pronunciation habits which are inevitable to Japanese learners can be easily found in the stress generation. In our previous study, a stressed syllable detector and a pronunciation habit estimator were developed[5], where the estimated habits of individual learners accorded well with their English pronunciation proficiency rated by English teachers. However, the habit could be estimated only after a learner pronounced several dozens of words because the habit estimator referred to stressed syllable detection rates. In this paper, using another criterion, a method of instantaneous estimation was proposed which required only a single word utterance. Results showed that an average pattern of the instantaneously estimated habits accorded well with the habits obtained in our previous study. Nobuaki Minematsu, Seiichi Nakagawa |
INTERSPEECH | 1 |
| 2000 | Development of a formant-based analysis-synthesis system and generation of high quality liquid sounds of JapaneseabstractAlthough flexible control of acoustic features is possible in formant-based speech synthesizers, their development requires precise estimation of parameters related to vocal tract and source. This requirement is difficult to satisfy and often results in limiting quality of the synthesized speech. The difficulty is derived from the fact that estimation of the parameters is a nonlinear problem. Therefore, the completely automatic estimation of the parameters is quite difficult and some approximations or manual modifications of parameters with aprioriknowledge are required in the development. In this study, mainly to make the estimation more efficient and/or to assist developers doing the manual modifications of parameters, a formant-based analysissynthesis system is build. The system introduces pitchsynchronous acoustic analysis to reduce fluctuation of the estimated parameters. Experiments show that quality of synthetic speech of Japanese /r / sounds is significantly improved by using the proposed system. 1. Nobuyuki Nishizawa, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 2 |
| 2000 | Data-driven intonation modeling using a neural network and a command response model
Atsuhiro Sakurai, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 2 |
| 2000 | IPA Japanese Dictation Free Software Project
Katunobu Itou, Kiyohiro Shikano, Tatsuya Kawahara, Kazuya Takeda, Atsushi Yamada, Akinori Ito, Takehito Utsuro, Tetsunori Kobayashi, Nobuaki Minematsu, Mikio Yamamoto, Shigeki Sagayama, Akinobu Lee |
LREC | 9 |
| 1998 | Evaluation of Japanese manners of generating word accent of English based on a stressed syllable detection technique
Yukiko Fujisawa, Nobuaki Minematsu, Seiichi Nakagawa |
ICSLP | 2 |
| 1998 | Continuous speech recognition using segmental unit input HMMs with a mixture of probability density functions and context dependency
Kengo Hanai, Kazumasa Yamamoto, Nobuaki Minematsu, Seiichi Nakagawa |
ICSLP | 3 |
| 1998 | Sharable software repository for Japanese large vocabulary continuous speech recognitionabstractThe project of Japanese LVCSR (Large Vocabulary Continuous Speech Recognition) platform is introduced. It is a collaboration of researchers of different academic institutes and intended to develop a sharable software repository of not only databases but also models and programs. The platform consists of a standard recognition engine, Japanese phone models and Japanese statistical language models. A set of Japanese phone HMMs are trained with ASJ (Acoustic Society of Japan) databases of 20K sentence utterances per each gender. Japanese word N-gram (2-gram and 3-gram) models are constructed with a corpus of Mainichi newspaper of four years. The recognition engine JULIUS is developed for assessment of both acoustic and language models. The modules are integrated as a Japanese LVCSR system and evaluated on 5000-word dictation task. The software repository is available to the public. Tatsuya Kawahara, Tetsunori Kobayashi, Kazuya Takeda, Nobuaki Minematsu, Katunobu Itou, Mikio Yamamoto, Atsushi Yamada, Takehito Utsuro, Kiyohiro Shikano |
ICSLP | 4 |
| 1998 | Modeling of variations in cepstral coefficients caused by F0 changes and its application to speech processingabstractIn this paper, the correlation between spectral variations and F0 changes in a vowel sound is rstly analyzed, where the variations are also compared to VQ distortions calculated in a ve-vowel space. It is shown that the F0 change approximately by a half octave produces the spectral variation comparable to the averaged VQ distortion when the codebook size is the number of the vowels. Next, a model to predict the cepstral coe cients' variations caused by the F0 changes is built based on the multivariate regression analysis. Experiments show that the generated frame by the model has a remarkably small distance to the target frame and that the distance is almost the same as the VQ distortion with the codebook size being 10 to 20. Furthermore, the model is evaluated separately in terms of a spectral envelope predictor with a given F0 and a mapping function of feature sub-spaces. It is indicated that, while the models should be built dependently on phonemes and speakers as a spectrum predictor, adequate selection of parameters can enable the speaker/phoneme-independent models to work e ectively as a mapping function. Nobuaki Minematsu, Seiichi Nakagawa |
ICSLP | 1 |
| 1997 | Automatic detection of accent in English words spoken by Japanese students
Nobuaki Minematsu, Nariaki Ohashi, Seiichi Nakagawa |
EUROSPEECH | 1 |
| 1996 | Automatic detection of accent nuclei at the head of words for speech recognitionabstractA new scheme is proposed to incorporate prosodic processing into speech recognition, where the accent n uclei at the head of words are detected automatically and used to limit the searching space in speech recognition, that is, to preselect candidate words.Especially in this paper, the proposed method for the automatic detection of the accent n uclei and its performance are described.Using this scheme, it is expected that the recognition speed is improved.This scheme is derived from a nding by perceptual experiments conducted previously by the rst author.Results of the experiments indicated that the accent n ucleus at the rst mora has acceleration eect on perceiving the word.This effect can be explained by the earlier identication of the word accent t ype as type 1 b y its nucleus at the rst mora.In other words, the accent n ucleus at the head of a word can limit the searching space eectively in the mental lexicon.This mechanism was implemented using HMMs and examined for isolated words on a machine, where the vowel detection by broad segmental features and the rejection of words with a devoiced vowel at the rst or second mora were introduced at the same time.Evaluation experiments showed 94.7% and 90.0% as recall factor and precision factor of the accent n ucleus detection respectively. Nobuaki Minematsu, Seiichi Nakagawa |
ICSLP | 1 |
| 1996 | Prosodic manipulation system of speech material for perceptual experiments
Nobuaki Minematsu, Seiichi Nakagawa, Keikichi Hirose |
ICSLP | 1 |
| 1994 | Speech recognition using HMM with decreased intra-group variation in the temporal structure
Nobuaki Minematsu, Keikichi Hirose |
ICSLP | 1 |
| 1994 | Role of prosodic features in the human process of speech perception
Nobuaki Minematsu, Keikichi Hirose |
ICSLP | 1 |
| 1992 | The influence of semantic and syntactic information on spoken sentence recognition
Nobuaki Minematsu, Sumio Ohno, Keikichi Hirose, Hiroya Fujisaki |
ICSLP | 1 |
| 1990 | Influence of context and knowledge on the perception of continuous speech
Hiroya Fujisaki, Keikichi Hirose, Sumio Ohno, Nobuaki Minematsu |
ICSLP | 4 |