VLDB 2026 Research / reviewers in the wild / expert
Shinsuke Sakai
dblp:57/448
· DBLP profile ↗
38ranked-venue papers
8as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 35 · 8 first-author · 5 since 2021Artificial intelligence and machine learning · 27 · 5 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-lingual and Zero-Shot Speech Recognition by Incorporating Classification of Language-Independent Articulatory Features
Ryo Magoshi, Shinsuke Sakai, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2024 | Efficient and Robust Long-Form Speech Recognition with Hybrid H3-ConformerabstractRecently, Conformer has achieved state-of-the-art performance in many speech recognition tasks.However, the Transformer-based models show significant deterioration for long-form speech, such as lectures, because the self-attention mechanism becomes unreliable with the computation of the square order of the input length.To solve the problem, we incorporate a kind of state-space model, Hungry Hungry Hippos (H3), to replace or complement the multi-head self-attention (MHSA).H3 allows for efficient modeling of long-form sequences with a linear-order computation.In experiments using two datasets of CSJ and LibriSpeech, our proposed H3-Conformer model performs efficient and robust recognition of long-form speech.Moreover, we propose a hybrid of H3 and MHSA and show that using H3 in higher layers and MHSA in lower layers provides significant improvement in online recognition.We also investigate a parallel use of H3 and MHSA in all layers, resulting in the best performance. Tomoki Honda, Shinsuke Sakai, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2022 | Non-autoregressive Error Correction for CTC-based ASR with Phone-conditioned Masked LM
Hayato Futami, Hirofumi Inaguma, Sei Ueno, Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
INTERSPEECH | 5 |
| 2021 | ASR Rescoring and Confidence Estimation with ElectraabstractIn automatic speech recognition (ASR) rescoring, the hypothesis with the fewest errors should be selected from the$n$-best list using a language model (LM). However, LMs are usually trained to maximize the likelihood of correct word sequences, not to detect ASR errors. We propose an ASR rescoring method for directly detecting errors with ELECTRA, which is originally a pre-training method for NLP tasks. ELECTRA is pre-trained to predict whether each word is replaced by BERT or not, which can simulate ASR error detection on large text corpora. To make this pre-training closer to ASR error detection, we further propose an extended version of ELECTRA called phone-attentive ELECTRA (P-ELECTRA). In the pre-training of P-ELECTRA, each word is replaced by a phone-to-word conversion model, which leverages phone information to generate acoustically similar words. Since our rescoring method is optimized for detecting errors, it can also be used for word-level confidence estimation. Experimental evaluations on the Librispeech and TED-LIUM2 corpora show that our rescoring method with ELECTRA is competitive with conventional rescoring methods with faster inference. ELECTRA also performs better in confidence estimation than BERT because it can learn to detect inappropriate words not only in fine-tuning but also in pre-training. Hayato Futami, Hirofumi Inaguma, Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
ASRU | 4 |
| 2021 | Data Augmentation for ASR Using TTS Via a Discrete RepresentationabstractWhile end-to-end automatic speech recognition (ASR) has achieved high performance, it requires a huge amount of paired speech and transcription data for training. Recently, data augmentation methods have actively been investigated. One method is to use a text-to-speech (TTS) system to gen-erate speech data from text-only data and use the generated speech for data augmentation, but it has been found that the synthesized log Mel-scale filterbank (lmfb) features could have a serious mismatch with the real speech features. In this study, we propose a data augmentation method via a discrete speech representation. The TTS model predicts discrete ID sequences instead of lmfb features, and the ASR also uses the ID sequences as training data. We expect that the use of a discrete representation based on vq-wav2vec not only makes TTS training easier but also mitigates the mismatch with real data. Experimental evaluations show that the pro-posed method outperforms the data augmentation method using the conventional TTS. We found that it reduces speaker dependency, and the generated features are distributed more closely to the real ones. Sei Ueno, Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
ASRU | 3 |
| 2020 | Distilling the Knowledge of BERT for Sequence-to-Sequence ASRabstractAttention-based sequence-to-sequence (seq2seq) models have achieved promising results in automatic speech recognition (ASR).However, as these models decode in a left-to-right way, they do not have access to context on the right.We leverage both left and right context by applying BERT as an external language model to seq2seq ASR through knowledge distillation.In our proposed method, BERT generates soft labels to guide the training of seq2seq ASR.Furthermore, we leverage context beyond the current utterance as input to BERT.Experimental evaluations show that our method significantly improves the ASR performance from the seq2seq baseline on the Corpus of Spontaneous Japanese (CSJ).Knowledge distillation from BERT outperforms that from a transformer LM that only looks at left context.We also show the effectiveness of leveraging context beyond the current utterance.Our method outperforms other LM application approaches such as n-best rescoring and shallow fusion, while it does not require extra inference cost. Hayato Futami, Hirofumi Inaguma, Sei Ueno, Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
INTERSPEECH | 5 |
| 2020 | Generative Adversarial Training Data Adaptation for Very Low-Resource Automatic Speech RecognitionabstractIt is important to transcribe and archive speech data of endangered languages for preserving heritages of verbal culture and automatic speech recognition (ASR) is a powerful tool to facilitate this process. However, since endangered languages do not generally have large corpora with many speakers, the performance of ASR models trained on them are considerably poor in general. Nevertheless, we are often left with a lot of recordings of spontaneous speech data that have to be transcribed. In this work, for mitigating this speaker sparsity problem, we propose to convert the whole training speech data and make it sound like the test speaker in order to develop a highly accurate ASR system for this speaker. For this purpose, we utilize a CycleGAN-based non-parallel voice conversion technology to forge a labeled training data that is close to the test speaker's speech. We evaluated this speaker adaptation approach on two low-resource corpora, namely, Ainu and Mboshi. We obtained 35-60% relative improvement in phone error rate on the Ainu corpus, and 40% relative improvement was attained on the Mboshi corpus. This approach outperformed two conventional methods namely unsupervised adaptation and multilingual training with these two corpora. Kohei Matsuura, Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
INTERSPEECH | 3 |
| 2020 | Speech Corpus of Ainu Folklore and End-to-end Speech Recognition for Ainu LanguageabstractAinu is an unwritten language that has been spoken by Ainu people who are one of the ethnic groups in Japan. It is recognized as critically endangered by UNESCO and archiving and documentation of its language heritage is of paramount importance. Although a considerable amount of voice recordings of Ainu folklore has been produced and accumulated to save their culture, only a quite limited parts of them are transcribed so far. Thus, we started a project of automatic speech recognition (ASR) for the Ainu language in order to contribute to the development of annotated language archives. In this paper, we report speech corpus development and the structure and performance of end-to-end ASR for Ainu. We investigated four modeling units (phone, syllable, word piece, and word) and found that the syllable-based model performed best in terms of both word and phone recognition accuracy, which were about 60% and over 85% respectively in speaker-open condition. Furthermore, word and phone accuracy of 80% and 90% has been achieved in a speaker-closed setting. We also found out that a multilingual ASR training with additional speech corpora of English and Japanese further improves the speaker-open test accuracy. Kohei Matsuura, Sei Ueno, Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
LREC | 4 |
| 2019 | Multi-speaker Sequence-to-sequence Speech Synthesis for Data Augmentation in Acoustic-to-word Speech RecognitionabstractThe acoustic-to-word (A2W) automatic speech recognition (ASR) realizes very fast decoding with a simple architecture and achieves state-of-the-art performance. However, the A2W model suffers from the out-of-vocabulary (OOV) word problem and cannot use text-only data to improve the language modeling capability. Meanwhile, sequence-to-sequence neural speech synthesis has also been developed and achieved naturalness comparable to human speech. We investigate leveraging sequence-to-sequence neural speech synthesis to augment training data for the ASR system in a target domain. While speech synthesis model is usually trained with single speaker data, ASR needs to cover a variety of speakers. In this work, we extend the speech synthesizer so that it can output speech of many speakers. The multi-speaker speech synthesizer is trained with a large corpus in the source domain, then used to generate acoustic features from texts of the target domain. These synthesized speech features are combined with real speech features of the source domain to train an attention-based A2W model. Experimental results show that the A2W model trained with the multi-speaker model achieved a significant improvement over the baseline and the single speaker model. Sei Ueno, Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
ICASSP | 3 |
| 2018 | Forward-Backward Attention Decoder
Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2018 | Encoder Transfer for Attention-based Acoustic-to-word Speech Recognition
Sei Ueno, Takafumi Moriya, Masato Mimura, Shinsuke Sakai, Yusuke Shinohara, Yoshikazu Yamaguchi, Yushi Aono, Tatsuya Kawahara |
INTERSPEECH | 4 |
| 2018 | Improving OOV Detection and Resolution with External Language Models in Acoustic-to-Word ASRabstractAcoustic-to-word (A2W) end-to-end automatic speech recognition (ASR) systems have attracted attention because of an extremely simplified architecture and fast decoding. To alleviate data sparseness issues due to infrequent words, the combination with an acoustic-to-character (A2C) model is investigated. Moreover, the A2C model can be used to recover-of-vocabulary (OOV) words that are not covered by the A2W model, but this requires accurate detection of OOV words. A2W models learn contexts with both acoustic and transcripts; therefore they tend to falsely recognize OOV words as words in the vocabulary. In this paper, we tackle this problem by using external language models (LM), which are trained only with transcriptions and have better linguistic information to detect OOV words. The A2C model is used to resolve these OOV words. Experimental evaluations show that external LMs have the effects of not only reducing errors but also increasing the number of detected OOV words, and the proposed method significantly improves performances in English conversational and Japanese lecture corpora, especially for-of-domain scenario. We also investigate the impact of the vocabulary size of A2W models and the data size for training LMs. Moreover, our approach can reduce the vocabulary size several times with marginal performance degradation. Hirofumi Inaguma, Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
SLT | 3 |
| 2018 | Leveraging Sequence-to-Sequence Speech Synthesis for Enhancing Acoustic-to-Word Speech RecognitionabstractEncoder-decoder models for acoustic-to-word (A2W) automatic speech recognition (ASR) are attractive for their simplicity of architecture and run-time latency while achieving state-of-the-art performances. However, word-based models commonly suffer from the-of-vocabulary (OOV) word problem. They also cannot leverage text data to improve their language modeling capability. Recently, sequence-to-sequence neural speech synthesis models trainable from corpora have been developed and shown to achieve naturalness com- parable to recorded human speech. In this paper, we explore how we can leverage the current speech synthesis technology to tailor the ASR system for a target domain by preparing only a relevant text corpus. From a set of target domain texts, we generate speech features using a sequence-to-sequence speech synthesizer. These artificial speech features together with real speech features from conventional speech corpora are used to train an attention-based A2W model. Experimental results show that the proposed approach improves the word accuracy significantly compared to the baseline trained only with the real speech, although synthetic part of the training data comes only from a single female speaker voice. Masato Mimura, Sei Ueno, Hirofumi Inaguma, Shinsuke Sakai, Tatsuya Kawahara |
SLT | 4 |
| 2017 | Cross-domain speech recognition using nonparallel corpora with cycle-consistent adversarial networksabstractAutomatic speech recognition (ASR) systems often does not perform well when it is used in a different acoustic domain from the training time, such as utterances spoken in noisy environments or in different speaking styles. We propose a novel approach to cross-domain speech recognition based on acoustic feature mappings provided by a deep neural network, which is trained using nonparallel speech corpora from two different domains and using no phone labels. For training a target domain acoustic model, we generate “fake” target speech features from the labeleld source domain features using a mapping Gf. We can also generate “fake” source features for testing from the target features using the backward mapping Gbwhich has been learned simultaneously with G f. The mappings G f and Gbare trained as adversarial networks using a conventional adversarial loss and a cycle-consistency loss criterion that encourages the backward mapping to bring the translated feature back to the original as much as possible such that Gb(Gf (x)) ≈ x. In a highly challenging task of model adaptation only using domain speech features, our method achieved up to 16 % relative improvements in WER in the evaluation using the CHiME3 real test data. The backward mapping was also confirmed to be effective with a speaking style adaptation task. Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
ASRU | 2 |
| 2017 | Semi-supervised ensemble DNN acoustic model trainingabstractIt is very important to exploit abundant unlabeled speech for improving the acoustic model training in automatic speech recognition (ASR). Semi-supervised training methods incorporate unlabeled data in addition to labeled data to enhance the model training, but it encounters the error-prone label problem. The ensemble training scheme trains a set of models and combines them to make the model more general and robust, but it has not been applied to the unlabeled data. In this work, we propose an effective semi-supervised training of deep neural network (DNN) acoustic models by incorporating the diversity among the ensemble of models. The resultant model improved the performance in the lecture transcription task. Moreover, the proposed method has also shown a potential for DNN adaptation. Sheng Li 0010, Xugang Lu, Shinsuke Sakai, Masato Mimura, Tatsuya Kawahara |
ICASSP | 3 |
| 2017 | Combined Multi-Channel NMF-Based Robust Beamforming for Noisy Speech Recognition
Masato Mimura, Yoshiaki Bando, Kazuki Shimada, Shinsuke Sakai, Kazuyoshi Yoshii, Tatsuya Kawahara |
INTERSPEECH | 4 |
| 2016 | Joint Optimization of Denoising Autoencoder and DNN Acoustic Model Based on Multi-Target Learning for Noisy Speech Recognition
Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2015 | Deep autoencoders augmented with phone-class feature for reverberant speech recognitionabstractThis paper addresses reverberant speech recognition based on front-end processing using DAE (Deep AutoEncoder) coupled with DNN (Deep Neural Network) acoustic model. DAE can effectively and flexibly learn mapping from corrupted speech to the original clean speech based on the deep learning scheme. While this mapping is conventionally conducted only with the acoustic information, we presume the mapping is also dependent on the phone information. Therefore, we propose a new scheme (pDAE), which augments a phone-class feature to the standard acoustic features as input. Two types of the phone-class feature are investigated. One is the hard recognition result of monophones, and the other is a soft representation derived from the posterior outputs of monophone DNN. In the evaluation on the Reverb Challenge 2014 task, the augmented feature in either type results in a significant improvement (7-8% relative) from the standard DAE. It is also shown that using the soft representation in the training phase is critical. Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
ICASSP | 2 |
| 2015 | Speech dereverberation using long short-term memoryabstractRecently, neural networks have been used for not only phone recognition but also denoising and dereverberation. However, the conventional denoising deep autoencoder (DAE) based on the feed-forward structure is not capable of handling very long speech frames of reverberation. LSTM can be effectively trained to reduce the average error between the enhanced signal and the original clean signal by considering the effect of the long past time frames. In this paper, we demonstrate that considering as long as the maximum reverberation time of the database is effective. Since the effect of reverberation varies depending on the phone-class of the whole speech context, we augment the input of the autoencoder with the phone-class information of the past frames as well as the current frame and call this version of the LSTM autoencoder pLSTM. In the speech recognition experiment using the data set of Reverb Challenge 2014, the LSTM front-end reduced the WER of the multicondition DNN-HMM by 14.5%, and the use of the phone class feature yielded in pLSTM further improvement of 7.5%. The performance with the pLSTM is comparable to that of pDAE, while the number of parameters is only 1/25-1/8. Index Terms: Speech Dereverberation, Long Short-Term Memory (LSTM), Deep Autoencoder (DAE) Masato Mimura, Shinsuke Sakai, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2013 | A-STAR: Toward translating Asian spoken languages
Sakriani Sakti, Michael Paul, Andrew M. Finch, Shinsuke Sakai, Thang Tat Vu, Noriyuki Kimura, Chiori Hori, Eiichiro Sumita, Satoshi Nakamura 0001, Jun Park, Chai Wutiwiwatchai, Bo Xu 0002, Hammam Riza, Karunesh Arora, Haizhou Li 0001 |
Comput. Speech Lang. | 4 |
| 2011 | A sampling-based environment population projection approach for rapid acoustic model adaptationabstractWe propose an environment population projection (EPP) approach for rapid acoustic model adaptation to reduce environment mismatches with limited amounts of adaptation data. This approach consists of two stages: population construction and projection. In the population construction stage, we apply a sampling scheme on the adaptation data to construct an environment population based on acoustic models prepared in the training phase. With this sampling procedure, the environment samples in the population characterize diverse acoustic information embedded in the adaptation data. Next, the projection stage estimates a function to map the environment population into one set of acoustic models that matches the testing condition. With a well constructed environment population, a simple projection function can enable the EPP approach to accurately characterize the testing environment even with a small amount of adaptation data. To examine the rapid adaptation ability of EPP, we used only one adaptation utterance and tested performance in both supervised and unsupervised adaptation modes on Aurora-2 and Aurora-2J tasks. It is found that EPP achieves satisfactory performance under both modes for both tasks. On the Aurora-2J task for example, EPP gives a clear improvement of a 13.87% (8.58% to 7.39%) word error rate (WER) reduction over our baseline in the unsupervised adaptation mode. Yu Tsao 0001, Shigeki Matsuda, Shinsuke Sakai, Ryosuke Isotani, Hisashi Kawai, Satoshi Nakamura 0001 |
ICASSP | 3 |
| 2010 | Improved training of excitation for HMM-based parametric speech synthesisabstractThis paper presents an improved method of training for the unvoiced filter that comprises an excitation model, within the framework of parametric speech synthesis based on hidden Markov models. The conventional approach calculates the unvoiced filter response from the differential signal of the residual and voiced excitation estimate. The differential signal, however, includes the error generated by the voiced excitation estimates. Contaminated by the error, the unvoiced filter tends to be overestimated, which causes the synthetic speech to be noisy. In order for unvoiced filter training to obtain targets that are free from the contamination, the improved approach first separates the non-periodic component of residual signal from the periodic component. The unvoiced filter is then trained from the non-periodic component signals. Experimental results show that unvoiced filter responses trained with the new approach are clearly noiseless, in contrast to the responses trained with the conventional approach. Yoshinori Shiga, Tomoki Toda, Shinsuke Sakai, Hisashi Kawai |
INTERSPEECH | 3 |
| 2009 | CART-based modeling of Chinese tonal patterns with a functional model tracing the fundamental frequency trajectoriesabstractWe propose an approach to modeling Chinese tonal patterns, focusing on the basic fundamental frequency (F0) patterns characterized by the contextual linguistic features that can be directly extracted from text. We analyze tonal patterns as sparse target points (tonal F0peaks and valleys) and represent them in parametric form within the framework of a functional F0model. The relationships between the target points and underlying linguistic features are trained using classification and regression tree analysis (CARTs), and this functional model is used to trace the F0trajectories when training the CARTs and to synthesize a tonal pattern from the target points predicted by the CARTs. Our experiments indicate that the proposed method has low F0prediction errors. Utilization of the F0ranges measured from training samples could significantly reduce the influences of differences in voice ranges on training a speaker-independent model. Furthermore, the most important roles in characterizing tonal patterns were played by a few linguistic features such as lexical tone context and the distinction between voiced from unvoiced initials. Jinfu Ni, Shinsuke Sakai, Tohru Shimizu, Satoshi Nakamura 0001 |
ICASSP | 2 |
| 2009 | Optimal learning of P-Layer additive F0 models with cross-validationabstractIn this paper, we present the derivation of the backfitting training algorithms for generic p-layer additive F0models for arbitrary positive integer p. We have presented the special cases of the algorithms with p = 2 and p = 3 that have been successfully applied to the modelings of Japanese and English F0contours, whereas the derivation of the algorithm was presented only for the two-layer case. The additive F0model have smoothing parameters that establish a trade-off between the fit to the training data and the smoothness of the fitted curves, which have been all set to unity in the previous works. In this paper, we also present an optimal approach to set the values of these parameters using cross validation. We performed the training using the Boston University Radio News Corpus and confirmed the effectiveness of the proposed method. Shinsuke Sakai, Tatsuya Kawahara, Tohru Shimizu, Satoshi Nakamura 0001 |
ICASSP | 1 |
| 2009 | A decision tree-based clustering approach to state definition in an excitation modeling framework for HMM-based speech synthesisabstractThis paper presents a decision tree-based algorithm to cluster residual segments assuming an excitation model based on statedependent filtering of pulse train and white noise. The decision tree construction principle is the same as the one applied to speech recognition. Here parent nodes are split using the residual maximum likelihood criterion. Once these excitation decision trees are constructed for residual signals segmented by full context models, using questions related to the full context of the training sentences, they can be utilized for excitation modeling in speech synthesis based on hidden Markov models (HMM). Experimental results have shown that the algorithm in question is very effective in terms of clustering residual signals given segmentation, pitch marks and full context questions, resulting in filters with good residual modeling properties. Ranniery Maia, Tomoki Toda, Keiichi Tokuda, Shinsuke Sakai, Satoshi Nakamura 0001 |
INTERSPEECH | 4 |
| 2009 | A close look into the probabilistic concatenation model for corpus-based speech synthesis
Shinsuke Sakai, Ranniery Maia, Hisashi Kawai, Satoshi Nakamura 0001 |
INTERSPEECH | 1 |
| 2008 | Development of Indonesian Large Vocabulary Continuous Speech Recognition System within A-STAR Project
Sakriani Sakti, Eka Kelana, Hammam Riza, Shinsuke Sakai, Konstantin Markov, Satoshi Nakamura 0001 |
IJCNLP | 4 |
| 2006 | Decision tree-based training of probabilistic concatenation models for corpus-based speech synthesisabstractThe measure of the goodness, or cost, of concatenating synthesis units plays an important role in concatenative speech synthesis. In this paper, we present a probabilistic approach to concatenation modeling in which the goodness of concatenation is represented as the conditional probability of observing the spectral shape of a unit given the previous unit and the current phonetic context. This conditional probability is modeled by a conditional Gaussian density whose mean vector has a form of linear transform of the past spectral shape. A phonetic decision-tree based parameter tying is performed to achieve a robust training that balances between model complexity and the amount of training data available. The concatenation models are implemented in a corpus-based speech synthesizer trained with a CMU Arctic database and the effectiveness of the proposed method was confirmed by a subjective listening test. Index Terms: speech synthesis, unit selection, join costs. Shinsuke Sakai, Tatsuya Kawahara |
INTERSPEECH | 1 |
| 2005 | Additive Modeling of English F0 Contour for Speech SynthesisabstractIn this paper, we present an approach to fundamental frequency contour modeling of English for speech synthesis, based on a statistical learning technique called additive models that was successfully applied to the modeling of Japanese F/sub 0/ contours previously. In an attempt to model English F/sub 0/ contours, we defined a three-layer additive model consisting of an intonational phrase component, a word-level component representing lexical stress types, and a pitch-accent component related to accented syllables. These component functions are estimated simultaneously using a backfitting algorithm derived from a regularized least-squares error criterion specified on the model with regard to the training data. The proposed method was trained and tested using the widely used ToBI-labeled speech corpus and promising results were obtained. Shinsuke Sakai |
ICASSP (1) | 1 |
| 2005 | A probabilistic approach to unit selection for corpus-based speech synthesisabstractIn this paper, we present a novel statistical approach to corpus-based speech synthesis. Unit selection is directed by probabilistic models for F0 contour, duration, and spectral characteristics of the synthesis units. The F0 targets for units are modeled by statistical additive models, and duration targets are modeled by regression trees. Spectral targets for a unit is modeled by Gaussian mixtures on MFCC-based features. Goodness of concatenation of two units is modeled by conditional Gaussian models on MFCC-based features. Although the system is in its early stage of development, we implemented an English speech synthesizer with CMU Arctic corpora and confirmed the effectiveness of this new framework. 1. Shinsuke Sakai, Han Shu |
INTERSPEECH | 1 |
| 2000 | Continuous speech recognition with parse filtering
Ken Hanazawa, Shinsuke Sakai |
INTERSPEECH | 2 |
| 2000 | An automatic interpretation system for travel conversation
Takao Watanabe, Akitoshi Okumura, Shinsuke Sakai, Kiyoshi Yamabana, Shinichi Doi, Ken Hanazawa |
INTERSPEECH | 3 |
| 1995 | Multilingual spoken-language understanding in the MIT Voyager system
James R. Glass, Giovanni Flammia, David Goodine, Michael S. Phillips 0001, Joseph Polifroni, Shinsuke Sakai, Stephanie Seneff, Victor Zue |
Speech Commun. | 6 |
| 1994 | An automatic voice dialing system developed on PC speech i/o platform
Jun Noguchi, Shinsuke Sakai, Kaichiro Hatazaki, Ken-ichi Iso, Takao Watanabe |
ICSLP | 2 |
| 1993 | A bilingual Voyager system
James R. Glass, David Goodine, Michael S. Phillips 0001, Shinsuke Sakai, Stephanie Seneff, Victor Zue |
EUROSPEECH | 4 |
| 1993 | J-SUMMIT: Japanese spontaneous speech recognition
Shinsuke Sakai, Michael S. Phillips 0001 |
EUROSPEECH | 1 |
| 1992 | J-SUMMIT: a Japanese segment-based speech recognition system
Shinsuke Sakai, Michael S. Phillips 0001 |
ICSLP | 1 |
| 1990 | From interlingua to speech: generating prosodic information from conceptual representationabstractA method for generating prosodic information for speech synthesis from conceptual representation is presented. A computational model of generating prosodic information from pragmatic, semantic, syntactic, and lexical information of the utterance is proposed. The various types of information are computed using the information extracted from the conceptual representation and the lexicon through the process of uttered sentence generation. An experimental system for speech synthesis from conceptual representation is developed, and better prosodic quality is obtained than with conventional synthetic speech, according to an informal listening test.> Shinsuke Sakai, Kazunori Muraki |
ICASSP | 1 |