VLDB 2026 Research / reviewers in the wild / expert
Masayuki Suzuki
dblp:46/5214
· DBLP profile ↗
47ranked-venue papers
16as first author
5since 2021 · last 2025
0000-0002-0436-1490ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 11 first-author · 5 since 2021Artificial intelligence and machine learning · 26 · 9 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-authorSystems, architecture and hardware · 3 · 2 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 first-authorTheory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Granite-speech: open-source speech-aware LLMs with strong English ASR capabilitiesabstractGranite-speech LLMs are compact and efficient speech language models specifically designed for English ASR1and automatic speech translation (AST). The models were trained by modality aligning granite-3.3-instruct to speech on publicly available open-source corpora. Comprehensive benchmarking on English ASR shows that they outperform several competitors’ models that were trained on orders of magnitude more proprietary data, and they keep pace on English-to-X AST for major European languages, Japanese, and Mandarin. The speech-specific components are: a conformer acoustic encoder using block attention and self-conditioning trained with connectionist temporal classification, a windowed query-transformer speech modality adapter used to do temporal downsampling of the acoustic embeddings and map them to the LLM text embedding space, and LoRA adapters to further fine-tune the text LLM. The models are freely available on HuggingFace2under a permissive Apache 2.0 license.1The latest models (revision 3.3.2) support multilingual ASR in English, French, German, Spanish and Portuguese and bidirectional speech translation to and from English. This paper covers the initial English-only release.2https://huggingface.co/ibm-granite/granite-speech-3.3-2b (and…-8b). George Saon, Avihu Dekel, Alexi Brooks, Tohru Nagano, Abraham Daniels, Aharon Satt, Ashish R. Mittal, Brian Kingsbury, David Haws, Edmilson da Silva Morais, Gakuto Kurata, Hagai Aronowitz, Ibrahim Ibrahim, Hong-Kwang Jeff Kuo, Kate Soule, Luis A. Lastras, Masayuki Suzuki, Ron Hoory, Samuel Thomas 0001, Sashi Novitasari, Takashi Fukuda, Vishal Sunder, Zvi Kons |
ASRU | 17 |
| 2024 | Multiple Representation Transfer from Large Language Models to End-to-End ASR SystemsabstractTransferring the knowledge of large language models (LLMs) is a promising technique to incorporate linguistic knowledge into end-to-end automatic speech recognition (ASR) systems. However, existing works only transfer a single representation of LLM (e.g. the last layer of pretrained BERT), while the representation of a text is inherently non-unique and can be obtained variously from different layers, contexts and models. In this work, we explore a wide range of techniques to obtain and transfer multiple representations of LLMs into a transducer-based ASR system. While being conceptually simple, we show that transferring multiple representations of LLMs can be an effective alternative to transferring only a single LLM representation. Takuma Udagawa, Masayuki Suzuki, Gakuto Kurata, Masayasu Muraoka, George Saon |
ICASSP | 2 |
| 2022 | Improving Generalization of Deep Neural Network Acoustic Models with Length Perturbation and N-best Based Label SmoothingabstractWe introduce two techniques, length perturbation and n-best based label smoothing, to improve generalization of deep neural network (DNN) acoustic models for automatic speech recognition (ASR).Length perturbation is a data augmentation algorithm that randomly drops and inserts frames of an utterance to alter the length of the speech feature sequence.N-best based label smoothing randomly injects noise to ground truth labels during training in order to avoid overfitting, where the noisy labels are generated from n-best hypotheses.We evaluate these two techniques extensively on the 300-hour Switchboard (SWB300) dataset and an in-house 500-hour Japanese (JPN500) dataset using recurrent neural network transducer (RNNT) acoustic models for ASR.We show that both techniques improve the generalization of RNNT models individually and they can also be complementary.In particular, they yield good improvements over a strong SWB300 baseline and give state-of-art performance on SWB300 using RNNT models. George Saon, Tohru Nagano, Masayuki Suzuki, Takashi Fukuda, Brian Kingsbury, Gakuto Kurata |
INTERSPEECH | 4 |
| 2022 | Global RNN Transducer Models For Multi-dialect Speech Recognition
Takashi Fukuda, Samuel Thomas 0001, Masayuki Suzuki, Gakuto Kurata, George Saon, Brian Kingsbury |
INTERSPEECH | 3 |
| 2022 | Effect and Analysis of Large-scale Language Model Rescoring on Competitive ASR Systems
Takuma Udagawa, Masayuki Suzuki, Gakuto Kurata, Nobuyasu Itoh, George Saon |
INTERSPEECH | 2 |
| 2020 | Converting Written Language to Spoken Language with Neural Machine Translation for Language ModelingabstractWhen building a language model (LM) for spontaneous speech, the ideal situation is to have a large amount of spoken, in-domain training data. Having such abundant data, however, is not realistic. We address this problem by generating texts in spoken language from those in written language by using a neural machine translation (NMT) model. We collected faithful transcripts of fully spontaneous speech and corresponding written versions and used them as a parallel corpus to train the NMT model. We used top-k random sampling, which generates a large variety of texts of higher quality as compared to other generation methods for NMT. We indicate that the NMT model is capable of converting written texts in a certain domain to spoken texts, and that the converted texts are effective for training LMs. Our experimental results show significant improvement of speech recognition accuracy with the LMs. Shintaro Ando, Masayuki Suzuki, Nobuyasu Itoh, Gakuto Kurata, Nobuaki Minematsu |
ICASSP | 2 |
| 2020 | Speaker Embeddings Incorporating Acoustic Conditions for DiarizationabstractWe present our work on training speaker embeddings, especially effective for speaker diarization. For various speaker recognition tasks, extracting speaker embeddings using Deep Neural Networks (DNNs) has become major methods. These embeddings are generally trained to be discriminate speakers and be robust with respect to different acoustic conditions. In speaker diarization, however, the acoustic conditions can be used as consistent information for discriminating speakers. Such information can include the distances to a microphone in a meeting, or the channels for each speaker in telephone conversation recorded in monaural. Hence, the proposed speaker-embedding network leverages differences in acoustic conditions to train effective speaker embeddings for speaker diarization. The information on acoustic conditions can be anything that contributes to distinguishing between recording environments; for example, we explore using i-vectors. Experiments conducted on a practical diarization system demonstrated that the proposed embeddings significantly improve performance over embeddings without information on acoustic conditions. Yosuke Higuchi, Masayuki Suzuki, Gakuto Kurata |
ICASSP | 2 |
| 2020 | New Advances in Speaker Diarization
Hagai Aronowitz, Weizhong Zhu, Masayuki Suzuki, Gakuto Kurata, Ron Hoory |
INTERSPEECH | 3 |
| 2019 | Semi-Supervised Training and Data Augmentation for Adaptation of Automatic Broadcast News Captioning SystemsabstractIn this paper we present a comprehensive study on building and adapting deep neural network based speech recognition systems for automatic closed captioning. We develop the proposed systems by first building base automatic speech recognition (ASR) systems that are not specific to any particular show or station. These models are trained on nearly 6000 hours of broadcast news data using conventional hybrid and more recent attention based end-to-end acoustic models. We then employ various adaptation and data augmentation strategies to further improve the trained base models. We use 535 hours of data from two independent BN sources to study how the base models can be customized. We observe up to 32% relative improvement using the proposed techniques on test sets related to, but independent of the adaptation data. At these low word error rates (WERs), we believe the customized BN ASR systems can be used effectively for automatic closed captioning. Samuel Thomas 0001, Masayuki Suzuki, Zoltán Tüske, Larry Sansone, Michael Picheny |
ASRU | 3 |
| 2019 | Data Augmentation Based on Vowel Stretch for Improving Children's Speech RecognitionabstractProlongation is a speech disfluency that lengthens some portions of speech utterances. It is frequently observed in children's spontaneous speech, while it is rare in read speech. To make acoustic models more robust to children's spontaneous speech, collecting a large amount of children's speech data containing prolongation is usually required, which is very impractical in many cases. To tackle this problem, we propose a novel data augmentation method that virtually generates additional data by simulating prolongation. The method inserts pseudo frames into specific positions of speech utterances to simulate prolongation. The acoustic features of the inserted frames are calculated from the original frames on both sides. This is based on our analysis that many of vowels are actually stretched in children's spontaneous speech. Our proposed procedure can generate partially stretched utterances with low computational costs, unlike a conventional speed or tempo perturbation method that extends and shrinks entire utterances at a uniform rate. The effectiveness of the proposed method were confirmed with the experiments of acoustic model adaptations, in which our proposed method focusing on vowel stretch showed consistent improvement compared with conventional speed and tempo perturbation approach. Tohru Nagano, Takashi Fukuda, Masayuki Suzuki, Gakuto Kurata |
ASRU | 3 |
| 2019 | Improvements to N-gram Language Model Using Text Generated from Neural Language ModelabstractAlthough neural language models have emerged, n-gram language models are still used for many speech recognition tasks. This paper proposes four methods to improve n-gram language models using text generated from a recurrent neural network language model (RNNLM). First, we use multiple RNNLMs from different domains instead of a single RNNLM. The final n-gram language model is obtained by interpolating generated n-gram models from each domain. Second, we use subwords instead of words for RNNLM to reduce the out-of-vocabulary rate. Third, we generate text templates using an RNNLM for template-based data augmentation for named entities. Fourth, we use both forward RNNLM and backward RNNLM to generate text. We found that these four methods improved performance of speech recognition up to 4% relative in various tasks. Masayuki Suzuki, Nobuyasu Itoh, Tohru Nagano, Gakuto Kurata, Samuel Thomas 0001 |
ICASSP | 1 |
| 2019 | English Broadcast News Speech Recognition by Humans and MachinesabstractWith recent advances in deep learning, considerable attention has been given to achieving automatic speech recognition performance close to human performance on tasks like conversational telephone speech (CTS) recognition. In this paper we evaluate the usefulness of these proposed techniques on broadcast news (BN), a similar challenging task. We also perform a set of recognition measurements to understand how close the achieved automatic speech recognition results are to human performance on this task. On two publicly available BN test sets, DEV04F and RT04, our speech recognition system using LSTM and residual network based acoustic models with a combination of n-gram and neural network language models performs at 6.5% and 5.9% word error rate. By achieving new performance milestones on these test sets, our experiments show that techniques developed on other related tasks, like CTS, can be transferred to achieve similar performance. In contrast, the best measured human recognition performance on these test sets is much lower, at 3.6% and 2.8% respectively, indicating that there is still room for new techniques and improvements in this space, to reach human performance levels. Samuel Thomas 0001, Masayuki Suzuki, Gakuto Kurata, Zoltán Tüske, George Saon, Brian Kingsbury, Michael Picheny, Tom Dibert, Alice Kaiser-Schatzlein, Bern Samko |
ICASSP | 2 |
| 2019 | Direct Neuron-Wise Fusion of Cognate Neural Networks
Takashi Fukuda, Masayuki Suzuki, Gakuto Kurata |
INTERSPEECH | 2 |
| 2018 | Inference-Invariant Transformation of Batch Normalization for Domain Adaptation of Acoustic Models
Masayuki Suzuki, Tohru Nagano, Gakuto Kurata, Samuel Thomas 0001 |
INTERSPEECH | 1 |
| 2017 | Harmonic feature fusion for robust neural network-based acoustic modelingabstractAcoustic modeling with deep learning has drastically improved the performance of automatic speech recognition (ASR) where the main stream of the acoustic feature is still log-Mel filtered one. While the log-Mel filtered features lose harmonic-structure information, they still include useful information for ASR. Several attempts have been made to integrate higher-resolution information into the network. In order to improve the ASR accuracy in noisy conditions, we propose new features integrated into acoustic modeling to represent which parts in the time-frequency domain have a distinct harmonic structure, since it is partially observed in noisy environments. The new features are combined with the standard acoustic features, and the network is trained with them using various noisy data. Through these operations, it learns the acoustic features with a kind of quality tag describing which parts are clean or degraded. Our model reduced the word error rate in an Aurora-4 task by 10.3% in DNN compared with the strong baseline while retaining the high accuracy in clean test cases. Osamu Ichikawa, Takashi Fukuda, Masayuki Suzuki, Gakuto Kurata, Bhuvana Ramabhadran |
ICASSP | 3 |
| 2017 | Efficient Knowledge Distillation from an Ensemble of Teachers
Takashi Fukuda, Masayuki Suzuki, Gakuto Kurata, Samuel Thomas 0001, Jia Cui, Bhuvana Ramabhadran |
INTERSPEECH | 2 |
| 2017 | Ensembles of Multi-Scale VGG Acoustic Models
Michael Heck, Masayuki Suzuki, Takashi Fukuda, Gakuto Kurata, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2017 | Symbol Sequence Search from Telephone Conversation
Masayuki Suzuki, Gakuto Kurata, Abhinav Sethy, Bhuvana Ramabhadran, Kenneth Church 0001, Mark Drake |
INTERSPEECH | 1 |
| 2016 | Speech recognition robust against speech overlapping in monaural recordings of telephone conversationsabstractMonaural (single-channel) recording is sometimes used for telephone conversations in call centers. Generally speaking, the accuracy of automatic speech recognition of a monaural recording is worse than that of the multi-channel recording of the same conversation where each speaker's voice is separately recorded. The major reason is that the recognition system fails not only at the overlapping segments where the voices of the multiple speakers overlap, but also at the neighboring segments surrounding the overlapping segments. In this paper, we tackle this problem by using a combination of garbage modeling and noise-robust monaural acoustic modeling. Our proposed method trains the models by making use of multi-channel recordings and transcripts, which are relatively easy to prepare than monaural recordings and transcripts. We present experimental results where the proposed methods reduced the error rates by approximately 3% relative to the baseline methods for both of GMM-HMM and CNN-HMM cases. Because the proposed method is quite simple, the proposed method is easy to deploy to wide range of ASR systems for monaural speech transcription. Masayuki Suzuki, Gakuto Kurata, Tohru Nagano, Ryuki Tachibana |
ICASSP | 1 |
| 2016 | Domain Adaptation of CNN Based Acoustic Models Under Limited Resource Settings
Masayuki Suzuki, Ryuki Tachibana, Samuel Thomas 0001, Bhuvana Ramabhadran, George Saon |
INTERSPEECH | 1 |
| 2015 | How learners use feedback information: Effects of social comparative information and achievement goals
Masayuki Suzuki, Tetsuya Toyota |
CogSci | 1 |
| 2015 | Discriminative re-ranking for automatic speech recognition by leveraging invariant structures
Masayuki Suzuki, Gakuto Kurata, Masafumi Nishimura, Nobuaki Minematsu |
Speech Commun. | 1 |
| 2014 | Cognitive Model of Generic Skill: Cognitive Processes in Search and Editing
Akira Fujita, Masayuki Suzuki, Noriko H. Arai |
CogSci | 2 |
| 2013 | Improved estimation of femininity using GMM supervectors and SVR for voice therapy of Gender Identity Disorder ClientsabstractThis paper proposes a new method of estimating perceptual femininity (PF) of an input utterance using Gaussian Mixture Model (GMM) supervectors and support vector regression (SVR). The method is used to develop a femininity estimation tool, which is introduced to voice therapy of Gender Identity Disorder (GID) clients, especially MtF (Male to Female) transsexuals. In our previous study [1], we developed a PF estimator, where a male GMM and a female GMM of spectral features and those of pitch features were built and their likelihood scores of an input utterance were combined by linear regression to estimate PF. In this work, inspired by recent speaker recognition models [2], we replace the four likelihood scores from the four GMMs with supervectors composed by a spectral GMM and a pitch GMM estimated from an input utterance. Further, instead of simple linear regression, we introduce SVR, which is discriminative linear regression. Experiments using an MtF speech corpus show that the proposed method improves correlation between human and machine scores of PF and also reduces squared prediction error. Chengshuo Wang, Masayuki Suzuki, Nobuaki Minematsu, Kyoko Sakuraba, Keikichi Hirose |
ICASSP | 2 |
| 2013 | Designing Effective Feedback for Cognitive Diagnostic Assessment in Web-based Learning EnvironmentabstractAssessment is useful for students to improve their learning and for teachers to adjust their teaching practice. However, most traditional assessments do not provide useful information to improve learning and teaching. Recently, cognitive diagnostic assessment (CDA) which is designed to measure specific knowledge structures and processing skills in students has attracted a great deal of attentions. In this paper, we apply a CDA approach in fraction problems to 144 sixth grade students in an elementary school in Japan. We show how CDA can provide detailed information about learners’ strengths and weaknesses and discuss the applicability of web-based CDA for providing effective feedback. Masayuki Suzuki, Tetsuya Toyota |
ICCE | 2 |
| 2013 | Artificial bandwidth extension based on regularized piecewise linear mapping with discriminative region weighting and long-Span featuresabstractArtificial Bandwidth Extension (ABE) has been introduced to improve perceived speech quality and intelligibility of narrowband telephone speech. Most of the existing algorithms divided ABE into 2 sub-problems, namely extension of the excitation signal and that of the spectral envelope. In this paper, we propose a new method for spectral envelope extension based on REgularized piecewise linear mapping with DIscriminative region weighting And Long-span features (REDIAL). REDIAL is a revised version of SPLICE, a well-known method for speech enhancement. In REDIAL, however, discriminative model is introduced for space division step of the original SPLICE. The proposed REDIAL-based method approximates non-linear transformation from narrowband features to their wideband counterpart by a summation of piecewise linear transformations. The proposed method was compared with the widely used GMM-based method, through objective and subjective evaluations in both speaker-dependent and speaker-independent conditions. Both evaluations showed that the proposed method significantly outperforms the conventional GMM-based method. Nguyen Duc Duy, Masayuki Suzuki, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 2 |
| 2013 | A free online accent and intonation dictionary for teachers and learners of Japanese
Hiroko Hirano, Ibuki Nakamura, Nobuaki Minematsu, Masayuki Suzuki, Chieko Nakagawa, Noriko Nakamura, Yukinori Tagawa, Keikichi Hirose, Hiroya Hashimoto |
INTERSPEECH | 4 |
| 2013 | Development of a web framework for teaching and learning Japanese prosody: OJAD (online Japanese accent dictionary)abstractThis paper introduces the first online and free framework for teaching and learning Japanese prosody including word accent and phrase intonation. This framework is called OJAD (Online Japanese Accent Dictionary) [1] and it provides three functions. 1) Visual, auditory, systematic, and comprehensive illustration of patterns of accent change (accent sandhi) of verbs and adjectives. Here only the changes caused by twelve kinds of fundamental conjugation are focused upon. 2) Visual illustration of the accent pattern of a given verbal expression, which is a combination of a verb and its postpositional auxiliary words. 3) Visual illustration of the pitch pattern of an any given sentence and the expected positions of accent nuclei in the sentence. The third function is implemented by using an accent change prediction module that we developed for Japanese text-to-speech (TTS) synthesizers [2, 3]. Experiments show that accent nucleus assignment to given texts by the proposed framework is much more accurate than that by native speakers. Subjective assessment and objective assessment by teachers and learners show very high pedagogical effectiveness of the framework. Index Terms: language education, Japanese prosody, accent sandhi, OJAD, speech synthesis, assessment experiments Ibuki Nakamura, Nobuaki Minematsu, Masayuki Suzuki, Hiroko Hirano, Chieko Nakagawa, Noriko Nakamura, Yukinori Tagawa, Keikichi Hirose, Hiroya Hashimoto |
INTERSPEECH | 3 |
| 2013 | Feature Enhancement With Joint Use of Consecutive Corrupted and Noise Feature Vectors With Discriminative Region WeightingabstractThis paper proposes a feature enhancement method that can achieve high speech recognition performance in a variety of noise environments with feasible computational cost. As the well-known Stereo-based Piecewise Linear Compensation for Environments (SPLICE) algorithm, the proposed method learns piecewise linear transformation to map corrupted feature vectors to the corresponding clean features, which enables efficient operation. To make the feature enhancement process adaptive to changes in noise, the piecewise linear transformation is performed by using a subspace of the joint space of corrupted and noise feature vectors, where the subspace is chosen such that classes (i.e., Gaussian mixture components) of underlying clean feature vectors can be best predicted. In addition, we propose utilizing temporally adjacent frames of corrupted and noise features in order to leverage dynamic characteristics of feature vectors. To prevent overfitting caused by the high dimensionality of the extended feature vectors covering the neighboring frames, we introduce regularized weighted minimum mean square error criterion. The proposed method achieved relative improvements of 34.2% and 22.2% over SPLICE under the clean and multi-style conditions, respectively, on the Aurora 2 task. Masayuki Suzuki, Takuya Yoshioka, Shinji Watanabe 0001, Nobuaki Minematsu, Keikichi Hirose |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | Unseen noise robust speech recognition using adaptive piecewise linear transformationabstractSPLICE is one of the speech enhancement methods based on feature conversion, which shows a high performance with a relatively small amount of calculation. After modeling noisy speech features as GMM, conversion functions are obtained for individual GMM components. The original SPLICE estimates clean feature vectors as a weighted summation of the converted versions of input vectors. Since the conversion functions are determined and fixed only by using training data, the effectiveness of the original SPLICE will be lower in the case of unseen noisy environments. In this paper, we propose a novel method to adapt the conversion functions to work well in unseen environments. First, to realize adaptive conversion functions, we characterize those functions using their super vectors. Then, we conduct PCA on the super vectors to reduce the number of parameters to be adapted. By representing the super vectors through their PCA-based base functions and weights, we implement an efficient adaptation method of conversion functions, which we call Eigen-SPLICE here after. Evaluation experiments show that Eigen-SPLICE has reduced word error rate by 21.0% relative to the conventional SPLICE, and by 24.1% relative to EMS SPLICE in the test set B of the AURORA-2 task. Keigo Chijiiwa, Masayuki Suzuki, Nobuaki Minematsu, Keikichi Hirose |
ICASSP | 2 |
| 2012 | MFCC enhancement using joint corrupted and noise feature space for highly non-stationary noise environmentsabstractOne of the most effective approaches to noise robust speech recognition is to remove the noise effect directly from corrupted MFCC vectors. However, VTS enhancement, which is a typical method for performing MFCC enhancement, provides limited improvement when the noise is highly non-stationary. This is because the VTS enhancement method cannot use a time-varying noise model to keep the computational cost at an acceptable level. This paper proposes a method that can enhance MFCC vectors and their dynamic parameters by using noise estimates that change on a frame-by-frame basis at a practical computational cost. The proposed method employs stereo data-based feature mapping like the well known SPLICE algorithm. The novelty of the proposed method lies in that it uses the joint space spanned by a concatenated vector of corrupted and noise features. It is also proposed to use linear discriminant analysis to effectively reduce the dimensionality of the joint space. The proposed method achieves 19.1% and 8.3% relative error reduction from the SPLICE and noise-mean normalized SPLICE algorithms, respectively. Masayuki Suzuki, Takuya Yoshioka, Shinji Watanabe 0001, Nobuaki Minematsu, Keikichi Hirose |
ICASSP | 1 |
| 2012 | Discriminative Reranking for LVCSR Leveraging Invariant StructureabstractAn invariant structure is one of the long-span acoustic represen-tations, where acoustic variations caused by non-linguistic fac-tors are effectively removed from speech. We present in this pa-per a new method to leverage the invariant structures as features of discriminative reranking for Large Vocabulary Continuous Speech Recognition (LVCSR). First we use a traditional HMM-based LVCSR system to get a list of N-best candidates with phone alignments and construct an invariant structure for each candidate using its phone alignment. Here, the invariant struc-ture is composed of lengths between every two phonemes in the candidate. Then we estimate a score of each phoneme-pair in the invariant structure, and rerank the N-best candidates using a weighted sum of the phoneme-pair scores, where the weights are trained discriminatively by averaged perceptron. Experi-mental results show a relative CER improvement of 6.69 % over the baseline HMM-based LVCSR system. Index Terms: Invariant Structure, LVCSR, Discriminative reranking 1. Masayuki Suzuki, Gakuto Kurata, Masafumi Nishimura, Nobuaki Minematsu |
INTERSPEECH | 1 |
| 2012 | Audio-visual feature integration based on piecewise linear transformation for noise robust automatic speech recognitionabstractMultimodal speech recognition is a promising approach to realize noise robust automatic speech recognition (ASR), and is currently gathering the attention of many researchers. Multimodal ASR utilizes not only audio features, which are sensitive to background noises, but also non-audio features such as lip shapes to achieve noise robustness. Although various methods have been proposed to integrate audio-visual features, there are still continuing discussions on how the vest integration of audio and visual features is realized. Weights of audio and visual features should be decided according to the noise features and levels: in general, larger weights to visual features when the noise level is low and vice versa, but how it can be controlled? In this paper, we propose a method based on piecewise linear transformation in feature integration. In contrast to other feature integration methods, our proposed method can appropriately change the weight depending on a state of an observed noisy feature, which has information both on uttered phonemes and environmental noise. Experiments on noisy speech recognition are conducted following to CENSREC-1-AV, and word error reduction rate around 24% is realized in average as compared to a decision fusion method. Yosuke Kashiwagi, Masayuki Suzuki, Nobuaki Minematsu, Keikichi Hirose |
SLT | 2 |
| 2012 | Performance improvement of automatic pronunciation assessment in a noisy classroomabstractIn recent years Computer-Assisted Language Learning (CALL) systems have been widely used in foreign language education. Some systems use automatic speech recognition (ASR) technologies to detect pronunciation errors and estimate the proficiency level of individual students. When speech recording is done in a CALL classroom, however, utterances of a student are always recorded with those of the others in the same class. The latter utterances are just background noise, and the performance of automatic pronunciation assessment is degraded especially when a student is surrounded with very active students. To solve this problem, we apply a noise reduction technique, Stereo-based Piecewise Linear Compensation for Environments (SPLICE), and the compensated feature sequences are input to a Goodness Of Pronunciation (GOP) assessment system. Results show that SPLICE-based noise reduction works very well as a means to improve the assessment performance in a noisy classroom. Yi Luan, Masayuki Suzuki, Yutaka Yamauchi, Nobuaki Minematsu, Shuhei Kato, Keikichi Hirose |
SLT | 2 |
| 2012 | Automatic Chinese pronunciation error detection using SVM trained with structural featuresabstractPronunciation errors are often made by learners of a foreign language. To build a Computer-Assisted Language Learning (CALL) system to support them, automatic error detection is essential. In this study, Japanese learners of Chinese are focused on. We investigated in automatic detection of their typical and frequent phoneme production errors. For this aim, four databases are newly created and we propose a detection method using Support Vector Machine (SVM) with structural features. The proposed method is compared to two baseline methods of Goodness Of Pronunciation (GOP) and Likelihood Ratio (LR) under the task of phoneme error detection. Experiments show that the proposed method performs much better than both of the two baseline methods. For example, the false rejection rate is reduced by as much as 82%. However, the results also indicate some drawbacks of using SVM with structural features. In this paper, we discuss merits and demerits of the proposed method and in what kind of real applications it works effectively. Tongmu Zhao, Akemi Hoshino, Masayuki Suzuki, Nobuaki Minematsu, Keikichi Hirose |
SLT | 3 |
| 2011 | Variable and clause elimination in SAT problems using an FPGAabstractThe satisfiability (SAT) problem is to find an assignment of binary values to the variables which satisfy a given clausal normal form (CNF). Many practical application problems can be transformed to SAT problems, and many SAT solvers have been developed. SAT problem is, however, NP-complete and its computational cost is very high. In order to reduce the computational cost, preprocessors are widely used by SAT solvers. In this paper, we describe an approach for implementing a preprocessor (SatELite) on FPGA. In SatELite, the variables and clauses whose values can be uniquely determined from other variables and clauses are eliminated to reduce the search space of the given SAT problem. The algorithms used in SatELite have inherent parallelism, but the data size of the SAT problems is very large, and the performance of the system is limited by the throughput of the off-chip DRAM banks. In our implementation, several clauses are held on the FPGA, and are compared in parallel with a sequence of new clauses. The sequence is cached on the FPGA, and reused in order to hide the access delay to the DRAM banks. The speedup by our system depends on the problem size, however it becomes higher for larger problems. Masayuki Suzuki, Tsutomu Maruyama |
FPT | 1 |
| 2011 | Continuous Digits Recognition Leveraging Invariant StructureabstractRecently, an invariant structure of speech was proposed, where the inevitable acoustic variations caused by non-linguistic fac-tors are effectively removed from speech. The invariant struc-ture was applied to isolated word recognition and the experi-mental results showed good performance. However, the pre-vious method can’t apply to continuous speech recognition di-rectly because there was no efficient decoding algorithm. In this paper, we propose a method to leverage the invariant structure in continuous digits recognition. We use a traditional HMM-based Automatic Speech Recognition (ASR) system to get N-best lists with phone alignments. Then we construct invariant structures using these phone alignments and re-rank the N-best lists by investigating which hypothesis is structurally more valid. Experimental results show a relative WER improvement of 17.4 % over the baseline HMM-based ASR system. Masayuki Suzuki, Gakuto Kurata, Masafumi Nishimura, Nobuaki Minematsu |
INTERSPEECH | 1 |
| 2010 | Detecting Patterns in Various Size and Angle Using FPGAabstractIn this paper, we describe an approach for detecting patterns in various size and angle using FPGA. In many approaches, features of a given pattern which are invariant to scaling and/or rotation are defined in advance, and those features are searched in a given image. These approaches make it possible to narrow down the candidate regions with less computational cost, but the sensitivity depends on how to define the features. In our approach, the image is downscaled by αkand αlalong the x and y axes (k, l = 0, 1, 2, ..., n), and the regions in the downscaled images are compared with several sequences of the templates which are generated from the given pattern using direct cross-correlation. This approach requires high computational cost, but by calculating the cross-correlations incrementally starting from the nonrotated pattern, it becomes possible to detect the rotated patterns in various size and angle with one FPGA. Masayuki Suzuki, Yoshifumi Tanida, Tsutomu Maruyama |
FPL | 1 |
| 2010 | Integration of multilayer regression analysis with structure-based pronunciation assessmentabstractAutomatic pronunciation assessment has several difficulties. Adequacy in controlling the vocal organs is often estimated from the spectral envelopes of input utterances but the envelope patterns are also affected by other factors such as speaker iden-tity. Recently, a new method of speech representation was pro-posed where these non-linguistic variations are effectively re-moved through modeling only the contrastive aspects of speech features. This speech representation is called speech struc-ture. However, the often excessively high dimensionality of the speech structure can degrade the performance of structure-based pronunciation assessment. To deal with this problem, we integrate multilayer regression analysis with the structure-based assessment. The results show higher correlation between hu-man and machine scores and also show much higher robustness to speaker differences compared to widely used GOP-based analysis. Index Terms: CALL, speech structure, regression, GOP 1. Masayuki Suzuki, Yu Qiao 0001, Nobuaki Minematsu, Keikichi Hirose |
INTERSPEECH | 1 |
| 2010 | A new type of omnidirectional wheelchair robot for walking support and power assistanceabstractUp to now, many robotic aids for the elderly's walking support or the disabled's walking rehabilitation are reported, and numerous electrical-powered wheelchairs are developed. In this paper, a new kind of omnidirectional wheelchair typed robot is developed. The robot not only can accomplish the walking support or walking rehabilitation as the elderly or the disabled walk, but also can realize the power assistance for a caregiver when he/she pushes the robot to move. The basic structure of the robot is described, and the omnidirectional mobility of the robot is analyzed. Further, an admittance based human-machine interaction controller is introduced for power assistance. Experiments are implemented, and the experimental results show that the pushing force can be reduced and well controlled arbitrarily as designed. The development purposes of the robot for walking support and power assistance are achieved. Chi Zhu 0001, Masashi Oda, Masayuki Suzuki, Xiang Luo 0001, Hideomi Watanabe, Yuling Yan |
IROS | 3 |
| 2010 | Admittance based control of wheelchair typed omnidirectional robot for walking support and power assistanceabstractIn this paper, a new type of omnidirectional mobile robot is developed, that not only can be used for the elderly's walking support, the disabled's walking rehabilitation, but also can be used for a caregiver's power assistance when he/she pushes the robot to move while the elderly or the disabled is sitting in the seat. The omnidirectional mobility of the robot is analyzed, and an admittance based human-machine interaction controller is introduced for power assistance. The experimental results show that the pushing force is reduced and well controlled as we planned. The purpose of walking support and power assistance is achieved. Masashi Oda, Chi Zhu 0001, Masayuki Suzuki, Xiang Luo 0001, Hideomi Watanabe, Yuling Yan |
RO-MAN | 3 |
| 2009 | A study on Hidden Structural Model and its application to labeling sequencesabstractThis paper proposes hidden structure model (HSM) for statistical modeling of sequence data. The HSM generalizes our previous proposal on structural representation by introducing hidden states and probabilistic models. Compared with the previous structural representation, HSM not only can solve the problem of misalignment of events, but also can conduct structure-based decoding, which allows us to apply HSM to general speech recognition tasks. Different from HMM, HSM accounts for the probability of both locally absolute and globally contrastive features. This paper focuses on the fundamental formulation and theories of HSM. We also develop methods for the problems of state inference, probability calculation and parameter estimation of HSM. Especially, we show that the state inference of HSM can be reduced to a quadratic programming problem. We carry out two experiments to examine the performance of HSM on labeling sequences. The first experiment tests HSM by using artificially transformed sequences, and the second experiment is based on a Japanese corpus of connected vowel utterances. The experimental results demonstrate the effectiveness of HSM. Yu Qiao 0001, Masayuki Suzuki, Nobuaki Minematsu |
ASRU | 2 |
| 2009 | Sub-structure-based estimation of pronunciation proficiency and classification of learnersabstractAutomatic estimation of pronunciation proficiency has its specific difficulty. Adequacy in controlling the vocal organs can be estimated from spectral envelopes of input utterances but the envelope patterns are also affected easily by different speakers. To develop a pedagogically sound method for automatic estimation, the envelope changes caused by linguistic factors and those by extra-linguistic factors should be properly separated. For this aim, in our previous study [1], we proposed a mathematically-guaranteed and linguistically-valid speaker-invariant representation of pronunciation, called speech structure. After the proposal, we have examined that representation also for ASR [2], [3], [4] and, through these works, we have learned better how to apply speech structures to various tasks. In this paper, we focus on a proficiency estimation experiment done in [1] and, based on our recently proposed techniques for the structures, we carry out that experiment again but under new and different conditions. Here, we use smaller units of structural analysis, speaker-invariant substructures, and relative structural distances between a learner and a teacher. Results show that correlations between human and machine rating are improved and also show extremely higher robustness to speaker differences compared to widely used GOP scores. Further, we also demonstrate that the proposed representation can classify learners purely based on their pronunciation proficiency, not affected by their age and gender. Masayuki Suzuki, Nobuaki Minematsu, Dean Luo, Keikichi Hirose |
ASRU | 1 |
| 2009 | Affine invariant features and their application to speech recognitionabstractThis paper proposes a set of affine invariant features (AIFs) for sequence data. The proposed AIFs can be calculated directly from the sequence data, and their invariance to affine transformation is proved mathematically through algebraic calculation. We apply the AIFs to speech recognition. Since the vocal tract length (VTL) difference causes to frequency warping which can be approximated well by affine transform on cepstral features, the AIFs of cepstral sequence provide robust features for VTL variations. We experimentally examine the invariance of AIFs of speech signals, and apply AIFs for Japanese isolated word recognition. The experimental results show that the combination of AIFs with MFCC or MFCC+Delta can lead to higher recognition rates than MFCC or MFCC+Delta only. Especially in the mismatched experiments, the combination with AIFs can reduce the error rates about 30% when compared to MFCC or MFCC+Delta only. The AIFs are expected to have other applications than speech recognition, since their invariance is general. Yu Qiao 0001, Masayuki Suzuki, Nobuaki Minematsu |
ICASSP | 2 |
| 2006 | Mechanisms of Autonomous Pipe-Surface Inspection Robot with Magnetic ElementsabstractIn this paper, we propose a pipe inspection robot to move automatically along the outside of piping. In recent years, such pipe inspection robots and climbing robots have been the subject of important research activity worldwide [1]-[3]. Inspection of industrial pipe is a well-known and highly practical application of robotic technology. Robots especially designed to inspect the surface of pipe have recently been realized by some companies, and are in practical use as mechanical pipe inspection devices in various plants. Some of the robots have the capability of self-propelled inspection. These robots are equipped with ultrasonic diagnostic equipment that can measure the thickness of the pipe along its surface [4]. However, they are able to move by manual operation only. Such robots cannot be applied to pipes with flanges and valves, etc. We have designed a mechanism that enables robots to traverse flanges, rise along vertical flanges, and move along the underside of the piping [5]-[8]. The proposed robot is composed of three connected units. Masayuki Suzuki, Toshihiro Yukawa, Yuichi Satoh, Hideharu Okano |
SMC | 1 |
| 2006 | Design of Magnetic Wheels in Pipe Inspection RobotabstractIn an industrial plant, straight piping, vertical piping, curved piping, valves, and flanges are arranged in complex patterns. Problems regarding automated pipe inspection originate in the fact that sensors cannot scan complex piping schemes involving flanges, valves, curved piping, and vertical piping. Various types of pipes are used in industrial applications, taking into consideration the temperature and the corrosive properties of the fluid. Oil plants, for example, use mainly carbon steel, stainless steel, and chloridized vinyl pipes. In light of this, we have focused on developing a robot that can inspect carbon steel pipes, having the ability to move automatically along the outside of a length of piping. This paper describes the structure of magnetic wheels used in a pipe inspection robot. We propose a robot that has a mechanism consisting of wheel-type magnets. Using the power generated by the magnets, the robot maintains its attachment to the pipe. The mechanism of the robot is composed of three units, and is able to automatically inspect a pipe's surface, traverse flanges, and rise along vertical piping. To achieve the robot's optimum performance, the design of the magnetic wheels is important. The authors have previously designed a similar robot composed of three connected units, though the adherence power of this robot was weak and could not achieve a high level of stability while advancing along a straight length of pipe. Toshihiro Yukawa, Masayuki Suzuki, Yuichi Satoh, Hideharu Okano |
SMC | 2 |
| 1992 | Three New Algorithms for Multivariate Polynomial GCD
Tateaki Sasaki, Masayuki Suzuki |
J. Symb. Comput. | 2 |