EDBT 2026 Demo / reviewers in the wild / expert
Tetsunori Kobayashi
dblp:33/569
· DBLP profile ↗
126ranked-venue papers
13as first author
20since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 106 · 13 first-author · 18 since 2021Artificial intelligence and machine learning · 78 · 5 first-author · 11 since 2021Systems, architecture and hardware · 4Databases, data management, data science and information retrieval · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Harnessing the Zero-Shot Power of Instruction-Tuned Large Language Model for Guiding End-to-End Speech RecognitionabstractWe propose to utilize an instruction-tuned large language model (LLM) for guiding the text generation process in automatic speech recognition (ASR). Modern LLMs are adept at performing various text generation tasks through zero-shot learning, prompted with instructions designed for specific objectives. This paper explores the potential of LLMs to derive linguistic information that can facilitate text generation in end-to-end ASR models. Specifically, we instruct an LLM to correct grammatical errors in an ASR hypothesis and use the LLM-derived representations to refine the output further. The proposed model is built on the joint CTC and attention architecture, with the LLM serving as a front-end feature extractor for the decoder. The ASR hypothesis, subject to correction, is obtained from the encoder via CTC decoding and fed into the LLM along with a specific instruction. The decoder subsequently takes as input the LLM output to perform token predictions, combining acoustic information from the encoder and the powerful linguistic information provided by the LLM. Experimental results show that the proposed LLM-guided model achieves a relative gain of approximately 13% in word error rates across major benchmarks. Yosuke Higuchi, Tetsuji Ogawa, Tetsunori Kobayashi |
ICASSP | 3 |
| 2025 | End-to-End Speech Translation Guided by Robust Translation Capability of Large Language Model
Yosuke Higuchi, Tetsuji Ogawa, Tetsunori Kobayashi |
INTERSPEECH | 3 |
| 2024 | Hierarchical Multi-Task Learning with CTC and Recursive Operation
Nahomi Kusunoki, Yosuke Higuchi, Tetsuji Ogawa, Tetsunori Kobayashi |
INTERSPEECH | 4 |
| 2024 | Preprocessing for acoustic-to-articulatory inversion using real-time MRI movies of Japanese speech
Anna Oura, Hideaki Kikuchi, Tetsunori Kobayashi |
INTERSPEECH | 3 |
| 2023 | A Single Speech Enhancement Model Unifying Dereverberation, Denoising, Speaker Counting, Separation, And ExtractionabstractWe propose a multi-task universal speech enhancement (MUSE) model that can perform five speech enhancement (SE) tasks: dereverberation, denoising, speech separation (SS), target speaker extraction (TSE), and speaker counting. This is achieved by integrating two modules into an SE model: 1) an internal separation module that does both speaker counting and separation; and 2) a TSE module that extracts the target speech from the internal separation outputs using target speaker cues. The model is trained to perform TSE if the target speaker cue is given and SS otherwise. By training the model to remove noise and reverberation, we allow the model to tackle the five tasks mentioned above with a single model, which has not been accomplished yet. Evaluation results demonstrate that the proposed MUSE model can successfully handle multiple tasks with a single model. Kohei Saijo, Wangyou Zhang, Zhongqiu Wang 0001, Shinji Watanabe 0001, Tetsunori Kobayashi, Tetsuji Ogawa |
ASRU | 5 |
| 2023 | Intermpl: Momentum Pseudo-Labeling With Intermediate CTC LossabstractThis paper presents InterMPL, a semi-supervised learning method of end-to-end automatic speech recognition (ASR) that performs pseudo-labeling (PL) with intermediate supervision. Momentum PL (MPL) trains a connectionist temporal classification (CTC)-based model on unlabeled data by continuously generating pseudo-labels on the fly and improving their quality. In contrast to autoregressive formulations, such as the attention-based encoder-decoder and transducer, CTC is well suited for MPL, or PL-based semi-supervised ASR in general, owing to its simple/fast inference algorithm and robustness against generating collapsed labels. However, CTC generally yields inferior performance than the autoregressive models due to the conditional independence assumption, thereby limiting the performance of MPL. We propose to enhance MPL by introducing intermediate loss, inspired by the recent advances in CTC-based modeling. Specifically, we focus on self-conditional and hierarchical conditional CTC, that apply auxiliary CTC losses to intermediate layers such that the conditional independence assumption is explicitly relaxed. We also explore how pseudo-labels should be generated and used as supervision for intermediate losses. Experimental results in different semi-supervised settings demonstrate that the proposed approach outperforms MPL and improves an ASR model by up to a 12.1% absolute performance gain. In addition, our detailed analysis validates the importance of the intermediate loss. Yosuke Higuchi, Tetsuji Ogawa, Tetsunori Kobayashi, Shinji Watanabe 0001 |
ICASSP | 3 |
| 2023 | BECTRA: Transducer-Based End-To-End ASR with Bert-Enhanced EncoderabstractWe present BERT-CTC-Transducer (BECTRA), a novel end-to-end automatic speech recognition (E2E-ASR) model formulated by the transducer with a BERT-enhanced encoder. Integrating a large-scale pre-trained language model (LM) into E2E-ASR has been actively studied, aiming to utilize versatile linguistic knowledge for generating accurate text. One crucial factor that makes this integration challenging lies in the vocabulary mismatch; the vocabulary constructed for a pre-trained LM is generally too large for E2E-ASR training and is likely to have a mismatch against a target ASR domain. To overcome such an issue, we propose BECTRA, an extended version of our previous BERT-CTC, that realizes BERT-based E2E-ASR using a vocabulary of interest. BECTRA is a transducer-based model, which adopts BERT-CTC for its encoder and trains an ASR-specific decoder using a vocabulary suitable for a target task. With the combination of the transducer and BERT-CTC, we also propose a novel inference algorithm for taking advantage of both autoregressive and non-autoregressive decoding. Experimental results on several ASR tasks, varying in amounts of data, speaking styles, and languages, demonstrate that BECTRA outperforms BERT-CTC by effectively dealing with the vocabulary mismatch while exploiting BERT knowledge. Yosuke Higuchi, Tetsuji Ogawa, Tetsunori Kobayashi, Shinji Watanabe 0001 |
ICASSP | 3 |
| 2023 | Conversation-Oriented ASR with Multi-Look-Ahead CBS ArchitectureabstractDuring conversations, humans are capable of inferring the intention of the speaker at any point of the speech to prepare the following action promptly. Such ability is also the key for conversational systems to achieve rhythmic and natural conversation. To perform this, the automatic speech recognition (ASR) used for transcribing the speech in real-time must achieve high accuracy without delay. In streaming ASR, high accuracy is assured by attending to look-ahead frames, which leads to delay increments. To tackle this trade-off issue, we propose a multiple latency streaming ASR to achieve high accuracy with zero look-ahead. The proposed system contains two encoders that operate in parallel, where a primary encoder generates accurate outputs utilizing look-ahead frames, and the auxiliary encoder recognizes the look-ahead portion of the primary encoder without look-ahead. The proposed system is constructed based on contextual block streaming (CBS) architecture, which leverages block processing and has a high affinity for the multiple latency architecture. Various methods are also studied for architecting the system, including shifting the network to perform as different encoders; as well as generating both encoders’ outputs in one encoding pass. Huaibo Zhao, Shinya Fujie, Tetsuji Ogawa, Jin Sakuma, Yusuke Kida, Tetsunori Kobayashi |
ICASSP | 6 |
| 2023 | Improving the response timing estimation for spoken dialogue systems by reducing the effect of speech recognition delay
Jin Sakuma, Shinya Fujie, Huaibo Zhao, Tetsunori Kobayashi |
INTERSPEECH | 4 |
| 2022 | PostMe: Unsupervised Dynamic Microtask Posting For Efficient and Reliable CrowdsourcingabstractEven after over a decade of many crowdsourcing researches, we have no standard framework for low-cost quality assurance in crowdsourced data annotation. This paper proposes an unsupervised learning method for dynamic microtask posting which allows each microtask to adjust their own number of collected responses based on the data difficulty. Since crowdsourced data labels are likely to contain errors, researchers often employ majority voting that aggregates responses from multiple workers to calculate a final l abel. T his t echnique, h owever, i nvolves a trade-off between label accuracy and cost. This paper presents a dynamic microtask posting model that reduces the total number of collected responses while maintaining the labeling accuracy; we also aim to obtain the model with an “unsupervised” approach, which does not require training through experience of microtask posting for data labeled with ground-truths. Our simulation in annotating livestock surveillance images demonstrated that our approach achieved i) comparable learning performance to that of the supervised approach that required model training with labeled data, and ii) a significant c ost r eduction without degrading accuracy in comparison to simple majority voting. Ryo Yanagisawa, Teppei Nakano, Tetsunori Kobayashi, Tetsuji Ogawa |
IEEE Big Data | 4 |
| 2022 | Phrase-Level Localization of Inconsistency Errors in Summarization by Weak SupervisionabstractAlthough the fluency of automatically generated abstractive summaries has improved significantly with advanced methods, the inconsistency that remains in summarization is recognized as an issue to be addressed. In this study, we propose a methodology for localizing inconsistency errors in summarization. A synthetic dataset that contains a variety of factual errors likely to be produced by a common summarizer is created by applying sentence fusion, compression, and paraphrasing operations. In creating the dataset, we automatically label erroneous phrases and the dependency relations between them as “inconsistent,” which can contribute to detecting errors more adequately than existing models that rely only on dependency arc-level labels. Subsequently, this synthetic dataset is employed as weak supervision to train a model called SumPhrase, which jointly localizes errors in a summary and their corresponding sentences in the source document. The empirical results demonstrate that our SumPhrase model can detect factual errors in summarization more effectively than existing weakly supervised methods owing to the phrase-level labeling. Moreover, the joint identification of error-corresponding original sentences is proven to be effective in improving error detection accuracy. Masato Takatsuka, Tetsunori Kobayashi, Yoshihiko Hayashi |
COLING | 2 |
| 2022 | Hierarchical Conditional End-to-End ASR with CTC and Multi-Granular Subword UnitsabstractIn end-to-end automatic speech recognition (ASR), a model is expected to implicitly learn representations suitable for recognizing a word-level sequence. However, the huge abstraction gap between input acoustic signals and output linguistic tokens makes it challenging for a model to learn the representations. In this work, to promote the word-level representation learning in end-to-end ASR, we propose a hierarchical conditional model that is based on connectionist temporal classification (CTC). Our model is trained by auxiliary CTC losses applied to intermediate layers, where the vocabulary size of each target subword sequence is gradually increased as the layer becomes close to the word-level output. Here, we make each level of sequence prediction explicitly conditioned on the previous sequences predicted at lower levels. With the proposed approach, we expect the proposed model to learn the word-level representations effectively by exploiting a hierarchy of linguistic structures. Experimental results on LibriSpeech-{100h, 960h} and TEDLIUM2 demonstrate that the proposed model improves over a standard CTCbased model and other competitive models from prior work. We further analyze the results to confirm the effectiveness of the intended representation learning with our model. Yosuke Higuchi, Keita Karube, Tetsuji Ogawa, Tetsunori Kobayashi |
ICASSP | 4 |
| 2022 | Confusion Detection for Adaptive Conversational Strategies of An Oral Proficiency Assessment Interview Agent
Mao Saeki, Kotoka Miyagi, Shinya Fujie, Shungo Suzuki, Tetsuji Ogawa, Tetsunori Kobayashi, Yoichi Matsuyama |
INTERSPEECH | 6 |
| 2022 | Response Timing Estimation for Spoken Dialog System using Dialog Act Estimation
Jin Sakuma, Shinya Fujie, Tetsunori Kobayashi |
INTERSPEECH | 3 |
| 2022 | Response Timing Estimation for Spoken Dialog Systems Based on Syntactic Completeness PredictionabstractAppropriate response timing is very important for achieving smooth dialog progression. Conventionally, prosodic, temporal and linguistic features have been used to determine timing. In addition to the conventional parameters, we propose to utilize the syntactic completeness after a certain time, which represents whether the other party is about to finish speaking. We generate the next token sequence from intermediate speech recognition results using a language model and obtain the probability of the end of utterance appearing$K$tokens ahead, where$K$varies from 1 to$M$. We obtain an$M$-dimensional vector, which we denote as estimates of syntactic completeness (ESC). We evaluated this method on a simulated dialog database of a restaurant information center. The results confirmed that considering ESC improves the performance of response timing estimation, especially the accuracy in quick responses, compared with the method using only conventional features. Jin Sakuma, Shinya Fujie, Tetsunori Kobayashi |
SLT | 3 |
| 2021 | Improved Mask-CTC for Non-Autoregressive End-to-End ASRabstractFor real-world deployment of automatic speech recognition (ASR), the system is desired to be capable of fast inference while relieving the requirement of computational resources. The recently proposed end-to-end ASR system based on mask-predict with connectionist temporal classification (CTC), Mask-CTC, fulfills this demand by generating tokens in a non-autoregressive fashion. While Mask-CTC achieves remarkably fast inference speed, its recognition performance falls behind that of conventional autoregressive (AR) systems. To boost the performance of Mask-CTC, we first propose to enhance the encoder network architecture by employing a recently proposed architecture called Conformer. Next, we propose new training and decoding methods by introducing auxiliary objective to predict the length of a partial target sequence, which allows the model to delete or insert tokens during inference. Experimental results on different ASR tasks show that the proposed approaches improve Mask-CTC significantly, outperforming a standard CTC model (15.5% → 9.1% WER on WSJ). Moreover, Mask-CTC now achieves competitive results to AR models with no degradation of inference speed (< 0.1 RTF using CPU). We also show a potential application of Mask-CTC to end-to-end speech translation. Yosuke Higuchi, Hirofumi Inaguma, Shinji Watanabe 0001, Tetsuji Ogawa, Tetsunori Kobayashi |
ICASSP | 5 |
| 2021 | Timing Generating Networks: Neural Network Based Precise Turn-Taking Timing Prediction in Multiparty Conversation
Shinya Fujie, Hayato Katayama, Jin Sakuma, Tetsunori Kobayashi |
Interspeech | 4 |
| 2021 | Efficient and Stable Adversarial Learning Using Unpaired Data for Unsupervised Multichannel Speech Separation
Yu Nakagome, Masahito Togami, Tetsuji Ogawa, Tetsunori Kobayashi |
Interspeech | 4 |
| 2021 | Analysis of Multimodal Features for Speaking Proficiency Scoring in an Interview DialogueabstractThis paper analyzes the effectiveness of different modalities in automated speaking proficiency scoring in an online dialogue task of non-native speakers. Conversational competence of a language learner can be assessed through the use of multimodal behaviors such as speech content, prosody, and visual cues. Although lexical and acoustic features have been widely studied, there has been no study on the usage of visual features, such as facial expressions and eye gaze. To build an automated speaking proficiency scoring system using multi-modal features, we first constructed an online video interview dataset of 210 Japanese English-learners with annotations of their speaking proficiency. We then examined two approaches for incorporating visual features and compared the effectiveness of each modality. Results show the end-to-end approach with deep neural networks achieves a higher correlation with human scoring than one with handcrafted features. Modalities are effective in the order of lexical, acoustic, and visual features. Mao Saeki, Yoichi Matsuyama, Satoshi Kobashikawa, Tetsuji Ogawa, Tetsunori Kobayashi |
SLT | 5 |
| 2021 | Personalized Extractive Summarization for a News Dialogue SystemabstractIn modern society, people's interests and preferences are diversifying. Along with this, the demand for personalized summarization technology is increasing. In this study, we propose a method for generating summaries tailored to each user's interests using profile features obtained from questionnaires administered to users of our spoken-dialogue news delivery system. We propose a method that collects and uses the obtained user profile features to generate a summary tailored to each user's interests, specifically, the sentence features obtained by BERT and user profile features obtained from the questionnaire result. In addition, we propose a method for extracting sentences by solving an integer linear programming problem that considers redundancy and context coherence, using the degree of interest in sentences estimated by the model. The results of our experiments confirmed that summaries generated based on the degree of interest in sentences estimated using user profile information can transmit information more efficiently than summaries based solely on the importance of sentences. Hiroaki Takatsu, Mayu Okuda, Yoichi Matsuyama, Hiroshi Honda, Shinya Fujie, Tetsunori Kobayashi |
SLT | 6 |
| 2020 | Sentiment Analysis for Emotional Speech Synthesis in a News Dialogue SystemabstractAs smart speakers and conversational robots become ubiquitous, the demand for expressive speech synthesis has increased.In this paper, to control the emotional parameters of the speech synthesis according to certain dialogue contents, we construct a news dataset with emotion labels ("positive," "negative," or "neutral") annotated for each sentence.We then propose a method to identify emotion labels using a model combining BERT and BiLSTM-CRF, and evaluate its effectiveness using the constructed dataset.The results showed that the classification model performance can be efficiently improved by preferentially annotating news articles with low confidence in the human-in-the-loop machine learning framework. Hiroaki Takatsu, Ryota Ando, Yoichi Matsuyama, Tetsunori Kobayashi |
COLING | 4 |
| 2020 | Exploiting Narrative Context and A Priori Knowledge of Categories in Textual Emotion ClassificationabstractRecognition of the mental state of a human character in text is a major challenge in natural language processing. In this study, we investigate the efficacy of the narrative context in recognizing the emotional states of human characters in text and discuss an approach to make use of a priori knowledge regarding the employed emotion category system. Specifically, we experimentally show that the accuracy of emotion classification is substantially increased by encoding the preceding context of the target sentence using a BERT-based text encoder. We also compare ways to incorporate a priori knowledge of emotion categories by altering the loss function used in training, in which our proposal of multi-task learning that jointly learns to classify positive/negative polarity of emotions is included. The experimental results suggest that, when using Plutchik’s Wheel of Emotions, it is better to jointly classify the basic emotion categories with positive/negative polarity rather than directly exploiting its characteristic structure in which eight basic categories are arranged in a wheel. Hikari Tanabe, Tetsuji Ogawa, Tetsunori Kobayashi, Yoshihiko Hayashi |
COLING | 3 |
| 2020 | Deep Speech Extraction with Time-Varying Spatial Filtering Guided By Desired Direction AttractorabstractIn this investigation, a deep neural network (DNN) based speech extraction method is proposed to enhance a speech signal propagating from the desired direction. The proposed method integrates knowledge based on a sound propagation model and the time-varying characteristics of a speech source, into a DNN-based separation framework. This approach outputs a separated speech source using time-varying spatial filtering, which achieves superior speech extraction performance compared with time-invariant spatial filtering. Given that the gradient of all modules can be calculated, back-propagation can be performed to maximize the speech quality of the output signal in an end-to-end manner. Guided information is also modeled based on the sound propagation model, which facilitates disentangled representations of the target speech source and noise signals. The experimental results demonstrate that the proposed method can extract the target speech source more accurately than conventional DNN-based speech source separation and conventional speech extraction using time-invariant spatial filtering. Yu Nakagome, Masahito Togami, Tetsuji Ogawa, Tetsunori Kobayashi |
ICASSP | 4 |
| 2020 | Exploring and Exploiting the Hierarchical Structure of a Scene for Scene Graph Generation
Ikuto Kurosawa, Tetsunori Kobayashi, Yoshihiko Hayashi |
ICPR | 2 |
| 2020 | Mask CTC: Non-Autoregressive End-to-End ASR with CTC and Mask PredictabstractWe present Mask CTC, a novel non-autoregressive end-to-end automatic speech recognition (ASR) framework, which generates a sequence by refining outputs of the connectionist temporal classification (CTC).Neural sequence-to-sequence models are usually autoregressive: each output token is generated by conditioning on previously generated tokens, at the cost of requiring as many iterations as the output length.On the other hand, non-autoregressive models can simultaneously generate tokens within a constant number of iterations, which results in significant inference time reduction and better suits end-toend ASR model for real-world scenarios.In this work, Mask CTC model is trained using a Transformer encoder-decoder with joint training of mask prediction and CTC.During inference, the target sequence is initialized with the greedy CTC outputs and low-confidence tokens are masked based on the CTC probabilities.Based on the conditional dependence between output tokens, these masked low-confidence tokens are then predicted conditioning on the high-confidence tokens.Experimental results on different speech recognition tasks show that Mask CTC outperforms the standard CTC model (e.g., 17.9% → 12.1% WER on WSJ) and approaches the autoregressive model, requiring much less inference time using CPUs (0.07 RTF in Python implementation).All of our codes are publicly available at https://github.com/espnet/espnet. Yosuke Higuchi, Shinji Watanabe 0001, Nanxin Chen, Tetsuji Ogawa, Tetsunori Kobayashi |
INTERSPEECH | 5 |
| 2020 | Mentoring-Reverse Mentoring for Unsupervised Multi-Channel Speech Source Separation
Yu Nakagome, Masahito Togami, Tetsuji Ogawa, Tetsunori Kobayashi |
INTERSPEECH | 4 |
| 2020 | Word Attribute Prediction Enhanced by Lexical Entailment TasksabstractHuman semantic knowledge about concepts acquired through perceptual inputs and daily experiences can be expressed as a bundle of attributes. Unlike the conventional distributed word representations that are purely induced from a text corpus, a semantic attribute is associated with a designated dimension in attribute-based vector representations. Thus, semantic attribute vectors can effectively capture the commonalities and differences among concepts. However, as semantic attributes have been generally created by psychological experimental settings involving human annotators, an automatic method to create or extend such resources is highly demanded in terms of language resource development and maintenance. This study proposes a two-stage neural network architecture, Word2Attr, in which initially acquired attribute representations are then fine-tuned by employing supervised lexical entailment tasks. The quantitative empirical results demonstrated that the fine-tuning was indeed effective in improving the performances of semantic/visual similarity/relatedness evaluation tasks. Although the qualitative analysis confirmed that the proposed method could often discover valid but not-yet human-annotated attributes, they also exposed future issues to be worked: we should refine the inventory of semantic attributes that currently relies on an existing dataset. Mika Hasegawa, Tetsunori Kobayashi, Yoshihiko Hayashi |
LREC | 2 |
| 2019 | Postfiltering Using an Adversarial Denoising Autoencoder with Noise-aware TrainingabstractAn adversarial denoising autoencoder (ADAE) with noise-aware training is proposed and successfully applied to post-filtering for linear noise reduction. The ADAE is effective for attenuating interference sounds, however, it is difficult to learn to handle its various unexpected harmful effects (e.g., various types of noise) using a single network. Legacy speech enhancement was introduced as a pre-processor to make it possible to efficiently train the ADAEs by reducing the unexpected variabilities in the inputs to the ADAEs. Time-frequency masking performed well to suppress the variabilities, however, it induced unpleasant distortion, which is difficult for the ADAE to complement. In this paper, a minimum variance distortionless response (MVDR) beam-former, which can avoid troublesome non-linear distortions, is exploited as a preprocessor, and the MVDR outputs are used as the inputs to the ADAE-based post-filter. In addition, noise-dominant signals derived from the MVDR beamformer can improve the accuracy of the ADAE-based post-filter because the residual noise depends on the original noise signals. Experimental comparisons conducted using multichannel speech enhancement demonstrate that ADAE-based post-filtering yields significant improvements over the MVDR-and ADAE-based speech enhancement systems, and noise-aware training of ADAE works well. Naohiro Tawara, Hikari Tanabe, Tetsunori Kobayashi, Masaru Fujieda, Kazuhiro Katagiri, Takashi Yazu, Tetsuji Ogawa |
ICASSP | 3 |
| 2019 | Speaker Adversarial Training of DPGMM-Based Feature Extractor for Zero-Resource Languages
Yosuke Higuchi, Naohiro Tawara, Tetsunori Kobayashi, Tetsuji Ogawa |
INTERSPEECH | 3 |
| 2019 | Recognition of Intentions of Users' Short Responses for Conversational News Delivery System
Hiroaki Takatsu, Katsuya Yokoyama, Yoichi Matsuyama, Hiroshi Honda, Shinya Fujie, Tetsunori Kobayashi |
INTERSPEECH | 6 |
| 2019 | Multi-Channel Speech Enhancement Using Time-Domain Convolutional Denoising Autoencoder
Naohiro Tawara, Tetsunori Kobayashi, Tetsuji Ogawa |
INTERSPEECH | 2 |
| 2019 | TurkScanner: Predicting the Hourly Wage of MicrotasksabstractWorkers in crowd markets struggle to earn a living. One reason for this is that it is difficult for workers to accurately gauge the hourly wages of microtasks, and they consequently end up performing labor with little pay. In general, workers are provided with little information about tasks, and are left to rely on noisy signals, such as textual description of the task or rating of the requester. This study explores various computational methods for predicting the working times (and thus hourly wages) required for tasks based on data collected from other workers completing crowd work. We provide the following contributions. (i) A data collection method for gathering real-world training data on crowd-work tasks and the times required for workers to complete them; (ii) TurkScanner: a machine learning approach that predicts the necessary working time to complete a task (and can thus implicitly provide the expected hourly wage). We collected 9,155 data records using a web browser extension installed by 84 Amazon Mechanical Turk workers, and explored the challenge of accurately recording working times both automatically and by asking workers. TurkScanner was created using ~ 150 derived features, and was able to predict the hourly wages of 69.6% of all the tested microtasks within a 75% error. Directions for future research include observing the effects of tools on people's working practices, adapting this approach to a requester tool for better price setting, and predicting other elements of work (e.g., the acceptance likelihood and worker task preferences.) Chun-Wei Chiang, Saiph Savage, Teppei Nakano, Tetsunori Kobayashi, Jeffrey P. Bigham |
WWW | 5 |
| 2018 | Answerable or Not: Devising a Dataset for Extending Machine Reading ComprehensionabstractMachine-reading comprehension (MRC) has recently attracted attention in the fields of natural language processing and machine learning. One of the problematic presumptions with current MRC technologies is that each question is assumed to be answerable by looking at a given text passage. However, to realize human-like language comprehension ability, a machine should also be able to distinguish not-answerable questions (NAQs) from answerable questions. To develop this functionality, a dataset incorporating hard-to-detect NAQs is vital; however, its manual construction would be expensive. This paper proposes a dataset creation method that alters an existing MRC dataset, the Stanford Question Answering Dataset, and describes the resulting dataset. The value of this dataset is likely to increase if each NAQ in the dataset is properly classified with the difficulty of identifying it as an NAQ. This difficulty level would allow researchers to evaluate a machine’s NAQ detection performance more precisely. Therefore, we propose a method for automatically assigning difficulty level labels, which measures the similarity between a question and the target text passage. Our NAQ detection experiments demonstrate that the resulting dataset, having difficulty level annotations, is valid and potentially useful in the development of advanced MRC models. Mao Nakanishi, Tetsunori Kobayashi, Yoshihiko Hayashi |
COLING | 2 |
| 2018 | Language Model Domain Adaptation Via Recurrent Neural Networks with Domain-Shared and Domain-Specific RepresentationsabstractTraining recurrent neural network language models (RNNLMs) requires a large amount of data, which is difficult to collect for specific domains such as multiparty conversations. Data augmentation using external resources and model adaptation, which adjusts a model trained on a large amount of data to a target domain, have been proposed for low-resource language modeling. While there are the commonalities and discrepancies between the source and target domains in terms of the statistics of words and their contexts, these methods for domain adaptation make the commonalities and discrepancies jumbled. We propose novel domain adaptation techniques for RNNLM by introducing domain-shared and domain-specific word embedding and contextual features. This explicit modeling of the commonalities and discrepancies would improve the language modeling performance. Experimental comparisons using multiparty conversation data as the target domain augmented by lecture data from the source domain demonstrate that the proposed domain adaptation method exhibits improvements in the perplexity and word error rate over the long short-term memory based language model (LSTMLM) trained using the source and target domain data. Tsuyoshi Morioka, Naohiro Tawara, Tetsuji Ogawa, Atsunori Ogawa, Tomoharu Iwata, Tetsunori Kobayashi |
ICASSP | 6 |
| 2018 | Speaker Invariant Feature Extraction for Zero-Resource Languages with Adversarial LearningabstractWe introduce a novel type of representation learning to obtain a speaker invariant feature for zero-resource languages. Speaker adaptation is an important technique to build a robust acoustic model. For a zero-resource language, however, conventional model-dependent speaker adaptation methods such as constrained maximum likelihood linear regression are insufficient because the acoustic model of the target language is not accessible. Therefore, we introduce a model-independent feature extraction based on a neural network. Specifically, we introduce a multi-task learning to a bottleneck feature-based approach to make bottleneck feature invariant to a change of speakers. The proposed network simultaneously tackles two tasks: phoneme and speaker classifications. This network trains a feature extractor in an adversarial manner to allow it to map input data into a discriminative representation to predict phonemes, whereas it is difficult to predict speakers. We conduct phone discriminant experiments in Zero Resource Speech Challenge 2017. Experimental results showed that our multi-task network yielded more discriminative features eliminating the variety in speakers. Taira Tsuchiya, Naohiro Tawara, Tetsuji Ogawa, Tetsunori Kobayashi |
ICASSP | 4 |
| 2018 | Sequential Fish Catch Forecasting Using Bayesian State Space ModelsabstractA new state space model suitable for fixed shore net fishing is proposed and successfully applied to daily fish catch forecasting. Accurate prediction of daily fish catches makes it possible to support fishery workers with decision-making for efficient operations. For that purpose, the predictive model should be intuitive to the fishery workers and provide an estimate with a confidence. In the present paper, a fish catch forecasting method is developed using a state space model that emulates the process of fixed shore net fishing. In this method, the parameter estimation and prediction are sequentially performed using the Hamiltonian Monte Carlo method. The experimental comparisons using actual fish catch data and public meteorological information demonstrated that the proposed forecasting system yielded significant reductions in predictive errors over the systems based on decision-trees and legacy state-space models. Yuya Kokaki, Naohiro Tawara, Tetsunori Kobayashi, Kazuo Hashimoto, Tetsuji Ogawa |
ICPR | 3 |
| 2018 | Fine-grained Video Retrieval using Query Phrases - Waseda_Meisei TRECVID 2017 AVS System -abstractIn this paper, a joint team from Waseda University and Meisei University (team name: Waseda_Meisei) report their efforts on the ad-hoc video search (AVS) task for the TRECVID benchmark, which is conducted annually by the National Institute of Standards and Technology (NIST). For the AVS task, a system is required to perform a fine-grained search of target videos from a large-scale video database using a query phrase including multiple keywords, such as objects, persons, scenes, and actions. The system we submitted has the following two characteristics. First, to improve the coverage rate of classes corresponding to keywords in query phrases, we prepared a large number of classifiers that can detect objects, persons, scenes, and actions, which were trained using various image and video datasets. Second, when choosing a concept classifier corresponding to a keyword, we introduced a mechanism that allows us to select additional concept classifiers by incorporating natural language processing techniques. We submitted multiple systems with these characteristics to the TRECVID 2017 AVS task and one of our systems ranked the highest among all the submitted systems from 22 teams. Kazuya Ueki, Koji Hirakawa, Kotaro Kikuchi, Tetsunori Kobayashi |
ICPR | 4 |
| 2018 | Social Image Tags as a Source of Word Embeddings: A Task-oriented Evaluation
Mika Hasegawa, Tetsunori Kobayashi, Yoshihiko Hayashi |
LREC | 2 |
| 2018 | Investigation of Users' Short Responses in Actual Conversation System and Automatic Recognition of their IntentionsabstractIn human-human conversations, listeners often convey intentions to speakers through feedback consisting of reflexive short responses. The speakers recognize these intentions and change the conversational plans to make communication more efficient. These functions are expected to be effective in human-system conversations also; however, there is only a few systems using these functions or a research corpus including such functions. We created a corpus that consists of users' short responses to an actual conversation system and developed a model for recognizing the intention of these responses. First, we categorized the intention of feedback that affects the progress of conversations. We then collected 15604 short responses of users from 2060 conversation sessions using our news-delivery conversation system. Twelve annotators labeled each utterance based on intention through a listening test. We then designed our deep-neural-network-based intention recognition model using the collected data. We found that feedback in the form of questions, which is the most frequently occurring expression, was correctly recognized and contributed to the efficiency of the conversation system. Katsuya Yokoyama, Hiroaki Takatsu, Hiroshi Honda, Shinya Fujie, Tetsunori Kobayashi |
SLT | 5 |
| 2017 | Prosody Control of Utterance Sequence for Information Delivering
Ishin Fukuoka, Kazuhiko Iwata, Tetsunori Kobayashi |
INTERSPEECH | 3 |
| 2017 | Associative Memory Model-Based Linear Filtering and Its Application to Tandem Connectionist Blind Source SeparationabstractWe propose a blind source separation method that yields high-quality speech with low distortion. Time-frequency (TF) masking can effectively reduce interference, but it produces nonlinear distortion. By contrast, linear filtering using a separation matrix such as independent vector analysis (IVA) can avoid nonlinear distortion, but the separation performance is reduced under reverberant conditions. The tandem connectionist approach combines several separation methods and it has been used frequently to compensate for the disadvantages of these methods. In this study, we propose associative memory model (AMM)-based linear filtering and a tandem connectionist framework, which applies TF masking followed by linear filtering. By using AMM trained with speech spectra to optimize the separation matrix, the proposed linear filtering method considers the properties of speech that are not considered explicitly in IVA, such as the harmonic components of spectra. TF masking is applied in the proposed tandem connectionist framework to reduce unwanted components that hinder the optimization of the separation matrix, and it is approximated by using a linear separation matrix to reduce nonlinear distortion. The results obtained in simultaneous speech separation experiments demonstrate that although the proposed linear filtering method can increase the signal-to-distortion ratio (SDR) and signal-to-interference ratio (SIR) compared with IVA, the proposed tandem connectionist framework can obtain greater increases in SDR and SIR, and it reduces the phoneme error rate more than the proposed linear filtering method. Motoi Omachi, Tetsuji Ogawa, Tetsunori Kobayashi |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | A Spoken Dialog System for Coordinating Information Consumption and ExplorationabstractPassive consumption of information is boring in most cases and even painful in some cases, especially when the information content is delivered by employing speech media. The user of a speech-based information delivery system, for example a text-to-speech system, usually cannot interrupt the ongoing information flow, inhibiting her/him to confirm some part of the content, or to pose an inquiry for further information exploration. We argue that a carefully designed spoken dialog system could remedy these undesirable situations, and further enable an enjoyable conversation with the users. The key technologies to realize such an attractive dialog system are: (1) pre-compilation of a dialog plan based on the analysis of a source content, and (2) the dynamic recognition of user's state of understanding and interests. This paper illustrates technical views to implement these functionalities, and discusses a dialog example to exemplify the technical merits of the proposed system. Shinya Fujie, Ishin Fukuoka, Asumi Mugita, Hiroaki Takatsu, Yoshihiko Hayashi, Tetsunori Kobayashi |
CHIIR | 6 |
| 2016 | Improving semantic video indexing: Efforts in Waseda TRECVID 2015 SIN systemabstractIn this paper, we propose a method for improving the performance of semantic video indexing. Our approach involves extracting features from multiple convolutional neural networks (CNNs), creating multiple classifiers, and integrating them. We employed four measures to accomplish this: (1) utilizing multiple evidences observed in each video and effectively compressing them into a fixed-length vector; (2) introducing gradient and motion features to CNNs; (3) enriching variations of the training and the testing sets; and (4) extracting features from several CNNs trained with various large-scale datasets. Using the test dataset from TRECVID's 2014 evaluation benchmark, we evaluated the performance of the proposal in terms of the mean extended inferred average precision measure. On this measure, our system's performance was 35.7, outperforming the state-of-the-art TRECVID 2014 benchmark performance of 33.2. Based on this work, our submission at TRECVID 2015 was ranked second among all submissions. Kazuya Ueki, Tetsunori Kobayashi |
ICASSP | 2 |
| 2015 | A comparative study of spectral clustering for i-vector-based speaker clustering under noisy conditionsabstractThe present paper dealt with speaker clustering for speech corrupted by noise. In general, the performance of speaker clustering significantly depends on how well the similarities between speech utterances can be measured. The recently proposed i-vector-based cosine similarity has yielded the state-of-the-art performance in speaker clustering systems. However, this similarity often fails to capture the speaker similarity under noisy conditions. Therefore, we attempted to examine the efficiency of spectral clustering on i-vector-based similarity for speech corrupted by noise because spectral clustering can yield robustness against noise by non-linear projection. Experimental comparisons demonstrated that spectral clustering yielded significant improvement from conventional methods, such as agglomerative clustering and k-means clustering, under non-stationary noise conditions. Naohiro Tawara, Tetsuji Ogawa, Tetsunori Kobayashi |
ICASSP | 3 |
| 2015 | Multiscale recurrent neural network based language model
Tsuyoshi Morioka, Tomoharu Iwata, Takaaki Hori, Tetsunori Kobayashi |
INTERSPEECH | 4 |
| 2015 | Bilinear map of filter-bank outputs for DNN-based speech recognition
Tetsuji Ogawa, Kenshiro Ueda, Kouichi Katsurada, Tetsunori Kobayashi, Tsuneo Nitta |
INTERSPEECH | 4 |
| 2015 | Four-participant group conversation: A facilitation robot controlling engagement density as the fourth participantabstractIn this paper, we present a framework for facilitation robots that regulate imbalanced engagement density in a four-participant conversation as the forth participant with proper procedures for obtaining initiatives. Four is the special number in multiparty conversations. In three-participant conversations, the minimum unit for multiparty conversations, social imbalance, in which a participant is left behind in the current conversation, sometimes occurs. In such scenarios, a conversational robot has the potential to objectively observe and control situations as the fourth participant. Consequently, we present model procedures for obtaining conversational initiatives in incremental steps to harmonize such four-participant conversations. During the procedures, a facilitator must be aware of both the presence of dominant participants leading the current conversation and the status of any participant that is left behind. We model and optimize these situations and procedures as a partially observable Markov decision process (POMDP), which is suitable for real-world sequential decision processes. The results of experiments conducted to evaluate the proposed procedures show evidence of their acceptability and feeling of groupness. Yoichi Matsuyama, Iwao Akiba, Shinya Fujie, Tetsunori Kobayashi |
Comput. Speech Lang. | 4 |
| 2015 | Automatic Expressive Opinion Sentence Generation for Enjoyable Conversational SystemsabstractIn terms of functional conversations, Grice's Maxim of Quantity suggests that responses should contain no more information than was explicitly asked for. However, in our daily conversations, more informative response skills are usually employed in order to hold enjoyable conversations with interlocutors. These responses are usually produced as forms of one's additional opinions, which usually contain their original viewpoints as well as novel means of expression, rather than simple and common responses characteristic of the general public. In this paper, we propose automatic expressive opinion sentence generation mechanisms for enjoyable conversational systems. The generated opinions are extracted from a large number of reviews on the web, and ranked in terms of contextual relevance, length of sentences, and amount of information represented by the frequency of adjectives. The sentence generator also has an additional phrasing skill. Three controlled lab experiments were conducted, where subjects were requested to read generated sentences and watch videos filmed about conversations between the robot and a person. The results implied that mechanisms effectively promote users' enjoyment and interests. Yoichi Matsuyama, Akihiro Saito, Shinya Fujie, Tetsunori Kobayashi |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2014 | Effect of frequency weighting on MLP-based speaker canonicalization
Yuichi Kubota, Motoi Omachi, Tetsuji Ogawa, Tetsunori Kobayashi, Tsuneo Nitta |
INTERSPEECH | 4 |
| 2013 | Speaker's intentions conveyed to listeners by sentence-final particles and their intonations in Japanese conversational speechabstractWe investigated listeners' perception of speaker's intention depending on sentence-final particles and their intonations in Japanese conversational speech in order to build a speech synthesis system that can express different intentions and subtle nuances. First, we clustered F0 contours derived from approximately 2000 sentence-final syllables and found the sentence-final F0 contours varied a great deal. Next, we selected six distinctive F0 contours that gave perceptually different intonations from among the cluster centroids, and subjectively evaluated synthesized sentence utterances that had various sentence-final particles and their intonations. Results showed that suitable combinations of a sentence-final particle and its intonation should be used to precisely convey the intention to the listeners, and whether the sentence was positive or negative also affected the listeners' perception of the intention. Kazuhiko Iwata, Tetsunori Kobayashi |
ICASSP | 2 |
| 2013 | A Four-Participant Group Facilitation Framework for Conversational Robots
Yoichi Matsuyama, Iwao Akiba, Akihiro Saito, Tetsunori Kobayashi |
SIGDIAL Conference | 4 |
| 2012 | Fully Bayesian inference of multi-mixture Gaussian model and its evaluation using speaker clusteringabstractThis study aims to verify effective optimization methods for estimating parametric, fully Bayesian models in speech processing. For that purpose, we investigate the impact of the difference in optimization methods for the multi-scale Gaussian mixture model, which is suitable for speaker clustering, on the clustering accuracy. The Markov chain Monte Carlo (MCMC)-based method was compared with the variational Bayesian method in the speaker clustering experiment; with a small amount of data, the MCMC-based method was more effective; with large scale data (more than one million samples), the difference between these methods in terms of the clustering accuracy decreased and the MCMC-based method was computationally efficient. Naohiro Tawara, Tetsuji Ogawa, Shinji Watanabe 0001, Tetsunori Kobayashi |
ICASSP | 4 |
| 2012 | Expressing Speaker's Intentions through Sentence-Final Intonations for Japanese Conversational Speech Synthesis
Kazuhiko Iwata, Tetsunori Kobayashi |
INTERSPEECH | 2 |
| 2012 | Fully Bayesian speaker clustering based on hierarchically structured utterance-oriented Dirichlet process mixture model
Naohiro Tawara, Tetsuji Ogawa, Shinji Watanabe 0001, Atsushi Nakamura, Tetsunori Kobayashi |
INTERSPEECH | 5 |
| 2011 | Subspace pursuit method for kernel-log-linear modelsabstractThis paper presents a novel method for reducing the dimensionality of kernel spaces. Recently, to maintain the convexity of training, log linear models without mixtures have been used as emission probability density functions in hidden Markov models for automatic speech recognition. In that framework, nonlinearly-transformed high-dimensional features are used to achieve the nonlinear classification of the original observation vectors without using mixtures. In this paper, with the goal of using high-dimensional features in kernel spaces, the cutting plane subspace pursuit method proposed for support vector machines is generalized and applied to log-linear models. The experimental results show that the proposed method achieved an efficient approximation of the feature space by using a limited number of basis vectors. Yotaro Kubo, Simon Wiesler, Ralf Schlüter, Hermann Ney, Shinji Watanabe 0001, Atsushi Nakamura, Tetsunori Kobayashi |
ICASSP | 7 |
| 2011 | Speaker recognition using multiple kernel learning based on conditional entropy minimizationabstractWe applied a multiple kernel learning (MKL) method based on information-theoretic optimization to speaker recognition. Most of the kernel methods applied to speaker recognition systems require a suitable kernel function and its parameters to be determined for a given data set. In contrast, MKL eliminates the need for strict determination of the kernel function and parameters by using a convex combination of element kernels. In the present paper, we describe an MKL algorithm based on conditional entropy minimization (MCEM). We experimentally verified the effectiveness of MCEM for speaker classification; this method reduced the speaker error rate as compared to conventional methods. Tetsuji Ogawa, Hideitsu Hino, Nima Reyhani, Noboru Murata, Tetsunori Kobayashi |
ICASSP | 5 |
| 2011 | Speaker Verification Robust to Talking Style Variation Using Multiple Kernel Learning Based on Conditional Entropy Minimization
Tetsuji Ogawa, Hideitsu Hino, Noboru Murata, Tetsunori Kobayashi |
INTERSPEECH | 4 |
| 2011 | Spatial Filter Calibration Based on Minimization of Modified LSD
Nobuaki Tanaka, Tetsuji Ogawa, Tetsunori Kobayashi |
INTERSPEECH | 3 |
| 2011 | Speaker Clustering Based on Utterance-Oriented Dirichlet Process Mixture ModelabstractThis paper provides the analytical solution and algorithm of UO-DPMM based on a non-parametric Bayesian manner, and thus realizes fully Bayesian speaker clustering. We carried out preliminary speaker clustering experiments by using a TIMIT database to compare the proposed method with the conventional Bayesian Information Criterion (BIC) based method, which is an approximate Bayesian approach. The results showed that the proposed method outperformed the conventional one in terms of both computational cost and robustness to changes in tuning parameters. Naohiro Tawara, Shinji Watanabe 0001, Tetsuji Ogawa, Tetsunori Kobayashi |
INTERSPEECH | 4 |
| 2010 | A regularized discriminative training method of acoustic models derived by minimum relative entropy discrimination
Yotaro Kubo, Shinji Watanabe 0001, Atsushi Nakamura, Tetsunori Kobayashi |
INTERSPEECH | 4 |
| 2010 | Psychological evaluation of a group communication activation robot in a party game
Yoichi Matsuyama, Shinya Fujie, Hikaru Taniyama, Tetsunori Kobayashi |
INTERSPEECH | 4 |
| 2009 | System design of group communication activator: an entertainment task for elderly careabstractOur community is facing serious Aging Society especially in Japan. Yoichi Matsuyama, Hikaru Taniyama, Shinya Fujie, Tetsunori Kobayashi |
HRI | 4 |
| 2009 | Conversation robot participating in and activating a group communication
Shinya Fujie, Yoichi Matsuyama, Hikaru Taniyama, Tetsunori Kobayashi |
INTERSPEECH | 4 |
| 2009 | Robot auditory system using head-mounted square microphone arrayabstractA new noise reduction method suitable for autonomous mobile robots was proposed and applied to preprocessing of a hands-free spoken dialogue system. When a robot talks with a conversational partner in real environments, not only speech utterances by the partner but also various types of noise, such as directional noise, diffuse noise, and noise from the robot, are observed at microphones. We attempted to remove these types of noise simultaneously with small and light-weighted devices and low-computational-cost algorithms. We assumed that the conversational partner of the robot was in front of the robot. In this case, the aim of the proposed method is extracting speech signals coming from the frontal direction of the robot. The proposed noise reduction system was evaluated in the presence of various types of noise: the number of word errors was reduced by 69% as compared to the conventional methods. The proposed robot auditory system can also cope with the case in which a conversational partner (i.e., a sound source) moves from the front of the robot: the sound source was localized by face detection and tracking using facial images obtained from a camera mounted on an eye of the robot. As a result, various types of noise could be reduced in real time, irrespective of the sound source positions, by combining speech information with image information. Kosuke Hosoya, Tetsuji Ogawa, Tetsunori Kobayashi |
IROS | 3 |
| 2009 | Upper-Body Contour Extraction Using Face and Body Shape Variance Information
Kazuki Hoshiai, Shinya Fujie, Tetsunori Kobayashi |
PSIVT | 3 |
| 2008 | An ASM fitting method based on machine learning that provides a robust parameter initialization for AAM fittingabstractDue to their use of information contained in texture, active appearance models (AAM) generally outperform active shape models (ASM) in terms of fitting accuracy. Although many extensions and improvements over the original AAM have been proposed, on of the main drawbacks of AAMs remains its dependence on good initial model parameters to achieve accurate fitting results. In this paper, we determine the initial model parameters for AAM fitting with ASM fitting, and use machine learning techniques to improve the scope and accuracy of ASM fitting. Combining the precision of AAM fitting with the large radius of convergence of learned ASM fitting improves the results by an order of magnitude, as our empirical evaluation on a database of publicly available benchmark images demonstrates. Matthias Wimmer, Shinya Fujie, Freek Stulp, Tetsunori Kobayashi, Bernd Radig |
FG | 4 |
| 2008 | Incorporation of phrase intonation to context clustering for average voice models in HMM-based Thai speech synthesisabstractThis paper describes a novel approach to the context clustering process in a speaker independent HMM-based Thai speech synthesis for improvement of the tone intelligibility of the average voice and also the speaker adapted voice. A couple of phrase intonation features from a generative model including a baseline value of fundamental frequency and a phrase command amplitude are extracted and thereafter exploited in the context clustering process of HMM training stage. In the experiments, subjective evaluations of both average voice and adapted voice in terms of the intelligibility of tone are conducted. The results show that the tone correctness of the synthesized speech is significantly improved. Suphattharachai Chomphan, Tetsunori Kobayashi |
ICASSP | 2 |
| 2008 | Speech enhancement using square microphone array for mobile devicesabstractIn this paper, we propose a new type of speech enhancement method that is suitable for mobile devices used in noisy environments. For the sake of achieving high-performance speech recognition and auditory perception in the mobile devices, disturbance noises have to be removed under the requirements of a space-saving microphone arrangement and a low computational cost. The proposed method can reduce both the directional and the diffuse noises under the requirements for the mobile devices by applying the square microphone array and the low-cost processing that consists of multiple null beamforming, their minimum power channel selection and Wiener filtering. The effectiveness of the proposed method is clarified for speech recognition accuracies and speech qualities under the condition in which both the directional and the diffuse noises exist simultaneously: it reduced 40% of recognition errors and improved PESQ-based MOS value by 0.75 point. Shintaro Takada, Tetsuji Ogawa, Kenzo Akagiri, Tetsunori Kobayashi |
ICASSP | 4 |
| 2008 | Design and formulation for speech interface based on flexible shortcuts
Teppei Nakano, Tomoyuki Kumai, Tetsunori Kobayashi, Yasushi Ishikawa |
INTERSPEECH | 3 |
| 2007 | Introduction of the METI project "development of fundamental speech recognition technology"abstractWaseda University, Tokyo Institute of Technology, and six companies, Asahi-kasei, Hitachi, Mitsubishi, NEC, Oki and Toshiba, initiated a three year project in 2006 supported by the Ministry of Economy, Industry and Trade (METI), Japan, for jointly developing fundamental automatic speech recognition (ASR) technology. The project focuses on utilizing ASR technology in car and home environments. Seven subtasks are being investigated: speech/non-speech separation using multiple microphones, speech/non-speech separation for a single audio stream, developing a high-performance WFST-based decoder, multi-lingual ASR modeling, higher-order language modeling, developing a system for assisting speech interface development, and overall technology evaluation. This talk will give an overview of the intermediate technological progress achieved by the project. Sadaoki Furui, Tetsunori Kobayashi |
ASRU | 2 |
| 2007 | Extensible speech recognition system using proxy-agentabstractThis paper presents an extension framework for a speech recognition system. This framework is designed to use “Proxy-Agent,” a software component located between applications, speech recognition engines, and input devices. By taking advantage of its structural characteristics, Proxy-Agent can provide supplementary services for speech recognition systems as well as user extensions. A monitoring capability, a feedback capability, and an extension capability are implemented and presented in this paper. For the first prototype, we developed a data collection application and an application control system using Proxy-Agent. Through these developments, we verified the effectiveness of the data collection capability of Proxy-Agent, and the framework extension capability. Teppei Nakano, Shinya Fujie, Tetsunori Kobayashi |
ASRU | 3 |
| 2007 | Adequacy Analysis of Simulation-Based Assessment of Speech Recognition SystemabstractThe adequacies of the simulation-based assessment of speech recognition systems under noisy conditions are investigated and discussed. To evaluate the speech recognition systems in various environments, it is desirable to collect the test data uttered in the corresponding environments but it is not realistic since enormous works are required. To conduct evaluations of the speech recognition systems properly, it is promising to simulate evaluation experiments in the target environments as described below: comparatively small test data are collected, and test data of the target environment are generated by computing convolution of the impulse response of the target environment with the collected data. However, it is well known that changes of the acoustic characteristics are caused by the Lombard effect, and so it is not necessarily obvious whether the simulation can precisely approximate the experiment in actual environment. This paper clarifies the condition to perform effective simulations of the noisy speech recognition, focusing on the influence of impulse response accuracies and Lombard effects on the speech recognition performance. Tetsuji Ogawa, Satoshi Kanba, Tetsunori Kobayashi |
ICASSP (4) | 3 |
| 2007 | Dynamic integration of multiple feature streams for robust real-time LVCSR
Shoei Sato, Kazuo Onoe, Akio Kobayashi, Shinichi Homma, Toru Imai, Tohru Takagi, Tetsunori Kobayashi |
INTERSPEECH | 7 |
| 2006 | MONEA: Message-oriented Networked-robot ArchitectureabstractThis paper proposes message-oriented networked-robot architecture, or MONEA, as an efficient development platform architecture for multifunctional robots. In order to avoid problems occurred in multifunctional robot developments, we design the architecture to fulfill the following three features. Firstly, it embodies the meta-architecture for networked-robots. Secondly, it supports bazaar-style development model. Finally, it doesn't require heavy weight middleware. To realize them, we developed an information sharing framework named networked-whiteboard model along with message passing framework via P2P virtual network. A development methodology using interest-oriented module groups and software patterns is also presented as a means to reduce complexity risks. A middleware is developed as an implementation of this architecture, and we verify the availability and effectivity of our platform through the development of dialogue robot for exhibition Teppei Nakano, Shinya Fujie, Tetsunori Kobayashi |
ICRA | 3 |
| 2006 | Manifold HLDA and its application to robust speech recognition
Toshiaki Kubo, Tetsuji Ogawa, Tetsunori Kobayashi |
INTERSPEECH | 3 |
| 2005 | Speech recognition in the blind condition based on multiple directivity patterns using a microphone arrayabstractThe proposed system is constructed by the cascade of the sound localization system, MUSIC, and the sound segregation system, SMDP (segregation using multiple directivity patterns) proposed in our previous paper. SMDP is characterized by using redundant directivity patterns. Usually, it is difficult for this sort of cascade system to achieve high performance because the sound localization stage cannot be perfect and errors occurring in this first stage cause serious damage to the segregation stage. Particularly, missing the sound source is critical. By arranging virtual sound sources, we deal with the excess sound sources. In the proposed method, contrarily, the errors in the localization stage hardly cause problems as long as they are insertions. SMDP uses redundant directivity patterns from the beginning, so it tolerates insertion errors. The proposed method achieved 70% word accuracy in a double-talk recognition experiment using a 20 K vocabulary, which is 18% better compared to ICA-based blind source separation, with the source-number-given condition. Toshiyuki Sekiya, Tetsunori Kobayashi |
ICASSP (1) | 2 |
| 2005 | Back-channel feedback generation using linguistic and nonlinguistic information and its application to spoken dialogue system
Shinya Fujie, Kenta Fukushima, Tetsunori Kobayashi |
INTERSPEECH | 3 |
| 2005 | Optimizing the structure of partly-hidden Markov models using weighted likelihood-ratio maximization criterion
Tetsuji Ogawa, Tetsunori Kobayashi |
INTERSPEECH | 2 |
| 2004 | A low-band spectrum envelope modeling for high quality pitch modificationabstractA low-band spectrum envelope reconstruction method was tested to see if it could improve the sound quality of speech modified by the PSOLA (pitch synchronous overlap add) method. In the conventional TD (time domain)-PSOLA method, the spectrum envelope extracted using a Hanning window with a two-pitch-period length had no reliable information in the band of frequencies lower than original F/sub 0/. This problem causes the sound degradation of the F/sub 0/ modified speech. In the proposed method, the low-band spectrum envelope was properly modified according to the F/sub 0/ modification rate. The amplitude of the F/sub 0/ harmonic components in the low-band was reproduced based on the spectral tilt of the spectrum envelope. Subjective listening test results suggest this proposed method yields better sound quality than the conventional TD-PSOLA method when the downward modification rate exceeds 0.4 octave. Ryo Mochizuki, Tetsunori Kobayashi |
ICASSP (1) | 2 |
| 2004 | Speech enhancement based on multiple directivity patterns using a microphone arrayabstractA novel speech segregation method using a microphone array with multiple directivities is proposed and applied to speech recognition under existence of disturbance speech. Conventional microphone array techniques use only single directivity of their own. It is very difficult for this kind of array technique to remove the influence of the disturbance. In our method, redundant simultaneous equations of the amplitudes of sound sources are generated by using these multiple directivities. The solution of these equations gives good estimates of disturbances. The spectral subtraction is applied with these estimates of disturbances, and the perfect enhancement of target speech is performed. The experimental results of double talk recognition with 20 K vocabulary show that the proposed enhancement technique is effective to achieve 45 % error reduction. Toshiyuki Sekiya, Tetsunori Kobayashi |
ICASSP (1) | 2 |
| 2004 | Prosody based attitude recognition with feature selection and its application to spoken dialog system as para-linguistic informationabstractIn this paper, prosody-based attitude recognition and its application to a spoken dialog system are proposed. Paralinguistic information plays a important role in the human communication. We aimed to recognize the user’s attitude by prosody, and apply it to a spoken dialog system as para-linguistic information. In order to find important features to recognize the attitude from automatically extracted features, we applied some feature selection methods. Experimental results show the stepwise method, a combination of the forward selection method and the backward selection method, achieved the best recognition rate. Finally, the dialog system using the recognition results as para-linguistic information is shown. Shinya Fujie, Tetsunori Kobayashi, Daizo Yagi, Hideaki Kikuchi |
INTERSPEECH | 2 |
| 2004 | Speech spotter: on-demand speech recognition in human-human conversation on the telephone or in face-to-face situationsabstractThis paper describes a novel speech-interface function, called “speech spotter”,whichenablesausertoentervoicecommands into a speech recognizer in the midst of natural human-human conversation. In the past, it has been difficult to use automatic speech recognition in human-human conversation since it was not easy to judge, from only microphone input, whether a user was speaking to another person or a speech recognizer. We solve this problem by using two kinds of nonverbal speech information: a filled pause (a vowel-lengthening hesitation like “er...”) and voice pitch. Only when a user utters a voice command with a high pitch just after a filled pause is the voice command accepted by the speech recognizer. By using this speechspotter function, we have built two application systems: an ondemand information system for assisting human-human conversation and a music-playback system for enriching telephone conversation. The results from using these systems have shown thatthespeech-spotter functionisrobustandconvenientenough to be used in face-to-face or cellular-phone conversations. Masataka Goto, Koji Kitayama, Katunobu Itou, Tetsunori Kobayashi |
INTERSPEECH | 4 |
| 2004 | Recognition of three simultaneous utterance of speech by four-line directivity microphone mounted on head of robotabstractA sound source separation method using four-line directivity microphones mounted on a head of a robot is proposed and applied to speech recognition under existence of two disturbances of speech. Sound source separation methods using microphones mounted on robot heads generally used strict head-related transfer functions(HRTF). We propose a robust sound source separation that does not require an estimate of a strict HRTF. Our method takes advantage of a sound pressure difference with the robot head acting as a sound barrier. The enhancement of the difference in the target speech is performed by signal processing of three layers:two-line SAFIA, twoline Spectral Subtraction and their integration. The experimental results of three simultaneous utterance recognition with vocabulary of 20K show that the proposed method is effective in achieving 71% error reduction. Naoya Mochiki, Tetsunori Kobayashi, Toshiyuki Sekiya, Tetsuji Ogawa |
INTERSPEECH | 2 |
| 2003 | Hybrid modeling of PHMM and HMM for speech recognitionabstractA hybrid acoustic model of partly hidden Markov model (PHMM) and HMM is proposed. PHMM was proposed in our previous work to deal with the complicated temporal changes of acoustic features (Ogawa, T. and Kobayashi, T, Proc. ICSLP2002, p.2673-6, 2002). It can realized observation dependent behaviors in both observations and state transitions. It achieved good performance but some errors with different trends from HMM still remained. We have designed a new acoustic model on the basis of PHMM, in which the observation and state transition probabilities are defined by the geometric means of PHMM-based ones and HMM-based ones. In this framework, if a word hypothesis is given a low score by either PHMM or HMM, it almost loses the possibility of being a probable candidate. Since many errors are due to high-scores of incorrect categories rather than low-score of the correct category, this property contributes to reducing errors. Moreover, the proposed model is more stable than PHMM because the higher order statistics of PHMM, which is generally accurate but sometimes less reliable, are smoothed by the lower order statistics of HMM, which is not so accurate, but robust. Experimental results show the effectiveness of the proposed model: it reduces the word errors by 25% compared with HMM. Tetsuji Ogawa, Tetsunori Kobayashi |
ICASSP (1) | 2 |
| 2003 | Speech shift: direct speech-input-mode switching through intentional control of voice pitchabstractThis paper describes a speech-input interface function, called speech shift, that enables a user to specify a speech-input mode by simply changing (shifting) voice pitch. While current speech-input interfaces have used only verbal information, we aimed at building a more user-friendly speech interface by making use of nonverbal information, the voice pitch. By intentionally controlling the pitch, a user can enter the same word with it having different meanings (functions) without explicitly changing the speech-input mode. Our speech-shift function implemented on a voice-enabled word processor, for example, can distinguish an utterance with a high pitch from one with a normal (low) pitch, and regard the former as voice-command-mode input(suchasfile-menuandedit-menucommands)andthelatter as regular dictation-mode text input. Our experimental results from twenty subjects showed that the speech-shift function is effective, easy to use, and a labor-saving input method. Masataka Goto, Yukihiro Omoto, Katunobu Itou, Tetsunori Kobayashi |
INTERSPEECH | 4 |
| 2003 | Speech starter: noise-robust endpoint detection by using filled pausesabstractIn this paper we propose a speech interface function, called speech starter, that enables noise-robust endpoint (utterance) detection for speech recognition. When current speech recognizers are used in a noisy environment, a typical recognition error is caused by incorrect endpoints because their automatic detection is likely to be disturbed by non-stationary noises. The speech starter function enables a user to specify the beginning of each utterance by uttering a filler with a filled pause, which is used as a trigger to start speech-recognition processes. Since filled pauses can be detected robustly in a noisy environment, practical endpoint detection is achieved. Speech starter also offers the advantage of providing a hands-free speech interface and it is user-friendly because a speaker tends to utter filled pauses (e.g., “er...”) at the beginning of utterances when hesitating in human-human communication. Experimental results from a 10-dB-SNR noisy environment show that the recognition error rate with speech starter was lower than with conventional endpoint-detection methods. 1. Koji Kitayama, Masataka Goto, Katunobu Itou, Tetsunori Kobayashi |
INTERSPEECH | 4 |
| 2003 | Speech recognition of double talk using SAFIA-based audio segregation
Toshiyuki Sekiya, Tetsuji Ogawa, Tetsunori Kobayashi |
INTERSPEECH | 3 |
| 2002 | Design and collection of acoustic sound data for hands-free speech recognition and sound scene understandingabstractThe sound data for open evaluation is necessary for studies such as sound source localization, sound retrieval, sound recognition and hands-free speech recognition in real acoustic environments. This paper reports on our project for acoustic data collection. There are many kinds of sound scenes in real environments. The sound scene is specified by sound sources and room acoustics. The number of combinations of the sound sources, source positions and rooms is huge in real acoustic environments. We assumed that the sound in the environments can be simulated by convolution of the isolated sound sources and impulse responses. As an isolated sound source, hundred kinds of environment sounds and speech sounds are collected. The impulse responses are collected in various acoustic environments. Additionally we collected sounds from a moving source. In this paper, progress of our sound scene database collection project and application to environment sound recognition and hands-free speech recognition are described. Satoshi Nakamura 0001, Kazuo Hiyane, Futoshi Asano, Yutaka Kaneda, Takeshi Yamada, Takanobu Nishiura, Tetsunori Kobayashi, Shiro Ise, Hiroshi Saruwatari |
ICME (2) | 7 |
| 2002 | Generalization of state-observation-dependency in partly hidden Markov models
Tetsuji Ogawa, Tetsunori Kobayashi |
INTERSPEECH | 2 |
| 2002 | Inter-module cooperation architecture for interactive robotabstractWe designed an inter-module cooperation architecture that enables the collaborative development of interactive robots. In the bazaar-like development model, each module is developed by an individual developer and the total system is developed by the cooperation of these developers. To realize smooth collaboration of these developers under bazaar-like model, inter-module cooperation architecture is required to avoid the confliction of modules and select the appropriate modules according to the situations and the tasks. For this aim, we introduce a priority based cooperation mechanism for modules. The priority for each module under the various situations is, described in a situated-priority description script (SPDS). Flexible module selection is realized by the modification of the SPDS under the various situations and the tasks. We also evaluate the efficiency of proposed architecture through the development of the multi-modal conversation robot. KyeongJu Kim, Yosuke Matsusaka, Tetsunori Kobayashi |
IROS | 3 |
| 2001 | Estimating positions of multiple adjacent speakers based on MUSIC spectra correlation using a microphone arrayabstractWe propose an improved method of estimating the positions of two speakers using a microphone array. A well-known method, MUSIC, can be used to estimate speaker positions with high precision. However, in the special case that the speakers are closely located, the conventional MUSIC-based method sometimes fails to identify the existence of certain speakers, because close peaks in the MUSIC spectrum cannot be resolved. To overcome this difficulty, we propose a new method utilizing a cross-correlation between space-spectra calculated by MUSIC. Experimental results in a real environment have shown that the proposed method is effective in resolving the approximate positions of adjacent speakers. Hidetomo Tanaka, Tetsunori Kobayashi |
ICASSP | 2 |
| 2001 | Modeling of conversational strategy for the robot participating in the group conversationabstractThis paper describes a strategy for the conversation system to take part in human-to-human group conversation. One big characteristic of the group conversation system is that it can choose whether to observe or to take turn in the conversation. We implement the computational model combined with speech and gaze recognizers to keep the rules in turn taking, and define an interruption decision strategy based on an analysis of human needs. And finally, we realized a humanfriendly group conversation system by combining multimodal information processing/expression abilities of humanoid robot ROBITA. Yosuke Matsusaka, Shinya Fujie, Tetsunori Kobayashi |
INTERSPEECH | 3 |
| 2000 | Dictation of multiparty conversation using statistical turn taking model and speaker modelabstractA new speech decoder dealing with multiparty conversation is proposed. Multiparty conversation denotes a situation in which many speakers talk to each other. Almost of all conventional speech recognition systems assume that the input data consist of single speaker's voice. However, some applications, such as dialogue dictation and voice interfaces for multi-users, have to deal with mixed speakers' voices. In such a situation, the system has to recognize not only the word sequence of the input speech but also the speaker of each part of them. Therefore, we propose a decoder utilizing not only an acoustic model and language model, which are the resources of a conventional single-user speech decoder, but also a statistic turn taking model and speakers models to recognize speech. This framework realizes simultaneous maximum likelihood estimation of spoken word sequence and the speaker sequence. Experimental results using a TV sports news show that the proposed method reduce the word error rate by 7.7% and speaker error rate by 97.8% compared to the conventional method. Noriyuki Murai, Tetsunori Kobayashi |
ICASSP | 2 |
| 2000 | Free software toolkit for Japanese large vocabulary continuous speech recognitionabstractA sharable software repository for Japanese LVCSR (Large Vocabulary Continuous Speech Recognition) is introduced. It is designed as a baseline platform for research and developed by researchers of different academic institutes under a governmental support. The repository consists of a recognition engine (Julius), Japanese acoustic models and statistical language models as well as Japanese morphological analysis tools. These modules can be easily integrated and replaced under a plug-and-play framework, which makes it possible to fairly evaluate components and to develop specific application systems. Assessment of these modules and systems in a 20000-word dictation task is reported. The software repository is freely available to the public. Tatsuya Kawahara, Akinobu Lee, Tetsunori Kobayashi, Kazuya Takeda, Nobuaki Minematsu, Shigeki Sagayama, Katunobu Itou, Akinori Ito, Mikio Yamamoto, Atsushi Yamada, Takehito Utsuro, Kiyohiro Shikano |
INTERSPEECH | 3 |
| 2000 | IPA Japanese Dictation Free Software Project
Katunobu Itou, Kiyohiro Shikano, Tatsuya Kawahara, Kazuya Takeda, Atsushi Yamada, Akinori Ito, Takehito Utsuro, Tetsunori Kobayashi, Nobuaki Minematsu, Mikio Yamamoto, Shigeki Sagayama, Akinobu Lee |
LREC | 8 |
| 2000 | A conversational robot utilizing facial and body expressionsabstractCommunication using non-verbal expression, such as gesture, eyeshot, posture and facial expression, plays an important role in human interaction. For example, in conveying spatial information of the position and/or size of an object, pointing actions also contribute to simplifying the verbal expressions. By nodding or changing facial expressions toward the speaker's utterance, humans convey their state of mind to the speaker and realize a conversable environment. No current conversational systems are equipped with such adequate communicational ability. The authors implement body and facial expressions and conversational function in a humanoid robot and try to realize a natural conversational system. As a result, effective use of the pointing action with verbal expressions "that one" or "this one" contribute to simplifying the verbal expression and realization of rhythmical conversation. In addition, proper feedback actions to the user with facial expressions improve the transparency of the system and the performance of the interface. Tsuyoshi Tojo, Yosuke Matsusaka, Tomotada Ishii, Tetsunori Kobayashi |
SMC | 4 |
| 1999 | Partly hidden Markov model and its application to speech recognitionabstractA new pattern matching method, the partly hidden Markov model, is proposed and applied to speech recognition. The hidden Markov model, which is widely used for speech recognition, can deal with only piecewise stationary stochastic process. We solved this problem by introducing the modified second order Markov model, in which the first state is hidden and the second one is observable. In this model, not only the feature parameter observations but also the state transitions are dependent on the previous feature observation. Therefore, even the complicated transient can be modeled precisely. Some simulation experiments showed the high potential of the proposed model. From the results of the word recognition test is was observed that the error rate was reduced by 39% compared with the normal HMM. Tetsunori Kobayashi, Junko Furuyama, Ken Masumitsu |
ICASSP | 1 |
| 1999 | Class-combined word n-gram for robust language modeling
Norihiko Kobayashi, Tetsunori Kobayashi |
EUROSPEECH | 2 |
| 1999 | Multi-person conversation via multi-modal interface - a robot who communicate with multi-user -
Yosuke Matsusaka, Tsuyoshi Tojo, Sentaro Kubota, Kenji Furukawa, Daisuke Tamiya, Keisuke Hayata, Yuichiro Nakano, Tetsunori Kobayashi |
EUROSPEECH | 8 |
| 1998 | The design of the newspaper-based Japanese large vocabulary continuous speech recognition corpusabstractIn this paper we present the first public Japanese speech corpus for large vocabulary continuous speech recognition (LVCSR) technology, which we have titled JNAS (Japanese Newspaper Article Sentences). We designed it to be comparable to the corpora used in the American and European LVCSR projects. The corpus contains speech recordings (60 hrs.) and their orthographic transcriptions for 306 speakers (153 males and 153 females) reading excerpts from the newspaper's articles and phonetically balanced (PB) sentences. This corpus contains utterances of about 45,000 sentences as a whole with each speaker reading about 150 sentences. JNAS is being distributed on 16 CD-ROMs. Katunobu Itou, Mikio Yamamoto, Kazuya Takeda, Toshiyuki Takezawa, Tatsuo Matsuoka, Tetsunori Kobayashi, Kiyohiro Shikano, Shuichi Itahashi |
ICSLP | 6 |
| 1998 | Sharable software repository for Japanese large vocabulary continuous speech recognitionabstractThe project of Japanese LVCSR (Large Vocabulary Continuous Speech Recognition) platform is introduced. It is a collaboration of researchers of different academic institutes and intended to develop a sharable software repository of not only databases but also models and programs. The platform consists of a standard recognition engine, Japanese phone models and Japanese statistical language models. A set of Japanese phone HMMs are trained with ASJ (Acoustic Society of Japan) databases of 20K sentence utterances per each gender. Japanese word N-gram (2-gram and 3-gram) models are constructed with a corpus of Mainichi newspaper of four years. The recognition engine JULIUS is developed for assessment of both acoustic and language models. The modules are integrated as a Japanese LVCSR system and evaluated on 5000-word dictation task. The software repository is available to the public. Tatsuya Kawahara, Tetsunori Kobayashi, Kazuya Takeda, Nobuaki Minematsu, Katunobu Itou, Mikio Yamamoto, Atsushi Yamada, Takehito Utsuro, Kiyohiro Shikano |
ICSLP | 2 |
| 1998 | Source-extended language model for large vocabulary continuous speech recognitionabstractInformation source extension is utilized to improve the language model for large vocabulary continuous speech recognition (LVCSR). McMillan's theory, source extension make the model entropy close to the real source entropy, implies that the better language model can be obtained by source extension (making new unit through word concatenations and using the new unit for the language modeling). In this paper, we examined the e ectiveness of this source extension. Here, we tested two methods of source extension: frequency-based extension and entropy-based extension. We tested the e ect in terms of perplexity and recognition accuracy using Mainichi newspaper articles and JNAS speech corpus. As the results, the bi-gram perplexity is improved from 98.6 to 70.8 and tri-gram perplexity is improved from 41.9 to 26.4. The bigram-based recognition accuracy is improved from 79.8% to 85.3%. Tetsunori Kobayashi, Yosuke Wada, Norihiko Kobayashi |
ICSLP | 1 |
| 1998 | Controlling gaze of humanoid in communication with humanabstractThis paper describes controlling robot's gaze which has relation to smoothness of turn-taking in communication. We considered the role of gaze in dialogues between human beings and examined it by simulation and our humanoid. Also we analyzed the features of gaze movement in dialogues by plural persons and confirmed that controlling gaze is efficient in confirmation of communication channel by implementing it on the humanoid. Hideaki Kikuchi, Masao Yokoyama, Keiichiro Hoashi, Yasuaki Hidaki, Tetsunori Kobayashi, Katsuhiko Shirai |
IROS | 5 |
| 1997 | Partly-hidden Markov model and its application to gesture recognitionabstractA new pattern matching method, the partly-hidden Markov model, is proposed for gesture recognition. The hidden Markov model, which is widely used for the time series pattern recognition, can deal with only piecewise stationary stochastic process. We solved this problem by introducing the modified second order Markov model, in which the first state is hidden and the second one is observable. As shown by the results of 6 sign-language recognition test, the error rate was improved by 73% compared with normal HMM. Tetsunori Kobayashi, Satoshi Haruyama |
ICASSP | 1 |
| 1996 | ALICE: acquisition of language in conversational environment - an approach to weakly supervised training of spoken language system for language porting
Tetsunori Kobayashi |
ICSLP | 1 |
| 1994 | Markov model based noise modeling and its application to noisy speech recognition using dynamical features of speechabstractIn this paper, some algorithms to recognize speech in time varying noise are proposed. In the proposed methods, spectral subtraction and Markov model based noise models are successfully utilized in the framework of spectral decomposition of noisy speech. Firstly, we considered the problem of the mis-subtraction noise which is caused in the subtraction based decomposition procedure. Then, the precise use of dynamical feature of speech such as delta cepstrum is discussed. Using the methods proposed here, recognition performance are improved more than 60% compared to no compensation method.> Tetsunori Kobayashi, Ryuji Mine, Katsuhiko Shirai |
ICASSP (2) | 1 |
| 1994 | Automatic training of phoneme dictionary based on mutual information criterionabstractProposes an automatic training mechanism for phoneme recognition using labelless speech data under the condition that only its orthographical phonemic symbol sequence is given. For the purpose of obtaining better recognition performance the authors attempt to realize an automatic labeling procedure based on a phoneme classification method by mutual information criterion. By iterative training of a phoneme dictionary for a large amount of speech data, one can investigate the performance and convergence properties of the dictionary. From experimental results, the percent correct of the labeling is over 98% after three iterations, and for the phoneme recognition, a very high accuracy is also obtained.> Shigeki Okawa, Tetsunori Kobayashi, Katsuhiko Shirai |
ICASSP (1) | 2 |
| 1994 | Multimodal drawing tool using speech, mouse and key-board
Takuya Nishirnoto, Nobutoshi Shida, Tetsunori Kobayashi, Katsuhiko Shirai |
ICSLP | 3 |
| 1994 | Phoneme recognition in various styles of utterance based on mutual information criterion
Shigeki Okawa, Tetsunori Kobayashi, Katsuhiko Shirai |
ICSLP | 2 |
| 1994 | Generation of prosody in speech synthesis using large speech data-base
Naohiro Sakurai, Takerni Mochida, Tetsunori Kobayashi, Katsuhiko Shirai |
ICSLP | 3 |
| 1993 | Speech recognition under the unstationary noise based on the noise Markov model and spectral-subtraction
Tetsunori Kobayashi, Ryuji Mine, Katsuhiko Shirai |
EUROSPEECH | 1 |
| 1993 | Word spotting in conversational speech based on phonemic unit likelihood by mutual information criterion
Shigeki Okawa, Tetsunori Kobayashi, Katsuhiko Shirai |
EUROSPEECH | 2 |
| 1992 | Speaker adaptive phoneme recognition based on feature mapping from spectral domain to probabilistic domainabstractA feature parameter space for speech recognition called PRPG (probability ratios between phoneme group pairs) is described, and speaker adaptive phoneme recognition is performed. In the coordinate system proposed, the area with the same information for speech recognition is compressed into one point. The mapping function from spectral coordinate system to the proposed one is realized using a neural network. The code-vectors designed on this coordinate system are guaranteed to be information-theoretically more efficient than that of spectral coordinate system. Moreover, by the definition of the coordinate system, the meaning of axes is equivalent among different speakers, so speaker adaptation can be easily performed without trajectory mapping. Experimental results show that errors are reduced by 40% by coordinate conversion in speaker-dependent tasks. The scores of speaker-adaptive tasks in the proposed feature domain are always superior to those of the speaker-dependent tasks in the spectral domain.> Tetsunori Kobayashi, Y. Uchiyama, J. Osada, Katsuhiko Shirai |
ICASSP | 1 |
| 1992 | Spectral mapping onto probabilistic domain using neural networks and its application to speaker adaptive phoneme recognition
Tetsunori Kobayashi, Katsuhiko Shirai |
ICSLP | 1 |
| 1992 | Phoneme recognition in continuous speech based on mutual information considering phonemic duration and connectivity
Katsuhiko Shirai, Shigeki Okawa, Tetsunori Kobayashi |
ICSLP | 3 |
| 1991 | Application of neural networks to articulatory motion estimationabstractThe authors discuss an application of neural networks (NNs) to the problem of estimating the motion of articulatory organs from speech waves. A four-layer feedforward network was successfully applied to the articulatory parameter estimation problem. The evaluation test was performed using the vowel data in 5200 tokens in the ATR word database. Results show that the difference in estimated articulatory parameter values between the conventional model matching method (MM) and NN is only 0.1, which is about 3% of the value range, on average. For a few data, large differences arise between MM and NN, but this is due to misestimation in MM rather than NN. The percentage of misestimates in NN is less than 50% of that for MM. As for calculation time, NN is 10 times faster than MM. Thus, a high-speed and stable articulatory parameter estimation technique can be realized using neural networks.> Tetsunori Kobayashi, Masayuki Yagyu, Katsuhiko Shirai |
ICASSP | 1 |
| 1991 | Text-to-speech synthesizer using superposition of sinusoidal waves generated by synchronized oscillators
Katsuhiko Shirai, Kazuo Hashimoto, Tetsunori Kobayashi |
EUROSPEECH | 3 |
| 1990 | Statistical properties of fluctuation of pitch intervals and its modeling for natural synthetic speechabstractStatistical properties of the fluctuation of pitch intervals are investigated, and pitch generation models considering fluctuation are discussed. Experimental results of natural speech analysis show that the distribution of pitch fluctuation can be approximated by shifted gamma distribution and that the correlation coefficients of 0th-5th and 30th-60th order show strong positive values. Several pitch generation models dealing with fluctuation are tested with the aim of realizing natural synthetic speech. The results of perceptual experiments recommend the fluctuation model using a 15th-order autoregressive filter excited by a uniform random number. The quality of the synthetic speech using the above fluctuation model is comparable to that of speech with the original fluctuation.> Tetsunori Kobayashi, Hidetoshi Sekine |
ICASSP | 1 |
| 1990 | Dependence of phonemic feature on contextabstractConventional quantification theory and a nonlinear quantification theory have been used to investigate the influence of phonemic context on the variation of vowel spectra. Using the distinctive features of the surrounding phonemes, the one-dimensional distribution of /a/,/i/ and /e/ and the two-dimensional distribution of /u/ and /o/ have been successfully modeled. As a result, some primal factors which affect the vowel spectra are apparent.> Tetsunori Kobayashi, Kazuhiro Watanabe |
ICASSP | 1 |
| 1989 | Contextual factor analysis of vowel distribution
Tetsunori Kobayashi, Toshiyuki Matsuda, Kazuhiro Watanabe |
EUROSPEECH | 1 |
| 1986 | A network model dealing with focus of conversation for speech understanding systemabstractA new network model is proposed by introducing a special node called ε-node. This model makes flexible path weight control of the network representing the acceptable sentences (ASN) possible by giving a score to the ε-node. The path weight control strategy is developed using a rule based system, which has the ability to provide an adequate ASN according to the flow of conversation. Furthermore, a state transition network based system is adopted in order to follow and adapt to the changing topics of the conversation. Thus, a high speed and high reliable conversational speech understanding system is realized. Tetsunori Kobayashi, Katsuhiko Shirai |
ICASSP | 1 |
| 1986 | Estimation of articulatory parameters by table look-up method and its application for speaker independent phoneme recognitionabstractEstimation of articulatory parameters by table look-up method and the phoneme recognition of unspecified speakers are described. Speech recognition using features in articulatory level is effective respecting adaptability to unspecified speakers and simplicity to compensate co-articulation. However, estimation of the parameter requires a lot of time. In the method, first the articulatory domain is quantized into finite centroids obtained by clustering technique and table search is carried out. It enabled shortening of estimation time but still maintain the same degree of preciseness obtained in the earlier method. Further, a recognition method considering compensation of co-articulation is examined. Phoneme recognition rate of 50 city names uttered by 8 speakers, 4 males and 4 females was 90.8%. Katsuhiko Shirai, Tetsunori Kobayashi, J. Yazawa |
ICASSP | 2 |
| 1986 | Estimating articulatory motion from speech wave
Katsuhiko Shirai, Tetsunori Kobayashi |
Speech Commun. | 2 |
| 1984 | Phrase speech recognition of large vocabulary using feature in articulatory domainabstractA phrase unit speech recognition system is discussed, which is applicable for a large vocabulary and is independent of the task. In the case of large vocabulary, it is desirable to express the words in the dictionary by the sequence of phonemes or phoneme-like units. Therefore, the recognition of phonemes in continuous speech is essential to achieve a flexible speech understanding system. In this paper, a technique to recognize phrases based on the phoneme recognition is introduced. The system is composed of the phoneme recognition part and the phrase recognition part. In the phoneme recognition part, the features in the articulatory domain are extracted and applied to compensate coarticulation. In the phrase recognition part, a word sequence corresponding to the phoneme sequence is determined by using two-level DP matching with automaton control, in which words are processed symbolically to attain the acceptable processing speed. Katsuhiko Shirai, Tetsunori Kobayashi |
ICASSP | 2 |
| 1983 | Considerations on articulatory dynamics for continuous speech recognitionabstractIn this paper, a new method is proposed to eliminate coarticulation effect in articulatory domain. Since coarticulation phenomena are due to the physiological and physical processes of speech production, it is effective to consider articulatory dynamics using a suitable model. This method is to estimate motor command based on the modeling of the transfer function from the command of phonemic unit to the articulatory motion. If second order system is adopted as the model, it is shown that parameters of the dynamics satisfy a simple equation for wide variety of data. And an effective algorithm is proposed to get optimal phonemic commands under the assumption of the above dynamics. When this alogrithm is applied to the phoneme recognition in continuous speech, it is found that the command can be estimated successfully and a few percent higher recognition rate can be obtained compared with the result by our previous method. Katsuhiko Shirai, Tetsunori Kobayashi |
ICASSP | 2 |
| 1982 | Recognition of semivowels and consonants in continuous speech using articulatory parametersabstractArticulatory parameters estimated from speech waves were used for the recognition of semivowels and consonants in continuous speech. It has been shown that introduction of the articulatory model in speech recognition is one effective method to solve the difficulties of coarticulation phenomena and speaker differences. In this paper, the recognition of semivowels and consonants is discussed. As for semivowels, it is found that the phase difference between the movement of the tongue and that of the jaw is important to characterize semivowels, and this can be effectively used in the recognition. In the case of consonants, it is possible to find the typical feature of each consonant which corresponds to its place of articulation in the transient parts of the articulatory parameters. A preliminary experiment adopting the DP matching technique in VCV contexts gave fairly hopeful results. And for nasal sounds, it is shown that introduction of the nasal model is useful. The nasal model consists of the nasal cavity and the velum parameter. Katsuhiko Shirai, Tetsunori Kobayashi |
ICASSP | 2 |