VLDB 2026 Research / reviewers in the wild / expert
Kengo Ohta
dblp:82/8162
· DBLP profile ↗
11ranked-venue papers
6as first author
5since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Backchannel prediction for natural spoken dialog systems using general speaker and listener information
Yoshinori Fukunaga, Ryota Nishimura, Kengo Ohta, Norihide Kitaoka |
INTERSPEECH | 3 |
| 2025 | Domain adaptation using non-parallel target domain corpus for self-supervised learning-based automatic speech recognitionabstractThe recognition accuracy of conventional automatic speech recognition (ASR) systems depends heavily on the amount of speech and associated transcription data available in the target domain for model training. However, preparing parallel speech and text data each time a model is trained for a new domain is costly and time-consuming. To solve this problem, we propose a method of domain adaptation that does not require the use of a large amount of parallel target domain training data, as most of the data used for model training is not from the target domain. Instead, only target domain speech is used for model training, along with non-target domain speech and its parallel text data, i.e., the domains and contents of the two types of training data do not correspond to one another. Collecting this type of training data is relatively inexpensive. Domain adaptation is performed in two steps: (1) A pre-trained wav2vec 2.0 model is further pre-trained using a large amount of target domain speech data and is then fine-tuned using a large amount of non-target domain speech and its transcriptions. (2) The density ratio approach (DRA) is applied during inference to a language model (LM) trained using target domain text unrelated to, and independently from, the wav2vec 2.0 training. Experimental evaluation illustrated that the proposed domain adaptation obtained character error rate (CER) 10.4 pts lower than baseline with wav2vec 2.0 and 3.9 pts with XLS-R under the situation that the parallel target domain data is unavailable against the target domain test set, achieving 34.4% and 16.2% reductions in relative CER. Takahiro Kinouchi, Atsunori Ogawa, Yukoh Wakabayashi, Kengo Ohta, Norihide Kitaoka |
Speech Commun. | 4 |
| 2024 | Boosting CTC-based ASR using inter-layer attention-based CTC lossabstractThis paper addresses improving the performance of CTC-based models, which leverage the intermediate outputs of all encoder layers with an attention mechanism.Several previous studies have used the intermediate outputs of the encoder layer to modify CTC-based models.Here, we focus on the role of the Transformer encoder layer, and each encoder layer is computed for two CTC losses by weighting the intermediate outputs of its lower and upper layers using an attention mechanism.By dividing the layer into two groups, it is expected to be possible to calculate the loss, taking into account both acoustic and linguistic features.Experimental results showed that the proposed method improved the baseline recognition performance of TEDLIUM2 speech data, achieving a WER of 9.9% on the dev set and 11.8% on the test set.Our method outperformed the conventional methods for WER with only slightly increased inference speed measured by RTF. Keigo Hojo, Yukoh Wakabayashi, Kengo Ohta, Atsunori Ogawa, Norihide Kitaoka |
INTERSPEECH | 3 |
| 2022 | End-to-End Spontaneous Speech Recognition Using Disfluency Labeling
Koharu Horii, Meiko Fukuda, Kengo Ohta, Ryota Nishimura, Atsunori Ogawa, Norihide Kitaoka |
INTERSPEECH | 3 |
| 2021 | Response type selection for chat-like spoken dialog systems based on LSTM and multi-task learningabstractWe propose a method of automatically selecting appropriate responses in conversational spoken dialog systems by explicitly determining the correct response type that is needed first, based on a comparison of the user’s input utterance with many other utterances. Response utterances are then generated based on this response type designation (back channel, changing the topic, expanding the topic, etc.). This allows the generation of more appropriate responses than conventional end-to-end approaches, which only use the user’s input to directly generate response utterances. As a response type selector, we propose an LSTM-based encoder–decoder framework utilizing acoustic and linguistic features extracted from input utterances. In order to extract these features more accurately, we utilize not only input utterances but also response utterances in the training corpus. To do so, multi-task learning using multiple decoders is also investigated. To evaluate our proposed method, we conducted experiments using a corpus of dialogs between elderly people and an interviewer. Our proposed method outperformed conventional methods using either a point-wise classifier based on Support Vector Machines, or a single-task learning LSTM. The best performance was achieved when our two response type selectors (one trained using acoustic features, and the other trained using linguistic features) were combined, and multi-task learning was also performed. Kengo Ohta, Ryota Nishimura, Norihide Kitaoka |
Speech Commun. | 1 |
| 2012 | Developing Partially-Transcribed Speech Corpus from Edited Transcriptions
Kengo Ohta, Masatoshi Tsuchiya, Seiichi Nakagawa |
LREC | 1 |
| 2011 | Detection of precisely transcribed parts from inexact transcribed corpusabstractAlthough large-scale spontaneous speech corpora are crucial resource for various domains of spoken language processing, they are usually limited due to their construction cost especially in transcribing precisely. On the other hand, inexact transcribed corpora like shorthand notes, meeting records and closed captions are widely available. Unfortunately, it is difficult to use them directly as speech corpora for learning acoustic models, because they contain two kinds of text, precisely transcribed parts and edited parts. In order to resolve this problem, this paper proposes an automatic detection method of precisely transcribed parts from inexact transcribed corpora. Our method consists of two steps: the first step is an automatic alignment between the inexact transcription and its corresponding utterance, and the second step is a support vector machine based detector of precisely transcribed parts using several features obtained by the first step. Experiments using the Japanese National Diet Record shows that automatic detection of precise parts is effective for lightly supervised speaker adaptation, and shows that it achieves reasonable performance to reduce the converting cost from inexact transcribed corpora into precisely transcribed ones. Kengo Ohta, Masatoshi Tsuchiya, Seiichi Nakagawa |
ASRU | 1 |
| 2009 | Effective use of pause information in language modelling for speech recognitionabstractThis paper addresses mismatch between speech processing units used by a speech recognizer and sentences of corpora. A standard speech recognizer divides an input speech into speech processing units based on its power information. On the other hand, training corpora of language models are divided into sentences based on punctuations. There is inevitable mismatch between speech processing units and sentences, and both of them are not optimal for a spontaneous speech recognition task. This paper presents two sub issues to address this problem. At first, the words of the preceding units are utilized to predict the words of the succeeding units, in order to address the mismatch between speech processing units and optimal units. Secondly, we propose a method to build a language model including short pause from a corpus with no short pause to address the mismatch between speech processing units and sentences. Their combination achieved a 4.5% relative improvement over the conventional method in the meeting speech recognition task. Index Terms: speech recognition, language model, pause information, processing unit for speech recognition Kengo Ohta, Masatoshi Tsuchiya, Seiichi Nakagawa |
INTERSPEECH | 1 |
| 2008 | Evaluating spoken language model based on filler prediction model in speech recognitionabstractWe propose a method that uses a filler prediction model for building a language model that includes fillers from a cor-pus without fillers. In our method, a filler prediction model is trained from a corpus that does not cover domain-relevant topics. It recovers fillers in inexact transcribed corpora in the target domain, and then a language model that includes fillers is built from the corpora. The results of an evaluation of the Japanese National Diet Record showed that a model using our method achieves higher recognition performance than conven-tional ones. Kengo Ohta, Masatoshi Tsuchiya, Seiichi Nakagawa |
INTERSPEECH | 1 |
| 2008 | Developing Corpus of Japanese Classroom Lecture Speech Contents
Masatoshi Tsuchiya, Satoru Kogure, Hiromitsu Nishizaki, Kengo Ohta, Seiichi Nakagawa |
LREC | 4 |
| 2007 | Construction of spoken language model including fillers using filler prediction modelabstractThis paper proposes a novel method to construct a spoken language model including fillers from a corpus including no fillers using a filler prediction model. It consists of two submodels: a filler insertion model which predicts places where fillers should be inserted, and a filler selection model which predicts appropriate fillers for given places. It converts a corpus that covers domain-relevant topics but includes no fillers into a corpus that contains fillers as well as domain-relevant topics. The experiment against the corpus of spontaneous Japanese shows that language models constructed by the proposed method achieve quite near performance of the traditional trigram language model constructed from the real spontaneous corpus including fillers. Kengo Ohta, Masatoshi Tsuchiya, Seiichi Nakagawa |
INTERSPEECH | 1 |