VLDB 2026 Research / reviewers in the wild / expert
Tomoya Mizumoto
dblp:67/9254
· DBLP profile ↗
18ranked-venue papers
5as first author
12since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating the Effect of Question Wording Variations on Answer Consistency in Large Language Models
Junya Takayama, Masaya Ohagi, Tomoya Mizumoto, Katsumasa Yoshikawa |
LREC | 3 |
| 2025 | Evaluating Japanese Dialect Robustness Across Speech and Text-based Large Language ModelsabstractDialogue systems based on large language models (LLMs) have advanced significantly in recent years. However, dialectal variation remains a major challenge, particularly for systems that process spoken input. LLM-based speech language models (SLMs), which integrate LLMs with speech processing components, show promise for spoken language tasks, yet their ability to comprehend dialects has not been sufficiently studied. Moreover, it remains unclear how the dialectal understanding of the base LLM affects SLM performance. This study investigates the dialectal robustness of both LLMs and SLMs using Japanese dialects as a test case. We define robustness as the ratio of performance on dialectal versus standard inputs, enabling fair comparisons. Our experiments show that SLM robustness correlates with that of their text-based counterparts. Furthermore, training with dialectal data and fine-tuning the speech encoder each improves robustness in SLMs. Tomoya Mizumoto, Yusuke Fujita, Lianbo Liu, Atsushi Kojima, Yui Sudo |
ASRU | 1 |
| 2025 | Serialized Output Prompting for Large Language Model-based Multi-Talker Speech RecognitionabstractPrompts are crucial for task definition and for improving the performance of large language models (LLM)-based systems. However, existing LLM-based multi-talker (MT) automatic speech recognition (ASR) systems either omit prompts or rely on simple task-definition prompts, with no prior work exploring the design of prompts to enhance performance. In this paper, we propose extracting serialized output prompts (SOP) and explicitly guiding the LLM using structured prompts to improve system performance (SOP-MT-ASR). A Separator and serialized Connectionist Temporal Classification (CTC) layers are inserted after the speech encoder to separate and extract MT content from the mixed speech encoding in a first-speaking-first-out manner. Subsequently, the SOP, which serves as a prompt for LLMs, is obtained by decoding the serialized CTC outputs using greedy search. To train the model effectively, we design a threestage training strategy, consisting of serialized output training (SOT) fine-tuning, serialized speech information extraction, and SOP-based adaptation. Experimental results on the LibriMix dataset show that, although the LLM-based SOT model performs well in the two-talker scenario, it fails to fully leverage LLMs under more complex conditions, such as the three-talker scenario. The proposed SOP approach significantly improved performance under both two- and three-talker conditions. Yusuke Fujita, Tomoya Mizumoto, Lianbo Liu, Atsushi Kojima, Yui Sudo |
ASRU | 3 |
| 2025 | Persona-Consistent Dialogue Generation via Pseudo Preference TuningabstractWe propose a simple yet effective method for enhancing persona consistency in dialogue response generation using Direct Preference Optimization (DPO). In our method, we generate responses from the response generation model using persona information that has been randomly swapped with data from other dialogues, treating these responses as pseudo-negative samples. The reference responses serve as positive samples, allowing us to create pseudo-preference data. Experimental results demonstrate that our model, fine-tuned with DPO on the pseudo preference data, produces more consistent and natural responses compared to models trained using supervised fine-tuning or reinforcement learning approaches based on entailment relations between personas and utterances. Junya Takayama, Masaya Ohagi, Tomoya Mizumoto, Katsumasa Yoshikawa |
COLING | 3 |
| 2025 | AC/DC: LLM-based Audio Comprehension via Dialogue Continuation
Yusuke Fujita, Tomoya Mizumoto, Atsushi Kojima, Lianbo Liu, Yui Sudo |
INTERSPEECH | 2 |
| 2025 | Is Synthetic Data Truly Effective for Training Speech Language Models?
Tomoya Mizumoto, Atsushi Kojima, Yusuke Fujita, Lianbo Liu, Yui Sudo |
INTERSPEECH | 1 |
| 2025 | OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary
Yui Sudo, Yusuke Fujita, Atsushi Kojima, Tomoya Mizumoto, Lianbo Liu |
INTERSPEECH | 4 |
| 2024 | Dialogue Systems Can Generate Appropriate Responses without the Use of Question Marks?- a Study of the Effects of "?" for Spoken Dialogue Systems -abstractWhen individuals engage in spoken discourse, various phenomena can be observed that differ from those that are apparent in text-based conversation. While written communication commonly uses a question mark to denote a query, in spoken discourse, queries are frequently indicated by a rising intonation at the end of a sentence. However, numerous speech recognition engines do not append a question mark to recognized queries, presenting a challenge when creating a spoken dialogue system. Specifically, the absence of a question mark at the end of a sentence can impede the generation of appropriate responses to queries in spoken dialogue systems. Hence, we investigate the impact of question marks on dialogue systems, with the results showing that they have a significant impact. Moreover, we analyze specific examples in an effort to determine which types of utterances have the impact on dialogue systems. Tomoya Mizumoto, Takato Yamazaki, Katsumasa Yoshikawa, Masaya Ohagi, Toshiki Kawamoto |
LREC/COLING | 1 |
| 2024 | Investigation of look-ahead techniques to improve response time in spoken dialogue system
Masaya Ohagi, Tomoya Mizumoto, Katsumasa Yoshikawa |
INTERSPEECH | 2 |
| 2023 | Reducing the Cost: Cross-Prompt Pre-finetuning for Short Answer Scoring
Hiroaki Funayama, Yuya Asazuma, Yuichiroh Matsubayashi, Tomoya Mizumoto, Kentaro Inui |
AIED | 4 |
| 2023 | An Open-Domain Avatar Chatbot by Exploiting a Large Language ModelabstractWith the ambition to create avatars capable of human-level casual conversation, we developed an open-domain avatar chatbot, situated in a virtual reality environment, that employs a large language model (LLM).Introducing the LLM posed several challenges for multimodal integration, such as developing techniques to align diverse outputs and avatar control, as well as addressing the issue of slow generation speed.To address these challenges, we integrated various external modules into our system.Our system is based on the award-winning model from the Dialogue System Live Competition 5. Through this work, we hope to stimulate discussions within the research community about the potential and challenges of multimodal dialogue systems enhanced with LLMs. Takato Yamazaki, Tomoya Mizumoto, Katsumasa Yoshikawa, Masaya Ohagi, Toshiki Kawamoto |
SIGDIAL | 2 |
| 2022 | Balancing Cost and Quality: An Exploration of Human-in-the-Loop Frameworks for Automated Short Answer Scoring
Hiroaki Funayama, Tasuku Sato, Yuichiroh Matsubayashi, Tomoya Mizumoto, Jun Suzuki 0001, Kentaro Inui |
AIED (1) | 4 |
| 2020 | Massive Exploration of Pseudo Data for Grammatical Error CorrectionabstractCollecting a large amount of training data for grammatical error correction (GEC) models has been an ongoing challenge in the field of GEC. Recently, it has become common to use data demanding deep neural models such as an encoder-decoder for GEC; thus, tackling the problem of data collection has become increasingly important. The incorporation of pseudo data in the training of GEC models is one of the main approaches for mitigating the problem of data scarcity. However, a consensus is lacking on experimental configurations, namely, (i) the methods for generating pseudo data, (ii) the seed corpora used as the source of the pseudo data, and (iii) the means of optimizing the model. In this study, these configurations are thoroughly explored through massive amount of experiments, with the aim of providing an improved understanding of pseudo data. Our main experimental finding is that pretraining a model with pseudo data generated by back-translation-based method is the most effective approach. Our findings are supported by the achievement of state-of-the-art performance on multiple benchmark test sets (the CoNLL-2014 test set and the official test set of the BEA-2019 shared task) without requiring any modifications to the model architecture. We also perform an in-depth analysis of our model with respect to the grammatical error type and proficiency level of the text. Finally, we suggest future directions for further improving model performance. Shun Kiyono, Jun Suzuki 0001, Tomoya Mizumoto, Kentaro Inui |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | An Empirical Study of Incorporating Pseudo Data into Grammatical Error CorrectionabstractShun Kiyono, Jun Suzuki, Masato Mita, Tomoya Mizumoto, Kentaro Inui. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Shun Kiyono, Jun Suzuki 0001, Masato Mita, Tomoya Mizumoto, Kentaro Inui |
EMNLP/IJCNLP (1) | 4 |
| 2016 | Discriminative Reranking for Grammatical Error Correction with Statistical Machine TranslationabstractResearch on grammatical error correction has received considerable attention. For dealing with all types of errors, grammatical error correction methods that employ statistical machine translation (SMT) have been proposed in recent years. An SMT system generates candidates with scores for all candidates and selects the sentence with the highest score as the correction result. However, the 1-best result of an SMT system is not always the best result. Thus, we propose a reranking approach for grammatical error correction. The reranking approach is used to re-score N-best results of the SMT and reorder the results. Our experiments show that our reranking system using parts of speech and syntactic features improves performance and achieves state-of-theart quality, with an F0.5 score of 40.0. Tomoya Mizumoto, Yuji Matsumoto 0001 |
HLT-NAACL | 1 |
| 2012 | Joint English Spelling Error Correction and POS Tagging for Language Learners Writing
Keisuke Sakaguchi, Tomoya Mizumoto, Mamoru Komachi, Yuji Matsumoto 0001 |
COLING | 2 |
| 2011 | The chanty bear: a new application for hri researchabstractThis paper presents yet another English-teaching robot, while putting emphasis on the merits which are offered by second language education to human robot interaction (HRI) research. The chanty bear, our prototype robot based on a rhythmic teaching method of English called Jazz Chants is introduced. Kotaro Funakoshi, Tomoya Mizumoto, Ryo Nagata, Mikio Nakano |
HRI | 2 |
| 2011 | Mining Revision Log of Language Learning SNS for Automated Japanese Error Correction of Second Language Learners
Tomoya Mizumoto, Mamoru Komachi, Masaaki Nagata, Yuji Matsumoto 0001 |
IJCNLP | 1 |