VLDB 2026 Research / reviewers in the wild / expert
Atsushi Kojima
dblp:192/3713
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0001-6829-5819ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Evaluating Japanese Dialect Robustness Across Speech and Text-based Large Language ModelsabstractDialogue systems based on large language models (LLMs) have advanced significantly in recent years. However, dialectal variation remains a major challenge, particularly for systems that process spoken input. LLM-based speech language models (SLMs), which integrate LLMs with speech processing components, show promise for spoken language tasks, yet their ability to comprehend dialects has not been sufficiently studied. Moreover, it remains unclear how the dialectal understanding of the base LLM affects SLM performance. This study investigates the dialectal robustness of both LLMs and SLMs using Japanese dialects as a test case. We define robustness as the ratio of performance on dialectal versus standard inputs, enabling fair comparisons. Our experiments show that SLM robustness correlates with that of their text-based counterparts. Furthermore, training with dialectal data and fine-tuning the speech encoder each improves robustness in SLMs. Tomoya Mizumoto, Yusuke Fujita, Lianbo Liu, Atsushi Kojima, Yui Sudo |
ASRU | 5 |
| 2025 | Serialized Output Prompting for Large Language Model-based Multi-Talker Speech RecognitionabstractPrompts are crucial for task definition and for improving the performance of large language models (LLM)-based systems. However, existing LLM-based multi-talker (MT) automatic speech recognition (ASR) systems either omit prompts or rely on simple task-definition prompts, with no prior work exploring the design of prompts to enhance performance. In this paper, we propose extracting serialized output prompts (SOP) and explicitly guiding the LLM using structured prompts to improve system performance (SOP-MT-ASR). A Separator and serialized Connectionist Temporal Classification (CTC) layers are inserted after the speech encoder to separate and extract MT content from the mixed speech encoding in a first-speaking-first-out manner. Subsequently, the SOP, which serves as a prompt for LLMs, is obtained by decoding the serialized CTC outputs using greedy search. To train the model effectively, we design a threestage training strategy, consisting of serialized output training (SOT) fine-tuning, serialized speech information extraction, and SOP-based adaptation. Experimental results on the LibriMix dataset show that, although the LLM-based SOT model performs well in the two-talker scenario, it fails to fully leverage LLMs under more complex conditions, such as the three-talker scenario. The proposed SOP approach significantly improved performance under both two- and three-talker conditions. Yusuke Fujita, Tomoya Mizumoto, Lianbo Liu, Atsushi Kojima, Yui Sudo |
ASRU | 5 |
| 2025 | AC/DC: LLM-based Audio Comprehension via Dialogue Continuation
Yusuke Fujita, Tomoya Mizumoto, Atsushi Kojima, Lianbo Liu, Yui Sudo |
INTERSPEECH | 3 |
| 2025 | Is Synthetic Data Truly Effective for Training Speech Language Models?
Tomoya Mizumoto, Atsushi Kojima, Yusuke Fujita, Lianbo Liu, Yui Sudo |
INTERSPEECH | 2 |
| 2025 | OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary
Yui Sudo, Yusuke Fujita, Atsushi Kojima, Tomoya Mizumoto, Lianbo Liu |
INTERSPEECH | 3 |
| 2024 | Sub-Table Rescorer for Table Question AnsweringabstractWe propose a sub-table rescorer (STR) to improve the performance of an inner table retriever (ITR)-based inference for the table question answering. Tabular language model (TLM) truncates the sequence of a long table due to their input token limits. It leads to accuracy degradation. To solve this problem, ITR extracts sub-table candidates, which correspond to a part of an entire greater original table on the basis of relevance scores to the question for each of the columns and rows. Then, the topN longest sub-tables are selected. Our proposed STR estimates the relevance score between a question and each sub-table. In this work, we explored two different methods to integrate STR to the ITR-based inference. In the first method, STR rescores sub-table candidates, and the topN sub-tables are chosen. Then, TLM outputs the most confident answer. In the second method, the score calculated by STR is interpolated with the score calculated by TLM. Then, the most confident answer is chosen. In the experiment, we evaluate the performance on the WikiTableQuestions dataset. By applying STR to the ITR-based inference, we observed 4.4% and 6.3% relative reductions in error rate in the rescoring- and score-fusion-based methods, respectively. Atsushi Kojima |
LREC/COLING | 1 |
| 2021 | Knowledge Distillation for Streaming Transformer-Transducer
Atsushi Kojima |
Interspeech | 1 |