VLDB 2026 Research / reviewers in the wild / expert
Ryo Fukuda
dblp:51/7438
· DBLP profile ↗
9ranked-venue papers
5as first author
7since 2021 · last 2025
0009-0005-6213-3241ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Predictive ASR and Turn-taking Prediction at Once: Towards More Responsive Spoken Dialog SystemabstractSpoken dialog systems usually wait for users to finish speaking before generating responses, resulting in response delays. A possible solution for reducing the response delay is to predict future words and/or turn-ends while the user is speaking. To realize this, we propose a method to jointly perform predictive automatic speech recognition and turn-taking prediction. Our model receives partial utterances as input and performs speech recognition, future word prediction, and turntaking prediction via autoregressive decoding. It enables turntaking prediction based on prosodic and linguistic cues of observed partial utterances and predicted future linguistic cues. We also incorporate dialogue contexts to improve the performance. Experiments on the Switchboard corpus showed that our multi-task model outperforms a single-task model in turn-taking prediction. We found that conditioning turn-taking prediction on predicted words improved performance when words were correctly predicted. Ryo Fukuda, Takatomo Kano, Naohiro Tawara, Marc Delcroix, Atsunori Ogawa, Yuya Chiba, Atsushi Ando |
ASRU | 1 |
| 2025 | Speech Emotion Recognition Based on Large-Scale Automatic Speech RecognizerabstractThis paper proposes a novel speech emotion recognition (SER) method that fully leverages the architecture of Whisper, a large-scale automatic speech recognition (ASR) model. The conventional SER models using a pre-trained speech encoder may fail to capture linguistic content since their decoders are too simple. Our proposed method addresses this shortcoming by adopting the decoder of Whisper, which has been discarded in conventional SER, to leverage its language modeling capability. The proposed method introduces special tokens corresponding to the target emotions and then fine-tunes the entire Whisper model. Furthermore, we also propose a new training scheme suitable for Whisper, named serialized multi-task learning (SerialMTL), to consider various speech information as context for the objective SER task. In SerialMTL, the model initially predicts subtask tokens, such as transcription and gender tokens, and then estimates the emotion token. An advantage of the proposed method is the simplicity of the model structure, even when adding any new subtasks. Experimental results show that our model, based on the entire Whisper, achieves better SER performance than the conventional model and further improves with SerialMTL training via ASR and gender recognition subtasks. Ryo Fukuda, Takatomo Kano, Atsushi Ando, Atsunori Ogawa |
ICASSP | 1 |
| 2025 | Bridging Speech and Text Foundation Models with ReShape AttentionabstractThis paper investigates cascade approaches bridging speech and text foundation models (FMs) for speech translation (ST). We address the limitations of cascade systems which suffer from the propagation of speech recognition errors and the lack of access to acoustic information. We propose a ReShape Attention (RSA) that bridges speech embeddings of Whisper, a speech FM, to LLaMA2, a text FM. Speech and text embeddings have temporal and dimensional gaps, which make merging them challenging. RSA reshapes the speech and text embeddings into a sequence of subvectors sharing the same feature dimension. RSA performs cross-attention in the LLaMA2 layers between these two sequences, which allows combining the two embeddings. The RSA allows text FM to directly access speech FM embeddings and optimize the entire ST system for input speech. RSA improves 8.5% relative BLEU score compared to the baseline ST system, which cascades Whisper and LLaMA2. Moreover, our analyses show that the proposed method could even improve performance with ground-truth transcriptions, which suggests that our bridging approach is not limited to mitigating the effect of recognition errors but can also exploit the benefit of acoustic information. Takatomo Kano, Atsunori Ogawa, Marc Delcroix, Ryo Fukuda, Kohei Matsuura, Takanori Ashihara, Shinji Watanabe 0001 |
ICASSP | 5 |
| 2025 | Pick and Summarize: Integrating Extractive and Abstractive Speech Summarization
Takatomo Kano, Atsunori Ogawa, Marc Delcroix, Ryo Fukuda, Shinji Watanabe 0001 |
INTERSPEECH | 4 |
| 2024 | NAIST-SIC-Aligned: An Aligned English-Japanese Simultaneous Interpretation CorpusabstractIt remains a question that how simultaneous interpretation (SI) data affects simultaneous machine translation (SiMT). Research has been limited due to the lack of a large-scale training corpus. In this work, we aim to fill in the gap by introducing NAIST-SIC-Aligned, which is an automatically-aligned parallel English-Japanese SI dataset. Starting with a non-aligned corpus NAIST-SIC, we propose a two-stage alignment approach to make the corpus parallel and thus suitable for model training. The first stage is coarse alignment where we perform a many-to-many mapping between source and target sentences, and the second stage is fine-grained alignment where we perform intra- and inter-sentence filtering to improve the quality of aligned pairs. To ensure the quality of the corpus, each step has been validated either quantitatively or qualitatively. This is the first open-sourced large-scale parallel SI dataset in the literature. We also manually curated a small test set for evaluation purposes. Our results show that models trained with SI data lead to significant improvement in translation quality and latency over baselines. We hope our work advances research on SI corpora construction and SiMT. Our data will be released upon the paper’s acceptance. Jinming Zhao, Katsuhito Sudoh, Satoshi Nakamura 0001, Yuka Ko, Kosuke Doi, Ryo Fukuda |
LREC/COLING | 6 |
| 2024 | Improving Speech Translation Accuracy and Time Efficiency With Fine-Tuned wav2vec 2.0-Based Speech SegmentationabstractSpeech translation (ST) automatically converts utterances in a source language into text in another language. Splitting continuous speech into shorter segments, known as speech segmentation, plays an important role in ST. Recent segmentation methods trained to mimic the segmentation of ST corpora have surpassed traditional approaches. Tsiamas et al. [1] proposed a segmentation frame classifier (SFC) based on a pre-trained speech encoder called wav2vec 2.0. Their method, named SHAS, retains 95-98% of the BLEU score for ST corpus segmentation. However, the segments generated by SHAS are very different from ST corpus segmentation and tend to be longer with multiple combined utterances. This is due to SHAS's reliance on length heuristics, i.e., it splits speech into segments of easily translatable length without fully considering the potential for ST improvement by splitting them into even shorter segments. Longer segments often degrade translation quality and ST's time efficiency. In this study, we extended SHAS to improve ST translation accuracy and efficiency by splitting speech into shorter segments that correspond to sentences. We introduced a simple segmentation avlgorithm using the moving average of SFC predictions without relying on length heuristics and explored wav2vec 2.0 fine-tuning for improved speech segmentation prediction. Our experimental results reveal that our speech segmentation method significantly improved the quality and the time efficiency of speech translation compared to SHAS. Ryo Fukuda, Katsuhito Sudoh, Satoshi Nakamura 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | Speech Segmentation Optimization using Segmented Bilingual Speech Corpus for End-to-end Speech TranslationabstractSpeech segmentation, which splits long speech into short segments, is essential for speech translation (ST).Popular VAD tools like WebRTC VAD 1 have generally relied on pause-based segmentation.Unfortunately, pauses in speech do not necessarily match sentence boundaries, and sentences can be connected by a very short pause that is difficult to detect by VAD.In this study, we propose a speech segmentation method using a binary classification model trained using a segmented bilingual speech corpus.We also propose a hybrid method that combines VAD and the above speech segmentation method.Experimental results reveal that the proposed method is more suitable for cascade and end-to-end ST systems than conventional segmentation methods.The hybrid approach further improves the translation performance. Ryo Fukuda, Katsuhito Sudoh, Satoshi Nakamura 0001 |
INTERSPEECH | 1 |
| 2011 | Practice and Evaluation with Planetary Simulator in Junior High School Science Classes
Ryo Fukuda, Mariko Suzuki, Kazuhiko Sawada, Masato Soga |
ICCE | 1 |
| 2009 | Log-aesthetic space curve segmentsabstractFor designing aesthetic surfaces, such as the car bodies, it is very important to use aesthetic curves as characteristic lines. In such curves, the curvature should be monotonically varying, since it dominates the distortion of reflected images on curved surfaces. In this paper, we present an interactive control method of log-aesthetic space curves. We define log-aesthetic space curves to be curves whose logarithmic curvature and torsion graphs are both linear. The linearity of these graphs constrains that the curvature and torsion are monotonically varying. We clarify the characteristics of log-aesthetic space curves and identify their family. Moreover, we present a novel method for drawing a log-aesthetic space curve segment by specifying two endpoints, their tangents, the slopes, α and β, of straight lines of the logarithmic curvature and torsion graphs, and the torsion parameter Ω. Our implementation shows that log-aesthetic curve segments can be controlled fully interactively. Norimasa Yoshida, Ryo Fukuda, Takafumi Saito |
Symposium on Solid and Physical Modeling | 2 |