VLDB 2026 Research / reviewers in the wild / expert
Takatomo Kano
dblp:140/2697
· DBLP profile ↗
20ranked-venue papers
12as first author
15since 2021 · last 2026
0000-0001-9693-3785ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 10 first-author · 14 since 2021Artificial intelligence and machine learning · 13 · 8 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Microphone array geometry-independent multi-talker distant ASR: NTT system for DASR task of the CHiME-8 challenge
Naoyuki Kamo, Naohiro Tawara, Atsushi Ando, Takatomo Kano, Hiroshi Sato 0002, Rintaro Ikeshita, Takafumi Moriya, Shota Horiguchi, Kohei Matsuura, Atsunori Ogawa, Alexis Plaquet, Takanori Ashihara, Tsubasa Ochiai, Masato Mimura, Marc Delcroix, Tomohiro Nakatani, Taichi Asami, Shoko Araki |
Comput. Speech Lang. | 4 |
| 2025 | Predictive ASR and Turn-taking Prediction at Once: Towards More Responsive Spoken Dialog SystemabstractSpoken dialog systems usually wait for users to finish speaking before generating responses, resulting in response delays. A possible solution for reducing the response delay is to predict future words and/or turn-ends while the user is speaking. To realize this, we propose a method to jointly perform predictive automatic speech recognition and turn-taking prediction. Our model receives partial utterances as input and performs speech recognition, future word prediction, and turntaking prediction via autoregressive decoding. It enables turntaking prediction based on prosodic and linguistic cues of observed partial utterances and predicted future linguistic cues. We also incorporate dialogue contexts to improve the performance. Experiments on the Switchboard corpus showed that our multi-task model outperforms a single-task model in turn-taking prediction. We found that conditioning turn-taking prediction on predicted words improved performance when words were correctly predicted. Ryo Fukuda, Takatomo Kano, Naohiro Tawara, Marc Delcroix, Atsunori Ogawa, Yuya Chiba, Atsushi Ando |
ASRU | 2 |
| 2025 | Speech Emotion Recognition Based on Large-Scale Automatic Speech RecognizerabstractThis paper proposes a novel speech emotion recognition (SER) method that fully leverages the architecture of Whisper, a large-scale automatic speech recognition (ASR) model. The conventional SER models using a pre-trained speech encoder may fail to capture linguistic content since their decoders are too simple. Our proposed method addresses this shortcoming by adopting the decoder of Whisper, which has been discarded in conventional SER, to leverage its language modeling capability. The proposed method introduces special tokens corresponding to the target emotions and then fine-tunes the entire Whisper model. Furthermore, we also propose a new training scheme suitable for Whisper, named serialized multi-task learning (SerialMTL), to consider various speech information as context for the objective SER task. In SerialMTL, the model initially predicts subtask tokens, such as transcription and gender tokens, and then estimates the emotion token. An advantage of the proposed method is the simplicity of the model structure, even when adding any new subtasks. Experimental results show that our model, based on the entire Whisper, achieves better SER performance than the conventional model and further improves with SerialMTL training via ASR and gender recognition subtasks. Ryo Fukuda, Takatomo Kano, Atsushi Ando, Atsunori Ogawa |
ICASSP | 2 |
| 2025 | Bridging Speech and Text Foundation Models with ReShape AttentionabstractThis paper investigates cascade approaches bridging speech and text foundation models (FMs) for speech translation (ST). We address the limitations of cascade systems which suffer from the propagation of speech recognition errors and the lack of access to acoustic information. We propose a ReShape Attention (RSA) that bridges speech embeddings of Whisper, a speech FM, to LLaMA2, a text FM. Speech and text embeddings have temporal and dimensional gaps, which make merging them challenging. RSA reshapes the speech and text embeddings into a sequence of subvectors sharing the same feature dimension. RSA performs cross-attention in the LLaMA2 layers between these two sequences, which allows combining the two embeddings. The RSA allows text FM to directly access speech FM embeddings and optimize the entire ST system for input speech. RSA improves 8.5% relative BLEU score compared to the baseline ST system, which cascades Whisper and LLaMA2. Moreover, our analyses show that the proposed method could even improve performance with ground-truth transcriptions, which suggests that our bridging approach is not limited to mitigating the effect of recognition errors but can also exploit the benefit of acoustic information. Takatomo Kano, Atsunori Ogawa, Marc Delcroix, Ryo Fukuda, Kohei Matsuura, Takanori Ashihara, Shinji Watanabe 0001 |
ICASSP | 1 |
| 2025 | Pick and Summarize: Integrating Extractive and Abstractive Speech Summarization
Takatomo Kano, Atsunori Ogawa, Marc Delcroix, Ryo Fukuda, Shinji Watanabe 0001 |
INTERSPEECH | 1 |
| 2024 | Train Long and Test Long: Leveraging Full Document Contexts in Speech ProcessingabstractThe quadratic memory complexity of self-attention has generally restricted Transformer-based models to utterance-based speech processing, preventing models from leveraging long-form contexts. A common solution has been to formulate long-form speech processing into a streaming problem, only using limited prior context. We propose a new and simple paradigm, encoding entire documents at once, which has been unexplored in Automatic Speech Recognition (ASR) and Speech Translation (ST) due to its technical infeasibility. We exploit developments in efficient attention mechanisms, such as Flash Attention, and show that Transformer-based models can be easily adapted to document-level processing. We experiment with methods to address the quadratic complexity of attention by replacing it with simpler alternatives. As such, our models can handle up to 30 minutes of speech during both training and testing. We evaluate our models on ASR, ST, and Speech Summarization (SSUM) using How2, TEDLIUM3, and SLUE-TED. With document-level context, our ASR models achieve 33.3% and 6.5% relative improvements in WER on How2 and TEDLIUM3 over prior work. Finally, we use our findings to propose a new attention-free self-supervised model, LongHuBERT, capable of handling long inputs. In doing so, we achieve state-of-the-art performance on SLUE-TED SSUM, outperforming cascaded systems that have dominated the benchmark. Takatomo Kano, Atsunori Ogawa, Marc Delcroix, Shinji Watanabe 0001 |
ICASSP | 2 |
| 2024 | Sentence-wise Speech Summarization: Task, Datasets, and End-to-End Modeling with LM Knowledge Distillation
Kohei Matsuura, Takanori Ashihara, Takafumi Moriya, Masato Mimura, Takatomo Kano, Atsunori Ogawa, Marc Delcroix |
INTERSPEECH | 5 |
| 2023 | Summarize While Translating: Universal Model With Parallel Decoding for Summarization and TranslationabstractRecently, multi-decoder and universal models have attracted increased interest in speech and language processing as they allow learning common representations across tasks. These models learn a common representation by sharing a part of or all network parameters. Moreover, such a universal model can handle tasks unseen during training (zero-shot tasks). However, these models do not fully exploit inter-dependencies between tasks during decoding since they usually perform decoding for each task independently. In this paper, we propose to address this issue by extending the universal model to perform multi-task parallel decoding with a cross-attention module between decoders to capture task inter-dependencies explicitly. We also introduce a novel multi-stream beam search algorithm to allow such parallel decoding. We test our proposed model on multi-lingual (English and Portuguese) text/speech translation and summarization, confirming its potential, especially in zero-shot tasks. Takatomo Kano, Atsunori Ogawa, Marc Delcroix, Kohei Matsuura, Takanori Ashihara, Shinji Watanabe 0001 |
ASRU | 1 |
| 2023 | Espnet-Summ: Introducing a Novel Large Dataset, Toolkit, and a Cross-Corpora Evaluation of Speech Summarization SystemsabstractSpeech summarization has garnered significant interest and progressed rapidly over the past few years. In particular, end-to-end models have recently emerged as a competitive alternative to cascade systems for abstractive video summarization. This paper aims to establish progress in this rapidly evolving research field, by introducing ESPNet-SUMM, a new open-source toolkit that facilitates a comprehensive comparison of end-to-end and cascade speech summarization models on 4 different speech summarization tasks spanning diverse applications. Experiments demonstrate that end-to-end models perform better for larger corpora with shorter inputs. This work also introduces Interview, the largest public open-domain multiparty interview corpus with $4400 \mathrm{~h}$ of conversations between radio hosts and guests. Finally, this work explores the use of multiple datasets to improve end-to-end summarization, and experiments demonstrate the benefit of multi-style training over fine-tuning. 1 Roshan S. Sharma, Takatomo Kano, Ruchira Sharma, Siddhant Arora, Shinji Watanabe 0001, Atsunori Ogawa, Marc Delcroix, Rita Singh, Bhiksha Raj |
ASRU | 3 |
| 2023 | Speech Summarization of Long Spoken Document: Improving Memory Efficiency of Speech/Text EncodersabstractSpeech summarization requires processing several minute-long speech sequences to allow exploiting the whole context of a spoken document. A conventional approach is a cascade of automatic speech recognition (ASR) and text summarization (TS). However, the cascade systems are sensitive to ASR errors. Moreover, the cascade system cannot be optimized for input speech and utilize para-linguistic information. Recently, there has been an increased interest in end-to-end (E2E) approaches optimized to output summaries directly from speech. Such systems can thus mitigate the ASR errors of cascade approaches. However, E2E speech summarization requires massive computational resources because it needs to encode long speech sequences. We propose a speech summarization system that enables E2E summarization from 100 seconds, which is the limit of the conventional method, to up to 10 minutes (i.e., the duration of typical instructional videos on YouTube). However, the modeling capability of this model for minute-long speech sequences is weaker than the conventional approach. We thus exploit auxiliary text information from ASR transcriptions to improve the modeling capabilities. The resultant system consists of a dual speech/text encoder decoder-based summarization system. We perform experiments on the How2 dataset showing the proposed system improved METEOR scores by up to 2.7 points by fully exploiting the long spoken documents. Takatomo Kano, Atsunori Ogawa, Marc Delcroix, Roshan S. Sharma, Kohei Matsuura, Shinji Watanabe 0001 |
ICASSP | 1 |
| 2023 | Transfer Learning from Pre-trained Language Models Improves End-to-End Speech Summarization
Kohei Matsuura, Takanori Ashihara, Takafumi Moriya, Tomohiro Tanaka, Takatomo Kano, Atsunori Ogawa, Marc Delcroix |
INTERSPEECH | 5 |
| 2022 | Integrating Multiple ASR Systems into NLP Backend with Attention FusionabstractSpoken language processing (SLP) systems such as speech summarization and translation can be achieved by cascade models. It combines an automatic speech recognition (ASR) frontend and a natural language processing (NLP) backend including machine translation (MT) or text summarization (TS). With this cascade approach, we can exploit large non-paired datasets to independently train state-of-the-art models for each module. However, ASR errors directly affect the performance of the NLP backend in the cascade approach. In this paper, we reduce the impact of ASR errors on the NLP back-end by combining transcriptions from various ASR systems. Recognizer output voting error reduction (ROVER) is a widely used technique for system combination. Although ROVER improves ASR performance, the combination process is not optimized for backend tasks. We propose a system combination that resembles ROVER using attention fusion to achieve the alignment and the combination of multiple ASR hypotheses. This allows the combination process to be optimized for the backend NLP task without changing the ASR frontend. Our proposed technique is general and can be applied to various SLP tasks. We confirm its effectiveness on both speech summarization and translation experiments. Takatomo Kano, Atsunori Ogawa, Marc Delcroix, Shinji Watanabe 0001 |
ICASSP | 1 |
| 2021 | Attention-Based Multi-Hypothesis Fusion for Speech SummarizationabstractSpeech summarization, which generates a text summary from speech, can be achieved by combining automatic speech recognition (ASR) and text summarization (TS). With this cascade approach, we can exploit state-of-the-art models and large training datasets for both subtasks, i.e., Transformer for ASR and Bidirectional Encoder Representations from Transformers (BERT) for TS. However, ASR errors directly affect the quality of the output summary in the cascade approach. We propose a cascade speech summarization model that is robust to ASR errors and that exploits multiple hypotheses generated by ASR to attenuate the effect of ASR errors on the summary. We investigate several schemes to combine ASR hypotheses. First, we propose using the sum of sub-word embedding vectors weighted by their posterior values provided by an ASR system as an input to a BERT-based TS system. Then, we introduce a more general scheme that uses an attention-based fusion module added to a pre-trained BERT module to align and combine several ASR hypotheses. Finally, we perform speech summarization experiments on the How2 dataset and a newly assembled TED-based dataset that we will release with this paper11https://github.com/nttcslab-sp-admin/TEDSummary. These experiments show that retraining the BERT-based TS system with these schemes can improve summarization performance and that the attention-based fusion module is particularly effective. Takatomo Kano, Atsunori Ogawa, Marc Delcroix, Shinji Watanabe 0001 |
ASRU | 1 |
| 2021 | BLSTM-Based Confidence Estimation for End-to-End Speech RecognitionabstractConfidence estimation, in which we estimate the reliability of each recognized token (e.g., word, sub-word, and character) in automatic speech recognition (ASR) hypotheses and detect incorrectly recognized tokens, is an important function for developing ASR applications. In this study, we perform confidence estimation for end-to-end (E2E) ASR hypotheses. Recent E2E ASR systems show high performance (e.g., around 5% token error rates) for various ASR tasks. In such situations, confidence estimation becomes difficult since we need to detect infrequent incorrect tokens from mostly correct token sequences. To tackle this imbalanced dataset problem, we employ a bidirectional long short-term memory (BLSTM)-based model as a strong binary-class (correct/incorrect) sequence labeler that is trained with a class balancing objective. We experimentally confirmed that, by utilizing several types of ASR decoding scores as its auxiliary features, the model steadily shows high confidence estimation performance under highly imbalanced settings. We also confirmed that the BLSTM-based model outperforms Transformer-based confidence estimation models, which greatly underestimate incorrect tokens. Atsunori Ogawa, Naohiro Tawara, Takatomo Kano, Marc Delcroix |
ICASSP | 3 |
| 2021 | Transformer-Based Direct Speech-To-Speech Translation with TranscoderabstractTraditional speech translation systems use a cascade manner that concatenates speech recognition (ASR), machine translation (MT), and text-to-speech (TTS) synthesis to translate speech from one language to another language in a step-by-step manner. Unfortunately, since those components are trained separately, MT often struggles to handle ASR errors, resulting in unnatural translation results. Recently, one work attempted to construct direct speech translation in a single model. The model used a multi-task scheme that learns to predict not only the target speech spectrograms directly but also the source and target phoneme transcription as auxiliary tasks. However, that work was only evaluated Spanish-English language pairs with similar syntax and word order. With syntactically distant language pairs, speech translation requires distant word order, and thus direct speech frame-to-frame alignments become difficult. Another direction was to construct a single deep-learning framework while keeping the step-by-step translation process. However, such studies focused only on speech-to-text translation. Furthermore, all of these works were based on a recurrent neural net-work (RNN) model. In this work, we propose a step-by-step scheme to a complete end-to-end speech-to-speech translation and propose a Transformer-based speech translation using Transcoder. We compare our proposed and multi-task model using syntactically similar and distant language pairs. Takatomo Kano, Sakriani Sakti, Satoshi Nakamura 0001 |
SLT | 1 |
| 2020 | End-to-End Speech Translation With Transcoding by Multi-Task Learning for Distant Language PairsabstractDirectly translating spoken utterances from a source language to a target language is challenging because it requires a fundamental transformation in both linguistic and para/non-linguistic features. Traditional speech-to-speech translation approaches concatenate automatic speech recognition (ASR), text-to-text machine translation (MT), and text-to-speech synthesizer (TTS) by text information. The current state-of-the-art models for ASR, MT, and TTS have mainly been built using deep neural networks, in particular, an attention-based encoder-decoder neural network with an attention mechanism. Recently, several works have constructed end-to-end direct speech-to-text translation by combining ASR and MT into a single model. However, the usefulness of these models has only been investigated on language pairs of similar syntax and word order (e.g., English-French or English-Spanish). For syntactically distant language pairs (e.g., English-Japanese), speech translation requires distant word reordering. Furthermore, parallel texts with corresponding speech utterances that are suitable for training end-to-end speech translation are generally unavailable. Collecting such corpora is usually time-consuming and expensive. This article proposes the first attempt to build an end-to-end direct speech-to-text translation system on syntactically distant language pairs that suffer from long-distance reordering. We train the model on English (subject-verb-object (SVO) word order) and Japanese (SOV word order) language pairs. To guide the attention-based encoder-decoder model on this difficult problem, we construct end-to-end speech translation with transcoding and utilize curriculum learning (CL) strategies that gradually train the network for end-to-end speech translation tasks by adapting the decoder or encoder parts. We use TTS for data augmentation to generate corresponding speech utterances from the existing parallel text data. Our experiment results show that the proposed approach provides significant improvements compared with conventional cascade models and the direct speech translation approach that uses a single model without transcoding and CL strategies. Takatomo Kano, Sakriani Sakti, Satoshi Nakamura 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Neural Machine Translation with Acoustic EmbeddingabstractNeural machine translation (NMT) has successfully redefined the state of the art in machine translation on several language pairs. One popular framework models the translation process end-to-end using attentional encoder-decoder architecture and treats each word in the vectors of intermediate representation. These embedding vectors are sensitive to the meaning of words and allow semantically similar words to be near each other in the vector spaces and share their statistical power. Unfortunately, the model often maps such similar words too closely, which complicates distinguishing them. Consequently, NMT systems often mistranslate words that seem natural in the context but do not reflect the content of the source sentence. Incorporating auxiliary information usually enhances the discriminability. In this research, we integrate acoustic information within NMT by multi-task learning. Here, our model learns how to embed and translate word sequences based on their acoustic and semantic differences by helping it choose the correct output word based on its meaning and pronunciation. Our experiment results show that our proposed approach provides more significant improvement than the standard text-based transformer NMT model in BLEU score evaluation. Takatomo Kano, Sakriani Sakti, Satoshi Nakamura 0001 |
ASRU | 1 |
| 2018 | An end-to-end model for cross-lingual transformation of paralinguistic information
Takatomo Kano, Shinnosuke Takamichi, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura 0001 |
Mach. Transl. | 1 |
| 2017 | Structured-Based Curriculum Learning for End-to-End English-Japanese Speech TranslationabstractSequence-to-sequence attentional-based neural network architectures have been shown to provide a powerful model for machine translation and speech recognition.Recently, several works have attempted to extend the models for end-to-end speech translation task.However, the usefulness of these models were only investigated on language pairs with similar syntax and word order (e.g., English-French or English-Spanish).In this work, we focus on end-to-end speech translation tasks on syntactically distant language pairs (e.g., English-Japanese) that require distant word reordering.To guide the encoder-decoder attentional model to learn this difficult problem, we propose a structured-based curriculum learning strategy.Unlike conventional curriculum learning that gradually emphasizes difficult data examples, we formalize learning strategies from easier network structures to more difficult network structures.Here, we start the training with end-to-end encoder-decoder for speech recognition or text-based machine translation task then gradually move to end-to-end speech translation task.The experiment results show that the proposed approach could provide significant improvements in comparison with the one without curriculum learning. Takatomo Kano, Sakriani Sakti, Satoshi Nakamura 0001 |
INTERSPEECH | 1 |
| 2013 | Generalizing continuous-space translation of paralinguistic informationabstractIn previous work, we proposed a model for speech-to-speech translation that is sensitive to paralinguistic information such as duration and power of spoken words [1]. This model uses linear regression to map source acoustic features to target acoustic features directly and in continuous space. However, while the model is effective, it faces scalability issues as a single model must be trained for every word, which makes it difficult to generalize to words for which we do not have parallel speech. In this work we first demonstrate that simply training a linear regression model on all words is not sufficient to express paralinguistic translation. We next describe a neural network model that has sufficient expressive power to perform paralinguistic translation with a single model. We evaluate the proposed method on a digit translation task and show that we achieve similar results with a single neural network-based model as previous work did using word-dependent models. Index Terms: speech translation, paralinguistic information, linear regression, neural network Takatomo Kano, Shinnosuke Takamichi, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura 0001 |
INTERSPEECH | 1 |