EDBT 2026 Demo / reviewers in the wild / expert
Yusuke Fujita
dblp:50/6352
· DBLP profile ↗
57ranked-venue papers
13as first author
26since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 49 · 11 first-author · 25 since 2021Artificial intelligence and machine learning · 42 · 11 first-author · 15 since 2021Systems, architecture and hardware · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Evaluating Japanese Dialect Robustness Across Speech and Text-based Large Language ModelsabstractDialogue systems based on large language models (LLMs) have advanced significantly in recent years. However, dialectal variation remains a major challenge, particularly for systems that process spoken input. LLM-based speech language models (SLMs), which integrate LLMs with speech processing components, show promise for spoken language tasks, yet their ability to comprehend dialects has not been sufficiently studied. Moreover, it remains unclear how the dialectal understanding of the base LLM affects SLM performance. This study investigates the dialectal robustness of both LLMs and SLMs using Japanese dialects as a test case. We define robustness as the ratio of performance on dialectal versus standard inputs, enabling fair comparisons. Our experiments show that SLM robustness correlates with that of their text-based counterparts. Furthermore, training with dialectal data and fine-tuning the speech encoder each improves robustness in SLMs. Tomoya Mizumoto, Yusuke Fujita, Lianbo Liu, Atsushi Kojima, Yui Sudo |
ASRU | 2 |
| 2025 | Serialized Output Prompting for Large Language Model-based Multi-Talker Speech RecognitionabstractPrompts are crucial for task definition and for improving the performance of large language models (LLM)-based systems. However, existing LLM-based multi-talker (MT) automatic speech recognition (ASR) systems either omit prompts or rely on simple task-definition prompts, with no prior work exploring the design of prompts to enhance performance. In this paper, we propose extracting serialized output prompts (SOP) and explicitly guiding the LLM using structured prompts to improve system performance (SOP-MT-ASR). A Separator and serialized Connectionist Temporal Classification (CTC) layers are inserted after the speech encoder to separate and extract MT content from the mixed speech encoding in a first-speaking-first-out manner. Subsequently, the SOP, which serves as a prompt for LLMs, is obtained by decoding the serialized CTC outputs using greedy search. To train the model effectively, we design a threestage training strategy, consisting of serialized output training (SOT) fine-tuning, serialized speech information extraction, and SOP-based adaptation. Experimental results on the LibriMix dataset show that, although the LLM-based SOT model performs well in the two-talker scenario, it fails to fully leverage LLMs under more complex conditions, such as the three-talker scenario. The proposed SOP approach significantly improved performance under both two- and three-talker conditions. Yusuke Fujita, Tomoya Mizumoto, Lianbo Liu, Atsushi Kojima, Yui Sudo |
ASRU | 2 |
| 2025 | Music Tagging with Classifier Group ChainsabstractWe propose music tagging with classifier chains that model the interplay of music tags. Most conventional methods estimate multiple tags independently by treating them as multiple independent binary classification problems. This treatment overlooks the conditional dependencies among music tags, leading to suboptimal tagging performance. Unlike most music taggers, the proposed method sequentially estimates each tag based on the idea of the classifier chains. Beyond the naive classifier chains, the proposed method groups the multiple tags by category, such as genre, and performs chains by unit of groups, which we call classifier group chains. Our method allows the modeling of the dependence between tag groups. We evaluate the effectiveness of the proposed method for music tagging performance through music tagging experiments using the MTG-Jamendo dataset. Furthermore, we investigate the effective order of chains for music tagging. Takuya Hasumi, Tatsuya Komatsu, Yusuke Fujita |
ICASSP | 3 |
| 2025 | Aligned Contrastive Learning for Text-to-Music RetrievalabstractThis paper proposes aligned contrastive learning for text-to-music retrieval. The proposed method introduces a new similarity measure, 'aligned similarity', which captures the frame-level and token-level correspondence within text and audio sequences. Unlike traditional approaches that aggregate sequence into clip-level and sentence-level embeddings, our method aligns the text token exhibiting the highest cosine similarity with each temporal frame of the audio sequence and averages these maximum similarity values across the entire sequence. This approach enables the capture of fine-grained relationships between audio and text that are often overlooked when sequences are aggregated into a single embedding. Retrieval experiments show significant performance improvements, with a notable gain being a 17.8% increase in Recall@5. Moreover, the alignment elucidates how specific audio frames correlate with textual tokens, enhancing the model's transparency and interpretability. Tatsuya Komatsu, Hokuto Munakata, Takuya Hasumi, Yusuke Fujita |
ICASSP | 4 |
| 2025 | AC/DC: LLM-based Audio Comprehension via Dialogue Continuation
Yusuke Fujita, Tomoya Mizumoto, Atsushi Kojima, Lianbo Liu, Yui Sudo |
INTERSPEECH | 1 |
| 2025 | DnR-nonverbal: Cinematic Audio Source Separation DatasetContaining Non-Verbal Sounds
Takuya Hasumi, Yusuke Fujita |
INTERSPEECH | 2 |
| 2025 | Is Synthetic Data Truly Effective for Training Speech Language Models?
Tomoya Mizumoto, Atsushi Kojima, Yusuke Fujita, Lianbo Liu, Yui Sudo |
INTERSPEECH | 3 |
| 2025 | OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary
Yui Sudo, Yusuke Fujita, Atsushi Kojima, Tomoya Mizumoto, Lianbo Liu |
INTERSPEECH | 2 |
| 2024 | Keep Decoding Parallel With Effective Knowledge Distillation From Language Models To End-To-End Speech RecognisersabstractThis study presents a novel approach for knowledge distillation (KD) from a BERT teacher model to an automatic speech recognition (ASR) model using intermediate layers. To distil the teacher’s knowledge, we use an attention decoder that learns from BERT’s token probabilities. Our method shows that language model (LM) information can be more effectively distilled into an ASR model using both the intermediate layers and the final layer. By using the intermediate layers as distillation target, we can more effectively distil LM knowledge into the lower network layers. Using our method, we achieve better recognition accuracy than with shallow fusion of an external LM, allowing us to maintain fast parallel decoding. Experiments on the LibriSpeech dataset demonstrate the effectiveness of our approach in enhancing greedy decoding with connectionist temporal classification (CTC). Michael Hentschel, Yuta Nishikawa, Tatsuya Komatsu, Yusuke Fujita |
ICASSP | 4 |
| 2024 | Audio Difference Learning for Audio CaptioningabstractThis study introduces a novel training paradigm, audio difference learning, for improving audio captioning. The fundamental concept of the proposed learning method is to create a feature representation space that preserves the relationship between audio, enabling the generation of captions that detail intricate audio information. This method employs a reference audio along with the input audio, both of which are transformed into feature representations via a shared encoder. Captions are then generated from these differential features to describe their differences. Furthermore, a unique technique is proposed that involves mixing the input audio with additional audio, and using the additional audio as a reference. This results in the difference between the mixed audio and the reference audio reverting back to the original input audio. This allows the original input’s caption to be used as the caption for their difference, eliminating the need for additional annotations for the differences. In the experiments using the Clotho and ESC50 datasets, the proposed method demonstrated an improvement in the SPIDEr score by 7% compared to conventional methods. Tatsuya Komatsu, Yusuke Fujita, Kazuya Takeda, Tomoki Toda |
ICASSP | 2 |
| 2024 | Audio Fingerprinting with Holographic Reduced Representations
Yusuke Fujita, Tatsuya Komatsu |
INTERSPEECH | 1 |
| 2024 | Song Data Cleansing for End-to-End Neural Singer Diarization Using Neural Analysis and Synthesis FrameworkabstractWe propose a data cleansing method that utilizes a neural analysis and synthesis (NANSY++) framework to train an end-to-end neural diarization model (EEND) for singer diarization.Our proposed model converts song data with choral singing commonly contained in popular music and unsuitable for generating a simulated dataset to the solo singing data.This cleansing is based on NANSY++, which is a framework trained to reconstruct an input non-overlapped audio signal.We exploit the pretrained NANSY++ to convert choral singing into clean, nonoverlapped audio.This cleansing process mitigates the mislabeling of choral singing to solo singing and helps the effective training of EEND models even when the majority of available song data contains choral singing sections.We experimentally evaluated the EEND model trained with a dataset using our proposed method using annotated popular duet songs.As a result, our proposed method improved 14.8 points in diarization error rate. Hokuto Munakata, Ryo Terashima, Yusuke Fujita |
INTERSPEECH | 3 |
| 2024 | Universal Score-based Speech Enhancement with High Content Preservation
Robin Scheibler, Yusuke Fujita, Yuma Shirahata, Tatsuya Komatsu |
INTERSPEECH | 2 |
| 2023 | Neural Diarization with Non-Autoregressive Intermediate AttractorsabstractEnd-to-end neural diarization (EEND) with encoder-decoder-based attractors (EDA) is a promising method to handle the whole speaker diarization problem simultaneously with a single neural network. While the EEND model can produce all frame-level speaker labels simultaneously, it disregards output label dependency. In this work, we propose a novel EEND model that introduces the label dependency between frames. The proposed method generates non-autoregressive intermediate attractors to produce speaker labels at the lower layers and conditions the subsequent layers with these labels. While the proposed model works in a non-autoregressive manner, the speaker labels are refined by referring to the whole sequence of intermediate labels. The experiments with the two-speaker CALLHOME dataset show that the intermediate labels with the proposed non-autoregressive intermediate attractors boost the diarization performance. The proposed method with the deeper net-work benefits more from the intermediate labels, resulting in better performance and training throughput than EEND-EDA. Yusuke Fujita, Tatsuya Komatsu, Robin Scheibler, Yusuke Kida, Tetsuji Ogawa |
ICASSP | 1 |
| 2023 | Target Vocabulary Recognition Based on Multi-Task Learning with Decomposed Teacher Sequences
Aoi Ito, Tatsuya Komatsu, Yusuke Fujita, Yusuke Kida |
INTERSPEECH | 3 |
| 2022 | Better Intermediates Improve CTC InferenceabstractThis paper proposes a method for improved CTC inference with searched intermediates and multi-pass conditioning.The paper first formulates self-conditioned CTC as a probabilistic model with an intermediate prediction as a latent representation and provides a tractable conditioning framework.We then propose two new conditioning methods based on the new formulation:(1) Searched intermediate conditioning that refines intermediate predictions with beam-search, (2) Multi-pass conditioning that uses predictions of previous inference for conditioning the next inference.These new approaches enable better conditioning than the original self-conditioned CTC during inference and improve the final performance.Experiments with the LibriSpeech dataset show relative 3%/12% performance improvement at the maximum in test clean/other sets compared to the original selfconditioned CTC. Tatsuya Komatsu, Yusuke Fujita, Jaesong Lee, Lukas Lee, Shinji Watanabe 0001, Yusuke Kida |
INTERSPEECH | 2 |
| 2022 | InterAug: Augmenting Noisy Intermediate Predictions for CTC-based ASRabstractThis paper proposes InterAug: a novel training method for CTC-based ASR using augmented intermediate representations for conditioning.The proposed method exploits the conditioning framework of self-conditioned CTC to train robust models by conditioning with "noisy" intermediate predictions.During the training, intermediate predictions are changed to incorrect intermediate predictions, and fed into the next layer for conditioning.The subsequent layers are trained to correct the incorrect intermediate predictions with the intermediate losses.By repeating the augmentation and the correction, iterative refinements, which generally require a special decoder, can be realized only with the audio encoder.To produce noisy intermediate predictions, we also introduce new augmentation: intermediate feature space augmentation and intermediate token space augmentation that are designed to simulate typical errors.The combination of the proposed InterAug framework with new augmentation allows explicit training of the robust audio encoders.In experiments using augmentations simulating deletion, insertion, and substitution error, we confirmed that the trained model acquires robustness to each error, boosting the speech recognition performance of the strong self-conditioned CTC baseline. Yu Nakagome, Tatsuya Komatsu, Yusuke Fujita, Shuta Ichimura, Yusuke Kida |
INTERSPEECH | 3 |
| 2022 | Alternate Intermediate Conditioning with Syllable-Level and Character-Level Targets for Japanese ASR
Yusuke Fujita, Tatsuya Komatsu, Yusuke Kida |
SLT | 1 |
| 2022 | Interdecoder: using Attention Decoders as Intermediate Regularization for CTC-Based Speech RecognitionabstractWe propose InterDecoder: a new non-autoregressive automatic speech recognition (NAR-ASR) training method that injects the advantage of token-wise autoregressive decoders while keeping the efficient non-autoregressive inference. The NAR-ASR models are often less accurate than autoregressive models such as Transformer decoder, which predict tokens conditioned on previously predicted tokens. The Inter-Decoder regularizes training by feeding intermediate encoder outputs into the decoder to compute the token-level prediction errors given previous ground-truth tokens, whereas the widely used Hybrid CTC/Attention model uses the decoder loss only at the final layer. In combination with Self-conditioned CTC, which uses the Intermediate CTC predictions to condition the encoder, performance is further improved. Experiments on the Librispeech and Tedlium2 dataset show that the proposed method shows a relative 6% WER improvement at the maximum compared to the conventional NAR-ASR methods. Tatsuya Komatsu, Yusuke Fujita |
SLT | 2 |
| 2022 | Encoder-Decoder Based Attractors for End-to-End Neural DiarizationabstractThis paper investigates an end-to-end neural diarization (EEND) method for an unknown number of speakers. In contrast to the conventional cascaded approach to speaker diarization, EEND methods are better in terms of speaker overlap handling. However, EEND still has a disadvantage in that it cannot deal with a flexible number of speakers. To remedy this problem, we introduce encoder-decoder-based attractor calculation module (EDA) to EEND. Once frame-wise embeddings are obtained, EDA sequentially generates speaker-wise attractors on the basis of a sequence-to-sequence method using an LSTM encoder-decoder. The attractor generation continues until a stopping condition is satisfied; thus, the number of attractors can be flexible. Diarization results are then estimated as dot products of the attractors and embeddings. The embeddings from speaker overlaps result in larger dot product values with multiple attractors; thus, this method can deal with speaker overlaps. Because the maximum number of output speakers is still limited by the training set, we also propose an iterative inference method to remove this restriction. Further, we propose a method that aligns the estimated diarization results with the results of an external speech activity detector, which enables fair comparison against cascaded approaches. Extensive evaluations on simulated and real datasets show that EEND-EDA outperforms the conventional cascaded approach. Shota Horiguchi, Yusuke Fujita, Shinji Watanabe 0001, Yawen Xue, L. Paola García-Perera |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | End-To-End Speaker Diarization as Post-ProcessingabstractThis paper investigates the utilization of an end-to-end diarization model as post-processing of conventional clustering-based diarization. Clustering-based diarization methods partition frames into clusters of the number of speakers; thus, they typically cannot handle overlapping speech because each frame is assigned to one speaker. On the other hand, some end-to-end diarization methods can handle overlapping speech by treating the problem as multi-label classification. Although some methods can treat a flexible number of speakers, they do not perform well when the number of speakers is large. To compensate for each other’s weakness, we propose to use a two-speaker end-to-end diarization method as post-processing of the results obtained by a clustering-based method. We iteratively select two speakers from the results and update the results of the two speakers to improve the overlapped region. Experimental results show that the proposed algorithm consistently improved the performance of the state-of-the-art methods across CALLHOME, AMI, and DIHARD II datasets. Shota Horiguchi, L. Paola García-Perera, Yusuke Fujita, Shinji Watanabe 0001, Kenji Nagamatsu |
ICASSP | 3 |
| 2021 | Semi-Supervised Training with Pseudo-Labeling for End-To-End Neural DiarizationabstractIn this paper, we present a semi-supervised training technique using pseudo-labeling for end-to-end neural diarization (EEND).The EEND system has shown promising performance compared with traditional clustering-based methods, especially in the case of overlapping speech.However, to get a welltuned model, EEND requires labeled data for all the joint speech activities of every speaker at each time frame in a recording.In this paper, we explore a pseudo-labeling approach that employs unlabeled data.First, we propose an iterative pseudolabel method for EEND, which trains the model using unlabeled data of a target condition.Then, we also propose a committeebased training method to improve the performance of EEND.To evaluate our proposed method, we conduct the experiments of model adaptation using labeled and unlabeled data.Experimental results on the CALLHOME dataset show that our proposed pseudo-label achieved a 37.4% relative diarization error rate reduction compared to a seed model.Moreover, we analyzed the results of semi-supervised adaptation with pseudo-labeling.We also show the effectiveness of our approach on the third DI-HARD dataset. Yuki Takashima, Yusuke Fujita, Shota Horiguchi, Shinji Watanabe 0001, L. Paola García-Perera, Kenji Nagamatsu |
Interspeech | 2 |
| 2021 | Online Streaming End-to-End Neural Diarization Handling Overlapping Speech and Flexible Numbers of Speakers
Yawen Xue, Shota Horiguchi, Yusuke Fujita, Yuki Takashima, Shinji Watanabe 0001, L. Paola García-Perera, Kenji Nagamatsu |
Interspeech | 3 |
| 2021 | Block-Online Guided Source SeparationabstractWe propose a block-online algorithm of guided source separation (GSS). GSS is a speech separation method that uses diarization information to update parameters of the generative model of observation signals. Previous studies have shown that GSS performs well in multi-talker scenarios. However, it requires a large amount of calculation time, which is an obstacle to the deployment of online applications. It is also a problem that the offline GSS is an utterance-wise algorithm so that it produces latency according to the length of the utterance. With the proposed algorithm, block-wise input samples and corresponding time annotations are concatenated with those in the preceding context and used to update the parameters. Using the context enables the algorithm to estimate time-frequency masks accurately only from one iteration of optimization for each block, and its latency does not depend on the utterance length but predetermined block length. It also reduces calculation cost by updating only the parameters of active speakers in each block and its context. Evaluation on the CHiME-6 corpus and a meeting corpus showed that the proposed algorithm achieved almost the same performance as the conventional offline GSS algorithm but with 32x faster calculation, which is sufficient for real-time applications. Shota Horiguchi, Yusuke Fujita, Kenji Nagamatsu |
SLT | 2 |
| 2021 | End-to-End Speaker Diarization Conditioned on Speech Activity and Overlap DetectionabstractIn this paper, we present a conditional multitask learning method for end-to-end neural speaker diarization (EEND). The EEND system has shown promising performance compared with traditional clustering-based methods, especially in the case of overlapping speech. In this paper, to further improve the performance of the EEND system, we propose a novel multitask learning framework that solves speaker diarization and a desired subtask while explicitly considering the task dependency. We optimize speaker diarization conditioned on speech activity and overlap detection that are subtasks of speaker diarization, based on the probabilistic chain rule. Experimental results show that our proposed method can leverage a subtask to effectively model speaker diarization, and outperforms conventional EEND systems in terms of diarization error rate. Yuki Takashima, Yusuke Fujita, Shinji Watanabe 0001, Shota Horiguchi, L. Paola García-Perera, Kenji Nagamatsu |
SLT | 2 |
| 2021 | Online End-To-End Neural Diarization with Speaker-Tracing BufferabstractThis paper proposes a novel online speaker diarization algorithm based on a fully supervised self-attention mechanism (SA-EEND). Online diarization inherently presents a speaker's permutation problem due to the possibility to assign speaker regions incorrectly across the recording. To circumvent this inconsistency, we proposed a speaker-tracing buffer mechanism that selects several input frames representing the speaker permutation information from previous chunks and stores them in a buffer. These buffered frames are stacked with the input frames in the current chunk and fed into a self-attention network. Our method ensures consistent diarization outputs across the buffer and the current chunk by checking the correlation between their corresponding outputs. Additionally, we trained SA-EEND with variable chunk-sizes to mitigate the mismatch between training and inference introduced by the speaker-tracing buffer mechanism. Experimental results, including online SA-EEND and variable chunk-size, achieved DERs of 12.54% for CALLHOME and 20.77% for CSJ with 1.4 s actual latency. Yawen Xue, Shota Horiguchi, Yusuke Fujita, Shinji Watanabe 0001, L. Paola García-Perera, Kenji Nagamatsu |
SLT | 3 |
| 2020 | Speaker Diarization with Region Proposal NetworkabstractSpeaker diarization is an important pre-processing step for many speech applications, and it aims to solve the "who spoke when" problem. Although the standard diarization systems can achieve satisfactory results in various scenarios, they are composed of several independently-optimized modules and cannot deal with the overlapped speech. In this paper, we propose a novel speaker diarization method: Region Proposal Network based Speaker Diarization (RPNSD). In this method, a neural network generates overlapped speech segment proposals, and compute their speaker embeddings at the same time. Compared with standard diarization systems, RPNSD has a shorter pipeline and can handle the overlapped speech. Experimental results on three diarization datasets reveal that RPNSD achieves remarkable improvements over the state-of-the-art x-vector baseline. Zili Huang, Shinji Watanabe 0001, Yusuke Fujita, L. Paola García-Perera, Yiwen Shao, Daniel Povey, Sanjeev Khudanpur |
ICASSP | 3 |
| 2020 | Speaker-Conditional Chain Model for Speech Separation and ExtractionabstractSpeech separation has been extensively explored to tackle the cocktail party problem. However, these studies are still far from having enough generalization capabilities for real scenarios. In this work, we raise a common strategy named Speaker-Conditional Chain Model to process complex speech recordings. In the proposed method, our model first infers the identities of variable numbers of speakers from the observation based on a sequence-to-sequence model. Then, it takes the information from the inferred speakers as conditions to extract their speech sources. With the predicted speaker information from whole observation, our model is helpful to solve the problem of conventional speech separation and speaker extraction for multi-round long recordings. The experiments from standard fully-overlapped speech separation benchmarks show comparable results with prior studies, while our proposed model gets better adaptability for multi-round long recordings. Jing Shi 0003, Jiaming Xu 0001, Yusuke Fujita, Shinji Watanabe 0001, Bo Xu 0002 |
INTERSPEECH | 3 |
| 2020 | End-to-End Speaker Diarization for an Unknown Number of Speakers with Encoder-Decoder Based AttractorsabstractEnd-to-end speaker diarization for an unknown number of speakers is addressed in this paper.Recently proposed end-toend speaker diarization outperformed conventional clusteringbased speaker diarization, but it has one drawback: it is less flexible in terms of the number of speakers.This paper proposes a method for encoder-decoder based attractor calculation (EDA), which first generates a flexible number of attractors from a speech embedding sequence.Then, the generated multiple attractors are multiplied by the speech embedding sequence to produce the same number of speaker activities.The speech embedding sequence is extracted using the conventional self-attentive end-to-end neural speaker diarization (SA-EEND) network.In a two-speaker condition, our method achieved a 2.69 % diarization error rate (DER) on simulated mixtures and a 8.07 % DER on the two-speaker subset of CALLHOME, while vanilla SA-EEND attained 4.56 % and 9.54 %, respectively.In unknown numbers of speakers conditions, our method attained a 15.29 % DER on CALLHOME, while the x-vectorbased clustering method achieved a 19.43 % DER. Shota Horiguchi, Yusuke Fujita, Shinji Watanabe 0001, Yawen Xue, Kenji Nagamatsu |
INTERSPEECH | 2 |
| 2020 | Utterance-Wise Meeting Transcription System Using Asynchronous Distributed MicrophonesabstractA novel framework for meeting transcription using asynchronous microphones is proposed in this paper.It consists of audio synchronization, speaker diarization, utterance-wise speech enhancement using guided source separation, automatic speech recognition, and duplication reduction.Doing speaker diarization before speech enhancement enables the system to deal with overlapped speech without considering sampling frequency mismatch between microphones.Evaluation on our real meeting datasets showed that our framework achieved a character error rate (CER) of 28.7 % by using 11 distributed microphones, while a monaural microphone placed on the center of the table had a CER of 38.2 %.We also showed that our framework achieved CER of 21.8 %, which is only 2.1 percentage points higher than the CER in headset microphone-based transcription. Shota Horiguchi, Yusuke Fujita, Kenji Nagamatsu |
INTERSPEECH | 2 |
| 2020 | Sequence to Multi-Sequence Learning via Conditional Chain Mapping for Mixture SignalsabstractNeural sequence-to-sequence models are well established for applications which can be cast as mapping a single input sequence into a single output sequence. In this work, we focus on one-to-many sequence transduction problems, such as extracting multiple sequential sources from a mixture sequence. We extend the standard sequence-to-sequence model to a conditional multi-sequence model, which explicitly models the relevance between multiple output sequences with the probabilistic chain rule. Based on this extension, our model can conditionally infer output sequences one-by-one by making use of both input and previously-estimated contextual output sequences. This model additionally has a simple and efficient stop criterion for the end of the transduction, making it able to infer the variable number of output sequences. We take speech data as a primary test field to evaluate our methods since the observed speech data is often composed of multiple sources due to the nature of the superposition principle of sound waves. Experiments on several different tasks including speech separation and multi-speaker speech recognition show that our conditional multi-sequence models lead to consistent improvements over the conventional non-conditional models. Jing Shi 0003, Xuankai Chang, Shinji Watanabe 0001, Yusuke Fujita, Jiaming Xu 0001, Bo Xu 0002, Lei Xie 0001 |
NeurIPS | 5 |
| 2019 | End-to-End Neural Speaker Diarization with Self-AttentionabstractSpeaker diarization has been mainly developed based on the clustering of speaker embeddings. However, the clustering-based approach has two major problems; i.e., (i) it is not optimized to minimize diarization errors directly, and (ii) it cannot handle speaker overlaps correctly. To solve these problems, the End-to-End Neural Diarization (EEND), in which a bidirectional long short-term memory (BLSTM) network directly outputs speaker diarization results given a multi-talker recording, was recently proposed. In this study, we enhance EEND by introducing self-attention blocks instead of BLSTM blocks. In contrast to BLSTM, which is conditioned only on its previous and next hidden states, self-attention is directly conditioned on all the other frames, making it much suitable for dealing with the speaker diarization problem. We evaluated our proposed method on simulated mixtures, real telephone calls, and real dialogue recordings. The experimental results revealed that the self-attention was the key to achieving good performance and that our proposed method performed significantly better than the conventional BLSTM-based method. Our method was even better than that of the state-of-the-art x-vector clustering-based method. Finally, by visualizing the latent representation, we show that the self-attention can capture global speaker characteristics in addition to local speech activity dynamics. Our source code is available online at https://github.com/hitachi-speech/EEND. Yusuke Fujita, Naoyuki Kanda, Shota Horiguchi, Yawen Xue, Kenji Nagamatsu, Shinji Watanabe 0001 |
ASRU | 1 |
| 2019 | Simultaneous Speech Recognition and Speaker Diarization for Monaural Dialogue Recordings with Target-Speaker Acoustic ModelsabstractThis paper investigates the use of target-speaker automatic speech recognition (TS-ASR) for simultaneous speech recognition and speaker diarization of single-channel dialogue recordings. TS-ASR is a technique to automatically extract and recognize only the speech of a target speaker given a short sample utterance of that speaker. One obvious drawback of TS-ASR is that it cannot be used when the speakers in the recordings are unknown because it requires a sample of the target speakers in advance of decoding. To remove this limitation, we propose an iterative method, in which (i) the estimation of speaker embeddings and (ii) TS-ASR based on the estimated speaker embeddings are alternately executed. We evaluated the proposed method by using very challenging dialogue recordings in which the speaker overlap ratio was over 20%. We confirmed that the proposed method significantly reduced both the word error rate (WER) and diarization error rate (DER). Our proposed method combined with i-vector speaker embeddings ultimately achieved a WER that differed by only 2.1 % from that of TS-ASR given oracle speaker embeddings. Furthermore, our method can solve speaker diarization simultaneously as a by-product and achieved better DER than that of the conventional clustering-based speaker diarization method based on i-vector. Naoyuki Kanda, Shota Horiguchi, Yusuke Fujita, Yawen Xue, Kenji Nagamatsu, Shinji Watanabe 0001 |
ASRU | 3 |
| 2019 | Acoustic Modeling for Distant Multi-talker Speech Recognition with Single- and Multi-channel BranchesabstractThis paper presents a novel heterogeneous-input multi-channel acoustic model (AM) that has both single-channel and multi-channel input branches. In our proposed training pipeline, a single-channel AM is trained first, then a multi-channel AM is trained starting from the single-channel AM with a randomly initialized multi-channel input branch. Our model uniquely uses the power of a complemen-tal speech enhancement (SE) module while exploiting the power of jointly trained AM and SE architecture. Our method was the foundation for the Hitachi/JHU CHiME-5 system that achieved the second-best result in the CHiME-5 competition, and this paper details various investigation results that we were not able to present during the competition period. We also evaluated and reconfirmed our method's effectiveness with the AMI Meeting Corpus. Our AM achieved a 30.12% word error rate (WER) for the development set and a 32.33% WER for the evaluation set for the AMI Corpus, both of which are the best results ever reported to the best of our knowledge. Naoyuki Kanda, Yusuke Fujita, Shota Horiguchi, Rintaro Ikeshita, Kenji Nagamatsu, Shinji Watanabe 0001 |
ICASSP | 2 |
| 2019 | Acoustic Modeling for Overlapping Speech Recognition: Jhu Chime-5 Challenge SystemabstractThis paper summarizes our acoustic modeling efforts in the Johns Hopkins University speech recognition system for the CHiME-5 challenge to recognize highly-overlapped dinner party speech recorded by multiple microphone arrays. We explore data augmentation approaches, neural network architectures, front-end speech dereverberation, beamforming and robust i-vector extraction with comparisons of our in-house implementations and publicly available tools. We finally achieved a word error rate of 69.4% on the development set, which is a 11.7% absolute improvement over the previous baseline of 81.1%, and release this improved baseline with refined techniques/tools as an advanced CHiME-5 recipe. Vimal Manohar, Szu-Jui Chen, Yusuke Fujita, Shinji Watanabe 0001, Sanjeev Khudanpur |
ICASSP | 4 |
| 2019 | End-to-End Neural Speaker Diarization with Permutation-Free ObjectivesabstractIn this paper, we propose a novel end-to-end neural-network-based speaker diarization method. Unlike most existing methods, our proposed method does not have separate modules for extraction and clustering of speaker representations. Instead, our model has a single neural network that directly outputs speaker diarization results. To realize such a model, we formulate the speaker diarization problem as a multi-label classification problem, and introduces a permutation-free objective function to directly minimize diarization errors without being suffered from the speaker-label permutation problem. Besides its end-to-end simplicity, the proposed method also benefits from being able to explicitly handle overlapping speech during training and inference. Because of the benefit, our model can be easily trained/adapted with real-recorded multi-speaker conversations just by feeding the corresponding multi-speaker segment labels. We evaluated the proposed method on simulated speech mixtures. The proposed method achieved diarization error rate of 12.28%, while a conventional clustering-based system produced diarization error rate of 28.77%. Furthermore, the domain adaptation with real-recorded speech provided 25.6% relative improvement on the CALLHOME dataset. Our source code is available online at https://github.com/hitachi-speech/EEND. Yusuke Fujita, Naoyuki Kanda, Shota Horiguchi, Kenji Nagamatsu, Shinji Watanabe 0001 |
INTERSPEECH | 1 |
| 2019 | Guided Source Separation Meets a Strong ASR Backend: Hitachi/Paderborn University Joint Investigation for Dinner Party ASRabstractIn this paper, we present Hitachi and Paderborn University's joint effort for automatic speech recognition (ASR) in a dinner party scenario.The main challenges of ASR systems for dinner party recordings obtained by multiple microphone arrays are (1) heavy speech overlaps, (2) severe noise and reverberation, (3) very natural conversational content, and possibly (4) insufficient training data.As an example of a dinner party scenario, we have chosen the data presented during the CHiME-5 speech recognition challenge, where the baseline ASR had a 73.3% word error rate (WER), and even the best performing system at the CHiME-5 challenge had a 46.1% WER.We extensively investigated a combination of the guided source separation-based speech enhancement technique and an already proposed strong ASR backend and found that a tight combination of these techniques provided substantial accuracy improvements.Our final system achieved WERs of 39.94% and 41.64% for the development and evaluation data, respectively, both of which are the best published results for the dataset.We also investigated with additional training data on the official small data in the CHiME-5 corpus to assess the intrinsic difficulty of this ASR task. Naoyuki Kanda, Christoph Böddeker, Jens Heitkaemper, Yusuke Fujita, Shota Horiguchi, Kenji Nagamatsu, Reinhold Häb-Umbach |
INTERSPEECH | 4 |
| 2019 | Auxiliary Interference Speaker Loss for Target-Speaker Speech RecognitionabstractIn this paper, we propose a novel auxiliary loss function for target-speaker automatic speech recognition (ASR). Our method automatically extracts and transcribes target speaker's utterances from a monaural mixture of multiple speakers speech given a short sample of the target speaker. The proposed auxiliary loss function attempts to additionally maximize interference speaker ASR accuracy during training. This will regularize the network to achieve a better representation for speaker separation, thus achieving better accuracy on the target-speaker ASR. We evaluated our proposed method using two-speaker-mixed speech in various signal-to-interference-ratio conditions. We first built a strong target-speaker ASR baseline based on the state-of-the-art lattice-free maximum mutual information. This baseline achieved a word error rate (WER) of 18.06% on the test set while a normal ASR trained with clean data produced a completely corrupted result (WER of 84.71%). Then, our proposed loss further reduced the WER by 6.6% relative to this strong baseline, achieving a WER of 16.87%. In addition to the accuracy improvement, we also showed that the auxiliary output branch for the proposed loss can even be used for a secondary ASR for interference speakers' speech. Naoyuki Kanda, Shota Horiguchi, Ryoichi Takashima, Yusuke Fujita, Kenji Nagamatsu, Shinji Watanabe 0001 |
INTERSPEECH | 4 |
| 2018 | Sequence Distillation for Purely Sequence Trained Acoustic ModelsabstractThis paper presents our exploration into teacher-student (TS) training for acoustic models (AMs) based on the lattice-free maximum mutual information technique. Whereas most previous studies of TS training used a frame-level distance between teacher and student models' distributions, we propose using the sequence-level temper-atured Kullback-Leibler divergence as a metric for TS training. In our experiment on the AMI meeting corpus, we prepared a strong teacher model consisting of a convolutional neural network, time delay neural network, and long short-term memory, which had 47.7M parameters and achieved a state-of-the-art word error rate (WER) of 18.05%. Whereas the small student AM (10.8M params. and 19.72% WER) trained by a frame-level TS training was able to fill only 43% of the WER gap between teacher and student AMs, the student AM trained by the proposed method achieved a 18.23% WER, filling 89% of the WER gap from the teacher AM. We also show that the frame-level TS training sometimes even degrades the performance of the student model whereas the proposed method consistently improved the accuracy. Naoyuki Kanda, Yusuke Fujita, Kenji Nagamatsu |
ICASSP | 2 |
| 2018 | Lattice-free State-level Minimum Bayes Risk Training of Acoustic Models
Naoyuki Kanda, Yusuke Fujita, Kenji Nagamatsu |
INTERSPEECH | 2 |
| 2017 | Investigation of lattice-free maximum mutual information-based acoustic models with sequence-level Kullback-Leibler divergenceabstractLattice-free maximum mutual information (LFMMI) was recently proposed as a mixture of the ideas of hidden-Markov-model-based acoustic models (AMs) and connectionist-temporal-classification-based AMs. In this paper, we investigate LFMMI from various perspectives of model combination, teacher-student training, and unsupervised speaker adaptation. Especially, we thoroughly investigate the use of the “sequence-level” Kullback-Leibler divergence with its novel and simple error derivation to enhance LFMMI-based AMs. In our experiment, we used the corpus of spontaneous Japanese (CSJ). Our best AM was an ensemble of three types of time delay neural networks and one long short-term memory-based network, and it finally achieved a WER of 6.94%, which is, to the best of our knowledge, the best published result for the CSJ. Naoyuki Kanda, Yusuke Fujita, Kenji Nagamatsu |
ASRU | 2 |
| 2016 | Training ROI Selection Based on MILBoost for Liver Cirrhosis Classification Using Ultrasound Images
Yusuke Fujita, Yoshihiro Mitani, Yoshihiko Hamamoto, Makoto Segawa, Shuji Terai, Isao Sakaida |
IEA/AIE | 1 |
| 2016 | Data Augmentation Using Multi-Input Multi-Output Source Separation for Deep Neural Network Based Acoustic Modeling
Yusuke Fujita, Ryoichi Takashima, Takeshi Homma, Masahito Togami |
INTERSPEECH | 1 |
| 2015 | Unified ASR system using LGM-based source separation, noise-robust feature extraction, and word hypothesis selectionabstractIn this paper, we propose a unified system that incorporates speech source separation and automatic speech recognition for various noise environments. There are three features in the proposed system. The first feature of the proposed method is the LGM (local Gaussian modeling) based source separation with the efficient permutation alignment method that integrates a power spectrum correlation based method and a direction-of-arrival (DOA) based method. Evaluation results show that using the separated speech with the baseline acoustic modeling method reduces the word error rate (WER) significantly. The second feature of the proposed method is multi-condition training with per-utterance normalized features and noise-aware features in the acoustic modeling step. In this paper, we show that the proposed training method is effective even when an input signal has been distorted through the source separation step. The third feature is the word hypothesis selection method for integrating multiple recognition results. The proposed selection method estimates correct words based on a recognizer's confidence and co-occurrence characteristics. The evaluation results show that the proposed selection method outperforms the conventional recognizer output voting error reduction (ROVER) method. The proposed system is evaluated using the third CHiME challenge dataset. Evaluation results show that the proposed system resulted in an improvement of 66.1% over the baseline system. Yusuke Fujita, Ryoichi Takashima, Takeshi Homma, Rintaro Ikeshita, Yohei Kawaguchi, Takashi Sumiyoshi, Takashi Endo, Masahito Togami |
ASRU | 1 |
| 2015 | Novel Architecture for Cellular Neural Network Suitable for High-Density Integration of Electron Devices-Learning of Multiple Logics
Mutsumi Kimura, Yusuke Fujita, Tomohiro Kasakawa, Tokiyoshi Matsuda |
ICONIP (1) | 2 |
| 2014 | Comparative Study of Classifiers for Prediction of Recurrence of Liver Cancer Using Binary Patterns
Hiroyuki Ogihara, Yusuke Fujita, Norio Iizuka, Masaaki Oka, Yoshihiko Hamamoto |
IEA/AIE (2) | 2 |
| 2014 | A Method of Bubble Removal for Computer-Assisted Diagnosis of Capsule Endoscopic Images
Masato Suenaga, Yusuke Fujita, Shinichi Hashimoto, Shuji Terai, Isao Sakaida, Yoshihiko Hamamoto |
IEA/AIE (2) | 2 |
| 2012 | The Use of a Local Histogram Feature Vector of Classifying Diffuse Lung Opacities in High-Resolution Computed Tomography
Yoshihiro Mitani, Yusuke Fujita, Naofumi Matsunaga, Yoshihiko Hamamoto |
IEA/AIE | 2 |
| 2011 | A robust automatic crack detection method from noisy concrete surfaces
Yusuke Fujita, Yoshihiko Hamamoto |
Mach. Vis. Appl. | 1 |
| 2010 | An Improved Method for Cirrhosis Detection Using Liver's Ultrasound ImagesabstractThis paper describes an improved method for cirrhosis detection in the liver using Gabor features from ultrasound images. There are three main contributions of our cirrhosis detection method. The first contribution of this method is to combine weak classifiers using the AdaBoost algorithm. The second one is to use an artificial dataset to avoid the problem of over fitting the limited training dataset. The third one is to apply a voting classification with use of multiple regions of interest (ROIs). Although the accuracy rate of a single classifier designed with only original dataset was 56%, that of the proposed method was 80% in cross-validation. Yusuke Fujita, Yoshihiko Hamamoto, Makoto Segawa, Shuji Terai, Isao Sakaida |
ICPR | 1 |
| 2009 | A Robust Method for Automatically Detecting Cracks on Noisy Concrete Surfaces
Yusuke Fujita, Yoshihiko Hamamoto |
IEA/AIE | 1 |
| 2008 | Visualization of transitions of developing of hepatitis C virus-associated hepatocellular carcinomaabstractIn our previous study, we visualized microarray data of hepatocellular carcinoma (HCC) by using self-organizing-map, and investigated molecular signature representing the development of HCC. In this study, we propose two visualization methods of microarray data with Euclidean distance classifiers and Sammonpsilas nonlinear mapping. Our proposed methods will serve as tool to discover molecular signature representing the development of HCC for molecular biologists or doctors. Takanobu Miyamoto, Yusuke Fujita, Shunji Uchimura, Yoshihiko Hamamoto, Norio Iizuka, Masaaki Oka |
ICPR | 2 |
| 2005 | Estimation of Physically and Physiologically Valid Somatosensory InformationabstractThe goal of this research is to enable precise estimation of human muscle forces in whole-body motions based not only on physiological muscle model but also on the equation of motion. The potential application areas include human-machine interface, medicine, biomechanics, and computer animation. Towards this goal, in this paper we discuss the inverse dynamics of musculoskeletal human model using the data from electromyogram (EMG) and force sensor. The inverse dynamics of musculoskeletal human models is formulated as an optimization problem subject to equality and inequality conditions taking into account of the equation of motion and muscle model from physiology literature. We evaluate the developed algorithm on a complex musculoskeletal model with 366 muscles driving a skeleton with 155 degrees of freedom. Katsu Yamane, Yusuke Fujita, Yoshihiko Nakamura |
ICRA | 2 |
| 2004 | CELP-based speaker verification: an evaluation under noisy conditionsabstractWe propose a text-independent speaker verification method based on a speech coding scheme. The proposed method utilizes CELP parameters which are used in speech coding schemes for mobile communication systems, and verifies a speaker only with the encoded speech information. The reliability of the proposed method under noisy conditions is mainly discussed with some simulation results. Yasushi Yamazaki, Yusuke Fujita, Naohisa Komatsu |
ICARCV | 2 |
| 2004 | Computing a Set of Local Optimal Paths through Cluttered Environments and over Open TerrainabstractThis paper describes an efficient algorithm to generate a set of local optimal paths between two given end points in cluttered environments or over open terrain. The local optimal paths are selected from the set of shortest constrained paths through every node (one for each path) in the graph, generated by running twice a "single-source" search. The initial set of the shortest constrained paths spans the entire search space and includes local optimal paths with costs equal or better than the longest constrained path in the set. The search for the optimal path is transformed to a search for the best path in each homotopy class generated by this search. The initial search is of complexity O(nlogn), and the pruning procedure is O(nmlogm), where n is the number of nodes and m is the number of homotopy classes generated by this search. The algorithm is demonstrated for motion planning on rough terrain. Zvi Shiller, Yusuke Fujita, Dan Ophir, Yoshihiko Nakamura |
ICRA | 2 |
| 2004 | A Study on Nonparametric Classifiers for a CAD System of Diffuse Lung Opacities in Thin-Section Computed Tomography Images
Yoshihiro Mitani, Yusuke Fujita, Naofumi Matsunaga, Yoshihiko Hamamoto |
KES | 2 |
| 2003 | Dual Dijkstra search for paths with different topologiesabstractThis paper describes a new search algorithm, the Dual Dijkstra Search. From a given initial and final configuration, Dual Dijkstra Search finds various paths which have different topologies simultaneously. This algorithm allows you to enumerate not only the optimal one but variety of meaningful candidates among local minimum paths. It is based on the algorithm of Dijkstra, which is popularly used to find an optimal solution. The method consists of two procedures: First computes local minima and ranks the paths in order of optimality. Then classify them with their topological properties and take out only the optimal paths in each groups. Computed examples include generating collision-free motion along 2D space and motion planning of 3-DOF robot. We also proposed the idea of motion compression, which simplifies the high dimensional motion planning problem. Together with this idea, we applied Dual Dijkstra Search to 7-DOF arm manipulation problem and succeeded in obtaining variety of motion candidates. Yusuke Fujita, Yoshihiko Nakamura, Zvi Shiller |
ICRA | 1 |