EDBT 2026 Demo / reviewers in the wild / expert
Yuya Fujita
dblp:140/2817
· DBLP profile ↗
24ranked-venue papers
5as first author
17since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 5 first-author · 16 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Exploring Speech Recognition, Translation, and Understanding with Discrete Speech Units: A Comparative StudyabstractSpeech signals, typically sampled at rates in the tens of thousands per second, contain redundancies, evoking inefficiencies in sequence modeling. High-dimensional speech features such as spectrograms are often used as the input for the subsequent model. However, they can still be redundant. Recent investigations proposed the use of discrete speech units derived from self-supervised learning representations, which significantly compresses the size of speech data. Applying various methods, such as de-duplication and subword modeling, can further compress the speech sequence length. Hence, training time is significantly reduced while retaining notable performance. In this study, we undertake a comprehensive and systematic exploration into the application of discrete units within end-to-end speech processing models. Experiments on 12 automatic speech recognition, 3 speech translation, and 1 spoken language understanding corpora demonstrate that discrete units achieve reasonably good results in almost all the settings. Our configurations and trained models are released in ESPnet to foster future research efforts. Xuankai Chang, Brian Yan, Kwanghee Choi, Jee-Weon Jung, Soumi Maiti, Roshan S. Sharma, Jiatong Shi, Jinchuan Tian, Shinji Watanabe 0001, Yuya Fujita, Takashi Maekaku, Yao-Fei Cheng, Pavel Denisov, Kohei Saijo, Hsiu-Hsuan Wang |
ICASSP | 11 |
| 2024 | Hubertopic: Enhancing Semantic Representation of Hubert Through Self-Supervision Utilizing Topic ModelabstractRecently, the usefulness of self-supervised representation learning (SSRL) methods has been confirmed in various downstream tasks. Many of these models, as exemplified by HuBERT and WavLM, use pseudo-labels generated from spectral features or the model’s own representation features. From previous studies, it is known that the pseudo-labels contain semantic information. However, the masked prediction task, the learning criterion of HuBERT, focuses on local contextual information and may not make effective use of global semantic information such as speaker, theme of speech, and so on. In this paper, we propose a new approach to enrich the semantic representation of HuBERT. We apply topic model to pseudo-labels to generate a topic label for each utterance. An auxiliary topic classification task is added to HuBERT by using topic labels as teachers. This allows additional global semantic information to be incorporated in an unsupervised manner. Experimental results demonstrate that our method achieves comparable or better performance than the baseline in most tasks, including automatic speech recognition and five out of the eight SUPERB tasks. Moreover, we find that topic labels include various information about utterance, such as gender, speaker, and its theme. This highlights the effectiveness of our approach in capturing multifaceted semantic nuances. Takashi Maekaku, Jiatong Shi, Xuankai Chang, Yuya Fujita, Shinji Watanabe 0001 |
ICASSP | 4 |
| 2024 | Cross-Modal Multi-Tasking for Speech-to-Text Translation via Hard Parameter SharingabstractRecent works in end-to-end speech-to-text translation (ST) have proposed multi-tasking methods with soft parameter sharing which leverage machine translation (MT) data via secondary encoders that map text inputs to an eventual cross-modal representation. In this work, we instead propose a ST/MT multi-tasking framework with hard parameter sharing in which all model parameters are shared cross-modally. Our method reduces the speech-text modality gap via a pre-processing stage which converts speech and text inputs into two discrete token sequences of similar length – this allows models to indiscriminately process both modalities simply using a joint vocabulary. With experiments on MuST-C, we demonstrate that our multi-tasking framework improves attentional encoder-decoder, Connectionist Temporal Classification (CTC), transducer, and joint CTC/attention models by an average of +0.5 BLEU without any external MT data. Further, we show that this framework incorporates external MT data, yielding +0.8 BLEU, and also improves transfer learning from pre-trained textual models, yielding +1.8 BLEU.1 Brian Yan, Xuankai Chang, Antonios Anastasopoulos, Yuya Fujita, Shinji Watanabe 0001 |
ICASSP | 4 |
| 2024 | MC-Whisper: Extending Speech Foundation Models to Multichannel Distant Speech RecognitionabstractDistant Automatic Speech Recognition (DASR) stands as a crucial aspect in the realm of speech and audio processing. Recent advancements have spotlighted the efficacy of pre-trained speech foundation models, exemplified by Whisper, garnering considerable attention in the speech-processing domain. These models, trained on hundreds of thousands of hours of speech data, exhibit notable strengths in performance and generalization across various zero-shot scenarios. However, a limitation arises from their exclusive handling of single-channel input due to challenges in accumulating extensive multi-channel speech data. The spatial information in the multi-channel input is important for the DASR task. This study introduces an innovation by enabling the incorporation of multi-channel (MC) signals into the pre-trained Whisper model, called MC-Whisper. The proposed model introduces a multi-channel speech processing branch as a sidecar, to maximize the utilization of the foundation model's ability to handle multi-channel input. Experimental results on the distant microphone speech recordings from AMI meeting corpus demonstrate substantial improvements through the proposed approach. Xuankai Chang, Yuya Fujita, Takashi Maekaku, Shinji Watanabe 0001 |
IEEE Signal Process. Lett. | 3 |
| 2023 | LV-CTC: Non-Autoregressive ASR With CTC and Latent Variable ModelsabstractNon-autoregressive (NAR) models for automatic speech recognition (ASR) aim to achieve high accuracy and fast inference by simplifying the autoregressive (AR) generation process of conventional models. Connectionist temporal classification (CTC) is one of the key techniques used in NAR ASR models. In this paper, we propose a new model combining CTC and a latent variable model, which is one of the state-of-the-art models in the neural machine translation research field. A new neural network architecture and formulation specialized for ASR application are introduced. In the proposed model, CTC alignment is assumed to be dependent on the latent variables that are expected to capture dependencies between tokens. Experimental results on a 100 hours subset of Librispeech corpus showed the best recognition accuracy among CTC-based NAR models. On the TED-LIUM2 corpus, the best recognition accuracy is achieved including AR E2E models with faster inference speed. Yuya Fujita, Shinji Watanabe 0001, Xuankai Chang, Takashi Maekaku |
ASRU | 1 |
| 2023 | Fully Unsupervised Topic Clustering of Unlabelled Spoken Audio Using Self-Supervised Representation Learning and Topic ModelabstractUnsupervised topic clustering of spoken audio is an important research topic for zero-resourced unwritten languages. A classical approach is to find a set of spoken terms from only the audio based on dynamic time warping or generative modeling (e.g., hidden Markov model), and apply a topic model to classify topics. The spoken term discovery is the most important and difficult part. In this paper, we propose to combine self-supervised representation learning (SSRL) methods as a component of spoken term discovery and probabilistic topic models. Most SSRL methods pre-train a model which predicts high-quality pseudo labels generated from an audio-only corpus. These pseudo labels can be used to produce a sequence of pseudo subwords by applying deduplication and a subword model. Then, we apply a topic model based on latent Dirichlet allocation for these pseudo-subword sequences in an unsupervised manner. The clustering performance is evaluated on the Fisher corpus using normalized mutual information. We confirm the improvement of the proposed method and its effectiveness compared to an existing approach using dynamic time warping and topic models although the experimental setups are not directly comparable. Takashi Maekaku, Yuya Fujita, Xuankai Chang, Shinji Watanabe 0001 |
ICASSP | 2 |
| 2023 | Align, Write, Re-Order: Explainable End-to-End Speech Translation via Operation Sequence GenerationabstractThe black-box nature of end-to-end speech-to-text translation (E2E ST) makes it difficult to understand how source language inputs are being mapped to the target language. To solve this problem, we propose to simultaneously generate automatic speech recognition (ASR) and ST predictions such that each source language word is explicitly mapped to a target language word. A major challenge arises from the fact that translation is a non-monotonic sequence transduction task due to word ordering differences between languages – this clashes with the monotonic nature of ASR. Therefore, we propose to generate ST tokens out-of-order while remembering how to re-order them later. We achieve this by predicting a sequence of tuples consisting of a source word, the corresponding target words, and post-editing operations dictating the correct insertion points for the target word. We examine two variants of such operation sequences which enable generation of monotonic transcriptions and non-monotonic translations from the same speech input simultaneously. We apply our approach to offline and real-time streaming models, demonstrating that we can provide explainable translations without sacrificing quality or latency. In fact, the delayed re-ordering ability of our approach improves performance during streaming. As an added benefit, our method performs ASR and ST simultaneously, making it faster than using two separate systems to perform these tasks. Motoi Omachi, Brian Yan, Siddharth Dalmia, Yuya Fujita, Shinji Watanabe 0001 |
ICASSP | 4 |
| 2023 | Exploration of Efficient End-to-End ASR using Discretized Input from Self-Supervised Learning
Xuankai Chang, Brian Yan, Yuya Fujita, Takashi Maekaku, Shinji Watanabe 0001 |
INTERSPEECH | 3 |
| 2022 | An Exploration of Hubert with Large Number of Cluster Units and Model Assessment Using Bayesian Information CriterionabstractSelf-supervised learning (SSL) has become one of the most important technologies to realize spoken dialogue systems for languages that do not have much audio data and its transcription available. Speech representation models are one of the keys to achieving this, and have been actively studied in recent years. Among them, Hidden-Unit BERT (HuBERT) has shown promising results in automatic speech recognition (ASR) tasks. However, previous studies have investigated with limited iterations and cluster units. We explore HuBERT with larger numbers of clusters and iterations in order to obtain better speech representation. Furthermore, we introduce the Bayesian Information Criterion (BIC) as the performance measure of the model. Experimental results show that our model achieves the best performance in 5 out of 8 scores in the 4 metrics for the Zero Resource Speech 2021 task. It also outperforms the HuBERT BASE model trained with 960-hour LibriSpeech (LS) even though our model is only trained with 100-hour LS. In addition, we report that BIC is useful as a clue for determining the appropriate number of clusters to improve performance on phonetic, lexical, and syntactic metrics. Finally, we show that these findings are also effective for the ASR task. Takashi Maekaku, Xuankai Chang, Yuya Fujita, Shinji Watanabe 0001 |
ICASSP | 3 |
| 2022 | Non-Autoregressive End-To-End Automatic Speech Recognition Incorporating Downstream Natural Language ProcessingabstractWe propose a fast and accurate end-to-end (E2E) model, which executes automatic speech recognition (ASR) and downstream natural language processing (NLP) simultaneously. The proposed approach predicts a single-aligned sequence of transcriptions and linguistic annotations such as part-of-speech (POS) tags and named entity (NE) tags from speech. We use non-autoregressive (NAR) decoding instead of autoregressive (AR) decoding to reduce execution time since NAR can output multiple tokens in parallel across time. We use the connectionist temporal classification (CTC) model with mask-predict, i.e., Mask-CTC, to predict the single-aligned sequence accurately. Mask-CTC improves performance by joint training of CTC and a conditioned masked language model and refining output tokens with low confidence conditioned on reliable output tokens and audio embeddings. The proposed method jointly performs the ASR and downstream NLP task, i.e., POS or NE tagging, in a NAR manner. Experiments using the Corpus of Spontaneous Japanese and Spoken Language Understanding Resource Package show that the proposed E2E model can predict transcriptions and linguistic annotations with consistently better performance than vanilla CTC using greedy decoding and 15–97x faster than Transformer-based AR model. Motoi Omachi, Yuya Fujita, Shinji Watanabe 0001, Tianzi Wang |
ICASSP | 2 |
| 2022 | End-to-End Integration of Speech Recognition, Speech Enhancement, and Self-Supervised Learning RepresentationabstractThis work presents our end-to-end (E2E) automatic speech recognition (ASR) model targetting at robust speech recognition, called Integraded speech Recognition with enhanced speech Input for Self-supervised learning representation (IRIS).Compared with conventional E2E ASR models, the proposed E2E model integrates two important modules including a speech enhancement (SE) module and a self-supervised learning representation (SSLR) module.The SE module enhances the noisy speech.Then the SSLR module extracts features from enhanced speech to be used for speech recognition (ASR).To train the proposed model, we establish an efficient learning scheme.Evaluation results on the monaural CHiME-4 task show that the IRIS model achieves the best performance reported in the literature for the single-channel CHiME-4 benchmark (2.0% for the real development and 3.9% for the real test) thanks to the powerful pre-trained SSLR module and the finetuned SE module. Xuankai Chang, Takashi Maekaku, Yuya Fujita, Shinji Watanabe 0001 |
INTERSPEECH | 3 |
| 2022 | Attention Weight Smoothing Using Prior Distributions for Transformer-Based End-to-End ASR
Takashi Maekaku, Yuya Fujita, Yifan Peng 0003, Shinji Watanabe 0001 |
INTERSPEECH | 2 |
| 2021 | A Comparative Study on Non-Autoregressive Modelings for Speech-to-Text GenerationabstractNon-autoregressive (NAR) models simultaneously generate multiple outputs in a sequence, which significantly reduces the inference speed at the cost of accuracy drop compared to autoregressive baselines. Showing great potential for real-time applications, an increasing number of NAR models have been explored in different fields to mitigate the performance gap against AR models. In this work, we conduct a comparative study of various NAR modeling methods for end-to-end automatic speech recognition (ASR). Experiments are performed in the state-of-the-art setting using ESPnet. The results on various tasks provide interesting findings for developing an understanding of NAR ASR, such as the accuracy-speed trade-off and robustness against long-form utterances. We also show that the techniques can be combined for further improvement and applied to NAR end-to-end speech translation. All the implementations are publicly available to encourage further research in NAR speech processing. Yosuke Higuchi, Nanxin Chen, Yuya Fujita, Hirofumi Inaguma, Tatsuya Komatsu, Jaesong Lee, Jumon Nozaki, Tianzi Wang, Shinji Watanabe 0001 |
ASRU | 3 |
| 2021 | Toward Streaming ASR with Non-Autoregressive Insertion-Based ModelabstractNeural end-to-end (E2E) models have become a promising technique to realize practical automatic speech recognition (ASR) systems.When realizing such a system, one important issue is the segmentation of audio to deal with streaming input or long recording.After audio segmentation, the ASR model with a small real-time factor (RTF) is preferable because the latency of the system can be faster.Recently, E2E ASR based on non-autoregressive models becomes a promising approach since it can decode an N -length token sequence with less than N iterations.We propose a system to concatenate audio segmentation and non-autoregressive ASR to realize high accuracy and low RTF ASR.As a non-autoregressive ASR, the insertion-based model is used.In addition, instead of concatenating separated models for segmentation and ASR, we introduce a new architecture that realizes audio segmentation and non-autoregressive ASR by a single neural network.Experimental results on Japanese and English dataset show that the method achieved a reasonable trade-off between accuracy and RTF compared with baseline autoregressive Transformer and connectionist temporal classification. Yuya Fujita, Tianzi Wang, Shinji Watanabe 0001, Motoi Omachi |
Interspeech | 1 |
| 2021 | Speech Representation Learning Combining Conformer CPC with Deep Cluster for the ZeroSpeech Challenge 2021abstractWe present a system for the Zero Resource Speech Challenge 2021, which combines a Contrastive Predictive Coding (CPC) with deep cluster. In deep cluster, we first prepare pseudo-labels obtained by clustering the outputs of a CPC network with k-means. Then, we train an additional autoregressive model to classify the previously obtained pseudo-labels in a supervised manner. Phoneme discriminative representation is achieved by executing the second-round clustering with the outputs of the final layer of the autoregressive model. We show that replacing a Transformer layer with a Conformer layer leads to a further gain in a lexical metric. Experimental results show that a relative improvement of 35% in a phonetic metric, 1.5% in the lexical metric, and 2.3% in a syntactic metric are achieved compared to a baseline method of CPC-small which is trained on LibriSpeech 460h data. We achieve top results in this challenge with the syntactic metric. Takashi Maekaku, Xuankai Chang, Yuya Fujita, Shinji Watanabe 0001, Alexander I. Rudnicky |
Interspeech | 3 |
| 2021 | Streaming End-to-End ASR Based on Blockwise Non-Autoregressive ModelsabstractNon-autoregressive (NAR) modeling has gained more and more attention in speech processing.With recent state-of-the-art attention-based automatic speech recognition (ASR) structure, NAR can realize promising real-time factor (RTF) improvement with only small degradation of accuracy compared to the autoregressive (AR) models.However, the recognition inference needs to wait for the completion of a full speech utterance, which limits their applications on low latency scenarios.To address this issue, we propose a novel end-to-end streaming NAR speech recognition system by combining blockwiseattention and connectionist temporal classification with maskpredict (Mask-CTC) NAR.During inference, the input audio is separated into small blocks and then processed in a blockwise streaming way.To address the insertion and deletion error at the edge of the output of each block, we apply an overlapping decoding strategy with a dynamic mapping trick that can produce more coherent sentences.Experimental results show that the proposed method improves online ASR recognition in low latency conditions compared to vanilla Mask-CTC.Moreover, it can achieve a much faster inference speed compared to the AR attention-based models.All of our codes will be publicly available at https://github.com/espnet/espnet. Tianzi Wang, Yuya Fujita, Xuankai Chang, Shinji Watanabe 0001 |
Interspeech | 2 |
| 2021 | End-to-end ASR to jointly predict transcriptions and linguistic annotationsabstractMotoi Omachi, Yuya Fujita, Shinji Watanabe, Matthew Wiesner. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Motoi Omachi, Yuya Fujita, Shinji Watanabe 0001, Matthew Wiesner |
NAACL-HLT | 2 |
| 2020 | Attention-Based ASR with Lightweight and Dynamic ConvolutionsabstractEnd-to-end (E2E) automatic speech recognition (ASR) with sequence-to-sequence models has gained attention because of its simple model training compared with conventional hidden Markov model based ASR. Recently, several studies report the state-of-the-art E2E ASR results obtained by Transformer. Compared to recurrent neural network (RNN) based E2E models, training of Transformer is more efficient and also achieves better performance on various tasks. However, self-attention used in Transformer requires computation quadratic in its input length. In this paper, we propose to apply lightweight and dynamic convolution to E2E ASR as an alternative architecture to the self-attention to make the computational order linear. We also propose joint training with connectionist temporal classification, convolution on the frequency axis, and combination with self-attention. With these techniques, the proposed architectures achieve better performance than RNN-based E2E model and performance competitive to state-of-the-art Transformer on various ASR benchmarks including noisy/reverberant tasks. Yuya Fujita, Aswin Shanmugam Subramanian, Motoi Omachi, Shinji Watanabe 0001 |
ICASSP | 1 |
| 2020 | End-to-End ASR with Adaptive Span Self-Attention
Xuankai Chang, Aswin Shanmugam Subramanian, Shinji Watanabe 0001, Yuya Fujita, Motoi Omachi |
INTERSPEECH | 5 |
| 2020 | Insertion-Based Modeling for End-to-End Automatic Speech RecognitionabstractEnd-to-end (E2E) models have gained attention in the research field of automatic speech recognition (ASR).Many E2E models proposed so far assume left-to-right autoregressive generation of an output token sequence except for connectionist temporal classification (CTC) and its variants.However, left-toright decoding cannot consider the future output context, and it is not always optimal for ASR.One of the non-left-to-right models is known as non-autoregressive Transformer (NAT) and has been intensively investigated in the area of neural machine translation (NMT) research.One NAT model, mask-predict, has been applied to ASR but the model needs some heuristics or additional component to estimate the length of the output token sequence.This paper proposes to apply another type of NAT called insertion-based models, that were originally proposed for NMT, to ASR tasks.Insertion-based models solve the above mask-predict issues and can generate an arbitrary generation order of an output sequence.In addition, we introduce a new formulation of joint training of the insertionbased models and CTC.This formulation reinforces CTC by making it dependent on insertion-based token generation in a non-autoregressive manner.We conducted experiments on three public benchmarks and achieved competitive performance to strong autoregressive Transformer with a similar decoding condition. Yuya Fujita, Shinji Watanabe 0001, Motoi Omachi, Xuankai Chang |
INTERSPEECH | 1 |
| 2018 | Multi Scale Feedback Connection for Noise Robust Acoustic ModelingabstractSimply feeding of a last hidden layer of the deep neural network (DNN) back to the input layer recently found to be effective for noise robust acoustic modeling. Such high level feature strengthens the robustness of DNN based acoustic model while paying approximately twice the computational cost. In this paper, we proposed to feed such high level feature iteratively back to lower layers, which is referred as multi-scale feedback connection. With this intention, we firstly extract the high level feature at the last hidden layer of DNN. Second, this high level feature feed back to a lower scale features, they then generates a subsequent prediction as well as a subsequent high level feature. This subsequent high level feature is further feed down to a lower layers. We evaluated the proposed approach on both TIMIT and a large scale internal dataset. The large scale internal dataset includes voice search and far field dataset. Our finding is two aspects. First, at equivalent computational costs, the multiscale feedback connection outperforms the DNN, the DNN with skip connection and the DNN with feedback connection. The improvement is larger on the far field dataset. Second, pair layers-wise pretraining helps the proposed approach to converge better. Dung T. Tran, Ken-ichi Iso, Motoi Omachi, Yuya Fujita |
ICASSP | 4 |
| 2018 | Speaker Selective Beamformer with Keyword Mask EstimationabstractThis paper addresses the problem of automatic speech recognition (ASR) of a target speaker in background speech. The novelty of our approach is that we focus on a wakeup keyword, which is usually used for activating ASR systems like smart speakers. The proposed method firstly utilizes a DNN-based mask estimator to separate the mixture signal into the keyword signal uttered by the target speaker and the remaining background speech. Then the separated signals are used for calculating a beamforming filter to enhance the subsequent utterances from the target speaker. Experimental evaluations show that the trained DNN-based mask can selectively separate the keyword and background speech from the mixture signal. The effectiveness of the proposed method is also verified with Japanese ASR experiments, and we confirm that the character error rates are significantly improved by the proposed method for both simulated and real recorded test sets. Yusuke Kida, Dung T. Tran, Motoi Omachi, Toru Taniguchi, Yuya Fujita |
SLT | 5 |
| 2016 | Robust DNN-Based VAD Augmented with Phone Entropy Based Rejection of Background Speech
Yuya Fujita, Ken-ichi Iso |
INTERSPEECH | 1 |
| 2013 | Lightly supervised training for risk-based discriminative language modelsabstractWe propose a lightly supervised training method for a discriminative language model (DLM) based on risk minimization criteria. In lightly supervised training, pseudo labels generated by automatic speech recognition (ASR) are used as references. However, as these labels usually include recognition errors, the discriminative models estimated from such faulty reference labels may degrade ASR performance. Therefore, an approach to prevent performance degradation is necessary for discriminative language modeling. In our proposed lightly supervised training, the DLM is estimated from a “fused” risk, which is a relaxed version of the conventional Bayes risk. The fused risk is computed in a supervised manner when pseudo labels are accepted as references with high confidence while computed in an unsupervised manner when the labels are rejected due to low confidence. Accordingly, minimizing the fused risk for the training lattices results in a DLM with smoothed model parameters. The experimental results show that our proposed lightly supervised training method significantly reduced the word error rate compared with DLMs trained in conventional lightly supervised manners. Akio Kobayashi, Takahiro Oku, Yuya Fujita, Shoei Sato |
INTERSPEECH | 3 |