VLDB 2026 Research / reviewers in the wild / expert
Yifan Peng 0003
dblp:150/7295-3
· DBLP profile ↗
44ranked-venue papers
11as first author
44since 2021 · last 2025
0000-0002-8581-8674ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 8 first-author · 37 since 2021Artificial intelligence and machine learning · 28 · 8 first-author · 28 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing Audiovisual Speech Recognition Through Bifocal Preference OptimizationabstractAudiovisual Automatic Speech Recognition (AV-ASR) aims to improve speech recognition accuracy by leveraging visual signals. It is particularly challenging in unconstrained real-world scenarios across various domains due to noisy acoustic environments, spontaneous speech, and the uncertain use of visual information. Most previous works fine-tune audio-only ASR models on audiovisual datasets, optimizing them for conventional ASR objectives. However, they often neglect visual features and common errors in unconstrained video scenarios. In this paper, we propose using a preference optimization strategy to improve speech recognition accuracy for real-world videos. First, we create preference data via simulating common errors that occurred in AV-ASR from two focals: manipulating the audio or vision input and rewriting the output transcript. Second, we propose BPO-AVASR, a Bifocal Preference Optimization method to improve AV-ASR models by leveraging both input-side and output-side preference. Extensive experiments demonstrate that our approach significantly improves speech recognition accuracy across various domains, outperforming previous state-of-the-art models on real-world video speech recognition. Yihan Wu 0008, Yifan Peng 0003, Xihua Wang 0002, Ruihua Song, Shinji Watanabe 0001 |
AAAI | 3 |
| 2025 | Open Full-duplex Voice Agent with Speech-to-Speech Language ModelabstractWe present the system demonstration and opensource code release of a novel, data-efficient framework that converts any standard text Large Language Model (LLM) into a full-duplex end-to-end (E2E) speech-to-speech (S2S) model, for building conversational voice agents. Our new modeling method enables any LLMs to simultaneously listen and speak without requiring extensive speech-text pretraining. Moreover, we demonstrate how to put together a low-latency and full-duplex voice agent with open-source modeling, inference optimization, and serving solutions. This work significantly lowers the barrier to entry for developing low-latency, human-like voice agents by providing a generalizable, end-to-end solution built on open-source technologies. Edresson Casanova, Chen Chen 0075, Kevin Hu, Ankita Pasad, Elena Rastorgueva, Seelan Lakshmi Narasimhan, Slyne Deng, Ehsan Hosseini-Asl, Piotr Zelasko, Valentin Mendelev, Subhankar Ghosh, Yifan Peng 0003, Zhehuai Chen, Jason Li 0007, Jagadeesh Balam, Vitaly Lavrukhin, Boris Ginsburg |
ASRU | 12 |
| 2025 | Unifying Diarization, Separation, and ASR with Multi-Speaker EncoderabstractThis paper presents a unified multi-speaker encoder (UME), a novel architecture that jointly learns representations for speaker diarization (SD), speech separation (SS), and multi-speaker automatic speech recognition (ASR) tasks using a shared speech foundational encoder. We leverage the hidden representations from multiple layers of UME as a residual weighted-sum encoding (RWSE) to effectively use information from different semantic levels, contributing to bottom-up alignment between tasks. This joint training approach captures the inherent inter-dependencies among the tasks, enhancing overall performance on overlapping speech data. Our evaluations demonstrate that UME substantially improves over the single-task baselines dedicated to SD, SS, and multi-speaker ASR on LibriMix evaluation sets. Notably, for SD, UME outperforms the previous studies, achieving diarization error rates of 1.37% and 2.29% on Libri2Mix and Libri3Mix evaluation sets, respectively. Muhammad Shakeel 0001, Yui Sudo, Yifan Peng 0003, Chyi-Jiunn Lin, Shinji Watanabe 0001 |
ASRU | 3 |
| 2025 | Context-aware Dynamic Pruning for Speech Foundation ModelsabstractFoundation models, such as large language models, have achieved remarkable success in natural language processing and are evolving into models capable of handling multiple modalities.
Listening ability, in particular, is crucial for many applications, leading to research on building speech foundation models. However, the high computational cost of these large models presents a significant challenge for real-world applications. Although substantial efforts have been made to reduce computational costs, such as through pruning techniques, the majority of these approaches are applied primarily during the training phase for specific downstream tasks. In this study, we hypothesize that optimal pruned networks may vary based on contextual factors such as speaker characteristics, languages, and tasks. To address this, we propose a dynamic pruning technique that adapts to these contexts during inference without altering the underlying model. We demonstrated that we could successfully reduce inference time by approximately 30\% while maintaining accuracy in multilingual/multi-task scenarios. We also found that the obtained pruned structure offers meaningful interpretations based on the context, e.g., task-related information emerging as the dominant factor for efficient pruning. Masao Someki, Yifan Peng 0003, Siddhant Arora, Athanasios Mouchtaris, Grant P. Strimel, Shinji Watanabe 0001 |
ICLR | 2 |
| 2025 | OWLS: Scaling Laws for Multilingual Speech Recognition and Translation ModelsabstractNeural scaling laws offer valuable insights for designing robust sequence processing architectures. While these laws have been extensively characterized in other modalities, their behavior in speech remains comparatively underexplored. In this work, we introduce OWLS, an open-access, reproducible suite of multilingual speech recognition and translation models spanning 0.25B to 18B parameters, with the 18B version being the largest speech model, to the best of our knowledge. OWLS leverages up to 360K hours of public speech data across 150 languages, enabling a systematic investigation into how data, model, and compute scaling each influence performance in multilingual speech tasks. We use OWLS to derive neural scaling laws, showing how final performance can be reliably predicted when scaling. Scaling to larger models can improve ASR performance across the board, in both low and high resource languages, improving the accessibility of speech technologies. Finally, we show how OWLS can be used to power new research directions by discovering emergent abilities in large-scale speech models. Model checkpoints will be released on https://huggingface.co/collections/espnet/owls-scaling-laws-for-speech-recognition-and-translation-67ab7f991c194065f057ce8d for future studies. Jinchuan Tian, Yifan Peng 0003, Brian Yan, Chao-Han Huck Yang, Shinji Watanabe 0001 |
ICML | 3 |
| 2025 | OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
Yifan Peng 0003, Muhammad Shakeel 0001, Yui Sudo, Jinchuan Tian, Chyi-Jiunn Lin, Shinji Watanabe 0001 |
INTERSPEECH | 1 |
| 2025 | Exploring Linear Variant Transformers and k-NN Memory Inference for Long-Form ASR
Carlos Carvalho 0003, Jinchuan Tian, Yifan Peng 0003, Alberto Abad, Shinji Watanabe 0001 |
INTERSPEECH | 4 |
| 2025 | Granary: Speech Recognition and Translation Dataset in 25 European Languages
Nithin Rao Koluguri, Monica Sekoyan, George Zelenfroynd, Sasha Meister, Shuoyang Ding, Sofia Kostandian, He Huang 0012, Nikolay Karpov, Jagadeesh Balam, Vitaly Lavrukhin, Yifan Peng 0003, Sara Papi, Marco Gaido, Alessio Brutti, Boris Ginsburg |
INTERSPEECH | 11 |
| 2025 | DYNAC: Dynamic Vocabulary-based Non-Autoregressive Contextualization for Speech Recognition
Yui Sudo, Yosuke Fukumoto, Muhammad Shakeel 0001, Yifan Peng 0003, Chyi-Jiunn Lin, Shinji Watanabe 0001 |
INTERSPEECH | 4 |
| 2025 | OpusLM: A Family of Open Unified Speech Language Models
Jinchuan Tian, Yifan Peng 0003, Jiatong Shi, Siddhant Arora, Shikhar Bharadwaj, Takashi Maekaku, Yusuke Shinohara, Keita Goto, Xiang Yue, Chao-Han Huck Yang, Shinji Watanabe 0001 |
INTERSPEECH | 3 |
| 2025 | Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with Data Augmentation and LID-Aware CTC
Qingzheng Wang, Jiancheng Sun, Yifan Peng 0003, Shinji Watanabe 0001 |
INTERSPEECH | 3 |
| 2025 | VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-TuningabstractYifan Peng, Krishna C Puvvada, Zhehuai Chen, Piotr Zelasko, He Huang, Kunal Dhawan, Ke Hu, Shinji Watanabe, Jagadeesh Balam, Boris Ginsburg. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yifan Peng 0003, Krishna C. Puvvada, Zhehuai Chen, Piotr Zelasko, He Huang 0012, Kunal Dhawan, Shinji Watanabe 0001, Jagadeesh Balam, Boris Ginsburg |
NAACL (Long Papers) | 1 |
| 2024 | OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language IdentificationabstractThere has been an increasing interest in large speech models that can perform multiple tasks in a single model.Such models usually adopt an encoder-decoder or decoder-only architecture due to their popularity and good performance in many domains.However, autoregressive models can be slower during inference compared to non-autoregressive models and also have potential risks of hallucination.Though prior studies observed promising results of non-autoregressive models for certain tasks at small scales, it remains unclear if they can be scaled to speech-to-text generation in diverse languages and tasks.Inspired by the Open Whisper-style Speech Model (OWSM) project, we propose OWSM-CTC, a novel encoder-only speech foundation model based on Connectionist Temporal Classification (CTC).It is trained on 180k hours of public audio data for multilingual automatic speech recognition (ASR), speech translation (ST), and language identification (LID).Compared to encoder-decoder OWSM, our OWSM-CTC achieves competitive results on ASR and up to 24% relative improvement on ST, while it is more robust and 3 to 4 times faster for inference.OWSM-CTC also improves the longform ASR result with 20x speed-up.We will publicly release our code, pre-trained model, and training logs to promote open science in speech foundation models. 1 Yifan Peng 0003, Yui Sudo, Muhammad Shakeel 0001, Shinji Watanabe 0001 |
ACL (1) | 1 |
| 2024 | Towards Robust Speech Representation Learning for Thousands of LanguagesabstractWilliam Chen, Wangyou Zhang, Yifan Peng, Xinjian Li, Jinchuan Tian, Jiatong Shi, Xuankai Chang, Soumi Maiti, Karen Livescu, Shinji Watanabe. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Wangyou Zhang, Yifan Peng 0003, Jinchuan Tian, Jiatong Shi, Xuankai Chang, Soumi Maiti, Karen Livescu, Shinji Watanabe 0001 |
EMNLP | 3 |
| 2024 | Dynamic-Superb: Towards a Dynamic, Collaborative, and Comprehensive Instruction-Tuning Benchmark For SpeechabstractText language models have shown remarkable zero-shot capability in generalizing to unseen tasks when provided with well-formulated instructions. However, existing studies in speech processing primarily focus on limited or specific tasks. Moreover, the lack of standardized benchmarks hinders a fair comparison across different approaches. Thus, we present Dynamic-SUPERB, a benchmark designed for building universal speech models capable of leveraging instruction tuning to perform multiple tasks in a zero-shot fashion. To achieve comprehensive coverage of diverse speech tasks and harness instruction tuning, we invite the community to collaborate and contribute, facilitating the dynamic growth of the benchmark. To initiate, Dynamic-SUPERB features 55 evaluation instances by combining 33 tasks and 22 datasets. This spans a broad spectrum of dimensions, providing a comprehensive platform for evaluation. Additionally, we propose several approaches to establish benchmark baselines. These include the utilization of speech models, text language models, and the multimodal encoder. Evaluation results indicate that while these baselines perform reasonably on seen tasks, they struggle with unseen ones. We release all materials to the public and welcome researchers to collaborate on the project, advancing technologies in the field together1. Chien-Yu Huang, Ke-Han Lu, Shih-Heng Wang, Chi-Yuan Hsiao, Chun-Yi Kuan, Siddhant Arora, Kai-Wei Chang 0001, Jiatong Shi, Yifan Peng 0003, Roshan S. Sharma, Shinji Watanabe 0001, Bhiksha Raj, Shady Shehata, Hung-yi Lee |
ICASSP | 10 |
| 2024 | VoxtLM: Unified Decoder-Only Models for Consolidating Speech Recognition, Synthesis and Speech, Text Continuation TasksabstractWe propose a decoder-only language model, VoxtLM, that can perform four tasks: speech recognition, speech synthesis, text generation, and speech continuation. VoxtLM integrates text vocabulary with discrete speech tokens from self-supervised speech features and uses special tokens to enable multitask learning. Compared to a single-task model, VoxtLM exhibits a significant improvement in speech synthesis, with improvements in both speech intelligibility from 28.9 to 5.6 and objective quality from 2.68 to 3.90. VoxtLM also improves speech generation and speech recognition performance over the single-task counterpart. Further, VoxtLM is trained with publicly available data and training recipes and model checkpoints are open-sourced to make fully reproducible work. Soumi Maiti, Yifan Peng 0003, Shukjae Choi, Jee-Weon Jung, Xuankai Chang, Shinji Watanabe 0001 |
ICASSP | 2 |
| 2024 | Contextualized Automatic Speech Recognition With Attention-Based Bias Phrase Boosted Beam SearchabstractEnd-to-end (E2E) automatic speech recognition (ASR) methods exhibit remarkable performance. However, since the performance of such methods is intrinsically linked to the context present in the training data, E2E-ASR methods do not perform as desired for unseen user contexts (e.g., technical terms, personal names, and playlists). Thus, E2E-ASR methods must be easily contextualized by the user or developer. This paper proposes an attention-based contextual biasing method that can be customized using an editable phrase list (referred to as a bias list). The proposed method can be trained effectively by combining a bias phrase index loss and special tokens to detect the bias phrases in the input speech data. In addition, to improve the contextualization performance during inference further, we propose a bias phrase boosted (BPB) beam search algorithm based on the bias phrase index probability. Experimental results demonstrate that the proposed method consistently improves the word error rate and the character error rate of the target phrases in the bias list on both the Librispeech-960 (English) and our in-house (Japanese) dataset, respectively. Yui Sudo, Muhammad Shakeel 0001, Yosuke Fukumoto, Yifan Peng 0003, Shinji Watanabe 0001 |
ICASSP | 4 |
| 2024 | Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
Muhammad Shakeel 0001, Yui Sudo, Yifan Peng 0003, Shinji Watanabe 0001 |
INTERSPEECH | 3 |
| 2024 | OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer
Yifan Peng 0003, Jinchuan Tian, Siddhant Arora, Brian Yan, Yui Sudo, Muhammad Shakeel 0001, Kwanghee Choi, Jiatong Shi, Xuankai Chang, Jee-Weon Jung, Shinji Watanabe 0001 |
INTERSPEECH | 1 |
| 2024 | MULTI-CONVFORMER: Extending Conformer with Multiple Convolution Kernels
Darshan Prabhu, Yifan Peng 0003, Preethi Jyothi, Shinji Watanabe 0001 |
INTERSPEECH | 2 |
| 2024 | On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models
Jinchuan Tian, Yifan Peng 0003, Kwanghee Choi, Karen Livescu, Shinji Watanabe 0001 |
INTERSPEECH | 2 |
| 2024 | UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language InstructionsabstractSiddhant Arora, Hayato Futami, Jee-weon Jung, Yifan Peng, Roshan Sharma, Yosuke Kashiwagi, Emiru Tsunoo, Karen Livescu, Shinji Watanabe. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Siddhant Arora, Hayato Futami, Jee-Weon Jung, Yifan Peng 0003, Roshan S. Sharma, Yosuke Kashiwagi, Emiru Tsunoo, Karen Livescu, Shinji Watanabe 0001 |
NAACL-HLT | 4 |
| 2024 | ESPnet-EZ: Python-Only ESPnet For Easy Fine-Tuning And IntegrationabstractWe introduce ESPnet-EZ, an extension of the open-source speech processing toolkit ESPnet, aimed at quick and easy development of speech models. ESPnet-EZ focuses on two major aspects: (i) easy fine-tuning and inference of existing ESPnet models on various tasks and (ii) easy integration with popular deep neural network frameworks such as PyTorch-Lightning, Hugging Face transformers and datasets, and Lhotse. By replacing ESPnet design choices inherited from Kaldi with a Python-only, Bash-free interface, we dramatically reduce the effort required to build, debug, and use a new model. For example, to fine-tune a speech foundation model, ESPnet-EZ, compared to ESPnet, reduces the number of newly written code by 2.7x and the amount of dependent code by 6.7x while dramatically reducing the Bash script dependencies. The codebase of ESPnet-EZ is publicly available.1 Masao Someki, Kwanghee Choi, Siddhant Arora, Samuele Cornell, Jionghao Han, Yifan Peng 0003, Jiatong Shi, Vaibhav Srivastav, Shinji Watanabe 0001 |
SLT | 7 |
| 2024 | Contextualized Automatic Speech Recognition With Dynamic VocabularyabstractDeep biasing (DB) enhances the performance of end-to-end automatic speech recognition (E2E-ASR) models for rare words or contextual phrases using a bias list. However, most existing methods treat bias phrases as sequences of subwords in a predefined static vocabulary. This naive sequence decomposition produces unnatural token patterns, significantly lowering their occurrence probability. More advanced techniques address this problem by expanding the vocabulary with additional modules, including the external language model shallow fusion or rescoring. However, they result in increasing the workload due to the additional modules. This paper proposes a dynamic vocabulary where bias tokens can be added during inference. Each entry in a bias list is represented as a single token, unlike a sequence of existing subword tokens. This approach eliminates the need to learn subword dependencies within the bias phrases. This method is easily applied to various architectures because it only expands the embedding and output layers in common E2E-ASR architectures. Experimental results demonstrate that the proposed method improves the bias phrase WER on English and Japanese datasets by $3.1-4.9$ points compared with the conventional DB method. Yui Sudo, Yosuke Fukumoto, Muhammad Shakeel 0001, Yifan Peng 0003, Shinji Watanabe 0001 |
SLT | 4 |
| 2024 | Robust Audiovisual Speech Recognition Models with Mixture-of-ExpertsabstractVisual signals can enhance audiovisual speech recognition accuracy by providing additional contextual information. Given the complexity of visual signals, an audiovisual speech recognition model requires robust generalization capabilities across diverse video scenarios, presenting a significant challenge. In this paper, we introduce EVA, leveraging the mixture-of-Experts for audioVisual ASR to perform robust speech recognition for “in-the-wild” videos. Specifically, we first encode visual information into visual tokens sequence and map them into speech space by a lightweight projection. Then, we build EVA upon a robust pretrained speech recognition model, ensuring its generalization ability. Moreover, to incorporate visual information effectively, we inject visual information into the ASR model through a mixture-of-experts module. Experiments show our model achieves state-of-the-art results on three benchmarks, which demonstrates the generalization ability of EVA across diverse video domains. Yihan Wu 0008, Yifan Peng 0003, Xuankai Chang, Ruihua Song, Shinji Watanabe 0001 |
SLT | 2 |
| 2023 | Joint Prediction and Denoising for Large-Scale Multilingual Self-Supervised LearningabstractMultilingual self-supervised learning (SSL) has often lagged behind state-of-the-art (SOTA) methods due to the expenses and complexity required to handle many languages. This further harms the reproducibility of SSL, which is already limited to few research groups due to its resource usage. We show that more powerful techniques can actually lead to more efficient pre-training, opening SSL to more research groups. We propose WavLabLM, which extends WavLM’s joint prediction and denoising to 40k hours of data across 136 languages. To build WavLabLM, we devise a novel multi-stage pre-training method, designed to address the language imbalance of multilingual data. WavLabLM achieves comparable performance to XLS-R on ML-SUPERB with less than $10 \%$ of the training data, making SSL realizable with academic compute. We show that further efficiency can be achieved with a vanilla HuBERT Base model, which can maintain $94 \%$ of XLS-R’s performance with only $3 \%$ of the data, 4 GPUs, and limited trials. We open-source all code and models in ESPnet. Jiatong Shi, Brian Yan, Dan Berrebbi, Wangyou Zhang, Yifan Peng 0003, Xuankai Chang, Soumi Maiti, Shinji Watanabe 0001 |
ASRU | 6 |
| 2023 | Reproducing Whisper-Style Training Using An Open-Source Toolkit And Publicly Available DataabstractPre-training speech models on large volumes of data has achieved remarkable success. OpenAI Whisper is a multilingual multitask model trained on 680k hours of supervised speech data. It generalizes well to various speech recognition and translation benchmarks even in a zero-shot setup. However, the full pipeline for developing such models (from data collection to training) is not publicly accessible, which makes it difficult for researchers to further improve its performance and address training-related issues such as efficiency, robustness, fairness, and bias. This work presents an Open Whisper-style Speech Model (OWSM), which reproduces Whisperstyle training using an open-source toolkit and publicly available data. OWSM even supports more translation directions and can be more efficient to train. We will publicly release all scripts used for data preparation, training, inference, and scoring as well as pretrained models and training logs to promote open science.11https://github.com/espnet/espnet Yifan Peng 0003, Jinchuan Tian, Brian Yan, Dan Berrebbi, Xuankai Chang, Jiatong Shi, Siddhant Arora, Roshan S. Sharma, Wangyou Zhang, Yui Sudo, Muhammad Shakeel 0001, Jee-Weon Jung, Soumi Maiti, Shinji Watanabe 0001 |
ASRU | 1 |
| 2023 | A Study on the Integration of Pipeline and E2E SLU Systems for Spoken Semantic Parsing Toward Stop Quality ChallengeabstractRecently there have been efforts to introduce new benchmark tasks for spoken language understanding (SLU), like semantic parsing. In this paper, we describe our proposed spoken semantic parsing system for the quality track (Track 1) in Spoken Language Understanding Grand Challenge which is part of ICASSP Signal Processing Grand Challenge 2023. We experiment with both end-to-end and pipeline systems for this task. Strong automatic speech recognition (ASR) models like Whisper and pretrained Language models (LM) like BART are utilized inside our SLU framework to boost performance. We also investigate the output level combination of various models to get an exact match accuracy of 80.8, which won the 1st place at the challenge. Siddhant Arora, Hayato Futami, Shih-Lun Wu, Jessica Huynh, Yifan Peng 0003, Yosuke Kashiwagi, Emiru Tsunoo, Brian Yan, Shinji Watanabe 0001 |
ICASSP | 5 |
| 2023 | Improving Massively Multilingual ASR with Auxiliary CTC ObjectivesabstractMultilingual Automatic Speech Recognition (ASR) models have extended the usability of speech technologies to a wide variety of languages. With how many languages these models have to handle, however, a key to understanding their imbalanced performance across different languages is to examine if the model actually knows which language it should transcribe. In this paper, we introduce our work on improving performance on FLEURS, a 102-language open ASR benchmark, by conditioning the entire model on language identity (LID). We investigate techniques inspired from recent Connectionist Temporal Classification (CTC) studies to help the model handle the large number of languages, conditioning on the LID predictions of auxiliary tasks. Our experimental results demonstrate the effectiveness of our technique over standard CTC/Attention-based hybrid models. Furthermore, our state-of-the-art systems using self-supervised models with the Conformer architecture improve over the results of prior work on FLEURS by a relative 28.4% CER. Trained models are reproducible recipes are available at https://github.com/espnet/espnet/tree/master/egs2/fleurs/asr1. Brian Yan, Jiatong Shi, Yifan Peng 0003, Soumi Maiti, Shinji Watanabe 0001 |
ICASSP | 4 |
| 2023 | The Pipeline System of ASR and NLU with MLM-based data Augmentation Toward Stop Low-Resource ChallengeabstractThis paper describes our system for the low-resource domain adaptation track (Track 3) in Spoken Language Understanding Grand Challenge, which is a part of ICASSP Signal Processing Grand Challenge 2023. In the track, we adopt a pipeline approach of ASR and NLU. For ASR, we fine-tune Whisper for each domain with upsampling. For NLU, we fine-tune BART on all the Track3 data and then on low-resource domain data. We apply masked LM (MLM) -based data augmentation, where some of input tokens and corresponding target labels are replaced using MLM. We also apply a retrieval-based approach, where model input is augmented with similar training samples. As a result, we achieved exact match (EM) accuracy 63.3/75.0 (average: 69.15) for reminder/weather domain, and won the 1st place at the challenge. Hayato Futami, Jessica Huynh, Siddhant Arora, Shih-Lun Wu, Yosuke Kashiwagi, Yifan Peng 0003, Brian Yan, Emiru Tsunoo, Shinji Watanabe 0001 |
ICASSP | 6 |
| 2023 | E-Branchformer-Based E2E SLU Toward Stop on-Device ChallengeabstractIn this paper, we report our team’s study on track 2 of the Spoken Language Understanding Grand Challenge, which is a component of the ICASSP Signal Processing Grand Challenge 2023. The task is intended for on-device processing and involves estimating semantic parse labels from speech using a model with 15 million parameters. We use E2E E-Branchformer-based spoken language understanding model, which is more parameter controllable than cascade models, and reduced the parameter size through sequential distillation and tensor decomposition techniques. On the STOP dataset, we achieved an exact match accuracy of 70.9% under the tight constraint of 15 million parameters. Yosuke Kashiwagi, Siddhant Arora, Hayato Futami, Jessica Huynh, Shih-Lun Wu, Yifan Peng 0003, Brian Yan, Emiru Tsunoo, Shinji Watanabe 0001 |
ICASSP | 6 |
| 2023 | Speechlmscore: Evaluating Speech Generation Using Speech Language ModelabstractWhile human evaluation is the most reliable metric for evaluating speech generation systems, it is generally costly and time-consuming. Previous studies on automatic speech quality assessment address the problem by predicting human evaluation scores with machine learning models. However, they rely on supervised learning and thus suffer from high annotation costs and domain-shift problems. We propose SpeechLMScore, an unsupervised metric to evaluate generated speech using a speech language model. SpeechLMScore computes the average log-probability of a speech signal by mapping it into discrete tokens and measures the average probability of generating the sequence of tokens. Therefore, it does not require human annotation and is a highly scalable framework. Evaluation results demonstrate that the proposed metric shows a promising correlation with human evaluation scores on different speech generation tasks including voice conversion, text-to-speech, and speech enhancement. Soumi Maiti, Yifan Peng 0003, Takaaki Saeki, Shinji Watanabe 0001 |
ICASSP | 2 |
| 2023 | Structured Pruning of Self-Supervised Pre-Trained Models for Speech Recognition and UnderstandingabstractSelf-supervised speech representation learning (SSL) has shown to be effective in various downstream tasks, but SSL models are usually large and slow. Model compression techniques such as pruning aim to reduce the model size and computation without degradation in accuracy. Prior studies focus on the pruning of Transformers; however, speech models not only utilize a stack of Transformer blocks, but also combine a frontend network based on multiple convolutional layers for low-level feature representation learning. This frontend has a small size but a heavy computational cost. In this work, we propose three task-specific structured pruning methods to deal with such heterogeneous networks. Experiments on LibriSpeech and SLURP show that the proposed method is more accurate than the original wav2vec2-base with 10% to 30% less computation, and is able to reduce the computation by 40% to 50% without any degradation. Yifan Peng 0003, Kwangyoun Kim, Felix Wu, Prashant Sridhar, Shinji Watanabe 0001 |
ICASSP | 1 |
| 2023 | I3D: Transformer Architectures with Input-Dependent Dynamic Depth for Speech RecognitionabstractTransformer-based end-to-end speech recognition has achieved great success. However, the large footprint and computational overhead make it difficult to deploy these models in some real-world applications. Model compression techniques can reduce the model size and speed up inference, but the compressed model has a fixed architecture which might be suboptimal. We propose a novel Transformer encoder with Input-Dependent Dynamic Depth (I3D) to achieve strong performance-efficiency trade-offs. With a similar number of layers at inference time, I3D-based models outperform the vanilla Transformer and the static pruned model via iterative layer pruning. We also present interesting analysis on the gate probabilities and the input-dependency, which helps us better understand deep encoders. Yifan Peng 0003, Jaesong Lee, Shinji Watanabe 0001 |
ICASSP | 1 |
| 2023 | Reducing Barriers to Self-Supervised Learning: HuBERT Pre-training with Academic Compute
Xuankai Chang, Yifan Peng 0003, Zhaoheng Ni, Soumi Maiti, Shinji Watanabe 0001 |
INTERSPEECH | 3 |
| 2023 | Tensor decomposition for minimization of E2E SLU model toward on-device processing
Yosuke Kashiwagi, Siddhant Arora, Hayato Futami, Jessica Huynh, Shih-Lun Wu, Yifan Peng 0003, Brian Yan, Emiru Tsunoo, Shinji Watanabe 0001 |
INTERSPEECH | 6 |
| 2023 | A Comparative Study on E-Branchformer vs Conformer in Speech Recognition, Translation, and Understanding Tasks
Yifan Peng 0003, Kwangyoun Kim, Felix Wu, Brian Yan, Siddhant Arora, Jiyang Tang, Suwon Shon, Prashant Sridhar, Shinji Watanabe 0001 |
INTERSPEECH | 1 |
| 2023 | DPHuBERT: Joint Distillation and Pruning of Self-Supervised Speech Models
Yifan Peng 0003, Yui Sudo, Muhammad Shakeel 0001, Shinji Watanabe 0001 |
INTERSPEECH | 1 |
| 2023 | Time-synchronous one-pass Beam Search for Parallel Online and Offline Transducers with Dynamic Block Training
Yui Sudo, Muhammad Shakeel 0001, Yifan Peng 0003, Shinji Watanabe 0001 |
INTERSPEECH | 3 |
| 2022 | ESPnet-SLU: Advancing Spoken Language Understanding Through ESPnetabstractAs Automatic Speech Processing (ASR) systems are getting better, there is an increasing interest of using the ASR output to do downstream Natural Language Processing (NLP) tasks. However, there are few open source toolkits that can be used to generate reproducible results on different Spoken Language Understanding (SLU) benchmarks. Hence, there is a need to build an open source standard that can be used to have a faster start into SLU research. We present ESPnet-SLU, which is designed for quick development of spoken language understanding in a single framework. ESPnet-SLU is a project inside end-to-end speech processing toolkit, ESPnet, which is a widely used open-source standard for various speech processing tasks like ASR, Text to Speech (TTS) and Speech Translation (ST). We enhance the toolkit to provide implementations for various SLU benchmarks that enable researchers to seamlessly mix-and-match different ASR and NLU models. We also provide pretrained models with intensively tuned hyper-parameters that can match or even outperform the current state-of-the-art performances. The toolkit is publicly available at https://github.com/espnet/espnet. Siddhant Arora, Siddharth Dalmia, Pavel Denisov, Xuankai Chang, Yushi Ueda, Yifan Peng 0003, Yuekai Zhang, Sujay Kumar, Karthik Ganesan 0003, Brian Yan, Ngoc Thang Vu, Alan W. Black, Shinji Watanabe 0001 |
ICASSP | 6 |
| 2022 | Branchformer: Parallel MLP-Attention Architectures to Capture Local and Global Context for Speech Recognition and UnderstandingabstractConformer has proven to be effective in many speech processing tasks. It combines the benefits of extracting local dependencies using convolutions and global dependencies using self-attention. Inspired by this, we propose a more flexible, interpretable and customizable encoder alternative, Branchformer, with parallel branches for modeling various ranged dependencies in end-to-end speech processing. In each encoder layer, one branch employs self-attention or its variant to capture long-range dependencies, while the other branch utilizes an MLP module with convolutional gating (cgMLP) to extract local relationships. We conduct experiments on several speech recognition and spoken language understanding benchmarks. Results show that our model outperforms both Transformer and cgMLP. It also matches with or outperforms state-of-the-art results achieved by Conformer. Furthermore, we show various strategies to reduce computation thanks to the two-branch architecture, including the ability to have variable inference complexity in a single trained model. The weights learned for merging branches indicate how local and global dependencies are utilized in different layers, which benefits model designing. Yifan Peng 0003, Siddharth Dalmia, Ian Lane, Shinji Watanabe 0001 |
ICML | 1 |
| 2022 | Attention Weight Smoothing Using Prior Distributions for Transformer-Based End-to-End ASR
Takashi Maekaku, Yuya Fujita, Yifan Peng 0003, Shinji Watanabe 0001 |
INTERSPEECH | 3 |
| 2022 | E-Branchformer: Branchformer with Enhanced Merging for Speech RecognitionabstractConformer, combining convolution and self-attention sequentially to capture both local and global information, has shown remarkable performance and is currently regarded as the state-of-the-art for automatic speech recognition (ASR). Several other studies have explored integrating convolution and self-attention but they have not managed to match Conformer's performance. The recently introduced Branchformer achieves comparable performance to Conformer by using dedicated branches of convolution and self-attention and merging local and global context from each branch. In this paper, we propose E-Branchformer, which enhances Branchformer by applying an effective merging method and stacking additional point-wise modules. E-Branchformer sets new state-of-the-art word error rates (WERs) 1.81% and 3.65% on LibriSpeech test-clean and test-other sets without using any external training data. Kwangyoun Kim, Felix Wu, Yifan Peng 0003, Prashant Sridhar, Kyu Jeong Han, Shinji Watanabe 0001 |
SLT | 3 |
| 2022 | A Study on the Integration of Pre-Trained SSL, ASR, LM and SLU Models for Spoken Language UnderstandingabstractCollecting sufficient labeled data for spoken language understanding (SLU) is expensive and time-consuming. Recent studies achieved promising results by using pre-trained models in low-resource scenarios. Inspired by this, we aim to ask: which (if any) pre-training strategies can improve performance across SLU benchmarks? To answer this question, we employ four types of pre-trained models and their combinations for SLU. We leverage self-supervised speech and language models (LM) pre-trained on large quantities of un-paired data to extract strong speech and text representations. We also explore using supervised models pre-trained on larger external automatic speech recognition (ASR) or SLU corpora. We conduct extensive experiments on the SLU Evaluation (SLUE) benchmark and observe self-supervised pre-trained models to be more powerful, with pre-trained LM and speech models being most beneficial for the Sentiment Analysis and Named Entity Recognition task, respectively.11Our code and models will be publicly available as part of the ESPnet-SLU toolkit. Yifan Peng 0003, Siddhant Arora, Yosuke Higuchi, Yushi Ueda, Sujay Kumar, Karthik Ganesan 0003, Siddharth Dalmia, Xuankai Chang, Shinji Watanabe 0001 |
SLT | 1 |