EDBT 2026 Demo / reviewers in the wild / expert
Chen Chen 0075
dblp:65/4423-75
· DBLP profile ↗
43ranked-venue papers
12as first author
43since 2021 · last 2026
0000-0003-4181-9285ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 7 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 9 first-author · 28 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TextBFGS: A Case-Based Reasoning Approach to Code Optimization via Error-Operator Retrieval
Zizheng Zhang, Yuyang Liao, Chen Chen 0075, Dun Wu, Qianjin Yu, Yanqin Gao, Kailai Zhang, Chng Eng Siong, Xionghu Zhong |
ICCBR | 3 |
| 2026 | Chronological Thinking in Full-Duplex Spoken Dialogue Language ModelsabstractRecent advances in spoken dialogue language models (SDLMs) reflect growing interest in shifting from turn-based to full-duplex systems, where the models continuously perceive user speech streams while generating responses. This simultaneous listening and speaking design enables real-time interaction and the agent can handle dynamic conversational behaviors like user barge-in. However, during the listening phase, existing systems keep the agent idle by repeatedly predicting the silence token, which departs from human behavior: we usually engage in lightweight thinking during conversation rather than remaining absent-minded. Inspired by this, we propose Chronological Thinking, an on-the-fly conversational thinking mechanism that aims to improve response quality in full-duplex SDLMs. Specifically, chronological thinking presents a paradigm shift from conventional LLM thinking approaches, such as Chain-of-Thought, purpose-built for streaming acoustic input. (1) Strictly causal: the agent reasons incrementally while listening, updating internal hypotheses only from past audio with no lookahead. (2) No additional latency: reasoning is amortized during the listening window; once the user stops speaking, the agent halts thinking and begins speaking without further delay. Experiments demonstrate the effectiveness of chronological thinking through both objective metrics and human evaluations show consistent improvements in response quality. Furthermore, chronological thinking robustly handles conversational dynamics and attains competitive performance on full-duplex interaction metrics. Donghang Wu, Chen Chen 0075, Xuerui Yang, Gang Yu 0002, Hexin Liu, Nana Hou, Chng Eng Siong |
SIGDIAL | 3 |
| 2025 | Beyond Output Matching: Bidirectional Alignment for Enhanced In-Context LearningabstractChengwei Qin, Wenhan Xia, Fangkai Jiao, Chen Chen, Yuchen Hu, Bosheng Ding, Ruirui Chen, Shafiq Joty. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Chengwei Qin, Wenhan Xia, Fangkai Jiao, Chen Chen 0075, Bosheng Ding, Ruirui Chen 0002, Shafiq R. Joty |
ACL (1) | 4 |
| 2025 | Open Full-duplex Voice Agent with Speech-to-Speech Language ModelabstractWe present the system demonstration and opensource code release of a novel, data-efficient framework that converts any standard text Large Language Model (LLM) into a full-duplex end-to-end (E2E) speech-to-speech (S2S) model, for building conversational voice agents. Our new modeling method enables any LLMs to simultaneously listen and speak without requiring extensive speech-text pretraining. Moreover, we demonstrate how to put together a low-latency and full-duplex voice agent with open-source modeling, inference optimization, and serving solutions. This work significantly lowers the barrier to entry for developing low-latency, human-like voice agents by providing a generalizable, end-to-end solution built on open-source technologies. Edresson Casanova, Chen Chen 0075, Kevin Hu, Ankita Pasad, Elena Rastorgueva, Seelan Lakshmi Narasimhan, Slyne Deng, Ehsan Hosseini-Asl, Piotr Zelasko, Valentin Mendelev, Subhankar Ghosh, Yifan Peng 0003, Zhehuai Chen, Jason Li 0007, Jagadeesh Balam, Vitaly Lavrukhin, Boris Ginsburg |
ASRU | 2 |
| 2025 | SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and SynthesisabstractIn this paper, we introduce SSR-Speech, a neural codec autoregressive model designed for stable, safe, and robust zero-shot text-based speech editing and text-to-speech synthesis. SSR-Speech is built on a Transformer decoder and incorporates classifier-free guidance to enhance the stability of the generation process. A watermark Encodec is proposed to embed frame-level watermarks into the edited regions of the speech so that which parts were edited can be detected. In addition, the waveform reconstruction leverages the original unedited speech segments, providing superior recovery compared to the Encodec model. Our approach achieves state-of-the-art performance in the RealEdit speech editing task and the LibriTTS text-to-speech task, surpassing previous methods. Furthermore, SSR-Speech excels in multi-span speech editing and also demonstrates remarkable robustness to background sounds. The source code1and demos2are released. Helin Wang, Meng Yu 0003, Jiarui Hai, Chen Chen 0075, Rilin Chen, Najim Dehak, Dong Yu 0001 |
ICASSP | 4 |
| 2025 | Audio Large Language Models Can Be Descriptive Speech Quality EvaluatorsabstractAn ideal multimodal agent should be aware of the quality of its input modalities. Recent advances have enabled large language models (LLMs) to incorporate auditory systems for handling various speech-related tasks. However, most audio LLMs remain unaware of the quality of the speech they process. This limitation arises because speech quality evaluation is typically excluded from multi-task training due to the lack of suitable datasets. To address this, we introduce the first natural language-based speech evaluation corpus, generated from authentic human ratings. In addition to the overall Mean Opinion Score (MOS), this corpus offers detailed analysis across multiple dimensions and identifies causes of quality degradation. It also enables descriptive comparisons between two speech samples (A/B tests) with human-like judgment. Leveraging this corpus, we propose an alignment approach with LLM distillation (ALLD) to guide the audio LLM in extracting relevant information from raw speech and generating meaningful responses. Experimental results demonstrate that ALLD outperforms the previous state-of-the-art regression model in MOS prediction, with a mean square error of 0.17 and an A/B test accuracy of 98.6%. Additionally, the generated responses achieve BLEU scores of 25.8 and 30.2 on two tasks, surpassing the capabilities of task-specific models. This work advances the comprehensive perception of speech signals by audio LLMs, contributing to the development of real-world auditory and sensory intelligent agents. Chen Chen 0075, Siyin Wang, Helin Wang, Zhehuai Chen, Chao Zhang 0031, Chao-Han Huck Yang, Chng Eng Siong |
ICLR | 1 |
| 2025 | GenSE: Generative Speech Enhancement via Language Models using Hierarchical ModelingabstractSemantic information refers to the meaning conveyed through words, phrases, and contextual relationships within a given linguistic structure. Humans can leverage semantic information, such as familiar linguistic patterns and contextual cues, to reconstruct incomplete or masked speech signals in noisy environments. However, existing speech enhancement (SE) approaches often overlook the rich semantic information embedded in speech, which is crucial for improving intelligibility, speaker consistency, and overall quality of enhanced speech signals. To enrich the SE model with semantic information, we employ language models as an efficient semantic learner and propose a comprehensive framework tailored for language model-based speech enhancement, called GenSE. Specifically, we approach SE as a conditional language modeling task rather than a continuous signal regression problem defined in existing works. This is achieved by tokenizing speech signals into semantic tokens using a pre-trained self-supervised model and into acoustic tokens using a custom-designed single-quantizer neural codec model. To improve the stability of language model predictions, we propose a hierarchical modeling method that decouples the generation of clean semantic tokens and clean acoustic tokens into two distinct stages. Moreover, we introduce a token chain prompting mechanism during the acoustic token generation stage to ensure timbre consistency throughout the speech enhancement process. Experimental results on benchmark datasets demonstrate that our proposed approach outperforms state-of-the-art SE systems in terms of speech quality and generalization capability. Codes and demos are publicly available at https://anonymous.4open.science/w/gen-se-7F52/. Jixun Yao, Hexin Liu, Chen Chen 0075, Chng Eng Siong, Lei Xie 0001 |
ICLR | 3 |
| 2025 | Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
Ehsan Hosseini-Asl, Chen Chen 0075, Edresson Casanova, Subhankar Ghosh, Piotr Zelasko, Zhehuai Chen, Jason Li 0007, Jagadeesh Balam, Boris Ginsburg |
INTERSPEECH | 3 |
| 2025 | From KAN to GR-KAN: Advancing Speech Enhancement with KAN-Based MethodologyabstractDeep neural network (DNN)-based speech enhancement (SE) usually uses conventional activation functions, which lack the expressiveness to capture complex multiscale structures needed for high-fidelity SE. Group-Rational KAN (GR-KAN), a variant of Kolmogorov-Arnold Networks (KAN), retains KAN's expressiveness while improving scalability on complex tasks. We adapt GR-KAN to existing DNN-based SE by replacing dense layers with GR-KAN layers in the time-frequency (T-F) domain MP-SENet and adapting GR-KAN's activations into the 1D CNN layers in the time-domain Demucs. Results on Voicebank-DEMAND show that GR-KAN requires up to 4× fewer parameters while improving PESQ by up to 0.1. In contrast, KAN, facing scalability issues, outperforms MLP on a small-scale signal modeling task but fails to improve MP-SENet. We demonstrate the first successful use of KAN-based methods for consistent improvement in both time- and SoTA TF-domain SE, establishing GR-KAN as a promising alternative for SE. Haoyang Li 0018, Chen Chen 0075, Sabato Marco Siniscalchi, Songting Liu, Chng Eng Siong |
INTERSPEECH | 3 |
| 2025 | CNVSRC 2024: The Second Chinese Continuous Visual Speech Recognition Challenge
Zehua Liu, Xiaolou Li, Chen Chen 0075, Lantian Li, Dong Wang 0013 |
INTERSPEECH | 3 |
| 2024 | GenTranslate: Large Language Models are Generative Multilingual Speech and Machine TranslatorsabstractYuchen Hu, Chen Chen, Chao-Han Huck Yang, Ruizhe Li, Dong Zhang, Zhehuai Chen, Eng Siong Chng. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Chen Chen 0075, Chao-Han Huck Yang, Ruizhe Li 0001, Zhehuai Chen, Chng Eng Siong |
ACL (1) | 2 |
| 2024 | Noise-Aware Speech Separation with Contrastive LearningabstractRecently, speech separation (SS) task has achieved remarkable progress driven by deep learning technique. However, it is still challenging to separate target speech from noisy mixture, as the neural model is vulnerable to assign background noise to each speaker. In this paper, we propose a noise-aware SS (NASS) method, which aims to improve the speech quality for separated signals under noisy conditions. Specifically, NASS views background noise as an additional output and predicts it along with other speakers in a mask-based manner. To effectively denoise, we introduce patch-wise contrastive learning (PCL) between noise and speaker representations from the decoder input and encoder output. PCL loss aims to minimize the mutual information between predicted noise and other speakers at multiple-patch level to suppress the noise information in separated signals. Experimental results show that NASS achieves 1 to 2dB SI-SNRi or SDRi over DPRNN and Sepformer on WHAM! and LibriMix noisy datasets, with less than 0.1M parameter increase. Zizheng Zhang, Chen Chen 0075, Hsin-Hung Chen, Chng Eng Siong |
ICASSP | 2 |
| 2024 | Cross-Modality and Within-Modality Regularization for Audio-Visual Deepfake DetectionabstractAudio-visual deepfake detection scrutinizes manipulations in public video using complementary multimodal cues. Current methods, which train on fused multimodal data for multimodal targets face challenges due to uncertainties and inconsistencies in learned representations caused by independent modality manipulations in deepfake videos. To address this, we propose cross-modality and within-modality regularization to preserve modality distinctions during multimodal representation learning. Our approach includes an audio-visual transformer module for modality correspondence and a cross-modality regularization module to align paired audio-visual signals, preserving modality distinctions. Simultaneously, a within-modality regularization module refines unimodal representations with modality-specific targets to retain modal-specific details. Experimental results on the public audio-visual dataset, FakeAVCeleb, demonstrate the effectiveness and competitiveness of our approach. Heqing Zou, Meng Shen 0002, Chen Chen 0075, Chng Eng Siong, Deepu Rajan |
ICASSP | 4 |
| 2024 | It's Never Too Late: Fusing Acoustic Information into Large Language Models for Automatic Speech RecognitionabstractRecent studies have successfully shown that large language models (LLMs) can be successfully used for generative error correction (GER) on top of the automatic speech recognition (ASR) output. Specifically, an LLM is utilized to carry out a direct mapping from the N-best hypotheses list generated by an ASR system to the predicted output transcription. However, despite its effectiveness, GER introduces extra data uncertainty since the LLM is trained without taking into account acoustic information available in the speech signal. In this work, we aim to overcome such a limitation by infusing acoustic information before generating the predicted transcription through a novel late fusion solution termed Uncertainty-Aware Dynamic Fusion (UADF). UADF is a multimodal fusion approach implemented into an auto-regressive decoding process and works in two stages: (i) It first analyzes and calibrates the token-level LLM decision, and (ii) it then dynamically assimilates the information from the acoustic modality. Experimental evidence collected from various ASR tasks shows that UADF surpasses existing fusion mechanisms in several ways. It yields significant improvements in word error rate (WER) while mitigating data uncertainty issues in LLM and addressing the poor generalization relied with sole modality during fusion. We also demonstrate that UADF seamlessly adapts to audio-visual speech recognition. Chen Chen 0075, Ruizhe Li 0001, Sabato Marco Siniscalchi, Chng Eng Siong, Chao-Han Huck Yang |
ICLR | 1 |
| 2024 | Large Language Models are Efficient Learners of Noise-Robust Speech RecognitionabstractRecent advances in large language models (LLMs) have promoted generative error correction (GER) for automatic speech recognition (ASR), which leverages the rich linguistic knowledge and powerful reasoning ability of LLMs to improve recognition results. The latest work proposes a GER benchmark with "HyPoradise" dataset to learn the mapping from ASR N-best hypotheses to ground-truth transcription by efficient LLM finetuning, which shows great effectiveness but lacks specificity on noise-robust ASR. In this work, we extend the benchmark to noisy conditions and investigate if we can teach LLMs to perform denoising for GER just like what robust ASR do, where one solution is introducing noise information as a conditioner into LLM. However, directly incorporating noise embeddings from audio encoder could harm the LLM tuning due to cross-modality gap. To this end, we propose to extract a language-space noise embedding from the N-best list to represent the noise conditions of source speech, which can promote the denoising process in GER. Furthermore, in order to enhance its representation ability of audio noise, we design a knowledge distillation (KD) approach via mutual information estimation to distill the real noise information in audio embeddings to our language embedding. Experiments on various latest LLMs demonstrate our approach achieves a new breakthrough with up to 53.9% correction improvement in terms of word error rate while with limited training data. Analysis shows that our language-space noise embedding can well represent the noise conditions of source speech, under which off-the-shelf LLMs show strong ability of language-space denoising. Chen Chen 0075, Chao-Han Huck Yang, Ruizhe Li 0001, Chao Zhang 0031, Chng Eng Siong |
ICLR | 2 |
| 2024 | CNVSRC 2023: The First Chinese Continuous Visual Speech Recognition Challenge
Chen Chen 0075, Zehua Liu, Xiaolou Li, Lantian Li, Dong Wang 0013 |
INTERSPEECH | 1 |
| 2024 | Noise-aware Speech Enhancement using Diffusion Probabilistic ModelabstractWith recent advances of diffusion model, generative speech enhancement (SE) has attracted a surge of research interest due to its great potential for unseen testing noises. However, existing efforts mainly focus on inherent properties of clean speech, underexploiting the varying noise information in real world. In this paper, we propose a noise-aware speech enhancement (NASE) approach that extracts noise-specific information to guide the reverse process in diffusion model. Specifically, we design a noise classification (NC) model to produce acoustic embedding as a noise conditioner to guide the reverse denoising process. Meanwhile, a multi-task learning scheme is devised to jointly optimize SE and NC tasks to enhance the noise specificity of conditioner. NASE is shown to be a plug-and-play module that can be generalized to any diffusion SE models. Experiments on VB-DEMAND dataset show that NASE effectively improves multiple mainstream diffusion SE models, especially on unseen noises. Chen Chen 0075, Ruizhe Li 0001, Qiushi Zhu, Chng Eng Siong |
INTERSPEECH | 2 |
| 2024 | Investigating ASR Error Correction with Large Language Model and Multilingual 1-best Hypotheses
Sheng Li 0010, Chen Chen 0075, Kwok Chin Yuen, Chenhui Chu, Chng Eng Siong, Hisashi Kawai |
INTERSPEECH | 2 |
| 2024 | Zero-Shot Fake Video Detection by Audio-Visual Consistency
Xiaolou Li, Zehua Liu, Chen Chen 0075, Lantian Li, Li Guo 0004, Dong Wang 0013 |
INTERSPEECH | 3 |
| 2024 | Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation ModelsabstractWe propose an unsupervised adaptation framework, Self-TAught Recognizer (STAR), which leverages unlabeled data to enhance the robustness of automatic speech recognition (ASR) systems in diverse target domains, such as noise and accents. STAR is developed for prevalent speech foundation models based on Transformer-related architecture with auto-regressive decoding (e.g., Whisper, Canary). Specifically, we propose a novel indicator that empirically integrates step-wise information during decoding to assess the token-level quality of pseudo labels without ground truth, thereby guiding model updates for effective unsupervised adaptation. Experimental results show that STAR achieves an average of 13.5% relative reduction in word error rate across 14 target domains, and it sometimes even approaches the upper-bound performance of supervised adaptation. Surprisingly, we also observe that STAR prevents the adapted model from the common catastrophic forgetting problem without recalling source-domain data. Furthermore, STAR exhibits high data efficiency that only requires less than one-hour unlabeled data, and seamless generality to alternative large speech models and speech translation tasks. Our code aims to open source to the research communities. Chen Chen 0075, Chao-Han Yang, Chengwei Qin, Chng Eng Siong, Chao Zhang 0031 |
NeurIPS | 2 |
| 2024 | Large Language Model Based Generative Error Correction: A Challenge and Baselines For Speech Recognition, Speaker Tagging, and Emotion RecognitionabstractGiven recent advances in generative AI technology, a key question is how large language models (LLMs) can enhance acoustic modeling tasks using text decoding results from a frozen, pretrained automatic speech recognition (ASR) model. To explore new capabilities in language modeling for speech processing, we introduce the generative speech transcription error correction (GenSEC) challenge. This challenge comprises three post-ASR language modeling tasks: (i) post-ASR transcription correction, (ii) speaker tagging, and (iii) emotion recognition. These tasks aim to emulate future LLM-based agents handling voice-based interfaces while remaining accessible to a broad audience by utilizing open pretrained language models or agent-based APIs. We also discuss insights from baseline evaluations, as well as lessons learned for designing future evaluations. Chao-Han Huck Yang, Taejin Park, Yuan Gong 0001, Yuanchao Li, Zhehuai Chen, Chen Chen 0075, Kunal Dhawan, Piotr Zelasko, Chao Zhang 0031, Yun-Nung Chen, Yu Tsao 0001, Jagadeesh Balam, Boris Ginsburg, Sabato Marco Siniscalchi, Chng Eng Siong, Peter Bell 0001, Catherine Lai, Shinji Watanabe 0001, Andreas Stolcke |
SLT | 7 |
| 2024 | Wav2code: Restore Clean Speech Representations via Codebook Lookup for Noise-Robust ASRabstractAutomatic speech recognition (ASR) has gained remarkable successes thanks to recent advances of deep learning, but it usually degrades significantly under real-world noisy conditions. Recent works introduce speech enhancement (SE) as front-end to improve speech quality, which is proved effective but may not be optimal for downstream ASR due to speech distortion problem. Based on that, latest works combine SE and currently popular self-supervised learning (SSL) to alleviate distortion and improve noise robustness. Despite the effectiveness, the speech distortion caused by conventional SE still cannot be cleared out. In this paper, we propose a self-supervised framework named Wav2code to implement a feature-level SE with reduced distortions for noise-robust ASR. First, in pre-training stage the clean speech representations from SSL model are sent to lookup a discrete codebook via nearest-neighbor feature matching, the resulted code sequence are then exploited to reconstruct the original clean representations, in order to store them in codebook as prior. Second, during finetuning we propose a Transformer-based code predictor to accurately predict clean codes by modeling global and local dependency of input noisy representations, which enables discovery and restoration of high-quality clean representations with reduced distortions. Furthermore, we propose an interactive feature fusion network to combine original noisy and the restored clean representations to consider both fidelity and quality, resulting in more informative features for downstream ASR. Finally, experiments on both synthetic and real noisy datasets demonstrate that Wav2code can solve the speech distortion and improve ASR performance under various noisy conditions, resulting in stronger robustness. Chen Chen 0075, Qiushi Zhu, Chng Eng Siong |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Leveraging Modality-Specific Representations for Audio-Visual Speech Recognition via Reinforcement LearningabstractAudio-visual speech recognition (AVSR) has gained remarkable success for ameliorating the noise-robustness of speech recognition. Mainstream methods focus on fusing audio and visual inputs to obtain modality-invariant representations. However, such representations are prone to over-reliance on audio modality as it is much easier to recognize than video modality in clean conditions. As a result, the AVSR model underestimates the importance of visual stream in face of noise corruption. To this end, we leverage visual modality-specific representations to provide stable complementary information for the AVSR task. Specifically, we propose a reinforcement learning (RL) based framework called MSRL, where the agent dynamically harmonizes modality-invariant and modality-specific representations in the auto-regressive decoding process. We customize a reward function directly related to task-specific metrics (i.e., word error rate), which encourages the MSRL to effectively explore the optimal integration strategy. Experimental results on the LRS3 dataset show that the proposed method achieves state-of-the-art in both clean and various noisy conditions. Furthermore, we demonstrate the better generality of MSRL system than other baselines when test set contains unseen noises. Chen Chen 0075, Heqing Zou, Beier Zhu, Chng Eng Siong |
AAAI | 1 |
| 2023 | MIR-GAN: Refining Frame-Level Modality-Invariant Representations with Adversarial Network for Audio-Visual Speech RecognitionabstractAudio-visual speech recognition (AVSR) attracts a surge of research interest recently by leveraging multimodal signals to understand human speech.Mainstream approaches addressing this task have developed sophisticated architectures and techniques for multi-modality fusion and representation learning.However, the natural heterogeneity of different modalities causes distribution gap between their representations, making it challenging to fuse them.In this paper, we aim to learn the shared representations across modalities to bridge their gap.Different from existing similar methods on other multimodal tasks like sentiment analysis, we focus on the temporal contextual dependencies considering the sequence-to-sequence task setting of AVSR.In particular, we propose an adversarial network to refine framelevel modality-invariant representations (MIR-GAN), which captures the commonality across modalities to ease the subsequent multimodal fusion process.Extensive experiments on public benchmarks LRS3 and LRS2 show that our approach outperforms the state-of-the-arts 1 . Chen Chen 0075, Ruizhe Li 0001, Heqing Zou, Chng Eng Siong |
ACL (1) | 2 |
| 2023 | Hearing Lips in Noise: Universal Viseme-Phoneme Mapping and Transfer for Robust Audio-Visual Speech RecognitionabstractAudio-visual speech recognition (AVSR) provides a promising solution to ameliorate the noise-robustness of audio-only speech recognition with visual information.However, most existing efforts still focus on audio modality to improve robustness considering its dominance in AVSR task, with noise adaptation techniques such as front-end denoise processing.Though effective, these methods are usually faced with two practical challenges: 1) lack of sufficient labeled noisy audio-visual training data in some real-world scenarios and 2) less optimal model generality to unseen testing noises.In this work, we investigate the noiseinvariant visual modality to strengthen robustness of AVSR, which can adapt to any testing noises while without dependence on noisy training data, a.k.a., unsupervised noise adaptation.Inspired by human perception mechanism, we propose a universal viseme-phoneme mapping (UniVPM) approach to implement modality transfer, which can restore clean audio from visual signals to enable speech recognition under any noisy conditions.Extensive experiments on public benchmarks LRS3 and LRS2 show that our approach achieves the state-of-the-art under various noisy as well as clean conditions.In addition, we also outperform previous stateof-the-arts on visual speech recognition task 1 . Ruizhe Li 0001, Chen Chen 0075, Chengwei Qin, Qiushi Zhu, Chng Eng Siong |
ACL (1) | 3 |
| 2023 | Lifelong Sequence Generation with Dynamic Module Expansion and AdaptationabstractLifelong sequence generation (LSG), a problem in continual learning, aims to continually train a model on a sequence of generation tasks to learn constantly emerging new generation patterns while avoiding the forgetting of previous knowledge.Existing LSG methods mainly focus on maintaining old knowledge while paying little attention to knowledge transfer across tasks.In contrast, humans can better learn new tasks by leveraging previously acquired knowledge from similar tasks.Inspired by the learning paradigm of humans, we propose Dynamic Module Expansion and Adaptation (DMEA), which enables the model to dynamically determine the architecture for acquiring new knowledge based on task correlation and select the most similar previous tasks to facilitate adaptation to new tasks.In addition, as the learning process can easily be biased towards the current task which might cause more severe forgetting of previously learned knowledge, we propose dynamic gradient scaling to balance the learning of the current task and replayed tasks.With extensive experiments, we demonstrate that DMEA can consistently outperform existing methods in different LSG settings. Chengwei Qin, Chen Chen 0075, Shafiq R. Joty |
EMNLP | 2 |
| 2023 | Metric-Oriented Speech Enhancement Using Diffusion Probabilistic ModelabstractDeep neural network based speech enhancement technique focuses on learning a noisy-to-clean transformation supervised by paired training data. However, the task-specific evaluation metric (e.g., PESQ) is usually non-differentiable and can not be directly constructed in the training criteria. This mismatch between the training objective and evaluation metric likely results in sub-optimal performance. To alleviate it, we propose a metric-oriented speech enhancement method (MOSE), which leverages the recent advances in the diffusion probabilistic model and integrates a metric-oriented training strategy into its reverse process. Specifically, we design an actor-critic based framework that considers the evaluation metric as a posterior reward, thus guiding the reverse process to the metric-increasing direction. The experimental results demonstrate that MOSE obviously benefits from metric-oriented training and surpasses the generative baselines in terms of all evaluation metrics. Chen Chen 0075, Weiwei Weng, Chng Eng Siong |
ICASSP | 1 |
| 2023 | Unsupervised Noise Adaptation Using Data SimulationabstractDeep neural network based speech enhancement approaches aim to learn a noisy-to-clean transformation using a supervised learning paradigm. However, such a trained-well transformation is vulnerable to unseen noises that are not included in training set. In this work, we focus on the unsupervised noise adaptation problem in speech enhancement, where the ground truth of target domain data is completely unavailable. Specifically, we propose a generative adversarial network based method to efficiently learn a converse clean-to-noisy transformation using a few minutes of unpaired target domain data. Then this transformation is utilized to generate sufficient simulated data for domain adaptation of the enhancement model. Experimental results show that our method effectively mitigates the domain mismatch between training and test sets, and surpasses the best baseline by a large margin. Chen Chen 0075, Heqing Zou, Linhui Sun, Chng Eng Siong |
ICASSP | 1 |
| 2023 | CN-CVS: A Mandarin Audio-Visual Dataset for Large Vocabulary Continuous Visual to Speech SynthesisabstractResearch on Video to Speech Synthesis (VTS) surges recently and the focus is gradually shifting from small-vocabulary short-phrase VTS to large-vocabulary continuous VTS (LVC-VTS). A large-scale dataset with sufficient speakers and utterances is a prerequisite for such research, and the database is certainly language dependent.In this paper, we introduce CN-CVS, a large-scale Mandarin continuous visual-speech dataset, to support LVC-VTS research. The dataset contains about 200k utterances from more than 2500 individuals, amounting to more than 300 hours of visual-speech data. We built a state-of-the-art VTS model with the new dataset and conducted preliminary studies. Our results show that models that achieve good performance on small vocabulary tasks may perform very poor on CN-CVS, indicating that continuous VTS is indeed a challenging task, and the main challenge comes from the unconstrained vocabulary. The dataset and baseline code can be downloaded for free from http://cncvs.cslt.org. Chen Chen 0075, Dong Wang 0013, Thomas Fang Zheng |
ICASSP | 1 |
| 2023 | Gradient Remedy for Multi-Task Learning in End-to-End Noise-Robust Speech RecognitionabstractSpeech enhancement (SE) is proved effective in reducing noise from noisy speech signals for downstream automatic speech recognition (ASR), where multi-task learning strategy is employed to jointly optimize these two tasks. However, the enhanced speech learned by SE objective may not always yield good ASR results. From the optimization view, there sometimes exists interference between the gradients of SE and ASR tasks, which could hinder the multi-task learning and finally lead to sub-optimal ASR performance. In this paper, we propose a simple yet effective approach called gradient remedy (GR) to solve interference between task gradients in noise-robust speech recognition, from perspectives of both angle and magnitude. Specifically, we first project the SE task's gradient onto a dynamic surface that is at acute angle to ASR gradient, in order to remove the conflict between them and assist in ASR optimization. Furthermore, we adaptively rescale the magnitude of two gradients to prevent the dominant ASR task from being misled by SE gradient. Experimental results show that the proposed approach well resolves the gradient interference and achieves relative word error rate (WER) reductions of 9.3% and 11.1% over multi-task learning baseline, on RATS and CHiME-4 datasets, respectively. Our code is available at GitHub1. Chen Chen 0075, Ruizhe Li 0001, Qiushi Zhu, Chng Eng Siong |
ICASSP | 2 |
| 2023 | Unifying Speech Enhancement and Separation with Gradient Modulation for End-to-End Noise-Robust Speech SeparationabstractRecent studies in neural network-based monaural speech separation (SS) have achieved a remarkable success thanks to increasing ability of long sequence modeling. However, they would degrade significantly when put under realistic noisy conditions, as the background noise could be mistaken for speaker’s speech and thus interfere with the separated sources. To alleviate this problem, we propose a novel network to unify speech enhancement and separation with gradient modulation to improve noise-robustness. Specifically, we first build a unified network by combining speech enhancement (SE) and separation modules, with multi-task learning for optimization, where SE is supervised by parallel clean mixture to reduce noise for downstream speech separation. Furthermore, in order to avoid suppressing valid speaker information when reducing noise, we propose a gradient modulation (GM) strategy to harmonize the SE and SS tasks from optimization view. Experimental results show that our approach achieves the state-of-the-art on large-scale Libri2Mix- and Libri3Mix-noisy datasets, with SI-SNRi results of 16.0 dB and 15.8 dB respectively. Our code is available at GitHub1. Chen Chen 0075, Heqing Zou, Xionghu Zhong, Chng Eng Siong |
ICASSP | 2 |
| 2023 | Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech RecognitionabstractAudio-visual speech recognition (AVSR) research has gained a great success recently by improving the noise-robustness of audio-only automatic speech recognition (ASR) with noise-invariant visual information. However, most existing AVSR approaches simply fuse the audio and visual features by concatenation, without explicit interactions to capture the deep correlations between them, which results in sub-optimal multimodal representations for downstream speech recognition task. In this paper, we propose a cross-modal global interaction and local alignment (GILA) approach for AVSR, which captures the deep audio-visual (A-V) correlations from both global and local perspectives. Specifically, we design a global interaction model to capture the A-V complementary relationship on modality level, as well as a local alignment approach to model the A-V temporal consistency on frame level. Such a holistic view of cross-modal correlations enable better multimodal representations for AVSR. Experiments on public benchmarks LRS3 and LRS2 show that our GILA outperforms the supervised learning state-of-the-art. Code is at https://github.com/YUCHEN005/GILA. Ruizhe Li 0001, Chen Chen 0075, Heqing Zou, Qiushi Zhu, Chng Eng Siong |
IJCAI | 3 |
| 2023 | A Neural State-Space Modeling Approach to Efficient Speech Separation
Chen Chen 0075, Chao-Han Huck Yang, Pin-Jui Ku, Chng Eng Siong |
INTERSPEECH | 1 |
| 2023 | Dual-Path Style Learning for End-to-End Noise-Robust Speech Recognition
Nana Hou, Chen Chen 0075, Chng Eng Siong |
INTERSPEECH | 3 |
| 2023 | CN-Celeb-AV: A Multi-Genre Audio-Visual Dataset for Person Recognition
Lantian Li, Xiaolou Li, Chen Chen 0075, Ruihai Hou, Dong Wang 0013 |
INTERSPEECH | 4 |
| 2023 | HyPoradise: An Open Baseline for Generative Speech Recognition with Large Language ModelsabstractAdvancements in deep neural networks have allowed automatic speech recognition (ASR) systems to attain human parity on several publicly available clean speech datasets. However, even state-of-the-art ASR systems experience performance degradation when confronted with adverse conditions, as a well-trained acoustic model is sensitive to variations in the speech domain, e.g., background noise. Intuitively, humans address this issue by relying on their linguistic knowledge: the meaning of ambiguous spoken terms is usually inferred from contextual cues thereby reducing the dependency on the auditory system. Inspired by this observation, we introduce the first open-source benchmark to utilize external large language models (LLMs) for ASR error correction, where N-best decoding hypotheses provide informative elements for true transcription prediction. This approach is a paradigm shift from the traditional language model rescoring strategy that can only select one candidate hypothesis as output transcription. The proposed benchmark contains a novel dataset, "HyPoradise" (HP), encompassing more than 316,000 pairs of N-best hypotheses and corresponding accurate transcriptions across prevalent speech domains. Given this dataset, we examine three types of error correction techniques based on LLMs with varying amounts of labeled hypotheses-transcription pairs, which gains significant word error rate (WER) reduction. Experimental evidence demonstrates the proposed technique achieves a breakthrough by surpassing the upper bound of traditional re-ranking based methods. More surprisingly, LLM with reasonable prompt design can even correct those tokens that are missing in N-best list. We make our results publicly accessible for reproducible pipelines with released pre-trained models, thus providing a new paradigm for ASR error correction with LLMs. Chen Chen 0075, Chao-Han Huck Yang, Sabato Marco Siniscalchi, Chng Eng Siong |
NeurIPS | 1 |
| 2023 | Random Cycle Loss and Its Application to Voice ConversionabstractSpeech disentanglement aims to decompose independent causal factors of speech signals into separate codes. Perfect disentanglement benefits to a broad range of speech processing tasks. This paper presents a simple but effective disentanglement approach based on cycle consistency loss and random factor substitution. This leads to a novel random cycle (RC) loss that enforces analysis-and-resynthesis consistency, a main principle of reductionism. We theoretically demonstrate that the proposed RC loss can achieve independent codes if well optimized, which in turn leads to superior disentanglement when combined with information bottleneck (IB). Extensive simulation experiments were conducted to understand the properties of the RC loss, and experimental results on voice conversion further demonstrate the practical merit of the proposal. Source code and audio samples can be found on the webpage http://rc.cslt.org. Dong Wang 0013, Lantian Li, Chen Chen 0075, Thomas Fang Zheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Self-Critical Sequence Training for Automatic Speech RecognitionabstractAlthough automatic speech recognition (ASR) task has gained remarkable success by sequence-to-sequence models, there are two main mismatches between its training and testing that might lead to performance degradation: 1) The typically used cross-entropy criterion aims to maximize log-likelihood of the training data, while the performance is evaluated by word error rate (WER), not log-likelihood; 2) The teacher-forcing method leads to the dependence on ground truth during training, which means that model has never been exposed to its own prediction before testing. In this paper, we propose an optimization method called self-critical sequence training (SCST) to make the training procedure much closer to the testing phase. As a reinforcement learning (RL) based method, SCST utilizes a customized reward function to associate the training criterion and WER. Furthermore, it removes the reliance on teacher-forcing and harmonizes the model with respect to its inference procedure. We conducted experiments on both clean and noisy speech datasets, and the results show that the proposed SCST respectively achieves 8.7% and 7.8% relative improvements over the baseline in terms of WER. Chen Chen 0075, Nana Hou, Xiaofeng Qi, Heqing Zou, Chng Eng Siong |
ICASSP | 1 |
| 2022 | Noise-Robust Speech Recognition With 10 Minutes Unparalleled In-Domain DataabstractNoise-robust speech recognition systems require large amounts of training data including noisy speech data and corresponding transcripts to achieve state-of-the-art performances in face of various practical environments. However, such plenty of in-domain data is not always available in the real-life world. In this paper, we propose a generative adversarial network to simulate noisy spectrum from the clean spectrum (SimuGAN), where only 10 minutes of unparalleled in-domain noisy speech data is required as labels. Furthermore, we also propose a dual-path speech recognition system to improve the robustness of the system under noisy conditions. Experimental results show that the proposed speech recognition system achieves 7.3% absolute improvement with simulated noisy data by Simu-GAN over the best baseline in terms of word error rate (WER). Chen Chen 0075, Nana Hou, Shashank Shirol, Chng Eng Siong |
ICASSP | 1 |
| 2022 | Interactive Feature Fusion for End-to-End Noise-Robust Speech RecognitionabstractSpeech enhancement (SE) aims to suppress the additive noise from noisy speech signals to improve the speech’s perceptual quality and intelligibility. However, the over-suppression phenomenon in the enhanced speech might degrade the performance of downstream automatic speech recognition (ASR) task due to the missing latent information. To alleviate such problem, we propose an interactive feature fusion network (IFF-Net) for noise-robust speech recognition to learn complementary information from the enhanced feature and original noisy feature. Experimental results show that the proposed method achieves absolute word error rate (WER) reduction of 4.1% over the best baseline on RATS Channel-A corpus. Our further analysis indicates that the proposed IFF-Net can complement some missing information in the over-suppressed enhanced feature. Nana Hou, Chen Chen 0075, Chng Eng Siong |
ICASSP | 3 |
| 2022 | Speech Emotion Recognition with Co-Attention Based Multi-Level Acoustic InformationabstractSpeech Emotion Recognition (SER) aims to help the machine to understand human’s subjective emotion from only audio in-formation. However, extracting and utilizing comprehensive in-depth audio information is still a challenging task. In this paper, we propose an end-to-end speech emotion recognition system using multi-level acoustic information with a newly designed co-attention module. We firstly extract multi-level acoustic information, including MFCC, spectrogram, and the embedded high-level acoustic information with CNN, BiL-STM and wav2vec2, respectively. Then these extracted features are treated as multimodal inputs and fused by the pro-posed co-attention mechanism. Experiments are carried on the IEMOCAP dataset, and our model achieves competitive performance with two different speaker-independent cross-validation strategies. Our code is available on GitHub. Heqing Zou, Yuke Si, Chen Chen 0075, Deepu Rajan, Chng Eng Siong |
ICASSP | 3 |
| 2022 | Interactive Auido-text Representation for Automated Audio Captioning with Contrastive Learning
Chen Chen 0075, Nana Hou, Heqing Zou, Xiaofeng Qi, Chng Eng Siong |
INTERSPEECH | 1 |
| 2022 | DENT-DDSP: Data-efficient noisy speech generator using differentiable digital signal processors for explicit distortion modelling and noise-robust speech recognitionabstractThe performances of automatic speech recognition (ASR) systems degrade drastically under noisy conditions.Explicit distortion modelling (EDM), as a feature compensation step, is able to enhance ASR systems under such conditions by simulating the in-domain noisy speeches from the clean counterparts.Yet, existing distortion models are either non-trainable or unexplainable and often lack controllability and generalization ability.In this paper, we propose a fully explainable and controllable model: DENT-DDSP to achieve EDM.DENT-DDSP utilizes novel differentiable digital signal processing (DDSP) components and requires only 10 seconds of training data to achieve high fidelity.The experiment shows that the simulated noisy data from DENT-DDSP achieves the highest simulation fidelity compared to other baseline models in terms of multi-scale spectral loss (MSSL).Moreover, to validate whether the data simulated by DENT-DDSP are able to replace the scarce in-domain noisy data in the noise-robust ASR tasks, several downstream ASR models with the same architecture are trained using the simulated data and the real data.The experiment shows that the model trained with the simulated noisy data from DENT-DDSP achieves similar performances to the benchmark with a 2.7% difference in terms of word error rate (WER).The code of the model is released online 1 . Zixun Guo, Chen Chen 0075, Chng Eng Siong |
INTERSPEECH | 2 |