VLDB 2026 Research / reviewers in the wild / expert
Qian Chen 0003
dblp:11/1394-3
· DBLP profile ↗
65ranked-venue papers
13as first author
53since 2021 · last 2026
0000-0001-6939-7438ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 10 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 41 · 8 first-author · 36 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language ModelsabstractThe Speaker Diarization and Recognition (SDR) task aims to predict ``who spoke when and what'' within an audio clip, which is a crucial task in various real-world multi-speaker scenarios such as meeting transcription and dialogue systems. Existing SDR systems typically adopt a cascaded framework, combining multiple modules such as speaker diarization (SD) and automatic speech recognition (ASR). The cascaded systems suffer from several limitations, such as error propagation, difficulty in handling overlapping speech, and lack of joint optimization for exploring the synergy between SD and ASR tasks. To address these limitations, we introduce SpeakerLM, a unified multimodal large language model for SDR that jointly performs SD and ASR in an end-to-end manner. Moreover, to facilitate diverse real-world scenarios, we incorporate a flexible speaker registration mechanism into SpeakerLM, enabling SDR under different speaker registration settings. SpeakerLM is progressively developed with a multi-stage training strategy on large-scale real data. Extensive experiments show that SpeakerLM demonstrates strong data scaling capability and generalizability, outperforming state-of-the-art cascaded baselines on both in-domain and out-of-domain public SDR benchmarks. Furthermore, experimental results show that the proposed speaker registration mechanism effectively ensures robust SDR performance of SpeakerLM across diverse speaker registration conditions and varying numbers of registered speakers. Han Yin, Yafeng Chen, Chong Deng, Luyao Cheng, Hui Wang 0030, Chao-Hong Tan, Qian Chen 0003, Wen Wang 0001, Xiangang Li |
AAAI | 7 |
| 2026 | Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration CodingabstractExisting speech tokenizers typically assign a fixed number of tokens per second, regardless of the varying information density or temporal fluctuations in the speech signal. This uniform token allocation mismatches the intrinsic structure of speech, where information is distributed unevenly over time. To address this, we propose VARSTok, a VAriable-frame-Rate Speech Tokenizer that adapts token allocation based on local feature similarity. VARSTok introduces two key innovations: (1) a temporal-aware density peak clustering algorithm that adaptively segments speech into variable-length units, and (2) a novel implicit duration coding scheme that embeds both content and temporal span into a single token index, eliminating the need for auxiliary duration predictors. Extensive experiments show that VARSTok significantly outperforms strong fixed-rate baselines. Notably, it achieves superior reconstruction naturalness while using up to 23% fewer tokens than a 40 Hz fixed-frame-rate baseline. VARSTok further yields lower word error rates and improved naturalness in zero-shot text-to-speech synthesis. To the best of our knowledge, this is the first work to demonstrate that a fully dynamic, variable-frame-rate acoustic speech tokenizer can be seamlessly integrated into downstream speech language models. Rui-Chen Zheng, Wenrui Liu 0003, Hui-Peng Du, Chong Deng, Qian Chen 0003, Wen Wang 0001, Yang Ai, Zhen-Hua Ling |
AAAI | 6 |
| 2026 | Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue ModelsabstractYifu Chen, Shengpeng Ji, Zhengqing Liu, Qian Chen, Wen Wang, Ziqing Wang, Yangzhuo Li, Tianle Liang, Zhou Zhao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shengpeng Ji, Zhengqing Liu, Qian Chen 0003, Wen Wang 0001, Yangzhuo Li, Tianle Liang, Zhou Zhao 0001 |
ACL (1) | 4 |
| 2026 | UniVocal: Unified Speech-Singing Code-Switching SynthesisabstractWe propose UniVocal, a unified framework that implicitly infers vocal modes from text context to pioneer Speech-Singing Code-Switching (SCS) Synthesis-a task where transitions are autonomously driven by textual semantics, akin to seamless human language blending.Unlike single-mode generation or systems relying on switching-control tags, our proposed UniVocal implicitly infers vocal modes solely from text context.To achieve this, we employ a data-efficient two-stage curriculum learning strategy that progressively trains a competitive TTS system to acquire the desired SCS capability.Addressing data scarcity, we introduce a scalable pipeline to synthesize diverse code-switching data that is both semantically and acoustically natural, alongside a new multi-scenario benchmark, SCSBench.To address limitations of semantic tokenizers in capturing acoustic details, we also introduce refined cent token and Chainof-Thought (CoT) generation for planning prosody before content generation, effectively enhancing empathetic speech generation and singing melody.Experimental results demonstrate that UniVocal achieves state-of-the-art performance on SCSBench while maintaining competitive performance on regular speech and singing tasks. Qian Chen 0003, Wen Wang 0001, Xiangang Li, Zhen-Hua Ling, Yang Ai |
ACL (1) | 2 |
| 2026 | GenesisFunc: Multi-Agent Data Generation for Accurate and Generalizable Function-CallingabstractHao-Xiang Xu, Chong Deng, Jiaqing Liu, Wen Wang, Qian Chen, Lujia Bao, Xiangang Li, Zhen-Hua Ling. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hao-Xiang Xu, Chong Deng, Jiaqing Liu, Wen Wang 0001, Qian Chen 0003, Lujia Bao, Xiangang Li, Zhen-Hua Ling |
ACL (1) | 5 |
| 2025 | Recording for Eyes, Not Echoing to Ears: Contextualized Spoken-to-Written Conversion of ASR TranscriptsabstractAutomatic Speech Recognition (ASR) transcripts exhibit recognition errors and various spoken language phenomena such as disfluencies, ungrammatical sentences, and incomplete sentences, hence suffering from poor readability. To improve readability, we propose a Contextualized Spoken-to-Written conversion (CoS2W) task to address ASR and grammar errors and also transfer the informal text into the formal style with content preserved, utilizing contexts and auxiliary information. This task naturally matches the in-context learning capabilities of Large Language Models (LLMs). To facilitate comprehensive comparisons of various LLMs, we construct a document-level Spoken-to-Written conversion of ASR Transcripts Benchmark (SWAB) dataset. Using SWAB, we study the impact of different granularity levels on the CoS2W performance, and propose methods to exploit contexts and auxiliary information to enhance the outputs. Experimental results reveal that LLMs have the potential to excel in the CoS2W task, particularly in grammaticality and formality, our methods achieve effective understanding of contexts and auxiliary information by LLMs. We further investigate the effectiveness of using LLMs as evaluators and find that LLM evaluators show strong correlations with human evaluations on rankings of faithfulness and formality, which validates the reliability of LLM evaluators for the CoS2W task. Jiaqing Liu, Chong Deng, Shilin Zhou 0002, Qian Chen 0003, Wen Wang 0001 |
AAAI | 5 |
| 2025 | Speech Recognition Meets Large Language Model: Benchmarking, Models, and ExplorationabstractIn this paper, we focus on prompting one of the most important tasks in the field of speech processing, i.e., automatic speech recognition (ASR), with speech foundation encoders and large language models (LLM). Despite the growing body of research in this area, we find that many crucial design decisions in LLM-based ASR systems are often inadequately justified. This lack of clarity impedes the field's progress, making it challenging to pinpoint which design choices truly improve model performance. To address these challenges, we conduct a comprehensive series of experiments that explore various aspects, leading to the optimal LLM-based ASR system. We found that delicate designs are not necessary, while a clean setup with little task-specific design is competent. The models achieve strong performance on the Librispeech and Gigaspeech datasets, compared to both LLM-based models and non-LLM-based models. Finally, we explore the capability emergence of LLM-based ASR in the process of modal alignment. We hope that our study can facilitate the research on extending LLM with cross-modality capacity and shed light on the LLM-based ASR community. Ziyang Ma 0001, Guanrou Yang, Yifan Yang 0005, Zhifu Gao, Jiaming Wang 0004, Zhihao Du, Fan Yu 0002, Qian Chen 0003, Shiliang Zhang, Xie Chen 0001 |
AAAI | 8 |
| 2025 | CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based VerificationabstractLarge Language Models (LLMs) have made significant progress in code generation, offering developers groundbreaking automated programming support. However, LLMs often generate code that is syntactically correct and even semantically plausible, but may not execute as expected or fulfill specified requirements. This phenomenon of hallucinations in the code domain has not been systematically explored. To advance the community's understanding and research on this issue, we introduce the concept of code hallucinations and propose a classification method for code hallucination based on execution verification. We categorize code hallucinations into four main types: mapping, naming, resource, and logic hallucinations, with each category further divided into different subcategories to understand and address the unique challenges faced by LLMs in code generation with finer granularity. Additionally, we present a dynamic detection algorithm called CodeHalu designed to detect and quantify code hallucinations. We also introduce the CodeHaluEval benchmark, which includes 8,883 samples from 699 tasks, to systematically and quantitatively evaluate code hallucinations. By evaluating 17 popular LLMs using this benchmark, we reveal significant differences in their accuracy and reliability in code generation, offering detailed insights for further improving the code generation capabilities of LLMs. Weixiang Yan, Qian Yang 0007, Xuandong Zhao, Qian Chen 0003, Wen Wang 0001, Dawn Song |
AAAI | 5 |
| 2025 | Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization on Multi-party ConversationabstractLuyao Cheng, Hui Wang, Chong Deng, Siqi Zheng, Yafeng Chen, Rongjie Huang, Qinglin Zhang, Qian Chen, Xihao Li, Wen Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Luyao Cheng, Hui Wang 0030, Chong Deng, Yafeng Chen, Rongjie Huang 0001, Qian Chen 0003, Xihao Li, Wen Wang 0001 |
ACL (1) | 8 |
| 2025 | ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style ControlabstractIn this paper, we present ControlSpeech, a text-to-speech (TTS) system capable of fully cloning the speaker’s voice and enabling arbitrary control and adjustment of speaking style. Prior zero-shot TTS models only mimic the speaker’s voice without further control and adjustment capabilities while prior controllable TTS models cannot perform speaker-specific voice generation. Therefore, ControlSpeech focuses on a more challenging task—a TTS system with controllable timbre, content, and style at the same time. ControlSpeech takes speech prompts, content prompts, and style prompts as inputs and utilizes bidirectional attention and mask-based parallel decoding to capture codec representations corresponding to timbre, content, and style in a discrete decoupling codec space. Moreover, we analyze the many-to-many issue in textual style control and propose the Style Mixture Semantic Density (SMSD) module, which is based on Gaussian mixture density networks, to resolve this problem. To facilitate empirical validations, we make available a new style controllable dataset called VccmDataset. Our experimental results demonstrate that ControlSpeech exhibits comparable or state-of-the-art (SOTA) performance in terms of controllability, timbre similarity, audio quality, robustness, and generalizability. Codes are available at https://github.com/jishengpeng/ControlSpeech. Shengpeng Ji, Qian Chen 0003, Wen Wang 0001, Jialong Zuo, Minghui Fang 0002, Ziyue Jiang 0004, Hai Huang 0013, Zehan Wang 0001, Xize Cheng, Zhou Zhao 0001 |
ACL (1) | 2 |
| 2025 | UniCodec: Unified Audio Codec with Single Domain-Adaptive CodebookabstractYidi Jiang, Qian Chen, Shengpeng Ji, Yu Xi, Wen Wang, Chong Zhang, Xianghu Yue, ShiLiang Zhang, Haizhou Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yidi Jiang, Qian Chen 0003, Shengpeng Ji, Yu Xi, Wen Wang 0001, Chong Zhang 0003, Xianghu Yue, Shiliang Zhang, Haizhou Li 0001 |
ACL (1) | 2 |
| 2025 | OmniFlatten: An End-to-end GPT Model for Seamless Voice ConversationabstractQinglin Zhang, Luyao Cheng, Chong Deng, Qian Chen, Wen Wang, Siqi Zheng, Jiaqing Liu, Hai Yu, Chao-Hong Tan, Zhihao Du, ShiLiang Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Luyao Cheng, Chong Deng, Qian Chen 0003, Wen Wang 0001, Jiaqing Liu, Chao-Hong Tan, Zhihao Du, Shiliang Zhang |
ACL (1) | 4 |
| 2025 | Self-Distillation Prototypes Network: Learning Robust Speaker Representations without SupervisionabstractTraining speaker-discriminative and robust speaker verification systems without explicit speaker labels remains a persistent challenge. In this paper, we propose a novel self-supervised speaker verification approach, Self-Distillation Prototypes Network (SDPN), which effectively facilitates self-supervised speaker representation learning. SDPN assigns the representation of the augmented views of an utterance to the same prototypes as the representation of the original view, thereby enabling effective knowledge transfer between the augmented and original views. Due to lack of negative pairs in the SDPN training process, the network tends to align positive pairs quite closely in the embedding space, a phenomenon known as model collapse. To mitigate this problem, we introduce a diversity regularization term to embeddings in SDPN. Comprehensive experiments on the VoxCeleb datasets demonstrate the superiority of SDPN among self-supervised speaker verification approaches. SDPN sets a new state-of-the-art on the VoxCeleb1 speaker verification evaluation benchmark, achieving Equal Error Rate 1.80%, 1.99%, and 3.62% for trial VoxCeleb1-O, VoxCeleb1-E and VoxCeleb1H respectively1, without using any speaker labels in training. Ablation studies show that both proposed learnable prototypes in self-distillation network and diversity regularization contribute to the verification performance. Yafeng Chen, Hui Wang 0030, Luyao Cheng, Qian Chen 0003, Chong Deng, Shiliang Zhang, Wen Wang 0001 |
ICASSP | 5 |
| 2025 | 3D-Speaker-Toolkit: An Open-Source Toolkit for Multimodal Speaker Verification and DiarizationabstractWe introduce 3D-Speaker-Toolkit, an open-source toolkit for multimodal speaker verification and diarization, designed for meeting the needs of academic researchers and industrial practitioners. The 3D-Speaker-Toolkit adeptly leverages the combined strengths of acoustic, semantic, and visual data, seamlessly fusing these modalities to offer robust speaker recognition capabilities. The acoustic module extracts speaker embeddings from acoustic features, employing both fully-supervised and self-supervised learning approaches. The semantic module leverages advanced language models to comprehend the substance and context of spoken language, thereby augmenting the system’s proficiency in distinguishing speakers through linguistic patterns. The visual module applies image processing technologies to scrutinize facial features, which bolsters the precision of speaker diarization in multi-speaker environments. Collectively, these modules empower the 3D-Speaker-Toolkit to achieve substantially improved accuracy and reliability in speaker-related tasks. With 3D-Speaker-Toolkit, we establish a new benchmark for multimodal speaker analysis. The toolkit also includes a handful of open-source state-of-the-art models and a large-scale dataset containing over 10,000 speakers. The toolkit is publicly available at https://github.com/modelscope/3D-Speaker. Yafeng Chen, Hui Wang 0030, Luyao Cheng, Tinglong Zhu, Rongjie Huang 0001, Chong Deng, Qian Chen 0003, Shiliang Zhang, Wen Wang 0001, Xihao Li |
ICASSP | 8 |
| 2025 | Unified Audio Event DetectionabstractSound Event Detection (SED) detects regions of sound events, while Speaker Diarization (SD) segments speech conversations attributed to individual speakers. In SED, all speaker segments are classified as a single speech event, while in SD, non-speech sounds are treated merely as background noise. Thus, both tasks provide only partial analysis in complex audio scenarios involving both speech conversation and non-speech sounds. In this paper, we introduce a novel task called Unified Audio Event Detection (UAED) for comprehensive audio analysis. UAED explores the synergy between SED and SD tasks, simultaneously detecting non-speech sound events and fine-grained speech events based on speaker identities. To tackle this task, we propose a Transformer-based UAED (T-UAED) framework and construct the UAED Data derived from the Librispeech dataset and DESED soundbank. Experiments demonstrate that the proposed framework effectively exploits task interactions and substantially outperforms the baseline that simply combines the outputs of SED and SD models. T-UAED also shows its versatility by performing comparably to specialized models for individual SED and SD tasks on DESED and CALLHOME datasets. Yidi Jiang, Ruijie Tao, Qian Chen 0003, Wen Wang 0001 |
ICASSP | 4 |
| 2025 | WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language ModelingabstractLanguage models have been effectively applied to modeling natural signals, such as images, video, speech, and audio. A crucial component of these models is the codec tokenizer, which compresses high-dimensional natural signals into lower-dimensional discrete tokens. In this paper, we introduce WavTokenizer, which offers several advantages over previous SOTA acoustic codec models in the audio domain: 1) extreme compression. By compressing the layers of quantizers and the temporal dimension of the discrete codec, one-second audio of 24kHz sampling rate requires only a single quantizer with 40 or 75 tokens. 2) improved subjective quality. Despite the reduced number of tokens, WavTokenizer achieves state-of-the-art reconstruction quality with outstanding UTMOS scores and inherently contains richer semantic information. Specifically, we achieve these results by designing a broader VQ space, extended contextual windows, and improved attention networks, as well as introducing a powerful multi-scale discriminator and an inverse Fourier transform structure. We conducted extensive reconstruction experiments in the domains of speech, audio, and music. WavTokenizer exhibited strong performance across various objective and subjective metrics compared to state-of-the-art models. We also tested semantic information, VQ utilization, and adaptability to generative models. Comprehensive ablation studies confirm the necessity of each module in WavTokenizer. The code is available at https://github.com/jishengpeng/WavTokenizer. Shengpeng Ji, Ziyue Jiang 0001, Wen Wang 0001, Minghui Fang 0002, Jialong Zuo, Qian Yang 0006, Xize Cheng, Zehan Wang 0001, Ruiqi Li 0002, Xiaoda Yang, Rongjie Huang 0001, Yidi Jiang, Qian Chen 0003, Zhou Zhao 0001 |
ICLR | 15 |
| 2025 | OmniAudio: Generating Spatial Audio from 360-Degree VideoabstractTraditional video-to-audio generation techniques primarily focus on perspective video and non-spatial audio, often missing the spatial cues necessary for accurately representing sound sources in 3D environments. To address this limitation, we introduce a novel task, 360V2SA, to generate spatial audio from 360-degree videos, specifically producing First-order Ambisonics (FOA) audio - a standard format for representing 3D spatial audio that captures sound directionality and enables realistic 3D audio reproduction. We first create Sphere360, a novel dataset tailored for this task that is curated from real-world data. We also design an efficient semi-automated pipeline for collecting and cleaning paired video-audio data. To generate spatial audio from 360-degree video, we propose a novel framework OmniAudio, which leverages self-supervised pre-training using both spatial audio data (in FOA format) and large-scale non-spatial data. Furthermore, OmniAudio features a dual-branch framework that utilizes both panoramic and perspective video inputs to capture comprehensive local and global information from 360-degree videos. Experimental results demonstrate that OmniAudio achieves state-of-the-art performance across both objective and subjective metrics on Sphere360. Code and datasets are available at https://github.com/liuhuadai/OmniAudio. The project website is available at https://OmniAudio-360V2SA.github.io. Huadai Liu, Tianyi Luo, Kaicheng Luo, Qikai Jiang, Peiwen Sun, Rongjie Huang 0001, Qian Chen 0003, Wen Wang 0001, Xiangtai Li, Shiliang Zhang, Zhijie Yan, Zhou Zhao 0001, Wei Xue 0002 |
ICML | 8 |
| 2025 | Pushing the Frontiers of Self-Distillation Prototypes Network with Dimension Regularization and Score Normalization
Yafeng Chen, Chong Deng, Hui Wang 0030, Yiheng Jiang, Han Yin, Qian Chen 0003, Wen Wang 0001 |
INTERSPEECH | 6 |
| 2025 | Speech Token Prediction via Compressed-to-fine Language Modeling for Speech GenerationabstractNeural audio codecs, used as speech tokenizers, have demonstrated remarkable potential in the field of speech generation. However, to ensure high-fidelity audio reconstruction, neural audio codecs typically encode audio into long sequences of speech tokens, posing a significant challenge for downstream language models in long-context modeling. We observe that speech token sequences exhibit short-range dependency: due to the monotonic alignment between text and speech in text-to-speech (TTS) tasks, the prediction of the current token primarily relies on its local context, while long-range tokens contribute less to the current token prediction and often contain redundant information. Inspired by this observation, we propose a compressed-to-fine language modeling approach to address the challenge of long sequence speech tokens within neural codec language models: (1) Fine-grained Initial and Short-range Information: Our approach retains the prompt and local tokens during prediction to ensure text alignment and the integrity of paralinguistic information; (2) Compressed Long-range Context: Our approach compresses long-range token spans into compact representations to reduce redundant information while preserving essential semantics. Extensive experiments on various neural audio codecs and downstream language models validate the effectiveness and generalizability of the proposed approach, highlighting the importance of token compression in improving speech generation within neural codec language models. The demo of audio samples will be available at https://anonymous.4open.science/r/SpeechTokenPredictionViaCompressedToFinedLM. Wenrui Liu 0003, Qian Chen 0003, Wen Wang 0019, Guanrou Yang, Minghui Fang 0002, Jialong Zuo, Xiaoda Yang, Tao Jin 0004, Jin Xu 0010, Yafeng Chen, Jionghao Bai, Zhifang Guo |
ACM Multimedia | 2 |
| 2025 | MelodyEdit: Zero-shot Music Editing with Disentangled Inversion ControlabstractText-guided diffusion models revolutionize audio generation by adapting source audio to specific text prompts. However, existing zero-shot audio editing methods such as DDIM inversion accumulate errors across diffusion steps, reducing the effectiveness. Moreover, existing editing methods struggle with conducting complex non-rigid music edits while maintaining content integrity and high fidelity. To address these challenges, we propose MelodyEdit, a novel zero-shot music editing system based on innovative Disentangled Inversion Control (DIC) technique, which comprises Harmonized Attention Control and Disentangled Inversion. Disentangled Inversion disentangles the diffusion process into triple branches to rectify the deviated path of the source branch caused by DDIM inversion. Harmonized Attention Control unifies the mutual self-attention control and the cross-attention control with an intermediate Harmonic Branch to progressively generate the desired harmonic and melodic information in the target music. We also introduce ZoME-Bench, a comprehensive music editing benchmark with 1,100 samples covering ten distinct editing categories. ZoME-Bench facilitates both zero-shot and instruction-based music editing tasks. Our method outperforms state-of-the-art inversion techniques in editing fidelity and content preservation. Huadai Liu, Xiangtai Li, Wen Wang 0019, Qian Chen 0003, Rongjie Huang 0001, Zhou Zhao 0001, Wei Xue 0002 |
ACM Multimedia | 5 |
| 2025 | EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text PromptingabstractHuman speech goes beyond the mere transfer of information; it is a profound exchange of emotions and a connection between individuals. While Text-to-Speech (TTS) models have made huge progress, they still face challenges in controlling the emotional expression in the generated speech. In this work, we propose EmoVoice, a novel emotion-controllable TTS model that exploits large language models (LLMs) to enable fine-grained freestyle natural language emotion control, and a phoneme boost variant design that makes the model output phoneme tokens and audio tokens in parallel to enhance content consistency, inspired by chain-of-thought (CoT) and modality-of-thought (CoM) techniques. Besides, we introduce EmoVoice-DB, a high-quality 40-hour English emotion dataset featuring expressive speech and fine-grained emotion labels with natural language descriptions. EmoVoice achieves state-of-the-art performance on the English EmoVoice-DB test set using only synthetic training data, and on the Chinese Secap test set using our in-house data. We further investigate the reliability of existing emotion evaluation metrics and their alignment with human perceptual preferences, and explore using SOTA multimodal LLMs GPT-4o-audio and Gemini to assess emotional speech. Dataset, code, checkpoints and demo samples are available at https://github.com/yanghaha0908/EmoVoice. Guanrou Yang, Qian Chen 0003, Ziyang Ma 0001, Wen Wang 0019, Tianrui Wang, Yifan Yang 0005, Zhikang Niu, Wenrui Liu 0003, Fan Yu 0002, Zhihao Du, Zhifu Gao, Shiliang Zhang, Xie Chen 0001 |
ACM Multimedia | 3 |
| 2025 | ThinkSound: Chain-of-Thought Reasoning in Multimodal LLMs for Audio Generation and EditingabstractWhile end-to-end video-to-audio generation has greatly improved, producing high-fidelity audio that authentically captures the nuances of visual content remains challenging. Like professionals in the creative industries, this generation requires sophisticated reasoning about items such as visual dynamics, acoustic environments, and temporal relationships. We present **ThinkSound**, a novel framework that leverages Chain-of-Thought (CoT) reasoning to enable stepwise, interactive audio generation and editing for videos. Our approach decomposes the process into three complementary stages: foundational foley generation that creates semantically coherent soundscapes, interactive object-centric refinement through precise user interactions, and targeted editing guided by natural language instructions. At each stage, a multimodal large language model generates contextually aligned CoT reasoning that guides a unified audio foundation model. Furthermore, we introduce **AudioCoT**, a comprehensive dataset with structured reasoning annotations that establishes connections between visual content, textual descriptions, and sound synthesis. Experiments demonstrate that ThinkSound achieves state-of-the-art performance in video-to-audio generation across both audio metrics and CoT metrics, and excels in the out-of-distribution Movie Gen Audio benchmark. The project page is available at https://ThinkSound-Project.github.io. Huadai Liu, Kaicheng Luo, Wen Wang 0001, Qian Chen 0003, Zhou Zhao 0001, Wei Xue 0002 |
NeurIPS | 5 |
| 2025 | ClinicalLab: Aligning Agents for Multi-Departmental Clinical Diagnostics in the Real WorldabstractLarge language models (LLMs) have achieved significant performance progress in various natural language processing applications. However, LLMs still struggle to meet the strict requirements for accuracy and reliability in the medical field and face many challenges in clinical applications. Existing clinical diagnostic evaluation benchmarks for evaluating medical agents powered by LLMs have severe limitations. To address these limitations, we introduce ClinicalLab, a comprehensive clinical diagnosis agent alignment suite. ClinicalLab includes ClinicalBench, an end-to-end multi-departmental clinical diagnostic evaluation benchmark for evaluating medical agents and LLMs. ClinicalBench is based on real cases that cover 24 departments and 150 diseases. We ensure that ClinicalBench does not have data leakage. ClinicalLab also includes four novel metrics (ClinicalMetrics) for evaluating the effectiveness of LLMs in clinical diagnostic tasks. We evaluate 17 general and medical-domain LLMs and find that their performance varies significantly across different departments. Based on these findings, in ClinicalLab, we propose ClinicalAgent, an end-to-end clinical agent that aligns with real-world clinical diagnostic practices. We systematically investigate the performance and applicable scenarios of variants of ClinicalAgent on ClinicalBench. Our findings demonstrate the importance of aligning with modern medical practices in designing medical agents. Weixiang Yan, Tengxiao Wu, Qian Chen 0003, Wen Wang 0001, Haoyuan Chai |
NeurIPS | 4 |
| 2024 | CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and GenerationabstractWeixiang Yan, Haitian Liu, Yunkun Wang, Yunzhe Li, Qian Chen, Wen Wang, Tingyu Lin, Weishan Zhao, Li Zhu, Hari Sundaram, Shuiguang Deng. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Weixiang Yan, Yunkun Wang, Yunzhe Li 0001, Qian Chen 0003, Wen Wang 0001, Tingyu Lin 0002, Weishan Zhao, Hari Sundaram, Shuiguang Deng |
ACL (1) | 5 |
| 2024 | Advancing Precise Outline-Conditioned Text Generation with Task Duality and Explicit Outline ControlabstractYunzhe Li, Qian Chen, Weixiang Yan, Wen Wang, Qinglin Zhang, Hari Sundaram. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yunzhe Li 0001, Qian Chen 0003, Weixiang Yan, Wen Wang 0001, Hari Sundaram |
EACL (1) | 2 |
| 2024 | Loss Masking Is Not Needed In Decoder-Only Transformer For Discrete-Token-Based ASRabstractRecently, unified speech-text models, such as SpeechGPT, VioLA, and AudioPaLM, have achieved remarkable performance on various speech tasks. These models discretize speech signals into tokens (speech discretization) and use a shared vocabulary for both text and speech tokens. Then they train a single decoder-only Transformer on a mixture of speech tasks. However, these models rely on the Loss Masking strategy for the ASR task, which ignores the dependency among speech tokens. In this paper, we propose to model speech tokens in an autoregressive way, similar to text. We find that applying the conventional cross-entropy loss on input speech tokens does not consistently improve the ASR performance over the Loss Masking approach. To address this issue, we propose a novel approach denoted Smoothed Label Distillation (SLD), which applies a KL divergence loss with smoothed labels on speech tokens. Our experiments show that SLD effectively models speech tokens and outperforms Loss Masking for decoder-only Transformers in ASR tasks with different speech discretization methods1. Qian Chen 0003, Wen Wang 0001, Shiliang Zhang, Chong Deng, Jiaqing Liu, Chong Zhang 0003 |
ICASSP | 1 |
| 2024 | Leveraging Speech PTM, Text LLM, And Emotional TTS For Speech Emotion RecognitionabstractIn this paper, we explored how to boost speech emotion recognition (SER) with the state-of-the-art speech pre-trained model (PTM), data2vec, text generation technique, GPT-4, and speech synthesis technique, Azure TTS. First, we investigated the representation ability of different speech self-supervised pre-trained models, and we found that data2vec has a good representation ability on the SER task. Second, we employed a powerful large language model (LLM), GPT-4, and emotional text-to-speech (TTS) model, Azure TTS, to generate emotionally congruent text and speech. We carefully designed the text prompt and dataset construction, to obtain the synthetic emotional speech data with high quality. Third, we studied different ways of data augmentation to promote the SER task with synthetic speech, including random mixing, adversarial training, transfer learning, and curriculum learning. Experiments and ablation studies on the IEMOCAP dataset demonstrate the effectiveness of our method, compared with other data augmentation methods, and data augmentation with other synthetic data. Ziyang Ma 0001, Wen Wu 0007, Zhisheng Zheng, Qian Chen 0003, Shiliang Zhang, Xie Chen 0001 |
ICASSP | 5 |
| 2024 | ERes2NetV2: Boosting Short-Duration Speaker Verification Performance with Computational Efficiency
Yafeng Chen, Hui Wang 0030, Luyao Cheng, Qian Chen 0003, Shiliang Zhang |
INTERSPEECH | 5 |
| 2024 | Tuning Large Language Model for Speech Recognition With Mixed-Scale Re-TokenizationabstractLarge Language Models (LLMs) have proven successful across a spectrum of speech-related tasks, such as speech recognition, text-to-speech, and spoken language understanding. Recently, the use of discretized speech features has gained attention as an efficient and compatible alternative to continuous features for LLMs. This is mainly due to their reduced storage requirements and better alignment of these features with LLM's input space. However, the typical practice of freezing the speech encoder during training poses challenges in bridging the modality gap between speech and text. To address this, we propose to use a mixed-scale re-tokenization layer, integrating multiple granularities in discretized speech features directly within the LLM's input module. Our experimental results demonstrated that the proposed method can effectively enhance the performance of ASR in the setting of continuous learning of an LLM, highlighting the importance of a meticulously designed input module for the integration of discretized speech features with an LLM. Chong Zhang 0003, Qian Chen 0003, Wen Wang 0001, Bin Ma 0001 |
IEEE Signal Process. Lett. | 3 |
| 2023 | The Second Multi-Channel Multi-Party Meeting Transcription Challenge (M2MeT 2.0): A Benchmark for Speaker-Attributed ASRabstractWith the success of the first Multi-channel Multi-party Meeting Transcription challenge (M2MeT), the second M2MeT challenge (M2MeT 2.0) held in ASRU2023 particularly aims to tackle the complex task of speaker-attributed ASR (SAASR), which directly addresses the practical and challenging problem of “who spoke what at when” at typical meeting scenario. We particularly established two sub-tracks. The fixed training condition sub-track, where the training data is constrained to predetermined datasets, but participants can use any open-source pre-trained model. The open training condition sub-track, which allows for the use of all available data and models without limitation. In addition, we release a new 10-hour test set for challenge ranking. This paper provides an overview of the dataset, track settings, results, and analysis of submitted systems, as a benchmark to show the current state of speaker-attributed ASR. Yuhao Liang, Mohan Shi, Fan Yu 0002, Yangze Li, Shiliang Zhang, Zhihao Du, Qian Chen 0003, Lei Xie 0001, Yanmin Qian, Jian Wu 0027, Zhuo Chen 0006, Kong-Aik Lee, Zhijie Yan, Hui Bu |
ASRU | 7 |
| 2023 | Ditto: A Simple and Efficient Approach to Improve Sentence EmbeddingsabstractPrior studies diagnose the anisotropy problem in sentence representations from pre-trained language models, e.g., BERT, without finetuning.Our analysis reveals that the sentence embeddings from BERT suffer from a bias towards uninformative words, limiting the performance in semantic textual similarity (STS) tasks.To address this bias, we propose a simple and efficient unsupervised approach, Diagonal Attention Pooling (Ditto), which weights words with model-based importance estimations and computes the weighted average of word representations from pre-trained models as sentence embeddings.Ditto can be easily applied to any pre-trained language model as a postprocessing operation.Compared to prior sentence embedding approaches, Ditto does not add parameters nor requires any learning.Empirical evaluations demonstrate that our proposed Ditto can alleviate the anisotropy problem and improve various pre-trained models on the STS benchmarks. 1 Qian Chen 0003, Wen Wang 0001, Chong Deng, Jiaqing Liu, Chong Zhang 0003 |
EMNLP | 1 |
| 2023 | Improving Long Document Topic Segmentation Models With Enhanced Coherence ModelingabstractTopic segmentation is critical for obtaining structured documents and improving downstream tasks such as information retrieval.Due to its ability of automatically exploring clues of topic shift from abundant labeled data, recent supervised neural models have greatly promoted the development of long document topic segmentation, but leaving the deeper relationship between coherence and topic segmentation underexplored.Therefore, this paper enhances the ability of supervised models to capture coherence from both logical structure and semantic similarity perspectives to further improve the topic segmentation performance, proposing Topic-aware Sentence Structure Prediction (TSSP) and Contrastive Semantic Similarity Learning (CSSL).Specifically, the TSSP task is proposed to force the model to comprehend structural information by learning the original relations between adjacent sentences in a disarrayed document, which is constructed by jointly disrupting the original document at topic and sentence levels.Moreover, we utilize inter-and intra-topic information to construct contrastive samples and design the CSSL objective to ensure that the sentences representations in the same topic have higher similarity, while those in different topics are less similar.Extensive experiments show that Longformer with our approach significantly outperforms state-of-the-art (SOTA) methods.Our approach improves F 1 of SOTA by 3.42 (73.74 → 77.16) and improves P k by 1.11 points (15.0 → 13.89) on WIKI-727K and achieves an average relative reduction of 4.3% on P k on WikiSection.The average relative P k drop of 8.38% on two out-of-domain datasets also demonstrates the robustness of our approach 1 . Chong Deng, Jiaqing Liu, Qian Chen 0003, Wen Wang 0001 |
EMNLP | 5 |
| 2023 | Pushing the Limits of Self-Supervised Speaker Verification using Regularized Distillation FrameworkabstractTraining robust speaker verification systems without speaker labels has long been a challenging task. Previous studies observed a large performance gap between self-supervised and fully supervised methods. In this paper, we apply a non-contrastive self-supervised learning framework called DIstillation with NO labels (DINO) and propose two regularization terms applied to embeddings in DINO. One regularization term guarantees the diversity of the embeddings, while the other regularization term decorrelates the variables of each embedding. The effectiveness of various data augmentation techniques are explored, on both time and frequency domain. A range of experiments conducted on the VoxCeleb datasets demonstrate the superiority of the regularized DINO framework in speaker verification. Our method achieves the stateof-the-art speaker verification performance under a singlestage self-supervised setting on VoxCeleb. Yafeng Chen, Hui Wang 0030, Luyao Cheng, Qian Chen 0003 |
ICASSP | 5 |
| 2023 | Meeting Action Item Detection with Regularized Context ModelingabstractMeetings are increasingly important for collaborations. Action items in meeting transcripts are crucial for managing post-meeting to-do tasks, which usually are summarized laboriously. The Action Item Detection task aims to automatically detect meeting content associated with action items. However, datasets manually annotated with action item detection labels are scarce and in small scale. We construct and release the first Chinese meeting corpus with manual action item annotations1. In addition, we propose a Context-Drop approach to utilize both local and global contexts by contrastive learning, and achieve better accuracy and robustness for action item detection. We also propose a Lightweight Model Ensemble method to exploit different pre-trained models2. Experimental results on our Chinese meeting corpus and the English AMI corpus demonstrate the effectiveness of the proposed approaches. Jiaqing Liu, Chong Deng, Qian Chen 0003, Wen Wang 0001 |
ICASSP | 4 |
| 2023 | Auxiliary Pooling Layer For Spoken Language UnderstandingabstractEnd-to-end spoken language understanding requires speech data annotated with semantic information and may suffer from the shortage of annotated data. Recent progresses leverage unlabelled speech data to pre-train a speech encoder. However, it remains a challenge for the pre-trained speech encoder to encode semantic information. Existing works explore transferring knowledge from a pre-trained text model with different alignment losses at a fixed granularity. In this paper, we address the variable granularity in transferring knowledge from texts to speech representation via APLY, an auxiliary pooling layer, that fuses the global information with the adaptively encoded local context. We demonstrate the effectiveness of APLY on three benchmarks of spoken language understanding. Trung Hieu Nguyen 0001, Jinjie Ni, Wen Wang 0001, Qian Chen 0003, Chong Zhang 0003, Bin Ma 0001 |
ICASSP | 5 |
| 2023 | Weighted Sampling for Masked Language ModelingabstractMasked Language Modeling (MLM) is widely used to pretrain language models. The standard random masking strategy in MLM causes the pre-trained language models (PLMs) to be biased towards high-frequency tokens. Representation learning of rare tokens is poor and PLMs have limited performance on downstream tasks. To alleviate this frequency bias issue, we propose two simple and effective Weighted Sampling strategies for masking tokens based on token frequency and training loss. We apply these two strategies to BERT and obtain Weighted-Sampled BERT (WSBERT). Experiments on the Semantic Textual Similarity benchmark (STS) show that WSBERT significantly improves sentence embeddings over BERT. Combining WSBERT with calibration methods and prompt learning further improves sentence embeddings. We also investigate fine-tuning WSBERT on the GLUE benchmark and show that Weighted Sampling also improves the transfer learning capability of the backbone PLM. We further analyze and provide insights into how WSBERT improves token embeddings. Linhan Zhang, Qian Chen 0003, Wen Wang 0001, Chong Deng, Xin Cao 0001, Kongzhang Hao, Wei Wang 0011 |
ICASSP | 2 |
| 2023 | Overview of the ICASSP 2023 General Meeting Understanding and Generation Challenge (MUG)abstractICASSP2023 General Meeting Understanding and Generation Challenge (MUG) focuses on prompting a wide range of spoken language processing (SLP) research on meeting transcripts, as SLP applications are critical to improve users’ efficiency in grasping important information in meetings. MUG includes five tracks, including topic segmentation, topic-level and session-level extractive summarization, topic title generation, keyphrase extraction, and action item detection. To facilitate MUG, we construct and release a large-scale meeting dataset, the AliMeeting4MUG Corpus. We review the dataset, track settings and baselines, and summarize the challenge results and major techniques used in the submissions. Chong Deng, Jiaqing Liu, Qian Chen 0003, Wen Wang 0001, Zhijie Yan, Jinglin Liu, Yi Ren 0006, Zhou Zhao 0001 |
ICASSP | 5 |
| 2023 | MUG: A General Meeting Understanding and Generation BenchmarkabstractListening to long video/audio recordings from video conferencing and online courses for acquiring information is extremely inefficient. Even after ASR systems transcribe recordings into long-form spoken language documents, reading ASR transcripts only partly speeds up seeking information. It has been observed that a range of NLP applications, such as keyphrase extraction, topic segmentation, and summarization, significantly improve users’ efficiency in grasping important information. The meeting scenario is among the most valuable scenarios for deploying these spoken language processing (SLP) capabilities. However, the lack of large-scale public meeting datasets annotated for these SLP tasks severely hinders their advancement. To prompt SLP advancement, we establish a large-scale general Meeting Understanding and Generation Benchmark (MUG) to benchmark the performance of a wide range of SLP tasks, including topic segmentation, topic-level and session-level extractive summarization and topic title generation, keyphrase extraction, and action item detection. To facilitate the MUG benchmark, we construct and release a large-scale meeting dataset for comprehensive long-form SLP development, the AliMeeting4MUG Corpus, which consists of 654 recorded Mandarin meeting sessions with diverse topic coverage, with manual annotations for SLP tasks on manual transcripts of meeting recordings. To the best of our knowledge, the AliMeeting4MUG Corpus is so far the largest meeting corpus in scale and facilitates most SLP tasks. In this paper, we provide a detailed introduction of this corpus, SLP tasks and evaluation methods, baseline systems and their performance1. Chong Deng, Jiaqing Liu, Qian Chen 0003, Wen Wang 0001, Zhijie Yan, Jinglin Liu, Yi Ren 0006, Zhou Zhao 0001 |
ICASSP | 5 |
| 2023 | An Enhanced Res2Net with Local and Global Feature Fusion for Speaker Verification
Yafeng Chen, Hui Wang 0030, Luyao Cheng, Qian Chen 0003, Jiajun Qi |
INTERSPEECH | 5 |
| 2023 | Personality-aware Training based Speaker Adaptation for End-to-end Speech Recognition
Zhihao Du, Shiliang Zhang, Qian Chen 0003, Jiqing Han 0001 |
INTERSPEECH | 4 |
| 2023 | BA-SOT: Boundary-Aware Serialized Output Training for Multi-Talker ASR
Yuhao Liang, Fan Yu 0002, Yangze Li, Shiliang Zhang, Qian Chen 0003, Lei Xie 0001 |
INTERSPEECH | 6 |
| 2023 | Adapter-tuning with Effective Token-dependent Representation Shift for Automatic Speech Recognition
Dianwen Ng, Chong Zhang 0003, Ruixi Zhang, Trung Hieu Nguyen 0001, Chongjia Ni, Shengkui Zhao, Qian Chen 0003, Wen Wang 0001, Chng Eng Siong, Bin Ma 0001 |
INTERSPEECH | 8 |
| 2023 | CASA-ASR: Context-Aware Speaker-Attributed ASR
Mohan Shi, Zhihao Du, Qian Chen 0003, Fan Yu 0002, Yangze Li, Shiliang Zhang, Jie Zhang 0042, Li-Rong Dai 0001 |
INTERSPEECH | 3 |
| 2023 | Semantic VAD: Low-Latency Voice Activity Detection for Speech Interaction
Mohan Shi, Yuchun Shu, Lingyun Zuo, Qian Chen 0003, Shiliang Zhang, Jie Zhang 0042, Li-Rong Dai 0001 |
INTERSPEECH | 4 |
| 2023 | CAM++: A Fast and Efficient Network for Speaker Verification Using Context-Aware Masking
Hui Wang 0030, Yafeng Chen, Luyao Cheng, Qian Chen 0003 |
INTERSPEECH | 5 |
| 2023 | Enhancing Multi-modal Multi-hop Question Answering via Structured Knowledge and Unified Retrieval-GenerationabstractMulti-modal multi-hop question answering involves answering a question by reasoning over multiple input sources from different modalities. Existing methods often retrieve evidences separately and then use a language model to generate an answer based on the retrieved evidences, and thus do not adequately connect candidates and are unable to model the interdependent relations during retrieval. Moreover, the pipelined approaches of retrieval and generation might result in poor generation performance when retrieval performance is low. To address these issues, we propose a Structured Knowledge and Unified Retrieval-Generation (SKURG) approach. SKURG employs an Entity-centered Fusion Encoder to align sources from different modalities using shared entities. It then uses a unified Retrieval-Generation Decoder to integrate intermediate retrieval results for answer generation and also adaptively determine the number of retrieval steps. Extensive experiments on two representative multi-modal multi-hop QA datasets MultimodalQA and WebQA demonstrate that SKURG outperforms the state-of-the-art models in both source retrieval and answer generation performance with fewer parameters1. Qian Yang 0007, Qian Chen 0003, Wen Wang 0001, Baotian Hu, Min Zhang 0005 |
ACM Multimedia | 2 |
| 2022 | Prosospeech: Enhancing Prosody with Quantized Vector Pre-Training in Text-To-SpeechabstractExpressive text-to-speech (TTS) has become a hot research topic recently, mainly focusing on modeling prosody in speech. Prosody modeling has several challenges: 1) the extracted pitch used in previous prosody modeling works have inevitable errors, which hurts the prosody modeling; 2) different attributes of prosody (e.g., pitch, duration and energy) are dependent on each other and produce the natural prosody together; and 3) due to high variability of prosody and the limited amount of high-quality data for TTS training, the distribution of prosody cannot be fully shaped. To tackle these issues, we propose ProsoSpeech, which enhances the prosody using quantized latent vectors pre-trained on large-scale unpaired and low-quality text and speech data. Specifically, we first introduce a word-level prosody encoder, which quantizes the low-frequency band of the speech and compresses prosody at-tributes in the latent prosody vector (LPV). Then we introduce an LPV predictor, which predicts LPV given word sequence. We pre-train the LPV predictor on large-scale text and low-quality speech data and fine-tune it on the high-quality TTS dataset. Finally, our model can generate expressive speech conditioned on the predicted LPV. Experimental results show that ProsoSpeech can generate speech with richer prosody compared with baseline methods. Yi Ren 0006, Zhiying Huang, Shiliang Zhang, Qian Chen 0003, Zhijie Yan, Zhou Zhao 0001 |
ICASSP | 5 |
| 2022 | PoNet: Pooling Network for Efficient Token Mixing in Long Sequences
Chao-Hong Tan, Qian Chen 0003, Wen Wang 0001, Zhen-Hua Ling |
ICLR | 2 |
| 2022 | PRISM: Pre-trained Indeterminate Speaker Representation Model for Speaker Diarization and Speaker VerificationabstractSpeaker embedding has been a fundamental feature for speaker-related tasks such as verification, clustering, and diarization. Traditionally, speaker embeddings are represented as fixed vectors in high-dimensional space. This could lead to biased estimations, especially when handling shorter utterances. In this paper we propose to represent a speaker utterance as "floating" vector whose state is indeterminate without knowing the context. The state of a speaker representation is jointly determined by itself, other speech from the same speaker, as well as other speakers it is being compared to. The content of the speech also contributes to determining the final state of a speaker representation. We pre-train an indeterminate speaker representation model that estimates the state of an utterance based on the context. The pre-trained model can be fine-tuned for downstream tasks such as speaker verification, speaker clustering, and speaker diarization. Substantial improvements are observed across all downstream tasks. Hongbin Suo, Qian Chen 0003 |
INTERSPEECH | 3 |
| 2021 | Sequence Model with Self-Adaptive Sliding Window for Efficient Spoken Document SegmentationabstractTranscripts generated by automatic speech recognition (ASR) systems for spoken documents lack structural annotations such as paragraphs, significantly reducing their readability. Automatically predicting paragraph segmentation for spoken documents may both improve readability and downstream NLP performance such as summarization and machine reading comprehension. We propose a sequence model with self-adaptive sliding window for accurate and efficient paragraph segmentation. We also propose an approach to exploit pho-netic information, which significantly improves robustness of spoken document segmentation to ASR errors. Evaluations are conducted on the English Wiki-727K document seg-mentation benchmark, a Chinese Wikipedia-based document segmentation dataset we created, and an in-house Chinese spoken document dataset. Our proposed model outperforms the state-of-the-art (SOTA) model based on the same BERT-Base, increasing segmentation F1 on the English benchmark by 4.2 points and on Chinese datasets by 4.3-10.1 points, while reducing inference time to less than 1/6 of inference time of the current SOTA. Qian Chen 0003, Yali Li 0001, Jiaqing Liu, Wen Wang 0001 |
ASRU | 2 |
| 2021 | Discriminative Self-Training for Punctuation PredictionabstractPunctuation prediction for automatic speech recognition (ASR) output transcripts plays a crucial role for improving the readability of the ASR transcripts and for improving the performance of downstream natural language processing applications. However, achieving good performance on punctuation prediction often requires large amounts of labeled speech transcripts, which is expensive and laborious. In this paper, we propose a Discriminative Self-Training approach with weighted loss and discriminative label smoothing to exploit unlabeled speech transcripts. Experimental results on the English IWSLT2011 benchmark test set and an internal Chinese spoken language dataset demonstrate that the proposed approach achieves significant improvement on punctuation prediction accuracy over strong baselines including BERT, RoBERTa, and ELECTRA models. The proposed Discriminative Self-Training approach outperforms the vanilla self-training approach. We establish a new state-of-the-art (SOTA) on the IWSLT2011 test set, outperforming the current SOTA model by 1.3% absolute gain on F$_1$. Qian Chen 0003, Wen Wang 0001, Mengzhe Chen |
Interspeech | 1 |
| 2021 | Pre-Training for Spoken Language Understanding with Joint Textual and Phonetic Representation LearningabstractIn the traditional cascading architecture for spoken language understanding (SLU), it has been observed that automatic speech recognition errors could be detrimental to the performance of natural language understanding. End-to-end (E2E) SLU models have been proposed to directly map speech input to desired semantic frame with a single model, hence mitigating ASR error propagation. Recently, pre-training technologies have been explored for these E2E models. In this paper, we propose a novel joint textual-phonetic pre-training approach for learning spoken language representations, aiming at exploring the full potentials of phonetic information to improve SLU robustness to ASR errors. We explore phoneme labels as high-level speech features, and design and compare pre-training tasks based on conditional masked language model objectives and inter-sentence relation objectives. We also investigate the efficacy of combining textual and phonetic information during fine-tuning. Experimental results on spoken language understanding benchmarks, Fluent Speech Commands and SNIPS, show that the proposed approach significantly outperforms strong baseline models and improves robustness of spoken language understanding to ASR errors. Qian Chen 0003, Wen Wang 0001 |
Interspeech | 1 |
| 2021 | TRS: Transferability Reduced Ensemble via Promoting Gradient Diversity and Model SmoothnessabstractAdversarial Transferability is an intriguing property - adversarial perturbation crafted against one model is also effective against another model, while these models are from different model families or training processes. To better protect ML systems against adversarial attacks, several questions are raised: what are the sufficient conditions for adversarial transferability, and how to bound it? Is there a way to reduce the adversarial transferability in order to improve the robustness of an ensemble ML model? To answer these questions, in this work we first theoretically analyze and outline sufficient conditions for adversarial transferability between models; then propose a practical algorithm to reduce the transferability between base models within an ensemble to improve its robustness. Our theoretical analysis shows that only promoting the orthogonality between gradients of base models is not enough to ensure low transferability; in the meantime, the model smoothness is an important factor to control the transferability. We also provide the lower and upper bounds of adversarial transferability under certain conditions. Inspired by our theoretical analysis, we propose an effective Transferability Reduced Smooth (TRS) ensemble training strategy to train a robust ensemble with low transferability by enforcing both gradient orthogonality and model smoothness between base models. We conduct extensive experiments on TRS and compare with 6 state-of-the-art ensemble baselines against 8 whitebox attacks on different datasets, demonstrating that the proposed TRS outperforms all baselines significantly. Linyi Li 0001, Shiliang Zuo, Qian Chen 0003, Benjamin I. P. Rubinstein, Ce Zhang 0001, Bo Li 0026 |
NeurIPS | 5 |
| 2020 | T3: Tree-Autoencoder Constrained Adversarial Text Generation for Targeted AttackabstractAdversarial attacks against natural language processing systems, which perform seemingly innocuous modifications to inputs, can induce arbitrary mistakes to the target models.Though raised great concerns, such adversarial attacks can be leveraged to estimate the robustness of NLP models.Compared with the adversarial example generation in continuous data domain (e.g., image), generating adversarial text that preserves the original meaning is challenging since the text space is discrete and non-differentiable.To handle these challenges, we propose a target-controllable adversarial attack framework T3, which is applicable to a range of NLP tasks.In particular, we propose a tree-based autoencoder to embed the discrete text data into a continuous representation space, upon which we optimize the adversarial perturbation.A novel tree-based decoder is then applied to regularize the syntactic correctness of the generated text and manipulate it on either sentence (T3(SENT)) or word (T3(WORD)) level.We consider two most representative NLP tasks: sentiment analysis and question answering (QA).Extensive experimental results and human studies show that T3 generated adversarial texts can successfully manipulate the NLP models to output the targeted incorrect answer without misleading the human.Moreover, we show that the generated adversarial texts have high transferability which enables the black-box attacks in practice.Our work sheds light on an effective and general way to examine the robustness of NLP models.Our code is publicly available at Boxin Wang, Hengzhi Pei, Boyuan Pan, Qian Chen 0003, Shuohang Wang, Bo Li 0026 |
EMNLP (1) | 4 |
| 2020 | Controllable Time-Delay Transformer for Real-Time Punctuation Prediction and Disfluency DetectionabstractWith the increased applications of automatic speech recognition (ASR) in recent years, it is essential to automatically insert punctuation marks and remove disfluencies in transcripts, to improve the readability of the transcripts as well as the performance of subsequent applications, such as machine translation, dialogue systems, and so forth. In this paper, we propose a Controllable Time-delay Transformer (CT-Transformer) model that jointly completes the punctuation prediction and disfluency detection tasks in real time. The CT-Transformer model facilitates freezing partial outputs with controllable time delay to fulfill the real-time constraints in partial decoding required by subsequent applications. We further propose a fast decoding strategy to minimize latency while maintaining competitive performance. Experimental results on the IWSLT2011 benchmark dataset and an in-house Chinese annotated dataset demonstrate that the proposed approach outperforms the previous state-of-the-art models on F-scores and achieves a competitive inference speed. Qian Chen 0003, Mengzhe Chen, Bo Li 0026, Wen Wang 0001 |
ICASSP | 1 |
| 2020 | Sequential neural networks for noetic end-to-end response selection
Qian Chen 0003, Wen Wang 0001 |
Comput. Speech Lang. | 1 |
| 2019 | Transfer Learning for Context-Aware Spoken Language UnderstandingabstractSpoken language understanding (SLU) is a key component of task-oriented dialogue systems. SLU parses natural language user utterances into semantic frames. Previous work has shown that incorporating context information significantly improves SLU performance for multi-turn dialogues. However, collecting a large-scale human-labeled multi-turn dialogue corpus for the target domains is complex and costly. To reduce dependency on the collection and annotation effort, we propose a Context Encoding Language Transformer (CELT) model facilitating exploiting various context information for SLU. We explore different transfer learning approaches to reduce dependency on data collection and annotation. In addition to unsupervised pre-training using large-scale general purpose unlabeled corpora, such as Wikipedia, we explore unsupervised and supervised adaptive training approaches for transfer learning to benefit from other in-domain and out-of-domain dialogue corpora. Experimental results demonstrate that the proposed model with the proposed transfer learning approaches achieves significant improvement on the SLU performance over state-of-the-art models on two large-scale single-turn dialogue benchmarks and one large-scale multi-turn dialogue benchmark. Qian Chen 0003, Zhu Zhuo, Wen Wang 0001, Qiuyun Xu |
ASRU | 1 |
| 2019 | Sequential Matching Model for End-to-end Multi-turn Response SelectionabstractMulti-turn conversation understanding is an important challenge for building intelligent dialogue systems, and end-to-end multi-turn response selection is one of the major tasks. Previous state-of-the-art models used hierarchy-based (utterance-level and token-level) neural networks to explicitly model the interactions among the different turns' utterances for context modeling. In this paper, we demonstrate that the potentials of sequential matching approaches have not yet been fully exploited in the past for multi-turn response selection. We investigate a sequential matching model based only on chain sequence for multi-turn response selection. The proposed model outperforms all previous models, including previous state-of-the-art hierarchy-based models, and achieves new state-of-the-art performances on two large-scale public multi-turn response selection benchmark datasets. Qian Chen 0003, Wen Wang 0001 |
ICASSP | 1 |
| 2018 | Neural Natural Language Inference Models Enhanced with External KnowledgeabstractModeling natural language inference is a very challenging task.With the availability of large annotated data, it has recently become feasible to train complex models such as neural-network-based inference models, which have shown to achieve the state-of-the-art performance.Although there exist relatively large annotated data, can machines learn all knowledge needed to perform natural language inference (NLI) from these data?If not, how can neural-network-based NLI models benefit from external knowledge and how to build NLI models to leverage it?In this paper, we enrich the state-of-the-art neural natural language inference models with external knowledge.We demonstrate that the proposed models improve neural NLI models to achieve the state-of-the-art performance on the SNLI and MultiNLI datasets. Qian Chen 0003, Xiaodan Zhu 0001, Zhen-Hua Ling, Diana Inkpen, Si Wei |
ACL (1) | 1 |
| 2018 | Enhancing Sentence Embedding with Generalized PoolingabstractPooling is an essential component of a wide variety of sentence representation and embedding models. This paper explores generalized pooling methods to enhance sentence embedding. We propose vector-based multi-head attention that includes the widely used max pooling, mean pooling, and scalar self-attention as special cases. The model benefits from properly designed penalization terms to reduce redundancy in multi-head attention. We evaluate the proposed model on three different tasks: natural language inference (NLI), author profiling, and sentiment classification. The experiments show that the proposed model achieves significant improvement over strong sentence-encoding-based methods, resulting in state-of-the-art performances on four datasets. The proposed approach can be easily implemented for more problems than we discuss in this paper. Qian Chen 0003, Zhen-Hua Ling, Xiaodan Zhu 0001 |
COLING | 1 |
| 2018 | A Sequential Neural Encoder With Latent Structured Description for Modeling SentencesabstractIn this paper, we propose a sequential neural encoder with latent structured description (SNELSD) for modeling sentences. This model introduces latent chunk-level representations into conventional sequential neural encoders, i.e., recurrent neural networks with long short-term memory (LSTM) units, to consider the compositionality of languages in semantic modeling. An SNELSD model has a hierarchical structure that includes a detection layer and a description layer. The detection layer predicts the boundaries of latent word chunks in an input sentence and derives a chunk-level vector for each word. The description layer utilizes modified LSTM units to process these chunk-level vectors in a recurrent manner and produces sequential encoding outputs. These output vectors are further concatenated with word vectors or the outputs of a chain LSTM encoder to obtain the final sentence representation. All the model parameters are learned in an end-to-end manner without a dependency on additional text chunking or syntax parsing. A natural language inference task and a sentiment analysis task are adopted to evaluate the performance of our proposed model. The experimental results demonstrate the effectiveness of the proposed SNELSD model on exploring task-dependent chunking patterns during the semantic modeling of sentences. Furthermore, the proposed method achieves better performance than conventional chain LSTMs and tree-structured LSTMs on both tasks. Yu-Ping Ruan, Qian Chen 0003, Zhen-Hua Ling |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Enhanced LSTM for Natural Language InferenceabstractReasoning and inference are central to human and artificial intelligence.Modeling inference in human language is very challenging.With the availability of large annotated data (Bowman et al., 2015), it has recently become feasible to train neural network based inference models, which have shown to be very effective.In this paper, we present a new state-of-the-art result, achieving the accuracy of 88.6% on the Stanford Natural Language Inference Dataset.Unlike the previous top models that use very complicated network architectures, we first demonstrate that carefully designing sequential inference models based on chain LSTMs can outperform all previous models.Based on this, we further show that by explicitly considering recursive architectures in both local inference modeling and inference composition, we achieve additional improvement.Particularly, incorporating syntactic parsing information contributes to our best result-it further improves the performance even when added to the already very strong model. Qian Chen 0003, Xiaodan Zhu 0001, Zhen-Hua Ling, Si Wei, Hui Jiang 0001, Diana Inkpen |
ACL (1) | 1 |
| 2016 | Distraction-Based Neural Networks for Modeling Document
Qian Chen 0003, Xiaodan Zhu 0001, Zhen-Hua Ling, Si Wei, Hui Jiang 0001 |
IJCAI | 1 |
| 2015 | Revisiting Word Embedding for Contrasting MeaningabstractZhigang Chen, Wei Lin, Qian Chen, Xiaoping Chen, Si Wei, Hui Jiang, Xiaodan Zhu. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Zhigang Chen 0003, Qian Chen 0003, Si Wei, Hui Jiang 0001, Xiaodan Zhu 0001 |
ACL (1) | 3 |
| 2015 | Automatic phrase boundary labeling of speech synthesis database using context-dependent HMMs and n-gram prior distributions
Qian Chen 0003, Zhen-Hua Ling, Chen-Yu Yang, Li-Rong Dai 0001 |
INTERSPEECH | 1 |