VLDB 2026 Research / reviewers in the wild / expert
Wen Wang 0001
dblp:29/4680-1
· DBLP profile ↗
99ranked-venue papers
15as first author
38since 2021 · last 2026
0000-0002-0356-1968ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 71 · 13 first-author · 21 since 2021Artificial intelligence and machine learning · 58 · 7 first-author · 25 since 2021Databases, data management, data science and information retrieval · 2Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language ModelsabstractThe Speaker Diarization and Recognition (SDR) task aims to predict ``who spoke when and what'' within an audio clip, which is a crucial task in various real-world multi-speaker scenarios such as meeting transcription and dialogue systems. Existing SDR systems typically adopt a cascaded framework, combining multiple modules such as speaker diarization (SD) and automatic speech recognition (ASR). The cascaded systems suffer from several limitations, such as error propagation, difficulty in handling overlapping speech, and lack of joint optimization for exploring the synergy between SD and ASR tasks. To address these limitations, we introduce SpeakerLM, a unified multimodal large language model for SDR that jointly performs SD and ASR in an end-to-end manner. Moreover, to facilitate diverse real-world scenarios, we incorporate a flexible speaker registration mechanism into SpeakerLM, enabling SDR under different speaker registration settings. SpeakerLM is progressively developed with a multi-stage training strategy on large-scale real data. Extensive experiments show that SpeakerLM demonstrates strong data scaling capability and generalizability, outperforming state-of-the-art cascaded baselines on both in-domain and out-of-domain public SDR benchmarks. Furthermore, experimental results show that the proposed speaker registration mechanism effectively ensures robust SDR performance of SpeakerLM across diverse speaker registration conditions and varying numbers of registered speakers. Han Yin, Yafeng Chen, Chong Deng, Luyao Cheng, Hui Wang 0030, Chao-Hong Tan, Qian Chen 0003, Wen Wang 0001, Xiangang Li |
AAAI | 8 |
| 2026 | Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration CodingabstractExisting speech tokenizers typically assign a fixed number of tokens per second, regardless of the varying information density or temporal fluctuations in the speech signal. This uniform token allocation mismatches the intrinsic structure of speech, where information is distributed unevenly over time. To address this, we propose VARSTok, a VAriable-frame-Rate Speech Tokenizer that adapts token allocation based on local feature similarity. VARSTok introduces two key innovations: (1) a temporal-aware density peak clustering algorithm that adaptively segments speech into variable-length units, and (2) a novel implicit duration coding scheme that embeds both content and temporal span into a single token index, eliminating the need for auxiliary duration predictors. Extensive experiments show that VARSTok significantly outperforms strong fixed-rate baselines. Notably, it achieves superior reconstruction naturalness while using up to 23% fewer tokens than a 40 Hz fixed-frame-rate baseline. VARSTok further yields lower word error rates and improved naturalness in zero-shot text-to-speech synthesis. To the best of our knowledge, this is the first work to demonstrate that a fully dynamic, variable-frame-rate acoustic speech tokenizer can be seamlessly integrated into downstream speech language models. Rui-Chen Zheng, Wenrui Liu 0003, Hui-Peng Du, Chong Deng, Qian Chen 0003, Wen Wang 0001, Yang Ai, Zhen-Hua Ling |
AAAI | 7 |
| 2026 | Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue ModelsabstractYifu Chen, Shengpeng Ji, Zhengqing Liu, Qian Chen, Wen Wang, Ziqing Wang, Yangzhuo Li, Tianle Liang, Zhou Zhao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shengpeng Ji, Zhengqing Liu, Qian Chen 0003, Wen Wang 0001, Yangzhuo Li, Tianle Liang, Zhou Zhao 0001 |
ACL (1) | 5 |
| 2026 | UniVocal: Unified Speech-Singing Code-Switching SynthesisabstractWe propose UniVocal, a unified framework that implicitly infers vocal modes from text context to pioneer Speech-Singing Code-Switching (SCS) Synthesis-a task where transitions are autonomously driven by textual semantics, akin to seamless human language blending.Unlike single-mode generation or systems relying on switching-control tags, our proposed UniVocal implicitly infers vocal modes solely from text context.To achieve this, we employ a data-efficient two-stage curriculum learning strategy that progressively trains a competitive TTS system to acquire the desired SCS capability.Addressing data scarcity, we introduce a scalable pipeline to synthesize diverse code-switching data that is both semantically and acoustically natural, alongside a new multi-scenario benchmark, SCSBench.To address limitations of semantic tokenizers in capturing acoustic details, we also introduce refined cent token and Chainof-Thought (CoT) generation for planning prosody before content generation, effectively enhancing empathetic speech generation and singing melody.Experimental results demonstrate that UniVocal achieves state-of-the-art performance on SCSBench while maintaining competitive performance on regular speech and singing tasks. Qian Chen 0003, Wen Wang 0001, Xiangang Li, Zhen-Hua Ling, Yang Ai |
ACL (1) | 3 |
| 2026 | GenesisFunc: Multi-Agent Data Generation for Accurate and Generalizable Function-CallingabstractHao-Xiang Xu, Chong Deng, Jiaqing Liu, Wen Wang, Qian Chen, Lujia Bao, Xiangang Li, Zhen-Hua Ling. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hao-Xiang Xu, Chong Deng, Jiaqing Liu, Wen Wang 0001, Qian Chen 0003, Lujia Bao, Xiangang Li, Zhen-Hua Ling |
ACL (1) | 4 |
| 2025 | Recording for Eyes, Not Echoing to Ears: Contextualized Spoken-to-Written Conversion of ASR TranscriptsabstractAutomatic Speech Recognition (ASR) transcripts exhibit recognition errors and various spoken language phenomena such as disfluencies, ungrammatical sentences, and incomplete sentences, hence suffering from poor readability. To improve readability, we propose a Contextualized Spoken-to-Written conversion (CoS2W) task to address ASR and grammar errors and also transfer the informal text into the formal style with content preserved, utilizing contexts and auxiliary information. This task naturally matches the in-context learning capabilities of Large Language Models (LLMs). To facilitate comprehensive comparisons of various LLMs, we construct a document-level Spoken-to-Written conversion of ASR Transcripts Benchmark (SWAB) dataset. Using SWAB, we study the impact of different granularity levels on the CoS2W performance, and propose methods to exploit contexts and auxiliary information to enhance the outputs. Experimental results reveal that LLMs have the potential to excel in the CoS2W task, particularly in grammaticality and formality, our methods achieve effective understanding of contexts and auxiliary information by LLMs. We further investigate the effectiveness of using LLMs as evaluators and find that LLM evaluators show strong correlations with human evaluations on rankings of faithfulness and formality, which validates the reliability of LLM evaluators for the CoS2W task. Jiaqing Liu, Chong Deng, Shilin Zhou 0002, Qian Chen 0003, Wen Wang 0001 |
AAAI | 7 |
| 2025 | CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based VerificationabstractLarge Language Models (LLMs) have made significant progress in code generation, offering developers groundbreaking automated programming support. However, LLMs often generate code that is syntactically correct and even semantically plausible, but may not execute as expected or fulfill specified requirements. This phenomenon of hallucinations in the code domain has not been systematically explored. To advance the community's understanding and research on this issue, we introduce the concept of code hallucinations and propose a classification method for code hallucination based on execution verification. We categorize code hallucinations into four main types: mapping, naming, resource, and logic hallucinations, with each category further divided into different subcategories to understand and address the unique challenges faced by LLMs in code generation with finer granularity. Additionally, we present a dynamic detection algorithm called CodeHalu designed to detect and quantify code hallucinations. We also introduce the CodeHaluEval benchmark, which includes 8,883 samples from 699 tasks, to systematically and quantitatively evaluate code hallucinations. By evaluating 17 popular LLMs using this benchmark, we reveal significant differences in their accuracy and reliability in code generation, offering detailed insights for further improving the code generation capabilities of LLMs. Weixiang Yan, Qian Yang 0007, Xuandong Zhao, Qian Chen 0003, Wen Wang 0001, Dawn Song |
AAAI | 6 |
| 2025 | Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization on Multi-party ConversationabstractLuyao Cheng, Hui Wang, Chong Deng, Siqi Zheng, Yafeng Chen, Rongjie Huang, Qinglin Zhang, Qian Chen, Xihao Li, Wen Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Luyao Cheng, Hui Wang 0030, Chong Deng, Yafeng Chen, Rongjie Huang 0001, Qian Chen 0003, Xihao Li, Wen Wang 0001 |
ACL (1) | 10 |
| 2025 | ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style ControlabstractIn this paper, we present ControlSpeech, a text-to-speech (TTS) system capable of fully cloning the speaker’s voice and enabling arbitrary control and adjustment of speaking style. Prior zero-shot TTS models only mimic the speaker’s voice without further control and adjustment capabilities while prior controllable TTS models cannot perform speaker-specific voice generation. Therefore, ControlSpeech focuses on a more challenging task—a TTS system with controllable timbre, content, and style at the same time. ControlSpeech takes speech prompts, content prompts, and style prompts as inputs and utilizes bidirectional attention and mask-based parallel decoding to capture codec representations corresponding to timbre, content, and style in a discrete decoupling codec space. Moreover, we analyze the many-to-many issue in textual style control and propose the Style Mixture Semantic Density (SMSD) module, which is based on Gaussian mixture density networks, to resolve this problem. To facilitate empirical validations, we make available a new style controllable dataset called VccmDataset. Our experimental results demonstrate that ControlSpeech exhibits comparable or state-of-the-art (SOTA) performance in terms of controllability, timbre similarity, audio quality, robustness, and generalizability. Codes are available at https://github.com/jishengpeng/ControlSpeech. Shengpeng Ji, Qian Chen 0003, Wen Wang 0001, Jialong Zuo, Minghui Fang 0002, Ziyue Jiang 0004, Hai Huang 0013, Zehan Wang 0001, Xize Cheng, Zhou Zhao 0001 |
ACL (1) | 3 |
| 2025 | UniCodec: Unified Audio Codec with Single Domain-Adaptive CodebookabstractYidi Jiang, Qian Chen, Shengpeng Ji, Yu Xi, Wen Wang, Chong Zhang, Xianghu Yue, ShiLiang Zhang, Haizhou Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yidi Jiang, Qian Chen 0003, Shengpeng Ji, Yu Xi, Wen Wang 0001, Chong Zhang 0003, Xianghu Yue, Shiliang Zhang, Haizhou Li 0001 |
ACL (1) | 5 |
| 2025 | OmniFlatten: An End-to-end GPT Model for Seamless Voice ConversationabstractQinglin Zhang, Luyao Cheng, Chong Deng, Qian Chen, Wen Wang, Siqi Zheng, Jiaqing Liu, Hai Yu, Chao-Hong Tan, Zhihao Du, ShiLiang Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Luyao Cheng, Chong Deng, Qian Chen 0003, Wen Wang 0001, Jiaqing Liu, Chao-Hong Tan, Zhihao Du, Shiliang Zhang |
ACL (1) | 5 |
| 2025 | Self-Distillation Prototypes Network: Learning Robust Speaker Representations without SupervisionabstractTraining speaker-discriminative and robust speaker verification systems without explicit speaker labels remains a persistent challenge. In this paper, we propose a novel self-supervised speaker verification approach, Self-Distillation Prototypes Network (SDPN), which effectively facilitates self-supervised speaker representation learning. SDPN assigns the representation of the augmented views of an utterance to the same prototypes as the representation of the original view, thereby enabling effective knowledge transfer between the augmented and original views. Due to lack of negative pairs in the SDPN training process, the network tends to align positive pairs quite closely in the embedding space, a phenomenon known as model collapse. To mitigate this problem, we introduce a diversity regularization term to embeddings in SDPN. Comprehensive experiments on the VoxCeleb datasets demonstrate the superiority of SDPN among self-supervised speaker verification approaches. SDPN sets a new state-of-the-art on the VoxCeleb1 speaker verification evaluation benchmark, achieving Equal Error Rate 1.80%, 1.99%, and 3.62% for trial VoxCeleb1-O, VoxCeleb1-E and VoxCeleb1H respectively1, without using any speaker labels in training. Ablation studies show that both proposed learnable prototypes in self-distillation network and diversity regularization contribute to the verification performance. Yafeng Chen, Hui Wang 0030, Luyao Cheng, Qian Chen 0003, Chong Deng, Shiliang Zhang, Wen Wang 0001 |
ICASSP | 8 |
| 2025 | 3D-Speaker-Toolkit: An Open-Source Toolkit for Multimodal Speaker Verification and DiarizationabstractWe introduce 3D-Speaker-Toolkit, an open-source toolkit for multimodal speaker verification and diarization, designed for meeting the needs of academic researchers and industrial practitioners. The 3D-Speaker-Toolkit adeptly leverages the combined strengths of acoustic, semantic, and visual data, seamlessly fusing these modalities to offer robust speaker recognition capabilities. The acoustic module extracts speaker embeddings from acoustic features, employing both fully-supervised and self-supervised learning approaches. The semantic module leverages advanced language models to comprehend the substance and context of spoken language, thereby augmenting the system’s proficiency in distinguishing speakers through linguistic patterns. The visual module applies image processing technologies to scrutinize facial features, which bolsters the precision of speaker diarization in multi-speaker environments. Collectively, these modules empower the 3D-Speaker-Toolkit to achieve substantially improved accuracy and reliability in speaker-related tasks. With 3D-Speaker-Toolkit, we establish a new benchmark for multimodal speaker analysis. The toolkit also includes a handful of open-source state-of-the-art models and a large-scale dataset containing over 10,000 speakers. The toolkit is publicly available at https://github.com/modelscope/3D-Speaker. Yafeng Chen, Hui Wang 0030, Luyao Cheng, Tinglong Zhu, Rongjie Huang 0001, Chong Deng, Qian Chen 0003, Shiliang Zhang, Wen Wang 0001, Xihao Li |
ICASSP | 10 |
| 2025 | Unified Audio Event DetectionabstractSound Event Detection (SED) detects regions of sound events, while Speaker Diarization (SD) segments speech conversations attributed to individual speakers. In SED, all speaker segments are classified as a single speech event, while in SD, non-speech sounds are treated merely as background noise. Thus, both tasks provide only partial analysis in complex audio scenarios involving both speech conversation and non-speech sounds. In this paper, we introduce a novel task called Unified Audio Event Detection (UAED) for comprehensive audio analysis. UAED explores the synergy between SED and SD tasks, simultaneously detecting non-speech sound events and fine-grained speech events based on speaker identities. To tackle this task, we propose a Transformer-based UAED (T-UAED) framework and construct the UAED Data derived from the Librispeech dataset and DESED soundbank. Experiments demonstrate that the proposed framework effectively exploits task interactions and substantially outperforms the baseline that simply combines the outputs of SED and SD models. T-UAED also shows its versatility by performing comparably to specialized models for individual SED and SD tasks on DESED and CALLHOME datasets. Yidi Jiang, Ruijie Tao, Qian Chen 0003, Wen Wang 0001 |
ICASSP | 5 |
| 2025 | WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language ModelingabstractLanguage models have been effectively applied to modeling natural signals, such as images, video, speech, and audio. A crucial component of these models is the codec tokenizer, which compresses high-dimensional natural signals into lower-dimensional discrete tokens. In this paper, we introduce WavTokenizer, which offers several advantages over previous SOTA acoustic codec models in the audio domain: 1) extreme compression. By compressing the layers of quantizers and the temporal dimension of the discrete codec, one-second audio of 24kHz sampling rate requires only a single quantizer with 40 or 75 tokens. 2) improved subjective quality. Despite the reduced number of tokens, WavTokenizer achieves state-of-the-art reconstruction quality with outstanding UTMOS scores and inherently contains richer semantic information. Specifically, we achieve these results by designing a broader VQ space, extended contextual windows, and improved attention networks, as well as introducing a powerful multi-scale discriminator and an inverse Fourier transform structure. We conducted extensive reconstruction experiments in the domains of speech, audio, and music. WavTokenizer exhibited strong performance across various objective and subjective metrics compared to state-of-the-art models. We also tested semantic information, VQ utilization, and adaptability to generative models. Comprehensive ablation studies confirm the necessity of each module in WavTokenizer. The code is available at https://github.com/jishengpeng/WavTokenizer. Shengpeng Ji, Ziyue Jiang 0001, Wen Wang 0001, Minghui Fang 0002, Jialong Zuo, Qian Yang 0006, Xize Cheng, Zehan Wang 0001, Ruiqi Li 0002, Xiaoda Yang, Rongjie Huang 0001, Yidi Jiang, Qian Chen 0003, Zhou Zhao 0001 |
ICLR | 3 |
| 2025 | OmniAudio: Generating Spatial Audio from 360-Degree VideoabstractTraditional video-to-audio generation techniques primarily focus on perspective video and non-spatial audio, often missing the spatial cues necessary for accurately representing sound sources in 3D environments. To address this limitation, we introduce a novel task, 360V2SA, to generate spatial audio from 360-degree videos, specifically producing First-order Ambisonics (FOA) audio - a standard format for representing 3D spatial audio that captures sound directionality and enables realistic 3D audio reproduction. We first create Sphere360, a novel dataset tailored for this task that is curated from real-world data. We also design an efficient semi-automated pipeline for collecting and cleaning paired video-audio data. To generate spatial audio from 360-degree video, we propose a novel framework OmniAudio, which leverages self-supervised pre-training using both spatial audio data (in FOA format) and large-scale non-spatial data. Furthermore, OmniAudio features a dual-branch framework that utilizes both panoramic and perspective video inputs to capture comprehensive local and global information from 360-degree videos. Experimental results demonstrate that OmniAudio achieves state-of-the-art performance across both objective and subjective metrics on Sphere360. Code and datasets are available at https://github.com/liuhuadai/OmniAudio. The project website is available at https://OmniAudio-360V2SA.github.io. Huadai Liu, Tianyi Luo, Kaicheng Luo, Qikai Jiang, Peiwen Sun, Rongjie Huang 0001, Qian Chen 0003, Wen Wang 0001, Xiangtai Li, Shiliang Zhang, Zhijie Yan, Zhou Zhao 0001, Wei Xue 0002 |
ICML | 9 |
| 2025 | Pushing the Frontiers of Self-Distillation Prototypes Network with Dimension Regularization and Score Normalization
Yafeng Chen, Chong Deng, Hui Wang 0030, Yiheng Jiang, Han Yin, Qian Chen 0003, Wen Wang 0001 |
INTERSPEECH | 7 |
| 2025 | ThinkSound: Chain-of-Thought Reasoning in Multimodal LLMs for Audio Generation and EditingabstractWhile end-to-end video-to-audio generation has greatly improved, producing high-fidelity audio that authentically captures the nuances of visual content remains challenging. Like professionals in the creative industries, this generation requires sophisticated reasoning about items such as visual dynamics, acoustic environments, and temporal relationships. We present **ThinkSound**, a novel framework that leverages Chain-of-Thought (CoT) reasoning to enable stepwise, interactive audio generation and editing for videos. Our approach decomposes the process into three complementary stages: foundational foley generation that creates semantically coherent soundscapes, interactive object-centric refinement through precise user interactions, and targeted editing guided by natural language instructions. At each stage, a multimodal large language model generates contextually aligned CoT reasoning that guides a unified audio foundation model. Furthermore, we introduce **AudioCoT**, a comprehensive dataset with structured reasoning annotations that establishes connections between visual content, textual descriptions, and sound synthesis. Experiments demonstrate that ThinkSound achieves state-of-the-art performance in video-to-audio generation across both audio metrics and CoT metrics, and excels in the out-of-distribution Movie Gen Audio benchmark. The project page is available at https://ThinkSound-Project.github.io. Huadai Liu, Kaicheng Luo, Wen Wang 0001, Qian Chen 0003, Zhou Zhao 0001, Wei Xue 0002 |
NeurIPS | 4 |
| 2025 | ClinicalLab: Aligning Agents for Multi-Departmental Clinical Diagnostics in the Real WorldabstractLarge language models (LLMs) have achieved significant performance progress in various natural language processing applications. However, LLMs still struggle to meet the strict requirements for accuracy and reliability in the medical field and face many challenges in clinical applications. Existing clinical diagnostic evaluation benchmarks for evaluating medical agents powered by LLMs have severe limitations. To address these limitations, we introduce ClinicalLab, a comprehensive clinical diagnosis agent alignment suite. ClinicalLab includes ClinicalBench, an end-to-end multi-departmental clinical diagnostic evaluation benchmark for evaluating medical agents and LLMs. ClinicalBench is based on real cases that cover 24 departments and 150 diseases. We ensure that ClinicalBench does not have data leakage. ClinicalLab also includes four novel metrics (ClinicalMetrics) for evaluating the effectiveness of LLMs in clinical diagnostic tasks. We evaluate 17 general and medical-domain LLMs and find that their performance varies significantly across different departments. Based on these findings, in ClinicalLab, we propose ClinicalAgent, an end-to-end clinical agent that aligns with real-world clinical diagnostic practices. We systematically investigate the performance and applicable scenarios of variants of ClinicalAgent on ClinicalBench. Our findings demonstrate the importance of aligning with modern medical practices in designing medical agents. Weixiang Yan, Tengxiao Wu, Qian Chen 0003, Wen Wang 0001, Haoyuan Chai |
NeurIPS | 5 |
| 2024 | CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and GenerationabstractWeixiang Yan, Haitian Liu, Yunkun Wang, Yunzhe Li, Qian Chen, Wen Wang, Tingyu Lin, Weishan Zhao, Li Zhu, Hari Sundaram, Shuiguang Deng. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Weixiang Yan, Yunkun Wang, Yunzhe Li 0001, Qian Chen 0003, Wen Wang 0001, Tingyu Lin 0002, Weishan Zhao, Hari Sundaram, Shuiguang Deng |
ACL (1) | 6 |
| 2024 | Advancing Precise Outline-Conditioned Text Generation with Task Duality and Explicit Outline ControlabstractYunzhe Li, Qian Chen, Weixiang Yan, Wen Wang, Qinglin Zhang, Hari Sundaram. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yunzhe Li 0001, Qian Chen 0003, Weixiang Yan, Wen Wang 0001, Hari Sundaram |
EACL (1) | 4 |
| 2024 | Loss Masking Is Not Needed In Decoder-Only Transformer For Discrete-Token-Based ASRabstractRecently, unified speech-text models, such as SpeechGPT, VioLA, and AudioPaLM, have achieved remarkable performance on various speech tasks. These models discretize speech signals into tokens (speech discretization) and use a shared vocabulary for both text and speech tokens. Then they train a single decoder-only Transformer on a mixture of speech tasks. However, these models rely on the Loss Masking strategy for the ASR task, which ignores the dependency among speech tokens. In this paper, we propose to model speech tokens in an autoregressive way, similar to text. We find that applying the conventional cross-entropy loss on input speech tokens does not consistently improve the ASR performance over the Loss Masking approach. To address this issue, we propose a novel approach denoted Smoothed Label Distillation (SLD), which applies a KL divergence loss with smoothed labels on speech tokens. Our experiments show that SLD effectively models speech tokens and outperforms Loss Masking for decoder-only Transformers in ASR tasks with different speech discretization methods1. Qian Chen 0003, Wen Wang 0001, Shiliang Zhang, Chong Deng, Jiaqing Liu, Chong Zhang 0003 |
ICASSP | 2 |
| 2024 | Tuning Large Language Model for Speech Recognition With Mixed-Scale Re-TokenizationabstractLarge Language Models (LLMs) have proven successful across a spectrum of speech-related tasks, such as speech recognition, text-to-speech, and spoken language understanding. Recently, the use of discretized speech features has gained attention as an efficient and compatible alternative to continuous features for LLMs. This is mainly due to their reduced storage requirements and better alignment of these features with LLM's input space. However, the typical practice of freezing the speech encoder during training poses challenges in bridging the modality gap between speech and text. To address this, we propose to use a mixed-scale re-tokenization layer, integrating multiple granularities in discretized speech features directly within the LLM's input module. Our experimental results demonstrated that the proposed method can effectively enhance the performance of ASR in the setting of continuous learning of an LLM, highlighting the importance of a meticulously designed input module for the integration of discretized speech features with an LLM. Chong Zhang 0003, Qian Chen 0003, Wen Wang 0001, Bin Ma 0001 |
IEEE Signal Process. Lett. | 4 |
| 2023 | Ditto: A Simple and Efficient Approach to Improve Sentence EmbeddingsabstractPrior studies diagnose the anisotropy problem in sentence representations from pre-trained language models, e.g., BERT, without finetuning.Our analysis reveals that the sentence embeddings from BERT suffer from a bias towards uninformative words, limiting the performance in semantic textual similarity (STS) tasks.To address this bias, we propose a simple and efficient unsupervised approach, Diagonal Attention Pooling (Ditto), which weights words with model-based importance estimations and computes the weighted average of word representations from pre-trained models as sentence embeddings.Ditto can be easily applied to any pre-trained language model as a postprocessing operation.Compared to prior sentence embedding approaches, Ditto does not add parameters nor requires any learning.Empirical evaluations demonstrate that our proposed Ditto can alleviate the anisotropy problem and improve various pre-trained models on the STS benchmarks. 1 Qian Chen 0003, Wen Wang 0001, Chong Deng, Jiaqing Liu, Chong Zhang 0003 |
EMNLP | 2 |
| 2023 | Improving Long Document Topic Segmentation Models With Enhanced Coherence ModelingabstractTopic segmentation is critical for obtaining structured documents and improving downstream tasks such as information retrieval.Due to its ability of automatically exploring clues of topic shift from abundant labeled data, recent supervised neural models have greatly promoted the development of long document topic segmentation, but leaving the deeper relationship between coherence and topic segmentation underexplored.Therefore, this paper enhances the ability of supervised models to capture coherence from both logical structure and semantic similarity perspectives to further improve the topic segmentation performance, proposing Topic-aware Sentence Structure Prediction (TSSP) and Contrastive Semantic Similarity Learning (CSSL).Specifically, the TSSP task is proposed to force the model to comprehend structural information by learning the original relations between adjacent sentences in a disarrayed document, which is constructed by jointly disrupting the original document at topic and sentence levels.Moreover, we utilize inter-and intra-topic information to construct contrastive samples and design the CSSL objective to ensure that the sentences representations in the same topic have higher similarity, while those in different topics are less similar.Extensive experiments show that Longformer with our approach significantly outperforms state-of-the-art (SOTA) methods.Our approach improves F 1 of SOTA by 3.42 (73.74 → 77.16) and improves P k by 1.11 points (15.0 → 13.89) on WIKI-727K and achieves an average relative reduction of 4.3% on P k on WikiSection.The average relative P k drop of 8.38% on two out-of-domain datasets also demonstrates the robustness of our approach 1 . Chong Deng, Jiaqing Liu, Qian Chen 0003, Wen Wang 0001 |
EMNLP | 6 |
| 2023 | Meeting Action Item Detection with Regularized Context ModelingabstractMeetings are increasingly important for collaborations. Action items in meeting transcripts are crucial for managing post-meeting to-do tasks, which usually are summarized laboriously. The Action Item Detection task aims to automatically detect meeting content associated with action items. However, datasets manually annotated with action item detection labels are scarce and in small scale. We construct and release the first Chinese meeting corpus with manual action item annotations1. In addition, we propose a Context-Drop approach to utilize both local and global contexts by contrastive learning, and achieve better accuracy and robustness for action item detection. We also propose a Lightweight Model Ensemble method to exploit different pre-trained models2. Experimental results on our Chinese meeting corpus and the English AMI corpus demonstrate the effectiveness of the proposed approaches. Jiaqing Liu, Chong Deng, Qian Chen 0003, Wen Wang 0001 |
ICASSP | 5 |
| 2023 | Auxiliary Pooling Layer For Spoken Language UnderstandingabstractEnd-to-end spoken language understanding requires speech data annotated with semantic information and may suffer from the shortage of annotated data. Recent progresses leverage unlabelled speech data to pre-train a speech encoder. However, it remains a challenge for the pre-trained speech encoder to encode semantic information. Existing works explore transferring knowledge from a pre-trained text model with different alignment losses at a fixed granularity. In this paper, we address the variable granularity in transferring knowledge from texts to speech representation via APLY, an auxiliary pooling layer, that fuses the global information with the adaptively encoded local context. We demonstrate the effectiveness of APLY on three benchmarks of spoken language understanding. Trung Hieu Nguyen 0001, Jinjie Ni, Wen Wang 0001, Qian Chen 0003, Chong Zhang 0003, Bin Ma 0001 |
ICASSP | 4 |
| 2023 | Adaptive Knowledge Distillation Between Text and Speech Pre-Trained ModelsabstractLearning on a massive amount of speech corpus leads to the recent success of many self-supervised speech models. With knowledge distillation, these models may also benefit from the knowledge encoded by language models that are pre-trained on rich sources of texts. The distillation process, however, is challenging due to the modal disparity between textual and speech embedding spaces. This paper studies metric-based distillation to align the embedding space of text and speech with only a small amount of data without modifying the model structure. Since the semantic and granularity gap between text and speech has been omitted in literature, which impairs the distillation, we propose the Prior-informed Adaptive knowledge Distillation (PAD) that adaptively leverages text/speech units of variable granularity and prior distributions to achieve better global and local alignments between text and speech pre-trained models. We evaluate on three spoken language understanding benchmarks to show that PAD is more effective in transferring linguistic knowledge than other metric-based distillation approaches. Jinjie Ni, Wen Wang 0001, Qian Chen 0033, Dianwen Ng, Han Lei, Trung Hieu Nguyen 0001, Chong Zhang 0003, Bin Ma 0001, Erik Cambria |
ICASSP | 3 |
| 2023 | Weighted Sampling for Masked Language ModelingabstractMasked Language Modeling (MLM) is widely used to pretrain language models. The standard random masking strategy in MLM causes the pre-trained language models (PLMs) to be biased towards high-frequency tokens. Representation learning of rare tokens is poor and PLMs have limited performance on downstream tasks. To alleviate this frequency bias issue, we propose two simple and effective Weighted Sampling strategies for masking tokens based on token frequency and training loss. We apply these two strategies to BERT and obtain Weighted-Sampled BERT (WSBERT). Experiments on the Semantic Textual Similarity benchmark (STS) show that WSBERT significantly improves sentence embeddings over BERT. Combining WSBERT with calibration methods and prompt learning further improves sentence embeddings. We also investigate fine-tuning WSBERT on the GLUE benchmark and show that Weighted Sampling also improves the transfer learning capability of the backbone PLM. We further analyze and provide insights into how WSBERT improves token embeddings. Linhan Zhang, Qian Chen 0003, Wen Wang 0001, Chong Deng, Xin Cao 0001, Kongzhang Hao, Wei Wang 0011 |
ICASSP | 3 |
| 2023 | Overview of the ICASSP 2023 General Meeting Understanding and Generation Challenge (MUG)abstractICASSP2023 General Meeting Understanding and Generation Challenge (MUG) focuses on prompting a wide range of spoken language processing (SLP) research on meeting transcripts, as SLP applications are critical to improve users’ efficiency in grasping important information in meetings. MUG includes five tracks, including topic segmentation, topic-level and session-level extractive summarization, topic title generation, keyphrase extraction, and action item detection. To facilitate MUG, we construct and release a large-scale meeting dataset, the AliMeeting4MUG Corpus. We review the dataset, track settings and baselines, and summarize the challenge results and major techniques used in the submissions. Chong Deng, Jiaqing Liu, Qian Chen 0003, Wen Wang 0001, Zhijie Yan, Jinglin Liu, Yi Ren 0006, Zhou Zhao 0001 |
ICASSP | 6 |
| 2023 | MUG: A General Meeting Understanding and Generation BenchmarkabstractListening to long video/audio recordings from video conferencing and online courses for acquiring information is extremely inefficient. Even after ASR systems transcribe recordings into long-form spoken language documents, reading ASR transcripts only partly speeds up seeking information. It has been observed that a range of NLP applications, such as keyphrase extraction, topic segmentation, and summarization, significantly improve users’ efficiency in grasping important information. The meeting scenario is among the most valuable scenarios for deploying these spoken language processing (SLP) capabilities. However, the lack of large-scale public meeting datasets annotated for these SLP tasks severely hinders their advancement. To prompt SLP advancement, we establish a large-scale general Meeting Understanding and Generation Benchmark (MUG) to benchmark the performance of a wide range of SLP tasks, including topic segmentation, topic-level and session-level extractive summarization and topic title generation, keyphrase extraction, and action item detection. To facilitate the MUG benchmark, we construct and release a large-scale meeting dataset for comprehensive long-form SLP development, the AliMeeting4MUG Corpus, which consists of 654 recorded Mandarin meeting sessions with diverse topic coverage, with manual annotations for SLP tasks on manual transcripts of meeting recordings. To the best of our knowledge, the AliMeeting4MUG Corpus is so far the largest meeting corpus in scale and facilitates most SLP tasks. In this paper, we provide a detailed introduction of this corpus, SLP tasks and evaluation methods, baseline systems and their performance1. Chong Deng, Jiaqing Liu, Qian Chen 0003, Wen Wang 0001, Zhijie Yan, Jinglin Liu, Yi Ren 0006, Zhou Zhao 0001 |
ICASSP | 6 |
| 2023 | Adapter-tuning with Effective Token-dependent Representation Shift for Automatic Speech Recognition
Dianwen Ng, Chong Zhang 0003, Ruixi Zhang, Trung Hieu Nguyen 0001, Chongjia Ni, Shengkui Zhao, Qian Chen 0003, Wen Wang 0001, Chng Eng Siong, Bin Ma 0001 |
INTERSPEECH | 9 |
| 2023 | Enhancing Multi-modal Multi-hop Question Answering via Structured Knowledge and Unified Retrieval-GenerationabstractMulti-modal multi-hop question answering involves answering a question by reasoning over multiple input sources from different modalities. Existing methods often retrieve evidences separately and then use a language model to generate an answer based on the retrieved evidences, and thus do not adequately connect candidates and are unable to model the interdependent relations during retrieval. Moreover, the pipelined approaches of retrieval and generation might result in poor generation performance when retrieval performance is low. To address these issues, we propose a Structured Knowledge and Unified Retrieval-Generation (SKURG) approach. SKURG employs an Entity-centered Fusion Encoder to align sources from different modalities using shared entities. It then uses a unified Retrieval-Generation Decoder to integrate intermediate retrieval results for answer generation and also adaptively determine the number of retrieval steps. Extensive experiments on two representative multi-modal multi-hop QA datasets MultimodalQA and WebQA demonstrate that SKURG outperforms the state-of-the-art models in both source retrieval and answer generation performance with fewer parameters1. Qian Yang 0007, Qian Chen 0003, Wen Wang 0001, Baotian Hu, Min Zhang 0005 |
ACM Multimedia | 3 |
| 2022 | PoNet: Pooling Network for Efficient Token Mixing in Long Sequences
Chao-Hong Tan, Qian Chen 0003, Wen Wang 0001, Zhen-Hua Ling |
ICLR | 3 |
| 2021 | Sequence Model with Self-Adaptive Sliding Window for Efficient Spoken Document SegmentationabstractTranscripts generated by automatic speech recognition (ASR) systems for spoken documents lack structural annotations such as paragraphs, significantly reducing their readability. Automatically predicting paragraph segmentation for spoken documents may both improve readability and downstream NLP performance such as summarization and machine reading comprehension. We propose a sequence model with self-adaptive sliding window for accurate and efficient paragraph segmentation. We also propose an approach to exploit pho-netic information, which significantly improves robustness of spoken document segmentation to ASR errors. Evaluations are conducted on the English Wiki-727K document seg-mentation benchmark, a Chinese Wikipedia-based document segmentation dataset we created, and an in-house Chinese spoken document dataset. Our proposed model outperforms the state-of-the-art (SOTA) model based on the same BERT-Base, increasing segmentation F1 on the English benchmark by 4.2 points and on Chinese datasets by 4.3-10.1 points, while reducing inference time to less than 1/6 of inference time of the current SOTA. Qian Chen 0003, Yali Li 0001, Jiaqing Liu, Wen Wang 0001 |
ASRU | 5 |
| 2021 | Discriminative Self-Training for Punctuation PredictionabstractPunctuation prediction for automatic speech recognition (ASR) output transcripts plays a crucial role for improving the readability of the ASR transcripts and for improving the performance of downstream natural language processing applications. However, achieving good performance on punctuation prediction often requires large amounts of labeled speech transcripts, which is expensive and laborious. In this paper, we propose a Discriminative Self-Training approach with weighted loss and discriminative label smoothing to exploit unlabeled speech transcripts. Experimental results on the English IWSLT2011 benchmark test set and an internal Chinese spoken language dataset demonstrate that the proposed approach achieves significant improvement on punctuation prediction accuracy over strong baselines including BERT, RoBERTa, and ELECTRA models. The proposed Discriminative Self-Training approach outperforms the vanilla self-training approach. We establish a new state-of-the-art (SOTA) on the IWSLT2011 test set, outperforming the current SOTA model by 1.3% absolute gain on F$_1$. Qian Chen 0003, Wen Wang 0001, Mengzhe Chen |
Interspeech | 2 |
| 2021 | Pre-Training for Spoken Language Understanding with Joint Textual and Phonetic Representation LearningabstractIn the traditional cascading architecture for spoken language understanding (SLU), it has been observed that automatic speech recognition errors could be detrimental to the performance of natural language understanding. End-to-end (E2E) SLU models have been proposed to directly map speech input to desired semantic frame with a single model, hence mitigating ASR error propagation. Recently, pre-training technologies have been explored for these E2E models. In this paper, we propose a novel joint textual-phonetic pre-training approach for learning spoken language representations, aiming at exploring the full potentials of phonetic information to improve SLU robustness to ASR errors. We explore phoneme labels as high-level speech features, and design and compare pre-training tasks based on conditional masked language model objectives and inter-sentence relation objectives. We also investigate the efficacy of combining textual and phonetic information during fine-tuning. Experimental results on spoken language understanding benchmarks, Fluent Speech Commands and SNIPS, show that the proposed approach significantly outperforms strong baseline models and improves robustness of spoken language understanding to ASR errors. Qian Chen 0003, Wen Wang 0001 |
Interspeech | 2 |
| 2021 | Reducing BERT Computation by Padding Removal and Curriculum LearningabstractBERT is very computationally expensive, which is a hurdle for its training and deployment. This work focuses on removing the unnecessary computation due to input padding in BERT. The input of BERT consists of two concatenated sentences. If the length of the two concatenated sentences is shorter than the maximum sequence length, padding must be added to the end of the sentences to fill the empty slots in the input. Because the lengths of sentences vary greatly, there can be a large amount of padding in input. For the English Wikipedia & BooksCorpus dataset, the percentage of padding among all the input tokens is 17% and 48%, respectively, when the max sequence length is set to 128 and 512. For the Chinese Wikipedia dataset, this percentage is 35% and 79%, respectively, when the max sequence length is 128 and 512. For SQuAD-v1.1 [2], padding accounts for 54% of the total input tokens when the max sequence length is 384. Thus, there is a lot of wasted computation on padding, which significantly increases the training and inference time. Wei Zhang 0127, Wei Wei 0021, Wen Wang 0001, Lingling Jin, Zheng Cao 0003 |
ISPASS | 3 |
| 2020 | Controllable Time-Delay Transformer for Real-Time Punctuation Prediction and Disfluency DetectionabstractWith the increased applications of automatic speech recognition (ASR) in recent years, it is essential to automatically insert punctuation marks and remove disfluencies in transcripts, to improve the readability of the transcripts as well as the performance of subsequent applications, such as machine translation, dialogue systems, and so forth. In this paper, we propose a Controllable Time-delay Transformer (CT-Transformer) model that jointly completes the punctuation prediction and disfluency detection tasks in real time. The CT-Transformer model facilitates freezing partial outputs with controllable time delay to fulfill the real-time constraints in partial decoding required by subsequent applications. We further propose a fast decoding strategy to minimize latency while maintaining competitive performance. Experimental results on the IWSLT2011 benchmark dataset and an in-house Chinese annotated dataset demonstrate that the proposed approach outperforms the previous state-of-the-art models on F-scores and achieves a competitive inference speed. Qian Chen 0003, Mengzhe Chen, Bo Li 0026, Wen Wang 0001 |
ICASSP | 4 |
| 2020 | Sequential neural networks for noetic end-to-end response selection
Qian Chen 0003, Wen Wang 0001 |
Comput. Speech Lang. | 2 |
| 2019 | Transfer Learning for Context-Aware Spoken Language UnderstandingabstractSpoken language understanding (SLU) is a key component of task-oriented dialogue systems. SLU parses natural language user utterances into semantic frames. Previous work has shown that incorporating context information significantly improves SLU performance for multi-turn dialogues. However, collecting a large-scale human-labeled multi-turn dialogue corpus for the target domains is complex and costly. To reduce dependency on the collection and annotation effort, we propose a Context Encoding Language Transformer (CELT) model facilitating exploiting various context information for SLU. We explore different transfer learning approaches to reduce dependency on data collection and annotation. In addition to unsupervised pre-training using large-scale general purpose unlabeled corpora, such as Wikipedia, we explore unsupervised and supervised adaptive training approaches for transfer learning to benefit from other in-domain and out-of-domain dialogue corpora. Experimental results demonstrate that the proposed model with the proposed transfer learning approaches achieves significant improvement on the SLU performance over state-of-the-art models on two large-scale single-turn dialogue benchmarks and one large-scale multi-turn dialogue benchmark. Qian Chen 0003, Zhu Zhuo, Wen Wang 0001, Qiuyun Xu |
ASRU | 3 |
| 2019 | Sequential Matching Model for End-to-end Multi-turn Response SelectionabstractMulti-turn conversation understanding is an important challenge for building intelligent dialogue systems, and end-to-end multi-turn response selection is one of the major tasks. Previous state-of-the-art models used hierarchy-based (utterance-level and token-level) neural networks to explicitly model the interactions among the different turns' utterances for context modeling. In this paper, we demonstrate that the potentials of sequential matching approaches have not yet been fully exploited in the past for multi-turn response selection. We investigate a sequential matching model based only on chain sequence for multi-turn response selection. The proposed model outperforms all previous models, including previous state-of-the-art hierarchy-based models, and achieves new state-of-the-art performances on two large-scale public multi-turn response selection benchmark datasets. Qian Chen 0003, Wen Wang 0001 |
ICASSP | 2 |
| 2018 | Articulatory Information and Multiview Features for Large Vocabulary Continuous Speech RecognitionabstractThis paper explores the use of multi-view features and their discriminative transforms in a convolutional deep neural network (CNN) architecture for a continuous large vocabulary speech recognition task. Mel-filterbank energies and perceptually motivated forced damped oscillator coefficient (DOC) features are used after feature-space maximum-likelihood linear regression (fMLLR) transforms, which are combined and fed as a multi-view feature to a single CNN acoustic model. Use of multi-view feature representation demonstrated significant reduction in word error rates (WERs) compared to the use of individual features by themselves. In addition, when articulatory information was used as an additional input to a fused deep neural network (DNN) and CNN acoustic model, it was found to demonstrate further reduction in WER for the Switchboard subset and the CallHome subset (containing partly non-native accented speech) of the NIST 2000 conversational telephone speech test set, reducing the error rate by 12% relative to the baseline in both cases. This work shows that multi-view features in association with articulatory information can improve speech recognition robustness to spontaneous and non-native speech. Vikramjit Mitra, Wen Wang 0001, Chris Bartels, Horacio Franco, Dimitra Vergyri |
ICASSP | 2 |
| 2017 | Joint modeling of articulatory and acoustic spaces for continuous speech recognition tasksabstractArticulatory information can effectively model variability in speech and can improve speech recognition performance under varying acoustic conditions. Learning speaker-independent articulatory models has always been challenging, as speaker-specific information in the articulatory and acoustic spaces increases the complexity of the speech-to-articulatory space inverse modeling, which is already an ill-posed problem due to its inherent nonlinearity and non-uniqueness. This paper investigates using deep neural networks (DNN) and convolutional neural networks (CNNs) for mapping speech data into its corresponding articulatory space. Our results indicate that the CNN models perform better than their DNN counterparts for speech inversion. In addition, we used the inverse models to generate articulatory trajectories from speech for three different standard speech recognition tasks. To effectively model the articulatory features' temporal modulations while retaining the acoustic features' spatiotemporal signatures, we explored a joint modeling strategy to simultaneously learn both the acoustic and articulatory spaces. The results from multiple speech recognition tasks indicate that articulatory features can improve recognition performance when the acoustic and articulatory spaces are jointly learned with one common objective function. Vikramjit Mitra, Ganesh Sivaraman, Chris Bartels, Hosung Nam, Wen Wang 0001, Carol Y. Espy-Wilson, Dimitra Vergyri, Horacio Franco |
ICASSP | 5 |
| 2016 | Fusion Strategies for Robust Speech Recognition and Keyword Spotting for Channel- and Noise-Degraded Speech
Vikramjit Mitra, Julien van Hout, Wen Wang 0001, Chris Bartels, Horacio Franco, Dimitra Vergyri, Abeer Alwan, Adam Janin, John H. L. Hansen, Richard M. Stern, Abhijeet Sangwan, Nelson Morgan |
INTERSPEECH | 3 |
| 2016 | Toward human-assisted lexical unit discovery without text resourcesabstractThis work addresses lexical unit discovery for languages without (usable) written resources. Previous work has addressed this problem using entirely unsupervised methodologies. Our approach in contrast investigates the use of linguistic and speaker knowledge which are often available even if text resources are not. We create a framework that benefits from such resources, not assuming orthographic representations and avoiding generation of word-level transcriptions. We adapt a universal phone recognizer to the target language and use it to convert audio into a searchable phone string for lexical unit discovery via fuzzy sub-string matching. Linguistic knowledge is used to constrain phone recognition output and to constrain lexical unit discovery on the phone recognizer output. Chris Bartels, Wen Wang 0001, Vikramjit Mitra, Colleen Richey, Andreas Kathol, Dimitra Vergyri, Harry Bratt, Chiachi Hung |
SLT | 2 |
| 2015 | Improving robustness against reverberation for automatic speech recognitionabstractReverberation is a phenomenon observed in almost all enclosed environments. Human listeners rarely experience problems in comprehending speech in reverberant environments, but automatic speech recognition (ASR) systems often suffer increased error rates under such conditions. In this work, we explore the role of robust acoustic features motivated by human speech perception studies, for building ASR systems robust to reverberation effects. Using the dataset distributed for the "Automatic Speech Recognition In Reverberant Environments" (ASpIRE-2015) challenge organized by IARPA, we explore Gaussian mixture models (GMMs), deep neural nets (DNNs) and convolutional deep neural networks (CDNN) as candidate acoustic models for recognizing continuous speech in reverberant environments. We demonstrate that DNN-based systems trained with robust features offer significant reduction in word error rates (WERs) compared to systems trained with baseline mel-filterbank features. We present a novel time-frequency convolution neural net (TFCNN) framework that performs convolution on the feature space across both the time and frequency scales, which we found to consistently outperform the CDNN systems for all feature sets across all testing conditions. Finally, we show that further WER reduction is achievable through system fusion of n-best lists from multiple systems. Vikramjit Mitra, Julien van Hout, Wen Wang 0001, Martin Graciarena, Mitchell McLaren, Horacio Franco, Dimitra Vergyri |
ASRU | 3 |
| 2015 | Name-aware language model adaptation and sparse features for statistical machine translationabstractWe propose approaches improving statistical machine translation (SMT) performance, by developing name-aware language model adaptations and sparse features, in addition to extracting name-aware translation grammar and rules, adding name phrase table, and name translation driven decoding. Chinese-English translation experiments showed that our proposed approaches produce an absolute gain of +2.3 BLEU on top of our previous high-performing, name-aware machine translation system. Wen Wang 0001, Haibo Li 0004, Heng Ji 0001 |
ASRU | 1 |
| 2015 | Combating reverberation in large vocabulary continuous speech recognition
Vikramjit Mitra, Julien van Hout, Mitchell McLaren, Wen Wang 0001, Martin Graciarena, Dimitra Vergyri, Horacio Franco |
INTERSPEECH | 4 |
| 2015 | Morphological Modeling for Machine Translation of English-Iraqi Arabic Spoken DialogsabstractKatrin Kirchhoff, Yik-Cheung Tam, Colleen Richey, Wen Wang. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Katrin Kirchhoff, Yik-Cheung Tam, Colleen Richey, Wen Wang 0001 |
HLT-NAACL | 4 |
| 2014 | Highly accurate phonetic segmentation using boundary correction models and system fusionabstractAccurate phone-level segmentation of speech remains an important task for many subfields of speech research. We investigate techniques for boosting the accuracy of automatic phonetic segmentation based on HMM acoustic-phonetic models. In prior work [25] we were able to improve on state-of-the-art alignment accuracy by employing special phone boundary HMM models, trained on phonetically segmented training data, in conjunction with a simple boundary-time correction model. Here we present further improved results by using more powerful statistical models for boundary correction that are conditioned on phonetic context and duration features. Furthermore, we find that combining multiple acoustic front-ends gives additional gains in accuracy, and that conditioning the combiner on phonetic context and side information helps. Overall, we reduce segmentation errors on the TIMIT corpus by almost one half, from 93.9% to 96.8% boundary accuracy with a 20-ms tolerance. Andreas Stolcke, Neville Ryant, Vikramjit Mitra, Jiahong Yuan, Wen Wang 0001, Mark Y. Liberman |
ICASSP | 5 |
| 2014 | ASR error detection using recurrent neural network language model and complementary ASRabstractDetecting automatic speech recognition (ASR) errors can play an important role for effective human-computer spoken dialogue system, as recognition errors can hinder accurate system understanding of user intents. Our goal is to locate errors in an utterance so that the dialogue manager can pose appropriate clarification questions to the users. We propose two approaches to improve ASR error detection: (1) using recurrent neural network language models to capture long-distance word context within and across previous utterances; (2) using a complementary ASR system. The intuition is that when two complementary ASR systems disagree on a region in an utterance, this region is most likely an error. We train a neural network predictor of errors using a variety of features. We performed experiments on both English and Iraqi Arabic ASR and observed significant improvement in error detection using the proposed methods. Yik-Cheung Tam, Yun Lei, Jing Zheng 0001, Wen Wang 0001 |
ICASSP | 4 |
| 2014 | Evaluating robust features on deep neural networks for speech recognition in noisy and channel mismatched conditionsabstractDeep Neural Network (DNN) based acoustic models have shown significant improvement over their Gaussian Mixture Model (GMM) counterparts in the last few years. While several studies exist that evaluate the performance of GMM systems under noisy and channel degraded conditions, noise robustness studies on DNN systems have been far fewer. In this work we present a study exploring both conventional DNNs and deep Convolutional Neural Networks (CNN) for noiseand channel-degraded speech recognition tasks using the Aurora4 dataset. We compare the baseline mel-filterbank energies with noise-robust features that we have proposed earlier and show that the use of robust features helps to improve the performance of DNNs or CNNs compared to melfilterbank energies. We also show that vocal tract length normalization has a positive role in improving the performance of the robust acoustic features. Finally, we show that by combining multiple systems together we can achieve even further improvement in recognition accuracy. Vikramjit Mitra, Wen Wang 0001, Horacio Franco, Yun Lei, Chris Bartels, Martin Graciarena |
INTERSPEECH | 2 |
| 2014 | ISOMER: Informative Segment Observations for Multimedia Event RecountingabstractThis paper describes a system for multimedia event detection and recounting. The goal is to detect a high level event class in unconstrained web videos and generate event oriented summarization for display to users. For this purpose, we detect informative segments and collect observations for them, leading to our ISOMER system. We combine a large collection of both low level and semantic level visual and audio features for event detection. For event recounting, we propose a novel approach to identify event oriented discriminative video segments and their descriptions with a linear SVM event classifier. User friendly concepts including objects, actions, scenes, speech and optical character recognition are used in generating descriptions. We also develop several mapping and filtering strategies to cope with noisy concept detectors. Our system performed competitively in the TRECVID 2013 Multimedia Event Detection task with near 100,000 videos and was the highest performer in TRECVID 2013 Multimedia Event Recounting task. Chen Sun 0002, J. Brian Burns, Ramakant Nevatia, Cees Snoek, Robert C. Bolles, Gregory K. Myers, Wen Wang 0001, Eric Yeh |
ICMR | 7 |
| 2014 | Deep convolutional nets and robust features for reverberation-robust speech recognitionabstractWhile human listeners can understand speech in reverberant conditions, indicating that the auditory system is robust to such degradations, reverberation leads to high word error rates for automatic speech recognition (ASR) systems. In this work, we present robust acoustic features motivated by human speech perception for use in a convolutional deep neural network (CDNN)-based acoustic model for recognizing continuous speech in a reverberant condition. Using a single-feature system trained with the single channel data distributed through the REVERB 2014 challenge on ASR in reverberant conditions, we show a substantial relative reduction in word error rates (WERs) compared to the conventional filterbank energy-based features for single-channel simulated and real reverberation conditions. The reduction is more pronounced when multiple features and systems were combined together. The proposed system outperforms the best system reported in REVERB-2014 challenge in single channel full-batch processing task. Vikramjit Mitra, Wen Wang 0001, Horacio Franco |
SLT | 2 |
| 2013 | Name-aware Machine Translation
Haibo Li 0004, Jing Zheng 0001, Heng Ji 0001, Qi Li 0014, Wen Wang 0001 |
ACL (1) | 5 |
| 2013 | Rich system combination for keyword spotting in noisy and acoustically heterogeneous audio streamsabstractWe address the problem of retrieving spoken information from noisy and heterogeneous audio archives using system combination with a rich and diverse set of noise-robust modules. Audio search applications so far have focused on constrained domains or genres and not-so-noisy and heterogeneous acoustic or channel conditions. In this paper, our focus is to improve the accuracy of a keyword spotting system in highly degraded and diverse channel conditions by employing multiple recognition systems in parallel with different robust frontends and modeling choices, as well as different representations during audio indexing and search (words vs. subword units). After aligning keyword hits from different systems, we employ system combination at the score level using a logistic-regression-based classifier. Side information such as the output of an acoustic condition identification module is used to guide system combination system that is trained on a held-out dataset. Lattice-based indexing and search is used in all keyword spotting systems. We present improvements in probability-miss at a fixed probability-false-alarm by employing our proposed rich system combination approach on DARPA Robust Automatic Transcription of Speech (RATS) Phase-I evaluation data that contains highly degraded channel recordings (signal-to-noise ratio levels as low as 0 dB) and different channel characteristics. Murat Akbacak, Lukás Burget, Wen Wang 0001, Julien van Hout |
ICASSP | 3 |
| 2013 | Using multiple versions of speech input in phone recognitionabstractThis study investigates the use of multiple versions of the same speech unit in automatic phone recognition. Two methods were applied to combine multiple utterance versions in decoding: cross forced-alignment and n-best ROVER. The phone error rate was reduced from 15% to 2% on isolated words and from 33% to 19% on TIMIT sentences. The error rate was reduced the most when the second version was added, and less so as each additional version was added. Depending on the language model weight, it might be better to use the language model only in n-best generation, but omit it in scoring the hypotheses applied to the combination methods. N-best ROVER effectiveness may be enhanced by lowering the language model weight. Mark Y. Liberman, Jiahong Yuan, Andreas Stolcke, Wen Wang 0001, Vikramjit Mitra |
ICASSP | 4 |
| 2013 | Articulatory trajectories for large-vocabulary speech recognitionabstractStudies have demonstrated that articulatory information can model speech variability effectively and can potentially help to improve speech recognition performance. Most of the studies involving articulatory information have focused on effectively estimating them from speech, and few studies have actually used such features for speech recognition. Speech recognition studies using articulatory information have been mostly confined to digit or medium vocabulary speech recognition, and efforts to incorporate them into large vocabulary systems have been limited. We present a neural network model to estimate articulatory trajectories from speech signals where the model was trained using synthetic speech signals generated by Haskins Laboratories' task-dynamic model of speech production. The trained model was applied to natural speech, and the estimated articulatory trajectories obtained from the models were used in conjunction with standard cepstral features to train acoustic models for large-vocabulary recognition systems. Two different large-vocabulary English datasets were used in the experiments reported here. Results indicate that employing articulatory information improves speech recognition performance not only under clean conditions but also under noisy background conditions. Perceptually motivated robust features were also explored in this study and the best performance was obtained when systems based on articulatory, standard cepstral and perceptually motivated feature were all combined. Vikramjit Mitra, Wen Wang 0001, Andreas Stolcke, Hosung Nam, Colleen Richey, Jiahong Yuan, Mark Y. Liberman |
ICASSP | 2 |
| 2013 | Automatic phonetic segmentation using boundary modelsabstractThis study attempts to improve automatic phonetic segmentation within the HMM framework. Experiments were conducted to investigate the use of phone boundary models, the use of precise phonetic segmentation for training HMMs, and the difference between context-dependent and contextindependent phone models in terms of forced alignment performance. Results show that the combination of special one-state phone boundary models and monophone HMMs can significantly improve forced alignment accuracy. HMM-based forced alignment systems can also benefit from using precise phonetic segmentation for training HMMs. Context-dependent phone models are not better than context-independent models when combined with phone boundary models. The proposed system achieves 93.92% agreement (of phone boundaries) within 20 ms compared to manual segmentation on the TIMIT corpus. This is the best reported result on TIMIT to our knowledge. Jiahong Yuan, Neville Ryant, Mark Y. Liberman, Andreas Stolcke, Vikramjit Mitra, Wen Wang 0001 |
INTERSPEECH | 6 |
| 2013 | A Cross-language Study on Automatic Speech Disfluency Detection
Wen Wang 0001, Andreas Stolcke, Jiahong Yuan, Mark Y. Liberman |
HLT-NAACL | 1 |
| 2012 | Joint bilingual name tagging for parallel corporaabstractTraditional isolated monolingual name taggers tend to yield inconsistent results across two languages. In this paper, we propose two novel approaches to jointly and consistently extract names from parallel corpora. The first approach uses standard linear-chain Conditional Random Fields (CRFs) as the learning framework, incorporating cross-lingual features propagated between two languages. The second approach is based on a joint CRFs model to jointly decode sentence pairs, incorporating bilingual factors based on word alignment. Experiments on Chinese-English parallel corpora demonstrated that the proposed methods significantly outperformed monolingual name taggers, were robust to automatic alignment noise and achieved state-of-the-art performance. With only 20%of the training data, our proposed methods can already achieve better performance compared to the baseline learned from the whole training set.1 Qi Li 0014, Haibo Li 0004, Heng Ji 0001, Wen Wang 0001, Jing Zheng 0001, Fei Huang 0002 |
CIKM | 4 |
| 2012 | Detecting leadership and cohesion in spoken interactionsabstractWe present a system for detecting leadership and group cohesion in multiparty dialogs and broadcast conversations in English and Mandarin. We systematically investigate the impact of features and designs of the prediction systems, the relationships between features and their individual significance in logistic regressions, and the contributions of feature groupings as predictors for leader and group cohesion, across genres and languages. We achieve 73.0% to 94.7% F1 accuracy for leader detection and around 80% F1 accuracy for group cohesion detection, on all data sets. Wen Wang 0001, Kristin Precoda, Raia Hadsell, Zsolt Kira, Colleen Richey, Gabriel Jiva |
ICASSP | 1 |
| 2011 | N-Best Rescoring Based on Pitch-accent Patterns
Je Hun Jeon, Wen Wang 0001, Yang Liu 0004 |
ACL | 2 |
| 2011 | Automatic identification of speaker role and agreement/disagreement in broadcast conversationabstractWe present supervised approaches for detecting speaker roles and agreement/disagreement between speakers in broadcast conversation shows in three languages: English, Arabic, and Mandarin. We develop annotation approaches for a variety of linguistic phenomena. Various lexical, structural, and social network analysis based features are explored, and feature importance is analyzed across the three languages. We also compare the performance when using features extracted from automatically generated annotations against that when using human annotations. The algorithms achieve speaker role labeling accuracy of more than 86% for all three languages. For agreement and disagreement detection, the algorithms achieve precision of 63% to 92% and 55% to 85%, respectively, across the three languages. Wen Wang 0001, Sibel Yaman, Kristin Precoda, Colleen Richey |
ICASSP | 1 |
| 2011 | Identifying Agreement/Disagreement in Conversational Speech: A Cross-Lingual StudyabstractThis paper presents models for detecting agreement/disagreement between speakers in English and Arabic broadcast conversation shows. We explore a variety of features, including lexical, structural, durational, and prosodic features. We experiment with these features using Conditional Random Fields models and conduct systematic investigations on efficacy of various feature groups across languages. Sampling approaches are examined for handling highly imbalanced data. Overall, we achieved 79.2 % (precision), 50.5 % (recall), 61.7 % (F1) for agreement detection and 69.2 % (precision), 46.9 % (recall), and 55.9 % (F1) for disagreement detection, on English broadcast conversation data; and 89.2 % (precision), 30.1 % (recall), 45.1 % (F1) for agreement detection and 75.9% (precision), 28.4 % (recall), and 41.3 % (F1) for disagreement detection, on Arabic broadcast conversation data. Index Terms: agreement/disagreement detection, sampling approaches, Conditional Random Fields, feature analysis Wen Wang 0001, Kristin Precoda, Colleen Richey, Geoffrey Raymond |
INTERSPEECH | 1 |
| 2010 | Automatic disfluency removal for improving spoken language translationabstractStatistical machine translation (SMT) systems for spoken languages suffer from conversational speech phenomena, in particular, the presence of speech disfluencies. We examine the impact of disfluencies from broadcast conversation data on our hierarchical phrase-based SMT system and implement automatic disfluency removal approaches for cleansing the MT input. We evaluate the efficacy of proposed approaches and investigate the impact of disfluency removal on SMT performance across different disfluency types. We show that for translating Mandarin broadcast conversational transcripts into English, our automatic disfluency removal approaches could produce significant improvement in BLEU and TER. Wen Wang 0001, Gökhan Tür, Jing Zheng 0001, Necip Fazil Ayan |
ICASSP | 1 |
| 2010 | Unsupervised domain adaptation with multiple acoustic modelsabstractWe investigate the problem of adapting a recognition system with multiple acoustic models to a new domain in unsupervised mode. We compare maximum likelihood and discriminative approaches for unsupervised domain adaptation. Different adaptation data selection methods and adaptation strategies are investigated, using a baseline meeting recognition system and adaptation data from a congressional committee web site. Experiments show that one should avoid adapting all acoustic models to the same recognition output, and that ASR confidence estimates improve results when used for rejecting low-quality ASR output. The results show 8% relative overall improvement from unsupervised adaptation. Wen Wang 0001, Andreas Stolcke |
SLT | 2 |
| 2010 | Implementing SRI's Pashto speech-to-speech translation system on a smart phoneabstractWe describe our recent effort implementing SRI's UMPC-based Pashto speech-to-speech (S2S) translation system on a smart phone running the Android operating system. In order to maintain very low latencies of system response on computationally limited smart phone platforms, we developed efficient algorithms and data structures and optimized model sizes for various system components. Our current Android-based S2S system requires less than one-fourth the system memory and significantly lower processor speed with a sacrifice of 15% relative loss of system accuracy, compared to a laptop-based platform. Jing Zheng 0001, Arindam Mandal, Michael W. Frandsen, Necip Fazil Ayan, Dimitra Vergyri, Wen Wang 0001, Murat Akbacak, Kristin Precoda |
SLT | 7 |
| 2009 | Recent advances in SRI'S IraqCommTM Iraqi Arabic-English speech-to-speech translation systemabstractWe summarize recent progress on SRI's IraqCommtrade Iraqi Arabic-English two-way speech-to-speech translation system. In the past year we made substantial developments in our speech recognition and machine translation technology, leading to significant improvements in both accuracy and speed of the IraqComm system. On the 2008 NIST-evaluation dataset our twoway speech-to-text (S2T) system achieved 6% to 8% absolute improvement in BLEU in both directions, compared to our previous year system. Murat Akbacak, Horacio Franco, Michael W. Frandsen, Sasa Hasan, Huda Jameel, Andreas Kathol, Shahram Khadivi, Arindam Mandal, Saab Mansour, Kristin Precoda, Colleen Richey, Dimitra Vergyri, Wen Wang 0001, Jing Zheng 0001 |
ICASSP | 14 |
| 2009 | Data-driven lexicon expansion for Mandarin broadcast news and conversation speech recognitionabstractWe present a data-driven framework for expanding the lexicon to improve Mandarin broadcast news and conversation speech recognition. The lexicon expansion includes the generation of pronunciation variants for frequent words and vocabulary augmentation with new words and phrases derived from the training data. To learn multiple pronunciations, we first generate all possible pronunciation candidates for a word from its character pronunciation network. The top pronunciation variants are then selected from forced alignment statistics. To augment the acoustic vocabulary, we propose an efficient algorithm that derives new words based on N-gram statistics. Experiments show that a dictionary expanded in this manner yields significant improvements on a Mandarin broadcast speech recognition task. Wen Wang 0001, Andreas Stolcke |
ICASSP | 2 |
| 2009 | Development of the 2008 SRI Mandarin speech-to-text system for broadcast news and conversationabstractWe describe the recent progress in SRI’s Mandarin speech-to-text system developed for 2008 evaluation in the DARPAGALE program. A data-driven lexicon expansion technique and lan-guage model adaptation methods contribute to the improvement in recognition performance. Our system yields 8.3 % character error rate on the GALE dev08 test set, and 7.5 % after combining with RWTH systems. Compared to our 2007 evaluation system, a significant improvement of 13 % relative has been achieved. Index Terms: lexicon expansion, language model adaptation, Mandarin speech recognition Wen Wang 0001, Arindam Mandal, Andreas Stolcke |
INTERSPEECH | 3 |
| 2009 | Multifactor adaptation for Mandarin broadcast news and conversation speech recognitionabstractWe explore the integration of multiple factors such as genre and speaker gender for acoustic model adaptation tasks to improve Mandarin ASR system performance on broadcast news and broadcast conversation audio. We investigate the use of multifactor clustering of acoustic model training data and the application of MPE-MAP and fMPE-MAP acoustic model adaptations. We found that by effectively combining these adaptation approaches, we achieve 6 % relative reduction in recognition error rate compared to a Mandarin recognition system that does not use genre-specific acoustic models, and 5 % relative improvement if the genre-adaptive system is combined with another, genre-independent state-of-the-art system. Index Terms: large vocabulary automatic speech recognition, broadcast news, broadcast conversation, genre classification, MAP adaptation, MPE-MAP, fMPE-MAP 1. Wen Wang 0001, Arindam Mandal, Andreas Stolcke, Jing Zheng 0001 |
INTERSPEECH | 1 |
| 2009 | Using syntax in large-scale audio document translationabstractRecently, the use of syntax has very effectively improved machine translation (MT) quality in many text translation tasks. However, using syntax in speech translation poses additional challenges because of disfluencies and other spoken language phenomena, and of errors introduced by automatic speech recognition (ASR). In this paper, we investigate the effect of using syntax in a large-scale audio document translation task targeting broadcast news and broadcast conversations. We do so by comparing the performance of three synchronous context-free grammar based translation approaches: 1) hierarchical phrase-based translation, 2) syntax-augmented MT, and 3) string-to-dependency MT. The results show a positive effect of explicitly using syntax when translating broadcast news, but no benefit when translating broadcast conversations. The results indicate that improving the robustness of syntactic systems against conversational language style is important to their success and requires future effort. Index Terms: syntax, machine translation, audio document Jing Zheng 0001, Necip Fazil Ayan, Wen Wang 0001, David Burkett |
INTERSPEECH | 3 |
| 2009 | Building A Highly Accurate Mandarin Speech Recognizer With Language-Independent Technologies and Language-Dependent ModulesabstractWe describe a system for highly accurate large-vocabulary Mandarin speech recognition. The prevailing hidden Markov model based technologies are essentially language independent and constitute the backbone of our system. These include minimum-phone-error discriminative training and maximum-likelihood linear regression adaptation, among others. Additionally, careful considerations are taken into account for Mandarin-specific issues including lexical word segmentation, tone modeling, phone set design, and automatic acoustic segmentation. Our system comprises two sets of acoustic models for the purposes of cross adaptation. The systems are designed to be complementary in terms of errors but with similar overall accuracy by using different phone sets and different combinations of discriminative learning. The outputs of the two subsystems are then rescored by an adapted n-gram language model. Final confusion network combination yielded 9.1% character error rate on the DARPA GALE 2007 official evaluation, the best Mandarin recognition system in that year. Mei-Yuh Hwang, Gang Peng 0001, Mari Ostendorf, Wen Wang 0001, Arlo Faria, Aaron Heidel |
IEEE Trans. Speech Audio Process. | 4 |
| 2008 | Improving Alignments for Better Confusion Networks for Combining Machine Translation Systems
Necip Fazil Ayan, Jing Zheng 0001, Wen Wang 0001 |
COLING | 3 |
| 2008 | Development of the SRI/nightingale Arabic ASR systemabstractWe describe the large vocabulary automatic speech recognition system developed for Modern Standard Arabic by the SRI/Nightingale team, and used for the 2007 GALE evaluation as part of the speech translation system. We show how system performance is affected by different development choices, ranging from text processing and lexicon to decoding system architecture design. Word error rate results are reported on broadcast news and conversational data from the GALE development and evaluation test sets. Index Terms: speech recognition, large vocabulary, Arabic 1. Dimitra Vergyri, Arindam Mandal, Wen Wang 0001, Andreas Stolcke, Jing Zheng 0001, Martin Graciarena, David Rybach, Christian Gollan, Ralf Schlüter, Katrin Kirchhoff, Arlo Faria, Nelson Morgan |
INTERSPEECH | 3 |
| 2008 | Development of SRI's translation systems for broadcast news and broadcast conversationsabstractWe present our recent work on developing large-vocabulary Arabic-to-English and Chinese-to-English speech-to-text translation systems for the January 2008 Global Autonomous Language Exploitation (GALE) retest evaluation. Two audio genres were involved in the evaluation: broadcast news and broadcast conversation. Our system, following the hierarchical phrase-based translation approach, has a two-pass decoding strategy, with the first-pass integrated search generating 3000 unique n-best lists, which are then reranked by several different language models in the second pass. We emphasize our work on adapting the system, which was mostly trained on text data, to the speech genres, including number tokenization, punctuation compensation, and various optimization techniques. We present our results on several different tuning and testing data sets used for system development. Index Terms : speech-to-text translation, hierarchical-phrasebased translation, n-best reranking Jing Zheng 0001, Wen Wang 0001, Necip Fazil Ayan |
INTERSPEECH | 2 |
| 2008 | Phonetic name matching for cross-lingual Spoken Sentence RetrievalabstractCross-lingual Spoken Sentence Retrieval (CLSSR) remains a challenge, especially for queries including OOV words such as person names. This paper proposes a simple method of fuzzy matching between query names and phones of candidate audio segments. This approach has the advantage of avoiding some word decoding errors in Automatic Speech Recognition (ASR). Experiments on Mandarin-English CLSSR show that phone-based searching and conventional translation-based searching are complementary. Adding phone matching achieved 26.29% improvement on F-measure over searching on state-of-the-art Machine Translation (MT) output and 8.83% over Entity Translation (ET) output. Heng Ji 0001, Ralph Grishman, Wen Wang 0001 |
SLT | 3 |
| 2008 | Efficient data selection for machine translationabstractPerformance of statistical machine translation (SMT) systems relies on the availability of a large parallel corpus which is used to estimate translation probabilities. However, the generation of such corpus is a long and expensive process. In this paper, we introduce two methods for efficient selection of training data to be translated by humans. Our methods are motivated by active learning and aim to choose new data that adds maximal information to the currently available data pool. The first method uses a measure of disagreement between multiple SMT systems, whereas the second uses a perplexity criterion. We performed experiments on Chinese-English data in multiple domains and test sets. Our results show that we can select only one-fifth of the additional training data and achieve similar or better translation performance, compared to that of using all available data. Arindam Mandal, Dimitra Vergyri, Wen Wang 0001, Jing Zheng 0001, Andreas Stolcke, Gökhan Tür, Dilek Hakkani-Tür, Necip Fazil Ayan |
SLT | 3 |
| 2007 | Building a highly accurate Mandarin speech recognizerabstractWe describe a highly accurate large-vocabulary continuous Mandarin speech recognizer, a collaborative effort among four research organizations. Particularly, we build two acoustic models (AMs) with significant differences but similar accuracy for the purposes of cross adaptation and system combination. This paper elaborates on the main differences between the two systems, where one recognizer incorporates a discriminatively trained feature while the other utilizes a discriminative feature transformation. Additionally we present an improved acoustic segmentation algorithm and topic-based language model (LM) adaptation. Coupled with increased acoustic training data, we reduced the character error rate (CER) of the DARPA GALE 2006 evaluation set to 15.3% from 18.4%. Mei-Yuh Hwang, Gang Peng 0001, Wen Wang 0001, Arlo Faria, Aaron Heidel, Mari Ostendorf |
ASRU | 3 |
| 2007 | Reranking machine translation hypotheses with structured and web-based language modelsabstractIn this paper, we investigate the use of linguistically motivated and computationally efficient structured language models for reranking N-best hypotheses in a statistical machine translation system. These language models, developed from Constraint Dependency Grammar parses, tightly integrate knowledge of words, morphological and lexical features, and syntactic dependency constraints. Two structured language models are applied for N-best rescoring, one is an almost-parsing language model, and the other utilizes more syntactic features by explicitly modeling syntactic dependencies between words. We also investigate effective and efficient language modeling methods to use N-grams extracted from up to 1 teraword of web documents. We apply all these language models for N-best re-ranking on the NIST and DARPA GALE program12006 and 2007 machine translation evaluation tasks and find that the combination of these language models increases the BLEU score up to 1.6% absolutely on blind test sets. Wen Wang 0001, Andreas Stolcke, Jing Zheng 0001 |
ASRU | 1 |
| 2007 | Mandarin Part-of-Speech Tagging and Discriminative Reranking
Zhongqiang Huang, Mary P. Harper, Wen Wang 0001 |
EMNLP-CoNLL | 3 |
| 2007 | Semi-Supervised Learning for Part-of-Speech Tagging of Mandarin Transcribed SpeechabstractIn this paper, we investigate bootstrapping part-of-speech (POS) taggers for Mandarin broadcast news (BN) transcripts using co-training, by iteratively retraining two competitive POS taggers from a small set of labeled training data and a large set of unlabeled data. We compare co-training with self-training and our results show that the performance using co-training is significantly better than that from self-training and these semi-supervised learning methods significantly improve tagging accuracy over training only on the small labeled seed corpus. We also investigate a variety of example selection approaches for co-training and find that the computationally expensive, agreement-based selection approach and a more efficient selection approach based on maximizing training utility produce comparable tagging performance from resulting POS taggers. By applying co-training, we are able to build effective POS taggers for Mandarin transcribed speech with the tagging accuracy comparable to that obtained on newswire text. Wen Wang 0001, Zhongqiang Huang, Mary P. Harper |
ICASSP (4) | 1 |
| 2007 | Advances in Mandarin broadcast speech recognitionabstractWe describe our continuing efforts to improve the UW-SRI-ICSI Mandarin broadcast speech recognizer. This includes increasing acoustic and text training data, adding discriminative features, incorporating frame-level discriminative training criterion, multiplepass acoustic model (AM) cross adaptation, language model (LM) genre adaptation and system combination. The net effect without LM adaptation was a 24%-64 % relative reduction in character error rates (CERs) on a variety of test sets. In addition, LM adaptation gave us another 6 % of relative CER reduction on broadcast conversations. Index Terms — Mandarin, character error rates, discriminative features, discriminative training, LM adaptation. 1. Mei-Yuh Hwang, Wen Wang 0001, Jing Zheng 0001, Özgür Çetin, Gang Peng 0001 |
INTERSPEECH | 2 |
| 2007 | The SRI/OGI 2006 spoken term detection systemabstractThis paper describes the system developed jointly at SRI and OGI for participation in the 2006 NIST Spoken Term Detection (STD) evaluation. We participated in the three genres of the English track: Broadcast News (BN), Conversational Telephone Speech (CTS), and Conference Meetings (MTG). The system consists of two phases. First, audio indexing, an offline phase, converts the input speech waveform into a searchable index. Second, term retrieval, possibly an online phase, returns a ranked list of occurrences for each search term. We used a word-based indexing approach, obtained with SRI’s large vocabulary Speech-to-Text (STT) system. Apart from describing the submitted system and its performance on the NIST evaluation metric, we study the tradeoffs between performance and system design. We examine performance versus indexing speed, effectiveness of different index ranking schemes on the NIST score, and the utility of approaches to deal with out-of-vocabulary (OOV) terms. Dimitra Vergyri, Izhak Shafran, Andreas Stolcke, Venkata Ramana Rao Gadde, Murat Akbacak, Brian Roark, Wen Wang 0001 |
INTERSPEECH | 7 |
| 2007 | Integrating MAP, marginals, and unsupervised language model adaptationabstractWe investigate the integration of various language model adaptation approaches for a cross-genre adaptation task to improve Mandarin ASR system performance on a recently introduced new genre, broadcast conversation (BC). Various language model adaptation strategies are investigated and their efficacies are evaluated based on ASR performance, including unsupervised language model adaptation from ASR transcripts and ways to integrate supervised Maximum A Posteriori (MAP) and marginal adaptation within the unsupervised adaptation framework. We found that by effectively combining these adaptation approaches, we can achieve as much as 1.3 % absolute gain (6% relative) on the final recognition error rate in the BC genre. Index Terms: automatic speech recognition, language model adaptation, supervised adaptation, unsupervised adaptation, broadcast conversation 1. Wen Wang 0001, Andreas Stolcke |
INTERSPEECH | 1 |
| 2006 | Speech Recognition Engineering Issues in Speech to Speech Translation System Design for Low Resource Languages and DomainsabstractEngineering automatic speech recognition (ASR) for speech to speech (S2S) translation systems, especially targeting languages and domains that do not have readily available spoken language resources, is immensely challenging due to a number of reasons. In addition to contending with the conventional data-hungry speech acoustic and language modeling needs, these designs have to accommodate varying requirements imposed by the domain needs and characteristics, target device and usage modality (such as phrase-based, or spontaneous free form interactions, with or without visual feedback) and huge spoken language variability arising due to socio-linguistic and cultural differences of the users. This paper, using case studies of creating speech translation systems between English and languages such as Pashto and Farsi, describes some of the practical issues and the solutions that were developed for multilingual ASR development. These include novel acoustic and language modeling strategies such as language adaptive recognition, active-learning based language modeling, class-based language models that can better exploit resource poor language data, efficient search strategies, including N-best and confidence generation to aid multiple hypotheses translation, use of dialog information and clever interface choices to facilitate ASR, and audio interface design for meeting both usability and robustness requirements Shri Narayanan, Panayiotis G. Georgiou, Abhinav Sethy, Dagen Wang, Murtaza Bulut, Shiva Sundaram, Emil Ettelaie, Sankaranarayanan Ananthakrishnan, Horacio Franco, Kristin Precoda, Dimitra Vergyri, Jing Zheng 0001, Wen Wang 0001, Venkata Ramana Rao Gadde, Martin Graciarena, Victor Abrash, Michael W. Frandsen, Colleen Richey |
ICASSP (5) | 13 |
| 2006 | The Use of Word N-Grams and Parts of Speech for Hierarchical Cluster Language ModelingabstractWe present extensions to the work of backoff hierarchical class n-gram language modeling of Zitouni et al. [1] by studying the efficacy of exploring the use of parts of speech (POS) information in hierarchical word clustering. We propose two approaches. One is to use POS n-gram contextual distribution of a target word for clustering. The other is to generate a class tree for each group of words sharing the same POS. The resulting class tree and a set of class trees, from the two approaches, respectively, are then employed in the hierarchical cluster language modeling. We evaluate the two approaches on SRI Arabic conversational telephone speech recognition system and show that the approach of building a set of POS-specific class trees achieves a 3% relative improvement on perplexity compared to the model of Zitouni et al. and a 8% relative improvement on perplexity over the baseline standard word n-grams. When used for N-best rescoring, our approach also outperforms the model of Zitouni et al. and the baseline and achieves significant word error rate (WER) reductions. Wen Wang 0001, Dimitra Vergyri |
ICASSP (1) | 1 |
| 2006 | Investigation on Mandarin broadcast news speech recognitionabstractThis paper describes the authors' efforts in developing a competitive Mandarin broadcast news speech recognizer. They have successfully incorporated the most popular speech technologies into their system. More importantly, they present two novel algorithms for smoothing pitch features and segmenting Chinese characters into word units. In addition, they propose to borrow the principle of point-wise mutual information for creating a Chinese word lexicon automatically. Their final system achieved a 6.0% character error rate (CER) on dev04 and a 16.0% CER on eval04 with simpler acoustic models, less training data, and simpler decoding architecture compared with other state-of-the-art systems. This system is equally competitive. Mei-Yuh Hwang, Wen Wang 0001, Takahiro Shinozaki |
INTERSPEECH | 3 |
| 2006 | Impact of Automatic Comma Prediction on POS/Name Tagging of speechabstractThis work looks at the impact of automatically predicted commas on part-of-speech (POS) and name tagging of speech recognition transcripts of Mandarin broadcast news. There is a significant gain in both POS and name tagging accuracy due to using automatically predicted commas over sentence boundary prediction alone. One difference between Mandarin and English is that there are two types of commas, and experiments here show that, while they can be reliably distinguished in automatic prediction, the distinction does not give a clear benefit for POS or name tagging. Dustin Hillard, Zhongqiang Huang, Heng Ji 0001, Ralph Grishman, Dilek Hakkani-Tür, Mary P. Harper, Mari Ostendorf, Wen Wang 0001 |
SLT | 8 |
| 2006 | Recent innovations in speech-to-text transcription at SRI-ICSI-UWabstractWe summarize recent progress in automatic speech-to-text transcription at SRI, ICSI, and the University of Washington. The work encompasses all components of speech modeling found in a state-of-the-art recognition system, from acoustic features, to acoustic modeling and adaptation, to language modeling. In the front end, we experimented with nonstandard features, including various measures of voicing, discriminative phone posterior features estimated by multilayer perceptrons, and a novel phone-level macro-averaging for cepstral normalization. Acoustic modeling was improved with combinations of front ends operating at multiple frame rates, as well as by modifications to the standard methods for discriminative Gaussian estimation. We show that acoustic adaptation can be improved by predicting the optimal regression class complexity for a given speaker. Language modeling innovations include the use of a syntax-motivated almost-parsing language model, as well as principled vocabulary-selection techniques. Finally, we address portability issues, such as the use of imperfect training transcripts, and language-specific adjustments required for recognition of Arabic and Mandarin Andreas Stolcke, Barry Y. Chen, Horacio Franco, Venkata Ramana Rao Gadde, Martin Graciarena, Mei-Yuh Hwang, Katrin Kirchhoff, Arindam Mandal, Nelson Morgan, Tim Ng, Mari Ostendorf, M. Kemal Sönmez, Anand Venkataraman, Dimitra Vergyri, Wen Wang 0001, Jing Zheng 0001, Qifeng Zhu 0001 |
IEEE Trans. Speech Audio Process. | 16 |
| 2005 | Speech translation for low-resource languages: the case of PashtoabstractWe present a number of challenges and solutions that have arisen in the development of a speech translation system for American English and Pashto, highlighting those specific to a very low resource language. In particular, we address issues posed by Pashto in the areas of written representation, corpus creation, speech recognition, speech synthesis, and grammar development for translation. Andreas Kathol, Kristin Precoda, Dimitra Vergyri, Wen Wang 0001, Susanne Z. Riehemann |
INTERSPEECH | 4 |
| 2004 | The use of a linguistically motivated language model in conversational speech recognitionabstractStructured language models have recently been shown to give significant improvements in large-vocabulary recognition relative to traditional word N-gram models, but typically imply a heavy computational burden and have not been applied to large training sets or complex recognition systems. Previously, we developed a linguistically motivated and computationally efficient almost-parsing language model, using a data structure derived from constraint dependency grammar parsing, that tightly integrates knowledge of words, lexical features, and syntactic constraints. We show that such a model can be used effectively and efficiently in all stages of a complex, multi-pass conversational telephone speech recognition system. Compared to a state-of-the-art 4-gram interpolated word- and class-based language model, we obtained a 6.2% relative word error reduction (a 1.6% absolute reduction) on a recent NIST evaluation set. Wen Wang 0001, Andreas Stolcke, Mary P. Harper |
ICASSP (1) | 1 |
| 2004 | An efficient repair procedure for quick transcriptionsabstractWe describe an efficient procedure for automatic repair of quickly transcribed (QT) speech. QT speech, typically closed captioned data from television broadcasts, usually has a significant number of deletions and misspellings, and has a characteristic absence of disfluencies such as filled pauses (for example, um, uh). Errors of these kinds often throw an acoustic model training program out of alignment and make it hard for it to resynchronize. At best the erroneous utterance is discarded and does not benefit the training procedure. At worst, it could misalign and end up sabotaging the training data. The procedure we propose in this paper aims to cleanse such quick transcriptions so that they align better with the acoustic evidence and thus provide for better acoustic models for automatic speech recognition (ASR). Results from comparing our transcripts with those from careful transcriptions on the same corpus, and from comparable state-of-the-art methods are also presented and discussed. Anand Venkataraman, Andreas Stolcke, Wen Wang 0001, Dimitra Vergyri, Jing Zheng 0001, Venkata Ramana Rao Gadde |
INTERSPEECH | 3 |
| 2003 | The robustness of an almost-parsing language model given errorful training dataabstractAn almost-parsing language model has been developed (Wang and Harper 2002) that provides a framework for tightly integrating multiple knowledge sources. Lexical features and syntactic constraints are integrated into a uniform linguistic structure (called a SuperARV) that is associated with words in the lexicon. The SuperARV language model has been found able to reduce perplexity and word error rate (WER) compared to trigram, part-of-speech-based, and parser-based language models on the DARPA Wall Street Journal (WSJ) CSR task. In this paper we further investigate the robustness of the language model to possibly inconsistent and flawed training data, as well as its ability to scale up to sophisticated LVCSR tasks by comparing performance on the DARPA WSJ and Hub4 (Broadcast News) CSR tasks. Wen Wang 0001, Mary P. Harper, Andreas Stolcke |
ICASSP (1) | 1 |
| 2003 | Techniques for effective vocabulary selectionabstractThe vocabulary of a continuous speech recognition (CSR) system is a significant factor in determining its performance. In this paper, we present three principled approaches to select the target vocabulary for a particular domain by trading off between the target out-of-vocabulary (OOV) rate and vocabulary size. We evaluate these approaches against an ad-hoc baseline strategy. Results are presented in the form of OOV rate graphs plotted against increasing vocabulary size for each technique. Anand Venkataraman, Wen Wang 0001 |
INTERSPEECH | 2 |
| 2002 | The SuperARV Language Model: Investigating the Effectiveness of Tightly Integrating Multiple Knowledge SourcesabstractA new almost-parsing language model incorporating multiple knowledge sources that is based upon the concept of constraint Dependency Grammars is presented in this paper. Lexical features and syntactic constraints are tightly integrated into a uniform linguistic structure called a SuperARV that is associated with a word in the lexicon. The SuperARV language model reduces perplexity and word error rate compared to trigram, part-of-speech-based, and parser-based language models. The relative contributions of the various knowledge sources to the strength of our model are also investigated by using constraint relaxation at the level of the knowledge sources. We have found that although each knowledge source contributes to language model quality, lexical features are an outstanding contributor when they are tightly integrated with word identity and syntactic constraints. Our investigation also suggests possible reasons for the reported poor performance of several probabilistic dependency grammar models in the literature. Wen Wang 0001, Mary P. Harper |
EMNLP | 1 |
| 2002 | Rescoring effectiveness of language models using different levels of knowledge and their integrationabstractIn this paper, we compare the efficacy of a variety of language models (LMs) for rescoring word graphs and N-best lists generated by a large vocabulary continuous speech recognizer. These LMs differ based on the level of knowledge used (word, lexical features, syntax) and the type of integration of that knowledge (tight or loose). The trigram LM incorporates word level information; our Part-of-Speech (POS) LM uses word and lexical class information in a tightly coupled way; our new SuperARV LM tightly integrates word, a richer set of lexical features than POS, and syntactic dependency information; and the Parser LM integrates some limited word information, POS, and syntactic information. We also investigate LMs created using a linear interpolation of LM pairs. When comparing each LM on the task of rescoring word graphs or N-best lists for the Wall Street Journal (WSJ) 5k- and 20k- vocabulary test sets, the SuperARV LM always achieves the greatest reduction in word error rate (WER) and the greatest increase in sentence accuracy (SAC). On the 5k test sets, the SuperARV LM obtains more than a 10% relative reduction in WER compared to the trigram LM, and on the 20k test set more than 2%. Additionally, the SuperARV LM performs comparably to or better than the interpolated LMs. Hence, we conclude that the tight coupling of knowledge from all three levels is an effective method of constructing high quality LMs. Wen Wang 0001, Yang Liu 0004, Mary P. Harper |
ICASSP | 1 |