Dong Yu 0001

dblp:71/4598-1 · DBLP profile ↗
← Back
394ranked-venue papers
46as first author
184since 2021 · last 2026
0000-0003-0520-6844ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 263 · 30 first-author · 142 since 2021Graphics, computer vision, multimedia, augmented reality and games · 244 · 32 first-author · 87 since 2021Applied, interdisciplinary, general and emerging computing · 3Computer networks · 1 · 1 first-authorSecurity and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DegVoC: Revisiting Neural Vocoder from a Degradation Perspective
abstract
Existing neural vocoders have demonstrated promising performance by leveraging Mel-spectrum as an acoustic feature for conditional audio generation. Nonetheless, they remain constrained by an inherent ``performance-cost'' dilemma that significantly hinders the development of this field. This paper revisits this foundational task from a novel degradation perspective, where Mel-spectrum is regarded as a special signal degradation process from the target spectrum. Drawing inspiration from traditional sparse signal recovery problems, we propose DegVoC, a GAN-based neural vocoder with a two-step solution procedure. First, by exploiting degradation priors, we attempt to retrieve the initial spectral structure from Mel-domain representations as an initial solution via a simple linear transformation. Based on that, we introduce a deep prior solver that accounts for the heterogeneous distribution of sub-bands in the time-frequency domain. A convolution-style attention module with a large kernel size is specially devised for efficient inter-frame and inter-band contextual modeling. With 3.89 M parameters and substantially reduced inference complexity, DegVoC achieves state-of-the-art performance across objective and subjective evaluations, outperforming existing GAN-, DDPM- and flow-matching-based baselines.
Andong Li, Lingling Dai, Rilin Chen, Meng Yu 0003, Xiaodong Li 0002, Dong Yu 0001, Chengshi Zheng
AAAI8
2026 Enhancing Stability and Fidelity for Zero-Shot TTS with a Multi-Level Evaluator
abstract
Recent advances in zero-shot text-to-speech (TTS), driven by language models, diffusion models and masked generation, have achieved impressive naturalness in speech synthesis. Nevertheless, stability and fidelity remain key challenges, manifesting as mispronunciations, audible noise, and quality degradation. To address these issues, we introduce Vox-Evaluator, a multi-level evaluator designed to guide the correction of erroneous speech segments and preference alignment for TTS systems. It is capable of identifying the temporal boundaries of erroneous segments and providing a holistic quality assessment of the generated speech. Specifically, to refine erroneous segments and enhance the robustness of the zero-shot TTS model, we propose to automatically identify acoustic errors with the evaluator, mask the erroneous segments, and finally regenerate speech conditioning on the correct portions. In addition, the fine-gained information obtained from Vox-Evaluator can guide the preference alignment for TTS model, thereby reducing the bad cases in speech synthesize. Due to the lack of suitable training datasets for the Vox-Evaluator, we also constructed a synthesized text-speech dataset annotated with fine-grained pronunciation errors or audio quality issues. The experimental results demonstrate the effectiveness of the proposed Vox-Evaluator in enhancing the stability and fidelity of TTS systems through the speech correction mechanism and preference optimization.
Hualei Wang, Na Li 0012, Chuke Wang, Zhifeng Li 0001, Dong Yu 0001
AAAI6
2026 UniCUE: Unified Recognition and Generation Framework for Chinese Cued Speech Video-to-Speech Generation
abstract
Cued Speech (CS) enhances lipreading via hand coding, offering visual phonemic cues that support precise speech perception for the hearing-impaired. The task of CS Video-to-Speech generation (CSV2S) aims to convert CS videos into intelligible speech signals. Most existing research focuses on CS Recognition (CSR), which transcribes video content into text. Consequently, a common solution for CSV2S is to integrate CSR with a text-to-speech (TTS) system. However, this pipeline relies on text as an intermediate medium, which may lead to error propagation and temporal misalignment between speech and CS video dynamics. In contrast, directly generating audio speech from CS video (direct CSV2S) often suffer from the inherent multimodal complexity and the limited availability of CS data. To address these challenges, we propose UniCUE, the first unified framework for CSV2S that directly generates speech from CS videos without relying on intermediate text. The core innovation of UniCUE lies in integrating a understanding task (CSR) that provides fine-grained CS visual-semantic cues to to guide the speech generation. Specifically, UniCUE incorporates a pose-aware visual processor, a semantic alignment pool that enables precise visual–semantic mapping, and a VisioPhonetic adapter to bridge the understanding and generation tasks within a unified architecture. To support this framework, we construct UniCUE-HI, a large-scale Mandarin CS dataset containing 11,282 videos from 14 cuers, including both hearing-impaired and normal-hearing individuals. Extensive experiments conducted on this dataset demonstrate that UniCUE achieves state-of-the-art (SOTA) performance across multiple evaluation metrics.
Jinting Wang, Shan Yang 0001, Chenxing Li, Dong Yu 0001, Li Liu 0036
AAAI4
2026 Audio-Thinker: Guiding Large Audio Language Model When and How to Think via Reinforcement Learning
abstract
Recent advancements in large language models, multimodal large language models, and large audio language models (LALMs) have significantly improved their reasoning capabilities through reinforcement learning utilizing rule-based rewards. However, the explicit reasoning process has not yet yielded substantial benefits for audio question answering, and effectively leveraging deep reasoning remains an open challenge, with LALMs still falling short of achieving human-level auditory-language reasoning. To address these limitations, we propose Audio-Thinker, a reinforcement learning framework designed to enhance the reasoning capabilities of LALMs through improved adaptability, consistency, and effectiveness. Our approach introduces an adaptive think accuracy reward, enabling the model to adjust its reasoning strategies based on task complexity. Furthermore, we incorporate an external reward model to evaluate the overall consistency and quality of the reasoning process, complemented by think-based rewards that assist the model in distinguishing between valid and flawed reasoning paths during training. Experimental results demonstrate that Audio-Thinker models outperform existing reasoning-oriented LALMs across various benchmark tasks, exhibiting superior reasoning and generalization capabilities.
Chenxing Li, Wenfu Wang, Hao Zhang 0112, Hualei Wang, Meng Yu 0003, Dong Yu 0001
AAAI7
2026 EconProver: Towards More Economical Test-Time Scaling for Automated Theorem Proving
abstract
Mukai Li, Linfeng Song, Zhenwen Liang, Jiahao Xu, Shansan Gong, Qi Liu, Haitao Mi, Dong Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Mukai Li, Linfeng Song, Zhenwen Liang, Shansan Gong, Qi Liu 0049, Haitao Mi, Dong Yu 0001
ACL (1)8
2026 Your Reasoning Model is Secretly a Reward Model - Optimization-Free Verification from Experience
abstract
Zhenwen Liang, Ruosen Li, Yujun Zhou, Linfeng Song, Dian Yu, Xinya Du, Haitao Mi, Dong Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhenwen Liang, Ruosen Li, Yujun Zhou 0002, Linfeng Song, Dian Yu 0001, Xinya Du, Haitao Mi, Dong Yu 0001
ACL (1)8
2026 Crossing the Reward Bridge: Expanding Reinforcement Learning with Verifiable Rewards Across Diverse Domains
abstract
Yi Su, Dian Yu, Linfeng Song, Juntao Li, Haitao Mi, Zhaopeng Tu, Min Zhang, Dong Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yi Su 0006, Dian Yu 0001, Linfeng Song, Juntao Li 0005, Haitao Mi, Zhaopeng Tu, Min Zhang 0005, Dong Yu 0001
ACL (1)8
2026 Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation
abstract
Audio-language pretraining (ALP) holds promise for learning general-purpose audio representation, yet remains underexplored.Crucially, there is no consensus on whether audio-language models can build effective general-purpose audio encoders, nor a systematic understanding of how pretraining objectives behave across diverse tasks and scales.We identify three key barriers: limited scale of audio-text corpora, limited coverage of audio attributes in existing caption corpora, and lack of systematic exploration and evaluation.To fill this gap, we present the first principled empirical study of ALP.We first introduce Cap-tionStew, a 10.7M caption dataset aggregating open-source audio-text corpora across multiple domains and captioning focuses.We then conduct the first comprehensive evaluation comparing contrastive and captioning objectives for learning audio representation across speech, music, and environmental sound tasks.Our results not only demonstrate that ALP yields competitive, transferable representations, but reveal critical trade-offs: contrastive learning offers superior data efficiency, while captioning exhibits better scalability.Furthermore, we find that the benefits of supervised initialization often diminish at larger scales, challenging common practices.By grounding these claims in empirical evidence, we establish a viable pathway toward general-purpose audio representation learning, guiding future research.
Wei-Cheng Tseng, Xuanru Zhou, Mingyue Huo, Yiwen Shao, Hao Zhang 0112, Dong Yu 0001
ACL (1)6
2026 WebAggregator: Enhancing Compositional Reasoning Capabilities of Deep Research Agent Foundation Models
abstract
Rui Wang, Ce Zhang, Jun-Yu Ma, Jianshu Zhang, Hongru Wang, Yi Chen, Boyang Xue, Tianqing Fang, Zhisong Zhang, Hongming Zhang, Haitao Mi, Dong Yu, Kam-Fai Wong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Rui Wang 0015, Ce Zhang 0009, Jun-Yu Ma, Hongru Wang 0003, Yi Chen 0007, Boyang Xue, Tianqing Fang, Zhisong Zhang, Hongming Zhang 0009, Haitao Mi, Dong Yu 0001, Kam-Fai Wong
ACL (1)12
2026 Lifelong Learning of Large Language Model Based Agents: A Roadmap
abstract
Lifelong learning, also known as continual or incremental learning, is a crucial component for advancing Artificial General Intelligence (AGI) by enabling systems to continuously adapt in dynamic environments. While large language models (LLMs) have demonstrated impressive capabilities in natural language processing, existing LLM agents are typically designed for static systems and lack the ability to adapt over time in response to new challenges. This survey is the first to systematically summarize the potential techniques for incorporating lifelong learning into LLM-based agents. We categorize the core components of these agents into three modules: the perception module for multimodal input integration, the memory module for storing and retrieving evolving knowledge, and the action module for grounded interactions with the dynamic environment. We highlight how these pillars collectively enable continuous adaptation, mitigate catastrophic forgetting, and improve long-term performance. This survey provides a roadmap for researchers and practitioners working to develop lifelong learning capabilities in LLM agents, offering insights into emerging trends, evaluation metrics, and application scenarios.
Junhao Zheng, Chengming Shi, Xidi Cai, Qiuke Li, Duzhen Zhang, Chenxing Li, Dong Yu 0001, Qianli Ma 0001
IEEE Trans. Pattern Anal. Mach. Intell.7
2026 TencentLLMEval: A Hierarchical Evaluation of Real-World Capabilities for Human-Aligned LLMs
abstract
Large language models (LLMs) have shown impressive capabilities across various natural language tasks. However, evaluating their alignment with human preferences remains a challenge. To this end, we propose a comprehensive human evaluation framework to assess LLMs’ proficiency in following instructions on diverse real-world tasks. We construct a hierarchical task tree encompassing seven major areas covering over 200 categories and over 800 tasks, which covers diverse capabilities such as question answering, reasoning, multi-turn dialogue, and text generation, to evaluate LLMs in a comprehensive and in-depth manner. We also design detailed evaluation standards and processes to facilitate consistent, unbiased judgments from human evaluators. A test set of over 3,000 instances is released, spanning different difficulty levels and knowledge domains. Our work provides a standardized methodology to evaluate human alignment in LLMs for both English and Chinese. We also analyze the feasibility of automating parts of evaluation with a strong LLM (GPT-4). Our framework supports a thorough assessment of LLMs as they are integrated into real-world applications. We have made publicly available the task tree, TencentLLMEval dataset, and evaluation methodology which have been demonstrated as effective in assessing the performance of Tencent Hunyuan LLMs. By doing so, we aim to facilitate the benchmarking of advances in the development of safe and human-aligned LLMs.
Shuyi Xie, Wenlin Yao, Yong Dai 0001, Zishan Xu, Fan Lin, Donglin Zhou, Lifeng Jin, Xinhua Feng, Pengzhi Wei, Zhichao Hu, Dong Yu 0001, Zhengyou Zhang
ACM Trans. Intell. Syst. Technol.13
2025 LiteSearch: Efficient Tree Search with Dynamic Exploration Budget for Math Reasoning
abstract
Recent research suggests that tree search algorithms (e.g. Monte Carlo Tree Search) can dramatically boost LLM performance on complex mathematical reasoning tasks. However, they often require more than 10 times the computational resources of greedy decoding due to wasteful search strategies, making them difficult to be deployed in practical applications. This study introduces a novel guided tree search algorithm with a goal-directed heuristic function and node-level exploration budget (maximum number of children) calculation to tackle this issue. By considering the search progress towards the final answer (history) and the guidance from a value network (future) trained without any step-wise annotations, our algorithm iteratively selects the most promising tree node before expanding it within the boundaries of the allocated computational budget. Experiments conducted on the GSM8K, TabMWP, and MATH datasets demonstrate that our method not only offers competitive performance but also enjoys significantly lower computational costs compared to baseline methods.
Ante Wang, Linfeng Song, Baolin Peng, Dian Yu 0001, Haitao Mi, Jinsong Su, Dong Yu 0001
AAAI8
2025 A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression
abstract
Chenlong Deng, Zhisong Zhang, Kelong Mao, Shuaiyi Li, Xinting Huang, Dong Yu, Zhicheng Dou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Chenlong Deng, Zhisong Zhang, Kelong Mao, Shuaiyi Li, Xinting Huang, Dong Yu 0001, Zhicheng Dou
ACL (1)6
2025 OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and Optimization
abstract
Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Hongming Zhang, Tianqing Fang, Zhenzhong Lan, Dong Yu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Hongliang He 0002, Wenlin Yao, Kaixin Ma, Wenhao Yu 0002, Hongming Zhang 0009, Tianqing Fang, Zhen-Zhong Lan, Dong Yu 0001
ACL (1)8
2025 Low-Bit Quantization Favors Undertrained LLMs
abstract
Low-bit quantization improves machine learning model efficiency but surprisingly favors undertrained large language models (LLMs).Larger models or those trained on fewer tokens exhibit less quantization-induced degradation (QiD), while smaller, well-trained models face significant performance losses.To gain deeper insights into this trend, we study over 1500+ quantized LLM checkpoints of various sizes and at different training levels (undertrained or fully trained) in a controlled setting, deriving scaling laws for understanding the relationship between QiD and factors: the number of training tokens, model size and bit width.With our derived scaling laws, we propose a novel perspective that we can use QiD to measure an LLM's training levels and determine the number of training tokens required for fully training LLMs of various sizes.Moreover, we use the scaling laws to predict the quantization performance of different-sized LLMs trained with 100 trillion tokens.Our projection shows that the low-bit quantization performance of future models, which are expected to be trained with over 100 trillion tokens, may NOT be desirable.This poses a potential challenge for low-bit quantization in the future and highlights the need for awareness of a model's training level when evaluating lowbit quantization research.To facilitate future research on this problem, we release all the 1500+ quantized checkpoints used in this work at https://huggingface.co/Xu-Ouyang.
Xu Ouyang, Tao Ge 0001, Thomas Hartvigsen, Zhisong Zhang, Haitao Mi, Dong Yu 0001
ACL (1)6
2025 Don't Get Lost in the Trees: Streamlining LLM Reasoning by Overcoming Tree Search Exploration Pitfalls
abstract
Recent advancements in tree search algorithms guided by verifiers have significantly enhanced the reasoning capabilities of large language models (LLMs), but at the cost of increased computational resources. In this work, we identify two key challenges contributing to this inefficiency: \textit{over-exploration} due to redundant states with semantically equivalent content, and \textit{under-exploration} caused by high variance in verifier scoring leading to frequent trajectory switching. To address these issues, we propose FETCH – an e{\bf f}fici{\bf e}nt {\bf t}ree sear{\bf ch} framework, which is a flexible, plug-and-play system compatible with various tree search algorithms.Our framework mitigates over-exploration by merging semantically similar states using agglomerative clustering of text embeddings obtained from a fine-tuned SimCSE model. To tackle under-exploration, we enhance verifiers by incorporating temporal difference learning with adjusted \lambda-returns during training to reduce variance, and employing a verifier ensemble to aggregate scores during inference. Experiments on GSM8K, GSM-Plus, and MATH datasets demonstrate that our methods significantly improve reasoning accuracy and computational efficiency across four different tree search algorithms, paving the way for more practical applications of LLM-based reasoning. The code is available at https://github.com/DeepLearnXMU/Fetch.
Ante Wang, Linfeng Song, Dian Yu 0001, Haitao Mi, Xiangyu Duan, Zhaopeng Tu, Jinsong Su, Dong Yu 0001
ACL (1)9
2025 LoGU: Long-form Generation with Uncertainty Expressions
abstract
Ruihan Yang, Caiqi Zhang, Zhisong Zhang, Xinting Huang, Sen Yang, Nigel Collier, Dong Yu, Deqing Yang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Ruihan Yang, Caiqi Zhang, Zhisong Zhang, Xinting Huang, Nigel Collier, Dong Yu 0001, Deqing Yang
ACL (1)7
2025 Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language Models
abstract
Large language models have shown remarkable performance across a wide range of language tasks, owing to their exceptional capabilities in context modeling. The most commonly used method of context modeling is full self-attention, as seen in standard decoder-only Transformers. Although powerful, this method can be inefficient for long sequences and may overlook inherent input structures. To address these problems, an alternative approach is parallel context encoding, which splits the context into sub-pieces and encodes them parallelly. Because parallel patterns are not encountered during training, naively applying parallel encoding leads to performance degradation. However, the underlying reasons and potential mitigations are unclear. In this work, we provide a detailed analysis of this issue and identify that unusually high attention entropy can be a key factor. Furthermore, we adopt two straightforward methods to reduce attention entropy by incorporating attention sinks and selective mechanisms. Experiments on various tasks reveal that these methods effectively lower irregular attention entropy and narrow performance gaps. We hope this study can illuminate ways to enhance context modeling mechanisms.
Zhisong Zhang, Yan Wang 0060, Xinting Huang, Tianqing Fang, Hongming Zhang 0009, Chenlong Deng, Shuaiyi Li, Dong Yu 0001
ACL (1)8
2025 Efficient Scaling for LLM-based ASR
abstract
Large language model (LLM)-based automatic speech recognition (ASR) achieves strong performance but often incurs high computational costs. This work investigates how to obtain the best LLM-ASR performance efficiently. Through comprehensive and controlled experiments, we find that pretraining the speech encoder before integrating it with the LLM leads to significantly better scaling efficiency than the standard practice of joint post-training of LLM-ASR. Based on this insight, we propose a new multi-stage LLM-ASR training strategy, EFIN: Encoder First Integration. Among all training strategies evaluated, EFIN consistently delivers better performance (relative to 21.1 % CERR) with significantly lower computation budgets ($49.9 \%$ FLOPs). Furthermore, we derive a scaling law that approximates ASR error rates as a computation function, providing practical guidance for LLM-ASR scaling.
Bingshen Mu, Yiwen Shao, Dong Yu 0001, Lei Xie 0001
ASRU4
2025 Continual Pre-training for Codec-Based Speech LLMs: Balancing Understanding and Generation
abstract
Recent advances in speech language models (LLMs) have extended textual LLMs to the speech domain, but balancing speech understanding and generation remains challenging, especially with codec-based representations. We propose a continual pre-training (CPT) framework that adapts a textual LLM to handle codec-discretized speech, mitigating modality mismatch and preserving linguistic reasoning. Our unified model supports both understanding and generation, achieving strong results across ASR, TTS, S2T-Trans, and S2S-Trans. Notably, we present the first end-to-end, single-pass S2S-Trans system using only neural codec tokens, without intermediate transcriptions, translations, or semantic tokens. CPT proves essential for crossmodal alignment and task generalization, making it a powerful tool for building robust, unified speech LLMs.
Jiatong Shi, Jinchuan Tian, Junrui Ni, Hao Zhang 0112, Shinji Watanabe 0001, Dong Yu 0001
ASRU7
2025 Entropy Guided Extrapolative Decoding to Improve Factuality in Large Language Models
abstract
Large language models (LLMs) exhibit impressive natural language capabilities but suffer from hallucination – generating content ungrounded in the realities of training data. Recent work has focused on decoding techniques to improve factuality in decoding by leveraging LLMs’ hierarchical representation of factual knowledge, manipulating the predicted distributions at inference time. Current state-of-the-art approaches refine decoding by contrasting logits from a lower layer with the final layer to exploit information related factuality within the model forward procedure. However, such methods often assume the final layer is most reliable one and the lower layer selection process depends on it. In this work, we first propose logit extrapolation of critical token probabilities beyond the last layer for more accurate contrasting. We additionally employ layer-wise entropy-guided lower layer selection, decoupling the selection process from the final layer. Experiments demonstrate strong performance - surpassing state-of-the-art on multiple different datasets by large margins. Analyses show different kinds of prompts respond to different selection strategies.
Lifeng Jin, Linfeng Song, Haitao Mi, Baolin Peng, Dong Yu 0001
COLING6
2025 WebEvolver: Enhancing Web Agent Self-Improvement with Co-evolving World Model
abstract
Agent self-improvement, where agents autonomously train their underlying Large Language Model (LLM) on self-sampled trajectories, shows promising results but often stagnates in web environments due to limited exploration and under-utilization of pretrained web knowledge.To improve the performance of self-improvement, we propose a novel framework that introduces a co-evolving World Model LLM.This world model predicts the next observation based on the current observation and action within the web environment.The World Model serves dual roles: (1) as a virtual web server generating self-instructed training data to continuously refine the agent's policy, and (2) as an imagination engine during inference, enabling look-ahead simulation to guide action selection for the agent LLM.Experiments in real-world web environments (Mind2Web-Live, WebVoyager, and GAIAweb) show a 10% performance gain over existing self-evolving agents, demonstrating the efficacy and generalizability of our approach, without using any distillation from more powerful close-sourced models 1 .
Tianqing Fang, Hongming Zhang 0009, Zhisong Zhang, Kaixin Ma, Wenhao Yu 0002, Haitao Mi, Dong Yu 0001
EMNLP7
2025 Router-Tuning: A Simple and Effective Approach for Dynamic Depth
abstract
The Mixture of Depths (MoD) was introduced to improve computational efficiency by dynamically skipping less important layers, reducing redundant computation while maintaining model capacity.Despite its promise, existing MoD approaches remain under-explored and face two main challenges: (1) high training costs due to the need to train the entire model along with the routers that determine which layers to skip, and (2) performance degradation when important layers are bypassed.In response to the first issue, we propose Router-Tuning, which fine-tunes only the routers on a small dataset, drastically reducing the computational overhead associated with full model training.For the second challenge, we investigate Router-Tuning across different architectures and granularities, demonstrating its effectiveness on Attention layers and MoE layers.This method preserves the model's performance while significantly enhancing computational and memory efficiency.Extensive experiments demonstrate that our approach delivers competitive results while dramatically improving the computation efficiency, e.g., 21% speedup and only a 0.2% performance drop.
Shwai He, Tao Ge 0001, Guoheng Sun, Bowei Tian, Xiaoyang Wang 0001, Dong Yu 0001
EMNLP6
2025 Recall with Reasoning: Chain-of-Thought Distillation for Mamba's Long-Context Memory and Extrapolation
abstract
Mamba's theoretical infinite-context potential is limited in practice when sequences far exceed training lengths.This work explores unlocking Mamba's long-context memory ability by a simple-yet-effective method, Recall with Reasoning (RwR), by distilling chain-ofthought (CoT) summarization from a teacher model.Specifically, RwR prepends these summarization as CoT prompts during fine-tuning, teaching Mamba to actively recall and reason over long contexts.Experiments on LONG-MEMEVAL and HELMET show that RwR outperforms existing long-term memory methods on the Mamba model.Furthermore, under similar pre-training conditions, RwR improves the long-context performance of Mamba relative to comparable Transformer/hybrid baselines while preserving short-context capabilities, all without changing the architecture.
Jun-Yu Ma, Tianqing Fang, Zhisong Zhang, Hongming Zhang 0009, Haitao Mi, Dong Yu 0001
EMNLP6
2025 Retrieval-augmented GUI Agents with Generative Guidelines
abstract
GUI agents powered by vision-language models (VLMs) show promise in automating complex digital tasks.However, their effectiveness in real-world applications is often limited by scarce training data and the inherent complexity of these tasks, which frequently require longtailed knowledge covering rare, unseen scenarios.We propose RAG-GUI , a lightweight VLM that leverages web tutorials at inference time.RAG-GUI is first warm-started via supervised finetuning (SFT) and further refined through self-guided rejection sampling finetuning (RSF).Designed to be model-agnostic, RAG-GUI functions as a generic plug-in that enhances any VLM-based agent.Evaluated across three distinct tasks, it consistently outperforms baseline agents and surpasses other inference baselines by 2.6% to 13.3% across two model sizes, demonstrating strong generalization and practical plug-and-play capabilities in real-world scenarios.
Ran Xu 0002, Kaixin Ma, Wenhao Yu 0002, Hongming Zhang 0009, Joyce C. Ho, Carl Yang 0001, Dong Yu 0001
EMNLP7
2025 UNCLE: Benchmarking Uncertainty Expressions in Long-Form Generation
abstract
Large Language Models (LLMs) are prone to hallucination, particularly in long-form generations.A promising direction to mitigate hallucination is to teach LLMs to express uncertainty explicitly when they lack sufficient knowledge.However, existing work lacks direct and fair evaluation of LLMs' ability to express uncertainty effectively in long-form generation.To address this gap, we first introduce UNCLE, a benchmark designed to evaluate uncertainty expression in both long-and short-form question answering (QA).UNCLE covers five domains and includes more than 1,000 entities, each with paired short-and long-form QA items.Our dataset is the first to directly link short-and long-form QA through aligned questions and gold-standard answers.Along with UNCLE, we propose a suite of new metrics to assess the models' capabilities to selectively express uncertainty.We then demonstrate that current models fail to convey uncertainty appropriately in long-form generation.We further explore both prompt-based and training-based methods to improve models' performance, with the training-based methods yielding greater gains.Further analysis of alignment gaps between short-and long-form uncertainty expression highlights promising directions for future research using UNCLE.
Ruihan Yang, Caiqi Zhang, Zhisong Zhang, Xinting Huang, Dong Yu 0001, Nigel Collier, Deqing Yang
EMNLP5
2025 Neural Ambisonic Encoding For Multi-Speaker Scenarios Using A Circular Microphone Array
abstract
Spatial audio formats like Ambisonics are playback device layout-agnostic and well-suited for applications such as teleconferencing and virtual reality. Conventional Ambisonic encoding methods often rely on spherical microphone arrays for efficient sound field capture, which limits their flexibility in practical scenarios. We propose a deep learning (DL)-based approach, leveraging a two-stage network architecture for encoding circular microphone array signals into second-order Ambisonics (SOA) in multi-speaker environments. In addition, we introduce: (i) a novel loss function based on spatial power maps to regularize inter-channel correlations of the Ambisonic signals, and (ii) a channel permutation technique to resolve the ambiguity of encoding vertical information using a horizontal circular array. Evaluation on simulated speech and noise datasets shows that our approach consistently outperforms traditional signal processing (SP) and DL-based methods, providing significantly better timbral and spatial quality and higher source localization accuracy. Binaural audio demos with visualizations are available at https://bridgoon97.github.io/NeuralAmbisonicEncoding/.
Vinay Kothapally, Meng Yu 0003, Dong Yu 0001
ICASSP4
2025 BANC: Towards Efficient Binaural Audio Neural Codec for Overlapping Speech
abstract
We introduce BANC, a neural binaural audio codec designed for efficient speech compression in single and two-speaker scenarios while preserving the spatial location information of each speaker. Our key contributions are as follows: 1) The ability of our proposed model to compress and decode overlapping speech. 2) A novel architecture that compresses speech content and spatial cues separately, ensuring the preservation of each speaker’s spatial context after decoding. 3) BANC’s proficiency in reducing the bandwidth required for compressing binaural speech by 48% compared to compressing individual binaural channels. In our evaluation, we employed speech enhancement, room acoustics, and perceptual metrics to assess the accuracy of BANC’s clean speech and spatial cue estimates.
Anton Ratnarajah, Dong Yu 0001
ICASSP3
2025 STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment
abstract
Visual and auditory perception are two crucial ways humans experience the world. Text-to-video generation has made remarkable progress over the past year, but the absence of harmonious audio in generated video limits its broader applications. In this paper, we propose Semantic and Temporal Aligned Video-to-Audio (STA-V2A), an approach that enhances audio generation from videos by extracting both local temporal and global semantic video features and combining these refined video features with text as cross-modal guidance. To address the issue of information redundancy in videos, we propose an onset prediction pretext task for local temporal feature extraction and an attentive pooling module for global semantic feature extraction. To supplement the insufficient semantic information in videos, we propose a Latent Diffusion Model with Text-to-Audio priors initialization and cross-modal guidance. We also introduce Audio-Audio Align, a new metric to assess audio-temporal alignment. Subjective and objective metrics demonstrate that our method surpasses existing Video-to-Audio models in generating audio with better quality, semantic consistency, and temporal alignment. The ablation experiment validated the effectiveness of each module. Audio samples are available at https://y-ren16.github.io/STAV2A.
Yong Ren 0006, Chenxing Li, Manjie Xu, Rilin Chen, Dong Yu 0001
ICASSP7
2025 Preference Alignment Improves Language Model-Based TTS
abstract
Recent advancements in text-to-speech (TTS) have shown that language model (LM)-based systems offer competitive performance to their counterparts. Further optimization can be achieved through preference alignment algorithms, which adjust LMs to align with the preferences of reward models, enhancing the desirability of the generated content. This study presents a thorough empirical evaluation of how preference alignment algorithms, particularly Direct Preference Optimization (DPO), enhance LM-based TTS. With a 1.15B parameter LM-based TTS model, we demonstrate that preference alignment consistently improves intelligibility, speaker similarity, and proxy subjective evaluation scores, with the latter two metrics surpassing even human speech in certain evaluations. We also show preference alignment is applicable to low-resource scenarios and effectively generalized to out-of-domain applications.
Jinchuan Tian, Jiatong Shi, Hao Zhang 0112, Jianwei Yu 0001, Shinji Watanabe 0001, Dong Yu 0001
ICASSP7
2025 SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
abstract
In this paper, we introduce SSR-Speech, a neural codec autoregressive model designed for stable, safe, and robust zero-shot text-based speech editing and text-to-speech synthesis. SSR-Speech is built on a Transformer decoder and incorporates classifier-free guidance to enhance the stability of the generation process. A watermark Encodec is proposed to embed frame-level watermarks into the edited regions of the speech so that which parts were edited can be detected. In addition, the waveform reconstruction leverages the original unedited speech segments, providing superior recovery compared to the Encodec model. Our approach achieves state-of-the-art performance in the RealEdit speech editing task and the LibriTTS text-to-speech task, surpassing previous methods. Furthermore, SSR-Speech excels in multi-span speech editing and also demonstrates remarkable robustness to background sounds. The source code1and demos2are released.
Helin Wang, Meng Yu 0003, Jiarui Hai, Chen Chen 0075, Rilin Chen, Najim Dehak, Dong Yu 0001
ICASSP8
2025 DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?
abstract
Large Language Models (LLMs) and Large Vision-Language Models (LVLMs) have demonstrated impressive language/vision reasoning abilities, igniting the recent trend of building agents for targeted applications such as shopping assistants or AI software engineers. Recently, many data science benchmarks have been proposed to investigate their performance in the data science domain. However, existing data science benchmarks still fall short when compared to real-world data science applications due to their simplified settings. To bridge this gap, we introduce DSBench, a comprehensive benchmark designed to evaluate data science agents with realistic tasks. This benchmark includes 466 data analysis tasks and 74 data modeling tasks, sourced from Eloquence and Kaggle competitions. DSBench offers a realistic setting by encompassing long contexts, multimodal task backgrounds, reasoning with large data files and multi-table structures, and performing end-to-end data modeling tasks. Our evaluation of state-of-the-art LLMs, LVLMs, and agents shows that they struggle with most tasks, with the best agent solving only 34.12% of data analysis tasks and achieving a 34.74% Relative Performance Gap (RPG). These findings underscore the need for further advancements in developing more practical, intelligent, and autonomous data science agents.
Liqiang Jing, Zhehui Huang, Xiaoyang Wang 0001, Wenlin Yao, Wenhao Yu 0002, Kaixin Ma, Hongming Zhang 0009, Xinya Du, Dong Yu 0001
ICLR9
2025 RepoGraph: Enhancing AI Software Engineering with Repository-level Code Graph
abstract
Large Language Models (LLMs) excel in code generation yet struggle with modern AI software engineering tasks. Unlike traditional function-level or file-level coding tasks, AI software engineering requires not only basic coding proficiency but also advanced skills in managing and interacting with code repositories. However, existing methods often overlook the need for repository-level code understanding, which is crucial for accurately grasping the broader context and developing effective solutions. On this basis, we present RepoGraph, a plug-in module that manages a repository-level structure for modern AI software engineering solutions. RepoGraph offers the desired guidance and serves as a repository-wide navigation for AI software engineers. We evaluate RepoGraph on the SWE-bench by plugging it into four different methods of two lines of approaches, where RepoGraph substantially boosts the performance of all systems, leading to a new state-of-the-art among open-source frameworks. Our analyses also demonstrate the extensibility and flexibility of RepoGraph by testing on another repo-level coding benchmark, CrossCodeEval. Our code is available at https://github.com/ozyyshr/RepoGraph.
Siru Ouyang, Wenhao Yu 0002, Kaixin Ma, Zilin Xiao, Zhihan Zhang 0001, Mengzhao Jia, Jiawei Han 0001, Hongming Zhang 0009, Dong Yu 0001
ICLR9
2025 LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
abstract
Recent large language model (LLM)-driven chat assistant systems have integrated memory components to track user-assistant chat histories, enabling more accurate and personalized responses. However, their long-term memory capabilities in sustained interactions remain underexplored. We introduce LongMemEval, a comprehensive benchmark designed to evaluate five core long-term memory abilities of chat assistants: information extraction, multi-session reasoning, temporal reasoning, knowledge updates, and abstention. With 500 meticulously curated questions embedded within freely scalable user-assistant chat histories, LongMemEval presents a significant challenge to existing long-term memory systems, with commercial chat assistants and long-context LLMs showing a 30% accuracy drop on memorizing information across sustained interactions. We then present a unified framework that breaks down the long-term memory design into three stages: indexing, retrieval, and reading. Built upon key experimental insights, we propose several memory design optimizations including session decomposition for value granularity, fact-augmented key expansion for indexing, and time-aware query expansion for refining the search scope. Extensive experiments show that these optimizations greatly improve both memory recall and downstream question answering on LongMemEval. Overall, our study provides valuable resources and guidance for advancing the long-term memory capabilities of LLM-based chat assistants, paving the way toward more personalized and reliable conversational AI. Our benchmark and code are publicly available at https://github.com/xiaowu0162/LongMemEval.
Di Wu 0054, Wenhao Yu 0002, Yuwei Zhang 0001, Kai-Wei Chang 0001, Dong Yu 0001
ICLR6
2025 DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories Search
abstract
Enhancing the capability of large language models (LLMs) in reasoning has gained significant attention in recent years. Previous studies have demonstrated the effectiveness of various prompting strategies in aiding LLMs in reasoning (called "reasoning actions"), such as step-by-step thinking, reflecting before answering, solving with programs, and their combinations. However, these approaches often applied static, predefined reasoning actions uniformly to all questions, without considering the specific characteristics of each question or the capability of the task-solving LLM. In this paper, we propose DOTS, an approach enabling LLMs to reason Dynamically via Optimal reasoning Trajectories Search, tailored to the specific characteristics of each question and the inherent capability of the task-solving LLM. Our approach involves three key steps: i) defining atomic reasoning action modules that can be composed into various reasoning action trajectories; ii) searching for the optimal action trajectory for each training question through iterative exploration and evaluation for the specific task-solving LLM; and iii) using the collected optimal trajectories to train an LLM to plan for the reasoning trajectories of unseen questions. In particular, we propose two learning paradigms, i.e., fine-tuning an external LLM as a planner to guide the task-solving LLM, or directly fine-tuning the task-solving LLM with an internalized capability for reasoning actions planning. Our experiments across eight reasoning tasks show that our method consistently outperforms static reasoning techniques and the vanilla instruction tuning approach. Further analysis reveals that our method enables LLMs to adjust their computation based on problem complexity, allocating deeper thinking and reasoning to harder problems.
Murong Yue, Wenlin Yao, Haitao Mi, Dian Yu 0001, Ziyu Yao 0002, Dong Yu 0001
ICLR6
2025 Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
abstract
Reinforcement Learning with Human Feedback (RLHF) has achieved great success in aligning large language models (LLMs) with human preferences. Prevalent RLHF approaches are reward-based, following the Bradley-Terry (BT) model assumption, which may not fully capture the complexity of human preferences. In this paper, we explore RLHF under a general preference framework and approach it from a game-theoretic perspective. Specifically, we formulate the problem as a two-player game and propose a novel online algorithm, iterative Nash policy optimization (INPO). The key idea is to let the policy play against itself via no- regret learning, thereby approximating the Nash policy. Unlike previous methods, INPO bypasses the need for estimating the expected win rate for individual responses, which typically incurs high computational or annotation costs. Instead, we introduce a new loss objective that is directly minimized over a preference dataset. We provide theoretical analysis for our approach and demonstrate its effectiveness through experiments on various representative benchmarks. With an LLaMA-3-8B-based SFT model, INPO achieves a 42.6% length-controlled win rate on AlpacaEval 2.0 and a 37.8% win rate on Arena-Hard, showing substantial improvement over the state-of-the-art online RLHF algorithms.
Dian Yu 0001, Baolin Peng, Linfeng Song, Mingyue Huo, Nan Jiang 0008, Haitao Mi, Dong Yu 0001
ICLR9
2025 Do NOT Think That Much for 2+3=? On the Overthinking of Long Reasoning Models
abstract
The remarkable performance of long reasoning models can be attributed to their ability to emulate human-like long-time thinking during inference. These models employ extended chain-of-thought (CoT) processes, exploring multiple strategies to enhance problem-solving capabilities. However, a critical question remains: How to intelligently and efficiently scale computational resources during testing. This paper presents the first comprehensive study on the prevalent issue of overthinking in these models, where long reasoning models generate redundant solutions that contribute minimally to accuracy and diversity, thereby wasting computational resources on simple problems with minimal benefit. We introduce novel efficiency metrics from both outcome and process perspectives to evaluate the rational use of computational resources by long reasoning models. Using a self-training paradigm, we propose strategies to mitigate overthinking, simplifying reasoning processes without compromising accuracy. Experimental results show that our approach successfully reduces computational overhead while preserving model performance across a range of testsets with varying difficulty levels, such as GSM8K, MATH500, GPQA, and AIME. Our code is open-source and available at https://github.com/galaxyChen/overthinking.
Zhiwei He 0002, Jianhui Pang, Dian Yu 0001, Linfeng Song, Qiuzhi Liu, Mengfei Zhou, Zhuosheng Zhang 0001, Rui Wang 0015, Zhaopeng Tu, Haitao Mi, Dong Yu 0001
ICML14
2025 BridgeVoC: Neural Vocoder with Schrödinger Bridge
abstract
While previous diffusion-based neural vocoders typically follow a noise-to-data generation pipe-line, the linear-degradation prior of the mel-spectrogram is often neglected, resulting in limited generation quality. By revisiting the vocoding task and excavating its connection with the signal restoration task, this paper proposes a time-frequency (T-F) domain-based neural vocoder with the Schrödinger Bridge, called BridgeVoC, which is the first to follow the data-to-data generation paradigm. Specifically, the mel-spectrogram can be projected into the target linear-scale domain and regarded as a degraded spectral representation with a deficient rank distribution. Based on this, the Schrödinger Bridge is leveraged to establish a connection between the degraded and target data distributions. During the inference stage, starting from the degraded representation, the target spectrum can be gradually restored rather than generated from a Gaussian noise process. Quantitative experiments on LJSpeech and LibriTTS show that BridgeVoC achieves faster inference and surpasses existing diffusion-based vocoder baselines, while also matching or exceeding non-diffusion state-of-the-art methods across evaluation metrics.
Rilin Chen, Meng Yu 0003, Chengshi Zheng, Dong Yu 0001, Andong Li
IJCAI7
2025 EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
Jiarui Hai, Yong Xu 0004, Hao Zhang 0112, Chenxing Li, Helin Wang, Mounya Elhilali, Dong Yu 0001
INTERSPEECH7
2025 Video-to-Audio Generation with Fine-grained Temporal Semantics
Chenxing Li, Rilin Chen, Dong Yu 0001
INTERSPEECH5
2025 Efficient Multilingual ASR Finetuning via LoRA Language Experts
Yiwen Shao, Jianheng Zhuo, Chenda Li, Liliang Tang, Dong Yu 0001, Yanmin Qian
INTERSPEECH6
2025 Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model
Yong Ren 0006, Chenxing Li, Duzhen Zhang, Yujie Chen 0006, Manjie Xu, Ruibo Fu, Shan Yang 0001, Dong Yu 0001
INTERSPEECH10
2025 Scaling beyond Denoising: Submitted System and Findings in URGENT Challenge 2025
Zhihang Sun, Andong Li, Rilin Chen, Meng Yu 0003, Chengshi Zheng, Yi Zhou 0014, Dong Yu 0001
INTERSPEECH8
2025 Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning
Chenxing Li, Yong Ren 0006, Yujie Chen 0006, Ruibo Fu, Shan Yang 0001, Dong Yu 0001
INTERSPEECH8
2025 Towards Diverse and Efficient Audio Captioning via Diffusion Models
Manjie Xu, Chenxing Li, Yong Ren 0006, Ruibo Fu, Dong Yu 0001
INTERSPEECH7
2025 WAKE: Watermarking Audio with Key Enrichment
Yaoxun Xu, Jianwei Yu 0001, Hangting Chen, Zhiyong Wu 0001, Xixin Wu, Dong Yu 0001, Rongzhi Gu, Yi Luo 0004
INTERSPEECH6
2025 VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining
Jianheng Zhuo, Yifan Yang 0005, Yiwen Shao, Yong Xu 0004, Dong Yu 0001, Kai Yu 0004, Xie Chen 0001
INTERSPEECH5
2025 From Continuous to Discrete: Cross-Domain Collaborative General Speech Enhancement via Hierarchical Language Models
abstract
This paper introduces OmniGSE, a novel general speech enhancement (GSE) framework designed to mitigate the diverse distortions that speech signals encounter in real-world scenarios. These distortions include background noise, reverberation, bandwidth limitations, signal clipping, and network packet loss. Existing methods typically focus on optimizing for a single type of distortion, often struggling to effectively handle the simultaneous presence of multiple distortions in complex scenarios. OmniGSE bridges this gap by integrating the strengths of discriminative and generative approaches through a two-stage architecture that enables cross-domain collaborative optimization. In the first stage, continuous features are enhanced using a lightweight channel-split NAC-RoFormer. In the second stage, discrete tokens are generated to reconstruct high-quality speech through language models. Specifically, we designed a hierarchical language model structure consisting of a RootLM and multiple BranchLMs. The RootLM models general acoustic features across codebook layers, while the BranchLMs explicitly capture the progressive relationships between different codebook levels. Experimental results demonstrate that OmniGSE surpasses existing models across multiple benchmarks, particularly excelling in scenarios involving compound distortions. These findings underscore the framework's potential for robust and versatile speech enhancement in real-world applications.
Zhaoxi Mu, Rilin Chen, Andong Li, Meng Yu 0003, Xinyu Yang 0001, Dong Yu 0001
ACM Multimedia6
2025 UniGist: Towards General and Hardware-aligned Sequence-level Long Context Compression
abstract
Large language models are increasingly capable of handling long-context inputs, but the memory overhead of KV cache remains a major bottleneck for general-purpose deployment. While many compression strategies have been explored, sequence-level compression is particularly challenging due to its tendency to lose important details. We present UniGist, a gist token-based long context compression framework that removes the need for chunk-wise training, enabling the model to learn how to compress and utilize long-range context during training. To fully exploit the sparsity, we introduce a gist shift trick that transforms the attention layout into a right-aligned block structure and develop a block-table-free sparse attention kernel based on it. UniGist further supports one-pass training and flexible chunk sizes during inference, allowing efficient and adaptive context processing. Experiments across multiple long-context tasks show that UniGist significantly improves compression quality, with especially strong performance in recalling details and long-range dependency modeling.
Chenlong Deng, Zhisong Zhang, Kelong Mao, Shuaiyi Li, Tianqing Fang, Hongming Zhang 0009, Haitao Mi, Dong Yu 0001, Zhicheng Dou
NeurIPS8
2025 The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models
abstract
Improving the reasoning capabilities of large language models (LLMs) typically requires supervised fine-tuning with labeled data or computationally expensive sampling. We introduce Unsupervised Prefix Fine-Tuning (UPFT), which leverages the observation of Prefix Self-Consistency -- the shared initial reasoning steps across diverse solution trajectories -- to enhance LLM reasoning efficiency. By training exclusively on the initial prefix substrings (as few as 8 tokens), UPFT removes the need for labeled data or exhaustive sampling. Experiments on reasoning benchmarks show that UPFT matches the performance of supervised methods such as Rejection Sampling Fine-Tuning, while reducing training time by 75\% and sampling cost by 99\%. Further analysis reveals that errors tend to appear in later stages of the reasoning process and that prefix-based training preserves the model’s structural knowledge. This work demonstrates how minimal unsupervised fine-tuning can unlock substantial reasoning gains in LLMs, offering a scalable and resource-efficient alternative to conventional approaches.
Ke Ji, Qiuzhi Liu, Zhiwei He 0002, Benyou Wang, Zhaopeng Tu, Haitao Mi, Dong Yu 0001
NeurIPS12
2025 LeVo: High-Quality Song Generation with Multi-Preference Alignment
abstract
Recent advances in large language models (LLMs) and audio language models have significantly improved music generation, particularly in lyrics-to-song generation. However, existing approaches still struggle with the complex composition of songs and the scarcity of high-quality data, leading to limitations in audio quality, musicality, instruction following, and vocal-instrument harmony. To address these challenges, we introduce LeVo, a language model based framework consisting of LeLM and Music Codec. LeLM is capable of parallel modeling of two types of tokens: mixed tokens, which represent the combined audio of vocals and accompaniment to achieve better vocal-instrument harmony, and dual-track tokens, which separately encode vocals and accompaniment for high-quality song generation. It employs two decoder-only transformers and a modular extension training strategy to prevent interference between different token types. To further enhance musicality and instruction following ability, we introduce a multi-preference alignment method based on Direct Preference Optimization (DPO). This method handles diverse human preferences through a semi-automatic data construction process and post-training. Experimental results demonstrate that LeVo significantly outperforms existing open-source methods in both objective and subjective metrics, while performing competitively with industry systems. Ablation studies further justify the effectiveness of our designs. Audio examples and source code are available at https://levo-demo.github.io and https://github.com/tencent-ailab/songgeneration.
Shun Lei, Yaoxun Xu, Huaicheng Zhang, Wei Tan 0011, Hangting Chen, Yixuan Zhang 0005, Haina Zhu, Shuai Wang 0016, Zhiyong Wu 0001, Dong Yu 0001
NeurIPS12
2025 MPS-Prover: Advancing Stepwise Theorem Proving by Multi-Perspective Search and Data Curation
abstract
Automated Theorem Proving (ATP) in formal languages remains a formidable challenge in AI, demanding rigorous logical deduction and navigating vast search spaces. While large language models (LLMs) have shown promising performance, existing stepwise provers often suffer from biased search guidance, leading to inefficiencies and suboptimal proof strategies. This paper introduces the Multi-Perspective Search Prover (MPS-Prover), a novel stepwise ATP system designed to overcome these limitations. MPS-Prover incorporates two key innovations: a highly effective post-training data curation strategy that prunes approximately 40\% of redundant training data without sacrificing performance, and a multi-perspective tree search mechanism. This search integrates a learned critic model with strategically designed heuristic rules to diversify tactic selection, prevent getting trapped in unproductive states, and enhance search robustness. Extensive evaluations demonstrate that MPS-Prover achieves state-of-the-art performance on multiple challenging benchmarks, including miniF2F and ProofNet, outperforming prior 7B parameter models. Furthermore, our analyses reveal that MPS-Prover generates significantly shorter and more diverse proofs compared to existing stepwise and whole-proof methods, highlighting its efficiency and efficacy. Our work advances the capabilities of LLM-based formal reasoning and offers a robust framework and a comprehensive analysis for developing more powerful theorem provers.
Zhenwen Liang, Linfeng Song, Tao Yang 0033, Haitao Mi, Dong Yu 0001
NeurIPS6
2025 Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
abstract
Large Language Models (LLMs) show great promise in complex reasoning, with Reinforcement Learning with Verifiable Rewards (RLVR) being a key enhancement strategy. However, a prevalent issue is ``superficial self-reflection'', where models fail to robustly verify their own outputs. We introduce RISE (Reinforcing Reasoning with Self-Verification), a novel online RL framework designed to tackle this. RISE explicitly and simultaneously trains an LLM to improve both its problem-solving and self-verification abilities within a single, integrated RL process. The core mechanism involves leveraging verifiable rewards from an outcome verifier to provide on-the-fly feedback for both solution generation and self-verification tasks. In each iteration, the model generates solutions, then critiques its own on-policy generated solutions, with both trajectories contributing to the policy update. Extensive experiments on diverse mathematical reasoning benchmarks show that RISE consistently improves model's problem-solving accuracy while concurrently fostering strong self-verification skills. Our analyses highlight the advantages of online verification and the benefits of increased verification compute. Additionally, RISE models exhibit more frequent and accurate self-verification behaviors during reasoning. These advantages reinforce RISE as a flexible and effective path towards developing more robust and self-aware reasoners.
Zhiwei He 0002, Wenxuan Wang 0001, Pinjia He, Zhaopeng Tu, Haitao Mi, Dong Yu 0001
NeurIPS9
2025 Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training
abstract
Mixture-of-Experts (MoE) architectures within Large Reasoning Models (LRMs) have achieved impressive reasoning capabilities by selectively activating experts to facilitate structured cognitive processes. Despite notable advances, existing reasoning models often suffer from cognitive inefficiencies like overthinking and underthinking. To address these limitations, we introduce a novel inference-time steering methodology called Reinforcing Cognitive Experts (RICE), designed to improve reasoning depth and efficiency without additional training or complex heuristics. Leveraging normalized Pointwise Mutual Information (nPMI), we systematically identify specialized experts, termed cognitive experts that orchestrate meta-level reasoning operations characterized by tokens like <think>. Empirical evaluations with leading MoE-based LRMs (DeepSeek-R1 and Qwen3-235B) on rigorous quantitative and scientific reasoning benchmarks (AIME and GPQA Diamond) demonstrate noticeable and consistent improvements in reasoning accuracy, cognitive efficiency, and cross-domain generalization. Crucially, our lightweight approach substantially outperforms prevalent reasoning-steering techniques, such as prompt design and decoding constraints, while preserving the model's general instruction-following skills. These results highlight reinforcing cognitive experts as a promising, practical, and interpretable direction to enhance cognitive efficiency within advanced reasoning models.
Yue Wang 0039, Zhiwei He 0002, Qiuzhi Liu, Yunzhi Yao, Wenxuan Wang 0001, Ruotian Ma, Haitao Mi, Ningyu Zhang 0001, Zhaopeng Tu, Dong Yu 0001
NeurIPS15
2025 Thoughts Are All Over the Place: On the Underthinking of Long Reasoning Models
abstract
Long reasoning models (LRMs) such as OpenAI's o1 and DeepSeek's R1 have demonstrated remarkable abilities in complex reasoning tasks by scaling test-time compute and exhibiting human-like deep thinking. However, we identify a phenomenon we term underthinking, where LRMs frequently switch between different reasoning thoughts without sufficiently exploring promising paths to reach a correct solution. This behavior leads to inadequate depth of reasoning and decreased performance, particularly on challenging mathematical problems. To systematically analyze this issue, we conduct experiments on three challenging test sets and two representative open-source LRMs, revealing that frequent thought switching correlates with incorrect responses. We introduce a novel metric to quantify underthinking by measuring token efficiency in incorrect answers. To address underthinking, we propose a decoding strategy with thought switching penalty (Tip) that discourages premature transitions between thoughts, encouraging deeper exploration of each reasoning path. Experimental results demonstrate that our approach improves accuracy across challenging datasets without requiring model fine-tuning. Our findings contribute to understanding reasoning inefficiencies in LRMs and offer a practical solution to enhance their problem-solving capabilities. Our code is open-source and available at https://github.com/wangyuenlp/underthinking.
Yue Wang 0039, Qiuzhi Liu, Zhiwei He 0002, Linfeng Song, Dian Yu 0001, Juntao Li 0005, Zhuosheng Zhang 0001, Rui Wang 0015, Zhaopeng Tu, Haitao Mi, Dong Yu 0001
NeurIPS14
2025 Improving LLM General Preference Alignment via Optimistic Online Mirror Descent
abstract
Reinforcement learning from human feedback (RLHF) has demonstrated remarkable effectiveness in aligning large language models (LLMs) with human preferences. Many existing alignment approaches rely on the Bradley-Terry (BT) model assumption, which assumes the existence of a ground-truth reward for each prompt-response pair. However, this assumption can be overly restrictive when modeling complex human preferences. In this paper, we drop the BT model assumption and study LLM alignment under general preferences, formulated as a two-player game. Drawing on theoretical insights from learning in games, we integrate optimistic online mirror descent into our alignment framework to approximate the Nash policy. Theoretically, we demonstrate that our approach achieves an $\mathcal{O}(T^{-1})$ bound on the duality gap, improving upon the previous $\mathcal{O}(T^{-1/2})$ result. Meanwhile, it enjoys a linear convergence rate in the last iterate, a property not achieved by previous methods. More importantly, we implement our method and show through experiments that it outperforms state-of-the-art RLHF algorithms across multiple representative benchmarks.
Dian Yu 0001, Tao Ge 0001, Linfeng Song, Zhichen Zeng 0001, Haitao Mi, Nan Jiang 0008, Dong Yu 0001
NeurIPS8
2025 Information-theoretic complementary prompts for improved continual text classification
Duzhen Zhang, Yong Ren 0006, Chenxing Li, Dong Yu 0001, Tielin Zhang
Neural Networks4
2024 WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
abstract
Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, Dong Yu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Hongliang He 0002, Wenlin Yao, Kaixin Ma, Wenhao Yu 0002, Yong Dai 0001, Hongming Zhang 0009, Zhen-Zhong Lan, Dong Yu 0001
ACL (1)8
2024 SportsMetrics: Blending Text and Numerical Data to Understand Information Fusion in LLMs
abstract
Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang, Hassan Foroosh, Dong Yu, Fei Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang 0001, Hassan Foroosh, Dong Yu 0001, Fei Liu 0004
ACL (1)6
2024 CLOMO: Counterfactual Logical Modification with Large Language Models
abstract
Yinya Huang, Ruixin Hong, Hongming Zhang, Wei Shao, Zhicheng Yang, Dong Yu, Changshui Zhang, Xiaodan Liang, Linqi Song. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yinya Huang, Ruixin Hong, Hongming Zhang 0009, Wei Shao 0009, Dong Yu 0001, Changshui Zhang, Xiaodan Liang, Linqi Song
ACL (1)6
2024 Make-A-Voice: Revisiting Voice Large Language Models as Scalable Multilingual and Multitask Learners
abstract
Rongjie Huang, Chunlei Zhang, Yongqi Wang, Dongchao Yang, Jinchuan Tian, Zhenhui Ye, Luping Liu, Zehan Wang, Ziyue Jiang, Xuankai Chang, Jiatong Shi, Chao Weng, Zhou Zhao, Dong Yu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Rongjie Huang 0001, Dongchao Yang, Jinchuan Tian, Zhenhui Ye, Luping Liu, Zehan Wang 0001, Ziyue Jiang 0001, Xuankai Chang, Jiatong Shi, Chao Weng, Zhou Zhao 0001, Dong Yu 0001
ACL (1)14
2024 Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer
abstract
While recent advancements in speech language models have achieved significant progress, they face remarkable challenges in modeling the long acoustic sequences of neural audio codecs.In this paper, we introduce Generative Pretrained Speech Transformer (GPST), a hierarchical transformer designed for efficient speech language modeling.GPST quantizes audio waveforms into two distinct types of discrete speech representations and integrates them within a hierarchical transformer architecture, allowing for a unified one-stage generation process and enhancing Hi-Res audio generation capabilities.By training on large corpora of speeches in an end-to-end unsupervised manner, GPST can generate syntactically consistent speech with diverse speaker identities.Given a brief 3-second prompt, GPST can produce natural and coherent personalized speech, demonstrating in-context learning abilities.Moreover, our approach can be easily extended to spoken cross-lingual speech generation by incorporating multi-lingual semantic tokens and universal acoustic tokens.Experimental results indicate that GPST significantly outperforms the existing speech language models in terms of word error rate, speech quality, and speaker similarity.See https://youngsheen.github. io/GPST/demo for demo samples.
Yongxin Zhu 0003, Dan Su 0002, Liqiang He, Linli Xu 0002, Dong Yu 0001
ACL (1)5
2024 A Knowledge Plug-and-Play Test Bed for Open-domain Dialogue Generation
abstract
Knowledge-based, open-domain dialogue generation aims to build chit-chat systems that talk to humans using mined support knowledge. Many types and sources of knowledge have previously been shown to be useful as support knowledge. Even in the era of large language models, response generation grounded in knowledge retrieved from additional up-to-date sources remains a practically important approach. While prior work using single-source knowledge has shown a clear positive correlation between the performances of knowledge selection and response generation, there are no existing multi-source datasets for evaluating support knowledge retrieval. Further, prior work has assumed that the knowledge sources available at test time are the same as during training. This unrealistic assumption unnecessarily handicaps models, as new knowledge sources can become available after a model is trained. In this paper, we present a high-quality benchmark named multi-source Wizard of Wikipedia (Ms.WoW) for evaluating multi-source dialogue knowledge selection and response generation. Unlike existing datasets, it contains clean support knowledge, grounded at the utterance level and partitioned into multiple knowledge sources. We further propose a new challenge, dialogue knowledge plug-and-play, which aims to test an already trained dialogue model on using new support knowledge from previously unseen sources in a zero-shot fashion.
Xiangci Li, Linfeng Song, Lifeng Jin, Haitao Mi, Jessica Ouyang 0001, Dong Yu 0001
LREC/COLING6
2024 MinT: Boosting Generalization in Mathematical Reasoning via Multi-view Fine-tuning
abstract
Reasoning in mathematical domains remains a significant challenge for relatively small language models (LMs). Many current methods focus on specializing LMs in mathematical reasoning and rely heavily on distilling knowledge from powerful yet inefficient large LMs (LLMs). In this work, we explore a new direction that avoids over-reliance on LLM teachers, introducing a multi-view fine-tuning method that efficiently exploits existing mathematical problem datasets with diverse annotation styles. Our approach uniquely considers the various annotation formats as different “views” that may help each other and leverage them in training the model. By postpending distinct instructions to input questions, models can learn to generate solutions in diverse formats in a flexible manner. Experimental results show that our strategy enables relatively small LMs to outperform prior approaches that heavily rely on knowledge distillation, as well as carefully established baselines. Additionally, the proposed method grants the models promising generalization ability across various views and datasets, and the capability to learn from inaccurate or incomplete noisy data. We hope our multi-view training paradigm could inspire future studies in other machine reasoning domains.
Zhenwen Liang, Dian Yu 0001, Xiaoman Pan, Wenlin Yao, Xiangliang Zhang 0001, Dong Yu 0001
LREC/COLING7
2024 Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language Models
abstract
Retrieval-augmented language model (RALM) represents a significant advancement in mitigating factual hallucination by leveraging external knowledge sources.However, the reliability of the retrieved information is not always guaranteed, and the retrieval of irrelevant data can mislead the response generation.Moreover, standard RALMs frequently neglect their intrinsic knowledge due to the interference from retrieved information.In instances where the retrieved information is irrelevant, RALMs should ideally utilize their intrinsic knowledge or, in the absence of both intrinsic and retrieved knowledge, opt to respond with "unknown" to avoid hallucination.In this paper, we introduces CHAIN-OF-NOTE (CON), a novel approach to improve robustness of RALMs in facing noisy, irrelevant documents and in handling unknown scenarios.The core idea of CON is to generate sequential reading notes for each retrieved document, enabling a thorough evaluation of their relevance to the given question and integrating this information to formulate the final answer.Our experimental results show that GPT-4, when equipped with CON, outperforms the CHAIN-OF-THOUGHT approach.Besides, we utilized GPT-4 to create 10K CON data, subsequently trained on LLaMa-2 7B model.Our experiments across four open-domain QA benchmarks show that fine-tuned RALMs equipped with CON significantly outperform standard fine-tuned RALMs.
Wenhao Yu 0002, Hongming Zhang 0009, Xiaoman Pan, Peixin Cao, Kaixin Ma, Hongwei Wang 0001, Dong Yu 0001
EMNLP8
2024 Dense X Retrieval: What Retrieval Granularity Should We Use?
abstract
Dense retrieval has become a prominent method to obtain relevant context or world knowledge in open-domain NLP tasks.When we use a learned dense retriever on a retrieval corpus at inference time, an often-overlooked design choice is the retrieval unit in which the corpus is indexed, e.g.document, passage, or sentence.We discover that the retrieval unit choice significantly impacts the performance of both retrieval and downstream tasks.Distinct from the typical approach of using passages or sentences, we introduce a novel retrieval unit, proposition, for dense retrieval.Propositions are defined as atomic expressions within text, each encapsulating a distinct factoid and presented in a concise, self-contained natural language format.We conduct an empirical comparison of different retrieval granularity.Our experiments reveal that indexing a corpus by fine-grained units such as propositions significantly outperforms passage-level units in retrieval tasks.Moreover, constructing prompts with fine-grained retrieved units for retrievalaugmented language models improves the performance of downstream QA tasks given a specific computation budget.
Hongwei Wang 0010, Wenhao Yu 0002, Kaixin Ma, Hongming Zhang 0009, Dong Yu 0001
EMNLP8
2024 When Reasoning Meets Information Aggregation: A Case Study with Sports Narratives
abstract
Reasoning is most powerful when an LLM accurately aggregates relevant information.We examine the critical role of information aggregation in reasoning by requiring the LLM to analyze sports narratives.To succeed at this task, an LLM must infer points from actions, identify related entities, attribute points accurately to players and teams, and compile key statistics to draw conclusions.We conduct comprehensive experiments with real NBA basketball data and present SPORTSGEN, a new method to synthesize game narratives.By synthesizing data, we can rigorously evaluate LLMs' reasoning capabilities under complex scenarios with varying narrative lengths and density of information.Our findings show that most models, including GPT-4o, often fail to accurately aggregate basketball scores due to frequent scoring patterns.Open-source models like Llama-3 further suffer from significant score hallucinations.Finally, the effectiveness of reasoning is influenced by narrative complexity, information density, and domain-specific terms, highlighting the challenges in analytical reasoning tasks. 1 * Work done during Yebowen Hu's internship; Kaiqiang Song and Sangwoo Cho were full-time researchers at Tencent AI Lab, Seattle, USA at the time of this work. https://github.com/YebowenHu/SportsGenAnalyze the team-player affiliations and play-by-play descriptions below to determine the total points scored by each team (player).Please explain your reasoning step by step and provide the final results in JSON format.Start with: {New York Knicks: 0, Denver Nuggets: 0} ({Andrea Bargnani: 0, Timofey Mozgov: 0, ...
Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang 0001, Wenlin Yao, Hassan Foroosh, Dong Yu 0001, Fei Liu 0004
EMNLP7
2024 Learn Beyond The Answer: Training Language Models with Reflection for Mathematical Reasoning
abstract
Supervised fine-tuning enhances the problemsolving abilities of language models across various mathematical reasoning tasks.To maximize such benefits, existing research focuses on broadening the training set with various data augmentation techniques, which is effective for standard single-round question-answering settings.Our work introduces a novel technique aimed at cultivating a deeper understanding of the training problems at hand, enhancing performance not only in standard settings but also in more complex scenarios that require reflective thinking.Specifically, we propose reflective augmentation, a method that embeds problem reflection into each training instance.It trains the model to consider alternative perspectives and engage with abstractions and analogies, thereby fostering a thorough comprehension through reflective reasoning.Extensive experiments validate the achievement of our aim, underscoring the unique advantages of our method and its complementary nature relative to existing augmentation techniques. 1 Question Answer Question
Zhihan Zhang 0001, Tao Ge 0001, Zhenwen Liang, Wenhao Yu 0002, Dian Yu 0001, Mengzhao Jia, Dong Yu 0001, Meng Jiang 0001
EMNLP7
2024 UniX-Encoder: A Universal X-Channel Speech Encoder for AD-HOC Microphone Array Speech Processing
abstract
The speech field is evolving to solve more challenging scenarios, such as multi-channel recordings with multiple simultaneous talkers. In response to the diversity of microphone configurations in use, we introduce the UniX-Encoder, a universal encoder for multi-channel speech recordings. The UniX-Encoder is versatile, catering to a variety of speech tasks, and seamlessly integrates with any microphone array, whether in single-talker or multi-talker environments. Our research enhances previous multi-channel speech processing efforts in four aspects: 1) Adaptability: Contrasting traditional models constrained to certain microphone array configurations, our encoder is universally compatible. 2) Multi-Task Capability: Contrasting previous systems that were designed for single-task applications, the UniX-Encoder serves as a versatile upstream model, capable of extracting features for diverse speech tasks. 3) Self-Supervised Training: The UniX-Encoder is pretrained without the need for labeled multi-channel data. 4) End-to-End Integration: In contrast to models that first beamform then process single-channels, our encoder offers an end-to-end solution, bypassing explicit beamforming or separation. To validate its effectiveness, we tested the UniX-Encoder on a synthetic multi-channel dataset from the LibriSpeech corpus. Across various tasks, including ASR and speaker diarization, our encoder consistently outperformed combinations such as the WavLM model with the BeamformIt frontend.
Zili Huang, Yiwen Shao, Shixiong Zhang 0001, Dong Yu 0001
ICASSP4
2024 SPATIALCODEC: Neural Spatial Speech Coding
abstract
In this work, we address the challenge of encoding speech captured by a microphone array using deep learning techniques with the aim of preserving and accurately reconstructing crucial spatial cues embedded in multi-channel recordings. We propose a neural spatial audio coding framework that achieves a high compression ratio, leveraging single-channel neural sub-band codec and SpatialCodec. Our approach encompasses two phases: (i) a neural sub-band codec is designed to encode the reference channel with low bit rates, and (ii), a SpatialCodec captures relative spatial information for accurate multi-channel reconstruction at the decoder end. In addition, we also propose novel evaluation metrics to assess the spatial cue preservation: (i) spatial similarity, which calculates cosine similarity on a spatially intuitive beamspace, and (ii), beamformed audio quality. Our system shows superior spatial performance compared with high bitrate baselines and black-box neural architecture. Demos are available at https://xzwy.github.io/SpatialCodecDemo. Codes and models are available at https://github.com/XZWY/SpatialCodec.
Zhongweiyang Xu, Yong Xu 0004, Vinay Kothapally, Heming Wang, Muqiao Yang, Dong Yu 0001
ICASSP6
2024 uSee: Unified Speech Enhancement And Editing with Conditional Diffusion Models
abstract
Speech enhancement aims to improve the quality of speech signals in terms of quality and intelligibility, and speech editing refers to the process of editing the speech according to specific user needs. In this paper, we propose a Unified Speech Enhancement and Editing (uSee) model with conditional diffusion models to handle various tasks at the same time in a generative manner. Specifically, by providing multiple types of conditions including self-supervised learning embeddings and proper text prompts to the score-based diffusion model, we can enable controllable generation of the unified speech enhancement and editing model to perform corresponding actions on the source speech. Our experiments show that our proposed uSee model can achieve superior performance in both speech denoising and dereverberation compared to other related generative speech enhancement models, and can perform speech editing given desired environmental sound text description, signal-to-noise ratios (SNR), and room impulse responses (RIR). Demos of the generated speech are available at https://muqiaoy.github.io/usee.
Muqiao Yang, Yong Xu 0004, Zhongweiyang Xu, Heming Wang, Bhiksha Raj, Dong Yu 0001
ICASSP7
2024 The Trickle-down Impact of Reward Inconsistency on RLHF
abstract
Standard practice within Reinforcement Learning from Human Feedback (RLHF) involves optimizing against a Reward Model (RM), which itself is trained to reflect human preferences for desirable generations. A notable subject that is understudied is the (in-)consistency of RMs --- whether they can recognize the semantic changes to different prompts and appropriately adapt their reward assignments --- and their impact on the downstream RLHF model. In this paper, we visit a series of research questions relevant to RM inconsistency: (1) How can we measure the consistency of reward models? (2) How consistent are the existing RMs and how can we improve them? (3) In what ways does reward inconsistency influence the chatbots resulting from the RLHF model training? We propose **Contrast Instruction** -- a benchmarking strategy for the consistency of RM. Each example in **Contrast Instruction** features a pair of lexically similar instructions with different ground truth responses. A consistent RM is expected to rank the corresponding instruction and response higher than other combinations. We observe that current RMs trained with the standard ranking objective fail miserably on \contrast{} compared to average humans. To show that RM consistency can be improved efficiently without using extra training budget, we propose two techniques **ConvexDA** and **RewardFusion**, which enhance reward consistency through extrapolation during the RM training and inference stage, respectively. We show that RLHF models trained with a more consistent RM yield more useful responses, suggesting that reward inconsistency exhibits a trickle-down effect on the downstream RLHF process.
Lingfeng Shen, Linfeng Song, Lifeng Jin, Baolin Peng, Haitao Mi, Daniel Khashabi, Dong Yu 0001
ICLR8
2024 Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment
abstract
We consider the problem of multi-objective alignment of foundation models with human preferences, which is a critical step towards helpful and harmless AI systems. However, it is generally costly and unstable to fine-tune large foundation models using reinforcement learning (RL), and the multi-dimensionality, heterogeneity, and conflicting nature of human preferences further complicate the alignment process. In this paper, we introduce Rewards-in-Context (RiC), which conditions the response of a foundation model on multiple rewards in its prompt context and applies supervised fine-tuning for alignment. The salient features of RiC are simplicity and adaptivity, as it only requires supervised fine-tuning of a single foundation model and supports dynamic adjustment for user preferences during inference time. Inspired by the analytical solution of an abstracted convex optimization problem, our dynamic inference-time adjustment method approaches the Pareto-optimal solution for multiple objectives. Empirical evidence demonstrates the efficacy of our method in aligning both Large Language Models (LLMs) and diffusion models to accommodate diverse rewards with only around 10% GPU hours compared with multi-objective RL baseline.
Rui Yang 0010, Xiaoman Pan, Feng Luo 0003, Han Zhong 0001, Dong Yu 0001, Jianshu Chen
ICML6
2024 Prompt-guided Precise Audio Editing with Diffusion Models
abstract
Audio editing involves the arbitrary manipulation of audio content through precise control. Although text-guided diffusion models have made significant advancements in text-to-audio generation, they still face challenges in finding a flexible and precise way to modify target events within an audio track. We present a novel approach, referred to as **PPAE**, which serves as a general module for diffusion models and enables precise audio editing. The editing is based on the input textual prompt only and is entirely training-free. We exploit the cross-attention maps of diffusion models to facilitate accurate local editing and employ a hierarchical local-global pipeline to ensure a smoother editing process. Experimental results highlight the effectiveness of our method in various editing tasks.
Manjie Xu, Chenxing Li, Duzhen Zhang, Dan Su 0002, Dong Yu 0001
ICML6
2024 RIR-SF: Room Impulse Response Based Spatial Feature for Target Speech Recognition in Multi-Channel Multi-Speaker Scenarios
Yiwen Shao, Shixiong Zhang 0001, Dong Yu 0001
INTERSPEECH3
2024 Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment
Yiwen Shao, Shixiong Zhang 0001, Yong Xu 0004, Meng Yu 0003, Dong Yu 0001, Daniel Povey, Sanjeev Khudanpur
INTERSPEECH5
2024 Comparing Discrete and Continuous Space LLMs for Speech Recognition
Yaoxun Xu, Shixiong Zhang 0001, Jianwei Yu 0001, Zhiyong Wu 0001, Dong Yu 0001
INTERSPEECH5
2024 Polarity Calibration for Opinion Summarization
abstract
Yuanyuan Lei, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang, Ruihong Huang, Dong Yu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yuanyuan Lei 0001, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang 0001, Ruihong Huang, Dong Yu 0001
NAACL-HLT6
2024 Sub-Sentence Encoder: Contrastive Learning of Propositional Semantic Representations
abstract
Sihao Chen, Hongming Zhang, Tong Chen, Ben Zhou, Wenhao Yu, Dian Yu, Baolin Peng, Hongwei Wang, Dan Roth, Dong Yu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Hongming Zhang 0009, Ben Zhou, Wenhao Yu 0002, Dian Yu 0001, Baolin Peng, Hongwei Wang 0010, Dan Roth 0001, Dong Yu 0001
NAACL-HLT10
2024 A Closer Look at the Self-Verification Abilities of Large Language Models in Logical Reasoning
abstract
Ruixin Hong, Hongming Zhang, Xinyu Pang, Dong Yu, Changshui Zhang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Ruixin Hong, Hongming Zhang 0009, Xinyu Pang, Dong Yu 0001, Changshui Zhang
NAACL-HLT4
2024 MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning
abstract
Fuxiao Liu, Xiaoyang Wang, Wenlin Yao, Jianshu Chen, Kaiqiang Song, Sangwoo Cho, Yaser Yacoob, Dong Yu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Fuxiao Liu, Xiaoyang Wang 0001, Wenlin Yao, Jianshu Chen, Kaiqiang Song, Sangwoo Cho, Yaser Yacoob, Dong Yu 0001
NAACL-HLT8
2024 From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
abstract
Xuansheng Wu, Wenlin Yao, Jianshu Chen, Xiaoman Pan, Xiaoyang Wang, Ninghao Liu, Dong Yu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Xuansheng Wu, Wenlin Yao, Jianshu Chen, Xiaoman Pan, Xiaoyang Wang 0001, Ninghao Liu 0001, Dong Yu 0001
NAACL-HLT7
2024 Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing
abstract
Despite the impressive capabilities of Large Language Models (LLMs) on various tasks, they still struggle with scenarios that involves complex reasoning and planning. Self-correction and self-learning emerge as viable solutions, employing strategies that allow LLMs to refine their outputs and learn from self-assessed rewards. Yet, the efficacy of LLMs in self-refining its response, particularly in complex reasoning and planning task, remains dubious. In this paper, we introduce AlphaLLM for the self-improvements of LLMs, which integrates Monte Carlo Tree Search (MCTS) with LLMs to establish a self-improving loop, thereby enhancing the capabilities of LLMs without additional annotations. Drawing inspiration from the success of AlphaGo, AlphaLLM addresses the unique challenges of combining MCTS with LLM for self-improvement, including data scarcity, the vastness search spaces of language tasks, and the subjective nature of feedback in language tasks. AlphaLLM is comprised of prompt synthesis component, an efficient MCTS approach tailored for language tasks, and a trio of critic models for precise feedback. Our experimental results in mathematical reasoning tasks demonstrate that AlphaLLM significantly enhances the performance of LLMs without additional annotations, showing the potential for self-improvement in LLMs.
Baolin Peng, Linfeng Song, Lifeng Jin, Dian Yu 0001, Haitao Mi, Dong Yu 0001
NeurIPS8
2024 Spatialemb: Extract and Encode Spatial Information for 1-Stage Multi-Channel Multi-Speaker ASR on Arbitrary Microphone Arrays
abstract
Spatial information is a critical clue for multi-channel multispeaker target speech recognition. Most state-of-the-art multi-channel Automatic Speech Recognition (ASR) systems extract spatial features only during the speech separation stage, followed by standard single-channel ASR on the separated speech. This approach results in an inefficient, lengthy pipeline and sub-optimal ASR performance due to the accumulated errors from preprocessing modules. Furthermore, most spatial feature extraction methods depend on the knowledge of speaker positions and microphone topology, making the systems reliant on specific settings and challenging to adapt to new equipment. In this work, we propose a solution to these issues with a lightweight embedding module named SpatialEmb, which extracts and encodes spatial information directly for the ASR model, supporting both fixed and arbitrary microphone topology. We conduct comprehensive experiments on AliMeeting, a real meeting corpus, to determine the optimal model design for SpatialEmb in terms of both performance and efficiency. Our best model trained with 105 hours Train-Ali-far achieves 17.04% and 20.32% character error rates (CER) on the Eval and Test sets, establishing a new state-of-the-art result with the same training data.
Yiwen Shao, Yong Xu 0004, Sanjeev Khudanpur, Dong Yu 0001
SLT4
2024 Advancing Multi-Talker ASR Performance With Large Language Models
abstract
Recognizing overlapping speech from multiple speakers in conversational scenarios is one of the most challenging problem for automatic speech recognition (ASR). Serialized output training (SOT) is a classic method to address multi-talker ASR, with the idea of concatenating transcriptions from multiple speakers according to the emission times of their speech for training. However, SOT-style transcriptions, derived from concatenating multiple related utterances in a conversation, depend significantly on modeling long contexts. Therefore, compared to traditional methods that primarily emphasize encoder performance in attention-based encoderdecoder (AED) architectures, a novel approach utilizing large language models (LLMs) that leverages the capabilities of pre-trained decoders may be better suited for such complex and challenging scenarios. In this paper, we propose an LLM-based SOT approach for multi-talker ASR, leveraging pre-trained speech encoder and LLM, fine-tuning them on multi-talker dataset using appropriate strategies. Experimental results demonstrate that our approach surpasses traditional AED-based methods on the simulated dataset LibriMix and achieves state-of-the-art performance on the evaluation set of the real-world dataset AMI, outperforming the AED model trained with 1000 times more supervised data in previous works.
Mohan Shi, Zengrui Jin, Yaoxun Xu, Yong Xu 0004, Shixiong Zhang 0001, Yiwen Shao, Dong Yu 0001
SLT9
2024 SMRU: Split-And-Merge Recurrent-Based UNet For Acoustic Echo Cancellation And Noise Suppression
abstract
The proliferation of deep neural networks has spawned the rapid development of acoustic echo cancellation and noise suppression, and plenty of prior arts have been proposed, which yield promising performance. Nevertheless, they rarely consider the deployment generality in different processing scenarios, such as edge devices, and cloud processing. To this end, this paper proposes a general model, termed SMRU, to cover different application scenarios. The novelty lies in two-fold. First, a multi-scale band split layer and band merge layer are proposed to effectively fuse local frequency bands for lower complexity modeling. Besides, by simulating the multi-resolution feature modeling characteristic of the classical UNet structure, a novel recurrent-dominated UNet is devised. It consists of multiple variable frame rate blocks, each of which involves the causal time down-/upsampling layer with varying compression ratios and the dualpath structure for inter- and intra-band modeling. The model is configured from $50 \mathrm{M} / \mathrm{s}$ to $6.8 \mathrm{G} / \mathrm{s}$ in terms of MACs, and the experimental results show that the proposed approach yields competitive or even better performance over existing baselines, and has the full potential to adapt to more general scenarios with varying complexity requirements.
Zhihang Sun, Andong Li, Rilin Chen, Hao Zhang 0112, Meng Yu 0003, Yi Zhou 0014, Dong Yu 0001
SLT7
2024 Enhanced Acoustic Howling Suppression via Hybrid Kalman Filter and Deep Learning Models
abstract
This paper presents a comprehensive study addressing the challenging problem of acoustic howling suppression (AHS) through the fusion of Kalman filter and deep learning techniques. We introduce two integration approaches: HybridAHS, which concatenates Kalman and neural networks (NN), and NeuralKalmanAHS, where NN modules are embedded inside the Kalman filter for signal and parameter estimation. In HybridAHS, we explore two implementation methods. One is trained offline using pre-processed signals with a light training burden, while the other employs a recursive training strategy with training signals generated adaptively. The offline model serves as an initialization for recursively training the other model. With NeuralKalmanAHS, we harness the power of NN modules to refine the reference signal and improve covariance matrices estimation in the Kalman filter, resulting in enhanced feedback suppression. Our methods capitalize on the strengths of traditional and deep learning-based AHS techniques. We have explored different variants of combining Kalman filter and NN and systematically compared their howling suppression performance, providing users with versatile solutions for addressing AHS. Furthermore, by employing the proposed recursive training, we effectively mitigate the mismatch issues that plagued previous NN-based AHS methods. Extensive experimental results show the superiority of our approach over baseline techniques.
Hao Zhang 0112, Yixuan Zhang 0005, Meng Yu 0003, Dong Yu 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Generating User-Engaging News Headlines
abstract
Pengshan Cai, Kaiqiang Song, Sangwoo Cho, Hongwei Wang, Xiaoyang Wang, Hong Yu, Fei Liu, Dong Yu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Pengshan Cai, Kaiqiang Song, Sangwoo Cho, Hongwei Wang 0010, Xiaoyang Wang 0001, Hong Yu 0001, Fei Liu 0004, Dong Yu 0001
ACL (1)8
2023 Faithful Question Answering with Monte-Carlo Planning
abstract
Although large language models demonstrate remarkable question-answering performances, revealing the intermediate reasoning steps that the models faithfully follow remains challenging.In this paper, we propose FAME (FAithful question answering with MontE-carlo planning) to answer questions based on faithful reasoning steps.The reasoning steps are organized as a structured entailment tree, which shows how premises are used to produce intermediate conclusions that can prove the correctness of the answer.We formulate the task as a discrete decision-making problem and solve it through the interaction of a reasoning environment and a controller.The environment is modular and contains several basic task-oriented modules, while the controller proposes actions to assemble the modules.Since the search space could be large, we introduce a Monte-Carlo planning algorithm to do a look-ahead search and select actions that will eventually lead to highquality steps.FAME achieves advanced performance on the standard benchmark.It can produce valid and faithful reasoning steps compared with large language models with a much smaller model size.
Ruixin Hong, Hongming Zhang 0009, Dong Yu 0001, Changshui Zhang
ACL (1)4
2023 SafeConv: Explaining and Correcting Conversational Unsafe Behavior
abstract
One of the main challenges open-domain endto-end dialogue systems, or chatbots, face is the prevalence of unsafe behavior, such as toxic languages and harmful suggestions.However, existing dialogue datasets do not provide enough annotation to explain and correct such unsafe behavior.In this work, we construct a new dataset called SAFECONV for the research of conversational safety: (1) Besides the utterancelevel safety labels, SAFECONV also provides unsafe spans in an utterance, information able to indicate which words contribute to the detected unsafe behavior; (2) SAFECONV provides safe alternative responses to continue the conversation when unsafe behavior detected, guiding the conversation to a gentle trajectory.By virtue of the comprehensive annotation of SAFECONV, we benchmark three powerful models for the mitigation of conversational unsafe behavior, including a checker to detect unsafe utterances, a tagger to extract unsafe spans, and a rewriter to convert an unsafe response to a safe version.Moreover, we explore the huge benefits brought by combining the models for explaining the emergence of unsafe behavior and detoxifying chatbots.Experiments show that the detected unsafe behavior could be well explained with unsafe spans and popular chatbots could be detoxified by a huge extent.The dataset is available at https://github.com/mianzhang/SafeConv.
Lifeng Jin, Linfeng Song, Haitao Mi, Wenliang Chen, Dong Yu 0001
ACL (1)6
2023 Neuralecho: Hybrid of Full-Band and Sub-Band Recurrent Neural Network For Acoustic Echo Cancellation and Speech Enhancement
abstract
This paper presents a hybrid of full-band and sub-band recurrent neural network (RNN) model, named NeuralEcho, to jointly solve echo and noise suppression. The full-band model part processes the signal’s entire frequency bands as a whole, while the sub-band model part divides the features into sub-bands and processes each sub-band separately. This approach allows the model to capture both the fine-grained local details of the sub-band processing and the global context of the full-band processing. The single-channel model is then generalized to accommodate a range of input channel numbers. Experimental results show that the hybrid model outperforms the conventional full-band models in terms of objective speech quality metrics and speech recognition accuracy. This suggests that the hybrid approach of full-band and sub-band processing can be a promising direction for future research in the field of speech enhancement.
Meng Yu 0003, Yong Xu 0004, Shixiong Zhang 0001, Dong Yu 0001
ASRU5
2023 Neuralkalman: A Learnable Kalman Filter for Acoustic Echo Cancellation
abstract
The robustness of the Kalman filter to double talk and its rapid convergence make it a popular approach for addressing acoustic echo cancellation (AEC) challenges. However, the inability to model nonlinearity and the need to tune control parameters cast limitations on such adaptive filtering algorithms. In this paper, we integrate the frequency domain Kalman filter (FDKF) and deep neural networks (DNNs) into a hybrid method, called NeuralKalman, to leverage the advantages of deep learning and adaptive filtering algorithms. Specifically, we employ a DNN to estimate nonlinearly distorted far-end signals, a transition factor, and the nonlinear transition function in the state equation of the FDKF algorithm. Experimental results show that the proposed NeuralKalman improves the performance of FDKF significantly and outperforms strong baseline methods.
Yixuan Zhang 0005, Meng Yu 0003, Hao Zhang 0112, Dong Yu 0001, DeLiang Wang
ASRU4
2023 How do Words Contribute to Sentence Semantics? Revisiting Sentence Embeddings with a Perturbation Method
abstract
Wenlin Yao, Lifeng Jin, Hongming Zhang, Xiaoman Pan, Kaiqiang Song, Dian Yu, Dong Yu, Jianshu Chen. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Wenlin Yao, Lifeng Jin, Hongming Zhang 0009, Xiaoman Pan, Kaiqiang Song, Dian Yu 0001, Dong Yu 0001, Jianshu Chen
EACL7
2023 Friend-training: Learning from Models of Different but Related Tasks
abstract
Current self-training methods such as standard self-training, co-training, tri-training, and others often focus on improving model performance on a single task, utilizing differences in input features, model architectures, and training processes.However, many tasks in natural language processing are about different but related aspects of language, and models trained for one task can be great teachers for other related tasks.In this work, we propose friendtraining, a cross-task self-training framework, where models trained to do different tasks are used in an iterative training, pseudo-labeling, and retraining process to help each other for better selection of pseudo-labels.With two dialogue understanding tasks, conversational semantic role labeling and dialogue rewriting, chosen for a case study, we show that the models trained with the friend-training framework achieve the best performance compared to strong baselines.
Lifeng Jin, Linfeng Song, Haitao Mi, Xiabing Zhou, Dong Yu 0001
EACL6
2023 Bridging Continuous and Discrete Spaces: Interpretable Sentence Representation Learning via Compositional Operations
abstract
Traditional sentence embedding models encode sentences into vector representations to capture useful properties such as the semantic similarity between sentences.However, in addition to similarity, sentence semantics can also be interpreted via compositional operations such as sentence fusion or difference.It is unclear whether the compositional semantics of sentences can be directly reflected as compositional operations in the embedding space.To more effectively bridge the continuous embedding and discrete text spaces, we explore the plausibility of incorporating various compositional properties into the sentence embedding space that allows us to interpret embedding transformations as compositional sentence operations.We propose INTERSENT, an end-toend framework for learning interpretable sentence embeddings that supports compositional sentence operations in the embedding space.Our method optimizes operator networks and a bottleneck encoder-decoder model to produce meaningful and interpretable sentence embeddings.Experimental results demonstrate that our method significantly improves the interpretability of sentence embeddings on four textual generation tasks over existing approaches while maintaining strong performance on traditional semantic similarity tasks. 1 .
James Y. Huang, Wenlin Yao, Kaiqiang Song, Hongming Zhang 0009, Muhao Chen 0001, Dong Yu 0001
EMNLP6
2023 More Than Spoken Words: Nonverbal Message Extraction and Generation
abstract
Nonverbal messages (NM) such as speakers' facial expressions and speed of speech are essential for face-to-face communication, and they can be regarded as implicit knowledge as they are usually not included in existing dialogue understanding or generation tasks.This paper introduces the task of extracting NMs in written text and generating NMs for spoken text.Previous studies merely focus on extracting NMs from relatively small-scale well-structured corpora such as movie scripts wherein NMs are enclosed in parentheses by scriptwriters, which greatly decreases the difficulty of extraction.To enable extracting NMs from unstructured corpora, we annotate the first NM extraction dataset for Chinese based on novels and develop three baselines to extract single-span or multi-span NM of a target utterance from its surrounding context.Furthermore, we use the extractors to extract 749K (context, utterance, NM) triples from Chinese novels and investigate whether we can use them to improve NM generation via semi-supervised learning.Experimental results demonstrate that the automatically extracted triples can serve as high-quality augmentation data of clean triples extracted from scripts to generate more relevant, fluent, valid, and factually consistent 1 NMs than the purely supervised generator, and the resulting generator can in turn help Chinese dialogue understanding tasks such as dialogue machine reading comprehension and emotion classification by simply adding the predicted "unspoken" NM to each utterance or narrative in inputs.
Dian Yu 0001, Xiaoyang Wang 0001, Wanshun Chen, Longyue Wang, Haitao Mi, Dong Yu 0001
EMNLP7
2023 Trinet: Stabilizing Self-Supervised Learning From Complete or Slow Collapse
abstract
Self-supervised learning (SSL) models confront challenges of abrupt informational collapse or slow dimensional collapse. We propose TriNet, which introduces a novel triple-branch architecture for preventing collapse and stabilizing the pretraining. TriNet learns the SSL latent embedding space and incorporates it to a higher level space for predicting pseudo target vectors generated by a frozen teacher. Our experimental results show that the proposed method notably stabilizes and accelerates pre-training and achieves a relative word error rate reduction (WERR) of 6.06% compared to the state-of- the-art (SOTA) Data2vec for a downstream benchmark ASR task. We will release our code at https://github.com/tencent-ailab/.
Lixin Cao, Jun Wang 0091, Ben Yang, Dan Su 0002, Dong Yu 0001
ICASSP5
2023 Deep Neural Mel-Subband Beamformer for in-Car Speech Separation
abstract
While current deep learning (DL)-based beamforming techniques have been proved effective in speech separation, they are often designed to process narrow-band (NB) frequencies independently which results in higher computational costs and inference times, making them unsuitable for real-world use. In this paper, we propose DL-based mel-subband spatio-temporal beamformer to perform speech separation in a car environment with reduced computation cost and inference time. As opposed to conventional subband (SB) approaches, our framework uses a mel-scale based subband selection strategy which ensures a fine-grained processing for lower frequencies where most speech formant structure is present, and coarse-grained processing for higher frequencies. In a recursive way, robust frame-level beamforming weights are determined for each speaker location/zone in a car from the estimated subband speech and noise covariance matrices. Furthermore, proposed framework also estimates and suppresses any echoes from the loudspeaker(s) by using the echo reference signals. We compare the performance of our proposed framework to several NB, SB, and full-band (FB) processing techniques in terms of speech quality and recognition metrics. Based on experimental evaluations on simulated and real-world recordings, we find that our proposed framework achieves better separation performance over all SB and FB approaches and achieves performance closer to NB processing techniques while requiring lower computing cost.
Vinay Kothapally, Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Dong Yu 0001
ICASSP5
2023 Knowledge-in-Context: Towards Knowledgeable Semi-Parametric Language Models
Xiaoman Pan, Wenlin Yao, Hongming Zhang 0009, Dian Yu 0001, Dong Yu 0001, Jianshu Chen
ICLR5
2023 Bayes Risk CTC: Controllable CTC Alignment in Sequence-to-Sequence Tasks
Jinchuan Tian, Brian Yan, Jianwei Yu 0001, Chao Weng, Dong Yu 0001, Shinji Watanabe 0001
ICLR5
2023 Zoneformer: On-device Neural Beamformer For In-car Multi-zone Speech Separation, Enhancement and Echo Cancellation
Yong Xu 0004, Vinay Kothapally, Meng Yu 0003, Shixiong Zhang 0001, Dong Yu 0001
INTERSPEECH5
2023 Bayes Risk Transducer: Transducer with Controllable Alignment Prediction
Jinchuan Tian, Jianwei Yu 0001, Hangting Chen, Brian Yan, Chao Weng, Dong Yu 0001, Shinji Watanabe 0001
INTERSPEECH6
2023 Multi-mode Neural Speech Coding Based on Deep Generative Networks
Shan Yang 0001, Yupeng Shi, Yuyong Kang, Dan Su 0002, Shidong Shang, Dong Yu 0001
INTERSPEECH9
2023 Compressed MoE ASR Model Based on Knowledge Distillation and Quantization
Yuping Yuan, Zhao You, Shulin Feng, Dan Su 0002, Yanchun Liang 0001, Xiaohu Shi, Dong Yu 0001
INTERSPEECH7
2023 Hybrid AHS: A Hybrid of Kalman Filter and Deep Learning for Acoustic Howling Suppression
Hao Zhang 0112, Meng Yu 0003, Yuzhong Wu, Dong Yu 0001
INTERSPEECH5
2023 Text-Only Domain Adaptation for End-to-End Speech Recognition through Down-Sampling Acoustic Representation
abstract
Mapping two modalities, speech and text, into a shared representation space, is a research topic of using text-only data to improve end-to-end automatic speech recognition (ASR) performance in new domains. However, the length of speech representation and text representation is inconsistent. Although the previous method up-samples the text representation to align with acoustic modality, it may not match the expected actual duration. In this paper, we proposed novel representations match strategy through down-sampling acoustic representation to align with text modality. By introducing a continuous integrate-and-fire (CIF) module generating acoustic representations consistent with token length, our ASR model can learn unified representations from both modalities better, allowing for domain adaptation using text-only data of the target domain. Experiment results of new domain data demonstrate the effectiveness of the proposed method.
Jiaxu Zhu, Weinan Tong, Yaoxun Xu, Changhe Song, Zhiyong Wu 0001, Zhao You, Dan Su 0002, Dong Yu 0001, Helen M. Meng
INTERSPEECH8
2023 Thrust: Adaptively Propels Large Language Models with External Knowledge
abstract
Although large-scale pre-trained language models (PTLMs) are shown to encode rich knowledge in their model parameters, the inherent knowledge in PTLMs can be opaque or static, making external knowledge necessary. However, the existing information retrieval techniques could be costly and may even introduce noisy and sometimes misleading knowledge. To address these challenges, we propose the instance-level adaptive propulsion of external knowledge (IAPEK), where we only conduct the retrieval when necessary. To achieve this goal, we propose to model whether a PTLM contains enough knowledge to solve an instance with a novel metric, Thrust, which leverages the representation distribution of a small amount of seen instances. Extensive experiments demonstrate that Thrust is a good measurement of models' instance-level knowledgeability. Moreover, we can achieve higher cost-efficiency with the Thrust score as the retrieval indicator than the naive usage of external knowledge on 88% of the evaluated tasks with 26% average performance improvement. Such findings shed light on the real-world practice of knowledge-enhanced LMs with a limited budget for knowledge seeking due to computation latency or costs.
Hongming Zhang 0009, Xiaoman Pan, Wenlin Yao, Dong Yu 0001, Jianshu Chen
NeurIPS5
2023 Search-engine-augmented dialogue response generation with cheaply supervised query production
Ante Wang, Linfeng Song, Qi Liu 0049, Haitao Mi, Longyue Wang, Zhaopeng Tu, Jinsong Su, Dong Yu 0001
Artif. Intell.8
2023 Discover, Explain, Improve: An Automatic Slice Detection Benchmark for Natural Language Processing
abstract
Abstract Pretrained natural language processing (NLP) models have achieved high overall performance, but they still make systematic errors. Instead of manual error analysis, research on slice detection models (SDMs), which automatically identify underperforming groups of datapoints, has caught escalated attention in Computer Vision for both understanding model behaviors and providing insights for future model training and designing. However, little research on SDMs and quantitative evaluation of their effectiveness have been conducted on NLP tasks. Our paper fills the gap by proposing a benchmark named “Discover, Explain, Improve (DEIm)” for classification NLP tasks along with a new SDM Edisa. Edisa discovers coherent and underperforming groups of datapoints; DEIm then unites them under human-understandable concepts and provides comprehensive evaluation tasks and corresponding quantitative metrics. The evaluation in DEIm shows that Edisa can accurately select error-prone datapoints with informative semantic features that summarize error patterns. Detecting difficult datapoints directly boosts model performance without tuning any original model parameters, showing that discovered slices are actionable for users.1
Wenyue Hua, Lifeng Jin, Linfeng Song, Haitao Mi, Dong Yu 0001
Trans. Assoc. Comput. Linguistics6
2023 OpenFact: Factuality Enhanced Open Knowledge Extraction
abstract
Abstract We focus on the factuality property during the extraction of an OpenIE corpus named OpenFact, which contains more than 12 million high-quality knowledge triplets. We break down the factuality property into two important aspects—expressiveness and groundedness—and we propose a comprehensive framework to handle both aspects. To enhance expressiveness, we formulate each knowledge piece in OpenFact based on a semantic frame. We also design templates, extra constraints, and adopt human efforts so that most OpenFact triplets contain enough details. For groundedness, we require the main arguments of each triplet to contain linked Wikidata1 entities. A human evaluation suggests that the OpenFact triplets are much more accurate and contain denser information compared to OPIEC-Linked (Gashteovski et al., 2019), one recent high-quality OpenIE corpus grounded to Wikidata. Further experiments on knowledge base completion and knowledge base question answering show the effectiveness of OpenFact over OPIEC-Linked as supplementary knowledge to Wikidata as the major KG.
Linfeng Song, Ante Wang, Xiaoman Pan, Hongming Zhang 0009, Dian Yu 0001, Lifeng Jin, Haitao Mi, Jinsong Su, Yue Zhang 0004, Dong Yu 0001
Trans. Assoc. Comput. Linguistics10
2023 Towards Unified All-Neural Beamforming for Time and Frequency Domain Speech Separation
abstract
Recently, frequency domain all-neural beamforming methods have achieved remarkable progress for multichannel speech separation. In parallel, the integration of time domain network structure and beamforming also gains significant attention. This study proposes a novel all-neural beamforming method in time domain and makes an attempt to unify the all-neural beamforming pipelines for time domain and frequency domain multichannel speech separation. The proposed model consists of two modules: separation and beamforming. Both modules perform temporal-spectral-spatial modeling and are trained from end-to-end using a joint loss function. The novelty of this study lies in two folds. Firstly, a time domain directional feature conditioned on the direction of the target speaker is proposed, which can be jointly optimized within the time domain architecture to enhance target signal estimation. Secondly, an all-neural beamforming network in time domain is designed to refine the pre-separated results. This module features with parametric time-variant beamforming coefficient estimation, without explicitly following the derivation of optimal filters that may lead to an upper bound. The proposed method is evaluated on simulated reverberant overlapped speech data derived from the AISHELL-1 corpus. Experimental results demonstrate significant performance improvements over frequency domain state-of-the-arts, ideal magnitude masks and existing time domain neural beamforming methods.
Rongzhi Gu, Shixiong Zhang 0001, Yuexian Zou, Dong Yu 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Unsupervised TTS Acoustic Modeling for TTS With Conditional Disentangled Sequential VAE
abstract
In this paper, we propose a novel unsupervised textto-speech acoustic model training scheme, named UTTS, which does not require text-audio pairs. UTTS is a multi-speaker speech synthesizer that supports zero-shot voice cloning, it is developed from a perspective of disentangled speech representation learning. The framework offers a flexible choice of a speaker's duration model, timbre feature (identity) and content for TTS inference. We leverage recent advancements in self-supervised speech representation learning as well as speech synthesis frontend techniques for system development. Specifically, we employ our recently formulated Conditional Disentangled Sequential Variational Auto-encoder (C-DSVAE) as the backbone UTTS AM, which offers well-structured content representations given unsupervised alignment (UA) as condition during training. For UTTS inference, we utilize a lexicon to map input text to the phoneme sequence, which is expanded to the frame-level forced alignment (FA) with a speaker-dependent duration model. Then, we develop an alignment mapping module that converts FA to UA. Finally, the C-DSVAE, serving as the self-supervised TTS AM, takes the predicted UA and a target speaker embedding to generate the mel spectrogram, which is ultimately converted to waveform with a neural vocoder. We show how our method enables speech synthesis without using a paired TTS corpus in AM development stage. Experiments demonstrate that UTTS can synthesize speech of high naturalness and intelligibility measured by human and objective evaluations. Audio samples are available at our demo page
Jiachen Lian, Gopala Krishna Anumanchipalli, Dong Yu 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Integrating Lattice-Free MMI Into End-to-End Speech Recognition
abstract
In automatic speech recognition (ASR) research, discriminative criteria have achieved superior performance in DNN-HMM systems. Given this success, the adoption of discriminative criteria is promising to boost the performance of end-to-end (E2E) ASR systems. With this motivation, previous works have introduced the minimum Bayesian risk (MBR, one of the discriminative criteria) into E2E ASR systems. However, the effectiveness and efficiency of the MBR-based methods are compromised: the MBR criterion is only used in system training, which creates a mismatch between training and decoding; the on-the-fly decoding process in MBR-based methods results in the need for pre-trained models and slow training speeds. To this end, novel algorithms are proposed in this work to integrate another widely used discriminative criterion, lattice-free maximum mutual information (LF-MMI), into E2E ASR systems not only in the training stage but also in the decoding process. The proposed LF-MMI training and decoding methods show their effectiveness on two widely used E2E frameworks: Attention-Based Encoder-Decoders (AEDs) and Neural Transducers (NTs). Compared with MBR-based methods, the proposed LF-MMI method: maintains the consistency between training and decoding; eschews the on-the-fly decoding process; trains from randomly initialized models with superior training efficiency. Experiments suggest that the LF-MMI method outperforms its MBR counterparts and consistently leads to statistically significant performance improvements on various frameworks and datasets from 30 hours to 14.3 k hours. The proposed method achieves state-of-the-art (SOTA) results on Aishell-1 (CER 4.10%) and Aishell-2 (CER 5.02%) datasets. Code is released1.
Jinchuan Tian, Jianwei Yu 0001, Chao Weng, Yuexian Zou, Dong Yu 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2023 D$^{2}$PSG: Multi-Party Dialogue Discourse Parsing as Sequence Generation
abstract
Conversational discourse analysis aims to extract the interactions between dialogue turns, which is crucial for modeling complex multi-party dialogues. As the benchmarks are still limited in size and human annotations are costly, the current standard approaches apply pretrained language models, but they still require randomly initialized classifiers to make predictions. These classifiers usually require massive data to work smoothly with the pretrained encoder, causing severe data hunger issue. We propose two convenient strategies to formulate this task as a sequence generation problem, where classifier decisions are carefully converted into sequence of tokens. We then adopt a pretrained T5 1 model to solve this task so that no parameters are randomly initialized. We also leverage the descriptions of the discourse relations to help model understand their meanings. Experiments on two popular benchmarks show that our approach outperforms previous state-of-the-art models by a large margin, and it is also more robust in zero-shot and few-shot settings.
Ante Wang, Linfeng Song, Lifeng Jin, Junfeng Yao, Haitao Mi, Chen Lin 0001, Jinsong Su, Dong Yu 0001
IEEE ACM Trans. Audio Speech Lang. Process.8
2023 Diffsound: Discrete Diffusion Model for Text-to-Sound Generation
abstract
Generating sound effects that people want is an important topic. However, there are limited studies in this area for sound generation. In this study, we investigate generating sound conditioned on a text prompt and propose a novel text-to-sound generation framework that consists of a text encoder, a Vector Quantized Variational Autoencoder (VQ-VAE), a token-decoder, and a vocoder. The framework first uses the token-decoder to transfer the text features extracted from the text encoder to a mel-spectrogram with the help of VQ-VAE, and then the vocoder is used to transform the generated mel-spectrogram into a waveform. We found that the token-decoder significantly influences the generation performance. Thus, we focus on designing a good token-decoder in this study. We begin with th21e traditional autoregressive (AR) token-decoder, which has shown state-of-the-art performance in previous sound generation works. However, the AR token-decoder always predicts the mel-spectrogram tokens one by one in order, which may introduce the unidirectional bias and accumulation of errors problems. Moreover, with the AR token-decoder, the sound generation time increases linearly with the sound duration. To overcome the shortcomings introduced by AR token-decoders, we propose a non-autoregressive token-decoder based on the discrete diffusion model, named Diffsound. Specifically, the Diffsound model predicts all of the mel-spectrogram tokens in one step and then refines the predicted tokens in the next step, so the best-predicted results can be obtained by iteration. Our experiments show that our proposed Diffsound model not only produces better text-to-sound generation results when compared with the AR token-decoder but also has a faster generation speed,i.e., MOS: 3.56v.s2.786, and the generation speed is five times faster than the AR decoder. Furthermore, to automatically assess the quality of generated samples, we define three different objective evaluation metricsi.e., Fréchet Inception Distance (FID), Kullback-Leibler (KL), and audio caption loss, which can comprehensively assess the relevance and fidelity of the generated samples.
Dongchao Yang, Jianwei Yu 0001, Helin Wang, Chao Weng, Yuexian Zou, Dong Yu 0001
IEEE ACM Trans. Audio Speech Lang. Process.7
2022 Hierarchical Context Tagging for Utterance Rewriting
abstract
Utterance rewriting aims to recover coreferences and omitted information from the latest turn of a multi-turn dialogue. Recently, methods that tag rather than linearly generate sequences have proven stronger in both in- and out-of-domain rewriting settings. This is due to a tagger's smaller search space as it can only copy tokens from the dialogue context. However, these methods may suffer from low coverage when phrases that must be added to a source utterance cannot be covered by a single context span. This can occur in languages like English that introduce tokens such as prepositions into the rewrite for grammaticality. We propose a hierarchical context tagger (HCT) that mitigates this issue by predicting slotted rules (e.g., "besides _") whose slots are later filled with context spans. HCT (i) tags the source string with token-level edit actions and slotted rules and (ii) fills in the resulting rule slots with spans from the dialogue context. This rule tagging allows HCT to add out-of-context tokens and multiple spans at once; we further cluster the rules to truncate the long tail of the rule distribution. Experiments on several benchmarks show that HCT can outperform state-of-the-art rewriting systems by ~2 BLEU points.
Lisa Jin, Linfeng Song, Lifeng Jin, Dong Yu 0001, Daniel Gildea
AAAI4
2022 Improving Machine Reading Comprehension with Contextualized Commonsense Knowledge
abstract
To perform well on a machine reading comprehension (MRC) task, machine readers usually require commonsense knowledge that is not explicitly mentioned in the given documents.This paper aims to extract a new kind of structured knowledge from scripts and use it to improve MRC.We focus on scripts as they contain rich verbal and nonverbal messages, and two relevant messages originally conveyed by different modalities during a short time period may serve as arguments of a piece of commonsense knowledge as they function together in daily communications.To save human efforts to name relations, we propose to represent relations implicitly by situating such an argument pair in a context and call it contextualized knowledge.To use the extracted knowledge to improve MRC, we compare several fine-tuning strategies to use the weakly-labeled MRC data constructed based on contextualized knowledge and further design a teacher-student paradigm with multiple teachers to facilitate the transfer of knowledge in weakly-labeled MRC data.Experimental results show that our paradigm outperforms other methods that use weaklylabeled data and improves a state-of-the-art baseline by 4.3% in accuracy on a Chinese multiple-choice MRC dataset C 3 , wherein most of the questions require unstated prior knowledge.We also seek to transfer the knowledge to other tasks by simply adapting the resulting student reader, yielding a 2.9% improvement in F1 on a relation extraction dataset DialogRE, demonstrating the potential usefulness of the knowledge for non-MRC tasks that require document comprehension.Interior.Runaway office.Day.
Kai Sun 0006, Dian Yu 0001, Jianshu Chen, Dong Yu 0001, Claire Cardie
ACL (1)4
2022 Variational Graph Autoencoding as Cheap Supervision for AMR Coreference Resolution
abstract
Coreference resolution over semantic graphs like AMRs aims to group the graph nodes that represent the same entity.This is a crucial step for making document-level formal semantic representations.With annotated data on AMR coreference resolution, deep learning approaches have recently shown great potential for this task, yet they are usually data hungry and annotating data is costly.We propose a general pretraining method using variational graph autoencoder (VGAE) for AMR coreference resolution, which can leverage any general AMR corpus and even automatically parsed AMR data.Experiments on benchmarks show that the pretraining approach achieves performance gains of up to 6% absolute F1 points.Moreover, our model significantly improves on the previous state-of-theart model by up to 11% F1 points.
Irene Li, Linfeng Song, Kun Xu 0005, Dong Yu 0001
ACL (1)4
2022 Towards Abstractive Grounded Summarization of Podcast Transcripts
abstract
Podcasts have shown a recent rise in popularity.Summarization of podcasts is of practical benefit to both content providers and consumers.It helps people quickly decide whether they will listen to a podcast and/or reduces the cognitive load of content providers to write summaries.Nevertheless, podcast summarization faces significant challenges including factual inconsistencies of summaries with respect to the inputs.The problem is exacerbated by speech disfluencies and recognition errors in transcripts of spoken language.In this paper, we explore a novel abstractive summarization method to alleviate these issues.Our approach learns to produce an abstractive summary while grounding summary segments in specific regions of the transcript to allow for full inspection of summary details.We conduct a series of analyses of the proposed approach on a large podcast dataset and show that the approach can achieve promising results.Grounded summaries bring clear benefits in locating the summary and transcript segments that contain inconsistent information, and hence improve summarization quality in terms of automatic and human evaluation.
Kaiqiang Song, Chen Li 0003, Xiaoyang Wang 0001, Dong Yu 0001, Fei Liu 0004
ACL (1)4
2022 Toward Unifying Text Segmentation and Long Document Summarization
abstract
Text segmentation is important for signaling a document's structure.Without segmenting a long document into topically coherent sections, it is difficult for readers to comprehend the text, let alone find important information.The problem is only exacerbated by a lack of segmentation in transcripts of audio/video recordings.In this paper, we explore the role that section segmentation plays in extractive summarization of written and spoken documents.Our approach learns robust sentence representations by performing summarization and segmentation simultaneously, which is further enhanced by an optimization-based regularizer to promote selection of diverse summary sentences.We conduct experiments on multiple datasets ranging from scientific articles to spoken transcripts to evaluate the model's performance.Our findings suggest that the model can not only achieve state-of-the-art performance on publicly available benchmarks, but demonstrate better crossgenre transferability when equipped with text segmentation.We perform a series of analyses to quantify the impact of section segmentation on summarizing written and spoken documents of substantial length and complexity.
Sangwoo Cho, Kaiqiang Song, Xiaoyang Wang 0001, Fei Liu 0004, Dong Yu 0001
EMNLP5
2022 MetaLogic: Logical Reasoning Explanations with Fine-Grained Structure
abstract
In this paper, we propose a comprehensive benchmark to investigate models' logical reasoning capabilities in complex real-life scenarios.Current explanation datasets often employ synthetic data with simple reasoning structures.Therefore, it cannot express more complex reasoning processes, such as the rebuttal to a reasoning step and the degree of certainty of the evidence.To this end, we propose a comprehensive logical reasoning explanation form.Based on the multi-hop chain of reasoning, the explanation form includes three main components:(1) The condition of rebuttal that the reasoning node can be challenged; (2) Logical formulae that uncover the internal texture of reasoning nodes; (3) Reasoning strength indicated by degrees of certainty.The fine-grained structure conforms to the real logical reasoning scenario, better fitting the human cognitive process but, simultaneously, is more challenging for the current models.We evaluate the current best models' performance on this new explanation form.The experimental results show that generating reasoning graphs remains a challenging task for current models, even with the help of giant pre-trained language models.
Yinya Huang, Hongming Zhang 0009, Ruixin Hong, Xiaodan Liang, Changshui Zhang, Dong Yu 0001
EMNLP6
2022 Salience Allocation as Guidance for Abstractive Summarization
abstract
Fei Wang, Kaiqiang Song, Hongming Zhang, Lifeng Jin, Sangwoo Cho, Wenlin Yao, Xiaoyang Wang, Muhao Chen, Dong Yu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Fei Wang 0060, Kaiqiang Song, Hongming Zhang 0009, Lifeng Jin, Sangwoo Cho, Wenlin Yao, Xiaoyang Wang 0001, Muhao Chen 0001, Dong Yu 0001
EMNLP9
2022 Z-LaVI: Zero-Shot Language Solver Fueled by Visual Imagination
abstract
Large-scale pretrained language models have made significant advances in solving downstream language understanding tasks.However, they generally suffer from reporting bias, the phenomenon describing the lack of explicit commonsense knowledge in written text, e.g., "an orange is orange".To overcome this limitation, we develop a novel approach, Z-LaVI, to endow language models with visual imagination capabilities.Specifically, we leverage two complementary types of "imaginations": (i) recalling existing images through retrieval and (ii) synthesizing nonexistent images via text-toimage generation.Jointly exploiting the language inputs and the imagination, a pretrained vision-language model (e.g., CLIP) eventually composes a zero-shot solution to the original language tasks.Notably, fueling language models with imagination can effectively leverage visual knowledge to solve plain language tasks.In consequence, Z-LaVI consistently improves the zero-shot performance of existing language models across a diverse set of language tasks. 1 * Work was done during the internship at Tencent AI Lab.(a) Word Sense Disambiguation (b) Science Question Answering (c) Topic Classification sense1: bank(institute) sense2: bank(geography) Input: The species prefers the {bank} of pond.
Wenlin Yao, Hongming Zhang 0009, Xiaoyang Wang 0001, Dong Yu 0001, Jianshu Chen
EMNLP5
2022 Learning a Grammar Inducer from Massive Uncurated Instructional Videos
abstract
Video-aided grammar induction aims to leverage video information for finding more accurate syntactic grammars for accompanying text.While previous work focuses on building systems for inducing grammars on text that are well-aligned with video content, we investigate the scenario, in which text and video are only in loose correspondence.Such data can be found in abundance online, and the weak correspondence is similar to the indeterminacy problem studied in language acquisition.Furthermore, we build a new model that can better learn video-span correlation without manually designed features adopted by previous work.Experiments show that our model trained only on large-scale YouTube data with no textvideo alignment reports strong and robust performances across three unseen datasets, despite domain shift and noisy label issues.Furthermore our model yields higher F1 scores than the previous state-of-the-art systems trained on in-domain data.
Songyang Zhang 0004, Linfeng Song, Lifeng Jin, Haitao Mi, Kun Xu 0005, Dong Yu 0001, Jiebo Luo 0001
EMNLP6
2022 FlowEval: A Consensus-Based Dialogue Evaluation Framework Using Segment Act Flows
abstract
Despite recent progress in open-domain dialogue evaluation, how to develop automatic metrics remains an open problem.We explore the potential of dialogue evaluation featuring dialog act information, which was hardly explicitly modeled in previous methods.However, defined at the utterance level in general, dialog act is of coarse granularity, as an utterance can contain multiple segments possessing different functions.Hence, we propose segment act, an extension of dialog act from utterance level to segment level, and crowdsource a largescale dataset for it.To utilize segment act flows, sequences of segment acts, for evaluation, we develop the first consensus-based dialogue evaluation framework, FlowEval.This framework provides a reference-free approach for dialog evaluation by finding pseudo-references.Extensive experiments against strong baselines on three benchmark datasets demonstrate the effectiveness and other desirable characteristics of our FlowEval, pointing out a potential path for better dialogue evaluation.
Jianqiao Zhao, Yanyang Li, Wanyu Du, Yangfeng Ji, Dong Yu 0001, Michael R. Lyu, Liwei Wang 0009
EMNLP5
2022 Robust Disentangled Variational Speech Representation Learning for Zero-Shot Voice Conversion
abstract
Traditional studies on voice conversion (VC) have made progress with parallel training data and known speakers. Good voice conversion quality is obtained by exploring better alignment modules or expressive mapping functions. In this study, we investigate zero-shot VC from a novel perspective of self-supervised disentangled speech representation learning. Specifically, we achieve the disentanglement by balancing the information flow between global speaker representation and time-varying content representation in a sequential variational autoencoder (VAE). A zero-shot voice conversion is performed by feeding an arbitrary speaker embedding and content embeddings to the VAE decoder. Besides that, an on-the-fly data augmentation training strategy is applied to make the learned representation noise invariant. On TIMIT and VCTK datasets, we achieve state-of-the-art performance on both objective evaluation, i.e., speaker verification (SV) on speaker embedding and content embedding, and subjective evaluation, i.e., voice naturalness and similarity, and remains to be robust even with noisy source/target utterances.
Jiachen Lian, Dong Yu 0001
ICASSP3
2022 Referee: Towards Reference-Free Cross-Speaker Style Transfer with Low-Quality Data for Expressive Speech Synthesis
abstract
Cross-speaker style transfer (CSST) in text-to-speech (TTS) synthesis aims at transferring a speaking style to the synthesised speech in a target speaker’s voice. Most previous CSST approaches rely on expensive high-quality data carrying desired speaking style during training and require a reference utterance to obtain speaking style descriptors as conditioning on the generation of a new sentence. This work presents Referee, a robust reference-free CSST approach for expressive TTS, which fully leverages low-quality data to learn speaking styles from text. Referee is built by cascading a text-to-style (T2S) model with a style-to-wave (S2W) model. Phonetic PosteriorGram (PPG), phoneme-level pitch and energy contours are adopted as fine-grained speaking style descriptors, which are predicted from text using the T2S model. A novel pretrain-refinement method is adopted to learn a robust T2S model by only using readily accessible low-quality data. The S2W model is trained with high-quality target data, which is adopted to effectively aggregate style descriptors and generate high-fidelity speech in the target speaker’s voice. Experimental results are presented, showing that Referee outperforms a global-style-token (GST)-based baseline approach in CSST.
Songxiang Liu, Shan Yang 0001, Dan Su 0002, Dong Yu 0001
ICASSP4
2022 DP-DWA: Dual-Path Dynamic Weight Attention Network With Streaming Dfsmn-San For Automatic Speech Recognition
abstract
In multi-channel far-field automatic speech recognition (ASR) scenarios, distortion is introduced when the speech signal is processed by the front end, which damages the recognition performance for the ASR tasks. In this paper, we propose a dual-path network for the far-field acoustic model, which uses voice processing (VP) signal and acoustic echo cancellation (AEC) signal as input. Specifically, we design a dynamic weight attention (DWA) module for combining two signals. Besides, we streamline our best deep feed-forward sequential memory network with self-attention (DFSMN-SAN) acoustic model for real-time requirements. Joint-training strategy is adopted to optimize the proposed approach. We find that with dual-path network, we can achieve a 54.5% relative improvement in character error rate (CER) on a 10,000-hour online conference task. In addition, our proposed method is not affected by the arrangement of different microphone arrays. We achieve a 23.56% relative improvement on a vehicle task, which has an array with two microphones.
Dongpeng Ma, Liqiang He, Mingjie Jin, Dan Su 0002, Dong Yu 0001
ICASSP6
2022 Fast-Rir: Fast Neural Diffuse Room Impulse Response Generator
abstract
We present a neural-network-based fast diffuse room impulse response generator (FAST-RIR) for generating room impulse responses (RIRs) for a given acoustic environment. Our FAST-RIR takes rectangular room dimensions, listener and speaker positions, and reverberation time (T60) as inputs and generates specular and diffuse reflections for a given acoustic environment. Our FAST-RIR is capable of generating RIRs for a given input T60with an average error of 0.02s. We evaluate our generated RIRs in automatic speech recognition (ASR) applications using Google Speech API, Microsoft Speech API, and Kaldi tools. We show that our proposed FAST-RIR with batch size 1 is 400 times faster than a state-of-the-art diffuse acoustic simulator (DAS) on a CPU and gives similar performance to DAS in ASR experiments. Our FAST-RIR is 12 times faster than an existing GPU-based RIR generator (gpuRIR). We show that our FAST-RIR outperforms gpuRIR by 2.5% in an AMI far-field ASR benchmark.
Anton Ratnarajah, Shixiong Zhang 0001, Meng Yu 0003, Zhenyu Tang 0001, Dinesh Manocha, Dong Yu 0001
ICASSP6
2022 Multi-Channel Multi-Speaker ASR Using 3D Spatial Feature
abstract
Automatic speech recognition (ASR) of multi-channel multi-speaker overlapped speech remains one of the most challenging tasks to the speech community. In this paper, we look into this challenge by utilizing the location information of target speakers in the 3D space for the first time. To explore the strength of proposed the 3D spatial feature, two paradigms are investigated. 1) a pipelined system with a multi-channel speech separation module followed by the state-of-the-art single-channel ASR module; 2) a "All-In-One" model where the 3D spatial feature is directly used as an input to ASR system without explicit separation modules. Both of them are fully differentiable and can be back-propagated end-to-end. We test them on simulated overlapped speech and real recordings. Experimental results show that 1) the proposed ALL-In-One model achieved a comparable error rate to the pipelined system while reducing the inference time by half; 2) the proposed 3D spatial feature significantly outperformed (31% CERR) all previous works of using the 1D directional information in both paradigms.
Yiwen Shao, Shixiong Zhang 0001, Dong Yu 0001
ICASSP3
2022 Consistent Training and Decoding for End-to-End Speech Recognition Using Lattice-Free MMI
abstract
Recently, End-to-End (E2E) frameworks have achieved remarkable results on various Automatic Speech Recognition (ASR) tasks. However, Lattice-Free Maximum Mutual Information (LF-MMI), as one of the discriminative training criteria that show superior performance in hybrid ASR systems, is rarely adopted in E2E ASR frameworks. In this work, we propose a novel approach to introduce LF-MMI criterion into E2E ASR frameworks in both training and decoding stages. The proposed approach shows its effectiveness on two of the most widely used E2E frameworks including Attention-Based Encoder-Decoders (AEDs) and Neural Transducers (NTs). Experiments suggest that the introduction of the LF-MMI criterion consistently leads to significant performance improvements on various datasets and different E2E ASR frameworks. The best of our models achieves competitive CER of 4.1% / 4.4% on Aishell-1 dev/test set; significant error reduction is also achieved on Aishell-2 and Librispeech datasets over strong baselines. Code is released1.
Jinchuan Tian, Jianwei Yu 0001, Chao Weng, Shixiong Zhang 0001, Dan Su 0002, Dong Yu 0001, Yuexian Zou
ICASSP6
2022 VCVTS: Multi-Speaker Video-to-Speech Synthesis Via Cross-Modal Knowledge Transfer from Voice Conversion
abstract
Though significant progress has been made for speaker-dependent Video-to-Speech (VTS) synthesis, little attention is devoted to multi-speaker VTS that can map silent video to speech, while allowing flexible control of speaker identity, all in a single system. This paper proposes a novel multi-speaker VTS system based on cross-modal knowledge transfer from voice conversion (VC), where vector quantization with contrastive predictive coding (VQCPC) is used for the content encoder of VC to derive discrete phoneme-like acoustic units, which are transferred to a Lip-to-Index (Lip2Ind) network to infer the index sequence of acoustic units. The Lip2Ind network can then substitute the content encoder of VC to form a multi-speaker VTS system to convert silent video to acoustic units for reconstructing accurate spoken content. The VTS system also inherits the advantages of VC by using a speaker encoder to produce speaker representations to effectively control the speaker identity of generated speech. Extensive evaluations verify the effectiveness of proposed approach, which can be applied in both constrained vocabulary and open vocabulary conditions, achieving state-of-the-art performance in generating high-quality speech with high naturalness, intelligibility and speaker similarity. Our demo page is released here1.
Disong Wang, Shan Yang 0001, Dan Su 0002, Xunying Liu, Dong Yu 0001, Helen M. Meng
ICASSP5
2022 Joint Modeling of Code-Switched and Monolingual ASR via Conditional Factorization
abstract
Conversational bilingual speech encompasses three types of utterances: two purely monolingual types and one intra-sententially code-switched type. In this work, we propose a general framework to jointly model the likelihoods of the monolingual and code-switch sub-tasks that comprise bilingual speech recognition. By defining the monolingual sub-tasks with label-to-frame synchronization, our joint modeling framework can be conditionally factorized such that the final bilingual output, which may or may not be code-switched, is obtained given only monolingual information. We show that this conditionally factorized joint framework can be modeled by an end-to-end differentiable neural network. We demonstrate the efficacy of our proposed model on bilingual Mandarin-English speech recognition across both monolingual and code-switched corpora.
Brian Yan, Meng Yu 0003, Shixiong Zhang 0001, Siddharth Dalmia, Dan Berrebbi, Chao Weng, Shinji Watanabe 0001, Dong Yu 0001
ICASSP9
2022 Speechmoe2: Mixture-of-Experts Model with Improved Routing
abstract
Mixture-of-experts based acoustic models with dynamic routing mechanisms have proved promising results for speech recognition. The design principle of router architecture is important for the large model capacity and high computational efficiency. Our previous work SpeechMoE only uses local grapheme embedding to help routers to make route decisions. To further improve speech recognition performance against varying domains and accents, we propose a new router architecture which integrates additional global domain and accent embedding into router input to promote adaptability. Experimental results show that the proposed SpeechMoE2 can achieve lower character error rate (CER) with comparable parameters than SpeechMoE on both multi-domain and multi-accent task. Primarily, the proposed method provides up to 1.6% ∼ 4.8% relative CER improvement for the multi-domain task and 1.9% ∼ 17.7% relative CER improvement for the multi-accent task respectively. Besides, increasing the number of experts also achieves consistent performance improvement and keeps the computational cost constant.
Zhao You, Shulin Feng, Dan Su 0002, Dong Yu 0001
ICASSP4
2022 Towards end-to-end Speaker Diarization with Generalized Neural Speaker Clustering
abstract
Speaker diarization consists of many components, e.g., front-end processing, speech activity detection (SAD), overlapped speech detection (OSD) and speaker segmentation/clustering. Conventionally, most of the involved components are separately developed and optimized. The resulting speaker diarization systems are complicated and sometimes lack of satisfying generalization capabilities. In this study, we present a novel speaker diarization system, with a generalized neural speaker clustering module as the backbone. The whole system can be simplified to contain only two major parts, a speaker embedding extractor followed by a clustering module. Both parts are implemented with neural networks. In the training phase, an on-the-fly spoken dialogue generator is designed to provide the system with audio streams and the corresponding annotations in categories of non-speech, overlapped speech and active speakers. The chunk-wise inference and a speaker verification based tracing module are conducted to handle the arbitrary number of speakers. We demonstrate that the proposed speaker diarization system is able to integrate SAD, OSD and speaker segmentation/clustering, and yield competitive results in the VoxConverse20 benchmarks.
Jiatong Shi, Chao Weng, Meng Yu 0003, Dong Yu 0001
ICASSP5
2022 BDDM: Bilateral Denoising Diffusion Models for Fast and High-Quality Speech Synthesis
Max W. Y. Lam, Jun Wang 0091, Dan Su 0002, Dong Yu 0001
ICLR4
2022 FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis
abstract
Denoising diffusion probabilistic models (DDPMs) have recently achieved leading performances in many generative tasks. However, the inherited iterative sampling process costs hindered their applications to speech synthesis. This paper proposes FastDiff, a fast conditional diffusion model for high-quality speech synthesis. FastDiff employs a stack of time-aware location-variable convolutions of diverse receptive field patterns to efficiently model long-term time dependencies with adaptive conditions. A noise schedule predictor is also adopted to reduce the sampling steps without sacrificing the generation quality. Based on FastDiff, we design an end-to-end text-to-speech synthesizer, FastDiff-TTS, which generates high-fidelity speech waveforms without any intermediate feature (e.g., Mel-spectrogram). Our evaluation of FastDiff demonstrates the state-of-the-art results with higher-quality (MOS 4.28) speech samples. Also, FastDiff enables a sampling speed of 58x faster than real-time on a V100 GPU, making diffusion models practically applicable to speech synthesis deployment for the first time. We further show that FastDiff generalized well to the mel-spectrogram inversion of unseen speakers, and FastDiff-TTS outperformed other competing methods in end-to-end text-to-speech synthesis. Audio samples are available at https://FastDiff.github.io/.
Rongjie Huang 0001, Max W. Y. Lam, Jun Wang 0091, Dan Su 0002, Dong Yu 0001, Yi Ren 0006, Zhou Zhao 0001
IJCAI5
2022 Automatic Prosody Annotation with Pre-Trained Text-Speech Model
abstract
Prosodic boundary plays an important role in text-to-speech synthesis (TTS) in terms of naturalness and readability.However, the acquisition of prosodic boundary labels relies on manual annotation, which is costly and time-consuming.In this paper, we propose to automatically extract prosodic boundary labels from text-audio data via a neural text-speech model with pre-trained audio encoders.This model is pre-trained on text and speech data separately and jointly fine-tuned on TTS data in a triplet format: {speech, text, prosody}.The experimental results on both automatic evaluation and human evaluation demonstrate that: 1) the proposed text-speech prosody annotation framework significantly outperforms text-only baselines; 2) the quality of automatic prosodic boundary annotations is comparable to human annotations; 3) TTS systems trained with model-annotated boundaries are slightly better than systems that use manual ones.Code is released 1 .
Ziqian Dai, Jianwei Yu 0001, Yan Wang 0060, Nuo Chen 0001, Yanyao Bian, Guangzhi Li, Deng Cai 0002, Dong Yu 0001
INTERSPEECH8
2022 Joint Neural AEC and Beamforming with Double-Talk Detection
abstract
Acoustic echo cancellation (AEC) in full-duplex communication systems eliminates acoustic feedback.However, nonlinear distortions induced by audio devices, background noise, reverberation, and double-talk reduce the efficiency of conventional AEC systems.Several hybrid AEC models were proposed to address this, which use deep learning models to suppress residual echo from standard adaptive filtering.This paper proposes deep learning-based joint AEC and beamforming model (JAECBF) building on our previous self-attentive recurrent neural network (RNN) beamformer.The proposed network consists of two modules: (i) multi-channel neural-AEC, and (ii) joint AEC-RNN beamformer with a double-talk detection (DTD) that computes time-frequency (T-F) beamforming weights.We train the proposed model in an end-to-end approach to eliminate background noise and echoes from far-end audio devices, which include nonlinear distortions.From experimental evaluations, we find the proposed network outperforms other multi-channel AEC and denoising systems in terms of speech recognition rate and overall speech quality.
Vinay Kothapally, Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Dong Yu 0001
INTERSPEECH5
2022 Towards Improved Zero-shot Voice Conversion with Conditional DSVAE
abstract
Disentangling content and speaking style information is essential for zero-shot non-parallel voice conversion (VC). Our previous study investigated a novel framework with disentangled sequential variational autoencoder (DSVAE) as the backbone for information decomposition. We have demonstrated that simultaneous disentangling content embedding and speaker embedding from one utterance is feasible for zero-shot VC. In this study, we continue the direction by raising one concern about the prior distribution of content branch in the DSVAE baseline. We find the random initialized prior distribution will force the content embedding to reduce the phonetic-structure information during the learning process, which is not a desired property. Here, we seek to achieve a better content embedding with more phonetic information preserved. We propose conditional DSVAE, a new model that enables content bias as a condition to the prior modeling and reshapes the content embedding sampled from the posterior distribution. In our experiment on the VCTK dataset, we demonstrate that content embeddings derived from the conditional DSVAE overcome the randomness and achieve a much better phoneme classification accuracy, a stabilized vocalization and a better zero-shot VC performance compared with the competitive DSVAE baseline.
Jiachen Lian, Gopala Krishna Anumanchipalli, Dong Yu 0001
INTERSPEECH4
2022 LAE: Language-Aware Encoder for Monolingual and Multilingual ASR
abstract
Despite the rapid progress in automatic speech recognition (ASR) research, recognizing multilingual speech using a unified ASR system remains highly challenging. Previous works on multilingual speech recognition mainly focus on two directions: recognizing multiple monolingual speech or recognizing code-switched speech that uses different languages interchangeably within a single utterance. However, a pragmatic multilingual recognizer is expected to be compatible with both directions. In this work, a novel language-aware encoder (LAE) architecture is proposed to handle both situations by disentangling language-specific information and generating frame-level language-aware representations during encoding. In the LAE, the primary encoding is implemented by the shared block while the language-specific blocks are used to extract specific representations for each language. To learn language-specific information discriminatively, a language-aware training method is proposed to optimize the language-specific blocks in LAE. Experiments conducted on Mandarin-English code-switched speech suggest that the proposed LAE is capable of discriminating different languages in frame-level and shows superior performance on both monolingual and multilingual ASR tasks. With either a real-recorded or simulated code-switched dataset, the proposed LAE achieves statistically significant improvements on both CTC and neural transducer systems. Code is released
Jinchuan Tian, Jianwei Yu 0001, Yuexian Zou, Dong Yu 0001
INTERSPEECH5
2022 DDAM '22: 1st International Workshop on Deepfake Detection for Audio Multimedia
abstract
Over the last few years, the technology of speech synthesis and voice conversion has made significant improvement with the development of deep learning. The models can generate realistic and human-like speech. It is difficult for most people to distinguish the generated audio from the real. However, this technology also poses a great threat to the global political economy and social stability if some attackers and criminals misuse it with the intent to cause harm. In this workshop, we aim to bring together researchers from the fields of audio deepfake detection, audio deep synthesis, audio fake game and adversarial attacks to further discuss recent research and future directions for detecting deepfake and manipulated audios in multimedia.
Jianhua Tao 0001, Jiangyan Yi, Cunhang Fan, Ruibo Fu, Shan Liang 0007, Pengyuan Zhang, Haizhou Li 0001, Helen M. Meng, Dong Yu 0001, Masato Akagi
ACM Multimedia9
2022 End-to-End Chinese Speaker Identification
abstract
Speaker identification (SI) in texts aims to identify the speaker(s) for each utterance in texts.Previous studies divide SI into several sub-tasks (e.g., quote extraction, named entity recognition, gender identification, and coreference resolution).However, we are still far from solving these sub-tasks, making SI systems that rely on them seriously suffer from error propagation.End-to-end SI systems, on the other hand, are not limited by individual modules, but suffer from insufficient training data from the existing small-scale datasets.To make large end-to-end models possible, we design a new annotation guideline that regards SI as span extraction from the local context, and we annotate by far the largest SI dataset for Chinese named CSI based on eighteen novels.Viewing SI as a span extraction task also introduces the possibility of applying existing storng extractive machine reading comprehension (MRC) baselines.Surprisingly, simply using such a baseline without human-annotated character names and carefully designed rules, we can already achieve performance comparable or better than those of previous state-of-the-art SI methods on all public SI datasets for Chinese.Furthermore, we show that our dataset can serve as additional training data for existing benchmarks, which leads to further gains (up to 6.5% in accuracy).Finally, using CSI as a clean source, we design an effective self-training paradigm to continuously leverage hundreds of unlabeled novels.
Dian Yu 0001, Ben Zhou, Dong Yu 0001
NAACL-HLT3
2022 Efficient Text Analysis with Pre-Trained Neural Network Models
abstract
This paper investigates the application of pre-trained BERT model in three classic text analysis tasks: Chinese grapheme-to-phoneme(G2P), text normalization(TN) and sentence punctuation annotation. Even though the full-sized BERT has prominent modeling power, there are two challenges for it in real applications: the requirement for annotated training data and the considerable computational cost. In this paper, we propose BERT-based low-latency solutions. To collect sufficient training corpus for G2P, we transfer knowledge from existing rule-based system to BERT through a large amount of unlabeled corpus. The new model could convert all characters directly from raw texts with higher accuracy. We also propose a hybrid two-stage text normalization pipeline which reduces the sentence error rate by 25% compared to the rule-based system. We offer both supervised and weakly supervised versions and find that the latter has only 1% accuracy drop from the former.
Jia Cui, Heng Lu 0004, Shiyin Kang, Liqiang He, Guangzhi Li, Dong Yu 0001
SLT7
2022 Meta-learning without data via Wasserstein distributionally-robust model fusion
abstract
Existing meta-learning works assume that each task has available training and testing data. However, there are many available pre-trained models without accessing their training data in practice. We often need a single model to solve different tasks simultaneously as this is much more convenient to deploy the models. Our work aims to meta-learn a model initialization from these pre-trained models without using corresponding training data. We name this challenging problem setting as Data-Free Learning To Learn (DFL2L). We propose a distributionally robust optimization (DRO) framework to learn a black-box model to fuse and compress all the pre-trained models into a single network to address this problem. To encourage good generalization to the unseen new tasks, the proposed DRO framework diversifies the learned task embedding associated with each pre-trained model to cover the diversity in the underlying training task distributions. A model initialization is sampled from the black-box network during meta-testing as the meta learned initialization. Extensive experiments on offline and online DFL2L settings and several real image datasets demonstrate the effectiveness of the proposed methods.
Zhenyi Wang 0001, Xiaoyang Wang 0001, Li Shen 0008, Qiuling Suo, Kaiqiang Song, Dong Yu 0001, Yan Shen 0002, Mingchen Gao
UAI6
2022 An investigation of neural uncertainty estimation for target speaker extraction equipped RNN transducer
Jiatong Shi, Chao Weng, Shinji Watanabe 0001, Meng Yu 0003, Dong Yu 0001
Comput. Speech Lang.6
2022 Deep learning based multi-source localization with source splitting and its effectiveness in multi-talker speech recognition
abstract
Multi-source localization is an important and challenging technique for multi-talker conversation analysis. This paper proposes a novel supervised learning method using deep neural networks to estimate the direction of arrival (DOA) of all the speakers simultaneously from the audio mixture. At the heart of the proposal is a source splitting mechanism that creates source-specific intermediate representations inside the network. This allows our model to give source-specific posteriors as the output unlike the traditional multi-label classification approach . Existing deep learning methods perform a frame level prediction, whereas our approach performs an utterance level prediction by incorporating temporal selection and averaging inside the network to avoid post-processing. We also experiment with various loss functions and show that a variant of earth mover distance (EMD) is very effective in classifying DOA at a very high resolution by modeling inter-class relationships. In addition to using the prediction error as a metric for evaluating our localization model, we also establish its potency as a frontend with automatic speech recognition (ASR) as the downstream task. We convert the estimated DOAs into a feature suitable for ASR and pass it as an additional input feature to a strong multi-channel and multi-talker speech recognition baseline. This added input feature drastically improves the ASR performance and gives a word error rate (WER) of 6.3% on the evaluation data of our simulated noisy two speaker mixtures, while the baseline which does not use explicit localization input has a WER of 11.5%. We also perform ASR evaluation on real recordings with the overlapped set of the MC-WSJ-AV corpus in addition to simulated mixtures.
Aswin Shanmugam Subramanian, Chao Weng, Shinji Watanabe 0001, Meng Yu 0003, Dong Yu 0001
Comput. Speech Lang.5
2022 Improving Mandarin End-to-End Speech Recognition With Word N-Gram Language Model
abstract
Despite the rapid progress of end-to-end (E2E) automatic speech recognition (ASR), it has been shown that incorporating external language models (LMs) into the decoding can further improve the recognition performance of E2E ASR systems. To align with the modelingunits adopted in E2E ASR systems, subword-level (e.g., characters, BPE) LMs are usually used to cooperate with current E2E ASR systems. However, the use of subword-level LMs will ignore the word-level information, which may limit the strength of the external LMs in E2E ASR. Although several methods have been proposed to incorporate word-level external LMs in E2E ASR, these methods are mainly designed for languages with clear word boundaries such as English and cannot be directly applied to languages like Mandarin, in which each character sequence can have multiple corresponding word sequences. To this end, we propose a novel decoding algorithm where a word-level lattice is constructed on-the-fly to consider all possible word sequences for each partial hypothesis. Then, the LM score of the hypothesis is obtained by intersecting the generated lattice with an external word N-gram LM. The proposed method is examined on both Attention-based Encoder-Decoder (AED) and Neural Transducer (NT) frameworks. Experiments suggest that our method consistently outperforms subword-level LMs, including N-gram LM and neural network LM. We achieve state-of-the-art results on both Aishell-1 (CER 4.18%) and Aishell-2 (CER 5.06%) datasets and reduce CER by 14.8% relatively on a 21K-hour Mandarin dataset. Code is released.
Jinchuan Tian, Jianwei Yu 0001, Chao Weng, Yuexian Zou, Dong Yu 0001
IEEE Signal Process. Lett.5
2022 High-Fidelity 3D Digital Human Head Creation from RGB-D Selfies
abstract
We present a fully automatic system that can produce high-fidelity, photo-realistic three-dimensional (3D) digital human heads with a consumer RGB-D selfie camera. The system only needs the user to take a short selfie RGB-D video while rotating his/her head and can produce a high-quality head reconstruction in less than 30 s. Our main contribution is a new facial geometry modeling and reflectance synthesis procedure that significantly improves the state of the art. Specifically, given the input video a two-stage frame selection procedure is first employed to select a few high-quality frames for reconstruction. Then a differentiable renderer-based 3D Morphable Model (3DMM) fitting algorithm is applied to recover facial geometries from multiview RGB-D data, which takes advantages of a powerful 3DMM basis constructed with extensive data generation and perturbation. Our 3DMM has much larger expressive capacities than conventional 3DMM, allowing us to recover more accurate facial geometry using merely linear basis. For reflectance synthesis, we present a hybrid approach that combines parametric fitting andConvolutional Neural Networks (CNNs)to synthesize high-resolution albedo/normal maps with realistic hair/pore/wrinkle details. Results show that our system can produce faithful 3D digital human faces with extremely realistic details. The main code and the newly constructed 3DMM basis is publicly available.
Linchao Bao, Xiangkai Lin, Haoxian Zhang, Xuefei Zhe, Hao-Zhi Huang 0001, Xinwei Jiang, Jue Wang 0001, Dong Yu 0001, Zhengyou Zhang
ACM Trans. Graph.11
2021 Tune-In: Training Under Negative Environments with Interference for Attention Networks Simulating Cocktail Party Effect
abstract
We study the cocktail party problem and propose a novel attention network called Tune-In, abbreviated for training under negative environments with interference. It firstly learns two separate spaces of speaker-knowledge and speech-stimuli based on a shared feature space, where a new block structure is designed as the building block for all spaces, and then cooperatively solves different tasks. Between the two spaces, information is cast towards each other via a novel cross- and dual-attention mechanism, mimicking the bottom-up and top-down processes of a human's cocktail party effect. It turns out that substantially discriminative and generalizable speaker representations can be learnt in severely interfered conditions via our self-supervised training. The experimental results verify this seeming paradox. The learnt speaker embedding has superior discriminative power than a standard speaker verification method; meanwhile, Tune-In achieves remarkably better speech separation performances in terms of SI-SNRi and SDRi consistently in all test modes, and especially at lower memory and computational consumption, than state-of-the-art benchmark systems.
Jun Wang 0091, Max W. Y. Lam, Dan Su 0002, Dong Yu 0001
AAAI4
2021 NaturalConv: A Chinese Dialogue Dataset Towards Multi-turn Topic-driven Conversation
abstract
In this paper, we propose a Chinese multi-turn topic-driven conversation dataset, NaturalConv, which allows the participants to chat anything they want as long as any element from the topic is mentioned and the topic shift is smooth. Our corpus contains 19.9K conversations from six domains, and 400K utterances with an average turn number of 20.1. These conversations contain in-depth discussions on related topics or widely natural transition between multiple topics. We believe either way is normal for human conversation. To facilitate the research on this corpus, we provide results of several benchmark models. Comparative results show that for this dataset, our current models are not able to provide significant improvement by introducing background knowledge/topic. Therefore, the proposed dataset should be a good benchmark for further research to evaluate the validity and naturalness of multi-turn conversation systems. Our dataset is available at https://ai.tencent.com/ailab/nlp/dialogue/#datasets.
Xiaoyang Wang 0001, Chen Li 0003, Jianqiao Zhao, Dong Yu 0001
AAAI4
2021 3D Spatial Features for Multi-Channel Target Speech Separation
abstract
The use of speaker's directional information for speech sepa-ration and speech recognition has demonstrated the state-of-the-art performances on multi-talker scenarios. One major limitation of previous approaches using speaker's directional information is the significant performance degradation when the coming directions of two sound sources are close. To address these challenges, this paper proposed a set of new three-dimensional (3D) spatial features for target speech sep-aration, by leveraging all the 3D location information of the target speaker, including azimuth, elevation, and the distance to the microphone array center. Previous works in this area are extended in two important directions. First, the traditional 1D directional features are generalized to 3D spatial features. Thus more discriminative spatial diversity between speakers is achieved. Second, to unleash the full power of these 3D spatial features, a microphone pair-wise attention model is also proposed. The proposed features and models were evaluated on both simulated reverberant datasets and real recordings under near and far-field conditions. Exper-imental results show that both proposed 3D spatial features and attention models can significantly improve the separation performance as well as reducing the recognition error rate.
Rongzhi Gu, Shixiong Zhang 0001, Meng Yu 0003, Dong Yu 0001
ASRU4
2021 Latency-Controlled Neural Architecture Search for Streaming Speech Recognition
abstract
Neural architecture search (NAS) has attracted much attention and has been explored for automatic speech recognition (ASR). In this work, we focus on streaming ASR scenarios and propose the latency-controlled NAS for acoustic modeling. First, based on the vanilla neural architecture, normal cells are altered to causal cells to control the total latency of the architecture. Second, a revised operation space with a smaller receptive field is proposed to generate the final architecture with low latency. Extensive experiments show that: 1) Based on the proposed neural architecture, the neural networks with a medium latency of 550ms (millisecond) and a low latency of 190ms can be learned in the vanilla and revised operation space respectively. 2) For the low latency setting, the evaluation network can achieve more than 19% (average on the four test sets) relative improvements compared with the hybrid CLDNN baseline, on a 10k-hour large-scale dataset.
Liqiang He, Shulin Feng, Dan Su 0002, Dong Yu 0001
ASRU4
2021 Improving Weakly Supervised Visual Grounding by Contrastive Knowledge Distillation
abstract
Weakly supervised phrase grounding aims at learning region-phrase correspondences using only image-sentence pairs. A major challenge thus lies in the missing links between image regions and sentence phrases during training. To address this challenge, we leverage a generic object detector at training time, and propose a contrastive learning framework that accounts for both region-phrase and image-sentence matching. Our core innovation is the learning of a region-phrase score function, based on which an image-sentence score function is further constructed. Importantly, our region-phrase score function is learned by distilling from soft matching scores between the detected object names and candidate phrases within an image-sentence pair, while the image-sentence score function is supervised by ground-truth image-sentence pairs. The design of such score functions removes the need of object detection at test time, thereby significantly reducing the inference cost. Without bells and whistles, our approach achieves state-of-the-art results on visual phrase grounding, surpassing previous methods that require expensive object detectors at test time.
Liwei Wang 0009, Jing Huang 0014, Yin Li 0003, Kun Xu 0005, Zhengyuan Yang, Dong Yu 0001
CVPR6
2021 RAST: Domain-Robust Dialogue Rewriting as Sequence Tagging
abstract
The task of dialogue rewriting aims to reconstruct the latest dialogue utterance by copying the missing content from the dialogue context.Until now, the existing models for this task suffer from the robustness issue, i.e., performances drop dramatically when testing on a different dataset.We address this robustness issue by proposing a novel sequence-taggingbased model so that the search space is significantly reduced, yet the core of this task is still well covered.As a common issue of most tagging models for text generation, the model's outputs may lack fluency.To alleviate this issue, we inject the loss signal from BLEU or GPT-2 under a REINFORCE framework.Experiments show huge improvements of our model over the current state-of-the-art systems when transferring to another dataset.
Linfeng Song, Liwei Wang 0009, Kun Xu 0005, Zhaopeng Tu, Dong Yu 0001
EMNLP (1)6
2021 Instance-adaptive training with noise-robust losses against noisy labels
abstract
In order to alleviate the huge demand for annotated datasets for different tasks, many recent natural language processing datasets have adopted automated pipelines for fast-tracking usable data.However, model training with such datasets poses a challenge because popular optimization objectives are not robust to label noise induced in the annotation generation process.Several noise-robust losses have been proposed and evaluated on tasks in computer vision, but they generally use a single dataset-wise hyperparamter to control the strength of noise resistance.This work proposes novel instance-adaptive training frameworks to change dataset-wise hyperparameters of noise resistance in such losses to be instance-specific.Such instance-specific noise resistance hyperparameters are predicted by special instance-level label quality predictors, which are trained along with the main models.Experiments on noisy and corrupted NLP datasets show that proposed instance-adaptive training frameworks help increase the noiserobustness provided by such losses, promoting the use of the frameworks and associated losses in training NLP models with noisy data.
Lifeng Jin, Linfeng Song, Kun Xu 0005, Dong Yu 0001
EMNLP (1)4
2021 Connect-the-Dots: Bridging Semantics between Words and Definitions via Aligning Word Sense Inventories
abstract
Word Sense Disambiguation (WSD) aims to automatically identify the exact meaning of one word according to its context.Existing supervised models struggle to make correct predictions on rare word senses due to limited training data and can only select the best definition sentence from one predefined word sense inventory (e.g., WordNet).To address the data sparsity problem and generalize the model to be independent of one predefined inventory, we propose a gloss alignment algorithm that can align definition sentences (glosses) with the same meaning from different sense inventories to collect rich lexical knowledge.We then train a model to identify semantic equivalence between a target word in context and one of its glosses using these aligned inventories, which exhibits strong transfer capability to many WSD tasks 1 .Experiments on benchmark datasets show that the proposed method improves predictions on both frequent and rare word senses, outperforming prior work by 1.2% on the All-Words WSD Task and 4.3% on the Low-Shot WSD Task.Evaluation on WiC Task also indicates that our method can better capture word meanings in context.
Wenlin Yao, Xiaoman Pan, Lifeng Jin, Jianshu Chen, Dian Yu 0001, Dong Yu 0001
EMNLP (1)6
2021 Exophoric Pronoun Resolution in Dialogues with Topic Regularization
abstract
Resolving pronouns to their referents has long been studied as a fundamental natural language understanding problem.Previous works on pronoun coreference resolution (PCR) mostly focus on resolving pronouns to mentions in text while ignoring the exophoric scenario.Exophoric pronouns are common in daily communications, where speakers may directly use pronouns to refer to some objects present in the environment without introducing the objects first.Although such objects are not mentioned in the dialogue text, they can often be disambiguated by the general topics of the dialogue.Motivated by this, we propose to jointly leverage the local context and global topics of dialogues to solve the out-of-text PCR problem.Extensive experiments demonstrate the effectiveness of adding topic regularization for resolving exophoric pronouns.
Xintong Yu 0002, Hongming Zhang 0009, Yangqiu Song, Changshui Zhang, Kun Xu 0005, Dong Yu 0001
EMNLP (1)6
2021 Learned Transferable Architectures Can Surpass Hand-Designed Architectures for Large Scale Speech Recognition
abstract
In this paper, we explore the neural architecture search (NAS) for automatic speech recognition (ASR) systems. We conduct the architecture search on the small proxy dataset, and then evaluate the network, constructed from the searched architecture, on the large dataset. Specially, we propose a revised search space that theoretically facilitates the search algorithm to explore the architectures with low complexity. Extensive experiments show that: (i) the architecture learned in the revised search space can greatly reduce the computational overhead and GPU memory usage with mild performance degradation. (ii) the searched architecture can achieve more than 15% (average on the four test sets) relative improvements on the large dataset, compared with our best hand-designed DFSMN-SAN architecture. To the best of our knowledge, this is the first report of NAS results with a large scale dataset (up to 10K hours), indicating the promising application of NAS to industrial ASR systems.
Liqiang He, Dan Su 0002, Dong Yu 0001
ICASSP3
2021 Sandglasset: A Light Multi-Granularity Self-Attentive Network for Time-Domain Speech Separation
abstract
One of the leading single-channel speech separation (SS) models is based on a TasNet with a dual-path segmentation technique, where the size of each segment remains unchanged throughout all layers. In contrast, our key finding is that multi-granularity features are essential for enhancing contextual modeling and computational efficiency. We introduce a self-attentive network with a novel sandglass-shape, namely Sandglasset, which advances the state-of-the-art (SOTA) SS performance at significantly smaller model size and computational cost. Forward along each block inside Sandglasset, the temporal granularity of the features gradually becomes coarser until reaching half of the network blocks, and then successively turns finer towards the raw signal level. We also unfold that residual connections between features with the same granularity are critical for preserving information after passing through the bottleneck layer. Experiments show our Sandglasset with only 2.3M parameters has achieved the best results on two benchmark SS datasets – WSJ0-2mix and WSJ0-3mix, where the SI-SNRi scores have been improved by absolute 0.6 dB and 2.4 dB, respectively, comparing to the prior SOTA results.
Max W. Y. Lam, Jun Wang 0091, Dan Su 0002, Dong Yu 0001
ICASSP4
2021 Replay and Synthetic Speech Detection with Res2Net Architecture
abstract
Existing approaches for replay and synthetic speech detection still lack generalizability to unseen spoofing attacks. This work proposes to leverage a novel model structure, so-called Res2Net, to improve the anti-spoofing countermeasure’s generalizability. Res2Net mainly modifies the ResNet block to enable multiple feature scales. Specifically, it splits the feature maps within one block into multiple channel groups and designs a residual-like connection across different channel groups. Such connection increases the possible receptive fields, resulting in multiple feature scales. This multiple scaling mechanism significantly improves the countermeasure’s generalizability to unseen spoofing attacks. It also decreases the model size compared to ResNet-based models. Experimental results show that the Res2Net model consistently outperforms ResNet34 and ResNet50 by a large margin in both physical access (PA) and logical access (LA) of the ASVspoof 2019 corpus. Moreover, integration with the squeeze-and-excitation (SE) block can further enhance performance. For feature engineering, we investigate the gen-eralizability of Res2Net combined with different acoustic features, and observe that the constant-Q transform (CQT) achieves the most promising performance in both PA and LA scenarios. Our best single system outperforms other state-of-the-art single systems in both PA and LA of the ASVspoof 2019 corpus.
Xu Li 0015, Na Li 0012, Chao Weng, Xunying Liu, Dan Su 0002, Dong Yu 0001, Helen M. Meng
ICASSP6
2021 Improving RNN Transducer with Target Speaker Extraction and Neural Uncertainty Estimation
abstract
Target-speaker speech recognition aims to recognize target-speaker speech from noisy environments with background noise and interfering speakers. This work presents a joint framework that combines time-domain target-speaker speech extraction and Recurrent Neural Network Transducer (RNN-T). To stabilize the joint-training, we propose a multi-stage training strategy that pre-trains and fine-tunes each module in the system before joint-training. Meanwhile, speaker identity and speech enhancement uncertainty measures are proposed to compensate for residual noise and artifacts from the target speech extraction module. Compared to a recognizer fine-tuned with a target speech extraction model, our experiments show that adding the neural uncertainty module significantly reduces 17% relative Character Error Rate (CER) on multi-speaker signals with background noise. The multi-condition experiments indicate that our method can achieve 9% relative performance gain in the noisy condition while maintaining the performance in the clean condition.
Jiatong Shi, Chao Weng, Shinji Watanabe 0001, Meng Yu 0003, Dong Yu 0001
ICASSP6
2021 Directional ASR: A New Paradigm for E2E Multi-Speaker Speech Recognition with Source Localization
abstract
This paper proposes a new paradigm for handling far-field multi-speaker data in an end-to-end (E2E) neural network manner, called directional automatic speech recognition (D-ASR), which explicitly models source speaker locations. In D-ASR, the azimuth angle of the sources with respect to the microphone array is defined as a latent variable. This angle controls the quality of separation, which in turn determines the ASR performance. All three functionalities of D-ASR: localization, separation, and recognition are connected as a single differentiable neural network and trained solely based on ASR error minimization objectives. The advantages of D-ASR over existing methods are threefold: (1) it provides explicit speaker locations, (2) it improves the explainability factor, and (3) it achieves better ASR performance as the process is more streamlined. In addition, D-ASR does not require explicit direction of arrival (DOA) supervision like existing data-driven localization models, which makes it more appropriate for realistic data. For the case of two source mixtures, D-ASR achieves an average DOA prediction error of less than three degrees. It also outperforms a strong far-field multi-speaker end-to-end system in both separation quality and ASR performance.
Aswin Shanmugam Subramanian, Chao Weng, Shinji Watanabe 0001, Meng Yu 0003, Yong Xu 0004, Shixiong Zhang 0001, Dong Yu 0001
ICASSP7
2021 Contrastive Separative Coding for Self-Supervised Representation Learning
abstract
To extract robust deep representations from long sequential modeling of speech data, we propose a self-supervised learning approach, namely Contrastive Separative Coding (CSC). Our key finding is to learn such representations by separating the target signal from contrastive interfering signals. First, a multi-task separative encoder is built to extract shared separable and discriminative embedding; secondly, we propose a powerful cross-attention mechanism performed over speaker representations across various interfering conditions, allowing the model to focus on and globally aggregate the most critical information to answer the "query" (current bottom-up embedding) while paying less attention to interfering, noisy, or irrelevant parts; lastly, we form a new probabilistic contrastive loss which estimates and maximizes the mutual information between the representations and the global speaker vector. While most prior unsupervised methods have focused on predicting the future, neighboring, or missing samples, we take a different perspective of predicting the interfered samples. Moreover, our contrastive separative loss is free from negative sampling. The experiment demonstrates that our approach can learn useful representations achieving a strong speaker verification performance in adverse conditions.
Jun Wang 0091, Max W. Y. Lam, Dan Su 0002, Dong Yu 0001
ICASSP4
2021 Self-Supervised Text-Independent Speaker Verification Using Prototypical Momentum Contrastive Learning
abstract
In this study, we investigate self-supervised representation learning for speaker verification (SV). First, we examine a simple contrastive learning approach (SimCLR) with a momentum contrastive (MoCo) learning framework, where the MoCo speaker embedding system utilizes a queue to maintain a large set of negative examples. We show that better speaker embeddings can be learned by momentum contrastive learning. Next, alternative augmentation strategies are explored to normalize extrinsic speaker variabilities of two random segments from the same speech utterance. Specifically, augmentation in the waveform largely improves the speaker representations for SV tasks. The proposed MoCo speaker embedding is further improved when a prototypical memory bank is introduced, which encourages the speaker embeddings to be closer to their assigned prototypes with an intermediate clustering step. In addition, we generalize the self-supervised framework to a semi-supervised scenario where only a small portion of the data is labeled. Comprehensive experiments on the Voxceleb dataset demonstrate that our proposed self-supervised approach achieves competitive performance compared with existing techniques, and can approach fully supervised results with partially labeled data.
Chao Weng, Meng Yu 0003, Dong Yu 0001
ICASSP5
2021 ADL-MVDR: All Deep Learning MVDR Beamformer for Target Speech Separation
abstract
Speech separation algorithms are often used to separate the target speech from other interfering sources. However, purely neural network based speech separation systems often cause nonlinear distortion that is harmful for automatic speech recognition (ASR) systems. The conventional mask-based minimum variance distortionless response (MVDR) beamformer can be used to minimize the distortion, but comes with high level of residual noise. Furthermore, the matrix operations (e.g., matrix inversion) involved in the conventional MVDR solution are sometimes numerically unstable when jointly trained with neural networks. In this paper, we propose a novel all deep learning MVDR framework, where the matrix inversion and eigenvalue decomposition are replaced by two recurrent neural networks (RNNs), to resolve both issues at the same time. The proposed method can greatly reduce the residual noise while keeping the target speech undistorted by leveraging on the RNN-predicted frame-wise beamforming weights. The system is evaluated on a Mandarin audio-visual corpus and compared against several state-of-the-art (SOTA) speech separation systems. Experimental results demonstrate the superiority of the proposed method across several objective metrics and ASR accuracy.
Zhuohuang Zhang, Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Lianwu Chen, Dong Yu 0001
ICASSP6
2021 Towards Robust Speaker Verification with Target Speaker Enhancement
abstract
This paper proposes the target speaker enhancement based speaker verification network (TASE-SVNet), an all neural model that couples target speaker enhancement and speaker embedding extraction for robust speaker verification (SV). Specifically, an enrollment speaker conditioned speech enhancement module is employed as the front-end for extracting target speaker from its mixture with interfering speakers and environmental noises. Compared with the conventional target speaker enhancement models, nontarget speaker/interference suppression should draw additional attention for SV. Therefore, an effective nontarget speaker sampling strategy is explored. To improve speaker embedding extraction with a light-weighted model, a teacher-student (T/S) training is proposed to distill speaker discriminative information from large models to small models. Iterative inference is investigated to address the noisy speaker enrollment problem. We evaluate the proposed method on two SV tasks, i.e., one heavily overlapped speech and the other one with comprehensive noise types in vehicle environments. Experiments show significant and consistent improvements in Equal Error Rate (EER) over the state-of-the-art baselines.
Meng Yu 0003, Chao Weng, Dong Yu 0001
ICASSP4
2021 Multi-Channel Speaker Verification for Single and Multi-Talker Speech
abstract
To improve speaker verification in real scenarios with interference speakers, noise, and reverberation, we propose to bring together advancements made in multi-channel speech features.Specifically, we combine spectral, spatial, and directional features, which includes inter-channel phase difference, multichannel sinc convolutions, directional power ratio features, and angle features.To maximally leverage supervised learning, our framework is also equipped with multi-channel speech enhancement and voice activity detection.On all simulated, replayed, and real recordings, we observe large and consistent improvements at various degradation levels.On real recordings of multi-talker speech, we achieve a 36% relative reduction in equal error rate w.r.t.single-channel baseline.We find the improvements from speaker-dependent directional features more consistent in multi-talker conditions than clean.Lastly, we investigate if the learned multi-channel speaker embedding space can be made more discriminative through a contrastive loss-based fine-tuning.With a simple choice of Triplet loss, we observe a further 8.3% relative reduction in EER.
Saurabh Kataria 0001, Shixiong Zhang 0001, Dong Yu 0001
Interspeech3
2021 Raw Waveform Encoder with Multi-Scale Globally Attentive Locally Recurrent Networks for End-to-End Speech Recognition
abstract
End-to-end speech recognition generally uses hand-engineered acoustic features as input and excludes the feature extraction module from its joint optimization.To extract learnable and adaptive features and mitigate information loss, we propose a new encoder that adopts globally attentive locally recurrent (GALR) networks and directly takes raw waveform as input.We observe improved ASR performance and robustness by applying GALR on different window lengths to aggregate fine-grain temporal information into multi-scale acoustic features.Experiments are conducted on a benchmark dataset AISHELL-2 and two large-scale Mandarin speech corpus of 5, 000 hours and 21, 000 hours.With faster speed and comparable model size, our proposed multi-scale GALR waveform encoder achieved consistent character error rate reductions (CERRs) from 7.9% to 28.1% relative over strong baselines, including Conformer and TDNN-Conformer.In particular, our approach demonstrated notable robustness than the traditional handcrafted features and outperformed the baseline MFCC-based TDNN-Conformer model by a 15.2% CERR on a music-mixed real-world speech test set.
Max W. Y. Lam, Jun Wang 0091, Chao Weng, Dan Su 0002, Dong Yu 0001
Interspeech5
2021 MIMO Self-Attentive RNN Beamformer for Multi-Speaker Speech Separation
abstract
Recently, our proposed recurrent neural network (RNN) based all deep learning minimum variance distortionless response (ADL-MVDR) beamformer method yielded superior performance over the conventional MVDR by replacing the matrix inversion and eigenvalue decomposition with two RNNs.In this work, we present a self-attentive RNN beamformer to further improve our previous RNN-based beamformer by leveraging on the powerful modeling capability of self-attention.Temporal-spatial self-attention module is proposed to better learn the beamforming weights from the speech and noise spatial covariance matrices.The temporal self-attention module could help RNN to learn global statistics of covariance matrices.The spatial self-attention module is designed to attend on the cross-channel correlation in the covariance matrices.Furthermore, a multi-channel input with multi-speaker directional features and multi-speaker speech separation outputs (MIMO) model is developed to improve the inference efficiency.The evaluations demonstrate that our proposed MIMO self-attentive RNN beamformer improves both the automatic speech recognition (ASR) accuracy and the perceptual estimation of speech quality (PESQ) against prior arts.
Xiyun Li, Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Jiaming Xu 0001, Bo Xu 0002, Dong Yu 0001
Interspeech7
2021 TeCANet: Temporal-Contextual Attention Network for Environment-Aware Speech Dereverberation
abstract
In this paper, we exploit the effective way to leverage contextual information to improve the speech dereverberation performance in real-world reverberant environments. We propose a temporal-contextual attention approach on the deep neural network (DNN) for environment-aware speech dereverberation, which can adaptively attend to the contextual information. More specifically, a FullBand based Temporal Attention approach (FTA) is proposed, which models the correlations between the fullband information of the context frames. In addition, considering the difference between the attenuation of high frequency bands and low frequency bands (high frequency bands attenuate faster than low frequency bands) in the room impulse response (RIR), we also propose a SubBand based Temporal Attention approach (STA). In order to guide the network to be more aware of the reverberant environments, we jointly optimize the dereverberation network and the reverberation time (RT60) estimator in a multi-task manner. Our experimental results indicate that the proposed method outperforms our previously proposed reverberation-time-aware DNN and the learned attention weights are fully physical consistent. We also report a preliminary yet promising dereverberation and recognition experiment on real test data.
Helin Wang, Bo Wu 0011, Lianwu Chen, Meng Yu 0003, Jianwei Yu 0001, Yong Xu 0004, Shixiong Zhang 0001, Chao Weng, Dan Su 0002, Dong Yu 0001
Interspeech10
2021 Generalized Spatio-Temporal RNN Beamformer for Target Speech Separation
abstract
Although the conventional mask-based minimum variance distortionless response (MVDR) could reduce the non-linear distortion, the residual noise level of the MVDR separated speech is still high. In this paper, we propose a spatio-temporal recurrent neural network based beamformer (RNN-BF) for target speech separation. This new beamforming framework directly learns the beamforming weights from the estimated speech and noise spatial covariance matrices. Leveraging on the temporal modeling capability of RNNs, the RNN-BF could automatically accumulate the statistics of the speech and noise covariance matrices to learn the frame-level beamforming weights in a recursive way. An RNN-based generalized eigenvalue (RNN-GEV) beamformer and a more generalized RNN beamformer (GRNN-BF) are proposed. We further improve the RNN-GEV and the GRNN-BF by using layer normalization to replace the commonly used mask normalization on the covariance matrices. The proposed GRNN-BF obtains better performance against prior arts in terms of speech quality (PESQ), speech-to-noise ratio (SNR) and word error rate (WER).
Yong Xu 0004, Zhuohuang Zhang, Meng Yu 0003, Shixiong Zhang 0001, Dong Yu 0001
Interspeech5
2021 SpeechMoE: Scaling to Large Acoustic Models with Dynamic Routing Mixture of Experts
abstract
Recently, Mixture of Experts (MoE) based Transformer has shown promising results in many domains. This is largely due to the following advantages of this architecture: firstly, MoE based Transformer can increase model capacity without computational cost increasing both at training and inference time. Besides, MoE based Transformer is a dynamic network which can adapt to the varying complexity of input instances in realworld applications. In this work, we explore the MoE based model for speech recognition, named SpeechMoE. To further control the sparsity of router activation and improve the diversity of gate values, we propose a sparsity L1 loss and a mean importance loss respectively. In addition, a new router architecture is used in SpeechMoE which can simultaneously utilize the information from a shared embedding network and the hierarchical representation of different MoE layers. Experimental results show that SpeechMoE can achieve lower character error rate (CER) with comparable computation cost than traditional static networks, providing 7.0%-23.0% relative CER improvements on four evaluation datasets.
Zhao You, Shulin Feng, Dan Su 0002, Dong Yu 0001
Interspeech4
2021 MetricNet: Towards Improved Modeling For Non-Intrusive Speech Quality Assessment
abstract
The objective speech quality assessment is usually conducted by comparing received speech signal with its clean reference, while human beings are capable of evaluating the speech quality without any reference, such as in the mean opinion score (MOS) tests. Non-intrusive speech quality assessment has attracted much attention recently due to the lack of access to clean reference signals for objective evaluations in real scenarios. In this paper, we propose a novel non-intrusive speech quality measurement model, MetricNet, which leverages label distribution learning and joint speech reconstruction learning to achieve significantly improved performance compared to the existing non-intrusive speech quality measurement models. We demonstrate that the proposed approach yields promisingly high correlation to the intrusive objective evaluation of speech quality on clean, noisy and processed speech data.
Meng Yu 0003, Yong Xu 0004, Shixiong Zhang 0001, Dong Yu 0001
Interspeech5
2021 Video-aided Unsupervised Grammar Induction
abstract
Songyang Zhang, Linfeng Song, Lifeng Jin, Kun Xu, Dong Yu, Jiebo Luo. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Songyang Zhang 0004, Linfeng Song, Lifeng Jin, Kun Xu 0005, Dong Yu 0001, Jiebo Luo 0001
NAACL-HLT5
2021 Distant Finetuning with Discourse Relations for Stance Classification
Lifeng Jin, Kun Xu 0005, Linfeng Song, Dong Yu 0001
NLPCC (2)4
2021 Effective Low-Cost Time-Domain Audio Separation Using Globally Attentive Locally Recurrent Networks
abstract
Recent research on the time-domain audio separation networks (TasNets) has brought great success to speech separation. Nevertheless, conventional TasNets struggle to satisfy the memory and latency constraints in industrial applications. In this regard, we design a low-cost high-performance architecture, namely, globally attentive locally recurrent (GALR) network. Alike the dual-path RNN (DPRNN), we first split a feature sequence into 2D segments and then process the sequence along both the intra- and inter-segment dimensions. Our main innovation lies in that, on top of features recurrently processed along the inter-segment dimensions, GALR applies a self-attention mechanism to the sequence along the inter-segment dimension, which aggregates context-aware information and also enables parallelization. Our experiments suggest that GALR is a notably more effective network than the prior work. On one hand, with only 1.5M parameters, it has achieved comparable separation performance at a much lower cost with 36.1% less runtime memory and 49.4% fewer computational operations, relative to the DPRNN. On the other hand, in a comparable model size with DPRNN, GALR has consistently outperformed DPRNN in three datasets, in particular, with a substantial margin of 2.4dB absolute improvement of SI-SNRi in the benchmark WSJ0-2mix task.
Max W. Y. Lam, Jun Wang 0091, Dan Su 0002, Dong Yu 0001
SLT4
2021 Neural Mask based Multi-channel Convolutional Beamforming for Joint Dereverberation, Echo Cancellation and Denoising
abstract
This paper proposes a new joint optimization framework for simultaneous dereverberation, acoustic echo cancellation, and denoising, which is motivated by the recently proposed con-volutional beamformer for simultaneous denoising and dereverberation. Using the echo aware mask based beamforming framework, the proposed algorithm could effectively deal with double-talk case and local inference, etc. The evaluations based on ERLE for echo only, and PESQ for double-talk demonstrate that the proposed algorithm could significantly improve the performance.
Meng Yu 0003, Yong Xu 0004, Chao Weng, Shixiong Zhang 0001, Lianwu Chen, Dong Yu 0001
SLT7
2021 WPD++: An Improved Neural Beamformer for Simultaneous Speech Separation and Dereverberation
abstract
This paper aims at eliminating the interfering speakers' speech, additive noise, and reverberation from the noisy multi-talker speech mixture that benefits automatic speech recognition (ASR) backend. While the recently proposed Weighted Power minimization Distortionless response (WPD) beamformer can perform separation and dereverberation simultaneously, the noise cancellation component still has the potential to progress. We propose an improved neural WPD beamformer called "WPD++" by an enhanced beamforming module in the conventional WPD and a multi-objective loss function for the joint training. The beamforming module is improved by utilizing the spatio-temporal correlation. A multi-objective loss, including the complex spectra domain scale-invariant signal-to-noise ratio (C-Si-SNR) and the magnitude domain mean square error (Mag-MSE), is properly designed to make multiple constraints on the enhanced speech and the desired power of the dry clean signal. Joint training is conducted to optimize the complex-valued mask estimator and the WPD++ beamformer in an end-to-end way. The results show that the proposed WPD++ outperforms several state-of-the-art beamformers on the enhanced speech quality and word error rate (WER) of ASR.
Zhaoheng Ni, Yong Xu 0004, Meng Yu 0003, Bo Wu 0011, Shixiong Zhang 0001, Dong Yu 0001, Michael I. Mandel
SLT6
2021 Complex Neural Spatial Filter: Enhancing Multi-Channel Target Speech Separation in Complex Domain
abstract
To date, mainstream target speech separation (TSS) approaches are formulated to estimate the complex ratio mask (cRM) of target speech in time-frequency domain under supervised deep learning framework. However, the existing methods are designed in the way that the real and imaginary parts of the cRM are separately modeled using real-valued training data pairs. The research motivation of this study is to design a deep model that fully exploits the temporal-spectral-spatial information of multi-channel signals for estimating cRM directly and efficiently in complex domain. As a result, a novel TSS network is designed consisting of two modules, a complex neural spatial filter (cNSF) and an MVDR. Essentially, cNSF is a cRM estimation model and an MVDR module is cascaded to the cNSF module to reduce the nonlinear speech distortions introduced by neural network. Specifically, to fit the cRM target, all input features of cNSF are reformulated into complex-valued representations. Then, to achieve good hierarchical feature abstraction, a complex deep neural network (cDNN) is delicately designed with U-Net structure. Experiments conducted on simulated multi-channel speech data demonstrate the proposed cNSF outperforms the baseline NSF by 12.1% scale-invariant signal-to-distortion ratio and 33.1% word error rate.
Rongzhi Gu, Shixiong Zhang 0001, Yuexian Zou, Dong Yu 0001
IEEE Signal Process. Lett.4
2021 An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and Separation
abstract
Speech enhancement and speech separation are two related tasks, whose purpose is to extract either one or more target speech signals, respectively, from a mixture of sounds generated by several sources. Traditionally, these tasks have been tackled using signal processing and machine learning techniques applied to the available acoustic signals. Since the visual aspect of speech is essentially unaffected by the acoustic environment, visual information from the target speakers, such as lip movements and facial expressions, has also been used for speech enhancement and speech separation systems. In order to efficiently fuse acoustic and visual information, researchers have exploited the flexibility of data-driven approaches, specifically deep learning, achieving strong performance. The ceaseless proposal of a large number of techniques to extract features and fuse multimodal information has highlighted the need for an overview that comprehensively describes and discusses audio-visual speech enhancement and separation based on deep learning. In this paper, we provide a systematic survey of this research topic, focusing on the main elements that characterise the systems in the literature: acoustic features; visual features; deep learning methods; fusion techniques; training targets and objective functions. In addition, we review deep-learning-based methods for speech reconstruction from silent videos and audio-visual sound source separation for non-speech signals, since these methods can be more or less directly applied to audio-visual speech enhancement and separation. Finally, we survey commonly employed audio-visual speech datasets, given their central role in the development of data-driven approaches, and evaluation methods, because they are generally used to compare different systems and determine their performance.
Daniel Michelsanti, Zheng-Hua Tan, Shixiong Zhang 0001, Yong Xu 0004, Meng Yu 0003, Dong Yu 0001, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.6
2021 Conversational Semantic Role Labeling
abstract
Semantic role labeling (SRL) aims to extract the arguments for each predicate in an input sentence. Traditional SRL can fail to analyze dialogues because it only works on every single sentence, while ellipsis and anaphora frequently occur in dialogues. To address this problem, we propose the conversational SRL task, where an argument can be the dialogue participants, a phrase in the dialogue history or the current sentence. As the existing SRL datasets are in the sentence level, we manually annotate semantic roles for 3000 chit-chat dialogues (27198 sentences) to boost the research in this direction. Experiments show that while traditional SRL systems (even with the help of coreference resolution or rewriting) perform poorly for analyzing dialogues, modeling dialogue histories and participants greatly helps the performance, indicating that adapting SRL to conversations is very promising for universal dialogue understanding. Our initial study by applying CSRL to two mainstream conversational tasks, dialogue response generation and dialogue context rewriting, also confirms the usefulness of CSRL.
Kun Xu 0005, Han Wu 0004, Linfeng Song, Haisong Zhang, Linqi Song, Dong Yu 0001
IEEE ACM Trans. Audio Speech Lang. Process.6
2021 Audio-Visual Multi-Channel Integration and Recognition of Overlapped Speech
abstract
Automatic speech recognition (ASR) technologies have been significantly advanced in the past few decades. However, recognition of overlapped speech remains a highly challenging task to date. To this end, multi-channel microphone array data are widely used in current ASR systems. Motivated by the invariance of visual modality to acoustic signal corruption and the additional cues they provide to separate the target speaker from the interfering sound sources, this paper presents an audio-visual multi-channel based recognition system for overlapped speech. It benefits from a tight integration between a speech separation front-end and recognition back-end, both of which incorporate additional video input. A series of audio-visual multi-channel speech separation front-end components based on TF masking, Filter&Sum and mask-based MVDR neural channel integration approaches are developed. To reduce the error cost mismatch between the separation and the recognition components, the entire system is jointly fine-tuned using a multi-task criterion interpolation of the scale-invariant signal to noise ratio (Si-SNR) with either the connectionist temporal classification (CTC), or lattice-free maximum mutual information (LF-MMI) loss function. Experiments suggest that: the proposed audio-visual multi-channel recognition system outperforms the baseline audio-only multi-channel ASR system by up to 8.04% (31.68% relative) and 22.86% (58.51% relative) absolute WER reduction on overlapped speech constructed using either simulation or replaying of the LRS2 dataset respectively. Consistent performance improvements are also obtained using the proposed audio-visual multi-channel recognition system when using occluded video input with the lip region randomly covered up to 60%.
Jianwei Yu 0001, Shixiong Zhang 0001, Bo Wu 0011, Shansong Liu, Shoukang Hu, Mengzhe Geng, Xunying Liu, Helen M. Meng, Dong Yu 0001
IEEE ACM Trans. Audio Speech Lang. Process.9
2021 Multi-Channel Multi-Frame ADL-MVDR for Target Speech Separation
abstract
Many purely neural network based speech separation approaches have been proposed to improve objective assessment scores, but they often introduce nonlinear distortions that are harmful to modern automatic speech recognition (ASR) systems. Minimum variance distortionless response (MVDR) filters are often adopted to remove nonlinear distortions, however, conventional neural mask-based MVDR systems still result in relatively high levels of residual noise. Moreover, the matrix inverse involved in the MVDR solution is sometimes numerically unstable during joint training with neural networks. In this study, we propose a multi-channel multi-frame (MCMF) all deep learning (ADL)-MVDR approach for target speech separation, which extends our preliminary multi-channel ADL-MVDR approach. The proposed MCMF ADL-MVDR system addresses linear and nonlinear distortions. Spatio-temporal cross correlations are also fully utilized in the proposed approach. The proposed systems are evaluated using a Mandarin audio-visual corpus and are compared with several state-of-the-art approaches. Experimental results demonstrate the superiority of our proposed systems under different scenarios and across several objective evaluation metrics, including ASR performance.
Zhuohuang Zhang, Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Lianwu Chen, Donald S. Williamson, Dong Yu 0001
IEEE ACM Trans. Audio Speech Lang. Process.7
2020 Relation Extraction Exploiting Full Dependency Forests
abstract
Dependency syntax has long been recognized as a crucial source of features for relation extraction. Previous work considers 1-best trees produced by a parser during preprocessing. However, error propagation from the out-of-domain parser may impact the relation extraction performance. We propose to leverage full dependency forests for this task, where a full dependency forest encodes all possible trees. Such representations of full dependency forests provide a differentiable connection between a parser and a relation extraction model, and thus we are also able to study adjusting the parser parameters based on end-task loss. Experiments on three datasets show that full dependency forests and parser adjustment give significant improvements over carefully designed baselines, showing state-of-the-art or competitive performances on biomedical or newswire benchmarks.
Lifeng Jin, Linfeng Song, Yue Zhang 0004, Kun Xu 0005, Wei-Yun Ma, Dong Yu 0001
AAAI6
2020 Joint Parsing and Generation for Abstractive Summarization
abstract
Sentences produced by abstractive summarization systems can be ungrammatical and fail to preserve the original meanings, despite being locally fluent. In this paper we propose to remedy this problem by jointly generating a sentence and its syntactic dependency parse while performing abstraction. If generating a word can introduce an erroneous relation to the summary, the behavior must be discouraged. The proposed method thus holds promise for producing grammatical sentences and encouraging the summary to stay true-to-original. Our contributions of this work are twofold. First, we present a novel neural architecture for abstractive summarization that combines a sequential decoder with a tree-based decoder in a synchronized manner to generate a summary sentence and its syntactic parse. Secondly, we describe a novel human evaluation protocol to assess if, and to what extent, a summary remains true to its original meanings. We evaluate our method on a number of summarization datasets and demonstrate competitive results against strong baselines.
Kaiqiang Song, Logan Lebanoff, Qipeng Guo, Xipeng Qiu, Xiangyang Xue 0001, Chen Li 0003, Dong Yu 0001, Fei Liu 0004
AAAI7
2020 Coordinated Reasoning for Cross-Lingual Knowledge Graph Alignment
abstract
Existing entity alignment methods mainly vary on the choices of encoding the knowledge graph, but they typically use the same decoding method, which independently chooses the local optimal match for each source entity. This decoding method may not only cause the “many-to-one” problem but also neglect the coordinated nature of this task, that is, each alignment decision may highly correlate to the other decisions. In this paper, we introduce two coordinated reasoning methods, i.e., the Easy-to-Hard decoding strategy and joint entity alignment algorithm. Specifically, the Easy-to-Hard strategy first retrieves the model-confident alignments from the predicted results and then incorporates them as additional knowledge to resolve the remaining model-uncertain alignments. To achieve this, we further propose an enhanced alignment model that is built on the current state-of-the-art baseline. In addition, to address the many-to-one problem, we propose to jointly predict entity alignments so that the one-to-one constraint can be naturally incorporated into the alignment prediction. Experimental results show that our model achieves the state-of-the-art performance and our reasoning methods can also significantly improve existing baselines.
Kun Xu 0005, Linfeng Song, Yansong Feng 0002, Yan Song 0003, Dong Yu 0001
AAAI5
2020 Recurrent Chunking Mechanisms for Long-Text Machine Reading Comprehension
abstract
In this paper, we study machine reading comprehension (MRC) on long texts, where a model takes as inputs a lengthy document and a question and then extracts a text span from the document as an answer.State-of-the-art models tend to use a pretrained transformer model (e.g., BERT) to encode the joint contextual information of document and question.However, these transformer-based models can only take a fixed-length (e.g., 512) text as its input.To deal with even longer text inputs, previous approaches usually chunk them into equally-spaced segments and predict answers based on each segment independently without considering the information from other segments.As a result, they may form segments that fail to cover the correct answer span or retain insufficient contexts around it, which significantly degrades the performance.Moreover, they are less capable of answering questions that need cross-segment information.We propose to let a model learn to chunk in a more flexible way via reinforcement learning: a model can decide the next segment that it wants to process in either direction.We also employ recurrent mechanisms to enable information to flow across segments.Experiments on three MRC datasets -CoQA, QuAC, and TriviaQA -demonstrate the effectiveness of our proposed recurrent chunking mechanisms: we can obtain segments that are more likely to contain complete answers and at the same time provide sufficient contexts around the ground truth answers for better predictions.
Hongyu Gong, Yelong Shen, Dian Yu 0001, Jianshu Chen, Dong Yu 0001
ACL5
2020 MART: Memory-Augmented Recurrent Transformer for Coherent Video Paragraph Captioning
abstract
Generating multi-sentence descriptions for videos is one of the most challenging captioning tasks due to its high requirements for not only visual relevance but also discoursebased coherence across the sentences in the paragraph.Towards this goal, we propose a new approach called Memory-Augmented Recurrent Transformer (MART), which uses a memory module to augment the transformer architecture.The memory module generates a highly summarized memory state from the video segments and the sentence history so as to help better prediction of the next sentence (w.r.t.coreference and repetition aspects), thus encouraging coherent paragraph generation.Extensive experiments, human evaluations, and qualitative analyses on two popular datasets ActivityNet Captions and YouCookII show that MART generates more coherent and less repetitive paragraph captions than baseline methods, while maintaining relevance to the input video events. 1
Jie Lei 0003, Liwei Wang 0009, Yelong Shen, Dong Yu 0001, Tamara L. Berg, Mohit Bansal
ACL4
2020 Structural Information Preserving for Graph-to-Text Generation
abstract
The task of graph-to-text generation aims at producing sentences that preserve the meaning of input graphs.As a crucial defect, the current state-of-the-art models may mess up or even drop the core structural information of input graphs when generating outputs.We propose to tackle this problem by leveraging richer training signals that can guide our model for preserving input information.In particular, we introduce two types of autoencoding losses, each individually focusing on different aspects (a.k.a.views) of input graphs.The losses are then back-propagated to better calibrate our model via multi-task training.Experiments on two benchmarks for graph-to-text generation show the effectiveness of our approach over a state-of-the-art baseline.Our code is available at http://github.com/ Soistesimmer/AMR-multiview.
Linfeng Song, Ante Wang, Jinsong Su, Yue Zhang 0004, Kun Xu 0005, Yubin Ge, Dong Yu 0001
ACL7
2020 ZPR2: Joint Zero Pronoun Recovery and Resolution using Multi-Task Learning and BERT
abstract
Zero pronoun recovery and resolution aim at recovering the dropped pronoun and pointing out its anaphoric mentions, respectively.We propose to better explore their interaction by solving both tasks together, while the previous work treats them separately.For zero pronoun resolution, we study this task in a more realistic setting, where no parsing trees or only automatic trees are available, while most previous work assumes gold trees.Experiments on two benchmarks show that joint modeling significantly outperforms our baseline that already beats the previous state of the arts.
Linfeng Song, Kun Xu 0005, Yue Zhang 0004, Jianshu Chen, Dong Yu 0001
ACL5
2020 Towards Faithful Neural Table-to-Text Generation with Content-Matching Constraints
abstract
Text generation from a knowledge base aims to translate knowledge triples to naturallanguage descriptions.Most existing methods ignore the faithfulness between a generated text description and the original table, leading to generated information that goes beyond the content of the table.In this paper, for the first time, we propose a novel Transformerbased generation framework to achieve the goal.The core techniques in our method to enforce faithfulness include a new table-text optimal-transport matching loss and a tabletext embedding similarity loss based on the Transformer model.Furthermore, to evaluate faithfulness, we propose a new automatic metric specialized to the table-to-text generation problem.We also provide detailed analysis on each component of our model in our experiments.Automatic and human evaluations show that our framework can significantly outperform state-of-the-art by a large margin.
Zhenyi Wang 0001, Xiaoyang Wang 0001, Bang An 0001, Dong Yu 0001, Changyou Chen
ACL4
2020 Dialogue-Based Relation Extraction
abstract
We present the first human-annotated dialoguebased relation extraction (RE) dataset Dialo-gRE, aiming to support the prediction of relation(s) between two arguments that appear in a dialogue.We further offer DialogRE as a platform for studying cross-sentence RE as most facts span multiple sentences.We argue that speaker-related information plays a critical role in the proposed task, based on an analysis of similarities and differences between dialogue-based and traditional RE tasks.Considering the timeliness of communication in a dialogue, we design a new metric to evaluate the performance of RE methods in a conversational setting and investigate the performance of several representative RE methods on DialogRE.Experimental results demonstrate that a speaker-aware extension on the best-performing model leads to gains in both the standard and conversational evaluation settings.DialogRE is available at https:// dataset.org/dialogre/.
Dian Yu 0001, Kai Sun 0006, Claire Cardie, Dong Yu 0001
ACL4
2020 Comprehensive Image Captioning via Scene Graph Decomposition
Yiwu Zhong, Liwei Wang 0009, Jianshu Chen, Dong Yu 0001, Yin Li 0003
ECCV (14)4
2020 Better Highlighting: Creating Sub-Sentence Summary Highlights
abstract
Amongst the best means to summarize is highlighting. In this paper, we aim to generate summary highlights to be overlaid on the original documents to make it easier for readers to sift through a large amount of text. The method allows summaries to be understood in context to prevent a summarizer from distorting the original meaning, of which abstractive summarizers usually fall short. In particular, we present a new method to produce self-contained highlights that are understandable on their own to avoid confusion. Our method combines determinantal point processes and deep contextualized representations to identify an optimal set of sub-sentence segments that are both important and non-redundant to form summary highlights. To demonstrate the flexibility and modeling power of our method, we conduct extensive experiments on summarization datasets. Our analysis provides evidence that highlighting is a promising avenue of research towards future summarization.
Sangwoo Cho, Kaiqiang Song, Chen Li 0003, Dong Yu 0001, Hassan Foroosh, Fei Liu 0004
EMNLP (1)4
2020 Semantic Role Labeling Guided Multi-turn Dialogue ReWriter
abstract
For multi-turn dialogue rewriting, the capacity of effectively modeling the linguistic knowledge in dialog context and getting rid of the noises is essential to improve its performance.Existing attentive models attend to all words without prior focus, which results in inaccurate concentration on some dispensable words.In this paper, we propose to use semantic role labeling (SRL), which highlights the core semantic information of who did what to whom, to provide additional guidance for the rewriter model.Experiments show that this information significantly improves a RoBERTa-based model that already outperforms previous stateof-the-art systems.
Kun Xu 0005, Haochen Tan, Linfeng Song, Han Wu 0004, Haisong Zhang, Linqi Song, Dong Yu 0001
EMNLP (1)7
2020 Code-Switched Speech Synthesis Using Bilingual Phonetic Posteriorgram with Only Monolingual Corpora
abstract
Synthesizing fluent code-switched (CS) speech with consistent voice using only monolingual corpora is still a challenging task, since language alternation seldom occurs during training and the speaker identity is directly correlated with language. In this paper, we present a bilingual phonetic posteriorgram (PPG) based CS speech synthesizer using only monolingual corpora. The bilingual PPG is used to bridge across speakers and languages, which is formed by stacking two monolingual PPGs extracted from two monolingual speaker-independent speech recognition systems. It is assumed that bilingual PPG can represent the articulation of speech sounds speaker-independently and captures accurate phonetic information of both languages in the same feature space. The proposed model first extracts bilingual PPGs from training data. Then an encoder- decoder based model is used to learn the relationship between input text and bilingual PPGs, and the bilingual PPGs are mapped to acoustic features using bidirectional long-short term memory based model conditioned on speaker embedding to control speaker identity. Experiments validate the effectiveness of the proposed model in terms of speech intelligibility, audio fidelity and speaker consistency of the generated code-switched speech.
Yuewen Cao, Songxiang Liu, Xixin Wu, Shiyin Kang, Zhiyong Wu 0001, Xunying Liu, Dan Su 0002, Dong Yu 0001, Helen M. Meng
ICASSP9
2020 Pitchnet: Unsupervised Singing Voice Conversion with Pitch Adversarial Network
abstract
Singing voice conversion is to convert a singer's voice to another one's voice without changing singing content. Recent work shows that unsupervised singing voice conversion can be achieved with an autoencoder-based approach [1]. However, the converted singing voice can be easily out of key, showing that the existing approach cannot model the pitch information precisely. In this paper, we propose to advance the existing unsupervised singing voice conversion method proposed in [1] to achieve more accurate pitch translation and flexible pitch manipulation. Specifically, the proposed Pitch-Net added an adversarially trained pitch regression network to enforce the encoder network to learn pitch invariant phoneme representation, and a separate module to feed pitch extracted from the source audio to the decoder network. Our evaluation shows that the proposed method can greatly improve the quality of the converted singing voice (2.92 vs 3.75 in MOS). We also demonstrate that the pitch of converted singing can be easily controlled during generation by changing the levels of the extracted pitch before passing it to the decoder network.
Chengqi Deng, Chengzhu Yu, Heng Lu 0004, Chao Weng, Dong Yu 0001
ICASSP5
2020 Enhancing End-to-End Multi-Channel Speech Separation Via Spatial Feature Learning
abstract
Hand-crafted spatial features (e.g., inter-channel phase difference, IPD) play a fundamental role in recent deep learning based multi-channel speech separation (MCSS) methods. However, these manually designed spatial features are hard to incorporate into the end-to-end optimized MCSS framework. In this work, we propose an integrated architecture for learning spatial features directly from the multi-channel speech waveforms within an end-to-end speech separation framework. In this architecture, time-domain filters spanning signal channels are trained to perform adaptive spatial filtering. These filters are implemented by a 2d convolution (conv2d) layer and their parameters are optimized using a speech separation objective function in a purely data-driven fashion. Furthermore, inspired by the IPD formulation, we design a conv2d kernel to compute the inter-channel convolution differences (ICDs), which are expected to provide the spatial cues that help to distinguish the directional sources. Evaluation results on simulated multi-channel reverberant WSJ0 2-mix dataset demonstrate that our proposed ICD based MCSS model improves the overall signal-to-distortion ratio by 10.4% over the IPD based MCSS model.
Rongzhi Gu, Shixiong Zhang 0001, Lianwu Chen, Yong Xu 0004, Meng Yu 0003, Dan Su 0002, Yuexian Zou, Dong Yu 0001
ICASSP8
2020 A Random Gossip BMUF Process for Neural Language Modeling
abstract
Neural network language model (NNLM) is an essential component of industrial ASR systems. One important challenge of training an NNLM is to leverage between scaling the learning process and handling big data. Conventional approaches such as block momentum provides a blockwise model update filtering (BMUF) process and achieves almost linear speedups with no performance degradation for speech recognition. However, it needs to calculate the model average from all computing nodes (e.g., GPUs) and when the number of computing nodes is large, the learning suffers from the severe communication latency. As a consequence, BMUF is not suitable under restricted network conditions. In this paper, we present a decentralized BMUF process, in which the model is split into different components, each of which is updated by communicating to some randomly chosen neighbor nodes with the same component, followed by a BMUF-like process. We apply this method to several LSTM language modeling tasks. Experimental results show that our approach achieves consistently better performance than conventional BMUF. In particular, we obtain a lower perplexity than the single-GPU baseline on the wiki-text-103 benchmark using 4 GPUs. In addition, no performance degradation is observed when scaling to 8 and 16 GPUs.
Jinchuan Tian, Guangsen Wang, Xingcheng Song, Dan Su 0002, Dong Yu 0001
ICASSP7
2020 Integration of Multi-Look Beamformers for Multi-Channel Keyword Spotting
abstract
Keyword spotting (KWS) is in great demand in smart devices in the era of Internet of Things. Albeit recent progresses, the performance of KWS, measured in false alarms and false rejects, may still degrade significantly under the far field and noisy conditions. In this paper, we propose integrating multiple beamformed signals and a microphone signal as input to an end-to-end KWS model and leveraging the attention mechanism to dynamically tune the model’s attention to the reliable input sources. We demonstrate, on our large simulated and recorded noisy and far-field evaluation sets, that our proposed approach significantly improves the KWS performance and reduces the computation cost against the baseline KWS systems.
Meng Yu 0003, Jie Chen 0057, Jimeng Zheng, Dan Su 0002, Dong Yu 0001
ICASSP6
2020 Speaker-Aware Target Speaker Enhancement by Jointly Learning with Speaker Embedding Extraction
abstract
Deep learning based speech separation approaches have received great interest, among which the recent speaker-aware speech enhancement methods are promising for solving difficulties such as arbitrary source permutation and unknown number of sources. In this paper, we propose a novel training framework which jointly learns the speaker-conditioned target speaker extraction model and its associated speaker embedding model. The resulting unified model directly learns the appropriate speaker embedding for improved target speech enhancement. We demonstrate, on our large simulated noisy and far-field evaluation sets of overlapped speech signals, that our proposed approach significantly improves the speech enhancement performance compared to the baseline speaker-aware speech enhancement models.
Meng Yu 0003, Dan Su 0002, Dong Yu 0001
ICASSP7
2020 Mixup-breakdown: A Consistency Training Method for Improving Generalization of Speech Separation Models
abstract
Deep-learning based speech separation models confront poor generalization problem that even the state-of-the-art models could abruptly fail when evaluating them in mismatch conditions. To address this problem, we propose an easy-to-implement yet effective consistency based semi-supervised learning (SSL) approach, namely Mixup-Breakdown training (MBT). It learns a teacher model to "breakdown" unlabeled inputs, and the estimated separations are interpolated to produce more useful pseudo "mixup" input-output pairs, on which the consistency regularization could apply for learning a student model. In our experiment, we evaluate MBT under various conditions with ascending degrees of mismatch, including unseen interfering speech, noise, and music, and compare MBT’s generalization capability against state-of-the-art supervised learning and SSL approaches. The result indicates that MBT significantly outperforms several strong baselines with up to 13.77% relative SI-SNRi improvement. Moreover, MBT only adds negligible computational overhead to standard training schemes.
Max W. Y. Lam, Jun Wang 0091, Dan Su 0002, Dong Yu 0001
ICASSP4
2020 Multi-Level Deep Neural Network Adaptation for Speaker Verification Using MMD and Consistency Regularization
abstract
Adapting speaker verification (SV) systems to a new environment is a very challenging task. Current adaptation methods in SV mainly focus on the backend, i.e, adaptation is carried out after the speaker embeddings have been created. In this paper, we present a DNN-based adaptation method using maximum mean discrepancy (MMD). Our method exploits two important aspects neglected by previous research. First, instead of minimizing domain discrepancy at utterance-level alone, our method minimizes domain discrepancy at both frame-level and utterance-level, which we believe will make the adaptation more robust to the duration discrepancy between training data and test data. Second, we introduce a consistency regularization for unlabelled target-domain data. The consistency regularization encourages the target speaker embeddings robust to adverse perturbations. Experiments on NIST SRE 2016 and 2018 show that our DNN adaptation works significantly better than the previously proposed DNN adaptation methods. What's more, our method works well with backend adaptation. By combining the proposed method with backend adaptation, we achieve a 9% improvement over backend adaptation in SRE18.
Weiwei Lin 0002, Man-Wai Mak, Na Li 0012, Dan Su 0002, Dong Yu 0001
ICASSP5
2020 End-To-End Accent Conversion Without Using Native Utterances
abstract
Techniques for accent conversion (AC) aim to convert non-native to native accented speech. Conventional AC methods try to convert only the speaker identity of a native speaker's voice to that of the non-native accented target speaker, leaving the underlying content and pronunciations unchanged. This hinders their practical use in real-world applications, because native-accented utterances are required at conversion stage. In this paper, we present an end-to-end framework, which is able to conduct AC from non-native-accented utterances without using any native-accented utterances during online conversion. We achieve this by independently extracting linguistic and speaker representations from non-native accented speech and condition a speech synthesis model on these representations to generate native-accented speech. Experiments on open-source data corpora show that the proposed system can convert Hindi-accented English speech into native American English speech with high naturalness, which is indistinguishable from native-accented recordings in terms of accent.
Songxiang Liu, Disong Wang, Yuewen Cao, Lifa Sun, Xixin Wu, Shiyin Kang, Zhiyong Wu 0001, Xunying Liu, Dan Su 0002, Dong Yu 0001, Helen M. Meng
ICASSP10
2020 Far-Field Location Guided Target Speech Extraction Using End-to-End Speech Recognition Objectives
abstract
Target speech extraction is a specific case of source separation where an auxiliary information like the location or some pre-saved anchor speech examples of the target speaker is used to resolve the permutation ambiguity. Traditionally such systems are optimized based on signal reconstruction objectives. Recently end-to-end automatic speech recognition (ASR) methods have enabled to optimize source separation systems with only the transcription based objective. This paper proposes a method to jointly optimize a location guided target speech extraction module along with a speech recognition module only with ASR error minimization criteria. Experimental comparisons with corresponding conventional pipeline systems verify that this task can be realized by end-to-end ASR training objectives without using parallel clean data. We show promising target speech recognition results in mixtures of two speakers and noise, and discuss interesting properties of the proposed system in terms of speech enhancement/separation objectives and word error rates. Finally, we design a system that can take both location and anchor speech as input at the same time and show that the performance can be further improved.
Aswin Shanmugam Subramanian, Chao Weng, Meng Yu 0003, Shixiong Zhang 0001, Yong Xu 0004, Shinji Watanabe 0001, Dong Yu 0001
ICASSP7
2020 Improving Reverberant Speech Training Using Diffuse Acoustic Simulation
abstract
We present an efficient and realistic geometric acoustic simulation approach for generating and augmenting training data in speech-related machine learning tasks. Our physically-based acoustic simulation method is capable of modeling occlusion, specular and diffuse reflections of sound in complicated acoustic environments, whereas the classical image method can only model specular reflections in simple room settings. We show that by using our synthetic training data, the same neural networks gain significant performance improvement on real test sets in far-field speech recognition by 1.58% and keyword spotting by 21%, without fine-tuning using real impulse responses.
Zhenyu Tang 0001, Lianwu Chen, Bo Wu 0011, Dong Yu 0001, Dinesh Manocha
ICASSP4
2020 Dfsmn-San with Persistent Memory Model for Automatic Speech Recognition
abstract
Self-attention networks (SAN) have been introduced into automatic speech recognition (ASR) and achieved state-of-the-art performance owing to its superior ability in capturing long term dependency. One of the key ingredients is the self-attention mechanism which can be effectively performed on the whole utterance level. In this paper, we try to investigate whether even more information beyond the whole utterance level can be exploited and beneficial. We propose to apply self-attention layer with augmented memory to ASR. Specifically, we first propose a variant model architecture which combines deep feed-forward sequential memory network (DFSMN) with self-attention layers to form a better baseline model compared with a purely self-attention network. Then, we propose and compare two kinds of additional memory structures added into self-attention layers. Experiments on large-scale LVCSR tasks show that on four individual test sets, the DFSMN-SAN architecture outperforms vanilla SAN encoder by 5% relatively in character error rate (CER). More importantly, the additional memory structure provides further 5% to 11% relative improvement in CER.
Zhao You, Dan Su 0002, Jie Chen 0057, Chao Weng, Dong Yu 0001
ICASSP5
2020 Audio-Visual Recognition of Overlapped Speech for the LRS2 Dataset
abstract
Automatic recognition of overlapped speech remains a highly challenging task to date. Motivated by the bimodal nature of human speech perception, this paper investigates the use of audio-visual technologies for overlapped speech recognition. Three issues associated with the construction of audio-visual speech recognition (AVSR) systems are addressed. First, the basic architecture designs i.e. end-to-end and hybrid of AVSR systems are investigated. Second, purposefully designed modality fusion gates are used to robustly integrate the audio and visual features. Third, in contrast to a traditional pipelined architecture containing explicit speech separation and recognition components, a streamlined and integrated AVSR system optimized consistently using the lattice-free MMI (LF-MMI) discriminative criterion is also proposed. The proposed LF-MMI time-delay neural network (TDNN) system establishes the state-of-the-art for the LRS2 dataset. Experiments on overlapped speech simulated from the LRS2 dataset suggest the proposed AVSR system outperformed the audio only baseline LF-MMI DNN system by up to 29.98% absolute in word error rate (WER) reduction, and produced recognition performance comparable to a more complex pipelined system. Consistent performance improvements of 4.89% absolute in WER reduction over the baseline AVSR system using feature fusion are also obtained.
Jianwei Yu 0001, Shixiong Zhang 0001, Jian Wu 0027, Shahram Ghorbani, Bo Wu 0011, Shiyin Kang, Shansong Liu, Xunying Liu, Helen M. Meng, Dong Yu 0001
ICASSP10
2020 End-to-End Multi-Look Keyword Spotting
abstract
The performance of keyword spotting (KWS), measured in false alarms and false rejects, degrades significantly under the far field and noisy conditions. In this paper, we propose a multi-look neural network modeling for speech enhancement which simultaneously steers to listen to multiple sampled look directions. The multi-look enhancement is then jointly trained with KWS to form an end-to-end KWS model which integrates the enhanced signals from multiple look directions and leverages an attention mechanism to dynamically tune the model's attention to the reliable sources. We demonstrate, on our large noisy and far-field evaluation sets, that the proposed approach significantly improves the KWS performance against the baseline KWS system and a recent beamformer based multi-beam KWS system.
Meng Yu 0003, Bo Wu 0011, Dan Su 0002, Dong Yu 0001
INTERSPEECH5
2020 Investigating Robustness of Adversarial Samples Detection for Automatic Speaker Verification
abstract
Recently adversarial attacks on automatic speaker verification (ASV) systems attracted widespread attention as they pose severe threats to ASV systems.However, methods to defend against such attacks are limited.Existing approaches mainly focus on retraining ASV systems with adversarial data augmentation.Also, countermeasure robustness against different attack settings are insufficiently investigated.Orthogonal to prior approaches, this work proposes to defend ASV systems against adversarial attacks with a separate detection network, rather than augmenting adversarial data into ASV training.A VGG-like binary classification detector is introduced and demonstrated to be effective on detecting adversarial samples.To investigate detector robustness in a realistic defense scenario where unseen attack settings may exist, we analyze various kinds of unseen attack settings' impact and observe that the detector is robust (6.27%EER det degradation in the worst case) against unseen substitute ASV systems, but it has weak robustness (50.37%EER det degradation in the worst case) against unseen perturbation methods.The weak robustness against unseen perturbation methods shows a direction for developing stronger countermeasures.
Xu Li 0015, Na Li 0012, Jinghua Zhong, Xixin Wu, Xunying Liu, Dan Su 0002, Dong Yu 0001, Helen M. Meng
INTERSPEECH7
2020 Transferring Source Style in Non-Parallel Voice Conversion
abstract
Voice conversion (VC) techniques aim to modify speaker identity of an utterance while preserving the underlying linguistic information.Most VC approaches ignore modeling of the speaking style (e.g.emotion and emphasis), which may contain the factors intentionally added by the speaker and should be retained during conversion.This study proposes a sequence-tosequence based non-parallel VC approach, which has the capability of transferring the speaking style from the source speech to the converted speech by explicitly modeling.Objective evaluation and subjective listening tests show superiority of the proposed VC approach in terms of speech naturalness and speaker similarity of the converted speech.Experiments are also conducted to show the source-style transferability of the proposed approach.
Songxiang Liu, Yuewen Cao, Shiyin Kang, Na Hu, Xunying Liu, Dan Su 0002, Dong Yu 0001, Helen M. Meng
INTERSPEECH7
2020 Minimum Bayes Risk Training of RNN-Transducer for End-to-End Speech Recognition
abstract
In this work, we propose minimum Bayes risk (MBR) training of RNN-Transducer (RNN-T) for end-to-end speech recognition.Specifically, initialized with a RNN-T trained model, MBR training is conducted via minimizing the expected edit distance between the reference label sequence and on-thefly generated N-best hypothesis.We also introduce a heuristic to incorporate an external neural network language model (NNLM) in RNN-T beam search decoding and explore MBR training with the external NNLM.Experimental results demonstrate an MBR trained model outperforms a RNN-T trained model substantially and further improvements can be achieved if trained with an external NNLM.Our best MBR trained system achieves absolute character error rate (CER) reductions of 1.2% and 0.5% on read and spontaneous Mandarin speech respectively over a strong convolution and transformer based RNN-T baseline trained on ∼21,000 hours of speech.
Chao Weng, Chengzhu Yu, Jia Cui, Dong Yu 0001
INTERSPEECH5
2020 Peking Opera Synthesis via Duration Informed Attention Network
abstract
Peking Opera has been the most dominant form of Chinese performing art since around 200 years ago.A Peking Opera singer usually exhibits a very strong personal style via introducing improvisation and expressiveness on stage which leads the actual rhythm and pitch contour to deviate significantly from the original music score.This inconsistency poses a great challenge in Peking Opera singing voice synthesis from a music score.In this work, we propose to deal with this issue and synthesize expressive Peking Opera singing from the music score based on the Duration Informed Attention Network (DurIAN) framework.To tackle the rhythm mismatch, Lagrange multiplier is used to find the optimal output phoneme duration sequence with the constraint of the given note duration from music score.As for the pitch contour mismatch, instead of directly inferring from music score, we adopt a pseudo music score generated from the real singing and feed it as input during training.The experiments demonstrate that with the proposed system we can synthesize Peking Opera singing voice with high-quality timbre, pitch and expressiveness.
Yusong Wu, Shengchen Li, Chengzhu Yu, Heng Lu 0004, Chao Weng, Dong Yu 0001
INTERSPEECH7
2020 Neural Spatio-Temporal Beamformer for Target Speech Separation
abstract
Purely neural network (NN) based speech separation and enhancement methods, although can achieve good objective scores, inevitably cause nonlinear speech distortions that are harmful for the automatic speech recognition (ASR).On the other hand, the minimum variance distortionless response (MVDR) beamformer with NN-predicted masks, although can significantly reduce speech distortions, has limited noise reduction capability.In this paper, we propose a multi-tap MVDR beamformer with complex-valued masks for speech separation and enhancement.Compared to the state-of-the-art NN-mask based MVDR beamformer, the multi-tap MVDR beamformer exploits the inter-frame correlation in addition to the intermicrophone correlation that is already utilized in prior arts.Further improvements include the replacement of the real-valued masks with the complex-valued masks and the joint training of the complex-mask NN.The evaluation on our multi-modal multi-channel target speech separation and enhancement platform demonstrates that our proposed multi-tap MVDR beamformer improves both the ASR accuracy and the perceptual speech quality against prior arts.
Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Lianwu Chen, Chao Weng, Dong Yu 0001
INTERSPEECH7
2020 DurIAN: Duration Informed Attention Network for Speech Synthesis
Chengzhu Yu, Heng Lu 0004, Na Hu, Meng Yu 0003, Chao Weng, Kun Xu 0005, Deyi Tuo, Shiyin Kang, Guangzhi Lei, Dan Su 0002, Dong Yu 0001
INTERSPEECH12
2020 Audio-Visual Multi-Channel Recognition of Overlapped Speech
abstract
Automatic speech recognition (ASR) of overlapped speech remains a highly challenging task to date. To this end, multi-channel microphone array data are widely used in state-of-the-art ASR systems. Motivated by the invariance of visual modality to acoustic signal corruption, this paper presents an audio-visual multi-channel overlapped speech recognition system featuring tightly integrated separation front-end and recognition back-end. A series of audio-visual multi-channel speech separation front-end components based on \textit{TF masking}, \textit{filter\&sum} and \textit{mask-based MVDR} beamforming approaches were developed. To reduce the error cost mismatch between the separation and recognition components, they were jointly fine-tuned using the connectionist temporal classification (CTC) loss function, or a multi-task criterion interpolation with scale-invariant signal to noise ratio (Si-SNR) error cost. Experiments suggest that the proposed multi-channel AVSR system outperforms the baseline audio-only ASR system by up to 6.81\% (26.83\% relative) and 22.22\% (56.87\% relative) absolute word error rate (WER) reduction on overlapped speech constructed using either simulation or replaying of the lipreading sentence 2 (LRS2) dataset respectively.
Jianwei Yu 0001, Bo Wu 0011, Rongzhi Gu, Shixiong Zhang 0001, Lianwu Chen, Yong Xu 0004, Meng Yu 0003, Dan Su 0002, Dong Yu 0001, Xunying Liu, Helen M. Meng
INTERSPEECH9
2020 DurIAN-SC: Duration Informed Attention Network Based Singing Voice Conversion System
abstract
Singing voice conversion is converting the timbre in the source singing to the target speaker's voice while keeping singing content the same. However, singing data for target speaker is much more difficult to collect compared with normal speech this http URL this paper, we introduce a singing voice conversion algorithm that is capable of generating high quality target speaker's singing using only his/her normal speech data. First, we manage to integrate the training and conversion process of speech and singing into one framework by unifying the features used in standard speech synthesis system and singing synthesis system. In this way, normal speech data can also contribute to singing voice conversion training, making the singing voice conversion system more robust especially when the singing database is small.Moreover, in order to achieve one-shot singing voice conversion, a speaker embedding module is developed using both speech and singing data, which provides target speaker identify information during conversion. Experiments indicate proposed sing conversion system can convert source singing to target speaker's high-quality singing with only 20 seconds of target speaker's enrollment speech data.
Chengzhu Yu, Heng Lu 0004, Chao Weng, Yusong Wu, Dong Yu 0001
INTERSPEECH9
2020 Building Digital Human
abstract
Digital humans find their applications in areas such as virtual companion, virtual reporter, and virtual narrator. As the global trend of digitalization continues, the value of digital humans continues to increase. For example, a virtual teacher may mimic human teachers to deliver personalized education to students spread all over the world at a lower cost. There are many technical difficulties yet to be solved to make digital humans truly valuable. In this talk, I report our recent progresses on addressing two of these difficulties: multi-modal text-to-speech synthesis and multi-modal voice separation and recognition. To address the multi-modal text-to-speech synthesis problem, we developed the duration informed attention network (DurIAN) [1]. DurIAN enhanced the attention-based alignment in the state-of-the-art (SOTA) end-to-end speech synthesis systems such as Tacotron2 [2] with duration information estimated from the rich text input. This technology, while generating high quality natural speech, avoids popular pitfalls such as word repetition and missing in the pure end-to-end systems. More importantly, the system can easily align the facial representation and synthesized speech through the duration model. To more robustly drive the facial expression and mouth movement, we developed a 3D-model guided framework for multi-modal synthesis. To solve the multi-modal voice separation and recognition problem, which is in need in many scenarios such as virtual receptionist, we developed an all deep learning beamformer [3] which integrates the conventional minimum variance distortionless response (MVDR) beamformer, the recurrent neural network-based statistics estimator, and the visual cue guided speaker tracing and diarization system [4]. Our novel approach significantly improved the quality of the separated speech.
Dong Yu 0001
ACM Multimedia1
2020 On the localness modeling for the self-attention based end-to-end speech synthesis
Shan Yang 0001, Heng Lu 0004, Shiyin Kang, Liumeng Xue, Jinba Xiao, Dan Su 0002, Lei Xie 0001, Dong Yu 0001
Neural Networks8
2020 Investigating Prior Knowledge for Challenging Chinese Machine Reading Comprehension
abstract
Machine reading comprehension tasks require a machine reader to answer questions relevant to the given document. In this paper, we present the first free-form multiple-Choice Chinese machine reading Comprehension dataset (C 3 ), containing 13,369 documents (dialogues or more formally written mixed-genre texts) and their associated 19,577 multiple-choice free-form questions collected from Chinese-as-a-second-language examinations. We present a comprehensive analysis of the prior knowledge (i.e., linguistic, domain-specific, and general world knowledge) needed for these real-world problems. We implement rule-based and popular neural methods and find that there is still a significant performance gap between the best performing model (68.5%) and human readers (96.0%), especiallyon problems that require prior knowledge. We further study the effects of distractor plausibility and data augmentation based on translated relevant datasets for English on model performance. We expect C 3 to present great challenges to existing systems as answering 86.8% of questions requires both knowledge within and beyond the accompanying document, and we hope that C 3 can serve as a platform to study how to leverage various kinds of prior knowledge to better understand a given written or orally oriented text. C 3 is available at https://dataset.org/c3/ .
Kai Sun 0006, Dian Yu 0001, Dong Yu 0001, Claire Cardie
Trans. Assoc. Comput. Linguistics3
2020 A Framework for Adapting DNN Speaker Embedding Across Languages
abstract
Language mismatch remains a major hindrance to the extensive deployment of speaker verification (SV) systems. Current language adaptation methods in SV mainly rely on linear projection in embedding space; i.e., adaptation is carried out after the speaker embeddings have been created, which underutilizes the powerful representation of deep neural networks. This article proposes a maximum mean discrepancy (MMD) based framework for adapting deep neural network (DNN) speaker embedding across languages, featuring multi-level domain loss, separate batch normalization, and consistency regularization. We refer to the framework as MSC. We show that (1) minimizing domain discrepancy at both frame- and utterance-levels performs significantly better than at utterance-level alone; (2) separating the source-domain data from the target-domain in batch normalization improves adaptation performance; and (3) data augmentation can be utilized in the unlabelled target-domain through consistency regularization. By combining these findings, we achieve an EER of 8.69% and 7.95% in NIST SRE 2016 and 2018, respectively, which are significantly better than the previously proposed DNN adaptation methods. Our framework also works well with backend adaptation. By combining the proposed framework with backend adaptation, we achieve an 11.8% improvement over the backend adaptation in SRE18. When applying our framework to a 121-layer Densenet, we achieved an EER of 7.81% and 7.02% in NIST SRE 2016 and 2018, respectively.
Weiwei Lin 0002, Man-Wai Mak, Na Li 0012, Dan Su 0002, Dong Yu 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2019 Reliability-aware Dynamic Feature Composition for Name Tagging
abstract
Word embeddings are widely used on a variety of tasks and can substantially improve the performance. However, their quality is not consistent throughout the vocabulary due to the long-tail distribution of word frequency. Without sufficient contexts, rare word embeddings are usually less reliable than those of common words. However, current models typically trust all word embeddings equally regardless of their reliability and thus may introduce noise and hurt the performance. Since names often contain rare and uncommon words, this problem is particularly critical for name tagging. In this paper, we propose a novel reliability-aware name tagging model to tackle this issue. We design a set of word frequency-based reliability signals to indicate the quality of each word embedding. Guided by the reliability signals, the model is able to dynamically select and compose features such as word embedding and character-level representation using gating mechanisms. For example, if an input word is rare, the model relies less on its word embedding and assigns higher weights to its character and contextual features. Experiments on OntoNotes 5.0 show that our model outperforms the baseline model by 2.7% absolute gain in F-score. In cross-genre experiments on five genres in OntoNotes, our model improves the performance for most genre pairs and obtains up to 5% absolute F-score gain.
Heng Ji 0001, Dong Yu 0001, Jiawei Han 0001
ACL (1)4
2019 Cross-lingual Knowledge Graph Alignment via Graph Matching Neural Network
abstract
Previous cross-lingual knowledge graph (KG) alignment studies rely on entity embeddings derived only from monolingual KG structural information, which may fail at matching entities that have different facts in two KGs. In this paper, we introduce the topic entity graph, a local sub-graph of an entity, to represent entities with their contextual information in KG. From this view, the KB-alignment task can be formulated as a graph matching problem; and we further propose a graph-attention based solution, which first matches all entities in two topic entity graphs, and then jointly model the local matching information to derive a graph-level matching vector. Experiments show that our model outperforms previous state-of-the-art methods by a large margin.
Kun Xu 0005, Liwei Wang 0009, Mo Yu, Yansong Feng 0002, Yan Song 0003, Zhiguo Wang 0006, Dong Yu 0001
ACL (1)7
2019 Knowledge-aware Pronoun Coreference Resolution
abstract
Resolving pronoun coreference requires knowledge support, especially for particular domains (e.g., medicine).In this paper, we explore how to leverage different types of knowledge to better resolve pronoun coreference with a neural model.To ensure the generalization ability of our model, we directly incorporate knowledge in the format of triplets, which is the most common format of modern knowledge graphs, instead of encoding it with features or rules as that in conventional approaches.Moreover, since not all knowledge is helpful in certain contexts, to selectively use them, we propose a knowledge attention module, which learns to select and use informative knowledge based on contexts, to enhance our model.Experimental results on two datasets from different domains prove the validity and effectiveness of our model, where it outperforms state-of-the-art baselines by a large margin.Moreover, since our model learns to use external knowledge rather than only fitting the training data, it also demonstrates superior performance to baselines in the cross-domain setting.
Hongming Zhang 0009, Yan Song 0003, Yangqiu Song, Dong Yu 0001
ACL (1)4
2019 Syllable-Dependent Discriminative Learning for Small Footprint Text-Dependent Speaker Verification
abstract
This study proposes a novel scheme of syllable-dependent discriminative speaker embedding learning for small footprint text-dependent speaker verification systems. To suppress undesired syllable variation and enhance the power of discrimination inherited in the frame-level features, we design a novel syllable-dependent clustering loss to optimize the network. Specifically, this loss function utilizes syllable labels as auxiliary supervision information to explicitly maximize inter-syllable divisibility and intra-syllable compactness between the learned frame-level features. Successively, we propose two syllable-dependent pooling mechanisms to aggregate the frame-level features to several syllable-level features by averaging those features corresponding to each syllable. The utterance-level speaker embeddings with powerful discrimination are then obtained by concatenating the syllable-level features. Experimental results on Tencent voice wake-up dataset show that our proposed scheme can accelerate the network convergence and achieve significant performance improvement against the state-of-the-art methods.
Junyi Peng, Yuexian Zou, Na Li 0012, Deyi Tuo, Dan Su 0002, Meng Yu 0003, Dong Yu 0001
ASRU8
2019 Time Domain Audio Visual Speech Separation
abstract
Audio-visual multi-modal modeling has been demonstrated to be effective in many speech related tasks, such as speech recognition and speech enhancement. This paper introduces a new time-domain audio-visual architecture for target speaker extraction from monaural mixtures. The architecture generalizes the previous TasNet (time-domain speech separation network) to enable multi-modal learning and at meanwhile it extends the classical audio-visual speech separation from frequency-domain to time-domain. The main components of proposed architecture include an audio encoder, a video encoder that extracts lip embedding from video streams, a multi-modal separation network and an audio decoder. Experiments on simulated mixtures based on recently released LRS2 dataset show that our method can bring 3dB+ and 4dB+ Si-SNR improvements on two- and three-speaker cases respectively, compared to audio-only TasNet and frequency-domain audio-visual networks.
Jian Wu 0027, Yong Xu 0004, Shixiong Zhang 0001, Lianwu Chen, Meng Yu 0003, Lei Xie 0001, Dong Yu 0001
ASRU7
2019 Improving Speech Enhancement with Phonetic Embedding Features
abstract
In this paper, we present a speech enhancement framework that leverages phonetic information obtained from the acoustic model. It consists of two separate components: (i) a long short-term memory recurrent neural network (LSTM-RNN) based speech enhancement model that takes the combination of log-power spectra (LPS) and phonetic embedding features as input to predict the complex ideal ratio mask (cIRM); and (ii) a convolutional, long short-term memory and fully connected deep neural network (CLDNN) based acoustic model that extracts the phonetic feature vector in the hidden units of its LSTM layer. Our experimental results show that the proposed framework outperforms both the conventional and phoneme-dependent speech enhancement systems under various noisy conditions, generalizes well to unseen conditions, and performs robustly to the speech interference. We further demonstrate its superior enhancement performance on unvoiced speech and report a preliminary yet promising recognition experiment on real test data.
Bo Wu 0011, Meng Yu 0003, Lianwu Chen, Mingjie Jin, Dan Su 0002, Dong Yu 0001
ASRU6
2019 Improving Pre-Trained Multilingual Model with Vocabulary Expansion
abstract
Recently, pre-trained language models have achieved remarkable success in a broad range of natural language processing tasks.However, in multilingual setting, it is extremely resource-consuming to pre-train a deep language model over large-scale corpora for each language.Instead of exhaustively pre-training monolingual language models independently, an alternative solution is to pre-train a powerful multilingual deep language model over large-scale corpora in hundreds of languages.However, the vocabulary size for each language in such a model is relatively small, especially for low-resource languages.This limitation inevitably hinders the performance of these multilingual models on tasks such as sequence labeling, wherein in-depth token-level or sentence-level understanding is essential.In this paper, inspired by previous methods designed for monolingual settings, we investigate two approaches (i.e., joint mapping and mixture mapping) based on a pre-trained multilingual model BERT for addressing the out-of-vocabulary (OOV) problem on a variety of tasks, including part-of-speech tagging, named entity recognition, machine translation quality estimation, and machine reading comprehension.Experimental results show that using mixture mapping is more promising.To the best of our knowledge, this is the first work that attempts to address and discuss the OOV issue in multilingual settings.
Hai Wang 0013, Dian Yu 0001, Kai Sun 0006, Jianshu Chen, Dong Yu 0001
CoNLL5
2019 Evidence Sentence Extraction for Machine Reading Comprehension
abstract
Remarkable success has been achieved in the last few years on some limited machine reading comprehension (MRC) tasks.However, it is still difficult to interpret the predictions of existing MRC models.In this paper, we focus on extracting evidence sentences that can explain or support the answers of multiplechoice MRC tasks, where the majority of answer options cannot be directly extracted from reference documents.Due to the lack of ground truth evidence sentence labels in most cases, we apply distant supervision to generate imperfect labels and then use them to train an evidence sentence extractor.To denoise the noisy labels, we apply a recently proposed deep probabilistic logic learning framework to incorporate both sentence-level and cross-sentence linguistic indicators for indirect supervision.We feed the extracted evidence sentences into existing MRC models and evaluate the end-to-end performance on three challenging multiplechoice MRC datasets: MultiRC, RACE, and DREAM, achieving comparable or better performance than the same models that take as input the full reference document.To the best of our knowledge, this is the first work extracting evidence sentences for multiple-choice MRC.
Hai Wang 0013, Dian Yu 0001, Kai Sun 0006, Jianshu Chen, Dong Yu 0001, David A. McAllester, Dan Roth 0001
CoNLL5
2019 Multiplex Word Embeddings for Selectional Preference Acquisition
abstract
Hongming Zhang, Jiaxin Bai, Yan Song, Kun Xu, Changlong Yu, Yangqiu Song, Wilfred Ng, Dong Yu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Hongming Zhang 0009, Jiaxin Bai, Yan Song 0003, Kun Xu 0005, Changlong Yu, Yangqiu Song, Wilfred Ng, Dong Yu 0001
EMNLP/IJCNLP (1)8
2019 Multi-band PIT and Model Integration for Improved Multi-channel Speech Separation
abstract
The recent exploration of deep learning for supervised speech separation has significantly accelerated the progress on the multi-talker speech separation problem. Multi-channel extension has attracted much research attention due to the benefit of spatial information in far-field acoustic environments. In this paper, We review the most recent models of multi-channel permutation invariant training (PIT), investigate spatial features formed by microphone pairs and their underlying impact and issue, present a multi-band architecture for effective feature encoding, and conduct a model integration between single-channel and multi-channel PIT for resolving the spatial overlapping problem in the conventional multi-channel PIT framework. The evaluation confirms the significant improvement achieved with the proposed model and training approach for the multi-channel speech separation.
Lianwu Chen, Meng Yu 0003, Dan Su 0002, Dong Yu 0001
ICASSP4
2019 Boundary Discriminative Large Margin Cosine Loss for Text-independent Speaker Verification
abstract
Deep neural network based speaker embeddings have attracted much attention in text-independent speaker verification task. In addition to the network architecture, an appropriate design of the loss function is crucial for the deep discriminative embedding extractor. Inspired by the success of Large Margin Cosine Loss (LMCL) in face recognition, we propose an enhanced LMCL named boundary discriminative LMCL (BD-LMCL) to emphasize the discriminative information inherited in the speaker boundaries. Unlike LMCL, where all training samples contribute equally for the objective function, only the samples around the speaker boundaries are considered during the network training with BD-LMCL. Specifically, those samples close to the boundaries are dynamically selected using top-k zero-one loss. Experimental results on a short duration corpus Android Cellphone and NIST SRE 2012 demonstrate better performance compared to LMCL and other popular loss functions.
Rongjin Li, Na Li 0012, Deyi Tuo, Meng Yu 0003, Dan Su 0002, Dong Yu 0001
ICASSP6
2019 Component Fusion: Learning Replaceable Language Model Component for End-to-end Speech Recognition System
abstract
Recently, attention-based end-to-end automatic speech recognition system (ASR) has shown promising results. One of the limitations of an attention-based ASR system is that its language model (LM) component has to be implicitly learned from transcribed speech data which prevents one from uti-lizing plenty of text corpora to improve language modeling. In this work, the Component Fusion method is proposed to incorporate externally trained neural network (NN) LM into an attention-based ASR system. During training stage we equip the attention-based system with an additional LM component which is replaced by an externally trained NN LM at decoding stage. Experimental results show that the proposed Component Fusion outperforms two prior LM fusion approaches, i.e., Shallow Fusion and Cold Fusion, in both out-of-domain and in-domain scenarios. Further improvements can be achieved when combining Component and Shallow Fusion.
Changhao Shan, Chao Weng, Guangsen Wang, Dan Su 0002, Dong Yu 0001, Lei Xie 0001
ICASSP6
2019 Investigating End-to-end Speech Recognition for Mandarin-english Code-switching
abstract
Code-switching is a common phenomenon in many multilingual communities and presents a challenge to automatic speech recognition (ASR). In this paper, three approaches are investigated to improve end-to-end speech recognition on Mandarin-English code-switching task. First, multi-task learning (MTL) is introduced which enables the language identity information to facilitate Mandarin-English code-switching ASR. Second, we explore wordpieces, as opposed to graphemes, as English modeling units to reduce the mod-eling unit gap between Mandarin and English. Third, we employ transfer learning to utilize larger amount of monolingual Mandarin and English data to compensate the data sparsity issue of a code-switching task. Significant improvements are observed from all three approaches. With all three approaches combined, the final system achieves a character error rate (CER) of 6.49% on a real Mandarin-English code-switching task.
Changhao Shan, Chao Weng, Guangsen Wang, Dan Su 0002, Dong Yu 0001, Lei Xie 0001
ICASSP6
2019 Token-wise Training for Attention Based End-to-end Speech Recognition
abstract
In attention based end-to-end (A-E2E) speech recognition systems, the dependency between output tokens is typically formulated as an input-output mapping in decoder. Due to such dependency, decoding errors can easily propagate along output sequence. In this paper, we propose a token-wise training (TWT) method for A-E2E models. The new method is flexible and can be combined with a variety of loss functions. Applying TWT to multiple hypotheses, we propose a novel TWT in beam (TWTiB) training scheme. Trained on the benchmark Switchboard (SWBD) 300h corpus, TWTiB outperforms the previous best training scheme on the SWBD evaluation subset.
Jia Cui, Chao Weng, Dong Yu 0001
ICASSP4
2019 Learning Discriminative Features in Sequence Training without Requiring Framewise Labelled Data
abstract
In this work, we try to answer two questions: Can deeply learned features with discriminative power benefit an ASR system’s robustness to acoustic variability? And how to learn them without requiring framewise labelled sequence training data? As existing methods usually require knowing where the labels occur in the input sequence, they have so far been limited to many real-world sequence learning tasks. We propose a novel method which simultaneously models both the sequence discriminative training and the feature discriminative learning within a single network architecture, so that it can learn discriminative deep features in sequence training that obviates the need for presegmented training data. Our experiment in a realistic industrial ASR task shows that, without requiring any specific fine-tuning or additional complexity, our proposed models have consistently outperformed state-of-the-art models and significantly reduced Word Error Rate (WER) under all test conditions, and especially with highest improvements under unseen noise conditions, by relative 12.94%, 8.66% and 5.80%, showing our proposed models can generalize better to acoustic variability.
Jun Wang 0091, Dan Su 0002, Jie Chen 0057, Shulin Feng, Dongpeng Ma, Na Li 0012, Dong Yu 0001
ICASSP7
2019 Quasi-fully Convolutional Neural Network with Variational Inference for Speech Synthesis
abstract
Recurrent neural networks, such as gated recurrent units (GRUs) and long short-term memory (LSTM), are widely used on acoustic modeling for speech synthesis. However, such sequential generating processes are not friendly to today’s massively parallel computing devices. We introduce a fully convolutional neural network (CNN) model, which can effiently run on parallel processers, for speech synthesis. To improve the quality of the generated acoustic features, we strengthen our model with variational inference. We also use quasi-recurrent neural networks (QRNNs) to smoothen the generated acoustic features. Finally, a high-quality parallel WaveNet model is used to generate audio samples. Our contributions are twofold. First, we show that CNNs with variational inference can generate highly natural speech on a par with end-to-end models; the use of QRNNs further improves the synthetic quality by reducing trembling of generated acoustic features and introduces very little run-time overheads. Second, we show some techniques to further speed up the sampling process of the parallel WaveNet model.
Xixin Wu, Zhiyong Wu 0001, Shiyin Kang, Deyi Tuo, Guangzhi Li, Dan Su 0002, Dong Yu 0001, Helen M. Meng
ICASSP8
2019 A Comparison of Lattice-free Discriminative Training Criteria for Purely Sequence-trained Neural Network Acoustic Models
abstract
In this work, three lattice-free (LF) discriminative training criteria for purely sequence-trained neural network acoustic models are compared on LVCSR tasks, namely maximum mutual information (MMI), boosted maximum mutual information (bMMI) and state-level minimum Bayes risk (sMBR). We demonstrate that, analogous to LF-MMI, a neural network acoustic model can also be trained from scratch using LF-bMMI or LF-sMBR criteria respectively without the need of cross-entropy pre-training. Furthermore, experimental results on Switchboard-300hrs and Switchboard+Fisher-2100hrs datasets show that models trained with LF-bMMI consistently outperform those trained with plain LF-MMI and achieve a relative word error rate (WER) reduction of ~5% over competitive temporal convolution projected LSTM (TDNN-LSTMP) LF-MMI baselines.
Chao Weng, Dong Yu 0001
ICASSP2
2019 Joint Training of Complex Ratio Mask Based Beamformer and Acoustic Model for Noise Robust Asr
abstract
In this paper, we present a joint training framework between the multi-channel beamformer and the acoustic model for noise robust automatic speech recognition (ASR). The complex ratio mask (CRM), demonstrated to be more effective than the ideal ratio mask (IRM), is proposed to estimate the covariance matrix for the beamformer. Minimum Variance Distortionless Response (MVDR) beamformer and Generalized Eigenvalue (GEV) beamformer are both investigated under the CRM-based joint training architecture. We also propose a robust mask pooling strategy among multiple channels. A long short-term memory (LSTM) based language model is utilized to re-score hypotheses which further improves the overall performance. We evaluate the proposed methods on CHiME-4 challenge dataset. The CRM based system achieves a relative 10% reduction on word error rate (WER) compared with the IRM based system. Without sequence discriminative training, our best single system already achieves an average WER 2.72% on the test set which is comparable to the state-of-the-art.
Yong Xu 0004, Chao Weng, Like Hui, Meng Yu 0003, Dan Su 0002, Dong Yu 0001
ICASSP7
2019 Enhancing Hybrid Self-attention Structure with Relative-position-aware Bias for Speech Synthesis
abstract
Compared with the conventional "front-end"-"back-end"- "vocoder" structure, based on the attention mechanism, end-to-end speech synthesis systems directly train and synthesize from text sequence to the acoustic feature sequence as a whole. Recently, a more calculation efficient end-to-end architecture named transformer, which is solely based on self-attention, was proposed to model global dependencies between the input and output sequences. However, although with many advantages, transformer lacks position information in its structure. Moreover, the weighted sum form in self-attention may disperse the attention to the whole input sequence other than focusing on the more important neighbouring positions. In order to solve the above problems, this paper introduces a hybrid self-attention structure which combines self-attention with the recurrent neural networks (RNNs). We further enhance the proposed structure with relative-position-aware biases. Mean opinion score (MOS) test results indicate that by enhancing hybrid self-attention structure with relative-position-aware biases, the proposed system achieves the best performance with only 0.11 MOS score lower than natural recording.
Shan Yang 0001, Heng Lu 0004, Shiying Kang, Lei Xie 0001, Dong Yu 0001
ICASSP5
2019 Teach an All-rounder with Experts in Different Domains
abstract
In many automatic speech recognition (ASR) tasks, an ideal model has to be applicable over multiple domains. In this paper, we propose to teach an all-rounder with experts in different domains. Concretely, we build a multi-domain acoustic model by applying the teacher-student training framework. First, for each domain, a teacher model (domain-dependent model) is trained by fine-tuning a multi-condition model with domain-specific subset. Then all these teacher models are used to teach one single student model simultaneously. We perform experiments on two predefined domain setups. One is domains with different speaking styles, the other is near-field, far-field and far-field with noise. Moreover, two types of models are examined: deep feedforward sequential memory network (DFSMN) and long short term memory (LSTM). Experimental results show that the model trained with this framework outperforms not only multi-condition model but also domain-dependent model. Specially, our training method provides up to 10.4% relative character error rate improvement over baseline model (multi-condition model).
Zhao You, Dan Su 0002, Dong Yu 0001
ICASSP3
2019 Encrypted Speech Recognition Using Deep Polynomial Networks
abstract
The cloud-based speech recognition/API provides developers or enterprises an easy way to create speech-enabled features in their applications. However, sending audios about personal or company internal information to the cloud, raises concerns about the privacy and security issues. The recognition results generated in cloud may also reveal some sensitive information. This paper proposes a deep polynomial network (DPN) that can be applied to the encrypted speech as an acoustic model. It allows clients to send their data in an encrypted form to the cloud to ensure that their data remains confidential, at mean while the DPN can still make frame-level predictions over the encrypted speech and return them in encrypted form. One good property of the DPN is that it can be trained on unencrypted speech features in the traditional way. To keep the cloud away from the raw audio and recognition results, a cloud-local joint decoding framework is also proposed. We demonstrate the effectiveness of model and framework on the Switchboard and Cortana voice assistant tasks with small performance degradation and latency increased comparing with the traditional cloud-based DNNs.
Shixiong Zhang 0001, Yifan Gong 0001, Dong Yu 0001
ICASSP3
2019 Seq2Seq Attentional Siamese Neural Networks for Text-dependent Speaker Verification
abstract
In this paper, we present a Sequence-to-Sequence Attentional Siamese Neural Network (Seq2Seq-ASNN) that leverages temporal alignment information for end-to-end speaker verification. In prior works of speaker discriminative neural networks, utterance-level evaluation/enrollment speaker representations are usually calculated. Our proposed model, utilizing a sequence-to-sequence (Seq2Seq) attention mechanism, maps the frame-level evaluation representation into enrollment feature domain and further generates an utterance-level evaluation-enrollment joint vector for final similarity measure. Feature learning, attention mechanism, and metric learning are jointly optimized using an end-to-end loss function. Experimental results show that our proposed model outperforms various baseline methods, including the traditional i-Vector/PLDA method, multi-enrollment end-to-end speaker verification models, d-vector approaches, and a self attention model, for text-dependent speaker verification on a Tencent internal voice wake-up dataset.
Meng Yu 0003, Na Li 0012, Chengzhu Yu, Jia Cui, Dong Yu 0001
ICASSP6
2019 A Fast and Accurate One-Stage Approach to Visual Grounding
abstract
We propose a simple, fast, and accurate one-stage approach to visual grounding, inspired by the following insight. The performances of existing propose-and-rank two-stage methods are capped by the quality of the region candidates they propose in the first stage - if none of the candidates could cover the ground truth region, there is no hope in the second stage to rank the right region to the top. To avoid this caveat, we propose a one-stage model that enables end-to-end joint optimization. The main idea is as straightforward as fusing a text query's embedding into the YOLOv3 object detector, augmented by spatial features so as to account for spatial mentions in the query. Despite being simple, this one-stage approach shows great potential in terms of both accuracy and speed for both phrase localization and referring expression comprehension, according to our experiments. Given these results along with careful investigations into some popular region proposals, we advocate for visual grounding a paradigm shift from the conventional two-stage methods to the one-stage framework.
Zhengyuan Yang, Boqing Gong, Liwei Wang 0009, Wenbing Huang 0001, Dong Yu 0001, Jiebo Luo 0001
ICCV5
2019 Unsupervised Speech Recognition via Segmental Empirical Output Distribution Matching
Chih-Kuan Yeh, Jianshu Chen, Chengzhu Yu, Dong Yu 0001
ICLR (Poster)4
2019 Unsupervised Neural Aspect Extraction with Sememes
abstract
Aspect extraction relies on identifying aspects by discovering coherence among words, which is challenging when word meanings are diversified and processing on short texts. To enhance the performance on aspect extraction, leveraging lexical semantic resources is a possible solution to such challenge. In this paper, we present an unsupervised neural framework that leverages sememes to enhance lexical semantics. The overall framework is analogous to an autoenoder which reconstructs sentence representations and learns aspects by latent variables. Two models that form sentence representations are proposed by exploiting sememes via (1) a hierarchical attention; (2) a context-enhanced attention. Experiments on two real-world datasets demonstrate the validity and the effectiveness of our models, which significantly outperforms existing baselines.
Xiang Ao 0001, Yan Song 0003, Jinyao Li, Qing He 0003, Dong Yu 0001
IJCAI7
2019 A Comprehensive Study of Speech Separation: Spectrogram vs Waveform Separation
abstract
Speech separation has been studied widely for single-channel close-talk microphone recordings over the past few years; developed solutions are mostly in frequency-domain.Recently, a raw audio waveform separation network (TasNet) is introduced for single-channel data, with achieving high Si-SNR (scale-invariant source-to-noise ratio) and SDR (sourceto-distortion ratio) comparing against the state-of-the-art solution in frequency-domain.In this study, we incorporate effective components of the TasNet into a frequency-domain separation method.We compare both for alternative scenarios.We introduce a solution for directly optimizing the separation criterion in frequency-domain networks.In addition to speech separation objective and subjective measurements, we evaluate the separation performance on a speech recognition task as well.We study the speech separation problem for far-field data (more similar to naturalistic audio streams) and develop multi-channel solutions for both frequency and time-domain separators with utilizing spectral, spatial and speaker location information.For our experiments, we simulated multi-channel spatialized reverberate WSJ0-2mix dataset.Our experimental results show that spectrogram separation can achieve competitive performance with better network design.Multi-channel framework as well is shown to improve the single-channel performance relatively up to +35.5% and +46% in terms of WER and SDR, respectively.
Fahimeh Bahmaninezhad, Jian Wu 0027, Rongzhi Gu, Shixiong Zhang 0001, Yong Xu 0004, Meng Yu 0003, Dong Yu 0001
INTERSPEECH7
2019 Disambiguation of Chinese Polyphones in an End-to-End Framework with Semantic Features Extracted by Pre-Trained BERT
Dongyang Dai, Zhiyong Wu 0001, Shiyin Kang, Xixin Wu, Jia Jia 0001, Dan Su 0002, Dong Yu 0001, Helen M. Meng
INTERSPEECH7
2019 Neural Spatial Filter: Target Speaker Speech Separation Assisted with Directional Information
Rongzhi Gu, Lianwu Chen, Shixiong Zhang 0001, Jimeng Zheng, Yong Xu 0004, Meng Yu 0003, Dan Su 0002, Yuexian Zou, Dong Yu 0001
INTERSPEECH9
2019 Extract, Adapt and Recognize: An End-to-End Neural Network for Corrupted Monaural Speech Recognition
Max W. Y. Lam, Jun Wang 0091, Xunying Liu, Helen M. Meng, Dan Su 0002, Dong Yu 0001
INTERSPEECH6
2019 Large Margin Training for Attention Based End-to-End Speech Recognition
abstract
A method of attention-based end-to-end (E2E) automatic speech recognition (ASR) training, includes performing cross-entropy training of a model, based on one or more input features of a speech signal, performing beam searching of the model of which the cross-entropy training is performed, to generate an n-best hypotheses list of output hypotheses, and determining a one-best hypothesis among the generated n-best hypotheses list. The method further includes determining a character-based gradient and a word-based gradient, based on the model of which the cross-entropy training is performed and a loss function in which a distance between a reference sequence and the determined one-best hypothesis is maximized, and performing backpropagation of the determined character-based gradient and the determined word-based gradient to the model, to update the model.
Jia Cui, Chao Weng, Dong Yu 0001
INTERSPEECH4
2019 Improved Speaker-Dependent Separation for CHiME-5 Challenge
abstract
This paper summarizes several follow-up contributions for improving our submitted NWPU speaker-dependent system for CHiME-5 challenge, which aims to solve the problem of multi-channel, highly-overlapped conversational speech recognition in a dinner party scenario with reverberations and nonstationary noises.We adopt a speaker-aware training method by using i-vector as the target speaker information for multi-talker speech separation.With only one unified separation model for all speakers, we achieve a 10% absolute improvement in terms of word error rate (WER) over the previous baseline of 80.28% on the development set by leveraging our newly proposed data processing techniques and beamforming approach.With our improved back-end acoustic model, we further reduce WER to 60.15% which surpasses the result of our submitted CHiME-5 challenge system without applying any fusion techniques.
Jian Wu 0027, Yong Xu 0004, Shixiong Zhang 0001, Lianwu Chen, Meng Yu 0003, Lei Xie 0001, Dong Yu 0001
INTERSPEECH7
2019 Erratum to: Past review, current progress, and challenges ahead on the cocktail party problem
abstract
In the original version of this article, there is a mistake about the result of DPCL++ (Isik et al., 2016) in Section 5.6 (Fig. 7). As reported in Isik et al. (2016), the SDR improvement was 10.3 dB, rather than 9.4 dB. For further information, the best performance in Isik et al. (2016) was 10.8 dB with the help of a more complicated architecture.
Yanmin Qian, Chao Weng, Xuankai Chang, Shuai Wang 0016, Dong Yu 0001
Frontiers Inf. Technol. Electron. Eng.5
2019 DREAM: A Challenge Dataset and Models for Dialogue-Based Reading Comprehension
abstract
We present DREAM, the first dialogue-based multiple-choice reading comprehension data set. Collected from English as a Foreign Language examinations designed by human experts to evaluate the comprehension level of Chinese learners of English, our data set contains 10,197 multiple-choice questions for 6,444 dialogues. In contrast to existing reading comprehension data sets, DREAM is the first to focus on in-depth multi-turn multi-party dialogue understanding. DREAM is likely to present significant challenges for existing reading comprehension systems: 84% of answers are non-extractive, 85% of questions require reasoning beyond a single sentence, and 34% of questions also involve commonsense knowledge. We apply several popular neural reading comprehension models that primarily exploit surface information within the text and find them to, at best, just barely outperform a rule-based approach. We next investigate the effects of incorporating dialogue structure and different kinds of general world knowledge into both rule-based and (neural and non-neural) machine learning-based reading comprehension models. Experimental results on the DREAM data set show the effectiveness of dialogue structure and general world knowledge. DREAM is available at https://dataset.org/dream/ .
Kai Sun 0006, Dian Yu 0001, Jianshu Chen, Dong Yu 0001, Yejin Choi 0001, Claire Cardie
Trans. Assoc. Comput. Linguistics4
2018 XL-NBT: A Cross-lingual Neural Belief Tracking Framework
abstract
Task-oriented dialog systems are becoming pervasive, and many companies heavily rely on them to complement human agents for customer service in call centers.With globalization, the need for providing cross-lingual customer support becomes more urgent than ever.However, cross-lingual support poses great challenges-it requires a large amount of additional annotated data from native speakers.In order to bypass the expensive human annotation and achieve the first step towards the ultimate goal of building a universal dialog system, we set out to build a cross-lingual state tracking framework.Specifically, we assume that there exists a source language with dialog belief tracking annotations while the target languages have no annotated dialog data of any form.Then, we pre-train a state tracker for the source language as a teacher, which is able to exploit easy-to-access parallel data.We then distill and transfer its own knowledge to the student state tracker in target languages.We specifically discuss two types of common parallel resources: bilingual corpus and bilingual dictionary, and design different transfer learning strategies accordingly.Experimentally, we successfully use English state tracker as the teacher to transfer its knowledge to both Italian and German trackers and achieve promising results.
Wenhu Chen, Jianshu Chen, Yu Su 0001, Xin Wang 0061, Dong Yu 0001, Xifeng Yan, William Yang Wang
EMNLP5
2018 Knowledge Transfer in Permutation Invariant Training for Single-Channel Multi-Talker Speech Recognition
abstract
This paper proposes a framework that combines teacher-student training and permutation invariant training (PIT) for single-channel multi-talker speech recognition. In contrast to most of conventional teacher-student training methods that aim at compressing the model, the proposed method distills knowledge from the single-talker model to improve the multi-talker model in the PIT framework. The inputs to the teacher and student networks are the single-talker clean speech and the multi-talker mixed speech, respectively. The knowledge is transferred to the student through the soft labels generated by the teacher. Furthermore, the ensemble of multiple teachers is exploited with a progressive training scheme to further improve the system. In this framework it is easy to take advantage of data augmentation and perform domain adaptation for multi-talker speech recognition using only untranscribed data. The proposed techniques were evaluated on artificially mixed two-talker AMI speech data. The experimental results show that the teacher-student training can cut the word error rate (WER) by relative 20% against the baseline PIT model. We also evaluated our unsupervised domain adaptation method on an artificially mixed WSJO corpus and achieved relative 30% WER reduction against the AMI PIT model.
Tian Tan 0002, Yanmin Qian, Dong Yu 0001
ICASSP3
2018 Adaptive Permutation Invariant Training with Auxiliary Information for Monaural Multi-Talker Speech Recognition
abstract
In this paper, we extend our previous work on direct recognition of single-channel multi-talker mixed speech using permutation invariant training (PIT). We propose to adapt the PIT models with auxiliary features such as pitch and i-vector, and to exploit the gender information with multi-task learning which jointly optimizes for the speech recognition and speaker-pair prediction. We also compare CNN-BLSTMs against BLSTM-RNNs used in our previous PIT-ASR model. The experimental results on the artificially mixed two-talker AMI data indicate that our proposed model improvements can reduce word error rate (WER) by ~ 10.0% relative to our previous work for both speakers in the mixed speech. Our results also confirm that PIT can be easily combined with advanced techniques to improve the performance on multi-talker speech recognition.
Xuankai Chang, Yanmin Qian, Dong Yu 0001
ICASSP3
2018 Monaural Multi-Talker Speech Recognition with Attention Mechanism and Gated Convolutional Networks
abstract
Provided are a speech recognition training processing method and an apparatus including the same. The speech recognition training processing method includes acquiring multi-talker mixed speech sequence data corresponding to a plurality of speakers, encoding the multi-speaker mixed speech sequence data into an embedded sequence data, generating speaker specific context vectors at each frame based on the embedded sequence, generating senone posteriors for each of the speaker based on the speaker specific context vectors and updating an acoustic model by performing permutation invariant training (PIT) model training based on the senone posteriors.
Xuankai Chang, Yanmin Qian, Dong Yu 0001
INTERSPEECH3
2018 Permutation Invariant Training of Generative Adversarial Network for Monaural Speech Separation
Lianwu Chen, Meng Yu 0003, Yanmin Qian, Dan Su 0002, Dong Yu 0001
INTERSPEECH5
2018 Deep Discriminative Embeddings for Duration Robust Speaker Verification
Na Li 0012, Deyi Tuo, Dan Su 0002, Zhifeng Li 0001, Dong Yu 0001
INTERSPEECH5
2018 Deep Extractor Network for Target Speaker Recovery from Single Channel Speech Mixtures
abstract
Speaker-aware source separation methods are promising workarounds for major difficulties such as arbitrary source permutation and unknown number of sources.However, it remains challenging to achieve satisfying performance provided a very short available target speaker utterance (anchor).Here we present a novel "deep extractor network" which creates an extractor point for the target speaker in a canonical high dimensional embedding space, and pulls together the time-frequency bins corresponding to the target speaker.The proposed model is different from prior works in that the canonical embedding space encodes knowledges of both the anchor and the mixture during an end-to-end training phase: First, embeddings for the anchor and mixture speech are separately constructed in a primary embedding space, and then combined as an input to feed-forward layers to transform to a canonical embedding space which we discover more stable than the primary one.Experimental results show that given a very short utterance, the proposed model can efficiently recover high quality target speech from a mixture, which outperforms various baseline models, with 5.2% and 6.6% relative improvements in SDR and PESQ respectively compared with a baseline oracle deep attracor model.Meanwhile, we show it can be generalized well to more than one interfering speaker.
Jun Wang 0091, Jie Chen 0057, Dan Su 0002, Lianwu Chen, Meng Yu 0003, Yanmin Qian, Dong Yu 0001
INTERSPEECH7
2018 Improving Attention Based Sequence-to-Sequence Models for End-to-End English Conversational Speech Recognition
Chao Weng, Jia Cui, Guangsen Wang, Jun Wang 0091, Chengzhu Yu, Dan Su 0002, Dong Yu 0001
INTERSPEECH7
2018 Rapid Style Adaptation Using Residual Error Embedding for Expressive Speech Synthesis
Xixin Wu, Yuewen Cao, Songxiang Liu, Shiyin Kang, Zhiyong Wu 0001, Xunying Liu, Dan Su 0002, Dong Yu 0001, Helen M. Meng
INTERSPEECH9
2018 Text-Dependent Speech Enhancement for Small-Footprint Robust Keyword Detection
Meng Yu 0003, Lianwu Chen, Jie Chen 0057, Jimeng Zheng, Dan Su 0002, Dong Yu 0001
INTERSPEECH8
2018 A Multistage Training Framework for Acoustic-to-Word Model
Chengzhu Yu, Chao Weng, Jia Cui, Dong Yu 0001
INTERSPEECH5
2018 Improving Attention-Based End-to-End ASR Systems with Sequence-Based Loss Functions
abstract
Acoustic model and language model (LM) have been two major components in conventional speech recognition systems. They are normally trained independently, but recently there has been a trend to optimize both components simultaneously in a unified end-to-end (E2E) framework. However, the performance gap between the E2E systems and the traditional hybrid systems suggests that some knowledge has not yet been fully utilized in the new framework. An observation is that the current attention-based E2E systems could produce better recognition results when decoded with LMs which are independently trained with the same resource. In this paper, we focus on how to improve attention-based E2E systems without increasing model complexity or resorting to extra data. A novel training strategy is proposed for multi-task training with the connectionist temporal classification (CTC) loss. The sequence-based minimum Bayes risk (MBR) loss is also investigated. Our experiments on SWB 300hrs showed that both loss functions could significantly improve the baseline model performance. The additional gain from joint-LM decoding remains the same for CTC trained model but is only marginal for MBR trained model. This implies that while CTC loss function is able to capture more acoustic knowledge, MBR loss function exploits more word/character dependency.
Jia Cui, Chao Weng, Guangsen Wang, Jun Wang 0091, Chengzhu Yu, Dan Su 0002, Dong Yu 0001
SLT8
2018 An Exploration of Directly Using Word as ACOUSTIC Modeling Unit for Speech Recognition
abstract
Conventional acoustic models for automatic speech recognition (ASR) are usually constructed from sub-word unit (e.g., context-dependent phoneme, grapheme, wordpiece etc.). Recent studies demonstrate that connectionist temporal classification (CTC) based acoustic-to-word (A2W) models are also promising for ASR. Such structures have drawn increasing attention as they can directly target words as output units, which simplify ASR pipeline by avoiding additional pronunciation lexicon, or even language model. In this study, we systematically explore to use word as acoustic modeling unit for conversational speech recognition. By replacing senone alignment with word alignment in a convolutional bidirectional LSTM architecture and employing a lexicon-free weighted finite-state transducer (WFST) based decoding, we greatly simplify conventional hybrid speech recognition system. On Hub5-2000 Switchboard/CallHome test sets with 300-hour training data, we achieve a WER that is close to the senone based hybrid systems with a WFST based decoding.
Chengzhu Yu, Chao Weng, Jia Cui, Dong Yu 0001
SLT5
2018 Past review, current progress, and challenges ahead on the cocktail party problem
abstract
The cocktail party problem, i.e., tracing and recognizing the speech of a specific speaker when multiple speakers talk simultaneously, is one of the critical problems yet to be solved to enable the wide application of automatic speech recognition (ASR) systems. In this overview paper, we review the techniques proposed in the last two decades in attacking this problem. We focus our discussions on the speech separation problem given its central role in the cocktail party environment, and describe the conventional single-channel techniques such as computational auditory scene analysis (CASA), non-negative matrix factorization (NMF) and generative models, the conventional multi-channel techniques such as beamforming and multi-channel blind source separation, and the newly developed deep learning-based techniques, such as deep clustering (DPCL), the deep attractor network (DANet), and permutation invariant training (PIT). We also present techniques developed to improve ASR accuracy and speaker identification in the cocktail party environment. We argue effectively exploiting information in the microphone array, the acoustic training set, and the language itself using a more powerful model. Better optimization objective and techniques will be the approach to solving the cocktail party problem.
Yanmin Qian, Chao Weng, Xuankai Chang, Shuai Wang 0016, Dong Yu 0001
Frontiers Inf. Technol. Electron. Eng.5
2018 Erratum to: Past review, current progress, and challenges ahead on the cocktail party problem
abstract
In the original version of this article, the affiliations are incorrect. The correct affiliations are given above. The corresponding author’s E-mail address should be [email protected].
Yanmin Qian, Chao Weng, Xuankai Chang, Shuai Wang 0016, Dong Yu 0001
Frontiers Inf. Technol. Electron. Eng.5
2018 Single-channel multi-talker speech recognition with permutation invariant training
Yanmin Qian, Xuankai Chang, Dong Yu 0001
Speech Commun.3
2017 The microsoft 2016 conversational speech recognition system
abstract
We describe Microsoft's conversational speech recognition system, in which we combine recent developments in neural-network-based acoustic and language modeling to advance the state of the art on the Switchboard recognition task. Inspired by machine learning ensemble techniques, the system uses a range of convolutional and recurrent neural networks. I-vector modeling and lattice-free MMI training provide significant gains for all acoustic model architectures. Language model rescoring with multiple forward and backward running RNNLMs, and word posterior-based system combination provide a 20% boost. The best single system uses a ResNet architecture acoustic model with RNNLM rescoring, and achieves a word error rate of 6.9% on the NIST 2000 Switchboard task. The combined system has an error rate of 6.2%, representing an improvement over previously reported results on this benchmark task.
Wayne Xiong, Jasha Droppo, Xuedong Huang 0001, Frank Seide, Mike Seltzer, Andreas Stolcke, Dong Yu 0001, Geoffrey Zweig
ICASSP7
2017 Permutation invariant training of deep models for speaker-independent multi-talker speech separation
abstract
We propose a novel deep learning training criterion, named permutation invariant training (PIT), for speaker independent multi-talker speech separation, commonly known as the cocktail-party problem. Different from the multi-class regression technique and the deep clustering (DPCL) technique, our novel approach minimizes the separation error directly. This strategy effectively solves the long-lasting label permutation problem, that has prevented progress on deep learning based techniques for speech separation. We evaluated PIT on the WSJ0 and Danish mixed-speech separation tasks and found that it compares favorably to non-negative matrix factorization (NMF), computational auditory scene analysis (CASA), and DPCL and generalizes well over unseen speakers and languages. Since PIT is simple to implement and can be easily integrated and combined with other advanced techniques, we believe improvements built upon PIT can eventually solve the cocktail-party problem.
Dong Yu 0001, Morten Kolbæk, Zheng-Hua Tan, Jesper Jensen 0001
ICASSP1
2017 Empirical Evaluation of Parallel Training Algorithms on Acoustic Modeling
abstract
Deep learning models (DLMs) are state-of-the-art techniques in speech recognition.However, training good DLMs can be time consuming especially for production-size models and corpora.Although several parallel training algorithms have been proposed to improve training efficiency, there is no clear guidance on which one to choose for the task in hand due to lack of systematic and fair comparison among them.In this paper we aim at filling this gap by comparing four popular parallel training algorithms in speech recognition, namely asynchronous stochastic gradient descent (ASGD), blockwise model-update filtering (BMUF), bulk synchronous parallel (BSP) and elastic averaging stochastic gradient descent (EASGD), on 1000-hour LibriSpeech corpora using feed-forward deep neural networks (DNNs) and convolutional, long short-term memory, DNNs (CLDNNs).Based on our experiments, we recommend using BMUF as the top choice to train acoustic models since it is most stable, scales well with number of GPUs, can achieve reproducible results, and in many cases even outperforms single-GPU SGD.ASGD can be used as a substitute in some cases.
Wenpeng Li, Lei Xie 0001, Dong Yu 0001
INTERSPEECH4
2017 Recognizing Multi-Talker Speech with Permutation Invariant Training
abstract
In this paper, we propose a novel technique for direct recognition of multiple speech streams given the single channel of mixed speech, without first separating them.Our technique is based on permutation invariant training (PIT) for automatic speech recognition (ASR).In PIT-ASR, we compute the average cross entropy (CE) over all frames in the whole utterance for each possible output-target assignment, pick the one with the minimum CE, and optimize for that assignment.PIT-ASR forces all the frames of the same speaker to be aligned with the same output layer.This strategy elegantly solves the label permutation problem and speaker tracing problem in one shot.Our experiments on artificially mixed AMI data showed that the proposed approach is very promising.
Dong Yu 0001, Xuankai Chang, Yanmin Qian
INTERSPEECH1
2017 Multitalker Speech Separation With Utterance-Level Permutation Invariant Training of Deep Recurrent Neural Networks
abstract
In this paper, we propose the utterance-level permutation invariant training (uPIT) technique. uPIT is a practically applicable, end-to-end, deep-learning-based solution for speaker independent multitalker speech separation. Specifically, uPIT extends the recently proposed permutation invariant training (PIT) technique with an utterance-level cost function, hence eliminating the need for solving an additional permutation problem during inference, which is otherwise required by frame-level PIT. We achieve this using recurrent neural networks (RNNs) that, during training, minimize the utterance-level separation error, hence forcing separated frames belonging to the same speaker to be aligned to the same output stream. In practice, this allows RNNs, trained with uPIT, to separate multitalker mixed speech without any prior knowledge of signal duration, number of speakers, speaker identity, or gender. We evaluated uPIT on the WSJ0 and Danish two- and three-talker mixed-speech separation tasks and found that uPIT outperforms techniques based on nonnegative matrix factorization and computational auditory scene analysis, and compares favorably with deep clustering, and the deep attractor network. Furthermore, we found that models trained with uPIT generalize well to unseen speakers and languages. Finally, we found that a single model, trained with uPIT, can handle both two-speaker, and three-speaker speech mixtures.
Morten Kolbæk, Dong Yu 0001, Zheng-Hua Tan, Jesper Jensen 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Toward Human Parity in Conversational Speech Recognition
abstract
Conversational speech recognition has served as a flagship speech recognition task since the release of the Switchboard corpus in the 1990s. In this paper, we measure a human error rate on the widely used NIST 2000 test set for commercial bulk transcription. The error rate of professional transcribers is 5.9% for the Switchboard portion of the data, in which newly acquainted pairs of people discuss an assigned topic, and 11.3% for the CallHome portion, where friends and family members have open-ended conversations. In both cases, our automated system edges past the human benchmark, achieving error rates of 5.8% and 11.0%, respectively. The key to our system's performance is the use of various convolutional and long-short-term memory acoustic model architectures, combined with a novel spatial smoothing method and lattice-free discriminative acoustic training, multiple recurrent neural network language modeling approaches, and a systematic use of system combination. Comparing frequent errors in our human and machine transcripts, we find them to be remarkably similar, and highly correlated as a function of the speaker. Human subjects find it very difficult to tell which errorful transcriptions come from humans. Overall, this suggests that, given sufficient matched training data, conversational speech transcription engines are approximating human parity in both quantitative and qualitative terms.
Wayne Xiong, Jasha Droppo, Xuedong Huang 0001, Frank Seide, Michael L. Seltzer, Andreas Stolcke, Dong Yu 0001, Geoffrey Zweig
IEEE ACM Trans. Audio Speech Lang. Process.7
2016 An investigation into using parallel data for far-field speech recognition
abstract
Far-field speech recognition is an important yet challenging task due to low signal to noise ratio. In this paper, three novel deep neural network architectures are explored to improve the far-field speech recognition accuracy by exploiting the parallel far-field and close-talk recordings. All three novel architectures use multi-task learning for the model optimization but focus on three different ideas: dereverberation and recognition joint-learning, close-talk and far-field model knowledge sharing, and environment-code aware training. Experiments on the AMI single distant microphone (SDM) task show that each of the proposed method can boost accuracy individually, and additional improvement can be obtained with appropriate integration of these models. Overall we reduced the error rate by 10% relatively on the SDM set by exploiting the IHM data.
Yanmin Qian, Tian Tan 0002, Dong Yu 0001
ICASSP3
2016 Integrated adaptation with multi-factor joint-learning for far-field speech recognition
abstract
Although great progress has been made in automatic speech recognition (ASR), significant performance degradation still exists in distant talking scenarios due to significantly lower signal power. In this paper, a novel adaptation framework, named integrated adaptation with multi-factor joint-learning, is proposed to improve the recognition accuracy for distant speech recognition. We explore and extract speaker, phone and environment factor representations using deep neural networks (DNNs), which are integrated into the main ASR DNN to improve classification accuracy. In addition, the hidden activations in the main ASR DNN are used to improve the factor extraction, which in turn helps the ASR DNN. All the model parameters, including those in the ASR DNN and factor extractor DNNs, are jointly optimized under the multi-task learning framework. Further more, unlike prior techniques, our novel approach requires no explicit separate stages for factor extraction and adaptation. Experiments on the AMI single distant microphone (SDM) task show that the proposed architecture can significantly reduce word error rate (WER) and additional improvement can be achieved by combining it with the i-vector adaptation. Our best configuration obtained more than 15% and 10% relative reduction on WER over the baselines using the SDM and close-talk data generated alignments, respectively.
Yanmin Qian, Tian Tan 0002, Dong Yu 0001, Yu Zhang 0033
ICASSP3
2016 Speaker-aware training of LSTM-RNNS for acoustic modelling
abstract
Long Short-Term Memory (LSTM) is a particular type of recurrent neural network (RNN) that can model long term temporal dynamics. Recently it has been shown that LSTM-RNNs can achieve higher recognition accuracy than deep feed-forword neural networks (DNNs) in acoustic modelling. However, speaker adaption for LSTM-RNN based acoustic models has not been well investigated. In this paper, we study the LSTM-RNN speaker-aware training that incorporates the speaker information during model training to normalise the speaker variability. We first present several speaker-aware training architectures, and then empirically evaluate three types of speaker representation: I-vectors, bottleneck speaker vectors and speaking rate. Furthermore, to factorize the variability in the acoustic signals caused by speakers and phonemes respectively, we investigate the speaker-aware and phone-aware joint training under the framework of multi-task learning. In AMI meeting speech transcription task, speaker-aware training of LSTM-RNNs reduces word error rates by 6.5% relative to a very strong LSTM-RNN baseline, which uses FMLLR features.
Tian Tan 0002, Yanmin Qian, Dong Yu 0001, Souvik Kundu 0003, Liang Lu 0001, Khe Chai Sim, Yu Zhang 0033
ICASSP3
2016 Deep beamforming networks for multi-channel speech recognition
abstract
Despite the significant progress in speech recognition enabled by deep neural networks, poor performance persists in some scenarios. In this work, we focus on far-field speech recognition which remains challenging due to high levels of noise and reverberation in the captured speech signals. We propose to represent the stages of acoustic processing including beamforming, feature extraction, and acoustic modeling, as three components of a single unified computational network. The parameters of a frequency-domain beam-former are first estimated by a network based on features derived from the microphone channels. These filter coefficients are then applied to the array signals to form an enhanced signal. Conventional features are then extracted from this signal and passed to a second network that performs acoustic modeling for classification. The parameters of both the beamforming and acoustic modeling networks are trained jointly using back-propagation with a common cross-entropy objective function. In experiments on the AMI meeting corpus, we observed improvements by pre-training each sub-network with a network-specific objective function before joint training of both networks. The proposed method obtained a 3.2% absolute word error rate reduction compared to a conventional pipeline of independent processing stages.
Shinji Watanabe 0001, Hakan Erdogan, Liang Lu 0001, John R. Hershey, Michael L. Seltzer, Guoguo Chen, Yu Zhang 0033, Michael I. Mandel, Dong Yu 0001
ICASSP10
2016 Prediction-adaptation-correction recurrent neural networks for low-resource language speech recognition
abstract
In this paper, we investigate the use of prediction-adaptation-correction recurrent neural networks (PAC-RNNs) for low-resource speech recognition. A PAC-RNN is comprised of a pair of neural networks in which a correction network uses auxiliary information given by a prediction network to help estimate the state probability. The information from the correction network is also used by the prediction network in a recurrent loop. Our model outperforms other state-of-the-art neural networks (DNNs, LSTMs) on IARPA-Babel tasks. Moreover, transfer learning from a language that is similar to the target language can help improve performance further.
Yu Zhang 0033, Ekapol Chuangsuwanich, James R. Glass, Dong Yu 0001
ICASSP4
2016 Highway long short-term memory RNNS for distant speech recognition
abstract
In this paper, we extend the deep long short-term memory (DL-STM) recurrent neural networks by introducing gated direct connections between memory cells in adjacent layers. These direct links, called highway connections, enable unimpeded information flow across different layers and thus alleviate the gradient vanishing problem when building deeper LSTMs. We further introduce the latency-controlled bidirectional LSTMs (BLSTMs) which can exploit the whole history while keeping the latency under control. Efficient algorithms are proposed to train these novel networks using both frame and sequence discriminative criteria. Experiments on the AMI distant speech recognition (DSR) task indicate that we can train deeper LSTMs and achieve better improvement from sequence training with highway LSTMs (HLSTMs). Our novel model obtains 43.9/47.7% WER on AMI (SDM) dev and eval sets, outperforming all previous works. It beats the strong DNN and DLSTM baselines with 15.7% and 5.3% relative improvement respectively.
Yu Zhang 0033, Guoguo Chen, Dong Yu 0001, Kaisheng Yao, Sanjeev Khudanpur, James R. Glass
ICASSP3
2016 Deep Convolutional Neural Networks with Layer-Wise Context Expansion and Attention
abstract
In this paper, we propose a deep convolutional neural network (CNN) with layer-wise context expansion and location-based attention, for large vocabulary speech recognition. In our model each higher layer uses information from broader contexts, along both the time and frequency dimensions, than its immediate lower layer. We show that both the layer-wise context expansion and the location-based attention can be implemented using the element-wise matrix product and the convolution operation. For this reason, contrary to other CNNs, no pooling operation is used in our model. Experiments on the 309hr Switchboard task and the 375hr short message dictation task indicates that our model outperforms both the DNN and LSTM significantly.
Dong Yu 0001, Wayne Xiong, Jasha Droppo, Andreas Stolcke, Guoli Ye, Jinyu Li 0001, Geoffrey Zweig
INTERSPEECH1
2016 Recurrent Support Vector Machines For Slot Tagging In Spoken Language Understanding
abstract
Yangyang Shi, Kaisheng Yao, Hu Chen, Dong Yu, Yi-Cheng Pan, Mei-Yuh Hwang. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Yangyang Shi, Kaisheng Yao, Dong Yu 0001, Yi-Cheng Pan, Mei-Yuh Hwang
HLT-NAACL4
2016 Neural Network Based Multi-Factor Aware Joint Training for Robust Speech Recognition
abstract
Although great progress has been made in automatic speech recognition (ASR), significant performance degradation still exists in noisy environments. In this paper, a novel factor-aware training framework, named neural network-based multifactor aware joint training, is proposed to improve the recognition accuracy for noise robust speech recognition. This approach is a structured model which integrates several different functional modules into one computational deep model. We explore and extract speaker, phone, and environment factor representations using deep neural networks (DNNs), which are integrated into the main ASR DNN to improve classification accuracy. In addition, the hidden activations in the main ASR DNN are used to improve factor extraction, which in turn helps the ASR DNN. All the model parameters, including those in the ASR DNN and factor extraction DNNs, are jointly optimized under the multitask learning framework. Unlike prior traditional techniques for the factor-aware training, our approach requires no explicit separate stages for factor extraction and adaptation. Moreover, the proposed neural network-based multifactor aware joint training can be easily combined with the conventional factor-aware training which uses the explicit factors, such as i-vector, noise energy, and T60 value to obtain additional improvement. The proposed method is evaluated on two main noise robust tasks: the AMI single distant microphone task in which reverberation is the main concern, and the Aurora4 task in which multiple noise types exist. Experiments on both tasks show that the proposed model can significantly reduce word error rate (WER). The best configuration achieved more than 15% relative reduction in WER over the baselines on these two tasks.
Yanmin Qian, Tian Tan 0002, Dong Yu 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 Deep bi-directional recurrent networks over spectral windows
abstract
Long short-term memory (LSTM) acoustic models have recently achieved state-of-the-art results on speech recognition tasks. As a type of recurrent neural network, LSTMs potentially have the ability to model long-span phenomena relating the spectral input to linguistic units. However, it has not been clear whether their observed performance is actually due to this capability, or instead if it is due to a better modeling of short term dynamics through the recurrence. In this paper. we answer this question by applying a windowed (truncated) LSTM to conversational speech transcription, and find that a limited context is adequate, and that it is not necessaary to scan the entire utterance. The sliding window approach allows not only incremental (online) recognition with a bidirectional model, but also frame-wise randomization (as opposed to utterance randomization), which results in faster convergence. On the SWBD/Fisher corpus, applying bidirectional LSTM RNNs to spectral windows of about 0.5s improves WER on the Hub5'00 benchmark set by 16% relative compared to our best sequence-trained DNN. On an extended 3850h training set that that also includes lectures, the relative gain becomes 28% (Hub5'00 WER 9.2%). In-house conversational data improves by 12 to 17% relative.
Abdel-rahman Mohamed, Frank Seide, Dong Yu 0001, Jasha Droppo, Andreas Stolcke, Geoffrey Zweig, Gerald Penn
ASRU3
2015 Improving speech recognition in reverberation using a room-aware deep neural network and multi-task learning
abstract
In this paper, we propose two approaches to improve deep neural network (DNN) acoustic models for speech recognition in reverberant environments. Both methods utilize auxiliary information in training the DNN but differ in the type of information and the manner in which it is used. The first method uses parallel training data for multi-task learning, in which the network is trained to perform both a primary senone classification task and a secondary feature enhancement task using a shared representation. The second method uses a parameterization of the reverberant environment extracted from the observed signal to train a room-aware DNN. Experiments were performed on the single microphone task of the REVERB Challenge corpus. The proposed approach obtained a word error rate of 7.8% on the SimData test set, which is lower than all reported systems using the same training data and evaluation conditions, and 27.5% on the mismatched RealData test set, which is lower than all but two systems.
Ritwik Giri, Michael L. Seltzer, Jasha Droppo, Dong Yu 0001
ICASSP4
2015 Speech recognition with prediction-adaptation-correction recurrent neural networks
abstract
We propose the prediction-adaptation-correction RNN (PAC-RNN), in which a correction DNN estimates the state posterior probability based on both the current frame and the prediction made on the past frames by a prediction DNN. The result from the main DNN is fed back to the prediction DNN to make better predictions for the future frames. In the PAC-RNN, we can consider that, given the new, current frame information, the main DNN makes a correction on the prediction made by the prediction DNN. Alternatively, it can be viewed as adapting the main DNN's behavior based on the prediction DNN's prediction. Experiments on the TIMIT phone recognition task indicate that the PAC-RNN outperforms DNN, RNN, and LSTM with 2.4%, 2.1%, and 1.9% absolute phone accuracy improvement, respectively. We found that incorporating the prediction objective and including the recurrent loop are both important to boost the performance of the PAC-RNN.
Yu Zhang 0033, Dong Yu 0001, Michael L. Seltzer, Jasha Droppo
ICASSP2
2015 Using Recurrent Neural Networks for Slot Filling in Spoken Language Understanding
abstract
Semantic slot filling is one of the most challenging problems in spoken language understanding (SLU). In this paper, we propose to use recurrent neural networks (RNNs) for this task, and present several novel architectures designed to efficiently model past and future temporal dependencies. Specifically, we implemented and compared several important RNN architectures, including Elman, Jordan, and hybrid variants. To facilitate reproducibility, we implemented these networks with the publicly available Theano neural network toolkit and completed experiments on the well-known airline travel information system (ATIS) benchmark. In addition, we compared the approaches on two custom SLU data sets from the entertainment and movies domains. Our results show that the RNN-based models outperform the conditional random field (CRF) baseline by 2% in absolute error reduction on the ATIS benchmark. We improve the state-of-the-art by 0.5% in the Entertainment domain, and 6.7% for the movies domain.
Grégoire Mesnil, Yann N. Dauphin, Kaisheng Yao, Yoshua Bengio, Li Deng 0001, Dilek Hakkani-Tür, Xiaodong He 0001, Larry Heck, Gökhan Tür, Dong Yu 0001, Geoffrey Zweig
IEEE ACM Trans. Audio Speech Lang. Process.10
2015 Deep Neural Networks for Single-Channel Multi-Talker Speech Recognition
abstract
We investigate techniques based on deep neural networks (DNNs) for attacking the single-channel multi-talker speech recognition problem. Our proposed approach contains five key ingredients: a multi-style training strategy on artificially mixed speech data, a separate DNN to estimate senone posterior probabilities of the louder and softer speakers at each frame, a weighted finite-state transducer (WFST)-based two-talker decoder to jointly estimate and correlate the speaker and speech, a speaker switching penalty estimated from the energy pattern change in the mixed-speech, and a confidence based system combination strategy. Experiments on the 2006 speech separation and recognition challenge task demonstrate that our proposed DNN-based system has remarkable noise robustness to the interference of a competing speaker. The best setup of our proposed systems achieves an average word error rate (WER) of 18.8% across different SNRs and outperforms the state-of-the-art IBM superhuman system by 2.8% absolute with fewer assumptions.
Chao Weng, Dong Yu 0001, Michael L. Seltzer, Jasha Droppo
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 Phone sequence modeling with recurrent neural networks
abstract
In this paper, we investigate phone sequence modeling with recurrent neural networks in the context of speech recognition. We introduce a hybrid architecture that combines a phonetic model with an arbitrary frame-level acoustic model and we propose efficient algorithms for training, decoding and sequence alignment. We evaluate the advantage of our phonetic model on the TIMIT and Switchboard-mini datasets in complementarity to a powerful context-dependent deep neural network (DNN) acoustic classifier and a higher-level 3-gram language model. Consistent improvements of 2-10% in phone accuracy and 3% in word error rate suggest that our approach can readily replace HMMs in current state-of-the-art systems.
Nicolas Boulanger-Lewandowski, Jasha Droppo, Mike Seltzer, Dong Yu 0001
ICASSP4
2014 On parallelizability of stochastic gradient descent for speech DNNS
abstract
This paper compares the theoretical efficiency of model-parallel and data-parallel distributed stochastic gradient descent training of DNNs. For a typical Switchboard DNN with 46M parameters, the results are not pretty: With modern GPUs and interconnects, model parallelism is optimal with only 3 GPUs in a single server, while data parallelism with a minibatch size of 1024 does not even scale to 2 GPUs. We further show that data-parallel training efficiency can be improved by increasing the minibatch size (through a combination of AdaGrad and automatic adjustments of learning rate and minibatch size) and data compression. We arrive at an estimated possible end-to-end speed-up of 5 times or more. We do not address issues of robustness to process failure or other issues that might occur during training, nor of speed of convergence differences between ASGD and SGD parameter update patterns.
Frank Seide, Jasha Droppo, Gang Li 0012, Dong Yu 0001
ICASSP5
2014 Single-channel mixed speech recognition using deep neural networks
abstract
In this work, we study the problem of single-channel mixed speech recognition using deep neural networks (DNNs). Using a multi-style training strategy on artificially mixed speech data, we investigate several different training setups that enable the DNN to generalize to corresponding similar patterns in the test data. We also introduce a WFST-based two-talker decoder to work with the trained DNNs. Experiments on the 2006 speech separation and recognition challenge task demonstrate that the proposed DNN-based system has remarkable noise robustness to the interference of a competing speaker. The best setup of our proposed systems achieves an overall WER of 19.7% which improves upon the results obtained by the state-of-the-art IBM superhuman system by 1.9% absolute, with fewer assumptions and lower computational complexity.
Chao Weng, Dong Yu 0001, Michael L. Seltzer, Jasha Droppo
ICASSP2
2014 Recurrent deep neural networks for robust speech recognition
abstract
In this work, we propose recurrent deep neural networks (DNNs) for robust automatic speech recognition (ASR). Full recurrent connections are added to certain hidden layer of a conventional feedforward DNN and allow the model to capture the temporal dependency in deep representations. A new backpropagation through time (BPTT) algorithm is introduced to make the minibatch stochastic gradient descent (SGD) on the proposed recurrent DNNs more efficient and effective. We evaluate the proposed recurrent DNN architecture under the hybrid setup on both the 2ndCHiME challenge (track 2) and Aurora-4 tasks. Experimental results on the CHiME challenge data show that the proposed system can obtain consistent 7% relative WER improvements over the DNN systems, achieving state-of-the-art performance without front-end preprocessing, speaker adaptive training or multiple decoding passes. For the experiments on Aurora-4, the proposed system achieves 4% relative WER improvement over a strong DNN baseline system.
Chao Weng, Dong Yu 0001, Shinji Watanabe 0001, Biing-Hwang Juang
ICASSP2
2014 Singular value decomposition based low-footprint speaker adaptation and personalization for deep neural network
abstract
The large number of parameters in deep neural networks (DNN) for automatic speech recognition (ASR) makes speaker adaptation very challenging. It also limits the use of speaker personalization due to the huge storage cost in large-scale deployments. In this paper we address DNN adaptation and personalization issues by presenting two methods based on the singular value decomposition (SVD). The first method uses an SVD to replace the weight matrix of a speaker independent DNN by the product of two low rank matrices. Adaptation is then performed by updating a square matrix inserted between the two low-rank matrices. In the second method, we adapt the full weight matrix but only store the delta matrix - the difference between the original and adapted weight matrices. We decrease the footprint of the adapted model by storing a reduced rank version of the delta matrix via an SVD. The proposed methods were evaluated on short message dictation task. Experimental results show that we can obtain similar accuracy improvements as the previously proposed Kullback-Leibler divergence (KLD) regularized method with far fewer parameters, which only requires 0.89% of the original model storage.
Jinyu Li 0001, Dong Yu 0001, Mike Seltzer, Yifan Gong 0001
ICASSP3
2014 Recurrent conditional random field for language understanding
abstract
Recurrent neural networks (RNNs) have recently produced record setting performance in language modeling and word-labeling tasks. In the word-labeling task, the RNN is used analogously to the more traditional conditional random field (CRF) to assign a label to each word in an input sequence, and has been shown to significantly outperform CRFs. In contrast to CRFs, RNNs operate in an online fashion to assign labels as soon as a word is seen, rather than after seeing the whole word sequence. In this paper, we show that the performance of an RNN tagger can be significantly improved by incorporating elements of the CRF model; specifically, the explicit modeling of output-label dependencies with transition features, its global sequence-level objective function, and offline decoding. We term the resulting model a “recurrent conditional random field” and demonstrate its effectiveness on the ATIS travel domain dataset and a variety of web-search language understanding datasets.
Kaisheng Yao, Baolin Peng, Geoffrey Zweig, Dong Yu 0001
ICASSP4
2014 Speech emotion recognition using deep neural network and extreme learning machine
abstract
Speech emotion recognition is a challenging problem partly because it is unclear what features are effective for the task. In this paper we propose to utilize deep neural networks (DNNs) to extract high level features from raw data and show that they are effective for speech emotion recognition. We first produce an emotion state probability distribution for each speech segment using DNNs. We then construct utterance-level features from segment-level probability distributions. These utterancelevel features are then fed into an extreme learning machine (ELM), a special simple and efficient single-hidden-layer neural network, to identify utterance-level emotions. The experimental results demonstrate that the proposed approach effectively learns emotional information from low-level features and leads to 20% relative accuracy improvement compared to the stateof-the-art approaches.
Dong Yu 0001, Ivan Tashev
INTERSPEECH2
2014 A comparative analytic study on the Gaussian mixture and context dependent deep neural network hidden Markov models
abstract
We conducted a comparative analytic study on the contextdependent Gaussian mixture hiddenMarkov model (CD-GMMHMM) and deep neural network hidden Markov model (CDDNN-HMM) with respect to the phone discrimination and the robustness performance. We found that the DNN can significantly improve the phone recognition performance for every phoneme with 15.6% to 39.8% relative phone error rate reduction (PERR). It is particularly good at discriminating certain consonants, which are found to be “hard” in the GMM. On the robustness side, the DNN outperforms the GMM at all SNR levels, across different devices, and under all speaking rate with nearly uniform improvement. The performance gap with respect to different SNR levels, distinct channels, and varied speaking rate remains large. For example, in CD-DNNHMM, we observed 1∼2% performance degradation per 1dB SNR drop; 20∼25% performance gap between the best and least well performed devices; 15∼30% relative word error rate increase when the speaking rate speeds up or slows down by 30% from the “sweet” spot. Therefore, we conclude the robustness remains to be a major challenge in the deep learning acoustic model. Speech enhancement, channel normalization, and speaking rate compensation are important research areas in order to further improve the DNN model accuracy.
Yan Huang 0028, Dong Yu 0001, Chaojun Liu, Yifan Gong 0001
INTERSPEECH2
2014 Multi-accent deep neural network acoustic model with accent-specific top layer using the KLD-regularized model adaptation
abstract
We propose a multi-accent deep neural network acoustic model with an accent-specific top layer and shared bottom hidden layers. The accent-specific top layer is used to model the distinct accent specific patterns. The shared bottom hidden layers allow maximum knowledge sharing between the native and the accent models. This design is particularly attractive when considering deploying such a system to a live speech service due to its computational efficiency. We applied the KL-divergence (KLD) regularized model adaptation to train the accent-specific top layer. On the mobile short message dictation task (SMD), with 1K, 10K, and 100K British or Indian accent adaptation utterances, the proposed approach achieves 18.1%, 26.0%, and 28.5% or 16.1%, 25.4%, and 30.6% word error rate reduction (WERR) for the British and the Indian accent respectively against a baseline cross entropy (CE) model trained from 400 hour data. On the 100K utterance accent adaptation setup, comparable performance gain can be obtained against a baseline CE model trained with 2000 hour data. We observe smaller yet significant WER reduction on a baseline model trained using the MMI sequence-level criterion.
Yan Huang 0028, Dong Yu 0001, Chaojun Liu, Yifan Gong 0001
INTERSPEECH2
2014 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs
abstract
We show empirically that in SGD training of deep neural networks, one can, at no or nearly no loss of accuracy, quantize the gradients aggressively—to but one bit per value—if the quantization error is carried forward across minibatches (error feedback). This size reduction makes it feasible to parallelize SGD through data-parallelism with fast processors like recent GPUs. We implement data-parallel deterministically distributed SGD by combining this finding with AdaGrad, automatic minibatch-size selection, double buffering, and model parallelism. Unexpectedly, quantization benefits AdaGrad, giving a small accuracy gain. For a typical Switchboard DNN with 46M parameters, we reach computation speeds of 27k frames per second (kfps) when using 2880 samples per minibatch, and 51kfps with 16k, on a server with 8 K20X GPUs. This corresponds to speed-ups over a single GPU of 3.6 and 6.3, respectively. 7 training passes over 309h of data complete in under 7h. A 160M-parameter model training processes 3300h of data in under 16h on 20 dual-GPU servers—a 10 times speed-up—albeit at a small accuracy loss.
Frank Seide, Jasha Droppo, Gang Li 0012, Dong Yu 0001
INTERSPEECH5
2014 An introduction to computational networks and the computational network toolkit (invited talk)
Dong Yu 0001, Adam Eversole, Michael L. Seltzer, Kaisheng Yao, Brian Guenter, Oleksii Kuchaiev, Frank Seide, Huaming Wang, Jasha Droppo, Zhiheng Huang, Geoffrey Zweig, Christopher J. Rossbach, Jon Currey
INTERSPEECH1
2014 Spoken language understanding using long short-term memory neural networks
abstract
Neural network based approaches have recently produced record-setting performances in natural language understanding tasks such as word labeling. In the word labeling task, a tagger is used to assign a label to each word in an input sequence. Specifically, simple recurrent neural networks (RNNs) and convolutional neural networks (CNNs) have shown to significantly outperform the previous state-of-the-art - conditional random fields (CRFs). This paper investigates using long short-term memory (LSTM) neural networks, which contain input, output and forgetting gates and are more advanced than simple RNN, for the word labeling task. To explicitly model output-label dependence, we propose a regression model on top of the LSTM un-normalized scores. We also propose to apply deep LSTM to the task. We investigated the relative importance of each gate in the LSTM by setting other gates to a constant and only learning particular gates. Experiments on the ATIS dataset validated the effectiveness of the proposed models.
Kaisheng Yao, Baolin Peng, Yu Zhang 0033, Dong Yu 0001, Geoffrey Zweig, Yangyang Shi
SLT4
2014 A fast maximum likelihood nonlinear feature transformation method for GMM-HMM speaker adaptation
Kaisheng Yao, Dong Yu 0001, Li Deng 0001, Yifan Gong 0001
Neurocomputing2
2014 Convolutional Neural Networks for Speech Recognition
abstract
Recently, the hybrid deep neural network (DNN)-hidden Markov model (HMM) has been shown to significantly improve speech recognition performance over the conventional Gaussian mixture model (GMM)-HMM. The performance improvement is partially attributed to the ability of the DNN to model complex correlations in speech features. In this paper, we show that further error rate reduction can be obtained by using convolutional neural networks (CNNs). We first present a concise description of the basic CNN and explain how it can be used for speech recognition. We further propose a limited-weight-sharing scheme that can better model speech features. The special structure such as local connectivity, weight sharing, and pooling in CNNs exhibits some degree of invariance to small shifts of speech features along the frequency axis, which is important to deal with speaker and environment variations. Experimental results show that CNNs reduce the error rate by 6%-10% compared with DNNs on the TIMIT phone recognition and the voice search large vocabulary speech recognition tasks.
Ossama Abdel-Hamid, Abdel-rahman Mohamed, Hui Jiang 0001, Li Deng 0001, Gerald Penn, Dong Yu 0001
IEEE ACM Trans. Audio Speech Lang. Process.6
2013 Large-scale malware classification using random projections and neural networks
abstract
Automatically generated malware is a significant problem for computer users. Analysts are able to manually investigate a small number of unknown files, but the best large-scale defense for detecting malware is automated malware classification. Malware classifiers often use sparse binary features, and the number of potential features can be on the order of tens or hundreds of millions. Feature selection reduces the number of features to a manageable number for training simpler algorithms such as logistic regression, but this number is still too large for more complex algorithms such as neural networks. To overcome this problem, we used random projections to further reduce the dimensionality of the original input space. Using this architecture, we train several very large-scale neural network systems with over 2.6 million labeled samples thereby achieving classification results with a two-class error rate of 0.49% for a single neural network and 0.42% for an ensemble of neural networks.
George E. Dahl, Jack W. Stokes, Li Deng 0001, Dong Yu 0001
ICASSP4
2013 A deep convolutional neural network using heterogeneous pooling for trading acoustic invariance with phonetic confusion
abstract
We develop and present a novel deep convolutional neural network architecture, where heterogeneous pooling is used to provide constrained frequency-shift invariance in the speech spectrogram while minimizing speech-class confusion induced by such invariance. The design of the pooling layer is guided by domain knowledge about how speech classes would change when formant frequencies are modified. The convolution and heterogeneous-pooling layers are followed by a fully connected multi-layer neural network to form a deep architecture interfaced to an HMM for continuous speech recognition. During training, all layers of this entire deep net are regularized using a variant of the “dropout” technique. Experimental evaluation demonstrates the effectiveness of both heterogeneous pooling and dropout regularization. On the TIMIT phonetic recognition task, we have achieved an 18.7% phone error rate, lowest on this standard task reported in the literature with a single system and with no use of information about speaker identity. Preliminary experiments on large vocabulary speech recognition in a voice search task also show error rate reduction using heterogeneous pooling in the deep convolutional neural network.
Li Deng 0001, Ossama Abdel-Hamid, Dong Yu 0001
ICASSP3
2013 Recent advances in deep learning for speech research at Microsoft
abstract
Deep learning is becoming a mainstream technology for speech recognition at industrial scale. In this paper, we provide an overview of the work by Microsoft speech researchers since 2009 in this area, focusing on more recent advances which shed light to the basic capabilities and limitations of the current deep learning technology. We organize this overview along the feature-domain and model-domain dimensions according to the conventional approach to analyzing speech systems. Selected experimental results, including speech recognition and related applications such as spoken dialogue and language modeling, are presented to demonstrate and analyze the strengths and weaknesses of the techniques described in the paper. Potential improvement of these techniques and future research directions are discussed.
Li Deng 0001, Jinyu Li 0001, Jui-Ting Huang, Kaisheng Yao, Dong Yu 0001, Frank Seide, Michael L. Seltzer, Geoffrey Zweig, Xiaodong He 0001, Jason D. Williams, Yifan Gong 0001, Alex Acero
ICASSP5
2013 Cross-language knowledge transfer using multilingual deep neural network with shared hidden layers
abstract
In the deep neural network (DNN), the hidden layers can be considered as increasingly complex feature transformations and the final softmax layer as a log-linear classifier making use of the most abstract features computed in the hidden layers. While the loglinear classifier should be different for different languages, the feature transformations can be shared across languages. In this paper we propose a shared-hidden-layer multilingual DNN (SHL-MDNN), in which the hidden layers are made common across many languages while the softmax layers are made language dependent. We demonstrate that the SHL-MDNN can reduce errors by 3-5%, relatively, for all the languages decodable with the SHL-MDNN, over the monolingual DNNs trained using only the language specific data. Further, we show that the learned hidden layers sharing across languages can be transferred to improve recognition accuracy of new languages, with relative error reductions ranging from 6% to 28% against DNNs trained without exploiting the transferred hidden layers. It is particularly interesting that the error reduction can be achieved for the target language that is in different families of the languages used to learn the hidden layers.
Jui-Ting Huang, Jinyu Li 0001, Dong Yu 0001, Li Deng 0001, Yifan Gong 0001
ICASSP3
2013 Modeling spectral envelopes using restricted Boltzmann machines for statistical parametric speech synthesis
abstract
This paper presents a new spectral modeling method for statistical parametric speech synthesis. In contrast to the conventional methods in which high-level spectral parameters, such as mel-cepstra or line spectral pairs, are adopted as the features for hidden Markov model (HMM) based parametric speech synthesis, our new method directly models the distribution of the lower-level, un-transformed or raw spectral envelopes. Instead of using single Gaussian distributions, we adopt restricted Boltzmann machines (RBM) to represent the distribution of the spectral envelopes at each HMM state. We anticipate these will give superior performance in modeling the joint distribution of high-dimensional stochastic vectors. The spectral parameters are derived from the spectral envelope corresponding to the estimated mode of each context-dependent RBM and act as the Gaussian mean vector in the parameter generation procedure at synthesis time. Our experimental results show that the RBM is able to model the distribution of the spectral envelopes with better accuracy and generalization ability than the Gaussian mixture model. As a result, our proposed method can significantly improve the naturalness of the conventional HMM-based speech synthesis system using mel-cepstra.
Zhen-Hua Ling, Li Deng 0001, Dong Yu 0001
ICASSP3
2013 An investigation of deep neural networks for noise robust speech recognition
abstract
Recently, a new acoustic model based on deep neural networks (DNN) has been introduced. While the DNN has generated significant improvements over GMM-based systems on several tasks, there has been no evaluation of the robustness of such systems to environmental distortion. In this paper, we investigate the noise robustness of DNN-based acoustic models and find that they can match state-of-the-art performance on the Aurora 4 task without any explicit noise compensation. This performance can be further improved by incorporating information about the environment into DNN training using a new method called noise-aware training. When combined with the recently proposed dropout training technique, a 7.5% relative improvement over the previously best published result on this task is achieved using only a single decoding pass and no additional decoding complexity compared to a standard DNN.
Michael L. Seltzer, Dong Yu 0001, Yongqiang Wang 0005
ICASSP2
2013 Error back propagation for sequence training of Context-Dependent Deep NetworkS for conversational speech transcription
abstract
We investigate back-propagation based sequence training of Context-Dependent Deep-Neural-Network HMMs, or CD-DNN-HMMs, for conversational speech transcription. Theoretically, sequence training integrates with backpropagation in a straight-forward manner. However, we find that to get reasonable results, heuristics are needed that point to a problem with lattice sparseness: The model must be adjusted to the updated numerator lattices by additional iterations of frame-based cross-entropy (CE) training; and to avoid distortions from “runaway” models, we can either add artificial silence arcs to the denominator lattices, or smooth the sequence objective with the frame-based one (F-smoothing). With the 309h Switchboard training set, the MMI objective achieves a relative word-error rate reduction of 11-15% over CE for matched test sets, and 10-17% for mismatched ones. This includes gains of 4-7% from realigned CE iterations. The BMMI and sMBR objectives gain less. With 2000h of data, gains are 2-9% after realigned CE iterations. Using GPGPUs, MMI is about 70% slower than CE training.
Gang Li 0012, Dong Yu 0001, Frank Seide
ICASSP3
2013 KL-divergence regularized deep neural network adaptation for improved large vocabulary speech recognition
abstract
We propose a novel regularized adaptation technique for context dependent deep neural network hidden Markov models (CD-DNN-HMMs). The CD-DNN-HMM has a large output layer and many large hidden layers, each with thousands of neurons. The huge number of parameters in the CD-DNN-HMM makes adaptation a challenging task, esp. when the adaptation set is small. The technique developed in this paper adapts the model conservatively by forcing the senone distribution estimated from the adapted model to be close to that from the unadapted model. This constraint is realized by adding Kullback-Leibler divergence (KLD) regularization to the adaptation criterion. We show that applying this regularization is equivalent to changing the target distribution in the conventional backpropagation algorithm. Experiments on Xbox voice search, short message dictation, and Switchboard and lecture speech transcription tasks demonstrate that the proposed adaptation technique can provide 2%-30% relative error reduction against the already very strong speaker independent CD-DNN-HMM systems using different adaptation sets under both supervised and unsupervised adaptation setups.
Dong Yu 0001, Kaisheng Yao, Gang Li 0012, Frank Seide
ICASSP1
2013 Exploring convolutional neural network structures and optimization techniques for speech recognition
abstract
Recently, convolutional neural networks (CNNs) have been shown to outperform the standard fully connected deep neu-ral networks within the hybrid deep neural network / hidden Markov model (DNN/HMM) framework on the phone recogni-tion task. In this paper, we extend the earlier basic form of the CNN and explore it in multiple ways. We first investigate sev-eral CNN architectures, including full and limited weight shar-ing, convolution along frequency and time axes, and stacking of several convolution layers. We then develop a novel weighted softmax pooling layer so that the size in the pooling layer can be automatically learned. Further, we evaluate the effect of CNN pretraining, which is achieved by using a convolutional version of the RBM. We show that all CNN architectures we have in-vestigated outperform the earlier basic form of the DNN on both the phone recognition and large vocabulary speech recog-nition tasks. The architecture with limited weight sharing pro-vides additional gains over the full weight sharing architecture. The softmax pooling layer performs as well as the best CNN with the manually tuned fixed-pooling size, and has a potential for further improvement. Finally, we show that CNN pretrain-ing produces significantly better results on a large vocabulary speech recognition task.
Ossama Abdel-Hamid, Li Deng 0001, Dong Yu 0001
INTERSPEECH3
2013 Deep segmental neural networks for speech recognition
abstract
Hybrid systems which integrate the deep neural network (DNN) and hidden Markov model (HMM) have recently achieved re-markable performance in many large vocabulary speech recog-nition tasks. These systems, however, remain to rely on the HMM and assume the acoustic scores for the (windowed) frames are independent given the state, suffering from the same difficulty as in the previous GMM-HMM systems. In this pa-per, we propose the deep segmental neural network (DSNN), a segmental model that uses DNNs to estimate the acoustic scores of phonemic or sub-phonemic segments with variable lengths. This allows the DSNN to represent each segment as a single unit, in which frames are made dependent on each other. We describe the architecture of the DSNN, as well as its learning and decoding algorithms. Our evaluation experiments demon-strate that the DSNN can outperform the DNN/HMM hybrid systems and two existing segmental models including the seg-mental conditional random field and the shallow segmental neu-ral network.
Ossama Abdel-Hamid, Li Deng 0001, Dong Yu 0001, Hui Jiang 0001
INTERSPEECH3
2013 Semi-supervised GMM and DNN acoustic model training with multi-system combination and confidence re-calibration
abstract
We present our study on semi-supervised Gaussian mixture model (GMM) hidden Markov model (HMM) and deep neural network (DNN) HMM acoustic model training. We analyze the impact of transcription quality and data sampling approach on the performance of the resulting model, and propose a multisystem combination and confidence re-calibration approach to improve the transcription inference and data selection. Compared to using a single system recognition result and confidence score, our proposed approach reduces the phone error rate of the inferred transcription by 23.8% relatively when top 60% of data are selected. Experiments were conducted on the mobile short message dictation (SMD) task. For the GMM-HMM model, we achieved 7.2% relative word error rate reduction (WERR) against a well-trained narrow-band fMPE+bMMI system by adding 2100 hours of untranscribed data, and 28.2% relative WERR over a wide-band MLE model trained from transcribed out-of-domain voice search data after adding 10K hours of untranscribed SMD data. For the CD-DNN-HMM model, 11.7% and 15.0% relative WERRs are achieved after adding 1K hours of untranscribed data using random and importance sampling, respectively. We also found using large amount of untranscribed data for pretraining does not help. Index Terms: semi-supervised acoustic model training, system combination, confidence re-calibration, importance sampling
Yan Huang 0028, Dong Yu 0001, Yifan Gong 0001, Chaojun Liu
INTERSPEECH2
2013 Recurrent neural networks for language understanding
abstract
Recurrent Neural Network Language Models (RNN-LMs) have recently shown exceptional performance across a variety of applications. In this paper, we modify the architecture to perform Language Understanding, and advance the state-of-the-art for the widely used ATIS dataset. The core of our approach is to take words as input as in a standard RNN-LM, and then to predict slot labels rather than words on the output side. We present several variations that differ in the amount of word context that is used on the input side, and in the use of non-lexical features. Remarkably, our simplest model produces state-of-the-art results, and we advance state-of-the-art through the use of bagof-words, word embedding, named-entity, syntactic, and wordclass features. Analysis indicates that the superior performance is attributable to the task-specific word representations learned by the RNN.
Kaisheng Yao, Geoffrey Zweig, Mei-Yuh Hwang, Yangyang Shi, Dong Yu 0001
INTERSPEECH5
2013 Exploiting deep neural networks for detection-based speech recognition
Sabato Marco Siniscalchi, Dong Yu 0001, Li Deng 0001, Chin-Hui Lee 0001
Neurocomputing2
2013 Tensor Deep Stacking Networks
abstract
A novel deep architecture, the tensor deep stacking network (T-DSN), is presented. The T-DSN consists of multiple, stacked blocks, where each block contains a bilinear mapping from two hidden layers to the output layer, using a weight tensor to incorporate higher order statistics of the hidden binary (½0; 1) features. A learning algorithm for the T-DSN’s weight matrices and tensors is developed and described in which the main parameter estimation burden is shifted to a convex subproblem with a closed-form solution. Using an efficient and scalable parallel implementation for CPU clusters, we train sets of T-DSNs in three popular tasks in increasing order of the data size: handwritten digit recognition using MNIST (60k), isolated state/phone classification and continuous phone recognition using TIMIT (1.1 m), and isolated phone classification using WSJ0 (5.2 m). Experimental results in all three tasks demonstrate the effectiveness of the T-DSN and the associated learning methods in a consistent manner. In particular, a sufficient depth of the T-DSN, a symmetry in the two hidden layers structure in each T-DSN block, our model parameter learning algorithm, and a softmax layer on top of T-DSN are shown to have all contributed to the low error rates observed in the experiments for all three tasks.
Brian Hutchinson, Li Deng 0001, Dong Yu 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2013 Speech Recognition Using Long-Span Temporal Patterns in a Deep Network Model
abstract
In recent years, there has been a renewed interest in the use of artificial neural networks (ANNs) for speech applications, and it seems that a new trend to move the speech technology forward has begun. Two main contributions have triggered such a new trend: 1) a major advance has been made in training the weights in deep neural networks (DNNs), and a pre-trained deep neural network hidden Markov model (DNN-HMM) hybrid architecture has outperformed a conventional Gaussian mixture model hidden Markov model (GMM-HMM) automatic speech recognition (ASR) system on a challenging business search dataset, and 2) it has been shown that phoneme classification can be boosted by using a hierarchical structure of multi-layer perceptrons (MLPs) trained to model long-span temporal patterns with beneficial effects on language recognition tasks. In this work, we combine these two lines of research and demonstrate that word recognition accuracy can be significantly enhanced by arranging DNNs in a hierarchical structure to model long-term energy trajectories. The proposed solution has been evaluated on the 5000-word Wall Street Journal task, resulting in consistent and significant improvements in both phone and word recognition accuracy rates. We have also analyzed the effects of various modeling choices on the system performance, and several architectural solutions have been compared.
Sabato Marco Siniscalchi, Dong Yu 0001, Li Deng 0001, Chin-Hui Lee 0001
IEEE Signal Process. Lett.2
2013 Modeling Spectral Envelopes Using Restricted Boltzmann Machines and Deep Belief Networks for Statistical Parametric Speech Synthesis
abstract
This paper presents a new spectral modeling method for statistical parametric speech synthesis. In the conventional methods, high-level spectral parameters, such as mel-cepstra or line spectral pairs, are adopted as the features for hidden Markov model (HMM)-based parametric speech synthesis. Our proposed method described in this paper improves the conventional method in two ways. First, distributions of low-level, un-transformed spectral envelopes (extracted by the STRAIGHT vocoder) are used as the parameters for synthesis. Second, instead of using single Gaussian distribution, we adopt the graphical models with multiple hidden variables, including restricted Boltzmann machines (RBM) and deep belief networks (DBN), to represent the distribution of the low-level spectral envelopes at each HMM state. At the synthesis time, the spectral envelopes are predicted from the RBM-HMMs or the DBN-HMMs of the input sentence following the maximum output probability parameter generation criterion with the constraints of the dynamic features. A Gaussian approximation is applied to the marginal distribution of the visible stochastic variables in the RBM or DBN at each HMM state in order to achieve a closed-form solution to the parameter generation problem. Our experimental results show that both RBM-HMM and DBN-HMM are able to generate spectral envelope parameter sequences better than the conventional Gaussian-HMM with superior generalization capabilities and that DBN-HMM and RBM-HMM perform similarly due possibly to the use of Gaussian approximation. As a result, our proposed method can significantly alleviate the over-smoothing effect and improve the naturalness of the conventional HMM-based speech synthesis system using mel-cepstra.
Zhen-Hua Ling, Li Deng 0001, Dong Yu 0001
IEEE Trans. Speech Audio Process.3
2013 The Deep Tensor Neural Network With Applications to Large Vocabulary Speech Recognition
abstract
The recently proposed context-dependent deep neural network hidden Markov models (CD-DNN-HMMs) have been proved highly promising for large vocabulary speech recognition. In this paper, we develop a more advanced type of DNN, which we call the deep tensor neural network (DTNN). The DTNN extends the conventional DNN by replacing one or more of its layers with a double-projection (DP) layer, in which each input vector is projected into two nonlinear subspaces, and a tensor layer, in which two subspace projections interact with each other and jointly predict the next layer in the deep architecture. In addition, we describe an approach to map the tensor layers to the conventional sigmoid layers so that the former can be treated and trained in a similar way to the latter. With this mapping we can consider a DTNN as the DNN augmented with DP layers so that not only the BP learning algorithm of DTNNs can be cleanly derived but also new types of DTNNs can be more easily developed. Evaluation on Switchboard tasks indicates that DTNNs can outperform the already high-performing DNNs with 4-5% and 3% relative word error reduction, respectively, using 30-hr and 309-hr training sets.
Dong Yu 0001, Li Deng 0001, Frank Seide
IEEE Trans. Speech Audio Process.1
2012 Scalable stacking and learning for building deep architectures
abstract
Deep Neural Networks (DNNs) have shown remarkable success in pattern recognition tasks. However, parallelizing DNN training across computers has been difficult. We present the Deep Stacking Network (DSN), which overcomes the problem of parallelizing learning algorithms for deep architectures. The DSN provides a method of stacking simple processing modules in buiding deep architectures, with a convex learning problem in each module. Additional fine tuning further improves the DSN, while introducing minor non-convexity. Full learning in the DSN is batch-mode, making it amenable to parallel training over many machines and thus be scalable over the potentially huge size of the training data. Experimental results on both the MNIST (image) and TIMIT (speech) classification tasks demonstrate that the DSN learning algorithm developed in this work is not only parallelizable in implementation but it also attains higher classification accuracy than the DNN.
Li Deng 0001, Dong Yu 0001, John C. Platt
ICASSP2
2012 A deep architecture with bilinear modeling of hidden representations: Applications to phonetic recognition
abstract
We develop and describe a novel deep architecture, the Tensor Deep Stacking Network (T-DSN), where multiple blocks are stacked one on top of another and where a bilinear mapping from hidden representations to the output in each block is used to incorporate higher-order statistics of the input features. A learning algorithm for the T-DSN is presented, in which the main parameter estimation burden is shifted to a convex sub-problem with a closed-form solution. Using an efficient and scalable parallel implementation, we train a T-DSN to discriminate standard three-state monophones in the TIMIT database. The T-DSN outperforms an alternative pretrained Deep Neural Network (DNN) architecture in frame-level classification (both state and phone) and in the cross-entropy measure. For continuous phonetic recognition, T-DSN performs equivalently to a DNN but without the need for a hard-to-scale, sequential fine-tuning step.
Brian Hutchinson, Li Deng 0001, Dong Yu 0001
ICASSP3
2012 Boosting attribute and phone estimation accuracies with deep neural networks for detection-based speech recognition
abstract
Generation of high-precision sub-phonetic attribute (also known as phonological features) and phone lattices is a key frontend component for detection-based bottom-up speech recognition. In this paper we employ deep neural networks (DNNs) to improve detection accuracy over conventional shallow MLPs (multi-layer perceptrons) with one hidden layer. A range of DNN architectures with five to seven hidden layers and up to 2048 hidden units per layer have been explored. Training on the SI84 and testing on the Nov92 WSJ data, the proposed DNNs achieve significant improvements over the shallow MLPs, producing greater than 90% frame-level attribute estimation accuracies for all 21 attributes tested for the full system. On the phone detection task, we also obtain excellent frame-level accuracy of 86.6%. With this level of high-precision detection of basic speech units we have opened the door to a new family of flexible speech recognition system design for both top-down and bottom-up, lattice-based search strategies and knowledge integration.
Dong Yu 0001, Sabato Marco Siniscalchi, Li Deng 0001, Chin-Hui Lee 0001
ICASSP1
2012 Exploiting sparseness in deep neural networks for large vocabulary speech recognition
abstract
Recently, we developed context-dependent deep neural network (DNN) hidden Markov models for large vocabulary speech recognition. While reducing errors by 33% compared to its discriminatively trained Gaussian-mixture counterpart on the switchboard benchmark task, DNN requires much more parameters. In this paper, we report our recent work on DNN for improved generalization, model size, and computation speed by exploiting parameter sparseness. We formulate the goal of enforcing sparseness as soft regularization and convex constraint optimization problems, and propose solutions under the stochastic gradient ascent setting. We also propose novel data structures to exploit the random sparseness patterns to reduce model size and computation time. The proposed solutions have been evaluated on the voice-search and switchboard datasets. They have decreased the number of nonzero connections to one third while reducing the error rate by 0.2-0.3% over the fully connected model on both datasets. The nonzero connections have been further reduced to only 12% and 19% on the two respective datasets without sacrificing speech recognition performance. Under these conditions we can reduce the model size to 18% and 29%, and computation to 14% and 23%, respectively, on these two datasets.
Dong Yu 0001, Frank Seide, Gang Li 0012, Li Deng 0001
ICASSP1
2012 Conversational Speech Transcription Using Context-Dependent Deep Neural Networks
Dong Yu 0001, Frank Seide, Gang Li 0012
ICML1
2012 Pipelined Back-Propagation for Context-Dependent Deep Neural Networks
abstract
The Context-Dependent Deep-Neural-Network HMM, or CD-DNN-HMM, is a recently proposed acoustic-modeling tech-nique for HMM-based speech recognition that can greatly out-perform conventional Gaussian-mixture based HMMs. For ex-ample, a CD-DNN-HMM trained on the 2000h Fisher corpus achieves 14.4 % word error rate on the Hub5’00-FSH speaker-independent phone-call transcription task, compared to 19.6% obtained by a state-of-the-art, conventional discriminatively trained GMM-based HMM. That CD-DNN-HMM, however, took 59 days to train on a modern GPGPU—the immense computational cost of the mini-batch based back-propagation (BP) training is a major road-block. Unlike the familiar Baum-Welch training for conven-tional HMMs, BP cannot be efficiently parallelized across data. In this paper we show that the pipelined approximation to BP, which parallelizes computation with respect to layers, is an efficient way of utilizing multiple GPGPU cards in a single server. Using 2 and 4 GPGPUs, we achieve a 1.9 and 3.3 times end-to-end speed-up, at parallelization efficiency of 0.95 and 0.82, respectively, at no loss of recognition accuracy. Index Terms: speech recognition, deep neural networks, paral-lelization, GPGPU
Xie Chen 0001, Adam Eversole, Gang Li 0012, Dong Yu 0001, Frank Seide
INTERSPEECH4
2012 Parallel Training for Deep Stacking Networks
abstract
The Deep Stacking Network (DSN) is a special type of deep architecture developed to enable and benefit from parallel learning of its model parameters on large CPU clusters. As a prospective key component of future speech recognizers, the architectural design of the DSN and its parallel training endow the DSN with scalability over a vast amount of training data. In this paper, we present our first parallel implementation of the DSN training algorithm. Particularly, we show the tradeoff between the time/memory saving via training parallelism and the associated cost arising from inter-CPU communication. Further, in phone classification experiments, we demonstrate a significantly lowered error rate using parallel full-batch training distributed over a CPU cluster, compared with sequential minibatch training implemented in a single CPU machine under otherwise identical experimental conditions and as exploited prior to the work reported in this paper. Index Terms: parallel and distributed computing, deep stacking networks, full-batch training, phone classification
Li Deng 0001, Brian Hutchinson, Dong Yu 0001
INTERSPEECH3
2012 Large Vocabulary Speech Recognition Using Deep Tensor Neural Networks
abstract
Recently, we proposed and developed the context-dependent deep neural network hidden Markov models (CD-DNN-HMMs) for large vocabulary speech recognition and achieved highly promising recognition results including over one third fewer word errors than the discriminatively trained, conventional HMM-based systems on the 300hr Switchboard benchmark task. In this paper, we extend DNNs to deep tensor neural networks (DTNNs) in which one or more layers are double-projection and tensor layers. The basic idea of the DTNN comes from our realization that many factors interact with each other to predict the output. To represent these interactions, we project the input to two nonlinear subspaces through the double-projection layer and model the interactions between these two subspaces and the output neurons through a tensor with three-way connections. Evaluation on 30hr Switchboard task indicates that DTNNs can outperform DNNs with similar number of parameters with 5% relative word error reduction. Index Terms: automatic speech recognition, tensor deep neural networks, CD-DNN-HMM, large vocabulary 1.
Dong Yu 0001, Li Deng 0001, Frank Seide
INTERSPEECH1
2012 Improving wideband speech recognition using mixed-bandwidth training data in CD-DNN-HMM
abstract
Context-dependent deep neural network hidden Markov model (CD-DNN-HMM) is a recently proposed acoustic model that significantly outperformed Gaussian mixture model (GMM)-HMM systems in many large vocabulary speech recognition (LVSR) tasks. In this paper we present our strategy of using mixed-bandwidth training data to improve wideband speech recognition accuracy in the CD-DNN-HMM framework. We show that DNNs provide the flexibility of using arbitrary features. By using the Mel-scale log-filter bank features we not only achieve higher recognition accuracy than using MFCCs, but also can formulate the mixed-bandwidth training problem as a missing feature problem, in which several feature dimensions have no value when narrowband speech is presented. This treatment makes training CD-DNN-HMMs with mixed-bandwidth data an easy task since no bandwidth extension is needed. Our experiments on voice search data indicate that the proposed solution not only provides higher recognition accuracy for the wideband speech but also allows the same CD-DNN-HMM to recognize mixed-bandwidth speech. By exploiting mixed-bandwidth training data CD-DNN-HMM outperforms fMPE+BMMI trained GMM-HMM, which cannot benefit from using narrowband data, by 18.4%.
Jinyu Li 0001, Dong Yu 0001, Jui-Ting Huang, Yifan Gong 0001
SLT2
2012 Context-dependent Deep Neural Networks for audio indexing of real-life data
abstract
We apply Context-Dependent Deep-Neural-Network HMMs, or CD-DNN-HMMs, to the real-life problem of audio indexing of data across various sources. Recently, we had shown that on the Switchboard benchmark on speaker-independent transcription of phone calls, CD-DNN-HMMs with 7 hidden layers reduce the word error rate by as much as one-third, compared to discriminatively trained Gaussian-mixture HMMs, and by one-fourth if the GMM-HMM also uses fMPE features. This paper takes CD-DNN-HMM based recognition into a real-life deployment for audio indexing. We find that for our best speaker-independent CD-DNN-HMM, with 32k senones trained on 2000h of data, the one-fourth reduction does carry over to inhomogeneous field data (video podcasts and talks). Compared to a speaker-adaptive GMM system, the relative improvement is 18%, at very similar end-to-end runtime. In system building, we find that DNNs can benefit from a larger number of senones than the GMM-HMM; and that DNN likelihood evaluation is a sizeable runtime factor even in our wide-beam context of generating rich lattices: Cutting the model size by 60% reduces runtime by one-third at a 5% relative WER loss.
Gang Li 0012, Huifeng Zhu, Gong Cheng 0005, Kit Thambiratnam, Behrooz Chitsaz, Dong Yu 0001, Frank Seide
SLT6
2012 Adaptation of context-dependent deep neural networks for automatic speech recognition
abstract
In this paper, we evaluate the effectiveness of adaptation methods for context-dependent deep-neural-network hidden Markov models (CD-DNN-HMMs) for automatic speech recognition. We investigate the affine transformation and several of its variants for adapting the top hidden layer. We compare the affine transformations against direct adaptation of the softmax layer weights. The feature-space discriminative linear regression (fDLR) method with the affine transformations on the input layer is also evaluated. On a large vocabulary speech recognition task, a stochastic gradient ascent implementation of the fDLR and the top hidden layer adaptation is shown to reduce word error rates (WERs) by 17% and 14%, respectively, compared to the baseline DNN performances. With a batch update implementation, the softmax layer adaptation technique reduces WERs by 10%. We observe that using bias shift performs as well as doing scaling plus bias shift.
Kaisheng Yao, Dong Yu 0001, Frank Seide, Li Deng 0001, Yifan Gong 0001
SLT2
2012 Efficient and effective algorithms for training single-hidden-layer neural networks
Dong Yu 0001, Li Deng 0001
Pattern Recognit. Lett.1
2012 Context-Dependent Pre-Trained Deep Neural Networks for Large-Vocabulary Speech Recognition
abstract
We propose a novel context-dependent (CD) model for large-vocabulary speech recognition (LVSR) that leverages recent advances in using deep belief networks for phone recognition. We describe a pre-trained deep neural network hidden Markov model (DNN-HMM) hybrid architecture that trains the DNN to produce a distribution over senones (tied triphone states) as its output. The deep belief network pre-training algorithm is a robust and often helpful way to initialize deep neural networks generatively that can aid in optimization and reduce generalization error. We illustrate the key components of our model, describe the procedure for applying CD-DNN-HMMs to LVSR, and analyze the effects of various modeling choices on performance. Experiments on a challenging business search dataset demonstrate that CD-DNN-HMMs can significantly outperform the conventional context-dependent Gaussian mixture model (GMM)-HMMs, with an absolute sentence accuracy improvement of 5.8% and 9.2% (or relative error reduction of 16.0% and 23.2%) over the CD-GMM-HMMs trained using the minimum phone error rate (MPE) and maximum-likelihood (ML) criteria, respectively.
George E. Dahl, Dong Yu 0001, Li Deng 0001, Alex Acero
IEEE Trans. Speech Audio Process.2
2012 Introduction to the Special Section on Deep Learning for Speech and Language Processing
abstract
Current speech recognition systems, for example, typically use Gaussian mixture models (GMMs), to estimate the observation (or emission) probabilities of hidden Markov models (HMMs), and GMMs are generative models that have only one layer of latent variables. Instead of developing more powerful models, most of the research effort has gone into finding better ways of estimating the GMM parameters so that error rates are decreased or the margin between different classes is increased. The same observation holds for natural language processing (NLP) in which maximum entropy (MaxEnt) models and conditional random fields (CRFs) have been popular for the last decade. Both of these approaches use shallow models whose success largely depends on the use of carefully handcrafted features.
Dong Yu 0001, Geoffrey E. Hinton, Nelson Morgan, Jen-Tzung Chien, Shigeki Sagayama
IEEE Trans. Speech Audio Process.1
2011 Feature engineering in Context-Dependent Deep Neural Networks for conversational speech transcription
abstract
We investigate the potential of Context-Dependent Deep-Neural-Network HMMs, or CD-DNN-HMMs, from a feature-engineering perspective. Recently, we had shown that for speaker-independent transcription of phone calls (NIST RT03S Fisher data), CD-DNN-HMMs reduced the word error rate by as much as one third-from 27.4%, obtained by discriminatively trained Gaussian-mixture HMMs with HLDA features, to 18.5%-using 300+ hours of training data (Switchboard), 9000+ tied triphone states, and up to 9 hidden network layers.
Frank Seide, Gang Li 0012, Xie Chen 0001, Dong Yu 0001
ASRU4
2011 Large vocabulary continuous speech recognition with context-dependent DBN-HMMS
abstract
The context-independent deep belief network (DBN) hidden Markov model (HMM) hybrid architecture has recently achieved promising results for phone recognition. In this work, we propose a context-dependent DBN-HMM system that dramatically outperforms strong Gaussian mixture model (GMM)-HMM baselines on a challenging, large vocabulary, spontaneous speech recognition dataset from the Bing mobile voice search task. Our system achieves absolute sentence accuracy improvements of 5.8% and 9.2% over GMM-HMMs trained using the minimum phone error rate (MPE) and maximum likelihood (ML) criteria, respectively, which translate to relative error reductions of 16.0% and 23.2%.
George E. Dahl, Dong Yu 0001, Li Deng 0001, Alex Acero
ICASSP2
2011 Conversational Speech Transcription Using Context-Dependent Deep Neural Networks
abstract
We apply the recently proposed Context-Dependent Deep-Neural-Network HMMs, or CD-DNN-HMMs, to speech-to-text transcription. For single-pass speaker-independent recognition on the RT03S Fisher portion of phone-call transcription benchmark (Switchboard), the word-error rate is reduced from 27.4%, obtained by discriminatively trained Gaussian-mixture HMMs, to 18.5%—a 33 % relative improvement. CD-DNN-HMMs combine classic artificial-neural-network HMMs with traditional tied-state triphones and deep-beliefnetwork pre-training. They had previously been shown to reduce errors by 16 % relatively when trained on tens of hours of data using hundreds of tied states. This paper takes CD-DNN-HMMs further and applies them to transcription using over 300 hours of training data, over 9000 tied states, and up to 9 hidden layers, and demonstrates how sparseness can be exploited. On four less well-matched transcription tasks, we observe relative error reductions of 22–28%. Index Terms: speech recognition, deep belief networks, deep neural networks
Frank Seide, Gang Li 0012, Dong Yu 0001
INTERSPEECH3
2011 Accelerated Parallelizable Neural Network Learning Algorithm for Speech Recognition
abstract
We describe a set of novel, batch-mode algorithms we developed recently as one key component in scalable, deep neural network based speech recognition. The essence of these algorithms is to structure the singlehidden-layer neural network so that the upper-layer’s weights can be written as a deterministic function of the lower-layer’s weights. This structure is effectively exploited during training by plugging in the deterministic function to the least square error objective function while calculating the gradients. Accelerating techniques are further exploited to make the weight updates move along the most promising directions. The experiments on TIMIT frame-level phone and phonestate classification show strong results. In particular, the error rate is strictly monotonically dropping as the minibatch size increases. This demonstrates the potential for the proposed batch-mode algorithms in large scale speech recognition since they are easily parallelizable across computers. Index Terms: neural network, scalability, structure, constraints, FISTA acceleration, optimization, pseudoinverse, weighted LSE, phone state classification, speech recognition, deep learning 1.
Dong Yu 0001, Li Deng 0001
INTERSPEECH1
2011 Deep Convex Net: A Scalable Architecture for Speech Pattern Classification
abstract
We recently developed context-dependent DNN-HMM (Deep-Neural-Net/Hidden-Markov-Model) for large-vocabulary speech recognition. While achieving impressive recognition error rate reduction, we face the insurmountable problem of scalability in dealing with virtually unlimited amount of training data available nowadays. To overcome the scalability challenge, we have designed the deep convex network (DCN) architecture. The learning problem in DCN is convex within each module. Additional structure-exploited fine tuning further improves the quality of DCN. The full learning in DCN is batch-mode based instead of stochastic, naturally lending it amenable to parallel training that can be distributed over many machines. Experimental results on both MNIST and TIMIT tasks evaluated thus far demonstrate superior performance of DCN over the DBN (Deep Belief Network) counterpart that forms the basis of the DNN. The superiority is reflected not only in training scalability and CPU-only computation, but more importantly in classification accuracy in both tasks. Index Terms: deep learning, scalability, convex optimization, neural network, deep belief network, phone state
Dong Yu 0001, Li Deng 0001
INTERSPEECH1
2011 Improved Bottleneck Features Using Pretrained Deep Neural Networks
abstract
Bottleneck features have been shown to be effective in improving the accuracy of automatic speech recognition (ASR) systems. Conventionally, bottleneck features are extracted from a multi-layer perceptron (MLP) trained to predict context-independent monophone states. The MLP typically has three hidden layers and is trained using the backpropagation algorithm. In this paper, we propose two improvements to the training of bottleneck features motivated by recent advances in the use of deep neural networks (DNNs) for speech recognition. First, we show how the use of unsupervised pretraining of a DNN enhances the network’s discriminative power and improves the bottleneck features it generates. Second, we show that a neural network trained to predict context-dependent senone targets produces better bottleneck features than one trained to predict monophone states. Bottleneck features trained using the proposed methods produced a 16% relative reduction in sentence error rate over conventional bottleneck features on a large vocabulary business search task.
Dong Yu 0001, Michael L. Seltzer
INTERSPEECH1
2011 Calibration of Confidence Measures in Speech Recognition
abstract
Most speech recognition applications in use today rely heavily on confidence measure for making optimal decisions. In this paper, we aim to answer the question: what can be done to improve the quality of confidence measure if we cannot modify the speech recognition engine? The answer provided in this paper is a post-processing step called confidence calibration, which can be viewed as a special adaptation technique applied to confidence measure. Three confidence calibration methods have been developed in this work: the maximum entropy model with distribution constraints, the artificial neural network, and the deep belief network. We compare these approaches and demonstrate the importance of key features exploited: the generic confidence-score, the application-dependent word distribution, and the rule coverage ratio. We demonstrate the effectiveness of confidence calibration on a variety of tasks with significant normalized cross entropy increase and equal error rate reduction.
Dong Yu 0001, Jinyu Li 0001, Li Deng 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2010 Semantic confidence calibration for spoken dialog applications
abstract
The success of spoken dialog applications depends strongly on the quality of the semantic confidence measure that determines the selection of the dialog strategy. However, the semantic confidence measure obtained from typical automatic speech recognition engines is not optimized for specific semantic slots and applications. We present our recent work on using a novel maximum entropy model with distribution constraints to calibrate the semantic confidence scores with the inputs of only the raw semantic confidence and the associated raw word confidence scores. We illustrate how features can be constructed from the raw confidence scores with a variable number of words and how the quality of the semantic confidence measure can be further improved by adding another calibration stage for the word confidence measure. We demonstrate the effectiveness of our approach for two types of semantic slots of practical significance. For the ZIP-code semantic slot, the new measure achieves relative 10.6% mean square error (MSE), 19.3% normalized negative log-likelihood (NNLL), and 38.5% equal error rate (EER) reduction. The counterpart of the date-time semantic slot is 37.8%, 38.7%, and 23.1%, respectively.
Dong Yu 0001, Li Deng 0001
ICASSP1
2010 Language recognition using deep-structured conditional random fields
abstract
We present a novel language identification technique using our recently developed deep-structured conditional random fields (CRFs). The deep-structured CRF is a multi-layer CRF model in which each higher layer's input observation sequence consists of the lower layer's observation sequence and the resulting lower layer's frame-level marginal probabilities. In this paper we extend the original deep-structured CRF by allowing for distinct state representations at different layers and demonstrate its benefits. We propose an unsupervised algorithm to pre-train the intermediate layers by casting it as a multi-objective programming problem that is aimed at minimizing the average frame-level conditional entropy while maximizing the state occupation entropy. Empirical evaluation on a seven-language/dialect voice mail routing task showed that our approach can achieve a routing accuracy (RA) of 86.4% and average equal error rate (EER) of 6.6%. These results are significantly better than the 82.5% RA and 7.5% average EER obtained using the Gaussian mixture model trained with the maximum mutual information criterion but slightly worse than the 87.7% RA and 6.4% EER achieved using the support vector machine with model pushing on the Gaussian super vector (GSV).
Dong Yu 0001, Shizhen Wang, Zahi N. Karam, Li Deng 0001
ICASSP1
2010 Word confidence calibration using a maximum entropy model with constraints on confidence and word distributions
abstract
It is widely known that the quality of confidence measure is critical for speech applications. In this paper, we present our recent work on improving word confidence scores by calibrating them using a small set of calibration data when only the recognized word sequence and associated raw confidence scores are made available. The core of our technique is the maximum entropy model with distribution constraints which naturally and effectively make use of the word distribution, the raw confidence-score distribution, and the context information. We demonstrate the effectiveness of our approach by showing that it can achieve relative 38% mean square error (MSE), 39% negative normalized likelihood (NNLL), and 23% equal error rate (EER) reduction on a voice mail transcription data set and relative 35% MSE, 45% NNLL, and 35% EER reduction on a command and control data set.
Dong Yu 0001, Shizhen Wang, Jinyu Li 0001, Li Deng 0001
ICASSP1
2010 Binary coding of speech spectrograms using a deep auto-encoder
abstract
This paper reports our recent exploration of the layer-by-layer learning strategy for training a multi-layer generative model of patches of speech spectrograms. The top layer of the generative model learns binary codes that can be used for efficient compression of speech and could also be used for scalable speech recognition or rapid speech content retrieval. Each layer of the generative model is fully connected to the layer below and the weights on these connections are pretrained efficiently by using the contrastive divergence approximation to the log likelihood gradient. After layer-bylayer pre-training we “unroll” the generative model to form a deep auto-encoder, whose parameters are then fine-tuned using back-propagation. To reconstruct the full-length speech spectrogram, individual spectrogram segments predicted by their respective binary codes are combined using an overlapand-add method. Experimental results on speech spectrogram coding demonstrate that the binary codes produce a logspectral distortion that is approximately 2 dB lower than a subband vector quantization technique over the entire frequency range of wide-band speech. Index Terms: deep learning, speech feature extraction, neural networks, auto-encoder, binary codes, Boltzmann machine
Li Deng 0001, Michael L. Seltzer, Dong Yu 0001, Alex Acero, Abdel-rahman Mohamed, Geoffrey E. Hinton
INTERSPEECH3
2010 Unscented transform with online distortion estimation for HMM adaptation
abstract
In this paper, we propose to improve our previously developed method for joint compensation of additive and convolutive distortions (JAC) applied to model adaptation. The improvement entails replacing the vector Taylor series (VTS) approximation with unscented transform (UT) in formulating both the static and dynamic model parameter adaptation. Our new JAC-UT method differentiates itself from other UT-based approaches in that it combines the online noise and channel distortion estimation and model parameter adaptation in a unified UT framework. Experimental results on the standard Aurora 2 task show that the new algorithm enjoys 20.0 % and 16.9 % relative word error rate reductions over the previous JAC-VTS algorithm when using the simple and complex backend models, respectively. Index Terms: unscented transform, vector Taylor series, additive and convolutive distortions, robust ASR, adaptation
Jinyu Li 0001, Dong Yu 0001, Yifan Gong 0001, Li Deng 0001
INTERSPEECH2
2010 Investigation of full-sequence training of deep belief networks for speech recognition
abstract
Recently, Deep Belief Networks (DBNs) have been proposed for phone recognition and were found to achieve highly competitive performance. In the original DBNs, only framelevel information was used for training DBN weights while it has been known for long that sequential or full-sequence information can be helpful in improving speech recognition accuracy. In this paper we investigate approaches to optimizing the DBN weights, state-to-state transition parameters, and language model scores using the sequential discriminative training criterion. We describe and analyze the proposed training algorithm and strategy, and discuss practical issues and how they affect the final results. We show that the DBNs learned using the sequence-based training criterion outperform those with frame-based criterion using both threelayer and six-layer models, but the optimization procedure for the deeper DBN is more difficult for the former criterion.
Abdel-rahman Mohamed, Dong Yu 0001, Li Deng 0001
INTERSPEECH2
2010 Deep-structured hidden conditional random fields for phonetic recognition
abstract
We extend our earlier work on deep-structured conditional random field (DCRF) and develop deep-structured hidden conditional random field (DHCRF). We investigate the use of this new sequential deep-learning model for phonetic recognition. DHCRF is a hierarchical model in which the final layer is a hidden conditional random field (HCRF) and the intermediate layers are zero-th-order conditional random fields (CRFs). Parameter estimation and sequence inference in the DHCRF are developed in this work. They are carried out layer by layer so that the time complexity is linear to the number of layers. In the DHCRF, the training label is available only at the final layer and the state boundary is unknown. This difficulty is addressed by using unsupervised learning for the intermediate layers and lattice-based supervised learning for the final layer. Experiments on the standard TIMIT phone recognition task show small performance improvement of a three-layer DHCRF over a two-layer DHCRF; both are significantly better than the single-layer DHCRF and are superior to the discriminatively trained tri-phone hidden Markov model (HMM) using identical input features.
Dong Yu 0001, Li Deng 0001
INTERSPEECH1
2010 Active learning and semi-supervised learning for speech recognition: A unified framework using the global entropy reduction maximization criterion
Dong Yu 0001, Balakrishnan Varadarajan, Li Deng 0001, Alex Acero
Comput. Speech Lang.1
2009 A study on multilingual acoustic modeling for large vocabulary ASR
abstract
We study key issues related to multilingual acoustic modeling for automatic speech recognition (ASR) through a series of large-scale ASR experiments. Our study explores shared structures embedded in a large collection of speech data spanning over a number of spoken languages in order to establish a common set of universal phone models that can be used for large vocabulary ASR of all the languages seen or unseen during training. Language-universal and language-adaptive models are compared with language-specific models, and the comparison results show that in many cases it is possible to build general-purpose language-universal and language-adaptive acoustic models that outperform language-specific ones if the set of shared units, the structure of shared states, and the shared acoustic-phonetic properties among different languages can be properly utilized. Specifically, our results demonstrate that when the context coverage is poor in language-specific training, we can use one tenth of the adaptation data to achieve equivalent performance in cross-lingual speech recognition.
Li Deng 0001, Dong Yu 0001, Yifan Gong 0001, Alex Acero, Chin-Hui Lee 0001
ICASSP3
2009 Using collective information in semi-supervised learning for speech recognition
abstract
Training accurate acoustic models typically requires a large amount of transcribed data, which can be expensive to obtain. In this paper, we describe a novel semi-supervised learning algorithm for automatic speech recognition. The algorithm determines whether a hypothesized transcription should be used in the training by taking into consideration collective information from all utterances available instead of solely based on the confidence from that utterance itself. It estimates the expected entropy reduction each utterance and transcription pair may cause to the whole unlabeled dataset and choose the ones with the positive gains. We compare our algorithm with existing confidence-based semi-supervised learning algorithm and show that the former can consistently outperform the latter when the same amount of utterances is selected into the training set. We also indicate that our algorithm may determine the cutoff-point in a principled way by demonstrating that the point it finds is very close to the achievable peak point.
Balakrishnan Varadarajan, Dong Yu 0001, Li Deng 0001, Alex Acero
ICASSP2
2009 Maximizing global entropy reduction for active learning in speech recognition
abstract
We propose a new active learning algorithm to address the problem of selecting a limited subset of utterances for transcribing from a large amount of unlabeled utterances so that the accuracy of the automatic speech recognition system can be maximized. Our algorithm differentiates itself from earlier work in that it uses a criterion that maximizes the lattice entropy reduction over the whole dataset. We introduce our criterion, show how it can be simplified and approximated, and describe the detailed algorithm to optimize the criterion. We demonstrate the effectiveness of our new algorithm with directory assistance data collected under the real usage scenarios and show that our new algorithm consistently outperforms the confidence based approach by a significant margin. Using the algorithm cuts the number of utterances needed for transcribing by 50% to achieve the same recognition accuracy obtained using the confidence-based approach, and by 60% compared to the random sampling approach.
Balakrishnan Varadarajan, Dong Yu 0001, Li Deng 0001, Alex Acero
ICASSP2
2009 Discriminative pronounciation learning using phonetic decoder and minimum-classification-error criterion
abstract
In this paper, we report our recent research aimed at improving the pronunciation-modeling component of a speech recognition system designed for mobile voice search. Our new discriminative learning technique overcomes the limitation of the traditional ways of introducing alternative pronunciations that often enlarge confusability across different lexical items. Instead, we make use of a phonetic recognizer to generate pronunciation candidates, which are then evaluated and selected using the global minimum-classification-error measure, guaranteeing a reduction of the training-set error rate after introducing alternative pronunciations. A maximum entropy approach is subsequently used to learn the weight parameters of the selected pronunciation candidates. Our experimental results demonstrate the effectiveness of the discriminative pronunciation learning technique in a real-world speech recognition task where pronunciation of business names presents special difficulty for high-accuracy speech recognition.
Oriol Vinyals, Li Deng 0001, Dong Yu 0001, Alex Acero
ICASSP3
2009 Cross-lingual speech recognition under runtime resource constraints
abstract
This paper proposes and compares four cross-lingual and bilingual automatic speech recognition techniques under the constraint that only the acoustic model (AM) of the native language is used at runtime. The first three techniques fall into the category of lexicon conversion where each phoneme sequence (PHS) in the foreign language (FL) lexicon is mapped into the native language (NL) phoneme sequence. The first technique determines the PHS mapping through the international phonetic alphabet (IPA) features; The second and third techniques are data-driven. They determine the mapping by converting the PHS into corresponding context-independent and context-dependent hidden Markov models (HMMs) respectively and searching for the NL PHS with the least Kullback-Leibler divergence (KLD) between the HMMs. The fourth technique falls into the category of AM merging where the FL's AM is merged into the NL's AM by mapping each senone in the FL's AM to the senone in the NL's AM with the minimum KLD. We discuss the strengths and limitations of each technique developed, report empirical evaluation results on recognizing English utterances with a Korean recognizer, and demonstrate the high correlation between the average KLD and the word error rate (WER). The results show that the AM merging technique performs the best, achieving 60% relative WER reduction over the IPA-based technique.
Dong Yu 0001, Li Deng 0001, Jian Wu 0027, Yifan Gong 0001, Alex Acero
ICASSP1
2009 Hidden conditional random field with distribution constraints for phone classification
abstract
We advance the recently proposed hidden conditional random field (HCRF) model by replacing the moment constraints (MCs) with the distribution constraints (DCs). We point out that the distribution constraints are the same as the traditional moment constraints for the binary features but are able to better regularize the probability distribution of the continuousvalued features than the moment constraints. We show that under the distribution constraints the HCRF model is no longer log-linear but embeds the model parameters in non-linear functions. We provide an effective solution to the resulting more difficult optimization problem by converting it to the traditional log-linear form at a higher-dimensional space of features exploiting cubic spline. We demonstrate that a 20.8% classification error rate (CER) can be achieved on the TIMIT phone classification task using the HCRF-DC model. This result is superior to any published single-system result on this heavily evaluated task including the HCRF-MC model, the discriminatively trained HMMs, and the large-margin HMMs using the same features. Index Terms: hidden conditional random field, maximum entropy, moment constraint, distribution constraint, phone
Dong Yu 0001, Li Deng 0001, Alex Acero
INTERSPEECH1
2009 A unified framework of HMM adaptation with joint compensation of additive and convolutive distortions
Jinyu Li 0001, Li Deng 0001, Dong Yu 0001, Yifan Gong 0001, Alex Acero
Comput. Speech Lang.3
2009 Using continuous features in the maximum entropy model
Dong Yu 0001, Li Deng 0001, Alex Acero
Pattern Recognit. Lett.1
2009 A Novel Framework and Training Algorithm for Variable-Parameter Hidden Markov Models
abstract
We propose a new framework and the associated maximum-likelihood and discriminative training algorithms for the variable-parameter hidden Markov model (VPHMM) whose mean and variance parameters vary as functions of additional environment-dependent conditioning parameters. Our framework differs from the VPHMM proposed by Cui and Gong (2007) in that piecewise spline interpolation instead of global polynomial regression is used to represent the dependency of the HMM parameters on the conditioning parameters, and a more effective functional form is used to model the variances. Our framework unifies and extends the conventional discrete VPHMM. It no longer requires quantization in estimating the model parameters and can support both parameter sharing and instantaneous conditioning parameters naturally. We investigate the strengths and weaknesses of the model on the Aurora-3 corpus. We show that under the well-matched condition the proposed discriminatively trained VPHMM outperforms the conventional HMM trained in the same way with relative word error rate (WER) reduction of 19% and 15%, respectively, when only mean is updated and when both mean and variances are updated.
Dong Yu 0001, Li Deng 0001, Yifan Gong 0001, Alex Acero
IEEE Trans. Speech Audio Process.1
2008 HMM adaptation using a phase-sensitive acoustic distortion model for environment-robust speech recognition
abstract
In this paper, we present a new approach to HMM adaptation that jointly compensates for additive and convolutive acoustic distortion in environment-robust speech recognition. The hallmark of our new approach is the use of a nonlinear, phase-sensitive model of acoustic distortion that captures phase asynchrony between clean speech and the mixing noise. In the first step of the developed algorithm, both the static and dynamic portions of the noise and channel parameters are estimated in the cepstral domain, using the speech recognizer’s “feedback” information and the vector-Taylor-series linearization technique on the nonlinear phase-sensitive model. In the second step, the estimated noise and channel parameters are used to effectively adapt the static and dynamic portions of the HMM means and variances also using the linearized phase-sensitive acoustic distortion model. In the experimental evaluation using the standard Aurora 2 task, the proposed new algorithm achieves 93.3% accuracy using the clean-trained complex HMM backend as the baseline system for unsupervised HMM adaptation. This reaches the highest performance number in the literature on this task with clean-trained HMM model. The experimental results show that the phase term, which was missing in all previous HMM-adaptation work, contributes significantly to the achieved high recognition accuracy.
Jinyu Li 0001, Li Deng 0001, Dong Yu 0001, Yifan Gong 0001, Alex Acero
ICASSP3
2008 Adaptation of compressed HMM parameters for resource-constrained speech recognition
abstract
Recently, we successfully developed and reported a new unsupervised online adaptation technique, which jointly compensates for additive and convolutive distortions with vector Taylor series (JAC/VTS), to adjust (uncompressed) HMMs under acoustically distorted environments [1]. In this paper, we extend that technique to adapt compressed HMMs using JAC/VTS where limited computation and/or memory resources are available for speech recognition (e.g., on mobile devices). Subspace coding (SSC) is developed and used to quantize each dimension of the multivariate Gaussians in the compressed HMMs. Three algorithmic design options are proposed and evaluated that combine SSC with JAC/VTS, where three different types of tradeoffs are made between recognition accuracy and the required computation/memory/storage resources. The strengths and weaknesses of these three options are discussed and shown on the Aurora2 task of noise-robust speech recognition. The first option greatly reduces the storage space and gives 93.2% accuracy, which is the same as the baseline accuracy but with little reduction in the run-time computation/memory cost. The second option reduces about 79.9% of the computation cost and about 33.5% of the memory requirement at a very small price of 0.5% decrease of accuracy (to 92.7%). The third option cuts about 89.2% of the computation cost and about 65.5% of the memory requirement while reducing recognition accuracy by 2.7% (to 90.5%).
Jinyu Li 0001, Li Deng 0001, Dong Yu 0001, Jian Wu 0027, Yifan Gong 0001, Alex Acero
ICASSP3
2008 A minimum-mean-square-error noise reduction algorithm on Mel-frequency cepstra for robust speech recognition
abstract
We present a non-linear feature-domain noise reduction algorithm based on the minimum mean square error (MMSE) criterion on Mel-frequency cepstra (MFCC) for environment-robust speech recognition. Distinguishing from the MMSE enhancement in log spectral amplitude proposed by Ephraim and Malah (E&M) [7], the new algorithm presented in this paper develops the suppression rule that applies to power spectral magnitude of the filter-banks’ outputs and to MFCC directly, making it demonstrably more effective in noise-robust speech recognition. The noise variance in the new algorithm contains a significant term resulting from instantaneous phase asynchrony between clean speech and mixing noise, missing in the E&M algorithm. Speech recognition experiments on the standard Aurora-3 task demonstrate a reduction of word error rate by 48% against the ICSLP02 baseline, by 26% against the cepstral mean normalization baseline, and by 13% against the conventional E&M log-MMSE noise suppressor. The new algorithm is also much more efficient than E&M noise suppressor since the number of the channels in the Mel-frequency filter bank is much smaller (23 in our case) than the number of bins in the FFT domain (256). The results also show that our algorithm performs slightly better than the ETSI AFE on the well-matched and mid-mismatched settings.
Dong Yu 0001, Li Deng 0001, Jasha Droppo, Jian Wu 0027, Yifan Gong 0001, Alex Acero
ICASSP1
2008 Discriminative training of variable-parameter HMMs for noise robust speech recognition
abstract
We propose a new type of variable-parameter hidden Markov model (VPHMM) whose mean and variance parameters vary each as a continuous function of additional environmentdependent parameters. Different from the polynomialfunction-based VPHMM proposed by Cui and Gong (2007), the new VPHMM uses cubic splines to represent the dependency of the means and variances of Gaussian mixtures on the environment parameters. Importantly, the new model no longer requires quantization in estimating the model parameters and it supports parameter sharing and instantaneous conditioning parameters directly. We develop and describe a growth-transformation algorithm that discriminatively learns the parameters in our cubic-splinebased VPHMM (CS-VPHMM), and evaluate the model on the Aurora-3 corpus with our recently developed MFCC-MMSE noise suppressor applied. Our experiments show that the proposed CS-VPHMM outperforms the discriminatively trained and maximum-likelihood trained conventional HMMs with relative word error rate (WER) reduction of 14 % and 20 % respectively under the well-matched conditions when both mean and variances are updated. Index Terms: speech recognition, variable-parameter hidden Markov model, discriminative training, cubic spline, growth transformation 1.
Dong Yu 0001, Li Deng 0001, Yifan Gong 0001, Alex Acero
INTERSPEECH1
2008 Parameter clustering and sharing in variable-parameter HMMs for noise robust speech recognition
abstract
Recently we proposed a cubic-spline-based variableparameter hidden Markov model (CS-VPHMM) whose mean and variance parameters vary according to some cubic spline functions of additional environment-dependent parameters. We have shown good properties of the CS-VPHMM and demonstrated on the Aurora-3 corpus that MCE-trained CSVPHMM greatly outperforms the MCE-trained conventional HMM at the cost of increased total number of model parameters. In this paper, we propose to share spline functions across different Gaussian mixture components to reduce the total number of model parameters and develop a clustering algorithm to do so. We demonstrate the effectiveness of our parameter clustering and sharing algorithm for the CSVPHMM on Aurora-3 corpus and show that proper parameter sharing can reduce the number of parameters from 4 times of that used in the conventional HMM to 1.13 times and still get 18% relative WER reduction over the MCE trained conventional HMM under the well-matched condition. Effective parameter sharing makes the CS-VPHMM an attractive model for noise robustness.
Dong Yu 0001, Li Deng 0001, Yifan Gong 0001, Alex Acero
INTERSPEECH1
2008 Large-margin minimum classification error training: A theoretical risk minimization perspective
Dong Yu 0001, Li Deng 0001, Xiaodong He 0001, Alex Acero
Comput. Speech Lang.1
2008 An Integrative and Discriminative Technique for Spoken Utterance Classification
abstract
Traditional methods of spoken utterance classification (SUC) adopt two independently trained phases. In the first phase, an automatic speech recognition (ASR) module returns the most likely sentence for the observed acoustic signal. In the second phase, a semantic classifier transforms the resulting sentence into the most likely semantic class. Since the two phases are isolated from each other, such traditional SUC systems are suboptimal. In this paper, we present a novel integrative and discriminative learning technique for SUC to alleviate this problem, and thereby, reduce the semantic classification error rate (CER). Our approach revolves around the effective use of theN-best lists generated by the ASR module to reduce semantic classification errors. TheN-best list sentences are first rescored using all the available knowledge sources. Then, the sentence that is most likely to helps reduce the CER are extracted from theN-best lists as well as those sentences that are most likely to increase the CER. These sentences are used to discriminatively train the language and semantic-classifier models to minimize the overall semantic CER. Our experiments resulted in a reduction of CER from its initial value of 4.92% to 4.04% in the standard ATIS task.
Sibel Yaman, Li Deng 0001, Dong Yu 0001, Ye-Yi Wang, Alex Acero
IEEE Trans. Speech Audio Process.3
2008 Robust Speech Recognition Using a Cepstral Minimum-Mean-Square-Error-Motivated Noise Suppressor
abstract
We present an efficient and effective nonlinear feature-domain noise suppression algorithm, motivated by the minimum-mean-square-error (MMSE) optimization criterion, for noise-robust speech recognition. Distinguishing from the log-MMSE spectral amplitude noise suppressor proposed by Ephraim and Malah (E&M), our new algorithm is aimed to minimize the error expressed explicitly for the Mel-frequency cepstra instead of discrete Fourier transform (DFT) spectra, and it operates on the Mel-frequency filter bank's output. As a consequence, the statistics used to estimate the suppression factor become vastly different from those used in the E&M log-MMSE suppressor. Our algorithm is significantly more efficient than the E&M's log-MMSE suppressor since the number of the channels in the Mel-frequency filter bank is much smaller (23 in our case) than the number of bins (256) in DFT. We have conducted extensive speech recognition experiments on the standard Aurora-3 task. The experimental results demonstrate a reduction of the recognition word error rate by 48% over the standard ICSLP02 baseline, 26% over the cepstral mean normalization baseline, and 13% over the popular E&M's log-MMSE noise suppressor. The experiments also show that our new algorithm performs slightly better than the ETSI advanced front end (AFE) on the well-matched and mid-mismatched settings, and has 8% and 10% fewer errors than our earlier SPLICE (stereo-based piecewise linear compensation for environments) system on these settings, respectively.
Dong Yu 0001, Li Deng 0001, Jasha Droppo, Jian Wu 0027, Yifan Gong 0001, Alex Acero
IEEE Trans. Speech Audio Process.1
2007 High-performance hmm adaptation with joint compensation of additive and convolutive distortions via Vector Taylor Series
abstract
In this paper, we present our recent development of a model-domain environment-robust adaptation algorithm, which demonstrates high performance in the standard Aurora 2 speech recognition task. The algorithm consists of two main steps. First, the noise and channel parameters are estimated using a nonlinear environment distortion model in the cepstral domain, the speech recognizer’s “feedback” information, and the Vector-Taylor-Series (VTS) linearization technique collectively. Second, the estimated noise and channel parameters are used to adapt the static and dynamic portions of the HMM means and variances. This two-step algorithm enables Joint compensation of both Additive and Convolutive distortions (JAC). In the experimental evaluation using the standard Aurora 2 task, the proposed JAC/VTS algorithm achieves 91.11% accuracy using the clean-trained simple HMM backend as the baseline system for the model adaptation. This represents high recognition performance on this task without discriminative training of the HMM system. Detailed analysis on the experimental results shows that adaptation of the dynamic portion of the HMM mean and variance parameters is critical to the success of our algorithm.
Jinyu Li 0001, Li Deng 0001, Dong Yu 0001, Yifan Gong 0001, Alex Acero
ASRU3
2007 Use of Differential Cepstra as Acoustic Features in Hidden Trajectory Modeling for Phonetic Recognition
abstract
The earlier version of the hidden trajectory model (HTM) for speech dynamics which predicts the "static" cepstra as the observed acoustic feature is generalized to one which predicts joint static cepstra and their temporal differentials (i.e., delta cepstra). The formulation of this generalized HTM is presented in the generative-modeling framework, enabling efficient computation of the joint likelihood for both static and delta cepstral sequences as the acoustic features given the model. The parameter estimation techniques for the new model are developed and presented, giving closed-form estimation formulas after the use of vector Taylor series approximation. We show principled generalization from the earlier static-cepstra HTM to the new static/delta-cepstra HTM not only in terms of model formulations but also in terms of their respective analytical forms in (monophone) parameter estimation. Experimental results on the standard TIMIT phonetic recognition task demonstrate recognition accuracy improvement over the earlier best HTM system, both significantly better than state-of-the-art triphone HMM systems.
Li Deng 0001, Dong Yu 0001
ICASSP (4)2
2007 A Discriminative Training Framework using N-Best Speech Recognition Transcriptions and Scores for Spoken Utterance Classification
abstract
In this paper, we propose a novel discriminative training approach to spoken utterance classification (SUC). The ultimate objective of the SUC task, originally developed to map a spoken speech utterance into the most appropriate semantic class, is to minimize the classification error rate (CER). Conventionally, a two-phase approach is adapted, in which the first phase is the ASR transcription phase, and the second phase is the semantic classification phase. In the proposed framework, the classification error rate is approximated as differentiable functions of the language and classifier model parameters. Furthermore, in order to exploit all the available information from the first phase, class-specific discriminant functions are defined based on score functions derived from the N-best lists. Our experimental results on the standard ATIS database indicate a notable reduction in CER from the earlier best result on the identical task. The proposed framework achieved a reduction of CER from 4.92% to 4.04%.
Sibel Yaman, Li Deng 0001, Dong Yu 0001, Ye-Yi Wang, Alex Acero
ICASSP (4)3
2007 Large-Margin Minimum Classification Error Training for Large-Scale Speech Recognition Tasks
abstract
Recently, we have developed a novel discriminative training method named large-margin minimum classification error (LM-MCE) training that incorporates the idea of discriminative margin into the conventional minimum classification error (MCE) training method. In our previous work, this novel approach was formulated specifically for the MCE training using the sigmoid loss function and its effectiveness was demonstrated on the TIDIGITS task alone. In this paper two additional contributions are made. First, we formulate LM-MCE as a Bayes risk minimization problem whose loss function not only includes empirical error rates but also a margin-bound risk. This new formulation allows us to extend the same technique to a wide variety of MCE based training. Second, we have successfully applied LM-MCE training approach to the Microsoft internal large vocabulary telephony speech recognition task (with 2000 hours of training data and 120K of vocabulary) and achieved significant recognition accuracy improvement across-the-board. To our best knowledge, this is the first time that the large-margin approach is demonstrated to be successful in large-scale speech recognition tasks.
Dong Yu 0001, Li Deng 0001, Xiaodong He 0001, Alex Acero
ICASSP (4)1
2007 Voicepedia: towards speech-based access to unstructured information
abstract
Currently there are no dialog systems that enable purely voice-based access to the unstructured information on websites such as Wikipedia. Such systems could be revolutionary for non-literate users in the developing world. To investigate interface issues in such a system, we developed VoicePedia, a telephone-based dialog system for searching and browsing Wikipedia. In this paper, we present the system, as well as a user study comparing the use of VoicePedia to SmartPedia, a Smartphone GUI-based alternative. Keyword entry through the voice interface was significantly faster, while search result navigation, and page browsing were significantly slower. Although users preferred the GUI-based interface, task success rates between both systems were comparable – a promising result for regions where Smartphones and data plans are not viable. Index Terms: dialog system, information access
J. Sherwani, Dong Yu 0001, Tim Paek, Mary Czerwinski, Yun-Cheng Ju, Alex Acero
INTERSPEECH2
2007 Confidence measures for voice search applications
abstract
Voice search is the technology underlying many spoken dialog applications that enable users to access information using spoken queries. This paper reviews voice search technology, and proposes a new and effective method for computing semantic confidence measures. It explores the use of maximum entropy classifiers as confidence models, and investigates a feature selection algorithm that leads to an effective subset of prominent features for the classifier. The experimental results on a directory assistance application show that the reduced feature set not only makes the model more effective in handling different recognition and search engine combinations, but also results in a very informative confidence measure that is closely correlated with the actual voice search accuracy. Index Terms: voice search, directory assistance, confidence measure, Tf-Idf vector space model, maximum entropy model. 1.
Ye-Yi Wang, Dong Yu 0001, Yun-Cheng Ju, Geoffrey Zweig, Alex Acero
INTERSPEECH2
2007 Handling phonetic context and speaker variation in a structure-based speech recognizer
abstract
Recently we have developed a novel type of structure-based speech recognizer, which uses parameterized, non-recursive “hidden ” trajectory model of vocal tract resonances (VTR) or formants to capture the dynamic structure of long-range speech coarticulation and reduction. The underlying model of this recognizer carries out bi-directional FIR filtering on the piecewise constant sequences of the VTR targets. In this paper, we elaborate on two key aspects of the model. First, the phonetic context controls the movement direction and thus the formation of the VTR trajectories. This provides “structured ” context dependency for speech acoustics without using context dependent parameters as required by HMMs. Second, VTR targets as the key context-independent parameters of the model vary across speakers. We describe an effective target-value normalization algorithm that can be applied to both training and unknown test speakers. We report experimental results demonstrating the effectiveness of the normalization algorithm in the context of structure-based speech recognition. We also provide computational analysis on the HTM-based speech decoder. Index Terms: hidden trajectory model, phonetic contexts, normalization, vocal tract resonance, targets
Dong Yu 0001, Li Deng 0001, Alex Acero
INTERSPEECH1
2007 Automated directory assistance system - from theory to practice
abstract
The automated directory assistance system (ADAS) is traditionally formulated as an automatic speech recognition (ASR) problem. Recently, it has been formulated as a voice search problem, where a spoken utterance is firstly converted into text, which in turn is used to search for the listing. In this paper, we focus on the design and development of the utterance-to-listing component of ADAS. We show that many theoretical and practical issues need to be resolved when applying the basic idea of voice search to the development of ADAS. We share our experiences in addressing these issues, especially in pre-processing the listing database, generating a high performance LM, and developing efficient, accurate, and robust search algorithms. Field tests of our prototype system indicate that an 81 % task completion rate can be achieved. Index Terms: speech recognition, directory assistance, voice search, TFIDF, spoken dialog system, vector space model
Dong Yu 0001, Yun-Cheng Ju, Ye-Yi Wang, Geoffrey Zweig, Alex Acero
INTERSPEECH1
2007 The voice-rate dialog system for consumer ratings
abstract
Voice-Rate is an automated dialog system which provides access to over one million ratings of products and businesses. By calling a toll-free number, consumers can access ratings for products, national businesses such as airlines, and local businesses such as restaurants. Voice-Rate also has a facility for recording and analyzing ratings that are given over the phone. The service has been primed with ratings taken from a variety of web sources, and we are augmenting these with user ratings. Voice-Rate can be accessed by dialing 1-877-456-DATA. 1
Geoffrey Zweig, Patrick Nguyen, Yun-Cheng Ju, Ye-Yi Wang, Dong Yu 0001, Alex Acero
INTERSPEECH5
2007 Improving the quality of alerts and predicting intruder's next goal with Hidden Colored Petri-Net
Dong Yu 0001, Deborah A. Frincke
Comput. Networks1
2007 Speaker-adaptive learning of resonance targets in a hidden trajectory model of speech coarticulation
Dong Yu 0001, Li Deng 0001, Alex Acero
Comput. Speech Lang.1
2006 N-Gram Based Filler Model for Robust Grammar Authoring
abstract
We propose a technique for rapid speech application development that generates robust semantic context-free grammars (CFG) given rigid CFGs as input. Users' speech does not always conform to rigid CFGs, so robust grammars improve the caller's experience. Our system takes a simple CFG and then generates a hybrid n-gram/CFG that is written in the W3C SRGS format and thus can run in many standard automatic speech recognition engines. The hybrid network leverages an application-independent word n-gram which can be shared across different applications. In addition, our tool allows developers to provide a few example sentences to adapt the n-gram for improved accuracy. Our experiments show the robust CFG has no loss in accuracy for test utterances that can be covered by the rigid CFG, but offers large improvements for cases where the user's sentence cannot be covered by the rigid CFG. It also has a much better rejection for utterances that contain no slot at all. With a few example sentences for adaptation, our robust CFG can achieve the recognition accuracy close to the class-based n-gram LM customized for the application.
Dong Yu 0001, Yun-Cheng Ju, Ye-Yi Wang, Alex Acero
ICASSP (1)1
2006 A time-synchronous phonetic decoder for a long-contextual-Span hidden trajectory model
abstract
A novel time-synchronous decoder, designed specifically for a Hidden Trajectory Model (HTM) whose likelihood score computation depends on long-span phonetic contexts, is presented. HTM is a recently developed acoustic model aimed to capture the underlying dynamic structure of speech coarticulation and reduction using a compact set of parameters. The long-span nature of the HTM had posed a great technical challenge for developing efficient search algorithms for full evaluation of the model. Taking on the challenge, the decoding algorithm is developed to deal effectively with the exponentially increased search space by HTMspecific techniques for hypothesis representation, word-ending recombination, and hypothesis pruning. Experimental results obtained on the TIMIT phonetic recognition task are reported, extending our earlier HTM evaluation paradigms based on N-best and A * lattice rescoring.
Li Deng 0001, Dong Yu 0001, Alex Acero
INTERSPEECH3
2006 Use of incrementally regulated discriminative margins in MCE training for speech recognition
abstract
In this paper, we report our recent development of a novel discriminative learning technique which embeds the concept of discriminative margin into the well established minimum classification error (MCE) method. The idea is to impose an incrementally adjusted “margin ” in the loss function of MCE algorithm so that not only error rates are minimized but also discrimination “robustness ” between training and test sets is maintained. Experimental evaluation shows that the use of the margin improves a state-of-the-art MCE method by reducing 17 % digit errors and 19 % string errors in the TIDigits recognition task. The string error rate of 0.55 % and digit error rate of 0.19 % we have obtained are the best-ever results reported on this task in the literature. Index Terms: discriminative training, margin, minimum error 1.
Dong Yu 0001, Li Deng 0001, Xiaodong He 0001, Alex Acero
INTERSPEECH1
2006 An effective and efficient utterance verification technology using word n-gram filler models
abstract
In this paper we propose a novel, effective, and efficient utterance verification (UV) technology for access control in the interactive voice response (IVR) systems. The key of our approach is to construct a context-free grammar by using the secret answer to a question and a word N-gram based filler model. The N-gram filler provides rich alternatives to the secret answer and can potentially improve the accuracy of the UV task. It can also absorb carrier words used by callers and thus can improve the robustness. We also propose using a predictor based on the best alternative to calculate the confidence. We show detailed experimental results on a tough UV test set that contains 930 positive and 930 negative cases and discuss types of questions that are suitable for the UV task. We demonstrate that our approach can achieve a 2.14% equal error rate (EER) on average and 0.8 % false accept rate if the false reject rate is 2.6 % and above. This is a 49 % EER reduction compared with the approaches using acoustic fillers, and a 72 % EER reduction compared with the posterior probability based confidence measurement. Index Terms: utterance verification, filler model, word spotting, confidence measure
Dong Yu 0001, Yun-Cheng Ju, Alex Acero
INTERSPEECH1
2006 A lattice search technique for a long-contextual-span hidden trajectory model of speech
Dong Yu 0001, Li Deng 0001, Alex Acero
Speech Commun.1
2006 A bidirectional target-filtering model of speech coarticulation and reduction: two-stage implementation for phonetic recognition
abstract
A structured generative model of speech coarticulation and reduction is described with a novel two-stage implementation. At the first stage, the dynamics of formants or vocal tract resonances (VTRs) in fluent speech is generated using prior information of resonance targets in the phone sequence, in absence of acoustic data. Bidirectional temporal filtering with finite-impulse response (FIR) is applied to the segmental target sequence as the FIR filter's input, where forward filtering produces anticipatory coarticulation and backward filtering produces regressive coarticulation. The filtering process is shown also to result in realistic resonance-frequency undershooting or reduction for fast-rate and low-effort speech in a contextually assimilated manner. At the second stage, the dynamics of speech cepstra are predicted analytically based on the FIR-filtered and speaker-adapted VTR targets, and the prediction residuals are modeled by Gaussian random variables with trainable parameters. The combined system of these two stages, thus, generates correlated and causally related VTR and cepstral dynamics, where phonetic reduction is represented explicitly in the hidden resonance space and implicitly in the observed cepstral space. We present details of model simulation demonstrating quantitative effects of speaking rate and segment duration on the magnitude of reduction, agreeing closely with experimental measurement results in the acoustic-phonetic literature. This two-stage model is implemented and applied to the TIMIT phonetic recognition task. Using the N-best (N=2000) rescoring paradigm, the new model, which contains only context-independent parameters, is shown to significantly reduce the phone error rate of a standard hidden Markov model (HMM) system under the same experimental conditions.
Li Deng 0001, Dong Yu 0001, Alex Acero
IEEE Trans. Speech Audio Process.2
2006 Structured speech modeling
abstract
Modeling dynamic structure of speech is a novel paradigm in speech recognition research within the generative modeling framework, and it offers a potential to overcome limitations of the current hidden Markov modeling approach. Analogous to structured language models where syntactic structure is exploited to represent long-distance relationships among words , the structured speech model described in this paper makes use of the dynamic structure in the hidden vocal tract resonance space to characterize long-span contextual influence among phonetic units. A general overview is provided first on hierarchically classified types of dynamic speech models in the literature. A detailed account is then given for a specific model type called the hidden trajectory model, and we describe detailed steps of model construction and the parameter estimation algorithms. We show how the use of resonance target parameters and their temporal filtering enables joint modeling of long-span coarticulation and phonetic reduction effects. Experiments on phonetic recognition evaluation demonstrate superior recognizer performance over a modern hidden Markov model-based system. Error analysis shows that the greatest performance gain occurs within the sonorant speech class.
Li Deng 0001, Dong Yu 0001, Alex Acero
IEEE Trans. Speech Audio Process.2
2005 A Hidden Trajectory Model with Bi-directional Target-Filtering: Cascaded vs. Integrated Implementation for Phonetic Recognition
abstract
We present a novel acoustic model of speech, based on statistical hidden trajectory modeling (HTM) with bi-directional vocal tract resonance (VTR) target filtering, for speech recognition. The HTM consists of two stages of the generative process of speech: from the phone sequence to VTR dynamics and then to the cepstrum-based acoustic observation. Two types of model implementation are detailed, one with straightforward two-stage cascading, and another which integrates over the statistical distribution of VTR in model construction and in computing acoustic likelihood. With the use of first-order Taylor series approximation to the nonlinearity in the VTR-to-cepstrum prediction component of HTM, the acoustic likelihood is established in an analytical form. It is a Gaussian with the time-varying mean that gives structured long-span context dependence over the entire utterance, and with the dynamically adjusted variance proportional to the squared "local slope" in the nonlinear mapping function from VTR to cepstrum. When the HTM parameters are trained via maximizing this "integrated" likelihood, dramatic reduction of an upper error bound is achieved in the standard TIMIT phonetic recognition task using a large-scale N-best rescoring paradigm.
Li Deng 0001, Dong Yu 0001, Alex Acero
ICASSP (1)3
2005 Maximum Entropy Based Generic Filter for Language Model Adaptation
abstract
Language model (LM) adaptation has been shown to be very important in reducing the word error rate (WER) in task specific speech recognition systems. Adaptation data collected in the real world, however, usually contain large amounts of non-dictated text, such as email headers, long URL, code fragments, included reply, signature, etc., that the user never dictates. Adapting with these data may corrupt the LM. We propose a maximum entropy (MaxEnt) based filter to remove a variety of non-dictated words from the adaptation data and improve the effectiveness of the LM adaptation. We argue that this generic filter is language independent and efficient. We describe the design of the filter, and show that the use of the filter can give us 10% relative WER reduction over LM adaptation without filtering, and 22% relative WER reduction over the unadapted LM in an English email dictation task.
Dong Yu 0001, Milind Mahajan, Peter Mau, Alex Acero
ICASSP (1)1
2005 Learning statistically characterized resonance targets in a hidden trajectory model of speech coarticulation and reduction
abstract
We report our new development of a hidden trajectory model for co-articulated, time-varying patterns of speech. The model uses bi-directional filtering of vocal tract resonance targets to jointly represent contextual variation and phonetic reduction in speech acoustics. A novel maximum-likelihood-based learning algorithm is presented that accurately estimates the distributional parameters of the resonance targets. The results of the estimates are analyzed and shown to be consistent with all the relevant acoustic-phonetic facts and intuitions. Phonetic recognition experiments demonstrate that the model with more rigorous target training outperforms the most recent earlier version of the model, producing 17.5% fewer errors in N-best rescoring.
Li Deng 0001, Dong Yu 0001, Alex Acero
INTERSPEECH2
2005 Evaluation of a long-contextual-Span hidden trajectory model and phonetic recognizer using a* lattice search
abstract
A long-contextual-span Hidden Trajectory Model (HTM) developed recently captures underlying dynamic structure of speech coarticulation and reduction using a highly compact set of context-independent parameters. However, the longspan nature of the HTM makes it difficult to develop efficient search algorithms for its full evaluation. In this paper, we describe our initial effort in meeting this challenge. The basic search algorithm is time-asynchronous A*. Given the structural complexity of the long-span HTM, special considerations are needed to take into account the fact that the HTM score for each frame depends on the model parameters associated with a variable number of adjacent phones. Specifically, we present details on how the nodes and links in the lattices are expanded via look-ahead, how the A* heuristics are estimated, and what pruning strategies are applied to speed up the search. The experiments on TIMIT phonetic recognition show the capability of our newly developed lattice search algorithm in evaluating billions of hypotheses based on long-span HTM scores. The results significantly extend our earlier work from N-best rescoring to A * search over lattices. 1.
Dong Yu 0001, Li Deng 0001, Alex Acero
INTERSPEECH1
2005 Semiautomatic Improvements of System-Initiative Spoken Dialog Applications Using Interactive Clustering
abstract
While many successful spoken dialog systems have been deployed over telephone networks in recent years, the high cost of developing such applications has led to limited adoption. Despite large research efforts in user-initiative and mixed-initiative systems, most commercial applications follow a system initiative approach because they are simpler to design and are found to work adequately. Yet, even designing such system-initiative spoken dialog systems has proven costly when compared with simpler touchtone systems. To address this issue, we describe in this paper our efforts in building diagnostics tools to let nonexperienced speech developers write usable applications without the need for transcribing calls. Our approach consists of two steps. In the first step, we cluster calls based on Question/Answer (QA) states and transitions, analyze the success rates associated with each QA state and transition, and identify the most problematic QA states and transitions based on a criterion we call Arc Cut Gain in Success Rate (ACGSR). In the second step, we cluster calls associated with problematic QA transitions through an approach we term Interactive Clustering (IC). The purpose of this step is to automatically cluster calls that are similar to those already labeled by the developers to maximize productivity. Experiments on an internal auto-attendant application show that our approach can significantly reduce the time and effort needed to identify problems in spoken dialog applications.
Dong Yu 0001, Alex Acero
IEEE Trans. Speech Audio Process.1
2004 A Novel Framework for Alert Correlation and Understanding
Dong Yu 0001, Deborah A. Frincke
ACNS1
2004 Unsupervised learning from users' error correction in speech dictation
abstract
We propose an approach to adapting automatic speech recognition systems used in dictation systems through unsupervised learning from users ’ error correction. Three steps are involved in the adaptation: 1) infer whether the user is correcting a speech recognition error or simply editing the text, 2) infer what the mostpossible cause of the error is, and 3) adapt the system accordingly. To adapt the system effectively, we introduce an enhanced two-pass pronunciation learning algorithm that utilizes the output from both an ngram phoneme recognizer and a Letter-to-Sound component. Our experiments show that we can obtain greater than 10% relative word error rate reduction using the approaches we proposed. Learning new words gives the largest performance gain while adapting pronunciations and using a cache language model also produce a small gain. 1.
Dong Yu 0001, Mei-Yuh Hwang, Peter Mau, Alex Acero, Li Deng 0001
INTERSPEECH1
2003 Improved name recognition with user modeling
abstract
Speech recognition of names in Personal Information Management (PIM) systems is an important yet difficult task. The difficulty arises from various sources: the large number of possible names that users may speak, different ways a person may be referred to, ambiguity when only first names are used, and mismatched pronunciations. In this paper we present our recent work on name recognition with User Modeling (UM), i.e., automatic modeling of user’s behavior patterns. We show that UM and our learning algorithm lead to significant improvement in the perplexity, Out Of Vocabulary rate, recognition speed, and accuracy of the top recognized candidate. The use of an exponential window reduces the perplexity by more than 30%. 1.
Dong Yu 0001, Kuansan Wang, Milind Mahajan, Peter Mau, Alex Acero
INTERSPEECH1