VLDB 2026 Research / reviewers in the wild / expert
Chao Zhang 0031
dblp:94/3019-31
· DBLP profile ↗
108ranked-venue papers
15as first author
75since 2021 · last 2025
0000-0002-7730-5131ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 79 · 14 first-author · 48 since 2021Artificial intelligence and machine learning · 68 · 6 first-author · 48 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and DescriptionsabstractThis paper explores a novel perspective to speech quality assessment by leveraging natural language descriptions, offering richer, more nuanced insights than traditional numerical scoring methods. Natural language feedback provides instructive recommendations and detailed evaluations, yet existing datasets lack the comprehensive annotations needed for this approach. To bridge this gap, we introduce QualiSpeech, a comprehensive low-level speech quality assessment dataset encompassing 11 key aspects and detailed natural language comments that include reasoning and contextual insights. Additionally, we propose the QualiSpeech Benchmark to evaluate the low-level speech understanding capabilities of auditory large language models (LLMs). Experimental results demonstrate that finetuned auditory LLMs can reliably generate detailed descriptions of noise and distortion, effectively identifying their types and temporal characteristics. The results further highlight the potential for incorporating reasoning to enhance the accuracy and reliability of quality assessments. The dataset can be found at https://huggingface.co/datasets/tsinghua-ee/QualiSpeech. Siyin Wang, Wenyi Yu, Xianzhao Chen, Xiaohai Tian, Jun Zhang 0066, Lu Lu 0015, Yu Tsao 0001, Junichi Yamagishi, Yuxuan Wang 0002, Chao Zhang 0031 |
ACL (1) | 10 |
| 2025 | DNCASR: End-to-End Training for Speaker-Attributed ASRabstractThis paper introduces DNCASR, a novel endto-end trainable system designed for joint neural speaker clustering and automatic speech recognition (ASR), enabling speaker-attributed transcription of long multi-party meetings.DNCASR uses two separate encoders to independently encode global speaker characteristics and local waveform information, along with two linked decoders to generate speaker-attributed transcriptions.The use of linked decoders allows the entire system to be jointly trained under a unified loss function.By employing a serialised training approach, DNCASR effectively addresses overlapping speech in real-world meetings, where the link improves the prediction of speaker indices in overlapping segments.Experiments on the AMI-MDM meeting corpus demonstrate that the jointly trained DNCASR outperforms a parallel system that does not have links between the speaker and ASR decoders.Using cpWER to measure the speaker-attributed word error rate, DNCASR achieves a 9.0% relative reduction on the AMI-MDM Eval set. Xianrui Zheng, Chao Zhang 0031, Philip C. Woodland |
ACL (1) | 2 |
| 2025 | Audio-centric Video Understanding Benchmark without Text ShortcutabstractYudong Yang, Jimin Zhuang, Guangzhi Sun, Changli Tang, Yixuan Li, Peihan Li, Yifan Jiang, Wei Li, Zejun Ma, Chao Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yudong Yang, Jimin Zhuang, Guangzhi Sun, Changli Tang, Peihan Li, Wei Li 0119, Zejun Ma 0001, Chao Zhang 0031 |
EMNLP | 10 |
| 2025 | Editing Music with Melody and Text: Using ControlNet for Diffusion TransformerabstractDespite the significant progress in controllable music generation and editing, challenges remain in the quality and length of generated music due to the use of Mel-spectrogram representations and UNet-based model structures. To address these limitations, we propose a novel approach using a Diffusion Transformer (DiT) augmented with an additional control branch using ControlNet. This allows for long-form and variable-length music generation and editing controlled by text and melody prompts. For more precise and fine-grained melody control, we introduce a novel top-k constant-Q Transform representation as the melody prompt, reducing ambiguity compared to previous representations (e.g., chroma), particularly for music with multiple tracks or a wide range of pitch values. To effectively balance the control signals from text and melody prompts, we adopt a curriculum learning strategy that progressively masks the melody prompt, resulting in a more stable training process. Experiments have been performed on text-to-music generation and music-style transfer tasks using open-source instrumental recording data. The results demonstrate that by extending StableAudio, a pre-trained text-controlled DiT model, our approach enables superior melody-controlled editing while retaining good text-to-music generation performance. These results outperform a strong MusicGen baseline in terms of both text-based generation and melody preservation for editing. Audio examples can be found at https://stable-audio-control.github.io. Siyuan Hou, Shansong Liu, Ruibin Yuan, Wei Xue 0002, Ying Shan, Mangsuo Zhao, Chao Zhang 0031 |
ICASSP | 7 |
| 2025 | Enabling Auditory Large Language Models for Automatic Speech Quality EvaluationabstractSpeech quality assessment typically requires evaluating audio from multiple aspects, such as mean opinion score (MOS) and speaker similarity (SIM) etc., which can be challenging to cover using one small model designed for a single task. In this paper, we propose leveraging recently introduced auditory large language models (LLMs) for automatic speech quality assessment. By employing task-specific prompts, auditory LLMs are finetuned to predict MOS, SIM and A/B testing results, which are commonly used for evaluating text-to-speech systems. Additionally, the finetuned auditory LLM is able to generate natural language descriptions assessing aspects like noisiness, distortion, discontinuity, and overall quality, providing more interpretable outputs. Extensive experiments have been performed on the NISQA, BVCC, SOMOS and VoxSim speech quality datasets, using open-source auditory LLMs such as SALMONN, Qwen-Audio, and Qwen2-Audio. For the natural language descriptions task, a commercial model Google Gemini 1.5 Pro is also evaluated. The results demonstrate that auditory LLMs achieve competitive performance compared to state-of-the-art task-specific small models in predicting MOS and SIM, while also delivering promising results in A/B testing and natural language descriptions. Our data processing scripts and finetuned model checkpoints can be found at https://github.com/bytedance/SALMONN. Siyin Wang, Wenyi Yu, Yudong Yang, Changli Tang, Jimin Zhuang, Xianzhao Chen, Xiaohai Tian, Guangzhi Sun, Lu Lu 0015, Chao Zhang 0031 |
ICASSP | 12 |
| 2025 | Audio Large Language Models Can Be Descriptive Speech Quality EvaluatorsabstractAn ideal multimodal agent should be aware of the quality of its input modalities. Recent advances have enabled large language models (LLMs) to incorporate auditory systems for handling various speech-related tasks. However, most audio LLMs remain unaware of the quality of the speech they process. This limitation arises because speech quality evaluation is typically excluded from multi-task training due to the lack of suitable datasets. To address this, we introduce the first natural language-based speech evaluation corpus, generated from authentic human ratings. In addition to the overall Mean Opinion Score (MOS), this corpus offers detailed analysis across multiple dimensions and identifies causes of quality degradation. It also enables descriptive comparisons between two speech samples (A/B tests) with human-like judgment. Leveraging this corpus, we propose an alignment approach with LLM distillation (ALLD) to guide the audio LLM in extracting relevant information from raw speech and generating meaningful responses. Experimental results demonstrate that ALLD outperforms the previous state-of-the-art regression model in MOS prediction, with a mean square error of 0.17 and an A/B test accuracy of 98.6%. Additionally, the generated responses achieve BLEU scores of 25.8 and 30.2 on two tasks, surpassing the capabilities of task-specific models. This work advances the comprehensive perception of speech signals by audio LLMs, contributing to the development of real-world auditory and sensory intelligent agents. Chen Chen 0075, Siyin Wang, Helin Wang, Zhehuai Chen, Chao Zhang 0031, Chao-Han Huck Yang, Chng Eng Siong |
ICLR | 6 |
| 2025 | Bayesian WeakS-to-Strong from Text Classification to GenerationabstractAdvances in large language models raise the question of how alignment techniques will adapt as models become increasingly complex and humans will only be able to supervise them weakly. Weak-to-Strong mimics such a scenario where weak model supervision attempts to harness the full capabilities of a much stronger model. This work extends Weak-to-Strong to WeakS-to-Strong by exploring an ensemble of weak models which simulate the variability in human opinions. Confidence scores are estimated using a Bayesian approach to guide the WeakS-to-Strong generalization. Furthermore, we extend the application of WeakS-to-Strong from text classification tasks to text generation tasks where more advanced strategies are investigated for supervision. Moreover, direct preference optimization is applied to advance the student model's preference learning, beyond the basic learning framework of teacher forcing. Results demonstrate the effectiveness of the proposed approach for the reliability of a strong student model, showing potential for superalignment. Ziyun Cui, Guangzhi Sun, Wen Wu 0007, Chao Zhang 0031 |
ICLR | 5 |
| 2025 | Improving LLM Video Understanding with 16 Frames Per SecondabstractHuman vision is dynamic and continuous. However, in video understanding with multimodal large language models (LLMs), existing methods primarily rely on static features extracted from images sampled at a fixed low frame rate of frame-per-second (FPS) $\leqslant$2, leading to critical visual information loss. In this paper, we introduce F-16, the first multimodal LLM designed for high-frame-rate video understanding. By increasing the frame rate to 16 FPS and compressing visual tokens within each 1-second clip, F-16 efficiently captures dynamic visual features while preserving key semantic information.
Experimental results demonstrate that higher frame rates considerably enhance video understanding across multiple benchmarks, providing a new approach to improving video LLMs beyond scaling model size or training data. F-16 achieves state-of-the-art performance among 7-billion-parameter video LLMs on both general and fine-grained video understanding benchmarks, such as Video-MME and TemporalBench. Furthermore, F-16 excels in complex spatiotemporal tasks, including high-speed sports analysis (*e.g.*, basketball, football, gymnastics, and diving), outperforming SOTA proprietary visual models like GPT-4o and Gemini-1.5-pro.
Additionally, we introduce a novel decoding method for F-16 that enables highly efficient low-frame-rate inference without requiring model retraining. We will release the source code, model checkpoints, and data at [https://github.com/bytedance/F-16](https://github.com/bytedance/F-16). Changli Tang, Jimin Zhuang, Yudong Yang, Guangzhi Sun, Wei Li 0119, Zejun Ma 0001, Chao Zhang 0031 |
ICML | 8 |
| 2025 | video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language ModelabstractWhile recent advancements in reasoning optimization have significantly enhanced the capabilities of large language models (LLMs), existing efforts to improve reasoning have been limited to solving mathematical problems and focusing on visual graphical inputs, neglecting broader applications in general video understanding. This paper proposes video-SALMONN-o1, the first open-source reasoning-enhanced audio-visual LLM designed for general video understanding tasks. To enhance its reasoning abilities, we develop a reasoning-intensive dataset featuring challenging audio-visual questions with step-by-step solutions. We also propose process direct preference optimization (pDPO), which leverages contrastive step selection to achieve efficient step-level reward modelling tailored for multimodal inputs. Additionally, we introduce RivaBench, the first reasoning-intensive video understanding benchmark, featuring over 4,000 high-quality, expert-curated question-answer pairs across scenarios such as standup comedy, academic presentations, and synthetic video detection. video-SALMONN-o1 achieves 3-8% accuracy improvements over the LLaVA-OneVision baseline across different video reasoning benchmarks. Besides, pDPO achieves 6-8% improvements compared to the supervised fine-tuning model on RivaBench. Enhanced reasoning enables video-SALMONN-o1 zero-shot synthetic video detection capabilities. Guangzhi Sun, Yudong Yang, Jimin Zhuang, Changli Tang, Wei Li 0119, Zejun Ma 0001, Chao Zhang 0031 |
ICML | 8 |
| 2025 | The 1st SpeechWellness Challenge: Detecting Suicide Risk Among AdolescentsabstractThe 1st SpeechWellness Challenge (SW1) aims to advance methods for detecting current suicide risk in adolescents using speech analysis techniques. Suicide among adolescents is a critical public health issue globally. Early detection of suicidal tendencies can lead to timely intervention and potentially save lives. Traditional methods of assessment often rely on self-reporting or clinical interviews, which may not always be accessible. The SW1 challenge addresses this gap by exploring speech as a non-invasive and readily available indicator of mental health. We release the SW1 dataset which contains speech recordings from 600 adolescents aged 10-18 years. By focusing on speech generated from natural tasks, the challenge seeks to uncover patterns and markers that correlate with current suicide risk. Wen Wu 0007, Ziyun Cui, Chang Lei, Yinan Duan, Diyang Qu, Ji Wu 0002, Bowen Zhou 0001, Runsen Chen, Chao Zhang 0031 |
INTERSPEECH | 9 |
| 2025 | BrainOmni: A Brain Foundation Model for Unified EEG and MEG SignalsabstractElectroencephalography (EEG) and magnetoencephalography (MEG) measure neural activity non-invasively by capturing electromagnetic fields generated by dendritic currents.
Although rooted in the same biophysics, EEG and MEG exhibit distinct signal patterns, further complicated by variations in sensor configurations across modalities and recording devices.
Existing approaches typically rely on separate, modality- and dataset-specific models, which limits the performance and cross-domain scalability.
This paper proposes BrainOmni, the first brain foundation model that generalises across heterogeneous EEG and MEG recordings.
To unify diverse data sources, we introduce BrainTokenizer, the first tokeniser that quantises spatiotemporal brain activity into discrete representations.
Central to BrainTokenizer is a novel Sensor Encoder that encodes sensor properties such as spatial layout, orientation, and type, enabling compatibility across devices and modalities.
Building upon the discrete representations, BrainOmni learns unified semantic embeddings of brain signals by self-supervised pretraining. To the best of our knowledge, it is the first foundation model to support both EEG and MEG signals, as well as the first to incorporate large-scale MEG pretraining.
A total of 1,997 hours of EEG and 656 hours of MEG data are curated and standardised from publicly available sources for pretraining.
Experiments show that BrainOmni outperforms both existing foundation models and state-of-the-art task-specific models on a range of downstream tasks. It also demonstrates strong generalisation to unseen EEG and MEG devices. Further analysis reveals that joint EEG-MEG (EMEG) training yields consistent improvements across both modalities. Code and checkpoints are publicly available at https://github.com/OpenTSLab/BrainOmni Qinfan Xiao, Ziyun Cui, Wen Wu 0007, Andrew Thwaites, Alexandra Woolgar, Bowen Zhou 0001, Chao Zhang 0031 |
NeurIPS | 9 |
| 2025 | SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex ConversationabstractIn order to enable fluid and natural human-machine speech interaction, existing full-duplex conversational systems often adopt modular architectures with auxiliary components such as voice activity detectors, interrupters, conversation state predictors, or multiple LLMs. These systems, however, suffer from error accumulation across modules and struggle with key challenges such as context-dependent barge-in and echo cancellation. Recent approaches, most notably Moshi, simplify the pipeline by injecting audio codecs into the token space of a single LLM. However, such methods still incur significant performance degradation when operating on the speech rather than text modality. In this paper, we introduce SALMONN-omni, the first single, standalone full-duplex speech LLM that operates without audio codecs in its token space. It features a novel dynamic thinking mechanism within the LLM backbone, enabling the model to learn when to transition between speaking and listening states. Experiments on widely used benchmarks for spoken question answering and open-domain dialogue show that SALMONN-omni achieves at least 30\% relative performance improvement over existing open-source full-duplex models and performs highly competitively to half-duplex and turn-based systems, despite using substantially less training data. Moreover, SALMONN-omni demonstrates strong performance in complex conversational scenarios, including turn-taking, backchanneling, echo cancellation and context-dependent barge-in, with further improvements achieved through reinforcement learning. Some demo conversations between user and SALMONN-omni are provided in the following repository https://github.com/bytedance/SALMONN. Wenyi Yu, Siyin Wang, Xianzhao Chen, Xiaohai Tian, Jun Zhang 0003, Guangzhi Sun, Lu Lu 0015, Yuxuan Wang 0002, Chao Zhang 0031 |
NeurIPS | 10 |
| 2025 | Knowledge-aware audio-grounded generative slot filling for limited annotated dataabstractManually annotating fine-grained slot-value labels for task-oriented dialogue (ToD) systems is an expensive and time-consuming endeavour. This motivates research into slot-filling methods that operate with limited amounts of labelled data. Moreover, the majority of current work on ToD is based solely on text as the input modality, neglecting the additional challenges of imperfect automatic speech recognition (ASR) when working with spoken language. In this work, we propose a Knowledge-Aware Audio-Grounded generative slot filling framework, termed KA2G, that focuses on few-shot and zero-shot slot filling for ToD with speech input. KA2G achieves robust and data-efficient slot filling for speech-based ToD by (1) framing it as a text generation task, (2) grounding text generation additionally in the audio modality, and (3) conditioning on available external knowledge ( e.g. a predefined list of possible slot values). We show that combining both modalities within the KA2G framework improves the robustness against ASR errors. Further, the knowledge-aware slot-value generator in KA2G, implemented via a pointer generator mechanism, particularly benefits few-shot and zero-shot learning. Experiments, conducted on the standard speech-based single-turn SLURP dataset and a multi-turn dataset extracted from a commercial ToD system , display strong and consistent gains over prior work, especially in few-shot and zero-shot setups. Guangzhi Sun, Chao Zhang 0031, Ivan Vulic, Pawel Budzianowski, Philip C. Woodland |
Comput. Speech Lang. | 2 |
| 2024 | M³AV: A Multimodal, Multigenre, and Multipurpose Audio-Visual Academic Lecture DatasetabstractZhe Chen, Heyang Liu, Wenyi Yu, Guangzhi Sun, Hongcheng Liu, Ji Wu, Chao Zhang, Yu Wang, Yanfeng Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhe Chen 0024, Heyang Liu, Wenyi Yu, Guangzhi Sun, Ji Wu 0002, Chao Zhang 0031, Yu Wang 0027, Yanfeng Wang 0001 |
ACL (1) | 7 |
| 2024 | Handling Ambiguity in Emotion: From Out-of-Domain Detection to Distribution EstimationabstractWen Wu, Bo Li, Chao Zhang, Chung-Cheng Chiu, Qiujia Li, Junwen Bai, Tara Sainath, Phil Woodland. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Wen Wu 0007, Bo Li 0028, Chao Zhang 0031, Chung-Cheng Chiu, Qiujia Li, Junwen Bai, Tara N. Sainath, Philip C. Woodland |
ACL (1) | 3 |
| 2024 | Bayesian Example Selection Improves In-Context Learning for Speech, Text and Visual ModalitiesabstractLarge language models (LLMs) can adapt to new tasks through in-context learning (ICL) based on a few examples presented in dialogue history without any model parameter update.Despite such convenience, the performance of ICL heavily depends on the quality of the incontext examples presented, which makes the in-context example selection approach a critical choice.This paper proposes a novel Bayesian in-Context example Selection method (ByCS) for ICL.Extending the inference probability conditioned on in-context examples based on Bayes' theorem, ByCS focuses on the inverse inference conditioned on test input.Following the assumption that accurate inverse inference probability (likelihood) will result in accurate inference probability (posterior), incontext examples are selected based on their inverse inference results.Diverse and extensive cross-tasking and cross-modality experiments are performed with speech, text, and image examples.Experimental results show the efficacy and robustness of our ByCS method on various models, tasks and modalities. Siyin Wang, Chao-Han Huck Yang, Ji Wu 0002, Chao Zhang 0031 |
EMNLP | 4 |
| 2024 | Extending Large Language Models for Speech and Audio CaptioningabstractMultimodal large language models (LLMs) have shown promising visual perception abilities by connecting with image encoders, but their performance on auditory tasks has not yet been widely investigated. Meanwhile, automatic speech recognition (ASR) and automatic audio captioning (AAC) are often achieved with separate systems, resulting in incomplete auditory perception abilities. To fill in these gaps, in this paper, we present the first study that achieves both ASR and AAC by connecting an LLM with auditory encoders. A dual auditory encoder structure is proposed, integrating the Whisper encoder for speech and the BEATs encoder for audio events with a high temporal resolution by using a Q-Former at the window level. Experiments for ASR and AAC are performed correspondingly on the widely used LibriSpeech, GigaSpeech, WavCaps, AudioCaps, and Clotho datasets and yield promising results. In particular, state-of-the-art results are achieved on GigaSpeech, AudioCaps and Clotho. Our model is also able to caption speech and audio events simultaneously from clips with mixed speech and background audio events, which is a step towards more complete machine auditory perception. Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan 0019, Wei Li 0119, Lu Lu 0015, Zejun Ma 0001, Chao Zhang 0031 |
ICASSP | 9 |
| 2024 | Can Whisper Perform Speech-Based In-Context Learning?abstractThis paper investigates the in-context learning abilities of the Whisper automatic speech recognition (ASR) models released by OpenAI. A novel speech-based in-context learning (SICL) approach is proposed for test-time adaptation, which can reduce the word error rates (WERs) with only a small number of labelled speech samples without gradient descent. Language-level adaptation experiments using Chinese dialects showed that when applying SICL to isolated word ASR, consistent and considerable relative WER reductions can be achieved using Whisper models of any size on two dialects, which is on average 32.3%. A k-nearest-neighbours-based in-context example selection technique can be applied to further improve the efficiency of SICL, which can increase the average relative WER reduction to 36.4%. The findings are verified using speaker adaptation or continuous speech recognition tasks, and both achieved considerable relative WER reductions. Detailed quantitative analyses are also provided to shed light on SICL’s adaptability to phonological variances and dialect-specific lexical nuances. Siyin Wang, Chao-Han Huck Yang, Ji Wu 0002, Chao Zhang 0031 |
ICASSP | 4 |
| 2024 | Connecting Speech Encoder and Large Language Model for ASRabstractThe impressive capability and versatility of large language models (LLMs) have aroused increasing attention in automatic speech recognition (ASR), with several pioneering studies attempting to build integrated ASR models by connecting a speech encoder with an LLM. This paper presents a comparative study of three commonly used structures as connectors, including fully connected layers, multi-head cross-attention, and Q-Former. Speech encoders from the Whisper model series as well as LLMs from the Vicuna model series with different model sizes were studied. Experiments were performed on the commonly used LibriSpeech, Common Voice, and GigaSpeech datasets, where the LLMs with Q-Formers demonstrated consistent and considerable word error rate (WER) reductions over LLMs with other connector structures. Q-Former-based LLMs can generalise well to out-of-domain datasets, where 12% relative WER reductions over the Whisper baseline ASR model were achieved on the Eval2000 test set without using any in-domain training data from Switchboard. Moreover, a novel segment-level Q-Former is proposed to enable LLMs to recognise speech segments with a duration exceeding the limitation of the encoders, which results in 17% relative WER reductions over other connector structures on 90-second-long speech data. Wenyi Yu, Changli Tang, Guangzhi Sun, Xianzhao Chen, Tian Tan 0019, Wei Li 0119, Lu Lu 0015, Zejun Ma 0001, Chao Zhang 0031 |
ICASSP | 9 |
| 2024 | Enhancing Quantised End-to-End ASR Models Via PersonalisationabstractRecent end-to-end automatic speech recognition (ASR) models have become increasingly larger, making them particularly challenging to be deployed on resource-constrained devices. Model quantisation is an effective solution that sometimes causes the word error rate (WER) to increase. In this paper, a novel strategy of personalisation for a quantised model (PQM) is proposed, which combines speaker adaptive training (SAT) with model quantisation to improve the performance of heavily compressed models. Specifically, PQM uses a 4-bit NormalFloat Quantisation (NF4) approach for model quantisation and low-rank adaptation (LoRA) for SAT. Experiments have been performed on the LibriSpeech and the TED-LIUM 3 corpora. Remarkably, with a 7x reduction in model size and 1% additional speaker-specific parameters, 15.1% and 23.3% relative WER reductions were achieved on quantised Whisper and Conformer-based attention-based encoder-decoder ASR models respectively, comparing to the original full precision models. Qiuming Zhao, Guangzhi Sun, Chao Zhang 0031, Mingxing Xu, Thomas Fang Zheng |
ICASSP | 3 |
| 2024 | Large Language Models are Efficient Learners of Noise-Robust Speech RecognitionabstractRecent advances in large language models (LLMs) have promoted generative error correction (GER) for automatic speech recognition (ASR), which leverages the rich linguistic knowledge and powerful reasoning ability of LLMs to improve recognition results. The latest work proposes a GER benchmark with "HyPoradise" dataset to learn the mapping from ASR N-best hypotheses to ground-truth transcription by efficient LLM finetuning, which shows great effectiveness but lacks specificity on noise-robust ASR. In this work, we extend the benchmark to noisy conditions and investigate if we can teach LLMs to perform denoising for GER just like what robust ASR do, where one solution is introducing noise information as a conditioner into LLM. However, directly incorporating noise embeddings from audio encoder could harm the LLM tuning due to cross-modality gap. To this end, we propose to extract a language-space noise embedding from the N-best list to represent the noise conditions of source speech, which can promote the denoising process in GER. Furthermore, in order to enhance its representation ability of audio noise, we design a knowledge distillation (KD) approach via mutual information estimation to distill the real noise information in audio embeddings to our language embedding. Experiments on various latest LLMs demonstrate our approach achieves a new breakthrough with up to 53.9% correction improvement in terms of word error rate while with limited training data. Analysis shows that our language-space noise embedding can well represent the noise conditions of source speech, under which off-the-shelf LLMs show strong ability of language-space denoising. Chen Chen 0075, Chao-Han Huck Yang, Ruizhe Li 0001, Chao Zhang 0031, Chng Eng Siong |
ICLR | 5 |
| 2024 | SALMONN: Towards Generic Hearing Abilities for Large Language ModelsabstractHearing is arguably an essential ability of artificial intelligence (AI) agents in the physical world, which refers to the perception and understanding of general auditory information consisting of at least three types of sounds: speech, audio events, and music. In this paper, we propose SALMONN, a speech audio language music open neural network, built by integrating a pre-trained text-based large language model (LLM) with speech and audio encoders into a single multimodal model. SALMONN enables the LLM to directly process and understand general audio inputs and achieve competitive performances on a number of speech and audio tasks used in training, such as
automatic speech recognition and translation, auditory-information-based question answering, emotion recognition, speaker verification, and music and audio captioning etc. SALMONN also has a diverse set of emergent abilities unseen in the training, which includes but is not limited to speech translation to untrained languages, speech-based slot filling, spoken-query-based question answering, audio-based storytelling, and speech audio co-reasoning etc. The presence of cross-modal emergent abilities is studied, and a novel few-shot activation tuning approach is proposed to activate such abilities. To our knowledge, SALMONN is the first model of its type and can be regarded as a step towards AI with generic hearing abilities. The source code, model checkpoints and data are available at https://github.com/bytedance/SALMONN. Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan 0019, Wei Li 0119, Lu Lu 0015, Zejun Ma 0001, Chao Zhang 0031 |
ICLR | 9 |
| 2024 | video-SALMONN: Speech-Enhanced Audio-Visual Large Language ModelsabstractSpeech understanding as an element of the more generic video understanding using audio-visual large language models (av-LLMs) is a crucial yet understudied aspect. This paper proposes video-SALMONN, a single end-to-end av-LLM for video processing, which can understand not only visual frame sequences, audio events and music, but speech as well. To obtain fine-grained temporal information required by speech understanding, while keeping efficient for other video elements, this paper proposes a novel multi-resolution causal Q-Former (MRC Q-Former) structure to connect pre-trained audio-visual encoders and the backbone large language model. Moreover, dedicated training approaches including the diversity loss and the unpaired audio-visual mixed training scheme are proposed to avoid frames or modality dominance. On the introduced audio-visual evaluation benchmark, video-SALMONN achieves more than 25% absolute accuracy improvements on the video-QA task and over 30% absolute accuracy improvements on audio-visual QA tasks with human speech. In addition, video-SALMONN demonstrates remarkable video comprehension and reasoning abilities on tasks that are unprecedented by other av-LLMs. Our training code and model checkpoints are available at https://github.com/bytedance/SALMONN/ Guangzhi Sun, Wenyi Yu, Changli Tang, Xianzhao Chen, Tian Tan 0019, Wei Li 0119, Lu Lu 0015, Zejun Ma 0001, Yuxuan Wang 0002, Chao Zhang 0031 |
ICML | 10 |
| 2024 | Confidence Estimation for Automatic Detection of Depression and Alzheimer's Disease Based on Clinical Interviews
Wen Wu 0007, Chao Zhang 0031, Philip C. Woodland |
INTERSPEECH | 2 |
| 2024 | Spontaneous Speech-Based Suicide Risk Detection Using Whisper and Large Language Models
Ziyun Cui, Chang Lei, Wen Wu 0007, Yinan Duan, Diyang Qu, Ji Wu 0002, Runsen Chen, Chao Zhang 0031 |
INTERSPEECH | 8 |
| 2024 | Can Large Language Models Understand Spatial Audio?
Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan 0019, Wei Li 0119, Jun Zhang 0066, Lu Lu 0015, Zejun Ma 0001, Yuxuan Wang 0002, Chao Zhang 0031 |
INTERSPEECH | 11 |
| 2024 | SAML: Speaker Adaptive Mixture of LoRA Experts for End-to-End ASR
Qiuming Zhao, Guangzhi Sun, Chao Zhang 0031, Mingxing Xu, Thomas Fang Zheng |
INTERSPEECH | 3 |
| 2024 | Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models
Shuai Wang 0016, Guangzhi Sun, Chao Zhang 0031, Mingxing Xu, Thomas Fang Zheng |
INTERSPEECH | 5 |
| 2024 | SOT Triggered Neural Clustering for Speaker Attributed ASR
Xianrui Zheng, Guangzhi Sun, Chao Zhang 0031, Philip C. Woodland |
INTERSPEECH | 3 |
| 2024 | Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation ModelsabstractWe propose an unsupervised adaptation framework, Self-TAught Recognizer (STAR), which leverages unlabeled data to enhance the robustness of automatic speech recognition (ASR) systems in diverse target domains, such as noise and accents. STAR is developed for prevalent speech foundation models based on Transformer-related architecture with auto-regressive decoding (e.g., Whisper, Canary). Specifically, we propose a novel indicator that empirically integrates step-wise information during decoding to assess the token-level quality of pseudo labels without ground truth, thereby guiding model updates for effective unsupervised adaptation. Experimental results show that STAR achieves an average of 13.5% relative reduction in word error rate across 14 target domains, and it sometimes even approaches the upper-bound performance of supervised adaptation. Surprisingly, we also observe that STAR prevents the adapted model from the common catastrophic forgetting problem without recalling source-domain data. Furthermore, STAR exhibits high data efficiency that only requires less than one-hour unlabeled data, and seamless generality to alternative large speech models and speech translation tasks. Our code aims to open source to the research communities. Chen Chen 0075, Chao-Han Yang, Chengwei Qin, Chng Eng Siong, Chao Zhang 0031 |
NeurIPS | 7 |
| 2024 | An Improved Empirical Fisher Approximation for Natural Gradient DescentabstractApproximate Natural Gradient Descent (NGD) methods are an important family of optimisers for deep learning models, which use approximate Fisher information matrices to pre-condition gradients during training. The empirical Fisher (EF) method approximates the Fisher information matrix empirically by reusing the per-sample gradients collected during back-propagation. Despite its ease of implementation, the EF approximation has its theoretical and practical limitations. This paper investigates the *inversely-scaled projection* issue of EF, which is shown to be a major cause of its poor empirical approximation quality. An improved empirical Fisher (iEF) method is proposed to address this issue, which is motivated as a generalised NGD method from a loss reduction perspective, meanwhile retaining the practical convenience of EF. The exact iEF and EF methods are experimentally evaluated using practical deep learning setups, including widely-used setups for parameter-efficient fine-tuning of pre-trained models (T5-base with LoRA and Prompt-Tuning on GLUE tasks, and ViT with LoRA for CIFAR100). Optimisation experiments show that applying exact iEF directly as an optimiser provides strong convergence and generalisation. It achieves the best test performance and the lowest training loss for the majority of the tasks, even when compared to well-tuned AdamW/Adafactor baselines. Additionally, under a novel empirical evaluation framework, the proposed iEF method shows consistently better approximation quality to exact Natural Gradient updates than both the EF and the more expensive sampled Fisher methods, meanwhile demonstrating the superior property of being robust to the choice of damping across tasks and training stages. Improving existing approximate NGD optimisers with iEF is expected to lead to better convergence and robustness. Furthermore, the iEF method also serves as a better approximation method to the Fisher information matrix itself, which enables the improvement of a variety of Fisher-based methods, not limited to the scope of optimisation. Wenyi Yu, Chao Zhang 0031, Philip C. Woodland |
NeurIPS | 3 |
| 2024 | Affect Recognition in Conversations Using Large Language ModelsabstractAffect recognition, encompassing emotions, moods, and feelings, plays a pivotal role in human communication.In the realm of conversational artificial intelligence, the ability to discern and respond to human affective cues is a critical factor for creating engaging and empathetic interactions.This study investigates the capacity of large language models (LLMs) to recognise human affect in conversations, with a focus on both open-domain chit-chat dialogues and task-oriented dialogues.Leveraging three diverse datasets, namely IEMOCAP (Busso et al., 2008), EmoWOZ (Feng et al., 2022), and DAIC-WOZ (Gratch et al., 2014), covering a spectrum of dialogues from casual conversations to clinical interviews, we evaluate and compare LLMs' performance in affect recognition.Our investigation explores the zero-shot and few-shot capabilities of LLMs through incontext learning as well as their model capacities through task-specific fine-tuning.Additionally, this study takes into account the potential impact of automatic speech recognition errors on LLM predictions.With this work, we aim to shed light on the extent to which LLMs can replicate human-like affect recognition capabilities in conversations. Shutong Feng, Guangzhi Sun, Nurul Lubis, Wen Wu 0007, Chao Zhang 0031, Milica Gasic |
SIGDIAL | 5 |
| 2024 | Hierarchical Multi-Path and Multi-Model Selection For Fake Speech DetectionabstractThe variety of spoofing algorithms used in generating speech poses obstacles to fake speech detection. Earlier methods have demonstrated complementary effects for detection. This paper proposes a novel hierarchical multi-path multi-model selection method for fake speech detection. It is designed to dynamically select and utilise the most suitable model from a set of complementary models. In our method, four basic detection models are incorporated, each offering partial but complementary detection abilities, to enhance balanced performance on diverse fake speech. The models are trained through a multi-path schema and the selection mechanism is structured hierarchically to improve the generalisation ability. Our method achieves an Equal Error Rate (EER) of 0.37% on the ASVspoof 2019 LA dataset, and outperforms other state-of-the-art method on the cross-domain and cross-dataset scenarios. A statistical analysis of EERs against thirteen unknown attacks reveals our method’s superiority, evidenced by the lowest standard deviation of 0.24, further underscoring our method’s robustness against a range of attacks. Chang Feng, Guangzhi Sun, Shuai Wang 0016, Chao Zhang 0031, Mingxing Xu, Thomas Fang Zheng |
SLT | 6 |
| 2024 | Automatic Time Alignment Generation For End-to-End ASR Using Acoustic Probability ModellingabstractEnd-to-end trainable (E2E) automatic speech recognition (ASR) models can achieve low error rates, but unlike hidden Markov model (HMM)-based systems they cannot naturally provide reliable timestamps. In this paper, an extra alignment module is proposed to enable pre-trained E2E ASR models to find accurate time-stamps by exploiting the connections between E2E and conventional HMMbased systems, which consists of a voice activity detection (VAD) and an acoustic scorer (AS). Full-sum training was adopted and the Viterbi algorithm used for alignment generation with HMM-like acoustic probabilities. Experiments using different E2E backbones and acoustic encoders with different output frame rates showed that for a jointly trained VAD and AS, the proposed methods can achieve high in-domain VAD performance and 99% alignment accuracy with a 200 ms collar for recognised words, with an average mismatch between 20 ms and 50 ms. Dongcheng Jiang, Chao Zhang 0031, Philip C. Woodland |
SLT | 2 |
| 2024 | Large Language Model Based Generative Error Correction: A Challenge and Baselines For Speech Recognition, Speaker Tagging, and Emotion RecognitionabstractGiven recent advances in generative AI technology, a key question is how large language models (LLMs) can enhance acoustic modeling tasks using text decoding results from a frozen, pretrained automatic speech recognition (ASR) model. To explore new capabilities in language modeling for speech processing, we introduce the generative speech transcription error correction (GenSEC) challenge. This challenge comprises three post-ASR language modeling tasks: (i) post-ASR transcription correction, (ii) speaker tagging, and (iii) emotion recognition. These tasks aim to emulate future LLM-based agents handling voice-based interfaces while remaining accessible to a broad audience by utilizing open pretrained language models or agent-based APIs. We also discuss insights from baseline evaluations, as well as lessons learned for designing future evaluations. Chao-Han Huck Yang, Taejin Park, Yuan Gong 0001, Yuanchao Li, Zhehuai Chen, Chen Chen 0075, Kunal Dhawan, Piotr Zelasko, Chao Zhang 0031, Yun-Nung Chen, Yu Tsao 0001, Jagadeesh Balam, Boris Ginsburg, Sabato Marco Siniscalchi, Chng Eng Siong, Peter Bell 0001, Catherine Lai, Shinji Watanabe 0001, Andreas Stolcke |
SLT | 11 |
| 2024 | SWIM: Short-Window CNN Integrated With Mamba for EEG-Based Auditory Spatial Attention DecodingabstractIn complex auditory environments, the human auditory system possesses the remarkable ability to focus on a specific speaker while disregarding others. In this study, a new model named SWIM, a short-window convolution neural network (CNN) integrated with Mamba, is proposed for identifying the locus of auditory attention (left or right) from electroencephalography (EEG) signals without relying on speech envelopes. SWIM consists of two parts. The first is a short-window CNN ($\mathrm{SW}_{\mathrm{CNN}}$), which acts as a short-term EEG feature extractor and achieves a final accuracy of 84.9% in the leave-one-speaker-out setup on the widely used KUL dataset. This improvement is due to the use of an improved CNN structure, data augmentation, multitask training, and model combination. The second part, Mamba, is a sequence model first applied to auditory spatial attention decoding to leverage the long-term dependency from previous $\mathrm{SW}_{\mathrm{CNN}}$ time steps. By joint training $\mathrm{SW}_{\mathrm{CNN}}$ and Mamba, the proposed SWIM structure uses both short-term and long-term information and achieves an accuracy of 86.2%, which reduces the classification errors by a relative 31.0% compared to the previous state-of-the-art result. The source code is available at https://github.com/windowso/SWIM-ASAD. Andrew Thwaites, Alexandra Woolgar, Chao Zhang 0031 |
SLT | 5 |
| 2024 | Cross-Utterance Conditioned VAE for Speech GenerationabstractSpeech synthesis systems powered by neural networks hold promise for multimedia production, but frequently face issues with producing expressive speech and seamless editing. In response, we present the Cross-Utterance Conditioned Variational Autoencoder speech synthesis (CUC-VAE S2) framework to enhance prosody and ensure natural speech generation. This framework leverages the powerful representational capabilities of pre-trained language models and the re-expression abilities of variational autoencoders (VAEs). The core component of the CUC-VAE S2 framework is the cross-utterance CVAE, which extracts acoustic, speaker, and textual features from surrounding sentences to generate context-sensitive prosodic features, more accurately emulating human prosody generation. We further propose two practical algorithms tailored for distinct speech synthesis applications: CUC-VAE TTS for text-to-speech and CUC-VAE SE for speech editing. The CUC-VAE TTS is a direct application of the framework, designed to generate audio with contextual prosody derived from surrounding texts. On the other hand, the CUC-VAE SE algorithm leverages real mel spectrogram sampling conditioned on contextual information, producing audio that closely mirrors real sound and thereby facilitating flexible speech editing based on text such as deletion, insertion, and replacement. Experimental results on the LibriTTS datasets demonstrate that our proposed models significantly enhance speech synthesis and editing, producing more natural and expressive speech. Yang Li 0116, Guangzhi Sun, Weiqin Zu, Zheng Tian 0002, Ying Wen 0001, Wei Pan 0004, Chao Zhang 0031, Jun Wang 0012, Yang Yang 0001, Fanglei Sun |
IEEE ACM Trans. Audio Speech Lang. Process. | 8 |
| 2024 | Graph Neural Networks for Contextual ASR With the Tree-Constrained Pointer GeneratorabstractIncorporating biasing words obtained through contextual knowledge is paramount in automatic speech recognition (ASR) applications. This paper proposes an innovative method for achieving end-to-end contextual ASR using graph neural network (GNN) encodings based on the tree-constrained pointer generator method. GNN node encodings facilitate lookahead for future word pieces in the process of ASR decoding at each tree node by incorporating information about all word pieces on the tree branches rooted from it. This results in a more precise prediction of the generation probability of the biasing words. The study explores three GNN encoding techniques: namely the tree recursive neural network (Tree-RNN), the graph convolutional network (GCN), and GraphSAGE, along with different combinations of the complementary GCN and GraphSAGE structures. The performance of the systems was evaluated using both Librispeech and the AMI corpus with a visual-grounded contextual ASR pipeline. The findings indicate that using GNN encodings achieved consistent and significant reductions in word error rate (WER), particularly for words that are rare or have not been seen during the training process. Notably, on LibriSpeech test sets, the combined GNN proposed in this paper achieved a 20% relative rare word error rate reduction compared to Tree-RNN, 30%-40% compared to standard TCPGen and 60% compared to standard ASR systems without TCPGen. Guangzhi Sun, Chao Zhang 0031, Philip C. Woodland |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Estimating the Uncertainty in Emotion Attributes using Deep Evidential RegressionabstractIn automatic emotion recognition (AER), labels assigned by different human annotators to the same utterance are often inconsistent due to the inherent complexity of emotion and the subjectivity of perception.Though deterministic labels generated by averaging or voting are often used as the ground truth, it ignores the intrinsic uncertainty revealed by the inconsistent labels.This paper proposes a Bayesian approach, deep evidential emotion regression (DEER), to estimate the uncertainty in emotion attributes.Treating the emotion attribute labels of an utterance as samples drawn from an unknown Gaussian distribution, DEER places an utterance-specific normal-inverse gamma prior over the Gaussian likelihood and predicts its hyper-parameters using a deep neural network model.It enables a joint estimation of emotion attributes along with the aleatoric and epistemic uncertainties.AER experiments on the widely used MSP-Podcast and IEMOCAP datasets showed DEER produced state-of-theart results for both the mean values and the distribution of emotion attributes 1 . Wen Wu 0007, Chao Zhang 0031, Philip C. Woodland |
ACL (1) | 2 |
| 2023 | Transferring Speech-Generic and Depression-Specific Knowledge for Alzheimer's Disease DetectionabstractThe detection of Alzheimer’s disease (AD) from spontaneous speech has attracted increasing attention while the sparsity of training data remains an important issue. This paper handles the issue by knowledge transfer, specifically from both speech-generic and depression-specific knowledge. The paper first studies sequential knowledge transfer from generic foundation models pretrained on large amounts of speech and text data. A block-wise analysis is performed for AD diagnosis based on the representations extracted from different intermediate blocks of different foundation models. Apart from the knowledge from speech-generic representations, this paper also proposes to simultaneously transfer the knowledge from a speech depression detection task based on the high comorbidity rates of depression and AD. A parallel knowledge transfer framework is studied that jointly learns the information shared between these two tasks. Experimental results show that the proposed method improves AD and depression detection, and produces a state-of-the-art F1 score of 0.928 for AD diagnosis on the commonly used ADReSSo dataset. Ziyun Cui, Wen Wu 0007, Weiqiang Zhang 0001, Ji Wu 0002, Chao Zhang 0031 |
ASRU | 5 |
| 2023 | Improving Speech Enhancement Using Audio Tagging Knowledge From Pre-Trained Representations and Multi-Task LearningabstractIn deep-learning-based speech enhancement (SE), an audio-knowledge-ignorant approach is often used, which estimates a denoising model to transform the noisy input speech into clean output speech without understanding the audio events that constitute the background noises. In this paper, an audio-knowledge-aware approach is proposed to improve SE, which explicitly leverages the knowledge from audio taggings to understand the background noises. Based on the recent progress in audio pattern analysis, the audio tagging knowledge is obtained using either additional input representations extracted by pre-trained audio tagging models, or from multi-task learning with extra audio event classification or regression tasks. Experimental results based on the DNS-2020 dataset and the pre-trained Wavegram-Logmel-CNN audio tagging model show that the proposed approach leads to considerable improvements in the STOI, SDR, and SI-SNR metrics. Shaoxiong Lin, Chao Zhang 0031, Yanmin Qian |
ASRU | 2 |
| 2023 | Context-Aware end-to-end ASR Using Self-Attentive Embedding and Tensor FusionabstractTypical automatic speech recognition (ASR) systems are built to recognize independent utterances without using the cross-utterance context. However, the context over multiple utterances often provides useful information for the ASR task. In this work, we propose a context-aware end-to-end ASR model that injects the self-attentive context embedding into the decoder of the recurrent neural network transducer (RNN-T). We also propose a factorised 3-way tensor fusion approach to fuse the context embedding with the acoustic representations extracted from the acoustic encoder and the text representations obtained using the prediction network based on the previous subword units. Experimental results on a long-form Youtube ASR task shows that the proposed approach achieves 10.8% relative word error rate reductions. Shuo-Yiin Chang, Chao Zhang 0031, Tara N. Sainath, Bo Li 0028, Trevor Strohman |
ICASSP | 2 |
| 2023 | Spectral Clustering-Aware Learning of Embeddings for Speaker DiarisationabstractIn speaker diarisation, speaker embedding extraction models often suffer from the mismatch between their training loss functions and the speaker clustering method. In this paper, we propose the method of spectral clustering-aware learning of embeddings (SCALE) to address the mismatch. Specifically, besides an angular prototypical (AP) loss, SCALE uses a novel affinity matrix loss which directly minimises the error between the affinity matrix estimated from speaker embeddings and the reference. SCALE also includes p-percentile thresholding and Gaussian blur as two important hyper-parameters for spectral clustering in training. Experiments on the AMI dataset showed that speaker embeddings obtained with SCALE achieved over 50% relative speaker error rate reductions using oracle segmentation, and over 30% relative diarisation error rate reductions using automatic segmentation when compared to a strong baseline with the AP-loss-based speaker embeddings. Evonne P. C. Lee, Guangzhi Sun, Chao Zhang 0031, Philip C. Woodland |
ICASSP | 3 |
| 2023 | End-to-End Spoken Language Understanding with Tree-Constrained Pointer GeneratorabstractEnd-to-end spoken language understanding (SLU) suffers from the long-tail word problem. This paper exploits contextual biasing, a technique to improve the speech recognition of rare words, in end-to-end SLU systems. Specifically, a tree-constrained pointer generator (TCPGen), a powerful and efficient biasing model component, is studied, which leverages a slot shortlist with corresponding entities to extract biasing lists. Meanwhile, to bias the SLU model output slot distribution, a slot probability biasing (SPB) mechanism is proposed to calculate a slot distribution from TCPGen. Experiments on the SLURP dataset showed consistent SLU-F1 improvements using TCPGen and SPB, especially on unseen entities. On a new split by holding out 5 slot types for the test, TCPGen with SPB achieved zero-shot learning with an SLU-F1 score over 50% compared to baselines which can not deal with it. In addition to slot filling, the intent classification accuracy was also improved. Guangzhi Sun, Chao Zhang 0031, Philip C. Woodland |
ICASSP | 2 |
| 2023 | Self-Supervised Representations in Speech-Based Depression DetectionabstractThis paper proposes handling training data sparsity in speech-based automatic depression detection (SDD) using foundation models pre-trained with self-supervised learning (SSL). An analysis of SSL representations derived from different layers of pre-trained foundation models is first presented for SDD, which provides insight to suitable indicator for depression detection. Knowledge transfer is then performed from automatic speech recognition (ASR) and emotion recognition to SDD by fine-tuning the foundation models. Results show that the uses of oracle and ASR transcriptions yield similar SDD performance when the hidden representations of the ASR model is incorporated along with the ASR textual information. By integrating representations from multiple foundation models, state-of-the-art SDD results based on real ASR were achieved on the DAIC-WOZ dataset. Wen Wu 0007, Chao Zhang 0031, Philip C. Woodland |
ICASSP | 2 |
| 2023 | UML: A Universal Monolingual Output Layer For Multilingual AsrabstractWord-piece models (WPMs) are commonly used subword units in state-of-the-art end-to-end automatic speech recognition (ASR) systems. For multilingual ASR, due to the differences in written scripts across languages, multilingual WPMs bring the challenges of having overly large output layers and scaling to more languages. In this work, we propose a universal monolingual output layer (UML) to address such problems. Instead of one output node for only one WPM, UML re-associates each output node with multiple WPMs, one for each language, and results in a smaller monolingual output layer shared across languages. Consequently, the UML enables to switch in the interpretation of each output node depending on the language of the input speech. Experimental results on an 11-language voice search task demonstrated the feasibility of using UML for high-quality and high-efficiency multilingual streaming ASR. Chao Zhang 0031, Bo Li 0028, Tara N. Sainath, Trevor Strohman, Shuo-Yiin Chang |
ICASSP | 1 |
| 2023 | A Neural Time Alignment Module for End-to-End Automatic Speech Recognition
Dongcheng Jiang, Chao Zhang 0031, Philip C. Woodland |
INTERSPEECH | 2 |
| 2023 | Pushing the Limits of Unsupervised Unit Discovery for SSL Speech Representation
Ziyang Ma 0001, Zhisheng Zheng, Guanrou Yang, Yu Wang 0027, Chao Zhang 0031, Xie Chen 0001 |
INTERSPEECH | 5 |
| 2023 | Can Contextual Biasing Remain Effective with Whisper and GPT-2?
Guangzhi Sun, Xianrui Zheng, Chao Zhang 0031, Philip C. Woodland |
INTERSPEECH | 3 |
| 2023 | Integrating Emotion Recognition with Speech Recognition and Speaker Diarisation for ConversationsabstractAlthough automatic emotion recognition (AER) has recently drawn significant research interest, most current AER studies use manually segmented utterances, which are usually unavailable for dialogue systems. This paper proposes integrating AER with automatic speech recognition (ASR) and speaker diarisation (SD) in a jointly-trained system. Distinct output layers are built for four sub-tasks including AER, ASR, voice activity detection and speaker classification based on a shared encoder. Taking the audio of a conversation as input, the integrated system finds all speech segments and transcribes the corresponding emotion classes, word sequences, and speaker identities. Two metrics are proposed to evaluate AER performance with automatic segmentation based on time-weighted emotion and speaker classification errors. Results on the IEMOCAP dataset show that the proposed system consistently outperforms two baselines with separately trained single-task systems on AER, ASR and SD. Wen Wu 0007, Chao Zhang 0031, Philip C. Woodland |
INTERSPEECH | 2 |
| 2023 | Obstructive Sleep Apnea Detection using Pre-trained Speech Representations
Kaibo Zhang, Lili Cao, Yanru Li, Chao Zhang 0031, Ji Wu 0002, Demin Han |
INTERSPEECH | 5 |
| 2023 | Combining hybrid DNN-HMM ASR systems with attention-based models using lattice rescoringabstractThe traditional hybrid deep neural network (DNN)–hidden Markov model (HMM) system and attention-based encoder–decoder (AED) model are both commonly used automatic speech recognition (ASR) approaches with distinct characteristics and advantages. While hybrid systems are per-frame-based and highly modularised to leverage external phonetic and linguistic knowledge, AED models operate on a per-label basis and jointly learn the acoustic and language information using a single model in an end-to-end trainable fashion. In this paper, we propose combining these two approaches in a two-pass rescoring framework. The first-pass uses hybrid ASR systems to facilitate streaming and controllable ASR, and the second-pass re-scores the N-best hypotheses or lattices produced by the first-pass hybrid DNN-HMM system with AED models. We also propose an improved algorithm for lattice rescoring with AED models. Experiments show the combined two-pass systems achieve competitive performance without using extra speech or text data on two standard ASR tasks. For the 80-hour AMI IHM dataset, the combined system has a 13.7% word error rate (WER) on the evaluation set and is up to a 29% relative WER reduction over the individual systems. For the 300-hour Switchboard dataset, the WERs of the combined system are 5.7% and 12.1% on Switchboard and CallHome subsets of Hub5’00, and 13.2% and 7.6% on Switchboard Cellular and Fisher subsets of RT03, and are up to a 33% relative reduction in WER over the individual systems. Qiujia Li, Chao Zhang 0031, Philip C. Woodland |
Speech Commun. | 2 |
| 2023 | Estimating the Uncertainty in Emotion Class Labels With Utterance-Specific Dirichlet PriorsabstractEmotion recognition is a key attribute for artificial intelligence systems that need to naturally interact with humans. However, the task definition is still an open problem due to the inherent ambiguity of emotions. In this paper, a novel Bayesian training loss based on per-utterance Dirichlet prior distributions is proposed for verbal emotion recognition, which models the uncertainty in one-hot labels created when human annotators assign the same utterance to different emotion classes. An additional metric is used to evaluate the performance by detecting test utterances with high labelling uncertainty. This removes a major limitation that emotion classification systems only consider utterances with labels where the majority of annotators agree on the emotion class. Furthermore, a frequentist approach is studied to leverage the continuous-valued “soft” labels obtained by averaging the one-hot labels. We propose a two-branch model structure for emotion classification on a per-utterance basis, which achieves state-of-the-art classification results on the widely used IEMOCAP dataset. Based on this, uncertainty estimation experiments were performed. The best performance in terms of the area under the precision-recall curve when detecting utterances with high uncertainty was achieved by interpolating the Bayesian training loss with the Kullback-Leibler divergence training loss for the soft labels. The generality of the proposed approach was verified using the MSP-Podcast dataset which yielded the same pattern of results. Wen Wu 0007, Chao Zhang 0031, Xixin Wu, Philip C. Woodland |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | Minimising Biasing Word Errors for Contextual ASR With the Tree-Constrained Pointer GeneratorabstractContextual knowledge is essential for reducing speech recognition errors on high-valued long-tail words. This paper proposes a novel tree-constrained pointer generator (TCPGen) component that enables end-to-end ASR models to bias towards a list of long-tail words obtained using external contextual information. With only a small overhead in memory use and computation cost, TCPGen can structure thousands of biasing words efficiently into a symbolic prefix-tree, and creates a neural shortcut between the tree and the final ASR output to facilitate the recognition of the biasing words. To enhance TCPGen, we further propose a novel minimum biasing word error (MBWE) loss that directly optimises biasing word errors during training, along with a biasing-word-driven language model discounting (BLMD) method during the test. All contextual ASR systems were evaluated on the public Librispeech audiobook corpus and the data from the dialogue state tracking challenges (DSTC) with the biasing lists extracted from the dialogue-system ontology. Consistent word error rate (WER) reductions were achieved with TCPGen, which were particularly significant on the biasing words with around 40% relative reductions in the recognition error rates. MBWE and BLMD further improved the effectiveness of TCPGen, and achieved more significant WER reductions on the biasing words. TCPGen also achieved zero-shot learning of words not in the audio training set with large WER reductions on the out-of-vocabulary words in the biasing list. Guangzhi Sun, Chao Zhang 0031, Philip C. Woodland |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Prosody Modelling With Pre-Trained Cross-Utterance Representations for Improved Speech SynthesisabstractWhen humans speak multiple utterances in a continuous manner, the prosodic features generated in each utterance are related to those in its neighbouring utterances. Such cross-utterance (CU) dependencies are often ignored by the current neural text-to-speech (TTS) systems, which reduces the naturalness and expressiveness of the synthesized speeches. In this paper, we propose to improve the prosody modelling ability of neural TTS systems using pre-trained CU acoustic and text representations. Such CU acoustic representations are derived using the Wav2Vec 2.0 model (W2V2) from the synthesized audios of the past utterances, while the CU text representations are extracted using the Bidirectional Encoder Representation from Transformers (BERT) model from the scripts of the future utterances. Experimental results on a Mandarin audiobook and an English audiobook showed the naturalness and expressiveness of the synthesized audios were significantly improved by incorporating such pre-trained W2V2 and BERT CU representations into the Fastspeech2 TTS framework. Ya-Jie Zhang, Chao Zhang 0031, Zhengchen Zhang, Youzheng Wu, Xiaodong He 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Improving the Fusion of Acoustic and Text Representations in RNN-TabstractThe recurrent neural network transducer (RNN-T) has recently become the mainstream end-to-end approach for streaming automatic speech recognition (ASR). To estimate the output distributions over subword units, RNN-T uses a fully connected layer as the joint network to fuse the acoustic representations extracted using the acoustic encoder with the text representations obtained using the prediction network based on the previous subword units. In this paper, we propose to use gating, bilinear pooling, and a combination of them in the joint network to produce more expressive representations to feed into the output layer. A regularisation method is also proposed to enable better acoustic encoder training by reducing the gradients back-propagated into the prediction network at the beginning of RNN-T training. Experimental results on a multilingual ASR setting for voice search over nine languages show that the joint use of the proposed methods can result in 4%–5% relative word error rate reductions with only a few million extra parameters. Chao Zhang 0031, Bo Li 0028, Zhiyun Lu, Tara N. Sainath, Shuo-Yiin Chang |
ICASSP | 1 |
| 2022 | Turn-Taking Prediction for Natural Conversational Speech
Shuo-Yiin Chang, Bo Li 0028, Tara N. Sainath, Chao Zhang 0031, Trevor Strohman, Qiao Liang 0001, Yanzhang He |
INTERSPEECH | 4 |
| 2022 | Tree-constrained Pointer Generator with Graph Neural Network Encodings for Contextual Speech RecognitionabstractIncorporating biasing words obtained as contextual knowledge is critical for many automatic speech recognition (ASR) applications.This paper proposes the use of graph neural network (GNN) encodings in a tree-constrained pointer generator (TCP-Gen) component for end-to-end contextual ASR.By encoding the biasing words in the prefix-tree with a tree-based GNN, lookahead for future wordpieces in end-to-end ASR decoding is achieved at each tree node by incorporating information about all wordpieces on the tree branches rooted from it, which allows a more accurate prediction of the generation probability of the biasing words.Systems were evaluated on the Librispeech corpus using simulated biasing tasks, and on the AMI corpus by proposing a novel visual-grounded contextual ASR pipeline that extracts biasing words from slides alongside each meeting.Results showed that TCPGen with GNN encodings achieved about a further 15% relative WER reduction on the biasing words compared to the original TCPGen, with a negligible increase in the computation cost for decoding. Guangzhi Sun, Chao Zhang 0031, Philip C. Woodland |
INTERSPEECH | 2 |
| 2022 | Streaming End-to-End Multilingual Speech Recognition with Joint Language Identification
Chao Zhang 0031, Bo Li 0028, Tara N. Sainath, Trevor Strohman, Sepand Mavandadi, Shuo-Yiin Chang, Parisa Haghani |
INTERSPEECH | 1 |
| 2022 | Tandem Multitask Training of Speaker Diarisation and Speech Recognition for Meeting TranscriptionabstractSelf-supervised-learning-based pre-trained models for speech data, such as Wav2Vec 2.0 (W2V2), have become the backbone of many speech tasks. In this paper, to achieve speaker diarisation and speech recognition using a single model, a tandem multitask training (TMT) method is proposed to fine-tune W2V2. For speaker diarisation, the tasks of voice activity detection (VAD) and speaker classification (SC) are required, and connectionist temporal classification (CTC) is used for ASR. The multitask framework implements VAD, SC, and ASR using an early layer, middle layer, and late layer of W2V2, which coincides with the order of segmenting the audio with VAD, clustering the segments based on speaker embeddings, and transcribing each segment with ASR. Experimental results on the augmented multi-party (AMI) dataset showed that using different W2V2 layers for VAD, SC, and ASR from the earlier to later layers for TMT not only saves computational cost, but also reduces diarisation error rates (DERs). Joint fine-tuning of VAD, SC, and ASR yielded 16%/17% relative reductions of DER with manual/automatic segmentation respectively, and consistent reductions in speaker attributed word error rate, compared to the baseline with separately fine-tuned models. Xianrui Zheng, Chao Zhang 0031, Philip C. Woodland |
INTERSPEECH | 2 |
| 2022 | Unified End-to-End Speech Recognition and Endpointing for Fast and Efficient Speech SystemsabstractAutomatic speech recognition (ASR) systems typically rely on an external endpointer (EP) model to identify speech boundaries. In this work, we propose a method to jointly train the ASR and EP tasks in a single end-to-end (E2E) multitask model, improving EP quality by optionally leveraging information from the ASR audio encoder. We introduce a “switch” connection, which trains the EP to consume either the audio frames directly or low-level latent representations from the ASR model. This results in a single E2E model that can be used during inference to perform frame filtering at low cost, and also make high quality end-of-query (EOQ) predictions based on ongoing ASR computation. We present results on a voice search test set showing that, compared to separate single-task models, this approach reduces median endpoint latency by 120 ms (30.8% reduction), and 90th percentile latency by 170 ms (23.0% reduction), without regressing word error rate. For continuous recognition, WER improves by 10.6% (relative). Shaan Bijwadia, Shuo-Yiin Chang, Bo Li 0028, Tara N. Sainath, Chao Zhang 0031, Yanzhang He |
SLT | 5 |
| 2022 | A Truly Multilingual First Pass and Monolingual Second Pass Streaming on-Device ASR SystemabstractAutomatic speech recognition (ASR) systems need to be accurate, have low latency, and effectively handle language switching in order to be useful for the 60% of the world population that speaks more than one language. Thus, we propose a truly multilingual first-pass and monolingual second-pass streaming on-device ASR system based on the recently developed Cascaded Encoders model. The streaming first-pass recognizes multilingual speech without needing language information, providing real-time transcription, even for code-switching speech. The second-pass uses a language dependent right context encoder to improve the recognition accuracy. On a 9 language Voice Search task, we find that a system combining shared causal encoder with decoders and non-causal encoders replicated per-language reduces word error rate (WER) by 4.4% relative to monolingual baselines. We further show this design to be parameter efficient, outperforming other architectures when matched in the number of parameters. Sepand Mavandadi, Bo Li 0028, Chao Zhang 0031, Brian Farris, Tara N. Sainath, Trevor Strohman |
SLT | 3 |
| 2022 | Distribution-Based Emotion Recognition in ConversationabstractAutomatic emotion recognition in conversation (ERC) is crucial for emotion-aware conversational artificial intelligence. This paper proposes a distribution-based framework that formulates ERC as a sequence-to-sequence problem for emotion distribution estimation. The inherent ambiguity of emotions and the subjectivity of human perception lead to disagreements in emotion labels, which is handled naturally in our framework from the perspective of uncertainty estimation in emotion distributions. A Bayesian training loss is introduced to improve the uncertainty estimation by conditioning each emotional state on an utterance-specific Dirichlet prior distribution. Experimental results on the IEMOCAP dataset show that ERC outperformed the single-utterance-based system, and the proposed distribution-based ERC methods have not only better classification accuracy, but also show improved uncertainty estimation. Wen Wu 0007, Chao Zhang 0031, Philip C. Woodland |
SLT | 2 |
| 2021 | Tree-Constrained Pointer Generator for End-to-End Contextual Speech RecognitionabstractContextual knowledge is important for real-world automatic speech recognition (ASR) applications. In this paper, a novel tree-constrained pointer generator (TCPGen) component is proposed that incorpo-rates such knowledge as a list of biasing words into both attention-based encoder-decoder and transducer end-to-end ASR models in a neural-symbolic way. TCPGen structures the biasing words into an efficient prefix tree to serve as its symbolic input and creates a neu-ral shortcut between the tree and the final ASR output distribution to facilitate recognising biasing words during decoding. Systems were trained and evaluated on the Librispeech corpus where biasing words were extracted at the scales of an utterance, a chapter, or a book to simulate different application scenarios. Experimental results showed that TCPGen consistently improved word error rates (WERs) compared to the baselines, and in particular, achieved sig-nificant WER reductions on the biasing words. TCPGen is highly efficient: it can handle 5,000 biasing words and distractors and only add a small overhead to memory use and computation cost. Guangzhi Sun, Chao Zhang 0031, Philip C. Woodland |
ASRU | 2 |
| 2021 | Adapting GPT, GPT-2 and BERT Language Models for Speech RecognitionabstractLanguage models (LMs) pre-trained on massive amounts of text, in particular bidirectional encoder representations from Transformers (BERT), generative pre-training (GPT), and GPT-2, have become a key technology for many natural language processing tasks. In this paper, we present results using fine-tuned GPT, GPT-2, and their combination for automatic speech recognition (ASR). Unlike unidirectional LM GPT and GPT-2, BERT is bidirectional whose direct product of the output probabilities is no longer a valid language prior probability. A conversion method is proposed to compute the correct language prior probability based on bidirectional LM outputs in a mathematically exact way. Experimental results on the widely used AMI and Switchboard ASR tasks showed that the combination of the fine-tuned GPT and GPT-2 outperformed the combination of three neural LMs with different architectures trained from scratch on the in-domain text by up to a 12% relative word error rate reduction (WERR). Furthermore, on the AMI corpus, the proposed conversion for language prior probabilities enables BERT to obtain an extra 3% relative WERR, and the combination of BERT, GPT and GPT-2 results in further improvements. Xianrui Zheng, Chao Zhang 0031, Philip C. Woodland |
ASRU | 2 |
| 2021 | Dian: Duration Informed Auto-Regressive Network for Voice CloningabstractIn this paper, we propose a novel end-to-end speech synthesis approach, Duration Informed Auto-regressive Network (DIAN), which consists of an acoustic model and a separate duration model. Un-like other auto-regressive TTS methods, the duration information of phonemes is provided as part of the input to the acoustic model, which enables the removal of the attention mechanism between its encoder and decoder parts. This eliminates the common seen skipping and repeating issues and improves speech intelligibility while ensuring high speech quality. A Transformer-based duration model is used to predict the duration of each phoneme for the attention-free acoustic model. We developed our TTS systems for the multi-speaker multi-style voice cloning challenge (M2VoC) using the proposed DIAN approach. In our procedure, a multi-speaker attention-free acoustic model and its Transformer-based duration model are first separately trained based on the training data released by M2VoC. Next, the multi-speaker models are adapted to form the speaker-specific models with the speaker-dependent data and transfer learning. At last, a speaker-specific LPCNet is estimated and used to synthesize the speech of the corresponding speaker. The M2VoC results showed that our proposed approach achieved the 3rd-place in the speech quality ranking and the 4th-place in the speaker similarity and style similarity ranking in the Track1-a task. Zhengchen Zhang, Chao Zhang 0031, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
ICASSP | 4 |
| 2021 | Transformer Language Models with LSTM-Based Cross-Utterance Information RepresentationabstractThe effective incorporation of cross-utterance information has the potential to improve language models (LMs) for automatic speech recognition (ASR). To extract more powerful and robust cross-utterance representations for the Transformer LM (TLM), this paper proposes the R-TLM which uses hidden states in a long short-term memory (LSTM) LM. To encode the cross-utterance information, the R-TLM incorporates an LSTM module together with a segment-wise recurrence in some of the Transformer blocks. In addition to the LSTM module output, a shortcut connection using a fusion layer which bypasses the LSTM module is also investigated. The proposed system was evaluated on the AMI meeting corpus, the Eval2000 and the RT03 telephone conversation evaluation sets. The best R-TLM achieved 0.9%, 0.6% and 0.8% absolute WER reductions over the single-utterance TLM baseline, and 0.5%, 0.3%, 0.2% absolute WER reductions over a strong cross-utterance TLM baseline on the AMI evaluation set, Eval2000 and RT03 respectively. Improvements on Eval2000 and RT03 were further supported by significance tests. R-TLMs were found to have better LM scores on words where recognition errors are more likely to occur. The R-TLM WER can be further reduced by interpolation with an LSTM-LM. Guangzhi Sun, Chao Zhang 0031, Philip C. Woodland |
ICASSP | 2 |
| 2021 | Content-Aware Speaker Embeddings for Speaker DiarisationabstractRecent speaker diarisation systems often convert variable length speech segments into fixed-length vector representations for speaker clustering, which are known as speaker embeddings. In this paper, the content-aware speaker embeddings (CASE) approach is proposed, which extends the input of the speaker classifier to include not only acoustic features but also their corresponding speech content, via phone, character, and word embeddings. Compared to alternative methods that leverage similar information, such as multitask or adversarial training, CASE factorises automatic speech recognition (ASR) from speaker recognition to focus on modelling speaker characteristics and correlations with the corresponding content units to derive more expressive representations. CASE is evaluated for speaker re-clustering with a realistic speaker diarisation setup using the AMI meeting transcription dataset, where the content information is obtained by performing ASR based on an automatic segmentation. Experimental results showed that CASE achieved a 17.8% relative speaker error rate reduction over conventional methods. Guangzhi Sun, Chao Zhang 0031, Philip C. Woodland |
ICASSP | 3 |
| 2021 | Emotion Recognition by Fusing Time Synchronous and Time Asynchronous RepresentationsabstractIn this paper, a novel two-branch neural network model structure is proposed for multimodal emotion recognition, which consists of a time synchronous branch (TSB) and a time asynchronous branch (TAB). To capture correlations between each word and its acoustic realisation, the TSB combines speech and text modalities at each input window frame and then uses pooling across time to form a single embedding vector. The TAB, by contrast, provides cross-utterance information by integrating sentence text embeddings from a number of context utterances into another embedding vector. The final emotion classification uses both the TSB and the TAB embeddings. Experimental results on the IEMOCAP dataset demonstrate that the two-branch structure achieves state-of-the-art results in 4-way classification with all common test setups. When using automatic speech recognition (ASR) output instead of manually transcribed reference text, it is shown that the cross-utterance information considerably improves robustness against ASR errors. Furthermore, by incorporating an extra class for all the other emotions, the final 5-way classification system with ASR hypotheses can be viewed as a prototype for more realistic emotion recognition systems. Wen Wu 0007, Chao Zhang 0031, Philip C. Woodland |
ICASSP | 2 |
| 2021 | Improving Prosody Modelling with Cross-Utterance Bert Embeddings for End-to-End Speech SynthesisabstractAlthough speech prosody is related to the linguistic information up to the discourse structure, most text-to-speech (TTS) systems only take into account the information within each sentence. This makes it challenging when converting a paragraph of text into natural and expressive speech. In this paper, we propose to use the text embeddings of the neighboring sentences to improve the prosody generation for each utterance of a paragraph in an end-to-end fashion without using any explicit prosody features. More specifically, cross-utterance (CU) context vectors, which are produced by an additional CU encoder based on the sentence embeddings extracted by a pretrained BERT model, are used to augment the input of the Tacotron2 decoder. Two types of BERT embeddings are investigated, which leads to the use of different CU encoder structures. Experimental results on a Mandarin audiobook dataset and the LJ-Speech English audiobook dataset demonstrate the use of CU information can improve the naturalness and expressiveness of the synthesized speech. Subjective listening testing shows most of the participants prefer the voice generated using the CU encoder over that generated using standard Tacotron2. It is also found that the prosody can be controlled indirectly by changing the neighbouring sentences. Zhengchen Zhang, Chao Zhang 0031, Xiaodong He 0001, Bowen Zhou 0001 |
ICASSP | 4 |
| 2021 | Neural Kalman Filtering for Speech EnhancementabstractConventional learning-based speech enhancement methods usually utilize existing building blocks to design the deep neural networks (DNNs), while how to effectively integrate the statistical signal processing based schemes, which are expert-knowledge driven and could ameliorate the over-fitting problem, into the network design remains an open issue. In this paper, we extend the conventional Kalman filtering (KF) and propose a supervised-learning based neural Kalman filter (NKF) for speech enhancement. Similar to KF, the proposed method first obtains a prediction from the speech evolution model and then integrates the short-term instantaneous observation by linear weighting, and the weights are calculated by comparing between the speech prediction residual error and the environmental noise level. An end-to-end network is designed to convert the speech linear prediction model in KF to non-linear, and to compact all other conventional linear filtering operations. Different with other DNN based methods, the proposed method provides a specialized network design inspired from the conventional signal processing, the backpropagation can be directly applied on the linear filtering operations integrated from KF. We conduct experiments in different noisy conditions, and the results demonstrate that the proposed method outperforms the baseline methods which are based on either signal processing or DNNs. Wei Xue 0002, Gang Quan, Chao Zhang 0031, Guo-Hong Ding, Xiaodong He 0001, Bowen Zhou 0001 |
ICASSP | 3 |
| 2021 | Variable Frame Rate Acoustic Models Using Minimum Error Reinforcement Learning
Dongcheng Jiang, Chao Zhang 0031, Philip C. Woodland |
Interspeech | 2 |
| 2021 | Discriminative Neural Clustering for Speaker DiarisationabstractIn this paper, we propose Discriminative Neural Clustering (DNC) that formulates data clustering with a maximum number of clusters as a supervised sequence-to-sequence learning problem. Com-pared to traditional unsupervised clustering algorithms, DNC learns clustering patterns from training data without requiring an explicit definition of a similarity measure. An implementation of DNC based on the Transformer architecture is shown to be effective on a speaker diarisation task using the challenging AMI dataset. Since AMI contains only 147 complete meetings as individual input sequences, data scarcity is a significant issue for training a Transformer model for DNC. Accordingly, this paper proposes three data augmentation schemes: sub-sequence randomisation, input vector randomisation, and Diaconis augmentation, which generates new data samples by rotating the entire input sequence of L2-normalised speaker embeddings. Experimental results on AMI show that DNC achieves a reduction in speaker error rate (SER) of 29.4% relative to spectral clustering. Qiujia Li, Florian Kreyssig, Chao Zhang 0031, Philip C. Woodland |
SLT | 3 |
| 2021 | A distributed optimisation framework combining natural gradient with Hessian-free for discriminative sequence training
Adnan Haider, Chao Zhang 0031, Florian Kreyssig, Philip C. Woodland |
Neural Networks | 2 |
| 2021 | Combination of deep speaker embeddings for diarisation
Guangzhi Sun, Chao Zhang 0031, Philip C. Woodland |
Neural Networks | 2 |
| 2020 | Improved Large-Margin Softmax Loss for Speaker DiarisationabstractSpeaker diarisation systems nowadays use embeddings generated from speech segments in a bottleneck layer, which are needed to be discriminative for unseen speakers. It is well-known that large-margin training can improve the generalisation ability to unseen data, and its use in such open-set problems has been widespread. Therefore, this paper introduces a general approach to the large-margin softmax loss without any approximations to improve the quality of speaker embeddings for diarisation. Furthermore, a novel and simple way to stabilise training, when large-margin softmax is used, is proposed. Finally, to combat the effect of overlapping speech, different training margins are used to reduce the negative effect overlapping speech has on creating discriminative embeddings. Experiments on the AMI meeting corpus show that the use of large-margin softmax significantly improves the speaker error rate (SER). By using all hyper parameters of the loss in a unified way, further improvements were achieved which reached a relative SER reduction of 24.6% over the baseline. However, by training overlapping and single speaker speech samples with different margins, the best result was achieved, giving overall a 29.5% SER reduction relative to the baseline. Yassir Fathullah, Chao Zhang 0031, Philip C. Woodland |
ICASSP | 2 |
| 2020 | Efficient WaveGlow: An Improved WaveGlow Vocoder with Enhanced Speed
Zhengchen Zhang, Chao Zhang 0031, Xiaodong He 0001, Bowen Zhou 0001 |
INTERSPEECH | 4 |
| 2020 | The JD AI Speaker Verification System for the FFSVC 2020 ChallengeabstractThis paper presents the development of our systems for the Interspeech 2020 Far-Field Speaker Verification Challenge (FFSVC). Our focus is the task 2 of the challenge, which is to perform far-field text-independent speaker verification using a single microphone array. The FFSVC training set provided by the challenge is augmented by pre-processing the far-field data with both beamforming, voice channel switching, and a combination of weighted prediction error (WPE) and beamforming. Two open-access corpora, CHData in Mandarin and VoxCeleb2 in English, are augmented using multiple methods and mixed with the augmented FFSVC data to form the final training data. Four different model structures are used to model speaker characteristics: ResNet, extended time-delay neural network (ETDNN), Transformer, and factorized TDNN (FTDNN), whose output values are pooled across time using the self-attentive structure, the statistic pooling structure, and the GVLAD structure. The final results are derived by fusing the adaptively normalized scores of the four systems with a two-stage fusion method, which achieves a minimum of the detection cost function (minDCF) of 0.3407 and an equal error rate (EER) of 2.67% on the development set of the challenge. Ying Tong, Wei Xue 0002, Shanluo Huang, Fan Lu 0003, Chao Zhang 0031, Guo-Hong Ding, Xiaodong He 0001 |
INTERSPEECH | 5 |
| 2020 | Sound Event Localization and Detection Based on Multiple DOA Beamforming and Multi-Task LearningabstractThe performance of sound event localization and detection (SELD) degrades in source-overlapping cases since features of different sources collapse with each other, and the network tends to fail to learn to separate these features effectively. In this paper, by leveraging the conventional microphone array signal processing to generate comprehensive representations for SELD, we propose a new SELD method based on multiple direction of arrival (DOA) beamforming and multi-task learning. By using multiple beamformers to extract the signals from different DOAs, the sound field is more diversely described, and specialised representations of target source and noises can be obtained. With labelled training data, the steering vector is estimated based on the cross-power spectra (CPS) and the signal presence probability (SPP), which eliminates the need of knowing the array geometry. We design two networks for sound event localization (SED) and sound source localization (SSL) and use a multi-task learning scheme for SED, in which the SSL-related task act as a regularization. Experimental results using the database of DCASE2019 SELD task show that the proposed method achieves the state-of-art performance. Wei Xue 0002, Ying Tong, Chao Zhang 0031, Guo-Hong Ding, Xiaodong He 0001, Bowen Zhou 0001 |
INTERSPEECH | 3 |
| 2019 | Integrating Source-Channel and Attention-Based Sequence-to-Sequence Models for Speech RecognitionabstractThis paper proposes a novel automatic speech recognition (ASR) framework called Integrated Source-Channel and Attention (ISCA) that combines the advantages of traditional systems based on the noisy source-channel model (SC) and end-to-end style systems using attention-based sequence-to-sequence models. The traditional SC system framework includes hidden Markov models and connectionist temporal classification (CTC) based acoustic models, language models (LMs), and a decoding procedure based on a lexicon, whereas the end-to-end style attention-based system jointly models the whole process with a single model. By rescoring the hypotheses produced by traditional systems using end-to-end style systems based on an extended noisy source-channel model, ISCA allows structured knowledge to be easily incorporated via the SC-based model while exploiting the complementarity of the attention-based model. Experiments on the AMI meeting corpus show that ISCA is able to give a relative word error rate reduction up to 21% over an individual system, and by 13% over an alternative method which also involves combining CTC and attention-based models. Qiujia Li, Chao Zhang 0031, Philip C. Woodland |
ASRU | 2 |
| 2019 | Speaker Diarisation Using 2D Self-attentive Combination of EmbeddingsabstractSpeaker diarisation systems often cluster audio segments using speaker embeddings such as i-vectors and d-vectors. Since different types of embeddings are often complementary, this paper proposes a generic framework to improve performance by combining them into a single embedding, referred to as a c-vector. This combination uses a 2-dimensional (2D) self-attentive structure, which extends the standard self-attentive layer by averaging not only across time but also across different types of embeddings. Two types of 2D self-attentive structure studied in this paper are simultaneous combination and consecutive combination, which adopt single and multiple self-attentive layers respectively. The penalty term in the original self-attentive layer, which is jointly minimised with the objective function to encourage diversity of annotation vectors, is also modified to obtain not only different local peaks but also the overall trends in the multiple annotation vectors. Experiments on the AMI meeting corpus show that our modified penalty term improves the d-vector relative speaker error rate (SER) by 6% and 21% for d-vector systems, and a 10% further relative SER reduction can be obtained using the c-vector from our best 2D self-attentive structure. Guangzhi Sun, Chao Zhang 0031, Philip C. Woodland |
ICASSP | 2 |
| 2019 | PyHTK: Python Library and ASR Pipelines for HTKabstractThis paper describes PyHTK, which is a Python-based library and associated pipeline to facilitate the construction of large-scale complex automatic speech recognition (ASR) systems using the hidden Markov model toolkit (HTK). PyHTK can be used to generate sophisticated artificial neural network (ANN) models with versatile architectures by converting a compact configuration file defining the ANN, into the form used by HTK tools, as well as supporting a range of capabilities to train and test ANN models. The ASR pipeline is divided into multiple steps, which can be arranged and customised for different ASR data sets, and allows for both step-by-step and fully automatic end-to-end operation. PyHTK is integrated with HTK 3.5.1 which includes an expanded range of ANN layer types and very flexible ways to connect them, together with capabilities for ASR training and testing. Some example systems are included to illustrate the flexibility and performance achievable. Chao Zhang 0031, Florian Kreyssig, Qiujia Li, Philip C. Woodland |
ICASSP | 1 |
| 2019 | Multi-Span Acoustic Modelling Using Raw Waveform SignalsabstractTraditional automatic speech recognition (ASR) systems often use an acoustic model (AM) built on handcrafted acoustic features, such as log Mel-filter bank (FBANK) values. Recent studies found that AMs with convolutional neural networks (CNNs) can directly use the raw waveform signal as input. Given sufficient training data, these AMs can yield a competitive word error rate (WER) to those built on FBANK features. This paper proposes a novel multi-span structure for acoustic modelling based on the raw waveform with multiple streams of CNN input layers, each processing a different span of the raw waveform signal. Evaluation on both the single channel CHiME4 and AMI data sets show that multi-span AMs give a lower WER than FBANK AMs by an average of about 5% (relative). Analysis of the trained multi-span model reveals that the CNNs can learn filters that are rather different to the log Mel filters. Furthermore, the paper shows that a widely used single span raw waveform AM can be improved by using a smaller CNN kernel size and increased stride to yield improved WERs. Patrick von Platen, Chao Zhang 0031, Philip C. Woodland |
INTERSPEECH | 2 |
| 2019 | Direct-Path Signal Cross-Correlation Estimation for Sound Source Localization in ReverberationabstractSound source localization (SSL) is challenging in presence of reverberation since the cross-correlation between the direct-path signals in different microphones, which indicates the spatial information of the sound source, is interfered by the reverberation signal components. A novel algorithm is proposed in this paper to estimate the cross-correlation of the direct-path speech signals, such that the robustness of SSL to reverberation can be improved. The proposed method follows a similar scheme to the multichannel linear prediction (MCLP), which is commonly used for speech dereverberation, while avoids the explicit estimation of the direct-path signal of each channel. This is achieved by revealing the relationship between the direct-path signal cross-correlation (DPCC) and the MCLP coefficient vector, and finally deriving the DPCC by using only the multichannel reverberant signals. It is also shown that the pre-whitening operation, which is widely used for SSL, can be inherently integrated into the estimated DPCC. An adaptive method is further derived to facilitate online frame-level SSL. The proposed method can be easily applied to conventional cross-correlation based SSL methods by using the DPCC rather than the full cross-correlation. Experiments conducted in various reverberant conditions demonstrate the effectiveness of the proposed method. Wei Xue 0002, Ying Tong, Guo-Hong Ding, Chao Zhang 0031, Xiaodong He 0001, Bowen Zhou 0001 |
INTERSPEECH | 4 |
| 2018 | Improved Tdnns Using Deep Kernels and Frequency Dependent Grid-RNNSabstractTime delay neural networks (TDNNs) are an effective acoustic model for large vocabulary speech recognition. The strength of the model can be attributed to its ability to effectively model long temporal contexts. However, current TDNN models are relatively shallow, which limits the modelling capability. This paper proposes a method of increasing the network depth by deepening the kernel used in the TDNN temporal convolutions. The best performing kernel consists of three fully connected layers with a residual (ResNet) connection from the output of the first to the output of the third. The addition of spectro-temporal processing as the input to the TDNN in the form of a convolutional neural network (CNN) and a newly designed Grid-RNN was investigated. The Grid-RNN strongly outperforms a CNN if different sets of parameters for different frequency bands are used and can be further enhanced by using a bi-directional Grid-RNN. Experiments using the multi-genre broadcast (MGB3) English data (275h) show that deep kernel TDNNs reduces the word error rate (WER) by 6% relative and when combined with the frequency dependent Grid-Rnn gives a relative WER reduction of 9%. Florian Kreyssig, Chao Zhang 0031, Philip C. Woodland |
ICASSP | 2 |
| 2018 | High Order Recurrent Neural Networks for Acoustic ModellingabstractVanishing long-term gradients are a major issue in training standard recurrent neural networks (RNNs), which can be alleviated by long short-term memory (LSTM) models with memory cells. However, the extra parameters associated with the memory cells mean an LSTM layer has four times as many parameters as an RNN with the same hidden vector size. This paper addresses the vanishing gradient problem using a high order RNN (HORNN) which has additional connections from multiple previous time steps. Speech recognition experiments using British English multi-genre broadcast (MGB3) data showed that the proposed HORNN architectures for rectified linear unit and sigmoid activation functions reduced word error rates (WER) by 4.2% and 6.3% over the corresponding RNNs, and gave similar WERs to a (projected) LSTM while using only 20%-50% of the recurrent layer parameters and computation. Chao Zhang 0031, Philip C. Woodland |
ICASSP | 1 |
| 2018 | Speaker Adaptation and Adaptive Training for Jointly Optimised Tandem SystemsabstractSpeaker independent (SI) Tandem systems trained by joint optimisation of bottleneck (BN) deep neural networks (DNNs) and Gaussian mixture models (GMMs) have been found to produce similar word error rates (WERs) to Hybrid DNN systems. A key advantage of using GMMs is that existing speaker adaptation methods, such as maximum likelihood linear regression (MLLR), can be used which to account for diverse speaker variations and improve system robustness. This paper investigates speaker adaptation and adaptive training (SAT) schemes for jointly optimised Tandem systems. Adaptation techniques investigated include constrained MLLR (CMLLR) transforms based on BN features for SAT as well as MLLR and parameterised sigmoid functions for unsupervised test-time adaptation. Experiments using English multi-genre broadcast (MGB3) data show that CMLLR SAT yields a 4% relative WER reduction over jointly trained Tandem and Hybrid SI systems, and further reductions in WER are obtained by system combination. Yu Wang 0027, Chao Zhang 0031, Mark J. F. Gales, Philip C. Woodland |
INTERSPEECH | 2 |
| 2018 | Semi-tied Units for Efficient Gating in LSTM and Highway NetworksabstractGating is a key technique used for integrating information from multiple sources by long short-term memory (LSTM) models and has recently also been applied to other models such as the highway network.Although gating is powerful, it is rather expensive in terms of both computation and storage as each gating unit uses a separate full weight matrix.This issue can be severe since several gates can be used together in e.g. an LSTM cell.This paper proposes a semi-tied unit (STU) approach to solve this efficiency issue, which uses one shared weight matrix to replace those in all the units in the same layer.The approach is termed "semi-tied" since extra parameters are used to separately scale each of the shared output values.These extra scaling factors are associated with the network activation functions and result in the use of parameterised sigmoid, hyperbolic tangent, and rectified linear unit functions.Speech recognition experiments using British English multi-genre broadcast data showed that using STUs can reduce the calculation and storage cost by a factor of three for highway networks and four for LSTMs, while giving similar word error rates to the original models. Chao Zhang 0031, Philip C. Woodland |
INTERSPEECH | 1 |
| 2017 | Joint optimisation of tandem systems using Gaussian mixture density neural network discriminative sequence trainingabstractThe use of deep neural networks (DNNs) for feature extraction and Gaussian mixture models (GMMs) for acoustic modelling is often termed a tandem system configuration and can be viewed as a Gaussian mixture density neural network (MDNN). Compared to the direct use of DNN output probabilities in the acoustic model, the tandem approach suffers from a major weakness in that the feature extraction stage and the final acoustic models are optimised separately. This paper proposes a joint optimisation approach to all the stages of the tandem acoustic model by using MDNN discriminative sequence training. A set of techniques is used to improve the training performance and stability. Experiments using the multi-genre broadcast (MGB) English data show that the proposed method produced a 6% relative lower word error rate (WER) than that of a traditional discriminatively trained tandem system. The resulting jointly optimised tandem systems are comparable in WER to hybrid DNN systems optimised using discriminative sequence training with the same number of parameters. Chao Zhang 0031, Philip C. Woodland |
ICASSP | 1 |
| 2017 | Relating dynamic brain states to dynamic machine states: Human and machine solutions to the speech recognition problemabstractThere is widespread interest in the relationship between the neurobiological systems supporting human cognition and emerging computational systems capable of emulating these capacities. Human speech comprehension, poorly understood as a neurobiological process, is an important case in point. Automatic Speech Recognition (ASR) systems with near-human levels of performance are now available, which provide a computationally explicit solution for the recognition of words in continuous speech. This research aims to bridge the gap between speech recognition processes in humans and machines, using novel multivariate techniques to compare incremental 'machine states', generated as the ASR analysis progresses over time, to the incremental 'brain states', measured using combined electro- and magneto-encephalography (EMEG), generated as the same inputs are heard by human listeners. This direct comparison of dynamic human and machine internal states, as they respond to the same incrementally delivered sensory input, revealed a significant correspondence between neural response patterns in human superior temporal cortex and the structural properties of ASR-derived phonetic models. Spatially coherent patches in human temporal cortex responded selectively to individual phonetic features defined on the basis of machine-extracted regularities in the speech to lexicon mapping process. These results demonstrate the feasibility of relating human and ASR solutions to the problem of speech recognition, and suggest the potential for further studies relating complex neural computations in human speech comprehension to the rapidly evolving ASR systems that address the same problem domain. Cai Wingfield, Li Su 0002, Xunying Liu, Chao Zhang 0031, Philip C. Woodland, Andrew Thwaites, Elisabeth Fonteneau, William D. Marslen-Wilson |
PLoS Comput. Biol. | 4 |
| 2016 | Improved DNN-based segmentation for multi-genre broadcast audioabstractAutomatic segmentation is a crucial initial processing step for processing multi-genre broadcast (MGB) audio. It is very challenging since the data exhibits a wide range of both speech types and background conditions with many types of non-speech audio. This paper describes a segmentation system for multi-genre broadcast audio with deep neural network (DNN) based speech/non-speech detection. A further stage of change-point detection and clustering is used to obtain homogeneous segments. Suitable DNN inputs, context window sizes and architectures are studied with a series of experiments using a large corpus of MGB television audio. For MGB transcription, the improved segmenter yields roughly half the increase in word error rate, over manual segmentation, compared to the baseline DNN segmenter supplied for the 2015 ASRU MGB challenge. Chao Zhang 0031, Philip C. Woodland, Mark J. F. Gales, Panagiota Karanasou, Pierre Lanchantin, Xunying Liu, Yanmin Qian |
ICASSP | 2 |
| 2016 | System combination with log-linear modelsabstractImproved speech recognition performance can often be obtained by combining multiple systems together. Joint decoding, where scores from multiple systems are combined during decoding rather than combining hypotheses, is one efficient approach for system combination. In standard joint decoding the frame log-likelihoods from each system are used as the scores. These scores are then weighted and summed to yield the final score for a frame. The system combination weights for this process are usually empirically set. In this paper, a recently proposed scheme for learning these system weights is investigated for a standard noise-robust speech recognition task, AURORA 4. High performance tandem and hybrid systems for this task are described. By applying state-of-the-art training approaches and configurations for the bottleneck features of the tandem system, the difference in performance between the tandem and hybrid systems is significantly smaller than usually observed on this task. A log-linear model is then used to estimate system weights between these systems. Training the system weights yields additional gains over empirically set system weights when used for decoding. Furthermore, when used in a lattice rescoring fashion, further gains can be obtained. Chao Zhang 0031, Anton Ragni, Mark J. F. Gales, Philip C. Woodland |
ICASSP | 2 |
| 2016 | DNN speaker adaptation using parameterised sigmoid and ReLU hidden activation functionsabstractThis paper investigates the use of parameterised sigmoid and rectified linear unit (ReLU) hidden activation functions in deep neural network (DNN) speaker adaptation. The sigmoid and ReLU parameterisation schemes from a previous study for speaker independent (SI) training are used. An adaptive linear factor associated with each sigmoid or ReLU hidden unit is used to scale the unit output value and create a speaker dependent (SD) model. Hence, DNN adaptation becomes re-weighting the importance of different hidden units for every speaker. This adaptation scheme is applied to both hybrid DNN acoustic modelling and DNN-based bottleneck (BN) feature extraction. Experiments using multi-genre British English television broadcast data show that the technique is effective in both directly adapting DNN acoustic models and the BN features, and combines well with other DNN adaptation techniques. Reductions in word error rate are consistently obtained using parameterised sigmoid and ReLU activation function for multiple hidden layer adaptation. Chao Zhang 0031, Philip C. Woodland |
ICASSP | 1 |
| 2016 | Selection of Multi-Genre Broadcast Data for the Training of Automatic Speech Recognition SystemsabstractThis paper compares schemes for the selection of multi-genre broadcast data and corresponding transcriptions for speech recognition model training. Selections of the same amount of data (700 hours) from lightly supervised alignments based on the same original subtitle transcripts are compared. Data segments were selected according to a maximum phone matched error rate between the lightly supervised decoding and the original transcript. The data selected with an improved lightly supervised system yields lower word error rates (WERs). Detailed comparisons of the data selected on carefully transcribed development data show how the selected portions match the true phone error rate for each genre. From a broader perspective, it is shown that for different genres, either the original subtitles or the lightly supervised output should be used for model training and a suitable combination yields further reductions in final WER. Pierre Lanchantin, Mark J. F. Gales, Panagiota Karanasou, Xunying Liu, Yanman Qian, Philip C. Woodland, Chao Zhang 0031 |
INTERSPEECH | 8 |
| 2015 | Structured discriminative models using deep neural-network featuresabstractState-of-the-art speech recognisers employ neural networks in various configurations. A standard (hybrid) speech recogniser computes the likelihood for one time frame and state, using only one out of thousands of possible neural-network outputs. However, the whole output vector carries information. In this paper, features from state-of-the-art speech recognisers are collected per phone given a particular context, and input to a discriminative log-linear model. The log-linear model is trained with conditional maximum likelihood or a large-margin criterion. A key element is the prior on the parameters of the log-linear model. The mean of the prior is set to the point where the performance of the original systems is attained. The log-linear model then provides an additional increase over the state-of-the-art performance of the individual systems. Rogier C. van Dalen, Anton Ragni, Chao Zhang 0031, Mark J. F. Gales |
ASRU | 5 |
| 2015 | Speaker diarisation and longitudinal linking in multi-genre broadcast dataabstractThis paper presents a multi-stage speaker diarisation system with longitudinal Linking developed on BBC multi-genre data for the 2015 Multi-Genre Broadcast (MGB) challenge. The basic speaker diarisation system draws on techniques from the Cambridge March 2005 system with a new deep neural network (DNN)-based speech/non speech segmenter. A newly developed linking stage is next added to the basic diarisation output aiming at the identification of speakers across multiple episodes of the same series. The longitudinal constraint imposes an incremental processing of the episodes, where speaker labels for each episode can be obtained using only material from the episode in question, and those broadcast earlier in time. The nature of the data as well as the longitudinal linking constraint position this diarisation task as a new open-research topic, and a particularly challenging one. Different linking clustering metrics are compared and the lowest within-episode and cross-episode DER scores are achieved on the MGB challenge evaluation set. Panagiota Karanasou, Mark J. F. Gales, Pierre Lanchantin, Xunying Liu, Yanmin Qian, Philip C. Woodland, Chao Zhang 0031 |
ASRU | 8 |
| 2015 | The development of the cambridge university alignment systems for the multi-genre broadcast challengeabstractWe describe the alignment systems developed both for the preparation of data for the Multi-Genre Broadcast (MGB) challenge and for our participation in the transcription and alignment tasks. Captions of varying quality are aligned with the audio of TV shows that range from few minutes long to more than six hours. Lightly supervised decoding is performed on the audio and the output text is aligned with the original text transcript. Reliable split points are found and the resulting text chunks are force-aligned with the corresponding audio segments. Confidence scores are associated with the aligned data. Multiple refinements — including audio segmentation based on deep neural networks (DNNs) and the use of DNN-based acoustic models — were used to improve the performance. The final MGB alignment system had the highest F-measure value on the evaluation data. Pierre Lanchantin, Mark J. F. Gales, Panagiota Karanasou, Xunying Liu, Yanmin Qian, Philip C. Woodland, Chao Zhang 0031 |
ASRU | 8 |
| 2015 | Cambridge university transcription systems for the multi-genre broadcast challengeabstractWe describe the development of our speech-to-text transcription systems for the 2015 Multi-Genre Broadcast (MGB) challenge. Key features of the systems are: a segmentation system based on deep neural networks (DNNs); the use of HTK 3.5 for building DNN-based hybrid and tandem acoustic models and the use of these models in a joint decoding framework; techniques for adaptation of DNN based acoustic models including parameterised activation function adaptation; alternative acoustic models built using Kaldi; and recurrent neural network language models (RNNLMs) and RNNLM adaptation. The same language models were used with both HTK and Kaldi acoustic models and various combined systems built. The final systems had the lowest error rates on the evaluation data. Philip C. Woodland, Xunying Liu, Yanmin Qian, Chao Zhang 0031, Mark J. F. Gales, Panagiota Karanasou, Pierre Lanchantin |
ASRU | 4 |
| 2015 | The Cambridge University 2014 BOLT conversational telephone Mandarin Chinese LVCSR system for speech translationabstractThis paper presents the development of the 2014 Cambridge University conversational telephone Mandarin Chinese LVCSR system for the DARPA BOLT speech translation evaluation. A range of advanced modelling techniques were employed to both improve the recognition performance and provide a suitable integration with the translation system. These include an improved system combination technique using frame level acoustic model combination via joint decoding. Sequence trained deep neural network (DNN) based hybrid and tandem systems were combined on-the-fly to produce a consistent decoding output during search. A multi-level paraphrastic recurrent neural network LM (RNNLM) modelling both alternative paraphrase expressions and character sequences while preserving a consistent character to word segmentation was also used. This system gave an overall character error rate (CER) of 29.1% on the BOLT dev14 development set. Xunying Liu, Federico Flego, Chao Zhang 0031, Mark J. F. Gales, Philip C. Woodland |
INTERSPEECH | 4 |
| 2015 | Joint decoding of tandem and hybrid systems for improved keyword spotting on low resource languagesabstractCopyright © 2015 ISCA. Keyword spotting (KWS) for low-resource languages has drawn increasing attention in recent years. The state-of-the-art KWS systems are based on lattices or Confusion Networks (CN) generated by Automatic Speech Recognition (ASR) systems. It has been shown that considerable KWS gains can be obtained by combining the keyword detection results from different forms of ASR systems, e.g., Tandem and Hybrid systems. This paper investigates an alternative combination scheme for KWS using joint decoding. This scheme treats a Tandem system and a Hybrid system as two separate streams, and makes a linear combination of individual acoustic model log-likelihoods. Joint decoding is more efficient as it requires just a single pass of decoding and a single pass of keyword search. Experiments on six Babel OP2 development languages show that joint decoding is capable of providing consistent gains over each individual system. Moreover, it is possible to efficiently rescore the joint decoding lattices with Tandem or Hybrid acoustic models, and further KWS gains can be obtained by merging the detection posting lists from the joint decoding lattices and rescored lattices. Anton Ragni, Mark J. F. Gales, Kate M. Knill, Philip C. Woodland, Chao Zhang 0031 |
INTERSPEECH | 6 |
| 2015 | Parameterised sigmoid and reLU hidden activation functions for DNN acoustic modellingabstractThe form of hidden activation functions has been always an important issue in deep neural network (DNN) design. The most common choices for acoustic modelling are the standard Sigmoid and rectified linear unit (ReLU), which are normally used with fixed function shapes and no adaptive parameters. Recently, there have been several papers that have studied the use of parameterised activation functions for both computer vision and speaker adaptation tasks. In this paper, we investigate generalised forms of both Sigmoid and ReLU with learnable parameters, as well as their integration with the standard DNN acoustic model training process. Experiments using conversational telephone speech (CTS) Mandarin data, result in an average of 3.4% and 2.0% relative word error rate (WER) reduction with Sigmoid and ReLU parameterisations. Chao Zhang 0031, Philip C. Woodland |
INTERSPEECH | 1 |
| 2015 | A general artificial neural network extension for HTKabstractThis paper describes the recently developed artificial neural network (ANN) modules in HTK hidden Markov model toolkit, which enables ANN models with very general feed-forward architectures to be used for either acoustic modelling or feature extraction. The HTK ANN extension includes many recent ANN-based speech processing techniques, such as sequence training, model stacking, speaker adaptation, and parameterised activation functions. The implementation allows efficient training by supporting GPUs and various types of data cache. The ANN modules are fully integrated into the rest of the HTK toolkit, which allows existing GMM-HMM methods to be easily used in the ANN-HMM framework. Speech recognition results on a 300 hours DARPA BOLT conversational Mandarin task show that HTK can produce tandem and hybrid systems with state-of-the-art performance on this very challenging task. Furthermore, the flexibility of the implementation is illustrated using demo systems for a Wall Street Journal (WSJ) task. The HTK ANN extension is planned for release in HTK version 3.5. Chao Zhang 0031, Philip C. Woodland |
INTERSPEECH | 1 |
| 2014 | Standalone training of context-dependent deep neural network acoustic modelsabstractRecently, context-dependent (CD) deep neural network (DNN) hidden Markov models (HMMs) have been widely used as acoustic models for speech recognition. However, the standard method to build such models requires target training labels from a system using HMMs with Gaussian mixture model output distributions (GMM-HMMs). In this paper, we introduce a method for training state-of-the-art CD-DNN-HMMs without relying on such a pre-existing system. We achieve this in two steps: build a context-independent (CI) DNN iteratively with word transcriptions, and then cluster the equivalent output distributions of the untied CD-DNN HMM states using the decision tree based state tying approach. Experiments have been performed on the Wall Street Journal corpus and the resulting system gave comparable word error rates (WER) to CD-DNNs built based on GMM-HMM alignments and state-clustering. Chao Zhang 0031, Philip C. Woodland |
ICASSP | 1 |
| 2013 | Investigation of multilingual deep neural networks for spoken term detectionabstractThe development of high-performance speech processing systems for low-resource languages is a challenging area. One approach to address the lack of resources is to make use of data from multiple languages. A popular direction in recent years is to use bottleneck features, or hybrid systems, trained on multilingual data for speech-to-text (STT) systems. This paper presents an investigation into the application of these multilingual approaches to spoken term detection. Experiments were run using the IARPA Babel limited language pack corpora (~10 hours/language) with 4 languages for initial multilingual system development and an additional held-out target language. STT gains achieved through using multilingual bottleneck features in a Tandem configuration are shown to also apply to keyword search (KWS). Further improvements in both STT and KWS were observed by incorporating language questions into the Tandem GMM-HMM decision trees for the training set languages. Adapted hybrid systems performed slightly worse on average than the adapted Tandem systems. A language independent acoustic model test on the target language showed that retraining or adapting of the acoustic models to the target language is currently minimally needed to achieve reasonable performance. Kate M. Knill, Mark J. F. Gales, Shakti P. Rath, Philip C. Woodland, Chao Zhang 0031, Shixiong Zhang 0001 |
ASRU | 5 |
| 2013 | Reliable Accent-Specific Unit Generation With Discriminative Dynamic Gaussian Mixture Selection for Multi-Accent Chinese Speech RecognitionabstractIn this paper, we propose a discriminative dynamic Gaussian mixture selection (DGMS) strategy to generate reliable accent-specific units (ASUs) for multi-accent speech recognition. Time-aligned phone recognition is used to generate the ASUs that model accent variations explicitly and accurately. DGMS reconstructs and adjusts a pre-trained set of hidden Markov model (HMM) state densities to build dynamic observation densities for each input speech frame. A discriminative minimum classification error criterion is adopted to optimize the sizes of the HMM state observation densities with a genetic algorithm (GA). To the author's knowledge, the discriminative optimization for DGMS accomplishes discriminative training of discrete variables that is first proposed. We found the proposed framework is able to cover more multi-accent changes, thus reduce some performance loss in pruned beam search, without increasing the model size of the original acoustic model set. Evaluation on three typical Chinese accents, Chuan, Yue and Wu, shows that our approach outperforms traditional acoustic model reconstruction techniques with a syllable error rate reduction of 8.0%, 5.5% and 5.0%, respectively, while maintaining a good performance on standard Putonghua speech. Chao Zhang 0031, Yi Liu 0050, Yunqing Xia |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | Discriminative dynamic Gaussian mixture selection with enhanced robustness and performance for multi-accent speech recognitionabstractWe propose a discriminative DGMS (dynamic Gaussian mixture selection) strategy to enhance restructuring of a pre-trained set of Gaussian mixture models to cover unexpected acoustic variations at run time in automatic speech recognition. The number of Gaussian components in each hidden Markov model (HMM) state set aside is determined by a minimum classification error criterion. We also use a genetic algorithm to solve the integer programing problem to find the globally optimal state size. This parameter is used to adjust the HMM state densities for each input speech frame, leading to both high robustness and good resolution for dynamic tracking to cover a diversity of temporal variations in speech. Tested on an accented speech recognition application, the proposed framework yields an improved syllable error rate reduction over the conventional DGMS and augmented HMM systems when evaluated on three typical Chinese accents, Chuan, Yue and Wu, while maintaining its performance for standard Putonghua. Chao Zhang 0031, Yi Liu 0050, Yunqing Xia |
ICASSP | 1 |
| 2011 | Detection-based accented speech recognition using articulatory featuresabstractWe propose an attribute-based approach to accented speech recognition based on automatic speech attribute transcription with high efficiency detection of articulatory features. In order to utilize appropriate and extensible phonetic and linguistic knowledge, conditional random field (CRF) is designed to take frame-level inputs with binary feature functions. The use of CRF with merely the state features to generate probabilistic phone lattices is then utilized to solve the phone under-generation problem. Finally an attribute discrimination module is incorporated to handle a diversity of accent changes without retraining any model, leading to flexible “plug `n' play” modular design. The effectiveness of the proposed approach is evaluated on three typical Chinese accents, namely Guanhua, Yue and Wu. Our method yields a significant absolute phone recognition accuracy improvement 5.04%, 4.68% and 6.06% for the corresponding three accent types over a conventional monophone HMM system. Compared to a context-dependent triphone HMM system, we achieve comparable phone accuracies at only less than 20% of the computation cost. In addition, our proposed method is equally applicable to speaker-independent systems handling multiple accents. Chao Zhang 0031, Yi Liu 0050 |
ASRU | 1 |
| 2011 | Reliable accent specific unit generation with dynamic Gaussian mixture selection for multi-accent speech recognitionabstractMultiple accents are often present in Mandarin speech, as most Chinese have learned Mandarin as a second language. We propose generating reliable accent specific unit together with dynamic Gaussian mixture selection for multi-accent speech recognition. Time alignment phoneme recognition is used to generate such unit and to model accent variations explicitly and accurately. Dynamic Gaussian mixture selection scheme builds a dynamical observation density for each specified frame in decoding, and leads to use Gaussian mixture component efficiently. This method increases the covering ability for a diversity of accent variations in multi-accent, and alleviates the performance degradation caused by pruned beam search without augmenting the model size. The effectiveness of this approach is evaluated on three typical Chinese accents Chuan, Yue and Wu. Our approach outperforms traditional acoustic model reconstruction approach significantly by 6.30%, 4.93% and 5.53%, respectively on Syllable Error Rate (SER) reduction, without degrading on standard speech. Chao Zhang 0031, Yi Liu 0050, Yunqing Xia, Thomas Fang Zheng, Jesper Ø. Olsen, Jilei Tian |
ICME | 1 |