Li-Rong Dai 0001

dblp:48/6462-1 · also Lirong Dai 0001 · DBLP profile ↗
← Back
241ranked-venue papers
2as first author
73since 2021 · last 2026
0000-0002-0859-2827ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 197 · 58 since 2021Artificial intelligence and machine learning · 120 · 36 since 2021Databases, data management, data science and information retrieval · 4Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Generative Diffusion Contrastive Network for Multi-View Clustering
abstract
In recent years, Multi-View Clustering (MVC) has been significantly advanced under the influence of deep learning. By integrating heterogeneous data from multiple views, MVC enhances clustering analysis, making multi-view fusion critical to clustering performance. However, multi-view fusion remains challenged by low-quality data, primarily stemming from two reasons: 1) Certain views are contaminated by noisy data. 2) Some views suffer from missing data. This paper proposes a novel Stochastic Generative Diffusion Fusion (SGDF) method to address this problem. SGDF leverages a multiple generative mechanism for the multi-view feature of each sample. It exhibits robustness against low-quality data. Building on SGDF, we further present the Generative Diffusion Contrastive Network (GDCN). Extensive experiments show that GDCN achieves the state-of-the-art results in deep MVC tasks. The source code is publicly available athttps://github.com/HackerHyper/GDCN.
Xin Zou 0001, Lei Liu 0029, Chang Tang, Li-Rong Dai 0001
IEEE Signal Process. Lett.6
2025 CSSinger: End-to-End Chunkwise Streaming Singing Voice Synthesis System Based on Conditional Variational Autoencoder
abstract
Singing Voice Synthesis (SVS) aims to generate singing voices of high fidelity and expressiveness. Conventional SVS systems usually utilize an acoustic model to transform a music score into acoustic features, followed by a vocoder to reconstruct the singing voice. It was recently shown that end-to-end modeling is effective in the fields of SVS and Text to Speech (TTS). In this work, we thus present a fully end-to-end SVS method together with a chunkwise streaming inference to address the latency issue for practical usages. Note that this is the first attempt to fully implement end-to-end streaming audio synthesis using latent representations in VAE. We have made specific improvements to enhance the performance of streaming SVS using latent representations. Experimental results demonstrate that the proposed method achieves synthesized audio with high expressiveness and pitch accuracy in both streaming SVS and TTS tasks.
Jianwei Cui 0003, Shihao Chen, Jie Zhang 0042, Li-Rong Dai 0001
AAAI6
2025 Sinba: Singing-To-Accompaniment Generation With Pitch Guidance Via Mamba-Based Language Model
abstract
In this paper, we propose Sinba, a system that can directly generate corresponding background accompaniment music from vocal input, allowing users to create complete songs using only sung vocals. Sinba adopts a decoder-only backbone network architecture. We utilize the Mamba model, which is a linear-time sequence modeling method with selective state spaces and has been proven to achieve more advanced performance than Transformers as a foundation model in long-sequence modeling tasks. However, the Mamba was initially applied to audio tasks by pre-training directly on raw audio waveform samples as the backbone model. In this paper, we convert both the training targets and inputs into discretized tokens for direct training. We also extract pitch information from the vocal input as an additional feature for the model. The proposed model is trained using source-separated data pairs. Subjective and objective experimental results demonstrate that the proposed model can generate high-quality accompaniment that matches the style and rhythm of the vocal input, outperforming the Transformerbased baseline. Synthesized audio samples are available at: https://sounddemos.github.io/sinba.
Jianwei Cui 0003, Shihao Chen, Jie Zhang 0042, Chengxing Li, Shan Yang 0001, Li-Rong Dai 0001
ASRU9
2025 Leveraging Boolean Directivity Embedding for Binaural Target Speaker Extraction
abstract
Direction-based target speaker extraction (TSE) attracts a constant attention due to the convenience of direction acquisition over assistive video or enrollment audio. The direction clue heavily affects the TSE performance, which might be more seriously in the case of binaural setups due to the small-sized and irregular microphone array. In this paper, we propose Boolean Directivity Embedding (BDE) as a new direction feature in order to precisely lock onto the target speaker independent on microphone array configurations for binaural TSE (BiTSE). We design an encoder that accurately aligns the BDE with the mixed audio signals for feature fusion. Considering that the Boolean representation may contain insufficient spatial and temporal information, we enhance the BDE by incorporating the previously-proposed spatiotemporal features, showing a compatibility and stronger capacity for BiTSE. The proposed BiTSE model, which is based on the narrow-band Conformer as the backbone, can adapt to the cases of target speaker switching and moving by simply modifying the frame-wise BDE. Experimental results demonstrate the efficacy of our method in both stationary and dynamic scenarios. The proposed method is open-sourced in https://github.com/ichi131/Direction-based-BiTSE.
Yichi Wang 0001, Jie Zhang 0042, Chengqian Jiang, Weitai Zhang, Zhongyi Ye, Li-Rong Dai 0001
ICASSP6
2025 Bridging Modality Gap with Large Speech and Language Models for End-to-End Speech-to-Text Translation
abstract
End-to-end speech-to-text translation (E2E ST) has increasingly aroused interest and attention recently, attempting to address the problem of data scarcity and modeling burden. Several attempts exploring the combination of Large Speech and Language Models into a unified model to improve E2E ST are carried out. However, the inherent differences between speech and text modalities often impede effective cross-modal and cross-lingual transfer. In this study, we introduce LaSaLM-ST, a novel model architecture built upon Pre-trained Large Speech and Language Models for improving E2E ST. Our speech encoder begins with processing the source speech sequence. An adaptor and speech decoder then project speech features into the compatible feature space for the decoder-only Large Language Model (LLM), which then aligns the representation spaces of speech and text modalities with attentive interactions. Besides, we also develop a multi-step fine-tuning method to preserve the pre-trained multilingual knowledge and keep ST fine-tuning stably. Experiments conducted on the IWSLT2023 offline ST task from English to German, Chinese and Japanese demonstrate that our methodology not only achieves state-of-the-art BLEU scores but also outperforms the highly competitive cascaded ST systems in an unrestricted setting.
Weitai Zhang, Simran Naagar, Zhongyi Ye, Peiwang Tang, Xinyuan Zhou, Li-Rong Dai 0001
ICASSP7
2025 Trusted Mamba Contrastive Network for Multi-View Clustering
abstract
Multi-view clustering can partition data samples into their categories by learning a consensus representation in an unsupervised way and has received more and more attention in recent years. However, there is an untrusted fusion problem. The reasons for this problem are as follows: 1) The current methods ignore the presence of noise or redundant information in the view; 2) The similarity of contrastive learning comes from the same sample rather than the same cluster in deep multi-view clustering. It causes multi-view fusion in the wrong direction. This paper proposes a novel multi-view clustering network to address this problem, termed as Trusted Mamba Contrastive Network (TMCN). Specifically, we present a new Trusted Mamba Fusion Network (TMFN), which achieves a trusted fusion of multi-view data through a selective mechanism. Moreover, we align the fused representation and the view-specific representation using the Average-similarity Contrastive Learning (AsCL) module. AsCL increases the similarity of view presentation from the same cluster, not merely from the same sample. Extensive experiments show that the proposed method achieves state-of-the-art results in deep multi-view clustering tasks. The source code is available at https://github.com/HackerHyper/TMCN.
Xin Zou 0001, Lei Liu 0029, Zhangmin Huang, Chang Tang, Li-Rong Dai 0001
ICASSP7
2025 Dynamic SRM Curriculum for Trustworthy Multi-modal Classification
abstract
Trustworthy multi-modal learning integrates multiple sources of data reliably. However, the current methods still focus on performance improvement by developing deep multi-modal networks. These approaches frequently encounter challenges due to the inherent non-convex nature of deep neural networks and their vulnerability to local minima, ultimately leading to a diminished ability for generalization. To address this problem, we present a novel curriculum termed the Dynamic SRM Curriculum (DSRMC). Within DSRMC, the deep trustworthy multi-modal networks undergo training with data provided sequentially, progressing from simple to complex samples. This training strategy mimics the human learning process, commencing with fundamental concepts and gradually advancing to tackle more complex and abstract ideas. Building upon DSRMC, we propose an innovative Curriculum Trustworthy Multi-modal Learning (CTML) method. CTML makes it easier to place the learned model in a flatter area, which improves its overall ability for generalization. Comprehensive experiments on three public datasets demonstrate that the proposed CTML performs better than state-of-the-art methods, achieving a maximum improvement of 6.7% on macroF1.
Cui Yu, Xin Zou 0001, Zhangmin Huang, Chenshu Hu, Jun Sun 0014, Bo Lyu, Lei Liu 0029, Chang Tang, Li-Rong Dai 0001
ICASSP10
2025 An Effective Anomalous Sound Detection Method Based on Global and Local Attribute Mining
Nan Jiang 0022, Yan Song 0001, Qing Gu 0002, Li-Rong Dai 0001, Ian McLoughlin 0001
INTERSPEECH5
2025 Finetune Large Pre-Trained Model Based on Frequency-Wise Multi-Query Attention Pooling for Anomalous Sound Detection
Nan Jiang 0022, Yan Song 0001, Qing Gu 0002, Li-Rong Dai 0001, Ian McLoughlin 0001
INTERSPEECH5
2025 MISP-QEKS: A Large-Scale Dataset with Multimodal Cues for Query-by-Example Keyword Spotting
Shifu Xiong, Hang Chen 0001, Shi Cheng 0001, Hengshun Zhou, Genshun Wan, Chenyue Zhang, Jun Du 0002, Li-Rong Dai 0001
ACM Multimedia10
2025 A graph neural network explainability strategy driven by key subgraph connectivity
Li-Rong Dai 0001, D. H. Xu, Y. F. Gao
J. Biomed. Informatics1
2025 Adaptive Confidence Multi-View Learning
abstract
Multi-view hashing is a crucial technology for multimedia retrieval because it transforms heterogeneous data from many viewpoints into binary hash codes. However, the existing approaches focus mostly on the complementarity among multiple views while being without confidence fusion. Furthermore, redundant noise is present in the single-view data in real-world application contexts. We present an innovativeAdaptive Confidence Multi-View Learning(ACMVL) method to perform confidence fusion and remove extraneous noise. Initially, a confidence network is constructed to eliminate noise data and extract useful information from various single-view features. Moreover, an adaptive confidence multi-view network is utilized to quantify the confidence of each view and further fuse multiple view features using a weighted summation. Here, we propose anAutomatic View Confidence Metric(AVCM) as a score for evaluating the confidence of views. Finally, to improve the semantic representation of the fused feature, a dilation network is created. Based on ACMVL, we introduce a novelAdaptive Confidence Multi-View Hashing(ACMVH) method. To our knowledge, we are the pioneers in using confidence learning for multimedia retrieval. Comprehensive experiments on three publicly available datasets demonstrate that our ACMVH outperforms the state-of-the-art methods (maximum improvement of 3.24% on mAP).
Lei Liu 0029, Chang Tang, Li-Rong Dai 0001
IEEE Trans. Multim.5
2024 Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech Representation
abstract
Self-supervised speech pre-training methods have developed rapidly in recent years, which show to be very effective for many near-field single-channel speech tasks. However, far-field multichannel speech processing is suffering from the scarcity of labeled multichannel data and complex ambient noises. The efficacy of self-supervised learning for far-field multichannel and multi-modal speech processing has not been well explored. Considering that visual information helps to improve speech recognition performance in noisy scenes, in this work we propose the multichannel multi-modal speech self-supervised learning framework AV-wav2vec2, which utilizes video and multichannel audio data as inputs. First, we propose a multi-path structure to process multi-channel audio streams and a visual stream in parallel, with intra-, and inter-channel contrastive as training targets to fully exploit the rich information in multi-channel speech data. Second, based on contrastive learning, we use additional single-channel audio data, which is trained jointly to improve the performance of multichannel multi-modal representation. Finally, we use a Chinese multichannel multi-modal dataset in real scenarios to validate the effectiveness of the proposed method on audio-visual speech recognition (AVSR), automatic speech recognition (ASR), visual speech recognition (VSR) and audio-visual speaker diarization (AVSD) tasks.
Qiushi Zhu, Jie Zhang 0042, Li-Rong Dai 0001
AAAI5
2024 Adversarial Speech for Voice Privacy Protection from Personalized Speech Generation
abstract
The rapid progress in personalized speech generation technology, including personalized text-to-speech (TTS) and voice conversion (VC), poses a challenge in distinguishing between generated and real speech for human listeners, resulting in an urgent demand in protecting speakers' voices from malicious misuse. In this regard, we propose a speaker protection method based on adversarial attacks. The proposed method perturbs speech signals by minimally altering the original speech while rendering downstream speech generation models unable to accurately generate the voice of the target speaker. For validation, we employ the open-source pre-trained YourTTS model for speech generation and protect the target speaker's speech in the white-box scenario. Automatic speaker verification (ASV) evaluations were carried out on the generated speech as the assessment of the voice protection capability. Our experimental results show that we successfully perturbed the speaker encoder of the YourTTS model using the gradient-based I-FGSM adversarial perturbation method. Furthermore, the adversarial perturbation is effective in preventing the YourTTS model from generating the speech of the target speaker. Audio samples can be found in https://voiceprivacy.github.io/Adeversarial-Speech-with-YourTTS.
Shihao Chen, Jie Zhang 0042, Kong-Aik Lee, Zhen-Hua Ling, Li-Rong Dai 0001
ICASSP6
2024 Sifisinger: A High-Fidelity End-to-End Singing Voice Synthesizer Based on Source-Filter Model
abstract
This paper presents an advanced end-to-end singing voice synthesis (SVS) system based on the source-filter mechanism that directly translates lyrical and melodic cues into expressive and high-fidelity human-like singing. Similarly to VISinger 2, the proposed system also utilizes training paradigms evolved from VITS and incorporates elements like the fundamental pitch (F0) predictor and waveform generation decoder. To address the issue that the coupling of mel-spectrogram features with F0 information may introduce errors during F0 prediction, we consider two strategies. Firstly, we leverage mel-cepstrum (mcep) features to decouple the intertwined mel-spectrogram and F0 characteristics. Secondly, inspired by the neural source-filter models, we introduce source excitation signals as the representation of F0 in the SVS system, aiming to capture pitch nuances more accurately. Meanwhile, differentiable mcep and F0 losses are employed as the waveform decoder supervision to fortify the prediction accuracy of speech envelope and pitch in the generated speech. Experiments on the Opencpop dataset demonstrate efficacy of the proposed model in synthesis quality and intonation accuracy. Synthesized audio samples are available at: https://sounddemos.github.io/sifisinger.
Jianwei Cui 0003, Chao Weng, Jie Zhang 0042, Li-Rong Dai 0001
ICASSP6
2024 A Study of Multichannel Spatiotemporal Features and Knowledge Distillation on Robust Target Speaker Extraction
abstract
Target speaker extraction (TSE) based on direction of arrival (DOA) has a wide range of applications in e.g., remote conferencing, hearing aids, in-car speech interaction. Due to the inherent phase uncertainty, existing TSE methods usually suffer from speaker confusion within specific frequency bands. Imprecise DOA measurements caused by e.g., the calibration of the microphone array and ambient noises, can also deteriorate the TSE performance. In order to improve the robustness of TSE, in this work we propose several new multichannel spatiotemporal features to represent the discriminability of the target speaker. The narrow-band Conformer model is applied in combination with the proposed features to facilitate the extraction of the target speaker. In addition, we consider knowledge distillation for improving the model robustness, particularly in the presence of DOA mis-match. Experimental results on a public dataset verify the efficacy of the proposed method.
Yichi Wang 0001, Jie Zhang 0042, Shihao Chen, Weitai Zhang, Zhongyi Ye, Xinyuan Zhou, Li-Rong Dai 0001
ICASSP7
2024 Pre-Trained Acoustic-and-Textual Modeling for End-To-End Speech-To-Text Translation
abstract
End-to-end paradigm has aroused more and more interests and attention for improving speech-to-text translation (ST) recently. Existing end-to-end models mainly attributes and attempts to address the problem of modeling burden and data scarcity, while always fail to maintain both cross-modal and cross-lingual mapping well at the same time. In this work, we investigate methods for improving endto-end ST with pre-trained acoustic-and-textual models. Our acoustic encoder and decoder begins with processing the source speech sequence as usual. A textual encoder and an adaptor module then obtain source acoustic and textual information respectively, alleviating the representation inconsistency with attentive interactions in the textual decoder. Also, we utilize pre-trained models, and develop an adaptation fine-tuning method to preserve the pre-training knowledge. Experimental results on the IWSLT2023 offline ST task from English to German, Japanese and Chinese show that our method achieves state-of-the-art BLEU scores and surpasses the strong cascaded ST counterparts in unrestricted setting.
Weitai Zhang, Hanyi Zhang, Chenxuan Liu, Zhongyi Ye, Xinyuan Zhou, Li-Rong Dai 0001
ICASSP7
2024 LDM-SVC: Latent Diffusion Model Based Zero-Shot Any-to-Any Singing Voice Conversion with Singer Guidance
Shihao Chen, Jie Zhang 0042, Rilin Chen, Li-Rong Dai 0001
INTERSPEECH7
2024 An Effective Local Prototypical Mapping Network for Speech Emotion Recognition
Yuxuan Xi, Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001
INTERSPEECH3
2024 Sketch-fusion: A gradient compression method with multi-layer fusion for communication-efficient distributed training
Li-Rong Dai 0001, Luqi Gong, Zhulin An, Yongjun Xu 0001, Boyu Diao
J. Parallel Distributed Comput.1
2024 Boosted Curriculum Multi-View Hashing for Multimedia Retrieval
abstract
The multi-view hash method plays a pivotal role in multimedia retrieval, transforming diverse data from multiple perspectives into binary hash codes. While existing methods primarily emphasize complementarity across multiple views, they often face challenges associated with the non-convex nature of deep neural networks, ultimately causing a decrease in generalization ability. To overcome this limitation, we propose a novel curriculum calledAutomatic Multiple Loss Curriculum(AMLC). In AMLC, the deep multi-view hashing network undergoes training with data presented sequentially, progressing from simple to complex samples. This training strategy mirrors the human learning process, commencing with fundamental concepts and progressively advancing to tackle more intricate and abstract ideas. Building upon AMLC, we propose theBoosted Curriculum Multi-View Hashing(BCMVH) method. BCMVH facilitates the positioning of the learned model in a more flat region, enhancing its overall generalization capability. Extensive experiments conducted on three public datasets demonstrate that the proposed BCMVH outperforms state-of-the-art methods, achieving a maximum improvement of 3.17% in terms of mean Average Precision.
Zhangmin Huang, Lei Liu 0029, Chang Tang, Li-Rong Dai 0001
IEEE Signal Process. Lett.5
2024 SpeechLM: Enhanced Speech Pre-Training With Unpaired Textual Data
abstract
How to boost speech pre-training with textual data is an unsolved problem due to the fact that speech and text are very different modalities with distinct characteristics. In this paper, we propose a cross-modalSpeechandLanguageModel (SpeechLM) to explicitly align speech and text pre-training with a pre-defined unified discrete representation. Specifically, we introduce two alternative discrete tokenizers to bridge the speech and text modalities, including phoneme-unit and hidden-unit tokenizers, which can be trained using unpaired speech or a small amount of paired speech-text data. Based on the trained tokenizers, we convert the unlabeled speech and text data into tokens of phoneme units or hidden units. The pre-training objective is designed to unify the speech and the text into the same discrete semantic space with a unified Transformer network. We evaluate SpeechLM on various spoken language processing tasks including speech recognition, speech translation, and universal representation evaluation framework SUPERB, demonstrating significant improvements on content-related tasks. Code and models are available athttps://aka.ms/SpeechLM.
Sanyuan Chen, Yu Wu 0012, Shuo Ren 0002, Shujie Liu 0001, Zhuoyuan Yao, Xun Gong 0005, Li-Rong Dai 0001, Jinyu Li 0001, Furu Wei
IEEE ACM Trans. Audio Speech Lang. Process.9
2024 VatLM: Visual-Audio-Text Pre-Training With Unified Masked Prediction for Speech Representation Learning
abstract
Although speech is a simple and effective way for humans to communicate with the outside world, a more realistic speech interaction contains multimodal information, e.g., vision, text. How to design a unified framework to integrate different modal information and leverage different resources (e.g., visual-audio pairs, audio-text pairs, unlabeled speech, and unlabeled text) to facilitate speech representation learning was not well explored. In this paper, we propose a unified cross-modal representation learning frameworkVatLM(Visual-Audio-Text Language Model). The proposedVatLMemploys a unified backbone network to model the modality-independent information and utilizes three simple modality-dependent modules to preprocess visual, speech, and text inputs. In order to integrate these three modalities into one shared semantic space,VatLMis optimized with a masked prediction task of unified tokens, given by our proposed unified tokenizer. We evaluate the pre-trainedVatLMon audio-visual related downstream tasks, including audio-visual speech recognition (AVSR), and visual speech recognition (VSR) tasks. Results show that the proposedVatLMoutperforms previous state-of-the-art models, such as the audio-visual pre-trained AV-HuBERT model, and analysis also demonstrates thatVatLMis capable of aligning different modalities into the same space. To facilitate future research, we release the code and pre-trained models athttps://aka.ms/vatlm.
Qiushi Zhu, Shujie Liu 0001, Binxing Jiao, Jie Zhang 0042, Li-Rong Dai 0001, Daxin Jiang, Jinyu Li 0001, Furu Wei
IEEE Trans. Multim.7
2024 MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services
abstract
While modern internet services, such as chatbots, search engines, and online advertising, demand the use of large-scale deep neural networks (DNNs), distributed training and inference over heterogeneous computing systems are desired to facilitate these DNN models. Mixture-of-Experts (MoE) is one the most common strategies to lower the cost of training subject to the overall size of models/data through gating and parallelism in a divide-and-conquer fashion. While DeepSpeed [1] has made efforts in carrying out large-scale MoE training over heterogeneous infrastructures, the efficiency of training and inference could be further improved from several system aspects, including load balancing, communication/computation efficiency, and memory footprint limits. In this work, we present a novel MoESys that boosts efficiency in both large-scale training and inference. Specifically, in the training procedure, the proposed MoESys adopts an Elastic MoE training strategy with 2D prefetch and Fusion communication over Hierarchical storage, so as to enjoy efficient parallelisms. For scalable inference in a single node, especially when the model size is larger than GPU memory, MoESys builds the CPU-GPU memory jointly into a ring of sections to load the model, and executes the computation tasks across the memory sections in a round-robin manner for efficient inference. We carried out extensive experiments to evaluate MoESys, where MoESys successfully trains a Unified Feature Optimization [2] (UFO) model with a Sparsely-Gated Mixture-of-Experts model of 12B parameters in 8 days on 48 A100 GPU cards. The comparison against the state-of-the-art shows that MoESys outperformed DeepSpeed with 33 unbalanced MoE Tasks, e.g., UFO, MoESys achieved 64
Dianhai Yu, Hongxiang Hao, Weibao Gong, HuaChao Wu, Jiang Bian 0003, Li-Rong Dai 0001, Haoyi Xiong
IEEE Trans. Serv. Comput.7
2023 Stargan-vc Based Cross-Domain Data Augmentation for Speaker Verification
abstract
Automatic speaker verification (ASV) faces domain shift caused by the mismatch of intrinsic and extrinsic factors, such as recording device and speaking style, in real-world applications, which leads to severe performance degradation. Since single-speaker multi-condition (SSMC) data is difficult to collect in practice, existing domain adaptation methods are hard to ensure the feature consistency of the same class but different domains. To this end, we propose a cross-domain data generation method to obtain a domain-invariant ASV system. Inspired by voice conversion (VC) task, a StarGAN based generative model first learns cross-domain mappings from SSMC data, and then generates missing domain data for all speakers, thus increasing the intra-class diversity of the training set. Considering the difference between ASV and VC task, we renovate the corresponding training objectives and network structure to make the adaptation task-specific. Evaluations on achieve a relative performance improvement of about 5-8% over the baseline in terms of minDCF and EER, outperforming the CNSRC winner’s system of the equivalent scale.
Hang-Rui Hu, Yan Song 0001, Jian-Tao Zhang, Li-Rong Dai 0001, Ian McLoughlin 0001, Zhu Zhuo, Yu-Hong Li
ICASSP4
2023 AST-SED: An Effective Sound Event Detection Method Based on Audio Spectrogram Transformer
abstract
In this paper, we propose an effective sound event detection (SED) method based on the audio spectrogram transformer (AST) model, pretrained on the large-scale AudioSet for audio tagging (AT) task, termed AST-SED. Pretrained AST models have recently shown promise on DCASE2022 challenge task4 where they help mitigate a lack of sufficient real annotated data. However, mainly due to differences between the AT and SED tasks, it is suboptimal to directly utilize outputs from a pretrained AST model. Hence the proposed AST-SED adopts an encoder-decoder architecture to enable effective and efficient fine-tuning without needing to redesign or retrain the AST model. Specifically, the Frequency-wise Transformer Encoder (FTE) consists of transformers with self attention along the frequency axis to address multiple overlapped audio events issue in a single clip. The Local Gated Recurrent Units Decoder (LGD) consists of nearest-neighbor interpolation (NNI) and Bidirectional Gated Recurrent Units (Bi-GRU) to compensate for temporal resolution loss in the pretrained AST model output. Experimental results on DCASE2022 task4 development set have demonstrated the superiority of the proposed AST-SED with FTE-LGD architecture. Specifically, the Event-Based F1-score (EB-F1) of 59.60% and Polyphonic Sound detection Score scenario1 (PSDS1) of 0.5140 significantly outperform CRNN and other pretrained AST-based systems.
Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001, Lin Liu 0017
ICASSP3
2023 A Multi-Scale Feature Aggregation Based Lightweight Network for Audio-Visual Speech Enhancement
abstract
Audio-visual speech enhancement (AVSE) was shown to be superior over conventional audio-only counterpart for improving the speech quality. However, most existing AVSE models are heavyweight in the sense of parameter count, which is inappropriate for the deployment and practical applications. In this paper, we therefore present a lightweight AVSE approach (called M3Net) by incorporating several multi-modality, multi-scale and multi-branch strategies. Three multi-scale techniques are designed for the visual and audio streams, including multi-scale average pooling (MSAP), multi-scale ResNet (MSResNet) and multi-scale short time Fourier transform (MSSTFT). It is shown that each multi-scale module positively contributes to the performance. Also, we consider four skip connections for the audio-visual feature aggregation, which have a great complementary effect on the designed multi-scale techniques. Experimental results show that these techniques are flexible in combination with existing approaches, and more importantly obtain a comparable performance with a smaller model size compared to the heavyweight networks.
Liangfa Wei, Jie Zhang 0042, Jianming Yang, Yannan Wang, Tian Gao 0005, Li-Rong Dai 0001
ICASSP8
2023 Joint Generative-Contrastive Representation Learning for Anomalous Sound Detection
abstract
In this paper, we propose a joint generative and contrastive representation learning method (GeCo) for anomalous sound detection (ASD). GeCo exploits a Predictive AutoEncoder (PAE) equipped with self-attention as a generative model to perform frame-level prediction. The output of the PAE together with original normal samples, are used for supervised contrastive representative learning in a multi-task framework. Besides cross-entropy loss between classes, contrastive loss is used to separate PAE output and original samples within each class. GeCo aims to better capture context information among frames, thanks to the self-attention mechanism for PAE model. Furthermore, GeCo combines generative and contrastive learning from which we aim to yield more effective and informative representations, compared to existing methods. Extensive experiments have been conducted on the DCASE2020 Task2 development dataset, showing that GeCo outperforms state-of-the-art generative and discriminative methods.
Xiao-Min Zeng, Yan Song 0001, Zhu Zhuo, Yu-Hong Li, Li-Rong Dai 0001, Ian McLoughlin 0001
ICASSP7
2023 Robust Data2VEC: Noise-Robust Speech Representation Learning for ASR by Combining Regression and Improved Contrastive Learning
abstract
Self-supervised pre-training methods based on contrastive learning or regression tasks can utilize more unlabeled data to improve the performance of automatic speech recognition (ASR). However, the robustness impact of combining the two pre-training tasks and constructing different negative samples for contrastive learning still remains unclear. In this paper, we propose a noise-robust data2vec for self-supervised speech representation learning by jointly optimizing the contrastive learning and regression tasks in the pre-training stage. Furthermore, we present two improved methods to facilitate contrastive learning. More specifically, we first propose to construct patch-based non-semantic negative samples to boost the noise robustness of the pre-training model, which is achieved by dividing the features into patches at different sizes (i.e., so-called negative samples). Second, by analyzing the distribution of positive and negative samples, we propose to remove the easily distinguishable negative samples to improve the discriminative capacity for pre-training models. Experimental results on the CHiME-4 dataset show that our method is able to improve the performance of the pre-trained model in noisy scenarios. We find that joint training of the contrastive learning and regression tasks can avoid the model collapse to some extent compared to only training the regression task.
Qiushi Zhu, Jie Zhang 0042, Shujie Liu 0001, Yu-Chen Hu, Li-Rong Dai 0001
ICASSP6
2023 Speech4Mesh: Speech-Assisted Monocular 3D Facial Reconstruction for Speech-Driven 3D Facial Animation
abstract
Recent audio2mesh-based methods have shown promising prospects for speech-driven 3D facial animation tasks. However, some intractable challenges are urgent to be settled. For example, the data-scarcity problem is intrinsically inevitable due to the difficulty of 4D data collection. Besides, current methods generally lack controllability on the animated face. To this end, we propose a novel framework named Speech4Mesh to consecutively generate 4D talking head data and train the audio2mesh network with the reconstructed meshes. In our framework, we first reconstruct the 4D talking head sequence based on the monocular videos. For precise capture of the talking-related variation on the face, we exploit the audio-visual alignment information from the video by employing a contrastive learning scheme. We next can train the audio2mesh network (e.g., FaceFormer) based on the generated 4D data. To get control of the animated talking face, we encode the speaking-unrelated factors (e.g., emotion, etc.) into an emotion embedding for manipulation. Finally, a differentiable renderer guarantees more accurate photometric details of the reconstruction and animation results. Empirical experiments demonstrate that the Speech4Mesh framework can not only outperform state-of-the-art reconstruction methods, especially on the lower-face part but also achieve better animation performance both perceptually and objectively after pre-trained on the synthesized data. Besides, we also verify that the proposed framework is able to explicitly control the emotion of the animated talking face.
Shuo Yang 0006, Pengcheng Xia 0002, Cong Liu 0006, Li-Rong Dai 0001, Chang Xu 0002
ICCV8
2023 Vision-Language Adaptive Mutual Decoder for OOV-STR
Jinshui Hu, Qiandong Yan, Xuyang Zhu, Jiajia Wu 0003, Jun Du 0002, Li-Rong Dai 0001
ICIG (2)7
2023 A Multimodal Text Block Segmentation Framework for Photo Translation
Jiajia Wu 0003, Anni Li, Zhengyan Yang, Cong Liu 0006, Li-Rong Dai 0001
ICIG (3)7
2023 End-to-End Multilingual Text Recognition Based on Byte Modeling
Jiajia Wu 0003, Zhengyan Yang, Cong Liu 0006, Li-Rong Dai 0001
ICIG (3)6
2023 Fine-tuning Audio Spectrogram Transformer with Task-aware Adapters for Sound Event Detection
abstract
In this paper, we present a task-aware fine-tuning method to transfer Patchout faSt Spectrogram Transformer (PaSST) model to sound event detection (SED) task. Pretrained PaSST has shown significant performance on audio tagging (AT) and SED tasks, but it is not optimal to fine-tune the model from a single layer as the local and semantic information have not been well exploited. To address this, we first introduce task-aware adapters including SED-adapter and AT-adapter to fine-tune PaSST for SED and AT task respectively, and then propose task-aware fine-tuning to combine local information from shallower layer with semantic information from deeper layer, based on task-aware adapters. Besides, we propose the self-distillated mean teacher (SdMT) to train a robust student model with soft pseudo labels from teacher. Experiments are conducted on DCASE2022 task4 development set, the EB-F1 of 64.85% and PSDS1 of 0.5548 are achieved which outperform previous state-of-the-art systems.
Yan Song 0001, Ian McLoughlin 0001, Lin Liu 0017, Li-Rong Dai 0001
INTERSPEECH6
2023 CASA-ASR: Context-Aware Speaker-Attributed ASR
Mohan Shi, Zhihao Du, Qian Chen 0003, Fan Yu 0002, Yangze Li, Shiliang Zhang, Jie Zhang 0042, Li-Rong Dai 0001
INTERSPEECH8
2023 Semantic VAD: Low-Latency Voice Activity Detection for Speech Interaction
Mohan Shi, Yuchun Shu, Lingyun Zuo, Qian Chen 0003, Shiliang Zhang, Jie Zhang 0042, Li-Rong Dai 0001
INTERSPEECH7
2023 Real-Time Causal Spectro-Temporal Voice Activity Detection Based on Convolutional Encoding and Residual Decoding
Jie Zhang 0042, Li-Rong Dai 0001
INTERSPEECH3
2023 Robust Prototype Learning for Anomalous Sound Detection
abstract
In this paper, we present a robust prototype learning framework for anomalous sound detection (ASD), where prototypical loss is exploited to measure the similarity between samples and prototypes. We show that existing generative and discriminative based ASD methods can be unified into this framework from the perspective of prototypical learning. For ASD in recent DCASE challenges, extensions related to imbalanced learning are proposed to improve the robustness of prototypes learned from source and target domains. Specifically, balanced sampling and multiple-prototype expansion (MPE) strategies are proposed to address imbalances across attributes of source and target domains. Furthermore, a novel negative-prototype expansion (NPE) method is used to construct pseudo-anomalies to learn a more compact and effective embedding space for normal sounds. Evaluation on the DCASE2022 Task2 development dataset demonstrates the validity of the proposed prototype learning framework.
Xiao-Min Zeng, Yan Song 0001, Ian McLoughlin 0001, Lin Liu 0017, Li-Rong Dai 0001
INTERSPEECH5
2023 Handwritten Chemical Structure Image to Structure-Specific Markup Using Random Conditional Guided Decoder
abstract
Satisfactory recognition performance has been achieved for simple and controllable printed molecular images. However, recognizing handwritten chemical structure images remains unresolved due to the inherent ambiguities in handwritten atoms and bonds, as well as the signifcant challenge of converting projected 2D molecular layouts into markup strings. Target to address these problems, this paper proposes an end-to-end framework for handwritten chemical structure images recognition, with novel structure-specific markup language (SSML) and random conditional guided decoder (RCGD). SSML alleviates ambiguity and complexity in Chemfig syntax by designing an innovative markup language to accurately depict molecular structures. Besides, we propose RCGD to address the issue of multiple path decoding of molecular structures, which is composed of conditional attention guidance, memory classification and path selection mechanisms. In order to fully confirm the effectiveness of the end-to-end method, a new database containing 50,000 handwritten chemical structure images (EDU-CHEMC) has been established. Experimental results demonstrate that compared to traditional SMILES sequences, our SSML can significantly reduces the semantic gap between chemical images and markup strings. It is worth noting that our method can also recognize invalid or non-existent organic molecular structures, making it highly applicable for tasks related to teaching evaluations in the fields of chemistry and biology education. The EDU-CHEMC will be released soon in https://github.com/iFLYTEK-CV/EDU-CHEMC.
Jinshui Hu, Hao Wu 0090, Mingjun Chen, Jiajia Wu 0003, Cong Liu 0006, Jun Du 0002, Li-Rong Dai 0001
ACM Multimedia11
2023 Energy-Efficient Sparsity-Driven Speech Enhancement in Wireless Acoustic Sensor Networks
abstract
Wireless acoustic sensor network (WASN) has shown a superiority over conventional microphone arrays in many aspects. There exists an important tradeoff between the performance and power consumption, as usually the sensors are power driven with a limited amount of battery resource. Given a prescribed performance bound, in literature sensor selection (SS) and rate allocation (RA) methods can be leveraged to optimize the energy efficiency. In this work, we propose a joint rate allocation and sensor selection (RASS) approach to simultaneously optimize the sensor subset and rate distribution, which is formulated by minimizing the total transmission power in terms of selection and bit-rate variables and constraining the residual noise power. It can be shown that under a set of linear constraints on beamforming, the linearly-constrained minimum variance (LCMV) beamformer is the optimal noise reduction filter. Based on this, the RASS reduces to a mixed semi-definite and bilinear programming problem, which is then solved using a two-step algorithm. As the selection and bit-rate unknowns are bilinear, we first consider to optimize their product, resulting in an upper bound of RASS. Then, we use McCormick envelopes to relax the bilinear constraint, resulting in a linear program. The final selection and bit-rate solutions are obtained by posterior randomized rounding. It can be shown that SS and RA are special cases of the proposed RASS. Numerical results using simulated WASNs validate the power efficiency of the proposed method as well as the robustness against dynamic factors.
Jie Zhang 0042, Jun Du 0002, Li-Rong Dai 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 SDW-SWF: Speech Distortion Weighted Single-Channel Wiener Filter for Noise Reduction
abstract
Speech enhancement shows an important necessity in many audio applications, particularly in noisy environments, where the speech quality needs to be improved. In this work, we consider the single-channel noise reduction (NR) problem from the conventional signal processing perspective. As conventional single-channel NR filters suffer from a serious speech distortion (SD) problem, we propose an SD weighted single-channel Wiener filter (SDW-SWF) in the short-time Fourier transform domain, which is obtained by minimizing the mean-square error (MSE) of the clean speech plus a$\mu$-weighted residual noise variance. Based on the generalized eigenvalue decomposition (GEVD) and rank-$r$approximation of the speech correlation matrix, the SDW-SWF can be written as a linear combination of eigenpairs, from which some special cases reduce to existing single-channel NR filters. As such, the proposed SDW-SWF has two parameters (i.e.,$\mu$and$r$) to tradeoff the MSE and SD. Then we theoretically analyze the impacts of the tradeoff parameters on the NR performance in SD, residual noise variance and the output signal-to-noise ratio (SNR). In addition, it is shown that the STFT-domain SDW-SWF can be further extended to the time domain, where the derived theorems still hold. Numerical results from several perspectives validate the effectiveness of the proposed method.
Jie Zhang 0042, Jun Du 0002, Li-Rong Dai 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 A Joint Speech Enhancement and Self-Supervised Representation Learning Framework for Noise-Robust Speech Recognition
abstract
Though speech enhancement (SE) can be used to improve speech quality in noisy environments, it may also cause distortions that degrade the performance of automatic speech recognition (ASR) models. Self-supervised pre-training, on the other hand, has been shown to improve the noise robustness of ASR models. However, the potential of the (optimal) integration of SE and self-supervised pre-training still remains unclear. In this paper, we propose a novel self-supervised pre-training framework that incorporates SE to improve ASR performance in noisy environments. First, in the pre-training phase the original noisy waveform or the waveform obtained by SE is fed into the self-supervised model to learn the contextual representation, where the quantized clean speech acts as the target. Second, we propose a dual-attention fusion method to fuse the features of noisy and enhanced speech, which can compensate for the information loss caused by separately using individual modules. Due to the flexible exploitation of clean/noisy/enhanced branches, the proposed method turns out to be a generalization of some existing noise-robust ASR models, e.g., enhanced wav2vec2.0. Finally, experimental results on both synthetic and real noisy datasets show that the proposed joint training approach can improve the ASR performance under various noisy settings, leading to a stronger noise robustness.
Qiushi Zhu, Jie Zhang 0042, Li-Rong Dai 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-training
abstract
The rapid development of single-modal pretraining has prompted researchers to pay more attention to cross-modal pre-training methods.In this paper, we propose a unified-modal speech-unit-text pre-training model, SpeechUT, to connect the representations of a speech encoder and a text decoder with a shared unit encoder.Leveraging hidden-unit as an interface to align speech and text, we can decompose the speech-to-text model into a speech-to-unit model and a unit-to-text model, which can be jointly pre-trained with unpaired speech and text data respectively.Our proposed SpeechUT is fine-tuned and evaluated on automatic speech recognition (ASR) and speech translation (ST) tasks.Experimental results show that SpeechUT gets substantial improvements over strong baselines, and achieves state-of-the-art performance on both the Lib-riSpeech ASR and MuST-C ST tasks.To better understand the proposed SpeechUT, detailed analyses are conducted.The code and pretrained models are available at https://aka. ms/SpeechUT.
Junyi Ao, Shujie Liu 0001, Li-Rong Dai 0001, Jinyu Li 0001, Furu Wei
EMNLP5
2022 Self-Supervised Representation Learning for Unsupervised Anomalous Sound Detection Under Domain Shift
abstract
In this paper, a self-supervised representation learning method is proposed for anomalous sound detection (ASD). ASD has received much research attention in recent DCASE challenges. It aims to identify whether a sound emitted from a machine is anomalous or not, given only normal sound data. This is a challenging task due to highly variable time-frequency characteristics of sounds from different machine types, and the fact that many attributes affect machine without being anomalous. This is especially true for domain shift tasks, where only a few training sound clips are available. From the perspective of self-supervised learning, each given sound clip can be considered as a transformation of an original clean sound, where the attribute of each clip may indicate different supervision signals. We propose a unified representation learning framework, equipped with a time-frequency attention mechanism, to perform ASD for different machine types and attributes. For domain shift, a centre imprinting method, which directly sets centres for target domain attributes, is presented. This provides immediate good representation and an initialization for further fine-tuning. Evaluation on DCASE2021 ASD task demonstrates the effectiveness of the proposed method.
Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001, Lin Liu 0017
ICASSP3
2022 Reference Microphone Selection and Low-Rank Approximation Based Multichannel Wiener Filter with Application to Speech Recognition
abstract
For multichannel speech recognition systems, it is necessary to use a speech enhancement module to suppress ambient noises. Given second-order statistics, the multichannel Wiener filter (MWF) can be designed for noise reduction. It was shown that the MWF noise reduction performance depends on the selection of reference microphone and the rank of the speech correlation matrix. It is questionable how the reference microphone and rank would affect the subsequent recognition accuracy. In this paper, we present an experimental study on the low-rank approximation and reference microphone selection based MWF with application to noisy speech recognition. Further, we propose to maximize the input signal-to-noise ratio (SNR) for reference selection in the sense of signal quality. Experimental results show that the output SNR of rank-1 MWF is independent of the reference, while the speech intelligibility is always related to both the rank and reference microphone. The word error rate is positively affected by the rank, and the proposed reference selection method can improve the performance in terms of both speech intelligibility and speech recognition.
Xing-Yu Chen, Jie Zhang 0042, Li-Rong Dai 0001
ICASSP3
2022 Supervised and Self-Supervised Pretraining Based Covid-19 Detection Using Acoustic Breathing/Cough/Speech Signals
abstract
A rapid-accurate detection method for COVID-19 is rather important for avoiding its pandemic. In this work, we propose a bi-directional long short-term memory (BiLSTM) network based COVID-19 detection method using breath/speech/cough signals. Three kinds of acoustic signals are taken to train the network and individual models for three tasks are built, respectively, whose parameters are averaged to obtain an average model, which is then used as the initialization for the BiLSTM model training of each task. It is shown that such an initialization method can significantly improve the detection performance on three tasks. This is called supervised pre-training based detection. Besides, we utilize an existing pre-trained wav2vec2.0 model and pre-train it using the DiCOVA dataset, which is utilized to extract a high-level representation as the model input to replace conventional mel-frequency cepstral coefficients (MFCC) features. This is called self-supervised pre-training based detection. To reduce the information redundancy contained in the recorded sounds, silent segment removal, amplitude normalization and time-frequency masking are also considered. The proposed detection model is evaluated on the DiCOVA dataset and results show that our method achieves an area under curve (AUC) score of 88.44% on blind test in the fusion track. It is shown that using high-level features together with MFCC features is helpful for diagnosing accuracy.
Xing-Yu Chen, Qiushi Zhu, Jie Zhang 0042, Li-Rong Dai 0001
ICASSP4
2022 Domain Robust Deep Embedding Learning for Speaker Recognition
abstract
This paper presents a domain robust deep embedding learning method for speaker verification (SV) tasks. Most recent methods utilize deep neural networks (DNN) to learn compact and discriminative speaker embeddings from large-scale labeled datasets such as VoxCeleb and the NIST SRE corpus. Despite the success of exiting methods, performance may degrade significantly for new target datasets, mainly due to the distribution discrepancy between training and test domains. Moreover, how corpora are collected, and the languages they contain differ, leading to them spanning multiple, perhaps mismatched, latent domains. To address this, a multi-task end-to-end framework is proposed to learn speaker embeddings from both labeled source and unlabeled target datasets. Motivated by label smoothing, a smoothed knowledge distillation (SKD) based self-supervised learning method is designed to exploit latent structural information from the unlabeled target domain. Furthermore, a domain-aware batch normalization (DABN) module aims to reduce the cross-domain distribution discrepancy, while a domain-agnostic instance normalization (DAIN) module aims to learn features that are robust to within-domain variance. Evaluation on NIST SRE16 demonstrates significant performance gains.
Hang-Rui Hu, Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001, Lin Liu 0017
ICASSP4
2022 Frontend Attributes Disentanglement for Speech Emotion Recognition
abstract
Speech emotion recognition (SER) with limited size dataset is a challenging task, since a spoken utterance contains various disturbing attributes besides emotion, including speaker, content, and language. However, due to a close relationship between speaker and emotion attributes, simply fine-tuning a linear model is enough to obtain a good SER performance on the utterance-level embeddings (i.e., i-vector and x-vectors) extracted from the pre-trained speaker recognition (SR) frontends. In this paper, we aim to perform frontend attributes disentanglement (AD) for SER task, using a pre-trained SR model. Specifically, the AD module consists of attribute normalization (AN) and attribute reconstruction (AR) phases. The AN filters out the variation information using instance normalization (IN), and AR reconstructs the emotion-relevant features from the residual space to ensure high emotion discrimination. For better disentanglement, a dual space loss is then designed to encourage the separability of emotion-relevant and emotion-irrelevant spaces. To introduce the long-range contextual information for emotion related reconstruction, a time-frequency (TF) attention is further proposed. Different from the style disentanglement of the extracted x-vectors, the proposed AD module can be applied on frontend feature extractor. Experiments on IEMOCAP benchmark demonstrate the effectiveness of the proposed method.
Yuxuan Xi, Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001, Lin Liu 0017
ICASSP3
2022 A Noise-Robust Self-Supervised Pre-Training Model Based Speech Representation Learning for Automatic Speech Recognition
abstract
Wav2vec2.0 is a popular self-supervised pre-training framework for learning speech representations in the context of automatic speech recognition (ASR). It was shown that wav2vec2.0 has a good robustness against the domain shift, while the noise robustness is still unclear. In this work, we therefore first analyze the noise robustness of wav2vec2.0 via experiments. We observe that wav2vec2.0 pre-trained on noisy data can obtain good representations and thus improve the ASR performance on the noisy test set, which however brings a performance degradation on the clean test set. To avoid this issue, in this work we propose an enhanced wav2vec2.0 model. Specifically, the noisy speech and the corresponding clean version are fed into the same feature encoder, where the clean speech provides training targets for the model. Experimental results reveal that the proposed method can not only improve the ASR performance on the noisy test set which surpasses the original wav2vec2.0, but also ensure a tiny performance decrease on the clean test set. In addition, the effectiveness of the proposed method is demonstrated under different types of noise conditions.
Qiushi Zhu, Jie Zhang 0042, Ming-Hui Wu, Li-Rong Dai 0001
ICASSP6
2022 Learning Contextually Fused Audio-Visual Representations For Audio-Visual Speech Recognition
abstract
With the advance in self-supervised learning for audio and visual modalities, it has become possible to learn a robust audio-visual speech representation. This would be beneficial for improving the audio-visual speech recognition (AVSR) performance, as the multi-modal inputs contain more fruitful information in principle. In this paper, based on existing self-supervised representation learning methods for audio modality, we therefore propose an audio-visual representation learning approach. The proposed approach explores both the complementarity of audio-visual modalities and long-term context dependency using a transformer-based fusion module and a flexible masking strategy. After pre-training, the model is able to extract fused representations required by AVSR. Without loss of generality, it can be applied to single-modal tasks, e.g., audio/visual speech recognition by simply masking out one modality in the fusion module. The proposed pre-trained model is evaluated on speech recognition and lipreading tasks using one or two modalities, where the superiority is revealed.
Jie Zhang 0042, Jianshu Zhang 0001, Ming-Hui Wu, Li-Rong Dai 0001
ICIP6
2022 An Experimental Comparison between Low-Resource Semi-Supervised and High-Resource Supervised Automatic Speech Recognition Models
abstract
Automatic speech recognition (ASR) is an important module in many multimedia applications. Recently, semi-supervised ASR has attracted increasing attention, which can be classified into self-supervised learning and self-training. It was shown that the combination of the two methods is beneficial for ASR on open datasets, e.g., LibriSpeech. However, it is still not completely clear how the semi-supervised model behaves on more challenging industrial datasets compared with supervised approaches using high-resource labeled data. In this paper, we therefore present an experimental study on the combination of self-supervised learning (e.g., wav2vec 2.0) and self-training (e.g., noisy student) in a low-resource industrial setting. The in-domain pre-trained and fine-tuned wav2vec 2.0 model is utilized to teach a baseline VGG-Transformer by pseudo-labeling in an industrial setting. Results reveal that without extra language models the combined semi-supervised acoustic model using much less labeled data can perform as well as the supervised counterpart, which would be rather beneficial for the application of semi-supervised ASR models in low-resource scenarios.
Ao-Ran Gan, Jie Zhang 0042, Ming-Hui Wu, Li-Rong Dai 0001
ICME5
2022 Structural String Decoder for Handwritten Mathematical Expression Recognition
abstract
Recently, recognition of handwritten mathematical expression has been greatly improved by employing sequence modeling methods such as encoder-decoder based methods. Existing encoder-decoder models use string decoders or tree decoders to generate markup of mathematical expression recognition. String decoders directly generate LaTeX strings and tree decoders decode expressions into tree structures. The generalization of string decoders is poor on mathematical expressions with complex hierarchical structures, but its language model is better. Tree decoders can deal with the complex hierarchical structures, but its language model is weakened. In order to take advantage of the above two decoders, we propose a novel structural string decoder (SSD) which not only has good generalization but also can make good use of language model. We demonstrate how the proposed SSD outperforms state-of-the-art string decoders and tree decoders through a set of experiments on CROHME database, which is currently the largest benchmark for online handwritten mathematical expression recognition.
Jiajia Wu 0003, Jinshui Hu, Mingjun Chen, Li-Rong Dai 0001, Xuejing Niu
ICPR4
2022 Pre-Training Transformer Decoder for End-to-End ASR Model with Unpaired Speech Data
abstract
This paper studies a novel pre-training technique with unpaired speech data, Speech2C, for encoder-decoder based automatic speech recognition (ASR).Within a multi-task learning framework, we introduce two pre-training tasks for the encoderdecoder network using acoustic units, i.e., pseudo codes, derived from an offline clustering model.One is to predict the pseudo codes via masked language modeling in encoder output, like HuBERT model, while the other lets the decoder learn to reconstruct pseudo codes autoregressively instead of generating textual scripts.In this way, the decoder learns to reconstruct original speech information with codes before learning to generate correct text.Comprehensive experiments on the LibriSpeech corpus show that the proposed Speech2C can relatively reduce the word error rate (WER) by 19.2% over the method without decoder pre-training, and also outperforms significantly the state-of-the-art wav2vec 2.0 and Hu-BERT on fine-tuning subsets of 10h and 100h.We release our code and model at https://github.com/microsoft/SpeechT5/tree/main/Speech2C.
Junyi Ao, Shujie Liu 0001, Haizhou Li 0001, Tom Ko, Li-Rong Dai 0001, Jinyu Li 0001, Yao Qian, Furu Wei
INTERSPEECH7
2022 A Complementary Joint Training Approach Using Unpaired Speech and Text A Complementary Joint Training Approach Using Unpaired Speech and Text
Ye-Qian Du, Jie Zhang 0042, Qiushi Zhu, Li-Rong Dai 0001, Ming-Hui Wu, Zhouwang Yang
INTERSPEECH4
2022 Class-Aware Distribution Alignment based Unsupervised Domain Adaptation for Speaker Verification
abstract
Existing speaker verification (SV) systems usually suffer from significant performance degradation when applied to a new domain that lies outside the training distribution. Given the unlabeled target-domain dataset, most Unsupervised Domain Adaptation (UDA) methods aim to minimize the distribution divergence between different domains. However, global distribution alignment strategies fail to consider the latent speaker label information and can hardly guarantee the feature discriminative capability in target domain. In this paper, we propose a novel UDA approach called WBDA (Within-class and Between-class Distribution Alignment), which aims to transfer the class-aware information (i.e., within- and between-class distributions) learned from the well-labeled source-domain to unlabeled target-domain. Motivated by the recent progress of self-supervised contrastive learning, the positive and negative pairs are constructed separately for source and target domains, from which the within- and between-class distribution can be estimated. And the SV system can then be learned by jointly optimizing the cross-domain class-aware distribution discrepancy loss and source-domain classification loss in an end-to-end manner. Evaluations on NIST SRE16 and SRE18 achieve a relative performance improvement of about 43.7% and 26.2% over the baseline in terms of Equal Error Rate (EER) separately, significantly outperforming the previous adaption methods based on global distribution alignment.
Hang-Rui Hu, Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001, Lin Liu 0017
INTERSPEECH3
2022 Differential Time-frequency Log-mel Spectrogram Features for Vision Transformer Based Infant Cry Recognition
Hai-tao Xu, Jie Zhang 0042, Li-Rong Dai 0001
INTERSPEECH3
2022 External Text Based Data Augmentation for Low-Resource Speech Recognition in the Constrained Condition of OpenASR21 Challenge
Guolong Zhong, Hongyu Song, Ruoyu Wang 0029, Lei Sun 0010, Diyuan Liu, Jun Du 0002, Jie Zhang 0042, Li-Rong Dai 0001
INTERSPEECH10
2022 A multimodal attention fusion network with a dynamic vocabulary for TextVQA
Jiajia Wu 0003, Jun Du 0002, Fengren Wang, Xinzhe Jiang, Jinshui Hu, Jianshu Zhang 0001, Li-Rong Dai 0001
Pattern Recognit.9
2022 Frequency-Invariant Sensor Selection for MVDR Beamforming in Wireless Acoustic Sensor Networks
abstract
Wireless acoustic sensor network (WASN) has a wide range of applications in internet of things, where signal estimation is one of the network design objectives. Due to the existence of ambient noises, the recorded audio signals are inevitably corrupted, resulting in a low signal-to-noise ratio (SNR), which triggers the necessity of signal enhancement. As using all sensor measurements brings a large amount of data transmissions and computational cost, the narrowband sensor selection was proposed to choose an informative subset of sensors to perform noise reduction in the audio context. However, the resulting frequency-dependent selection status has to be switched across frequencies. In order to avoid the complicated switching operations, we consider frequency-invariant sensor selection in this work. We propose to minimize the total power consumption over the WASN by constraining the broadband SNR, which can be solved using broadband semi-definite optimization (BroadOpt) or narrowband voting (NaVo) approaches. In order to further reduce the time complexity, we propose two near-optimal greedy methods, including gradient removal (GradR) and weighted input SNR removal (SnrR). As comparison, we also show a broadband energy removal (EnergyR) method. The greedy methods remove one sensor at each iteration from the complete network until the performance constraint is not satisfied. Numerical results using a simulated large-scale WASN show that the greedy methods can achieve a comparable performance compared to the optimization based counterparts, while the corresponding time complexity is much lower. In general, the sensors around the target source and the fusion center are more likely to be selected.
Jie Zhang 0042, Li-Rong Dai 0001
IEEE Trans. Wirel. Commun.3
2021 TaLNet: Voice Reconstruction from Tongue and Lip Articulation with Transfer Learning from Text-to-Speech Synthesis
abstract
This paper presents TaLNet, a model for voice reconstruction with ultrasound tongue and optical lip videos as inputs. TaLNet is based on an encoder-decoder architecture. Separate encoders are dedicated to processing the tongue and lip data streams respectively. The decoder predicts acoustic features conditioned on encoder outputs and speaker codes.To mitigate for having only relatively small amounts of dual articulatory-acoustic data available for training, and since our task here shares with text-to-speech (TTS) the common goal of speech generation, we propose a novel transfer learning strategy to exploit the much larger amounts of acoustic-only data available to train TTS models. For this, a Tacotron 2 TTS model is first trained, and then the parameters of its decoder are transferred to the TaLNet decoder. We have evaluated our approach on an unconstrained multi-speaker voice recovery task. Our results show the effectiveness of both the proposed model and the transfer learning strategy. Speech reconstructed using our proposed method significantly outperformed all baselines (DNN, BLSTM and without transfer learning) in terms of both naturalness and intelligibility. When using an ASR model decoding the recovery speech, the WER of our proposed method is relatively reduced over 30% compared to baselines.
Jing-Xuan Zhang, Korin Richmond, Zhen-Hua Ling, Li-Rong Dai 0001
AAAI4
2021 An Effective Deep Embedding Learning Method Based on Dense-Residual Networks for Speaker Verification
abstract
In this paper, we present an effective end-to-end deep embedding learning method based on Dense-Residual networks, which combine the advantages of a densely connected convolutional network (DenseNet) and a residual network (ResNet), for speaker verification (SV). Unlike a model ensemble strategy which merges the results of multiple systems, the proposed Dense-Residual networks perform feature fusion on every basic DenseR building block. Specifically, two types of DenseR blocks are designed. A sequential-DenseR block is constructed by densely connecting stacked basic units in a residual block of ResNet. A parallel-DenseR comprises split and concatenation operations on residual and dense components via corresponding skip connections. These building blocks are stacked into deep networks to exploit the complementary information with different receptive field sizes and growth rates. Extensive experiments have been conducted on the VoxCeleb1 dataset to evaluate the proposed methods. The SV performance achieved by the proposed Dense-Residual networks is shown to outperform corresponding ResNet, DenseNet or fusions of them, with similar model complexity, by a significant margin.
Yan Song 0001, Ian McLoughlin 0001, Lin Liu 0017, Li-Rong Dai 0001
ICASSP5
2021 An Improved Mean Teacher Based Method for Large Scale Weakly Labeled Semi-Supervised Sound Event Detection
abstract
This paper presents an improved mean teacher (MT) based method for large-scale weakly labeled semi-supervised sound event detection (SED), by focusing on learning a better student model. Two main improvements are proposed based on the authors’ previous perturbation based MT method. Firstly, an event-aware module is de-signed to allow multiple branches with different kernel sizes to be fused via an attention mechanism. By inserting this module after the convolutional layer, each neuron can adaptively adjust its receptive field to suit different sound events. Secondly, instead of using the teacher model to provide a consistency cost term, we propose using a stochastic inference of unlabeled examples to generate high quality pseudo-targets by averaging multiple predictions from the perturbed student model. MixUp of both labeled and unlabeled data is further exploited to improve the effectiveness of student model. Finally, the teacher model can be obtained via exponential moving average (EMA) of the student model, which generates final predictions for SED during inference. Experiments on the DCASE2018 task4 dataset demonstrate the ability of the proposed method. Specifically, an F1-score of 42.1% is achieved, significantly outperforming the 32.4% achieved by the winning system, or the 39.3% by the previous perturbation based method.
Yan Song 0001, Ian McLoughlin 0001, Lin Liu 0017, Li-Rong Dai 0001
ICASSP5
2021 Automatic Lip-Reading with Hierarchical Pyramidal Convolution and Self-Attention for Image Sequences with No Word Boundaries
Hang Chen 0001, Jun Du 0002, Yu Hu 0003, Li-Rong Dai 0001, Chin-Hui Lee 0001
Interspeech4
2021 A Weight Moving Average Based Alternate Decoupled Learning Algorithm for Long-Tailed Language Identification
abstract
Language identification (LID) research has made tremendous progress in recent years, especially with the introduction of deep learning techniques. However, for real-world applications where the distribution of different language data is highly imbalanced, the performance of existing LID systems is still far from satisfactory. This raises the challenge of long-tailed LID. In this paper, we propose an effective weight moving average (WMA) based alternate decoupled learning algorithm, termed WADCL, for long-tailed LID. The system is divided into two components, a frontend feature extractor and a backend classifier. These are then alternately learned in an end-to-end manner using different sampling schemes to alleviate the distribution mismatch between training and test datasets. Furthermore, our WMA method aims to mitigate the side-effects of re-sampling schemes, by fusing the model parameters learned along the trajectory of stochastic gradient descent (SGD) optimization. To validate the effectiveness of the proposed WADCL algorithm, we evaluate and compare several systems over a language dataset constructed to match a long-tailed distribution based on real world application [1]. The experimental results from the long-tailed language dataset demonstrate that the proposed algorithm is able to achieve significant performance gains over existing state-of-the-art x-vector based LID methods.
Lin Liu 0017, Yan Song 0001, Ian McLoughlin 0001, Li-Rong Dai 0001
Interspeech6
2021 An Effective Mutual Mean Teaching Based Domain Adaptation Method for Sound Event Detection
abstract
In this paper, we present a novel mutual mean teaching based domain adaptation (MMT-DA) method for sound event detection (SED) task, which can effectively exploit synthetic data to improve the SED performance. Existing methods simply treat the synthetic data as strongly-labeled data in semi-supervised learning (SSL) framework. Benefiting from the strong labels of synthetic data, superior SED performance can be achieved. However, a distribution mismatch between synthetic and real data raises an evident challenge for domain adaptation (DA). In MMT-DA, convolutional recurrent neural networks (CRNN) learned from different datasets (i.e. total data:real+synthetic, and real data) are exploited for DA. Specifically, mean teacher method using CRNN is employed for utilizing the unlabeled real data. To compensate the domain diversity, an additional domain classifier with gradient reverse layer(GRL) is used for training a mean teacher for total data. The student CRNNs are mutually taught using the soft predictions of unlabeled data obtained from different teachers. Furthermore, a strip pooling based attention module is exploited to model the inter-dependencies between channels and time-frequency dimensions to exploit the structure information. Experimental results on Task4 of DCASE2020 demonstrate the ability of the proposed method, achieving 52.0% F1-score on the validation dataset, which outperforms the winning system’s 50.6%.
Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001, Lin Liu 0017
Interspeech3
2021 UnitNet-Based Hybrid Speech Synthesis
Xiao Zhou 0024, Zhen-Hua Ling, Li-Rong Dai 0001
Interspeech3
2021 An Improved Wav2Vec 2.0 Pre-Training Approach Using Enhanced Local Dependency Modeling for Speech Recognition
Qiushi Zhu, Jie Zhang 0042, Ming-Hui Wu, Li-Rong Dai 0001
Interspeech5
2021 Correlating subword articulation with lip shapes for embedding aware audio-visual speech enhancement
Hang Chen 0001, Jun Du 0002, Yu Hu 0003, Li-Rong Dai 0001, Chin-Hui Lee 0001
Neural Networks4
2021 Multi-Granularity Sequence Alignment Mapping for Encoder-Decoder Based End-to-End ASR
abstract
Encoder-decoder based automatic speech recognition (ASR) methods are increasingly popular due to their simplified processing stages and low reliance on prior knowledge. Conventional encoder-decoder based approaches usually learn a sequence-to-sequence mapping function from the source speech to target units (e.g., subwords, characters) in an end-to-end manner. However, it is still unclear how to choose the optimal target unit, or granularity of multiple units. In general, as increasing the information available for learning sequence-to-sequence mapping functions can improve modeling effectiveness, we therefore propose a multi-granularity sequence alignment (MGSA) approach. This aims to enhance cross-sequence interactions between different granularity units for both modeling and inference stages in the encoder-decoder based ASR. Specifically, a decoder module is designed to generate multi-granularity sequence predictions. We then exploit the latent alignment mapping among units having different levels of granularity, by utilizing the decoded multi-level sequences as input for model prediction. The cross-sequence interaction can also be employed to re-calibrate output probabilities in the proposed post-inference algorithm. Experimental results on both WSJ-80 hrs and Switchboard-300 hrs datasets show the superiority of the proposed method compared to traditional multi-task methods as well as to single granularity baseline systems.
Jie Zhang 0042, Yan Song 0001, Ian McLoughlin 0001, Li-Rong Dai 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2021 A Study on Reference Microphone Selection for Multi-Microphone Speech Enhancement
abstract
Multi-microphone speech enhancement methods typically require a reference position with respect to which the target signal is estimated. Often, this reference position is arbitrarily chosen as one of the reference microphones. However, it has been shown that the choice of the reference microphone can have a significant impact on the final noise reduction performance. In this paper, we therefore theoretically analyze the impact of selecting a reference on the noise reduction performance with near-end noise being taken into account. Following the generalized eigenvalue decomposition (GEVD) based optimal variable span filtering framework, we find that for any linear beamformer, the output signal-to-noise ratio (SNR) taking both the near-end and far-end noise into account is reference dependent. Only when the near-end noise is neglected, the output SNR of rank-1 beamformers does not depend on the reference position. However, in general for rank-r beamformers with r > 1 (e.g., the multichannel Wiener filter) the performance does depend on the reference position. Based on these, we propose an optimal algorithm for microphone reference selection that maximizes the output SNR. In addition, we propose a lower-complexity algorithm that is still optimal for rank-1 beamformers, but sub-optimal for the general r > 1 rank beamformers. Experiments using a simulated microphone array validate the effectiveness of both proposed methods and show that in terms of quality, several dB can be gained by selecting the proper reference microphone.
Jie Zhang 0042, Li-Rong Dai 0001, Richard C. Hendriks
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Sensor Selection for Relative Acoustic Transfer Function Steered Linearly-Constrained Beamformers
abstract
For multi-microphone speech enhancement, different microphones might have different contributions, assome are even marginal. This is more likely to happen in wireless acoustic sensor networks (WASNs), where somesensors might be distant. In this work, we therefore consider sensor selection for linearly-constrained beamformers. Theproposed sensor selection approach is formulated by minimizing the total output noise power and constraining thenumber of selected sensors. As the considered sensor selection problem requires the relative acoustic transfer function(RTF), the covariance whitening based RTF estimation or a direct-path RTF approximation is exploited. For a singletarget source, we can thus substitute the estimated RTF or the assumed RTF to the original problem formulation in orderto design a minimum variance distortionless response (MVDR) beamformer. Alternatively, we can integrate the two RTFsto design a linearly constrained minimum variance (LCMV) beamformer in order to alleviate the effects of RTFestimation/approximation errors. By leveraging the superiority of LCMV beamformers, the proposed approach can beapplied to the multi-source case. An evaluation using a simulated large-scale WASN demonstrates that the integration ofRTFs for the sensor selection based LCMV beamformer can be beneficial as opposed to relying on either of theindividual RTF steered sensor selection based MVDR beamformers. We conclude that the sensors that are close to thetarget source(s) and also some around the coherent interferers are more informative.
Jie Zhang 0042, Jun Du 0002, Li-Rong Dai 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 UnitNet: A Sequence-to-Sequence Acoustic Model for Concatenative Speech Synthesis
abstract
This paper presents UnitNet, a sequence-to-sequence (Seq2Seq) acoustic model for concatenative speech synthesis. Comparing with the Tacotron2 model for Seq2Seq speech synthesis, UnitNet utilizes the phone boundaries of training data and its decoder contains autoregressive structures at both phone and frame levels. This hierarchical architecture can not only extract embedding vectors for representing phone-sized units in the corpus but also measure the dependency among consecutive units, which makes the UnitNet model capable of guiding the selection of phone-sized units for concatenative speech synthesis. A byproduct of this model is that it can also be applied to statistical parametric speech synthesis (SPSS) and improve the robustness of Seq2Seq acoustic feature prediction since it adopts interpretable transition probability prediction rather than attention mechanism for frame-level alignment. Experimental results show that our UnitNet-based concatenative speech synthesis method not only outperforms the unit selection methods using hidden Markov models and Tacotron-based unit embeddings, but also achieves better naturalness and faster inference speed than the SPSS method using FastSpeech and Parallel WaveGAN. Besides, the UnitNet-based SPSS method makes fewer synthesis errors than Tacotron2 and FastSpeech without naturalness degradation.
Xiao Zhou 0024, Zhen-Hua Ling, Li-Rong Dai 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 SRD: A Tree Structure Based Decoder for Online Handwritten Mathematical Expression Recognition
abstract
Recently, recognition of online handwritten mathe- matical expression has been greatly improved by employing encoder-decoder based methods. Existing encoder-decoder models use string decoders to generate LaTeX strings for mathematical expression recognition. However, in this paper, we importantly argue that string representations might not be the most natural for mathematical expressions – mathematical expressions are inherently tree structures other than flat strings. For this purpose, we propose a novel sequential relation decoder (SRD) that aims to decode expressions into tree structures for online handwritten mathematical expression recognition. At each step of tree construction, a sub-tree structure composed of a relation node and two symbol nodes is computed based on previous sub-tree structures. This is the first work that builds a tree structure based decoder for encoder-decoder based mathematical expression recognition. Compared with string decoders, a decoder that better understands tree structures is crucial for mathematical expression recognition as it brings a more reasonable learning objective and improves overall generalization ability. We demonstrate how the proposed SRD outperforms state-of-the-art string decoders through a set of experiments on CROHME database, which is currently the largest benchmark for online handwritten mathematical expression recognition.
Jianshu Zhang 0001, Jun Du 0002, Yongxin Yang, Yi-Zhe Song, Li-Rong Dai 0001
IEEE Trans. Multim.5
2020 Attention-Based Gated Scaling Adaptive Acoustic Model for CTC-Based Speech Recognition
abstract
In this paper, we propose a novel adaptive technique that uses an attention-based gated scaling (AGS) scheme to improve deep feature learning for connectionist temporal classification (CTC) acoustic modeling. In AGS, the outputs of each hidden layer of the main network are scaled by an auxiliary gate matrix extracted from the lower layer by using an attention mechanism. Furthermore, the auxiliary AGS layer and the main network are jointly trained without requiring second-pass model training or additional speaker information, such as i-vector. On the Mandarin AISHELL-1 dataset, the proposed AGS yields a 7.94% character error rate (CER). To the best of our knowledge, the results obtained when training on the full AISHELL-1 training set, are the best published currently for the end-to-end systems.
Fenglin Ding, Wu Guo, Li-Rong Dai 0001, Jun Du 0002
ICASSP3
2020 An Improved Deep Neural Network for Modeling Speaker Characteristics at Different Temporal Scales
abstract
This paper presents an improved deep embedding learning method based on a convolutional neural network (CNN) for text-independent speaker verification. Two improvements are proposed for x-vector embedding learning: (1) a multiscale convolution (MSCNN) is adopted in the frame-level layers to capture the complementary speaker information in different receptive fields; (2) a Baum-Welch statistics attention (BWSA) mechanism is applied in the pooling layer, which can integrate more useful long-term speaker characteristics in the temporal pooling layer. Experiments are carried out on the NIST SRE16 evaluation set. The results demonstrate the effectiveness of the MSCNN and show that the proposed BWSA can further improve the performance of the DNN embedding system.
Bin Gu 0004, Wu Guo, Li-Rong Dai 0001, Jun Du 0002
ICASSP3
2020 An Online Speaker-aware Speech Separation Approach Based on Time-domain Representation
abstract
Despite the significant progress of deep learning based speech separation methods, it remains challenging to extract and track the speech from target speakers, especially in a single-channel multiple speaker situation. Previously, the authors proposed a source-aware context network to exploit the temporal context in mixtures and estimated sources for online speech separation. In this paper, we propose a speaker-aware approach based on the source-aware context network structure, in which the speaker information is explicitly modeled by an auxiliary speaker identification branch. Then speech separation and speaker tracking can be jointly optimized by multi-task learning. Furthermore, we study the effectiveness of time-domain representation by proposing a raw sparse waveform encoder to preserve discriminative information. Experimental results on the WSJ0-2mix benchmark show that the proposed system significantly improves Signal-to-Distortion Ratio (SDR) performance.
Yan Song 0001, Zengxi Li, Ian McLoughlin 0001, Li-Rong Dai 0001
ICASSP5
2020 Task-Aware Mean Teacher Method for Large Scale Weakly Labeled Semi-Supervised Sound Event Detection
abstract
Weakly labeled semi-supervised learning methods have recently drawn increasing attention from the research community for sound event detection tasks. Due to the weakness of the labelling, neural networks are often designed to perform sound event detection (SED) and audio tagging (AT) at the same time. In this paper, we propose a task-aware mean teacher method using a convolutional recurrent neural network (CRNN) with multi-branch structure to solve the SED and AT tasks differently. Specifically, a branch with coarse-level temporal resolution is designed for the AT task, while a branch with fine-level temporal resolution is designed for the SED task. The mean teacher based semi-supervised learning method is first adopted to improve the performance of the coarse-level AT branch by exploiting unlabeled data. Then the coarse-level AT branch is introduced as a teacher to guide the aggregated AT output of the fine-level SED branch, yielding an improvement in the SED performance. To further improve the AT and SED performance, information from multiple layers is exploited in the form of a multi-resolution feature. Experimental results on Task4 of the DCASE2018 challenge demonstrate the superiority of the proposed method, achieving 37.7% F1-score, which outperforms the winning system's 32.4%.
Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001
ICASSP3
2020 Extracting Unit Embeddings Using Sequence-To-Sequence Acoustic Models for Unit Selection Speech Synthesis
abstract
This paper presents a method of using the intermediate representations between linguistic and acoustic features in a Tacotron model to derive the cost functions for unit selection speech synthesis. By extracting the outputs of the Tacotron encoder, each phone-sized candidate unit in the corpus is represented by a fixed-length unit vector. Similarly, each target unit to be synthesized is also converted into a unit vector of the same dimension by encoding the input phone sequence. The normalized Euclidean distances between these two vectors are utilized to fulfill unit pre-selection and to calculate the target cost for unit selection. Then, another DNN which predicts the unit vector of each phone from its preceding ones is constructed to derive the concatenation cost function. Experimental results demonstrate that the unit vectors extracted from Tacotron contain both duration and acoustic information of phone units. Comparing with our previous work, which learned unit vectors using a DNN and only acoustic features, the method proposed in this paper further improves the naturalness of unit selection speech synthesis in our experiments.
Xiao Zhou 0024, Zhen-Hua Ling, Li-Rong Dai 0001
ICASSP3
2020 A Tree-Structured Decoder for Image-to-Markup Generation
abstract
Recent encoder-decoder approaches typically employ string decoders to convert images into serialized strings for image-to-markup. However, for tree-structured representational markup, string representations can hardly cope with the structural complexity. In this work, we first show via a set of toy problems that string decoders struggle to decode tree structures, especially as structural complexity increases, we then propose a tree-structured decoder that specifically aims at generating a tree-structured markup. Our decoders works sequentially, where at each step a child node and its parent node are simultaneously generated to form a sub-tree. This sub-tree is consequently used to construct the final tree structure in a recurrent manner. Key to the success of our tree decoder is twofold, (i) it strictly respects the parent-child relationship of trees, and (ii) it explicitly outputs trees as oppose to a linear string. Evaluated on both math formula recognition and chemical formula recognition, the proposed tree decoder is shown to greatly outperform strong string decoder baselines.
Jianshu Zhang 0001, Jun Du 0002, Yongxin Yang, Yi-Zhe Song, Si Wei, Li-Rong Dai 0001
ICML6
2020 An Effective Speaker Recognition Method Based on Joint Identification and Verification Supervisions
abstract
Deep embedding learning based speaker verification methods have attracted significant recent research interest due to their superior performance. Existing methods mainly focus on designing frame-level feature extraction structures, utterance-level aggregation methods and loss functions to learn discriminative speaker embeddings. The scores of verification trials are then computed using cosine distance or Probabilistic Linear Discriminative Analysis (PLDA) classifiers. This paper proposes an effective speaker recognition method which is based on joint identification and verification supervisions, inspired by multi-task learning frameworks. Specifically, a deep architecture with convolutional feature extractor, attentive pooling and two classifier branches is presented. The first, an identification branch, is trained with additive margin softmax loss (AM-Softmax) to classify the speaker identities. The second, a verification branch, trains a discriminator with binary cross entropy loss (BCE) to optimize a new triplet-based mutual information. To balance the two losses during different training stages, a ramp-up/ramp-down weighting scheme is employed. Furthermore, an attentive bilinear pooling method is proposed to improve the effectiveness of embeddings. Extensive experiments have been conducted on VoxCeleb1 to evaluate the proposed method, demonstrating results that relatively reduce the equal error rate (EER) by 22% compared to the baseline system using identification supervision only.
Yan Song 0001, Yiheng Jiang, Ian McLoughlin 0001, Lin Liu 0017, Li-Rong Dai 0001
INTERSPEECH6
2020 Recognition-Synthesis Based Non-Parallel Voice Conversion with Adversarial Learning
abstract
This paper presents an adversarial learning method for recognition-synthesis based non-parallel voice conversion.A recognizer is used to transform acoustic features into linguistic representations while a synthesizer recovers output features from the recognizer outputs together with the speaker identity.By separating the speaker characteristics from the linguistic representations, voice conversion can be achieved by replacing the speaker identity with the target one.In our proposed method, a speaker adversarial loss is adopted in order to obtain speaker-independent linguistic representations using the recognizer.Furthermore, discriminators are introduced and a generative adversarial network (GAN) loss is used to prevent the predicted features from being over-smoothed.For training model parameters, a strategy of pre-training on a multi-speaker dataset and then fine-tuning on the source-target speaker pair is designed.Our method achieved higher similarity than the baseline model that obtained the best performance in Voice Conversion Challenge 2018.
Jing-Xuan Zhang, Zhen-Hua Ling, Li-Rong Dai 0001
INTERSPEECH3
2020 Semi-Supervised End-to-End ASR via Teacher-Student Learning with Conditional Posterior Distribution
abstract
Encoder-decoder based methods have become popular for automatic speech recognition (ASR), thanks to their simplified processing stages and low reliance on prior knowledge. However, large amounts of acoustic data with paired transcriptions is generally required to train an effective encoder-decoder model, which is expensive, time-consuming to be collected and not always readily available. However unpaired speech data is abundant, hence several semi-supervised learning methods, such as teacher-student (T/S) learning and pseudo-labeling, have recently been proposed to utilize this potentially valuable resource. In this paper, a novel T/S learning with conditional posterior distribution for encoder-decoder based ASR is proposed. Specifically, the 1-best hypotheses and the conditional posterior distribution from the teacher are exploited to provide more effective supervision. Combined with model perturbation techniques, the proposed method reduces WER by 19.2% relatively on the LibriSpeech benchmark, compared with a system trained using only paired data. This outperforms previous reported 1-best hypothesis results on the same task.
Yan Song 0001, Jianshu Zhang 0001, Ian McLoughlin 0001, Li-Rong Dai 0001
INTERSPEECH5
2020 An Effective Perturbation Based Semi-Supervised Learning Method for Sound Event Detection
abstract
Mean teacher based methods are increasingly achieving state-of-the-art performance for large-scale weakly labeled and unlabeled sound event detection (SED) tasks in recent DCASE challenges. By penalizing inconsistent predictions under different perturbations, mean teacher methods can exploit large-scale unlabeled data in a self-ensembling manner. In this paper, an effective perturbation based semi-supervised learning (SSL) method is proposed based on the mean teacher method. Specifically, a new independent component (IC) module is proposed to introduce perturbations for different convolutional layers, designed as a combination of batch normalization and dropblock operations. The proposed IC module can reduce correlation between neurons to improve performance. A global statistics pooling based attention module is further proposed to explicitly model inter-dependencies between the time-frequency domain and channels, using statistics information (e.g. mean, standard deviation, max) along different dimensions. This can provide an effective attention mechanism to adaptively re-calibrate the output feature map. Experimental results on Task 4 of the DCASE2018 challenge demonstrate the superiority of the proposed method, achieving about 39.8% F1-score, outperforming the previous winning system’s 32.4% by a significant margin.
Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001, Lin Liu 0017
INTERSPEECH4
2020 Radical analysis network for learning hierarchies of Chinese characters
Jianshu Zhang 0001, Jun Du 0002, Li-Rong Dai 0001
Pattern Recognit.3
2020 Learning and Modeling Unit Embeddings Using Deep Neural Networks for Unit-Selection-Based Mandarin Speech Synthesis
abstract
A method of learning and modeling unit embeddings using deep neutral networks (DNNs) is presented in this article for unit-selection-based Mandarin speech synthesis. Here, a unit embedding is defined as a fixed-length embedding vector for a phone-sized unit candidate in a corpus. Modeling phone-sized embedding vectors instead of frame-sized acoustic features can better measure the long-term dependencies among consecutive units in an utterance. First, a DNN with an embedding layer is built to learn the embedding vectors of all unit candidates in the corpus from scratch. In order to enable the extracted embedding vectors to carry both acoustic and linguistic information of unit candidates, a multitarget learning strategy is designed for the DNN. Its optional prediction targets include frame-level acoustic features, unit durations, monophone and tone identifiers, and context classes. Then, another two DNNs are constructed to map linguistic features toward the extracted embedding vectors. One of them employs the unit vectors of preceding phones besides the linguistic features of current phone as its input. At synthesis time, the distances between the unit vectors predicted by these two DNNs and the ones derived from unit candidates are used as a part of the target cost and a part of the concatenation cost, respectively. Our experiments on a Mandarin speech synthesis corpus demonstrate that learning and modeling unit embeddings improve the naturalness of hidden Markov model (HMM)-based unit selection speech synthesis. Furthermore, integrating multiple targets for learning unit embeddings achieves better performance than using only acoustic targets according to our subjective evaluation results.
Xiao Zhou 0024, Zhen-Hua Ling, Li-Rong Dai 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2020 Non-Parallel Sequence-to-Sequence Voice Conversion With Disentangled Linguistic and Speaker Representations
abstract
This article presents a method of sequence-to-sequence (seq2seq) voice conversion using non-parallel training data. In this method, disentangled linguistic and speaker representations are extracted from acoustic features, and voice conversion is achieved by preserving the linguistic representations of source utterances while replacing the speaker representations with the target ones. Our model is built under the framework of encoder-decoder neural networks. A recognition encoder is designed to learn the disentangled linguistic representations with two strategies. First, phoneme transcriptions of training data are introduced to provide the references for leaning linguistic representations of audio signals. Second, an adversarial training strategy is employed to further wipe out speaker information from the linguistic representations. Meanwhile, speaker representations are extracted from audio signals by a speaker encoder. The model parameters are estimated by two-stage training, including a pre-training stage using a multi-speaker dataset and a fine-tuning stage using the dataset of a specific conversion pair. Since both the recognition encoder and the decoder for recovering acoustic features are seq2seq neural networks, there are no constrains of frame alignment and frame-by-frame conversion in our proposed method. Experimental results showed that our method obtained higher similarity and naturalness than the best non-parallel voice conversion method in Voice Conversion Challenge 2018. Besides, the performance of our proposed method was closed to the state-of-the-art parallel seq2seq voice conversion method.
Jing-Xuan Zhang, Zhen-Hua Ling, Li-Rong Dai 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 A Region Based Attention Method for Weakly Supervised Sound Event Detection and Classification
abstract
Recently, an attention based convolutional recurrent neural network (CRNN) with learnable gated linear units (GLUs) has achieved state-of-the-art performance for audio tagging (AT) and sound event detection (SED) tasks in the Detection and Classification of Acoustic Scenes and Events (DCASE) challenges. The introduction of GLU and temporal attention-based localization mechanisms plays an important role for both AT and SED tasks. In this paper, we propose a novel region based attention method to further boost the representation power of the existing GLU based CRNN. Specifically, we insert a feature selection (FS) structure after each GLU to create what we term a GLU-F. block, to exploit channel relationships. Furthermore, we extract region features (or the prototypes of certain sound events) from multi-scale sliding windows over higher convolutional layers, which are fed into an attention-based recurrent neural network to model their context information for AT and SED tasks. To evaluate the proposed region based attention method, we conduct extensive experiments on SED and AT tasks in DCASE2017. We achieve 59.5% and 60.1% AT F1-score, 51.3% and 55.1% SED F1-score for development and evaluation sets respectively, significantly outperforming state-of-the-art results.
Yan Song 0001, Wu Guo, Li-Rong Dai 0001, Ian McLoughlin 0001
ICASSP4
2019 Improving Sequence-to-sequence Voice Conversion by Adding Text-supervision
abstract
This paper presents methods of making using of text supervision to improve the performance of sequence-to-sequence (seq2seq) voice conversion. Compared with conventional frame-to-frame voice conversion approaches, the seq2seq acoustic modeling method proposed in our previous work achieved higher naturalness and similarity. In this paper, we further improve its performance by utilizing the text transcriptions of parallel training data. First, a multi-task learning structure is designed which adds auxiliary classifiers to the middle layers of the seq2seq model and predicts linguistic labels as a secondary task. Second, a data-augmentation method is proposed which utilizes text alignment to produce extra parallel sequences for model training. Experiments are conducted to evaluate our proposed method with training sets at different sizes. Experimental results show that the multi-task learning with linguistic labels is effective at reducing the errors of seq2seq voice conversion. The data-augmentation method can further improve the performance of seq2seq voice conversion when only 50 or 100 training utterances are available.
Jing-Xuan Zhang, Zhen-Hua Ling, Yuan Jiang 0006, Li-Juan Liu, Li-Rong Dai 0001
ICASSP6
2019 Neural Text Clustering with Document-Level Attention Based on Dynamic Soft Labels
Wu Guo, Li-Rong Dai 0001, Zhen-Hua Ling, Jun Du 0002
INTERSPEECH3
2019 A Chinese Dataset for Identifying Speakers in Novels
Jia-Xiang Chen, Zhen-Hua Ling, Li-Rong Dai 0001
INTERSPEECH3
2019 Improving Aggregation and Loss Function for Better Embedding Learning in End-to-End Speaker Verification System
Zhifu Gao, Yan Song 0001, Ian McLoughlin 0001, Yiheng Jiang, Li-Rong Dai 0001
INTERSPEECH6
2019 An Effective Deep Embedding Learning Architecture for Speaker Verification
Yiheng Jiang, Yan Song 0001, Ian McLoughlin 0001, Zhifu Gao, Li-Rong Dai 0001
INTERSPEECH5
2019 Singing Voice Synthesis Using Deep Autoregressive Neural Networks for Acoustic Modeling
abstract
This paper presents a method of using autoregressive neural networks for the acoustic modeling of singing voice synthesis (SVS).Singing voice differs from speech and it contains more local dynamic movements of acoustic features, e.g., vibratos.Therefore, our method adopts deep autoregressive (DAR) models to predict the F0 and spectral features of singing voice in order to better describe the dependencies among the acoustic features of consecutive frames.For F0 modeling, discretized F0 values are used and the influences of the history length in DAR are analyzed by experiments.An F0 post-processing strategy is also designed to alleviate the inconsistency between the predicted F0 contours and the F0 values determined by music notes.Furthermore, we extend the DAR model to deal with continuous spectral features, and a prenet module with self-attention layers is introduced to process historical frames.Experiments on a Chinese singing voice corpus demonstrate that our method using DARs can produce F0 contours with vibratos effectively, and can achieve better objective and subjective performance than the conventional method using recurrent neural networks (RNNs).
Yuan-Hao Yi, Yang Ai, Zhen-Hua Ling, Li-Rong Dai 0001
INTERSPEECH4
2019 Multi-Task Learning with High-Order Statistics for x-Vector Based Text-Independent Speaker Verification
abstract
The x-vector based deep neural network (DNN) embedding systems have demonstrated effectiveness for text-independent speaker verification.This paper presents a multi-task learning architecture for training the speaker embedding DNN with the primary task of classifying the target speakers, and the auxiliary task of reconstructing the first-and higher-order statistics of the original input utterance.The proposed training strategy aggregates both the supervised and unsupervised learning into one framework to make the speaker embeddings more discriminative and robust.Experiments are carried out using the NIST SRE16 evaluation dataset and the VOiCES dataset.The results demonstrate that our proposed method outperforms the original x-vector approach with very low additional complexity added.
Lanhua You, Wu Guo, Li-Rong Dai 0001, Jun Du 0002
INTERSPEECH3
2019 Deep Neural Network Embeddings with Gating Mechanisms for Text-Independent Speaker Verification
abstract
In this paper, gating mechanisms are applied in deep neural network (DNN) training for x-vector-based text-independent speaker verification. First, a gated convolution neural network (GCNN) is employed for modeling the frame-level embedding layers. Compared with the time-delay DNN (TDNN), the GCNN can obtain more expressive frame-level representations through carefully designed memory cell and gating mechanisms. Moreover, we propose a novel gated-attention statistics pooling strategy in which the attention scores are shared with the output gate. The gated-attention statistics pooling combines both gating and attention mechanisms into one framework; therefore, we can capture more useful information in the temporal pooling layer. Experiments are carried out using the NIST SRE16 and SRE18 evaluation datasets. The results demonstrate the effectiveness of the GCNN and show that the proposed gated-attention statistics pooling can further improve the performance.
Lanhua You, Wu Guo, Li-Rong Dai 0001, Jun Du 0002
INTERSPEECH3
2019 Listening and Grouping: An Online Autoregressive Approach for Monaural Speech Separation
abstract
This paper proposes an autoregressive approach to harness the power of deep learning for multi-speaker monaural speech separation. It exploits a causal temporal context in both mixture and past estimated separated signals and performs online separation that is compatible with real-time applications. The approach adopts a learned listening and grouping architecture motivated by computational auditory scene analysis, with a grouping stage that effectively addresses the label permutation problem at both frame and segment levels. Experimental results on the WSJ0-2mix benchmark show that the new approach can achieve better signal-to-distortion ratio and perceptual evaluation of speech quality scores than most of the state-of-the-art methods for both closed-set and open-set evaluations, even methods that exploit whole-utterance statistics for separation. It achieves this while requiring fewer model parameters.
Zengxi Li, Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 Sequence-to-Sequence Acoustic Modeling for Voice Conversion
abstract
In this paper, a neural network named sequence-to-sequence ConvErsion NeTwork (SCENT) is presented for acoustic modeling in voice conversion. At training stage, a SCENT model is estimated by aligning the feature sequences of source and target speakers implicitly using attention mechanism. At the conversion stage, acoustic features and durations of source utterances are converted simultaneously using the unified acoustic model. Mel-scale spectrograms are adopted as acoustic features, which contain both excitation and vocal tract descriptions of speech signals. The bottleneck features extracted from source speech using an automatic speech recognition model are appended as an auxiliary input. A WaveNet vocoder conditioned on Mel-spectrograms is built to reconstruct waveforms from the outputs of the SCENT model. It is worth noting that our proposed method can achieve appropriate duration conversion, which is difficult in conventional methods. Experimental results show that our proposed method obtained better objective and subjective performance than the baseline methods using Gaussian mixture models and deep neural networks as acoustic models. This proposed method also outperformed our previous work, which achieved the top rank in Voice Conversion Challenge 2018. Ablation tests further confirmed the effectiveness of several components in our proposed method.
Jing-Xuan Zhang, Zhen-Hua Ling, Li-Juan Liu, Yuan Jiang 0006, Li-Rong Dai 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2019 Track, Attend, and Parse (TAP): An End-to-End Framework for Online Handwritten Mathematical Expression Recognition
abstract
In this paper, we introduce Track, Attend, and Parse (TAP), an end-to-end approach based on neural networks for online handwritten mathematical expression recognition (OHMER). The architecture of TAP consists of a tracker and a parser. The tracker employs a stack of bidirectional recurrent neural networks with gated recurrent units (GRU) to model the input handwritten traces, which can fully utilize the dynamic trajectory information in OHMER. Followed by the tracker, the parser adopts a GRU equipped with guided hybrid attention (GHA) to generate notations. The proposed GHA is composed of a coverage-based spatial attention, a temporal attention, and an attention guider. Moreover, we demonstrate the strong complementarity between offline information with static-image input and online information with ink-trajectory input by blending a fully convolutional networks-based watcher into TAP. Inherently, unlike traditional methods, this end-to-end framework does not require the explicit symbol segmentation and a predefined expression grammar for parsing. Validated on a benchmark published by the CROHME competition, the proposed approach outperforms the state-of-the-art methods and achieves the best reported results with an expression recognition accuracy of 61.16% on CROHME 2014 and 57.02% on CROHME 2016, using only official training dataset.
Jianshu Zhang 0001, Jun Du 0002, Li-Rong Dai 0001
IEEE Trans. Multim.3
2018 Pseudo-Supervised Approach for Text Clustering Based on Consensus Analysis
abstract
In recent years, neural networks (NN) have achieved remarkable performance improvement in text classification due to their powerful ability to encode discriminative features by incorporating label information into model training. Inspired by the success of NN in text classification, we propose a pseudo-supervised neural network approach for text clustering. The neural network is trained in a supervised fashion with pseudo-labels, which are provided by the cluster labels of pre-clustering on unsupervised document representations. To enhance the quality of pseudo-labels, a consensus analysis is employed to select training samples for the neural network. The experimental results demonstrate that the proposed approach can improve the clustering performance significantly.
Peixin Chen, Wu Guo, Li-Rong Dai 0001, Zhen-Hua Ling
ICASSP3
2018 Densely Connected Progressive Learning for LSTM-Based Speech Enhancement
abstract
Recently, we proposed a novel progressive learning (PL) framework for deep neural network (DNN) based speech enhancement to improve the performance in low signal-to-noise ratio (SNR) environments. In this study, several new contributions are made to this framework. First, the advanced long short-term memory (LSTM) architecture is adopted to achieve better results, namely LSTM-PL, where each LSTM layer is guided to explicitly learn an intermediate target with a specific SNR gain. However, we observe that the performance of LSTM-PL architecture is easily degraded by increasing the number of intermediate targets due to the possible information loss when involving more target layers. Accordingly, we propose densely connected progressive learning in which the input and the estimations of intermediate targets are spliced together to learn the next target. This new structure can fully utilize the rich set of information from the multiple learning targets and alleviate the information loss problem. Experimental results demonstrate that the dense structure with deeper LSTM layers can yield significant gains of speech intelligibility measure for all noise types and levels. Moreover, the post-processing with more targets tends to achieve better performance.
Tian Gao 0005, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001
ICASSP3
2018 Source-Aware Context Network for Single-Channel Multi-Speaker Speech Separation
abstract
Deep learning based approaches have achieved promising performance in speaker-dependent single-channel multispeaker speech separation. However, partly due to the label permutation problem, they may encounter difficulties in speaker-independent conditions. Recent methods address this problem by some assignment operations. Different from them, we propose a novel source-aware context network, which explicitly inputs speech sources as well as mixture signal. By exploiting the temporal dependency and continuity of the same source signal, the permutation order of outputs can be easily determined without any additional post-processing. Furthermore, a Multi-time-step Prediction Training strategy is proposed to address the mismatch between training and inference stages. Experimental results on benchmark WSJ0-2mix dataset revealed that our network achieved comparable or better results than state-of-the-art methods in both closed-set and open-set conditions, in terms of Signal-to-Distortion Ratio (SDR) improvement.
Zengxi Li, Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001
ICASSP3
2018 Forward Attention in Sequence- To-Sequence Acoustic Modeling for Speech Synthesis
abstract
This paper proposes a forward attention method for the sequence-to-sequence acoustic modeling of speech synthesis. This method is motivated by the nature of the monotonic alignment from phone sequences to acoustic sequences. Only the alignment paths that satisfy the monotonic condition are taken into consideration at each decoder timestep. The modified attention probabilities at each timestep are computed recursively using a forward algorithm. A transition agent for forward attention is further proposed, which helps the attention mechanism to make decisions whether to move forward or stay at each decoder timestep. Experimental results show that the proposed forward attention method achieves faster convergence speed and higher stability than the baseline attention method. Besides, the method of forward attention with transition agent can also help improve the naturalness of synthetic speech and control the speed of synthetic speech effectively.
Jing-Xuan Zhang, Zhen-Hua Ling, Li-Rong Dai 0001
ICASSP3
2018 Deep-FSMN for Large Vocabulary Continuous Speech Recognition
abstract
In this paper, we present an improved feedforward sequential memory networks (FSMN) architecture, namely Deep-FSMN (DFSMN), by introducing skip connections between memory blocks in adjacent layers. These skip connections enable the information flow across different layers and thus alleviate the gradient vanishing problem when building very deep structure. As a result, DFSMN significantly benefits from these skip connections and deep structure. We have compared the performance of DFSMN to BLSTM both with and without lower frame rate (LFR) on several large speech recognition tasks, including English and Mandarin. Experimental results shown that DFSMN can consistently outperform BLSTM with dramatic gain, especially trained with LFR using CD-Phone as modeling units. In the 20000 hours Fisher (FSH) task, the proposed DFSMN can achieve a word error rate of 9.4% by purely using the cross-entropy criterion and decoding with a 3-gram language model, which achieves a 1.5% absolute improvement compared to the BLSTM. In a 20000 hours Mandarin recognition task, the LFR trained DFSMN can achieve more than 20% relative improvement compared to the LFR trained BLSTM. Moreover, we can easily design the lookahead filter order of the memory blocks in DFSMN to control the latency for real-time applications.
Shiliang Zhang, Zhijie Yan, Li-Rong Dai 0001
ICASSP4
2018 Radical Analysis Network for Zero-Shot Learning in Printed Chinese Character Recognition
abstract
Chinese characters have a huge set of character categories, more than 20, 000 and the number is still increasing as more and more novel characters continue being created. However, the enormous characters can be decomposed into a compact set of about 500 fundamental and structural radicals. This paper introduces a novel radical analysis network (RAN) to recognize printed Chinese characters by identifying radicals and analyzing two-dimensional spatial structures among them. The proposed RAN first extracts visual features from input by employing convolutional neural networks as an encoder. Then a decoder based on recurrent neural networks is employed, aiming at generating captions of Chinese characters by detecting radicals and two-dimensional structures through a spatial attention mechanism. The manner of treating a Chinese character as a composition of radicals rather than a single character class largely reduces the size of vocabulary and enables RAN to possess the ability of recognizing unseen Chinese character classes, namely zero-shot learning.
Jianshu Zhang 0001, Yixing Zhu, Jun Du 0002, Li-Rong Dai 0001
ICME4
2018 Multi-Scale Attention with Dense Encoder for Handwritten Mathematical Expression Recognition
abstract
Handwritten mathematical expression recognition is a challenging problem due to the complicated two-dimensional structures, ambiguous handwriting input and variant scales of handwritten math symbols. To settle this problem, recently we propose the attention based encoder-decoder model that recognizes mathematical expression images from two-dimensional layouts to one-dimensional LaTeX strings. In this study, we improve the encoder by employing densely connected convolutional networks as they can strengthen feature extraction and facilitate gradient propagation especially on a small training set. We also present a novel multi-scale attention model which is employed to deal with the recognition of math symbols in different scales and restore the fine-grained details dropped by pooling operations. Validated on the CROHME competition task, the proposed method significantly outperforms the state-of-the-art methods with an expression recognition accuracy of 52.8% on CROHME 2014 and 50.1% on CROHME 2016, by only using the official training dataset.
Jianshu Zhang 0001, Jun Du 0002, Li-Rong Dai 0001
ICPR3
2018 Trajectory-based Radical Analysis Network for Online Handwritten Chinese Character Recognition
abstract
Recently, great progress has been made for online handwritten Chinese character recognition due to the emergence of deep learning techniques. However, previous research mostly treated each Chinese character as one class without explicitly considering its inherent structure, namely the radical components with complicated geometry. In this study, we propose a novel trajectory-based radical analysis network (TRAN) to firstly identify radicals and analyze two-dimensional structures among radicals simultaneously, then recognize Chinese characters by generating captions of them based on the analysis of their internal radicals. The proposed TRAN employs recurrent neural networks (RNNs) as both an encoder and a decoder. The RNN encoder makes full use of online information by directly transforming handwriting trajectory into high-level features. The RNN decoder aims at generating the caption by detecting radicals and spatial structures through an attention model. The manner of treating a Chinese character as a two-dimensional composition of radicals can reduce the size of vocabulary and enable TRAN to possess the capability of recognizing unseen Chinese character classes, only if the corresponding radicals have been seen. Evaluated on CASIA-OLHWDB database, the proposed approach significantly outperforms the state-of-the-art whole-character modeling approach with a relative character error rate (CER) reduction of 10%. Meanwhile, for the case of recognition of 500 unseen Chinese characters, TRAN can achieve a character accuracy of about 60 % while the traditional whole-character method has no capability to handle them.
Jianshu Zhang 0001, Yixing Zhu, Jun Du 0002, Li-Rong Dai 0001
ICPR4
2018 An Improved Deep Embedding Learning Method for Short Duration Speaker Verification
abstract
This paper presents an improved deep embedding learning method based on convolutional neural networks (CNN) for short-duration speaker verification (SV). Existing deep learning-based SV methods generally extract frontend embeddings from a feed-forward deep neural network, in which the long-term speaker characteristics are captured via a pooling operation over the input speech. The extracted embeddings are then scored via a backend model, such as Probabilistic Linear Discriminative Analysis (PLDA). Two improvements are proposed for frontend embedding learning based on the CNN structure: (1) Motivated by the WaveNet for speech synthesis, dilated filters are designed to achieve a tradeoff between computational efficiency and receptive-filter size; and (2) A novel cross-convolutional-layer pooling method is exploited to capture $1^{st}$-order statistics for modelling long-term speaker characteristics. Specifically, the activations of one convolutional layer are aggregated with the guidance of the feature maps from the successive layer. To evaluate the effectiveness of our proposed methods, extensive experiments are conducted on the modified female portion of NIST SRE 2010 evaluations, with conditions ranging from 10s-10s to 5s-4s. Excellent performance has been achieved on each evaluation condition, significantly outperforming existing SV systems using i-vector and d-vector embeddings.
Zhifu Gao, Yan Song 0001, Ian McLoughlin 0001, Wu Guo, Li-Rong Dai 0001
INTERSPEECH5
2018 An Attention Pooling Based Representation Learning Method for Speech Emotion Recognition
abstract
This paper proposes an attention pooling based representation learning method for speech emotion recognition (SER). The emotional representation is learned in an end-to-end fashion by applying a deep convolutional neural network (CNN) directly to spectrograms extracted from speech utterances. Motivated by the success of GoogleNet, two groups of filters with different shapes are designed to capture both temporal and frequency domain context information from the input spectrogram. The learned features are concatenated and fed into the subsequent convolutional layers. To learn the final emotional representation, a novel attention pooling method is further proposed. Compared with the existing pooling methods, such as max-pooling and average-pooling, the proposed attention pooling can effectively incorporate class-agnostic bottom-up, and class-specific top-down, attention maps. We conduct extensive evaluations on benchmark IEMOCAP data to assess the effectiveness of the proposed representation. Results demonstrate a recognition performance of 71.8% weighted accuracy (WA) and 68% unweighted accuracy (UA) over four emotions, which outperforms the state-of-the-art method by about 3% absolute for WA and 4% for UA.
Yan Song 0001, Ian McLoughlin 0001, Wu Guo, Li-Rong Dai 0001
INTERSPEECH5
2018 WaveNet Vocoder with Limited Training Data for Voice Conversion
Li-Juan Liu, Zhen-Hua Ling, Yuan Jiang 0006, Li-Rong Dai 0001
INTERSPEECH5
2018 Acoustic Modeling with Densely Connected Residual Network for Multichannel Speech Recognition
abstract
Motivated by recent advances in computer vision research, this paper proposes a novel acoustic model called Densely Connected Residual Network (DenseRNet) for multichannel speech recognition. This combines the strength of both DenseNet and ResNet. It adopts the basic "building blocks" of ResNet with different convolutional layers, receptive field sizes and growth rates as basic components that are densely connected to form so-called denseR blocks. By concatenating the feature maps of all preceding layers as inputs, DenseRNet can not only strengthen gradient back-propagation for the vanishing-gradient problem, but also exploit multi-resolution feature maps. Preliminary experimental results on CHiME-3 have shown that DenseRNet achieves a word error rate (WER) of 7.58% on beamforming-enhanced speech with six channel real test data by cross entropy criteria training while WER is 10.23% for the official baseline. Besides, additional experimental results are also presented to demonstrate that DenseRNet exhibits the robustness to beamforming-enhanced speech as well as near and far-field speech.
Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001
INTERSPEECH3
2018 Learning and Modeling Unit Embeddings for Improving HMM-based Unit Selection Speech Synthesis
Xiao Zhou 0024, Zhen-Hua Ling, Zhi-Ping Zhou, Li-Rong Dai 0001
INTERSPEECH4
2018 Articulatory-to-acoustic conversion using BLSTM-RNNs with augmented input representation
Zheng-Chen Liu, Zhen-Hua Ling, Li-Rong Dai 0001
Speech Commun.3
2018 Statistical Parametric Speech Synthesis Using Generalized Distillation Framework
abstract
This letter proposes an improved statistical parametric speech synthesis (SPSS) method which utilizes auxiliary information for acoustic modeling under generalized distillation framework. In conventional SPSS, acoustic models are trained using context features as input and acoustic features as output. In our proposed method, two acoustic models, so-called teacher and student, are involved. Both of them are recurrent neural networks (RNN) with bidirectional long short-term memory (BLSTM) units. The teacher, which aims to provide the student with additional knowledge, employs auxiliary features (e.g., articulatory features or spectra of short-time Fourier transform) in addition to the conventional input and output of acoustic models. The student, which serves as the final acoustic model for synthesis, adopts a multitask learning architecture which uses the outcome of the teacher as the target of its secondary task. Experimental results show that this method can achieve better accuracy of acoustic feature prediction and produce more natural synthetic speech than conventional BLSTM-RNN-based acoustic modeling with single-task or multitask learning.
Zheng-Chen Liu, Zhen-Hua Ling, Li-Rong Dai 0001
IEEE Signal Process. Lett.3
2018 LID-Senones and Their Statistics for Language Identification
abstract
Recent research on end-to-end training structures for language identification has raised the possibility that intermediate language-sensitive feature units exist which are analogous to phonetically sensitive senones in automatic speech recognition systems. Termed language identification (LID)-senones, the statistics derived from these feature units have been shown to be beneficial in discriminating between languages, particularly for short utterances. This paper examines the evidence for the existence of LID-senones before designing and evaluating LID systems based on low- and high-level statistics of LID-senones with both generative and discriminative models. For the standard NIST LRE 2009 task on 23 languages, LID-senone-based systems are shown to outperform state-of-the-art deep neural network/i-vector methods both when LID-senones are used directly for classification and when LID-senone statistics are used for i-vector formation.
Ma Jin, Yan Song 0001, Ian McLoughlin 0001, Li-Rong Dai 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2018 Waveform Modeling and Generation Using Hierarchical Recurrent Neural Networks for Speech Bandwidth Extension
abstract
This paper presents a waveform modeling and generation method using hierarchical recurrent neural networks (HRNN) for speech bandwidth extension (BWE). Different from conventional BWE methods that predict spectral parameters for reconstructing wideband speech waveforms, this BWE method models and predicts waveform samples directly without using vocoders. Inspired by SampleRNN, which is an unconditional neural audio generator, the HRNN model represents the distribution of each wideband or high-frequency waveform sample conditioned on the input narrowband waveform samples using a neural network composed of long short-term memory (LSTM) layers and feed-forward layers. The LSTM layers form a hierarchical structure and each layer operates at a specific temporal resolution to efficiently capture long-span dependencies between temporal sequences. Furthermore, additional conditions, such as the bottleneck features derived from narrowband speech using a deep neural network based state classifier, are employed as auxiliary input to further improve the quality of generated wideband speech. The experimental results of comparing several waveform modeling methods show that the HRNN-based method can achieve better speech quality and run-time efficiency than the dilated convolutional neural network based method and the plain sample-level recurrent neural network based method. Our proposed method also outperforms the conventional vocoder-based BWE method using LSTM-RNNs in terms of the subjective quality of the reconstructed wideband speech.
Zhen-Hua Ling, Yang Ai, Li-Rong Dai 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2018 A Multiobjective Learning and Ensembling Approach to High-Performance Speech Enhancement With Compact Neural Network Architectures
abstract
In this study, we propose a novel deep neural network (DNN) architecture for speech enhancement (SE) via a multiobjective learning and ensembling (MOLE) framework to achieve a compact and lowlatency design, while maintaining good performance in quality evaluations. MOLE follows the boosting concept when combining weak models into a strong classifier and consists of two compact DNNs. The first, called the multiobjective learning DNN (MOL-DNN), takes multiple features, such as log-power spectra (LPS), mel-frequency cepstral coefficients (MFCCs) and Gammatone frequency cepstral coefficients (GFCCs) to predict a multiobjective set that includes clean speech feature, dynamic noise feature, and ideal ratio mask (IRM). The second, called the multiobjective ensembling DNN (MOE-DNN), takes the learned features from MOL-DNN as inputs and separately predicts clean LPS and IRM, clean MFCC and IRM, and clean GFCC and IRM using three sets of weak regression functions. Finally, a postprocessing operation can be applied to the estimated clean features by leveraging the multiple targets learned from both the MOL-DNN and the MOE-DNN. On speech corrupted by 15 noise types not seen in model training the SE results show that the MOLE approach, which features a small model size and low run-time latency, can achieve consistent improvements over both DNN- and long short-term memory (LSTM)-based techniques in terms of all the objective metrics evaluated in this study for all three cases (the input contexts contain 1-frame, 4-frame and 7-frame instances). The 1-frame MOLE-based SE system outperforms the DNN-based SE system with a 7-frame input expansion at a 3-frame delay and also achieves better performance than the LSTM-based SE system with 4-frame, no delay expansion by including only 3 previous frames, and with 170 times less processing latency.
Qing Wang 0008, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2017 The USTC system for blizzard machine learning challenge 2017-ES2
abstract
The Blizzard Machine Learning Challenge (BMLC) aims to liberate participants from speech-specific processing when building speech synthesis systems. This paper describes the USTC system for the ES2 sub-task in BMLC2017, which requires participants to train a model to directly predict waveforms from linguistic features. We investigate three aspects of waveform modeling when preparing our system for this task. First, two different model structures for waveform modeling, i.e., WaveNet and SampleRNN, are compared on this task. Second, a strategy of using features extracted from waveforms as intermediate representations for waveform modeling is studied. Experimental results show that using low-level features (STFT amplitude spectra) as intermediate representations can achieve similar performance as using high-level features (mel-cepstra and F0). Third, the feasibility of applying WaveNet to wideband speech signals with more than 256 quantization levels is verified by experiments. Finally, a system which adopts STFT amplitude spectra as intermediate representations to model 24kHz speech waveforms with 1024 mu-law quantization levels is submitted for evaluation. The evaluation results of BMLC2017 demonstrate the effectiveness of our proposed methods.
Ya-Jun Hu, Li-Juan Liu, Chuang Ding, Zhen-Hua Ling, Li-Rong Dai 0001
ASRU5
2017 Adaptation of PLDA for multi-source text-independent speaker verification
abstract
Probabilistic linear discriminant analysis (PLDA) is widely described as an effective model for text-independent speaker verification in the i-vector space. The PLDA scoring function is typically formulated as the likelihood ratio between the speaker-adapted and the universal PLDAs. In this case, the adaptation of PLDA was performed through the speaker factors. In this paper, we show that the channel factors of the PLDA could be equivalently exploited to deal with the multi-source conditions. In speaker verification, with the proposed method, a PLDAmodel trained on conversational telephone speech could be adequately adapted for interview-style microphone recordings. Experimental results on NIST SRE'08 and SRE'10 datasets confirm that the proposed method is effective, especially for the case whereby enrollment and test utterances were captured from different sources.
Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001, Li-Rong Dai 0001
ICASSP6
2017 Extracting structural spectral features using what-where auto-encoders for statistical parametric speech synthesis
abstract
This paper presents a method to extract structural spectral features from spectral envelopes using what-where autoencoders (WWAE) for statistical parametric speech synthesis (SPSS). A WWAE is constructed by concatenating a convolutional net for input encoding and a deconvolutional net for reconstruction. The output values of the max-pooling layer in the encoder and the positions of the max-pooling switches are utilized as the what and where features respectively. Considering the intrinsic formant structures in the spectral envelopes of voiced speech frames, the WWAE model is adopted in this paper to detect, locate, and reconstruct the formants and other local structures in spectral envelopes. Here, the what and where features describe the prominences and positions of specific local spectral structures within a pooling frequency window. Then, the extracted what and where features are modeled as separate streams under the hidden Markov model (HMM)-based SPSS framework. Experimental results show that the speech synthesis system built using our proposed spectral features can produce synthetic speech with sharper formant structures and better naturalness than the systems using mel-cepstra and conventional auto-encoder-based spectral features.
Ya-Jun Hu, Zhen-Hua Ling, Li-Rong Dai 0001
ICASSP3
2017 A GRU-Based Encoder-Decoder Approach with Attention for Online Handwritten Mathematical Expression Recognition
abstract
In this study, we present a novel end-to-end approach based on the encoder-decoder framework with the attention mechanism for online handwritten mathematical expression recognition (OHMER). First, the input two-dimensional ink trajectory information of handwritten expression is encoded via the gated recurrent unit based recurrent neural network (GRU-RNN). Then the decoder is also implemented by the GRU-RNN with a coverage-based attention model. The proposed approach can simultaneously accomplish the symbol recognition and structural analysis to output a character sequence in LaTeX format. Validated on the CROHME 2014 competition task, our approach significantly outperforms the state-of-the-art with an expression recognition accuracy of 52.43% by only using the official training dataset. Furthermore, the alignments between the input trajectories of handwritten expressions and the output LaTeX sequences are visualized by the attention mechanism to show the effectiveness of the proposed method.
Jianshu Zhang 0001, Jun Du 0002, Li-Rong Dai 0001
ICDAR3
2017 An investigation of high-resolution modeling units of deep neural networks for acoustic scene classification
abstract
In this paper, we investigate high-resolution modeling units of deep neural networks (DNNs) from concrete to abstract for acoustic scene classification based on Gaussian mixture model (GMM) and ergodic hidden Markov model (HMM). A direct modeling strategy for DNN to classify acoustic scenes is to map each frame feature of an audio to one scene category. However, all frames tagged with the same label may not be the best choice because the representative pattern of an audio is sparse. GMM is also often employed to model each acoustic scene directly as a generative model. Because the multiple Gaussians in a GMM model have different levels of contribution, and each Gaussian can be seen as a subclass of the scene category, so we can utilize the subclass of GMM as a bit abstract modeling unit to adopt DNN-GMM system. When single scene category is subdivided into various subclasses, prior scores for each subclass calculated from training set are stored as one part of model to response the sparseness of representative pattern. Ergodic HMM should be more appropriate to model the acoustic scenes than GMM due to the uncertain structure of scene audio. Using HMM states as modeling units, we build DNN-HMM hybrid system. By comparison, we find high-resolution modeling units are more effective than direct modeling. The final system is obtained by performing system combination to take advantage of the complementarity of different-level modeling units. Experiments on acoustic scene classification task of DCASE2016 challenge show that our final system yields 25.9% relative error rate reduction compared with a GMM baseline on evaluation set.
Xiao Bao, Tian Gao 0005, Jun Du 0002, Li-Rong Dai 0001
IJCNN4
2017 Gaussian Prediction Based Attention for Online End-to-End Speech Recognition
Junfeng Hou, Shiliang Zhang, Li-Rong Dai 0001
INTERSPEECH3
2017 End-to-End Language Identification Using High-Order Utterance Representation with Bilinear Pooling
abstract
A key problem in spoken language identification (LID) is how to design effective representations which are specific to language information. Recent advances in deep neural networks have led to significant improvements in results, with deep end-to-end methods proving effective. This paper proposes a novel network which aims to model an effective representation for high (first and second)-order statistics of LID-senones, defined as being LID analogues of senones in speech recognition. The high-order information extracted through bilinear pooling is robust to speakers, channels and background noise. Evaluation with NIST LRE 2009 shows improved performance compared to current state-of-the-art DBF/i-vector systems, achieving over 33% and 20% relative equal error rate (EER) improvement for 3s and 10s utterances and over 40% relative Cavg improvement for all durations.
Ma Jin, Yan Song 0001, Ian McLoughlin 0001, Wu Guo, Li-Rong Dai 0001
INTERSPEECH5
2017 A Maximum Likelihood Approach to Deep Neural Network Based Nonlinear Spectral Mapping for Single-Channel Speech Separation
Yannan Wang, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001
INTERSPEECH3
2017 An information fusion framework with multi-channel feature concatenation and multi-perspective system combination for the deep-learning-based robust recognition of microphone array speech
Yanhui Tu, Jun Du 0002, Qing Wang 0008, Xiao Bao, Li-Rong Dai 0001, Chin-Hui Lee 0001
Comput. Speech Lang.5
2017 Towards human-like and transhuman perception in AI 2.0: a review
abstract
Perception is the interaction interface between an intelligent system and the real world. Without sophisticated and flexible perceptual capabilities, it is impossible to create advanced artificial intelligence (AI) systems. For the next-generation AI, called ‘AI 2.0’, one of the most significant features will be that AI is empowered with intelligent perceptual capabilities, which can simulate human brain’s mechanisms and are likely to surpass human brain in terms of performance. In this paper, we briefly review the state-of-the-art advances across different areas of perception, including visual perception, auditory perception, speech perception, and perceptual information processing and learning engines. On this basis, we envision several R&D trends in intelligent perception for the forthcoming era of AI 2.0, including: (1) human-like and transhuman active vision; (2) auditory perception and computation in an actual auditory setting; (3) speech perception and computation in a natural interaction setting; (4) autonomous learning of perceptual information; (5) large-scale perceptual information processing and learning platforms; and (6) urban omnidirectional intelligent perception and reasoning engines. We believe these research directions should be highlighted in the future plans for AI 2.0.
Yonghong Tian 0001, Xilin Chen 0001, Hongkai Xiong, Li-Rong Dai 0001, Jing Chen 0002, Junliang Xing, Jing Chen 0003, Xihong Wu, Weiming Hu 0004, Yu Hu 0003, Tiejun Huang 0001, Wen Gao 0001
Frontiers Inf. Technol. Electron. Eng.5
2017 Watch, attend and parse: An end-to-end neural network based approach to handwritten mathematical expression recognition
Jianshu Zhang 0001, Jun Du 0002, Shiliang Zhang, Dan Liu 0008, Yulong Hu, Jin-Shui Hu, Si Wei, Li-Rong Dai 0001
Pattern Recognit.8
2017 A unified DNN approach to speaker-dependent simultaneous speech enhancement and speech separation in low SNR environments
Tian Gao 0005, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001
Speech Commun.3
2017 A Gender Mixture Detection Approach to Unsupervised Single-Channel Speech Separation Based on Deep Neural Networks
abstract
We propose an unsupervised speech separation framework for mixtures of two unseen speakers in a single-channel setting based on deep neural networks (DNNs). We rely on a key assumption that two speakers could be well segregated if they are not too similar to each other. A dissimilarity measure between two speakers is first proposed to characterize the separation ability between competing speakers. We then show that speakers with the same or different genders can often be separated if two speaker clusters, with large enough distances between them, for each gender group could be established, resulting in four speaker clusters. Next, a DNN-based gender mixture detection algorithm is proposed to determine whether the two speakers in the mixture are females, males, or from different genders. This detector is based on a newly proposed DNN architecture with four outputs, two of them representing the female speaker clusters and the other two characterizing the male groups. Finally, we propose to construct three independent speech separation DNN systems, one for each of the female-female, male-male, and female-male mixture situations. Each DNN gives dual outputs, one representing the target speaker group and the other characterizing the interfering speaker cluster. Trained and tested on the speech separation challenge corpus, our experimental results indicate that the proposed DNN-based approach achieves large performance gains over the state-of-the-art unsupervised techniques without using any specific knowledge about the mixed target and interfering speakers being segregated.
Yannan Wang, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2017 Nonrecurrent Neural Structure for Long-Term Dependence
abstract
In this paper, we propose a novel neural network structure, namely feedforward sequential memory networks (FSMN), to model long-term dependence in time series without using recurrent feedback. The proposed FSMN is a standard fully connected feedforward neural network equipped with some learnable memory blocks in its hidden layers. The memory blocks use a tapped-delay line structure to encode the long context information into a fixed-size representation as short-term memory mechanism which are somehow similar to the time-delay neural networks layers. We have evaluated the FSMNs in several standard benchmark tasks, including speech recognition and language modeling. Experimental results have shown that FSMNs outperform the conventional recurrent neural networks (RNN) while can be learned much more reliably and faster in modeling sequential signals like speech or language. Moreover, we also propose a compact feedforward sequential memory networks (cFSMN) by combining FSMN with low-rank matrix factorization and make a slight modification to the encoding method used in FSMNs in order to further simplify the network architecture. On the speech recognition Switchboard task, the proposed cFSMN structures can reduce the model size by 60% and speed up the learning by more than seven times while the model can still significantly outperform the popular bidirectional LSTMs for both frame-level cross-entropy criterion-based training and MMI-based sequence training.
Shiliang Zhang, Cong Liu 0006, Hui Jiang 0001, Si Wei, Li-Rong Dai 0001, Yu Hu 0003
IEEE ACM Trans. Audio Speech Lang. Process.5
2016 Content-aware local variability vector for speaker verification with short utterance
abstract
I-vector has shown to be very effective in speaker verification with long-duration speech utterances. But when test utterances are of short duration, content mismatch between the enrollment and test utterances limit the performance of i-vector system. This paper proposes to extract local session variability vectors on different phonetic classes from the utterances instead of estimating the session variability across the whole utterance as i-vector does. Using the posteriors given by a deep neural network (DNN) trained for phone state classification, the local vectors represent the session variability contained in specific phonetic content. Our experiments show that the content-aware local vectors are better at coping with the content mismatch between training and test utterances of short durations for text-independent, text-constrained and text-dependent tasks.
Kong-Aik Lee, Chng Eng Siong, Bin Ma 0001, Haizhou Li 0001, Li-Rong Dai 0001
ICASSP6
2016 Deep belief network-based post-filtering for statistical parametric speech synthesis
abstract
The speech synthesized by statistical parametric speech synthesis (SPSS) always sounds muffled. One important reason is that the generated spectral envelopes are over-smoothed and many detailed spectral structures in natural speech are lost. This paper presents a deep belief network (DBN)-based post-filtering method for hidden Markov model (HMM)-based SPSS to address this issue. At training time, a DBN is estimated using the spectral envelopes extracted from natural speech. This DBN serves as a generatively trained postfilter which processes the spectral envelopes recovered from the predicted spectral features at synthesis time. Experimental results show that the effectiveness of this method depends on the sampling strategy used to generate the training data of the restricted Boltzmann machines (RBM) which forms the higher layers of the DBN. When binary samples are adopted instead of mean-filed approximation, the DBN post-filter can alleviate the over-smoothing effect of parameter generation and improve the naturalness of synthetic speech significantly when either mel-cepstra or line spectral pairs (LSP) are used as spectral features. Its performance is comparative with the parameter generation method with global variance (GV) modeling for mel-cepstra and better than the LSP-based formant enhancement method used in previous work.
Ya-Jun Hu, Zhen-Hua Ling, Li-Rong Dai 0001
ICASSP3
2016 Speaker adaptation OF RNN-BLSTM for speech recognition based on speaker code
abstract
Recently, recurrent neural network with bidirectional Long Short-Term Memory (RNN-BLSTM) acoustic model has been shown to give great performance on the TIMIT [1] and other speech recognition tasks. Meanwhile, the speaker code based adaptation method has been demonstrated as a valid adaptation method for Deep Neural Network (DNN) acoustic model [2]. However, whether the speaker code based adaptation method is also valid for RNN-BLSTM has not been reported to the best our knowledge. In this paper, we study how to conduct effective speaker code based speaker adaptation on RNN-BLSTM and demonstrate that the speaker code based adaptation method is also a valid adaptation method for RNN-BLSTM. Experimental results on TIMIT have shown that the adaptation of RNN-LSTM can achieve over 10% relative reduction in phone error rate (PER) compared to without adaptation. Then, a set of comparative experiments are implemented to analyze the different contribution of the adaptation on cell input and each gate activation function of the BLSTM. It's found that the adaptation on cell input activation function is more effective than the adaptation on each gate activation function.
Zhiying Huang, Shaofei Xue, Li-Rong Dai 0001
ICASSP4
2016 Compact convolutional neural network transfer learning for small-scale image classification
abstract
Transfer learning methods have demonstrated state-of-the-art performance on various small-scale image classification tasks. This is generally achieved by exploiting the information from an ImageNet convolution neural network (ImageNet CNN). However, the transferred CNN model is generally with high computational complexity and storage requirement. It raises the issue for real-world applications, especially for some portable devices like phones and tablets without high-performance GPUs. Several approximation methods have been proposed to reduce the complexity by reconstructing the linear or non-linear filters (responses) in convolutional layers with a series of small ones., In this paper, we present a compact CNN transfer learning method for small-scale image classification. Specifically, it can be decomposed into fine-tuning and joint learning stages. In fine-tuning stage, a high-performance target CNN is trained by transferring information from the ImageNet CNN. In joint learning stage, a compact target CNN is optimized based on ground-truth labels, jointly with the predictions of the high-performance target CNN. The experimental results on CIFAR-10 and MIT Indoor Scene demonstrate the effectiveness and efficiency of our proposed method.
Zengxi Li, Yan Song 0001, Ian McLoughlin 0001, Li-Rong Dai 0001
ICASSP4
2016 Modulation spectrum compensation for HMM-based speech synthesis using line spectral pairs
abstract
In previous work, a method to compensate the divergence between the distributions of natural and generated modulation spectra (MS) has been proposed for hidden Markov model (HMM) based speech synthesis. This method can alleviate the over-smoothing effect of parameter generation when Mel-cepstral coefficients (MCC) are used as spectral features. This paper further investigates the MS compensation method for line spectral pairs (LSP). Four approaches to extract MS from LSPs are implemented and compared. These approaches calculate MS vectors using original LSP sequences, log power spectra (LPS) derived from LSPs, MCCs derived from LSPs, and MCCs derived from speech waveforms, respectively. Experimental results show that the naturalness of synthetic speech gets improved after MS compensation when LSPs are used as spectral features for HMM modeling. The degree of improvement depends on the type of spectral features for MS calculation significantly. MCCs derived from LSPs are more suitable for MS compensation than original LSPs and LPS derived from LSPs. Besides, using MCCs derived from speech waveforms also achieves satisfactory performance. This means that MS compensation can also be implemented as a post-filter to synthetic waveforms which does not rely on the type of spectral features and vocoders adopted in the synthesis system.
Zhen-Hua Ling, Xiao-Hui Sun, Li-Rong Dai 0001, Yu Hu 0003
ICASSP3
2016 Modeling spectral envelopes using deep conditional restricted Boltzmann machines for statistical parametric speech synthesis
abstract
This paper proposes a spectral modeling method using a deep conditional restricted Boltzmann machine (DCRBM) for statistical parametric speech synthesis. In this method, a DCRBM, which combines a deep neural network (DNN) with a conditional restricted Boltzmann machine (CRBM), is utilized to describe the conditional distribution of spectral envelopes given linguistic features. Compared with DNN and deep mixture density network (DMDN), DCRBM is better at describing the multimodal distribution of high-dimensional acoustic features with cross-dimension correlations. At training stage, the DNN part and the CRBM part of the DCRBM are pre-trained successively and then a unified fine-tuning of all model parameters is conducted. At synthesis time, spectral envelopes are generated from the estimated DCRBM model by iterative sampling and dynamic-feature-constrained parameter generation given linguistic features of input text. Experimental results show that our proposed method can produce more natural speech sounds than the hidden Markov model (HMM)-based, DNN-based, and DMDN-based synthesis methods. This method also outperforms previous work which adopts restricted Boltzmann machines (RBM) to model the distributions of spectral envelopes at HMM states.
Xiang Yin 0002, Zhen-Hua Ling, Ya-Jun Hu, Li-Rong Dai 0001
ICASSP4
2016 The USTC System for Voice Conversion Challenge 2016: Neural Network Based Approaches for Spectrum, Aperiodicity and F0 Conversion
Linghui Chen, Li-Juan Liu, Zhen-Hua Ling, Yuan Jiang 0006, Li-Rong Dai 0001
INTERSPEECH5
2016 SNR-Based Progressive Learning of Deep Neural Network for Speech Enhancement
Tian Gao 0005, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001
INTERSPEECH3
2016 Speech Bandwidth Extension Using Bottleneck Features and Deep Recurrent Neural Networks
Zhen-Hua Ling, Li-Rong Dai 0001
INTERSPEECH3
2016 Articulatory-to-Acoustic Conversion with Cascaded Prediction of Spectral and Excitation Features Using Neural Networks
Zheng-Chen Liu, Zhen-Hua Ling, Li-Rong Dai 0001
INTERSPEECH3
2016 Future Context Attention for Unidirectional LSTM Based Acoustic Model
Shiliang Zhang, Si Wei, Li-Rong Dai 0001
INTERSPEECH4
2016 Compact Feedforward Sequential Memory Networks for Large Vocabulary Continuous Speech Recognition
Shiliang Zhang, Hui Jiang 0001, Shifu Xiong, Si Wei, Li-Rong Dai 0001
INTERSPEECH5
2016 RNN-BLSTM Based Multi-Pitch Estimation
Jianshu Zhang 0001, Li-Rong Dai 0001
INTERSPEECH3
2016 Image classification with CNN-based Fisher vector coding
abstract
Fisher vector coding methods have been demonstrated to be effective for image classification. With the help of convolutional neural networks (CNN), several Fisher vector coding methods have shown state-of-the-art performance by adopting the activations of a single fully-connected layer as region features. These methods generally exploit a diagonal Gaussian mixture model (GMM) to describe the generative process of region features. However, it is difficult to model the complex distribution of high-dimensional feature space with a limited number of Gaussians obtained by unsupervised learning. Simply increasing the number of Gaussians turns out to be inefficient and computationally impractical. To address this issue, we re-interpret a pre-trained CNN as the probabilistic discriminative model, and present a CNN based Fisher vector coding method, termed CNN-FVC. Specifically, activations of the intermediate fully-connected and output soft-max layers are exploited to derive the posteriors, mean and covariance parameters for Fisher vector coding implicitly. To further improve the efficiency, we convert the pre-trained CNN to a fully convolutional one to extract the region features. Extensive experiments have been conducted on two standard scene benchmarks (i.e. SUN397 and MIT67) to evaluate the effectiveness of the proposed method. Classification accuracies of 60.7% and 82.1% are achieved on the SUN397 and MIT67 benchmarks respectively, outperforming previous state-of-the-art approaches. Furthermore, the method is complementary to GMM-FVC methods, allowing a simple fusion scheme to further improve performance to 61.1% and 83.1% respectively.
Yan Song 0001, Xinhai Hong, Ian McLoughlin 0001, Li-Rong Dai 0001
VCIP4
2016 Concept-to-Speech generation with knowledge sharing for acoustic modelling and utterance filtering
Xin Wang 0037, Zhen-Hua Ling, Li-Rong Dai 0001
Comput. Speech Lang.3
2016 Hybrid Orthogonal Projection and Estimation (HOPE): A New Framework to Learn Neural Networks
abstract
In this paper, we propose a novel model for high-dimensional data, called the Hybrid Orthogonal Projection and Estimation (HOPE) model, which combines a linear orthogonal projection and a finite mixture model under a unified generative modeling framework. The HOPE model itself can be learned unsupervised from unlabelled data based on the maximum likelihood estimation as well as discriminatively from labelled data. More interestingly, we have shown the proposed HOPE models are closely related to neural networks (NNs) in a sense that each hidden layer can be reformulated as a HOPE model. As a result, the HOPE framework can be used as a novel tool to probe why and how NNs work, more importantly, to learn NNs in either supervised or unsupervised ways. In this work, we have investigated the HOPE framework to learn NNs for several standard tasks, including image recognition on MNIST and speech recognition on TIMIT. Experimental results have shown that the HOPE framework yields significant performance gains over the current state-of-the-art methods in various types of NN learning problems, including unsupervised feature learning, supervised or semi-supervised learning.
Shiliang Zhang, Hui Jiang 0001, Li-Rong Dai 0001
J. Mach. Learn. Res.3
2016 Modeling F0 trajectories in hierarchically structured deep neural networks
Xiang Yin 0002, Yao Qian, Frank K. Soong, Lei He 0005, Zhen-Hua Ling, Li-Rong Dai 0001
Speech Commun.7
2016 A Regression Approach to Single-Channel Speech Separation Via High-Resolution Deep Neural Networks
abstract
We propose a novel data-driven approach to single-channel speech separation based on deep neural networks (DNNs) to directly model the highly nonlinear relationship between speech features of a mixed signal containing a target speaker and other interfering speakers. We focus our discussion on a semisupervised mode to separate speech of the target speaker from an unknown interfering speaker, which is more flexible than the conventional supervised mode with known information of both the target and interfering speakers. Two key issues are investigated. First, we propose a DNN architecture with dual outputs of the features of both the target and interfering speakers, which is shown to achieve a better generalization capability than that with output features of only the target speaker. Second, we propose using a set of multiple DNNs, each intending to be signal-noise-dependent (SND), to cope with the difficulty that one single general DNN could not well accommodate all the speaker mixing variabilities at different signal-to-noise ratio (SNR) levels. Experimental results on the speech separation challenge (SSC) data demonstrate that our proposed framework achieves better separation results than other conventional approaches in a supervised or semisupervised mode. SND-DNNs could also yield significant performance improvements over a general DNN for speech separation in low SNR cases. Furthermore, for automatic speech recognition (ASR) following speech separation, this purely front-end processing with a single set of speaker-independent ASR acoustic models, achieves a relative word error rate (WER) reduction of 11.6% over a state-of-the-art separation and recognition system where a complicated joint back-end decoding framework with multiple sets of speaker-dependent ASR acoustic models needs to be implemented. When speaker-adaptive ASR acoustic models for the target speakers are adopted for the enhanced signals, another 12.1% WER reduction over our best speaker-independent ASR system is achieved.
Jun Du 0002, Yanhui Tu, Li-Rong Dai 0001, Chin-Hui Lee 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 An information fusion approach to recognizing microphone array speech in the CHiME-3 challenge based on a deep learning framework
abstract
We present an information fusion approach to robust recognition of microphone array speech for the recently launched 3rd CHiME Challenge. It is based on a deep learning framework with a large neural network consisting of subnets with different architectures. Multiple knowledge sources are integrated via an early fusion of normalized noisy features with different beamforming techniques, speech enhanced features, speaker related features, and other auxiliary features concatenated as the input to each subnet, and a late fusion by combining the outputs of all subnets to produce one single output set. Our experiments demonstrate that all information sources are complementary in our proposed framework. Our best system achieves an average word error rate reduction of 68% from the officially released baseline results on the test set of real data.
Jun Du 0002, Qing Wang 0008, Yanhui Tu, Xiao Bao, Li-Rong Dai 0001, Chin-Hui Lee 0001
ASRU5
2015 Channel adaptation of plda for text-independent speaker verification
abstract
Probabilistic linear discriminant analysis (PLDA) has shown to be effective for modeling channel variability in the i-vector space for text-independent speaker verification. Speaker verification is a binary hypothesis testing. Given a test segment, the verification score could be computed as the log-likelihood ratio between a speaker-adapted PLDA and the universal PLDA model. This work proposes to infer the channel factor specific to each test segment and to include the channel estimate in the PLDA models, which essentially shifts the scoring function to better match that of the test channel. We also explore the influence of covariance adaptation in both speaker and channel adaptations. Experimental results on NIST SRE'08 and SRE'10 dataset confirm that the proposed channel adaptation can be effective when the covariance is kept un-adapted, while the covariance adaptation is necessary in the speaker adaptation.
Kong-Aik Lee, Bin Ma 0001, Wu Guo, Haizhou Li 0001, Li-Rong Dai 0001
ICASSP6
2015 Joint training of front-end and back-end deep neural networks for robust speech recognition
abstract
Based on the recently proposed speech pre-processing front-end with deep neural networks (DNNs), we first investigate different feature mapping directly from noisy speech via DNN for robust speech recognition. Next, we propose to jointly train a single DNN for both feature mapping and acoustic modeling. In the end, we show that the word error rate (WER) of the jointly trained system could be significantly reduced by the fusion of multiple DNN pre-processing systems which implies that features obtained from different domains of the DNN-enhanced speech signals are strongly complementary. Testing on the Aurora4 noisy speech recognition task our best system with multi-condition training can achieves an average WER of 10.3%, yielding a relative reduction of 16.3% over our previous DNN pre-processing only system with a WER of 12.3%. To the best of our knowledge, this represents the best published result on the Aurora4 task without using any adaptation techniques.
Tian Gao 0005, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001
ICASSP3
2015 Spectral conversion using deep neural networks trained with multi-source speakers
abstract
This paper presents a method for voice conversion using deep neural networks (DNNs) trained with multiple source speakers. The proposed DNNs can be used in two ways for different scenarios: 1) in the absence of training data for source speaker, the DNNs can be treated as source-speaker-independent models and perform conversions directly from arbitrary source speakers to certain target speaker; 2) the DNNs can also be used as initial models for further fine-tuning of source-speaker-dependent DNNs when parallel training data for both source and target speakers are available. Experimental results show that, as source-speaker-independent models, the proposed DNNs can achieve comparable performance to conventional source-speaker-dependent models. On the other hand, the proposed method outperforms the conventional initialization method with restricted Boltzmann machines (RBMs).
Li-Juan Liu, Linghui Chen, Zhen-Hua Ling, Li-Rong Dai 0001
ICASSP4
2015 Improved language identification using deep bottleneck network
abstract
Effective representation plays an important role in automatic spoken language identification (LID). Recently, several representations that employ a pre-trained deep neural network (DNN) as the front-end feature extractor, have achieved state-of-the-art performance. However the performance is still far from satisfactory for dialect and short-duration utterance identification tasks, due to the deficiency of existing representations. To address this issue, this paper proposes the improved representations to exploit the information extracted from different layers of the DNN structure. This is conceptually motivated by regarding the DNN as a bridge between low-level acoustic input and high-level phonetic output features. Specifically, we employ deep bottleneck network (DBN), a DNN with an internal bottleneck layer acting as a feature extractor. We extract representations from two layers of this single network, i.e. DBN-TopLayer and DBN-MidLayer. Evaluations on the NIST LRE2009 dataset, as well as the more specific dialect recognition task, show that each representation can achieve an incremental performance gain. Furthermore, a simple fusion of the representations is shown to exceed current state-of-the-art performance.
Yan Song 0001, Ruilian Cui, Xinhai Hong, Ian McLoughlin 0001, Jiong Shi, Li-Rong Dai 0001
ICASSP6
2015 Speech Separation based on signal-noise-dependent deep neural networks for robust speech recognition
abstract
In this paper, we propose a new signal-noise-dependent (SND) deep neural network (DNN) framework to further improve the separation and recognition performance of the recently developed technique for general DNN-based speech separation. We adopt a divide and conquer strategy to design the proposed SND-DNNs with higher resolutions that a single general DNN could not well accommodate for all the speaker mixing variabilities at different levels of signal-to-noise ratios (SNRs). In this study two kinds of SNR-dependent DNNs, namely positive and negative DNNs, are trained to cover the mixed speech signals with positive and negative SNR levels, respectively. At the separation stage, a first-pass separation using a general DNN can give an accurate SNR estimation for a model selection. Experimental results on the Speech Separation Challenge (SSC) task show that SND-DNNs could yield significant performance improvements for both speech separation and recognition over a general DNN. Furthermore, this purely front-end processing method achieves a relative word error rate reduction of 11.6% over a state-of-the-art recognition system where a complicated joint decoding framework needs to be implemented in the back-end.
Yanhui Tu, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001
ICASSP3
2015 Unsupervised speaker adaptation of deep neural network based on the combination of speaker codes and singular value decomposition for speech recognition
abstract
Recently, we have proposed a general adaptation scheme for deep neural network based on discriminant condition codes and applied it to supervised speaker adaptation in speech recognition based on either frame-level cross-entropy or sequence-level maximum mutual information training criterion [1, 2, 3, 4]. In this case, each condition code is associated with one speaker in data, which is thus called speaker code for convenience. Our previous work has shown that speaker code based methods are quite effective in adapting DNNs even when only a very small amount of adaptation data is available. However, we have to use a large speaker code size and complex processes to obtain the best ASR performance since good initializations of speaker codes and connection weights are very important. In this paper, we propose a method using singular value decomposition (SVD) as in [5] to initialize speaker codes and connection weights to obtain a comparable ASR performance as before but with a smaller speaker code size and much less computation complexity. Meanwhile, we have evaluated unsupervised speaker adaptation with the proposed method in large vocabulary speech recognition in the Switchboard task. Experimental results have shown that it is effective for providing well initializations and suitable in adapting large DNN models.
Shaofei Xue, Hui Jiang 0001, Li-Rong Dai 0001, Qingfeng Liu
ICASSP3
2015 Writer adaptive feature extraction based on convolutional neural networks for online handwritten Chinese character recognition
abstract
This paper presents a novel approach to writer adaptation based on convolutional neural network (CNN) as a feature extractor and improved discriminative linear regression for online handwritten Chinese character recognition. First, the proposed recognizer consisting of CNN-based feature extractor and prototype-based classifier can achieve comparable performance with the state-of-the-art CNN-based classifier while it could be designed more compact and efficient as a practical solution. Second, the writer adaption is performed via a linear transformation of the extracted feature from CNN. The transformation parameters are optimized with a so-called sample separation margin based minimum classification error criterion, which can be further improved by using more synthesized adaptation data and a simple regularization method. The experiments on the data collected from user inputs of Smartphones with a vocabulary of 20,936 characters demonstrate that our writer adaptation approach can yield significant improvements of recognition accuracy over a high-performance baseline system and also outperform a state-of-the-art approach based on style transfer mapping especially with increased adaptation data.
Jun Du 0002, Jian-Fang Zhai, Jin-Shui Hu, Si Wei, Li-Rong Dai 0001
ICDAR6
2015 Phone-centric local variability vector for text-constrained speaker verification
Kong-Aik Lee, Bin Ma 0001, Wu Guo, Haizhou Li 0001, Li-Rong Dai 0001
INTERSPEECH6
2015 Automatic phrase boundary labeling of speech synthesis database using context-dependent HMMs and n-gram prior distributions
Qian Chen 0003, Zhen-Hua Ling, Chen-Yu Yang, Li-Rong Dai 0001
INTERSPEECH4
2015 Deep bottleneck network based i-vector representation for language identification
abstract
This paper presents a unified i-vector framework for language identification (LID) based on deep bottleneck networks (DBN) trained for automatic speech recognition (ASR). The framework covers both front-end feature extraction and back-end modeling stages.The output from different layers of a DBN are exploited to improve the effectiveness of the i-vector representation through incorporating a mixture of acoustic and phonetic information. Furthermore, a universal model is derived from the DBN with a LID corpus. This is a somewhat inverse process to the GMM-UBM method, in which the GMM of each language is mapped from a GMM-UBM. Evaluations on specific dialect recognition tasks show that the DBN based i-vector can achieve significant and consistent performance gains over conventional GMM-UBM and DNN based i-vector methods. The generalization capability of this framework is also evaluated using DBNs trained on Mandarin and English corpuses. Index Terms: Language Identification, Deep Neural Network, Deep Bottleneck Feature, i-vector representation
Yan Song 0001, Xinhai Hong, Bing Jiang, Ruilian Cui, Ian McLoughlin 0001, Li-Rong Dai 0001
INTERSPEECH6
2015 A universal VAD based on jointly trained deep neural networks
Qing Wang 0008, Jun Du 0002, Xiao Bao, Zi-Rui Wang, Li-Rong Dai 0001, Chin-Hui Lee 0001
INTERSPEECH5
2015 High-resolution acoustic modeling and compact language modeling of language-universal speech attributes for spoken language identification
Yannan Wang, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001
INTERSPEECH3
2015 Multi-objective learning and mask-based post-processing for deep neural network based speech enhancement
abstract
We propose a multi-objective framework to learn both secondary targets not directly related to the intended task of speech enhancement (SE) and the primary target of the clean log-power spectra (LPS) features to be used directly for constructing the enhanced speech signals.In deep neural network (DNN) based SE we introduce an auxiliary structure to learn secondary continuous features, such as mel-frequency cepstral coefficients (MFCCs), and categorical information, such as the ideal binary mask (IBM), and integrate it into the original DNN architecture for joint optimization of all the parameters.This joint estimation scheme imposes additional constraints not available in the direct prediction of LPS, and potentially improves the learning of the primary target.Furthermore, the learned secondary information as a byproduct can be used for other purposes, e.g., the IBM-based post-processing in this work.A series of experiments show that joint LPS and MFCC learning improves the SE performance, and IBM-based post-processing further enhances listening quality of the reconstructed speech.
Yong Xu 0004, Jun Du 0002, Zhen Huang 0001, Li-Rong Dai 0001, Chin-Hui Lee 0001
INTERSPEECH4
2015 Rectified linear neural networks with tied-scalar regularization for LVCSR
Shiliang Zhang, Hui Jiang 0001, Si Wei, Li-Rong Dai 0001
INTERSPEECH4
2015 Deep Bottleneck Feature for Image Classification
abstract
Effective image representation plays an important role for image classification and retrieval. Bag-of-Features (BoF) is well known as an effective and robust visual representation. However, on large datasets, convolutional neural networks (CNN) tend to perform much better, aided by the availability of large amounts of training data. In this paper, we propose a bag of Deep Bottleneck Features (DBF) for image classification, effectively combining the strengths of a CNN within a BoF framework. The DBF features, obtained from a previously well-trained CNN, form a compact and low-dimensional representation of the original inputs, effective for even small datasets. We will demonstrate that the resulting BoDBF method has a very powerful and discriminative capability that is generalisable to other image classification tasks.
Yan Song 0001, Ian McLoughlin 0001, Li-Rong Dai 0001
ICMR3
2015 Statistical parametric speech synthesis using a hidden trajectory model
Ming-Qi Cai, Zhen-Hua Ling, Li-Rong Dai 0001
Speech Commun.3
2015 Quasi-Factorial Prior for i-vector Extraction
abstract
We analyze the i-vector extraction from the perspective of the prior distribution exerted on the mean supervector of Gaussian mixture model (GMM). To this end, we start off with the analysis of the subspace prior which leads to the compressed representation in the standard i-vector extraction. We then propose the use of quasi-factorial prior and show how it impacts the total variability space and its application for i-vector extraction. The quasi-factorial prior could be used in a standalone manner, or in combination with a subspace prior. In the latter context, we found that the performance of the standard i-vector can be greatly improved with the use of quasi-factorial prior followed by a subspace prior. This assertion is confirmed through experiments conducted on the NIST 2010 Speaker Recognition Evaluation (SRE10) dataset.
Kong-Aik Lee, Li-Rong Dai 0001, Haizhou Li 0001
IEEE Signal Process. Lett.3
2015 A Regression Approach to Speech Enhancement Based on Deep Neural Networks
abstract
In contrast to the conventional minimum mean square error (MMSE)-based noise reduction techniques, we propose a supervised method to enhance speech by means of finding a mapping function between noisy and clean speech signals based on deep neural networks (DNNs). In order to be able to handle a wide range of additive noises in real-world situations, a large training set that encompasses many possible combinations of speech and noise types, is first designed. A DNN architecture is then employed as a nonlinear regression function to ensure a powerful modeling capability. Several techniques have also been proposed to improve the DNN-based speech enhancement system, including global variance equalization to alleviate the over-smoothing problem of the regression model, and the dropout and noise-aware training strategies to further improve the generalization capability of DNNs to unseen noise conditions. Experimental results demonstrate that the proposed framework can achieve significant improvements in both objective and subjective measures over the conventional MMSE based technique. It is also interesting to observe that the proposed DNN approach can well suppress highly nonstationary noise, which is tough to handle in general. Furthermore, the resulting DNN model, trained with artificial synthesized data, is also effective in dealing with noisy speech data recorded in real-world scenarios without the generation of the annoying musical artifact commonly observed in conventional enhancement methods.
Yong Xu 0004, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 State-Clustering Based Multiple Deep Neural Networks Modeling Approach for Speech Recognition
abstract
The hybrid deep neural network (DNN) and hidden Markov model (HMM) has recently achieved dramatic performance gains in automatic speech recognition (ASR). The DNN-based acoustic model is very powerful but its learning process is extremely time-consuming. In this paper, we propose a novel DNN-based acoustic modeling framework for speech recognition, where the posterior probabilities of HMM states are computed from multiple DNNs (mDNN), instead of a single large DNN, for the purpose of parallel training towards faster turnaround. In the proposed mDNN method all tied HMM states are first grouped into several disjoint clusters based on data-driven methods. Next, several hierarchically structured DNNs are trained separately in parallel for these clusters using multiple computing units (e.g. GPUs). In decoding, the posterior probabilities of HMM states can be calculated by combining outputs from multiple DNNs. In this work, we have shown that the training procedure of the mDNN under popular criteria, including both frame-level cross-entropy and sequence-level discriminative training, can be parallelized efficiently to yield significant speedup. The training speedup is mainly attributed to the fact that multiple DNNs are parallelized over multiple GPUs and each DNN is smaller in size and trained by only a subset of training data. We have evaluated the proposed mDNN method on a 64-hour Mandarin transcription task and the 320-hour Switchboard task. Compared to the conventional DNN, a 4-cluster mDNN model with similar size can yield comparable recognition performance in Switchboard (only about 2% performance degradation) with a greater than 7 times speed improvement in CE training and a 2.9 times improvement in sequence training, when 4 GPUs are used.
Hui Jiang 0001, Li-Rong Dai 0001, Yu Hu 0003, Qingfeng Liu
IEEE ACM Trans. Audio Speech Lang. Process.3
2014 Minimum divergence estimation of speaker prior in multi-session PLDA scoring
abstract
Probabilistic linear discriminant analysis (PLDA) has shown to be effective for modeling speaker and channel variability in the i-vector space for text-independent speaker verification. This paper shows that the PLDA scoring function could be formulated as model comparison between an adapted PLDA model and the universal PLDA. Based on this formulation, we show that a more robust adaptation could be attained by adapting the PLDA model through the use of minimum divergence estimate of speaker prior in the latent subspace. Experimental results on NIST SRE'10 and SRE'12 dataset confirm that the proposed method is effective in handling multi-session task. Notably, it is free from the covariance shrinkage problem typically found in the standard multi-session PLDA scoring.
Kong-Aik Lee, Bin Ma 0001, Wu Guo, Haizhou Li 0001, Li-Rong Dai 0001
ICASSP6
2014 Synthesized stereo mapping via deep neural networks for noisy speech recognition
abstract
In our previous work, we extend the traditional stereo-based stochastic mapping by relaxing the constraint of stereo-data, which is not practical in real applications, via HMM-based speech synthesis to construct the “clean” channel data for noisy speech recognition. In this paper, we propose to use deep neural networks (DNNs) for stereo mapping compared with the joint Gaussian mixture model (GMM). The experimental results on Aurora3 databases show that our proposed DNN based synthesized stereo mapping can achieve consistently significant improvements of recognition performance over joint GMM based synthesized stereo mapping in the well-matched (WM) condition among four different European languages.
Jun Du 0002, Li-Rong Dai 0001, Qiang Huo
ICASSP2
2014 Using bidirectional associative memories for joint spectral envelope modeling in voice conversion
abstract
The spectral envelope is the most natural representation of speech signal. But in voice conversion, it is difficult to directly model the raw spectral envelope space, which is high dimensional and strongly cross-dimensional correlated, with conventional Gaussian distributions. Bidirectional associative memory (BAM) is a two-layer feedback neural network that can better model the cross-dimensional correlations in high dimensional vectors. In this paper, we propose to reformulate BAMs as Gaussian distributions in order to model the spectral envelope space. The parameters of BAMs are estimated using the contrastive divergence algorithm. The evaluations on likelihood show that BAMs have better modeling ability than Gaussians with diagonal covariance. And the subjective tests on voice conversion indicate that the performance of the proposed method is significantly improved comparing with the conventional GMM based method.
Li-Juan Liu, Linghui Chen, Zhen-Hua Ling, Li-Rong Dai 0001
ICASSP4
2014 Lattice based optimization of bottleneck feature extractor with linear transformation
abstract
This paper proposes a lattice-based sequential discriminative training method to extract more discriminative bottleneck features. In our method, the bottleneck neural network is first trained with cross entropy criteria, and then only the weights of bottleneck layer are retrained with sequential criteria. If the outputs of the layer before bottleneck are treated as the raw features, the new method is an equivalent to a linear feature transformation algorithm. This linearity makes the optimization much easier than updating the whole neural network. Just like the fMPE and RDLT, the neural network is retrained with batch mode gradient descent, making the training to be easily implemented in parallel. Meanwhile, batch mode optimization can naturally deal with the indirect gradient to make the optimization more precise. Experimental results on a Mandarin transcription task and the Switchboard task have shown the effectiveness of the proposed method with the CER decreases from 12.2% to 11.3% and the WER from 16.1% to 15.0%, respectively.
Diyuan Liu, Si Wei, Wu Guo, Yebo Bao, Shifu Xiong, Li-Rong Dai 0001
ICASSP6
2014 Direct adaptation of hybrid DNN/HMM model for fast speaker adaptation in LVCSR based on speaker code
abstract
Recently an effective fast speaker adaptation method using discriminative speaker code (SC) has been proposed for the hybrid DNN-HMM models in speech recognition [1]. This adaptation method depends on a joint learning of a large generic adaptation neural network for all speakers as well as multiple small speaker codes using the standard back-propagation algorithm. In this paper, we propose an alternative direct adaptation in model space, where speaker codes are directly connected to the original DNN models through a set of new connection weights, which can be estimated very efficiently from all or part of training data. As a result, the proposed method is more suitable for large scale speech recognition tasks since it eliminates the time-consuming training process to estimate another adaptation neural networks. In this work, we have evaluated the proposed direct SC-based adaptation method in the large scale 320-hr Switchboard task. Experimental results have shown that the proposed SC-based rapid adaptation method is very effective not only for small recognition tasks but also for very large scale tasks. For example, it has shown that the proposed method leads to up to 8% relative reduction in word error rate in Switchboard by using only a very small number of adaptation utterances per speaker (from 10 to a few dozens). Moreover, the extra training time required for adaptation is also significantly reduced from the method in [1].
Shaofei Xue, Ossama Abdel-Hamid, Hui Jiang 0001, Li-Rong Dai 0001
ICASSP4
2014 Spectral modeling using neural autoregressive distribution estimators for statistical parametric speech synthesis
abstract
This paper describes a new approach which utilizes neural autoregressive distribution estimators (NADE) for the spectral modeling in statistical parametric speech synthesis. In order to alleviate the over-smoothing effect on the generated spectral structures, a restricted Boltzmann machine (RBM) modeling method has been proposed in our previous work, where the RBM is adopted to represent the joint distribution of high-dimensional and physically meaningful spectral envelopes. However, the RBM can not provide a tractable partition function even in a moderate size. In this paper, we introduce NADE to model the distribution of mel-cepstra and spectral envelopes at each HMM state considering its simplicity in evaluating the probability of given observations. At the stage of synthesis, the spectral parameters derived from the mode of each context-dependent NADE are used to replace the Gaussian mean vector in the parameter generation process. Experimental results show that the NADE is able to model the distribution of the spectral features with better accuracy than the RBM model. Furthermore, our proposed method improves the naturalness of the conventional HMM-based speech synthesis system using mel-cepstra significantly and outperforms the RBM-based spectral modeling.
Xiang Yin 0002, Zhen-Hua Ling, Li-Rong Dai 0001
ICASSP3
2014 Improving deep neural networks for LVCSR using dropout and shrinking structure
abstract
Recently, the hybrid deep neural networks and hidden Markov models (DNN/HMMs) have achieved dramatic gains over the conventional GMM/HMMs method on various large vocabulary continuous speech recognition (LVCSR) tasks. In this paper, we propose two new methods to further improve the hybrid DNN/HMMs model: i) use dropout as pre-conditioner (DAP) to initialize DNN prior to back-propagation (BP) for better recognition accuracy; ii) employ a shrinking DNN structure (sDNN) with hidden layers decreasing in size from bottom to top for the purpose of reducing model size and expediting computation time. The proposed DAP method is evaluated in a 70-hour Mandarin transcription (PSC) task and the 309-hour Switchboard (SWB) task. Compared with the traditional greedy layer-wise pre-trained DNN, it can achieve about 10% and 6.8% relative recognition error reduction for PSC and SWB tasks respectively. In addition, we also evaluate sDNN as well as its combination with DAP on the SWB task. Experimental results show that these methods can reduce model size to 45% of original size and accelerate training and test time by 55%, without losing recognition accuracy.
Shiliang Zhang, Yebo Bao, Hui Jiang 0001, Li-Rong Dai 0001
ICASSP5
2014 Sequence training of multiple deep neural networks for better performance and faster training speed
abstract
Recently, sequence level discriminative training methods have been proposed to fine-tune deep neural networks (DNN) after the framelevel cross entropy (CE) training to further improve recognition performance of DNNs. In our previous work, we have proposed a new cluster-based multiple DNNs structure and its parallel training algorithm based on the frame-level cross entropy criterion, which can significantly expedite CE training with multiple GPUs. In this paper, we extend to full sequence training for the multiple DNNs structure for better performance and meanwhile we also consider a partial parallel implementation of sequence training using multiple GPUs for faster training speed. In this work, it is shown that sequence training can be easily extended to multiple DNNs by slightly modifying error signals in output layer. Many implementation steps in sequence training of multiple DNNs can still be parallelized across multiple GPUs for better efficiency. Experiments on the Switchboard task have shown that both frame-level CE training and sequence training of multiple DNNs can lead to massive training speedup with little degradation in recognition performance. Comparing with the state-of-the-art DNN, 4-cluster multiple DNNs model with similar size can achieve more than 7 times faster in CE training and about 1.5 times faster in sequence training when using 4 GPUs.
Li-Rong Dai 0001, Hui Jiang 0001
ICASSP2
2014 Writer Adaptation Using Bottleneck Features and Discriminative Linear Regression for Online Handwritten Chinese Character Recognition
abstract
This paper presents a novel approach to writer adaptation using bottleneck features and discriminative linear regression for the recognition of online handwritten Chinese characters. First, bottleneck features extracted from a bottleneck layer of a deep neural network representing a nonlinear and discriminative transformation of the input features are verified to be much more effective in adaptation of writing styles than the conventional features after linear discriminant analysis transformation. Second, discriminative linear regression via a so-called sample separation margin based minimum classification error criterion is adopted for writer adaptation. The experiments on an in-house developed online Chinese handwriting corpus with a vocabulary of 15,167 characters and testing data collected from user inputs of Smartphones show that our proposed approach can achieve very significant improvements of recognition accuracy compared with a state-of-the-art adaptation approach for writer adaptation.
Jun Du 0002, Jin-Shui Hu, Si Wei, Li-Rong Dai 0001
ICFHR5
2014 A Study of Designing Compact Classifiers Using Deep Neural Networks for Online Handwritten Chinese Character Recognition
abstract
This paper presents a study of designing compact classifiers using deep neural networks for recognition of online handwritten Chinese characters. Two schemes are investigated based on practical considerations. First, deep neural networks are adopted purely as a classifier with a state-of-the-art feature extractor of online handwritten Chinese characters. Second, the so-called bottleneck features extracted from a bottleneck layer of deep neural networks are fed to the prototype-based classifier. The experiments on an in-house developed online Chinese handwriting corpus with a vocabulary of 15,167 characters show that compared with prototype-based classifier widely developed on the mobile device, deep neural network based classifier can yield significant improvements of recognition accuracy with acceptably increased footprint and latency while the bottleneck-feature approach can bring a more compact classifier with an observable performance gain.
Jun Du 0002, Jin-Shui Hu, Si Wei, Li-Rong Dai 0001
ICPR5
2014 Formant-controlled speech synthesis using hidden trajectory model
Ming-Qi Cai, Zhen-Hua Ling, Li-Rong Dai 0001
INTERSPEECH3
2014 Voice conversion using generative trained deep neural networks with multiple frame spectral envelopes
Linghui Chen, Zhen-Hua Ling, Li-Rong Dai 0001
INTERSPEECH3
2014 Robust speech recognition with speech enhanced deep neural networks
Jun Du 0002, Qing Wang 0008, Tian Gao 0005, Yong Xu 0004, Li-Rong Dai 0001, Chin-Hui Lee 0001
INTERSPEECH5
2014 Task-aware deep bottleneck features for spoken language identification
abstract
Recently, deep bottleneck features (DBF) extracted from a deep neural network (DNN) containing a narrow bottleneck lay-er, have been applied for language identification (LID), and yield significant performance improvement over state-of-the-art methods on NIST LRE 2009. However, the DNN is trained us-ing a large corpus of specific language which is not directly related to the LID task. More recently, lattice based discrimi-native training methods for extracting more targeted DBF were proposed for ASR. Inspired by this, this paper proposes to tune the post-trained DNN parameters using an LID-specific train-ing corpus, which may make the resulting DBF, termed a Dis-criminative DBF (D2BF), more discriminative and task-aware. Specifically, the maximum mutual information (MMI) criteri-on, with gradient descent, is applied to update the DNN param-eters of the bottleneck layer in an iterative fashion. We evaluate the performance of the proposed D2BF using different back-end models, including GMM-MMI and ivector, over the most con-fused 6-languages selected from NIST LRE 2009. The results show that the proposed D2BF is more appropriate and effective than the original DBF. Index Terms: language identification, deep bottleneck feature, deep neural network, discriminative training, Gaussian mixture model, maximum mutual information 1.
Bing Jiang, Yan Song 0001, Si Wei, Ian McLoughlin 0001, Li-Rong Dai 0001
INTERSPEECH5
2014 Concept-to-speech generation by integrating syntagmatic features into HMM-based speech synthesis
Xin Wang 0037, Zhen-Hua Ling, Li-Rong Dai 0001
INTERSPEECH3
2014 Dynamic noise aware training for speech enhancement based on deep neural networks
abstract
We propose three algorithms to address the mismatch problem in deep neural network (DNN) based speech enhancement. First, we investigate noise aware training by incorporating noise informationin the testutterance with anideal binary maskbased dynamic noise estimation approach to improve DNN’s speech separation ability from the noisy signal. Next, a set of more than 100 noise types is adopted to enrich the generalization capabilities of the DNN to unseen and non-stationary noise conditions. Finally, the quality of the enhanced speech can further be improved by global variance equalization. Empirical results show that each of the three proposed techniques contributes to the performance improvement. Compared to the conventional logarithmic minimum mean squared error speech enhancement method, our DNN system achieves 0.32 PESQ (perceptual evaluation of speech quality) improvement across six signal-tonoise ratio levels ranging from -5dB to 20dB on a test set with unknown noise types. We also observe that the combined strategies can well suppress highly non-stationary noise better than all the competing state-of-the-art techniques we have evaluated. Index Terms: Speech enhancement, deep neural networks, noise aware training, ideal binary mask, non-stationary noise
Yong Xu 0004, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001
INTERSPEECH3
2014 Modeling DCT parameterized F0 trajectory at intonation phrase level with DNN or decision tree
Xiang Yin 0002, Yao Qian, Frank K. Soong, Lei He 0005, Zhen-Hua Ling, Li-Rong Dai 0001
INTERSPEECH7
2014 HMM-based unit selection speech synthesis using log likelihood ratios derived from perceptual data
Xian-Jun Xia, Zhen-Hua Ling, Yuan Jiang 0006, Li-Rong Dai 0001
Speech Commun.4
2014 An Experimental Study on Speech Enhancement Based on Deep Neural Networks
abstract
This letter presents a regression-based speech enhancement framework using deep neural networks (DNNs) with a multiple-layer deep architecture. In the DNN learning process, a large training set ensures a powerful modeling capability to estimate the complicated nonlinear mapping from observed noisy speech to desired clean signals. Acoustic context was found to improve the continuity of speech to be separated from the background noises successfully without the annoying musical artifact commonly observed in conventional speech enhancement algorithms. A series of pilot experiments were conducted under multi-condition training with more than 100 hours of simulated speech data, resulting in a good generalization capability even in mismatched testing conditions. When compared with the logarithmic minimum mean square error approach, the proposed DNN-based algorithm tends to achieve significant improvements in terms of various objective quality measures. Furthermore, in a subjective preference evaluation with 10 listeners, 76.35% of the subjects were found to prefer DNN-based enhanced speech to that obtained with other conventional technique.
Yong Xu 0004, Jun Du 0002, Li-Rong Dai 0001, Chin-Hui Lee 0001
IEEE Signal Process. Lett.3
2014 Voice conversion using deep neural networks with layer-wise generative training
abstract
This paper presents a new spectral envelope conversion method using deep neural networks (DNNs). The conventional joint density Gaussian mixture model (JDGMM) based spectral conversion methods perform stably and effectively. However, the speech generated by these methods suffer severe quality degradation due to the following two factors: 1) inadequacy of JDGMM in modeling the distribution of spectral features as well as the non-linear mapping relationship between the source and target speakers, 2) spectral detail loss caused by the use of high-level spectral features such as mel-cepstra. Previously, we have proposed to use the mixture of restricted Boltzmann machines (MoRBM) and the mixture of Gaussian bidirectional associative memories (MoGBAM) to cope with these problems. In this paper, we propose to use a DNN to construct a global non-linear mapping relationship between the spectral envelopes of two speakers. The proposed DNN is generatively trained by cascading two RBMs, which model the distributions of spectral envelopes of source and target speakers respectively, using a Bernoulli BAM (BBAM). Therefore, the proposed training method takes the advantage of the strong modeling ability of RBMs in modeling the distribution of spectral envelopes and the superiority of BAMs in deriving the conditional distributions for conversion. Careful comparisons and analysis among the proposed method and some conventional methods are presented in this paper. The subjective results show that the proposed method can significantly improve the performance in terms of both similarity and naturalness compared to conventional methods.
Linghui Chen, Zhen-Hua Ling, Li-Juan Liu, Li-Rong Dai 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2014 Fast adaptation of deep neural network based on discriminant codes for speech recognition
abstract
Fast adaptation of deep neural networks (DNN) is an important research topic in deep learning. In this paper, we have proposed a general adaptation scheme for DNN based on discriminant condition codes, which are directly fed to various layers of a pre-trained DNN through a new set of connection weights. Moreover, we present several training methods to learn connection weights from training data as well as the corresponding adaptation methods to learn new condition code from adaptation data for each new test condition. In this work, the fast adaptation scheme is applied to supervised speaker adaptation in speech recognition based on either frame-level cross-entropy or sequence-level maximum mutual information training criterion. We have proposed three different ways to apply this adaptation scheme based on the so-called speaker codes: i) Nonlinear feature normalization in feature space; ii) Direct model adaptation of DNN based on speaker codes; iii) Joint speaker adaptive training with speaker codes. We have evaluated the proposed adaptation methods in two standard speech recognition tasks, namely TIMIT phone recognition and large vocabulary speech recognition in the Switchboard task. Experimental results have shown that all three methods are quite effective to adapt large DNN models using only a small amount of adaptation data. For example, the Switchboard results have shown that the proposed speaker-code-based adaptation methods may achieve up to 8-10% relative error reduction using only a few dozens of adaptation utterances per speaker. Finally, we have achieved very good performance in Switchboard (12.1% in WER) after speaker adaptation using sequence training criterion, which is very close to the best performance reported in this task (“Deep convolutional neural networks for LVCSR,” T. N. Sainath , Proc. IEEE Acoust., Speech, Signal Process., 2013).
Shaofei Xue, Ossama Abdel-Hamid, Hui Jiang 0001, Li-Rong Dai 0001, Qingfeng Liu
IEEE ACM Trans. Audio Speech Lang. Process.4
2013 Incoherent training of deep neural networks to de-correlate bottleneck features for speech recognition
abstract
Recently, the hybrid model combining deep neural network (DNN) with context-dependent HMMs has achieved some dramatic gains over the conventional GMM/HMM method in many speech recognition tasks. In this paper, we study how to compete with the state-of-the-art DNN/HMM method under the traditional GMM/HMM framework. Instead of using DNN as acoustic model, we use DNN as a front-end bottleneck (BN) feature extraction method to decorrelate long feature vectors concatenated from several consecutive speech frames. More importantly, we have proposed two novel incoherent training methods to explicitly de-correlate BN features in learning of DNN. The first method relies on minimizing coherence of weight matrices in DNN while the second one attempts to minimize correlation coefficients of BN features calculated in each mini-batch data in DNN training. Experimental results on a 70-hr Mandarin transcription task and the 309-hr Switchboard task have shown that the traditional GMM/HMMs using BN features can yield comparable performance as DNN/HMM. The proposed incoherent training can produce 2-3% additional gain over the baseline BN features. At last, the discriminatively trained GMM/HMMs using incoherently trained BN features have consistently surpassed the state-of-the-art DNN/HMMs in all evaluated tasks.
Yebo Bao, Hui Jiang 0001, Li-Rong Dai 0001, Cong Liu 0006
ICASSP3
2013 Phoneme variation based synthesized speech discrimination for speaker verification
abstract
How to discriminate the synthesized speech from the natural speech for speaker verification is addressed in this paper. With the development of HMM-based speech synthesis, it is easy to obtain high quality synthesized speech which sounds like target speaker, the robustness of synthesized speech become important for speaker verification. In this paper, a method based on the phoneme variation is proposed to discriminate synthesized speech from natural speech, which could be used as front-end module to detect the synthesized speech for speaker verification system. The experimental results show the effectiveness of the proposed method.
LianWu Chen, Wu Guo, Yan Song 0001, Li-Rong Dai 0001
ICASSP4
2013 Exemplar based language recognition method for short-duration speech segments
abstract
This paper proposes a novel exemplar-based language recognition method for short duration speech segments. It is known that language identity is a kind of weak information that can be deduced from the speech content. For short duration speech segments, the limited content also leads to a large intra-language variability. To address this issue, we propose a new method. This borrows a vector quantization based representation from image classification methods, and constructs the exemplar space using the popular i-vector representation of short duration speech segments. A mapping function is then defined to build the new representation. To evaluate the effectiveness of our proposed method, we conduct extensive experiments on the NIST LRE2007 dataset. The experimental results demonstrate improved performance for short duration speech segments.
Meng-Ge Wang, Yan Song 0001, Bing Jiang, Li-Rong Dai 0001, Ian McLoughlin 0001
ICASSP4
2013 Unsupervised prosodic phrase boundary labeling of Mandarin speech synthesis database using context-dependent HMM
abstract
In this paper, an automatic and unsupervised method based on context-dependent hidden Markov model (CD-HMM) is proposed for labeling the phrase boundary positions of a Mandarin speech synthesis database. The initial phrase boundary labels are predicted by clustering the durations of the pauses between every two prosodic words in an unsupervised way. Then, the CD-HMMs for the spectrum, F0 and phone duration are estimated by a means similar to the HMM-based parametric speech synthesis using the initial phrase boundary labels. These labels are further updated by Viterbi decoding under the maximum likelihood criterion given the acoustic feature sequences and the trained CD-HMMs. The model training and Viterbi decoding procedures are conducted iteratively until convergence. Experimental results on a Mandarin speech synthesis database show that this method is able to label the phrase boundary positions much more accurately than the text-analysis-based method without requiring any manually labeled training data. The unit selection speech synthesis system constructed using the phrase boundary labels generated by our proposed method achieves similar performance to that using the manual labels.
Chen-Yu Yang, Zhen-Hua Ling, Li-Rong Dai 0001
ICASSP3
2013 A cluster-based multiple deep neural networks method for large vocabulary continuous speech recognition
abstract
Recently a pre-trained context-dependent hybrid deep neural network (DNN) and HMM method has achieved significant performance gain in many large-scale automatic speech recognition (ASR) tasks. However, the error back-propagation (BP) algorithm for training neural networks is sequential in nature and is hard to parallelize into multiple computing threads. Therefore, training a deep neural network is extremely time-consuming even with a modern GPU board. In this paper we have proposed a new acoustic modelling framework to use multiple DNNs instead of a single DNN to compute the posterior probabilities of tied HMM states. In our method, all tied states of context-dependent HMMs are first grouped into several disjoined clusters based on the training data associated with these HMM states. Then, several hierarchically structured DNNs are trained separately for these disjoined clusters of data using multiple GPUs. In decoding, the final posterior probability of each tied HMM state can be calculated based on output posteriors from multiple DNNs. We have evaluated the proposed method on a 64-hour Mandarin transcription task and 309-hour Switchboard Hub5 task. Experimental results have shown that the new method using clusterbased multiple DNNs can achieve over 5 times reduction in total training time with only negligible performance degradation (about 1-2% in average) when using 3 or 4 GPUs respectively.
Cong Liu 0006, Qingfeng Liu, Li-Rong Dai 0001, Hui Jiang 0001
ICASSP4
2013 Joint spectral distribution modeling using restricted boltzmann machines for voice conversion
Linghui Chen, Zhen-Hua Ling, Yan Song 0001, Li-Rong Dai 0001
INTERSPEECH4
2012 Exemplar-Based Sparse Representation for Language Recognition on I-Vectors
abstract
In this paper, a new automatic language identification method using sparse representation on i-vectors in low-dimensional total variability space is proposed. It is mainly based on the recently proposed i-vector based language recognition systems. In our proposed method, an over-complete dictionary is first constructed by randomly sampling of the low-dimensional total variability space after Within-Class Covariance Normalization (WCCN) and Linear Discriminate Analysis (LDA). And then for each test sample, the classification score is derived from sparse linear representation with respect to the over-complete dictionary. Furthermore, a random subspace method, which combines different sparse representation classifiers, is introduced to address the possible over-fitting issue. Evaluations on NIST LRE 2007 dataset show that the proposed method outperforms the state-of-the-art i-vector based language recognition system. Especially for 30s test condition, our proposed method achieves relative reduction of 29.6 % on Equal Error Rate (EER) compared with the baseline system.
Bing Jiang, Yan Song 0001, Wu Guo, Li-Rong Dai 0001
INTERSPEECH4
2012 Considering Global Variance of the Log Power Spectrum Derived from Mel-Cepstrum in HMM-based Parametric Speech Synthesis
abstract
This paper utilizes global variance (GV) of the log power spectrum (LPS) derived from mel-cepstrum to improve hidden Markov model (HMM) based parametric speech synthesis. In order to alleviate over-smoothing of the generated spectral structures, an LPS-GV modeling method using line spectral pairs (LSPs) has been proposed in our previous work, where the estimated distribution of LPS-GV was combined with the trained acoustic model to determine the optimal spectral features at synthesis time. In this paper, we extend this method to the condition where mel-cepstral coefficients are used as spectral features. Further, a method of integrating LPS-GV distortions into the criterion of minimum generation error (MGE) model training is proposed in order to avoid high computational complexity of the parameter generation algorithm with GV model. Experimental results show that the parameter generation algorithm using LPS-GV model produces more natural acoustic features than the conventional GV modeling method when mel-cepstrum features are adopted. Besides, integrating LPS-GV distortions into model training criterion achieves similar performance as applying LPS-GV model at synthesis time.
Xiang Yin 0002, Zhen-Hua Ling, Li-Rong Dai 0001
INTERSPEECH4
2012 Minimum Kullback-Leibler Divergence Parameter Generation for HMM-Based Speech Synthesis
abstract
This paper presents a parameter generation method for hidden Markov model (HMM)-based statistical parametric speech synthesis that uses a similarity measure for probability distributions. In contrast to conventional maximum output probability parameter generation (MOPPG), the method we propose derives a parameter generation criterion from the distribution characteristics of the generated acoustic features. Kullback-Leibler (KL) divergence between the sentence HMM used for parameter generation and the HMM estimated from the generated features is calculated by upper bound approximation. During parameter generation, this KL divergence is minimized either by optimizing the generated acoustic parameters directly or by applying a linear transform to the MOPPG outputs. Our experiments show both these approaches are effective for alleviating over-smoothing in the generated spectral features and for improving the naturalness of synthetic speech. Compared with the direct optimization approach, which is susceptible to over-fitting, the feature transform approach gives better performance. In order to reduce the computational complexity of transform estimation, an offline training method is further developed to estimate a global transform under the minimum KL divergence criterion for the training set. Experimental results show that this global transform is as effective as the transform estimated for each sentence at synthesis stage.
Zhen-Hua Ling, Li-Rong Dai 0001
IEEE Trans. Speech Audio Process.2
2011 Non-parallel training for voice conversion based on FT-GMM
abstract
This paper presents a non-parallel training algorithm for voice con version based on feature transform Gaussian mixture model (FT GMM), which is a mixture model of joint density space of source speaker and target speaker with explicit feature transform modeling. In FT-GMM, the correlations between the distributions of two speakers in each component of the mixture model are not directly modeled, but absorbed into these explicit feature transformations. This makes it possible to extend this model to non-parallel training by simply decomposing it into two sub-models, one for each speaker and optimizing them separatively. A frequency warping process is adopted to compensate performance degradation caused by original spectral distance between source and target speakers. Cross-gender experimental results show that the proposed method achieves comparable performance as parallel training.
Linghui Chen, Zhen-Hua Ling, Li-Rong Dai 0001
ICASSP3
2011 Preserve ordering property of generated LSPS for minimum generation error training in HMM-based speech synthesis
abstract
Ordering property is an important property of LSP and closely connected with the naturalness of reconstructed speech. When LSP is adopted as spectrum feature in HMM-based parametric speech synthesis, the ordering property cannot be guaranteed because diagonal covariance matrix is used in conventional system and the cross dimension correlation of LSP vector is ignored. It will cause un stable issue in synthesized speech. In this paper, we propose some methods to preserve the ordering property of generated LSPs for MGE training by introducing mis-ordering related distance measurements into model training criterion. Experimental results show that two methods can alleviate the mis-orderings significantly without degrading the MGE performance, and one of which, the minimum mis-ordering counting method, requires no acoustic observations for model optimization.
Zhen-Hua Ling, Li-Rong Dai 0001
ICASSP3
2011 Speaker characterization using spectral subband energy ratio based on Harmonic plus Noise Model
abstract
This paper proposes a feature extraction for speaker characterization by exploring the relationship between the two distinct components of the speech signal, one is harmonics accounting for the periodicity of the signal and the other is modulated noise accounting for the turbulences of the glottal airflow. The harmonic and noise parts of the speech signal are decomposed based on the Harmonic plus Noise Model approach. We estimate the spectral subband energy ratios (SSERs) as the speaker characteristic features, which are expected to reflect the interaction property of the vocal tract and glottal airflow of individual speakers for speaker verification. The speaker verification experiments based on a GMM-UBM system have shown the efficiency of the SSER features, reducing the error equal rate by 27.2% by combining with the conventional MFCC features.
Yanhua Long, Zhijie Yan, Frank K. Soong, Li-Rong Dai 0001, Wu Guo
ICASSP4
2011 Building HMM based unit-selection speech synthesis system using synthetic speech naturalness evaluation score
abstract
This paper proposes a unit-selection and waveform concatenation speech synthesis system based on synthetic speech naturalness evaluation. A Support Vector Machine (SVM) and Log Likelihood Ratio (LLR) based synthetic speech naturalness evaluation system was introduced in our previous work. In this paper, the evaluation system is improved in three aspects. Finally, a unit-selection and concatenation waveform speech synthesis system is built on the base of the synthetic speech naturalness evaluation system. Optimum unit sequence is chosen through the re-scoring for the N-best path. Subjective listening tests show the proposed synthetic speech evaluation based speech synthesis system significantly outperforms the traditional unit-selection speech synthesis system.
Heng Lu 0002, Zhen-Hua Ling, Li-Rong Dai 0001, Renhua Wang
ICASSP3
2011 Factored covariance modeling for text-independent speaker verification
abstract
Gaussian mixture models (GMMs) are commonly used to model the spectral distribution of speech signals for text-independent speaker verification. Mean vectors of the GMM, used in conjunction with support vector machine (SVM), have shown to be effective in characterizing speaker information. In addition to the mean vectors, covariance matrices capture the correlation between spectral features, which also represent some salient information about speaker identity. This paper investigates the use of local correlation between different dimensions of acoustic vector by using factor analysis and linear Gaussian model. Log-Euclidean inner product kernel is used to measure the similarity between two speech utterances in the form of covariance matrices. Experiments carried on NIST 2006 speaker verification tasks shows promising results.
Eryu Wang, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001, Wu Guo, Li-Rong Dai 0001
ICASSP6
2011 Estimation of Window Coefficients for Dynamic Feature Extraction for HMM-Based Speech Synthesis
Linghui Chen, Yoshihiko Nankaku, Heiga Zen, Keiichi Tokuda, Zhen-Hua Ling, Li-Rong Dai 0001
INTERSPEECH6
2011 Formant-Controlled HMM-Based Speech Synthesis
abstract
This paper proposes a novel framework that enables us to manipulate and control formants in HMM-based speech synthesis. In this framework, the dependency between formants and spectral features is modelled by piecewise linear transforms; formant parameters are effectively mapped by these to the means of Gaussian distributions over the spectral synthesis parameters. The spectral envelope features generated under the influence of formants in this way may then be passed to high-quality vocoders to generate the speech waveform. This provides two major advantages over conventional frameworks. First, we can achieve spectral modification by changing formants only in those parts where we want control, whereas the user must specify all formants manually in conventional formant synthesisers (e.g. Klatt). Second, this can produce high-quality speech. Our results show the proposed method can control vowels in the synthesized speech by manipulating F 1 and F 2 without any degradation in synthesis quality.
Junichi Yamagishi, Korin Richmond, Zhen-Hua Ling, Simon King 0001, Li-Rong Dai 0001
INTERSPEECH6
2011 Improvements in Speaker Characterization Using Spectral Subband Energy Based on Harmonic plus Noise Model
Yanhua Long, Zhijie Yan, Frank K. Soong, Li-Rong Dai 0001, Wu Guo
INTERSPEECH4
2011 Trust Region-Based Optimization for Maximum Mutual Information Estimation of HMMs in Speech Recognition
abstract
In this paper, we have proposed two novel optimization methods for discriminative training (DT) of hidden Markov models (HMMs) in speech recognition based on an efficient global optimization algorithm used to solve the so-called trust region (TR) problem, where a quadratic function is minimized under a spherical constraint. In the first method, maximum mutual information estimation (MMIE) of Gaussian mixture HMMs is formulated as a standard TR problem so that the efficient global optimization method can be used in each iteration to maximize the auxiliary function of discriminative training for speech recognition. In the second method, we propose to construct a new auxiliary function for DT of HMMs by adding a quadratic penalty term. The new auxiliary function is constructed to serve as first-order approximation as well as lower bound of the original discriminative objective function within a locality constraint. Due to the lower-bound property, the found optimal point of the new auxiliary function is guaranteed to improve the original discriminative objective function until it converges to a local optimum or stationary point of the objective function. Both TR-based optimization methods have been investigated on two standard large-vocabulary continuous speech recognition tasks, using the WSJ0 and Switchboard databases. Experimental results have shown that the proposed TR methods outperform the conventional EBW method in terms of convergence behavior as well as recognition performance.
Cong Liu 0006, Yu Hu 0003, Li-Rong Dai 0001, Hui Jiang 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2010 HMM-based pseudo-clean speech synthesis for splice algorithm
abstract
In this paper, we present a novel approach to relax the constraint of stereo-data which is needed in a series of algorithms for noise-robust speech recognition. As a demonstration in SPLICE algorithm, we generate the pseudo-clean features to replace the ideal clean features from one of the stereo channels, by using HMM-based speech synthesis. Experimental results on aurora2 database show that the performance of our approach is comparable with that of SPLICE. Further improvements are achieved by concatenating a bias adaptation algorithm to handle unknown environments. Relative word error rate reductions of 66% and 24% are achieved over the baseline systems in the clean-training and multi-training conditions, respectively.
Jun Du 0002, Yu Hu 0003, Li-Rong Dai 0001, Renhua Wang
ICASSP3
2010 N-gram nearest neighbor algorithm for voice password system
abstract
A specific issue in the voice password system is addressed in this paper: When the text content of target speaker's enrollment password has been already known by imposters, they can do a well-behaved impersonation using the same text content as the target speaker. This results in a much higher false acceptance than the traditional voice password system. N-gram based nearest neighbor algorithm is proposed here to improve the speaker detection accuracy. Furthermore, correlation coefficient is adopted as the distance measurement between two acoustic features instead of the traditional Euclidean distance. Experimental results show that the proposed method outperforms the DTW and GMM-UBM algorithms.
Wu Guo, Yanhua Long, Li-Rong Dai 0001
ICASSP4
2010 Minimum generation error training with weighted Euclidean distance on LSP for HMM-based speech synthesis
abstract
This paper presents a minimum generation error (MGE) training method using weighted Euclidean distance measure on line spectral pairs (LSP) for HMM-based speech synthesis. In this paper, weighted Euclidean distance on LSP is introduced as the measurement of generation error to improve the consistency between the model training criterion and the subjective perception on the distortion of synthetic speech. Several common weighting techniques are investigated and compared within the MGE training framework. The experimental results show that the formant bounded weighting (FBW) method achieves the best performance, which improves the naturalness of synthetic speech significantly compared with the Euclidean LSP distance measure. Compared with the MGE training using log spectral distortion (LSD) measure, the FBW criterion can achieve similar performance on naturalness with much less computation complexity of model training.
Zhen-Hua Ling, Li-Rong Dai 0001
ICASSP3
2010 A bounded trust region optimization for discriminative training of HMMS in speech recognition
abstract
In this paper, we have proposed a new method to construct an auxiliary function for the discriminative training of HMMs in speech recognition. The new auxiliary function serves as a first-order approximation of the original objective function but more importantly it remains as a lower bound of the original objective function as well. Furthermore, the trust region (TR) method in [1] is applied to find the globally optimal point of the new auxiliary function. Due to its lower-bound property, the found optimal point is theoretically guaranteed to increase the original discriminative objective function. The proposed bounded trust region method has been investigated on two LVCSR tasks, namely WSJ-5k and Switchboard 60-hour subset tasks. Experimental results show that the bounded TR method yields much better convergence behavior than both the conventional EBW method and the original TR method.
Cong Liu 0006, Yu Hu 0003, Hui Jiang 0001, Li-Rong Dai 0001
ICASSP4
2010 Multiple instance learning using visual phrases for object classification
abstract
Recently, bag of words (BoW) model has led to many significant results in visual object classification. However, due to the limited descriptive and discriminative ability of visual words, the resulting performance of visual object classification is still incomparable to its analogy in text domain, i.e. document categorization. Furthermore, for weakly labeled image data, where we only know whether an object is present or not, traditional learning based methods may suffer from background clutters and large appearance variations. To address these issues, we propose a novel visual phrase based Multiple Instance Learning (MIL) method. In this method, the visual phrase is first generated from over-segmented image regions of homogeneous appearance and visual words within each region, which may provide enhanced descriptive ability by enforcing the spatial coherency. Then a MIL algorithm is applied to efficiently learn from the weakly labeled image data. The experiments on benchmark datasets show that our proposed method always significantly outperforms several state-of-the-art algorithms, such as Spatial Pyramid Matching (SPM) and Spatial-LTM.
Yan Song 0001, Qi Tian 0001, Mengyue Wang, Li-Rong Dai 0001
ICME5
2010 A hierarchical F0 modeling method for HMM-based speech synthesis
Yi-Jian Wu, Frank K. Soong, Zhen-Hua Ling, Li-Rong Dai 0001
INTERSPEECH5
2010 Global variance modeling on the log power spectrum of LSPs for HMM-based speech synthesis
Zhen-Hua Ling, Yu Hu 0003, Li-Rong Dai 0001
INTERSPEECH3
2010 Effects of the phonological relevance in speaker verification
Yanhua Long, Li-Rong Dai 0001, Bin Ma 0001, Wu Guo
INTERSPEECH2
2010 Automatic error detection for unit selection speech synthesis using log likelihood ratio based SVM classifier
abstract
This paper proposes a method to detect the errors in synthetic speech of a unit selection speech synthesis system automatically using log likelihood ratio and support vector machine (SVM). For SVM training, a set of synthetic speech are firstly generated by a given speech synthesis system and their synthetic errors are labeled by manually annotating the segments that sound unnatural. Then, two context-dependent acoustic models are trained using the natural and unnatural segments of labeled synthetic speech respectively. The log likelihood ratio of acoustic features between these two models is adopted to train the SVM classifier for error detection. Experimental results show the proposed method is effective in detecting the errors of pitch contour within a word for a Mandarin speech synthesis system. The proposed SVM method using log likelihood ratio between context-dependent acoustic models outperforms the SVM classifier trained on acoustic features directly.
Heng Lu 0002, Zhen-Hua Ling, Si Wei, Li-Rong Dai 0001, Renhua Wang
INTERSPEECH4
2010 The estimation and kernel metric of spectral correlation for text-independent speaker verification
Eryu Wang, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001, Wu Guo, Li-Rong Dai 0001
INTERSPEECH6
2009 iFLY system for the NIST 2008 speaker recognition evaluation
abstract
The description of iFLY system submitted for NIST 2008 speaker recognition evaluation (SRE), which has achieved excellent performance in the 2008 SRE evaluation, is presented in this paper. Our primary system is a fusion of two subsystems GMM-UBM and GMM-SVM. For each sub-system, two kinds of short-time acoustic features PLP and LPCC are adopted. We focus on three key issues in this evaluation: channel compensation, multi-lingual or bi-lingual cues and the voice activity detection. We also point out that data selection and factor analysis play key roles in the system improvement.
Wu Guo, Yanhua Long, Yijie Li 0001, Eryu Wang, Li-Rong Dai 0001
ICASSP6
2009 The I4U system in NIST 2008 speaker recognition evaluation
abstract
This paper describes the performance of the I4U speaker recognition system in the NIST 2008 Speaker Recognition Evaluation. The system consists of seven subsystems, each with different cepstral features and classifiers. We describe the I4U Primary system and report on its core test results as they were submitted, which were among the best-performing submissions. The I4U effort was led by the Institute for Infocomm Research, Singapore (IIR), with contributions from the University of Science and Technology of China (USTC), the University of New South Wales, Australia (UNSW), Nanyang Technological University, Singapore (NTU) and Carnegie Mellon University, USA (CMU).
Haizhou Li 0001, Bin Ma 0001, Kong-Aik Lee, Hanwu Sun, Donglai Zhu, Khe Chai Sim, Chang Huai You, Rong Tong, Ismo Kärkkäinen, Chien-Lin Huang, Vladimir Pervouchine, Wu Guo, Yijie Li 0001, Li-Rong Dai 0001, Mohaddeseh Nosratighods, Tharmarajah Thiruvaran, Julien Epps, Eliathamby Ambikairajah, Chng Eng Siong, Tanja Schultz, Qin Jin
ICASSP14
2009 Exploiting prosodic information for Speaker Recognition
abstract
In this paper, we study speaker characterization using prosodic supervectors with negative within-class covariance normalization (NWCCN) projection and speaker modeling with support vector regression (SVR). We also propose a segmental weight fusion (SWF) technique that combines acoustic and prosodic subsystems effectively, despite the big performance gap between the subsystems. We validate the effectiveness of our proposed techniques on the NIST 2006 Speaker Recognition Evaluation (SRE) in comparison with other prominent solutions. The experiments have reported competitive results of 17.72% Equal Error Rate for the prosodic subsystem alone and 4.50% for the fusion system on NIST 2006 SRE core test condition.
Yanhua Long, Bin Ma 0001, Haizhou Li 0001, Wu Guo, Chng Eng Siong, Li-Rong Dai 0001
ICASSP6
2009 Full covariance state duration modeling for HMM-based speech synthesis
abstract
This paper proposes a state duration modeling method using full covariance matrix for HMM-based speech synthesis. In this method, a full covariance matrix instead of the conventional diagonal covariance matrix is adopted in the multi-dimensional Gaussian distribution to model the state duration of each context-dependent phoneme. At synthesis stage, the state durations are predicted using the clustered context-dependent distributions with full covariance matrices. Experimental results show that the synthesized speech using full-covariance state duration models is more natural than the conventional method when we change the speaking rate of synthesized speech.
Heng Lu 0002, Yi-Jian Wu, Keiichi Tokuda, Li-Rong Dai 0001, Renhua Wang
ICASSP4
2009 An automatic language identification method based on subspace analysis
abstract
Gaussian mixture models (GMM) have become one of the standard acoustic approaches for language identification. Furthermore, the GMM-SVM is proven to work well by introducing the discriminative method into the GMM-based acoustic systems. In these systems, the intersession variability within language has become an important adverse factor that degrades the system performance. To tackle this problem, we propose a subspace analysis method, termed as Intra-language Difference Subspace Estimatio (IDSE), under the GMM-SVM framework. In IDSE method, the difference vector is modeled with three components: Extra-language difference, Intra-language difference and noise difference. Then the Intra-language and noise difference are effectively estimated and eliminated from the difference vector. The experiments on NIST 07 evaluation tasks show effectiveness of the proposed method.
Yan Song 0001, Li-Rong Dai 0001, Renhua Wang
ICME2
2009 Asynchronous F0 and spectrum modeling for HMM-based speech synthesis
Cheng-Cheng Wang, Zhen-Hua Ling, Li-Rong Dai 0001
INTERSPEECH3
2009 Semi-supervised kernel density estimation for video annotation
Meng Wang 0001, Xian-Sheng Hua 0001, Tao Mei 0001, Richang Hong, Guo-Jun Qi, Yan Song 0001, Li-Rong Dai 0001
Comput. Vis. Image Underst.7
2008 Minumum generation error linear regression based model adaptation for HMM-based speech synthesis
abstract
Due to the inconsistency between the maximum likelihood (ML) based training and the synthesis application in HMM-based speech synthesis, a minimum generation error (MGE) criterion had been proposed for HMM training. This paper continues to apply the MGE criterion to model adaptation for HMM-based speech synthesis. We propose a MGE linear regression (MGELR) based model adaptation algorithm, where the regression matrices used to transform source models to target models are optimized to minimize the generation errors for the input speech data uttered by the target speaker. The proposed MGELR approach was compared with the maximum likelihood linear regression (MLLR) based model adaptation. Experimental results indicate that the generation errors were reduced after the MGELR-based model adaptation. And from the subjective listening test, the discrimination and the quality of the synthesized speech using MGELR were better than the results using MLLR.
Yi-Jian Wu, Zhen-Hua Ling, Renhua Wang, Li-Rong Dai 0001
ICASSP5
2008 Minimum generation error criterion considering global/local variance for HMM-based speech synthesis
abstract
Two techniques, including minimum generation error (MGE) criterion for HMM training, and the parameter generation algorithm considering global variance (GV), had been proposed to improve the quality of HMM-based speech synthesis. In this paper, we incorporate the GV technique into MGE criterion, where an additional generation error component considering global/local variance (GV/LV) is introduced for generation error definition, and the model parameters are optimized to minimize the new generation error function. From the experimental results, the quality of synthesized speech was improved after MGE-GV/LV training, which is similar to the effectiveness of considering GV in parameter generation, however, without introducing any extra computational cost in synthesis process.
Yi-Jian Wu, Zhen-Hua Ling, Renhua Wang, Li-Rong Dai 0001
ICASSP5
2007 An Interactive Video Annotation Frameowrk with Multiple Modalities
abstract
Active learning and semi-supervised learning methods are frequently applied in multimedia annotation tasks in order to reduce human labeling effort. However, in most of these methods only single modality is applied. This paper presents an interactive video annotation framework, which is based on semi-supervised learning and active learning with multiple multimodalities. In the proposed framework, unlabeled samples are iteratively selected to be annotated manually according to certain strategy which has taken the potentials of different modalities into account, and then a graph-based semi-supervised learning algorithm is conducted on each modality. This process repeats for several rounds, and the results obtained from multiple modalities are then fused to generate final output. The proposed framework is computationally efficient, and the experimental results on TRECVID 2005 benchmark show that the proposed framework considerably outperforms previous approaches.
Meng Wang 0001, Xian-Sheng Hua 0001, Yan Song 0001, Li-Rong Dai 0001, Renhua Wang
ICASSP (1)4
2007 Lazy Learning Based Efficient Video Annotation
abstract
Eager learning methods, such as SVM, are widely applied in video annotation task for their substantial performance. However, their computational costs are usually prohibitive when a large dataset is faced, especially when annotating a large lexicon of semantic concepts. This paper proposes a video annotation scheme based on lazy learning, and shows that this scheme is much more computationally efficient and flexible. Based on a recently proposed improved Parzen window method, we provide a lazy learning based video annotation scheme. After building the pairwise relationships in dataset, the annotation can be finished rapidly for each concept. Experiments show that the proposed method is much more efficient than SVM while retaining comparable performance.
Meng Wang 0001, Xian-Sheng Hua 0001, Yan Song 0001, Richang Hong, Li-Rong Dai 0001
ICME5
2007 Multi-Graph Semi-Supervised Learning for Video Semantic Feature Extraction
abstract
This paper proposes a video semantic feature extraction approach based on multi-graph semi-supervised learning, which aims to simultaneously deal with the insufficiency of training data and the curse of dimensionality. In contrast to traditional graph-based semi-supervised learning, which generates graph from high-dimensional low-level features, we separate the original low-level features into multiple modalities with minimum correlations, and thus multiple graphs are obtained from these modalities. This way can tackle the curse of dimensionality brought by the high-dimensional feature space. We then propose a criterion to optimally fuse these graphs based on the pairwise relationships among training samples, and implement semi-supervised learning on the fused graph. Experimental results have demonstrated the effectiveness of the proposed approach.
Meng Wang 0001, Xian-Sheng Hua 0001, Xun Yuan 0001, Yan Song 0001, Li-Rong Dai 0001
ICME5
2007 Optimizing multi-graph learning: towards a unified video annotation scheme
abstract
Learning based semantic video annotation is a promising approach for enabling content-based video search. However, severe difficulties, such as insufficiency of training data and curse of dimensionality, are frequently encountered. This paper proposes a novel unified scheme, Optimized Multi-Graph-based Semi-Supervised Learning (OMG-SSL), to simultaneously attack these difficulties. Instead of only using a single graph, OMG-SSL integrates multiple graphs into a regularization and optimization framework to sufficiently explore their complementary nature. We then show that various crucial factors in video annotation, including multiple modalities, multiple distance metrics, and temporal consistency, in fact all correspond to different correlations among samples, and hence they can be represented by different graphs. Therefore, OMG-SSL is able to simultaneously deal with these factors within a unified framework. Experiments on the TRECVID benchmark demonstrate the effectiveness of our proposed approach.
Meng Wang 0001, Xian-Sheng Hua 0001, Xun Yuan 0001, Yan Song 0001, Li-Rong Dai 0001
ACM Multimedia5
2007 Video annotation by graph-based learning with neighborhood similarity
abstract
Graph-based semi-supervised learning methods have been proven effective in tackling the difficulty of training data insufficiency in many practical applications such as video annotation. These methods are all based on an assumption that the labels of similar samples are close. However, as a crucial factor of these algorithms, the estimation of pairwise similarity has not been sufficiently studied. Usually, the similarity of two samples is estimated based on the Euclidean distance between them. But we will show that similarities are not merely related to distances but also related to the structures around the samples. It is shown that distance-based similarity measure may lead to high classification error rates even on several simple datasets. In this paper we propose a novel neighborhood similarity measure, which simultaneously takes into account both thse distance between samples and the difference between the structures around the corresponding samples. Experiments on synthetic dataset and TRECVID benchmark demonstrate that the neighborhood similarity is superior to existing distance based similarity.
Meng Wang 0001, Tao Mei 0001, Xun Yuan 0001, Yan Song 0001, Li-Rong Dai 0001
ACM Multimedia5
2007 An Efficient Automatic Video Shot Size Annotation Scheme
Meng Wang 0001, Xian-Sheng Hua 0001, Yan Song 0001, Li-Rong Dai 0001, Renhua Wang
MMM (1)5
2006 An Automatic Video Semantic Annotation Scheme Based on Combination of Complementary Predictors
abstract
Given a large set of video database, how to connect video segments with a certain set of semantic concepts with least manual labors is an elementary step for video indexing and searching. Due to the large gap between high-level semantics and low-level features, automatic video annotation with high accuracy is a challenging task. In this paper, we propose a novel automatic video annotation framework, which improves the annotation performance by learning from unlabeled samples and exploring temporal relationship in video sequences. To effectively learn from unlabeled data, a sample selection scheme based on combining a set of complementary predictors is proposed, which iteratively refines the performance of the initial predictors. A filtering-based method is applied to further improve the annotation accuracy as well, in which video temporal relationship is sufficiently exploited. Experiment results show that the proposed automatic video annotation method performs superior to general supervised learning methods and co-training
Yan Song 0001, Xian-Sheng Hua 0001, Li-Rong Dai 0001, Meng Wang 0001, Renhua Wang
ICASSP (5)3
2006 Semi-Supervised Kernel Regression
abstract
Insufficiency of training data is a major obstacle in machine learning and data mining applications. Many different semi-supervised learning algorithms have been proposed to tackle this difficulty by leveraging a large amount of unlabeled data. However, most of them focus on semi-supervised classification. In this paper we propose a semi-supervised regression algorithm named semi-supervised kernel regression (SSKR). While classical kernel regression is only based on labeled examples, our approach extends it to all observed examples using a weighting factor to modulate the effect of unlabeled examples. Experimental results prove that SSKR significantly outperforms traditional kernel regression and graph-based semi-supervised regression methods.
Meng Wang 0001, Xian-Sheng Hua 0001, Yan Song 0001, Li-Rong Dai 0001, HongJiang Zhang
ICDM4
2006 Video Annotation by Active Learning and Semi-Supervised Ensembling
abstract
Supervised and semi-supervised learning are frequently applied methods to annotate videos by mapping low-level features into semantic concepts. Due to the large semantic gap, the main constraint of these methods is that the information contained in a limited-size labeled dataset can hardly represent the distributions of the semantic concepts. In this paper, we propose a novel semi-automatic video annotation framework, active learning with semi-supervised ensembling, which tries to tackle the disadvantages of current video annotation solutions. Firstly the initial training set is constructed based on distribution analysis of the entire video dataset and then an active learning scheme is combined into a semi-supervised ensembling framework, which selects the samples to maximize the margin of the ensemble classifier based on both labeled and unlabeled data. Experimental results show that the proposed method performs superior to general semi-supervised learning algorithms and typical active learning algorithms in terms of annotation accuracy and stability
Yan Song 0001, Guo-Jun Qi, Xian-Sheng Hua 0001, Li-Rong Dai 0001, Renhua Wang
ICME4
2006 Enhanced Semi-Supervised Learning for Automatic Video Annotation
abstract
For automatic semantic annotation of large-scale video database, the insufficiency of labeled training samples is a major obstacle. General semi-supervised learning algorithms can help solve the problem but the improvement is limited. In this paper, two semi-supervised learning algorithms, self-training and co-training, are enhanced by exploring the temporal consistency of semantic concepts in video sequences. In the enhanced algorithms, instead of individual shots, time-constraint shot clusters are taken as the basic sample units, in which most mis-classifications can be corrected before they are applied for re-training, thus more accurate statistical models can be obtained. Experiments show that enhanced self-training/co-training significantly improves the performance of video annotation
Meng Wang 0001, Xian-Sheng Hua 0001, Li-Rong Dai 0001, Yan Song 0001
ICME3
2006 Automatic video annotation based on co-adaptation and label correction
abstract
As there is a large gap between high-level semantics and low-level features, it is difficult to obtain high-accuracy video semantic annotation through automatic methods. In this paper, we propose a novel automatic video annotation method, which greatly improves the annotation performance by learning from unlabeled video data, as well as exploring temporal consistency of video sequences. To effectively learn from unlabeled data, a scheme called co-adaptation is proposed to progressively refine two pre-trained complementary classifiers, and then a minimum entropy based method is applied to sufficiently explore the video temporal consistency, which further improves the annotation accuracy. Experiments show that the proposed automatic video annotation method performs superior than both general learning-based and co-training-based methods
Meng Wang 0001, Xian-Sheng Hua 0001, Yan Song 0001, Li-Rong Dai 0001, Shipeng Li 0001
ISCAS4
2005 Sliding Window Smoothing For Maximum Entropy Based Intonational Phrase Prediction In Chinese
abstract
In Chinese TTS (text-to-speech) systems, intonational phrase prediction has a great influence on the naturalness of synthesized speech. Different kinds of statistical models have been applied to this domain, and achieved good performance. We first build a maximum entropy model to yield the probability of each word boundary to be an intonational phrase break, and then a sliding window smoothing algorithm is proposed, in which the length distribution curve of the intonational phrase acts as the sliding window. The maximum entropy model and the distribution curve are trained from 19,000 sentences and tested on a test set of 1,000 sentences. Experimental results show that the sliding window smoothing algorithm makes an improvement of 5.3% in terms of F-score, 10.0% in terms of average score, and 55.6% in terms of unacceptable rate. From the results, we draw the conclusion that the length distribution information is of great usefulness for intonational phrase break prediction, and the sliding window smoothing method is quite effective in improving the performance significantly.
Renhua Wang, Li-Rong Dai 0001
ICASSP (1)4
2005 An Improved Spectral and Prosodic Transformation Method in STRAIGHT-based Voice Conversion
abstract
The paper presents a novel spectral conversion method by considering the glottal effect on the spectrum of the STRAIGHT (speech transformation and representation using adaptive interpolation of weighted spectral contour) speech synthesizer to improve the performance of a former voice conversion system based on codebook mapping. By introducing a MoG (mixture of Gaussians) model into the spectral representation, the STRAIGHT spectrum is decomposed into excitation-dependent and excitation-independent components, which are transformed separately. Besides, an SFC model is adopted to measure the prosodic characteristics of different speakers and realize prosodic conversion. Listening tests prove that the proposed method can effectively improve the discrimination and speech quality of converted speech at the same time.
Gao Peng Chen, Zhen-Hua Ling, Li-Rong Dai 0001
ICASSP (1)4
2004 A complexity reduction of ETSI advanced front-end for DSR
abstract
In October 2002, the advanced front-end (AFE) for distributed speech recognition (DSR) was standardized by ETSI. In order to use the AFE feature on low computational resource devices, we propose a novel approach to improve the computational efficiency. In our new algorithm, the structure of the two-stage mel-warped Wiener filtering algorithm, which is the main part of AFE, is modified. A Wiener filter is constructed and applied directly in the mel-warped filter-bank domain. The measures we take make many time-consuming operations in the original algorithm completely unnecessary, including the re-calculations of power spectrum and the time-domain convolution operations. Consequently, a large amount of computations are saved. Experiments show that the new approach can substantially reduce the computation load while preserving the excellent performance of the ETSI AFE.
Jinyu Li 0001, Renhua Wang, Li-Rong Dai 0001
ICASSP (1)4
2004 A region based multiple frame-rate tradeoff of video streaming
Renhua Wang, Li-Rong Dai 0001, HongJiang Zhang
ICIP4