EDBT 2026 Demo / reviewers in the wild / expert
Ning Cheng 0001
dblp:86/797-1
· DBLP profile ↗
95ranked-venue papers
4as first author
85since 2021 · last 2025
0000-0002-0988-5023ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 55 · 3 first-author · 48 since 2021Graphics, computer vision, multimedia, augmented reality and games · 55 · 3 first-author · 46 since 2021Computer networks · 6 · 6 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GRASP: Replace Redundant Layers with Adaptive Singular Parameters for Efficient Model CompressionabstractRecent studies have demonstrated that many layers are functionally redundant in large language models (LLMs), enabling model compression by removing these layers to reduce inference cost.While such approaches can improve efficiency, indiscriminate layer pruning often results in significant performance degradation.In this paper, we propose GRASP (Gradient-based Retention of Adaptive Singular Parameters), a novel compression framework that mitigates this issue by preserving sensitivity-aware singular values.Unlike direct layer pruning, GRASP leverages gradient-based attribution on a small calibration dataset to adaptively identify and retain critical singular components.By replacing redundant layers with only a minimal set of parameters, GRASP achieves efficient compression while maintaining strong performance with minimal overhead.Experiments across multiple LLMs show that GRASP consistently outperforms existing compression methods, achieving 90% of the original model's performance under 20% compression ratio.The source code is available at https://github.com/LyoAI/GRASP. Kainan Liu, Yong Zhang 0058, Ning Cheng 0001, Zhitao Li 0002, Jing Xiao 0006 |
EMNLP | 3 |
| 2025 | Improving Contextual ASR with Enhanced Phrase-Level Representation Based on MCTC LossabstractContextual biasing is essential for addressing scenario-specific challenges in End-to-End (E2E) Automatic Speech Recognition (ASR) systems. Prior contextual E2E ASR methods, such as the contextual bias with CPP Network, have utilized bias CTC loss for explicit supervision of bias tasks, However, the alignment between the ASR and bias tasks has been largely neglected. This paper introduces a novel contextual biasing approach that employs the Multi-label Synchronous Output CTC (MCTC) algorithm to enhance the synchronization between ASR and bias task outputs. Furthermore, we propose an enhancement to phrase-level contextual representation. Our proposed method demonstrates significant improvements in Word Error Rate (WER). Specifically, our experiments reveal a 15.0% reduction in WER on the Librispeech-960 dataset compared to the CPPN method, with an impressive 29.0% reduction in WER for context phrases. Tao Wei 0003, Ziyang Zhuang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 6 |
| 2025 | Token-Level Contextual Network with Ladder-Shaped Attention for End-to-End ASRabstractContextual automatic speech recognition (ASR) plays an increasingly important role in addressing the long-tail issues of general ASR. In the past, contextual ASR mainly focused on phrase-level discussions, providing a convenient way to handle biasing phrases. This paper introduces a new contextual network for extracting context tokens, with a focus on optimizing contextual knowledge using token-level information. We enhance the fusion of context information and ASR acoustic features, recognizing that textual knowledge inherently represents a distinct modality compared to acoustic information. More importantly, we propose a creative approach to address the challenge of the increasing size of expanding token list compared to phrase list. Our method is tested on the LibriSpeech and AISHELL-2 datasets, the results demonstrate a 31.0%/23.3% Word Error Rate (WER) reduction on LibriSpeech and a 26.9% reduction Character Error Rate (CER) on the named entity (NE) set from AISHELL-2. Tao Wei 0003, Ziyang Zhuang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 6 |
| 2025 | LEF-TTS: Lightweight and Efficient End-to-End Text-to-Speech Synthesis With Multi-Stream GeneratorabstractRecently, the field of Text-to-speech synthesis has been predominantly characterized by end-to-end models, with the quality of speech generated by these models becoming increasingly comparable to that of human speech. In this work, we propose a Lightweight and Efficient Text-to-speech model, a fast end-to-end framework based on EfficientTTS 2 with fully differentiable. We utilize Fast Linear Attention with a Single Head instead of the standard stacked Transformer, which decreases computational complexity and reduces parameters. Additionally, we improve a network architecture ConvWaveNet to further decrease model parameters, and accelerate inference through a multi-stream inverse short-time Fourier Transform generator. These improvements significantly reduce model parameters and increase inference speed, thereby achieving the objectives of faster inference and lightweight modeling. Experimental results show that the proposed model achieves speech quality comparable to that of the baseline models, while also offering improved inference speed and reduced model size. Minchuan Chen, Chenfeng Miao, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 6 |
| 2025 | Self-Enhanced Reasoning Training: Activating Latent Reasoning in Small Models for Enhanced Reasoning DistillationabstractThe rapid advancement of large language models (LLMs) has significantly enhanced their reasoning abilities, enabling increasingly complex tasks. However, these capabilities often diminish in smaller, more computationally efficient models like GPT-2. Recent research shows that reasoning distillation can help small models acquire reasoning capabilities, but most existing methods focus primarily on improving teacher-generated reasoning paths. Our observations reveal that small models can generate high-quality reasoning paths during sampling, even without chain-of-thought prompting, though these paths are often latent due to their low probability under standard decoding strategies. To address this, we propose Self-Enhanced Reasoning Training (SERT), which activates and leverages latent reasoning capabilities in small models through self-training on filtered, self-generated reasoning paths under zero-shot conditions. Experiments using OpenAI’s GPT-3.5 as the teacher model and GPT-2 models as the student models demonstrate that SERT enhances the reasoning abilities of small models, improving their performance in reasoning distillation. Yong Zhang 0058, Zhitao Li 0002, Ming Li 0010, Ning Cheng 0001, Minchuan Chen, Tao Wei 0003, Jun Ma 0018, Jing Xiao 0006 |
ICASSP | 5 |
| 2025 | EffectiveASR: A Single-Step Non-Autoregressive Mandarin Speech Recognition Architecture with High Accuracy and Inference SpeedabstractNon-autoregressive (NAR) automatic speech recognition (ASR) models predict tokens independently and simultaneously, bringing high inference speed. However, there is still a gap in the accuracy of the NAR models compared to the autoregressive (AR) models. In this paper, we propose a single-step NAR ASR architecture with high accuracy and inference speed, called EffectiveASR. It uses an Index Mapping Vector (IMV) based alignment generator to generate alignments during training, and an alignment predictor to learn the alignments for inference. It can be trained end-to-end (E2E) with cross-entropy loss combined with alignment loss. The proposed EffectiveASR achieves competitive results on the AISHELL-1 and AISHELL-2 Mandarin benchmarks compared to the leading models. Specifically, it achieves character error rates (CER) of 4.26%/4.62% on the AISHELL-1 dev/test dataset, which outperforms the AR Conformer with about 30x inference speedup. Ziyang Zhuang, Chenfeng Miao, Tao Wei 0003, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 7 |
| 2025 | Enhancing Serialized Output Training for Multi-Talker ASR with Soft Monotonic Alignment and Utterance-level Timestamp
Fengyun Tan, Tao Wei 0003, Ning Cheng 0001, Jing Xiao 0006 |
INTERSPEECH | 4 |
| 2025 | Towards Efficiently Whisper Fine-tuning with Monotonic Alignments
Ziyang Zhuang, Tao Wei 0003, Ning Cheng 0001, Jing Xiao 0006 |
INTERSPEECH | 4 |
| 2024 | Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-TuningabstractMing Li, Yong Zhang, Shwai He, Zhitao Li, Hongyu Zhao, Jianzong Wang, Ning Cheng, Tianyi Zhou. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Ming Li 0010, Yong Zhang 0058, Shwai He, Zhitao Li 0002, Jianzong Wang, Ning Cheng 0001, Tianyi Zhou 0001 |
ACL (1) | 7 |
| 2024 | Medical Speech Symptoms Classification via Disentangled RepresentationabstractIntent is defined for understanding spoken language in existing works. Both textual features and acoustic features involved in medical speech contain intent, which is important for symptomatic diagnosis. In this paper, we propose a medical speech classification model named DRSC that automatically learns to disentangle intent and content representations from textual-acoustic data for classification. The intent representations of the text domain and the Mel-spectrogram domain are extracted via intent encoders, and then the reconstructed text feature and the Mel-spectrogram feature are obtained through two exchanges. After combining the intent from two domains into a joint representation, the integrated intent representation is fed into a decision layer for classification. Experimental results show that our model obtains an average accuracy rate of 95% in detecting 25 different medical symptoms. Jianzong Wang, Pengcheng Li 0013, Xulong Zhang 0001, Ning Cheng 0001, Jing Xiao 0006 |
CSCWD | 4 |
| 2024 | Learning Disentangled Speech Representations with Contrastive Learning and Time-Invariant RetrievalabstractVoice conversion refers to transferring speaker identity with well-preserved content. Better disentanglement of speech representations leads to better voice conversion. Recent studies have found that phonetic information from input audio has the potential ability to well represent content. Besides, the speaker-style modeling with pre-trained models making the process more complex. To tackle these issues, we introduce an new method named "CTVC" which utilizes disen-tangled speech representations with contrastive learning and time-invariant retrieval.Specifically, a similarity-based compression module is used to facilitate a more intimate connection between the frame-level hidden features and linguistic information at phoneme-level. Additionally, a time-invariant retrieval is proposed for timbre extraction based on multiple segmentation and mutual information. Experimental results demonstrate that "CTVC" outperforms previous studies and improves the sound quality and similarity of converted results. Huaizhen Tang, Xulong Zhang 0001, Ning Cheng 0001, Jing Xiao 0006, Jianzong Wang |
ICASSP | 4 |
| 2024 | ED-TTS: Multi-Scale Emotion Modeling Using Cross-Domain Emotion Diarization for Emotional Speech SynthesisabstractExisting emotional speech synthesis methods often utilize an utterance-level style embedding extracted from reference audio, neglecting the inherent multi-scale property of speech prosody. We introduce ED-TTS, a multi-scale emotional speech synthesis model that leverages Speech Emotion Diarization (SED) and Speech Emotion Recognition (SER) to model emotions at different levels. Specifically, our proposed approach integrates the utterance-level emotion embedding extracted by SER with fine-grained frame-level emotion embedding obtained from SED. These embeddings are used to condition the reverse process of the denoising diffusion probabilistic model (DDPM). Additionally, we employ cross-domain SED to accurately predict soft labels, addressing the challenge of a scarcity of fine-grained emotion-annotated datasets for supervising emotional TTS training. Haobin Tang, Xulong Zhang 0001, Ning Cheng 0001, Jing Xiao 0006, Jianzong Wang |
ICASSP | 3 |
| 2024 | EmoTalker: Emotionally Editable Talking Face Generation via Diffusion ModelabstractIn recent years, the field of talking faces generation has attracted considerable attention, with certain methods adept at generating virtual faces that convincingly imitate human expressions. However, existing methods face challenges related to limited generalization, particularly when dealing with challenging identities. Furthermore, methods for editing expressions are often confined to a singular emotion, failing to adapt to intricate emotions. To overcome these challenges, this paper proposes EmoTalker, an emotionally editable portraits animation approach based on the diffusion model. EmoTalker modifies the denoising process to ensure preservation of the original portrait’s identity during inference. To enhance emotion comprehension from text input, Emotion Intensity Block is introduced to analyze fine-grained emotions and strengths derived from prompts. Additionally, a crafted dataset is harnessed to enhance emotion comprehension within prompts. Experiments show the effectiveness of EmoTalker in generating high-quality, emotionally customizable facial expressions. Xulong Zhang 0001, Ning Cheng 0001, Jun Yu 0001, Jing Xiao 0006, Jianzong Wang |
ICASSP | 3 |
| 2024 | Leveraging Biases in Large Language Models: "bias-kNN" for Effective Few-Shot LearningabstractLarge Language Models (LLMs) have shown significant promise in various applications, including zero-shot and few-shot learning. However, their performance can be hampered by inherent biases. Instead of traditionally sought methods that aim to minimize or correct these biases, this study introduces a novel methodology named "bias-kNN". This approach capitalizes on the biased outputs, harnessing them as primary features for kNN and supplementing with gold labels. Our comprehensive evaluations, spanning diverse domain text classification datasets and different GPT-2 model sizes, indicate the adaptability and efficacy of the "bias-kNN" method. Remarkably, this approach not only outperforms conventional in-context learning in few-shot scenarios but also demonstrates robustness across a spectrum of samples, templates and verbalizers. This study, therefore, presents a unique perspective on harnessing biases, transforming them into assets for enhanced model performance. Yong Zhang 0058, Hanzhang Li, Zhitao Li 0002, Ning Cheng 0001, Ming Li 0010, Jing Xiao 0006, Jianzong Wang |
ICASSP | 4 |
| 2024 | Learning Expressive Disentangled Speech Representations with Soft Speech Units and Adversarial Style AugmentationabstractVoice conversion is the task to transform voice characteristics of source speech while preserving content information. Nowadays, self-supervised representation learning models are increasingly utilized in content extraction. However, in these representations, a lot of hidden speaker information leads to timbre leakage while the prosodic information of hidden units lacks use. To address these issues, we propose a novel framework for expressive voice conversion called “SAVC” based on soft speech units from HuBert-soft. Taking soft speech units as input, we design an attribute encoder to extract content and prosody features respectively. Specifically, we first introduce statistic perturbation imposed by adversarial style augmentation to eliminate speaker information. Then the prosody is implicitly modeled on soft speech units with knowledge distillation. Experiment results show that the intelligibility and naturalness of converted speech outperform previous work. Jianzong Wang, Xulong Zhang 0001, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 4 |
| 2024 | MAIN-VC: Lightweight Speech Representation Disentanglement for One-shot Voice ConversionabstractOne-shot voice conversion aims to change the timbre of any source speech to match that of the unseen target speaker with only one speech sample. Existing methods face difficulties in satisfactory speech representation disentanglement and suffer from sizable networks as some of them leverage numerous complex modules for disentanglement. In this paper, we propose a model named MAIN-VC to effectively disentangle via a concise neural network. The proposed model utilizes Siamese encoders to learn clean representations, further enhanced by the designed mutual information estimator. The Siamese structure and the newly designed convolution module contribute to the lightweight of our model while ensuring performance in diverse voice conversion tasks. The experimental results show that the proposed model achieves comparable subjective scores and exhibits improvements in objective metrics compared to existing methods in a one-shot voice conversion scenario. Pengcheng Li 0013, Jianzong Wang, Xulong Zhang 0001, Yong Zhang 0058, Jing Xiao 0006, Ning Cheng 0001 |
IJCNN | 6 |
| 2024 | EAD-VC: Enhancing Speech Auto-Disentanglement for Voice Conversion with IFUB Estimator and Joint Text-Guided Consistent LearningabstractUsing unsupervised learning to disentangle speech into content, rhythm, pitch, and timbre for voice conversion has become a hot research topic. Existing works generally take into account disentangling speech components through human-crafted bottleneck features which can not achieve sufficient information disentangling, while pitch and rhythm may still be mixed together. There is a risk of information overlap in the disentangling process which results in less speech naturalness. To overcome such limits, we propose a two-stage model to disentangle speech representations in a self-supervised manner without a human-crafted bottleneck design, which uses the Mutual Information (MI) with the designed upper bound estimator (IFUB) to separate overlapping information between speech components. Moreover, we design a Joint Text-Guided Consistent (TGC) module to guide the extraction of speech content and eliminate timbre leakage issues. Experiments show that our model can achieve a better performance than the baseline, regarding disentanglement effectiveness, speech naturalness, and similarity. Audio samples can be found at https://largeaudiomodel.com/eadvc. Ziqi Liang, Jianzong Wang, Xulong Zhang 0001, Yong Zhang 0058, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 5 |
| 2024 | QLSC: A Query Latent Semantic Calibrator for Robust Extractive Question AnsweringabstractExtractive Question Answering (EQA) in Machine Reading Comprehension (MRC) often faces the challenge of dealing with semantically identical but format-variant inputs. Our work introduces a novel approach, called the "Query Latent Semantic Calibrator (QLSC)", designed as an auxiliary module for existing MRC models. We propose a unique scaling strategy to capture latent semantic center features of queries. These features are then seamlessly integrated into traditional query and passage embeddings using an attention mechanism. By deepening the comprehension of the semantic queries-passage relationship, our approach diminishes sensitivity to variations in text format and boosts the model’s capability in pinpointing accurate answers. Experimental results on robust Question-Answer datasets confirm that our approach effectively handles format-variant but semantically identical queries, highlighting the effectiveness and adaptability of our proposed method. Sheng Ouyang, Jianzong Wang, Yong Zhang 0058, Zhitao Li 0002, Ziqi Liang, Xulong Zhang 0001, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 7 |
| 2024 | ConTuner: Singing Voice Beautifying with Pitch and Expressiveness ConditionabstractSinging voice beautifying is a novel task that has application value in people’s daily life, aiming to correct the pitch of the singing voice and improve the expressiveness without changing the original timbre and content. Existing methods rely on paired data or only concentrate on the correction of pitch. However, professional songs and amateur songs from the same person are hard to obtain, and singing voice beautifying doesn’t only contain pitch correction but other aspects like emotion and rhythm. Since we propose a fast and high-fidelity singing voice beautifying system called ConTuner, a diffusion model combined with the modified condition to generate the beautified Mel-spectrogram, where the modified condition is composed of optimized pitch and expressiveness. For pitch correction, we establish a mapping relationship from MIDI, spectrum envelope to pitch. To make amateur singing more expressive, we propose the expressiveness enhancer in the latent space to convert amateur vocal tone to professional. ConTuner achieves a satisfactory beautification effect on both Mandarin and English songs. Ablation study demonstrates that the expressiveness enhancer and generator-based accelerate method in ConTuner are effective. Jianzong Wang, Pengcheng Li 0013, Xulong Zhang 0001, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 4 |
| 2024 | EfficientASR: Speech Recognition Network Compression via Attention Redundancy and Chunk-Level FFN OptimizationabstractIn recent years, Transformer networks have shown remarkable performance in speech recognition tasks. However, their deployment poses challenges due to high computational and storage resource requirements. To address this issue, a lightweight model called EfficientASR is proposed in this paper, aiming to enhance the versatility of Transformer models. EfficientASR employs two primary modules: Shared Residual Multi-Head Attention (SRMHA) and Chunk-Level Feedforward Networks (CFFN). The SRMHA module effectively reduces redundant computations in the network, while the CFFN module captures spatial knowledge and reduces the number of parameters. The effectiveness of the EfficientASR model is validated on two public datasets, namely Aishell-1 and HKUST. Experimental results demonstrate a 36% reduction in parameters compared to the baseline Transformer network, along with improvements of 0.3% and 0.2% in Character Error Rate (CER) on the Aishell-1 and HKUST datasets, respectively. Jianzong Wang, Ziqi Liang, Xulong Zhang 0001, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 4 |
| 2024 | From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction TuningabstractMing Li, Yong Zhang, Zhitao Li, Jiuhai Chen, Lichang Chen, Ning Cheng, Jianzong Wang, Tianyi Zhou, Jing Xiao. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Ming Li 0010, Yong Zhang 0058, Zhitao Li 0002, Jiuhai Chen, Lichang Chen, Ning Cheng 0001, Jianzong Wang, Tianyi Zhou 0001, Jing Xiao 0006 |
NAACL-HLT | 6 |
| 2023 | On the Calibration and Uncertainty with Pólya-Gamma Augmentation for Dialog Retrieval ModelsabstractDeep neural retrieval models have amply demonstrated their power but estimating the reliability of their predictions remains challenging. Most dialog response retrieval models output a single score for a response on how relevant it is to a given question. However, the bad calibration of deep neural network results in various uncertainty for the single score such that the unreliable predictions always misinform user decisions. To investigate these issues, we present an efficient calibration and uncertainty estimation framework PG-DRR for dialog response retrieval models which adds a Gaussian Process layer to a deterministic deep neural network and recovers conjugacy for tractable posterior inference by Pólya-Gamma augmentation. Finally, PG-DRR achieves the lowest empirical calibration error (ECE) in the in-domain datasets and the distributional shift task while keeping R10@1 and MAP performance. Shijing Si, Jianzong Wang, Ning Cheng 0001, Zhitao Li 0002, Jing Xiao 0006 |
AAAI | 4 |
| 2023 | Machine Unlearning Methodology Based on Stochastic Teacher Network
Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Yifu Sun, Chuanyao Zhang, Jing Xiao 0006 |
ADMA (5) | 3 |
| 2023 | Voice Conversion with Denoising Diffusion Probabilistic GAN Models
Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ADMA (4) | 3 |
| 2023 | Symbolic and Acoustic: Multi-domain Music Emotion Modeling for Instrumental Music
Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ADMA (4) | 4 |
| 2023 | PRCA: Fitting Black-Box Large Language Models for Retrieval Question Answering via Pluggable Reward-Driven Contextual AdapterabstractThe Retrieval Question Answering (ReQA) task employs the retrieval-augmented framework, composed of a retriever and generator.The generator formulates the answer based on the documents retrieved by the retriever.Incorporating Large Language Models (LLMs) as generators is beneficial due to their advanced QA capabilities, but they are typically too large to be fine-tuned with budget constraints while some of them are only accessible via APIs.To tackle this issue and further improve ReQA performance, we propose a trainable Pluggable Reward-Driven Contextual Adapter (PRCA), keeping the generator as a black box.Positioned between the retriever and generator in a Pluggable manner, PRCA refines the retrieved information by operating in a tokenautoregressive strategy via maximizing rewards of the reinforcement learning phase.Our experiments validate PRCA's effectiveness in enhancing ReQA performance on three datasets by up to 20% improvement to fit black-box LLMs into existing frameworks, demonstrating its considerable potential in the LLMs era. Zhitao Li 0002, Yong Zhang 0058, Jianzong Wang, Ning Cheng 0001, Ming Li 0010, Jing Xiao 0006 |
EMNLP | 5 |
| 2023 | Improving Music Genre Classification from multi-modal Properties of Music and Genre Correlations PerspectiveabstractMusic genre classification has been widely studied in past few years for its various applications in music information retrieval. Previous works tend to perform unsatisfactorily, since those methods only use audio content or jointly use audio content and lyrics content inefficiently. In addition, as genres normally co-occur in a music track, it is desirable to capture and model the genre correlations to improve the performance of multi-label music genre classification. To solve these issues, we present a novel multi-modal method leveraging audio-lyrics contrastive loss and two symmetric cross-modal attention, to align and fuse features from audio and lyrics. Furthermore, based on the nature of the multi-label classification, a genre correlations extraction module is presented to capture and model potential genre correlations. Extensive experiments demonstrate that our proposed method significantly surpasses other multi-label music genre classification methods and achieves state-of-the-art result on Music4All dataset. Ganghui Ru, Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 4 |
| 2023 | Learning Speech Representations with Flexible Hidden Feature DimensionsabstractNon-parallel many-to-many voice conversion is a kind of style transfer task in speech. Recently, AutoVC has been applied in this field as a popular solution, as it can achieve distribution-matching style transfer by training only the re- construction loss. However, in order to strike a good balance between timbre disentanglement and sound quality, AutoVC requires imposing very strict constraints on the dimensionality of the latent representation. This constraint affects the quality of the converted speech while making it challenging to apply to other datasets directly. This paper proposes a new voice conversion framework that uses only one encoder to obtain timbre and content information by partitioning the latent space in the channel dimension. Furthermore, two different types of classifiers and two additional reconstruction losses are proposed to ensure that different parts of the latent space contain only separated content and timbre information, respectively. Experiments on the VCTK dataset show that the proposed model achieves state-of-the-art results in terms of the naturalness and similarity of converted speech. In addition, we experimentally show that for different division proportions of latent space, the content and timbre information will always be well separated. Huaizhen Tang, Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 4 |
| 2023 | VQ-CL: Learning Disentangled Speech Representations with Contrastive Learning and Vector QuantizationabstractVoice Conversion(VC) refers to converting the voice characteristics of audio to another one as it is said by other people. Recently, more and more studies have focused on disentangle-based VC, which separates the timbre and linguistic content information from an audio signal to effectively achieve VC tasks. However, It’s still challenging to extract phoneme-level features from frame-level hidden representations. This paper proposed a novel zero-shot voice conversion framework that utilizes contrastive learning and vector quantization to encourage the frame-level hidden features closer to the phoneme-level linguistic information, called VQ-CL. All objective and subjective experiment results show that VQ-CL has better performance than previous studies in separating content and voice characteristics to improve the sound quality of generated speech. Huaizhen Tang, Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 4 |
| 2023 | QI-TTS: Questioning Intonation Control for Emotional Speech SynthesisabstractRecent expressive text to speech (TTS) models focus on synthesizing emotional speech, but some fine-grained styles such as intonation are neglected. In this paper, we propose QI-TTS which aims to better transfer and control intonation to further deliver the speaker’s questioning intention while transferring emotion from reference speech. We propose a multi-style extractor to extract style embedding from two different levels. While the sentence level represents emotion, the final syllable level represents intonation. For fine-grained intonation control, we use relative attributes to represent intonation intensity at the syllable level. Experiments have validated the effectiveness of QI-TTS for improving intonation expressiveness in emotional speech synthesis. Haobin Tang, Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 4 |
| 2023 | Efficient Uncertainty Estimation with Gaussian Process for Reliable Dialog Response RetrievalabstractDeep neural networks have achieved remarkable performance in retrieval-based dialogue systems, but they are shown to be ill calibrated. Though basic calibration methods like Monte Carlo Dropout and Ensemble can calibrate well, these methods are time-consuming in the training or inference stages. To tackle these challenges, we propose an efficient uncertainty calibration framework GPF-BERT for BERT-based conversational search, which employs a Gaussian Process layer and the focal loss on top of the BERT architecture to achieve a high-quality neural ranker. Extensive experiments are conducted to verify the effectiveness of our method. In comparison with basic calibration methods, GPF-BERT achieves the lowest empirical calibration error (ECE) in three in-domain datasets and the distributional shift tasks, while yielding the highest R10@1 and MAP performance on most cases. In terms of time consumption, our GPF-BERT has an 8× speedup. Zhitao Li 0002, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 4 |
| 2023 | Dynamic Alignment Mask CTC: Improved Mask CTC With Aligned Cross EntropyabstractBecause of predicting all the target tokens in parallel, the non-autoregressive models greatly improve the decoding efficiency of speech recognition compared with traditional autoregressive models. In this work, we present dynamic alignment Mask CTC, introducing two methods: (1) Aligned Cross Entropy (AXE), finding the monotonic alignment that minimizes the cross-entropy loss through dynamic programming, (2) Dynamic Rectification, creating new training samples by replacing some masks with model predicted tokens. The AXE ignores the absolute position alignment between prediction and ground truth sentence and focuses on tokens matching in relative order. The dynamic rectification method makes the model capable of simulating the non-mask but possible wrong tokens, even if they have high confidence. Our experiments on WSJ dataset demonstrated that not only AXE loss but also the rectification method could improve the WER performance of Mask CTC. Xulong Zhang 0001, Haobin Tang, Jianzong Wang, Ning Cheng 0001, Jian Luo 0007, Jing Xiao 0006 |
ICASSP | 4 |
| 2023 | Improving EEG-based Emotion Recognition by Fusing Time-Frequency and Spatial RepresentationsabstractUsing deep learning methods to classify EEG signals can accurately identify people’s emotions. However, existing studies have rarely considered the application of the information in another domain’s representations to feature selection in the time-frequency domain. We propose a classification network of EEG signals based on the cross-domain feature fusion method, which makes the network more focused on the features most related to brain activities and thinking changes by using the multi-domain attention mechanism. In addition, we propose a two-step fusion method and apply these methods to the EEG emotion recognition network. Experimental results show that our proposed network, which combines multiple representations in the time-frequency domain and spatial domain, outperforms previous methods on public datasets and achieves state-of-the-art at present. Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 4 |
| 2023 | Contrastive Latent Space Reconstruction Learning for Audio-Text RetrievalabstractCross-modal retrieval (CMR) has been extensively applied in various domains, such as multimedia search engines and recommendation systems. Most existing CMR methods focus on image-to-text retrieval, whereas audio-to-text retrieval, a less explored domain, has posed a great challenge due to the difficulty to uncover discriminative features from audio clips and texts. Existing studies are restricted in the following two ways: 1) Most researchers utilize contrastive learning to construct a common subspace where similarities among data can be measured. However, they considers only cross-modal transformation, neglecting the intra-modal separability. Besides, the temperature parameter is not adaptively adjusted along with semantic guidance, which degrades the performance. 2) These methods do not take latent representation reconstruction into account, which is essential for semantic alignment. This paper introduces a novel audio-text oriented CMR approach, termed Contrastive Latent Space Reconstruction Learning (CLSR). CLSR improves contrastive representation learning by taking intra-modal separability into account and adopting an adaptive temperature control strategy. Moreover, the latent representation reconstruction modules are embedded into the CMR framework, which improves modal interaction. Experiments in comparison with some state-of-the-art methods on two audio-text datasets have validated the superiority of CLSR. Kaiyi Luo, Xulong Zhang 0001, Jianzong Wang, Huaxiong Li, Ning Cheng 0001, Jing Xiao 0006 |
ICTAI | 5 |
| 2023 | AOSR-Net: All-in-One Sandstorm Removal NetworkabstractMost existing sandstorm image enhancement methods are based on traditional theory and prior knowledge, which often restrict their applicability in real-world scenarios. In addition, these approaches often adopt a strategy of color correction followed by dust removal, which makes the algorithm structure too complex. To solve the issue, we introduce a novel image restoration model, named all-in-one sandstorm removal network (AOSR-Net). This model is developed based on a re-formulated sandstorm scattering model, which directly establishes the image mapping relationship by integrating intermediate parameters. Such integration scheme effectively addresses the problems of over-enhancement and weak generalization in the field of sand dust image enhancement. Experimental results on synthetic and real-world sandstorm images demonstrate the superiority of the proposed AOSR-Net over state-of-the-art (SOTA) algorithms. Yazhong Si, Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICTAI | 5 |
| 2023 | FastGraphTTS: An Ultrafast Syntax-Aware Speech Synthesis FrameworkabstractThis paper integrates graph-to-sequence into an end-to-end text-to-speech framework for syntax-aware modelling with syntactic information of input text. Specifically, the input text is parsed by a dependency parsing module to form a syntactic graph. The syntactic graph is then encoded by a graph encoder to extract the syntactic hidden information, which is concatenated with phoneme embedding and input to the alignment and flow-based decoding modules to generate the raw audio waveform. The model is experimented on two languages, English and Mandarin, using single-speaker, few samples of target speakers, and multi-speaker datasets, respectively. Experimental results show better prosodic consistency performance between input text and generated audio, and also get higher scores in the subjective prosodic evaluation, and show the ability of voice conversion. Besides, the efficiency of the model is largely boosted through the design of the AI chip operator with 5x acceleration. Jianzong Wang, Xulong Zhang 0001, Aolan Sun, Ning Cheng 0001, Jing Xiao 0006 |
ICTAI | 4 |
| 2023 | SAR: Self-Supervised Anti-Distortion Representation for End-To-End Speech ModelabstractIn recent Text-to-Speech (TTS) systems, a neural vocoder often generates speech samples by solely conditioning on acoustic features predicted from an acoustic model. However, there are always distortions existing in the predicted acoustic features, compared to those of the groundtruth, especially in the common case of poor acoustic modeling due to low-quality training data. To overcome such limits, we propose a Self-supervised learning framework to learn an Anti-distortion acoustic Representation (SAR) to replace human-crafted acoustic features by introducing distortion prior to an auto-encoder pre-training process. The learned acoustic representation from the proposed framework is proved anti-distortion compared to the most commonly used mel-spectrogram through both objective and subjective evaluation. Jianzong Wang, Xulong Zhang 0001, Haobin Tang, Aolan Sun, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 5 |
| 2023 | Boosting Chinese ASR Error Correction with Dynamic Error Scaling Mechanism
Yong Zhang 0058, Hanzhang Li, Jianzong Wang, Zhitao Li 0002, Sheng Ouyang, Ning Cheng 0001, Jing Xiao 0006 |
INTERSPEECH | 7 |
| 2023 | Investigation of Music Emotion Recognition Based on Segmented Semi-Supervised Learning
Yifu Sun, Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Kaiyu Hu, Jing Xiao 0006 |
INTERSPEECH | 4 |
| 2023 | EmoMix: Emotion Mixing via Diffusion Models for Emotional Speech Synthesis
Haobin Tang, Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
INTERSPEECH | 4 |
| 2023 | Prompt Guided Copy Mechanism for Conversational Question AnsweringabstractConversational Question Answering (CQA) is a challenging task that aims to generate natural answers for conversational flow questions.In this paper, we propose a pluggable approach for extractive methods that introduces a novel prompt-guided copy mechanism to improve the fluency and appropriateness of the extracted answers.Our approach uses prompts to link questions to answers and employs attention to guide the copy mechanism to verify the naturalness of extracted answers, making necessary edits to ensure that the answers are fluent and appropriate.The three prompts, including a question-rationale relationship prompt, a question description prompt, and a conversation history prompt, enhance the copy mechanism's performance.Our experiments demonstrate that this approach effectively promotes the generation of natural answers and achieves good results in the CoQA challenge. Yong Zhang 0058, Zhitao Li 0002, Jianzong Wang, Yiming Gao 0010, Ning Cheng 0001, Fengying Yu, Jing Xiao 0006 |
INTERSPEECH | 5 |
| 2023 | PMVC: Data Augmentation-Based Prosody Modeling for Expressive Voice ConversionabstractVoice conversion as the style transfer task applied to speech, refers to converting one person's speech into a new speech that sounds like another person's. Up to now, there has been a lot of research devoted to better implementation of VC tasks. However, a good voice conversion model should not only match the timbre information of the target speaker, but also expressive information such as prosody, pace, pause, etc. In this context, prosody modeling is crucial for achieving expressive voice conversion that sounds natural and convincing. Unfortunately, prosody modeling is important but challenging, especially without text transcriptions. In this paper, we firstly propose a novel voice conversion framework named 'PMVC', which effectively separates and models the content, timbre, and prosodic information from the speech without text transcriptions. Specially, we introduce a new speech augmentation algorithm for robust prosody extraction. And building upon this, mask and predict mechanism is applied in the disentanglement of prosody and content information. The experimental results on the AIShell-3 corpus supports our improvement of naturalness and similarity of converted speech. Huaizhen Tang, Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ACM Multimedia | 5 |
| 2022 | Avqvc: One-Shot Voice Conversion By Vector Quantization With Applying Contrastive LearningabstractVoice Conversion(VC) refers to changing the timbre of a speech while retaining the discourse content. Recently, many works have focused on disentangle-based learning techniques to separate the timbre and the linguistic content information from a speech signal. Once successful, voice conversion will be feasible and straightforward. This paper proposed a novel one-shot voice conversion framework based on vector quantization voice conversion (VQVC) and AutoVC, called AVQVC. A new training method is applied to VQVC to separate content and timbre information from speech more effectively. The result shows that this approach has better performance than VQVC in separating content and timbre to improve the sound quality of generated speech. Huaizhen Tang, Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 4 |
| 2022 | DRVC: A Framework of Any-to-Any Voice Conversion with Self-Supervised LearningabstractAny-to-any voice conversion problem aims to convert voices for source and target speakers, which are out of the training data. Previous works wildly utilize the disentangle-based models. The disentangle-based model assumes the speech consists of content and speaker style information and aims to untangle them to change the style information for conversion. Previous works focus on reducing the dimension of speech to get the content information. But the size is hard to determine to lead to the untangle overlapping problem. We propose the Disentangled Representation Voice Conversion (DRVC) model to address the issue. DRVC model is an end-to-end self-supervised model consisting of the content encoder, timbre encoder, and generator. Instead of the previous work for reducing speech size to get content, we propose a cycle for restricting the disentanglement by the Cycle Reconstruct Loss and Same Loss. The experiments show there is an improvement for converted speech on quality and voice similarity. Qiqi Wang 0005, Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 4 |
| 2022 | VU-BERT: A Unified Framework for Visual DialogabstractThe visual dialog task attempts to train an agent to answer multi-turn questions given an image, which requires the deep understanding of interactions between the image and dialog history. Existing researches tend to employ the modality-specific modules to model the interactions, which might be troublesome to use. To fill in this gap, we propose a unified framework for image-text joint embedding, named VU-BERT, and apply patch projection to obtain vision embedding firstly in visual dialog tasks to simplify the model. The model is trained over two tasks: masked language modeling and next utterance retrieval. These tasks help in learning visual concepts, utterances dependence, and the relationships between these two modalities. Finally, our VU-BERT achieves competitive performance (0.7287 NDCG scores) on VisDial v1.0 Datasets. Shijing Si, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 5 |
| 2022 | Self-Attention for Incomplete Utterance RewritingabstractIncomplete utterance rewriting (IUR) has recently become an essential task in NLP, aiming to complement the incomplete utterance with sufficient context information for comprehension. In this paper, we propose a novel method by directly extracting the coreference and omission relationship from the self-attention weight matrix of the transformer in-stead of word embeddings and edit the original text accordingly to generate the complete utterance. Benefiting from the rich information in the self-attention weight matrix, our method achieved competitive results on public IUR datasets. Yong Zhang 0058, Zhitao Li 0002, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 4 |
| 2022 | nnSpeech: Speaker-Guided Conditional Variational Autoencoder for Zero-Shot Multi-speaker text-to-speechabstractMulti-speaker text-to-speech (TTS) using a few adaption data is a challenge in practical applications. To address that, we propose a zero-shot multi-speaker TTS, named nnSpeech, that could synthesis a new speaker voice without fine-tuning and using only one adaption utterance. Compared with using a speaker representation module to extract the characteristics of new speakers, our method bases on a speaker-guided conditional variational autoencoder and can generate a variable Z, which contains both speaker characteristics and content information. The latent variable Z distribution is approximated by another variable conditioned on reference mel-spectrogram and phoneme. Experiments on the English corpus, Mandarin corpus, and cross-dataset proves that our model could generate natural and similar speech with only one adaption speech. Botao Zhao 0001, Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 4 |
| 2022 | Boosting StarGANs for Voice Conversion with Contrastive Discriminator
Shijing Si, Jianzong Wang, Xulong Zhang 0001, Xiaoyang Qu, Ning Cheng 0001, Jing Xiao 0006 |
ICONIP (2) | 5 |
| 2022 | Blur the Linguistic Boundary: Interpreting Chinese Buddhist Sutra in English via Neural Machine TranslationabstractBuddhism is an influential religion with a long-standing history and profound philosophy. Nowadays, more and more people worldwide aspire to learn the essence of Buddhism, attaching importance to Buddhism dissemination. However, Buddhist scriptures written in classical Chinese are obscure to most people and machine translation applications. For instance, general Chinese-English neural machine translation (NMT) fails in this domain. In this paper, we proposed a novel approach to building a practical NMT model for Buddhist scriptures. The performance of our translation pipeline acquired highly promising results in ablation experiments under three criteria. Denghao Li, Yuqiao Zeng, Jianzong Wang, Lingwei Kong, Zhangcheng Huang 0002, Ning Cheng 0001, Xiaoyang Qu, Jing Xiao 0006 |
ICTAI | 6 |
| 2022 | Pre-Avatar: An Automatic Presentation Generation Framework Leveraging Talking AvatarabstractSince the beginning of the COVID-19 pandemic, remote conferencing and school-teaching have become important tools. The previous applications aim to save the commuting cost with real-time interactions. However, our application is going to lower the production and reproduction costs when preparing the communication materials. This paper proposes a system called Pre-Avatar, generating a presentation video with a talking face of a target speaker with 1 front-face photo and a 3-minute voice recording. Technically, the system consists of three main modules, user experience interface (UEI), talking face module and few-shot text-to-speech (TTS) module. The system firstly clones the target speaker's voice, and then generates the speech, and finally generate an avatar with appropriate lip and head movements. Under any scenario, users only need to replace slides with different notes to generate another new video. The demo has been released here11https://pre-avatar.github.io/ and will be published as free software for use. Aolan Sun, Xulong Zhang 0001, Tiandong Ling, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICTAI | 5 |
| 2022 | Speech Augmentation Based Unsupervised Learning for Keyword SpottingabstractIn this paper, we investigated a speech augmentation based unsupervised learning approach for keyword spotting (KWS) task. KWS is a useful speech application, yet also heavily depends on the labeled data. We designed a CNN-Attention architecture to conduct the KWS task. CNN layers focus on the local acoustic features, and attention layers model the long-time dependency. To improve the robustness of KWS model, we also proposed an unsupervised learning method. The unsupervised loss is based on the similarity between the original and augmented speech features, as well as the audio reconstructing information. Two speech augmentation methods are explored in the unsupervised learning: speed and intensity. The experiments on Google Speech Commands V2 Dataset demonstrated that our CNN-Attention model has competitive results. Moreover, the augmentation based unsupervised learning could further improve the classification accuracy of KWS task. In our experiments, with augmentation based unsupervised learning, our KWS model achieves better performance than other unsupervised methods, such as CPC, APC, and MPC. Jian Luo 0007, Jianzong Wang, Ning Cheng 0001, Haobin Tang, Jing Xiao 0006 |
IJCNN | 3 |
| 2022 | Adaptive Activation Network for Low Resource Multilingual Speech RecognitionabstractLow resource automatic speech recognition (ASR) is a useful but thorny task, since deep learning ASR models usually need huge amounts of training data. The existing models mostly established a bottleneck (BN) layer by pre-training on a large source language, and transferring to the low resource target language. In this work, we introduced an adaptive activation network to the upper layers of ASR model, and applied different activation functions to different languages. We also proposed two approaches to train the model: (1) cross-lingual learning, replacing the activation function from source language to target language, (2) multilingual learning, jointly training the Connectionist Temporal Classification (CTC) loss of each language and the relevance of different languages. Our experiments on IARPA Babel datasets demonstrated that our approaches outperform the from-scratch training and traditional bottleneck feature based methods. In addition, combining the cross-lingual learning and multilingual learning together could further improve the performance of multilingual speech recognition. Jian Luo 0007, Jianzong Wang, Ning Cheng 0001, Zhenpeng Zheng, Jing Xiao 0006 |
IJCNN | 3 |
| 2022 | SUSing: SU-net for Singing Voice SynthesisabstractSinging voice synthesis is a generative task that involves multi-dimensional control of the singing model, including lyrics, pitch, and duration, and includes the timbre of the singer and singing skills such as vibrato. In this paper, we proposed SU-net for singing voice synthesis named SUSing. Synthesizing singing voice is treated as a translation task between lyrics and music score and spectrum. The lyrics and music score information is encoded into a two-dimensional feature representation through the convolution layer. The two-dimensional feature and its frequency spectrum are mapped to the target spectrum in an autoregressive manner through a SU-net network. Within the SU-net the stripe pooling method is used to replace the alternate global pooling method to learn the vertical frequency relationship in the spectrum and the changes of frequency in the time domain. The experimental results on the public dataset Kiritan show that the proposed method can synthesize more natural singing voices. Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 3 |
| 2022 | MDCNN-SID: Multi-scale Dilated Convolution Network for Singer IdentificationabstractMost singer identification methods are processed in the frequency domain, which potentially leads to information loss during the spectral transformation. In this paper, instead of the frequency domain, we propose an end-to-end architecture that addresses this problem in the waveform domain. An en-coder based on Multi-scale Dilated Convolution Neural Networks (MDCNN) was introduced to generate wave embedding from the raw audio signal. Specifically, dilated convolution layers are used in the proposed method to enlarge the receptive field, aiming to extract song-level features. Furthermore, skip connection in the backbone network integrates the multi-resolution acoustic features learned by the stack of convolution layers. Then, the obtained wave embedding is passed into the following networks for singer identification. In experiments, the proposed method achieves comparable performance on the benchmark dataset of Artist20, which significantly improves related works. Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 3 |
| 2022 | TDASS: Target Domain Adaptation Speech Synthesis Framework for Multi-speaker Low-Resource TTSabstractRecently, synthesizing personalized speech by text-to-speech (TTS) application is highly demanded. But the previous TTS models require a mass of target speaker speeches for training. It is a high-cost task, and hard to record lots of utterances from the target speaker. Data augmentation of the speeches is a solution but leads to the low-quality synthesis speech problem. Some multi-speaker TTS models are proposed to address the issue. But the quantity of utterances of each speaker imbalance leads to the voice similarity problem. We propose the Target Domain Adaptation Speech Synthesis Network (TDASS) to address these issues. Based on the backbone of the Tacotron2 model, which is the high-quality TTS model, TDASS introduces a self-interested classifier for reducing the non-target influence. Besides, a special gradient reversal layer with different operations for target and non-target is added to the classifier. We evaluate the model on a Chinese speech corpus, the experiments show the proposed method outperforms the baseline method in terms of voice quality and voice similarity. Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 3 |
| 2022 | Singer Identification for Metaverse with Timbral and Middle-Level Perceptual FeaturesabstractMetaverse is an interactive world that combines reality and virtuality, where participants can be virtual avatars. Anyone can hold a concert in a virtual concert hall, and users can quickly identify the real singer behind the virtual idol through the singer identification. Most singer identification methods are processed using the frame-level features. However, expect the singer's timbre, the music frame includes music information, such as melodiousness, rhythm, and tonal. It means the music information is noise for using frame-level features to identify the singers. In this paper, instead of only the frame-level features, we propose to use another two features that address this problem. Middle-level feature, which represents the music's melodiousness, rhythmic stability, and tonal stability, and is able to capture the perceptual features of music. The timbre feature, which is used in speaker identification, represents the singers' voice features. Furthermore, we propose a convolutional recurrent neural network (CRNN) to combine three features for singer identification. The model firstly fuses the frame-level feature and timbre feature and then combines middle-level features to the mix features. In experiments, the proposed method achieves comparable performance on an average F1 score of 0.81 on the benchmark dataset of Artist20, which significantly improves related works. Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 3 |
| 2022 | MetaSID: Singer Identification with Domain Adaptation for MetaverseabstractMetaverse has stretched the real world into unlimited space. There will be more live concerts in Metaverse. The task of singer identification is to identify the song belongs to which singer. However, there has been a tough problem in singer identification, which is the different live effects. The studio version is different from the live version, the data distribution of the training set and the test set are different, and the performance of the classifier decreases. This paper proposes the use of the domain adaptation method to solve the live effect in singer identification. Three methods of domain adaptation combined with Convolutional Recurrent Neural Network (CRNN) are designed, which are Maximum Mean Discrepancy (MMD), gradient reversal (Revgrad), and Contrastive Adaptation Network (CAN). MMD is a distance-based method, which adds domain loss. Revgrad is based on the idea that learned features can represent different domain samples. CAN is based on class adaptation, it takes into account the correspondence between the categories of the source domain and target domain. Experimental results on the public dataset of Artist20 show that CRNN-MMD leads to an improvement over the baseline CRNN by 0.14. The CRNN-RevGrad outperforms the baseline by 0.21. The CRNN-CAN achieved state of the art with the F1 measure value of 0.83 on album split. Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 3 |
| 2022 | Tiny-Sepformer: A Tiny Time-Domain Transformer Network For Speech SeparationabstractTime-domain Transformer neural networks have proven their superiority in speech separation tasks.However, these models usually have a large number of network parameters, thus often encountering the problem of GPU memory explosion.In this paper, we proposed Tiny-Sepformer, a tiny version of Transformer network for speech separation.We present two techniques to reduce the model parameters and memory consumption: (1) Convolution-Attention (CA) block, spliting the vanilla Transformer to two paths, multi-head attention and 1D depthwise separable convolution, (2) parameter sharing, sharing the layer parameters within the CA block.In our experiments, Tiny-Sepformer could greatly reduce the model size, and achieves comparable separation performance with vanilla Sepformer on WSJ0-2/3Mix datasets. Jian Luo 0007, Jianzong Wang, Ning Cheng 0001, Edward Xiao, Xulong Zhang 0001, Jing Xiao 0006 |
INTERSPEECH | 3 |
| 2022 | Speech Representation Disentanglement with Adversarial Mutual Information Learning for One-shot Voice Conversion
Methawee Tantrawenith, Haolin Zhuang, Zhiyong Wu 0001, Aolan Sun, Jianzong Wang, Ning Cheng 0001, Huaizhen Tang, Xintao Zhao, Helen M. Meng |
INTERSPEECH | 7 |
| 2022 | Uncertainty Calibration for Deep Audio ClassifiersabstractAlthough deep Neural Networks (DNNs) have achieved tremendous success in audio classification tasks, their uncertainty calibration are still under-explored.A well-calibrated model should be accurate when it is certain about its prediction and indicate high uncertainty when it is likely to be inaccurate.In this work, we investigate the uncertainty calibration for deep audio classifiers.In particular, we empirically study the performance of popular calibration methods: (i) Monte Carlo Dropout, (ii) ensemble, (iii) focal loss, and (iv) spectralnormalized Gaussian process (SNGP), on audio classification datasets.To this end, we evaluate (i-iv) for the tasks of environment sound and music genre classification.Results indicate that uncalibrated deep audio classifiers may be over-confident, and SNGP performs the best and is very efficient on the two datasets of this paper. Shijing Si, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
INTERSPEECH | 4 |
| 2022 | Adapitch: Adaption Multi-Speaker Text-to-Speech Conditioned on Pitch Disentangling with Untranscribed DataabstractIn this paper, we proposed Adapitch, a multi-speaker TTS method that makes adaptation of the supervised module with untranscribed data. We design two self supervised modules to train the text encoder and mel decoder separately with untranscribed data to enhance the representation of text and mel. To better handle the prosody information in a synthesized voice, a supervised TTS module is designed conditioned on content disentangling of pitch, text, and speaker. The training phase was separated into two parts, pretrained and fixed the text encoder and mel decoder with unsupervised mode, then the supervised mode on the disentanglement of TTS. Experiment results show that the Adaptich achieved much better quality than baseline methods. Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
MSN | 3 |
| 2022 | MetaSpeech: Speech Effects Switch Along with Environment for MetaverseabstractMetaverse expands the physical world to a new dimension, and the physical environment and Metaverse environment can be directly connected and entered. Voice is an indispensable communication medium in the real world and Metaverse. Fusion of the voice with environment effects is important for user immersion in Metaverse. In this paper, we proposed using the voice conversion based method for the conversion of target environment effect speech. The proposed method was named MetaSpeech, which introduces an environment effect module containing an effect extractor to extract the environment information and an effect encoder to encode the environment effect condition, in which gradient reversal layer was used for adversarial training to keep the speech content and speaker information while disentangling the environmental effects. From the experiment results on the public dataset of LJSpeech with four environment effects, the proposed model could complete the specific environment effect conversion and outperforms the baseline methods from the voice conversion task. Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
MSN | 3 |
| 2022 | Semi-Supervised Learning Based on Reference Model for Low-resource TTSabstractMost previous neural text-to-speech (TTS) methods are mainly based on supervised learning methods, which means they depend on a large training dataset and hard to achieve comparable performance under low-resource conditions. To ad-dress this issue, we propose a semi-supervised learning method for neural TTS in which labeled target data is limited, which can also resolve the problem of exposure bias in the previous auto-regressive models. Specifically, we pre-train the reference model based on Fastspeech2 with much source data, fine-tuned on a limited target dataset. Meanwhile, pseudo labels generated by the original reference model are used to guide the fine-tuned model's training further, achieve a regularization effect, and reduce the overfitting of the fine-tuned model during training on the limited target data. Experimental results show that our proposed semi-supervised learning scheme with limited target data significantly improves the voice quality for test data to achieve naturalness and robustness in speech synthesis. Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
MSN | 3 |
| 2022 | Improving Imbalanced Text Classification with Dynamic Curriculum LearningabstractRecent advances in pre-trained language models have improved the performance for text classification tasks. However, little attention is paid to the priority scheduling strategy on the samples during training. Humans acquire knowledge gradually from easy to complex concepts, and the difficulty of the same material can also vary significantly in different learning stages. Inspired by this insights, we proposed a novel self-paced dynamic curriculum learning (SPDCL) method for imbalanced text classification, which evaluates the sample difficulty by both linguistic character and model capacity. Meanwhile, rather than using static curriculum learning as in the existing research, our SPDCL can reorder and resample training data by difficulty criterion with an adaptive from easy to hard pace. The extensive experiments on several classification tasks show the effectiveness of SPDCL strategy, especially for the imbalanced dataset. Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
MSN | 3 |
| 2022 | Improving Speech Representation Learning via Speech-level and Phoneme-level Masking ApproachabstractRecovering the masked speech frames is widely applied in speech representation learning. However, most of these models use random masking in the pre-training. In this work, we proposed two kinds of masking approaches: (1) speech-level masking, making the model to mask more speech segments than silence segments, (2) phoneme-level masking, forcing the model to mask the whole frames of the phoneme, instead of phoneme pieces. We pre-trained the model via these two approaches, and evaluated on two downstream tasks, phoneme classification and speaker recognition. The experiments demonstrated that the proposed masking approaches are beneficial to improve the performance of speech representation. Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
MSN | 3 |
| 2022 | Linguistic-Enhanced Transformer with CTC Embedding for Speech RecognitionabstractThe recent emergence of joint CTC-Attention model shows significant improvement in automatic speech recognition (ASR). The improvement largely lies in the modeling of linguistic information by decoder. The decoder joint-optimized with an acoustic encoder renders the language model from ground-truth sequences in an auto-regressive manner during training. However, the training corpus of the decoder is limited to the speech transcriptions, which is far less than the corpus needed to train an acceptable language model. This leads to poor robustness of decoder. To alleviate this problem, we propose linguistic-enhanced transformer, which introduces refined CTC information to decoder during training process, so that the decoder can be more robust. Our experiments on AISHELL-1 speech corpus show that the character error rate (CER) is relatively reduced by up to 7 %. We also find that in joint CTC-Attention ASR model, decoder is more sensitive to linguistic information than acoustic information. Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
MSN | 3 |
| 2021 | Reconstructing Dual Learning for Neural Voice Conversion Using Relatively Few SamplesabstractThis paper introduces a dual learning system for neural voice conversion (DualVC) using relatively few samples based on the symmetry of the speech conversion task. The system contains a pair of sequence-to-sequence neural networks that have the same structure but are trained in opposite directions. The objective function of the dual model training is the sum of paired conversion loss and reconstruction loss during the dual training circle. The models in the two directions are trained alternately to guide each other by the corresponding reconstruction loss. Furthermore, curriculum learning techniques are used to load models in existing fields into the current task to accelerate the rapid iteration and convergence of the model. The experiment on the voice conversion task with the proposed DualVC and curriculum learning strategy obtained a comparable naturalness and similarity with only a 30% dataset than the BaseVC model trained on the full dataset. Aolan Sun, Jianzong Wang, Ning Cheng 0001, Methawee Tantrawenith, Zhiyong Wu 0001, Helen M. Meng, Edward Xiao, Jing Xiao 0006 |
ASRU | 3 |
| 2021 | TGAVC: Improving Autoencoder Voice Conversion with Text-Guided and Adversarial TrainingabstractNon-parallel many-to-many voice conversion remains an interesting but challenging speech processing task. Recently, AutoVC, a conditional autoencoder based method, achieved excellent conversion results by disentangling the speaker identity and the speech content using information-constraining bottlenecks. However, due to the pure autoencoder training method, it is difficult to evaluate the separation effect of content and speaker identity. In this paper, a novel voice conversion framework, named Text Guided AutoVC(TGAVC), is proposed to more effectively separate content and timbre from speech, where an expected content embedding produced based on the text transcriptions is designed to guide the extraction of voice content. In addition, the adversarial training is applied to eliminate the speaker identity information in the estimated content embedding extracted from speech. Under the guidance of the expected content embedding and the adversarial training, the content encoder is trained to extract speaker-independent content embedding from speech. Experiments on AIShell-3 dataset show that the proposed model outperforms AutoVC in terms of naturalness and similarity of converted speech. Huaizhen Tang, Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Edward Xiao, Jing Xiao 0006 |
ASRU | 4 |
| 2021 | Cyclegean: Cycle Generative Enhanced Adversarial Network for Voice ConversionabstractCycle Generative Adversarial Network (CycleGAN) for voice conversion (VC) task only used discriminators to identify whether the input voice is generated or real. It means the confrontational does not check the similarity with the target voice, leading the generated voice not much similar to the target. In this paper, instead of vocal checking, we propose to enhance the confrontation to target similarity checking that addresses this problem. A Cycle Generative Enhanced Adversarial Network (CycleGEAN) was introduced to make the original two discriminators to target classifier and non-target classifier. The target classifier aims to identify whether the target speaks the input voice or not. Similarly, the non-target classifier identifies the non-target voice. Furthermore, we add a gradient reversal layer with different operations for target and non-target. Then in each GAN, we used both classifiers. One is the discriminator, and the other is trained for using in another GAN. In experiments, the proposed method compare to CycleGAN improves Mean Opinion Score (MOS) of 0.1 and Voice Similarity Score (VSS) of 0.2 on the Voice Conversion Challenge 2018 (VCC2018) dataset. Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Edward Xiao, Jing Xiao 0006 |
ASRU | 3 |
| 2021 | Joint Intent Detection and Slot Filling Based on Continual Learning ModelabstractSlot filling and intent detection have become a significant theme in the field of natural language understanding. Even though slot filling is intensively associated with intent detection, the characteristics of the information required for both tasks are different while most of those approaches may not fully aware of this problem. In addition, balancing the accuracy of two tasks effectively is an inevitable problem for the joint learning model. In this paper, a Continual Learning Interrelated Model (CLIM) is proposed to consider semantic information with different characteristics and balance the accuracy between intent detection and slot filling effectively. The experimental results show that CLIM achieves state-of-the-art performace on slot filling and intent detection on ATIS and Snips. Yanfei Hui, Jianzong Wang, Ning Cheng 0001, Fengying Yu, Tianbo Wu, Jing Xiao 0006 |
ICASSP | 3 |
| 2021 | Unidirectional Memory-Self-Attention Transducer for Online Speech RecognitionabstractSelf-attention models have been successfully applied in end-to-end speech recognition systems, which greatly improve the performance of recognition accuracy. However, such attention-based models cannot be used in online speech recognition, because these models usually have to utilize a whole acoustic sequences as inputs. A common method is restricting the field of attention sights by a fixed left and right window, which makes the computation costs manageable yet also introduces performance degradation. In this paper, we propose Memory-Self-Attention (MSA), which adds history information into the Restricted-Self-Attention unit. MSA only needs localtime features as inputs, and efficiently models long temporal contexts by attending memory states. Meanwhile, recurrent neural network transducer (RNN-T) has proved to be a great approach for online ASR tasks, because the alignments of RNN-T are local and monotonic. We propose a novel network structure, called Memory-Self-Attention (MSA) Transducer. Both encoder and decoder of the MSA Transducer contain the proposed MSA unit. The experiments demonstrate that our proposed models improve WER results than Restricted-Self-Attention models by 13.5% on WSJ and 7.1% on SWBD datasets relatively, and without much computation costs increase. Jian Luo 0007, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 3 |
| 2021 | LVCNet: Efficient Condition-Dependent Modeling Network for Waveform GenerationabstractIn this paper, we propose a novel conditional convolution network, named location-variable convolution, to model the dependencies of the waveform sequence. Different from the use of unified convolution kernels in WaveNet to capture the dependencies of arbitrary waveform, the location-variable convolution uses convolution kernels with different coefficients to perform convolution operations on different waveform intervals, where the coefficients of kernels is predicted according to conditioning acoustic features, such as Mel-spectrograms. Based on location-variable convolutions, we design LVCNet for waveform generation, and apply it in Parallel WaveGAN to design more efficient vocoder. Experiments on the LJSpeech dataset show that our proposed model achieves a four-fold increase in synthesis speed compared to the original Parallel WaveGAN without any degradation in sound quality, which verifies the effectiveness of location-variable convolutions. Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 3 |
| 2021 | Cross-Language Transfer Learning and Domain Adaptation for End-to-End Automatic Speech RecognitionabstractIn this paper, we demonstrate the efficacy of transfer learning and continuous learning for various automatic speech recognition (ASR) tasks using end-to-end models trained with CTC loss. We start with a large pre-trained English ASR model and show that transfer learning can be effectively and easily performed on: (1) different English accents, (2) different languages (from English to German, Spanish, Russian, or from Mandarin to Cantonese) and (3) application-specific domains. Our extensive set of experiments demonstrate that in all three cases, transfer learning from a good base model has higher accuracy than a model trained from scratch. Our results indicate that, for fine-tuning, larger pre-trained models are better than small pre-trained models, even if the dataset for fine-tuning is small. We also show that transfer learning significantly speeds up convergence, which could result in significant cost savings when training with large datasets. Jian Luo 0007, Jianzong Wang, Ning Cheng 0001, Edward Xiao, Jing Xiao 0006, Georg Kucsko, Patrick K. O'Neill, Jagadeesh Balam, Slyne Deng, Adriana Flores, Boris Ginsburg, Jocelyn Huang, Oleksii Kuchaiev, Vitaly Lavrukhin, Jason Li 0007 |
ICME | 3 |
| 2021 | Loss Prediction: End-to-End Active Learning Approach For Speech RecognitionabstractEnd-to-end speech recognition systems usually require huge amounts of labeling resource, while annotating the speech data is complicated and expensive. Active learning is the solution by selecting the most valuable samples for annotation. In this paper, we proposed to use a predicted loss that estimates the uncertainty of the sample. The CTC (Connectionist Temporal Classification) and attention loss are informative for speech recognition since they are computed based on all decoding paths and alignments. We defined an end-to-end active learning pipeline, training an ASR/LP (Automatic Speech Recognition/Loss Prediction) joint model. The proposed approach was validated on an English and a Chinese speech recognition task. The experiments show that our approach achieves competitive results, outperforming random selection, least confidence, and estimated loss method. Jian Luo 0007, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 3 |
| 2021 | Transfer Ability of Monolingual Wav2vec2.0 for Low-resource Speech RecognitionabstractRecently, there are several domains that have their own feature extractors, such as ResNet, BERT, and GPT-x, which are widely used for various down-stream tasks. These models are pre-trained on large amounts of unlabeled data by self-supervision. In the speech domain, wav2vec2.0 starts to show its powerful representation ability and feasibility for ultra-low resource speech recognition tasks. This speech feature extractor is pre-trained on the monolingual audiobook corpus, whereas it has not been thoroughly examined in real spoken scenarios and other languages. In this work, we endeavor to transfer the knowledge from the pre-trained monolingual wav2vec2.0 to cross-lingual spoken ASR tasks with less than 20 hours of labeled data. We achieve more than 20% relative improvements in all the six languages compared with previous methods, establishing a strong benchmark on CALLHOME datasets. Compared with supervised pre-training, self-supervision training used in wav2vec2.0 has a better transfer ability. We also find that using coarse-grained modeling units, such as subword or character, usually achieves better results than fine-grained modeling units, such as phone or letter. Jianzong Wang, Ning Cheng 0001, Bo Xu 0002 |
IJCNN | 3 |
| 2021 | A Language Model Based Pseudo-Sample Deliberation for Semi-supervised Speech RecognitionabstractEnd-to-end modeling requires tremendous amounts of transcribed speech to achieve an automatic speech recognition (ASR) model with high performance. For low-resource ASR tasks, it is a promising approach to utilize the highly accessible unlabeled speech and text corpus. Previous works have shown that training with pseudo samples, which are the inferring results given the unlabeled speech, can substantially improve the accuracy of a baseline ASR model. Besides the common data filtering to improve pseudo-label quality, we propose an alternative pseudo-sample deliberation method that operates on the output of the ASR model through a pre-trained bidirectional language model (BERT). It fixes the unreasonable tokens in the inference by substitution, which can distill knowledge from the large text corpus. Experiments on Librispeech show that assisted with our fixing operation, self-training on additional unlabeled samples can bridge up to 82.3 % of the gap with the supervised training. Jianzong Wang, Ning Cheng 0001, Bo Xu 0002 |
IJCNN | 3 |
| 2021 | Semantic Extraction for Sentence Representation via Reinforcement LearningabstractMany modern Natural Language Processing(NLP) systems rely on word embeddings, previously trained in an unsupervised manner on large corpora, as base features. Several attempts of learning unsupervised sentence representations have not achieved satisfactory performance and have not been widely adopted. In this work, we present a Semantic Extraction Reinforcement learning Model(SERM) for encoding sentences into embedding vectors, which transforms the learning process into intent detection and named entity recognition tasks. Intent detection is mainly related to the semantics of the whole sentence, while named entity recognition pays more attention to local entities. Given a sentence, SERM builds a sentence representation by extracting the most important words and removing irrelevant words in a sentence. Unlike the mainstream approach, SERM compresses sentences from semantic and entity perspectives and this allows us to efficiently learn different types of encoding functions. The experimental results show that SERM can learn high quality sentence representation. This paper demonstrates that our sentence representations have sufficient competitiveness with the best performing model on text classification task. Fengying Yu, Dewei Tao, Jianzong Wang, Yanfei Hui, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 5 |
| 2021 | CACnet: Cube Attentional CNN for Automatic Speech RecognitionabstractEnd-to-end models have been widely used in Automatic Speech Recognition (ASR). Convolutional Neural Networks (CNNs) can effectively use spectrum information to model acoustic models. However, the convolution layers have limitations on the receptive field leading to restrictions for long speech signals. Inspired by this, we propose a Cube Attention CNN network(CACnet) that uses two different attention blocks to integrate the feature information of different dimensions for extending context information. Thereinto, the Global Deep Attention Block utilizes non-local operations to compute interactions between any two positions on feature maps and enables the acquirement of global feature representations while the Cross-Channel Attention Block adaptively recalibrates channel-wise feature responses. Then, outputs of the above two attention modules will be added up to further improve the feature representation which contributes to enrich contextual information. Finally, the performance of our proposed architecture will be explored under ASR tasks in English circumstances. Experiments on LibriSpeech indicate that CACnet achieves a word error rate (WER) of 3.78%/9.56% without language model (LM), and 2.84%/6.97% with LM, which is near state-of-the-art accuracy. CACnet on WSJ with 4.4% WER obtains better performance, compared to CTC-based CNN models, such as QuartzNet and Jasper, with the same language model. The proposed network achieves competitive accuracy while having fewer parameters. Moreover, CACnet can be easily incorporated into any existed network since it has the same input and output dimensions. Jianzong Wang, Xiaoyang Qu, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 5 |
| 2021 | Dropout Regularization for Self-Supervised Learning of Transformer Encoder Speech RepresentationabstractPredicting the altered acoustic frames is an effective way of self-supervised learning for speech representation.However, it is challenging to prevent the pretrained model from overfitting.In this paper, we proposed to introduce two dropout regularization methods into the pretraining of transformer encoder: (1) attention dropout, (2) layer dropout.Both of the two dropout methods encourage the model to utilize global speech information, and avoid just copying local spectrum features when reconstructing the masked frames.We evaluated the proposed methods on phoneme classification and speaker recognition tasks.The experiments demonstrate that our dropout approaches achieve competitive results, and improve the performance of classification accuracy on downstream tasks. Jian Luo 0007, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
Interspeech | 3 |
| 2021 | Speech2Video: Cross-Modal Distillation for Speech to Video GenerationabstractThis paper investigates a novel task of talking face video generation solely from speeches.The speech-to-video generation technique can spark interesting applications in entertainment, customer service, and human-computer-interaction industries.Indeed, the timbre, accent and speed in speeches could contain rich information relevant to speakers' appearance.The challenge mainly lies in disentangling the distinct visual attributes from audio signals.In this article, we propose a light-weight, cross-modal distillation method to extract disentangled emotional and identity information from unlabelled video inputs.The extracted features are then integrated by a generative adversarial network into talking face video clips.With carefully crafted discriminators, the proposed framework achieves realistic generation results.Experiments with observed individuals demonstrated that the proposed framework captures the emotional expressions solely from speeches, and produces spontaneous facial motion in the video output.Compared to the baseline method where speeches are combined with a static image of the speaker, the results of the proposed framework is almost indistinguishable.User studies also show that the proposed method outperforms the existing algorithms in terms of emotion expression in the generated videos. Shijing Si, Jianzong Wang, Xiaoyang Qu, Ning Cheng 0001, Xinghua Zhu, Jing Xiao 0006 |
Interspeech | 4 |
| 2021 | Variational Information Bottleneck for Effective Low-Resource Audio ClassificationabstractLarge-scale deep neural networks (DNNs) such as convolutional neural networks (CNNs) have achieved impressive performance in audio classification for their powerful capacity and strong generalization ability. However, when training a DNN model on low-resource tasks, it is usually prone to overfitting the small data and learning too much redundant information. To address this issue, we propose to use variational information bottleneck (VIB) to mitigate overfitting and suppress irrelevant information. In this work, we conduct experiments on a 4-layer CNN. However, the VIB framework is ready-to-use and could be easily utilized with many other state-of-the-art network architectures. Evaluation on a few audio datasets shows that our approach significantly outperforms baseline methods, yielding _ 5:0% improvement in terms of classification accuracy in some low-source settings. Copyright © 2021 ISCA. Shijing Si, Jianzong Wang, Huiming Sun, Jianhan Wu 0001, Chuanyao Zhang, Xiaoyang Qu, Ning Cheng 0001, Jing Xiao 0006 |
Interspeech | 7 |
| 2021 | Multi-Quartznet: Multi-Resolution Convolution for Speech Recognition with Multi-Layer Feature FusionabstractIn this paper, we propose an end-to-end speech recognition network based on Nvidia's previous QuartzNet [1] model. We try to promote the model performance, and design three components: (1) Multi-Resolution Convolution Module, re-places the original 1D time-channel separable convolution with multi-stream convolutions. Each stream has a unique dilated stride on convolutional operations. (2) Channel-Wise Attention Module, calculates the attention weight of each convolutional stream by spatial channel-wise pooling. (3) Multi-Layer Feature Fusion Module, reweights each convolutional block by global multi-layer feature maps. Our experiments demonstrate that Multi-QuartzNet model achieves CER 6.77% on AISHELL-1 data set, which outperforms original QuartzNet and is close to state-of-art result. Jian Luo 0007, Jianzong Wang, Ning Cheng 0001, Guilin Jiang, Jing Xiao 0006 |
SLT | 3 |
| 2021 | End-To-End Silent Speech Recognition with Acoustic SensingabstractSilent speech interfaces (SSI) has been an exciting area of recent interest. In this paper, we present a non-invasive silent speech interface that uses inaudible acoustic signals to capture people's lip movements when they speak. We exploit the speaker and microphone of the smartphone to emit signals and listen to their reflections, respectively. The extracted phase features of these reflections are fed into the deep learning networks to recognize speech. And we also propose an end-to-end recognition framework, which combines the CNN and attention-based encoder-decoder network. Evaluation results on a limited vocabulary (54 sentences) yield word error rates of 8.4% in speaker-independent and environment-independent settings, and 8.1% for unseen sentence testing. Jian Luo 0007, Jianzong Wang, Ning Cheng 0001, Guilin Jiang, Jing Xiao 0006 |
SLT | 3 |
| 2021 | GraphPB: Graphical Representations of Prosody Boundary in Speech SynthesisabstractThis paper introduces a graphical representation approach of prosody boundary (GraphPB) in the task of Chinese speech synthesis, intending to parse the semantic and syntactic relationship of input sequences in a graphical domain for improving the prosody performance. The nodes of the graph embedding are formed by prosodic words, and the edges are formed by the other prosodic boundaries, namely prosodic phrase boundary (PPH) and intonation phrase boundary (IPH). Different Graph Neural Networks (GNN) like Gated Graph Neural Network (GGNN) and Graph Long Short-term Memory (G-LSTM) are utilised as graph encoders to exploit the graphical prosody boundary information. Graph-to-sequence model is proposed and formed by a graph encoder and an attentional decoder. Two techniques are proposed to embed sequential information into the graph-to-sequence text-to-speech model. The experimental results show that this proposed approach can encode the phonetic and prosody rhythm of an utterance. The mean opinion score (MOS) of these GNN models shows comparative results with the state-of-the-art sequence-to-sequence models with better performance in the aspect of prosody. This provides an alternative approach for prosody modelling in end-to-end speech synthesis. Aolan Sun, Jianzong Wang, Ning Cheng 0001, Huayi Peng, Lingwei Kong, Jing Xiao 0006 |
SLT | 3 |
| 2021 | MelGlow: Efficient Waveform Generative Network Based On Location-Variable ConvolutionabstractRecent neural vocoders usually use a WaveNet-like network to capture the long-term dependencies of the waveform, but a large number of parameters are required to obtain good modeling capabilities. In this paper, an efficient network, named location-variable convolution, is proposed to model the dependencies of waveforms. Different from the use of unified convolution kernels in WaveNet to capture the dependencies of arbitrary waveforms, location-variable convolutions utilizes a kernel predictor to generate multiple sets of convolution kernels based on the melspectrum, where each set of convolution kernels is used to perform convolution operations on the associated waveform intervals. Combining WaveGlow and location-variable convolutions, an efficient vocoder, named MelGlow, is designed. Experiments on the LJSpeech dataset show that MelGlow achieves better performance than WaveGlow at small model sizes, which verifies the effectiveness and potential optimization space of location-variable convolutions. Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
SLT | 3 |
| 2020 | GraphTTS: Graph-to-Sequence Modelling in Neural Text-to-SpeechabstractThis paper leverages the graph-to-sequence method in neural text-to-speech (GraphTTS), which maps the graph embedding of the input sequence to spectrograms. The graphical inputs consist of node and edge representations constructed from input texts. The encoding of these graphical inputs incorporates syntax information by a GNN encoder module. Besides, applying the encoder of GraphTTS as a graph auxiliary encoder (GAE) can analyse prosody information from the semantic structure of texts. This can remove the manual selection of reference audios process and makes prosody modelling an end-to-end procedure. Experimental analysis shows that GraphTTS outperforms the state-of-the-art sequence-to-sequence models by 0.24 in Mean Opinion Score (MOS). GAE can adjust the pause, ventilation and tones of synthesised audios automatically. This experimental conclusion may give some inspiration to researchers working on improving speech synthesis prosody. Aolan Sun, Jianzong Wang, Ning Cheng 0001, Huayi Peng, Jing Xiao 0006 |
ICASSP | 3 |
| 2020 | Aligntts: Efficient Feed-Forward Text-to-Speech System Without Explicit AlignmentabstractTargeting at both high efficiency and performance, we propose AlignTTS to predict the mel-spectrum in parallel. AlignTTS is based on a Feed-Forward Transformer which generates mel-spectrum from a sequence of characters, and the duration of each character is determined by a duration predictor. Instead of adopting the attention mechanism in Transformer TTS to align text to mel-spectrum, the alignment loss is presented to consider all possible alignments in training by use of dynamic programming. Experiments on the LJSpeech dataset show that our model achieves not only state-of-the-art performance which outperforms Transformer TTS by 0.03 in mean option score (MOS), but also a high efficiency which is more than 50 times faster than real-time. Jianzong Wang, Ning Cheng 0001, Tian Xia 0004, Jing Xiao 0006 |
ICASSP | 3 |
| 2020 | Large-Scale Transfer Learning for Low-Resource Spoken Language UnderstandingabstractEnd-to-end Spoken Language Understanding (SLU) models are made increasingly large and complex to achieve the state-ofthe-art accuracy.However, the increased complexity of a model can also introduce high risk of over-fitting, which is a major challenge in SLU tasks due to the limitation of available data.In this paper, we propose an attention-based SLU model together with three encoder enhancement strategies to overcome data sparsity challenge.The first strategy focuses on the transferlearning approach to improve feature extraction capability of the encoder.It is implemented by pre-training the encoder component with a quantity of Automatic Speech Recognition annotated data relying on the standard Transformer architecture and then fine-tuning the SLU model with a small amount of target labelled data.The second strategy adopts multitask learning strategy, the SLU model integrates the speech recognition model by sharing the same underlying encoder, such that improving robustness and generalization ability.The third strategy, learning from Component Fusion (CF) idea, involves a Bidirectional Encoder Representation from Transformer (BERT) model and aims to boost the capability of the decoder with an auxiliary network.It hence reduces the risk of over-fitting and augments the ability of the underlying encoder, indirectly.Experiments on the FluentAI dataset show that cross-language transfer learning and multi-task strategies have been improved by up to 4.52% and 3.89% respectively, compared to the baseline. Xueli Jia, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
INTERSPEECH | 4 |
| 2020 | A Real-Time Robot-Based Auxiliary System for Risk Evaluation of COVID-19 InfectionabstractIn this paper, we propose a real-time robot-based auxiliary system for risk evaluation of COVID-19 infection.It combines real-time speech recognition, temperature measurement, keyword detection, cough detection and other functions in order to convert live audio into actionable structured data to achieve the COVID-19 infection risk assessment function.In order to better evaluate the COVID-19 infection, we propose an end-to-end method for cough detection and classification for our proposed system.It is based on real conversation data from human-robot, which processes speech signals to detect cough and classifies it if detected.The structure of our model are maintained concise to be implemented for real-time applications.And we further embed this entire auxiliary diagnostic system in the robot and it is placed in the communities, hospitals and supermarkets to support COVID-19 testing.The system can be further leveraged within a business rules engine, thus serving as a foundation for real-time supervision and assistance applications.Our model utilizes a pretrained, robust training environment that allows for efficient creation and customization of customer-specific health states. Jianzong Wang, Jiteng Ma, Ning Cheng 0001, Jing Xiao 0006 |
INTERSPEECH | 4 |
| 2020 | Prosody Learning Mechanism for Speech Synthesis System Without Text Length LimitabstractRecent neural speech synthesis systems have gradually focused on the control of prosody to improve the quality of synthesized speech, but they rarely consider the variability of prosody and the correlation between prosody and semantics together.In this paper, a prosody learning mechanism is proposed to model the prosody of speech based on TTS system, where the prosody information of speech is extracted from the melspectrum by a prosody learner and combined with the phoneme sequence to reconstruct the mel-spectrum.Meanwhile, the sematic features of text from the pre-trained language model is introduced to improve the prosody prediction results.In addition, a novel self-attention structure, named as local attention, is proposed to lift this restriction of input text length, where the relative position information of the sequence is modeled by the relative position matrices so that the position encodings is no longer needed.Experiments on English and Mandarin show that speech with more satisfactory prosody has obtained in our model.Especially in Mandarin synthesis, our proposed model outperforms baseline model with a MOS gap of 0.08, and the overall naturalness of the synthesized speech has been significantly improved. Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
INTERSPEECH | 3 |
| 2020 | MLNET: An Adaptive Multiple Receptive-Field Attention Neural Network for Voice Activity DetectionabstractVoice activity detection (VAD) makes a distinction between speech and non-speech and its performance is of crucial importance for speech based services.Recently, deep neural network (DNN)-based VADs have achieved better performance than conventional signal processing methods.The existed DNNbased models always handcrafted a fixed window to make use of the contextual speech information to improve the performance of VAD.However, the fixed window of contextual speech information can't handle various unpredictable noise environments and highlight the critical speech information to VAD task.In order to solve this problem, this paper proposed an adaptive multiple receptive-field attention neural network, called MLNET, to finish VAD task.The MLNET leveraged multi-branches to extract multiple contextual speech information and investigated an effective attention block to weight the most crucial parts of the context for final classification.Experiments in real-world scenarios demonstrated that the proposed MLNET-based model outperformed other baselines. Zhenpeng Zheng, Jianzong Wang, Ning Cheng 0001, Jian Luo 0007, Jing Xiao 0006 |
INTERSPEECH | 3 |
| 2011 | Generalized Variable Parameter HMMs for Noise Robust Speech RecognitionabstractHandling variable ambient noise is a challenging task for au-tomatic speech recognition (ASR) systems. To address this is-sue, multi-style, noise condition independent (CI) model train-ing using speech data collected in diverse noise environments, or uncertainty decoding techniques can be used. An alternative approach is to explicitly approximate the continuous trajectory of Gaussian component mean and variance parameters against the varying noise level, for example, using variable parameter HMMs (VP-HMM). This paper investigates a more generalized form of variable parameter HMMs (GVP-HMM). In addition to Gaussian component means and variances, it can also provide a more compact trajectory modelling for tied linear transforma-tions. An alternative noise condition dependent (CD) training algorithm is also proposed to handle the bias to training noise condition distribution. Consistent error rate gains were obtained over conventional VP-HMM mean and variance only trajectory modelling on a medium vocabulary Mandarin Chinese in-car navigation command recognition task. Index Terms: variable noise, generalized variable parameter HMM, noise robust speech recognition Ning Cheng 0001, Xunying Liu |
INTERSPEECH | 1 |
| 2011 | A flexible framework for HMM based noise robust speech recognition using generalized parametric space polynomial regression
Ning Cheng 0001, Xunying Liu |
Sci. China Inf. Sci. | 1 |
| 2010 | Masking property based microphone array post-filter design
Ning Cheng 0001 |
INTERSPEECH | 1 |
| 2008 | An effective microphone array post-filter in arbitrary environmentsabstractThe theoretic foundation of traditional microphone array post-filters is the signal model in which the noise between sensors is assumed to be uncorrelated. However, this model is inaccurate in real environments since the correlated noise exists. In this paper, a more generalized signal model which considers both the correlated and uncorrelated noise is introduced. A general expression of the microphone array post-filter is proposed for this model. For better residual noise shaping, the human auditory property is incorporated into the post-filter estimation process. In experiments with real noise microphone array recordings, the proposed technique has shown to produce impressive results in terms of quality measures of the enhanced speech. Index Terms: post-filter, generalized signal model, human auditory property, speech enhancement Ning Cheng 0001, Peng Li 0030, Bo Xu 0002 |
INTERSPEECH | 1 |