VLDB 2026 Research / reviewers in the wild / expert
Xulong Zhang 0001
dblp:132/9890-1
· DBLP profile ↗
57ranked-venue papers
17as first author
57since 2021 · last 2026
0000-0001-7005-992XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 9 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 3 first-author · 23 since 2021Computer networks · 7 · 6 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Inheritance to Saturation: Disentangling the Evolution of Visual Redundancy for Architecture-Aware MLLM Inference AccelerationabstractHigh-resolution Multimodal Large Language Models (MLLMs) face prohibitive computational costs during inference due to the explosion of visual tokens. Existing acceleration strategies, such as token pruning or layer sparsity, suffer from severe “backbone dependency”, performing well on Vicuna or Mistral architectures (e.g., LLaVA) but causing significant performance degradation when transferred to architectures like Qwen. To address this, we leverage truncated matrix entropy to uncover a universal three-stage inference lifecycle, decoupling visual redundancy into universal Intrinsic Visual Redundancy (IVR) and architecture-dependent Secondary Saturation Redundancy (SSR). Guided by this insight, we propose HalfV, a framework that first mitigates IVR via a unified pruning strategy and then adaptively handles SSR based on its specific manifestation. Experiments demonstrate that HalfV achieves superior efficiency-performance trade-offs across diverse backbones. Notably, on Qwen25-VL, it retains 96.8% performance at a 4.1\times FLOPs speedup, significantly outperforming state-of-the-art baselines. Our code is available at https://github.com/civilizwa/HalfV. Xulong Zhang 0001, Yuechan Li, Xiaoyang Qu, Jianzong Wang |
ACL (1) | 2 |
| 2025 | Graph Contrastive Learning with Decoupled AugmentationabstractGraph contrastive learning based on augmentation strategies has recently demonstrated remarkable performance. Existing methods typically jointly leverage attribute and structural augmentations to generate graph views, learning data invariance information through contrasting sample pairs. However, this joint approach may deviate from the expectation of semantically similar before and after augmentation. The propagation of attribute information in graphs usually occurs through their structure, meaning that structural and attribute augmentations can interfere with each other and potentially distort the graph’s semantics. To address this, we propose a decoupled augmentation framework for graph contrastive learning, which eliminates the mutual interference between the two levels of augmentation while fully exploring graph information. Specifically, our framework employs separate encoders to learn data invariance under different augmentation levels, and it considers the positive gains generated between these levels. Experimental results on five public datasets show that the proposed method is more competitive than state-of-the-art approaches. Shihao Gao, Caoshuo Li, Cunli Mao, Xulong Zhang 0001, Xiaoyang Qu, Taisong Jin, Jianzong Wang |
ICASSP | 4 |
| 2025 | Homogeneous Graph Extraction: An Approach to Learning Heterogeneous Graph EmbeddingabstractHeterogeneous Graph Neural Networks (HGNNs) aim to embed rich structural and semantic information of heterogeneous graphs into low-dimensional node representations. While HGNNs extend the foundational work of homogeneous Graph Neural Networks, the methodology for effectively transforming heterogeneous graphs into homogeneous graphs and then learning node representations remains under-explored. In this paper, we propose a novel heterogeneous graph embedding method via the Homogeneous Graph Extraction strategy, termed HGE. Specifically, the proposed method ingeniously harnesses information clusters and metapaths to extract tailored homogeneous graphs from the complex heterogeneous graph. Subsequently, these distilled homogeneous graphs are fed into a weight-shared homogeneous graph encoder to obtain embeddings with diverse semantic information. Finally, we employ an attention mechanism, which adeptly fuses embeddings derived from distinct homogeneous graphs, resulting in the more expressive capability of the nodes. The effectiveness of the proposed architecture was demonstrated through experiments on three real heterogeneous graph datasets. Shihao Gao, Xulong Zhang 0001, Jianzong Wang, Taisong Jin |
ICASSP | 4 |
| 2025 | CycleFlow: Leveraging Cycle Consistency in Flow Matching for Speaker Style AdaptationabstractVoice Conversion (VC) aims to convert the style of a source speaker, such as timbre and pitch, to the style of any target speaker while preserving the linguistic content. However, the ground truth of the converted speech does not exist in a non-parallel VC scenario, which induces the train-inference mismatch problem. Moreover, existing methods still have an inaccurate pitch and low speaker adaptation quality, there is a significant disparity in pitch between the source and target speaker style domains. As a result, the models tend to generate speech with hoarseness, posing challenges in achieving high-quality voice conversion. In this study, we propose CycleFlow, a novel VC approach that leverages cycle consistency in conditional flow matching (CFM) for speaker timbre adaptation training on non-parallel data. Furthermore, we design a Dual-CFM based on VoiceCFM and PitchCFM to generate speech and improve speaker pitch adaptation quality. Experiments show that our method can significantly improve speaker similarity, generating natural and higher-quality speech. Ziqi Liang, Xulong Zhang 0001, Xiaoyang Qu, Weifeng Zhao, Jianzong Wang |
ICASSP | 2 |
| 2025 | Turbo-TTS: Enhancing Diffusion Model TTS with an Improved ODE Solver
Xulong Zhang 0001, Xiaoyang Qu, Hui Tian 0002, Jianzong Wang |
ICONIP (1) | 1 |
| 2025 | Rano: Restorable Speaker Anonymization via Conditional Invertible Neural NetworkabstractSpeech contains ample information, including the primary semantic content and information about the speaker, such as gender, age and health status. Speaker-dependent information partially carries personal privacy and has raised concerns about the protection of voice privacy. Speaker anonymization aims to conceal the speaker’s identity in speech, while preserving speaker-independent information to the greatest extent possible, and has become an increasingly important task in the field of speech. Existing research generally treats speaker anonymization as a downstream task of voice conversion, often employing speech representation disentanglement-based methods to separate speaker-dependent and speaker-independent information. However, speech representation disentanglement, especially for speaker-independent information, faces challenges such as information leakage or excessive disentangling, resulting in quality degradation. In this paper, we propose a speaker anonymization model called Rano, which does not rely on precise disentanglement. Rano employs a generative invertible neural network to forge anonymous speaker identities from keys and then uses the speaker embeddings as conditions to guide the speaker anonymization process via a conditional invertible neural network. Moreover, when the key is provided, lossless restoration from anonymized speech to original speech can be achieved via the reverse network, thereby expanding the application scenarios of Rano. Experiments demonstrate that the proposed model achieves comparable performance to existing state-of-the-art models. We also verify the security guaranteed by the key in the restoration process. Jianzong Wang, Xulong Zhang 0001, Xiaoyang Qu |
IJCNN | 2 |
| 2025 | Bridging the Modality Gap: Semantic-Calibrated Zero-shot Speech Emotion CaptioningabstractSpeech Emotion Captioning (SEC) has emerged as an increasingly prominent research area. The emotional content expressed through human speech is often intricate, making it difficult to fully capture with fixed categorical labels. Instead, describing these emotions using natural language can offer a more comprehensive representation. However, obtaining well-matched speech-caption datasets is challenging in practical scenarios, and existing SEC techniques often generate hallucinated content or miss fine-grained details when working with cross-domain, unpaired data. To address these challenges, we introduce SeCCap, a Semantic-Calibrated Zero-shot Speech Emotion Captioning framework built upon large language models (LLMs). SeCCap exhibits three key features: 1) Zero-shot inference: It generates speech emotion captions without requiring training on paired speech-caption datasets. 2) Bridging the modality gap: It employs caption-only training and semantic-calibrated cross-modal mapping to enhance fine-grained content and reduce factual hallucinations during zero-shot SEC. Experimental results show that SeCCap outperforms other state-of-the-art models in zero-shot SEC tasks. Jianzong Wang, Xulong Zhang 0001, Xiaoyang Qu |
IJCNN | 2 |
| 2025 | Logic Consistency Makes Large Language Models Personalized Reasoning TeachersabstractLarge Language Models (LLMs) have advanced natural language processing, particularly through Chain-of-Thought (CoT) reasoning, but their high computational costs limit deployment. We propose Personalized Chain-of-Thought Distillation (PeCoTD), a method that transfers CoT reasoning from LLMs to smaller models by addressing the distribution gap—the difference in how large and small models process information. To bridge this gap, PeCoTD introduces the Self Logic Consistency (SLC) metric, which helps small models evaluate and select LLM-generated rationales that align better with their reasoning abilities. PeCoTD iteratively refines these rationales, adjusting them to better fit the learning patterns of small models while preserving their original meaning. Experiments show PeCoTD significantly enhances the reasoning abilities of small models across datasets, making CoT distillation more practical and effective. Xulong Zhang 0001, Yong Zhang 0058, Jun Yu 0001, Jianzong Wang |
IJCNN | 2 |
| 2025 | Knowledge distillation for financial large language models: a systematic review of strategies, applications, and evaluationabstractFinancial large language models (FinLLMs) offer immense potential for financial applications. While excessive deployment expenditures and considerable inference latency constitute major obstacles, as a prominent compression methodology, knowledge distillation (KD) offers an effective solution to these difficulties. A comprehensive survey is conducted in this work on how KD interacts with FinLLMs, covering three core aspects: strategy, application, and evaluation. At the strategy level, this review introduces a structured taxonomy to comparatively analyze existing distillation pathways. At the application level, this review puts forward a logical upstream–midstream–downstream framework to systematically explain the practical value of distilled models in the financial field. At the evaluation level, to tackle the absence of standards in the financial field, this review constructs a comprehensive evaluation framework that proceeds from multiple dimensions such as financial accuracy, reasoning fidelity, and robustness. In summary, this research aims to provide a clear roadmap for this interdisciplinary field, to accelerate the development of distilled FinLLMs. Xulong Zhang 0001, Xiaoyang Qu, Junfei Xie, Jianzong Wang |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2024 | Medical Speech Symptoms Classification via Disentangled RepresentationabstractIntent is defined for understanding spoken language in existing works. Both textual features and acoustic features involved in medical speech contain intent, which is important for symptomatic diagnosis. In this paper, we propose a medical speech classification model named DRSC that automatically learns to disentangle intent and content representations from textual-acoustic data for classification. The intent representations of the text domain and the Mel-spectrogram domain are extracted via intent encoders, and then the reconstructed text feature and the Mel-spectrogram feature are obtained through two exchanges. After combining the intent from two domains into a joint representation, the integrated intent representation is fed into a decision layer for classification. Experimental results show that our model obtains an average accuracy rate of 95% in detecting 25 different medical symptoms. Jianzong Wang, Pengcheng Li 0013, Xulong Zhang 0001, Ning Cheng 0001, Jing Xiao 0006 |
CSCWD | 3 |
| 2024 | IDEAW: Robust Neural Audio Watermarking with Invertible Dual-EmbeddingabstractThe audio watermarking technique embeds messages into audio and accurately extracts messages from the watermarked audio.Traditional methods develop algorithms based on expert experience to embed watermarks into the time-domain or transform-domain of signals.With the development of deep neural networks, deep learning-based neural audio watermarking has emerged.Compared to traditional algorithms, neural audio watermarking achieves better robustness by considering various attacks during training.However, current neural watermarking methods suffer from low capacity and unsatisfactory imperceptibility.Additionally, the issue of watermark locating, which is extremely important and even more pronounced in neural audio watermarking, has not been adequately studied.In this paper, we design a dual-embedding watermarking model for efficient locating.We also consider the impact of the attack layer on the invertible neural network in robustness training, improving the model to enhance both its reasonableness and stability.Experiments show that the proposed model, IDEAW, can withstand various attacks with higher capacity and more efficient locating ability compared to existing methods.The code is available at https://github.com/PecholaL/IDEAW. Pengcheng Li 0013, Xulong Zhang 0001, Jing Xiao 0006, Jianzong Wang |
EMNLP | 2 |
| 2024 | Learning Disentangled Speech Representations with Contrastive Learning and Time-Invariant RetrievalabstractVoice conversion refers to transferring speaker identity with well-preserved content. Better disentanglement of speech representations leads to better voice conversion. Recent studies have found that phonetic information from input audio has the potential ability to well represent content. Besides, the speaker-style modeling with pre-trained models making the process more complex. To tackle these issues, we introduce an new method named "CTVC" which utilizes disen-tangled speech representations with contrastive learning and time-invariant retrieval.Specifically, a similarity-based compression module is used to facilitate a more intimate connection between the frame-level hidden features and linguistic information at phoneme-level. Additionally, a time-invariant retrieval is proposed for timbre extraction based on multiple segmentation and mutual information. Experimental results demonstrate that "CTVC" outperforms previous studies and improves the sound quality and similarity of converted results. Huaizhen Tang, Xulong Zhang 0001, Ning Cheng 0001, Jing Xiao 0006, Jianzong Wang |
ICASSP | 3 |
| 2024 | ED-TTS: Multi-Scale Emotion Modeling Using Cross-Domain Emotion Diarization for Emotional Speech SynthesisabstractExisting emotional speech synthesis methods often utilize an utterance-level style embedding extracted from reference audio, neglecting the inherent multi-scale property of speech prosody. We introduce ED-TTS, a multi-scale emotional speech synthesis model that leverages Speech Emotion Diarization (SED) and Speech Emotion Recognition (SER) to model emotions at different levels. Specifically, our proposed approach integrates the utterance-level emotion embedding extracted by SER with fine-grained frame-level emotion embedding obtained from SED. These embeddings are used to condition the reverse process of the denoising diffusion probabilistic model (DDPM). Additionally, we employ cross-domain SED to accurately predict soft labels, addressing the challenge of a scarcity of fine-grained emotion-annotated datasets for supervising emotional TTS training. Haobin Tang, Xulong Zhang 0001, Ning Cheng 0001, Jing Xiao 0006, Jianzong Wang |
ICASSP | 2 |
| 2024 | EmoTalker: Emotionally Editable Talking Face Generation via Diffusion ModelabstractIn recent years, the field of talking faces generation has attracted considerable attention, with certain methods adept at generating virtual faces that convincingly imitate human expressions. However, existing methods face challenges related to limited generalization, particularly when dealing with challenging identities. Furthermore, methods for editing expressions are often confined to a singular emotion, failing to adapt to intricate emotions. To overcome these challenges, this paper proposes EmoTalker, an emotionally editable portraits animation approach based on the diffusion model. EmoTalker modifies the denoising process to ensure preservation of the original portrait’s identity during inference. To enhance emotion comprehension from text input, Emotion Intensity Block is introduced to analyze fine-grained emotions and strengths derived from prompts. Additionally, a crafted dataset is harnessed to enhance emotion comprehension within prompts. Experiments show the effectiveness of EmoTalker in generating high-quality, emotionally customizable facial expressions. Xulong Zhang 0001, Ning Cheng 0001, Jun Yu 0001, Jing Xiao 0006, Jianzong Wang |
ICASSP | 2 |
| 2024 | Learning Expressive Disentangled Speech Representations with Soft Speech Units and Adversarial Style AugmentationabstractVoice conversion is the task to transform voice characteristics of source speech while preserving content information. Nowadays, self-supervised representation learning models are increasingly utilized in content extraction. However, in these representations, a lot of hidden speaker information leads to timbre leakage while the prosodic information of hidden units lacks use. To address these issues, we propose a novel framework for expressive voice conversion called “SAVC” based on soft speech units from HuBert-soft. Taking soft speech units as input, we design an attribute encoder to extract content and prosody features respectively. Specifically, we first introduce statistic perturbation imposed by adversarial style augmentation to eliminate speaker information. Then the prosody is implicitly modeled on soft speech units with knowledge distillation. Experiment results show that the intelligibility and naturalness of converted speech outperform previous work. Jianzong Wang, Xulong Zhang 0001, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 3 |
| 2024 | MAIN-VC: Lightweight Speech Representation Disentanglement for One-shot Voice ConversionabstractOne-shot voice conversion aims to change the timbre of any source speech to match that of the unseen target speaker with only one speech sample. Existing methods face difficulties in satisfactory speech representation disentanglement and suffer from sizable networks as some of them leverage numerous complex modules for disentanglement. In this paper, we propose a model named MAIN-VC to effectively disentangle via a concise neural network. The proposed model utilizes Siamese encoders to learn clean representations, further enhanced by the designed mutual information estimator. The Siamese structure and the newly designed convolution module contribute to the lightweight of our model while ensuring performance in diverse voice conversion tasks. The experimental results show that the proposed model achieves comparable subjective scores and exhibits improvements in objective metrics compared to existing methods in a one-shot voice conversion scenario. Pengcheng Li 0013, Jianzong Wang, Xulong Zhang 0001, Yong Zhang 0058, Jing Xiao 0006, Ning Cheng 0001 |
IJCNN | 3 |
| 2024 | EAD-VC: Enhancing Speech Auto-Disentanglement for Voice Conversion with IFUB Estimator and Joint Text-Guided Consistent LearningabstractUsing unsupervised learning to disentangle speech into content, rhythm, pitch, and timbre for voice conversion has become a hot research topic. Existing works generally take into account disentangling speech components through human-crafted bottleneck features which can not achieve sufficient information disentangling, while pitch and rhythm may still be mixed together. There is a risk of information overlap in the disentangling process which results in less speech naturalness. To overcome such limits, we propose a two-stage model to disentangle speech representations in a self-supervised manner without a human-crafted bottleneck design, which uses the Mutual Information (MI) with the designed upper bound estimator (IFUB) to separate overlapping information between speech components. Moreover, we design a Joint Text-Guided Consistent (TGC) module to guide the extraction of speech content and eliminate timbre leakage issues. Experiments show that our model can achieve a better performance than the baseline, regarding disentanglement effectiveness, speech naturalness, and similarity. Audio samples can be found at https://largeaudiomodel.com/eadvc. Ziqi Liang, Jianzong Wang, Xulong Zhang 0001, Yong Zhang 0058, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 3 |
| 2024 | QLSC: A Query Latent Semantic Calibrator for Robust Extractive Question AnsweringabstractExtractive Question Answering (EQA) in Machine Reading Comprehension (MRC) often faces the challenge of dealing with semantically identical but format-variant inputs. Our work introduces a novel approach, called the "Query Latent Semantic Calibrator (QLSC)", designed as an auxiliary module for existing MRC models. We propose a unique scaling strategy to capture latent semantic center features of queries. These features are then seamlessly integrated into traditional query and passage embeddings using an attention mechanism. By deepening the comprehension of the semantic queries-passage relationship, our approach diminishes sensitivity to variations in text format and boosts the model’s capability in pinpointing accurate answers. Experimental results on robust Question-Answer datasets confirm that our approach effectively handles format-variant but semantically identical queries, highlighting the effectiveness and adaptability of our proposed method. Sheng Ouyang, Jianzong Wang, Yong Zhang 0058, Zhitao Li 0002, Ziqi Liang, Xulong Zhang 0001, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 6 |
| 2024 | ConTuner: Singing Voice Beautifying with Pitch and Expressiveness ConditionabstractSinging voice beautifying is a novel task that has application value in people’s daily life, aiming to correct the pitch of the singing voice and improve the expressiveness without changing the original timbre and content. Existing methods rely on paired data or only concentrate on the correction of pitch. However, professional songs and amateur songs from the same person are hard to obtain, and singing voice beautifying doesn’t only contain pitch correction but other aspects like emotion and rhythm. Since we propose a fast and high-fidelity singing voice beautifying system called ConTuner, a diffusion model combined with the modified condition to generate the beautified Mel-spectrogram, where the modified condition is composed of optimized pitch and expressiveness. For pitch correction, we establish a mapping relationship from MIDI, spectrum envelope to pitch. To make amateur singing more expressive, we propose the expressiveness enhancer in the latent space to convert amateur vocal tone to professional. ConTuner achieves a satisfactory beautification effect on both Mandarin and English songs. Ablation study demonstrates that the expressiveness enhancer and generator-based accelerate method in ConTuner are effective. Jianzong Wang, Pengcheng Li 0013, Xulong Zhang 0001, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 3 |
| 2024 | EfficientASR: Speech Recognition Network Compression via Attention Redundancy and Chunk-Level FFN OptimizationabstractIn recent years, Transformer networks have shown remarkable performance in speech recognition tasks. However, their deployment poses challenges due to high computational and storage resource requirements. To address this issue, a lightweight model called EfficientASR is proposed in this paper, aiming to enhance the versatility of Transformer models. EfficientASR employs two primary modules: Shared Residual Multi-Head Attention (SRMHA) and Chunk-Level Feedforward Networks (CFFN). The SRMHA module effectively reduces redundant computations in the network, while the CFFN module captures spatial knowledge and reduces the number of parameters. The effectiveness of the EfficientASR model is validated on two public datasets, namely Aishell-1 and HKUST. Experimental results demonstrate a 36% reduction in parameters compared to the baseline Transformer network, along with improvements of 0.3% and 0.2% in Character Error Rate (CER) on the Aishell-1 and HKUST datasets, respectively. Jianzong Wang, Ziqi Liang, Xulong Zhang 0001, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 3 |
| 2023 | Machine Unlearning Methodology Based on Stochastic Teacher Network
Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Yifu Sun, Chuanyao Zhang, Jing Xiao 0006 |
ADMA (5) | 1 |
| 2023 | Voice Conversion with Denoising Diffusion Probabilistic GAN Models
Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ADMA (4) | 1 |
| 2023 | Symbolic and Acoustic: Multi-domain Music Emotion Modeling for Instrumental Music
Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ADMA (4) | 2 |
| 2023 | Improving Music Genre Classification from multi-modal Properties of Music and Genre Correlations PerspectiveabstractMusic genre classification has been widely studied in past few years for its various applications in music information retrieval. Previous works tend to perform unsatisfactorily, since those methods only use audio content or jointly use audio content and lyrics content inefficiently. In addition, as genres normally co-occur in a music track, it is desirable to capture and model the genre correlations to improve the performance of multi-label music genre classification. To solve these issues, we present a novel multi-modal method leveraging audio-lyrics contrastive loss and two symmetric cross-modal attention, to align and fuse features from audio and lyrics. Furthermore, based on the nature of the multi-label classification, a genre correlations extraction module is presented to capture and model potential genre correlations. Extensive experiments demonstrate that our proposed method significantly surpasses other multi-label music genre classification methods and achieves state-of-the-art result on Music4All dataset. Ganghui Ru, Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 2 |
| 2023 | Learning Speech Representations with Flexible Hidden Feature DimensionsabstractNon-parallel many-to-many voice conversion is a kind of style transfer task in speech. Recently, AutoVC has been applied in this field as a popular solution, as it can achieve distribution-matching style transfer by training only the re- construction loss. However, in order to strike a good balance between timbre disentanglement and sound quality, AutoVC requires imposing very strict constraints on the dimensionality of the latent representation. This constraint affects the quality of the converted speech while making it challenging to apply to other datasets directly. This paper proposes a new voice conversion framework that uses only one encoder to obtain timbre and content information by partitioning the latent space in the channel dimension. Furthermore, two different types of classifiers and two additional reconstruction losses are proposed to ensure that different parts of the latent space contain only separated content and timbre information, respectively. Experiments on the VCTK dataset show that the proposed model achieves state-of-the-art results in terms of the naturalness and similarity of converted speech. In addition, we experimentally show that for different division proportions of latent space, the content and timbre information will always be well separated. Huaizhen Tang, Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 2 |
| 2023 | VQ-CL: Learning Disentangled Speech Representations with Contrastive Learning and Vector QuantizationabstractVoice Conversion(VC) refers to converting the voice characteristics of audio to another one as it is said by other people. Recently, more and more studies have focused on disentangle-based VC, which separates the timbre and linguistic content information from an audio signal to effectively achieve VC tasks. However, It’s still challenging to extract phoneme-level features from frame-level hidden representations. This paper proposed a novel zero-shot voice conversion framework that utilizes contrastive learning and vector quantization to encourage the frame-level hidden features closer to the phoneme-level linguistic information, called VQ-CL. All objective and subjective experiment results show that VQ-CL has better performance than previous studies in separating content and voice characteristics to improve the sound quality of generated speech. Huaizhen Tang, Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 2 |
| 2023 | QI-TTS: Questioning Intonation Control for Emotional Speech SynthesisabstractRecent expressive text to speech (TTS) models focus on synthesizing emotional speech, but some fine-grained styles such as intonation are neglected. In this paper, we propose QI-TTS which aims to better transfer and control intonation to further deliver the speaker’s questioning intention while transferring emotion from reference speech. We propose a multi-style extractor to extract style embedding from two different levels. While the sentence level represents emotion, the final syllable level represents intonation. For fine-grained intonation control, we use relative attributes to represent intonation intensity at the syllable level. Experiments have validated the effectiveness of QI-TTS for improving intonation expressiveness in emotional speech synthesis. Haobin Tang, Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 2 |
| 2023 | Dynamic Alignment Mask CTC: Improved Mask CTC With Aligned Cross EntropyabstractBecause of predicting all the target tokens in parallel, the non-autoregressive models greatly improve the decoding efficiency of speech recognition compared with traditional autoregressive models. In this work, we present dynamic alignment Mask CTC, introducing two methods: (1) Aligned Cross Entropy (AXE), finding the monotonic alignment that minimizes the cross-entropy loss through dynamic programming, (2) Dynamic Rectification, creating new training samples by replacing some masks with model predicted tokens. The AXE ignores the absolute position alignment between prediction and ground truth sentence and focuses on tokens matching in relative order. The dynamic rectification method makes the model capable of simulating the non-mask but possible wrong tokens, even if they have high confidence. Our experiments on WSJ dataset demonstrated that not only AXE loss but also the rectification method could improve the WER performance of Mask CTC. Xulong Zhang 0001, Haobin Tang, Jianzong Wang, Ning Cheng 0001, Jian Luo 0007, Jing Xiao 0006 |
ICASSP | 1 |
| 2023 | Improving EEG-based Emotion Recognition by Fusing Time-Frequency and Spatial RepresentationsabstractUsing deep learning methods to classify EEG signals can accurately identify people’s emotions. However, existing studies have rarely considered the application of the information in another domain’s representations to feature selection in the time-frequency domain. We propose a classification network of EEG signals based on the cross-domain feature fusion method, which makes the network more focused on the features most related to brain activities and thinking changes by using the multi-domain attention mechanism. In addition, we propose a two-step fusion method and apply these methods to the EEG emotion recognition network. Experimental results show that our proposed network, which combines multiple representations in the time-frequency domain and spatial domain, outperforms previous methods on public datasets and achieves state-of-the-art at present. Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 2 |
| 2023 | Contrastive Latent Space Reconstruction Learning for Audio-Text RetrievalabstractCross-modal retrieval (CMR) has been extensively applied in various domains, such as multimedia search engines and recommendation systems. Most existing CMR methods focus on image-to-text retrieval, whereas audio-to-text retrieval, a less explored domain, has posed a great challenge due to the difficulty to uncover discriminative features from audio clips and texts. Existing studies are restricted in the following two ways: 1) Most researchers utilize contrastive learning to construct a common subspace where similarities among data can be measured. However, they considers only cross-modal transformation, neglecting the intra-modal separability. Besides, the temperature parameter is not adaptively adjusted along with semantic guidance, which degrades the performance. 2) These methods do not take latent representation reconstruction into account, which is essential for semantic alignment. This paper introduces a novel audio-text oriented CMR approach, termed Contrastive Latent Space Reconstruction Learning (CLSR). CLSR improves contrastive representation learning by taking intra-modal separability into account and adopting an adaptive temperature control strategy. Moreover, the latent representation reconstruction modules are embedded into the CMR framework, which improves modal interaction. Experiments in comparison with some state-of-the-art methods on two audio-text datasets have validated the superiority of CLSR. Kaiyi Luo, Xulong Zhang 0001, Jianzong Wang, Huaxiong Li, Ning Cheng 0001, Jing Xiao 0006 |
ICTAI | 2 |
| 2023 | AOSR-Net: All-in-One Sandstorm Removal NetworkabstractMost existing sandstorm image enhancement methods are based on traditional theory and prior knowledge, which often restrict their applicability in real-world scenarios. In addition, these approaches often adopt a strategy of color correction followed by dust removal, which makes the algorithm structure too complex. To solve the issue, we introduce a novel image restoration model, named all-in-one sandstorm removal network (AOSR-Net). This model is developed based on a re-formulated sandstorm scattering model, which directly establishes the image mapping relationship by integrating intermediate parameters. Such integration scheme effectively addresses the problems of over-enhancement and weak generalization in the field of sand dust image enhancement. Experimental results on synthetic and real-world sandstorm images demonstrate the superiority of the proposed AOSR-Net over state-of-the-art (SOTA) algorithms. Yazhong Si, Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICTAI | 2 |
| 2023 | FastGraphTTS: An Ultrafast Syntax-Aware Speech Synthesis FrameworkabstractThis paper integrates graph-to-sequence into an end-to-end text-to-speech framework for syntax-aware modelling with syntactic information of input text. Specifically, the input text is parsed by a dependency parsing module to form a syntactic graph. The syntactic graph is then encoded by a graph encoder to extract the syntactic hidden information, which is concatenated with phoneme embedding and input to the alignment and flow-based decoding modules to generate the raw audio waveform. The model is experimented on two languages, English and Mandarin, using single-speaker, few samples of target speakers, and multi-speaker datasets, respectively. Experimental results show better prosodic consistency performance between input text and generated audio, and also get higher scores in the subjective prosodic evaluation, and show the ability of voice conversion. Besides, the efficiency of the model is largely boosted through the design of the AI chip operator with 5x acceleration. Jianzong Wang, Xulong Zhang 0001, Aolan Sun, Ning Cheng 0001, Jing Xiao 0006 |
ICTAI | 2 |
| 2023 | SAR: Self-Supervised Anti-Distortion Representation for End-To-End Speech ModelabstractIn recent Text-to-Speech (TTS) systems, a neural vocoder often generates speech samples by solely conditioning on acoustic features predicted from an acoustic model. However, there are always distortions existing in the predicted acoustic features, compared to those of the groundtruth, especially in the common case of poor acoustic modeling due to low-quality training data. To overcome such limits, we propose a Self-supervised learning framework to learn an Anti-distortion acoustic Representation (SAR) to replace human-crafted acoustic features by introducing distortion prior to an auto-encoder pre-training process. The learned acoustic representation from the proposed framework is proved anti-distortion compared to the most commonly used mel-spectrogram through both objective and subjective evaluation. Jianzong Wang, Xulong Zhang 0001, Haobin Tang, Aolan Sun, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 2 |
| 2023 | Investigation of Music Emotion Recognition Based on Segmented Semi-Supervised Learning
Yifu Sun, Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Kaiyu Hu, Jing Xiao 0006 |
INTERSPEECH | 2 |
| 2023 | EmoMix: Emotion Mixing via Diffusion Models for Emotional Speech Synthesis
Haobin Tang, Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
INTERSPEECH | 2 |
| 2023 | PMVC: Data Augmentation-Based Prosody Modeling for Expressive Voice ConversionabstractVoice conversion as the style transfer task applied to speech, refers to converting one person's speech into a new speech that sounds like another person's. Up to now, there has been a lot of research devoted to better implementation of VC tasks. However, a good voice conversion model should not only match the timbre information of the target speaker, but also expressive information such as prosody, pace, pause, etc. In this context, prosody modeling is crucial for achieving expressive voice conversion that sounds natural and convincing. Unfortunately, prosody modeling is important but challenging, especially without text transcriptions. In this paper, we firstly propose a novel voice conversion framework named 'PMVC', which effectively separates and models the content, timbre, and prosodic information from the speech without text transcriptions. Specially, we introduce a new speech augmentation algorithm for robust prosody extraction. And building upon this, mask and predict mechanism is applied in the disentanglement of prosody and content information. The experimental results on the AIShell-3 corpus supports our improvement of naturalness and similarity of converted speech. Huaizhen Tang, Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ACM Multimedia | 3 |
| 2023 | Melody Generation from Lyrics with Local InterpretabilityabstractMelody generation aims to learn the distribution of real melodies to generate new melodies conditioned on lyrics, which has been a very interesting topic in the area of artificial intelligence and music. However, a challenging issue still limits the quality and reliability of melody generation conditioned on lyrics: how to enhance the interpretability between the input lyrics and generated melodies so humans can understand their relationships. To solve this issue, in this article, we propose a model for melody generation from lyrics with local interpretability, which contains two significant contributions: (i) Mutual information between input lyrics and generated melody is exploited to instruct the training of the network, which avoids the loss of content consistency during the training stage. (ii) Transformer is explored to efficiently extract semantic features from lyrics sequences, which provides more interpretable correlations between different syllables in lyrics. Experiments on a large-scale dataset with paired lyrics-melodies demonstrate that the proposed approach can generate higher-quality melodies from lyrics compared with existing methods. Wei Duan 0004, Yi Yu 0001, Xulong Zhang 0001, Suhua Tang, Wei Li 0012, Keizo Oyama |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | Avqvc: One-Shot Voice Conversion By Vector Quantization With Applying Contrastive LearningabstractVoice Conversion(VC) refers to changing the timbre of a speech while retaining the discourse content. Recently, many works have focused on disentangle-based learning techniques to separate the timbre and the linguistic content information from a speech signal. Once successful, voice conversion will be feasible and straightforward. This paper proposed a novel one-shot voice conversion framework based on vector quantization voice conversion (VQVC) and AutoVC, called AVQVC. A new training method is applied to VQVC to separate content and timbre information from speech more effectively. The result shows that this approach has better performance than VQVC in separating content and timbre to improve the sound quality of generated speech. Huaizhen Tang, Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 2 |
| 2022 | DRVC: A Framework of Any-to-Any Voice Conversion with Self-Supervised LearningabstractAny-to-any voice conversion problem aims to convert voices for source and target speakers, which are out of the training data. Previous works wildly utilize the disentangle-based models. The disentangle-based model assumes the speech consists of content and speaker style information and aims to untangle them to change the style information for conversion. Previous works focus on reducing the dimension of speech to get the content information. But the size is hard to determine to lead to the untangle overlapping problem. We propose the Disentangled Representation Voice Conversion (DRVC) model to address the issue. DRVC model is an end-to-end self-supervised model consisting of the content encoder, timbre encoder, and generator. Instead of the previous work for reducing speech size to get content, we propose a cycle for restricting the disentanglement by the Cycle Reconstruct Loss and Same Loss. The experiments show there is an improvement for converted speech on quality and voice similarity. Qiqi Wang 0005, Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 2 |
| 2022 | nnSpeech: Speaker-Guided Conditional Variational Autoencoder for Zero-Shot Multi-speaker text-to-speechabstractMulti-speaker text-to-speech (TTS) using a few adaption data is a challenge in practical applications. To address that, we propose a zero-shot multi-speaker TTS, named nnSpeech, that could synthesis a new speaker voice without fine-tuning and using only one adaption utterance. Compared with using a speaker representation module to extract the characteristics of new speakers, our method bases on a speaker-guided conditional variational autoencoder and can generate a variable Z, which contains both speaker characteristics and content information. The latent variable Z distribution is approximated by another variable conditioned on reference mel-spectrogram and phoneme. Experiments on the English corpus, Mandarin corpus, and cross-dataset proves that our model could generate natural and similar speech with only one adaption speech. Botao Zhao 0001, Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 2 |
| 2022 | Boosting StarGANs for Voice Conversion with Contrastive Discriminator
Shijing Si, Jianzong Wang, Xulong Zhang 0001, Xiaoyang Qu, Ning Cheng 0001, Jing Xiao 0006 |
ICONIP (2) | 3 |
| 2022 | Pre-Avatar: An Automatic Presentation Generation Framework Leveraging Talking AvatarabstractSince the beginning of the COVID-19 pandemic, remote conferencing and school-teaching have become important tools. The previous applications aim to save the commuting cost with real-time interactions. However, our application is going to lower the production and reproduction costs when preparing the communication materials. This paper proposes a system called Pre-Avatar, generating a presentation video with a talking face of a target speaker with 1 front-face photo and a 3-minute voice recording. Technically, the system consists of three main modules, user experience interface (UEI), talking face module and few-shot text-to-speech (TTS) module. The system firstly clones the target speaker's voice, and then generates the speech, and finally generate an avatar with appropriate lip and head movements. Under any scenario, users only need to replace slides with different notes to generate another new video. The demo has been released here11https://pre-avatar.github.io/ and will be published as free software for use. Aolan Sun, Xulong Zhang 0001, Tiandong Ling, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICTAI | 2 |
| 2022 | SUSing: SU-net for Singing Voice SynthesisabstractSinging voice synthesis is a generative task that involves multi-dimensional control of the singing model, including lyrics, pitch, and duration, and includes the timbre of the singer and singing skills such as vibrato. In this paper, we proposed SU-net for singing voice synthesis named SUSing. Synthesizing singing voice is treated as a translation task between lyrics and music score and spectrum. The lyrics and music score information is encoded into a two-dimensional feature representation through the convolution layer. The two-dimensional feature and its frequency spectrum are mapped to the target spectrum in an autoregressive manner through a SU-net network. Within the SU-net the stripe pooling method is used to replace the alternate global pooling method to learn the vertical frequency relationship in the spectrum and the changes of frequency in the time domain. The experimental results on the public dataset Kiritan show that the proposed method can synthesize more natural singing voices. Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 1 |
| 2022 | MDCNN-SID: Multi-scale Dilated Convolution Network for Singer IdentificationabstractMost singer identification methods are processed in the frequency domain, which potentially leads to information loss during the spectral transformation. In this paper, instead of the frequency domain, we propose an end-to-end architecture that addresses this problem in the waveform domain. An en-coder based on Multi-scale Dilated Convolution Neural Networks (MDCNN) was introduced to generate wave embedding from the raw audio signal. Specifically, dilated convolution layers are used in the proposed method to enlarge the receptive field, aiming to extract song-level features. Furthermore, skip connection in the backbone network integrates the multi-resolution acoustic features learned by the stack of convolution layers. Then, the obtained wave embedding is passed into the following networks for singer identification. In experiments, the proposed method achieves comparable performance on the benchmark dataset of Artist20, which significantly improves related works. Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 1 |
| 2022 | TDASS: Target Domain Adaptation Speech Synthesis Framework for Multi-speaker Low-Resource TTSabstractRecently, synthesizing personalized speech by text-to-speech (TTS) application is highly demanded. But the previous TTS models require a mass of target speaker speeches for training. It is a high-cost task, and hard to record lots of utterances from the target speaker. Data augmentation of the speeches is a solution but leads to the low-quality synthesis speech problem. Some multi-speaker TTS models are proposed to address the issue. But the quantity of utterances of each speaker imbalance leads to the voice similarity problem. We propose the Target Domain Adaptation Speech Synthesis Network (TDASS) to address these issues. Based on the backbone of the Tacotron2 model, which is the high-quality TTS model, TDASS introduces a self-interested classifier for reducing the non-target influence. Besides, a special gradient reversal layer with different operations for target and non-target is added to the classifier. We evaluate the model on a Chinese speech corpus, the experiments show the proposed method outperforms the baseline method in terms of voice quality and voice similarity. Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 1 |
| 2022 | Singer Identification for Metaverse with Timbral and Middle-Level Perceptual FeaturesabstractMetaverse is an interactive world that combines reality and virtuality, where participants can be virtual avatars. Anyone can hold a concert in a virtual concert hall, and users can quickly identify the real singer behind the virtual idol through the singer identification. Most singer identification methods are processed using the frame-level features. However, expect the singer's timbre, the music frame includes music information, such as melodiousness, rhythm, and tonal. It means the music information is noise for using frame-level features to identify the singers. In this paper, instead of only the frame-level features, we propose to use another two features that address this problem. Middle-level feature, which represents the music's melodiousness, rhythmic stability, and tonal stability, and is able to capture the perceptual features of music. The timbre feature, which is used in speaker identification, represents the singers' voice features. Furthermore, we propose a convolutional recurrent neural network (CRNN) to combine three features for singer identification. The model firstly fuses the frame-level feature and timbre feature and then combines middle-level features to the mix features. In experiments, the proposed method achieves comparable performance on an average F1 score of 0.81 on the benchmark dataset of Artist20, which significantly improves related works. Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 1 |
| 2022 | MetaSID: Singer Identification with Domain Adaptation for MetaverseabstractMetaverse has stretched the real world into unlimited space. There will be more live concerts in Metaverse. The task of singer identification is to identify the song belongs to which singer. However, there has been a tough problem in singer identification, which is the different live effects. The studio version is different from the live version, the data distribution of the training set and the test set are different, and the performance of the classifier decreases. This paper proposes the use of the domain adaptation method to solve the live effect in singer identification. Three methods of domain adaptation combined with Convolutional Recurrent Neural Network (CRNN) are designed, which are Maximum Mean Discrepancy (MMD), gradient reversal (Revgrad), and Contrastive Adaptation Network (CAN). MMD is a distance-based method, which adds domain loss. Revgrad is based on the idea that learned features can represent different domain samples. CAN is based on class adaptation, it takes into account the correspondence between the categories of the source domain and target domain. Experimental results on the public dataset of Artist20 show that CRNN-MMD leads to an improvement over the baseline CRNN by 0.14. The CRNN-RevGrad outperforms the baseline by 0.21. The CRNN-CAN achieved state of the art with the F1 measure value of 0.83 on album split. Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 1 |
| 2022 | Tiny-Sepformer: A Tiny Time-Domain Transformer Network For Speech SeparationabstractTime-domain Transformer neural networks have proven their superiority in speech separation tasks.However, these models usually have a large number of network parameters, thus often encountering the problem of GPU memory explosion.In this paper, we proposed Tiny-Sepformer, a tiny version of Transformer network for speech separation.We present two techniques to reduce the model parameters and memory consumption: (1) Convolution-Attention (CA) block, spliting the vanilla Transformer to two paths, multi-head attention and 1D depthwise separable convolution, (2) parameter sharing, sharing the layer parameters within the CA block.In our experiments, Tiny-Sepformer could greatly reduce the model size, and achieves comparable separation performance with vanilla Sepformer on WSJ0-2/3Mix datasets. Jian Luo 0007, Jianzong Wang, Ning Cheng 0001, Edward Xiao, Xulong Zhang 0001, Jing Xiao 0006 |
INTERSPEECH | 5 |
| 2022 | Adapitch: Adaption Multi-Speaker Text-to-Speech Conditioned on Pitch Disentangling with Untranscribed DataabstractIn this paper, we proposed Adapitch, a multi-speaker TTS method that makes adaptation of the supervised module with untranscribed data. We design two self supervised modules to train the text encoder and mel decoder separately with untranscribed data to enhance the representation of text and mel. To better handle the prosody information in a synthesized voice, a supervised TTS module is designed conditioned on content disentangling of pitch, text, and speaker. The training phase was separated into two parts, pretrained and fixed the text encoder and mel decoder with unsupervised mode, then the supervised mode on the disentanglement of TTS. Experiment results show that the Adaptich achieved much better quality than baseline methods. Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
MSN | 1 |
| 2022 | MetaSpeech: Speech Effects Switch Along with Environment for MetaverseabstractMetaverse expands the physical world to a new dimension, and the physical environment and Metaverse environment can be directly connected and entered. Voice is an indispensable communication medium in the real world and Metaverse. Fusion of the voice with environment effects is important for user immersion in Metaverse. In this paper, we proposed using the voice conversion based method for the conversion of target environment effect speech. The proposed method was named MetaSpeech, which introduces an environment effect module containing an effect extractor to extract the environment information and an effect encoder to encode the environment effect condition, in which gradient reversal layer was used for adversarial training to keep the speech content and speaker information while disentangling the environmental effects. From the experiment results on the public dataset of LJSpeech with four environment effects, the proposed model could complete the specific environment effect conversion and outperforms the baseline methods from the voice conversion task. Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
MSN | 1 |
| 2022 | Semi-Supervised Learning Based on Reference Model for Low-resource TTSabstractMost previous neural text-to-speech (TTS) methods are mainly based on supervised learning methods, which means they depend on a large training dataset and hard to achieve comparable performance under low-resource conditions. To ad-dress this issue, we propose a semi-supervised learning method for neural TTS in which labeled target data is limited, which can also resolve the problem of exposure bias in the previous auto-regressive models. Specifically, we pre-train the reference model based on Fastspeech2 with much source data, fine-tuned on a limited target dataset. Meanwhile, pseudo labels generated by the original reference model are used to guide the fine-tuned model's training further, achieve a regularization effect, and reduce the overfitting of the fine-tuned model during training on the limited target data. Experimental results show that our proposed semi-supervised learning scheme with limited target data significantly improves the voice quality for test data to achieve naturalness and robustness in speech synthesis. Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
MSN | 1 |
| 2022 | Improving Imbalanced Text Classification with Dynamic Curriculum LearningabstractRecent advances in pre-trained language models have improved the performance for text classification tasks. However, little attention is paid to the priority scheduling strategy on the samples during training. Humans acquire knowledge gradually from easy to complex concepts, and the difficulty of the same material can also vary significantly in different learning stages. Inspired by this insights, we proposed a novel self-paced dynamic curriculum learning (SPDCL) method for imbalanced text classification, which evaluates the sample difficulty by both linguistic character and model capacity. Meanwhile, rather than using static curriculum learning as in the existing research, our SPDCL can reorder and resample training data by difficulty criterion with an adaptive from easy to hard pace. The extensive experiments on several classification tasks show the effectiveness of SPDCL strategy, especially for the imbalanced dataset. Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
MSN | 1 |
| 2022 | Improving Speech Representation Learning via Speech-level and Phoneme-level Masking ApproachabstractRecovering the masked speech frames is widely applied in speech representation learning. However, most of these models use random masking in the pre-training. In this work, we proposed two kinds of masking approaches: (1) speech-level masking, making the model to mask more speech segments than silence segments, (2) phoneme-level masking, forcing the model to mask the whole frames of the phoneme, instead of phoneme pieces. We pre-trained the model via these two approaches, and evaluated on two downstream tasks, phoneme classification and speaker recognition. The experiments demonstrated that the proposed masking approaches are beneficial to improve the performance of speech representation. Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
MSN | 1 |
| 2022 | Linguistic-Enhanced Transformer with CTC Embedding for Speech RecognitionabstractThe recent emergence of joint CTC-Attention model shows significant improvement in automatic speech recognition (ASR). The improvement largely lies in the modeling of linguistic information by decoder. The decoder joint-optimized with an acoustic encoder renders the language model from ground-truth sequences in an auto-regressive manner during training. However, the training corpus of the decoder is limited to the speech transcriptions, which is far less than the corpus needed to train an acceptable language model. This leads to poor robustness of decoder. To alleviate this problem, we propose linguistic-enhanced transformer, which introduces refined CTC information to decoder during training process, so that the decoder can be more robust. Our experiments on AISHELL-1 speech corpus show that the character error rate (CER) is relatively reduced by up to 7 %. We also find that in joint CTC-Attention ASR model, decoder is more sensitive to linguistic information than acoustic information. Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
MSN | 1 |
| 2021 | TGAVC: Improving Autoencoder Voice Conversion with Text-Guided and Adversarial TrainingabstractNon-parallel many-to-many voice conversion remains an interesting but challenging speech processing task. Recently, AutoVC, a conditional autoencoder based method, achieved excellent conversion results by disentangling the speaker identity and the speech content using information-constraining bottlenecks. However, due to the pure autoencoder training method, it is difficult to evaluate the separation effect of content and speaker identity. In this paper, a novel voice conversion framework, named Text Guided AutoVC(TGAVC), is proposed to more effectively separate content and timbre from speech, where an expected content embedding produced based on the text transcriptions is designed to guide the extraction of voice content. In addition, the adversarial training is applied to eliminate the speaker identity information in the estimated content embedding extracted from speech. Under the guidance of the expected content embedding and the adversarial training, the content encoder is trained to extract speaker-independent content embedding from speech. Experiments on AIShell-3 dataset show that the proposed model outperforms AutoVC in terms of naturalness and similarity of converted speech. Huaizhen Tang, Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Edward Xiao, Jing Xiao 0006 |
ASRU | 2 |
| 2021 | Cyclegean: Cycle Generative Enhanced Adversarial Network for Voice ConversionabstractCycle Generative Adversarial Network (CycleGAN) for voice conversion (VC) task only used discriminators to identify whether the input voice is generated or real. It means the confrontational does not check the similarity with the target voice, leading the generated voice not much similar to the target. In this paper, instead of vocal checking, we propose to enhance the confrontation to target similarity checking that addresses this problem. A Cycle Generative Enhanced Adversarial Network (CycleGEAN) was introduced to make the original two discriminators to target classifier and non-target classifier. The target classifier aims to identify whether the target speaks the input voice or not. Similarly, the non-target classifier identifies the non-target voice. Furthermore, we add a gradient reversal layer with different operations for target and non-target. Then in each GAN, we used both classifiers. One is the discriminator, and the other is trained for using in another GAN. In experiments, the proposed method compare to CycleGAN improves Mean Opinion Score (MOS) of 0.1 and Voice Similarity Score (VSS) of 0.2 on the Voice Conversion Challenge 2018 (VCC2018) dataset. Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Edward Xiao, Jing Xiao 0006 |
ASRU | 1 |
| 2021 | Singer Identification Using Deep Timbre Feature Learning with KNN-NETabstractIn this paper, we study the issue of automatic singer identification (SID) in popular music recordings, which aims to recognize who sang a given piece of song. The main challenge for this investigation lies in the fact that a singer’s singing voice changes and intertwines with the signal of background accompaniment in time domain. To handle this challenge, we propose the KNN-Net for SID, which is a deep neural network model with the goal of learning local timbre feature representation from the mixture of singer voice and background music. Unlike other deep neural networks using the softmax layer as the output layer, we instead utilize the KNN as a more interpretable layer to output target singer labels. Moreover, attention mechanism is first introduced to highlight crucial timbre features for SID. Experiments on the existing artist20 dataset show that the proposed approach outperforms the state-of-the-art method by 4%. We also create singer32 and singer60 datasets consisting of Chinese pop music to evaluate the reliability of the proposed method. The more extensive experiments additionally indicate that our proposed model achieves a significant performance improvement compared to the state-of-the-art methods. Xulong Zhang 0001, Jiale Qian, Yi Yu 0001, Yifu Sun, Wei Li 0012 |
ICASSP | 1 |