VLDB 2026 Research / reviewers in the wild / expert
Pengyuan Zhang
dblp:65/6794
· DBLP profile ↗
99ranked-venue papers
1as first author
72since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 76 · 1 first-author · 54 since 2021Artificial intelligence and machine learning · 50 · 35 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Data-Efficient Semi-Supervised Few-Shot Speaker Verification via Prototype Space OptimizationabstractSpeaker verification technology has widespread applications across many domains, benefiting from deep learning advancements. However, due to the high cost of acquiring labeled data, semi-supervised learning has emerged as a prominent research focus. Current semi-supervised learning frameworks commonly suffer from two limitations: (1) the labeled data distribution is often restricted, and (2) they still rely on a considerable amount of labeled data. To address these issues, we propose three different distribution scenarios of labeled data and construct a general semi-supervised framework. Furthermore, to enhance the guidance efficacy of limited labeled data, we innovatively employ prototype space optimization to strengthen the model's discriminative capability under low-resource scenarios. Experimental results demonstrate that on the Vox1-o test set, our approach achieves a 41.7% relative reduction in equal error rate compared to self-supervised baselines, and a 28.9% improvement over conventional semi-supervised framework baselines. Zhenduo Zhao, Shisong Wu, Pengyuan Zhang, Xueshuai Zhang, Yonghong Yan 0002 |
IEEE Signal Process. Lett. | 5 |
| 2026 | SIM-IDE-FFM: Similarity-Based Residual-Identity Feature Fusion for Speaker VerificationabstractRecent advances in speaker verification have focused on two primary areas: deepening or widening backbone networks and introducing different attention mechanisms in the main residual branch to enhance discriminative power. However, the shortcut branches in residual architectures are still treated as simple identity mappings, leaving their representational potential underexplored. In this paper, we propose SIM-IDE-FFM, a similarity-based identity feature enhancement framework which enables the trivial shortcut branch to explore correlations between identity feature maps through a similarity-based attention mechanism. SIM-IDE introduces an attention mechanism at the shortcut branch that models inter-channel similarity of identity inputs, enhancing channel-wise representations. A non-linear Feature Fusion Module (FFM) is further designed to efficiently combine original and enhanced features through nonlinear depthwise interactions. Without altering the backbones, SIM-IDE-FFM can be seamlessly integrated into ResNet, EcapaTDNN, and other state-of-the-art (SOTA) architectures. Experiments on VoxCeleb1 and cross-age subsets demonstrate consistent relative improvements across different frameworks, validating its robustness and generalization. Baizhu Li, Xueshuai Zhang, Pengyuan Zhang, Yonghong Yan 0002 |
IEEE Signal Process. Lett. | 4 |
| 2026 | Coarse Labels Matter: Revisiting the Role of Coarse-Grained Supervision in Fine-Grained LearningabstractThe prohibitive cost of acquiring high-quality fine-grained annotations has spurred significant interest in leveraging readily available coarse labels for fine-grained learning. However, prevailing approaches tend to rely on increasingly sophisticated unsupervised methods to define fine-grained proxy tasks, with coarse labels often playing an auxiliary role. In this paper, we propose CSer, a framework designed to maximize the utility of coarse label information for Coarse-to-Fine learning. Specifically, to reconcile the conflict between preserving fine-grained feature diversity and maintaining strong coarse-grained supervision, our coarse-grained self-distillation strategy fortifies the backbone's discriminative power by distilling knowledge from the final classifier to intermediate layers. Concurrently, we introduce dense supervision on common component features within each coarse class, which are decoupled using Non-negative Matrix Factorization. This enhances responses to distinct components, thereby mitigating the simplicity bias in embeddings that can arise under coarse supervision. Moreover, we leverage relationships among intra-class samples to dynamically adjust the negative sampling strategy in contrastive learning, thereby constructing distinct fine-grained class relationships tailored to different coarse classes. Extensive experiments conducted on multiple benchmark datasets demonstrate the effectiveness of our method, yielding state-of-the-art results surpassing competing methods. Xin-Yang Zhao, Pengyuan Zhang, Qiyuan Zhuang, Yazhou Yao, Xiu-Shen Wei |
IEEE Trans. Image Process. | 2 |
| 2025 | Pitch-Assistant Harmonic Recovery for Efficient Speech EnhancementabstractWith the rapid development of low-resource online speech enhancement models, noise suppression can now be achieved with significantly fewer model parameters. However, these models often suffer from limited effectiveness in preserving speech quality. In particular, most existing online speech enhancement methods tend to distort the harmonic structure of speech while performing noise reduction, leading to noticeable degradation in perceptual quality. In this paper, we propose a novel model architecture called Pitch-Assistant Harmonic Recovery for Efficient Speech Enhancement (PHRSE). The model operates on low-dimensional Bark-scale spectral features to perform noise suppression, while leveraging estimated fundamental frequency information to guide the reconstruction of harmonic components. This pitch-guided strategy enables the model to preserve the speech’s natural harmonic structure more effectively. Experimental results demonstrate that PHRSE not only achieves higher perceptual speech quality compared to existing benchmarks, but also maintains real-time performance with significantly lower computational overhead, making it suitable for online and resource-constrained scenarios. Zengqiang Shang, Haoyuan Xie, Mou Wang, Pengyuan Zhang |
ASRU | 6 |
| 2025 | SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue GenerationabstractRecently, "textless" speech language models (SLMs) based on speech units have made huge progress in generating naturalistic speech, including non-verbal vocalizations. However, the generated speech samples often lack semantic coherence. In this paper, we propose SLM and LLM Integration for spontaneous spoken Dialogue gEneration (SLIDE). Specifically, we first utilize an LLM to generate the textual content of spoken dialogue. Next, we convert the textual dialogues into phoneme sequences and use a two-tower transformer-based duration predictor to predict the duration of each phoneme. Finally, an SLM conditioned on the spoken phoneme sequences is used to vocalize the textual dialogue. Experimental results on the Fisher dataset demonstrate that our system can generate naturalistic spoken dialogue while maintaining high semantic coherence. Haitian Lu, Gaofeng Cheng, Liuping Luo, Leying Zhang, Yanmin Qian, Pengyuan Zhang |
ICASSP | 6 |
| 2025 | Debiased Training For Semi-supervised Sound Event DetectionabstractRecently, semi-supervised sound event detection has attracted increasing research interest due to the scarcity of labeled data. However, traditional semi-supervised learning methods can lead to training instability and confirmation bias because of potentially incorrect pseudo labels. To address this issue, we propose the debiased training, a novel approach to reduce the inherent bias of pseudo labels. Debiased training can effectively decouple the generation and utilization of pseudo labels to mitigate the error accumulation and promote model’s robustness against biased pseudo labels. In addition, we introduce the channel restruction module (CRM) to decrease redundant computing and facilitate representation ability. Experimental results on DCASE 2023 task4 dataset show that the proposed methods significantly enhance the performance of semi-supervised methods while maintaining relatively low computational complexity. Shengchang Xiao, Xueshuai Zhang, Pengyuan Zhang, Yonghong Yan 0002 |
ICASSP | 3 |
| 2025 | Restoring Harmonics: Enhancing Speech Quality with Deep Mask and Harmonic Restoration Network
Zengqiang Shang, Mou Wang, Pengyuan Zhang |
INTERSPEECH | 5 |
| 2025 | SMIIP-NV: A Multi-Annotation Non-Verbal Expressive Speech Corpus in Mandarin for LLM-Based Speech SynthesisabstractIn natural language communication, emotions are often conveyed through non-verbal sounds (NVs), such as laughter, crying, cough and so on. However, most existing text-to-speech (TTS) corpora lack annotations for these non-verbal sounds, leading to a scarcity of systems capable of generating them. To address this gap, we introduce SMIIP-NV, a non-verbal speech synthesis corpus annotated with both emotions and non-verbal sounds, including laughter, crying, and cough. To the best of our knowledge, SMIIP-NV is the largest publicly available open-source expressive speech corpus that includes non-verbal speech and rich annotations. It comprises 33 hours of speech data, covering five distinct emotions and three types of non-verbal sounds, with detailed transcriptions and precise timestamps for each occurrence of non-verbal sounds. Additionally, the corpus provides annotations for speech segments that contain laughter or crying. To demonstrate the utility of this dataset, we establish a baseline for non-verbal speech synthesis by employing a lightweight large language model (LLM). The SMIIP-NV dataset and static audio demonstrations are publicly available at https://axunyii.github.io/SMIIP-NV. The interactive real-time demonstrations can be accessed at https://huggingface.co/spaces/xunyi/SMIIP-NV_Finetuned_CosyVoice2. Zhuojun Wu, Dong Liu 0028, Juan Liu 0007, Yechen Wang, Hui Bu, Pengyuan Zhang, Ming Li 0026 |
ACM Multimedia | 8 |
| 2025 | Leveraging distance information for generalized spoofing speech detection
Jingze Lu, Zengqiang Shang, Pengyuan Zhang |
Comput. Speech Lang. | 6 |
| 2025 | A comprehensive validation study on the influencing factors of cough-based COVID-19 detection through multi-center data with abundant metadata
Jiakun Shen, Xueshuai Zhang, Yanfen Tang, Pengyuan Zhang, Yonghong Yan 0002, Pengfei Ye, Shaoxing Zhang |
J. Biomed. Informatics | 4 |
| 2025 | Flexpéro: Flexible Expressive Zero-Shot Speech Refinement via In-Context LearningabstractControlling speech expressiveness has emerged as a critical research frontier in speech generation, focusing on synthesizing natural, human-like speech that accurately conveys intended psychological and emotional states. While many large-scale models have demonstrated sufficient zero-shot capability by conditioning the acoustic model on reference speech—which provides cues on speaker identity and style—they often fall short of meeting desired emotional or prosodic targets at fine-grained levels. To address this challenge, we propose a novel speech refinement method based on a zero-shot voice synthesis model that can flexibly and interactively enhance expressiveness on unsatisfactory speech segments. It supports emotion modulation through chunk-wise valence/arousal and flexible keyframe-based prediction of pitch and energy, allowing for the creation of any prosodic patterns. Experimental results show that our method achieves fine-grained control, thus enriching the expressiveness of zero-shot synthetic speech. Hua Hua, Zengqiang Shang, Xuyuan Li, Pengyuan Zhang |
IEEE Signal Process. Lett. | 5 |
| 2025 | Multi-Branch Coordinate Attention With Channel Dynamic Difference for Speaker VerificationabstractIn prior studies, researchers have proved the excellent performance of various deep neural networks on speaker verification (SV). However, most of the improvements of SV systems are aimed at modifying the specific network structure to enhance its robustness but with limited flexibility. In this paper, MCA-CDD, which is a novel universal residual block module and can easily replace the original residual block without increasing the block number, is proposed with adaptive multi-branch coordinate attention (MCA) and channel-level dynamic difference (CDD). The design of multiple branches enables CA to model time and frequency at multiple scales. In addition, CDD fusion is applied into the feature fusion process of the residual block. The CDD fusion at shallow positions of the model enables the model to learn the detailed speaker-related dynamic texture information of speech. Experiments are conducted on several ResNet backbones and the results on different seen and unseen test sets show significant improvements, outperforming the baseline by about relatively 15-20%. Zhenduo Zhao, Shisong Wu, Xueshuai Zhang, Pengyuan Zhang, Yonghong Yan 0002 |
IEEE Signal Process. Lett. | 5 |
| 2024 | Fetal Heart Sounds Classification Using Time-Cyclic Frequency Spectrogram and Hybrid Attention NetworkabstractFetal heart monitoring is a crucial method for assessing fetal health status. However, commonly used techniques may pose potential medical risks and are not suitable for long-term monitoring. In this paper, we propose utilizing fetal heart sounds (FHS) for classifying fetal health status, taking advantage of its non-invasive, safe, straightforward, and cost-effective properties. Firstly, we introduce a novel acoustic feature for fetal heart sounds, termed the time-cyclic frequency spectrogram. This feature emphasizes the periodicity of heartbeats and effectively captures the changes in fetal heart rate. Additionally, we implement a frequency band energy-weighted algorithm to mitigate interference from periodic noises. Secondly, we propose a hybrid attention network that integrates both global-local attention and time-cyclic frequency attention. This network leverages medical prior knowledge to focus on the most critical aspects of the time-cyclic frequency spectrogram. Experimental results demonstrate that the proposed feature effectively characterizes variations in fetal heart rate, and the hybrid attention network can accurately capture spectral line variations, leading to improved classification of fetal health conditions. Shengchang Xiao, Xueshuai Zhang, Pengyuan Zhang, Yonghong Yan 0002 |
BIBM | 3 |
| 2024 | Make Audio Solely Drive Lip in Talking Face Video Synthesis
Xing Bai, Jun Zhou 0024, Pengyuan Zhang, Ruipeng Hao |
ICANN (3) | 3 |
| 2024 | Snore Sound Features Based on Percussive Enhancing and Positional Encoding Combined with Multi-Task Learning for Osahs DetectionabstractObstructive sleep apnea hypopnea syndrome (OSAHS) is a serious sleep disorder. As the typical symptom of OSAHS, snoring has been proved effective in OSAHS diagnosis and potential to replace the current laborious and expensive polysomnography. However, the lack of analysis on the characteristics of pathological snoring sounds limits the diagnosing performance. In this paper, we propose novel sound features for the classification of OSA, hypopnea and normal snores. The proposed features are based on percussive enhancing and positional encoding as the snores exhibit different percussive properties and temporal traits due to the disease generation mechanisms. To enhance the classification performance, we propose a multi-task learning framework to aid the main classification task by simultaneous learning of two related simple tasks. Experiments on real-recorded snoring sounds show that the proposed methods can greatly improve the classification AUC and ACC and the proposed system performs better than those in other literatures. Aolin Hu, Xueshuai Zhang, Shaoxing Zhang, Pengyuan Zhang, Pengfei Ye, Qingwei Zhao, Yonghong Yan 0002 |
ICASSP | 4 |
| 2024 | One-Class Knowledge Distillation for Spoofing Speech DetectionabstractThe detection of spoofing speech generated by unseen algorithms remains an unresolved challenge. One reason for the lack of generalization ability is that traditional detecting systems follow the binary classification paradigm, which inherently assumes the possession of prior knowledge of spoofing speech. One-class methods attempt to learn the distribution of bonafide speech and are inherently suited to the task where spoofing speech exhibits significant differences. However, training a one-class system using only bonafide speech is challenging. In this paper, we introduce a teacher-student framework to provide guidance for the training of a one-class model. The proposed one-class knowledge distillation method outperforms other state-of-the-art methods on the ASVspoof 21DF and InTheWild datasets, demonstrating its superior generalization ability. Jingze Lu, Zengqiang Shang, Pengyuan Zhang |
ICASSP | 5 |
| 2024 | One-Epoch Training with Single Test Sample in Test Time for Better Generalization of Cough-Based Covid-19 Detection ModelabstractThe outbreak of COVID-19 has raised researchers’ attention to audio-based rapid disease detection. Most of the previous studies have obtained competitive detection performance. However, these results are usually obtained by testing data from the same source offline. When making cross-dataset testing, the performance may deteriorate dramatically due to inconsistent data distribution between different datasets. In addition, in practical application, the model has to make a prediction for the current test audio without any prior information, which requires good model generalization under limited training data. To address the above issues, we adopt a test-time training framework to achieve a cough-based COVID-19 detection model with better generalizability. In the model development stage, resnet18 serves as the backbone network and a self-supervised learning branch is added as an auxiliary task. In testing stage, the model parameters are first fine-tuned by the self-supervised branch with the single test audio as input, and then the classification head outputs predictions. The proposed method is validated on three open-source datasets using a variety of hyperparameters. In cross-dataset testing, AUC and UAR increase by 3.65% and 3% on average absolutely, respectively. The results show that the proposed framework is applicable to improve the model performance in practical application. Jiakun Shen, Xueshuai Zhang, Pengyuan Zhang, Yonghong Yan 0002, Qingwei Zhao, Ta Li, Yanfen Tang, Shaoxing Zhang |
ICASSP | 3 |
| 2024 | Improving Short Utterance Anti-Spoofing with Aasist2abstractThe wav2vec 2.0 and integrated spectro-temporal graph attention network (AASIST) based countermeasure achieves great performance in speech anti-spoofing. However, current spoof speech detection systems have fixed training and evaluation durations, while the performance degrades significantly during short utterance evaluation. To solve this problem, AASIST can be improved to AASIST2 by modifying the residual blocks to Res2Net blocks. The modified Res2Net blocks can extract multi-scale features and improve the detection performance for speech of different durations, thus improving the short utterance evaluation performance. On the other hand, adaptive large margin fine-tuning (ALMFT) has achieved performance improvement in short utterance speaker verification. Therefore, we apply Dynamic Chunk Size (DCS) and ALMFT training strategies in speech anti-spoofing to further improve the performance of short utterance evaluation. Experiments demonstrate that the proposed AASIST2 improves the performance of short utterance evaluation while maintaining the performance of regular evaluation on different datasets. Jingze Lu, Zengqiang Shang, Pengyuan Zhang |
ICASSP | 5 |
| 2024 | Expressive paragraph text-to-speech synthesis with multi-step variational autoencoder
Xuyuan Li, Zengqiang Shang, Peiyang Shi, Hua Hua, Ta Li, Pengyuan Zhang |
INTERSPEECH | 6 |
| 2024 | Improving Copy-Synthesis Anti-Spoofing Training Method with Rhythm and Speaker Perturbation
Jingze Lu, Zengqiang Shang, Pengyuan Zhang |
INTERSPEECH | 6 |
| 2024 | Emilia: An Extensive, Multilingual, and Diverse Speech Dataset For Large-Scale Speech GenerationabstractRecent advancements in speech generation models have been significantly driven by the use of large-scale training data. However, producing highly spontaneous, human-like speech remains a challenge due to the scarcity of large, diverse, and spontaneous speech datasets. In response, we introduce Emilia, the first large-scale, multilingual, and diverse speech generation dataset. Emilia starts with over 101k hours of speech across six languages, covering a wide range of speaking styles to enable more natural and spontaneous speech generation. To facilitate the scale-up of Emilia, we also present Emilia-Pipe, the first open-source preprocessing pipeline designed to efficiently transform raw, in-the-wild speech data into high-quality training data with speech annotations. Experimental results demonstrate the effectiveness of both Emilia and Emilia-Pipe. Demos are available at: https://emilia-dataset.github.io/Emilia-Demo-Page/. Haorui He, Zengqiang Shang, Chaoren Wang, Xuyuan Li, Yicheng Gu, Hua Hua, Liwei Liu 0008, Jiaqi Li 0030, Peiyang Shi, Yuancheng Wang, Kai Chen 0026, Pengyuan Zhang, Zhizheng Wu 0001 |
SLT | 13 |
| 2024 | The Database and Benchmark For the Source Speaker Tracing Challenge 2024abstractVoice conversion (VC) systems can transform audio to mimic another speaker’s voice, thereby attacking speaker verification (SV) systems. However, ongoing studies on source speaker verification (SSV) are hindered by limited data availability and methodological constraints. This paper presents the Source Speaker Tracking Challenge (SSTC) on STL 2024, which aims to fill the gap in the database and benchmark for the SSV task. In this study, we generate a large-scale converted speech database with 16 common VC methods and train a batch of baseline systems based on the MFA-Conformer architecture. In addition, we introduced a related task called conversion method recognition, with the aim of assisting the SSV task. We expect SSTC to be a platform for advancing the development of the SSV task and provide further insights into the performance and limitations of current SV systems against VC attacks. Further details about SSTC can be found here1.1https://sstc-challenge.github.io/ Ze Li 0003, Yuke Lin, Hongbin Suo, Pengyuan Zhang, Yanzhen Ren, Zexin Cai, Hiromitsu Nishizaki, Ming Li 0026 |
SLT | 5 |
| 2024 | An efficient loss function and deep learning approach for ranking stock returns in the absence of prior knowledge
Wenkai Zhang 0001, Ming Zhang 0035, Jun Zhou 0024, Pengyuan Zhang |
Inf. Process. Manag. | 6 |
| 2024 | Synthetic Speech Detection Based on the Temporal Consistency of Speaker FeaturesabstractCurrent synthetic speech detection (SSD) methods perform well on specific datasets but require improvement in interpretability and robustness. One possible reason is the lack of interpretability analysis of synthetic speech defects. In this paper, the flaws in the temporal consistency (TC) of speaker features inherent in the speech synthesis process are analyzed. Differences in the TC of intra-utterance speaker features arise due to limited control over speaker features during speech synthesis. The speech generated by text-to-speech algorithms exhibits higher TC, while the speech generated by voice conversion algorithms yeilds slightly lower TC compared to bonafide speech. Based on this finding, a new SSD method based on the TC of speaker features is proposed. Modeling the TC of intra-utterance speaker features extracted by a pre-trained ASV system can be used for SSD. The proposed method achieves equal error rates of 0.84%, 3.93%, 12.98% and 24.66% on the ASVspoof 2019 LA, 2021 LA, 2021 DF and IntheWild evaluation datasets, respectively, demonstrating strong interpretability and robustness. Zhuo Li 0020, Jingze Lu, Pengyuan Zhang |
IEEE Signal Process. Lett. | 5 |
| 2024 | Prototype Division for Self-Supervised Speaker VerificationabstractSelf-supervised learning has shown promising performance on speaker verification tasks, among which Self DIstillation with NO labels (DINO) is currently a widely adopted framework. As one of the unsupervised deep clustering methods, the number of valid prototypes in DINO is far less than the speakers in practical applications and remains unchanged throughout the training period, leading to severe speaker confusion and performance degradation. Therefore, a strategy named prototype division (PD) is proposed to iteratively generate fine-grained prototypes in the projection space based on the converged model to separate confused categories, where new prototypes are derived from the neighborhood of the existing valid prototypes by clustering or sampling. The results on Vox1O achieve significant improvements, relatively outperforming the baseline by 31.1% without any auxiliary loss. Further experiments on CN-Celeb also show stable improvement, proving the consistency of the proposed method. Zhenduo Zhao, Zhuo Li 0020, Xueshuai Zhang, Pengyuan Zhang |
IEEE Signal Process. Lett. | 5 |
| 2024 | Interrelate Training and Clustering for Online Speaker DiarizationabstractIn clustering-based speaker diarization systems, the embedding clusters for distinctive speakers exhibit wide variability in size and density, posing difficulty for clustering accuracy. In spite of this, with the assistance of the overall distance relationships among speaker embeddings, most of the embeddings can be grouped to the correct cluster by sophisticated offline clustering algorithms. However, in online scenarios, such a complete distance relationships of the embeddings can not be obtained due to the incremental arrival of embeddings. Consequently, determining the number of clusters and then correctly grouping the embeddings become challenging in an online fashion. Furthermore, errors would accumulate quickly over time if the online clustering algorithm assigns the embeddings into clusters erroneously in the beginning. To address these problems, we designed a novel framework for online clustering. To reduce the high variability of speaker embeddings, we proposed the clustering guided embedding extractor training (CGEET) algorithm to encourage similarity between the size of the embedding space for different speakers in attempt to simplify the distance relationships of embeddings. The CGEET algorithm can grasp the distance information of the entire speaker embedding space and provide it to the online clustering algorithm. With this preliminary information, the distance thresholds guided online clustering (DTGOC) algorithm then processes incoming embeddings using a divide-and-conquer approach. It first handles the embeddings with explicit distance relationships and then searches for possible path combination they have with remaining embeddings in an online fashion. Moreover, in order to utilize the distance relationships of embeddings that are far apart in time, an online re-clustering strategy is incorporated in our DTGOC algorithm, which can alleviate error accumulation during online clustering. By implementing the above innovations, our proposed online clustering system achieves 14.00% DER with collar 0.25 at 2.5 s latency on the AISHELL-4, while the DER of the offline agglomerative hierarchical clustering system is 14.54%. Gaofeng Cheng, Runyan Yang, Pengyuan Zhang, Yonghong Yan 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2024 | Boosting Cross-Domain Speech Recognition With Self-SupervisionabstractThe cross-domain performance of automatic speech recognition (ASR) could be severely hampered due to the mismatch between training and testing distributions. Since the target domain usually lacks labeled data, and domain shifts exist at acoustic and linguistic levels, it is challenging to perform unsupervised domain adaptation (UDA) for ASR. Previous work has shown that self-supervised learning (SSL) or pseudo-labeling (PL) is effective in UDA by exploiting the self-supervisions of unlabeled data. However, these self-supervisions also face performance degradation in mismatched domain distributions, which previous work fails to address. This work presents a systematic UDA framework to fully utilize the unlabeled data with self-supervision in the pre-training and fine-tuning paradigm. On the one hand, we apply continued pre-training and data replay techniques to mitigate the domain mismatch of the SSL pre-trained model. On the other hand, we propose a domain-adaptive fine-tuning approach based on the PL technique with three unique modifications: Firstly, we design a dual-branch PL method to decrease the sensitivity to the erroneous pseudo-labels; Secondly, we devise an uncertainty-aware confidence filtering strategy to improve pseudo-label correctness; Thirdly, we introduce a two-step PL approach to incorporate target domain linguistic knowledge, thus generating more accurate target domain pseudo-labels. Experimental results on various cross-domain scenarios demonstrate that the proposed approach effectively boosts the cross-domain performance and significantly outperforms previous approaches. Han Zhu 0004, Gaofeng Cheng, Jindong Wang 0001, Wenxin Hou, Pengyuan Zhang, Yonghong Yan 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2024 | Modality-collaborative Transformer with Hybrid Feature Reconstruction for Robust Emotion RecognitionabstractAs a vital aspect of affective computing, Multimodal Emotion Recognition has been an active research area in the multimedia community. Despite recent progress, this field still confronts two major challenges in real-world applications: (1) improving the efficiency of constructing joint representations from unaligned multimodal features and (2) relieving the performance decline caused by random modality feature missing. In this article, we propose a unified framework, Modality-Collaborative Transformer with Hybrid Feature Reconstruction (MCT-HFR), to address these issues. The crucial component of MCT is a novel attention-based encoder that concurrently extracts and dynamically balances the intra- and inter-modality relations for all associated modalities. With additional modality-wise parameter sharing, a more compact representation can be encoded with less time and space complexity. To improve the robustness of MCT, we further introduce HFR, which consists of two modules: Local Feature Imagination (LFI) and Global Feature Alignment (GFA). During model training, LFI leverages complete features as supervisory signals to recover local missing features, while GFA is designed to reduce the global semantic gap between pairwise-complete and -incomplete representations. Experimental evaluations on two popular benchmark datasets demonstrate that our proposed method consistently outperforms advanced baselines in both complete and incomplete data scenarios. Chengxin Chen, Pengyuan Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Piecewise Position Encoding in Convolutional Neural Network for Cough-Based Covid-19 DetectionabstractA fast and efficient COVID-19 detection method is of vital importance to control the spread of the epidemic. Many studies have achieved good performance on cough-based COVID19 detection in the past two years. However, the effect of position information in time-frequency features of cough audio has been less considered in previous studies. Even the convolutional neural networks that are capable to learn position information may be affected by small transformations of input features. Therefore, we propose piecewise position encoding added to time-frequency features to provide supplementary position information explicitly. Considering the differences in recording devices among different people, we use modified instance normalization to achieve better generalization. The proposed methods are validated on three open-sourced datasets and achieve significant improvements in AUC and UAR. The proposed model also shows competitive results in detecting asymptomatic patients. Jiakun Shen, Xueshuai Zhang, Pengyuan Zhang, Yonghong Yan 0002, Shaoxing Zhang, Yanfen Tang, Fujie Zhang, Aijun Sun |
ICASSP | 3 |
| 2023 | Multi-Dimensional Frequency Dynamic Convolution with Confident Mean Teacher for Sound Event DetectionabstractRecently, convolutional neural networks (CNNs) have been widely used in sound event detection (SED). However, traditional convolution is deficient in learning time-frequency domain representation of different sound events. To address this issue, we propose multi-dimensional frequency dynamic convolution (MFDConv), a new design that endows convolutional kernels with frequency-adaptive dynamic properties along multiple dimensions. MFDConv utilizes a novel multi-dimensional attention mechanism with a parallel strategy to learn complementary frequency-adaptive attentions, which substantially strengthen the feature extraction ability of convolutional kernels. Moreover, in order to promote the performance of mean teacher, we propose the confident mean teacher to increase the accuracy of pseudo-labels from the teacher and train the student with high confidence labels. Experimental results show that the proposed methods achieve 0.470 and 0.692 of PSDS1 and PSDS2 on the DESED real validation dataset. Shengchang Xiao, Xueshuai Zhang, Pengyuan Zhang |
ICASSP | 3 |
| 2023 | PCF: ECAPA-TDNN with Progressive Channel Fusion for Speaker VerificationabstractECAPA-TDNN is currently the most popular TDNN-series model for speaker verification, which refreshed the state-of-the-art (SOTA) performance of TDNN models. However, one-dimensional convolution has a global receptive field over the feature channel. It destroys the time-frequency relevance of the spectrogram. Besides, as ECAPA-TDNN only has five layers, a much shallower structure compared to ResNet restricts the capability to generate deep representations. To further improve ECAPA-TDNN, we propose a progressive channel fusion strategy that splits the spectrogram across the feature channel and gradually expands the receptive field through the network. Secondly, we enlarge the model by extending the depth and adding branches. Our proposed model achieves EER with 0.718 and minDCF(0.01) with 0.0858 on vox1o, relatively improved 16.1% and 19.5% compared with ECAPA-TDNN(C=1024). Zhenduo Zhao, Pengyuan Zhang |
ICASSP | 4 |
| 2023 | How to make embeddings suitable for PLDA
Zhuo Li 0020, Runqiu Xiao, Hangting Chen, Zhenduo Zhao, Pengyuan Zhang |
Comput. Speech Lang. | 6 |
| 2023 | SFA: Searching faster architectures for end-to-end automatic speech recognition models
Ta Li, Pengyuan Zhang, Yonghong Yan 0002 |
Comput. Speech Lang. | 3 |
| 2023 | Enhancing stock movement prediction with market index and curriculum learning
Wenkai Zhang 0001, Xuejun Zhang 0002, Jun Zhou 0024, Pengyuan Zhang |
Expert Syst. Appl. | 5 |
| 2023 | First coarse, fine afterward: A lightweight two-stage complex approach for monaural speech enhancement
Feng Dang, Hangting Chen, Pengyuan Zhang, Yonghong Yan 0002 |
Speech Commun. | 4 |
| 2023 | So-DAS: A Two-Step Soft-Direction-Aware Speech Separation FrameworkabstractMost existing direction-aware speech separation systems lead to performance degradation when the angle difference between speakers is small due to the low spatial discrimination. To address this issue, we propose a two-step soft-direction-aware speech separation (So-DAS) framework, which consists of a direction of arrival (DOA) estimation module and a speech separation module. First, the two modules are individually optimized, and directional features (DFs) derived from ground-truth DOAs are utilized as spatial information to facilitate the separation module. Next, the two modules are cascaded and optimized with only separation loss, and the DFs are generated using the estimator outputs. By this means, the consistency between the two modules is strengthened, and thus spatial cues that are more beneficial to the separation task can be exploited by the network itself. The experimental results show that compared to the baselines, DFs extracted by our proposed method provides clearer superiority, especially when the angle difference between speakers is small. In addition, our approach yields a state-of-the-art word error rate of 3.4% on the real-recorded utterance-wise LibriCSS dataset. Yi Yang 0057, Qingwei Zhao, Pengyuan Zhang |
IEEE Signal Process. Lett. | 4 |
| 2023 | The Impact of Silence on Speech Anti-SpoofingabstractThe current speech anti-spoofing countermeasures (CMs) show excellent performance on specific datasets. However, removing the silence of test speech through Voice Activity Detection (VAD) can severely degrade performance. In this paper, the impact of silence on speech anti-spoofing is analyzed. First, the reasons for the impact are explored, including the proportion of silence duration and the content of silence. The proportion of silence duration in spoof speech generated by text-to-speech (TTS) algorithms is lower than that in bonafide speech. And the content of silence generated by different waveform generators varies compared to bonafide speech. Then the impact of silence on model prediction is explored. Even after retraining, the spoof speech generated by neural network based end-to-end TTS algorithms suffers a significant rise in error rates when the silence is removed. To demonstrate the reasons for the impact of silence on CMs, the attention distribution of a CM is visualized through class activation mapping (CAM). Furthermore, the implementation and analysis of the experiments masking silence or non-silence demonstrates the significance of the proportion of silence duration for detecting TTS and the importance of silence content for detecting voice conversion (VC). Based on the experimental results, improving the robustness of CMs against unknown spoofing attacks by masking silence is also proposed. Finally, the attacks on anti-spoofing CMs through concatenating silence, and the mitigation of VAD and silence attack through low-pass filtering are introduced. Zhuo Li 0020, Jingze Lu, Hua Hua, Pengyuan Zhang |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2023 | Alternative Pseudo-Labeling for Semi-Supervised Automatic Speech RecognitionabstractWhen labeled data is insufficient, semi-supervised learning with the pseudo-labeling technique can significantly improve the performance of automatic speech recognition. However, pseudo-labels are often noisy, containing numerous incorrect tokens. Taking noisy labels as ground-truth in the loss function results in suboptimal performance. Previous works attempted to mitigate this issue by either filtering out the nosiest pseudo-labels or improving the overall quality of pseudo-labels. While these methods are effective to some extent, it is unrealistic to entirely eliminate incorrect tokens in pseudo-labels. In this work, we propose a novel framework named alternative pseudo-labeling to tackle the issue of noisy pseudo-labels from the perspective of the training objective. The framework comprises several components. Firstly, a generalized CTC loss function is introduced to handle noisy pseudo-labels by accepting alternative tokens in the positions of incorrect tokens. Applying this loss function in pseudo-labeling requires detecting incorrect tokens in the predicted pseudo-labels. In this work, we adopt a confidence-based error detection method that identifies the incorrect tokens by comparing their confidence scores with a given threshold, thus necessitating the confidence score to be discriminative. Hence, the second proposed technique is the contrastive CTC loss function that widens the confidence gap between the correctly and incorrectly predicted tokens, thereby improving the error detection ability. Additionally, obtaining satisfactory performance with confidence-based error detection typically requires extensive threshold tuning. Instead, we propose an automatic thresholding method that uses labeled data as a proxy for determining the threshold, thus saving the pain of manual tuning. Experiments demonstrate that alternative pseudo-labeling outperforms existing pseudo-labeling approaches on datasets in various domains and languages. Han Zhu 0004, Dongji Gao, Gaofeng Cheng, Daniel Povey, Pengyuan Zhang, Yonghong Yan 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2022 | DPT-FSNet: Dual-Path Transformer Based Full-Band and Sub-Band Fusion Network for Speech EnhancementabstractSub-band models have achieved promising results due to their ability to model local patterns in the spectrogram. Some studies further improve the performance by fusing sub-band and full-band information. However, the structure for the full-band and sub-band fusion model was not fully explored. This paper proposes a dual-path transformer-based full-band and sub-band fusion network (DPT-FSNet) for speech enhancement in the frequency domain. The intra and inter parts of the dual-path transformer model sub-band and full-band information, respectively. The features utilized by our proposed method are more interpretable than those utilized by the time-domain dual-path transformer. We conducted experiments on the Voice Bank + DEMAND and Interspeech 2020 Deep Noise Suppression (DNS) datasets to evaluate the proposed method. Experimental results show that the proposed method outperforms the current state-of-the-art. Feng Dang, Hangting Chen, Pengyuan Zhang |
ICASSP | 3 |
| 2022 | Improving CTC-Based Speech Recognition Via Knowledge Transferring from Pre-Trained Language ModelsabstractRecently, end-to-end automatic speech recognition models based on connectionist temporal classification (CTC) have achieved impressive results, especially when fine-tuned from wav2vec2.0 models. Due to the conditional independence assumption, CTC-based models are always weaker than attention-based encoder-decoder models and require the assistance of external language models (LMs). To solve this issue, we propose two knowledge transferring methods that leverage pre-trained LMs, such as BERT and GPT2, to improve CTC-based models. The first method is based on representation learning, in which the CTC-based models use the representation produced by BERT as an auxiliary learning target. The second method is based on joint classification learning, which combines GPT2 for text modeling with a hybrid CTC/attention architecture. Experiment on AISHELL-1 corpus yields a character error rate (CER) of 4.2% on the test set. When compared to the vanilla CTC-based models fine-tuned from the wav2vec2.0 models, our knowledge transferring method reduces CER by 16.1% relatively without external LMs. Keqi Deng, Songjun Cao, Gaofeng Cheng, Pengyuan Zhang |
ICASSP | 7 |
| 2022 | Improving Non-Autoregressive End-to-End Speech Recognition with Pre-Trained Acoustic and Language ModelsabstractWhile Transformers have achieved promising results in end-to-end (E2E) automatic speech recognition (ASR), their autoregressive (AR) structure becomes a bottleneck for speeding up the decoding process. For real-world deployment, ASR systems are desired to be highly accurate while achieving fast inference. Non-autoregressive (NAR) models have become a popular alternative due to their fast inference speed, but they still fall behind AR systems in recognition accuracy. To fulfill the two demands, in this paper, we propose a NAR CTC/attention model utilizing both pre-trained acoustic and language models: wav2vec2.0 and BERT. To bridge the modality gap between speech and text representations obtained from the pre-trained models, we design a novel modality conversion mechanism, which is more suitable for logographic languages. During inference, we employ a CTC branch to generate a target length, which enables the BERT to predict tokens in parallel. We also design a cache-based CTC/attention joint decoding method to improve the recognition accuracy while keeping the decoding speed fast. Experimental results show that the proposed NAR model greatly outperforms our strong wav2vec2.0 CTC baseline (15.1% relative CER reduction on AISHELL-1). The proposed NAR model significantly surpasses previous NAR systems on the AISHELL-1 benchmark and shows a potential for English tasks. Keqi Deng, Zehui Yang, Shinji Watanabe 0001, Yosuke Higuchi, Gaofeng Cheng, Pengyuan Zhang |
ICASSP | 6 |
| 2022 | Interrelate Training and Searching: A Unified Online Clustering Framework for Speaker DiarizationabstractFor online speaker diarization, samples arrive incrementally, and the overall distribution of the samples is invisible.Moreover, in most existing clustering-based methods, the training objective of the embedding extractor is not designed specially for clustering.To improve online speaker diarization performance, we propose a unified online clustering framework, which provides an interactive manner between embedding extractors and clustering algorithms.Specifically, the framework consists of two highly coupled parts: clustering-guided recurrent training (CGRT) and truncated beam searching clustering (TBSC).The CGRT introduces the clustering algorithm into the training process of embedding extractors, which could provide not only cluster-aware information for the embedding extractor, but also crucial parameters for the clustering process afterward.And with these parameters, which contain preliminary information of the metric space, the TBSC penalizes the probability score of each cluster, in order to output more accurate clustering results in online fashion with low latency.With the above innovations, our proposed online clustering system achieves 14.48% DER with collar 0.25 at 2.5s latency on the AISHELL-4, while the DER of the offline agglomerative hierarchical clustering is 14.57%. Qingxuan Li, Gaofeng Cheng, Pengyuan Zhang, Yonghong Yan 0002 |
INTERSPEECH | 5 |
| 2022 | Beam-Guided TasNet: An Iterative Speech Separation Framework with Multi-Channel OutputabstractTime-domain audio separation network (TasNet) has achieved remarkable performance in blind source separation (BSS).Classic multi-channel speech processing framework employs signal estimation and beamforming.For example, Beam-TasNet links multi-channel convolutional TasNet (MC-Conv-TasNet) with minimum variance distortionless response (MVDR) beamforming, which leverages the strong modeling ability of data-driven network and boosts the performance of beamforming with an accurate estimation of speech statistics.Such integration can be viewed as a directed acyclic graph by accepting multi-channel input and generating multi-source output.In this paper, we design a "multi-channel input, multi-channel multi-source output" (MIMMO) speech separation system entitled "Beam-Guided TasNet", where MC-Conv-TasNet and MVDR can interact and promote each other more compactly under a directed cyclic flow.Specifically, the first stage uses Beam-TasNet to generate estimated single-speaker signals, which favors the separation in the second stage.The proposed framework facilitates iterative signal refinement with the guide of beamforming and seeks to reach the upper bound of the MVDR-based methods.Experimental results on the spatialized WSJ0-2MIX demonstrate that the Beam-Guided TasNet has achieved an SDR of 21.5 dB, exceeding the baseline Beam-TasNet by 4.1 dB under the same model size and narrowing the gap with the oracle signal-based MVDR to 2 dB. Hangting Chen, Yi Yang 0057, Feng Dang, Pengyuan Zhang |
INTERSPEECH | 4 |
| 2022 | CTA-RNN: Channel and Temporal-wise Attention RNN leveraging Pre-trained ASR Embeddings for Speech Emotion RecognitionabstractPrevious research has looked into ways to improve speech emotion recognition (SER) by utilizing both acoustic and linguistic cues of speech.However, the potential association between state-of-the-art ASR models and the SER task has yet to be investigated.In this paper, we propose a novel channel and temporal-wise attention RNN (CTA-RNN) architecture based on the intermediate representations of pre-trained ASR models.Specifically, the embeddings of a large-scale pre-trained endto-end ASR encoder contain both acoustic and linguistic information, as well as the ability to generalize to different speakers, making them well suited for downstream SER task.To further exploit the embeddings from different layers of the ASR encoder, we propose a novel CTA-RNN architecture to capture the emotional salient parts of embeddings in both the channel and temporal directions.We evaluate our approach on two popular benchmark datasets, IEMOCAP and MSP-IMPROV, using both within-corpus and cross-corpus settings.Experimental results show that our proposed method can achieve excellent performance in terms of accuracy and robustness. Chengxin Chen, Pengyuan Zhang |
INTERSPEECH | 2 |
| 2022 | NAS-SCAE: Searching Compact Attention-based Encoders For End-to-end Automatic Speech Recognition
Ta Li, Pengyuan Zhang, Yonghong Yan 0002 |
INTERSPEECH | 3 |
| 2022 | Open Source MagicData-RAMC: A Rich Annotated Mandarin Conversational(RAMC) Speech DatasetabstractThis paper introduces a high-quality rich annotated Mandarin conversational (RAMC) speech dataset called MagicData-RAMC.The MagicData-RAMC corpus contains 180 hours of conversational speech data recorded from native speakers of Mandarin Chinese over mobile phones with a sampling rate of 16 kHz.The dialogs in MagicData-RAMC are classified into 15 diversified domains and tagged with topic labels, ranging from science and technology to ordinary life.Accurate transcription and precise speaker voice activity timestamps are manually labeled for each sample.Speakers' detailed information is also provided.As a Mandarin speech dataset designed for dialog scenarios with high quality and rich annotations, MagicData-RAMC enriches the data diversity in the Mandarin speech community and allows extensive research on a series of speechrelated tasks, including automatic speech recognition, speaker diarization, topic detection, keyword search, text-to-speech, etc.We also conduct several relevant tasks and provide experimental results to help evaluate the dataset. Zehui Yang, Runyan Yang, Lingxuan Ye, Gaofeng Cheng, Yaohui Jin, Pengyuan Zhang, Lei Xie 0001, Yonghong Yan 0002 |
INTERSPEECH | 10 |
| 2022 | Improving Recognition of Out-of-vocabulary Words in E2E Code-switching ASR by Fusing Speech Generation Methods
Lingxuan Ye, Gaofeng Cheng, Runyan Yang, Zehui Yang, Sanli Tian, Pengyuan Zhang, Yonghong Yan 0002 |
INTERSPEECH | 6 |
| 2022 | SASV Based on Pre-trained ASV System and Integrated Scoring ModuleabstractBased on the assumption that there is a correlation between anti-spoofing and speaker verification, a Total-Divide-Total integrated Spoofing-Aware Speaker Verification (SASV) system based on pre-trained automatic speaker verification (ASV) system and integrated scoring module is proposed and submitted to the SASV 2022 Challenge.The training and scoring of ASV and anti-spoofing countermeasure (CM) in current SASV systems are relatively independent, ignoring the correlation.In this paper, by leveraging the correlation between the two tasks, an integrated SASV system can be obtained by simply training a few more layers on the basis of the baseline pre-trained ASV subsystem.The features in pre-trained ASV system are utilized for logical access spoofing speech detection.Further, speaker embeddings extracted by the pre-trained ASV system are used to improve the performance of the CM.The integrated scoring module takes the embeddings of the ASV and anti-spoofing branches as input and preserves the correlation between the two tasks through matrix operations to produce integrated SASV scores.Submitted primary system achieved equal error rate (EER) of 3.07% on the development dataset of the SASV 2022 Challenge and 4.30% on the evaluation part, which is a 25% improvement over the baseline systems. Zhuo Li 0020, Pengyuan Zhang |
INTERSPEECH | 4 |
| 2022 | Robust Cough Feature Extraction and Classification Method for COVID-19 Cough Detection Based on Vocalization CharacteristicsabstractA fast, efficient and accurate detection method of COVID-19 remains a critical challenge.Many cough-based COVID-19 detection researches have shown competitive results through artificial intelligence.However, the lack of analysis on vocalization characteristics of cough sounds limits the further improvement of detection performance.In this paper, we propose two novel acoustic features of cough sounds and a convolutional neural network structure for COVID-19 detection.First, a time-frequency differential feature is proposed to characterize dynamic information of cough sounds in time and frequency domain.Then, an energy ratio feature is proposed to calculate the energy difference caused by the phonation characteristics in different cough phases.Finally, a convolutional neural network with two parallel branches which is pre-trained on a large amount of unlabeled cough data is proposed for classification.Experiment results show that our proposed method achieves state-of-the-art performance on Coswara dataset for COVID-19 detection.The results on an external clinical dataset Virufy also show the better generalization ability of our proposed method. Xueshuai Zhang, Jiakun Shen, Jun Zhou 0024, Pengyuan Zhang, Yonghong Yan 0002, Yanfen Tang, Fujie Zhang, Shaoxing Zhang, Aijun Sun |
INTERSPEECH | 4 |
| 2022 | Wav2vec-S: Semi-Supervised Pre-Training for Low-Resource ASRabstractSelf-supervised pre-training could effectively improve the performance of low-resource automatic speech recognition (ASR).However, existing self-supervised pre-training are taskagnostic, i.e., could be applied to various downstream tasks.Although it enlarges the scope of its application, the capacity of the pre-trained model is not fully utilized for the ASR task, and the learned representations may not be optimal for ASR.In this work, in order to build a better pre-trained model for low-resource ASR, we propose a pre-training approach called wav2vec-S, where we use task-specific semi-supervised pretraining to refine the self-supervised pre-trained model for the ASR task thus more effectively utilize the capacity of the pretrained model to generate task-specific representations for ASR.Experiments show that compared to wav2vec 2.0, wav2vec-S only requires a marginal increment of pre-training time but could significantly improve ASR performance on in-domain, cross-domain and cross-lingual datasets.Average relative WER reductions are 24.5% and 6.6% for 1h and 10h fine-tuning, respectively.Furthermore, we show that semi-supervised pretraining could close the representation gap between the selfsupervised pre-trained model and the corresponding fine-tuned model through canonical correlation analysis. Han Zhu 0004, Gaofeng Cheng, Jindong Wang 0001, Pengyuan Zhang, Yonghong Yan 0002 |
INTERSPEECH | 5 |
| 2022 | Decoupled Federated Learning for ASR with Non-IID DataabstractAutomatic speech recognition (ASR) with federated learning (FL) makes it possible to leverage data from multiple clients without compromising privacy. The quality of FL-based ASR could be measured by recognition performance, communication and computation costs. When data among different clients are not independently and identically distributed (non-IID), the performance could degrade significantly. In this work, we tackle the non-IID issue in FL-based ASR with personalized FL, which learns personalized models for each client. Concretely, we propose two types of personalized FL approaches for ASR. Firstly, we adapt the personalization layer based FL for ASR, which keeps some layers locally to learn personalization models. Secondly, to reduce the communication and computation costs, we propose decoupled federated learning (DecoupleFL). On one hand, DecoupleFL moves the computation burden to the server, thus decreasing the computation on clients. On the other hand, DecoupleFL communicates secure high-level features instead of model parameters, thus reducing communication cost when models are large. Experiments demonstrate two proposed personalized FL-based ASR approaches could reduce WER by 2.3% - 3.4% compared with FedAvg. Among them, DecoupleFL has only 11.4% communication and 75% computation cost compared with FedAvg, which is also significantly less than the personalization layer based FL. Han Zhu 0004, Jindong Wang 0001, Gaofeng Cheng, Pengyuan Zhang, Yonghong Yan 0002 |
INTERSPEECH | 4 |
| 2022 | DDAM '22: 1st International Workshop on Deepfake Detection for Audio MultimediaabstractOver the last few years, the technology of speech synthesis and voice conversion has made significant improvement with the development of deep learning. The models can generate realistic and human-like speech. It is difficult for most people to distinguish the generated audio from the real. However, this technology also poses a great threat to the global political economy and social stability if some attackers and criminals misuse it with the intent to cause harm. In this workshop, we aim to bring together researchers from the fields of audio deepfake detection, audio deep synthesis, audio fake game and adversarial attacks to further discuss recent research and future directions for detecting deepfake and manipulated audios in multimedia. Jianhua Tao 0001, Jiangyan Yi, Cunhang Fan, Ruibo Fu, Shan Liang 0007, Pengyuan Zhang, Haizhou Li 0001, Helen M. Meng, Dong Yu 0001, Masato Akagi |
ACM Multimedia | 6 |
| 2022 | An IBC Reference Block Enhancement Model Based on GAN for Screen Content Video Coding
Pengjian Yang, Jun Wang 0015, Guangyu Zhong, Pengyuan Zhang, Lai Zhang, Fan Liang 0001, Jianxin Yang |
MMM (2) | 4 |
| 2022 | An Adversarial Domain Adaptation Framework With KL-Constraint for Remote Sensing Land Cover ClassificationabstractLand cover classification plays a crucial role in land resource monitoring and planning. Recently, deep learning-based methods are becoming the dominating method for precise land cover mapping. However, the large-scale application of them is deeply hindered by the domain shift between different images, which is easily caused by illumination, climate, regional divergence, and so on. With the aim to cope with the problem of domain shift, many domain adaptation (DA) methods have been provided and great achievements have been made, especially the newborn adversarial DA, which usually contains a generator and a discriminator. Among these methods, the pixel-level methods are of high memory consumption, whereas feature-level methods are found hard to decode the structured information for semantic segmentation tasks due to the lack of low-dimensional information. Therefore, we propose an adversarial domain adaptation framework with Kullback–Leibler constraint (KL-ADDA) for remote sensing land cover classification. A state-of-the-art (SOTA) semantic segmentation network is utilized as the generator, which directly outputs the segmentation results to the discriminator to retain more low-level information. Besides, a Kullback–Leibler (KL)-divergence is calculated to improve the discriminative ability of the discriminator and thus enhance the generator’s performance. Experiments on the international society for photogrammetry and remote sensing (ISPRS) data set and two simulated target data sets have shown the effectiveness of KL-ADDA for DA. Mengxi Liu 0001, Pengyuan Zhang, Qian Shi 0001, Mengwei Liu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | An E2E-ASR-Based Iteratively-Trained Timestamp EstimatorabstractText-to-speech alignment, also known as time alignment, is essential for automatic speech recognition (ASR) systems used for speech retrieval tasks, such as keyword search and speech segment extraction. Previous works have used the Gaussian mixture model-hidden Markov model (GMM-HMM) forced alignment to improve the alignment performance. However, when used with end-to-end (E2E) ASR, GMM-HMM forced alignment causes extra reliance on expertise such as pronunciation lexica. It also increases the system complexity because GMM-HMMs are very dissimilar to E2E models. To tackle these two problems, we propose an E2E-ASR-based iteratively-trained timestamp estimator (ITSE), which performs alignment between token-level transcription and speech. We train ITSE first with coarse initial alignment targets generated using connectionist temporal classification (CTC) posteriors. During training, we iteratively perform realignment to update the targets. We attribute the effectiveness of the iterative training to ITSE’s two vital features. First, ITSE performs alignment using similarities between token and speech embeddings instead of frame-wise token classification posteriors. Second, ITSE uses speech embeddings that are aware of left context rather than global context. ITSE significantly outperforms CTC-based baselines in word alignment accuracy and is comparable to a GMM-HMM forced aligner. In short, ITSE is an accurate, lightweight text-to-speech alignment module implemented without expertise such as pronunciation lexica. Runyan Yang, Gaofeng Cheng, Pengyuan Zhang, Yonghong Yan 0002 |
IEEE Signal Process. Lett. | 3 |
| 2022 | Self-Supervised Pre-Training for Attention-Based Encoder-Decoder ASR ModelabstractEnd-to-end (E2E) models, including the attention-based encoder-decoder (AED) models, have achieved promising performance on the automatic speech recognition (ASR) task. However, the supervised training process of the E2E model needs a large amount of speech-text paired data. In contrast, self-supervised pre-training can pre-train the model on the unlabeled data and then fine-tune it on the limited labeled data to realize better performance. Most of the previous self-supervised pre-training methods focus on learning hidden representations from speech but ignore how to utilize the unpaired text. As a result, previous works often pre-train an acoustic encoder and then fine-tune it as a classification based ASR model, such as Connectionist Temporal Classification (CTC) based model, rather than an AED model. In this paper, we propose a self-supervised pre-training method for the AED model (SP-AED). The SP-AED method contains acoustic pre-training for the encoder, linguistic pre-training for the decoder, and an adaptive combination fine-tuning for the whole system. We first design a linguistic pre-training method for decoder by utilizing the text-only data. The decoder will be pre-trained as a noise-condition language model to learn the prior distribution of the text. Then, we pre-train the AED encoder with the wav2vec2.0 method with some modifications. Finally, we combine the pre-trained encoder and decoder and fine-tune them on the limited labeled data. We design an adaptive combination method during fine-tuning by modifying the decoder’s input and output to prevent catastrophic forgetting. Experiments prove that compared with the random initialized models, the SP-AED pre-trained models can realize up to 17% relative improvement. And with similar model size or computational cost, we can get comparable results to other classification-based models on both English and Chinese corpus. Changfeng Gao, Gaofeng Cheng, Ta Li, Pengyuan Zhang, Yonghong Yan 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Far-Field Speech Recognition Based on Complex-Valued Neural Networks and Inter-Frame Similarity Difference MethodabstractFar-field automatic speech recognition (ASR) is a challenging task due to the background noise and reverberation. To address this issue, we introduce a novel end-to-end multi-channel far-field ASR architecture. First, we use a complex-valued CNN based architecture designed for speech tasks as a neural beamformer. Second, we propose an auxiliary mod-ule called absolute position regression module (APRM) with a position prediction loss to help the neural beamformer be better aware of the corresponding frequencies of each input time-frequency (T-F) bin. Third, inspired by the short-term stationarity of human speech, we propose an approach called the Inter-Frame Similarity Difference (IFSD) method to au-tomatically select useful channels as the inputs of the ASR backend from the outputs of the neural beamformer. We also implement a complex-valued attention module for the output channels of the neural beamformer to utilize each other's in-formation, thereby preventing the final outputs from information loss. With the above innovations, our proposed model achieves 9.7% and 11.1% relative WER reductions over a DNN-MVDR baseline on the CHiME4 dataset and a dataset simulated using the Librispeech corpus. Gaofeng Cheng, Pengyuan Zhang, Yonghong Yan 0002 |
ASRU | 4 |
| 2021 | History Utterance Embedding Transformer LM for Speech RecognitionabstractHistory utterances contain rich contextual information; however, better extracting information from the history utterances and using it to improve the language model (LM) is still challenging. In this paper, we propose the history utterance embedding Transformer LM (HTLM), which includes an embedding generation network for extracting contextual information contained in the history utterances and a main Transformer LM for current prediction. In addition, the two-stage attention (TSA) is proposed to encode richer contextual information into the embedding of history utterances (h-emb) while supporting GPU parallel training. Furthermore, we combine the extracted h-emb and embedding of current utterance (c-emb) through the dot-product attention and a fusion method for HTLM's current prediction. Experiments are conducted on the HKUST dataset and achieve a 23.4% character error rate (CER) on the test set. Compared with the baseline, the proposed method yields 12.86 absolute perplexity reduction and 0.8% absolute CER reduction. Keqi Deng, Gaofeng Cheng, Haoran Miao, Pengyuan Zhang, Yonghong Yan 0002 |
ICASSP | 4 |
| 2021 | Pre-Training Transformer Decoder for End-to-End ASR Model with Unpaired Text DataabstractThis paper presents a method to pre-train transformer-based encoder-decoder automatic speech recognition (ASR) models using sufficient target-domain text. During pre-training, we train the transformer decoder as a conditional language model with empty or artifical states, rather than the real encoder states. By this pre-training strategy, the decoder can learn how to generate grammatical text sequence before learning how to generate correct transcriptions. Contrast to other methods which utilize text only data to improve the ASR performance, our method does not change the network architecture of the ASR model or introduce extra component like text-to-speech (TTS) or text-to-encoder (TTE). Experimental results on LibriSpeech corpus show that the proposed method can relatively reduce the word error rate over 10%, using 960 hours transcriptions. Changfeng Gao, Gaofeng Cheng, Runyan Yang, Han Zhu 0004, Pengyuan Zhang, Yonghong Yan 0002 |
ICASSP | 5 |
| 2021 | RNN-T Based Open-Vocabulary Keyword Spotting in Mandarin with Multi-Level DetectionabstractDespite the recent prevalence of keyword spotting (KWS) in smart-home, open-vocabulary KWS remains a keen but unmet need among the users. In this paper, we propose an RNN Transducer (RNN-T) based keyword spotting system with a constrained attention mechanism biasing module that biases the RNN-T model towards a specific keyword of interest. The atonal syllables are adopted as the modeling units, which addresses the out-of-vocabulary (OOV) problem. A multi-level detection is applied to the posterior probabilities for the judgement. Evaluating on the AISHELL-2 dataset shows our proposed method outperforms the RNN-T-based approach by 2.70% in false reject rate (FRR) at 1 false alarm (FA) per hour. We further provide insights into the role of each stage of the detection cascade, where most negative samples are filtered out by the first stage with high computational efficiency. Zuozhen Liu, Ta Li, Pengyuan Zhang |
ICASSP | 3 |
| 2021 | The Thinkit System for Icassp2021 M2voc ChallengeabstractIn this paper, we introduce the low resource text-to-speech system from the ThinkIT team submitted to Multi-Speaker Multi-Style Voice Cloning Challenge (M2VoC). The challenge has two tasks: few-shot track1 provides 100 samples for each person and one-shot track2 offers 5 samples only. Each track contains two sub-tracks A and B. Instead of sub-track A, sub-track B can use extra public data besides the released data. But we participate in the sub-track A only. We choose the finetune as our backbone strategy. Our submitted systems include BERT based prosody boundary prediction module, FastSpeech based acoustic model to generate acoustic features from text input, and HIFIGAN based vocoder to generate waveform from acoustic features. Among them, acoustic models are susceptible to low resource speakers. To prevent over-fitting, we modified the acoustic model and split out validation set to assist the manual model selection. Evaluation results provided by the challenges organizers demonstrate the effectiveness of our system. Zengqiang Shang, Bolin Zhou, Pengyuan Zhang |
ICASSP | 5 |
| 2021 | Power Pooling: An Adaptive Pooling Function for Weakly Labelled Sound Event DetectionabstractAccess to large corpora with strongly labelled sound events is expensive and difficult in engineering applications. Many researches turn to address the problem of how to detect both the types and the timestamps of sound events with weak labels that only specify the types. This task can be treated as a multiple instance learning (MIL) problem, and a key to it in the sound event detection (SED) task is the design of a pooling function. The linear softmax pooling function achieves state-of-the-art performance since it can vary both the signs and the magnitudes of gradients. However, linear softmax pooling cannot flexibly deal with sound events of different time scales. In this paper, we propose a power pooling function which can automatically adapt to various sound events. By adding a trainable parameter to each event, power pooling can provide more accurate gradients for frames in a clip than other pooling functions. On both weakly supervised and semi-supervised SED datasets, the proposed power pooling function outperforms linear softmax pooling on both coarse-grained and fine-grained metrics. Specifically, it improves the event-based F1 score by 11.4% and 10.2% relatively on the two datasets. While this paper focuses on SED applications, the proposed method can be applied to MIL tasks in other domains. Yuzhuo Liu, Hangting Chen, Pengyuan Zhang |
IJCNN | 4 |
| 2021 | TVQVC: Transformer Based Vector Quantized Variational Autoencoder with CTC Loss for Voice Conversion
Pengyuan Zhang |
Interspeech | 2 |
| 2021 | Improved Speech Enhancement Using a Complex-Domain GAN with Fused Time-Domain and Time-Frequency Domain Constraints
Feng Dang, Pengyuan Zhang, Hangting Chen |
Interspeech | 2 |
| 2021 | Incorporating Cross-Speaker Style Transfer for Multi-Language Text-to-Speech
Zengqiang Shang, Pengyuan Zhang, Yonghong Yan 0002 |
Interspeech | 4 |
| 2021 | Adaptive Margin Circle Loss for Speaker VerificationabstractDeep-Neural-Network (DNN) based speaker verification systems use the angular softmax loss with margin penalties to enhance the intra-class compactness of speaker embeddings, which achieved remarkable performance.In this paper, we propose a novel angular loss function called adaptive margin circle loss for speaker verification.The stage-based margin and chunk-based margin are applied to improve the angular discrimination of circle loss on the training set.The analysis on gradients shows that, compared with the previous angular loss like Additive Margin Softmax(Am-Softmax), circle loss has flexible optimization and definite convergence status.Experiments are carried out on the Voxceleb and SITW.By applying adaptive margin circle loss, our best system achieves 1.31%EER on Voxceleb1 and 2.13% on SITW core-core. Runqiu Xiao, Xiaoxiao Miao, Pengyuan Zhang, Liuping Luo |
Interspeech | 4 |
| 2021 | LinearSpeech: Parallel Text-to-Speech with Linear Complexity
Zengqiang Shang, Pengyuan Zhang, Yonghong Yan 0002 |
Interspeech | 4 |
| 2021 | The Effect of Silence and Dual-Band Fusion in Anti-Spoofing System
Pengyuan Zhang |
Interspeech | 3 |
| 2021 | A dual-stream deep attractor network with multi-domain learning for speech dereverberation and separation
Hangting Chen, Pengyuan Zhang |
Neural Networks | 2 |
| 2021 | D-MONA: A dilated mixed-order non-local attention network for speaker and language recognition
Xiaoxiao Miao, Ian McLoughlin 0001, Pengyuan Zhang |
Neural Networks | 4 |
| 2021 | A unified system for multilingual speech recognition and language identification
Pengyuan Zhang, Yonghong Yan 0002 |
Speech Commun. | 3 |
| 2021 | Keyword Search Using Attention-Based End-to-End ASR and Frame-Synchronous Phoneme AlignmentsabstractAttention-based end-to-end (E2E) automatic speech recognition (ASR) architectures are now the state-of-the-art in terms of recognition performance. However, despite their effectiveness, they have not been widely applied in keyword search (KWS) tasks yet. In this paper, we propose the Att-E2E-KWS architecture, an attention-based E2E ASR framework for KWS that can afford accurate and reliable keyword retrieval results. First, we design a basic framework to carry out KWS based on attention-based E2E ASR. We adopt the connectionist temporal classification and attention (CTC/Att) joint E2E ASR architecture and exploit the spike posterior property of CTC to provide the keywords time stamps. Second, we introduce the frame-synchronous phonemes modeling and use the dynamic programming (DP) algorithm to provide alignments between E2E grapheme outputs and phoneme outputs. We call this alignment procedure dynamic time alignment (DTA), which can provide the proposed Att-E2E-KWS system with more accurate time stamps and reliable confidence scores. Third, we use the Transformer, a self-attention-based encoder-decoder neural network, in place of conventional recurrent neural networks in order to yield more parallelizable models and increased training speed. We conduct comprehensive experiments on English and Mandarin Chinese. To the best of our knowledge, this is the first practical Att-E2E-KWS framework, and experimental results on Switchboard and HKUST corpora show that our proposed Att-E2E-KWS systems significantly outperform the CTC E2E ASR based KWS baselines. Runyan Yang, Gaofeng Cheng, Haoran Miao, Ta Li, Pengyuan Zhang, Yonghong Yan 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2020 | CN-Celeb: A Challenging Chinese Speaker Recognition DatasetabstractRecently, researchers set an ambitious goal of conducting speaker recognition in unconstrained conditions where the variations on ambient, channel and emotion could be arbitrary. However, most publicly available datasets are collected under constrained environments, i.e., with little noise and limited channel variation. These datasets tend to deliver over-optimistic performance and do not meet the request of research on speaker recognition in unconstrained conditions.In this paper, we present CN-Celeb, a large-scale speaker recognition dataset collected ‘in the wild’. This dataset contains more than 130,000 utterances from 1,000 Chinese celebrities, and covers 11 different genres in real world. Experiments conducted with two state-of-the-art speaker recognition approaches (i-vector and x-vector) show that the performance on CN-Celeb is far inferior to the one obtained on Vox-Celeb, a widely used speaker recognition dataset. This result demonstrates that in real-life conditions, the performance of existing techniques might be much worse than it was thought. Our database is free for researchers and can be downloaded from http://project.cslt.org. Jiawen Kang 0002, Lantian Li, Kaicheng Li, Sitong Cheng, Pengyuan Zhang, Ziya Zhou, Yunqi Cai, Dong Wang 0013 |
ICASSP | 7 |
| 2020 | Transformer-Based Online CTC/Attention End-To-End Speech Recognition ArchitectureabstractRecently, Transformer has gained success in automatic speech recognition (ASR) field. However, it is challenging to deploy a Transformer-based end-to-end (E2E) model for online speech recognition. In this paper, we propose the Transformer-based online CTC/attention E2E ASR architecture, which contains the chunk self-attention encoder (chunk-SAE) and the monotonic truncated attention (MTA) based self-attention decoder (SAD). Firstly, the chunk-SAE splits the speech into isolated chunks. To reduce the computational cost and improve the performance, we propose the state reuse chunk-SAE. Sencondly, the MTA based SAD truncates the speech features monotonically and performs attention on the truncated features. To support the online recognition, we integrate the state reuse chunk-SAE and the MTA based SAD into online CTC/attention architecture. We evaluate the proposed online models on the HKUST Mandarin ASR benchmark and achieve a 23.66% character error rate (CER) with a 320 ms latency. Our online model yields as little as 0.19% absolute CER degradation compared with the offline baseline, and achieves significant improvement over our prior work on Long Short-Term Memory (LSTM) based online E2E models. Haoran Miao, Gaofeng Cheng, Changfeng Gao, Pengyuan Zhang, Yonghong Yan 0002 |
ICASSP | 4 |
| 2020 | Improved Guided Source Separation Integrated with a Strong Back-End for the CHiME-6 Dinner Party Scenario
Hangting Chen, Pengyuan Zhang, Qian Shi 0001, Zuozhen Liu |
INTERSPEECH | 2 |
| 2020 | Speaker Diarization System Based on DPCA Algorithm for Fearless Steps Challenge Phase-2
Xueshuai Zhang, Pengyuan Zhang |
INTERSPEECH | 3 |
| 2020 | Domain Adaptation Using Class Similarity for Robust Speech RecognitionabstractWhen only limited target domain data is available, domain adaptation could be used to promote performance of deep neural network (DNN) acoustic model by leveraging well-trained source model and target domain data. However, suffering from domain mismatch and data sparsity, domain adaptation is very challenging. This paper proposes a novel adaptation method for DNN acoustic model using class similarity. Since the output distribution of DNN model contains the knowledge of similarity among classes, which is applicable to both source and target domain, it could be transferred from source to target model for the performance improvement. In our approach, we first compute the frame level posterior probabilities of source samples using source model. Then, for each class, probabilities of this class are used to compute a mean vector, which we refer to as mean soft labels. During adaptation, these mean soft labels are used in a regularization term to train the target model. Experiments showed that our approach outperforms fine-tuning using one-hot labels on both accent and noise adaptation task, especially when source and target domain are highly mismatched. Han Zhu 0004, Jiangjiang Zhao, Yuling Ren, Pengyuan Zhang |
INTERSPEECH | 5 |
| 2020 | Domain Adaption for Fine-Grained Urban Village Extraction From Satellite ImagesabstractUrban villages (UVs) are distinctive products formed in the process of rapid urbanization. The fine-grained mapping of UVs from satellite images has always been a considerable challenge because of the complex urban structures and the insufficiency of labeled samples. In this letter, we propose using the domain adaptation strategy to tackle the domain shift problem by employing adversarial learning to tune the semantic segmentation network so as to adaptively obtain similar outputs for input images from different domains. The proposed method was coupled with several segmentation networks, including U-Net, RefineNet, and DeepLab v3+, and the results show that domain adaptation can significantly improve the pixel-level mapping of UVs. Qian Shi 0001, Mengxi Liu 0001, Xiaoping Liu 0001, Penghua Liu, Pengyuan Zhang, Jinxing Yang, Xia Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2020 | Online Hybrid CTC/Attention End-to-End Automatic Speech Recognition ArchitectureabstractRecently, there has been increasing progress in end-to-end automatic speech recognition (ASR) architecture, which transcribes speech to text without any pre-trained alignments. One popular end-to-end approach is the hybrid Connectionist Temporal Classification (CTC) and attention (CTC/attention) based ASR architecture, which utilizes the advantages of both CTC and attention. The hybrid CTC/attention ASR systems exhibit performance comparable to that of the conventional deep neural network (DNN)/ hidden Markov model (HMM) ASR systems. However, how to deploy hybrid CTC/attention systems for online speech recognition is still a non-trivial problem. This article describes our proposed online hybrid CTC/attention end-to-end ASR architecture, which replaces all the offline components of conventional CTC/attention ASR architecture with their corresponding streaming components. Firstly, we propose stable monotonic chunk-wise attention (sMoChA) to stream the conventional global attention, and further propose monotonic truncated attention (MTA) to simplify sMoChA and solve the training-and-decoding mismatch problem of sMoChA. Secondly, we propose truncated CTC (T-CTC) prefix score to stream CTC prefix score calculation. Thirdly, we design dynamic waiting joint decoding (DWJD) algorithm to dynamically collect the predictions of CTC and attention in an online manner. Finally, we use latency-controlled bidirectional long short-term memory (LC-BLSTM) to stream the widely-used offline bidirectional encoder network. Experiments with LibriSpeech English and HKUST Mandarin tasks demonstrate that, compared with the offline CTC/attention model, our proposed online CTC/attention model improves the real time factor in human-computer interaction services and maintains its performance with moderate degradation. To the best of our knowledge, this is the first work to provide the full-stack online solution for CTC/attention end-to-end ASR architecture. Haoran Miao, Gaofeng Cheng, Pengyuan Zhang, Yonghong Yan 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | An Audio Scene Classification Framework with Embedded Filters and a DCT-based Temporal ModuleabstractDeep convolutional neural network (DCNN) has recently improved the performance of acoustic scene classification. However, the input features of the network are usually based on predefined hand-tailored filters, which may not apply to the specific tasks. To overcome this, we propose a hybrid framework that jointly trains the front-end filters and the back-end DCNN. Also, a novel temporal module based on the discrete cosine transform (DCT) is inserted after the high-level feature map of the network, thus enabling us to utilize time information without a reduction of training samples. Our single system, composed of the fine-tuned wavelet front-end and the DCNN back-end, with the integrated DCT-based temporal module, has achieved an accuracy of 79.20% in the evaluation set in DCASE17, gaining around 3% and 8% accuracy improvement compared with scalogram-DCNN and FBank-DCNN systems, respectively. Hangting Chen, Pengyuan Zhang, Yonghong Yan 0002 |
ICASSP | 2 |
| 2019 | Self-attention Based Prosodic Boundary Prediction for Chinese Speech SynthesisabstractPredicting prosodic boundaries from input text plays an important role in Chinese text-to-speech (TTS) system, which directly influences the naturalness and intelligibility of synthesized speech. In this paper, we propose to combine self-attention with multitask learning for prosodic boundary prediction. Self-attention is used to capture the dependency between two arbitrary characters in the input sentence, while multitask learning models the relationships between prosodic boundaries and lexicon words by setting word segmentation as an auxiliary task. The proposed method can generate prosodic boundary labels directly from Chinese characters and achieve the whole process end-to-end. Experimental results show the effectiveness of our proposed model and prove that the performance can be further improved by pretraining the model with extra word segmentation data. Chunhui Lu, Pengyuan Zhang, Yonghong Yan 0002 |
ICASSP | 2 |
| 2019 | Target Speaker Recovery and Recognition Network with Average x-Vector and Global Training
Pengyuan Zhang, Yonghong Yan 0002 |
INTERSPEECH | 2 |
| 2019 | Character-Aware Sub-Word Level Language Modeling for Uyghur and Turkish ASR
Pengyuan Zhang, Yonghong Yan 0002 |
INTERSPEECH | 3 |
| 2019 | Online Hybrid CTC/Attention Architecture for End-to-End Speech Recognition
Haoran Miao, Gaofeng Cheng, Pengyuan Zhang, Ta Li, Yonghong Yan 0002 |
INTERSPEECH | 3 |
| 2019 | Speaker-Invariant Feature-Mapping for Distant Speech Recognition via Adversarial Teacher-Student Learning
Long Wu, Hangting Chen, Pengyuan Zhang, Yonghong Yan 0002 |
INTERSPEECH | 4 |
| 2019 | Multi-Accent Adaptation Based on Gate MechanismabstractWhen only a limited amount of accented speech data is available, to promote multi-accent speech recognition performance, the conventional approach is accent-specific adaptation, which adapts the baseline model to multiple target accents independently. To simplify the adaptation procedure, we explore adapting the baseline model to multiple target accents simultaneously with multi-accent mixed data. Thus, we propose using accent-specific top layer with gate mechanism (AST-G) to realize multi-accent adaptation. Compared with the baseline model and accent-specific adaptation, AST-G achieves 9.8% and 1.9% average relative WER reduction respectively. However, in real-world applications, we can't obtain the accent category label for inference in advance. Therefore, we apply using an accent classifier to predict the accent label. To jointly train the acoustic model and the accent classifier, we propose the multi-task learning with gate mechanism (MTL-G). As the accent label prediction could be inaccurate, it performs worse than the accent-specific adaptation. Yet, in comparison with the baseline model, MTL-G achieves 5.1% average relative WER reduction. Han Zhu 0004, Pengyuan Zhang, Yonghong Yan 0002 |
INTERSPEECH | 3 |
| 2019 | Weighted Feature Fusion Based Emotional Recognition for Variable-length Speech using DNNabstractEmotion recognition plays an increasingly important role in human-computer interaction systems, which is a key technology in multimedia communication. Because neural networks can automatically learn the intermediate representation of raw speech signal, currently, most methods use Convolutional Neural Network (CNN) to extract information directly from spectrograms, but this may result in the ineffective use of information in hand-crafted features. In this work, a model based on weighted feature fusion method is proposed for emotion recognition of variable-length speech. Since the Chroma-based features are closely related to speech emotions, our model can effectively utilize the useful information in Chromaticity map to improve the performance by combining CNN-based features and Chroma-based features. We evaluated the model on the Interactive Emotional Motion Capture (IEMOCAP) dataset and achieved more than 5% increase in weighted accuracy (WA) and unweighted accuracy (UA), comparing with the existing state-of-the-art methods. Fei Li 0014, Pengyuan Zhang |
IWCMC | 3 |
| 2019 | Tailoring an Interpretable Neural Language ModelabstractNeural networks have shown great potential in language modeling. Currently, the dominant approach to language modeling is based on recurrent neural networks (RNNs) and convolutional neural networks (CNNs). Nonetheless, it is not clear why RNNs and CNNs are suitable for the language modeling task since these neural models are lack of interpretability. The goal of this paper is to tailor an interpretable neural model as an alternative to RNNs and CNNs for the language modeling task. This paper proposes a unified framework for language modeling, which can partly interpret the rationales behind existing language models (LMs). Based on the proposed framework, an interpretable neural language model (INLM) is proposed, including a tailored architectural structure and a tailored learning method for the language modeling task. The proposed INLM can be approximated as a parameterized auto-regressive moving average model and provides interpretability in two aspects: component interpretability and prediction interpretability. Experiments demonstrate that the proposed INLM outperforms some typical neural LMs on several language modeling datasets and on the switchboard speech recognition task. Further experiments also show that the proposed INLM is competitive with the state-of-the-art long short-term memory LMs on the Penn Treebank and WikiText-2 datasets. Pengyuan Zhang, Yonghong Yan 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Long/Short-Term Utility Aware Optimal Selection of Manufacturing Service Composition Toward Industrial Internet PlatformsabstractAs numerous Industrial Internet platforms emerge, manufacturing services are shared among multiple stakeholders more frequently than ever before. The optimal selection of shared manufacturing service composition (MSC) should promise both the task completion and the stakeholders' satisfaction. However, as commercial entities, stakeholders concentrate on not only the temporary benefits but also the long-term acquisitions. Most of the existing MSC problems neglect the stakeholders' prospect on the manufacturing service sharing. This leads to the disappointment and dissatisfaction of the stakeholders with long-term expectations, who will abandon the participation in Industrial Internet platforms. Therefore, the long/short-term preferences of various stakeholders should be satisfied and balanced. In this paper, the long/short-term utilities of three parties (provider, consumer, and operator) are first defined and discussed, and the models considering short-term utility of a consumer and long-term utility of providers are established. The potential tasks assigned to providers are taken into account to estimate the long-term utility if the current task is accepted. Then, to solve the biobjective optimization problem, an improved Nondominated Sorting Genetic Algorithm-II algorithm, combining Tabu search and improved K-means mechanism, is proposed to find the optimal solution set. Finally, the effectiveness of the method is verified by the experimental results in terms of solution diversity, astringency, and stability, in which a finding is further observed that the changes of consumers' preferences have little impact on the long-term utility of providers. Fei Tao 0001, Yang Liu 0034, Pengyuan Zhang, Ying Cheng 0001, Ying Zuo |
IEEE Trans. Ind. Informatics | 4 |
| 2018 | Improving Multichannel Speech Recognition with Generalized Cross Correlation Inputs and Multitask LearningabstractAcoustic signals from microphone arrays are used to improve performance in distant speech recognition due to the availability of spatial information. And multichannel automatic speech recognition (ASR) systems often separate speech enhancement module from acoustic modeling, which may be not optimal for improving recognition accuracy. In this work, we propose to improve multichannel speech recognition by supplying the generalized cross correlation (GCC) between microphones, which encodes spatial information, as input features to a long short-term memory (LSTM) acoustic model in parallel with the regular acoustic features. Moreover, multitask learning architecture is incorporated and shows its ability to improve the robustness of the model. We performed experiments on the AMI and ICSI meeting corpora, with results indicating that the proposed model outperforms the model trained directly on the concatenation of multiple microphone outputs and the model trained on a beamformed channel. Pengyuan Zhang, Yonghong Yan 0002 |
ICASSP | 3 |
| 2018 | Deep Convolutional Neural Network with Scalogram for Audio Scene Modeling
Hangting Chen, Pengyuan Zhang, Haichuan Bai, Qingsheng Yuan, Xiuguo Bao, Yonghong Yan 0002 |
INTERSPEECH | 2 |
| 2018 | Investigation on the Combination of Batch Normalization and Dropout in BLSTM-based Acoustic Modeling for ASR
Gaofeng Cheng, Fengpei Ge, Pengyuan Zhang, Yonghong Yan 0002 |
INTERSPEECH | 4 |
| 2018 | Improving Language Modeling with an Adversarial Critic for Automatic Speech Recognition
Pengyuan Zhang, Yonghong Yan 0002 |
INTERSPEECH | 2 |
| 2018 | Multichannel ASR with Knowledge Distillation and Generalized Cross Correlation FeatureabstractMulti-channel signal processing techniques have played an important role in the far-field automatic speech recognition (ASR) as the separate front-end enhancement part. However, they often meet the mismatch problem. In this paper, we proposed a novel architecture of acoustic model, in which the multi-channel speech without preprocessing was utilized directly. Besides the strategy of knowledge distillation and the generalized cross correlation (GCC) adaptation were employed. We use knowledge distillation to transfer knowledge from a well-trained close-talking model to distant-talking scenarios in every frame of the multichannel distant speech. Moreover, the GCC between microphones, which contains the spatial information, is supplied as an auxiliary input to the neural network. We observe good compensation of those two techniques. Evaluated with the AMI and ICSI meeting corpora, the proposed methods achieve relative WER improvement of 7.7% and 7.5% over the model trained directly on the concatenated multi-channel speech. Pengyuan Zhang, Fengpei Ge |
SLT | 3 |
| 2017 | Attention-Based LSTM with Multi-Task Learning for Distant Speech Recognition
Pengyuan Zhang, Yonghong Yan 0002 |
INTERSPEECH | 2 |
| 2016 | An unsupervised vocabulary selection technique for Chinese automatic speech recognitionabstractThe vocabulary is a vital component of automatic speech recognition(ASR) systems. For a specific Chinese speech recognition task, using a large general vocabulary not only leads to a much longer time to decode, but also hurts the recognition accuracy. In this paper, we proposed an unsupervised algorithm to select task-specific words from a large general vocabulary. The out-of-vocabulary(OOV) rate is a measure of vocabularies, and it is related to the recognition accuracy. However, it is hard to compute OOV rate for a Chinese vocabulary, since OOVs are often segmented into single Chinese characters and most Chinese vocabularies contain all the single Chinese characters. To deal with this problem, we proposed a novel method to estimate the OOV rate of Chinese vocabularies. In experiments, we found that our estimated OOV rate is related to the character error rate(CER) of recognition. Our proposed vocabulary selection method provided both the lowest OOV rate and CER on two Chinese conversational telephone speech(CTS) evaluation sets compared to the general vocabulary and frequency based vocabulary selection method. In addition, our proposed method significantly reduced the size of the language model(LM) and the corresponding weighted finite state transducer(WFST) network, which led to a more efficient decoding. Pengyuan Zhang, Ta Li, Yonghong Yan 0002 |
SLT | 2 |
| 2014 | Using neural network front-ends on far field multiple microphones based speech recognitionabstractThis paper presents an investigation of far field speech recognition using beamforming and channel concatenation in the context of Deep Neural Network (DNN) based feature extraction. While speech enhancement with beamforming is attractive, the algorithms are typically signal-based with no information about the special properties of speech. A simple alternative to beamforming is concatenating multiple channel features. Results presented in this paper indicate that channel concatenation gives similar or better results. On average the DNN front-end yields a 25% relative reduction in Word Error Rate (WER). Further experiments aim at including relevant information in training adapted DNN features. Augmenting the standard DNN input with the bottleneck feature from a Speaker Aware Deep Neural Network (SADNN) shows a general advantage over the standard DNN based recognition system, and yields additional improvements for far field speech recognition. Yulan Liu, Pengyuan Zhang, Thomas Hain |
ICASSP | 2 |
| 2014 | Semi-supervised DNN training in meeting recognitionabstractTraining acoustic models for ASR requires large amounts of labelled data which is costly to obtain. Hence it is desirable to make use of unlabelled data. While unsupervised training can give gains for standard HMM training, it is more difficult to make use of unlabelled data for discriminative models. This paper explores semi-supervised training of Deep Neural Networks (DNN) in a meeting recognition task. We first analyse the impact of imperfect transcription on the DNN and the ASR performance. As labelling error is the source of the problem, we investigate two options available to reduce that: selecting data with fewer errors, and changing the dependence on noise by reducing label precision. Both confidence based data selection and label resolution change are explored in the context of two scenarios of matched and unmatched unlabelled data. We introduce improved DNN based confidence score estimators and show their performance on data selection for both scenarios. Confidence score based data selection was found to yield up to 14.6% relative WER reduction, while better balance between label resolution and recognition hypothesis accuracy allowed further WER reductions by 16.6% relative in the mismatched scenario. Pengyuan Zhang, Yulan Liu, Thomas Hain |
SLT | 1 |
| 2007 | A fast fuzzy keyword spotting algorithm based on syllable confusion network
Jian Shao 0001, Qingwei Zhao, Pengyuan Zhang, Zhaojie Liu, Yonghong Yan 0002 |
INTERSPEECH | 3 |