Zhengqi Wen

dblp:39/10650 · DBLP profile ↗
← Back
86ranked-venue papers
4as first author
48since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 65 · 4 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 65 · 4 first-author · 30 since 2021
YearPublicationVenuePosition
2026 AStar: Boosting Multimodal Reasoning with Automated Structured Thinking
abstract
Multimodal large language models excel across diverse domains but struggle with complex visual reasoning tasks. To enhance their reasoning capabilities, current approaches typically rely on explicit search or post-training techniques. However, search-based methods suffer from computational inefficiency due to extensive solution space exploration, while post-training methods demand substantial data, computational resources, and often exhibit training instability. To address these challenges, we propose **AStar**, a training-free, **A**utomatic **S**tructured **t**hinking paradigm for multimod**a**l **r**easoning. Specifically, we introduce novel "thought cards", a lightweight library of high-level reasoning patterns abstracted from prior samples. For each test problem, AStar adaptively retrieves the optimal thought cards and seamlessly integrates these external explicit guidelines with the model’s internal implicit reasoning capabilities. Compared to previous methods, AStar eliminates computationally expensive explicit search and avoids additional complex post-training processes, enabling a more efficient reasoning approach. Extensive experiments demonstrate that our framework achieves 53.9% accuracy on MathVerse (surpassing GPT-4o's 50.2%) and 32.7% on MathVision (outperforming GPT-4o's 30.4%). Further analysis reveals the remarkable transferability of our method: thought cards generated from mathematical reasoning can also be applied to other reasoning tasks, even benefiting general visual perception and understanding. AStar serves as a plug-and-play test-time inference method, compatible with other post-training techniques, providing an important complement to existing multimodal reasoning approaches.
Mingkuan Feng, Guocheng Zhai, Shuai Zhang 0014, Zheng Lian 0004, Fangrui Lv, Pengpeng Shao, Ruihan Jin, Zhengqi Wen, Jianhua Tao 0001
AAAI9
2026 PSA-MF: Personality-Sentiment Aligned Multi-Level Fusion for Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis (MSA) is a research field that recognizes human sentiments by combining textual, visual, and audio modalities. The main challenge lies in integrating sentiment-related information from different modalities, which typically arises during the unimodal feature extraction phase and the multimodal feature fusion phase. Existing methods extract only shallow information from unimodal features during the extraction phase, neglecting sentimental differences across different personalities. During the fusion phase, they directly merge the feature information from each modality without considering differences at the feature level. This ultimately affects the model's recognition performance. To address this problem, we propose a personality-sentiment aligned multi-level fusion framework. We introduce personality traits during the feature extraction phase and propose a novel personality-sentiment alignment method to obtain personalized sentiment embeddings from the textual modality for the first time. In the fusion phase, we introduce a novel multi-level fusion method. This method gradually integrates sentimental information from textual, visual, and audio modalities through multimodal pre-fusion and a multi-level enhanced fusion strategy. Our method has been evaluated through multiple experiments on two commonly used datasets, achieving state-of-the-art results.
Kang Zhu, Zhengqi Wen, Jianhua Tao 0001, Xuefei Liu, Ruibo Fu
AAAI3
2026 ReFL: Reflective Feedback Learning for Hallucination Detection of Large Language Models
abstract
Cunhang Fan, Jun Zhang, Xue Zhang, Shuai Zhang, Zhao Lv, Jianhua Tao, Zhengqi Wen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Cunhang Fan, Shuai Zhang 0014, Zhao Lv, Jianhua Tao 0001, Zhengqi Wen
ACL (1)7
2026 Two-Stage Regularization-Based Structured Pruning for LLMs
abstract
Mingkuan Feng, Jinyang Wu, Siyuan Liu, Shuai Zhang, Hongjian Fang, Ruihan Jin, Feihu Che, Pengpeng Shao, Zhengqi Wen, Jianhua Tao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Mingkuan Feng, Shuai Zhang 0014, Hongjian Fang, Ruihan Jin, Feihu Che, Pengpeng Shao, Zhengqi Wen, Jianhua Tao 0001
ACL (1)9
2026 Calibration-Aware Policy Optimization for Reasoning LLMs
abstract
Group Relative Policy Optimization (GRPO) enhances LLM reasoning but often induces overconfidence, where incorrect responses yield lower perplexity than correct ones, degrading relative calibration as described by the Area Under the Curve (AUC).Existing approaches either yield limited improvements in calibration or sacrifice gains in reasoning accuracy.We first prove that this degradation in GRPO-style algorithms stems from their uncertainty-agnostic advantage estimation, which inevitably misaligns optimization gradients with calibration.This leads to improved accuracy at the expense of degraded calibration.We then propose Calibration-Aware Policy Optimization (CAPO).It adopts a logistic AUC surrogate loss that is theoretically consistent and admits regret bound, enabling uncertainty-aware advantage estimation.By further incorporating a noise masking mechanism, CAPO achieves stable learning dynamics that jointly optimize calibration and accuracy.Experiments on multiple mathematical reasoning benchmarks show that CAPO-1.5Bsignificantly improves calibration by up to 15% while achieving accuracy comparable to or better than GRPO, and further boosts accuracy on downstream inference-time scaling tasks by up to 5%.Moreover, when allowed to abstain under low-confidence conditions, CAPO achieves a Pareto-optimal precision-coverage trade-off, highlighting its practical value for hallucination mitigation.
Xingzhou Lou, Meiqi Wu, Zhengqi Wen, Junge Zhang
ACL (1)4
2026 Beyond Examples: Towards Automated Thought-level In-Context Reasoning for Large Language Models
abstract
Jinyang Wu, Mingkuan Feng, Shuai Zhang, Feihu Che, Zhengqi Wen, Chonghua Liao, Ling Yang, Haoran Luo, Zheng Lian, Jianhua Tao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Mingkuan Feng, Shuai Zhang 0014, Feihu Che, Zhengqi Wen, Chonghua Liao, Zheng Lian 0004, Jianhua Tao 0001
ACL (1)5
2026 SPARK: Strategic Policy-Aware Exploration via Dynamic Branching for Long-Horizon Agentic Learning
abstract
Reinforcement learning has empowered large language models to act as intelligent agents, yet training them for long-horizon tasks remains challenging due to the scarcity of highquality trajectories, especially under limited resources.Existing methods typically scale up rollout sizes and indiscriminately allocate computational resources among intermediate steps.Such attempts inherently waste substantial computation budget on trivial steps while failing to guarantee sample quality.To address this, we propose SPARK (Strategic Policy-Aware ex-ploRation via Key-state dynamic branching), a novel framework that selectively branches at critical decision states for resource-efficient exploration.Our key insight is to activate adaptive branching exploration at critical decision points to probe promising trajectories, thereby achieving precise resource allocation that prioritizes sampling quality over blind coverage.This design leverages the agent's intrinsic decisionmaking signals to reduce dependence on human priors, enabling the agent to autonomously expand exploration and achieve stronger generalization.Experiments across diverse tasks (e.g., embodied planning), demonstrate that SPARK achieves superior success rates with significantly fewer training samples, exhibiting robust generalization even in unseen scenarios.Our code and checkpoints are available at https://github.com/jinyangwu/SPARK.
Shuai Zhang 0014, Zhengqi Wen, Jianhua Tao 0001
ACL (1)5
2026 OpenST: Toward open-set source tracing for neural codec deepfake audio
Yuankun Xie, Ruibo Fu, Zhengqi Wen, Songjun Cao, Chenxing Li, Haonan Cheng, Long Ye
Neurocomputing5
2026 SpeechPalette: A Comprehensive Speech Editing Method for Text-Based Speech Editing, One-Shot TTS and Attributes Editing
abstract
Speech editing has garnered more and more attention due to its diverse applications. However, existing systems often require substantial manual effort or have limited capabilities in attribute editing, imposing significant constraints. In this work, we present SpeechPalette, a comprehensive high-quality speech editing method that allows users to easily modify various attributes of the selected speech segment according to their preferences. Specifically, the proposed model approaches speech editing from a decoupling perspective, disentangling critical information such as text, pitch, duration and more from the input speech. Then, reconstruction is achieved through a mask and prediction mechanism. Furthermore, we leverage a diffusion model to predict the residuals between the real and predicted speech, further enhancing synthesis quality. The proposed method not only excels at text-based speech editing but also handles tasks involving pitch and speed rate adjustments. Moreover, it also demonstrates remarkable performance in one-shot text-to-speech scenarios. While recent large-scale models achieve impressive synthesis quality through massive computational resources, SpeechPalette offers a balanced approach with explicit fine-grained control over speech attributes, practical deployment requirements, and competitive performance relative to similarly-sized systems. Experimental results across a range of tasks consistently demonstrate the superior performance of our method compared to baseline systems. Additionally, comprehensive ablation studies validate the effectiveness of our proposed approach.
Tao Wang 0074, Jiangyan Yi, Ruibo Fu, Chunyu Qiang, Dading Chong, Dongyang Dai, Zhengqi Wen, Jianhua Tao 0001
IEEE Trans. Pattern Anal. Mach. Intell.8
2026 Personality-aware Multimodal Deception Detection with multimodal large language model
Cong Cai, Zhengqi Wen, Xuefei Liu, Jianhua Tao 0001, Bin Liu 0041
Pattern Recognit.2
2026 CMDPAD: A Chinese multimodal dynamic personality and affect dataset for affect prediction in conversations
Zisen Zhou, Chang Wen, Xuefei Liu, Jianhua Tao 0001, Zhengqi Wen, Zheng Lian 0004, Jinming Zhao, Bingsen Xiong, Shaozheng Qin
Pattern Recognit.6
2026 MSSG: Multi-Scale Speaker Graph Network for Active Speaker Detection
abstract
The active speaker detection task is to determine whether a person is speaking or not across a series of video frames. Existing methods heavily rely on facial information within the annotated face bounding boxes for cross-modal learning with audio. This leads to a substantial decline in detection performance when facial cues are unclear, such as in cases of face occlusion or low-resolution facial appearances. In this paper, we extend the perception scale using only face bounding box annotations to model both facial and gestural cues, addressing the over-reliance on facial cues in active speaker detection. We propose a novel graph neural network that models inter-speaker interactions and integrates various cues from individual speakers. The final detection results are obtained through a binary graph node classification task. Our method achieves state-of-the-art performance on the AVA-ActiveSpeaker dataset (mAP: 95.6%) and the ASW dataset (mAP: 99.4%), with a model size only 21% that of the second-best method. Additionally, when facial cues are of poor quality, our method demonstrates a significant performance advantage over existing approaches. The code and model weights will be available athttps://github.com/sdqdlgj/MSSG.
Guanjun Li, Jiangyan Yi, Zhengqi Wen, Ruibo Fu, Yuwang Wang, Jianhua Tao 0001
IEEE Trans. Multim.3
2025 Code-switching Mediated Sentence-level Semantic Learning
abstract
Code-switching is a linguistic phenomenon in which different languages are used interactively during conversation. It poses significant performance challenges to natural language processing (NLP) tasks due to the often monolingual nature of the underlying system. We focus on sentence-level semantic associations between the different code-switching expressions. And we propose an innovative task-free semantic learning method based on the semantic property. Specifically, there are many different ways of languages switching for a sentence with the same meaning. We refine this into a semantic computational method by designing the loss of semantic invariant constraint during the model optimization. In this work, we conduct thorough experiments on speech recognition, speech translation, and language modeling tasks. The experimental results fully demonstrate that the proposed method can widely improve the performance of code-switching related tasks.
Shuai Zhang 0014, Jiangyan Yi, Zhengqi Wen, Jianhua Tao 0001, Feihu Che, Ruibo Fu
AAAI3
2025 Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio Generation
abstract
Mainstream Text-to-Audio (TTA) models that rely on Mel-spectrograms often struggle to generate audio with rich content, leading to blurred or incoherent outputs. This stems from an inability to model intricate spectral details and textures. We investigate the role of U-Net components in generation and find that high-frequency components in skip-connections and the backbone are crucial for texture, while low-frequency backbone components are vital for the denoising process. Based on this, we propose “Mel-Refine,” a plug-and-play approach that enhances Mel-spectrogram quality by adjusting component weights during inference. Our method requires no additional training or fine-tuning and is fully compatible with any diffusion-based TTA architecture. Experiments show that Mel-Refine boosts the performance of the latest TTA model, Tango2, by $25 \%$, demonstrating its effectiveness.
Hongming Guo, Ruibo Fu, Yizhong Geng, Shuchen Shi, Tao Wang 0074, Chunyu Qiang, Ya Li 0001, Zhengqi Wen, Xuefei Liu, Chenxing Li
ASRU8
2025 ImViD: Immersive Volumetric Videos for Enhanced VR Engagement
abstract
User engagement is greatly enhanced by fully immersive multi-modal experiences that combine visual and auditory stimuli. Consequently, the next frontier in VR/AR technologies lies in immersive volumetric videos with complete scene capture, large 6-DoF interaction space, multimodal feedback, and high resolution & frame-rate contents. To stimulate the reconstruction of immersive volumetric videos, we introduce ImViD, a multi-view, multi-modal dataset featuring complete space-oriented data capture and various indoor/outdoor scenarios. Our capture rig supports multi-view video-audio capture while on the move, a capability absent in existing datasets, significantly enhancing the completeness, flexibility, and efficiency of data capture.The captured multi-view videos (with synchronized audios) are in 5K resolution at 60FPS, lasting from 1-5 minutes, and include rich foreground-background elements, and complex dynamics. We benchmark existing methods using our dataset and establish a base pipeline for constructing immersive volumetric videos from multi-view audiovisual inputs for 6-DoF multi-modal immersive VR experiences. The benchmark and the reconstruction and interaction results demonstrate the effectiveness of our dataset and baseline method, which we believe will stimulate future research on immersive volumetric video production. Project Page: https://yzxqh.github.io/ImViD/
Zhengxian Yang, Shi Pan, Shengqi Wang, Guanjun Li, Zhengqi Wen, Borong Lin, Jianhua Tao 0001
CVPR7
2025 DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
abstract
In recent years, speech diffusion models have advanced rapidly. Alongside the widely used U-Net architecture, transformer-based models such as the Diffusion Transformer (DiT) have also gained attention. However, current DiT speech models treat Mel spectrograms as general images, which overlooks the specific acoustic properties of speech. To address these limitations, we propose a method called Directional Patch Interaction for Text-to-Speech (DPI-TTS), which builds on DiT and achieves fast training without compromising accuracy. Notably, DPI-TTS employs a low-to-high frequency, frame-by-frame progressive inference approach that aligns more closely with acoustic properties, enhancing the naturalness of the generated speech. Additionally, we introduce a fine-grained style temporal modeling method that further improves speaker style similarity. Experimental results demonstrate that our method increases the training speed by nearly 2 times and significantly outperforms the baseline models.
Ruibo Fu, Zhengqi Wen, Tao Wang 0074, Chunyu Qiang, Jianhua Tao 0001, Chenxing Li, Shuchen Shi, Yuankun Xie, Xuefei Liu, Guanjun Li
ICASSP3
2025 Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0
abstract
Speech synthesis technology has posed a serious threat to speaker verification systems. Currently, the most effective fake audio detection methods utilize pretrained models, and integrating features from various layers of pretrained model further enhances detection performance. However, most of the previously proposed fusion methods require fine-tuning the pretrained models, resulting in excessively long training times and hindering model iteration when facing new speech synthesis technology. To address this issue, this paper proposes a feature fusion method based on the Mixture of Experts, which extracts and integrates features relevant to fake audio detection from layer features, guided by a gating network based on the last layer feature, while freezing the pretrained model. Experiments conducted on the ASVspoof2019 and ASVspoof2021 datasets demonstrate that the proposed method achieves competitive performance compared to those requiring fine-tuning.
Ruibo Fu, Zhengqi Wen, Jianhua Tao 0001, Yuankun Xie, Shuchen Shi, Chenxing Li, Xuefei Liu, Guanjun Li
ICASSP3
2025 DetailTTS: Learning Residual Detail Information for Zero-shot Text-to-speech
abstract
Traditional text-to-speech (TTS) systems often face challenges in aligning text and speech, leading to the omission of critical linguistic and acoustic details. This misalignment creates an information gap, which existing methods attempt to address by incorporating additional inputs, but these often introduce data inconsistencies and increase complexity. To address these issues, we propose DetailTTS, a zero-shot TTS system based on a conditional variational autoencoder. It incorporates two key components: the Prior Detail Module and the Duration Detail Module, which capture residual detail information missed during alignment. These modules effectively enhance the model’s ability to retain fine-grained details, significantly improving speech quality while simplifying the model by obviating the need for additional inputs. Experiments on the WenetSpeech4TTS dataset show that DetailTTS outperforms traditional TTS systems in both naturalness and speaker similarity, even in zero-shot scenarios. Our source code and demo page are available at https://detailtts.github.io/.
Yichen Han, Yizhong Geng, Yingming Gao, Fengping Wang, Bingsong Bai, Jinlong Xue, Yayue Deng, Zhengqi Wen, Ya Li 0001
ICASSP10
2025 MTPareto: A MultiModal Targeted Pareto Framework for Fake News Detection
abstract
Multimodal fake news detection is essential for maintaining the authenticity of Internet multimedia information. Significant differences in form and content of multimodal information lead to intensified optimization conflicts, hindering effective model training as well as reducing the effectiveness of existing fusion methods for bimodal. To address this problem, we propose the MTPareto framework to optimize multimodal fusion, using a Targeted Pareto(TPareto) optimization algorithm for fusion-level-specific objective learning with a certain focus. Based on the designed hierarchical fusion network, the algorithm defines three fusion levels with corresponding losses and implements all-modal-oriented Pareto gradient integration for each. This approach accomplishes superior multimodal fusion by utilizing the information obtained from intermediate fusion to provide positive effects to the entire process. Experiment results on FakeSV and FVC datasets show that the proposed framework outperforms baselines and the TPareto optimization algorithm achieves 2.40% and 1.89% accuracy improvement respectively.
Kaiying Yan, Moyang Liu, Ruibo Fu, Zhengqi Wen, Jianhua Tao 0001, Xuefei Liu, Guanjun Li
ICASSP5
2025 M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction
abstract
The brain-assisted target speaker extraction (TSE) aims to extract the attended speech from mixed speech by utilizing the brain neural activities, for example Electroencephalography (EEG). However, existing models overlook the issue of temporal misalignment between speech and EEG modalities, which hampers TSE performance. In addition, the speech encoder in current models typically uses basic temporal operations (e.g., one-dimensional convolution), which are unable to effectively extract target speaker information. To address these issues, this paper proposes a multi-scale and multi-modal alignment network (M3ANet) for brain-assisted TSE. Specifically, to eliminate the temporal inconsistency between EEG and speech modalities, the modal alignment module that uses a contrastive learning strategy is applied to align the temporal features of both modalities. Additionally, to fully extract speech information, multi-scale convolutions with GroupMamba modules are used as the speech encoder, which scans speech features at each scale from different directions, enabling the model to capture deep sequence information. Experimental results on three publicly available datasets show that the proposed model outperforms current state-of-the-art methods across various evaluation metrics, highlighting the effectiveness of our proposed method. The source code is available at: https://github.com/fchest/M3ANet.
Cunhang Fan, Jian Zhou 0006, Zexu Pan, Youdian Gao, Xiaoke Yang, Zhengqi Wen, Zhao Lv
IJCAI8
2025 MDPE: A Multimodal Deception Dataset with Personality and Emotional Characteristics
abstract
Deception detection has garnered increasing attention in recent years due to the significant growth of digital media and heightened ethical and security concerns. It has been extensively studied using multimodal methods, including video, audio, and text. In addition, individual differences in deception production and detection are believed to play a crucial role. Although some studies have utilized individual information such as personality traits to enhance the performance of deception detection, current systems remain limited, partly due to a lack of sufficient datasets for evaluating performance. To address this issue, we introduce a multimodal deception dataset MDPE. Besides deception features, this dataset also includes individual differences information in personality and emotional expression characteristics. It can explore the impact of individual differences on deception behavior. It comprises over 104 hours of deception and emotional videos from 193 subjects. Furthermore, we conducted numerous experiments to provide valuable insights for future deception detection research. MDPE not only supports deception detection, but also provides conditions for tasks such as personality recognition and emotion recognition, and can even study the relationships between them. We believe that MDPE will become a valuable resource for promoting research in the field of affective computing.
Cong Cai, Shan Liang 0007, Xuefei Liu, Kang Zhu, Zhengqi Wen, Jianhua Tao 0001, Jizhou Cui, Zhenhua Cheng, Hanzhe Xu, Ruibo Fu, Bin Liu 0041
ACM Multimedia5
2025 ALLM4ADD: Unlocking the Capabilities of Audio Large Language Models for Audio Deepfake Detection
abstract
Audio deepfake detection (ADD) has grown increasingly important due to the rise of high-fidelity audio generative models and their potential for misuse. Given that audio large language models (ALLMs) have made significant progress in various audio processing tasks, a heuristic question arises: Can ALLMs be leveraged to solve ADD?. In this paper, we first conduct a comprehensive zero-shot evaluation of ALLMs on ADD, revealing their ineffectiveness. To this end, we propose ALLM4ADD, an ALLM-driven framework for ADD. Specifically, we reformulate ADD task as an audio question answering problem, prompting the model with the question: ''Is this audio fake or real?''. We then perform supervised fine-tuning to enable the ALLM to assess the authenticity of query audio. Extensive experiments are conducted to demonstrate that our ALLM-based method can achieve superior performance in fake audio detection, particularly in data-scarce scenarios. As a pioneering study, we anticipate that this work will inspire the research community to leverage ALLMs to develop more effective ADD systems. Code is available at https://github.com/ucas-hao/qwen_audio_for_add.git.
Jiangyan Yi, Chenglong Wang 0001, Jianhua Tao 0001, Zheng Lian 0004, Yong Ren 0006, Yujie Chen 0006, Zhengqi Wen
ACM Multimedia9
2025 DashFusion: Dual-Stream Alignment With Hierarchical Bottleneck Fusion for Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis (MSA) integrates various modalities, such as text, image, and audio, to provide a more comprehensive understanding of sentiment. However, effective MSA is challenged by alignment and fusion issues. Alignment requires synchronizing both temporal and semantic information across modalities, while fusion involves integrating these aligned features into a unified representation. Existing methods often address alignment or fusion in isolation, leading to limitations in performance and efficiency. To tackle these issues, we propose a novel framework called dual-stream alignment with hierarchical bottleneck fusion (DashFusion). First, the dual-stream alignment module synchronizes multimodal features through temporal and semantic alignment. Temporal alignment employs cross-modal attention (CA) to establish frame-level correspondences among multimodal sequences. Semantic alignment ensures consistency across the feature space through contrastive learning. Second, supervised contrastive learning (SCL) leverages label information to refine the modality features. Finally, hierarchical bottleneck fusion (HBF) progressively integrates multimodal information through compressed bottleneck tokens, which achieves a balance between performance and computational efficiency. We evaluate DashFusion on three datasets: CMU-MOSI, CMU-MOSEI, and CH-SIMS. Experimental results demonstrate that DashFusion achieves state-of-the-art (SOTA) performance across various metrics, and ablation studies confirm the effectiveness of our alignment and fusion techniques. The codes for our experiments are available at https://github.com/ultramarineX/DashFusion.
Yuhua Wen, Yingying Zhou, Yingming Gao, Zhengqi Wen, Jianhua Tao 0001, Ya Li 0001
IEEE Trans. Neural Networks Learn. Syst.5
2024 Dual-View Multimodal Interaction in Multimodal Sentiment Analysis
abstract
Outstanding performance in sentiment analysis not only relies on the design of sophisticated fusion methods but also on the crucial step of designing excellent modal interaction methods. To the best of our knowledge, there are few methods addressing the capture of multimodal spatial features. Majority of feature interactions have been primarily focused on temporal aspects, with less attention given to the combined spatiotemporal feature interaction (SFI). In this paper, we design a dual-view multimodal interaction method, named DVMI, primarily consisting of two parts. In the first part, a triangular convolutional module is proposed for ample temporal interaction between modalities, implicit local and global SFI, and capturing global spatial representations. Building upon the foundation laid in the first part, the second part employs an attention mechanism for explicit global SFI. To demonstrate the effectiveness of the DVMI framework,we conduct extensive experiments on three datasets, achieving state-of-the-art experimental results.
Kang Zhu, Cunhang Fan, Jianhua Tao 0001, Jun Xue 0001, Xuefei Liu, Zhengqi Wen, Zhao Lv
ICME8
2024 Codecfake: An Initial Dataset for Detecting LLM-based Deepfake Audio
Yuankun Xie, Ruibo Fu, Zhengqi Wen, Jianhua Tao 0001, Xuefei Liu, Shuchen Shi
INTERSPEECH4
2024 PPPR: Portable Plug-in Prompt Refiner for Text to Audio Generation
Shuchen Shi, Ruibo Fu, Zhengqi Wen, Jianhua Tao 0001, Tao Wang 0074, Chunyu Qiang, Xuefei Liu
INTERSPEECH3
2024 Genuine-Focused Learning using Mask AutoEncoder for Generalized Fake Audio Detection
Ruibo Fu, Zhengqi Wen, Yuankun Xie, Jianhua Tao 0001, Xuefei Liu, Shuchen Shi
INTERSPEECH3
2024 Generalized Fake Audio Detection via Deep Stable Learning
Ruibo Fu, Zhengqi Wen, Yuankun Xie, Xuefei Liu, Jianhua Tao 0001, Shuchen Shi
INTERSPEECH3
2024 Generalized Source Tracing: Detecting Novel Audio Deepfake Algorithm with Real Emphasis and Fake Dispersion Strategy
Yuankun Xie, Ruibo Fu, Zhengqi Wen, Haonan Cheng, Long Ye, Jianhua Tao 0001
INTERSPEECH3
2024 Residual Speaker Representation for One-Shot Voice Conversion
abstract
International audience
Jiangyan Yi, Tao Wang 0074, Yong Ren 0006, Rongxiu Zhong, Zhengqi Wen, Jianhua Tao 0001
INTERSPEECH6
2024 TraceableSpeech: Towards Proactively Traceable Text-to-Speech with Watermarking
Junzuo Zhou, Jiangyan Yi, Tao Wang 0074, Jianhua Tao 0001, Ye Bai 0001, Chu Yuan Zhang, Yong Ren 0006, Zhengqi Wen
INTERSPEECH8
2024 Emotion selectable end-to-end text-based speech editing
Tao Wang 0074, Jiangyan Yi, Ruibo Fu, Jianhua Tao 0001, Zhengqi Wen, Chu Yuan Zhang
Artif. Intell.5
2024 VLP2MSA: Expanding vision-language pre-training to multimodal sentiment analysis
Guofeng Yi, Cunhang Fan, Kang Zhu, Zhao Lv, Shan Liang 0007, Zhengqi Wen, Guanxiong Pei, Taihao Li, Jianhua Tao 0001
Knowl. Based Syst.6
2024 Dual-Branch Knowledge Distillation for Noise-Robust Synthetic Speech Detection
abstract
Most research in synthetic speech detection (SSD) focuses on improving performance on standard noise-free datasets. However, in actual situations, noise interference is usually present, causing significant performance degradation in SSD systems. To improve noise robustness, this paper proposes a dual-branch knowledge distillation synthetic speech detection (DKDSSD) method. Specifically, a parallel data flow of the clean teacher branch and the noisy student branch is designed, and interactive fusion module and response-based teacher-student paradigms are proposed to guide the training of noisy data from both the data distribution and decision-making perspectives. In the noisy student branch, speech enhancement is introduced initially for denoising, aiming to reduce the interference of strong noise. The proposed interactive fusion combines denoised features and noisy features to mitigate the impact of speech distortion and ensure consistency with the data distribution of the clean branch. The teacher-student paradigm maps the student's decision space to the teacher's decision space, enabling noisy speech to behave similarly to clean speech. Additionally, a joint training method is employed to optimize both branches for achieving global optimality. Experimental results based on multiple datasets demonstrate that the proposed method performs effectively in noisy environments and maintains its performance in cross-dataset experiments. Source code is available athttps://github.com/fchest/DKDSSD.
Cunhang Fan, Mingming Ding, Jianhua Tao 0001, Ruibo Fu, Jiangyan Yi, Zhengqi Wen, Zhao Lv
IEEE ACM Trans. Audio Speech Lang. Process.6
2023 Learning From Yourself: A Self-Distillation Method For Fake Speech Detection
abstract
In this paper, we propose a novel self-distillation method for fake speech detection (FSD), which can significantly improve the performance of FSD without increasing the model complexity. For FSD, some fine-grained information is very important, such as spectrogram defects, mute segments, and so on, which are often perceived by shallow networks. However, shallow networks have much noise, which can not capture this very well. To address this problem, we propose using the deepest network instruct shallow network for enhancing shallow networks. Specifically, the networks of FSD are divided into several segments, the deepest network being used as the teacher model, and all shallow networks become multiple student models by adding classifiers. Meanwhile, the distillation path between the deepest network feature and shallow network features is used to reduce the feature difference. A series of experimental results on the ASVspoof 2019 LA and PA datasets show the effectiveness of the proposed method, with significant improvements compared to the baseline.
Jun Xue 0001, Cunhang Fan, Jiangyan Yi, Chenglong Wang 0001, Zhengqi Wen, Dan Zhang 0014, Zhao Lv
ICASSP5
2022 Context-Aware Mask Prediction Network for End-to-End Text-Based Speech Editing
abstract
The text-based speech editor allows the editing of speech through intuitive cutting, copying, and pasting operations to speed up the process of editing speech. However, the major drawback of current systems is that edited speech often sounds unnatural and it is not obvious how to synthesize records according to a new word not appearing in the transcript. This paper proposes a novel end-to-end text-based speech editing method called context-aware mask prediction network (CampNet), which avoids the unnatural phenomenon caused by cut-copy-paste operation in the traditional method and can synthesize a new word not appearing in the transcript. Besides, three text-based speech editing operations based on CampNet are designed: deletion, replacement, and insertion. These operations can comprehensively cover different kinds of situations that text-based speech editing can face. The subjective and objective experiments on VCTK and LibriTTS data sets show that the speech editing results based on CampNet are better than TTS technology, manual editing, and VoCo method (the combination of speech synthesis and speech conversion). We also conducted detailed ablation experiments to explore the effect of the CampNet structure on its performance. Examples of generated speech can be found at https://hairuo55.github.io/CampNet-demo.
Tao Wang 0074, Jiangyan Yi, Liqun Deng, Ruibo Fu, Jianhua Tao 0001, Zhengqi Wen
ICASSP6
2022 ADD 2022: the first Audio Deep Synthesis Detection Challenge
abstract
Audio deepfake detection is an emerging topic, which was included in the ASVspoof 2021. However, the recent shared tasks have not covered many real-life and challenging scenarios. The first Audio Deep synthesis Detection challenge (ADD) was motivated to fill in the gap. The ADD 2022 includes three tracks: low-quality fake audio detection (LF), partially fake audio detection (PF) and audio fake game (FG). The LF track focuses on dealing with bona fide and fully fake utterances with various real-world noises etc. The PF track aims to distinguish the partially fake audio from the real. The FG track is a rivalry game, which includes two tasks: an audio generation task and an audio fake detection task. In this paper, we describe the datasets, evaluation metrics, and protocols. We also report major findings that reflect the recent advances in audio deepfake detection tasks.
Jiangyan Yi, Ruibo Fu, Jianhua Tao 0001, Shuai Nie 0001, Haoxin Ma, Chenglong Wang 0001, Tao Wang 0074, Zhengkun Tian, Ye Bai 0001, Cunhang Fan, Shan Liang 0007, Shuai Zhang 0014, Xinrui Yan, Zhengqi Wen, Haizhou Li 0001
ICASSP16
2022 Hybrid Autoregressive and Non-Autoregressive Transformer Models for Speech Recognition
abstract
The autoregressive (AR) models, such as attention-based encoder-decoder models and RNN-Transducer, have achieved great success in speech recognition. They predict the output sequence conditioned on the previous tokens and acoustic encoded states, which is inefficient on GPUs. The non-autoregressive (NAR) models can get rid of the temporal dependency between the output tokens and predict the entire output tokens in one inference step. However, the NAR model still faces two major problems. Firstly, there is still a great gap in performance between the NAR models and the advanced AR models. Secondly, it’s difficult for most of the NAR models to train and converge. We propose a hybrid autoregressive and non-autoregressive transformer (HANAT) model, which integrates AR and NAR models deeply by sharing parameters. We assume that the AR model will assist the NAR model to learn some linguistic dependencies and accelerate the convergence. Furthermore, the two-stage hybrid inference is applied to improve the model performance. All the experiments are conducted on a mandarin dataset ASIEHLL-1 and a english dataset librispeech-960 h. The results show that the HANAT can achieve a competitive performance with the AR model and outperform many complicated NAR models. Besides, the RTF is only 1/5 of the AR model.
Zhengkun Tian, Jiangyan Yi, Jianhua Tao 0001, Shuai Zhang 0014, Zhengqi Wen
IEEE Signal Process. Lett.5
2022 NeuralDPS: Neural Deterministic Plus Stochastic Model With Multiband Excitation for Noise-Controllable Waveform Generation
abstract
The traditional vocoders have the advantages of high synthesis efficiency, strong interpretability, and speech editability, while the neural vocoders have the advantage of high synthesis quality. To combine the advantages of two vocoders, inspired by the traditional deterministic plus stochastic model, this paper proposes a novel neural vocoder named NeuralDPS which can retain high speech quality and acquire high synthesis efficiency and noise controllability. Firstly, this framework contains four modules: a deterministic source module, a stochastic source module, a neural V/UV decision module and a neural filter module. The input required by the vocoder is just the spectral parameter, which avoids the error caused by estimating additional parameters, such as F0. Secondly, to solve the problem that different frequency bands may have different proportions of deterministic components and stochastic components, a multiband excitation strategy is used to generate a more accurate excitation signal and reduce the neural filter’s burden. Thirdly, a method to control noise components of speech is proposed. In this way, the signal-to-noise ratio (SNR) of speech can be adjusted easily. Objective and subjective experimental results show that our proposed NeuralDPS vocoder can obtain similar performance with the WaveNet and it generates waveforms at least 280 times faster than the WaveNet vocoder. It is also 28% faster than WaveGAN’s synthesis efficiency on a single CPU core. We have also verified through experiments that this method can effectively control the noise components in the predicted speech and adjust the SNR of speech.
Tao Wang 0074, Ruibo Fu, Jiangyan Yi, Jianhua Tao 0001, Zhengqi Wen
IEEE ACM Trans. Audio Speech Lang. Process.5
2022 CampNet: Context-Aware Mask Prediction for End-to-End Text-Based Speech Editing
abstract
The text-based speech editor allows the editing of speech through intuitive cutting, copying, and pasting operations to speed up the process of editing speech. However, the major drawback of current systems is that edited speech often sounds unnatural due to cut-copy-paste operation. In addition, it is not obvious how to synthesize records according to a new word not appearing in the transcript, which often needs the help of text-to-speech (TTS) and voice conversion (VC) technology at the same time. This paper first proposes a novel end-to-end text-based speech editing method called context-aware mask prediction network (CampNet). The model can simulate the text-based speech editing process by randomly masking part of speech and then predicting the masked region by sensing the speech context. It can solve unnatural prosody in the edited region and synthesize the speech corresponding to the unseen words in the transcript. Secondly, for the possible operation of text-based speech editing, we design three text-based operations based on CampNet: deletion, insertion, and replacement. These operations can cover various situations of speech editing. Thirdly, to synthesize the speech corresponding to long text in insertion and replacement operations, a word-level autoregressive generation method is proposed, which can synthesize the speech of arbitrary length text. Fourthly, we propose a speaker adaptation method using only one sentence for CampNet and explore the ability of few-shot learning based on CampNet, which provides a new idea for speech forgery tasks. The subjective and objective experiments11Examples of generated speech can be found athttps://hairuo55.github.io/CampNet.on VCTK and LibriTTS datasets show that the speech editing results based on CampNet are better than TTS technology, manual editing, and VoCo method (the combination of TTS and VC). We also conduct detailed ablation experiments to explore the effect of the CampNet structure on its performance. Finally, the experiment shows that speaker adaptation with only one sentence can further improve the naturalness of speech editing for one-shot learning.
Tao Wang 0074, Jiangyan Yi, Ruibo Fu, Jianhua Tao 0001, Zhengqi Wen
IEEE ACM Trans. Audio Speech Lang. Process.5
2021 Decoupling Pronunciation and Language for End-to-End Code-Switching Automatic Speech Recognition
abstract
Despite the recent significant advances witnessed in end-to-end (E2E) ASR system for code-switching, hunger for audio-text paired data limits the further improvement of the models’ performance. In this paper, we propose a decoupled transformer model to use mono-lingual paired data and unpaired text data to alleviate the problem of code-switching data shortage. The model is decoupled into two parts: audio-to-phoneme (A2P) network and phoneme-to-text (P2T) network. The A2P network can learn acoustic pattern scenarios using large-scale monolingual paired data. Meanwhile, it generates multiple phoneme sequence candidates for single audio data in real time during the training process. Then the generated phoneme-text paired data is used to train the P2T network. This network can be pre-trained with large amounts of external unpaired text data. By using monolingual data and unpaired text data, the decoupled transformer model reduces the high dependency on code-switching paired training data of E2E model to a certain extent. Finally, the two networks are optimized jointly through attention fusion. We evaluate the proposed method on the public Mandarin-English code-switching dataset. Compared with our transformer baseline, the proposed method achieves 18.14% relative mix error rate reduction.
Shuai Zhang 0014, Jiangyan Yi, Zhengkun Tian, Ye Bai 0001, Jianhua Tao 0001, Zhengqi Wen
ICASSP6
2021 Bi-Level Style and Prosody Decoupling Modeling for Personalized End-to-End Speech Synthesis
abstract
End-to-end framework can generate high-quality and high-similarity speech in the personalized speech synthesis task. However, the generalization of out-of-domain texts is still a challenging task. Limited target data leads to unacceptable errors and poor prosody and similarity performance of the synthetic speech. In this paper, we present a bi-level function decoupling framework to realise separate modeling and controlling for solving above problems. Firstly, on the style representation modeling level, compared with the conventional methods that use single embedding to model all the text dependent discrepancies, it is proposed that the speaker embedding and prosody embedding are modeled separately based on the reference audio and phonetic posteriorgram (PPG) by a multi-head attention mechanism. Secondly, on the model structure level, the decoder model structure is factored into average-net and adaptation-net, where the duration prosody controlling and speaker timbre imitation are mainly designed in relatively separate areas. Experimental results on Mandarin dataset show that the proposed methods lead to an improvement on both robustness, naturalness and similarity.
Ruibo Fu, Jianhua Tao 0001, Zhengqi Wen, Jiangyan Yi, Tao Wang 0074, Chunyu Qiang
ICASSP3
2021 Prosody and Voice Factorization for Few-Shot Speaker Adaptation in the Challenge M2voc 2021
abstract
The paper describes the CASIA speech synthesis system entry for challenge M2VoC 2021. The low similarity and naturalness of synthesized speech remains a challenging problem for speaker adaptation with few resources. Since the end-to-end acoustic model is too complex to interpret, overfitting will occur when training with few data. To prevent the model from overfitting, this paper proposes a novel speaker adaptation framework that decomposes the prosody and voice characteristics in the end-to-end model. A prosody control attention is proposed to control the phonemes’ duration of different speakers. To make the attention controlled by the prosody information, a set of phoneme-level transition tokens is auto-learned from the prosody encoder in our framework and these transition tokens can determine the duration of phonemes in the attention mechanism. Secondly, when we need to use small data set for speaker adaptation, we just need to adapt the speaker related prosody model and decoder, which can prevent the model from overfitting. Further, we use a data puring model to automatically optimize the quality of datasets. Experiments demonstrate the effectiveness of speaker adaptation based on our method, and we (team identifier is T03) get the top three results in competition M2VoC by using this framework.
Tao Wang 0074, Ruibo Fu, Jiangyan Yi, Jianhua Tao 0001, Zhengqi Wen, Chunyu Qiang
ICASSP5
2021 End-to-End Spelling Correction Conditioned on Acoustic Feature for Code-Switching Speech Recognition
Shuai Zhang 0014, Jiangyan Yi, Zhengkun Tian, Ye Bai 0001, Jianhua Tao 0001, Xuefei Liu, Zhengqi Wen
Interspeech7
2021 FSR: Accelerating the Inference Process of Transducer-Based Models by Applying Fast-Skip Regularization
abstract
Transducer-based models, such as RNN-Transducer and transformer-transducer, have achieved great success in speech recognition. A typical transducer model decodes the output sequence conditioned on the current acoustic state and previously predicted tokens step by step. Statistically, The number of blank tokens in the prediction results accounts for nearly 90\% of all tokens. It takes a lot of computation and time to predict the blank tokens, but only the non-blank tokens will appear in the final output sequence. Therefore, we propose a method named fast-skip regularization, which tries to align the blank position predicted by a transducer with that predicted by a CTC model. During the inference, the transducer model can predict the blank tokens in advance by a simple CTC project layer without many complicated forward calculations of the transducer decoder and then skip them, which will reduce the computation and improve the inference speed greatly. All experiments are conducted on a public Chinese mandarin dataset AISHELL-1. The results show that the fast-skip regularization can indeed help the transducer model learn the blank position alignments. Besides, the inference with fast-skip can be speeded up nearly 4 times with only a little performance degradation.
Zhengkun Tian, Jiangyan Yi, Ye Bai 0001, Jianhua Tao 0001, Shuai Zhang 0014, Zhengqi Wen
Interspeech6
2021 Fast End-to-End Speech Recognition Via Non-Autoregressive Models and Cross-Modal Knowledge Transferring From BERT
abstract
Attention-based encoder-decoder (AED) models have achieved promising performance in speech recognition. However, because the decoder predicts text tokens (such as characters or words) in an autoregressive manner, it is difficult for an AED model to predict all tokens in parallel. This makes the inference speed relatively slow. In contrast, we propose an end-to-end non-autoregressive speech recognition model called LASO (Listen Attentively, and Spell Once). The model aggregates encoded speech features into the hidden representations corresponding to each token with attention mechanisms. Thus, the model can capture the token relations by self-attention on the aggregated hidden representations from the whole speech signal rather than autoregressive modeling on tokens. Without explicitly autoregressive language modeling, this model predicts all tokens in the sequence in parallel so that the inference is efficient. Moreover, we propose a cross-modal transfer learning method to use a text-modal language model to improve the performance of speech-modal LASO by aligning token semantics. We conduct experiments on two scales of public Chinese speech datasets AISHELL-1 and AISHELL-2. Experimental results show that our proposed model achieves a speedup of about 50× and competitive performance, compared with the autoregressive transformer models. And the cross-modal knowledge transferring from the text-modal model can improve the performance of the speech-modal model.
Ye Bai 0001, Jiangyan Yi, Jianhua Tao 0001, Zhengkun Tian, Zhengqi Wen, Shuai Zhang 0014
IEEE ACM Trans. Audio Speech Lang. Process.5
2021 Integrating Knowledge Into End-to-End Speech Recognition From External Text-Only Data
abstract
Attention-based encoder-decoder (AED) models have achieved promising performance in speech recognition. However, because of the end-to-end training, an AED model is usually trained with speech-text paired data. It is challenging to incorporate external text-only data into AED models. Another issue of the AED model is that it does not use the right context of a text token while predicting the token. To alleviate the above two issues, we propose a unified method called LST (Learn Spelling from Teachers) to integrate knowledge into an AED model from the external text-only data and leverage the whole context in a sentence. The method is divided into two stages. First, in the representation stage, a language model is trained on the text. It can be seen as that the knowledge in the text is compressed into the LM. Then, at the transferring stage, the knowledge is transferred to the AED model via teacher-student learning. To further use the whole context of the text sentence, we propose an LM called causal cloze completer (COR), which estimates the probability of a token, given both the left context and the right context of it. Therefore, with LST training, the AED model can leverage the whole context in the sentence. Different from fusion based methods, which use LM during decoding, the proposed method does not increase any extra complexity at the inference stage. We conduct experiments on two scales of public Chinese datasets AISHELL-1 and AISHELL-2. The experimental results demonstrate the effectiveness of leveraging external text-only data and the whole context in a sentence with our proposed method, compared with baseline hybrid systems and AED model based systems.
Ye Bai 0001, Jiangyan Yi, Jianhua Tao 0001, Zhengqi Wen, Zhengkun Tian, Shuai Zhang 0014
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 Gated Recurrent Fusion With Joint Training Framework for Robust End-to-End Speech Recognition
abstract
The joint training framework for speech enhancement and recognition methods have obtained quite good performances for robust end-to-end automatic speech recognition (ASR). However, these methods only utilize the enhanced feature as the input of the speech recognition component, which are affected by the speech distortion problem. In order to address this problem, this paper proposes a gated recurrent fusion (GRF) method with joint training framework for robust end-to-end ASR. The GRF algorithm is used to dynamically combine the noisy and enhanced features. Therefore, the GRF can not only remove the noise signals from the enhanced features, but also learn the raw fine structures from the noisy features so that it can alleviate the speech distortion. The proposed method consists of speech enhancement, GRF and speech recognition. Firstly, the mask based speech enhancement network is applied to enhance the input speech. Secondly, the GRF is applied to address the speech distortion problem. Thirdly, to improve the performance of ASR, the state-of-the-art speech transformer algorithm is used as the speech recognition component. Finally, the joint training framework is utilized to optimize these three components, simultaneously. Our experiments are conducted on an open-source Mandarin speech corpus called AISHELL-1. Experimental results show that the proposed method achieves the relative character error rate (CER) reduction of 10.04% over the conventional joint enhancement and transformer method only using the enhanced features. Especially for the low signal-to-noise ratio (0 dB), our proposed method can achieves better performances with 12.67% CER reduction, which suggests the potential of our proposed method.
Cunhang Fan, Jiangyan Yi, Jianhua Tao 0001, Zhengkun Tian, Bin Liu 0041, Zhengqi Wen
IEEE ACM Trans. Audio Speech Lang. Process.6
2020 Focusing on Attention: Prosody Transfer and Adaptative Optimization Strategy for Multi-Speaker End-to-End Speech Synthesis
abstract
End-to-end speech synthesis can generate high-quality synthetic speech and achieve high similarity scores with low-resource adaptation data. However, the generalization of out-domain texts is still a challenging task. The limited adaptation data leads to unacceptable errors and the poor prosody performance of the synthetic speech. In this paper, we present two novel methods to handle the above problems by focusing on the attention. Firstly, compared with the conventional methods that extract prosody embeddings for conditioning input, a duration controller with feedback mechanism is proposed, which can control the states transition in the sequence-to-sequence model more directly and precisely. Secondly, to alleviate the unmatching text-audio pairs' impact on model, an adaptative optimization strategy which would consider the matching degree of the training sample is also proposed. Experimental results on Mandarin dataset show that proposed methods lead to an improvement on both robustness and overall naturalness.
Ruibo Fu, Jianhua Tao 0001, Zhengqi Wen, Jiangyan Yi, Tao Wang 0074
ICASSP3
2020 Synchronous Transformers for end-to-end Speech Recognition
abstract
For most of the attention-based sequence-to-sequence models, the decoder predicts the output sequence conditioned on the entire input sequence processed by the encoder. The asynchronous problem between the encoding and decoding makes these models difficult to be applied for online speech recognition. In this paper, we propose a model named synchronous transformer to address this problem, which can predict the output sequence chunk by chunk. Once a fixed-length chunk of the input sequence is processed by the encoder, the decoder begins to predict symbols immediately. During training, a forward-backward algorithm is introduced to optimize all the possible alignment paths. Our model is evaluated on a Mandarin dataset AISHELL-1. The experiments show that the synchronous transformer is able to perform encoding and decoding synchronously, and achieves a character error rate of 8.91% on the test set.
Zhengkun Tian, Jiangyan Yi, Ye Bai 0001, Jianhua Tao 0001, Shuai Zhang 0014, Zhengqi Wen
ICASSP6
2020 Listen Attentively, and Spell Once: Whole Sentence Generation via a Non-Autoregressive Architecture for Low-Latency Speech Recognition
abstract
Although attention based end-to-end models have achieved promising performance in speech recognition, the multi-pass forward computation in beam-search increases inference time cost, which limits their practical applications.To address this issue, we propose a non-autoregressive end-to-end speech recognition system called LASO (listen attentively, and spell once).Because of the non-autoregressive property, LASO predicts a textual token in the sequence without the dependence on other tokens.Without beam-search, the one-pass propagation much reduces inference time cost of LASO.And because the model is based on the attention based feedforward structure, the computation can be implemented in parallel efficiently.We conduct experiments on publicly available Chinese dataset AISHELL-1.LASO achieves a character error rate of 6.4%, which outperforms the state-of-the-art autoregressive transformer model (6.7%).The average inference latency is 21 ms, which is 1/50 of the autoregressive transformer model.
Ye Bai 0001, Jiangyan Yi, Jianhua Tao 0001, Zhengkun Tian, Zhengqi Wen, Shuai Zhang 0014
INTERSPEECH5
2020 Gated Recurrent Fusion of Spatial and Spectral Features for Multi-Channel Speech Separation with Deep Embedding Representations
Cunhang Fan, Jianhua Tao 0001, Bin Liu 0041, Jiangyan Yi, Zhengqi Wen
INTERSPEECH5
2020 Joint Training for Simultaneous Speech Denoising and Dereverberation with Deep Embedding Representations
Cunhang Fan, Jianhua Tao 0001, Bin Liu 0041, Jiangyan Yi, Zhengqi Wen
INTERSPEECH5
2020 Dynamic Soft Windowing and Language Dependent Style Token for Code-Switching End-to-End Speech Synthesis
Ruibo Fu, Jianhua Tao 0001, Zhengqi Wen, Jiangyan Yi, Chunyu Qiang, Tao Wang 0074
INTERSPEECH3
2020 Dynamic Speaker Representations Adjustment and Decoder Factorization for Speaker Adaptation in End-to-End Speech Synthesis
Ruibo Fu, Jianhua Tao 0001, Zhengqi Wen, Jiangyan Yi, Tao Wang 0074, Chunyu Qiang
INTERSPEECH3
2020 ARVC: An Auto-Regressive Voice Conversion System Without Parallel Training Data
Zheng Lian 0004, Zhengqi Wen, Xinyong Zhou, Songbai Pu, Shengkai Zhang, Jianhua Tao 0001
INTERSPEECH2
2020 Spike-Triggered Non-Autoregressive Transformer for End-to-End Speech Recognition
abstract
Non-autoregressive transformer models have achieved extremely fast inference speed and comparable performance with autoregressive sequence-to-sequence models in neural machine translation.Most of the non-autoregressive transformers decode the target sequence from a predefined-length mask sequence.If the predefined length is too long, it will cause a lot of redundant calculations.If the predefined length is shorter than the length of the target sequence, it will hurt the performance of the model.To address this problem and improve the inference speed, we propose a spike-triggered non-autoregressive transformer model for end-to-end speech recognition, which introduces a CTC module to predict the length of the target sequence and accelerate the convergence.All the experiments are conducted on a public Chinese mandarin dataset AISHELL-1.The results show that the proposed model can accurately predict the length of the target sequence and achieve a competitive performance with the advanced transformers.What's more, the model even achieves a real-time factor of 0.0056, which exceeds all mainstream speech recognition models.
Zhengkun Tian, Jiangyan Yi, Jianhua Tao 0001, Ye Bai 0001, Shuai Zhang 0014, Zhengqi Wen
INTERSPEECH6
2020 Non-Autoregressive End-to-End TTS with Coarse-to-Fine Decoding
Tao Wang 0074, Xuefei Liu, Jianhua Tao 0001, Jiangyan Yi, Ruibo Fu, Zhengqi Wen
INTERSPEECH6
2020 Bi-Level Speaker Supervision for One-Shot Speech Synthesis
Tao Wang 0074, Jianhua Tao 0001, Ruibo Fu, Jiangyan Yi, Zhengqi Wen, Chunyu Qiang
INTERSPEECH5
2020 Spoken Content and Voice Factorization for Few-Shot Speaker Adaptation
Tao Wang 0074, Jianhua Tao 0001, Ruibo Fu, Jiangyan Yi, Zhengqi Wen, Rongxiu Zhong
INTERSPEECH5
2020 End-to-End Post-Filter for Speech Separation With Deep Attention Fusion Features
abstract
In this article, we propose an end-to-end post-filter method with deep attention fusion features for monaural speaker-independent speech separation. At first, a time-frequency domain speech separation method is applied as the pre-separation stage. The aim of pre-separation stage is to separate the mixture preliminarily. Although this stage can separate the mixture, it still contains the residual interference. In order to enhance the pre-separated speech and improve the separation performance further, the end-to-end post-filter (E2EPF) with deep attention fusion features is proposed. The E2EPF can make full use of the prior knowledge of the pre-separated speech, which contributes to speech separation. It is a fully convolutional speech separation network and uses the waveform as the input features. Firstly, the 1-D convolutional layer is utilized to extract the deep representation features for the mixture and pre-separated signals in the time domain. Secondly, to pay more attention to the outputs of the pre-separation stage, an attention module is applied to acquire deep attention fusion features, which are extracted by computing the similarity between the mixture and the pre-separated speech. These deep attention fusion features are conducive to reduce the interference and enhance the pre-separated speech. Finally, these features are sent to the post-filter to estimate each target signals. Experimental results on the WSJ0-2mix dataset show that the proposed method outperforms the state-of-the-art speech separation method. Compared with the pre-separation method, our proposed method can acquire 64.1%, 60.2%, 25.6% and 7.5% relative improvements in scale-invariant source-to-noise ratio (SI-SNR), the signal-to-distortion ratio (SDR), the perceptual evaluation of speech quality (PESQ) and the short-time objective intelligibility (STOI) measures, respectively.
Cunhang Fan, Jianhua Tao 0001, Bin Liu 0041, Jiangyan Yi, Zhengqi Wen, Xuefei Liu
IEEE ACM Trans. Audio Speech Lang. Process.5
2019 Phoneme Dependent Speaker Embedding and Model Factorization for Multi-speaker Speech Synthesis and Adaptation
abstract
This paper presents an architecture to perform speaker adaption in long short-term memory (LSTM) based Mandarin statistical parametric speech synthesis system. Compared with the conventional methods that focused on using fixed global speaker representations in utterance level for speaker recognition task, the proposed method extracts speaker representations in utterance and phoneme level, which can describe more pronunciation characteristics in phoneme level. And an attention mechanism is deployed to combine each level representations dynamically to train a task-specific phoneme dependent speaker embedding. To handle the unbalanced database and avoid over-fitting, the model is factored into an average model and an adaptation model and combined by an attention mechanism. We investigate the performance of speaker representations extracted by different methods. Experimental results confirm the adaptability of our proposed speaker embedding and model factorization structure. And listening tests demonstrate that our proposed method can achieve better adaptation performance than baselines in terms of naturalness and speaker similarity.
Ruibo Fu, Jianhua Tao 0001, Zhengqi Wen, Yibin Zheng
ICASSP3
2019 Learn Spelling from Teachers: Transferring Knowledge from Language Models to Sequence-to-Sequence Speech Recognition
abstract
Integrating an external language model into a sequence-tosequence speech recognition system is non-trivial.Previous works utilize linear interpolation or a fusion network to integrate external language models.However, these approaches introduce external components, and increase decoding computation.In this paper, we instead propose a knowledge distillation based training approach to integrating external language models into a sequence-to-sequence model.A recurrent neural network language model, which is trained on large scale external text, generates soft labels to guide the sequence-to-sequence model training.Thus, the language model plays the role of the teacher.This approach does not add any external component to the sequence-to-sequence model during testing.And this approach is flexible to be combined with shallow fusion technique together for decoding.The experiments are conducted on public Chinese datasets AISHELL-1 and CLMAD.Our approach achieves a character error rate of 9.3%, which is relatively reduced by 18.42% compared with the vanilla sequenceto-sequence model.
Ye Bai 0001, Jiangyan Yi, Jianhua Tao 0001, Zhengkun Tian, Zhengqi Wen
INTERSPEECH5
2019 A Time Delay Neural Network with Shared Weight Self-Attention for Small-Footprint Keyword Spotting
Ye Bai 0001, Jiangyan Yi, Jianhua Tao 0001, Zhengqi Wen, Zhengkun Tian, Chenghao Zhao, Cunhang Fan
INTERSPEECH4
2019 Discriminative Learning for Monaural Speech Separation Using Deep Embedding Features
abstract
Deep clustering (DC) and utterance-level permutation invariant training (uPIT) have been demonstrated promising for speakerindependent speech separation.DC is usually formulated as two-step processes: embedding learning and embedding clustering, which results in complex separation pipelines and a huge obstacle in directly optimizing the actual separation objectives.As for uPIT, it only minimizes the chosen permutation with the lowest mean square error, doesn't discriminate it with other permutations.In this paper, we propose a discriminative learning method for speaker-independent speech separation using deep embedding features.Firstly, a DC network is trained to extract deep embedding features, which contain each source's information and have an advantage in discriminating each target speakers.Then these features are used as the input for uPIT to directly separate the different sources.Finally, uPIT and DC are jointly trained, which directly optimizes the actual separation objectives.Moreover, in order to maximize the distance of each permutation, the discriminative learning is applied to fine tuning the whole model.Our experiments are conducted on WSJ0-2mix dataset.Experimental results show that the proposed models achieve better performances than DC and uPIT for speaker-independent speech separation.
Cunhang Fan, Bin Liu 0041, Jianhua Tao 0001, Jiangyan Yi, Zhengqi Wen
INTERSPEECH5
2019 Self-Attention Transducers for End-to-End Speech Recognition
abstract
Recurrent neural network transducers (RNN-T) have been successfully applied in end-to-end speech recognition. However, the recurrent structure makes it difficult for parallelization . In this paper, we propose a self-attention transducer (SA-T) for speech recognition. RNNs are replaced with self-attention blocks, which are powerful to model long-term dependencies inside sequences and able to be efficiently parallelized. Furthermore, a path-aware regularization is proposed to assist SA-T to learn alignments and improve the performance. Additionally, a chunk-flow mechanism is utilized to achieve online decoding. All experiments are conducted on a Mandarin Chinese dataset AISHELL-1. The results demonstrate that our proposed approach achieves a 21.3% relative reduction in character error rate compared with the baseline RNN-T. In addition, the SA-T with chunk-flow mechanism can perform online decoding with only a little degradation of the performance.
Zhengkun Tian, Jiangyan Yi, Jianhua Tao 0001, Ye Bai 0001, Zhengqi Wen
INTERSPEECH5
2019 Forward-Backward Decoding for Regularizing End-to-End TTS
abstract
Neural end-to-end TTS can generate very high-quality synthesized speech, and even close to human recording within similar domain text. However, it performs unsatisfactory when scaling it to challenging test sets. One concern is that the encoder-decoder with attention-based network adopts autoregressive generative sequence model with the limitation of exposure bias To address this issue, we propose two novel methods, which learn to predict future by improving agreement between forward and backward decoding sequence. The first one is achieved by introducing divergence regularization terms into model training objective to reduce the mismatch between two directional models, namely L2R and R2L (which generates targets from left-to-right and right-to-left, respectively). While the second one operates on decoder-level and exploits the future information during decoding. In addition, we employ a joint training strategy to allow forward and backward decoding to improve each other in an interactive process. Experimental results show our proposed methods especially the second one (bidirectional decoder regularization), leads a significantly improvement on both robustness and overall naturalness, as outperforming baseline (the revised version of Tacotron2) with a MOS gap of 0.14 in a challenging test, and achieving close to human quality (4.42 vs. 4.49 in MOS) on general test.
Yibin Zheng, Xi Wang 0016, Lei He 0005, Shifeng Pan, Frank K. Soong, Zhengqi Wen, Jianhua Tao 0001
INTERSPEECH6
2019 Language-Adversarial Transfer Learning for Low-Resource Speech Recognition
abstract
The acoustic model trained using the knowledge from the shared hidden layer (SHL) model outperforms the model trained only by using the target language, especially under low resource conditions. However, the shared features may contain some unnecessary language dependent information. It will degrade the performance of the target model. Therefore, this paper proposes language-adversarial transfer learning to alleviate this problem. Adversarial learning is used to ensure that the shared layers of the SHL-model can learn more language invariant features. Experiments are conducted on IARPA Babel datasets. The results show that the target model trained using the knowledge transferred from the adversarial SHL-model achieves up to 10.1% relative word error rate reduction when compared with the target model trained using the knowledge transferred from the SHL-model.
Jiangyan Yi, Jianhua Tao 0001, Zhengqi Wen, Ye Bai 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 Forward-Backward Decoding Sequence for Regularizing End-to-End TTS
abstract
Neural end-to-end TTS such as Tacotron like network can generate very high-quality synthesized speech, and even close to human recording for similar domain text. However, it performs unsatisfactory when scaling it to some challenging test sets. One concern is that the encoder-decoder with attention-based network adopts autoregressive generative sequence model with the limitation of “exposure bias”: errors made early could be quickly amplified, harming subsequent sequence generation. To address this issue, we propose two novel methods, which aim at predicting future by improving the agreement between forward and backward decoding sequence. The first one (denoted as MRBA) is achieved by adding divergence regularization terms to model training objective to maximize the agreement between two directional models, namely L2R (which generates targets from left-to-right) and R2L (which generates targets from right-to-left). While the second one (denoted as BDR) operates on decoder-level and exploits the future information during decoding. By introducing regularization term into the training objective of forward-backward decoders, the forward-decoder's hidden states are forced to be close to the backward-decoder's. Thus, the hidden representations of a unidirectional decoder are encouraged to embed some useful information about the future. Moreover, in order to make forward and backward decoding to improve each other in an interactive process, a joint training method is designed. Experimental results on both English and Mandarin dataset show that our proposed methods especially the second one (BDR), lead to a significantly improvement on both robustness and overall naturalness, as achieving obvious preference advantages in a challenging test, and achieving state-of-the-art performance (outperforming baseline “the revised version of Tacotron2” with a gap of 0.13 and 0.12 for English and Mandarin in MOS, respectively) on a general test.
Yibin Zheng, Jianhua Tao 0001, Zhengqi Wen, Jiangyan Yi
IEEE ACM Trans. Audio Speech Lang. Process.3
2018 Adversarial Multilingual Training for Low-Resource Speech Recognition
abstract
This paper proposes an adversarial multilingual training to train bottleneck (BN) networks for the target language. A parallel shared-exclusive model is also proposed to train the BN network. Adversarial training is used to ensure that the shared layers can learn language-invariant features. Experiments are conducted on IARPA Babel datasets. The results show that the proposed adversarial multilingual BN model outperforms the baseline BN model by up to 8.9% relative word error rate (WER) reduction. The results also show that the proposed parallel shared-exclusive model achieves up to 1.7% relative WER reduction when compared with the stacked share-exclusive model.
Jiangyan Yi, Jianhua Tao 0001, Zhengqi Wen, Ye Bai 0001
ICASSP3
2018 Transfer Learning Based Progressive Neural Networks for Acoustic Modeling in Statistical Parametric Speech Synthesis
Ruibo Fu, Jianhua Tao 0001, Yibin Zheng, Zhengqi Wen
INTERSPEECH4
2018 Deep Metric Learning for the Target Cost in Unit-Selection Speech Synthesizer
Ruibo Fu, Jianhua Tao 0001, Yibin Zheng, Zhengqi Wen
INTERSPEECH4
2018 On the Application and Compression of Deep Time Delay Neural Network for Embedded Statistical Parametric Speech Synthesis
Yibin Zheng, Jianhua Tao 0001, Zhengqi Wen, Ruibo Fu
INTERSPEECH3
2018 BLSTM-CRF Based End-to-End Prosodic Boundary Prediction with Context Sensitive Embeddings in a Text-to-Speech Front-End
Yibin Zheng, Jianhua Tao 0001, Zhengqi Wen, Ya Li 0001
INTERSPEECH3
2017 Distilling Knowledge from an Ensemble of Models for Punctuation Prediction
Jiangyan Yi, Jianhua Tao 0001, Zhengqi Wen, Ya Li 0001
INTERSPEECH3
2017 Investigating Efficient Feature Representation Methods and Training Objective for BLSTM-Based Phone Duration Prediction
Yibin Zheng, Jianhua Tao 0001, Zhengqi Wen, Ya Li 0001, Bin Liu 0041
INTERSPEECH3
2016 Long short term memory recurrent neural network based encoding method for emotion recognition in video
abstract
Human emotion is a temporally dynamic event which can be inferred from both audio and video feature sequences. In this paper we investigate the long short term memory recurrent neural network (LSTM-RNN) based encoding method for category emotion recognition in the video. LSTM-RNN is able to incorporate knowledge about how emotion evolves over long range successive frames and emotion clues from isolated frame. After encoding, each video clip can be represented by a vector for each input feature sequence. The vectors contain both frame level and sequence level emotion information. These vectors are then concatenated and fed into support vector machine (SVM) to get the final prediction result. Extensive evaluations on Emotion Challenge in the Wild (EmotiW2015) dataset show the efficiency of the proposed encoding method and competitive results are obtained. The final recognition accuracy achieves 46.38% for audio-video emotion recognition sub-challenge, where the challenge baseline is 39.33%.
Linlin Chao, Jianhua Tao 0001, Ya Li 0001, Zhengqi Wen
ICASSP5
2016 The Parameterized Phoneme Identity Feature as a Continuous Real-Valued Vector for Neural Network Based Speech Synthesis
Zhengqi Wen, Ya Li 0001, Jianhua Tao 0001
INTERSPEECH1
2016 Improving Prosodic Boundaries Prediction for Mandarin Speech Synthesis by Using Enhanced Embedding Feature and Model Fusion Approach
Yibin Zheng, Ya Li 0001, Zhengqi Wen, Xingguang Ding, Jianhua Tao 0001
INTERSPEECH3
2015 A novel method of artificial bandwidth extension using deep architecture
Bin Liu 0041, Jianhua Tao 0001, Zhengqi Wen, Ya Li 0001, Danish Bukhari
INTERSPEECH3
2015 Combining extreme learning machine and decision tree for duration prediction in HMM based speech synthesis
abstract
Hidden Markov Model (HMM) based speech synthesis using Decision Tree (DT) for duration prediction is known to produce over-averaged rhythm. To alleviate this problem, this paper proposes a two level duration prediction method together with outlier removal. This method takes advantages of accurate regression capability by Extreme Learning Machine (ELM) for phone level duration prediction, and the capability of distributing state durations by DT for state level duration prediction. Experimental results showed that the method decreased RMSE of phone duration, increased the fluctuation of syllable duration, and achieved 63.75% in preference evaluation. Furthermore, this method does not incur laborious manual alignment on training corpus.
Zhengqi Wen, Jianhua Tao 0001
INTERSPEECH3
2014 A novel hybrid mandarin speech synthesis system using different base units for model training and concatenation
abstract
The hybrid speech synthesis system, which uses the acoustic model trained according to the criterion of Maximum Likelihood to select the proper candidates from the corpus, has become a hot topic in recent days. For this hybrid system, the performance is affected by the size of the base training unit and the base candidate unit. Most of existed hybrid systems use the same kind of base unit such as syllable or phone for both model training and concatenation. In Mandarin, initials and finals form the fundamental elements of pronunciation, and are always chosen as the base training unit for statistical parametric TTS system. In this paper a new hybrid Mandarin TTS system is proposed, which uses initial/final for model training and syllable for concatenation. Objective and subjective evaluations are conducted and the comparison results show that the hybrid system we proposed outperforms the traditional systems which use the same base unit for both processes with 4000 and 6000 sentences' corpus.
Jianhua Tao 0001, Ya Li 0001, Zhengqi Wen
ICASSP4
2014 A hierarchical viterbi algorithm for Mandarin hybrid speech synthesis system
abstract
The hybrid speech synthesis system, which combines the hidden Markov model and unit selection method, has become an additional main stream in state-of-the-art TTS systems. However, traditional Viterbi algorithm is based on global minimization of a cost function and the procedure can end up selecting some poor-quality units with larger local errors, which can hardly be tolerated by the listeners. In Mandarin and many other languages, the naturalness of the region of consecutive voiced speech segments (CVS) is more essential to the overall quality of the synthetic speech. Consequently, in this paper, we proposed to use a hierarchical Viterbi algorithm which involves two rounds of Viterbi search: one is for the sub-paths in the CVS regions; the other is for the utterance path connecting all the sub-paths. In the proposed technique, we defined CVS Region as a region which is formed by two or more voiced phones, and whose observation of pitch has a continuous value. Subjective evaluations suggest that the use of hierarchical Viterbi algorithm in the Mandarin hybrid speech synthesis system outperforms the use of traditional algorithm in both the naturalness and speech quality of synthetic speech.
Zhengqi Wen, Jianhua Tao 0001, Ya Li 0001, Xiaoyan Lou
INTERSPEECH2
2012 Pitch-Scaled Analysis based Residual Reconstruction for Speech Analysis and Synthesis
abstract
The typical problem in LPC-like vocoder is buzzing sound which is mainly due to the simple pulse train or noise excitation model. One way to improve it is to reconstruct the residual obtained from inverse filtering. So a new parametric representation of speech based on pitch-scaled analysis is proposed in this paper. Pitch-scaled analysis is used to extract the periodic spectrum of residual with half pitch period length. Then these periodic spectrums are de-correlated by principal component analysis (PCA) to reduce their dimension. Aperiodic measure is defined as the harmonic-to-noise ratio in the frequency domain where voicing cut-off frequency (VCO) is used to control the smoothness of aperiodicity. Periodic spectrum and aperiodic measure together with F0 are indicated as excitation parameters in the proposed LPC vocoder. Experimental results show that this proposed vocoder can get a mean opinion score (MOS) of 4.1 for a female voice before dimensionality reduction and keep the high-quality property after parameter compression.
Zhengqi Wen, Hideki Kawahara, Jianhua Tao 0001
INTERSPEECH1
2012 Amplitude Spectrum based Excitation Model for HMM-based Speech Synthesis
abstract
This paper describes an excitation model based on amplitude spectrum for hidden Markov model (HMM)-based speech synthesis system (HTS). Residual signal obtained from inverse filtering is decomposed into periodic and aperiodic spectrums in frequency domain. Amplitude spectrum of half pitch period length is reserved as periodic component in synthesis stage and zero-phase criterion and pitch synchronous overlap add method (PSOLA) are adopted to reconstruct the residual signal. Before integrating this excitation model into HTS, these periodic spectrums are normalized and Linde-Buzo-Gray (LBG) algorithm is adopted to construct codebooks for every Mandarin final1. Then index parameters from these codebooks which are indicated as excitation information are taken into HTS training together with spectral, F0 and aperiodic parameters. Listening test showed that for female voice the analysis-synthesis result of the vocoder based on proposed excitation model is comparable with that of STRAIGHT and when integrating into HTS, the quality of generated speech is also improved.
Zhengqi Wen, Jianhua Tao 0001
INTERSPEECH1
2011 Inverse Filtering Based Harmonic Plus Noise Excitation Model for HMM-Based Speech Synthesis
abstract
In this paper, a new Voicing Cut-Off Frequency (VCO) estimation method based on inverse filtering is presented. The spectrum of residual signal got from inverse filtering is split into sub-bands which are clustered into two classes by using K-means algorithm. And then, the Viterbi algorithm is used to search a smoothed VCO contour. Based on this new VCO estimation method, an adaptation of Harmonic Noise Model is also proposed to reconstruct the residual signal with both harmonic and noise components. The proposed excitation model can reduce the buzziness of speech generated by normal vocoders using simple pulse train, and has been integrated into a HMM-based speech synthesis system (HTS). The listening test showed that the HTS with our new method gives better quality of synthesized speech than the traditional HTS which only uses simple pulse train excitation model.
Zhengqi Wen, Jianhua Tao 0001
INTERSPEECH1