Cunhang Fan

dblp:240/7820 · DBLP profile ↗
← Back
57ranked-venue papers
24as first author
48since 2021 · last 2026
0000-0001-6318-8803ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 40 · 17 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 14 first-author · 29 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Trainable EEG Interpolation and Structure-Sharing Dual-Path Encoders for Brain-Assisted Target Speaker Extraction
abstract
Brain-assisted target speaker extraction (TSE) isolates a target speaker's voice from a mixture by leveraging task-specific representations in Electroencephalogram (EEG) signals. However, existing methods rely on fixed interpolation for EEG-audio alignment, introducing redundant computations. They also employ single-path encoders that extract only target-relevant features while neglecting complementary, irrelevant ones, limiting discriminability. To address these limitations, this paper proposes a Trainable EEG Interpolation and Structure-sharing Dual-path Encoders network (TIDENet). The proposed Trainable EEG Interpolation (TEI) uses a neural network module to leverage cross-sample EEG information during resampling by parameters updating, thereby overcoming the limitations of fixed interpolation. The Structure-sharing Dual-path Encoders (SSDPE) extend existing speech and EEG encoders by introducing dual paths that separately process features relevant and irrelevant to the target speaker and incorporates interactive fusion between them, which enhances the encoder's ability to capture task-relevant information. Experimental results on public datasets demonstrate that TIDENet achieves relative improvements of up to 20.47%, 22.22%, 2.91%, 6.20%, and 15.84% in signal-to-distortion ratio (SDR), scale-invariant SDR (SI-SDR), short-time objective intelligibility (STOI), extended STOI (ESTOI), and perceptual evaluation of speech quality (PESQ), respectively, compared to the state-of-the-art. These significant gains validate the effectiveness of the proposed TEI method and SSDPE architecture.
Zhao Lv, Youdian Gao, Ruibo Fu, Cunhang Fan
AAAI7
2026 ReFL: Reflective Feedback Learning for Hallucination Detection of Large Language Models
abstract
Cunhang Fan, Jun Zhang, Xue Zhang, Shuai Zhang, Zhao Lv, Jianhua Tao, Zhengqi Wen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Cunhang Fan, Shuai Zhang 0014, Zhao Lv, Jianhua Tao 0001, Zhengqi Wen
ACL (1)1
2026 Two-stream attentive spatial-temporal graph convolutional network for P300 detection in brain-computer interface
Jincen Wang, Yan Zhao 0037, Cunhang Fan, Yong Li 0032, Fan Liu 0003, Hailun Lian, Cheng Lu 0005
Expert Syst. Appl.3
2025 Region-Based Optimization in Continual Learning for Audio Deepfake Detection
abstract
Rapid advancements in speech synthesis and voice conversion bring convenience but also new security risks, creating an urgent need for effective audio deepfake detection. Although current models perform well, their effectiveness diminishes when confronted with the diverse and evolving nature of real-world deepfakes. To address this issue, we propose a continual learning method named Region-Based Optimization (RegO) for audio deepfake detection. Specifically, we use the Fisher information matrix to measure important neuron regions for real and fake audio detection, dividing them into four regions. First, we directly fine-tune the less important regions to quickly adapt to new tasks. Next, we apply gradient optimization in parallel for regions important only to real audio detection, and in orthogonal directions for regions important only to fake audio detection. For regions that are important to both, we use sample proportion-based adaptive gradient optimization. This region-adaptive optimization ensures an appropriate trade-off between memory stability and learning plasticity. Additionally, to address the increase of redundant neurons from old tasks, we further introduce the Ebbinghaus forgetting mechanism to release them, thereby promoting the model’s ability to learn more generalized discriminative features. Experimental results show our method achieves a 21.3 percent improvement in EER over the state-of-the-art continual learning approach RWM for audio deepfake detection. Moreover, the effectiveness of RegO extends beyond the audio deepfake detection domain, showing potential significance in other tasks, such as image recognition.
Yujie Chen 0006, Jiangyan Yi, Cunhang Fan, Jianhua Tao 0001, Yong Ren 0006, Siding Zeng, Chu Yuan Zhang, Xinrui Yan, Jun Xue 0001, Chenglong Wang 0001, Zhao Lv, Xiaohui Zhang 0006
AAAI3
2025 BSDB-Net: Band-Split Dual-Branch Network with Selective State Spaces Mechanism for Monaural Speech Enhancement
abstract
Although the complex spectrum-based speech enhancement (SE) methods have achieved significant performance, coupling amplitude and phase can lead to a compensation effect, where amplitude information is sacrificed to compensate for the phase that is harmful to SE. In addition, to further improve the performance of SE, many modules are stacked onto SE, resulting in increased model complexity that limits the application of SE. To address these problems, we proposed a dual-path network based on compressed frequency using Mamba. First, we extract amplitude and phase information through parallel dual branches. This approach leverages structured complex spectra to implicitly capture phase information and solves the compensation effect by decoupling amplitude and phase, and the network incorporates an interaction module to suppress unnecessary parts and recover missing components from the other branch. Second, to reduce network complexity, the network introduces a band-split strategy to compress the frequency dimension. To further reduce complexity while maintaining good performance, we designed a Mamba-based module that models the time and frequency dimensions under linear complexity. Finally, compared to baselines, our model achieves an average 8.3 times reduction in computational complexity while maintaining superior performance. Furthermore, it achieves a 25 times reduction in complexity compared to transformer-based models.
Cunhang Fan, Enrui Liu, Andong Li, Jianhua Tao 0001, Jian Zhou 0006, Chengshi Zheng, Zhao Lv
AAAI1
2025 Improved Feature Extraction Network for Neuro-Oriented Target Speaker Extraction
abstract
The recent rapid development of auditory attention decoding (AAD) offers the possibility of using electroencephalography (EEG) as auxiliary information for target speaker extraction. However, effectively modeling long sequences of speech and resolving the identity of the target speaker from EEG signals remains a major challenge. In this paper, an improved feature extraction network (IFENet) is proposed for neuro-oriented target speaker extraction, which mainly consists of a speech encoder with dual-path Mamba and an EEG encoder with Kolmogorov-Arnold Networks (KAN). We propose SpeechBiMamba, which makes use of dual-path Mamba in modeling local and global speech sequences to extract speech features. In addition, we propose EEGKAN to effectively extract EEG features that are closely related to the auditory stimuli and locate the target speaker through the subject’s attention information. Experiments on the KUL and AVED datasets show that IFENet outperforms the state-of-the-art model, achieving 36% and 29% relative improvements in terms of scale-invariant signal-to-distortion ratio (SI-SDR) under an open evaluation condition.
Cunhang Fan, Youdian Gao, Zexu Pan, Jie Zhang 0042, Zhao Lv
ICASSP1
2025 SSM2Mel: State Space Model to Reconstruct Mel Spectrogram from the EEG
abstract
Decoding speech from brain signals is a challenging research problem that holds significant importance for studying speech processing in the brain. Although breakthroughs have been made in reconstructing the mel spectrograms of audio stimuli perceived by subjects at the word or letter level using non-invasive electroencephalography (EEG), there is still a critical gap in precisely reconstructing continuous speech features, especially at the minute level. To address this issue, this paper proposes a State Space Model (SSM) to reconstruct the mel spectrogram of continuous speech from EEG, named SSM2Mel. This model introduces a novel Mamba module to effectively model the long sequence of EEG signals for imagined speech. In the SSM2Mel model, the S4-UNet structure is used to enhance the extraction of local features of EEG signals, and the Embedding Strength Modulator (ESM) module is used to incorporate subject-specific information. Experimental results show that our model achieves a Pearson correlation of 0.069 on the SparrKULee dataset, which is a 38% improvement over the previous baseline.
Cunhang Fan, Zexu Pan, Zhao Lv
ICASSP1
2025 LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech Enhancement
abstract
Speech enhancement (SE) aims to extract the clean waveform from noise-contaminated measurements to improve the speech quality and intelligibility. Although learning-based methods can perform much better than traditional counterparts, the large computational complexity and model size heavily limit the deployment on latency-sensitive and low-resource edge devices. In this work, we propose a lightweight SE network (LiSenNet) for real-time applications. We design sub-band downsampling and upsampling blocks and a dual-path recurrent module to capture band-aware features and time-frequency patterns, respectively. A noise detector is developed to detect noisy regions in order to perform SE adaptively and save computational costs. Compared to recent higher-resource-dependent baseline models, the proposed LiSenNet can achieve a competitive performance with only 37k parameters (half of the state-of-the-art model) and 56M multiply-accumulate (MAC) operations per second.
Haoyin Yan, Jie Zhang 0042, Cunhang Fan, Yeping Zhou, Peiqi Liu
ICASSP3
2025 M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction
abstract
The brain-assisted target speaker extraction (TSE) aims to extract the attended speech from mixed speech by utilizing the brain neural activities, for example Electroencephalography (EEG). However, existing models overlook the issue of temporal misalignment between speech and EEG modalities, which hampers TSE performance. In addition, the speech encoder in current models typically uses basic temporal operations (e.g., one-dimensional convolution), which are unable to effectively extract target speaker information. To address these issues, this paper proposes a multi-scale and multi-modal alignment network (M3ANet) for brain-assisted TSE. Specifically, to eliminate the temporal inconsistency between EEG and speech modalities, the modal alignment module that uses a contrastive learning strategy is applied to align the temporal features of both modalities. Additionally, to fully extract speech information, multi-scale convolutions with GroupMamba modules are used as the speech encoder, which scans speech features at each scale from different directions, enabling the model to capture deep sequence information. Experimental results on three publicly available datasets show that the proposed model outperforms current state-of-the-art methods across various evaluation metrics, highlighting the effectiveness of our proposed method. The source code is available at: https://github.com/fchest/M3ANet.
Cunhang Fan, Jian Zhou 0006, Zexu Pan, Youdian Gao, Xiaoke Yang, Zhengqi Wen, Zhao Lv
IJCAI1
2025 ListenNet: A Lightweight Spatio-Temporal Enhancement Nested Network for Auditory Attention Detection
abstract
Auditory attention detection (AAD) aims to identify the direction of the attended speaker in multi-speaker environments from brain signals, such as Electroencephalography (EEG) signals. However, existing EEG-based AAD methods overlook the spatio-temporal dependencies of EEG signals, limiting their decoding and generalization abilities. To address these issues, this paper proposes a Lightweight Spatio-Temporal Enhancement Nested Network (ListenNet) for AAD. The ListenNet has three key components: Spatio-temporal Dependency Encoder (STDE), Multi-scale Temporal Enhancement (MSTE), and Cross-Nested Attention (CNA). The STDE reconstructs dependencies between consecutive time windows across channels, improving the robustness of dynamic pattern extraction. The MSTE captures temporal features at multiple scales to represent both fine-grained and long-range temporal patterns. In addition, the CNA integrates hierarchical features more effectively through novel dynamic attention mechanisms to capture deep spatio-temporal correlations. Experimental results on three public datasets demonstrate the superiority of ListenNet over state-of-the-art methods in both subject-dependent and challenging subject-independent settings, while reducing the trainable parameter count by approximately 7 times. Code is available at:https://github.com/fchest/ListenNet.
Cunhang Fan, Xiaoke Yang, Jian Zhou 0006, Zhao Lv
IJCAI1
2025 MHANet: Multi-scale Hybrid Attention Network for Auditory Attention Detection
abstract
Auditory attention detection (AAD) aims to detect the target speaker in a multi-talker environment from brain signals, such as electroencephalography (EEG), which has made great progress. However, most AAD methods solely utilize attention mechanisms sequentially and overlook valuable multi-scale contextual information within EEG signals, limiting their ability to capture long-short range spatiotemporal dependencies simultaneously. To address these issues, this paper proposes a multi-scale hybrid attention network (MHANet) for AAD, which consists of the multi-scale hybrid attention (MHA) module and the spatiotemporal convolution (STC) module. Specifically, MHA combines channel attention and multi-scale temporal and global attention mechanisms. This effectively extracts multi-scale temporal patterns within EEG signals and captures long-short range spatiotemporal dependencies simultaneously. To further improve the performance of AAD, STC utilizes temporal and spatial convolutions to aggregate expressive spatiotemporal representations. Experimental results show that the proposed MHANet achieves state-of-the-art performance with fewer trainable parameters across three datasets, 3 times lower than that of the most advanced model. Code is available at: https://github.com/fchest/MHANet.
Cunhang Fan, Xiaoke Yang, Jian Zhou 0006, Zhao Lv
IJCAI2
2025 ID-RemovalNet: Identity Removal Network for EEG Privacy Protection with Enhancing Decoding Tasks
abstract
Electroencephalogram (EEG) contains not only decoding task information but also personal identity privacy information. If it is stolen or attacked, the user's brain-computer interaction behavior may be maliciously manipulated. Existing EEG identity privacy protection generally adopts generative or adding tiny perturbation methods, which can protect the identity privacy in EEG signals to some extent. However, these methods also damage the performance of decoding task. In order to solve these problems, this paper proposes an identity removal network (ID-RemovalNet) to achieve EEG privacy protection while improving the classification accuracy of decoding task. Firstly, an identity decorrelation separation module is constructed to accurately remove the identity features to achieve privacy protection while reducing the interference with the task decoding features. Secondly, a multi-domain multi-level fusion feature extraction module is designed to extract the high-quality EEG time-frequency features. Finally, the feature enhancement module is used to compensate for the loss of task decoding features and excitation of dominant feature selection during identity feature removal. The experimental results show that ID-RemoveNet removes identity information to 0.43% on four EEG datasets with two different paradigms, and significantly improves the EEG task decoding accuracy by 3.28%, and achieves the state-of-the-art performance in cross-subject EEG experiment.
Jie Ruan, Cunhang Fan, Yingfan Cheng, Zhao Lv
IJCAI3
2025 REB-former: RWKV-enhanced E-branchformer for Speech Recognition
Wang Xiang, Jian Zhou 0006, Cunhang Fan, Zhao Lv
INTERSPEECH4
2025 SSF-DST: A Spectro-Spatial Features Enhanced Deep Spatiotemporal Network for EEG-Based Auditory Attention Detection
Xiaoke Yang, Jian Zhou 0006, Zhao Lv, Cunhang Fan
INTERSPEECH6
2025 DHGCN: Dual HyperGraph Convolutional Network for EEG-Based Auditory Attention Detection
abstract
Auditory attention detection (AAD) aims to identify the attended speaker in multi-talker environments by analyzing brain activity recorded through neural monitoring techniques. Recent AAD approaches have achieved great progress in improving detection accuracy. However, they still face challenges in capturing complex spatio-temporal dependencies and high-order nonlinear relationships across brain regions. To address these challenges, this paper proposes DHGCN, a dual hypergraph convolutional network that integrates a hypergraph modeling module, a dual-branch hypergraph learning (DHGL) module, and a feature fusion module. Specifically, the hypergraph modeling module constructs spatial and temporal hypergraphs from EEG signals, enabling the representation of high-order relationships among channels and time points. The DHGL module comprises two parallel branches: a spatial branch that learns high-order spatial dependencies across EEG channels, and a temporal branch that captures complex temporal dependencies. Each branch uses its corresponding hypergraph structure, which is established during the modeling phase. The feature fusion module then aggregates spatial and temporal representations from both branches to support robust auditory attention classification. Extensive experiments on multiple benchmark datasets demonstrate that DHGCN consistently outperforms state-of-the-art AAD models. It achieves superior classification performance while reducing the trainable parameters count by over 50% compared to the state-of-the-art models. Code is available at: https://github.com/nobody1219/DHGCN.git.
Jian Zhou 0006, Yingjie Xie, Cunhang Fan, Zhao Lv
ACM Multimedia3
2025 DMF2Mel: A Dynamic Multiscale Fusion Network for EEG-Driven Mel Spectrogram Reconstruction
abstract
Decoding speech from brain signals is a challenging research problem. Although existing technologies have made progress in reconstructing the mel spectrograms of auditory stimuli at the word or letter level, there remain core challenges in the precise reconstruction of minute-level continuous imagined speech: traditional models struggle to balance the efficiency of temporal dependency modeling and information retention in long-sequence decoding. To address this issue, this paper proposes the Dynamic Multiscale Fusion Network (DMF2Mel), which consists of four core components: the Dynamic Contrastive Feature Aggregation Module (DC-FAM), the Hierarchical Attention-Guided Multi-Scale Network (HAMS-Net), the SplineMap attention mechanism, and the bidirectional state space module (convMamba). Specifically, the DC-FAM separates speech-related ''foreground features'' from noisy ''background features'' through local convolution and global attention mechanisms, effectively suppressing interference and enhancing the representation of transient signals. HAMS-Net, based on the U-Net framework, achieves cross-scale fusion of high-level semantics and low-level details. The SplineMap attention mechanism integrates the Adaptive Gated Kolmogorov-Arnold Network (AGKAN) to combine global context modeling with spline-based local fitting. The convMamba captures long-range temporal dependencies with linear complexity and enhances nonlinear dynamic modeling capabilities. Results on the SparrKULee dataset show that DMF2Mel achieves a Pearson correlation coefficient of 0.074 in mel spectrogram reconstruction for known subjects (a 48% improvement over the baseline) and 0.048 for unknown subjects (a 35% improvement over the baseline).Code is available at: https://github.com/fchest/DMF2Mel.
Cunhang Fan, Enrui Liu, Gangming Zhao, Zhao Lv
ACM Multimedia1
2025 Multi-Level Contrastive Learning: Hierarchical Alleviation of Heterogeneity in Multimodal Sentiment Analysis
abstract
Recently, multimodal fusion efforts have achieved remarkable success in Multimodal Sentiment Analysis (MSA). However, most of the existing methods are based on model-level fusion, and the challenge of heterogeneity between modalities is not well resolved. Heterogeneity lies in the different feature distributions and distinct representation spaces among different modalities. To mitigate this problem, we propose that fusion is a progressive process, and we introduce a novel multi-level contrastive learning and multi-layer convolution fusion (MCL-MCF) method for MSA. Due to the relationships among multimodal data, the fusion process that involves single-modal to single-modal, single-modal to bimodal or trimodal, and higher-level fused modality semantic consistency is divided into three levels. The first-level contrast learning alleviates heterogeneity between unimodal modalities at the early level of multimodal feature fusion. The second-level contrast learning mitigates heterogeneity between unimodal and fused modalities. At the third level, we introduce a tensor convolution fusion (TCF) module that extracts high-level semantic features from the fused modalities and mitigates heterogeneity at the higher feature level through contrastive learning. To simulate fusion as a progressive process, MCF is proposed to fuse shallow and deep features to model complex relationships among modalities. Experiments on three public datasets show our approach's state-of-the-art performance.
Cunhang Fan, Kang Zhu, Jianhua Tao 0001, Guofeng Yi, Jun Xue 0001, Zhao Lv
IEEE Trans. Affect. Comput.1
2024 Progressive Distillation Based on Masked Generation Feature Method for Knowledge Graph Completion
abstract
In recent years, knowledge graph completion (KGC) models based on pre-trained language model (PLM) have shown promising results. However, the large number of parameters and high computational cost of PLM models pose challenges for their application in downstream tasks. This paper proposes a progressive distillation method based on masked generation features for KGC task, aiming to significantly reduce the complexity of pre-trained models. Specifically, we perform pre-distillation on PLM to obtain high-quality teacher models, and compress the PLM network to obtain multi-grade student models. However, traditional feature distillation suffers from the limitation of having a single representation of information in teacher models. To solve this problem, we propose masked generation of teacher-student features, which contain richer representation information. Furthermore, there is a significant gap in representation ability between teacher and student. Therefore, we design a progressive distillation method to distill student models at each grade level, enabling efficient knowledge transfer from teachers to students. The experimental results demonstrate that the model in the pre-distillation stage surpasses the existing state-of-the-art methods. Furthermore, in the progressive distillation stage, the model significantly reduces the model parameters while maintaining a certain level of performance. Specifically, the model parameters of the lower-grade student model are reduced by 56.7\% compared to the baseline.
Cunhang Fan, Yujie Chen 0006, Jun Xue 0001, Yonghui Kong, Jianhua Tao 0001, Zhao Lv
AAAI1
2024 Pruning via Merging: Compressing LLMs via Manifold Alignment Based Layer Merging
abstract
Deyuan Liu, Zhanyue Qin, Hairu Wang, Zhao Yang, Zecheng Wang, Fangying Rong, Qingbin Liu, Yanchao Hao, Bo Li, Xi Chen, Cunhang Fan, Zhao Lv, Dianhui Chu, Zhiying Tu, Dianbo Sui. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Deyuan Liu, Zhanyue Qin, Hairu Wang 0002, Zhao Yang 0004, Zecheng Wang, Fangying Rong, Qingbin Liu, Yanchao Hao, Xi Chen 0003, Cunhang Fan, Zhao Lv, Zhiying Tu, Dianbo Sui
EMNLP11
2024 UNO Arena for Evaluating Sequential Decision-Making Capability of Large Language Models
abstract
Zhanyue Qin, Haochuan Wang, Deyuan Liu, Ziyang Song, Cunhang Fan, Zhao Lv, Jinlin Wu, Zhen Lei, Zhiying Tu, Dianhui Chu, Xiaoyan Yu, Dianbo Sui. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Zhanyue Qin, Deyuan Liu, Cunhang Fan, Zhao Lv, Jinlin Wu, Zhen Lei 0001, Zhiying Tu, Dianbo Sui
EMNLP5
2024 Dual-View Multimodal Interaction in Multimodal Sentiment Analysis
abstract
Outstanding performance in sentiment analysis not only relies on the design of sophisticated fusion methods but also on the crucial step of designing excellent modal interaction methods. To the best of our knowledge, there are few methods addressing the capture of multimodal spatial features. Majority of feature interactions have been primarily focused on temporal aspects, with less attention given to the combined spatiotemporal feature interaction (SFI). In this paper, we design a dual-view multimodal interaction method, named DVMI, primarily consisting of two parts. In the first part, a triangular convolutional module is proposed for ample temporal interaction between modalities, implicit local and global SFI, and capturing global spatial representations. Building upon the foundation laid in the first part, the second part employs an attention mechanism for explicit global SFI. To demonstrate the effectiveness of the DVMI framework,we conduct extensive experiments on three datasets, achieving state-of-the-art experimental results.
Kang Zhu, Cunhang Fan, Jianhua Tao 0001, Jun Xue 0001, Xuefei Liu, Zhengqi Wen, Zhao Lv
ICME2
2024 DBPNet: Dual-Branch Parallel Network with Temporal-Frequency Fusion for Auditory Attention Detection
Qinke Ni, Cunhang Fan, Shengbing Pei, Zhao Lv
IJCAI3
2024 A Debiased Domain Adaptation Framework with Minimum Class Confusion for Motor Imagery Decoding
abstract
Recently, motor imagery decoding technology based on electroencephalogram (EEG) signals has made significant progress. However, there are still challenges in adapting to new sessions, mainly due to changes in data distribution between different sessions and confusion problems caused by similar oscillation patterns in various categories of EEG signals. To address these problems, this paper proposes a debiased domain adaptation framework with minimum class confusion to learn unbiased representations in motor imagery tasks. Specially, unlike the feature alignment and adversarial training methods, we explore the class predictions for domain adaptation, applying the minimum class confusion loss criterion in the target domain to reduce inter-class confusion. This approach aims to alleviate the bias issues inherent in classifiers trained on the source domain for making predictions in the target domain, achieving class-level alignment. Consequently, it enhances the model’s ability to adapt to the data distribution of the target domain. Experimental results on two public EEG datasets (BCI Competition IV datasets IIa and IIb) show that the method for cross-session decoding is significantly improved compared to the baseline, with average classification accuracy reaching 81.01% and 82.92%, respectively.
Cunhang Fan, Zhen Chen 0022, Xun Song, Jun Xue 0001, Ping Li 0020, Zhao Lv
IJCNN1
2024 RawBMamba: End-to-End Bidirectional State Space Model for Audio Deepfake Detection
abstract
Fake artefacts for discriminating between bonafide and fake audio can exist in both short-and long-range segments.Therefore, combining local and global feature information can effectively discriminate between bonafide and fake audio.This paper proposes an end-to-end bidirectional state space model, named RawBMamba, to capture both short-and long-range discriminative information for audio deepfake detection.Specifically, we use sinc Layer and multiple convolutional layers to capture short-range features, and then design a bidirectional Mamba to address Mamba's unidirectional modelling problem and further capture long-range feature information.Moreover, we develop a bidirectional fusion module to integrate embeddings, enhancing audio context representation and combining shortand long-range information.The results show that our proposed RawBMamba achieves a 34.1% improvement over Rawformer on ATSVspoof2021 LA dataset, and demonstrates competitive performance on other datasets.Codes will be released on https://github.com/cyjie429/RawBMamba.
Yujie Chen 0006, Jiangyan Yi, Jun Xue 0001, Chenglong Wang 0001, Xiaohui Zhang 0006, Shunbo Dong, Siding Zeng, Jianhua Tao 0001, Zhao Lv, Cunhang Fan
INTERSPEECH10
2024 Frequency-mix Knowledge Distillation for Fake Speech Detection
Cunhang Fan, Shunbo Dong, Jun Xue 0001, Yujie Chen 0006, Jiangyan Yi, Zhao Lv
INTERSPEECH1
2024 Prompt Link Multimodal Fusion in Multimodal Sentiment Analysis
Kang Zhu, Cunhang Fan, Jianhua Tao 0001, Zhao Lv
INTERSPEECH2
2024 MSFNet: Multi-Scale Fusion Network for Brain-Controlled Speaker Extraction
abstract
Speaker extraction aims to selectively extract the target speaker from the multi-talker environment under the guidance of auxiliary reference. Recent studies have shown that the attended speaker's information can be decoded by the auditory attention decoding from the listener's brain activity. However, how to more effectively utilize the common information about the target speaker contained in both electroencephalography (EEG) and speech is still an unresolved problem. In this paper, we propose a multi-scale fusion network (MSFNet) for brain-controlled speaker extraction, which utilizes the EEG recorded from the listener to extract the target speech. In order to make full use of the speech information, the mixed speech is encoded with multiple time scales so that the multi-scale embeddings are acquired. In addition, to effectively extract the non-Euclidean data of EEG, the graph convolutional networks are used as the EEG encoder. Finally, these multi-scale embeddings are separately fused with the EEG features. To facilitate research related to auditory attention decoding and further validate the effectiveness of the proposed method, we also construct the AVED dataset, a new EEG-Audio dataset. Experimental results on both the public Cocktail Party dataset and the newly proposed AVED dataset in this paper show that our MSFNet model significantly outperforms the state-of-the-art method in certain objective evaluation metrics.
Cunhang Fan, Wang Xiang, Jianhua Tao 0001, Jiangyan Yi, Dianbo Sui, Zhao Lv
ACM Multimedia1
2024 DARNet: Dual Attention Refinement Network with Spatiotemporal Construction for Auditory Attention Detection
abstract
At a cocktail party, humans exhibit an impressive ability to direct their attention. The auditory attention detection (AAD) approach seeks to identify the attended speaker by analyzing brain signals, such as EEG signals. However, current AAD algorithms overlook the spatial distribution information within EEG signals and lack the ability to capture long-range latent dependencies, limiting the model's ability to decode brain activity. To address these issues, this paper proposes a dual attention refinement network with spatiotemporal construction for AAD, named DARNet, which consists of the spatiotemporal construction module, dual attention refinement module, and feature fusion \& classifier module. Specifically, the spatiotemporal construction module aims to construct more expressive spatiotemporal feature representations, by capturing the spatial distribution characteristics of EEG signals. The dual attention refinement module aims to extract different levels of temporal patterns in EEG signals and enhance the model's ability to capture long-range latent dependencies. The feature fusion \& classifier module aims to aggregate temporal patterns and dependencies from different levels and obtain the final classification results. The experimental results indicate that DARNet achieved excellent classification performance, particularly under short decision windows. While maintaining excellent classification performance, DARNet significantly reduces the number of required parameters. Compared to the state-of-the-art models, DARNet reduces the parameter count by 91\%. Code is available at: https://github.com/fchest/DARNet.git.
Cunhang Fan, Xiaoke Yang, Jianhua Tao 0001, Zhao Lv
NeurIPS2
2024 Light-weight residual convolution-based capsule network for EEG emotion recognition
Cunhang Fan, Jinqin Wang, Xiaoke Yang, Guanxiong Pei, Taihao Li, Zhao Lv
Adv. Eng. Informatics1
2024 DuaPIN: Auxiliary task enhanced dual path interaction network for civil court view generation
Nayu Liu, Yiquan Wu 0001, Kaiwen Wei, Cunhang Fan
Knowl. Based Syst.5
2024 VLP2MSA: Expanding vision-language pre-training to multimodal sentiment analysis
Guofeng Yi, Cunhang Fan, Kang Zhu, Zhao Lv, Shan Liang 0007, Zhengqi Wen, Guanxiong Pei, Taihao Li, Jianhua Tao 0001
Knowl. Based Syst.2
2024 Spatial reconstructed local attention Res2Net with F0 subband for fake speech detection
Cunhang Fan, Jun Xue 0001, Jianhua Tao 0001, Jiangyan Yi, Chenglong Wang 0001, Chengshi Zheng, Zhao Lv
Neural Networks1
2024 DGSD: Dynamical graph self-distillation for EEG-based auditory spatial attention detection
Cunhang Fan, Jun Xue 0001, Jianhua Tao 0001, Jiangyan Yi, Zhao Lv, Xiaopei Wu
Neural Networks1
2024 Multimodal Cross-Lingual Summarization for Videos: A Revisit in Knowledge Distillation Induced Triple-Stage Training Method
abstract
Multimodal summarization (MS) for videos aims to generate summaries from multi-source information (e.g., video and text transcript), showing promising progress recently. However, existing works are limited to monolingual scenarios, neglecting non-native viewers' needs to understand videos in other languages. It stimulates us to introduce multimodal cross-lingual summarization for videos (MCLS), which aims to generate cross-lingual summaries from multimodal input of videos. Considering the challenge of high annotation cost and resource constraints in MCLS, we propose a knowledge distillation (KD) induced triple-stage training method to assist MCLS by transferring knowledge from abundant monolingual MS data to those data with insufficient volumes. In the triple-stage training method, a video-guided dual fusion network (VDF) is designed as the backbone network to integrate multimodal and cross-lingual information through diverse fusion strategies in the encoder and decoder; What's more, we propose two cross-lingual knowledge distillation strategies: adaptive pooling distillation and language-adaptive warping distillation (LAWD), designed for encoder-level and vocab-level distillation objects to facilitate effective knowledge transfer across cross-lingual sequences of varying lengths between MS and MCLS models. Specifically, to tackle lingual sequences of varying lengths between MS and MCLS models. Specifically, to tackle the challenge of unequal length of parallel cross-language sequences in KD, LAWD can directly conduct cross-language distillation while keeping the language feature shape unchanged to reduce potential information loss. We meticulously annotated the How2-MCLS dataset based on the How2 dataset to simulate MCLS scenarios. Experimental results show that the proposed method achieves competitive performance compared to strong baselines, and can bring substantial performance improvements to MCLS models by transferring knowledge from the MS model.
Nayu Liu, Kaiwen Wei, Yong Yang 0001, Jianhua Tao 0001, Xian Sun 0001, Fanglong Yao, Li Jin 0001, Zhao Lv, Cunhang Fan
IEEE Trans. Pattern Anal. Mach. Intell.10
2024 SceneFake: An initial dataset and benchmarks for scene fake audio detection
Jiangyan Yi, Chenglong Wang 0001, Jianhua Tao 0001, Chuyuan Zhang, Cunhang Fan, Zhengkun Tian, Haoxin Ma, Ruibo Fu
Pattern Recognit.5
2024 Dynamic Ensemble Teacher-Student Distillation Framework for Light-Weight Fake Audio Detection
abstract
In recent years, fake audio detection (FAD) has made great progress, and lightweight is important to achieve fast and reliable audio authenticity verification on resource-limited devices. However, most of the researchers ignore lightweight when improving the performance of FAD. To develop the application of FAD for small-end devices, this paper proposes a novel light-weight network named Light-ECA2Net. Given that networks with different depths have different abilities in capturing fake speech artifacts, this paper proposes a dynamic ensemble teacher-student distillation framework to fully transfer distillation knowledge. The dynamic ensemble distillation is divided into two aspects. First, we adopt one-to-one feature mapping to perceive the multidimensional feature knowledge and dynamically adjust every dimension feature weight by using ground truth labels, which can enable students to receive feature knowledge efficiently. Secondly, different network layers also have their strengths of predicting, further dynamically predicting weight can improve the learning ability of the student. Experimental results on the ASVspoof 2019 LA and PA datasets show that compared to the baseline, our system further improves performance by reducing the model complexity by 45%.
Jun Xue 0001, Cunhang Fan, Jiangyan Yi, Jian Zhou 0006, Zhao Lv
IEEE Signal Process. Lett.2
2024 Multi-Level Information Aggregation Based Graph Attention Networks Towards Fake Speech Detection
abstract
It is widely acknowledged that distinguishing genuine speech from spoofed speech encompasses various subbands and temporal segments within speech signals. However, prevailing spoofing detection methods tend to oversimplify the relationships between these cues by employing linear models. In this paper, we introduce a multi-level information aggregation Graph Attention Networks (MiaGATs) to generate highly discriminative features for fake speech detection (FSD). In MiaGATs, each subband and temporal segment of a speech signal is represented as distinct nodes. MiaGATs incorporates channel information aggregation within each node to effectively harness the unique spectral and temporal characteristics during the feature encoding stage. In particular, MiaGATs address the interactions between nodes through indirect node aggregation and integrates both indirect and direct node aggregation by max-pooling operation. Experimental results on ASVspoof2019 and ASVspoof2021 LA databases show significant relative improvement compared to the current state-of-the-art. In comparison to the leading integrated spectro-temporal graph attention networks, MiaGATs gains an impressive performance improvement in various conditions, underscoring MiaGATs's position as a new benchmark in spoofing detection performance.
Jian Zhou 0006, Yong Li 0032, Cunhang Fan, Hon Keung Kwan
IEEE Signal Process. Lett.3
2024 Dual-Branch Knowledge Distillation for Noise-Robust Synthetic Speech Detection
abstract
Most research in synthetic speech detection (SSD) focuses on improving performance on standard noise-free datasets. However, in actual situations, noise interference is usually present, causing significant performance degradation in SSD systems. To improve noise robustness, this paper proposes a dual-branch knowledge distillation synthetic speech detection (DKDSSD) method. Specifically, a parallel data flow of the clean teacher branch and the noisy student branch is designed, and interactive fusion module and response-based teacher-student paradigms are proposed to guide the training of noisy data from both the data distribution and decision-making perspectives. In the noisy student branch, speech enhancement is introduced initially for denoising, aiming to reduce the interference of strong noise. The proposed interactive fusion combines denoised features and noisy features to mitigate the impact of speech distortion and ensure consistency with the data distribution of the clean branch. The teacher-student paradigm maps the student's decision space to the teacher's decision space, enabling noisy speech to behave similarly to clean speech. Additionally, a joint training method is employed to optimize both branches for achieving global optimality. Experimental results based on multiple datasets demonstrate that the proposed method performs effectively in noisy environments and maintains its performance in cross-dataset experiments. Source code is available athttps://github.com/fchest/DKDSSD.
Cunhang Fan, Mingming Ding, Jianhua Tao 0001, Ruibo Fu, Jiangyan Yi, Zhengqi Wen, Zhao Lv
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 Learning From Yourself: A Self-Distillation Method For Fake Speech Detection
abstract
In this paper, we propose a novel self-distillation method for fake speech detection (FSD), which can significantly improve the performance of FSD without increasing the model complexity. For FSD, some fine-grained information is very important, such as spectrogram defects, mute segments, and so on, which are often perceived by shallow networks. However, shallow networks have much noise, which can not capture this very well. To address this problem, we propose using the deepest network instruct shallow network for enhancing shallow networks. Specifically, the networks of FSD are divided into several segments, the deepest network being used as the teacher model, and all shallow networks become multiple student models by adding classifiers. Meanwhile, the distillation path between the deepest network feature and shallow network features is used to reduce the feature difference. A series of experimental results on the ASVspoof 2019 LA and PA datasets show the effectiveness of the proposed method, with significant improvements compared to the baseline.
Jun Xue 0001, Cunhang Fan, Jiangyan Yi, Chenglong Wang 0001, Zhengqi Wen, Dan Zhang 0014, Zhao Lv
ICASSP2
2023 CompNet: Complementary network for single-channel speech enhancement
Cunhang Fan, Andong Li, Wang Xiang, Chengshi Zheng, Zhao Lv, Xiaopei Wu
Neural Networks1
2023 Subband fusion of complex spectrogram for fake speech detection
Cunhang Fan, Jun Xue 0001, Shunbo Dong, Mingming Ding, Jiangyan Yi, Jinpeng Li 0002, Zhao Lv
Speech Commun.1
2023 Transfer knowledge for punctuation prediction via adversarial training
Jiangyan Yi, Jianhua Tao 0001, Ye Bai 0001, Zhengkun Tian, Cunhang Fan
Speech Commun.5
2022 Csenet: Complex Squeeze-and-Excitation Network for Speech Depression Level Prediction
abstract
Automatic speech depression level prediction (SDLP) is a very challenging problem in affective computing. There are many studies that have acquired quite good performances for SDLP. However, most of the input speech features of these studies are based on the amplitude spectrogram, which loses the phase spectrogram information. Therefore, these speech features may lose some important information related to depression. In order to make full use of speech information, this paper proposes a complex squeeze-and-excitation network (CSENet) for SDLP. The complex spectrogram is used as the input speech feature, which contains both amplitude and phase spectrogram. In addition, to acquire a discriminative feature, the squeeze-and-excitation residual network is employed to extract deep speech feature. Finally, the attentive temporal pooling is utilized to dynamically select more important information according to the attention mechanisms. Experimental results on the AVEC 2013 and AVEC 2014 datasets prove the effectiveness of our proposed method. As for the mean absolute error (MAE) evaluation metric on AVEC 2013, our proposed method acquires state-of-the-art performance.
Cunhang Fan, Zhao Lv, Shengbing Pei, Mingyue Niu
ICASSP1
2022 ADD 2022: the first Audio Deep Synthesis Detection Challenge
abstract
Audio deepfake detection is an emerging topic, which was included in the ASVspoof 2021. However, the recent shared tasks have not covered many real-life and challenging scenarios. The first Audio Deep synthesis Detection challenge (ADD) was motivated to fill in the gap. The ADD 2022 includes three tracks: low-quality fake audio detection (LF), partially fake audio detection (PF) and audio fake game (FG). The LF track focuses on dealing with bona fide and fully fake utterances with various real-world noises etc. The PF track aims to distinguish the partially fake audio from the real. The FG track is a rivalry game, which includes two tasks: an audio generation task and an audio fake detection task. In this paper, we describe the datasets, evaluation metrics, and protocols. We also report major findings that reflect the recent advances in audio deepfake detection tasks.
Jiangyan Yi, Ruibo Fu, Jianhua Tao 0001, Shuai Nie 0001, Haoxin Ma, Chenglong Wang 0001, Tao Wang 0074, Zhengkun Tian, Ye Bai 0001, Cunhang Fan, Shan Liang 0007, Shuai Zhang 0014, Xinrui Yan, Zhengqi Wen, Haizhou Li 0001
ICASSP10
2022 DDAM '22: 1st International Workshop on Deepfake Detection for Audio Multimedia
abstract
Over the last few years, the technology of speech synthesis and voice conversion has made significant improvement with the development of deep learning. The models can generate realistic and human-like speech. It is difficult for most people to distinguish the generated audio from the real. However, this technology also poses a great threat to the global political economy and social stability if some attackers and criminals misuse it with the intent to cause harm. In this workshop, we aim to bring together researchers from the fields of audio deepfake detection, audio deep synthesis, audio fake game and adversarial attacks to further discuss recent research and future directions for detecting deepfake and manipulated audios in multimedia.
Jianhua Tao 0001, Jiangyan Yi, Cunhang Fan, Ruibo Fu, Shan Liang 0007, Pengyuan Zhang, Haizhou Li 0001, Helen M. Meng, Dong Yu 0001, Masato Akagi
ACM Multimedia3
2022 Dynamic Domain Adaptation for Class-Aware Cross-Subject and Cross-Session EEG Emotion Recognition
abstract
It is vital to develop general models that can be shared across subjects and sessions in the real-world deployment of electroencephalogram (EEG) emotion recognition systems. Many prior studies have exploited domain adaptation algorithms to alleviate the inter-subject and inter-session discrepancies of EEG distributions. However, these methods only aligned the global domain divergence, but overlooked the local domain divergence with respect to each emotion category. This degenerates the emotion-discriminating ability of the domain invariant features. In this paper, we argue that aligning the EEG data within the same emotion categories is important for generalizable and discriminative features. Hence, we propose the dynamic domain adaptation (DDA) algorithm where the global and local divergences are disposed by minimizing the global domain discrepancy and local subdomain discrepancy, respectively. To tackle the absence of emotion labels in the target domain, we introduce a dynamic training strategy where the model focuses on optimizing the global domain discrepancy in the early training steps, and then gradually switches to the local subdomain discrepancy. The DDA algorithm is formally implemented as an unsupervised version and a semi-supervised version for different experimental settings. Based on the coarse-to-fine alignment, our model achieves the average peak accuracy of 91.08%, 92.89% on SEED, and 81.58%, 80.82% on SEED-IV in the cross-subject and cross-session scenarios, respectively.
Zhunan Li, Enwei Zhu, Ming Jin 0006, Cunhang Fan, Huiguang He, Ting Cai 0001, Jinpeng Li 0002
IEEE J. Biomed. Health Informatics4
2021 Gated Recurrent Fusion With Joint Training Framework for Robust End-to-End Speech Recognition
abstract
The joint training framework for speech enhancement and recognition methods have obtained quite good performances for robust end-to-end automatic speech recognition (ASR). However, these methods only utilize the enhanced feature as the input of the speech recognition component, which are affected by the speech distortion problem. In order to address this problem, this paper proposes a gated recurrent fusion (GRF) method with joint training framework for robust end-to-end ASR. The GRF algorithm is used to dynamically combine the noisy and enhanced features. Therefore, the GRF can not only remove the noise signals from the enhanced features, but also learn the raw fine structures from the noisy features so that it can alleviate the speech distortion. The proposed method consists of speech enhancement, GRF and speech recognition. Firstly, the mask based speech enhancement network is applied to enhance the input speech. Secondly, the GRF is applied to address the speech distortion problem. Thirdly, to improve the performance of ASR, the state-of-the-art speech transformer algorithm is used as the speech recognition component. Finally, the joint training framework is utilized to optimize these three components, simultaneously. Our experiments are conducted on an open-source Mandarin speech corpus called AISHELL-1. Experimental results show that the proposed method achieves the relative character error rate (CER) reduction of 10.04% over the conventional joint enhancement and transformer method only using the enhanced features. Especially for the low signal-to-noise ratio (0 dB), our proposed method can achieves better performances with 12.67% CER reduction, which suggests the potential of our proposed method.
Cunhang Fan, Jiangyan Yi, Jianhua Tao 0001, Zhengkun Tian, Bin Liu 0041, Zhengqi Wen
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 Two Heads are Better Than One: A Two-Stage Complex Spectral Mapping Approach for Monaural Speech Enhancement
abstract
For challenging acoustic scenarios as low signal-to-noise ratios, current speech enhancement systems usually suffer from performance bottleneck in extracting the target speech from the mixtures within one step. To address this issue, we propose a novel complex spectral mapping approach with a two-stage pipeline for monaural speech enhancement in the time-frequency domain. The proposed algorithm aims to decouple the primal problem into multiple sub-problems, which follows the classic proverb, “two heads are better than one”. More specifically, in the first stage, only magnitude is estimated, which is incorporated with the noisy phase to obtain a coarse complex spectrum estimation. To facilitate the previous estimation, in the second stage, an auxiliary network serves as the post-processing module, where residual noise is further suppressed and the phase information is effectively modified. The global residual connection strategy is adopted in the second stage to accelerate the training convergence speed. To alleviate the parameter burden caused by the multi-stage pipeline, we propose a light-weight temporal convolutional module, which substantially decreases the trainable parameters and obtains even better objective performance over the original version. We conduct extensive experiments on three standard corpora, including WSJ0-SI84, DNS Challenge dataset, and Voice Bank + DEMAND dataset. Objective test results demonstrate that our proposed approach achieves state-of-the-art performance over previous advanced systems under various conditions. Meanwhile, subjective listening test results further validate the superiority of our proposed method in terms of subjective quality.
Andong Li, Chengshi Zheng, Cunhang Fan, Xiaodong Li 0002
IEEE ACM Trans. Audio Speech Lang. Process.4
2020 Gated Recurrent Fusion of Spatial and Spectral Features for Multi-Channel Speech Separation with Deep Embedding Representations
Cunhang Fan, Jianhua Tao 0001, Bin Liu 0041, Jiangyan Yi, Zhengqi Wen
INTERSPEECH1
2020 Joint Training for Simultaneous Speech Denoising and Dereverberation with Deep Embedding Representations
Cunhang Fan, Jianhua Tao 0001, Bin Liu 0041, Jiangyan Yi, Zhengqi Wen
INTERSPEECH1
2020 A Recursive Network with Dynamic Attention for Monaural Speech Enhancement
abstract
For continuous speech processing, dynamic attention is helpful in preferential processing, which has already been shown by the auditory dynamic attending theory.Accordingly, we propose a framework combining dynamic attention and recursive learning together for monaural speech enhancement.Apart from a major noise reduction network, we design a separated sub-network, which adaptively generates the attention distribution to control the information flow throughout the major network.Recursive learning is introduced to dynamically reduce the number of trainable parameters by reusing a network for multiple stages, where the intermediate output in each stage is corrected with a memory mechanism.By doing so, a more flexible and better estimation can be obtained.We conduct experiments on TIMIT corpus.Experimental results show that the proposed architecture obtains consistently better performance than recent state-of-the-art models in terms of both PESQ and STOI scores.The code is provided at https://github.com/Andong-Li-speech/DARCN.
Andong Li, Chengshi Zheng, Cunhang Fan, Renhua Peng, Xiaodong Li 0002
INTERSPEECH3
2020 Focal Loss for Punctuation Prediction
Jiangyan Yi, Jianhua Tao 0001, Zhengkun Tian, Ye Bai 0001, Cunhang Fan
INTERSPEECH5
2020 Deep imitator: Handwriting calligraphy imitation via deep attention networks
Bocheng Zhao, Jianhua Tao 0001, Zhengkun Tian, Cunhang Fan, Ye Bai 0001
Pattern Recognit.5
2020 End-to-End Post-Filter for Speech Separation With Deep Attention Fusion Features
abstract
In this article, we propose an end-to-end post-filter method with deep attention fusion features for monaural speaker-independent speech separation. At first, a time-frequency domain speech separation method is applied as the pre-separation stage. The aim of pre-separation stage is to separate the mixture preliminarily. Although this stage can separate the mixture, it still contains the residual interference. In order to enhance the pre-separated speech and improve the separation performance further, the end-to-end post-filter (E2EPF) with deep attention fusion features is proposed. The E2EPF can make full use of the prior knowledge of the pre-separated speech, which contributes to speech separation. It is a fully convolutional speech separation network and uses the waveform as the input features. Firstly, the 1-D convolutional layer is utilized to extract the deep representation features for the mixture and pre-separated signals in the time domain. Secondly, to pay more attention to the outputs of the pre-separation stage, an attention module is applied to acquire deep attention fusion features, which are extracted by computing the similarity between the mixture and the pre-separated speech. These deep attention fusion features are conducive to reduce the interference and enhance the pre-separated speech. Finally, these features are sent to the post-filter to estimate each target signals. Experimental results on the WSJ0-2mix dataset show that the proposed method outperforms the state-of-the-art speech separation method. Compared with the pre-separation method, our proposed method can acquire 64.1%, 60.2%, 25.6% and 7.5% relative improvements in scale-invariant source-to-noise ratio (SI-SNR), the signal-to-distortion ratio (SDR), the perceptual evaluation of speech quality (PESQ) and the short-time objective intelligibility (STOI) measures, respectively.
Cunhang Fan, Jianhua Tao 0001, Bin Liu 0041, Jiangyan Yi, Zhengqi Wen, Xuefei Liu
IEEE ACM Trans. Audio Speech Lang. Process.1
2019 A Time Delay Neural Network with Shared Weight Self-Attention for Small-Footprint Keyword Spotting
Ye Bai 0001, Jiangyan Yi, Jianhua Tao 0001, Zhengqi Wen, Zhengkun Tian, Chenghao Zhao, Cunhang Fan
INTERSPEECH7
2019 Discriminative Learning for Monaural Speech Separation Using Deep Embedding Features
abstract
Deep clustering (DC) and utterance-level permutation invariant training (uPIT) have been demonstrated promising for speakerindependent speech separation.DC is usually formulated as two-step processes: embedding learning and embedding clustering, which results in complex separation pipelines and a huge obstacle in directly optimizing the actual separation objectives.As for uPIT, it only minimizes the chosen permutation with the lowest mean square error, doesn't discriminate it with other permutations.In this paper, we propose a discriminative learning method for speaker-independent speech separation using deep embedding features.Firstly, a DC network is trained to extract deep embedding features, which contain each source's information and have an advantage in discriminating each target speakers.Then these features are used as the input for uPIT to directly separate the different sources.Finally, uPIT and DC are jointly trained, which directly optimizes the actual separation objectives.Moreover, in order to maximize the distance of each permutation, the discriminative learning is applied to fine tuning the whole model.Our experiments are conducted on WSJ0-2mix dataset.Experimental results show that the proposed models achieve better performances than DC and uPIT for speaker-independent speech separation.
Cunhang Fan, Bin Liu 0041, Jianhua Tao 0001, Jiangyan Yi, Zhengqi Wen
INTERSPEECH1
2019 Automatic Depression Level Detection via ℓp-Norm Pooling
Mingyue Niu, Jianhua Tao 0001, Bin Liu 0041, Cunhang Fan
INTERSPEECH4