Zhao Lv

dblp:38/6148 · DBLP profile ↗
← Back
66ranked-venue papers
4as first author
61since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 39 · 3 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 2 first-author · 32 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 7 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Dual-stream Relation-modeling Disentanglement for Cloth-Changing Person Re-Identification
abstract
Cloth-changing person re-identification (CC-ReID) aims to identify individuals across non-overlapping cameras despite clothing variations. Existing methods are often constrained by two primary limitations: approaches using auxiliary modalities typically rely on a single specific cue, limiting their robustness, while feature disentanglement methods struggle with discrete labels that create inconsistencies between ground truth labels and modality semantic similarity. To overcome these limitations, we propose DRDnet, a unified framework that synergistically integrates dual auxiliary cues and advanced relation modeling. Specifically, our Dual-Stream Disentanglement (DSD) module leverages textual descriptions and parsing images to decouple clothing factors through high-level semantic supervision and pixel-level operations, yielding robust clothing-agnostic features. Simultaneously, our Modal Relation Modeling (MRM) module constructs feature memory banks and employs adaptive soft label smoothing, effectively enhancing image-text semantic alignment and reinforcing identity consistency across clothing changes. We evaluate DRDnet on several CC-ReID benchmarks to demonstrate its effectiveness and provide state-of-the-art performance across all benchmarks.
Shijuan Huang, Zongyi Li, Zhao Lv
AAAI5
2026 Trainable EEG Interpolation and Structure-Sharing Dual-Path Encoders for Brain-Assisted Target Speaker Extraction
abstract
Brain-assisted target speaker extraction (TSE) isolates a target speaker's voice from a mixture by leveraging task-specific representations in Electroencephalogram (EEG) signals. However, existing methods rely on fixed interpolation for EEG-audio alignment, introducing redundant computations. They also employ single-path encoders that extract only target-relevant features while neglecting complementary, irrelevant ones, limiting discriminability. To address these limitations, this paper proposes a Trainable EEG Interpolation and Structure-sharing Dual-path Encoders network (TIDENet). The proposed Trainable EEG Interpolation (TEI) uses a neural network module to leverage cross-sample EEG information during resampling by parameters updating, thereby overcoming the limitations of fixed interpolation. The Structure-sharing Dual-path Encoders (SSDPE) extend existing speech and EEG encoders by introducing dual paths that separately process features relevant and irrelevant to the target speaker and incorporates interactive fusion between them, which enhances the encoder's ability to capture task-relevant information. Experimental results on public datasets demonstrate that TIDENet achieves relative improvements of up to 20.47%, 22.22%, 2.91%, 6.20%, and 15.84% in signal-to-distortion ratio (SDR), scale-invariant SDR (SI-SDR), short-time objective intelligibility (STOI), extended STOI (ESTOI), and perceptual evaluation of speech quality (PESQ), respectively, compared to the state-of-the-art. These significant gains validate the effectiveness of the proposed TEI method and SSDPE architecture.
Zhao Lv, Youdian Gao, Ruibo Fu, Cunhang Fan
AAAI1
2026 BrainHGT: A Hierarchical Graph Transformer for Interpretable Brain Network Analysis
abstract
Graph Transformer shows remarkable potential in brain network analysis due to its ability to model graph structures and complex node relationships. Most existing methods typically model the brain as a flat network, ignoring its modular structure, and their attention mechanisms treat all brain region connections equally, ignoring distance-related node connection patterns. However, brain information processing is a hierarchical process that involves local and long-range interactions between brain regions, interactions between regions and sub-functional modules, and interactions among functional modules themselves. This hierarchical interaction mechanism enables the brain to efficiently integrate local computations and global information flow, supporting the execution of complex cognitive functions. To address this issue, we propose BrainHGT, a hierarchical Graph Transformer that simulates the brain’s natural information processing from local regions to global communities. Specifically, we design a novel long-short range attention encoder that utilizes parallel pathways to handle dense local interactions and sparse long-range connections, thereby effectively alleviating the over-globalizing issue. To further capture the brain’s modular architecture, we designe a prior-guided clustering module that utilizes a cross-attention mechanism to group brain regions into functional communities and leverage neuroanatomical prior to guide the clustering process, thereby improving the biological plausibility and interpretability. Experimental results indicate that our proposed method significantly improves performance of disease identification, and can reliably capture the sub-functional modules of the brain, demonstrating its interpretability.
Chao Zhang 0047, Zhao Lv, Shengbing Pei
AAAI4
2026 ReFL: Reflective Feedback Learning for Hallucination Detection of Large Language Models
abstract
Cunhang Fan, Jun Zhang, Xue Zhang, Shuai Zhang, Zhao Lv, Jianhua Tao, Zhengqi Wen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Cunhang Fan, Shuai Zhang 0014, Zhao Lv, Jianhua Tao 0001, Zhengqi Wen
ACL (1)5
2026 A Shared Look: Detecting Deepfakes with Inter-Subject Neural Synchrony
abstract
The rapid evolution of generative AI presents a significant challenge for Deepfake detection. While most research focuses on face-swapping, the emerging threat of "image-to-video" (I2V) forgeries is harder to detect and poses a greater risk. Traditional computer vision detectors rely on transient digital artifacts, which often lack interpretability and robustness against the new generation techniques. This study introduces a neuro-cognitive method, using dyadic electroencephalogram (EEG) to decode the human perception of authenticity. We recorded inter-brain synchrony via EEG hyperscanning from 15 participant pairs as they viewed a balanced set of authentic and AI-generated videos. Results showed that these shared neural response can classify video authenticity with an accuracy of up to 89.23% using our proposed Hyper-FusionNet. In addition, the biomarkers exhibited distinct patterns for different emotional valences, highlighting their versatility. These findings highlight the potential of inter-brain synchrony for detecting emerging deepfakes, offering a new perspective for enhancing user trust and digital literacy.
Shiang Hu, Zhiwen Zha, Dongdong Jia, Zhao Lv
CHI8
2026 EEGCo-Diff: Manifold-Guided Diffusion and an Oscillation-Aware Classifier for Sample-Efficient EEG Depression Detection
abstract
Depression is a serious mental disorder, and timely detection and treatment are crucial. Electroencephalography (EEG), as a noninvasive tool directly reflecting brain activity, is suitable for objective depression detection. However, due to difficulty in acquiring depression EEG data, public datasets suffer from limited sample sizes and insufficient exploitation of features, limiting current approaches in training sufficiency and generalization. To this end, this paper proposes an EEG collaborative diffusion framework (EEGCo‐Diff) to address small‐sample learning and sufficient feature utilization. Specifically, this study constructs a multicondition guided diffusion model conditioned on Riemannian manifold features and ground‐truth labels to generate high‐quality EEG samples while maintaining geometric consistency of the original distribution, effectively alleviating sample scarcity and distribution shift. Second, this study proposes a CNN interaction transformer network (CITNet) that fuses multilayer convolution and a transformer to model local details and global temporal dependencies and uses an oscillation‐aware module (OAM) to highlight key channel features. EEGCo‐Diff couples geometry‐prior‐driven data generation with structured temporal modeling, significantly improving discriminative power and generalization in small‐sample, heterogeneous settings. On two public EEG depression datasets, our method achieves accuracies of 89.44% and 95.03%, outperforming the strongest baselines by 8.34% and 1.03%, respectively, and establishing state‐of‐the‐art performance.
Minchao Wu, Guowang Zhuang, Zhao Lv
Int. J. Intell. Syst.6
2026 Beyond element-level understanding: Explicit relational understanding for GUI agents
Longhui Ma, Siwei Wang 0001, Zhao Lv
Pattern Recognit.4
2025 Region-Based Optimization in Continual Learning for Audio Deepfake Detection
abstract
Rapid advancements in speech synthesis and voice conversion bring convenience but also new security risks, creating an urgent need for effective audio deepfake detection. Although current models perform well, their effectiveness diminishes when confronted with the diverse and evolving nature of real-world deepfakes. To address this issue, we propose a continual learning method named Region-Based Optimization (RegO) for audio deepfake detection. Specifically, we use the Fisher information matrix to measure important neuron regions for real and fake audio detection, dividing them into four regions. First, we directly fine-tune the less important regions to quickly adapt to new tasks. Next, we apply gradient optimization in parallel for regions important only to real audio detection, and in orthogonal directions for regions important only to fake audio detection. For regions that are important to both, we use sample proportion-based adaptive gradient optimization. This region-adaptive optimization ensures an appropriate trade-off between memory stability and learning plasticity. Additionally, to address the increase of redundant neurons from old tasks, we further introduce the Ebbinghaus forgetting mechanism to release them, thereby promoting the model’s ability to learn more generalized discriminative features. Experimental results show our method achieves a 21.3 percent improvement in EER over the state-of-the-art continual learning approach RWM for audio deepfake detection. Moreover, the effectiveness of RegO extends beyond the audio deepfake detection domain, showing potential significance in other tasks, such as image recognition.
Yujie Chen 0006, Jiangyan Yi, Cunhang Fan, Jianhua Tao 0001, Yong Ren 0006, Siding Zeng, Chu Yuan Zhang, Xinrui Yan, Jun Xue 0001, Chenglong Wang 0001, Zhao Lv, Xiaohui Zhang 0006
AAAI12
2025 BSDB-Net: Band-Split Dual-Branch Network with Selective State Spaces Mechanism for Monaural Speech Enhancement
abstract
Although the complex spectrum-based speech enhancement (SE) methods have achieved significant performance, coupling amplitude and phase can lead to a compensation effect, where amplitude information is sacrificed to compensate for the phase that is harmful to SE. In addition, to further improve the performance of SE, many modules are stacked onto SE, resulting in increased model complexity that limits the application of SE. To address these problems, we proposed a dual-path network based on compressed frequency using Mamba. First, we extract amplitude and phase information through parallel dual branches. This approach leverages structured complex spectra to implicitly capture phase information and solves the compensation effect by decoupling amplitude and phase, and the network incorporates an interaction module to suppress unnecessary parts and recover missing components from the other branch. Second, to reduce network complexity, the network introduces a band-split strategy to compress the frequency dimension. To further reduce complexity while maintaining good performance, we designed a Mamba-based module that models the time and frequency dimensions under linear complexity. Finally, compared to baselines, our model achieves an average 8.3 times reduction in computational complexity while maintaining superior performance. Furthermore, it achieves a 25 times reduction in complexity compared to transformer-based models.
Cunhang Fan, Enrui Liu, Andong Li, Jianhua Tao 0001, Jian Zhou 0006, Chengshi Zheng, Zhao Lv
AAAI8
2025 Self-supervised fMRI Outlier Detection via Graph Reachability Modeling
abstract
High-quality neuroimaging data is vital for robust brain disease identification, yet the presence of outliers in functional magnetic resonance imaging (fMRI) datasets often degrades performance and reliability of identification models. Existing unsupervised outlier detection methods struggle with parameter sensitivity and data distributional complexity, while supervised detection is hindered by the scarcity of labeled outliers. To address this challenge, a novel self-supervised frame-work named Variational Autoencoder Graph Outlier Detection (VAGOD) is proposed, in which pseudo-labeled normal and outlier samples are first generated by a conditional variational autoencoder with hierarchical batch-wise attention, and then a reachability-based graph neural network uses these labels to learn local structures and find subtle outliers in the functional connectivity network. Extensive experiments on the ADHD-200 and ADNI2 datasets demonstrate that removing outliers with VAGOD significantly improves downstream brain disease classification accuracy and stability. Furthermore, group-level analysis reveals that detected outliers exhibit distinct neurobiological signatures, validating the method's interpretability and its practical value for clinical neuroimaging applications.
Shengbing Pei, Wencong Jiang, Chao Zhang 0047, Zhao Lv
BIBM6
2025 Fair-Esi: Feature Adaptive Importance Refinement for Electrophysiological Source Imaging
abstract
An essential technique for diagnosing brain disorders is electrophysiological source imaging (ESI). While modelbased optimization and deep learning methods have achieved promising results in this field, the accurate selection and refinement of features remains a central challenge for precise ESI. This paper proposes FAIR-ESI, a novel framework that adaptively refines feature importance across different views, including FFTbased spectral feature refinement, weighted temporal feature refinement, and self-attention-based patch-wise feature refinement. Extensive experiments on two simulation datasets with diverse configurations and two real-world clinical datasets validate our framework's efficacy, highlighting its potential to advance brain disorder diagnosis and offer new insights into brain function.
Linyong Zou, Xiongfei Wang, Jia-Hong Gao 0001, Shurong Sheng, Kuntao Xiao, Pengfei Teng, Guoming Luan, Zhao Lv
BIBM11
2025 Detection of AI-Generated Contents Based on Dyadic-Brain Neural Synchronization
Shiang Hu, Lihao Fu, Piqiang Zhang, Juan Hou, Zhao Lv
CogSci7
2025 COLA: Collaborative Multi-Agent Framework with Dynamic Task Scheduling for GUI Automation
abstract
With the rapid advancements in Large Language Models (LLMs), an increasing number of studies have leveraged LLMs as the cognitive core of agents to address complex task decision-making challenges.Specially, recent research has demonstrated the potential of LLM-based agents on automating GUI operations.However, existing methodologies exhibit two critical challenges: (1) static agent architectures struggle to adapt to diverse GUI application scenarios, leading to inadequate scenario generalization; (2) the agent workflows lack fault tolerance mechanism, necessitating complete process re-execution for GUI agent decision error.To address these limitations, we introduce COLA, a collaborative multi-agent framework for automating GUI operations.In this framework, a scenario-aware agent Task Scheduler decomposes task requirements into atomic capability units, dynamically selects the optimal agent from a decision agent pool, effectively responds to the capability requirements of diverse scenarios.Furthermore, we develop an interactive backtracking mechanism that enables human to intervene to trigger state rollbacks for non-destructive process repair.Experiments on the GAIA dataset show that COLA achieves competitive performance among GUI Agent methods, with an average accuracy of 31.89%.On WindowsAgentArena, it performs particularly well in Web Browser (33.3%),Media & Video (33.3%), and Windows Utils (25.0%), suggesting the effectiveness of specialized agent design and dynamic strategy allocation.The code is available at https://github.com/Alokia/COLA-demo.
Longhui Ma, Siwei Wang 0001, Zhao Lv
EMNLP5
2025 Improved Feature Extraction Network for Neuro-Oriented Target Speaker Extraction
abstract
The recent rapid development of auditory attention decoding (AAD) offers the possibility of using electroencephalography (EEG) as auxiliary information for target speaker extraction. However, effectively modeling long sequences of speech and resolving the identity of the target speaker from EEG signals remains a major challenge. In this paper, an improved feature extraction network (IFENet) is proposed for neuro-oriented target speaker extraction, which mainly consists of a speech encoder with dual-path Mamba and an EEG encoder with Kolmogorov-Arnold Networks (KAN). We propose SpeechBiMamba, which makes use of dual-path Mamba in modeling local and global speech sequences to extract speech features. In addition, we propose EEGKAN to effectively extract EEG features that are closely related to the auditory stimuli and locate the target speaker through the subject’s attention information. Experiments on the KUL and AVED datasets show that IFENet outperforms the state-of-the-art model, achieving 36% and 29% relative improvements in terms of scale-invariant signal-to-distortion ratio (SI-SDR) under an open evaluation condition.
Cunhang Fan, Youdian Gao, Zexu Pan, Jie Zhang 0042, Zhao Lv
ICASSP7
2025 SSM2Mel: State Space Model to Reconstruct Mel Spectrogram from the EEG
abstract
Decoding speech from brain signals is a challenging research problem that holds significant importance for studying speech processing in the brain. Although breakthroughs have been made in reconstructing the mel spectrograms of audio stimuli perceived by subjects at the word or letter level using non-invasive electroencephalography (EEG), there is still a critical gap in precisely reconstructing continuous speech features, especially at the minute level. To address this issue, this paper proposes a State Space Model (SSM) to reconstruct the mel spectrogram of continuous speech from EEG, named SSM2Mel. This model introduces a novel Mamba module to effectively model the long sequence of EEG signals for imagined speech. In the SSM2Mel model, the S4-UNet structure is used to enhance the extraction of local features of EEG signals, and the Embedding Strength Modulator (ESM) module is used to incorporate subject-specific information. Experimental results show that our model achieves a Pearson correlation of 0.069 on the SparrKULee dataset, which is a 38% improvement over the previous baseline.
Cunhang Fan, Zexu Pan, Zhao Lv
ICASSP5
2025 EEG Correlation Analysis-guided Graph Local Enhanced Feature Learning For Emotion Recognition
abstract
EEG-based emotion recognition is a key technology in brain-computer interfaces. Many previous studies have applied deep learning methods to mine emotion-related features in EEG to decode emotions. However, they overlooked the importance of electrode correlations and varying brain region activation during emotional processes, which are critical for emotion recognition. In this paper, we propose a method named EEG correlation analysis-guided graph local enhanced feature learning network (CAGLE-net). In CAGLE-net, we use the correlation analysis to guide the learning of the dynamic directed connection matrix to capture topological features, which are then fed into the locally enhanced embedding layer to generate enhanced features for each brain region. Subsequently, the cross-attention fusion mechanism is employed to fully leverage these locally enhanced features, yielding more discriminative representations. Experiment results on the SEED dataset show that CAGLE-net outperforms existing baseline methods. This study offers a promising solution for EEG-based emotion recognition.
Guowang Zhuang, Minchao Wu, Zhao Lv
ICASSP4
2025 Transformer Based Multi-view Learning for Integrating Static and Dynamic Complementarity of Brain Function
abstract
Dynamic temporal information and static connectivity information derived from functional magnetic resonance imaging (fMRI) can assist in the diagnosis of neurological disorders. However, existing disease diagnosis methods primarily rely on information from a single view, neglecting the advantages of multi-view information fusion. In this work, we propose an end-to-end multi-view fusion method that pre-trains on one view of fMRI data and fine-tunes on another view for disease identification. First, the dynamic temporal information and static connectivity information are integrated during the pre-training stage based on the consistency between the two views, effectively combining complementary information from both data types to improve disease identification accuracy. Finally, in the fine-tuning stage, for different fine-tuning datasets, we combine the residual connections in the model with the self-attention mechanism through the hadamard product. This guides the learning process and can be seen as a form of regularization or inductive bias, enhancing the models ability to learn from the data. Experiments conducted on the ADHD-200 dataset demonstrate that: 1) our method effectively fuses temporal and connectivity information from fMRI, improving the accuracy of brain disorder identification; 2) analyzing the consistency between the two views validates the effectiveness of the pre-training strategy and its positive impact on accuracy; 3) the residual attention maps of the model fine-tuned with functional connectivity networks (FCN) capture distinct symmetrical connections, which align with the inherent symmetry of FCN, supporting the rationale for using the hadamard product.
Shengbing Pei, Zhao Lv, Chao Zhang 0047
ICASSP3
2025 M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction
abstract
The brain-assisted target speaker extraction (TSE) aims to extract the attended speech from mixed speech by utilizing the brain neural activities, for example Electroencephalography (EEG). However, existing models overlook the issue of temporal misalignment between speech and EEG modalities, which hampers TSE performance. In addition, the speech encoder in current models typically uses basic temporal operations (e.g., one-dimensional convolution), which are unable to effectively extract target speaker information. To address these issues, this paper proposes a multi-scale and multi-modal alignment network (M3ANet) for brain-assisted TSE. Specifically, to eliminate the temporal inconsistency between EEG and speech modalities, the modal alignment module that uses a contrastive learning strategy is applied to align the temporal features of both modalities. Additionally, to fully extract speech information, multi-scale convolutions with GroupMamba modules are used as the speech encoder, which scans speech features at each scale from different directions, enabling the model to capture deep sequence information. Experimental results on three publicly available datasets show that the proposed model outperforms current state-of-the-art methods across various evaluation metrics, highlighting the effectiveness of our proposed method. The source code is available at: https://github.com/fchest/M3ANet.
Cunhang Fan, Jian Zhou 0006, Zexu Pan, Youdian Gao, Xiaoke Yang, Zhengqi Wen, Zhao Lv
IJCAI9
2025 ListenNet: A Lightweight Spatio-Temporal Enhancement Nested Network for Auditory Attention Detection
abstract
Auditory attention detection (AAD) aims to identify the direction of the attended speaker in multi-speaker environments from brain signals, such as Electroencephalography (EEG) signals. However, existing EEG-based AAD methods overlook the spatio-temporal dependencies of EEG signals, limiting their decoding and generalization abilities. To address these issues, this paper proposes a Lightweight Spatio-Temporal Enhancement Nested Network (ListenNet) for AAD. The ListenNet has three key components: Spatio-temporal Dependency Encoder (STDE), Multi-scale Temporal Enhancement (MSTE), and Cross-Nested Attention (CNA). The STDE reconstructs dependencies between consecutive time windows across channels, improving the robustness of dynamic pattern extraction. The MSTE captures temporal features at multiple scales to represent both fine-grained and long-range temporal patterns. In addition, the CNA integrates hierarchical features more effectively through novel dynamic attention mechanisms to capture deep spatio-temporal correlations. Experimental results on three public datasets demonstrate the superiority of ListenNet over state-of-the-art methods in both subject-dependent and challenging subject-independent settings, while reducing the trainable parameter count by approximately 7 times. Code is available at:https://github.com/fchest/ListenNet.
Cunhang Fan, Xiaoke Yang, Jian Zhou 0006, Zhao Lv
IJCAI7
2025 MHANet: Multi-scale Hybrid Attention Network for Auditory Attention Detection
abstract
Auditory attention detection (AAD) aims to detect the target speaker in a multi-talker environment from brain signals, such as electroencephalography (EEG), which has made great progress. However, most AAD methods solely utilize attention mechanisms sequentially and overlook valuable multi-scale contextual information within EEG signals, limiting their ability to capture long-short range spatiotemporal dependencies simultaneously. To address these issues, this paper proposes a multi-scale hybrid attention network (MHANet) for AAD, which consists of the multi-scale hybrid attention (MHA) module and the spatiotemporal convolution (STC) module. Specifically, MHA combines channel attention and multi-scale temporal and global attention mechanisms. This effectively extracts multi-scale temporal patterns within EEG signals and captures long-short range spatiotemporal dependencies simultaneously. To further improve the performance of AAD, STC utilizes temporal and spatial convolutions to aggregate expressive spatiotemporal representations. Experimental results show that the proposed MHANet achieves state-of-the-art performance with fewer trainable parameters across three datasets, 3 times lower than that of the most advanced model. Code is available at: https://github.com/fchest/MHANet.
Cunhang Fan, Xiaoke Yang, Jian Zhou 0006, Zhao Lv
IJCAI7
2025 Community-Aware Graph Transformer for Brain Disorder Identification
abstract
Abnormal brain functional network is an effective biomarker for brain disease diagnosis. Most existing methods focus on mining discriminative information from whole-brain connectivity patterns. However, multi-level collaboration is the foundation of efficient brain function, in addition to the whole-brain network, there are multiple sub-networks that can quickly integrate and process specific cognitive functions, forming the modular community structure of the brain. To address this gap, we propose a novel method, community-aware graph Transformer (CAGT), that integrates the community information of sub-networks and the topological information of brain graph into the Transformer architecture for better brain disorder identification. CAGT enhances information exchange within and between functional communities through dual-scale feature fusion, capturing interactive information across various scales. Additionally, it incorporates prior knowledge to design brain region position encoding and guide the self-attention, thereby enhancing the spatial awareness of the Transformer and aligning it with the brain's natural information transfer process. Experimental results indicate that our proposed method significantly improves performance on both large and small datasets, and can reliably capture the interactions between sub-networks, demonstrating its generalization and interpretability.
Shengbing Pei, Zhao Lv, Chao Zhang 0047, Jihong Guan
IJCAI3
2025 ID-RemovalNet: Identity Removal Network for EEG Privacy Protection with Enhancing Decoding Tasks
abstract
Electroencephalogram (EEG) contains not only decoding task information but also personal identity privacy information. If it is stolen or attacked, the user's brain-computer interaction behavior may be maliciously manipulated. Existing EEG identity privacy protection generally adopts generative or adding tiny perturbation methods, which can protect the identity privacy in EEG signals to some extent. However, these methods also damage the performance of decoding task. In order to solve these problems, this paper proposes an identity removal network (ID-RemovalNet) to achieve EEG privacy protection while improving the classification accuracy of decoding task. Firstly, an identity decorrelation separation module is constructed to accurately remove the identity features to achieve privacy protection while reducing the interference with the task decoding features. Secondly, a multi-domain multi-level fusion feature extraction module is designed to extract the high-quality EEG time-frequency features. Finally, the feature enhancement module is used to compensate for the loss of task decoding features and excitation of dominant feature selection during identity feature removal. The experimental results show that ID-RemoveNet removes identity information to 0.43% on four EEG datasets with two different paradigms, and significantly improves the EEG task decoding accuracy by 3.28%, and achieves the state-of-the-art performance in cross-subject EEG experiment.
Jie Ruan, Cunhang Fan, Yingfan Cheng, Zhao Lv
IJCAI5
2025 REB-former: RWKV-enhanced E-branchformer for Speech Recognition
Wang Xiang, Jian Zhou 0006, Cunhang Fan, Zhao Lv
INTERSPEECH5
2025 SSF-DST: A Spectro-Spatial Features Enhanced Deep Spatiotemporal Network for EEG-Based Auditory Attention Detection
Xiaoke Yang, Jian Zhou 0006, Zhao Lv, Cunhang Fan
INTERSPEECH5
2025 DHGCN: Dual HyperGraph Convolutional Network for EEG-Based Auditory Attention Detection
abstract
Auditory attention detection (AAD) aims to identify the attended speaker in multi-talker environments by analyzing brain activity recorded through neural monitoring techniques. Recent AAD approaches have achieved great progress in improving detection accuracy. However, they still face challenges in capturing complex spatio-temporal dependencies and high-order nonlinear relationships across brain regions. To address these challenges, this paper proposes DHGCN, a dual hypergraph convolutional network that integrates a hypergraph modeling module, a dual-branch hypergraph learning (DHGL) module, and a feature fusion module. Specifically, the hypergraph modeling module constructs spatial and temporal hypergraphs from EEG signals, enabling the representation of high-order relationships among channels and time points. The DHGL module comprises two parallel branches: a spatial branch that learns high-order spatial dependencies across EEG channels, and a temporal branch that captures complex temporal dependencies. Each branch uses its corresponding hypergraph structure, which is established during the modeling phase. The feature fusion module then aggregates spatial and temporal representations from both branches to support robust auditory attention classification. Extensive experiments on multiple benchmark datasets demonstrate that DHGCN consistently outperforms state-of-the-art AAD models. It achieves superior classification performance while reducing the trainable parameters count by over 50% compared to the state-of-the-art models. Code is available at: https://github.com/nobody1219/DHGCN.git.
Jian Zhou 0006, Yingjie Xie, Cunhang Fan, Zhao Lv
ACM Multimedia5
2025 DMF2Mel: A Dynamic Multiscale Fusion Network for EEG-Driven Mel Spectrogram Reconstruction
abstract
Decoding speech from brain signals is a challenging research problem. Although existing technologies have made progress in reconstructing the mel spectrograms of auditory stimuli at the word or letter level, there remain core challenges in the precise reconstruction of minute-level continuous imagined speech: traditional models struggle to balance the efficiency of temporal dependency modeling and information retention in long-sequence decoding. To address this issue, this paper proposes the Dynamic Multiscale Fusion Network (DMF2Mel), which consists of four core components: the Dynamic Contrastive Feature Aggregation Module (DC-FAM), the Hierarchical Attention-Guided Multi-Scale Network (HAMS-Net), the SplineMap attention mechanism, and the bidirectional state space module (convMamba). Specifically, the DC-FAM separates speech-related ''foreground features'' from noisy ''background features'' through local convolution and global attention mechanisms, effectively suppressing interference and enhancing the representation of transient signals. HAMS-Net, based on the U-Net framework, achieves cross-scale fusion of high-level semantics and low-level details. The SplineMap attention mechanism integrates the Adaptive Gated Kolmogorov-Arnold Network (AGKAN) to combine global context modeling with spline-based local fitting. The convMamba captures long-range temporal dependencies with linear complexity and enhances nonlinear dynamic modeling capabilities. Results on the SparrKULee dataset show that DMF2Mel achieves a Pearson correlation coefficient of 0.074 in mel spectrogram reconstruction for known subjects (a 48% improvement over the baseline) and 0.048 for unknown subjects (a 35% improvement over the baseline).Code is available at: https://github.com/fchest/DMF2Mel.
Cunhang Fan, Enrui Liu, Gangming Zhao, Zhao Lv
ACM Multimedia7
2025 MGBF: Multi-GNNs Bridge Framework for Brain Diseases Classification via Information Sharing and Denoising
Honghao Li, Zhao Lv, Chao Zhang 0047, Shengbing Pei
PRCV (13)3
2025 Multi-Level Contrastive Learning: Hierarchical Alleviation of Heterogeneity in Multimodal Sentiment Analysis
abstract
Recently, multimodal fusion efforts have achieved remarkable success in Multimodal Sentiment Analysis (MSA). However, most of the existing methods are based on model-level fusion, and the challenge of heterogeneity between modalities is not well resolved. Heterogeneity lies in the different feature distributions and distinct representation spaces among different modalities. To mitigate this problem, we propose that fusion is a progressive process, and we introduce a novel multi-level contrastive learning and multi-layer convolution fusion (MCL-MCF) method for MSA. Due to the relationships among multimodal data, the fusion process that involves single-modal to single-modal, single-modal to bimodal or trimodal, and higher-level fused modality semantic consistency is divided into three levels. The first-level contrast learning alleviates heterogeneity between unimodal modalities at the early level of multimodal feature fusion. The second-level contrast learning mitigates heterogeneity between unimodal and fused modalities. At the third level, we introduce a tensor convolution fusion (TCF) module that extracts high-level semantic features from the fused modalities and mitigates heterogeneity at the higher feature level through contrastive learning. To simulate fusion as a progressive process, MCF is proposed to fuse shallow and deep features to model complex relationships among modalities. Experiments on three public datasets show our approach's state-of-the-art performance.
Cunhang Fan, Kang Zhu, Jianhua Tao 0001, Guofeng Yi, Jun Xue 0001, Zhao Lv
IEEE Trans. Affect. Comput.6
2025 CrossConvPyramid: Deep Multimodal Fusion for Epileptic Magnetoencephalography Spike Detection
abstract
Magnetoencephalography (MEG) is a vital non-invasive tool for epilepsy analysis, as it captures high-resolution signals that reflect changes in brain activity over time. The automated detection of epileptic spikes within these signals can significantly reduce the labor and time required for manual annotation of MEG recording data, thereby aiding clinicians in identifying epileptogenic foci and evaluating treatment prognosis. Research in this domain often utilizes the raw, multi-channel signals from MEG scans for spike detection, commonly neglecting the multi-channel spiking patterns from spatially adjacent channels. Moreover, epileptic spikes share considerable morphological similarities with artifact signals within the recordings, posing a challenge for models to differentiate between the two. In this paper, we introduce a multimodal fusion framework that addresses these two challenges collectively. Instead of relying solely on the signal recordings, our framework also mines knowledge from their corresponding topography-map images, which encapsulate the spatial context and amplitude distribution of the input signals. To facilitate more effective data fusion, we present a novel multimodal feature fusion technique called CrossConvPyramid, built upon a convolutional pyramid architecture augmented by an attention mechanism. It initially employs cross-attention and a convolutional pyramid to encode inter-modal correlations within the intermediate features extracted by individual unimodal networks. Subsequently, it utilizes a self-attention mechanism to refine and select the most salient features from both inter-modal and unimodal features, specifically tailored for the spike classification task. Our method achieved the average F1 scores of 92.88% and 95.23% across two distinct real-world MEG datasets from separate centers, respectively outperforming the current state-of-the-art by 2.31% and 0.88%. We plan to release the code on GitHub later.
Shurong Sheng, Xiongfei Wang, Jia-Hong Gao 0001, Kuntao Xiao, Pengfei Teng, Guoming Luan, Zhao Lv
IEEE J. Biomed. Health Informatics10
2025 SAMCL: Subgraph-Aligned Multiview Contrastive Learning for Graph Anomaly Detection
abstract
Graph anomaly detection (GAD) has gained increasing attention in various attribute graph applications, i.e., social communication and financial fraud transaction networks. Recently, graph contrastive learning (GCL)-based methods have been widely adopted as the mainstream for GAD with remarkable success. However, existing GCL strategies in GAD mainly focus on node-node and node-subgraph contrast and fail to explore subgraph-subgraph level comparison. Furthermore, the different sizes or component node indices of the sampled subgraph pairs may cause the "nonaligned" issue, making it difficult to accurately measure the similarity of subgraph pairs. In this article, we propose a novel subgraph-aligned multiview contrastive approach for graph anomaly detection, named SAMCL, which fills the subgraph-subgraph contrastive-level blank for GAD tasks. Specifically, we first generate the multiview augmented subgraphs by capturing different neighbors of target nodes forming contrasting subgraph pairs. Then, to fulfill the nonaligned subgraph pair contrast, we propose a subgraph-aligned strategy that estimates similarities with the Earth mover's distance (EMD) of both considering the node embedding distributions and typology awareness. With the newly established similarity measure for subgraphs, we conduct the interview subgraph-aligned contrastive learning module to better detect changes for nodes with different local subgraphs. Moreover, we conduct intraview node-subgraph contrastive learning to supplement richer information on abnormalities. Finally, we also employ the node reconstruction task for the masked subgraph to measure the local change of the target node. Finally, the anomaly score for each node is jointly calculated by these three modules. Extensive experiments conducted on benchmark datasets verify the effectiveness of our approach compared to existing state-of-the-art (SOTA) methods with significant performance gains (up to 6.36% improvement on ACM). Our code can be verified at https://github.com/hujingtao/SAMCL.
Jingtao Hu, Bin Xiao 0002, Hu Jin 0005, Jingcan Duan, Siwei Wang 0001, Zhao Lv, Siqi Wang 0001, Xinwang Liu 0002, En Zhu
IEEE Trans. Neural Networks Learn. Syst.6
2025 Location-aware Inaudible Attack Defense Towards Smart Speakers
abstract
Recent studies show that inaudible attacks pose a non-negligible security risk to smart speakers. While several countermeasures have been proposed to detect the occurrence of the inaudible attack passively, accurately locating the attack source in 3D free space remains an unresolved challenge. Arrow is designed to bridge this gap by attempting to detect the occurrence of inaudible attacks and determine their localization simultaneously. Instead of relying on dedicated hardware components, Arrow is implemented with the microphone array widely deployed on COTS (Commercial Off-The-Shelf) smart speakers. Throughout the spatial information captured by the microphone array, Arrow establishes a spatial mapping model and derives orientation-related features to pinpoint the location of the attack source. Furthermore, to improve the robustness against co-channel interference, Arrow adopt carefully-modulated ultrasonic waveforms to achieve noise-robust attack detection. Through the above technical mechanism, Arrow can significantly improve the security level of voice assistants on smart speakers with nearly zero deployment cost. We implement a prototype of Arrow and conduct a comprehensive performance evaluation. The results show Arrow can achieve 2.5 ○ and 7 ○ error in DoA estimation for horizontal and vertical angles, respectively.
Ping Li 0020, Xinrui He, Zhenfei Zhang, Feiyu Han, Panlong Yang, Zhao Lv
ACM Trans. Sens. Networks6
2024 Progressive Distillation Based on Masked Generation Feature Method for Knowledge Graph Completion
abstract
In recent years, knowledge graph completion (KGC) models based on pre-trained language model (PLM) have shown promising results. However, the large number of parameters and high computational cost of PLM models pose challenges for their application in downstream tasks. This paper proposes a progressive distillation method based on masked generation features for KGC task, aiming to significantly reduce the complexity of pre-trained models. Specifically, we perform pre-distillation on PLM to obtain high-quality teacher models, and compress the PLM network to obtain multi-grade student models. However, traditional feature distillation suffers from the limitation of having a single representation of information in teacher models. To solve this problem, we propose masked generation of teacher-student features, which contain richer representation information. Furthermore, there is a significant gap in representation ability between teacher and student. Therefore, we design a progressive distillation method to distill student models at each grade level, enabling efficient knowledge transfer from teachers to students. The experimental results demonstrate that the model in the pre-distillation stage surpasses the existing state-of-the-art methods. Furthermore, in the progressive distillation stage, the model significantly reduces the model parameters while maintaining a certain level of performance. Specifically, the model parameters of the lower-grade student model are reduced by 56.7\% compared to the baseline.
Cunhang Fan, Yujie Chen 0006, Jun Xue 0001, Yonghui Kong, Jianhua Tao 0001, Zhao Lv
AAAI6
2024 A Non-parametric Graph Clustering Framework for Multi-View Data
abstract
Multi-view graph clustering (MVGC) derives encouraging grouping results by seamlessly integrating abundant information inside heterogeneous data, and has captured surging focus recently. Nevertheless, the majority of current MVGC works involve at least one hyper-parameter, which not only requires additional efforts for tuning, but also leads to a complicated solving procedure, largely harming the flexibility and scalability of corresponding algorithms. To this end, in the article we are devoted to getting rid of hyper-parameters, and devise a non-parametric graph clustering (NpGC) framework to more practically partition multi-view data. To be specific, we hold that hyper-parameters play a role in balancing error item and regularization item so as to form high-quality clustering representations. Therefore, under without the assistance of hyper-parameters, how to acquire high-quality representations becomes the key. Inspired by this, we adopt two types of anchors, view-related and view-unrelated, to concurrently mine exclusive characteristics and common characteristics among views. Then, all anchors' information is gathered together via a consensus bipartite graph. By such ways, NpGC extracts both complementary and consistent multi-view features, thereby obtaining superior clustering results. Also, linear complexities enable it to handle datasets with over 120000 samples. Numerous experiments reveal NpGC's strong points compared to lots of classical approaches.
Shengju Yu, Siwei Wang 0001, Zhibin Dong, Wenxuan Tu, Suyuan Liu, Zhao Lv, En Zhu
AAAI6
2024 Automated Detection of Epileptic Spikes and Seizures Incorporating a Novel Spatial Clustering Prior
abstract
A Magnetoencephalography (MEG) time-series recording consists of multi-channel signals collected by superconducting sensors, with each signal’s intensity reflecting magnetic field changes over time at the sensor location. Automating epileptic MEG spike detection significantly reduces manual assessment time and effort, yielding substantial clinical benefits. Existing research addresses MEG spike detection by encoding neural network inputs with signals from all channel within a time segment, followed by classification. However, these methods overlook simultaneous spiking occurred from nearby sensors. We introduce a simple yet effective paradigm that first clusters MEG channels based on their sensor’s spatial position. Next, a novel convolutional input module is designed to integrate the spatial clustering and temporal changes of the signals. This module is fed into a custom MEEG-ResNet3D developed by the authors, which learns to extract relevant features and classify the input as a spike clip or not. Our method achieves an F1 score of 94.73% on a large real-world MEG dataset Sanbo-CMR collected from two centers, outperforming state-of-the-art approaches by 1.85%. Moreover, it demonstrates efficacy and stability in the Electroencephalographic (EEG) seizure detection task, yielding an improved weighted F1 score of 1.4% compared to current state-of-the-art techniques evaluated on TUSZ, whch is the largest EEG seizure dataset.
Hanyang Dong, Shurong Sheng, Xiongfei Wang, Jia-Hong Gao 0001, Kuntao Xiao, Pengfei Teng, Guoming Luan, Zhao Lv
BIBM10
2024 Integrating Low-order and High-order Functional Connectivity for Meta-stable State Transition based Brain Disorder Identification
abstract
Dynamic functional connectivity network (FCN) can effectively mine meta-stable state transition within the period of data acquisition time, which is related to neurological diseases. However, conventional FCN directly describes the correlation between two brain regions in a meta-stable state, it is low-order FCN. In fact, the connection between two brain regions within the meta-stable state also has a changing pattern, which can reveal the functional consistency between two connections over time, we denote the changing pattern between two regions as high-order FCN. Here, we propose an end-to-end method that integrates low-order and high-order dynamic FCNs for better brain disorder identification. First, a sliding window operation is adopted to capture meta-stable states. Then, a matrix variate normal distribution based approach is employed to construct the low-order and high-order FCNs for each meta-stable state. Finally, a two-stage Transformer is designed to extract meta-stable state transition feature for classification. Experimental results on the ADHD-200 and ABIDE datasets indicate that: 1) our proposed method integrate multi-level connectivity information of dynamic brain functions, thereby effectively improving the identification of brain disorders; 2) the proposed two-stage Transformer is more effective in feature extraction than the well-known CNN-LSTM architecture; 3) high-order FCN help locate biomarkers that low-order FCN cannot be determine, including brain regions as well as functional connections between brain regions, which contribute significantly to the diagnosis of brain disorders.
Shengbing Pei, Zhao Lv, Chao Zhang 0047, Jihong Guan
BIBM4
2024 DL2G: Anatomical Landmark Detection with Deep Local Features and Geometric Global Constraint
abstract
Anatomical landmark detection, a pivotal research area in medical image processing, holds immense value in surgical navigation, image registration, and related fields. Traditional machine learning methods struggle with generalization and robustness. Current supervised end-to-end approaches lacks in-terpretability in capturing global information, and GPU memory constraints restrict their application to extensive 3D medical image datasets. Moreover, existing methods overlook intrinsic geometric cues within images and point sets. Herein, we introduce a novel anatomical landmark detection framework that integrates deep learning’s representation capabilities with global geometric information derived from images and point sets (DL2G). This approach facilitates anatomical landmark localization in a local-to-global fashion. Initially, we train a deep feature descriptor based on self-supervised contrastive learning, which is further used to screen candidate points for the target one by comparing local patch embeddings. Subsequently, geometric constraints constructed from labeled template point sets enable the selection of geometrically consistent matching point sets. Ultimately, the framework utilizes the global mutual information of images to perform landmark localization through iterative optimization. Adhering to the AFID data annotation protocol, we evaluated our method on two public MRI head datasets, OASIS and HCP, for detecting 32 anatomical landmarks. The results convincingly demonstrate the superiority of our method over current state-of-the-art approaches in terms of both accuracy and computational efficiency.
Kuntao Xiao, Shurong Sheng, Zhao Lv, Jia-Hong Gao 0001
BIBM6
2024 Pruning via Merging: Compressing LLMs via Manifold Alignment Based Layer Merging
abstract
Deyuan Liu, Zhanyue Qin, Hairu Wang, Zhao Yang, Zecheng Wang, Fangying Rong, Qingbin Liu, Yanchao Hao, Bo Li, Xi Chen, Cunhang Fan, Zhao Lv, Dianhui Chu, Zhiying Tu, Dianbo Sui. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Deyuan Liu, Zhanyue Qin, Hairu Wang 0002, Zhao Yang 0004, Zecheng Wang, Fangying Rong, Qingbin Liu, Yanchao Hao, Xi Chen 0003, Cunhang Fan, Zhao Lv, Zhiying Tu, Dianbo Sui
EMNLP12
2024 UNO Arena for Evaluating Sequential Decision-Making Capability of Large Language Models
abstract
Zhanyue Qin, Haochuan Wang, Deyuan Liu, Ziyang Song, Cunhang Fan, Zhao Lv, Jinlin Wu, Zhen Lei, Zhiying Tu, Dianhui Chu, Xiaoyan Yu, Dianbo Sui. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Zhanyue Qin, Deyuan Liu, Cunhang Fan, Zhao Lv, Jinlin Wu, Zhen Lei 0001, Zhiying Tu, Dianbo Sui
EMNLP6
2024 Dual-View Multimodal Interaction in Multimodal Sentiment Analysis
abstract
Outstanding performance in sentiment analysis not only relies on the design of sophisticated fusion methods but also on the crucial step of designing excellent modal interaction methods. To the best of our knowledge, there are few methods addressing the capture of multimodal spatial features. Majority of feature interactions have been primarily focused on temporal aspects, with less attention given to the combined spatiotemporal feature interaction (SFI). In this paper, we design a dual-view multimodal interaction method, named DVMI, primarily consisting of two parts. In the first part, a triangular convolutional module is proposed for ample temporal interaction between modalities, implicit local and global SFI, and capturing global spatial representations. Building upon the foundation laid in the first part, the second part employs an attention mechanism for explicit global SFI. To demonstrate the effectiveness of the DVMI framework,we conduct extensive experiments on three datasets, achieving state-of-the-art experimental results.
Kang Zhu, Cunhang Fan, Jianhua Tao 0001, Jun Xue 0001, Xuefei Liu, Zhengqi Wen, Zhao Lv
ICME9
2024 DBPNet: Dual-Branch Parallel Network with Temporal-Frequency Fusion for Auditory Attention Detection
Qinke Ni, Cunhang Fan, Shengbing Pei, Zhao Lv
IJCAI6
2024 A Debiased Domain Adaptation Framework with Minimum Class Confusion for Motor Imagery Decoding
abstract
Recently, motor imagery decoding technology based on electroencephalogram (EEG) signals has made significant progress. However, there are still challenges in adapting to new sessions, mainly due to changes in data distribution between different sessions and confusion problems caused by similar oscillation patterns in various categories of EEG signals. To address these problems, this paper proposes a debiased domain adaptation framework with minimum class confusion to learn unbiased representations in motor imagery tasks. Specially, unlike the feature alignment and adversarial training methods, we explore the class predictions for domain adaptation, applying the minimum class confusion loss criterion in the target domain to reduce inter-class confusion. This approach aims to alleviate the bias issues inherent in classifiers trained on the source domain for making predictions in the target domain, achieving class-level alignment. Consequently, it enhances the model’s ability to adapt to the data distribution of the target domain. Experimental results on two public EEG datasets (BCI Competition IV datasets IIa and IIb) show that the method for cross-session decoding is significantly improved compared to the baseline, with average classification accuracy reaching 81.01% and 82.92%, respectively.
Cunhang Fan, Zhen Chen 0022, Xun Song, Jun Xue 0001, Ping Li 0020, Zhao Lv
IJCNN7
2024 Personal Identification and Authentication in Multi-Task EEG Database Using EEGNet and Siamese Network
abstract
Currently, there is a growing global concern regarding data privacy and security, particularly in the field of personal identification and authentication. Traditional biometric identification technologies are highly favored for their ease of use and high accuracy, but they fall short in ensuring liveness detection, making them susceptible to deception and forgery threats. This study focuses on personal identification and authentication based on a multi-task electroencephalogram (EEG) database, proposing an innovative model framework for these purposes. To validate the effectiveness of this model, we established a multitask EEG database containing data from 24 subjects engaged in five mental tasks. Each subject underwent four sessions, with each session consisting of 125 trials, and session intervals ranging from days to months. In this framework, we employed the EEGNet model for personal identification. It directly utilized preprocessed EEG data as input, mapping input signals to a new embedding space to extract identity features, ultimately achieving accurate individual personal identification. For the personal authentication module, we proposed the SiamEEGNet model, combining concepts from EEGNet and Siamese networks. This model comprised two EEGNet sub-networks with identical model parameters. Unlike the personal identification module, we removed the classification module from the EEGNet model in the SiamEEGNet model. Instead, we introduced a distance measurement layer to calculate the distance or similarity between the outputs of the two sub-networks, thereby achieving personal authentication. We conducted experiments for both personal identification and authentication. In the identification experiments, the proposed EEGNet model demonstrated outstanding performance with an average recognition accuracy of 99.84%. The model achieved a balance between precision and sensitivity, as reflected in high F1 score values. In the personal authentication experiments, the new SiamEEGNet model achieved a False Rejection Rate (FRR) of 2.29% and a False Acceptance Rate (FAR) of 4.75%. The experimental results collectively indicate the significant effectiveness of the proposed model framework.
Rui Ouyang, Xiaopei Wu, Zhao Lv
IJCNN3
2024 RawBMamba: End-to-End Bidirectional State Space Model for Audio Deepfake Detection
abstract
Fake artefacts for discriminating between bonafide and fake audio can exist in both short-and long-range segments.Therefore, combining local and global feature information can effectively discriminate between bonafide and fake audio.This paper proposes an end-to-end bidirectional state space model, named RawBMamba, to capture both short-and long-range discriminative information for audio deepfake detection.Specifically, we use sinc Layer and multiple convolutional layers to capture short-range features, and then design a bidirectional Mamba to address Mamba's unidirectional modelling problem and further capture long-range feature information.Moreover, we develop a bidirectional fusion module to integrate embeddings, enhancing audio context representation and combining shortand long-range information.The results show that our proposed RawBMamba achieves a 34.1% improvement over Rawformer on ATSVspoof2021 LA dataset, and demonstrates competitive performance on other datasets.Codes will be released on https://github.com/cyjie429/RawBMamba.
Yujie Chen 0006, Jiangyan Yi, Jun Xue 0001, Chenglong Wang 0001, Xiaohui Zhang 0006, Shunbo Dong, Siding Zeng, Jianhua Tao 0001, Zhao Lv, Cunhang Fan
INTERSPEECH9
2024 Frequency-mix Knowledge Distillation for Fake Speech Detection
Cunhang Fan, Shunbo Dong, Jun Xue 0001, Yujie Chen 0006, Jiangyan Yi, Zhao Lv
INTERSPEECH6
2024 Prompt Link Multimodal Fusion in Multimodal Sentiment Analysis
Kang Zhu, Cunhang Fan, Jianhua Tao 0001, Zhao Lv
INTERSPEECH4
2024 MSFNet: Multi-Scale Fusion Network for Brain-Controlled Speaker Extraction
abstract
Speaker extraction aims to selectively extract the target speaker from the multi-talker environment under the guidance of auxiliary reference. Recent studies have shown that the attended speaker's information can be decoded by the auditory attention decoding from the listener's brain activity. However, how to more effectively utilize the common information about the target speaker contained in both electroencephalography (EEG) and speech is still an unresolved problem. In this paper, we propose a multi-scale fusion network (MSFNet) for brain-controlled speaker extraction, which utilizes the EEG recorded from the listener to extract the target speech. In order to make full use of the speech information, the mixed speech is encoded with multiple time scales so that the multi-scale embeddings are acquired. In addition, to effectively extract the non-Euclidean data of EEG, the graph convolutional networks are used as the EEG encoder. Finally, these multi-scale embeddings are separately fused with the EEG features. To facilitate research related to auditory attention decoding and further validate the effectiveness of the proposed method, we also construct the AVED dataset, a new EEG-Audio dataset. Experimental results on both the public Cocktail Party dataset and the newly proposed AVED dataset in this paper show that our MSFNet model significantly outperforms the state-of-the-art method in certain objective evaluation metrics.
Cunhang Fan, Wang Xiang, Jianhua Tao 0001, Jiangyan Yi, Dianbo Sui, Zhao Lv
ACM Multimedia9
2024 DARNet: Dual Attention Refinement Network with Spatiotemporal Construction for Auditory Attention Detection
abstract
At a cocktail party, humans exhibit an impressive ability to direct their attention. The auditory attention detection (AAD) approach seeks to identify the attended speaker by analyzing brain signals, such as EEG signals. However, current AAD algorithms overlook the spatial distribution information within EEG signals and lack the ability to capture long-range latent dependencies, limiting the model's ability to decode brain activity. To address these issues, this paper proposes a dual attention refinement network with spatiotemporal construction for AAD, named DARNet, which consists of the spatiotemporal construction module, dual attention refinement module, and feature fusion \& classifier module. Specifically, the spatiotemporal construction module aims to construct more expressive spatiotemporal feature representations, by capturing the spatial distribution characteristics of EEG signals. The dual attention refinement module aims to extract different levels of temporal patterns in EEG signals and enhance the model's ability to capture long-range latent dependencies. The feature fusion \& classifier module aims to aggregate temporal patterns and dependencies from different levels and obtain the final classification results. The experimental results indicate that DARNet achieved excellent classification performance, particularly under short decision windows. While maintaining excellent classification performance, DARNet significantly reduces the number of required parameters. Compared to the state-of-the-art models, DARNet reduces the parameter count by 91\%. Code is available at: https://github.com/fchest/DARNet.git.
Cunhang Fan, Xiaoke Yang, Jianhua Tao 0001, Zhao Lv
NeurIPS6
2024 Light-weight residual convolution-based capsule network for EEG emotion recognition
Cunhang Fan, Jinqin Wang, Xiaoke Yang, Guanxiong Pei, Taihao Li, Zhao Lv
Adv. Eng. Informatics7
2024 VLP2MSA: Expanding vision-language pre-training to multimodal sentiment analysis
Guofeng Yi, Cunhang Fan, Kang Zhu, Zhao Lv, Shan Liang 0007, Zhengqi Wen, Guanxiong Pei, Taihao Li, Jianhua Tao 0001
Knowl. Based Syst.4
2024 Spatial reconstructed local attention Res2Net with F0 subband for fake speech detection
Cunhang Fan, Jun Xue 0001, Jianhua Tao 0001, Jiangyan Yi, Chenglong Wang 0001, Chengshi Zheng, Zhao Lv
Neural Networks7
2024 DGSD: Dynamical graph self-distillation for EEG-based auditory spatial attention detection
Cunhang Fan, Jun Xue 0001, Jianhua Tao 0001, Jiangyan Yi, Zhao Lv, Xiaopei Wu
Neural Networks7
2024 Multimodal Cross-Lingual Summarization for Videos: A Revisit in Knowledge Distillation Induced Triple-Stage Training Method
abstract
Multimodal summarization (MS) for videos aims to generate summaries from multi-source information (e.g., video and text transcript), showing promising progress recently. However, existing works are limited to monolingual scenarios, neglecting non-native viewers' needs to understand videos in other languages. It stimulates us to introduce multimodal cross-lingual summarization for videos (MCLS), which aims to generate cross-lingual summaries from multimodal input of videos. Considering the challenge of high annotation cost and resource constraints in MCLS, we propose a knowledge distillation (KD) induced triple-stage training method to assist MCLS by transferring knowledge from abundant monolingual MS data to those data with insufficient volumes. In the triple-stage training method, a video-guided dual fusion network (VDF) is designed as the backbone network to integrate multimodal and cross-lingual information through diverse fusion strategies in the encoder and decoder; What's more, we propose two cross-lingual knowledge distillation strategies: adaptive pooling distillation and language-adaptive warping distillation (LAWD), designed for encoder-level and vocab-level distillation objects to facilitate effective knowledge transfer across cross-lingual sequences of varying lengths between MS and MCLS models. Specifically, to tackle lingual sequences of varying lengths between MS and MCLS models. Specifically, to tackle the challenge of unequal length of parallel cross-language sequences in KD, LAWD can directly conduct cross-language distillation while keeping the language feature shape unchanged to reduce potential information loss. We meticulously annotated the How2-MCLS dataset based on the How2 dataset to simulate MCLS scenarios. Experimental results show that the proposed method achieves competitive performance compared to strong baselines, and can bring substantial performance improvements to MCLS models by transferring knowledge from the MS model.
Nayu Liu, Kaiwen Wei, Yong Yang 0001, Jianhua Tao 0001, Xian Sun 0001, Fanglong Yao, Li Jin 0001, Zhao Lv, Cunhang Fan
IEEE Trans. Pattern Anal. Mach. Intell.9
2024 Dynamic Ensemble Teacher-Student Distillation Framework for Light-Weight Fake Audio Detection
abstract
In recent years, fake audio detection (FAD) has made great progress, and lightweight is important to achieve fast and reliable audio authenticity verification on resource-limited devices. However, most of the researchers ignore lightweight when improving the performance of FAD. To develop the application of FAD for small-end devices, this paper proposes a novel light-weight network named Light-ECA2Net. Given that networks with different depths have different abilities in capturing fake speech artifacts, this paper proposes a dynamic ensemble teacher-student distillation framework to fully transfer distillation knowledge. The dynamic ensemble distillation is divided into two aspects. First, we adopt one-to-one feature mapping to perceive the multidimensional feature knowledge and dynamically adjust every dimension feature weight by using ground truth labels, which can enable students to receive feature knowledge efficiently. Secondly, different network layers also have their strengths of predicting, further dynamically predicting weight can improve the learning ability of the student. Experimental results on the ASVspoof 2019 LA and PA datasets show that compared to the baseline, our system further improves performance by reducing the model complexity by 45%.
Jun Xue 0001, Cunhang Fan, Jiangyan Yi, Jian Zhou 0006, Zhao Lv
IEEE Signal Process. Lett.5
2024 Dual-Branch Knowledge Distillation for Noise-Robust Synthetic Speech Detection
abstract
Most research in synthetic speech detection (SSD) focuses on improving performance on standard noise-free datasets. However, in actual situations, noise interference is usually present, causing significant performance degradation in SSD systems. To improve noise robustness, this paper proposes a dual-branch knowledge distillation synthetic speech detection (DKDSSD) method. Specifically, a parallel data flow of the clean teacher branch and the noisy student branch is designed, and interactive fusion module and response-based teacher-student paradigms are proposed to guide the training of noisy data from both the data distribution and decision-making perspectives. In the noisy student branch, speech enhancement is introduced initially for denoising, aiming to reduce the interference of strong noise. The proposed interactive fusion combines denoised features and noisy features to mitigate the impact of speech distortion and ensure consistency with the data distribution of the clean branch. The teacher-student paradigm maps the student's decision space to the teacher's decision space, enabling noisy speech to behave similarly to clean speech. Additionally, a joint training method is employed to optimize both branches for achieving global optimality. Experimental results based on multiple datasets demonstrate that the proposed method performs effectively in noisy environments and maintains its performance in cross-dataset experiments. Source code is available athttps://github.com/fchest/DKDSSD.
Cunhang Fan, Mingming Ding, Jianhua Tao 0001, Ruibo Fu, Jiangyan Yi, Zhengqi Wen, Zhao Lv
IEEE ACM Trans. Audio Speech Lang. Process.7
2023 Learning From Yourself: A Self-Distillation Method For Fake Speech Detection
abstract
In this paper, we propose a novel self-distillation method for fake speech detection (FSD), which can significantly improve the performance of FSD without increasing the model complexity. For FSD, some fine-grained information is very important, such as spectrogram defects, mute segments, and so on, which are often perceived by shallow networks. However, shallow networks have much noise, which can not capture this very well. To address this problem, we propose using the deepest network instruct shallow network for enhancing shallow networks. Specifically, the networks of FSD are divided into several segments, the deepest network being used as the teacher model, and all shallow networks become multiple student models by adding classifiers. Meanwhile, the distillation path between the deepest network feature and shallow network features is used to reduce the feature difference. A series of experimental results on the ASVspoof 2019 LA and PA datasets show the effectiveness of the proposed method, with significant improvements compared to the baseline.
Jun Xue 0001, Cunhang Fan, Jiangyan Yi, Chenglong Wang 0001, Zhengqi Wen, Dan Zhang 0014, Zhao Lv
ICASSP7
2023 Arrow: Capture the Inaudible Attacker in 3D Space via Smart-speaker
abstract
Recent works have shown that inaudible signals (at ultrasound frequencies) can become audible to the microphone by exploiting the nonlinear effects. With a well-designed inaudible signal, an adversary can control Amazon Echo and Google Homelike devices in people’s rooms silently and remotely. A voice command like “Alexa, open the door“ can be a serious treat. Although recent works design various methods against such inaudible attacks, one important issue remains open: there is no clear solution to locate the attack source accurately. Obviously, the only way to completely eliminate such inaudible threats is to locate and remove the attack source. This paper is an attempt to close this gap. We propose Arrow, an effective method to help users locate the ultrasound attack source in 3D space indoors. Arrow establishes the relationship between inaudible signals and the recorded sounds of the microphone, and then explores the architecture of the embedded microphone array on smart speaker for extracting a 3D direction-specific signature. By learning such directional signature, Arrow can accurately estimate the spatial orientation of the inaudible attack source and help users to locate and remove it. We implement a prototype of Arrow and conduct comprehensive experiments to validate its performance. The results show Arrow can achieve 2.5° and 7° error in DoA(Direction of Arrival) estimation for horizontal and vertical angles, respectively.
Zhenfei Zhang, Ping Li 0020, Biaokai Zhu, Tao Wu 0011, Panlong Yang, Zhao Lv
MSN6
2023 Packet rank-aware active queue management for programmable flow scheduling
Ziyong Li, Yuxiang Hu 0004, Le Tian 0002, Zhao Lv
Comput. Networks4
2023 CompNet: Complementary network for single-channel speech enhancement
Cunhang Fan, Andong Li, Wang Xiang, Chengshi Zheng, Zhao Lv, Xiaopei Wu
Neural Networks6
2023 Subband fusion of complex spectrogram for fake speech detection
Cunhang Fan, Jun Xue 0001, Shunbo Dong, Mingming Ding, Jiangyan Yi, Jinpeng Li 0002, Zhao Lv
Speech Commun.7
2022 Csenet: Complex Squeeze-and-Excitation Network for Speech Depression Level Prediction
abstract
Automatic speech depression level prediction (SDLP) is a very challenging problem in affective computing. There are many studies that have acquired quite good performances for SDLP. However, most of the input speech features of these studies are based on the amplitude spectrogram, which loses the phase spectrogram information. Therefore, these speech features may lose some important information related to depression. In order to make full use of speech information, this paper proposes a complex squeeze-and-excitation network (CSENet) for SDLP. The complex spectrogram is used as the input speech feature, which contains both amplitude and phase spectrogram. In addition, to acquire a discriminative feature, the squeeze-and-excitation residual network is employed to extract deep speech feature. Finally, the attentive temporal pooling is utilized to dynamically select more important information according to the attention mechanisms. Experimental results on the AVEC 2013 and AVEC 2014 datasets prove the effectiveness of our proposed method. As for the mean absolute error (MAE) evaluation metric on AVEC 2013, our proposed method acquires state-of-the-art performance.
Cunhang Fan, Zhao Lv, Shengbing Pei, Mingyue Niu
ICASSP2
2021 Research of Robust Video Object Tracking Algorithm Based on Jetson Nano Embedded Platform
Chao Zhang 0047, Zhao Lv
PRCV (1)3
2020 An improved SIFT algorithm for robust emotion recognition under various face poses and illuminations
Zhao Lv, Ning Bi, Chao Zhang 0047
Neural Comput. Appl.2
2020 To Explore the Potentials of Independent Component Analysis in Brain-Computer Interface of Motor Imagery
abstract
This paper is focused on the experimental approach to explore the potential of independent component analysis (ICA) in the context of motor imagery (MI)-based brain-computer interface (BCI). We presented a simple and efficient algorithmic framework of ICA-based MI BCI (ICA-MIBCI) for the evaluation of four classical ICA algorithms (Infomax, FastICA, Jade, and Sobi) as well as a simplified Infomax (sInfomax). Two novel performance indexes, self-test accuracy and the number of invalid ICA filters, were employed to assess the performance of MIBCI based on different ICA variants. As a reference method, common spatial pattern (CSP), a commonly-used spatial filtering method, was employed for the comparative study between ICA-MIBCI and CSP-MIBCI. The experimental results showed that sInfomax-based spatial filters exhibited significantly better transferability in session to session and subject to subject transfer as compared to CSP-based spatial filters. The online experiment was also introduced to demonstrate the practicability and feasibility of sInfomax-based MIBCI. However, four classical ICA variants, especially FastICA, Jade, and Sobi, performed much worse as compared to sInfomax and CSP in terms of classification accuracy and stability. We consider that conventional ICA-based spatial filtering methods tend to be overfitting while applied to real-life electroencephalogram data. Nevertheless, the sInfomax-based experimental results indicate that ICA methods have a great space for improvement in the application of MIBCI. We believe that this paper could bring forth new ideas for the practical implementation of ICA-MIBCI.
Xiaopei Wu, Bangyan Zhou, Zhao Lv, Chao Zhang 0047
IEEE J. Biomed. Health Informatics3
2018 Design and implementation of an eye gesture perception system based on electrooculography
Zhao Lv, Chao Zhang 0047, Bangyan Zhou, Xiangping Gao, Xiaopei Wu
Expert Syst. Appl.1
2017 A permutation algorithm based on dynamic time warping in speech frequency-domain blind source separation
Zhao Lv, Xiaopei Wu, Chao Zhang 0047, Bangyan Zhou
Speech Commun.1
2010 A novel eye movement detection algorithm for EOG driven human computer interface
Zhao Lv, Xiaopei Wu, Dexiang Zhang
Pattern Recognit. Lett.1