Gaoyan Zhang

dblp:85/5661 · DBLP profile ↗
← Back
24ranked-venue papers
0as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Improved ship-radiated noise recognition using a pre-trained noise reduction module combined with feature optimization network
Yunye Feng, Haonan Xing, Meng Ge, Tianrui Wang, Lin Gan 0003, Gaoyan Zhang
Knowl. Based Syst.8
2025 Multi-level disparity-guided transformers for light field spatial super-resolution
Yigeng Liao, Gaoyan Zhang, Fengzhou Fang
Pattern Recognit.3
2024 CEDNet: A Continuous Emotion Detection Network for Naturalistic Stimuli Using MEG Signals
abstract
Emotional detection is important for brain-computer interface or diagnosis of affective disorders. Traditional methods mainly focused on the recognition of brief stimuli evoked emotion, which cannot fully represent the complexity of real-life emotional changes. In this study, we proposed a continuous emotion detection network (CEDNet) based on magnetoencephalography (MEG) to detect the time-varying emotions evoked by a 2-hour movie. Two brain graphs constructed by functional connectivity and spatial location of brain regions were input as two views of the brain. An adaptive spatio-temporal graph convolutional network with an attention mechanism was adopted to extract the emotion-related high-level features. Considering the impact of unbalanced emotion labels, a label distribution smoother was introduced. Furthermore, we added a domain discriminator to enhance the generalization capability of the model. Experimental results show that the proposed model outperforms the state-of-the-art baselines and provides a deep insight into the human emotional processes.
Zeming He, Gaoyan Zhang
ICASSP2
2024 EEG-Based Fast Auditory Attention Detection in Real-Life Scenarios Using Time-Frequency Attention Mechanism
abstract
Auditory attention detection (AAD) based on electroencephalogram (EEG) helps recognize the target speaker in a cocktail party scenario, advancing auditory brain-computer interface development. Previous EEG studies on AAD were largely based on data collected in laboratory settings. In this study, we investigated the AAD with EEG data collected when subjects were walking and sitting in real-life scenarios. To improve the detection accuracy, we proposed the time-frequency attention mechanism to the convolution neural network on EEG data. Experimental results show that the proposed model outperforms the state-of-the-art models, with an accuracy of 98.1% on a decision window of 2s. When we used a 0.1s time window for fast decoding, the accuracy remained at 91.8%, suggesting the potential for real application. Further study on ablation experiments demonstrates the effectiveness of the proposed time-frequency attention mechanism. Analysis of the key EEG features indicates that the β band plays a vital role in AAD.
Zhuang Xie, Jianguo Wei, Wenhuan Lu, Chunli Wang, Gaoyan Zhang
ICASSP6
2024 MBCFNet: A Multimodal Brain-Computer Fusion Network for human intention recognition
Gaoyan Zhang, Shogo Okada, Longbiao Wang, Jianwu Dang 0001
Knowl. Based Syst.2
2024 Enhancing Major Depressive Disorder Diagnosis With Dynamic-Static Fusion Graph Neural Networks
abstract
Major Depressive Disorder (MDD) is a debilitating, complex mental condition with unclear mechanisms hindering diagnostic progress. Research links MDD to abnormal brain connectivity using functional magnetic resonance imaging (fMRI). Yet, existing fMRI-based MDD models suffer from limitations, including neglecting dynamic network traits, lacking interpretability, and struggling with small datasets. We present DSFGNN, a novel graph neural network framework addressing these issues for improved MDD diagnosis. DSFGNN employs a graph isomorphism encoder to model static and dynamic brain networks, achieving effective fusion of temporal and spatial information through a spatiotemporal attention mechanism, thereby enhancing interpretability. Furthermore, we incorporate a causal disentangling module and orthogonal regularization module to augment the model's expressiveness. We evaluate DSFGNN on the Rest-meta-MDD dataset, yielding superior results compared to the best baseline. Besides, extensive ablation studies and interpretability analysis confirm DSFGNN's effectiveness and potential for biomarker discovery.
Tianyi Zhao 0007, Gaoyan Zhang
IEEE J. Biomed. Health Informatics2
2023 Brain Network Features Differentiate Intentions from Different Emotional Expressions of the Same Text
abstract
Intent differentiation in speech communication relies not only on linguistic information but also on paralinguistic information. The same textual content, when pronounced with different prosodies and emotions, may express totally different intentions. The true intentions in this condition can be easily grasped by our brain. Therefore, combining text, speech, and electroencephalography (EEG) for intent discrimination on the same text may be an effective approach. Before fusing speech and text modalities, the current study focused on exploring effective EEG-based features for Chinese intent recognition as no previous research has utilized EEG signals for this purpose. To tackle this issue, we first created a Chinese multimodal spoken language intention understanding (CMSLIU) dataset, in which the same texts were pronounced with varying prosodies to express different intents. To identify effective brain features that were most relevant to intent recognition improvement, we compared the event-related spectral perturbation and effective brain connectivity patterns on two intent conditions (praise vs. irony). It was found that the praise expression tended to elicit stronger high-frequency brain activities while the irony expression involved a more suppressive network connection in the right hemisphere. These features were trained on the CMSLIU dataset and achieved an intention classification accuracy of 78.66%, which indicated a great potential of the EEG features in intent discrimination on the same text.
Gaoyan Zhang, Jianwu Dang 0001
ICASSP3
2023 Locate and Beamform: Two-dimensional Locating All-neural Beamformer for Multi-channel Speech Separation
Yanjie Fu, Meng Ge, Honglong Wang, Longbiao Wang, Gaoyan Zhang, Jianwu Dang 0001, Chengyun Deng
INTERSPEECH7
2023 Discrimination of the Different Intents Carried by the Same Text Through Integrating Multimodal Information
Gaoyan Zhang, Longbiao Wang, Jianwu Dang 0001
INTERSPEECH2
2023 SDNet: Stream-attention and Dual-feature Learning Network for Ad-hoc Array Speech Separation
Honglong Wang, Chengyun Deng, Yanjie Fu, Meng Ge, Longbiao Wang, Gaoyan Zhang, Jianwu Dang 0001
INTERSPEECH6
2023 Auditory Attention Detection in Real-Life Scenarios Using Common Spatial Patterns from EEG
Zhuang Xie, Di Zhou 0008, Longbiao Wang, Gaoyan Zhang
INTERSPEECH5
2022 A Bi-hemisphere Capsule Network Model for Cross-Subject EEG Emotion Recognition
Xueying Luan, Gaoyan Zhang
ICONIP (6)2
2022 An Improved Stimulus Reconstruction Method for EEG-Based Short-Time Auditory Attention Detection
Gaoyan Zhang, Masashi Unoki, Jianwu Dang 0001, Longbiao Wang
ICONIP (5)3
2022 Detecting Major Depressive Disorder by Graph Neural Network Exploiting Resting-State Functional MRI
Gaoyan Zhang
ICONIP (5)2
2022 Iterative Sound Source Localization for Unknown Number of Sources
abstract
Sound source localization aims to seek the direction of arrival (DOA) of all sound sources from the observed multichannel audio.For the practical problem of unknown number of sources, existing localization algorithms attempt to predict a likelihood-based coding (i.e., spatial spectrum) and employ a pre-determined threshold to detect the source number and corresponding DOA value.However, these threshold-based algorithms are not stable since they are limited by the careful choice of threshold.To address this problem, we propose an iterative sound source localization approach called ISSL, which can iteratively extract each source's DOA without threshold until the termination criterion is met.Unlike threshold-based algorithms, ISSL designs an active source detector network based on binary classifier to accept residual spatial spectrum and decide whether to stop the iteration.By doing so, our ISSL can deal with an arbitrary number of sources, even more than the number of sources seen during the training stage.The experimental results show that our ISSL achieves significant performance improvements in both DOA estimation and source number detection compared with the existing threshold-based algorithms.
Yanjie Fu, Meng Ge, Xinyuan Qian 0001, Longbiao Wang, Gaoyan Zhang, Jianwu Dang 0001
INTERSPEECH6
2022 MIMO-DoAnet: Multi-channel Input and Multiple Outputs DoA Network with Unknown Number of Sound Sources
abstract
Recent neural network based Direction of Arrival (DoA) estimation algorithms have performed well on unknown number of sound sources scenarios.These algorithms are usually achieved by mapping the multi-channel audio input to the single output (i.e.overall spatial pseudo-spectrum (SPS) of all sources), that is called MISO.However, such MISO algorithms strongly depend on empirical threshold setting and the angle assumption that the angles between the sound sources are greater than a fixed angle.To address these limitations, we propose a novel multi-channel input and multiple outputs DoA network called MIMO-DoAnet.Unlike the general MISO algorithms, MIMO-DoAnet predicts the SPS coding of each sound source with the help of the informative spatial covariance matrix.By doing so, the threshold task of detecting the number of sound sources becomes an easier task of detecting whether there is a sound source in each output, and the serious interaction between sound sources disappears during inference stage.Experimental results show that MIMO-DoAnet achieves relative 18.6% and absolute 13.3%, relative 34.4% and absolute 20.2% F1 score improvement compared with the MISO baseline system in 3, 4 sources scenes.The results also demonstrate MIMO-DoAnet alleviates the threshold setting problem and solves the angle assumption problem effectively.
Meng Ge, Yanjie Fu, Gaoyan Zhang, Longbiao Wang, Jianwu Dang 0001
INTERSPEECH4
2022 Constructing Accurate and Efficient Deep Spiking Neural Networks With Double-Threshold and Augmented Schemes
abstract
Spiking neural networks (SNNs) are considered as a potential candidate to overcome current challenges, such as the high-power consumption encountered by artificial neural networks (ANNs); however, there is still a gap between them with respect to the recognition accuracy on various tasks. A conversion strategy was, thus, introduced recently to bridge this gap by mapping a trained ANN to an SNN. However, it is still unclear that to what extent this obtained SNN can benefit both the accuracy advantage from ANN and high efficiency from the spike-based paradigm of computation. In this article, we propose two new conversion methods, namely TerMapping and AugMapping. The TerMapping is a straightforward extension of a typical threshold-balancing method with a double-threshold scheme, while the AugMapping additionally incorporates a new scheme of augmented spike that employs a spike coefficient to carry the number of typical all-or-nothing spikes occurring at a time step. We examine the performance of our methods based on the MNIST, Fashion-MNIST, and CIFAR10 data sets. The results show that the proposed double-threshold scheme can effectively improve the accuracies of the converted SNNs. More importantly, the proposed AugMapping is more advantageous for constructing accurate, fast, and efficient deep SNNs compared with other state-of-the-art approaches. Our study, therefore, provides new approaches for further integration of advanced techniques in ANNs to improve the performance of SNNs, which could be of great merit to applied developments with spike-based neuromorphic computing.
Qiang Yu 0005, Chenxiang Ma, Shiming Song 0001, Gaoyan Zhang, Jianwu Dang 0001, Kay Chen Tan
IEEE Trans. Neural Networks Learn. Syst.4
2021 Simultaneous Progressive Filtering-Based Monaural Speech Enhancement
Longbiao Wang, Luya Qiang, Sheng Li 0010, Meng Ge, Gaoyan Zhang, Jianwu Dang 0001
ICONIP (5)7
2021 Multi-Modal Emotion Recognition Based On deep Learning Of EEG And Audio Signals
abstract
Automatic recognition of human emotional states has attracted many researchers' attention in Human-Computer Interactions and emotional brain-computer interface recently. However, the accuracy of emotion recognition is not satisfying. Considering the advantage of information supplement based on deep learning of multi-modal signals related to emotion, this study proposed a novel emotion recognition architecture to fuse emotional features from brain electroencephalography (EEG) signal and the corresponding audio signal in emotion recognition on DEAP dataset. We used convolutional neural network (CNN) to extract EEG features and bidirectional long short term memory (BiLSTM) neural networks to extract audio features. After that, we combine the multi-modal features into a deep learning architecture to recognize arousal and valence levels. Results showed an improved accuracy compared with previous studies that merely used the EEG signals in both arousal level and valence level, which suggests the effectiveness of our proposed multi-modal fused emotion recognition model. In future work, multi-modal data from nature interaction scenes will be collected and inputted into this architecture to further validate the effectiveness of the method.
Gaoyan Zhang, Jianwu Dang 0001, Longbiao Wang, Jianguo Wei
IJCNN2
2020 EEG-Based Short-Time Auditory Attention Detection Using Multi-Task Deep Learning
Gaoyan Zhang, Jianwu Dang 0001, Di Zhou 0008, Longbiao Wang
INTERSPEECH2
2020 Cortical Oscillatory Hierarchy for Natural Sentence Processing
Jianwu Dang 0001, Gaoyan Zhang, Masashi Unoki
INTERSPEECH3
2020 Neural Entrainment to Natural Speech Envelope Based on Subject Aligned EEG Signals
Di Zhou 0008, Gaoyan Zhang, Jianwu Dang 0001
INTERSPEECH2
2018 Revealing Spatiotemporal Brain Dynamics of Speech Production Based on EEG and Eye Movement
Gaoyan Zhang, Jianwu Dang 0001, Minbo Chen, YingjianFu, Longbiao Wang
INTERSPEECH3
2017 A Neuro-Experimental Evidence for the Motor Theory of Speech Perception
Jianwu Dang 0001, Gaoyan Zhang
INTERSPEECH3