Yonghao Song

dblp:169/7096 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0003-1700-1133ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 EEGTune: A data-efficient fine-tuning framework for EEG foundation models
abstract
Electroencephalography (EEG) foundation models, pre-trained on large-scale unlabeled data via self-supervised learning, have demonstrated strong generalization capabilities across various brain-computer interface (BCI) tasks. However, their practical deployment remains constrained by the expensive cost of data annotation required for task-specific fine-tuning. To mitigate this limitation, we propose EEGTune, a data-efficient fine-tuning framework for EEG foundation models that integrates active learning with consistency-based semi-supervised learning. Specifically, EEGTune first selects the most informative samples for expert annotation under a limited labeling budget and performs initial fine-tuning on this augmented annotated dataset. It then employs a consistency-based pseudo-labeling strategy, which enhances robustness to noise and eliminates manual confidence thresholding by leveraging prediction stability under stochastic augmentations. The model is finally fine-tuned on both the labeled and pseudo-labeled data, maximizing the use of all available samples. To evaluate the efficacy of EEGTune, we have conducted extensive experiments involving two state-of-the-art EEG foundation models across four representative downstream BCI tasks, including sleep staging, seizure detection, emotion recognition, and motor imagery. Experimental results demonstrate that our method achieves superior fine-tuning performance with significantly reduced labeling effort. Notably, on sleep staging and seizure detection tasks, EEGTune attains performance within 1-2% of fully supervised fine-tuning using only 15% of the labeled data. These results underscore the potential of our framework to substantially reduce annotation requirements while maintaining high performance, facilitating broader applications of EEG foundation models in real-world scenarios. The source code will be released at https://github.com/zer02her0/EEGTune .
Zelin Liao, Yonghao Song, Chengjian Xu, Qingqing Zheng
Pattern Recognit.2
2026 Fuzzy Alignment Resolves Visual Representations From 1024-Channel Brain Recordings
Yonghao Song, Chengjian Xu, Qingqing Zheng, Nanlin Shi, Yijun Wang 0001, Xiaorong Gao
IEEE Trans. Fuzzy Syst.1
2025 Neuro-3D: Towards 3D Visual Decoding from EEG Signals
abstract
Human’s perception of the visual world is shaped by the stereo processing of 3D information. Understanding how the brain perceives and processes 3D visual stimuli in the real world has been a longstanding endeavor in neuroscience. Towards this goal, we introduce a new neuroscience task: decoding 3D visual perception from EEG signals, a neuroimaging technique that enables real-time monitoring of neural dynamics enriched with complex visual cues. To provide the essential benchmark, we first present EEG-3D, a pioneering dataset featuring multimodal analysis data and extensive EEG recordings from 12 subjects viewing 72 categories of 3D objects rendered in both videos and images. Furthermore, we propose Neuro-3D, a 3D visual decoding framework based on EEG signals. This framework adaptively integrates EEG features derived from static and dynamic stimuli to learn complementary and robust neural representations, which are subsequently utilized to recover both the shape and color of 3D objects through the proposed diffusion-based colored point cloud decoder. To the best of our knowledge, we are the first to explore EEG-based 3D visual decoding. Experiments indicate that Neuro-3D not only reconstructs colored 3D objects with high fidelity, but also learns effective neural representations that enable insightful brain region analysis. The code and dataset are available at https://github.com/gzq17/neuro-3D.
Zhanqiang Guo, Yonghao Song, Jiahui Bu, Weijian Mai, Qihao Zheng, Wanli Ouyang, Chunfeng Song
CVPR3
2025 Unsupervised multi-source domain adaptation via contrastive learning for EEG classification
Chengjian Xu, Yonghao Song, Qingqing Zheng, Qiong Wang 0001, Pheng-Ann Heng
Expert Syst. Appl.2
2025 Frequency-Aware Spatial-Temporal Attention Explainable Network for EEG Decoding
abstract
Representation learning in spatial and temporal domains has shown significant potential in EEG decoding, advancing the field of brain-computer interfaces (BCIs). However, the critical role of frequency information, closely tied to the brain's neurological mechanism, has been largely neglected. In this paper, we propose FSTNet, which integrates frequency-spatial-temporal domains synergistically. The network allows broadband EEG signals as input and adaptively learns informative frequency signatures. A frequency-aware module emphasizes the importance of frequency information by selectively assigning weights to latent representations in the frequency space. Subsequently, self-attention captures spatial and temporal dependencies, extracting discriminative neural signatures for EEG decoding. We conducted extensive experiments on EEG datasets for motor imagery and emotion recognition, achieving superior results on SEED, PhysioNet, and OpenBMI datasets in both individual and cross-subject scenarios. Additionally, visualization reveals that the network captures informative frequency ranges and spatial patterns associated with specific tasks, aligning with known physiological mechanisms. This enhances the transparency of the network's learning process. In conclusion, our method exhibits the potential for decoding EEG and advancing the understanding of neurological processes in the brain.
Luyao Jin, Yonghao Song, Huan Zhao 0006, Junyi Cao, Vincent C. K. Cheung, Wei-Hsin Liao
IEEE J. Biomed. Health Informatics2
2025 Recognizing Natural Images From EEG With Language-Guided Contrastive Learning
abstract
Electroencephalography (EEG), known for its convenient noninvasive acquisition but moderate signal-to-noise ratio, has recently gained much attention due to the potential to decode image information. However, previous works have not delivered sufficient evidence of this task, primarily limited by performance and biological plausibility. In this work, we first introduce a self-supervised framework to demonstrate the feasibility of recognizing images from EEG signals. Contrastive learning is leveraged to align the representations of EEG responses with image stimuli. Then, language descriptions of the stimuli generated by large language models (LLMs) help guide learning core semantic information. With the framework, we attain significantly above-chance results on the THINGS-EEG2 dataset, achieving a top-1 accuracy of 19.7% and a top-5 accuracy of 51.5% in challenging 200-way zero-shot tasks. Furthermore, we conduct thorough experiments to resolve the human visual responses with EEG from temporal, spatial, spectral, and semantic perspectives. These results provide evidence of feasibility and plausibility regarding EEG-based image recognition, substantiated by comparative studies with the THINGS-Magnetoencephalography (MEG) dataset. The findings offer valuable insights for neural decoding and real-world applications of brain-computer interfaces (BCIs), such as health care and robot control. The code is available at https://github.com/eeyhsong/NICE-LLM.
Yonghao Song, Yijun Wang 0001, Huiguang He, Xiaorong Gao
IEEE Trans. Neural Networks Learn. Syst.1
2024 Decoding Natural Images from EEG for Object Recognition
abstract
Electroencephalography (EEG) signals, known for convenient non-invasive acquisition but low signal-to-noise ratio, have recently gained substantial attention due to the potential to decode natural images. This paper presents a self-supervised framework to demonstrate the feasibility of learning image representations from EEG signals, particularly for object recognition. The framework utilizes image and EEG encoders to extract features from paired image stimuli and EEG responses. Contrastive learning aligns these two modalities by constraining their similarity. Our approach achieves state-of-the-art results on a comprehensive EEG-image dataset, with a top-1 accuracy of 15.6% and a top-5 accuracy of 42.8% in 200-way zero-shot tasks. Moreover, we perform extensive experiments to explore the biological plausibility by resolving the temporal, spatial, spectral, and semantic aspects of EEG signals. Besides, we introduce attention modules to capture spatial correlations, providing implicit evidence of the brain activity perceived from EEG data. These findings yield valuable insights for neural decoding and brain-computer interfaces in real-world scenarios. Code available at https://github.com/eeyhsong/NICE-EEG.
Yonghao Song, Bingchuan Liu, Nanlin Shi, Yijun Wang 0001, Xiaorong Gao
ICLR1
2024 High-performance c-VEP-BCI under minimal calibration
abstract
The ultimate goal of brain-computer interfaces (BCIs) based on visual modulation paradigms is to achieve high-speed performance without the burden of extensive calibration. Code-modulated visual evoked potential-based BCIs (c-VEP-BCIs) modulated by broadband white noise (WN) offer various advantages, including increased communication speed, expanded encoding target capabilities, and enhanced coding flexibility. However, the complexity of the spatial-temporal patterns under broadband stimuli necessitates extensive calibration for effective target identification in c-VEP-BCIs. Consequently, the information transfer rate (ITR) of c-VEP-BCI under limited calibration usually stays around 100 bits per minute (bpm), significantly lagging behind state-of-the-art steady-state visual evoked potential-based BCIs (SSVEP-BCIs), which achieve rates above 200 bpm. To enhance the performance of c-VEP-BCIs with minimal calibration, we devised an efficient calibration stage involving a brief single-target flickering, lasting less than a minute, to extract generalizable spatial-temporal patterns. Leveraging the calibration data, we developed two complementary methods to construct c-VEP temporal patterns: the linear modeling method based on the stimulus sequence and the transfer learning techniques using cross-subject data. As a result, we achieved the highest ITR of 250 bpm under a minute of calibration, which has been shown to be comparable to the state-of-the-art SSVEP paradigms. In summary, our work significantly improved the c-VEP performance under few-shot learning, which is expected to expand the practicality and usability of c-VEP-BCIs.
Yining Miao, Nanlin Shi, Changxing Huang, Yonghao Song, Yijun Wang 0001, Xiaorong Gao
Expert Syst. Appl.4
2022 A novel hybrid CNN-Transformer model for EEG Motor Imagery classification
abstract
Motor imagery is one of the most popular brain-computer interface (BCI) paradigms with good potential to help disabled people. However, EEG decoding is still challenging because of its low signal-to-noise ratio, with which valuable features are difficult to perceive. In this paper, we propose a hybrid model that combines a convolutional neural network (CNN) with the Transformer for decoding motor imagery EEG signals. The CNN is used to extract local features, while the Transformer is utilized to perceive global dependencies. Moreover, we exploit spatial-spectral-temporal features to improve classification performance. To validate the effectiveness and superiority of the proposed method, we conduct experiments on the BCICIV dataset 2a and compare it with other efficient approaches. The results show that our algorithm outperforms the state-of-the-art methods, indicating that the CNN-Transformer model is a competing strategy.
Yaxin Ma 0001, Yonghao Song, Fei Gao 0010
IJCNN2
2021 Improving Sequential Recommendation Consistency with Self-Supervised Imitation
abstract
Most sequential recommendation models capture the features of consecutive items in a user-item interaction history. Though effective, their representation expressiveness is still hindered by the sparse learning signals. As a result, the sequential recommender is prone to make inconsistent predictions. In this paper, we propose a model, SSI, to improve sequential recommendation consistency with Self-Supervised Imitation. Precisely, we extract the consistency knowledge by utilizing three self-supervised pre-training tasks, where temporal consistency and persona consistency capture user-interaction dynamics in terms of the chronological order and persona sensitivities, respectively. Furthermore, to provide the model with a global perspective, global session consistency is introduced by maximizing the mutual information among global and local interaction sequences. Finally, to comprehensively take advantage of all three independent aspects of consistency-enhanced knowledge, we establish an integrated imitation learning framework. The consistency knowledge is effectively internalized and transferred to the student model by imitating the conventional prediction logit as well as the consistency-enhanced item representations. In addition, the flexible self-supervised imitation framework can also benefit other student recommenders. Experiments on four real-world datasets show that SSI effectively outperforms the state-of-the-art sequential recommendation methods.
Xu Yuan 0006, Hongshen Chen, Yonghao Song, Zhuoye Ding
IJCAI3
2021 Text is NOT Enough: Integrating Visual Impressions into Open-domain Dialogue Generation
abstract
Open-domain dialogue generation in natural language processing (NLP) is by default a pure-language task, which aims to satisfy human need for daily communication on open-ended topics by producing related and informative responses. In this paper, we point out that hidden images, named as visual impressions (VIs), can be explored from the text-only data to enhance dialogue understanding and help generate better responses. Besides, the semantic dependency between an dialogue post and its response is complicated, e.g., few word alignments and some topic transitions. Therefore, the visual impressions of them are not shared, and it is more reasonable to integrate the response visual impressions (RVIs) into the decoder, rather than the post visual impressions (PVIs). However, both the response and its RVIs are not given directly in the test process. To handle the above issues, we propose a framework to explicitly construct VIs based on pure-language dialogue datasets and utilize them for better dialogue understanding and generation. Specifically, we obtain a group of images (PVIs) for each post based on a pre-trained word-image mapping model. These PVIs are used in a co-attention encoder to get a post representation with both visual and textual information. Since the RVIs are not provided during testing, we design a cascade decoder that consists of two sub-decoders. The first sub-decoder predicts the content words in response, and applies the word-image mapping model to get corresponding RVIs. Then, the second sub-decoder generates the response based on the post and RVIs. Experimental results on two open-domain dialogue datasets show that our proposed approach achieves superior performance over competitive baselines in terms of fluency, relatedness, and diversity.
Lei Shen 0001, Haolan Zhan, Yonghao Song
ACM Multimedia4
2020 Learning from Easy to Complex: Adaptive Multi-Curricula Learning for Neural Dialogue Generation
abstract
Current state-of-the-art neural dialogue systems are mainly data-driven and are trained on human-generated responses. However, due to the subjectivity and open-ended nature of human conversations, the complexity of training dialogues varies greatly. The noise and uneven complexity of query-response pairs impede the learning efficiency and effects of the neural dialogue generation models. What is more, so far, there are no unified dialogue complexity measurements, and the dialogue complexity embodies multiple aspects of attributes—specificity, repetitiveness, relevance, etc. Inspired by human behaviors of learning to converse, where children learn from easy dialogues to complex ones and dynamically adjust their learning progress, in this paper, we first analyze five dialogue attributes to measure the dialogue complexity in multiple perspectives on three publicly available corpora. Then, we propose an adaptive multi-curricula learning framework to schedule a committee of the organized curricula. The framework is established upon the reinforcement learning paradigm, which automatically chooses different curricula at the evolving learning process according to the learning status of the neural dialogue generation model. Extensive experiments conducted on five state-of-the-art models demonstrate its learning efficiency and effectiveness with respect to 13 automatic evaluation metrics and human judgments.
Hengyi Cai, Hongshen Chen, Yonghao Song, Yangxi Li, Dongsheng Duan, Dawei Yin 0001
AAAI4
2020 Data Manipulation: Towards Effective Instance Learning for Neural Dialogue Generation via Learning to Augment and Reweight
abstract
Current state-of-the-art neural dialogue models learn from human conversations following the data-driven paradigm.As such, a reliable training corpus is the crux of building a robust and well-behaved dialogue model.However, due to the open-ended nature of human conversations, the quality of user-generated training data varies greatly, and effective training samples are typically insufficient while noisy samples frequently appear.This impedes the learning of those data-driven neural dialogue models.Therefore, effective dialogue learning requires not only more reliable learning samples, but also fewer noisy samples.In this paper, we propose a data manipulation framework to proactively reshape the data distribution towards reliable samples by augmenting and highlighting effective learning samples as well as reducing the effect of inefficient samples simultaneously.In particular, the data manipulation model selectively augments the training samples and assigns an importance weight to each instance to reform the training data.Note that, the proposed data manipulation framework is fully data-driven and learnable.It not only manipulates training samples to optimize the dialogue generation model, but also learns to increase its manipulation skills through gradient descent with validation samples.Extensive experiments show that our framework can improve the dialogue generation performance with respect to various automatic evaluation metrics and human judgments.
Hengyi Cai, Hongshen Chen, Yonghao Song, Dawei Yin 0001
ACL3
2020 Exemplar Guided Neural Dialogue Generation
abstract
Humans benefit from previous experiences when taking actions. Similarly, related examples from the training data also provide exemplary information for neural dialogue models when responding to a given input message. However, effectively fusing such exemplary information into dialogue generation is non-trivial: useful exemplars are required to be not only literally-similar, but also topic-related with the given context. Noisy exemplars impair the neural dialogue models understanding the conversation topics and even corrupt the response generation. To address the issues, we propose an exemplar guided neural dialogue generation model where exemplar responses are retrieved in terms of both the text similarity and the topic proximity through a two-stage exemplar retrieval model. In the first stage, a small subset of conversations is retrieved from a training set given a dialogue context. These candidate exemplars are then finely ranked regarding the topical proximity to choose the best-matched exemplar response. To further induce the neural dialogue generation model consulting the exemplar response and the conversation topics more faithfully, we introduce a multi-source sampling mechanism to provide the dialogue model with both local exemplary semantics and global topical guidance during decoding. Empirical evaluations on a large-scale conversation dataset show that the proposed approach significantly outperforms the state-of-the-art in terms of both the quantitative metrics and human evaluations.
Hengyi Cai, Hongshen Chen, Yonghao Song, Dawei Yin 0001
IJCAI3
2020 Low Latency Auditory Attention Detection with Common Spatial Pattern Analysis of EEG Signals
Siqi Cai 0002, Enze Su, Yonghao Song, Longhan Xie, Haizhou Li 0001
INTERSPEECH3
2020 Intelligent video analysis: A Pedestrian trajectory extraction method for the whole indoor space without blind areas
Lie Yang, Guanghua Hu, Yonghao Song, Guofeng Li, Longhan Xie
Comput. Vis. Image Underst.3
2020 A novel feature separation model exchange-GAN for facial expression recognition
Lie Yang, Yonghao Song, Nachuan Yang, Ke Ma 0007, Longhan Xie
Knowl. Based Syst.3
2019 Adaptive Parameterization for Neural Dialogue Generation
abstract
Hengyi Cai, Hongshen Chen, Cheng Zhang, Yonghao Song, Xiaofang Zhao, Dawei Yin. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Hengyi Cai, Hongshen Chen, Yonghao Song, Dawei Yin 0001
EMNLP/IJCNLP (1)4
2017 FTGWS: Forming Optimal Tutor Group for Weak Students Discovered in Educational Settings
Yonghao Song, Hengyi Cai, Xiaohui Zheng
DEXA (1)1
2015 PSFK: A Student Performance Prediction Scheme for First-Encounter Knowledge in ITS
abstract
As a user modeling method, Bayesian Knowledge Tracing (BKT) has been extensively used in the area of Intelligent Tutoring Systems (ITS). Thereafter the various schemes based on BKT are proposed to model student knowledge state and learning process. However, these schemes seldom consider the situation when a student first encounters a knowledge component (KC). That is, the existing models cannot be applied directly to predict student performance on a first-encounter KC. To solve this issue, combined user-based collaborative filtering and BKT model, a novel student performance prediction scheme PSFK is proposed in this paper. The PSFK scheme contains three major steps: first, building BKT models for each student and each KC he or she has encountered; then, finding the top-k similar students for a specified student S ; finally, predicting S ’s response on first-encounter KC. We evaluate our scheme on a real-world data set (which contains 4883 students and 177 KCs). The experiments show that the student performance prediction results of the proposed PSFK are acceptable (the RMSE can be decreased to 0.403).
Yonghao Song, Xiaohui Zheng, Haiyan Han, Yunqin Zhong
KSEM1