Xuanyu Jin

dblp:252/6291 · DBLP profile ↗
← Back
15ranked-venue papers
1as first author
15since 2021 · last 2026
0000-0001-5542-3340ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Tensile property prediction of titanium and aluminum alloys dissimilar joint by plasma plume characteristics based on a multi-stage cascade model
Chuang Cai, Fashuai Xiong, Zilin Chen, Xuanyu Jin, Zejun Xian, Hua Tang
Eng. Appl. Artif. Intell.6
2026 DRFNet: Enhancing Identity Discriminability and Feature Robustness for Cross-Session VEP-Based EEG Biometrics
abstract
Biometric recognition using visually evoked potentials (VEPs), a type of neural response to visual stimuli recorded via electroencephalography (EEG), has shown great promise. However, the non-stationary nature of EEG signals poses a major challenge in cross-session scenarios, where data collected on different days often leads to performance degradation. To address this, we propose the Discriminative Robust Feature Network (DRFNet) to enhance the robustness and inter-subject discriminability of identity representations across sessions. DRFNet incorporates two key components: (1) A log-power transformation that amplifies inter-individual differences by capturing non-linear energy patterns from VEP features via signal squaring and logarithmic scaling; and (2) A hierarchical normalization strategy with adaptive attention to balance discriminative identity cues with inter-session invariance by stabilizing feature distributions across multiple levels (feature map, batch, and sample). On two public multi-session SSVEP datasets (Dataset A: 30 subjects, 6 s trials; Dataset B: 54 subjects, 4 s trials), our model outperformed state-of-the-art methods, achieving identification accuracies of 92.92% and 86.30%, and equal error rates of 3.92% and 4.09%, respectively. Further analysis demonstrates that filter bank processing and a reduced set of parietal-occipital electrodes can provide more discriminative features while offering a practical path toward system lightweighting.
Honggang Liu, Han Yang 0003, Dongjun Liu, Xuanyu Jin, Yong Peng 0001, Wanzeng Kong
IEEE J. Biomed. Health Informatics4
2026 Cognition-driven Adaptive Semantic Decoding Framework for Multimodal Sentiment Analysis
abstract
In real-world scenarios, multimodal sentiment analysis faces significant challenges, particularly in cross-scenario generalization. Existing works fail to effectively deal with the variability in evaluation frameworks and modality combinations, which results in poor transfer performance across different application contexts. In this article, the cognition-driven adaptive semantic decoding framework (CASDF) is proposed to realize an evaluation system and modality-independent multimodal sentiment analysis. Specifically, the adaptive modality association module is proposed to construct the adaptive modality mapping space, which allows us to dynamically adapt to arbitrary modality combinations. This indeed breaks through the limitation of the modality number and effectively deals with the modality gap. Furthermore, similar to the human hierarchical cognition (“perception-concept-decision”), the evaluation system progressive alignment module is presented to establish the unified evaluation system. This consists of the perception, concept, and decision analysis, which contributes to the adaptive cross-task analysis from the discrete sentiment space to the continuous sentiment space. The above joint analysis of the evaluation system and modality number indeed leads to the more flexible and generable multimodal sentiment semantic decoding paradigm. The experiments demonstrate that our sentiment semantic analysis network can achieve state-of-the-art performance.
Jiajia Tang, Honggang Liu, Xuanyu Jin, Wanzeng Kong
ACM Trans. Multim. Comput. Commun. Appl.4
2025 RPW-EEG: An Unified Framework for Robust and Practical Watermark of EEG
Tianyang Qin, Hangjie Yi, Jingsheng Qian, Xuanyu Jin, Honggang Liu, Wanzeng Kong
CogSci4
2025 PBFE-DAN:Personal Biological Feature Enhanced Domain Adaptation Network for Cross-Session Brainprint Recognition
abstract
In recent years, electroencephalogram (EEG) signal analysis has made significant strides, offering novel solutions for bolstering the security of brain–computer interfaces (BCI) and addressing long-standing vulnerabilities in traditional biometric authentication. Nonetheless, the low signal-to-noise ratio of EEG signals and their susceptibility to environmental influences lead to performance degradation when the model encounters data from subsequent unseen sessions. Targeting the crucial challenge of background noise in cross-session brainprint recognition, this paper presents the Personal Biological Feature Enhanced Domain Adaptation Network (PBFE-DAN). This Transformer-Based framework incorporates depth-ßwise separable convolutions with a Personal Biological Feature Enhancement Unit (PBFEU), which employs a gating mechanism to selectively amplify individual-specific EEG patterns while suppressing background noise. By emphasizing these unique EEG features, PBFE-DAN markedly improves robustness and generalizability across different sessions. Additionally, the self-attention module extracts global, session-invariant features that further enhance the stability and accuracy of cross-session brainprint recognition. Notably, PBFE-DAN can achieve state-of-the-art classification accuracies on both DSIRSVP and SEED-IV datasets, thereby underscoring the effectiveness of the proposed framework in mitigating background noise for cross-session brainprint recognition.
Xuanyu Jin, Wanzeng Kong
IJCNN2
2025 QELDBA: Query-Efficient and Low Distortion Black-Box Attack for Brainprint Recognition
abstract
While various deep learning techniques for electroencephalogram (EEG)-based brainprint recognition have achieved considerable success, these models remain vulnerable to adversarial attacks. However, existing black-box attack methods suffer from an inherent trade-off between query efficiency and distortion level. To address this challenge and further investigate the security risks of brainprint recognition systems in real-world black-box scenarios, we propose a query-efficient, low-distortion black-box attack method that targets the high-frequency components of EEG signals. Our approach innovatively selects sparse sampling points to estimate more accurate gradient information and leverages historical gradients to guide the prioritization of important points, thereby accelerating the attack process. The perturbations are applied in the high-frequency domain of the EEG signal to enhance stealth and effectiveness. Extensive experiments under black-box settings demonstrate that our method achieves state-of-the-art performance across two datasets and four models. Compared to existing methods, our approach significantly improves attack success rates while reducing the number of queries and minimizing distortion to imperceptible levels, thus achieving a superior balance between query efficiency and perturbation stealth.
Jingsheng Qian, Hangjie Yi, Honggang Liu, Xuanyu Jin, Wanzeng Kong
IEEE Signal Process. Lett.4
2025 ID-ProtoFormer: A Dynamic Identity Prototype-Infused Transformer for SSVEP-Based Biometric Recognition
Jiabin Zhu, Xuanyu Jin, Wanzeng Kong
IEEE Signal Process. Lett.2
2024 The Impact of Dynamic Icons on Mobile APP Interfaces: Evidence from EEG and Eye-tracking Signals
abstract
This study investigates the cognitive impact of dynamic icons in mobile interfaces by integrating electroencephalography (EEG) and eye-tracking technologies. Traditional research on mobile app interface design has relied mainly on questionnaire and eye-tracking methods for behavioral analysis. This research adds a new dimension by examining the neural mechanisms associated with dynamic icons. We employed EEG to analyze channel-wise power spectrum density (PSD), focusing on the alpha and theta frequency bands related to attention and working memory. Concurrently, eye-tracking data were analyzed through Areas of Interest (AOIs) and fixation metrics to assess visual attention patterns. The results indicate that dynamic icons significantly enhance neural activity, with a 15% increase in alpha band power and a 20% increase in theta band power compared to static icons. Additionally, eye-tracking data show a 30% increase in total fixation duration on AOIs containing dynamic icons, particularly in the left-top quarter of the mobile interface. This effect was observed without changes in the first fixation duration, suggesting that dynamic icons have a stronger impact on sustained attention rather than on initial capture. These findings highlight that dynamic icons not only attract and maintain visual attention more effectively but also enhance cognitive processing efficiency. This study provides valuable insights for optimizing mobile app interface design, emphasizing the benefits of incorporating dynamic elements to improve user engagement and interface effectiveness.
Ruizhe Yang, Jiaxuan Qin, Letao Fang, Haojie Tao, Li Zhu 0005, Xuanyu Jin, Wanzeng Kong
BIBM7
2024 Comparisons on Perception Mechanism of Mental Rotation Between Health and Stroke Groups with EEG Indicators
abstract
Clinically, motor rehabilitation and mechanism research have garnered increasing attention. However, cognitive impairments often accompany stroke patients. Mental rotation is a crucial task for cognitive evaluation and training. In our study, we proposed a method to compare mental rotation perception mechanisms between healthy individuals and stroke patients using EEG indicators. The experiment was designed to accommodate stroke patients. In the data analysis module, we utilized power spectral density (PSD) to explore single-channel frequency domain indicators and phase-locked value (PLV) to analyze connectivity between channels. Mechanism analysis based on EEG indicators considered three main conditions: counter-clockwise and clockwise rotation, lesion area versus functional area across strokes, and mental rotation perception sub-stages. Our experimental results show significant differences between the stroke group and health group in the θ and α frequency bands across counter-clock and clock-wise perception and lesion connectivity conditions. The stroke group reveals a compensatory effect in the lesion areas, with higher PLV values (averaging 0.2101) compared to those in common cognitive functional areas of the health group and the connectivity within lesion areas was significantly lower (averaging 0.0310) than that outside the lesion areas. These findings could assist in stroke treatment and monitoring rehabilitation progress.
Lingmin Zhou, Li Zhu 0005, Haibin Xia, Guifen Yang, Xuanyu Jin, Wanzeng Kong
BIBM6
2024 Efficient and Verifiable Multi-server Framework for Secure Information Classification and Storage
Ziqing Guo, Xuanyu Jin, Xiuhua Wang 0009, Yueyue Dai
Inscrypt (2)3
2024 Unbiased Semantic Representation Learning Based on Causal Disentanglement for Domain Generalization
abstract
Domain generalization primarily mitigates domain shift among multiple source domains, generalizing the trained model to an unseen target domain. However, the spurious correlation usually caused by context prior (e.g., background) makes it challenging to get rid of the domain shift. Therefore, it is critical to model the intrinsic causal mechanism. The existing domain generalization methods only attend to disentangle the semantic and context-related features by modeling the causation between input and labels, which totally ignores the unidentifiable but important confounders. In this article, a Causal Disentangled Intervention Model (CDIM) is proposed for the first time, to the best of our knowledge, to construct confounders via causal intervention. Specifically, a generative model is employed to disentangle the semantic and context-related features. The contextual information of each domain from generative model can be considered as a confounder layer, and the center of all context-related features is utilized for fine-grained hierarchical modeling of confounders. Then the semantic and confounding features from each layer are combined to train an unbiased classifier, which exhibits both transferability and robustness across an unknown distribution domain. CDIM is evaluated on three widely recognized benchmark datasets, namely, Digit-DG, PACS, and NICO, through extensive ablation studies. The experimental results clearly demonstrate that the proposed model achieves state-of-the-art performance.
Xuanyu Jin, Wanzeng Kong, Jiajia Tang
ACM Trans. Multim. Comput. Commun. Appl.1
2023 Brain-Machine Coupled Learning Method for Facial Emotion Recognition
abstract
Neural network models of machine learning have shown promising prospects for visual tasks, such as facial emotion recognition (FER). However, the generalization of the model trained from a dataset with a few samples is limited. Unlike the machine, the human brain can effectively realize the required information from a few samples to complete the visual tasks. To learn the generalization ability of the brain, in this article, we propose a novel brain-machine coupled learning method for facial emotion recognition to let the neural network learn the visual knowledge of the machine and cognitive knowledge of the brain simultaneously. The proposed method utilizes visual images and electroencephalogram (EEG) signals to couple training the models in the visual and cognitive domains. Each domain model consists of two types of interactive channels, common and private. Since the EEG signals can reflect brain activity, the cognitive process of the brain is decoded by a model following reverse engineering. Decoding the EEG signals induced by the facial emotion images, the common channel in the visual domain can approach the cognitive process in the cognitive domain. Moreover, the knowledge specific to each domain is found in each private channel using an adversarial strategy. After learning, without the participation of the EEG signals, only the concatenation of both channels in the visual domain is used to classify facial emotion images based on the visual knowledge of the machine and the cognitive knowledge learned from the brain. Experiments demonstrate that the proposed method can produce excellent performance on several public datasets. Further experiments show that the proposed method trained from the EEG signals has good generalization ability on new datasets and can be applied to other network models, illustrating the potential for practical applications.
Dongjun Liu, Weichen Dai 0001, Hangkui Zhang, Xuanyu Jin, Jianting Cao, Wanzeng Kong
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 BAFN: Bi-Direction Attention Based Fusion Network for Multimodal Sentiment Analysis
abstract
Attention-based networks currently identify their effectiveness in multimodal sentiment analysis. However, existing methods ignore the redundancy of auxiliary modalities. More importantly, existing methods only attend to top-down attention (static process) or down-top attention (implicit process), leading to the coarse-grained multimodal sentiment context. In this paper, during the preprocessing period, we first propose the multimodal dynamic enhanced block to capture the intra-modality sentiment context. This can effectively decrease the intra-modality redundancy of auxiliary modalities. Furthermore, the bi-direction attention block is proposed to capture fine-grained multimodal sentiment context via the novel bi-direction multimodal dynamic routing mechanism. Specifically, the bi-direction attention block first highlights the explicit and low-level multimodal sentiment context. Then, the low-level multimodal context is transmitted to a carefully designed bi-direction multimodal dynamic routing procedure. This allows us to dynamically update and investigate high-level and much more fine-grained multimodal sentiment contexts. The experiments demonstrate that our fusion network can achieve state-of-the-art performance. Notably, our model outperforms the best baseline on the metric ‘Acc-7’ with an improvement of 6.9%.
Jiajia Tang, Dongjun Liu, Xuanyu Jin, Yong Peng 0001, Qibin Zhao, Yu Ding 0001, Wanzeng Kong
IEEE Trans. Circuits Syst. Video Technol.3
2022 MMT: Multi-way Multi-modal Transformer for Multimodal Learning
abstract
The heart of multimodal learning research lies the challenge of effectively exploiting fusion representations among multiple modalities.However, existing two-way cross-modality unidirectional attention could only exploit the intermodal interactions from one source to one target modality. This indeed fails to unleash the complete expressive power of multimodal fusion with restricted number of modalities and fixed interactive direction.In this work, the multiway multimodal transformer (MMT) is proposed to simultaneously explore multiway multimodal intercorrelations for each modality via single block rather than multiple stacked cross-modality blocks. The core idea of MMT is the multiway multimodal attention, where the multiple modalities are leveraged to compute the multiway attention tensor. This naturally benefits us to exploit comprehensive many-to-many multimodal interactive paths. Specifically, the multiway tensor is comprised of multiple interconnected modality-aware core tensors that consist of the intramodal interactions. Additionally, the tensor contraction operation is utilized to investigate intermodal dependencies between distinct core tensors.Essentially, our tensor-based multiway structure allows for easily extending MMT to the case associated with an arbitrary number of modalities. Taking MMT as the basis, the hierarchical network is further established to recursively transmit the low-level multiway multimodal interactions to high-level ones. The experiments demonstrate that MMT can achieve state-of-the-art or comparable performance.
Jiajia Tang, Xuanyu Jin, Wanzeng Kong, Yu Ding 0001, Qibin Zhao
IJCAI4
2021 CTFN: Hierarchical Learning for Multimodal Sentiment Analysis Using Coupled-Translation Fusion Network
abstract
Jiajia Tang, Kang Li, Xuanyu Jin, Andrzej Cichocki, Qibin Zhao, Wanzeng Kong. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Jiajia Tang, Xuanyu Jin, Andrzej Cichocki, Qibin Zhao, Wanzeng Kong
ACL/IJCNLP (1)3