Xusheng Yang

dblp:50/6512 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
14since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Chinese Automatic Readability Assessment Using Adaptive Pre-training and Linguistic Feature Fusion
abstract
Chinese Automatic Readability Assessment (ARA) aims to classify the reading difficulty of Chinese texts. To address the issues of insufficient high-quality training data and underutilization of linguistic features in existing methods, we propose a method that combines adaptive pre-training with feature fusion based on an interactive attention mechanism. First, we enhance the model’s ability to capture different text difficulties through domain- and task-specific adaptive pre-training. Then, we propose an Adaptive Task-guided Corpus Filtering (ATCF) method, utilizing embeddings generated by the pre-trained model and applying nearest-neighbor search along with a sample balancing mechanism to ensure comprehensive learning across various difficulty levels. Finally, we propose an Interactive Attention-Driven Feature Fusion method that integrates linguistic and deep features, providing rich difficulty information to the model. Experiments on Chinese textbook dataset demonstrate that our method achieves state-of-the-art (SOTA) performance. Transfer learning experiments further indicate that our approach generalizes well to extracurricular reading and Chinese as a Foreign Language (CFL) ARA tasks.
Xusheng Yang, Jincai Yang
COLING1
2025 Audio-Faces Intra-Frame Alignment with Graph Attention Networks for Active Speaker Detection
abstract
Audio-Visual Active Speaker Detection(ASD) is the task of identifying, at any given moment, who is actively speaking in a multi-person scene by using audio and visual cues. Current main stream ASD methods separately encode audio and facial features, then adopt post-feature fusion approach where the acoustic features fused with the facial features from the same frame in the manner of vector concatenating or simple projecting. Such solution faces the challenges when there are more than one faces in the frame or overlapping speeches occur since there are lack of information alignment between active speech and the face of taking person. Based on this observation, in this study, we adopt a new solution to establish the relationships between the audio and face information using a heterogeneous graph explicitly. Specifically, we propose AFs-Net, which is able to capture both the relationships between the audio and each candidate’s face, and also the interactions between the faces of the candidates themselves within the same intra-frame. As the result, the graph with attention is trained to learn the importance (attention coefficient) between adjacent nodes. Additionally, we impose consistency constraints that bring speech features closer to speaker characteristics, while aligning non-speech features with non-speaker characteristics, further enhancing the audio-faces alignment. Our frame-level modeling approach supports both streaming applications and real-time operation. Experiments show that our method achieves state-of-the-art (SOTA) performance across multiple datasets.
Yongkang Yin, Xusheng Yang, Yuexian Zou
ICASSP2
2025 Applying Large Language Models with Active Learning and Ensemble Learning for Sentiment Recognition in Children's Reading
abstract
Sentiment recognition in children’s reading plays a crucial role in supporting their mental health and overall development. However, existing datasets often suffer from issues such as limited sample sizes, cross-linguistic ambiguities, and subjective human annotations, hindering the accurate modeling of complex emotional expressions in real-world contexts. These challenges result in suboptimal performance and limited generalization in sentiment recognition models. To address these issues, we leverage the extensive domain knowledge of Large Language Models (LLMs) to improve sentiment recognition performance. First, to mitigate data scarcity, we design a pretraining strategy that utilizes publicly available sentiment recognition datasets in both Chinese and English, enhancing LLMs’ ability to adapt to sentiment recognition tasks. Second, we propose a trustworthy ensemble strategy that integrates the outputs of LLMs across different languages. This strategy employs an uncertainty-aware weighted fusion mechanism to handle semantic ambiguities introduced by linguistic variations. Third, we develop an active learning framework for LLMs that iteratively resolves label inconsistencies, improving data quality and model performance. To validate the proposed methods, we construct a high-quality sentiment recognition dataset for children’s reading and incorporate the publicly available 40-Thai-Children dataset. Extensive experiments show that our method achieves an F1 score of 96.71% and accuracy of 97.42% on the collected dataset, and 86.31% F1 score and 87.02% accuracy on the 40-Thai-Children dataset. Furthermore, the active learning framework demonstrates its potential for generalization, improving accuracy from 82.05% to 86.32% in cross-domain testing by reducing data bias. The collected dataset and code are available at https://anonymous.4open.science/r/Children-s-Reading-Sentiment-Recognition-Dataset-745D.
Hongtao Mao, Jincai Yang, Xusheng Yang, Zhanghao Qin, Zeyu Zhang 0006
IJCNN3
2025 Deep multi-attribute spatial-temporal graph convolutional recurrent neural network-based multivariable spatial-temporal information fusion for short-term probabilistic forecast of multi-site photovoltaic power
Mingliang Bai, Guowen Zhou, Fuxiang Dong 0002, Yunxiao Chen, Xusheng Yang, Daren Yu
Expert Syst. Appl.7
2025 Enhancing automatic readability assessment with document and word knowledge fusion
Xusheng Yang, Jincai Yang, Qingchuang Guo
Expert Syst. Appl.1
2025 Progressive Gaussian filtering for nonlinear uncertain systems based on Gaussian process models
Xiaolei Zhuge, Xusheng Yang, Wen-An Zhang 0001
Signal Process.3
2024 Robust Heterophily Graph Learning via Uniformity Augmentation
abstract
Graphs serve as fundamental representations for a diverse array of complex systems, capturing intricate relationships and interactions between entities. In many real-world scenarios, graphs exhibit non-homophilous, or heterophilous, characteristics, challenging traditional graph analysis methods rooted in homophily assumptions. Recent heterophilous methods frequently struggle with noise in node attributes, which can degrade the quality of graph representations and affect downstream task performance. Common graph augmentations, while useful, often introduce bias and irrelevant noise. This paper proposes a novel method, Robust Heterophily Graph Learning via Uniformity Augmentation (RHGL-UA), which incorporates uniformity in the augmentation process through controlled random perturbations. This approach ensures a more uniform distribution of representations across different layers of the model. By adapting to data variations and learning more diverse information, RHGL-UA significantly improves performance on downstream tasks and stands out as the first practical robust heterophily graph method using representation augmentation with a theoretical guarantee. Extensive experiments demonstrate the merit of our proposed method.
Xusheng Yang, Zhengyu Chen 0001, Yuexian Zou
CIKM1
2024 DiffATR: Diffusion-based Generative Modeling for Audio-Text Retrieval
Yifei Xin, Xuxin Cheng, Zhihong Zhu 0001, Xusheng Yang, Yuexian Zou
INTERSPEECH4
2024 Audio-text Retrieval with Transformer-based Hierarchical Alignment and Disentangled Cross-modal Representation
Yifei Xin, Zhihong Zhu 0001, Xuxin Cheng, Xusheng Yang, Yuexian Zou
INTERSPEECH4
2024 Effects of Emotional Olfactory Stimuli on Modulating Angry Driving Based on an EEG Connectivity Study
abstract
Effectively regulating anger driving has become critical in ensuring road safety. The existing research lacks a feasible exploration of anger-driving regulation. This paper delves into the effect and neural mechanisms of emotional olfactory stimuli (EOS) on regulating anger driving based on EEG. First, this study designed an angry driving regulation experiment based on EOS to record EEG signals. Second, brain activation patterns under various EOS conditions are explored by analyzing functional brain networks (FBNs). Additionally, the paper analyzes dynamic alterations in anger-related characteristics to explore the intensity and persistence of regulating anger driving under different EOS. Finally, the paper studies the frequency energy of EEG changes under EOS through time-frequency analysis. The results indicate that EOS can effectively regulate a driver's anger emotions, especially with the banana odor showing superior effects. Under banana odor stimulus, synchronization between the parietal and temporal lobes significantly decreased. Notably, the regulatory effect of banana odor is optimal and exhibits sustained efficacy. The regulatory effect of banana odor on anger emotions is persistent. Furthermore, the impact of banana odor significantly reduces the distribution of high-energy activation states in the parietal lobe region. Our findings provide new insights into the dynamic characterization of functional connectivity during anger-driving regulation and demonstrate the potential of using EOS as a reliable tool for regulating angry driving.
Fo Hu, Peipei Yao, Kailun He, Xusheng Yang, Mohamed Amin Gouda, Lekai Zhang
Int. J. Neural Syst.4
2024 A Gaussian approximation filter for nonlinear systems with compound noises
Shu Yin 0002, Li Yu 0001, Xusheng Yang
Signal Process.4
2024 Gaussian Particle Filtering for Nonlinear Systems With Heavy-Tailed Noises: A Progressive Transform-Based Approach
abstract
The Gaussian particle filter (GPF) is a type of particle filter that employs the Gaussian filter approximation as the proposal distribution. However, the linearization errors are introduced during the calculation of the proposal distribution. In this article, a progressive transform-based GPF (PT-GPF) is proposed to solve this problem. First, a progressive transformation is applied to the measurement model to circumvent the necessity of linearization in the calculation of the proposal distribution, thereby ensuring the generation of optimal Gaussian proposal distributions in sense of linear minimum mean-square error (LMMSE). Second, to mitigate the potential impact of outliers, a supplementary screening process is employed to enhance the Monte Carlo approximation of the posterior probability density function. Finally, simulations of a target tracking example demonstrate the effectiveness and superiority of the proposed method.
Wen-An Zhang 0001, Ling Shi 0001, Xusheng Yang
IEEE Trans. Cybern.4
2024 EEG-Based Driver Fatigue Detection Using Spatio-Temporal Fusion Network With Brain Region Partitioning Strategy
abstract
Detecting driver fatigue is critical for ensuring traffic safety. Electroencephalography (EEG) is the golden standard for brain activity measurement and is considered a good indicator of detecting driver fatigue. However, the current driver fatigue detection algorithm has limitations in mining and fusing the spatiotemporal characteristics of EEG signals. In this paper, we propose a multi-branch deep learning network named spatio-temporal fusion network with brain region partitioning strategy (STFN-BRPS) to improve the accuracy and robustness of driver fatigue recognition. Initially, we develop a recurrent multi-scale convolution module (RMSCM) comprising a multi-scale convolution sub-module, a CNN-Bi-LSTM sub-module, and a residual structure branch. RMSCM effectively extracts highly discriminative long short-term temporal feature information. Secondly, we propose a dynamic graph convolution module and a spatial graph edges’ importance weight assignment method based on brain region partitioning strategy, which can acquire intrinsic spatial feature information between electrodes. Thirdly, we design a feature fusion module (FFM) that utilizes channel attention to fuse long short-term temporal and spatial features. FFM learns and prioritizes the significance and relevance of each channel in the fused features. Finally, the fused spatio-temporal features are passed into the classification module to obtain the predicted driver fatigue state. Extensive comparison and ablation studies are conducted on EEG signals collected from real-world driving scenarios. The results demonstrate that the proposed STFN-BRPS model delivers superior classification performance compared to the mainstream methods. This study establishes a benchmark for EEG-based driver fatigue detection and related deep-learning modeling work.
Fo Hu, Lekai Zhang, Xusheng Yang, Wen-An Zhang 0001
IEEE Trans. Intell. Transp. Syst.3
2022 Convergence analysis and improvement of progressive Gaussian approximation filters
Wenxiu He, Xusheng Yang
Signal Process.2
2020 Linear Fusion Estimation for Range-Only Target Tracking With Nonlinear Transformation
abstract
This article is concerned with the multisensor fusion estimation for target tracking with range-only wireless sensor networks. By employing a nonlinear transformation and a measurement fusion, the nonlinear distance measurements are transformed into a linear measurement with respect to the position of the target, which avoids the instability problem of nonlinear filtering. However, after the transformation, the new measurement noises are no longer Gaussian and cross uncorrelated. Taking the unmodeled disturbances into account, as well as the new noise properties, an adaptive factor is introduced by hypothesis test based on the posterior residual to improve the estimation performance, where only the root of a quadratic equation is required to be solved. Finally, both simulations and experiments of a target tracking example are presented to show the effectiveness of the proposed methods.
Xusheng Yang, Wen-An Zhang 0001, Andong Liu, Li Yu 0001
IEEE Trans. Ind. Informatics1
2020 A Bank of Decentralized Extended Information Filters for Target Tracking in Event-Triggered WSNs
abstract
This paper presents a hierarchical estimation method for maneuvering target tracking in event-triggered wireless sensor networks. First, several process noise covariances are chosen to characterize the dynamic characteristic of the target in the presence of maneuvers, and a bank of decentralized extended information filters (DEIFs) are used to generate state estimates of the target. Second, the estimates from the DEIFs are combined by covariance intersection (CI) to obtain an improved state estimate while still maintaining a consistent estimate. Thus, the DEIF and the CI methods form complementary advantages by satisfying the requirement of the consistency in the hierarchical estimation framework. Finally, both simulations and experiments of a target tracking example demonstrate that the proposed method is more suitable for applications to the maneuvering target tracking and it achieves a more satisfactory performance than the conventional DEIF method.
Xusheng Yang, Wen-An Zhang 0001, Li Yu 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2019 Progressive information filtering fusion for multi-sensor nonlinear systems
Liyan Zhao, Xusheng Yang, Wen-An Zhang 0001, Li Yu 0001
Signal Process.2
2016 Sequential Fusion Estimation for RSS-Based Mobile Robots Localization With Event-Driven WSNs
abstract
This paper is concerned with the sequential fusion estimation for mobile sensor node localizations with received signal strength measurements in mobile wireless sensor networks (MWSNs). The modeling errors induced by the communication uncertainties are considered and the process noise covariance is assumed to follow a uniform distribution. A sequential fusion estimation method based on a novel square root cubature Kalman filter is presented, where the process noise covariance is generated randomly. Moreover, a lower bound of the distribution is given to improve the stability and performance of the estimator. An E-puck robot-based MWSN experiment platform is designed, and both simulations and experiments show that the proposed sequential fusion estimation method help simplify the determination of the process noise covariance while maintaining a satisfactory estimation performance.
Wen-An Zhang 0001, Xusheng Yang, Li Yu 0001
IEEE Trans. Ind. Informatics2
2005 A Hybrid Method and Its Application for Power System
Xusheng Yang, Yong You, Wanxing Sheng, Sunan Wang
ISNN (3)1