Guangyu Fan

dblp:219/9330 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0002-1404-0009ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Resilient Secondary Frequency Control for Islanded Microgrids via a SSA-Optimized Multi-Instant Adaptive Cooperative Deployment Scheme
Yu Shan, Jiayue Sun, Guangyu Fan, Zhongyang Ming
IEEE Trans. Fuzzy Syst.3
2025 LLM-KGPlan: Long-Horizon Task Planning via Knowledge-Guided Reasoning
Dingyu Yang, Niansheng Chen, Guangyu Fan, Lei Rao, Songlin Cheng, Xiaoyong Song, Yingzhou Yu
PRICAI4
2025 MFBPNet: A Multi-Scale Fusion and Boundary Perception Network for Real-Time Semantic Segmentation in Autonomous Driving
abstract
Semantic segmentation is crucial in practical applications, especially in autonomous driving. Despite significant advancements in existing semantic segmentation methods, the performance of real-time segmentation approaches remains suboptimal. To address the trade-off between computational efficiency and accuracy in current methods, we propose a novel lightweight real-time semantic segmentation network named MFBPNet. Specifically, this paper introduces three core modules: (1) the Depthwise Separable Convolutional Pyramid Module (DSCPM), which expands the global receptive field and enhances deep feature representation; (2) the Local Attention Refinement Module (LARM), employing channel-wise attention to refine local feature discriminability, particularly for fine-grained objects; and (3) the Boundary Perception Feature Fusion Module (BPFM), which strengthens the feature representation of boundary regions through a multi-level feature fusion mechanism, effectively enhancing the clarity of object boundaries and mitigating boundary blurring issues. Extensive experiments on Cityscapes and CamVid datasets demonstrate that MFBPNet achieves state-of-the-art performance, attaining 75.3% mIoU at 67.3 FPS and 74.3% mIoU at 68.2 FPS, respectively. Compared to existing methods, MFBPNet achieves a superior balance between segmentation accuracy and real-time performance, rendering it highly suitable for autonomous driving systems requiring both real-time processing and high segmentation quality.
Guangyu Fan, Lei Rao, Songlin Cheng, Niansheng Chen, Xiaoyong Song, Dingyu Yang
SMC2
2025 Security optimization and beamforming design for active RIS-assisted UAV relaying NOMA networks
Songlin Cheng, Niansheng Chen, Guangyu Fan, Lei Rao, Xiaoyong Song, Dingyu Yang
Comput. Commun.4
2024 An Emotion Recognition with Online Learning Data Using Deep Learning Approaches
abstract
Emotion recognition helps teachers to obtain students' emotional states and their learning progress in their study. Many deep learning-based technologies, such as convolutional neural networks (CNN) and large language models (LLM) models are used for emotion recognition and prediction with learning data of students. An emotion recognition and prediction for online learning students using deep learning approaches is proposed, namely EPLLM. The method transforming the students' data, such as forum texts, questions, and answers, to a sequential form and extracts feature information from students' data using deep learning models with their behavioral changes in online learning. The student emotions are recognized and predicted based online learning data with combination of the deep learning models. To verify the performances of the model, many experiments are conducted using the public dataset. The results verify that, the model achieved better effectiveness in emotion prediction, compared with existing models. The values of F1 score of the proposed model is improved 0.051 compared that of the exist SGNN-LLM model.
Guangyu Fan, Songlin Cheng, Dingyu Yang
BIBM1
2024 A Student Emotions Recognition Based Online Texts with Large Language Models for Learning Prediction
abstract
with the rapid development of artificial intelligence (AI) and Large Language Models (LLM) technologies, it offers innovative approaches to model and analyze to predict student performance with educational behavior and learning data. But the challenges of data diversity, technical complex, and lack of semantic comprehension ability that limited the use of AI-based tools for learning performance prediction. In the paper, student emotions recognition based online learning data especially forum texts using LLMs is proposed, and an improved online learning performance prediction with online learning data, including student emotions using Signed Graph Neural Networks (SGNN) and LLMs, namely ER-SGNN-LLM, is constructed. In the model, the keywords are extracted from both forum texts and answers generated by students using LLM, and the keywords extracted from student’s forum texts are used for student emotions recognition, the keywords extracted from student’s answers, as well as the student emotions are used for student learning prediction. The relationships between each keyword extracted from texts and knowledge points of the course are encoded with SGNN. A graphic contrastive learning model is used to handle the noise in the dataset caused by students' subjective reasons. The combination of GNN and LLM is used to student emotions recognition and learning prediction. To verify the performances of the model, many experiments are conducted using the public dataset. The results demonstrate that, the model achieved better effectiveness in learning prediction, compared with other models. The values of F1 score of the proposed model is improved 0.032 compared that of the exist SGNN-LLM model.
Guangyu Fan, Shikun Zhang, Songlin Cheng, Dingyu Yang
BIBM1
2024 SemGO: Goal-Oriented Semantic Policy Based on MHSA for Object Goal Navigation
abstract
Object Goal Navigation is a task that seeks to allow intelligent agents to locate and navigate to a particular object goal in an unfamiliar environment. However, the current goal-oriented semantic policy, which is based on deep reinforcement learning (DRL), has difficulty in retaining long-term object semantic information and lacks adequate goal-oriented ability. This leads to intelligent agents having to extensively explore their environment in order to locate a goal, resulting in inefficient navigation. To address these challenges, this paper proposes a goal-oriented semantic policy based on multi-headed self-attention (MHSA) to improve the efficiency of object navigation. By using the self-attention mechanism, the policy can automatically learn and extract features relevant to goal navigation without the need for manual feature extractor design. Multiple attention heads can simultaneously focus on various semantic features to extract vital information about the objective goal. We propose a novel object-goal navigation model called SemGO based on this policy. The SemGO model is proficient at managing environments with intricate semantic structures. It can detect the correlation between global and local information, which improves navigation accuracy significantly. Additionally, it has superior generalization capabilities, making it adaptable to changes in different object goals and environments. The experimental results show that the SemGO model achieves a SPL of 0.324, a success rate of 0.635, and a reduction of DTS to 1.601m in the Gibson dataset for object-goal navigation.
Niansheng Chen, Lei Rao, Guangyu Fan, Dingyu Yang, Songlin Cheng, Xiaoyong Song, Yiping Ma 0008
CSCWD4
2024 DEUFormer: High-precision semantic segmentation for urban remote sensing images
abstract
Abstract Urban remote sensing image semantic segmentation has a wide range of applications, such as urban planning, resource exploration, intelligent transportation, and other scenarios. Although UNetFormer performs well by introducing the self‐attention mechanism of Transformer, it still faces challenges arising from relatively low segmentation accuracy and significant edge segmentation errors. To this end, this paper proposes DEUFormer by employing a special weighted sum method to fuse the features of the encoder and the decoder, thus capturing both local details and global context information. Moreover, an Enhanced Feature Refinement Head is designed to finely re‐weight features on the channel dimension and narrow the semantic gap between shallow and deep features, thereby enhancing multi‐scale feature extraction. Additionally, an Edge‐Guided Context Module is introduced to enhance edge areas through effective edge detection, which can improve edge information extraction. Experimental results show that DEUFormer achieves an average Mean Intersection over Union (mIoU) of 53.8% on the LoveDA dataset and 69.1% on the UAVid dataset. Notably, the mIoU of buildings in the LoveDA dataset is 5.0% higher than that of UNetFormer. The proposed model outperforms methods such as UNetFormer on multiple datasets, which demonstrates its effectiveness.
Xinqi Jia, Xiaoyong Song, Lei Rao, Guangyu Fan, Songlin Cheng, Niansheng Chen
IET Comput. Vis.4
2024 NAVS: A Neural Attention-Based Visual SLAM for Autonomous Navigation in Unknown 3D Environments
abstract
Abstract Navigation in unknown 3D environments aims to progressively find an efficient path to a given target goal in unseen scenarios. A challenge is how to explore the navigation quickly and effectively. An end-to-end learning approach has been proposed to extract geometric shapes from RGB images, but it is not suitable for large environments due to its exhaustive exploration with exponential search space. Active Neural SLAM (ANS) presents a Neural SLAM module to maximize the exploration coverage to tackle the active SLAM task. However, ANS still frequently visits the explored areas due to the inappropriate local target selection. In this paper, we propose a Neural Attention-based Visual SLAM (NAVS) model to explore unknown 3D environments. Spatial attention is provided to quickly identify obstacles (such as similarly colored tea table or floor). We also leverage the priority of unknown regions in the short-term goal decision to avoid frequent exploration with a channel attention. The experimental results show that our model can build a more accurate map than ANS and other baseline methods with less running time. In terms of relative coverage, NAVS achieves a 0.5 $$\%$$ % improvement over ANS in overall and a 1.1 $$\%$$ % improvement over ANS in large environments.
Niansheng Chen, Guangyu Fan, Dingyu Yang, Lei Rao, Songlin Cheng, Xiaoyong Song, Yiping Ma 0008
Neural Process. Lett.3
2024 Performance analysis of UAV-assisted DF relaying network with hardware impairments and energy harvesting
Jielin Chen, Niansheng Chen, Songlin Cheng, Guangyu Fan, Lei Rao, Xiaoyong Song, Wenjing Lv, Dingyu Yang
Wirel. Networks4
2023 ASKCC-DCNN-CTC: A Multi-Core Two Dimensional Causal Convolution Fusion Network with Attention Mechanism for End-to-End Speech Recognition
abstract
Aiming at the problems of difficulty in extracting key features and low prediction accuracy of traditional convolutional neural networks in Chinese speech recognition, we analyze the impacts of information leakage and unstandardized phoneme features on its performance, based on the deep convolutional neural network (DCNN)-connectionist temporal classification (CTC) model. In addition, a multi-core two dimensional causal convolution fusion network layer structure of SKNet is constructed, and we propose a DCNN-CTC model for fusion of attention mechanism and SKNet multi-core 2D causal convolution network (ASKCC-DCNN-CTC), which effectively improves the accuracy and training speed of Chinese speech recognition. The simulation results show that the error rate of our model on the ST-CMDS dataset is 12.201% lower than that of the DCNN-CTC model, the performance on the THCHS30 dataset is also improved, which reveals a good generalization ability.
Rongchuang Lv, Niansheng Chen, Songlin Cheng, Guangyu Fan, Lei Rao, Xiaoyong Song, Dingyu Yang
CSCWD4
2023 An End-to-End Robotic Visual Localization Algorithm Based on Deep Learning
abstract
Efficient localization plays a significant role in mobile autonomous robots’ navigation systems. Traditional visual simultaneous localization systems based on point feature matching suffer from two shortcomings. First one is that the method of tracking features is not robust for the environments with frequent changes in brightness. Another one is the large of consecutive visual keyframes consume expensive computation and storage resources in complex environments. To solve these problems, we propose an end-to-end visual localization algorithm to solve the robust and efficiency challenges via a deep learning mode. Firstly, we perform preprocessing operations such as cropping, averaging, and timestamp alignment on datasets to reduce computational cost and time. Secondly, we use CNN networks to autonomously learn the correct features for localization, which is robust to illumination changes. Finally, we utilize LSTM networks to memorize global trajectories to improve the localization precision. We performed a broad range of experiments on both indoor and outdoor datasets. The experimental results demonstrate that the translation and orientation accuracy in outdoor scenes improved by 32.9% and 31.4%, respectively. The average improvement of translation positioning accuracy in indoor scenes is 38.4%, and the orientation improvement is 13.1%. Moreover, the effectiveness of predicting the global motion trajectories of sequential images algorithm has been verified and is superior to other CNN methods.
Niansheng Chen, Guangyu Fan, Dingyu Yang, Lei Rao, Songlin Cheng
IJCNN3
2023 KS-Autoformer: An Autoformer-Based SOC Prediction Framework for Electric Vehicles
Yaoyidi Wang, Niansheng Chen, Lei Rao, Dingyu Yang, Guangyu Fan, Songlin Cheng, Xiaoyong Song
MobiQuitous (1)5