Xiaoyong Song

dblp:227/5585 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
11since 2021 · last 2026
0009-0005-8621-2315ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CD-PIM: A High-Bandwidth and Compute-Efficient LPDDR5-Based PIM for Low-Batch LLM Acceleration on Edge-Device
abstract
Edge deployment of low-batch large language models (LLMs) faces critical memory bandwidth bottlenecks when executing memory-intensive general matrix-vector multiplications (GEMV) operations. While digital processing-in-memory (PIM) architectures promise to accelerate GEMV operations, existing PIM-equipped edge devices still suffer from three key limitations: limited bandwidth improvement, component under-utilization in mixed workloads, and low compute capacity of computing units (CUs). In this paper, we propose CD-PIM to address these challenges through three key innovations. First, we introduce a high-bandwidth compute-efficient mode (HBCEM) that enhances bandwidth by dividing each bank into four pseudo-banks through segmented global bitlines. Second, we propose a low-batch interleaving mode (LBIM) to improve component utilization by overlapping GEMV operations with GEMM operations. Third, we design a compute-efficient CU that performs enhanced GEMV operations in a pipelined manner by serially feeding weight data into the computing core. Forth, we adopt a column-wise mapping for the key-cache matrix and row-wise mapping for the value-cache matrix, which fully utilizes CU resources. Our evaluation shows that compared to a GPU-only baseline and state-of-the-art PIM designs, our CD-PIM achieves 11.42× and 4.25× speedup on average within a single batch in HBCEM mode, respectively. Moreover, for low-batch sizes, the CD-PIM achieves an average speedup of 1.12× in LBIM compared to HBCEM.
Chao Fang 0005, Xiaoyong Song, Anying Jiang, Yichuan Bai
DATE3
2025 LLM-KGPlan: Long-Horizon Task Planning via Knowledge-Guided Reasoning
Dingyu Yang, Niansheng Chen, Guangyu Fan, Lei Rao, Songlin Cheng, Xiaoyong Song, Yingzhou Yu
PRICAI7
2025 MFBPNet: A Multi-Scale Fusion and Boundary Perception Network for Real-Time Semantic Segmentation in Autonomous Driving
abstract
Semantic segmentation is crucial in practical applications, especially in autonomous driving. Despite significant advancements in existing semantic segmentation methods, the performance of real-time segmentation approaches remains suboptimal. To address the trade-off between computational efficiency and accuracy in current methods, we propose a novel lightweight real-time semantic segmentation network named MFBPNet. Specifically, this paper introduces three core modules: (1) the Depthwise Separable Convolutional Pyramid Module (DSCPM), which expands the global receptive field and enhances deep feature representation; (2) the Local Attention Refinement Module (LARM), employing channel-wise attention to refine local feature discriminability, particularly for fine-grained objects; and (3) the Boundary Perception Feature Fusion Module (BPFM), which strengthens the feature representation of boundary regions through a multi-level feature fusion mechanism, effectively enhancing the clarity of object boundaries and mitigating boundary blurring issues. Extensive experiments on Cityscapes and CamVid datasets demonstrate that MFBPNet achieves state-of-the-art performance, attaining 75.3% mIoU at 67.3 FPS and 74.3% mIoU at 68.2 FPS, respectively. Compared to existing methods, MFBPNet achieves a superior balance between segmentation accuracy and real-time performance, rendering it highly suitable for autonomous driving systems requiring both real-time processing and high segmentation quality.
Guangyu Fan, Lei Rao, Songlin Cheng, Niansheng Chen, Xiaoyong Song, Dingyu Yang
SMC6
2025 Security optimization and beamforming design for active RIS-assisted UAV relaying NOMA networks
Songlin Cheng, Niansheng Chen, Guangyu Fan, Lei Rao, Xiaoyong Song, Dingyu Yang
Comput. Commun.6
2024 SemGO: Goal-Oriented Semantic Policy Based on MHSA for Object Goal Navigation
abstract
Object Goal Navigation is a task that seeks to allow intelligent agents to locate and navigate to a particular object goal in an unfamiliar environment. However, the current goal-oriented semantic policy, which is based on deep reinforcement learning (DRL), has difficulty in retaining long-term object semantic information and lacks adequate goal-oriented ability. This leads to intelligent agents having to extensively explore their environment in order to locate a goal, resulting in inefficient navigation. To address these challenges, this paper proposes a goal-oriented semantic policy based on multi-headed self-attention (MHSA) to improve the efficiency of object navigation. By using the self-attention mechanism, the policy can automatically learn and extract features relevant to goal navigation without the need for manual feature extractor design. Multiple attention heads can simultaneously focus on various semantic features to extract vital information about the objective goal. We propose a novel object-goal navigation model called SemGO based on this policy. The SemGO model is proficient at managing environments with intricate semantic structures. It can detect the correlation between global and local information, which improves navigation accuracy significantly. Additionally, it has superior generalization capabilities, making it adaptable to changes in different object goals and environments. The experimental results show that the SemGO model achieves a SPL of 0.324, a success rate of 0.635, and a reduction of DTS to 1.601m in the Gibson dataset for object-goal navigation.
Niansheng Chen, Lei Rao, Guangyu Fan, Dingyu Yang, Songlin Cheng, Xiaoyong Song, Yiping Ma 0008
CSCWD7
2024 DEUFormer: High-precision semantic segmentation for urban remote sensing images
abstract
Abstract Urban remote sensing image semantic segmentation has a wide range of applications, such as urban planning, resource exploration, intelligent transportation, and other scenarios. Although UNetFormer performs well by introducing the self‐attention mechanism of Transformer, it still faces challenges arising from relatively low segmentation accuracy and significant edge segmentation errors. To this end, this paper proposes DEUFormer by employing a special weighted sum method to fuse the features of the encoder and the decoder, thus capturing both local details and global context information. Moreover, an Enhanced Feature Refinement Head is designed to finely re‐weight features on the channel dimension and narrow the semantic gap between shallow and deep features, thereby enhancing multi‐scale feature extraction. Additionally, an Edge‐Guided Context Module is introduced to enhance edge areas through effective edge detection, which can improve edge information extraction. Experimental results show that DEUFormer achieves an average Mean Intersection over Union (mIoU) of 53.8% on the LoveDA dataset and 69.1% on the UAVid dataset. Notably, the mIoU of buildings in the LoveDA dataset is 5.0% higher than that of UNetFormer. The proposed model outperforms methods such as UNetFormer on multiple datasets, which demonstrates its effectiveness.
Xinqi Jia, Xiaoyong Song, Lei Rao, Guangyu Fan, Songlin Cheng, Niansheng Chen
IET Comput. Vis.2
2024 NAVS: A Neural Attention-Based Visual SLAM for Autonomous Navigation in Unknown 3D Environments
abstract
Abstract Navigation in unknown 3D environments aims to progressively find an efficient path to a given target goal in unseen scenarios. A challenge is how to explore the navigation quickly and effectively. An end-to-end learning approach has been proposed to extract geometric shapes from RGB images, but it is not suitable for large environments due to its exhaustive exploration with exponential search space. Active Neural SLAM (ANS) presents a Neural SLAM module to maximize the exploration coverage to tackle the active SLAM task. However, ANS still frequently visits the explored areas due to the inappropriate local target selection. In this paper, we propose a Neural Attention-based Visual SLAM (NAVS) model to explore unknown 3D environments. Spatial attention is provided to quickly identify obstacles (such as similarly colored tea table or floor). We also leverage the priority of unknown regions in the short-term goal decision to avoid frequent exploration with a channel attention. The experimental results show that our model can build a more accurate map than ANS and other baseline methods with less running time. In terms of relative coverage, NAVS achieves a 0.5 $$\%$$ % improvement over ANS in overall and a 1.1 $$\%$$ % improvement over ANS in large environments.
Niansheng Chen, Guangyu Fan, Dingyu Yang, Lei Rao, Songlin Cheng, Xiaoyong Song, Yiping Ma 0008
Neural Process. Lett.7
2024 An Implementation of Reconfigurable Match Table for FPGA-Based Programmable Switches
abstract
Match table is the key part to perform packet processing and forwarding for programmable switches in a software-defined network (SDN). However, the match table in current field-programmable gate array (FPGA)-based switches is inflexible or undisclosed. When the network function changes, the match table on FPGA needs to be redesigned or reset size parameters, and after recompilation and reimplementation, it could work again; this time-consuming and labor-intensive operation seriously reduces the flexibility and configurability of the switch. To address this issue, this article presents a design of reconfigurable match table (RMT) for FPGA-based programmable switches. A three-layer table structure is introduced to realize the reconfiguration and hardware-plane mapping of user-defined tables, and the logical tables in packet processing pipeline are interconnected with the physical tables in memory pool by the designed resource-efficient segment crossbar. To the best of our knowledge, this article is the first to publicly present the entire FPGA-based RMT design scheme and implementation details. The proposed design implements reconfigurable ternary content addressable memory (TCAM) based and static random access memory (SRAM) based match tables on Xilinx FPGA and verifies them with a packet filter system. In the proposed RMT system, a user could reconfigure the number, depth, and width of user-defined match tables (UMTs) in pipeline via control plane without modifying hardware, which enhances the flexibility of the data plane of FPGA-based switch greatly.
Xiaoyong Song, Zhichuan Guo
IEEE Trans. Very Large Scale Integr. Syst.1
2024 Performance analysis of UAV-assisted DF relaying network with hardware impairments and energy harvesting
Jielin Chen, Niansheng Chen, Songlin Cheng, Guangyu Fan, Lei Rao, Xiaoyong Song, Wenjing Lv, Dingyu Yang
Wirel. Networks6
2023 ASKCC-DCNN-CTC: A Multi-Core Two Dimensional Causal Convolution Fusion Network with Attention Mechanism for End-to-End Speech Recognition
abstract
Aiming at the problems of difficulty in extracting key features and low prediction accuracy of traditional convolutional neural networks in Chinese speech recognition, we analyze the impacts of information leakage and unstandardized phoneme features on its performance, based on the deep convolutional neural network (DCNN)-connectionist temporal classification (CTC) model. In addition, a multi-core two dimensional causal convolution fusion network layer structure of SKNet is constructed, and we propose a DCNN-CTC model for fusion of attention mechanism and SKNet multi-core 2D causal convolution network (ASKCC-DCNN-CTC), which effectively improves the accuracy and training speed of Chinese speech recognition. The simulation results show that the error rate of our model on the ST-CMDS dataset is 12.201% lower than that of the DCNN-CTC model, the performance on the THCHS30 dataset is also improved, which reveals a good generalization ability.
Rongchuang Lv, Niansheng Chen, Songlin Cheng, Guangyu Fan, Lei Rao, Xiaoyong Song, Dingyu Yang
CSCWD6
2023 KS-Autoformer: An Autoformer-Based SOC Prediction Framework for Electric Vehicles
Yaoyidi Wang, Niansheng Chen, Lei Rao, Dingyu Yang, Guangyu Fan, Songlin Cheng, Xiaoyong Song
MobiQuitous (1)7