Xianlei Long

dblp:236/5659 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0002-2623-8768ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 10 since 2021Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MambaSeg: Harnessing Mamba for Accurate and Efficient Image-Event Semantic Segmentation
abstract
Semantic segmentation is a fundamental task in computer vision with wide-ranging applications, including autonomous driving and robotics. While RGB-based methods have achieved strong performance with CNNs and Transformers, their effectiveness degrades under fast motion, low-light, or high dynamic range conditions due to limitations of frame cameras. Event cameras offer complementary advantages such as high temporal resolution and low latency, yet lack color and texture, making them insufficient on their own. To address this, recent research has explored multimodal fusion of RGB and event data; however, many existing approaches are computationally expensive and focus primarily on spatial fusion, neglecting the temporal dynamics inherent in event streams. In this work, we propose MambaSeg, a novel dual-branch semantic segmentation framework that employs parallel Mamba encoders to efficiently model RGB images and event streams. To reduce cross-modal ambiguity, we introduce the Dual-Dimensional Interaction Module (DDIM), comprising a Cross-Spatial Interaction Module (CSIM) and a Cross-Temporal Interaction Module (CTIM), which jointly perform fine-grained fusion along both spatial and temporal dimensions. This design improves cross-modal alignment, reduces ambiguity, and leverages the complementary properties of each modality. Extensive experiments on the DDD17 and DSEC datasets demonstrate that MambaSeg achieves state-of-the-art segmentation performance while significantly reducing computational cost, showcasing its promise for efficient, scalable, and robust multimodal perception.
Fuqiang Gu, Yuanke Li, Xianlei Long, Kangping Ji, Chao Chen 0004, Qingyi Gu, Zhen-Liang Ni
AAAI3
2026 Policy distillation-based multiagent actor-critic for cooperative UAV path planning in complex environments
Huidong Liu, Jiangshan Ai, Xianlei Long, Yong Li 0023, Xiangwei Zhu, Fuqiang Gu
Knowl. Based Syst.4
2025 OVA-Fields: Weakly Supervised Open-Vocabulary Affordance Fields for Robot Operational Part Detection
Heng Su, Mengying Xie, Nieqing Cao, Yan Ding 0002, Beichen Shao, Xianlei Long, Fuqiang Gu, Chao Chen 0004
ICCV6
2025 Heteroscedastic Bayesian Optimization-Based Dynamic PID Tuning for Accurate and Robust UAV Trajectory Tracking
abstract
Unmanned Aerial Vehicles (UAVs) play an important role in various applications, where precise trajectory tracking is crucial. However, conventional control algorithms for trajectory tracking often exhibit limited performance due to the underactuated, nonlinear, and highly coupled dynamics of quadrotor systems. To address these challenges, we propose HBO-PID, a novel control algorithm that integrates the Heteroscedastic Bayesian Optimization (HBO) framework with the classical PID controller to achieve accurate and robust trajectory tracking. By explicitly modeling input-dependent noise variance, the proposed method can better adapt to dynamic and complex environments, and therefore improve the accuracy and robustness of trajectory tracking. To accelerate the convergence of optimization, we adopt a two-stage optimization strategy that allow us to more efficiently find the optimal controller parameters. Through experiments in both simulation and real-world scenarios, we demonstrate that the proposed method significantly outperforms state-of-the-art (SOTA) methods. Compared to SOTA methods, it improves the position accuracy by 24.7% to 42.9%, and the angular accuracy by 40.9% to 78.4%.
Fuqiang Gu, Jiangshan Ai, Xianlei Long, Yan Li 0037, Tao Jiang 0018, Chao Chen 0004, Huidong Liu
IROS4
2025 SLTNet: Efficient Event-based Semantic Segmentation with Spike-driven Lightweight Transformer-based Networks
abstract
Event-based semantic segmentation has great potential in autonomous driving and robotics due to the advantages of event cameras, such as high dynamic range, low latency, and low power cost. Unfortunately, current artificial neural network (ANN)-based segmentation methods suffer from high computational demands, the requirements for image frames, and massive energy consumption, limiting their efficiency and application on resource-constrained edge/mobile platforms. To address these problems, we introduce SLTNet, a Spike-driven Lightweight Transformer-based Network designed for event-based semantic segmentation. Specifically, SLTNet is built on efficient spike-driven convolution blocks (SCBs) to extract rich semantic features while reducing the model’s parameters. Then, to enhance the long-range contextual feature interaction, we propose novel spike-driven transformer blocks (STBs) with binary mask operations. Based on these basic blocks, SLTNet employs a high-efficiency single-branch architecture while maintaining the low energy consumption of the Spiking Neural Network (SNN). Finally, extensive experiments on DDD17 and DSEC-Semantic datasets demonstrate that SLTNet outperforms state-of-the-art (SOTA) SNN-based methods by at most 9.06% and 9.39% mIoU, respectively, with extremely 4.58× lower energy consumption and 114 FPS inference speed. Our code is open-sourced and available at https://github.com/longxianlei/SLTNet-v1.0.
Xianlei Long, Xiaxin Zhu, Fangming Guo, Wanyi Zhang, Qingyi Gu, Chao Chen 0004, Fuqiang Gu
IROS1
2025 Adaptive multi-UAV cooperative path planning based on novel rotation artificial potential fields
Huidong Liu, Xianlei Long, Yong Li 0023, Jinjin Yan, Chao Chen 0004, Fuqiang Gu, Huayan Pu, Jun Luo 0006
Knowl. Based Syst.2
2025 Spike-BRGNet: Efficient and Accurate Event-Based Semantic Segmentation With Boundary Region-Guided Spiking Neural Networks
abstract
Event-based semantic segmentation in traffic scenes has attracted considerable attention in autonomous driving systems due to the advantages of event cameras such as high dynamic range, low latency, and low energy consumption. However, existing Artificial Neural Network (ANN)-based methods rely on conventional image frames, often neglecting the spatial-temporal dynamics inherent in event streams and consuming higher energy costs, significantly limiting their applicability in energy-constrained environments. In this study, we introduce Spike-BRGNet, a Spike-driven Boundary Region-Guided Network that efficiently extracts boundary information utilizing only events to guide the segmentation encoder, while preserving the energy efficiency of Spiking Neural Networks (SNNs). Specifically, to explore the implicit information from events, we design a three-branch spiking encoder that consists of semantic detail (SD), context aggregation (CA), and boundary aware (BA) branches to capture specific features. Then, a spiking multi-scale context aggregation (SMSCA) module is proposed to enhance the semantics of the CA branch. Finally, a novel boundary region-guided loss function and a dynamic surrogate gradient function, EvAF, are designed to optimize the model. Extensive experiments show that our model outperforms state-of-the-art (SOTA) SNN-based methods on DDD17 (+1.57%) and DSEC dataset (+1.91%). Furthermore, Spike-BRGNet consumes$17.76\times $less energy than ANN-based models, showing superior energy-saving performance.
Xianlei Long, Xiaxin Zhu, Fangming Guo, Chao Chen 0004, Xiangwei Zhu, Fuqiang Gu, Songyu Yuan, Chunlong Zhang
IEEE Trans. Circuits Syst. Video Technol.1
2025 Toward Accurate, Efficient, and Robust RGB-D Simultaneous Localization and Mapping in Challenging Environments
abstract
Visual Simultaneous Localization and Mapping (SLAM) is crucial to many applications such as self-driving vehicles and robot tasks. However, it is still challenging for existing visual SLAM approaches to achieve good performance in low-texture or illumination-changing scenes. In recent years, some researchers have turned to edge-based SLAM approaches to deal with the challenging scenes, which are more robust than feature-based and direct SLAM methods. Nevertheless, existing edge-based methods are computationally expensive and inferior than other visual SLAM systems in terms of accuracy. In this study, we propose EdgeSLAM, a novel RGB-D edge-based SLAM approach to deal with challenging scenarios that is efficient, accurate, and robust. EdgeSLAM is built on two innovative modules: efficient edge selection and adaptive robust motion estimation. The edge selection module can efficiently select a small set of edge pixels, which significantly improves the computational efficiency without sacrificing the accuracy. The motion estimation module improves the system's accuracy and robustness by adaptively handling outliers in motion estimation. Extensive experiments were conducted on TUM RGBD, ICL-NUIM and ETH3D datasets, and experimental results show that EdgeSLAM significantly outperforms five state-of-the-art (SOTA) methods in terms of efficiency, accuracy, and robustness, which achieves 29.17% accuracy improvements with a high processing speed of up to 120 FPS and a high positioning success rate of 97.06%.
Fuqiang Gu, Jianga Shang, Xianlei Long, Jiarui Dou, Chao Chen 0004, Huayan Pu, Jun Luo 0006
IEEE Trans. Robotics4
2024 A Novel Wide-Area Multiobject Detection System with High-Probability Region Searching
abstract
In recent years, wide-area visual surveillance systems have been widely applied in various industrial and transportation scenarios. These systems, however, face significant challenges when implementing multi-object detection due to conflicts arising from the need for high-resolution imaging, efficient object searching, and accurate localization. To address these challenges, this paper presents a hybrid system that incorporates a wide-angle camera, a high-speed search camera, and a galvano-mirror. In this system, the wide-angle camera offers panoramic images as prior information, which helps the search camera capture detailed images of the targeted objects. This integrated approach enhances the overall efficiency and effectiveness of wide-area visual detection systems. Specifically, in this study, we introduce a wide-angle camera-based method to generate a panoramic probability map (PPM) for estimating high-probability regions of target object presence. Then, we propose a probability searching module that uses the PPM-generated prior information to dynamically adjust the sampling range and refine target coordinates based on uncertainty variance computed by the object detector. Finally, the integration of PPM and the probability searching module yields an efficient hybrid vision system capable of achieving 120 fps multi-object search and detection. Extensive experiments are conducted to verify the system’s effectiveness and robustness.
Xianlei Long, Chao Chen 0004, Fuqiang Gu, Qingyi Gu
ICRA1
2024 MobileHAR: A Lightweight and Efficient Human Activity Recognition Model based on Inverted Residual Inception Block
abstract
With the increasing demand for high precision and low power consumption in Human Activity Recognition (HAR) techniques, deep learning-based HAR models have emerged as the hottest research topics. Due to the excellent feature extraction and modeling abilities of deep learning models, which enable them to fit a wide variety of complex patterns. However, these models often require a large number of parameters, leading to high computational costs and longer processing time. These inherent factors pose significant challenges for resource-constraint edge devices to perform efficient HAR. To address these issues, we propose MobileHAR, which combines depthwise separable convolutions and novel Inverted Residual Inception Blocks (IRIB). This combination significantly reduces computational load and frequent memory access while maintaining high recognition accuracy. Then, we design a special class imbalance loss to supervise the model to pay more attention to imbalance classes. Finally, extensive experiments on several public datasets demonstrate that our method improves accuracy by 3.15% compared to traditional methods and requires only 0.15M parameters, which is at least four times fewer than the compared methods.
Fangming Guo, Fuqiang Gu, Xianlei Long
MSN5
2024 Improving Anomaly Scene Recognition with Large Vision-Language Models
Xianlei Long, Yan Li 0037, Chao Chen 0004, Fuqiang Gu, Songyu Yuan, Chunlong Zhang
WASA (3)2
2024 HTQ: Exploring the High-Dimensional Trade-Off of mixed-precision quantization
Zhikai Li, Xianlei Long, Junrui Xiao, Qingyi Gu
Pattern Recognit.2
2023 Efficient and Accurate Indoor/Outdoor Detection with Deep Spiking Neural Networks
abstract
Sensor-rich smartphones have facilitated a lot of services and applications. Indoor/Outdoor (IO) status serves as a critical foundation for various upstream tasks, including seamless pedestrian navigation, power management, and activity recognition. Nevertheless, achieving robust, efficient, and accurate IO detection remains challenging due to environmental complexities and device heterogeneity. To tackle this challenge, some researchers have turned to deep learning for IO detection, which can deal with complex scenarios and achieve high detection accuracy. However, deep learning methods are often blamed for their expensive computational cost. Therefore, in this paper, we introduce a novel efficient IO detection method-DeepSIO, which can detect IO status accurately and efficiently. Specifically, different from existing IO detection methods, DeepSIO is developed based on spiking neural networks (SNN) that are more biologically plausible and computationally efficient than other deep neural networks. To better capture useful features, we propose to utilize dense connections between SNN layers. Extensive experiments are conducted in three typical scenarios, and experimental results demonstrate that DeepSIO outperforms state-of-the-art methods, achieving an accuracy of about 99.7%. Moreover, it has better generalization ability and can adapt well to new environments and devices.
Fangming Guo, Xianlei Long, Kai Liu 0001, Chao Chen 0004, Haiyong Luo, Jianga Shang, Fuqiang Gu
GLOBECOM2
2022 Dual-discriminator adversarial framework for data-free quantization
Zhikai Li, Liping Ma, Xianlei Long, Junrui Xiao, Qingyi Gu
Neurocomputing3
2022 Boosting semi-supervised face recognition with raw faces
Yunze Chen, Junjie Huang 0005, Xianlei Long, Qingyi Gu
Image Vis. Comput.4
2020 Natural Scene Facial Expression Recognition with Dimension Reduction Network
abstract
As an external manifestation of human emotions, expression recognition plays an important role in human-computer interaction. Although existing expression recognition methods performs perfectly on constrained frontal faces, there are still many challenges in expression recognition in natural scenes due to different unrestricted conditions. Expression classification belongs to a pattern recognition problem where intra-class distance is greater than the inter-class distance, which leads to severe over-fitting when using neural networks for expression recognition. This paper proposes a novel net-work structure called Dimension Reduction Network which can effectively reduce generalization error. By adding a data dimension reduction module before the general classification network, a lot of redundant information is filtered, and only useful information is left. This can reduce the interference by irrelevant information when performing classification tasks and reduce generalization error. The proposed method does not require any modification to the classification network, only a small dimension reduction module needs to be added in front of the classification network. However, it can effectively reduce generalization error. We designed big and tiny versions of Dimension Reduction Network, both exceeds our baseline on AffectNet data set. The big version of our proposed method surpassed the state-of-the-art methods by more than 1.2% on AffectNet data set. Our code will open source3when the paper is accepted.
Shenhua Hu, Yiming Hu, Xianlei Long, Mengjuan Chen, Qingyi Gu
ICRA4
2019 Cluster Regularized Quantization for Deep Networks Compression
abstract
Deep neural networks (DNNs) have achieved great success in a wide range of computer vision areas, but the applications to mobile devices is limited due to their high storage and computational cost. Much efforts have been devoted to compress DNNs. In this paper, we propose a simple yet effective method for deep networks compression, named Cluster Regularized Quantization (CRQ), which can reduce the presentation precision of a full-precision model to ternary values without significant accuracy drop. In particular, the proposed method aims at reducing the quantization error by introducing a cluster regularization term, which is imposed on the full-precision weights to enable them naturally concentrate around the target values. Through explicitly regularizing the weights during the re-training stage, the full-precision model can achieve the smooth transition to the low-bit one. Comprehensive experiments on benchmark datasets demonstrate the effectiveness of the proposed method.
Yiming Hu, Xianlei Long, Shenhua Hu, Jiagang Zhu, Xingang Wang 0003, Qingyi Gu
ICIP3