Fuqiang Gu

dblp:183/5568 · DBLP profile ↗
← Back
34ranked-venue papers
9as first author
28since 2021 · last 2026
0000-0002-3408-982XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 13 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 12 · 4 first-author · 11 since 2021Systems, architecture and hardware · 7 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
YearPublicationVenuePosition
2026 MambaSeg: Harnessing Mamba for Accurate and Efficient Image-Event Semantic Segmentation
abstract
Semantic segmentation is a fundamental task in computer vision with wide-ranging applications, including autonomous driving and robotics. While RGB-based methods have achieved strong performance with CNNs and Transformers, their effectiveness degrades under fast motion, low-light, or high dynamic range conditions due to limitations of frame cameras. Event cameras offer complementary advantages such as high temporal resolution and low latency, yet lack color and texture, making them insufficient on their own. To address this, recent research has explored multimodal fusion of RGB and event data; however, many existing approaches are computationally expensive and focus primarily on spatial fusion, neglecting the temporal dynamics inherent in event streams. In this work, we propose MambaSeg, a novel dual-branch semantic segmentation framework that employs parallel Mamba encoders to efficiently model RGB images and event streams. To reduce cross-modal ambiguity, we introduce the Dual-Dimensional Interaction Module (DDIM), comprising a Cross-Spatial Interaction Module (CSIM) and a Cross-Temporal Interaction Module (CTIM), which jointly perform fine-grained fusion along both spatial and temporal dimensions. This design improves cross-modal alignment, reduces ambiguity, and leverages the complementary properties of each modality. Extensive experiments on the DDD17 and DSEC datasets demonstrate that MambaSeg achieves state-of-the-art segmentation performance while significantly reducing computational cost, showcasing its promise for efficient, scalable, and robust multimodal perception.
Fuqiang Gu, Yuanke Li, Xianlei Long, Kangping Ji, Chao Chen 0004, Qingyi Gu, Zhen-Liang Ni
AAAI1
2026 Pedestrian and Router Colocalization Framework Using Distributed-IMU-Based VDR and Wi-Fi RTT
abstract
Inertial navigation and WiFi are two common approaches for pedestrian localization. However, conventional pedestrian dead reckoning (PDR) and WiFi fingerprinting suffer from limited adaptability to different users and poor robustness to environmental changes, respectively. Recent deep-learning-based methods address pedestrian localization by modeling sequential dependencies in inertial data, but they typically rely on a single inertial measurement unit (IMU), which is insufficient to capture the spatial correlations of human skeletal motion. In parallel, the Fine Time Measurement (FTM) procedure in IEEE 802.11mc enables round-trip time (RTT)–based ranging and localization, yet the coordinates of WiFi routers still require labor-intensive prior surveying, limiting deployment flexibility. This paper presents a pedestrian and router colocalization framework that jointly estimates pedestrian trajectories and WiFi router positions. The proposed framework employs multiple body-worn IMUs and a long short-term memory (LSTM) network to learn both spatial and temporal dependencies in human motion, thereby enabling velocity dead reckoning (VDR). The VDR-estimated pedestrian velocity is then fused with WiFi RTT measurements through factor graph optimization (FGO), in which both pedestrian and router coordinates are treated as unknown variables. Experimental results demonstrate that the multi-IMU-based VDR effectively models pedestrian velocity, while WiFi RTT ranging constrains the long-term drift of VDR. The combined VDR/WiFi RTT framework achieves meter-level positioning accuracy in both indoor and outdoor environments, without requiring pre-surveyed router coordinates, and thus provides a promising solution for pedestrian localization in the Internet of Things (IoT) era.
Mingxi Wang, Fuqiang Gu, Liang Chen 0007, Ruizhi Chen, Shikai Jin
IEEE Internet Things J.4
2026 Policy distillation-based multiagent actor-critic for cooperative UAV path planning in complex environments
Huidong Liu, Jiangshan Ai, Xianlei Long, Yong Li 0023, Xiangwei Zhu, Fuqiang Gu
Knowl. Based Syst.7
2026 HierTFL: Hierarchical Scheduling for Industrial TSN-Enhanced Federated Learning System
abstract
The adoption of federated learning (FL) in industrial IoT (IIoT) facilitates the deployment of field-level industrial intelligence by multi-node collaborative and distributed learning. Numerous studies on FL primarily concentrate on enhancing model accuracy under non-independent and identically distributed data. Nevertheless, this focus is inadequate for time-critical industrial systems, as these systems necessitate real-time processing during the FL model training phase and any strategy that prioritizes incremental accuracy gains at the expense of latency risks violating real-time deadlines. However, the system heterogeneity of computing and communication capabilities will impose a formidable bottleneck to the overall time consumed for FL model training. In this paper, we present a FL-enabled Time-Sensitive IIoT (FETI) framework that integrates FL with Time-Sensitive Networking (TSN) to support the deterministic forwarding of FL flows in industry. Aiming to speed up FL convergence within a targeted accuracy gap, we formulate a heterogeneity-aware FL-TSN joint optimization problem, which is theoretically transformed into a stochastic mixed- integer programming problem solved at each FL round. To address this problem, we propose a hierarchical reinforcement learning-based scheduling scheme, called HierTFL, with two interacting layers of policies. With the assistance of a proposed spatial-temporal state encoder, the high-level policy dynamically selects client participants based on data quality and resource availability in each FL round, while the low-level policy optimizes TSN flow scheduling of each selected client using non-cumulative Bellman updates. Experimental results under industrial monitoring datasets have shown the effectivity of HierTFL in achieving a balanced trade-off between model precision and convergence time compared to existing benchmarks.
Songtao Guo, Fuqiang Gu, Pengzhan Zhou, Weiting Zhang
IEEE Trans. Mob. Comput.4
2025 SKE-Layout: Spatial Knowledge Enhanced Layout Generation with LLMs
abstract
Generating layouts from textual descriptions by large language models (LLMs) plays a crucial role in precise spatial reasoning-induced domains such as robotic object rearrangement and text-to-image generation. However, current methods face challenges in limited real-world examples, handling diverse layout descriptions and varying levels of granularity. To address these issues, a novel framework named Spatial Knowledge Enhanced Layout (SKE-Layout), is introduced. SKE-Layout integrates mixed spatial knowledge sources, leveraging both real and synthetic data to enhance spatial contexts. It utilizes diverse representations tailored to specific tasks and employs contrastive learning and multitask learning techniques for accurate spatial knowledge retrieval. This framework generates more accurate and fine-grained visual layouts for object rearrangement and text-to-image generation tasks, achieving improvements of 5%-30% compared to existing methods.
Nieqing Cao, Yan Ding 0002, Mengying Xie, Fuqiang Gu, Chao Chen 0004
CVPR5
2025 MoMa-Pos: An Efficient Object-Kinematic-Aware Base Placement Determination Framework for Mobile Manipulation
Beichen Shao, Nieqing Cao, Yan Ding 0002, Fuqiang Gu, Chao Chen 0004
ICA3PP (3)5
2025 OVA-Fields: Weakly Supervised Open-Vocabulary Affordance Fields for Robot Operational Part Detection
Heng Su, Mengying Xie, Nieqing Cao, Yan Ding 0002, Beichen Shao, Xianlei Long, Fuqiang Gu, Chao Chen 0004
ICCV7
2025 Heteroscedastic Bayesian Optimization-Based Dynamic PID Tuning for Accurate and Robust UAV Trajectory Tracking
abstract
Unmanned Aerial Vehicles (UAVs) play an important role in various applications, where precise trajectory tracking is crucial. However, conventional control algorithms for trajectory tracking often exhibit limited performance due to the underactuated, nonlinear, and highly coupled dynamics of quadrotor systems. To address these challenges, we propose HBO-PID, a novel control algorithm that integrates the Heteroscedastic Bayesian Optimization (HBO) framework with the classical PID controller to achieve accurate and robust trajectory tracking. By explicitly modeling input-dependent noise variance, the proposed method can better adapt to dynamic and complex environments, and therefore improve the accuracy and robustness of trajectory tracking. To accelerate the convergence of optimization, we adopt a two-stage optimization strategy that allow us to more efficiently find the optimal controller parameters. Through experiments in both simulation and real-world scenarios, we demonstrate that the proposed method significantly outperforms state-of-the-art (SOTA) methods. Compared to SOTA methods, it improves the position accuracy by 24.7% to 42.9%, and the angular accuracy by 40.9% to 78.4%.
Fuqiang Gu, Jiangshan Ai, Xianlei Long, Yan Li 0037, Tao Jiang 0018, Chao Chen 0004, Huidong Liu
IROS1
2025 SLTNet: Efficient Event-based Semantic Segmentation with Spike-driven Lightweight Transformer-based Networks
abstract
Event-based semantic segmentation has great potential in autonomous driving and robotics due to the advantages of event cameras, such as high dynamic range, low latency, and low power cost. Unfortunately, current artificial neural network (ANN)-based segmentation methods suffer from high computational demands, the requirements for image frames, and massive energy consumption, limiting their efficiency and application on resource-constrained edge/mobile platforms. To address these problems, we introduce SLTNet, a Spike-driven Lightweight Transformer-based Network designed for event-based semantic segmentation. Specifically, SLTNet is built on efficient spike-driven convolution blocks (SCBs) to extract rich semantic features while reducing the model’s parameters. Then, to enhance the long-range contextual feature interaction, we propose novel spike-driven transformer blocks (STBs) with binary mask operations. Based on these basic blocks, SLTNet employs a high-efficiency single-branch architecture while maintaining the low energy consumption of the Spiking Neural Network (SNN). Finally, extensive experiments on DDD17 and DSEC-Semantic datasets demonstrate that SLTNet outperforms state-of-the-art (SOTA) SNN-based methods by at most 9.06% and 9.39% mIoU, respectively, with extremely 4.58× lower energy consumption and 114 FPS inference speed. Our code is open-sourced and available at https://github.com/longxianlei/SLTNet-v1.0.
Xianlei Long, Xiaxin Zhu, Fangming Guo, Wanyi Zhang, Qingyi Gu, Chao Chen 0004, Fuqiang Gu
IROS7
2025 Adaptive multi-UAV cooperative path planning based on novel rotation artificial potential fields
Huidong Liu, Xianlei Long, Yong Li 0023, Jinjin Yan, Chao Chen 0004, Fuqiang Gu, Huayan Pu, Jun Luo 0006
Knowl. Based Syst.7
2025 Spike-BRGNet: Efficient and Accurate Event-Based Semantic Segmentation With Boundary Region-Guided Spiking Neural Networks
abstract
Event-based semantic segmentation in traffic scenes has attracted considerable attention in autonomous driving systems due to the advantages of event cameras such as high dynamic range, low latency, and low energy consumption. However, existing Artificial Neural Network (ANN)-based methods rely on conventional image frames, often neglecting the spatial-temporal dynamics inherent in event streams and consuming higher energy costs, significantly limiting their applicability in energy-constrained environments. In this study, we introduce Spike-BRGNet, a Spike-driven Boundary Region-Guided Network that efficiently extracts boundary information utilizing only events to guide the segmentation encoder, while preserving the energy efficiency of Spiking Neural Networks (SNNs). Specifically, to explore the implicit information from events, we design a three-branch spiking encoder that consists of semantic detail (SD), context aggregation (CA), and boundary aware (BA) branches to capture specific features. Then, a spiking multi-scale context aggregation (SMSCA) module is proposed to enhance the semantics of the CA branch. Finally, a novel boundary region-guided loss function and a dynamic surrogate gradient function, EvAF, are designed to optimize the model. Extensive experiments show that our model outperforms state-of-the-art (SOTA) SNN-based methods on DDD17 (+1.57%) and DSEC dataset (+1.91%). Furthermore, Spike-BRGNet consumes$17.76\times $less energy than ANN-based models, showing superior energy-saving performance.
Xianlei Long, Xiaxin Zhu, Fangming Guo, Chao Chen 0004, Xiangwei Zhu, Fuqiang Gu, Songyu Yuan, Chunlong Zhang
IEEE Trans. Circuits Syst. Video Technol.6
2025 Toward Accurate, Efficient, and Robust RGB-D Simultaneous Localization and Mapping in Challenging Environments
abstract
Visual Simultaneous Localization and Mapping (SLAM) is crucial to many applications such as self-driving vehicles and robot tasks. However, it is still challenging for existing visual SLAM approaches to achieve good performance in low-texture or illumination-changing scenes. In recent years, some researchers have turned to edge-based SLAM approaches to deal with the challenging scenes, which are more robust than feature-based and direct SLAM methods. Nevertheless, existing edge-based methods are computationally expensive and inferior than other visual SLAM systems in terms of accuracy. In this study, we propose EdgeSLAM, a novel RGB-D edge-based SLAM approach to deal with challenging scenarios that is efficient, accurate, and robust. EdgeSLAM is built on two innovative modules: efficient edge selection and adaptive robust motion estimation. The edge selection module can efficiently select a small set of edge pixels, which significantly improves the computational efficiency without sacrificing the accuracy. The motion estimation module improves the system's accuracy and robustness by adaptively handling outliers in motion estimation. Extensive experiments were conducted on TUM RGBD, ICL-NUIM and ETH3D datasets, and experimental results show that EdgeSLAM significantly outperforms five state-of-the-art (SOTA) methods in terms of efficiency, accuracy, and robustness, which achieves 29.17% accuracy improvements with a high processing speed of up to 120 FPS and a high positioning success rate of 97.06%.
Fuqiang Gu, Jianga Shang, Xianlei Long, Jiarui Dou, Chao Chen 0004, Huayan Pu, Jun Luo 0006
IEEE Trans. Robotics2
2024 End-to-end Flow Scheduling Optimization for Industrial 5G and TSN Integrated Networks
abstract
The integrated of the fifth generation (5G) and time-sensitive networking (TSN) is a promising approach to meet the requirements of deterministic forwarding with extremely low latency and high flexibility for the Industrial Internet of Things (IIoT). However, due to the large dissimilarity of protocol stacks, the 5G system normally serves as a logical bridge for TSN, performing hold-and-forward operations in Base Stations (BSs) under the current 5G-TSN frameworks. To provide ubiquitous and seamless connectivity for IIoT devices, this paper focuses on the integrated enhancement of 5G and TSN by optimizing the scheduling of time-sensitive flows with end-to-end latency of 5G-TSN transmission taken into consideration. Specifically, we propose a novel architecture named GF-CQF, which combines 5G Grant-free (GF) access with TSN Cyclic Queuing and Forwarding (CQF). Subsequently, the distributed flow scheduling problem based on GF-CQF architecture is established due to the high timeliness demands. To alleviate the impact of the uncertainty inherent in 5G channels on end-to-end deterministic transmission, a feature-aware decentralized real-time scheduling (FDRS) policy based on Multi-agent Reinforcement Learning is proposed. FDRS allows each agent at BS to adaptively allocate TSN injection slots for flows mainly based on dynamic 5G transmission performance and TSN network resource state so that TSN queue overflow can be avoided and flow delay constraints can be guaranteed. Simulations show that FDRS offers superior scheduling capabilities under the limited TSN queue resources.
Houling Liu, Fuqiang Gu, Qihao Li, Weiting Zhang, Songtao Guo
GLOBECOM3
2024 A Novel Wide-Area Multiobject Detection System with High-Probability Region Searching
abstract
In recent years, wide-area visual surveillance systems have been widely applied in various industrial and transportation scenarios. These systems, however, face significant challenges when implementing multi-object detection due to conflicts arising from the need for high-resolution imaging, efficient object searching, and accurate localization. To address these challenges, this paper presents a hybrid system that incorporates a wide-angle camera, a high-speed search camera, and a galvano-mirror. In this system, the wide-angle camera offers panoramic images as prior information, which helps the search camera capture detailed images of the targeted objects. This integrated approach enhances the overall efficiency and effectiveness of wide-area visual detection systems. Specifically, in this study, we introduce a wide-angle camera-based method to generate a panoramic probability map (PPM) for estimating high-probability regions of target object presence. Then, we propose a probability searching module that uses the PPM-generated prior information to dynamically adjust the sampling range and refine target coordinates based on uncertainty variance computed by the object detector. Finally, the integration of PPM and the probability searching module yields an efficient hybrid vision system capable of achieving 120 fps multi-object search and detection. Extensive experiments are conducted to verify the system’s effectiveness and robustness.
Xianlei Long, Chao Chen 0004, Fuqiang Gu, Qingyi Gu
ICRA4
2024 MobileHAR: A Lightweight and Efficient Human Activity Recognition Model based on Inverted Residual Inception Block
abstract
With the increasing demand for high precision and low power consumption in Human Activity Recognition (HAR) techniques, deep learning-based HAR models have emerged as the hottest research topics. Due to the excellent feature extraction and modeling abilities of deep learning models, which enable them to fit a wide variety of complex patterns. However, these models often require a large number of parameters, leading to high computational costs and longer processing time. These inherent factors pose significant challenges for resource-constraint edge devices to perform efficient HAR. To address these issues, we propose MobileHAR, which combines depthwise separable convolutions and novel Inverted Residual Inception Blocks (IRIB). This combination significantly reduces computational load and frequent memory access while maintaining high recognition accuracy. Then, we design a special class imbalance loss to supervise the model to pay more attention to imbalance classes. Finally, extensive experiments on several public datasets demonstrate that our method improves accuracy by 3.15% compared to traditional methods and requires only 0.15M parameters, which is at least four times fewer than the compared methods.
Fangming Guo, Fuqiang Gu, Xianlei Long
MSN3
2024 Improving Anomaly Scene Recognition with Large Vision-Language Models
Xianlei Long, Yan Li 0037, Chao Chen 0004, Fuqiang Gu, Songyu Yuan, Chunlong Zhang
WASA (3)5
2024 ORB-NeuroSLAM: A Brain-Inspired 3-D SLAM System Based on ORB Features
abstract
Intelligent navigation is a fundamental technology that enables unmanned systems to achieve autonomy in the intelligent era. However, existing navigation schemes suffer from high computational complexity and power consumption, as well as low robustness in complex or unknown environments. To address these challenges, this paper proposes a novel 3D brain-inspired simultaneous localization and mapping (SLAM) method, called ORB-NeuroSLAM, based on the oriented FAST and rotated BRIEF (ORB) features. The proposed method takes inspiration from the robust and low-power navigation capabilities of humans and animals. The ORB-NeuroSLAM leverages the ORB features of camera images to compute robot self-motion and visual cues. Then, continuous attractor neural networks (CANNs) model multilayered head direction cells and three-dimensional grid cells that exist in animal brains. These cells are utilized jointly to represent the robot poses. Efficient and robust methods for loop closure detection and experience map construction were also developed. The proposed method was verified on 10 KITTI datasets, and experimental results demonstrate that it outperforms state-of-the-art brain-inspired SLAM methods in terms of accuracy and efficiency. Additionally, it is comparable to state-of-the-art visual SLAM method ORB-SLAM3.
Dan Shen 0003, Gelu Liu, Fangwen Yu, Fuqiang Gu, Xiangwei Zhu
IEEE Internet Things J.5
2024 Coupling Makes Better: An Intertwined Neural Network for Taxi and Ridesourcing Demand Co-Prediction
abstract
While a variety of innovative travel modes, such as taxi service and ridesourcing service, have been launched to improve the transportation efficiency, people still encounter travel problems in real life. The major cause is the imbalance between transportation supply and demand. To strike a balance, it is well-recognized that an accurate and timely passenger demand prediction model is the foundation to enable high-level human intelligence (i.e., taxi drivers) or machine intelligence (i.e., ride-hailing platforms) to allocate resources in advance. Although quite a lot of deep models have been designed to model the complicated spatial and temporal dependencies in a data-driven way, they focus on the demand prediction of a single mode and ignore the fact that passengers may shift between different modes, especially between taxis and ridesourcing cars. In this paper, we target a co-prediction problem that considers the prediction of taxi and ridesourcing as two coupled and associated tasks, and propose a novel Temporal and Spatial Intertwined Network (TSIN) that consists of two twin components and an intertwined component. Each twin in the TSIN model is able to extract spatial and temporal dependencies from its corresponding travel mode separately (i.e., intra-mode features), and the in-between intertwined component is designed to bridge the twins and allow them to exchange information (i.e., inter-mode features), thus enabling better prediction. We first evaluate our model on four real-world datasets. Results demonstrate the outstanding performance of our model and the necessity to take into account the influence between modes. Based on an additional demand data from bike in NYC, we then discuss the generalizability in coupling more transportation modes. Further results demonstrate that our proposed intertwined neural network is highly flexible and extendable, and can yield better prediction performance.
Jie Zhao 0022, Chao Chen 0004, Wanyi Zhang, Fuqiang Gu, Songtao Guo, Jun Luo 0006, Yu Zheng 0004
IEEE Trans. Intell. Transp. Syst.5
2023 Deep Fingerprint Metric Learning for KNN-Based Indoor Localization
abstract
WiFi fingerprinting is a widely used technique for indoor localization, leveraging existing infrastructure to estimate a user's location based on received signal strength (RSS) measurements. The popular used WiFi fingerprinting is K-nearest Neighbors (KNN), which assumes that there is a linear relationship between the WiFi signal distance and real space distance. However, such assumption often fails to hold in complex indoor environments, resulting in significant positioning errors of KNN methods. In this paper, we propose DeepFML, a novel deep learning-based approach for KNN-based fingerprinting positioning, which learns a mapping function to transform raw RSS measurements into features, improving the consistency between the spatial distance in location space and the distance in feature space. Our experiments in complex indoor environments show that DeepFML outperforms state-of-the-art methods, improving the positioning accuracy by about 10% compared to the popular KNN method.
Fuqiang Gu, Jianga Shang
GLOBECOM1
2023 Efficient and Accurate Indoor/Outdoor Detection with Deep Spiking Neural Networks
abstract
Sensor-rich smartphones have facilitated a lot of services and applications. Indoor/Outdoor (IO) status serves as a critical foundation for various upstream tasks, including seamless pedestrian navigation, power management, and activity recognition. Nevertheless, achieving robust, efficient, and accurate IO detection remains challenging due to environmental complexities and device heterogeneity. To tackle this challenge, some researchers have turned to deep learning for IO detection, which can deal with complex scenarios and achieve high detection accuracy. However, deep learning methods are often blamed for their expensive computational cost. Therefore, in this paper, we introduce a novel efficient IO detection method-DeepSIO, which can detect IO status accurately and efficiently. Specifically, different from existing IO detection methods, DeepSIO is developed based on spiking neural networks (SNN) that are more biologically plausible and computationally efficient than other deep neural networks. To better capture useful features, we propose to utilize dense connections between SNN layers. Extensive experiments are conducted in three typical scenarios, and experimental results demonstrate that DeepSIO outperforms state-of-the-art methods, achieving an accuracy of about 99.7%. Moreover, it has better generalization ability and can adapt well to new environments and devices.
Fangming Guo, Xianlei Long, Kai Liu 0001, Chao Chen 0004, Haiyong Luo, Jianga Shang, Fuqiang Gu
GLOBECOM7
2023 EdgeVO: An Efficient and Accurate Edge-based Visual Odometry
abstract
Visual odometry is important for plenty of applications such as autonomous vehicles, and robot navigation. It is challenging to conduct visual odometry in textureless scenes or environments with sudden illumination changes where popular feature-based methods or direct methods cannot work well. To address this challenge, some edge-based methods have been proposed, but they usually struggle between the efficiency and accuracy. In this work, we propose a novel visual odometry approach called EdgeVO, which is accurate, efficient, and robust. By efficiently selecting a small set of edges with certain strategies, we significantly improve the computational efficiency without sacrificing the accuracy. Compared to existing edge-based method, our method can significantly reduce the computational complexity while maintaining similar accuracy or even achieving better accuracy. This is attributed to that our method removes useless or noisy edges. Experimental results on the TUM datasets indicate that EdgeVO significantly outperforms other methods in terms of efficiency, accuracy and robustness.
Jianga Shang, Kai Liu 0001, Chao Chen 0004, Fuqiang Gu
ICRA5
2023 An Enhanced Indoor Positioning Solution Using Dynamic Radio Fingerprinting Spatial Context Recognition
abstract
Radio fingerprinting positioning is widely used for smartphone location-based services and Internet of Things applications given its high availability and low cost. Radio fingerprinting-based algorithms, however, are subject to the forced matching problem and often yield estimated positions even when a user is actually located outside of the fingerprint region. A positioning solution in multistory buildings should be able to locate positions accurately on the current floor; but these methods may generate unreasonable positioning trajectories, such as irrationally passing through a wall, when fingerprinting positioning is fused the with inertial measurement unit to further improve the accuracy. A radio map must be surveyed dynamically on-the-move, as a dynamic fingerprint, to reduce the time costs and on-site workload. Unlike static fingerprint-based methods, dynamic fingerprinting samples are sparse. Thus, we propose an enhanced indoor positioning solution using spatial context knowledge, extracted from the sparse dynamic fingerprints. In the offline stage of radio map calculation, we extract the dynamic fingerprint features and store them in a spatial features database to reduce the computational time and storage space complexity. The proposed floor detection, region recognition, and path correction algorithms identify the online spatial contexts from the stored spatial features to improve positioning performance. This solution was applied on smartphones combining WiFi and Bluetooth low-energy radio signals in two typical scenarios. The experimental results show that the floor detection accuracy reached 99% while region recognition accuracy reached 90.75%. The positioning path correction method enhances the accuracy of smartphone indoor positioning from 3.27 to 2.56 m.
Xiaodong Gong, Jingbin Liu, Fuqiang Gu, Gege Huang
IEEE Internet Things J.4
2023 Enriching Large-Scale Trips With Fine-Grained Travel Purposes: A Semi-Supervised Deep Graph Embedding Framework
abstract
Knowing why people travel is meaningful for human mobility understanding and smart services development. Unfortunately, in real-world scenarios, trip purpose cannot be automatically collected on a large scale, thus calling for effective prediction models. Nevertheless, since passengers’ trip purposes in the city are diverse and complicated, the prediction is very difficult especially at a fine-grained level. Worse still, the informative data sources and real purpose-labels about trips are commonly limited for model learning. To resolve the dilemma, we propose a semi-supervised deep embedding framework for predicting fine-grained trip purposes on a large scale. Specifically, we first derive augmented trip contexts from the vehicle’s GPS trajectory and public POI check-in data, then convert POI contexts into the graph structure. We further establish aDual-AttentionGraphEmbedding Network withAutoencoder architecture (DAGE-A) to accomplish prediction and reconstruction simultaneously, in which category-aware graph attention networks are devised to model the POI semantics at trip’s origin/destination and extract complementary knowledge from unlabeled trips; and soft-attention is employed to aggregate different trip semantics appropriately for the final prediction. We conduct extensive experiments in Beijing and Shanghai, and results show our framework outperforms state-of-the-arts and could reduce labelling efforts by up to 20%. We also find that our model is generalized at different times and locations, and the performance varies for different trip purposes.
Chengwu Liao, Chao Chen 0004, Suiming Guo, Leye Wang, Fuqiang Gu, Ke Xu 0002
IEEE Trans. Intell. Transp. Syst.5
2023 Towards Robust WiFi Fingerprint-Based Vehicle Tracking in Dynamic Indoor Parking Environments: An Online Learning Framework
abstract
The variation of wireless signal in dynamic indoor parking environments may seriously compromise the performance of fingerprint-based localization methods. In this regard, this paper investigates the problem of robust WiFi fingerprint-based vehicle tracking in dynamic indoor parking environments, aiming at designing an online learning framework to continuously train the localization model and counteract the effect of signal variation. Specifically, a Hidden Markov Model (HMM) based Online Evaluation (HOE) method is firstly proposed to assess the accuracy of localization results by measuring the inconsistency of locations inferred by WiFi fingerprinting and Dead Reckoning (DR). Further, an Online Transfer Learning (OTL) algorithm is designed to improve the robustness of the fingerprinting localization, which consists of a weight allocation scheme to combine two classification models (i.e., the batch model and the online model) and an instance-based transferring scheme to resample the offline fingerprints and retrain the batch model. Finally, we implement the system prototype and give comprehensive performance evaluation, which demonstrates that the proposed solutions can outperform the state-of-the-art localization algorithms around 28%$\sim$58% on vehicle tracking accuracy in dynamic indoor parking environments.
Kai Liu 0001, Feiyu Jin, Junbo Hu, Ruitao Xie, Fuqiang Gu, Songtao Guo, Jiangtao Luo
IEEE Trans. Mob. Comput.5
2022 Apache ShardingSphere: A Holistic and Pluggable Platform for Data Sharding
abstract
Traditional relational databases are nowadays over-whelmed by the increasing data volume and concurrent access. NoSQL databases can manage large-scale data, but most of them do not support complete transactions and standard SQL languages. NewSQL is proposed for both high scalability and transactional properties with SQL languages support. One type of NewSQL builds distributed systems from scratch, which is too radical for some critical applications. The other type of NewSQL, i.e., data sharding among relational databases, is a better option for these scenarios. This paper presents Apache ShardingSphere, the first top-level open-source platform for data sharding in Apache, which enables developers to use sharded databases like one database. Specifically Apache, ShardingSphere integrates six databases and designs and implements a complete SQL engine to route requests correctly and intelligently. Additionally it encapsulates three types of distributed transactions and provides two adaptors for different scenarios. Moreover it proposes a novel AutoTable strategy and a query language i.e DistSQL allowing database maintainers to easily configure the sharded databases. Further-more it provides many other pluggable features to better shard data. Extensive experiments are conducted using two famous benchmarking tools proving that Apache ShardingSphere is more efficient than eight state-of-the-art systems in our settings. All experimental source codes are publicly released. More than 170 companies are currently using Apache ShardingSphere.
Juan Pan, Junwen Liu, Nianjun Sun, Shanmin Wang, Chao Chen 0004, Fuqiang Gu, Songtao Guo
ICDE9
2021 Distributed Spatio-Temporal k Nearest Neighbors Join
abstract
The rapid development of positioning technology produces an extremely large volume of spatio-temporal data with various geometry types such as point, line string, polygon, or a mixed combination of them. As one of the most basic but time-consuming operations, k nearest neighbors join (kNN join) has attracted much attention. However, most existing works for kNN join either ignore temporal information or consider point data only.
Rubin Wang, Junwen Liu, Zisheng Yu, Huajun He, Tianfu He, Sijie Ruan, Jie Bao 0003, Chao Chen 0004, Fuqiang Gu, Liang Hong 0001, Yu Zheng 0004
SIGSPATIAL/GIS10
2021 EventDrop: Data Augmentation for Event-based Learning
abstract
The advantages of event-sensing over conventional sensors (e.g., higher dynamic range, lower time latency, and lower power consumption) have spurred research into machine learning for event data. Unsurprisingly, deep learning has emerged as a competitive methodology for learning with event sensors; in typical setups, discrete and asynchronous events are first converted into frame-like tensors on which standard deep networks can be applied. However, over-fitting remains a challenge, particularly since event datasets remain small relative to conventional datasets (e.g., ImageNet). In this paper, we introduce EventDrop, a new method for augmenting asynchronous event data to improve the generalization of deep models. By dropping events selected with various strategies, we are able to increase the diversity of training data (e.g., to simulate various levels of occlusion). From a practical perspective, EventDrop is simple to implement and computationally low-cost. Experiments on two event datasets (N-Caltech101 and N-Cars) demonstrate that EventDrop can significantly improve the generalization performance across a variety of deep networks.
Fuqiang Gu, Weicong Sng, Xuke Hu, Fangwen Yu
IJCAI1
2021 Tagging the main entrances of public buildings based on OpenStreetMap and binary imbalanced learning
abstract
Determining the location of a building’s entrance is crucial to location-based services, such as wayfinding for pedestrians. Unfortunately, entrance information is often missing from current mainstream map providers such as Google Maps. Frequently, automatic approaches for detecting building entrances are based on street-level images that are not widely available. To address this issue, we propose a more general approach for inferring the main entrances of public buildings based on the association between spatial elements extracted from OpenStreetMap. In particular, we adopt three binary classification approaches, weighted random forest, balanced random forest, and smooth-boost to model the association relationship. There are two types of features considered in the classification: intrinsic features derived from building footprints and extrinsic features derived from spatial contexts, such as roads, green spaces, bicycle parking areas, and neighboring buildings. We conducted extensive experiments on 320 public buildings with an average perimeter of 350 m. The experimental results showed that the locations of building entrances estimated by the weighted random forest and balanced random forest models have a mean linear distance error of 21 m and a mean path distance error of 22 m, ruling out 90% of the incorrect locations of the main entrance of buildings.
Xuke Hu, Alexey Noskov, Hongchao Fan, Tessio Novack, Hao Li 0019, Fuqiang Gu, Jianga Shang, Alexander Zipf
Int. J. Geogr. Inf. Sci.6
2020 TactileSGNet: A Spiking Graph Neural Network for Event-based Tactile Object Recognition
abstract
Tactile perception is crucial for a variety of robot tasks including grasping and in-hand manipulation. New advances in flexible, event-driven, electronic skins may soon endow robots with touch perception capabilities similar to humans. These electronic skins respond asynchronously to changes (e.g., in pressure, temperature), and can be laid out irregularly on the robot's body or end-effector. However, these unique features may render current deep learning approaches such as convolutional feature extractors unsuitable for tactile learning. In this paper, we propose a novel spiking graph neural network for event-based tactile object recognition. To make use of local connectivity of taxels, we present several methods for organizing the tactile data in a graph structure. Based on the constructed graphs, we develop a spiking graph convolutional network. The event-driven nature of spiking neural network makes it arguably more suitable for processing the event-based data. Experimental results on two tactile datasets show that the proposed method outperforms other state-of-the-art spiking methods, achieving high accuracies of approximately 90% when classifying a variety of different household objects.
Fuqiang Gu, Weicong Sng, Tasbolat Taunyazov, Harold Soh
IROS1
2020 Path distance-based map matching for Wi-Fi fingerprinting positioning
Pan Chen 0004, Xiaoping Zheng, Fuqiang Gu, Jianga Shang
Future Gener. Comput. Syst.3
2020 Landmark Graph-Based Indoor Localization
abstract
Indoor localization is important for a variety of applications, such as location-based services, mobile social networks, and emergency response. Fusing spatial information is an effective way to achieve accurate indoor localization with little or with no need for extra hardware. However, the existing indoor localization methods that make use of spatial information are either computationally expensive or sensitive to the completeness of landmarks. In this article, we propose a novel, low-cost, high-accuracy indoor localization method based on a landmark graph. The experimental results show that the proposed method outperforms the state-of-the-art methods.
Fuqiang Gu, Shahrokh Valaee, Kourosh Khoshelham, Jianga Shang, Rui Zhang 0003
IEEE Internet Things J.1
2019 ZeeFi: Zero-Effort Floor Identification with Deep Learning for Indoor Localization
abstract
The knowledge of the floor-level location of a user in a multi-storey building is important for many applications, especially for emergency response. Existing floor identification systems suffer from a variety of limitations such as low accuracy, the need for a time-consuming site survey, assumption of user encounters, knowledge of the initial floor, and/or poor applicability. In this paper, we propose a novel, zero-effort, deep learning-based floor identification system, called ZeeFi. The proposed system uses the widely-available smartphone sensing to identify on which floor a user is located. By recognizing the ground floor automatically, the proposed system does not require site survey, initial floor knowledge, and other assumptions. To achieve accurate floor identification performance, we have developed a deep learning-based method. Experimental results show that the proposed system outperforms the state-of-the-art systems, and is very promising for large-scale deployment.
Fuqiang Gu, Jörg Blankenbach, Kourosh Khoshelham, Jan Grottke, Shahrokh Valaee
GLOBECOM1
2018 Locomotion Activity Recognition Using Stacked Denoising Autoencoders
abstract
Locomotion activity recognition (LAR) is important for a number of applications, such as indoor localization, fitness tracking, and aged care. Existing methods usually use handcrafted features, which requires expert knowledge and is laborious, and the achieved result might still be suboptimal. To relieve the burden of designing and selecting features, we propose a deep learning method for LAR by using data from multiple sensors available on most smart devices. Experimental results show that the proposed method, which learns useful features automatically, outperforms conventional classifiers that require the hand-engineering of features. We also show that the combination of sensor data from four sensors (accelerometer, gyroscope, magnetometer, and barometer) achieves a higher accuracy than other combinations or individual sensors.
Fuqiang Gu, Kourosh Khoshelham, Shahrokh Valaee, Jianga Shang, Rui Zhang 0003
IEEE Internet Things J.1
2017 Locomotion activity recognition: A deep learning approach
abstract
Human activity recognition is important for a large number of applications including indoor localization. Existing methods usually involve manually-designed features, which require expert knowledge and are laborious. Also, previous works use only the accelerometer for activity recognition, which may fail to recognize some complex activities. In this paper, we propose a deep learning-based method for locomotion activity recognition by using the combination of data from multiple smartphone built-in sensors. Eight types of locomotion activities are identified including the new `False Motion' activity introduced for the first time in this work. Experimental results show that the proposed method, which learns useful features automatically, outperforms conventional classifiers that require hand-engineering of features. Also, using data from multiple sensors helps to improve recognition accuracy by about 10% compared to that using accelerometer data only.
Fuqiang Gu, Kourosh Khoshelham, Shahrokh Valaee
PIMRC1