Ling Pei

dblp:39/10180 · DBLP profile ↗
← Back
45ranked-venue papers
2as first author
34since 2021 · last 2026
0000-0003-1025-1260ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 18 since 2021Systems, architecture and hardware · 16 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 9 since 2021Computer networks · 9 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RadarLLM: Empowering Large Language Models to Understand Human Motion from Millimeter-wave Point Cloud Sequence
abstract
Millimeter-wave radar offers a privacy-preserving and environment-robust alternative to vision-based sensing, enabling human motion analysis in challenging conditions such as low light, occlusions, rain, or smoke. However, its sparse point clouds pose significant challenges for semantic understanding. We present RadarLLM, the first framework that leverages large language models (LLMs) for human motion understanding from radar signals. RadarLLM introduces two key innovations: (1) a motion-guided radar tokenizer based on our Aggregate VQ-VAE architecture, integrating deformable body templates and masked trajectory modeling to convert spatial-temporal radar sequences into compact semantic tokens; and (2) a radar-aware language model that establishes cross-modal alignment between radar and text in a shared embedding space. To overcome the scarcity of paired radar-text data, we generate a realistic radar-text dataset from motion-text datasets with a physics-aware synthesis pipeline. Extensive experiments on both synthetic and real-world benchmarks show that RadarLLM achieves state-of-the-art performance, enabling robust and interpretable motion understanding under privacy and visibility constraints, even in adverse environments.
Zengyuan Lai, Songpengcheng Xia, Lizhou Lin, Renwen Wang, Jianran Liu, Qi Wu 0007, Ling Pei
AAAI9
2026 DMAF-Net: A dual-modal attention fusion network for high-accuracy dense laser stripe matching
Junlin Lai, Ling Pei, Zhanzheng Ren, Zhenming Lv
Adv. Eng. Informatics3
2026 Geometric Unscented Particle Filters on Lie Groups for State Estimation
abstract
This article proposes two types of unscented particle filters (UPFs) that leverage unscented transformation (UT) from a geometric perspective to compute the proposal distribution. An UPF on Lie groups is first developed. Specifically, both the propagation of the sigma points and the computation of the mean and covariance are performed on the Lie groups, while the weight update and resampling are conducted on the Lie algebra. Second, we introduce the log-linear property of group elements to streamline particle propagation by reducing redundant operations, thereby optimizing the proposed UPF framework. In the update process, intermittent measurements that are caused by factors such as packet dropouts and stochastic sensor scheduling are considered. While lowering computational demands, these measurements pose challenges to filter stability. To this end, the introduced property is used to prove that the estimation error remains bounded under certain assumptions. We further establish a critical threshold for the arrival rate of intermittent measurements and derive an upper bound for the expected state error covariance. Moreover, a detailed computational complexity analysis is conducted to evaluate the efficiency of the proposed method. Finally, with the original method serving as a benchmark, simulation and real-world GNSS/INS integrated navigation experiments confirm that the redesigned approach delivers comparable performance and significantly improved computational efficiency.
Tao Li 0052, Yuqiang Jin, Ling Pei, Wen-An Zhang 0001, Trieu-Kien Truong
IEEE Trans. Cybern.5
2025 EnvPoser: Environment-aware Realistic Human Motion Estimation from Sparse Observations with Uncertainty Modeling
abstract
Estimating full-body motion using the tracking signals of head and hands from VR devices holds great potential for various applications. However, the sparsity and unique distribution of observations present a significant challenge, resulting in an ill-posed problem with multiple feasible solutions (i.e., hypotheses). This amplifies uncertainty and ambiguity in full-body motion estimation, especially for the lower-body joints. Therefore, we propose a new method, EnvPoser, that employs a two-stage framework to perform full-body motion estimation using sparse tracking signals and pre-scanned environment from VR devices. EnvPoser models the multi-hypothesis nature of human motion through an uncertainty-aware estimation module in the first stage. In the second stage, we refine these multi-hypothesis estimates by integrating semantic and geometric environmental constraints, ensuring that the final motion estimation aligns realistically with both the environmental context and physical interactions. Qualitative and quantitative experiments on two public datasets demonstrate that our method achieves state-of-the-art performance, highlighting significant improvements in human motion estimation within motion-environment interaction scenarios. Project page: https://xspc.github.io/EnvPoser/.
Songpengcheng Xia, Zhuo Su 0006, Xiaozheng Zheng, Guidong Wang, Qi Wu 0007, Ling Pei
CVPR10
2025 Spatiotemporal Decoupling for Efficient Vision-Based Occupancy Forecasting
abstract
The task of occupancy forecasting (OCF) involves utilizing past and present perception data to predict future occupancy states of autonomous vehicle surrounding environments, which is critical for downstream tasks such as obstacle avoidance and path planning. Existing 3D OCF approaches struggle to predict plausible spatial details for movable objects and suffer from slow inference speeds due to neglecting the bias and uneven distribution of changing occupancy states in both space and time. In this paper, we propose a novel spatiotemporal decoupling vision-based paradigm to explicitly tackle the bias and achieve both effective and efficient 3D OCF. To tackle spatial bias in empty areas, we introduce a novel spatial representation that decouples the conventional dense 3D format into 2D bird’s-eye view (BEV) occupancy with corresponding height values, enabling 3D OCF derived only from 2D predictions thus enhancing efficiency. To reduce temporal bias on static voxels, we design temporal decoupling to improve end-to-end OCF by temporally associating instances via predicted flows. We develop an efficient multi-head network EfficientOCF to achieve 3D OCF with our devised spatiotemporally decoupled representation. A new metric, conditional IoU (C-IoU), is also introduced to provide a robust 3D OCF performance assessment, especially in datasets with missing or incomplete annotations. The experimental results demonstrate that EfficientOCF surpasses existing baseline methods on accuracy and efficiency, achieving state-of-the-art performance with a fast inference time of 82.33 ms with a single GPU. Our code is released at: https://github.com/BIT-XJY/EfficientOCF.
Xieyuanli Chen, Junyi Ma, Jintao Xu 0001, Yue Wang 0020, Ling Pei
CVPR7
2025 IMOST: Incremental Memory Mechanism with Online Self-Supervision for Continual Traversability Learning
abstract
Traversability estimation is the foundation of path planning for a general navigation system. However, complex and dynamic environments pose challenges for the latest methods using self-supervised learning (SSL) technique. Firstly, existing SSL-based methods generate sparse annotations lacking detailed boundary information. Secondly, their strategies focus on hard samples for rapid adaptation, leading to forgetting and biased predictions. In this work, we propose IMOST, a continual traversability learning framework composed of two key modules: incremental dynamic memory (IDM) and self-supervised annotation (SSA). By mimicking human memory mechanisms, IDM allocates novel data samples to new clusters according to information expansion criterion. It also updates clusters based on diversity rule, ensuring a representative characterization of new scene. This mechanism enhances scene-aware knowledge diversity while maintaining a compact memory capacity. The SSA module, integrating FastSAM, utilizes point prompts to generate complete annotations in real time which reduces training complexity. Furthermore, IMOST has been successfully deployed on the quadruped robot, with performance evaluated during the online learning process. Experimental results on both public and self-collected datasets demonstrate that our IMOST outperforms current state-of-the-art method, maintains robust recognition capabilities and adaptability across various scenarios. The code is available at https://github.com/SJTUMKH/OCLTrav.
Kehui Ma, Chaoran Xiong, Qiumin Zhu, Kewei Wang 0009, Ling Pei
ICRA6
2025 Suite-IN: Aggregating Motion Features from Apple Suite for Robust Inertial Navigation
abstract
With the rapid development of wearable technology, devices like smartphones, smartwatches, and headphones equipped with IMUs have become essential for applications such as pedestrian positioning. However, traditional pedestrian dead reckoning (PDR) methods struggle with diverse motion patterns, while recent data-driven approaches, though improving accuracy, often lack robustness due to reliance on a single device. In our work, we attempt to enhance the positioning performance using the low-cost commodity IMUs embedded in the wearable devices. We propose a multi-device deep learning framework named Suite-IN, aggregating motion data from Apple Suite for inertial navigation. Motion data captured by sensors on different body parts contains both local and global motion information, making it essential to reduce the negative effects of localized movements and extract global motion representations from multiple devices. Our model innovatively introduces a contrastive learning module to disentangle motionshared and motion-private latent representations, enhancing positioning accuracy. We validate our method on a self-collected dataset consisting of Apple Suite: iPhone, Apple Watch and Airpods, which supports a variety of movement patterns and flexible device configurations. Experimental results demonstrate that our approach outperforms state-of-the-art models while maintaining robustness across diverse sensor configurations.
Songpengcheng Xia, Junyuan Deng, Zengyuan Lai, Qi Wu 0007, Ling Pei
ICRA7
2025 mmDEAR: mmWave Point Cloud Density Enhancement for Accurate Human Body Reconstruction
abstract
Millimeter-wave (mmWave) radar offers robust sensing capabilities in diverse environments, making it a highly promising solution for human body reconstruction due to its privacy-friendly and non-intrusive nature. However, the significant sparsity of mm Wave point clouds limits the estimation accuracy. To overcome this challenge, we propose a two-stage deep learning framework that enhances mm Wave point clouds and improves human body reconstruction accuracy. Our method includes a mm Wave point cloud enhancement module that densifies the raw data by leveraging temporal features and a multi-stage completion network, followed by a 2D-3D fusion module that extracts both 2D and 3D motion features to refine SMPL parameters. The mm Wave point cloud enhancement module learns the detailed shape and posture information from 2D human masks in single-view images. However, image-based supervision is involved only during the training phase, and the inference relies solely on sparse point clouds to maintain privacy. Experiments on multiple datasets demonstrate that our approach outperforms state-of-the-art methods, with the enhanced point clouds further improving performance when integrated into existing models.
Songpengcheng Xia, Zengyuan Lai, Qi Wu 0007, Wenxian Yu, Ling Pei
ICRA7
2025 A2I-Calib: An Anti-Noise Active Multi-IMU Spatial-Temporal Calibration Framework for Legged Robots
abstract
Recently, multi-node inertial measurement unit (IMU)-based odometry for legged robots has gained attention due to its cost-effectiveness, power efficiency, and high accuracy. However, the spatial and temporal misalignment between foot-end motion derived from forward kinematics and foot IMU measurements can introduce inconsistent constraints, resulting in odometry drift. Therefore, accurate spatial-temporal calibration is crucial for the multi-IMU systems. Although existing multi-IMU calibration methods have addressed passive single-rigid-body sensor calibration, they are inadequate for legged systems. This is due to the insufficient excitation from traditional gaits for calibration, and enlarged sensitivity to IMU noise during kinematic chain transformations. To address these challenges, we propose A2I-Calib, an anti-noise active multi-IMU calibration framework enabling autonomous spatial-temporal calibration for arbitrary foot-mounted IMUs. Our A2I-Calib includes: 1) an anti-noise trajectory generator leveraging a proposed basis function selection theorem to minimize the condition number in correlation analysis, thus reducing noise sensitivity, and 2) a reinforcement learning (RL)-based controller that ensures robust execution of calibration motions. Furthermore, A2I-Calib is validated on simulation and real-world quadruped robot platforms with various multi-IMU settings, which demonstrates a significant reduction in noise sensitivity and calibration errors, thereby improving the overall multi-IMU odometry performance.
Chaoran Xiong, Fangyu Jiang, Kehui Ma, Ling Pei
IROS6
2025 THE-SEAN: A Heart Rate Variation-Inspired Temporally High-Order Event-Based Visual Odometry with Self-Supervised Spiking Event Accumulation Networks
abstract
Event-based visual odometry has recently gained attention for its high accuracy and real-time performance in fast-motion systems. Unlike traditional synchronous estimators that rely on constant-frequency (zero-order) triggers, event-based visual odometry can actively accumulate information to generate temporally high-order estimation triggers. However, existing methods primarily focus on adaptive event representation after estimation triggers, neglecting the decision-making process for efficient temporal triggering itself. This oversight leads to the computational redundancy and noise accumulation. In this paper, we introduce a temporally high-order event-based visual odometry with spiking event accumulation networks (THE-SEAN). To the best of our knowledge, it is the first event-based visual odometry capable of dynamically adjusting its estimation trigger decision in response to motion and environmental changes. Inspired by biological systems that regulate hormone secretion to modulate heart rate, a self-supervised spiking neural network is designed to generate estimation triggers. This spiking network extracts temporal features to produce triggers, with rewards based on block matching points and Fisher information matrix (FIM) trace acquired from the estimator itself. Finally, THE-SEAN is evaluated across several open datasets, thereby demonstrating average improvements of 13% in estimation accuracy, 9% in smoothness, and 38% in triggering efficiency compared to the state-of-the-art methods.
Chaoran Xiong, Litao Wei, Kehui Ma, Zihan Nan, Trieu-Kien Truong, Ling Pei
IROS8
2025 360Recon: An Accurate Reconstruction Method based on Depth Fusion from 360 Images
abstract
Accurate 3D reconstruction is crucial for AR and VR applications. Compared with traditional pinhole camera-based methods, 360° image-based reconstruction can achieve higher precision with fewer input images, making it especially effective in low-texture environments. However, the severe distortion resulting from the wide field of view complicates feature extraction and matching, leading to geometric inconsistencies in multi-view reconstruction. To address these challenges, we propose 360Recon, a novel multi-view stereo (MVS) algorithm specifically designed for equirectangular projection (ERP) images. With the proposed spherical feature extraction module mitigating distortion, 360Recon integrates a 3D cost volume with multi-scale ERP features to deliver high-precision scene reconstruction while preserving local geometric consistency. Experimental results demonstrate that 360Recon outperforms existing methods in terms of accuracy, computational efficiency, and generalization capability. The source code will be released at https://github.com/LeonATP/360Recon.
Zhongmiao Yan, Qi Wu 0007, Songpengcheng Xia, Junyuan Deng, Xiang Mu, Renbiao Jin, Changchun Ye, Ling Pei
IROS8
2025 SMART: Scene-Motion-Aware Human Action Recognition Framework for Mental Disorder Group
abstract
Patients with mental disorders often exhibit risky abnormal actions, such as climbing walls or hitting windows, necessitating intelligent video behavior monitoring for smart healthcare with the rising Internet of Things (IoT) technology. However, the development of vision-based human action recognition (HAR) for these actions is hindered by the lack of specialized algorithms and datasets. In this article, we innovatively propose to build a vision-based HAR dataset, including abnormal actions often occurring in the mental disorder group and then introduce a novel scene-motion-aware action recognition technology framework, named SMART, consisting of two technical modules. First, we propose a scene perception module to extract human motion trajectory and human-scene interaction features, which introduces additional scene information for a supplementary semantic representation of the above actions. Second, the multistage fusion module fuses the skeleton motion, motion trajectory, and human-scene interaction features, enhancing the semantic association between the skeleton motion and the above supplementary representation, thus generating a comprehensive representation with both human motion and scene information. The effectiveness of our proposed method has been validated on our self-collected HAR dataset (MentalHAD), achieving 94.9% and 93.1% accuracy in un-seen subjects and scenes and outperforming state-of-the-art approaches by 6.5% and 13.2%, respectively. The demonstrated subject- and scene- generalizability makes it possible for SMART’s migration to practical deployment in smart healthcare systems for mental disorder patients in medical settings. The code and dataset will be released publicly for further research:https://github.com/Inowlzy/SMART.git.
Zengyuan Lai, Songpengcheng Xia, Qi Wu 0007, Wenxian Yu, Ling Pei
IEEE Internet Things J.7
2025 In-P3VINS: Tightly-Coupled PPP/INS/Visual SLAM Based on Invariant Optimization Approach
abstract
The state estimation on$SE_{2}(3)$Lie Group has been proven to have the ability to improve the consistency of the estimated results. They are called invariant state estimation approaches, including filter-based ones and optimization-based ones. Precise Point Positioning (PPP) is a Global Navigation Satellite System (GNSS) positioning technology which can achieve high-precision positioning without commercial base stations. Visual-Inertial Odometry (VIO) combines Visual-SLAM and IMU, realizing a more robust local pose estimation than either of the two. In this paper, the invariant optimization approach has been applied to fuse PPP/INS/Visual-SLAM. The proposed positioning system in our paper is called In-P3VINS. All raw data of the In-P3VINS is modeled and optimized under an invariant factor graph framework. In particular, the carrier phase measurement is utilized by adding the phase ambiguity into the estimated states. Finally, In-P3VINS is evaluated in both simulation experiments and real-world experiments. In the simulation experiments, the accuracy and consistency of In-P3VINS are superior to the other compared methods. In the real-world experiments, In-P3VINS has the most accurate results.
Tao Li 0052, Tong Hua, Minglei Fu, Wen-An Zhang 0001, Ling Pei, Wenxian Yu, Trieu-Kien Truong
IEEE Trans. Intell. Transp. Syst.5
2025 Suite-IN++: A FlexiWear BodyNet Integrating Global and Local Motion Features From Apple Suite for Robust Inertial Navigation
abstract
The proliferation of wearable technology has established multi-device ecosystems comprising smartphones, smartwatches, and headphones as critical enablers for ubiquitous pedestrian localization. However, traditional pedestrian dead reckoning (PDR) struggles with diverse motion modes, while data-driven methods, despite improving accuracy, often lack robustness due to their reliance on a single-device setup. Therefore, a promising solution is to fully leverage existing wearable devices to form a flexiwear bodynet for robust and accurate pedestrian localization. This paper presents Suite-IN++, a deep learning framework for flexiwear bodynet-based pedestrian localization. Suite-IN++ integrates motion data from wearable devices on different body parts, using contrastive learning to separate global and local motion features. It fuses global features based on the data reliability of each device to capture overall motion trends and employs an attention mechanism to uncover cross-device correlations in local features, extracting motion details helpful for accurate localization. To evaluate our method, we construct a real-life flexiwear bodynet dataset, incorporating Apple Suite (iPhone, Apple Watch, and AirPods) across diverse walking modes and device configurations. Experimental results demonstrate that Suite-IN++ achieves superior localization accuracy and robustness, significantly outperforming state-of-the-art models in real-life pedestrian tracking scenarios.
Songpengcheng Xia, Ling Pei
IEEE Trans. Mob. Comput.4
2024 Dynamic Inertial Poser (DynaIP): Part-Based Motion Dynamics Learning for Enhanced Human Pose Estimation with Sparse Inertial Sensors
abstract
This paper introduces a novel human pose estimation approach using sparse inertial sensors, addressing the short-comings of previous methods reliant on synthetic data. It leverages a diverse array of real inertial motion capture data from different skeleton formats to improve motion di-versity and model generalization. This method features two innovative components: a pseudo-velocity regression model for dynamic motion capture with inertial sensors, and a part-based model dividing the body and sensor data into three regions, each focusing on their unique characteristics. The approach demonstrates superior performance over state-of-the-art models across five public datasets, notably reducing pose error by 19% on the DIP-IMU dataset, thus representing a significant improvement in inertial sensor-based human pose estimation. Our codes are available at https://github.com/dx118/dynaip
Songpengcheng Xia, Qi Wu 0007, Ling Pei
CVPR6
2024 A Learning-Based Multi-Node Fusion Positioning Method Using Wearable Inertial Sensors
abstract
This study presents a novel approach to enhance the accuracy and adaptability of pedestrian positioning by fusing data from multiple Inertial Measurement Units (IMUs) attached to the human body. Leveraging the temporal and spatial richness of IMU data, our proposed multi-node sensors fusion strategy accommodates multiple sensor placements and various motion modes, a challenge that has posed difficulties for previous methods. The methodology is based on data-driven principles, directly estimating position information from IMU data. The core of our method is the strategic allocation of weights to optimize the contribution of each sensor based on reliability and relevance. Two restraint strategies, Gumbel Softmax Resampling and Empirical Risk Minimization under Fairness Constraints, are introduced to enhance fusion accuracy and robustness. Experimental results show improvements in positioning accuracy and robust performance in various situations compared to existing methods. Our approach is expected to be applied in domains such as indoor navigation and motion analysis, highlighting the key role of wearable sensors in advancing positioning accuracy and addressing complex scenarios.
Songpengcheng Xia, Ling Pei
ICASSP4
2024 mmBaT: A Multi-Task Framework for Mmwave-Based Human Body Reconstruction and Translation Prediction
abstract
Human body reconstruction with Millimeter Wave (mmWave) radar point clouds has gained significant interest due to its ability to work in adverse environments and its capacity to mitigate privacy concerns associated with traditional camera-based solutions. Despite pioneering efforts in this field, two challenges persist. Firstly, raw point clouds contain massive noise points, usually caused by the ambient objects and multi-path effects of Radio Frequency (RF) signals. Recent approaches typically rely on prior knowledge or elaborate preprocessing methods, limiting their applicability. Secondly, even after noise removal, the sparse and inconsistent body-related points pose an obstacle to accurate human body reconstruction. To address these challenges, we introduce mmBaT, a novel multi-task deep learning framework that concurrently estimates the human body and predicts body translations in subsequent frames to extract body-related point clouds. Our method is evaluated on two public datasets that are collected with different radar devices and noise levels. A comprehensive comparison against other state-of-the-art methods demonstrates that our method has a superior reconstruction performance and generalization ability from noisy raw data, even when compared to methods provided with body-related point clouds.
Songpengcheng Xia, Qi Wu 0007, Ling Pei
ICASSP5
2024 Explicit Interaction for Fusion-Based Place Recognition
abstract
Fusion-based place recognition is an emerging technique jointly utilizing multi-modal perception data, to recognize previously visited places in GPS-denied scenarios for robots and autonomous vehicles. Recent fusion-based place recognition methods combine multi-modal features in implicit manners. While achieving remarkable results, they do not explicitly consider what the individual modality affords in the fusion system. Therefore, the benefit of multi-modal feature fusion may not be fully explored. In this paper, we propose a novel fusion-based network, dubbed EINet, to achieve explicit interaction of the two modalities. EINet uses LiDAR ranges to supervise more robust vision features for long time spans, and simultaneously uses camera RGB data to improve the discrimination of LiDAR point clouds. In addition, we develop a new benchmark for the place recognition task based on the nuScenes dataset. To establish this benchmark for future research with comprehensive comparisons, we introduce both supervised and self-supervised training schemes alongside evaluation protocols. We conduct extensive experiments on the proposed benchmark, and the experimental results show that our EINet exhibits better recognition performance as well as solid generalization ability compared to the state-of-the-art fusion-based place recognition approaches. Our open-source code and benchmark are released at: https://github.com/BIT-XJY/EINet.
Junyi Ma, Qi Wu 0007, Yue Wang 0020, Xieyuanli Chen, Wenxian Yu, Ling Pei
IROS8
2024 Thermal-NeRF: Neural Radiance Fields from an Infrared Camera
abstract
In recent years, Neural Radiance Fields (NeRFs) have demonstrated significant potential in encoding highly-detailed 3D geometry and environmental appearance, positioning themselves as a promising alternative to traditional explicit representation for 3D scene reconstruction. However, the predominant reliance on RGB imaging presupposes ideal lighting conditions—a premise frequently unmet in robotic applications plagued by poor lighting or visual obstructions. This limitation overlooks the capabilities of infrared (IR) cameras, which excel in low-light detection and present a robust alternative under such adverse scenarios. To tackle these issues, we introduce Thermal-NeRF, the first method that estimates a volumetric scene representation in the form of a NeRF solely from IR imaging. By leveraging a thermal mapping and structural thermal constraint derived from the thermal characteristics of IR imaging, our method showcases unparalleled proficiency in recovering NeRFs in visually degraded scenes where RGB-based methods fall short. We conduct extensive experiments to demonstrate that Thermal-NeRF can achieve superior quality compared to existing methods. Furthermore, we contribute a dataset for IR-based NeRF applications, paving the way for future research in IR NeRF reconstruction, see https://github.com/Cerf-Volant425/Thermal-NeRF.
Tianxiang Ye, Qi Wu 0007, Junyuan Deng, Liu Liu 0012, Songpengcheng Xia, Wenxian Yu, Ling Pei
IROS9
2024 Unsupervised Domain Adaptation Depth Estimation Based on Self-attention Mechanism and Edge Consistency Constraints
abstract
Abstract In the unsupervised domain adaptation (UDA) (Akada et al. Self-supervised learning of domain invariant features for depth estimation, in: 2022 IEEE/CVF winter conference on applications of computer vision (WACV), pp 3377–3387 (2022). 10.1109/WACV51458.2022.00107 ) depth estimation task, a new adaptive approach is to use the bidirectional transformation network to transfer the style between the target and source domain inputs, and then train the depth estimation network in their respective domains. However, the domain adaptation process and the style transfer may result in defects and biases, often leading to depth holes and instance edge depth missing in the target domain’s depth output. To address these issues, We propose a training network that has been improved in terms of model structure and supervision constraints. First, we introduce a edge-guided self-attention mechanism in the task network of each domain to enhance the network’s attention to high-frequency edge features, maintain clear boundaries and fill in missing areas of depth. Furthermore, we utilize an edge detection algorithm to extract edge features from the input of the target domain. Then we establish edge consistency constraints between inter-domain entities in order to narrow the gap between domains and make domain-to-domain transfers easier. Our experimental demonstrate that our proposed method effectively solve the aforementioned problem, resulting in a higher quality depth map and outperforming existing state-of-the-art methods.
Shuguo Pan, Ling Pei, Baoguo Yu
Neural Process. Lett.4
2024 TextSLAM: Visual SLAM With Semantic Planar Text Features
abstract
We propose a novel visual SLAM method that integrates text objects tightly by treating them as semantic features via fully exploring their geometric and semantic prior. The text object is modeled as a texture-rich planar patch whose semantic meaning is extracted and updated on the fly for better data association. With the full exploration of locally planar characteristics and semantic meaning of text objects, the SLAM system becomes more accurate and robust even under challenging conditions such as image blurring, large viewpoint changes, and significant illumination variations (day and night). We tested our method in various scenes with the ground truth data. The results show that integrating texture features leads to a more superior SLAM system that can match images across day and night. The reconstructed semantic 3D text map could be useful for navigation and scene understanding in robotic and mixed reality applications.
Boying Li, Danping Zou, Yuan Huang 0011, Xinghan Niu, Ling Pei, Wenxian Yu
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Timestamp-Supervised Wearable-Based Activity Segmentation and Recognition With Contrastive Learning and Order-Preserving Optimal Transport
abstract
Human activity recognition (HAR) with wearables is one of the serviceable technologies in ubiquitous and mobile computing applications. The sliding-window scheme is widely adopted while suffering from the multi-class windows problem. As a result, there is a growing focus on joint segmentation and recognition with deep-learning methods, aiming at simultaneously dealing with HAR and time-series segmentation issues. However, obtaining the full activity annotations of wearable data sequences is resource-intensive or time-consuming, while unsupervised methods yield poor performance. To address these challenges, we propose a novel method for joint activity segmentation and recognition with timestamp supervision, in which only a single annotated sample is needed in each activity segment. However, the limited information of sparse annotations exacerbates the gap between recognition and segmentation tasks, leading to sub-optimal model performance. Therefore, the prototypes are estimated by class-activation maps to form a sample-to-prototype contrast module for well-structured embeddings. Moreover, with the optimal transport theory, our approach generates the sample-level pseudo-labels that take advantage of unlabeled data between timestamp annotations for further performance improvement. Comprehensive experiments on four public HAR datasets demonstrate that our model trained with timestamp supervision is superior to the state-of-the-art weakly-supervised methods and achieves comparable performance to the fully-supervised approaches.
Songpengcheng Xia, Ling Pei, Wenxian Yu, Robert C. Qiu
IEEE Trans. Mob. Comput.3
2023 NeRF-LOAM: Neural Implicit Representation for Large-Scale Incremental LiDAR Odometry and Mapping
abstract
Simultaneously odometry and mapping using LiDAR data is an important task for mobile systems to achieve full autonomy in large-scale environments. However, most existing LiDAR-based methods prioritize tracking quality over reconstruction quality. Although the recently developed neural radiance fields (NeRF) have shown promising advances in implicit reconstruction for indoor environments, the problem of simultaneous odometry and mapping for large-scale scenarios using incremental LiDAR data remains unexplored. To bridge this gap, in this paper, we propose a novel NeRF-based LiDAR odometry and mapping approach, NeRF-LOAM, consisting of three modules neural odometry, neural mapping, and mesh reconstruction. All these modules utilize our proposed neural signed distance function, which separates LiDAR points into ground and non-ground points to reduce Z-axis drift, optimizes odometry and voxel embeddings concurrently, and in the end generates dense smooth mesh maps of the environment. Moreover, this joint optimization allows our NeRF-LOAM to be pre-trained free and exhibit strong generalization abilities when applied to different environments. Extensive evaluations on three publicly available datasets demonstrate that our approach achieves state-of-the-art odometry and mapping performance, as well as a strong generalization in large-scale environments utilizing LiDAR data. Furthermore, we perform multiple ablation studies to validate the effectiveness of our network design. The implementation of our approach will be made available at https://github.com/JunyuanDeng/NeRF-LOAM.
Junyuan Deng, Qi Wu 0007, Xieyuanli Chen, Songpengcheng Xia, Wenxian Yu, Ling Pei
ICCV8
2023 PIEKF-VIWO: Visual-Inertial-Wheel Odometry using Partial Invariant Extended Kalman Filter
abstract
Invariant Extended Kalman Filter (IEKF) has been successfully applied in Visual-inertial Odometry (VIO) as an advanced achievement of Kalman filter, showing great potential in sensor fusion. In this paper, we propose partial IEKF (PIEKF), which only incorporates rotation-velocity state into the Lie group structure and apply it for Visual-Inertial-Wheel Odometry (VIWO) to improve positioning accuracy and consistency. Specifically, we derive the rotation-velocity measurement model, which combines wheel measurements with kinematic constraints. The model circumvents the wheel odometer's 3D integration and covariance propagation, which is essential for filter consistency. And a plane constraint is also introduced to enhance the position accuracy. A dynamic outlier detection method is adopted, leveraging the velocity state output. Through the simulation and real-world test, we validate the effectiveness of our approach, which outperforms the standard Multi-State Constraint Kalman Filter (MSCKF) based VIWO in consistency and accuracy.
Tong Hua, Tao Li 0052, Ling Pei
ICRA3
2023 A Boundary Consistency-Aware Multitask Learning Framework for Joint Activity Segmentation and Recognition With Wearable Sensors
abstract
With the development of industrial and sensing technology, sensor-based activity recognition has become a promising technology for informatics applications. However, in a typical activity recognition procedure, sensory data segmentation, usually considered a preprocess with sliding windows, rarely has been investigated and significantly affected the recognition performance. In this article, we propose a novel deep-learning method to jointly segment and recognize activities with wearable sensors. Our contributions are three-fold: First, we introduce a multistage temporal convolutional network for sample-level activity prediction to overcome the multiclass windows problem. Second, for alleviating oversegmentation errors, our model forms a multitask learning framework with a boundary prediction module to adjust the entire model’s gradients. Third, we innovatively propose a boundary consistency loss to enforce the consistency of the activity and boundary prediction. Our method shows impressive performance on three public datasets, especially achieving 16% improvement over very recently advanced competing methods with class-average F1-score on the Hospital dataset. The code of this work will be open source onhttps://github.com/xspc/Segmentation-Sensor-based-HAR.
Songpengcheng Xia, Ling Pei, Wenxian Yu, Robert C. Qiu
IEEE Trans. Ind. Informatics3
2022 Multi-level Contrast Network for Wearables-based Joint Activity Segmentation and Recognition
abstract
Human activity recognition (HAR) with wearables is promising research that can be widely adopted in many smart healthcare applications. In recent years, the deep learning-based HAR models have achieved impressive recognition performance. However, most HAR algorithms are susceptible to the multi-class windows problem that is essential yet rarely exploited. In this paper, we propose to relieve this challenging problem by introducing the segmentation technology into HAR, yielding joint activity segmentation and recognition. Especially, we introduce the Multi-Stage Temporal Convolutional Network (MS-TCN) architecture for sample-level activity prediction to joint segment and recognize the activity sequence. Furthermore, to enhance the robustness of HAR against the inter-class similarity and intra-class heterogeneity, a multi-level contrastive loss, containing the sample-level and segment-level contrast, has been proposed to learn a well-structured embedding space for better activity segmentation and recognition performance. Finally, with comprehensive experiments, we verify the effectiveness of the proposed method on two public HAR datasets, achieving significant improvements in the various evaluation metrics.
Songpengcheng Xia, Ling Pei, Wenxian Yu, Robert C. Qiu
GLOBECOM3
2022 MWNET: A Tracking Method for Frequently Occluded Scenes Based on Matter Waves
abstract
One of the main issues of current detection-based tracking methods is object annihilation, which is the object lost in one or more frames caused by errors of the object detector, especially in adverse weather or occluded scenes. In this work, the matter waves network (MWNET) based on quantum theory is proposed to improve tracking robustness. The object detection results are encoded into the matter waves by quantum measurement to utilize the features of the complex value. The object is regenerated from annihilation by quantum evolution to repair the detection errors. The network can improve the information flow between front and back frames and the continuity of tracking trajectory. Finally, on the MOT20 test set, 77.3 MOTA and 76.2 IDF1 are implemented by MWNET. Our work proves the effectiveness of quantum evolution in reconstructing information, and the model can realize continuous tracking of multiple objects in frequent occlusion scenes.
Ling Pei, Zheng-De Zhang
ICIP2
2022 CDNet: a real-time and robust crosswalk detection network on Jetson nano based on YOLOv5
Zheng-De Zhang, Meng-Lu Tan, Zhi-Cai Lan, Hai-Chun Liu, Ling Pei, Wen-Xian Yu
Neural Comput. Appl.5
2021 Open Set Mixed-Reality Human Activity Recognition
abstract
Sensor-based human activity recognition (HAR) is a fundamental problem that can have a broad impact on many research/industrial fields. The deep learning methods pave the way for extracting robust and informative features from the non-stationary HAR data, achieving high-accuracy HAR. Most of the works in the literature consider the closed set activity recognition, which assumes all classes of activities are known in both the training and test stages. However, it is a challenging way to apply deep learning models to real-world applications, which can contain unseen activities. These activities will sharply decrease the performance of the deep learning models. To this end, we introduce the problem of Open Set Mixed-Reality (OSM) HAR, which aims to recognize unseen activities while classify seen samples. Furthermore, we propose a novel balanced open set backpropagation method to realize accurate and robust OSM-HAR. Lastly, we verify the effectiveness of the proposed method with a publicly available and our newly collected dataset.
Songpengcheng Xia, Ling Pei
GLOBECOM4
2021 Grayscale And Normal Guided Depth Completion With A Low-Cost Lidar
abstract
In this paper, we introduce DenseLivox, a dataset with dense and accurate depth as ground truth. To our best knowledge, it is the first dataset with dense ground truth designed for LiDAR depth completion using a low-cost LiDAR. Also, we develop a simple yet effective multi-task learning network to tackle the problem of depth completion. Compared to the works in the literature, our model’s uniqueness is that it completes a depth map, a normal map, and a grayscale image simultaneously. To address the area with heavy noises, we use modified Huber loss to smooth these outliers’ effect. We evaluate our method on DenseLivox and show that accuracy is greatly improved with the grayscale and normal guidance. Our method outperforms other depth-only methods and is comparable to the methods that take RGB and depth as input.
Qingyang Yu, Qi Wu 0007, Ling Pei
ICIP4
2021 Ahed: A Heterogeneous-Domain Deep Learning Model for IoT-Enabled Smart Health With Few-Labeled EEG Data
abstract
Recent years have witnessed the successful development of health-related Internet-of-Things (IoT) devices, i.e., electroencephalographic (EEG), paving a ground-breaking way for better understanding the functionals in our brains. Despite massive research on EEG, there lacks an effective way to interpret complex EEG signals due to the shortage of informative EEG data, the challenge in capturing sophisticated connectivity patterns of EEG signals, and unfavorable results of the inherent noise associated with data collection. In this article, novel heterogeneous-domain deep learning is proposed to address these issues. Especially, we first propose a new scheme to extract the multilevel latent features using hybrid networks and provide two pathways for reconstructing the EEG signals. As a core part of the proposed method, the scheme considers the complex dependencies among adjacent EEG channels, spatiotemporal connectivity, and signal denoising. In addition, a novel consistency regularization method is proposed to enhance information sharing among the multilevel latent feature obtained from the labeled and unlabeled EEG samples, which is beneficial for both the information transfer and accelerating the training. Finally, we provide comprehensive case studies on the Lomonosov Moscow State University EEG data set, demonstrating that the proposed methods achieve superior performance than all the competing ones over a wide range of experimental settings.
Ling Pei, Robert C. Qiu
IEEE Internet Things J.2
2021 Toward Location-Enabled IoT (LE-IoT): IoT Positioning Techniques, Error Sources, and Error Mitigation
abstract
Localization techniques are becoming key to add location context to the Internet-of-Things (IoT) data without human perception and intervention. Meanwhile, the newly emerged low-power wide-area network (LPWAN) and 5G technologies have become strong candidates for mass-market localization applications. However, various error sources have limited localization performance by using such IoT signals. This article reviews the IoT localization system through the following sequence: IoT localization system review, localization data sources, localization algorithms, localization error sources and mitigation, and localization performance evaluation. Compared to the related surveys, this article has a more comprehensive and state-of-the-art review on IoT localization methods, an original review on IoT localization error sources and mitigation, an original review on IoT localization performance evaluation, and a more comprehensive review of IoT localization applications, opportunities, and challenges. Thus, this survey provides comprehensive guidance for peers who are interested in enabling localization ability in the existing IoT systems, using IoT systems for localization, or integrating IoT signals with the existing localization sensors.
You Li 0001, Yuan Zhuang 0001, Xin Hu 0006, Zhouzheng Gao, Jia Hu 0001, Long Chen 0005, Zhe He 0002, Ling Pei, Kejie Chen, Maosong Wang, Xiaoji Niu, Ruizhi Chen, John S. Thompson, Fadhel M. Ghannouchi, Naser El-Sheimy
IEEE Internet Things J.8
2021 MARS: Mixed Virtual and Real Wearable Sensors for Human Activity Recognition With Multidomain Deep Learning Model
abstract
Together with the rapid development of the Internet of Things, human activity recognition (HAR) using wearable inertial measurement units (IMUs) becomes a promising technology for many research areas. Recently, deep-learning-based methods pave a new way of understanding and performing analysis of the complex data in the HAR system. However, the performance of these methods is mostly based on the quality and quantity of the collected data. In this article, we innovatively propose to build a large data set based on virtual IMUs and then address technical issues by introducing a multiple-domain deep learning framework consisting of three technical parts. In the first part, we propose to learn the single-frame human activity from the noisy IMU data with hybrid convolutional neural networks in the semisupervised form. For the second part, the extracted data features are fused according to the principle of uncertainty-aware consistency, which reduces the uncertainty by weighting the importance of the features. The transfer learning is performed in the last part based on the newly released archive of motion capture as surface shapes data set, containing abundant synthetic human poses, which enhances the variety and diversity of the training data set and is beneficial for the process of training and feature transfer in the proposed method. The efficiency and effectiveness of the proposed method have been demonstrated in the real deep inertial poser data set. The experimental results show that the proposed methods can surprisingly converge within a few iterations and outperform all competing methods.
Ling Pei, Songpengcheng Xia, Fanyi Xiao, Qi Wu 0007, Wenxian Yu, Robert C. Qiu
IEEE Internet Things J.1
2021 A Novel Wireless Localization Approach Using Twice Receiving Array Spectra Fusions and ASSR Networks
abstract
Path fading and non-line-of-sight (NLOS) signals constitute serious problems in the wireless localization process. These problems cause unpredictable degradation in the localization precision. In this paper, a novel wireless localization approach, which is based on twice receiving array signal spectra fusions and asymmetric second-order stochastic resonance (ASSR) networks, is proposed. By combining and repetitively processing the above two techniques, the receiving signal-to-noise ratio (SNR) can be enhanced. Additionally, the receiving array signal without the line-of-sight (LOS) component can be determined and removed from the spectra fusion process. The theoretical analyses presented verify the unbiasedness and asymptotic efficiency of the proposed twice receiving spectra fusion approach. Computer simulations demonstrate that the fused spectra can significantly improve the wireless localization precision compared with conventional and up-to-date localization methods, especially under low SNR conditions.
Di He 0002, Ling Pei, Xin Chen 0017, Ling-ge Jiang, Jiaqing Qu, Wenxian Yu
IEEE Trans. Commun.2
2020 TextSLAM: Visual SLAM with Planar Text Features
abstract
We propose to integrate text objects in man-made scenes tightly into the visual SLAM pipeline. The key idea of our novel text-based visual SLAM is to treat each detected text as a planar feature which is rich of textures and semantic meanings. The text feature is compactly represented by three parameters and integrated into visual SLAM by adopting the illumination-invariant photometric error. We also describe important details involved in implementing a full pipeline of text-based visual SLAM. To our best knowledge, this is the first visual SLAM method tightly coupled with the text features. We tested our method in both indoor and outdoor environments. The results show that with text features, the visual SLAM system becomes more robust and produces much more accurate 3D text maps that could be useful for navigation and scene understanding in robotic or augmented reality applications.
Boying Li, Danping Zou, Daniele Sartori, Ling Pei, Wenxian Yu
ICRA4
2019 A Novel Wireless Positioning Approach Based on Distributed Stochastic-Resonance-Enhanced Power Spectrum Fusion Technique
abstract
In the wireless positioning, the performances of most traditional methods are influenced by the low signal-to-noise (SNR) and none-line-of-sight (NLOS) problems seriously. So in this study, a kind of nonlinear stochastic-resonance (SR) signal enhancement technique combined with the distributed receiving array power spectrum fusion approach is proposed. By utilizing the signal power improvement property, the fitness of distributed SR system with the array structure, and the information fusion of the spectrum, the corresponding wireless positioning can be realized satisfactorily. Computer simulations also show the advantages over conventional methods in the direction of arrival (DoA) estimation and positioning estimation.
Di He 0002, Xin Chen 0017, Danping Zou, Ling Pei, Ling-ge Jiang
ISCAS4
2019 A Method of 2D Semantic Map Generation for Autonomous Flight of MAVs
abstract
In many applications, MAVs(Micro Aerial Vehicle) need to recognize the ground targets and locate them precisely, or in other words, to generate a semantic map of those targets, to guide the MAVs to complete mission automatically. In this paper, we introduce a method of fast 2D semantic map generation using only the images captured by the downward-looking camera in an unknown environment. The map is firstly stitched from each individual image by feature matching. Then the contour and area of each target are extracted through color detection in the global map. Finally, neural network is utilized for recognition of ground targets marked by different printed numbers. We have tested our method for automatic task performing as an MAV competition challenge in a virtual environment. By taking the generated 2D semantic map into the control loop, the MAV can localize itself, realize autonomous flight, detect and explore the environment.
Ruochen Yao, Danping Zou, Daniele Sartori, Ling Pei, Ling Gong
ISCAS4
2019 StructVIO: Visual-Inertial Odometry With Structural Regularity of Man-Made Environments
abstract
In this paper, we propose a novel visual-inertial odometry (VIO) approach that adopts structural regularity in man-made environments. Instead of using Manhattan world assumption, we use Atlanta world model to describe such regularity. An Atlanta world is a world that contains multiple local Manhattan worlds with different heading directions. Each local Manhattan world is detected on the fly, and their headings are gradually refined by the state estimator when new observations are received. With full exploration of structural lines that aligned with each local Manhattan worlds, our VIO method becomes more accurate and robust, as well as more flexible to different kinds of complex man-made environments. Through benchmark tests and real-world tests, the results show that the proposed approach outperforms existing visual-inertial systems in large-scale man-made environments.
Danping Zou, Yuanxin Wu, Ling Pei, Haibin Ling, Wenxian Yu
IEEE Trans. Robotics3
2018 A Revisited Approach to Lateral Acceleration Modeling for Quadrotor UAVs State Estimation
abstract
Quadrotor state estimation generally relies on the vehicle aerodynamics modeling to achieve improved performance. In this paper the effects of the rotors angular speeds on the quadrotor drag, and therefore on the lateral accelerations, are investigated. While these effects are usually disregarded, we analyze their modeling starting from the Blade Element Theory and flight test data. Two lateral acceleration formulations are proposed. They are adopted within a velocity and attitude state estimator and validated in real-world flights. The EKF-based estimator fuses measurements from low-cost sensors present in the majority of quadrotors (IMU, magnetometer, ultrasonic sensor, optical flow) with the accelerations of the vehicle predicted from the revisited models. Experimental results show the benefits of adopting these innovative models in the estimator when compared with the existing modeling approach.
Daniele Sartori, Danping Wou, Ling Pei, Wenxian Yu
IROS3
2018 An Improved Kernel Clustering Algorithm Used in Computer Network Intrusion Detection
abstract
In the computer network intrusion detection system, data objects, mapped from original space, are analyzed based on kernel clustering algorithm. During the process of kernel clustering, some representative points are introduced to represent a cluster. In just one iteration, the distance between a data object and representative points of a cluster is computed to partition the data objects. The clustering results contain normal data and abnormal data, which achieve the goal of intrusion detection. At the same time, KDD CUP 1999 dataset is used to make simulations. The results show that the proposed algorithm has higher detecting probability under the condition of low constant false alarm rate compared with K-Means clustering algorithm and SVM.
Di He 0002, Xin Chen 0017, Danping Zou, Ling Pei, Ling-ge Jiang
ISCAS4
2018 Uplink Power Control Approach Based on Adaptive Chaotic Simulated Annealing
abstract
A kind of uplink power control approach of asynchronous code division multiple access (CDMA) mobile communication system based on the adaptive chaotic simulated annealing (CSA) is proposed. The particular influence of near-far effect can be solved by using the global optimal searching function of CSA method, and the multiple access interference (MAI) problem can be overcome by introducing adaptive MAI cancellation network in corresponding close loop control strategy. Computer simulations show the comparison results between proposed approach and conventional method with heterogeneous signal-to-interference ratio (SIR) threshold decision. It shows that this new approach can achieve higher channel capability and lower average outrage probability with some small increase of iteration steps, which can improve the performance of CDMA mobile communication system effectively.
Di He 0002, Xin Chen 0017, Danping Zou, Ling Pei, Ling-ge Jiang
ISCAS4
2017 An aerodynamic model-aided state estimator for multi-rotor UAVs
abstract
A robust state estimator is presented by fusing the aerodynamic model of multi-rotor UAVs with measurements from optical flow and other low-cost sensors such as IMU, magnetometer, and ultrasonic sensor. Due to the particular aerodynamics of multi-rotor UAVs, the body velocity in the rotor plane is able to be measured by the accelerometer. We therefore propose a novel state estimator by fully exploring the characteristic of aerodynamics of multi-rotor UAV. Our state estimator is fast and easy to be implemented. We have tested our estimator with different platforms in different scenes. Experimental results show that our estimator performs robustly in low light conditions where existing methods usually fail.
Rongzhi Wang, Danping Zou, Ling Pei, Wenxian Yu
IROS4
2016 Multidimensional Scaling-Based TDOA Localization Scheme Using an Auxiliary Line
abstract
This work deals with source localization with time-difference-of-arrival (TDOA) measurements in two-dimensional (2-D) scenarios. Although the celebrated two-step weighted least squares (2WLS) method is quite successful, its drawback lies in an ill-conditioning problem when the sensor array is quasi-linear. This work presents a multidimensional scaling (MDS)-based localization scheme. Based on the subspace analysis of the scalar product matrix, an auxiliary line is defined in the plane, close to the global minimizer of the cost function. Then, the minimizer on the auxiliary line is found as the estimation of the source position. Simulations show that the proposed scheme achieves high localization accuracy for all kinds of sensor arrays including quasi-linear arrays.
Wuyang Jiang, Ling Pei, Wenxian Yu
IEEE Signal Process. Lett.3
2013 Sound positioning using a small-scale linear microphone array
abstract
Microphone arrays, also known as acoustic antennas, have been extensively used for sound localization. Small-scale microphone arrays have especially been used in teleconferences and game consoles due to their small dimension and easy deployment. In this article, we present an approach to locating a sound source using a small linear microphone array. We describe the fundamentals of linear microphone arrays and analyze the impact of geometry in terms of positioning accuracy using the dilution of precision (DOP) concept. The generalized cross-correlation (GCC) based on the phase transform (PHAT) weighting function is used to estimate the time difference of arrivals in a microphone array. Given the time differences, we use both closed-form and iterative optimization solutions to calculate the coordinates of the sound source. In order to evaluate the performances of the solutions applied in this paper, simulations and field tests were conducted. Simulation results show that the closed-form algorithm gives a positioning error of less than 5 cm in a 10-by-10 meter room when the geometry of a microphone array is good and the signal to noise ratio (SNR) is high. Linear small microphone arrays have lower performances compared to a non-linear distributed array. When the scale of a linear array is reduced, the positioning accuracy decreases dramatically. With a small linear array, the iterative optimization algorithm gives much better performance compared to the closed-form algorithm. Field tests were conducted in an 11-by-5.6 meter room using a linear array with a length of 0.23 meters. Positioning results show an average error of 0.25 meters along the axis parallel to the linear array and 0.53 meters error along the axis which is perpendicular to the linear array.
Ling Pei, Liang Chen 0007, Robert Guinness, Jingbin Liu, Heidi Kuusniemi, Yuwei Chen 0005, Ruizhi Chen, Stefan Söderholm
IPIN1
2013 An improved indoor localization method using smartphone inertial sensors
abstract
In this paper, an improved indoor localization method based on smartphone inertial sensors is presented. Pedestrian dead reckoning (PDR), which determines the relative location change of a pedestrian without additional infrastructure supports, is combined with a floor plan for a pedestrian positioning in our work. To address the challenges of low sampling frequency and limited processing power in smartphones, reliable and efficient PDR algorithms have been proposed. A robust step detection technique leaves out the preprocessing of raw signal and reduces complex computation. Given the fact that the precision of the stride length estimation is influenced by different pedestrians and motion modes, an adaptive stride length estimation algorithm based on the motion mode classification is developed. Heading estimation is carried out by applying the principal component analysis (PCA) to acceleration measurements projected to the global horizontal plane, which is independent of the orientation of a smartphone. In addition, to eliminate the sensor drift due to the inaccurate distance and direction estimations, a particle filter is introduced to correct the drift and guarantee the localization accuracy. Extensive field tests have been conducted in a laboratory building to verify the performance of proposed algorithm. A pedestrian held a smartphone with arbitrary orientation in the tests. Test results show that the proposed algorithm can achieve significant performance improvements in terms of efficiency, accuracy and reliability.
Jiuchao Qian, Jiabin Ma, Rendong Ying, Ling Pei
IPIN5