Ruizhi Chen

dblp:120/4143 · DBLP profile ↗
← Back
81ranked-venue papers
5as first author
66since 2021 · last 2026
0000-0001-6683-2342ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 32 · 29 since 2021Artificial intelligence and machine learning · 29 · 3 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 6 since 2021Systems, architecture and hardware · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel Generation
abstract
Developing high-performance GPU kernels is critical for AI and scientific computing, but remains challenging due to its reliance on expert crafting and poor portability. While large language models (LLMs) offer promise for automation, both general-purpose and finetuned LLMs suffer from two fundamental and conflicting limitations: correctness and efficiency. The key reason is that existing LLM-based approaches directly generate the entire optimized low-level programs, requiring exploration of an extremely vast space encompassing both optimization policies and implementation codes. To address the challenge of exploring an intractable space, we propose Macro Thinking Micro Coding (MTMC), a hierarchical framework inspired by the staged optimization strategy of human experts. It decouples optimization strategy from implementation details, ensuring efficiency through high-level strategy and correctness through low-level implementation. Specifically, Macro Thinking employs reinforcement learning to guide lightweight LLMs in efficiently exploring and learning semantic optimization strategies that maximize hardware utilization. Micro Coding leverages general-purpose LLMs to incrementally implement the stepwise optimization proposals from Macro Thinking, avoiding full-kernel generation errors. Together, they effectively navigate the vast optimization space and intricate implementation details, enabling LLMs for high-performance GPU kernel generation. Comprehensive results on widely adopted benchmarks demonstrate the superior performance of MTMC on GPU kernel generation in both accuracy and running time. On KernelBench, MTMC achieves near 100% and 70% accuracy at Levels 1-2 and 3, over 50% than SOTA general-purpose and domain-finetuned LLMs, with up to 7.3× speedup over LLMs, and 2.2× over expert-optimized PyTorch Eager kernels. On the more challenging TritonBench, MTMC attains up to 59.64% accuracy and 34× speedup. All models and datasets will be made publicly available.
Xinguo Zhu, Shaohui Peng, Jiaming Guo, Yunji Chen, Qi Guo 0001, Yuanbo Wen 0001, Hang Qin, Ruizhi Chen, Qirui Zhou, Ke Gao 0012, Ling Li 0001
AAAI8
2026 Deployable uplink SRS-based 5G NR positioning in mixed LoS/NLoS Environments
Wenxin Dong, Zhanghai Ju, Hongjian Jiao, Ruizhi Chen, Liang Chen 0007
Expert Syst. Appl.4
2026 Learning-based analysis of 5G and WiFi CSI for indoor localization: Feature stability, model generalization, and performance trade-offs
Yanlin Ruan, Xin Zhou 0006, Zhaoliang Liu, Ruizhi Chen, Liang Chen 0007
Neurocomputing5
2026 Pedestrian and Router Colocalization Framework Using Distributed-IMU-Based VDR and Wi-Fi RTT
abstract
Inertial navigation and WiFi are two common approaches for pedestrian localization. However, conventional pedestrian dead reckoning (PDR) and WiFi fingerprinting suffer from limited adaptability to different users and poor robustness to environmental changes, respectively. Recent deep-learning-based methods address pedestrian localization by modeling sequential dependencies in inertial data, but they typically rely on a single inertial measurement unit (IMU), which is insufficient to capture the spatial correlations of human skeletal motion. In parallel, the Fine Time Measurement (FTM) procedure in IEEE 802.11mc enables round-trip time (RTT)–based ranging and localization, yet the coordinates of WiFi routers still require labor-intensive prior surveying, limiting deployment flexibility. This paper presents a pedestrian and router colocalization framework that jointly estimates pedestrian trajectories and WiFi router positions. The proposed framework employs multiple body-worn IMUs and a long short-term memory (LSTM) network to learn both spatial and temporal dependencies in human motion, thereby enabling velocity dead reckoning (VDR). The VDR-estimated pedestrian velocity is then fused with WiFi RTT measurements through factor graph optimization (FGO), in which both pedestrian and router coordinates are treated as unknown variables. Experimental results demonstrate that the multi-IMU-based VDR effectively models pedestrian velocity, while WiFi RTT ranging constrains the long-term drift of VDR. The combined VDR/WiFi RTT framework achieves meter-level positioning accuracy in both indoor and outdoor environments, without requiring pre-surveyed router coordinates, and thus provides a promising solution for pedestrian localization in the Internet of Things (IoT) era.
Mingxi Wang, Fuqiang Gu, Liang Chen 0007, Ruizhi Chen, Shikai Jin
IEEE Internet Things J.6
2026 AFLoc: A Deep Learning Indoor Localization Method With Trainable Activation Functions
abstract
With the advancement of the Internet of Things (IoT) and the rise of smart cities, achieving efficient and accurate indoor localization has become essential for indoor location-based services. However, current deep learning-based indoor localization methods often rely on fixed activation functions during RSSI feature extraction and learning, which limits their ability to capture complex geospatial information embedded in RSSI signals. This results in reduced localization accuracy and weak robustness. To address this problem, we propose AFLoc, a deep learning indoor localization method with trainable activation functions. Specifically, we first design AFLayer, a trainable activation function module grounded in the universal approximation theorem of single-layer neural networks. Then, we integrate convolutional neural network with Transformer to extract both global and local features from RSSI Sequence data. Finally, by incorporating AFLayer into the CNN-Transformer feature extraction framework, each layer is equipped with a trainable activation function, enhancing the network’s ability to learn discriminative RSSI features and capture richer geospatial information, ultimately improving localization accuracy. The experimental results convincingly demonstrate that AFLoc achieves superior performance compared to existing localization methods across both the SODIndoorLoc and WIFIne public datasets.
Shuai Zhang 0016, Ruizhi Chen, Guobing Pan, Huarong Li
IEEE Internet Things J.3
2026 Attentional Graph Meta-Learning for Indoor Localization Using Extremely Sparse Fingerprints
abstract
Fingerprint-based indoor localization is often labor-intensive due to the need for dense grids and repeated measurements across time and space. Maintaining high localization accuracy with extremely sparse fingerprints remains a persistent challenge. Existing benchmark methods primarily rely on the measured fingerprints, while neglecting valuable spatial and environmental characteristics. To address this issue, we propose a systematic integration of an Attentional Graph Neural Network (AGNN) model, capable of learning spatial adjacency relationships and aggregating information from neighboring fingerprints, and a meta-learning framework that utilizes datasets with similar environmental characteristics to enhance model training. To minimize the labor required for fingerprint collection, we introduce two novel data augmentation strategies: 1) unlabeled fingerprint augmentation using moving platforms, which enables the semi-supervised AGNN model to incorporate information from unlabeled fingerprints, and 2) synthetic labeled fingerprint augmentation through environmental digital twins, which enhances the meta-learning framework through a practical distribution alignment, which can minimize the feature discrepancy between synthetic and real-world fingerprints effectively. By integrating these novel modules, we propose the Attentional Graph Meta-Learning (AGML) model. This novel model combines the strengths of the AGNN model and the meta-learning framework to address the challenges posed by extremely sparse fingerprints. To validate our approach, we collected multiple datasets from both consumer-grade WiFi devices and professional equipment across diverse environments. These datasets can also serve as a valuable resource for benchmarking fingerprint-based indoor localization methods. Extensive experiments conducted on both synthetic and real-world datasets demonstrate that the AGML model-based localization method consistently outperforms all baseline methods using sparse fingerprints across all evaluated metrics.
Wenzhong Yan, Feng Yin 0001, Ruizhi Chen
IEEE Trans. Mob. Comput.6
2026 Beam-Switching-Based Time-of-Arrival Ranging on Commercial 5G NR Signals for Outdoor Positioning
abstract
The widespread adoption of 5G wireless communication devices has significantly increased the demand for precise 5G-based positioning services. This study presents a beam-switching time-of-arrival (TOA) ranging technique that leverages commercial 5G new radio (NR) signals and channel state information (CSI) extracted from the physical broadcast channel (PBCH). Initially, theoretical conditions for consistent multi-beam TOA estimation are derived, along with a multi-beam demodulation method designed to meet these conditions. To overcome the challenge of unknown base station (BS) radiation patterns, a signal quality-based beam-switching strategy is developed. Furthermore, a TOA tracking framework is introduced, integrating orthogonal matching pursuit (OMP) for multipath resolution, second-order frequency-locked loop (FLL)-assisted third-order delay-locked loop (DLL) for robust TOA tracking, and total variation regularization for anomaly removal. A software-defined radio (SDR) 5G receiver tailored for commercial beamforming BSs is implemented to validate the proposed system. With clock effects removed, the proposed system yields root-mean-square ranging errors of 3.73m and 4.99m in two distinct complex scenarios, corresponding to an average improvement of 69% in ranging accuracy over single-beam methods.
Wenxin Dong, Liang Chen 0007, Zhanghai Ju, Zhaoliang Liu, Ruizhi Chen
IEEE Trans. Wirel. Commun.6
2026 Beam-Switching-Based Joint DOA and TOA Acquisition and Tracking Using Commercial 5G NR Signals for Outdoor Positioning
abstract
The proliferation of 5G wireless devices creates a pressing need for high-precision positioning services. This paper presents an integrated acquisition–and–tracking framework that leverages physical broadcast channel (PBCH) transmissions from commercial 5G New Radio (NR) base stations (BSs). To select the strongest downlink beam without prior knowledge of the BS radiation pattern, we use an adaptive beam-switching strategy based on reference signal received power (RSRP), reference signal received quality (RSRQ), and signal-to-noise ratio (SNR). For coarse acquisition, direction-of-arrival (DOA) and time-of-arrival (TOA) estimates are produced by a two-stage procedure that combines three-dimensional sparse Bayesian learning (SBL) with gradient-ascent off-grid refinement. A closed-loop tracker then continuously refines these estimates through azimuth-locked loop (ALL), elevation-locked loop (ELL), and delay-locked loop (DLL) modules. We further derive information-theoretic limits that bound acquisition and tracking and elucidate the key factors shaping these limits. The complete framework is implemented on a software-defined radio (SDR) 5G receiver and validated in outdoor field trials. Results indicate robust performance, with an average TOA root-mean-square error (RMSE) of 3.58 m, an azimuth RMSE of 5.77°, and an elevation RMSE of 2.74°, while requiring only 2.64 switching events per minute.
Wenxin Dong, Zhanghai Ju, Hongjian Jiao, Zhaoliang Liu, Ruizhi Chen, Liang Chen 0007
IEEE Trans. Wirel. Commun.5
2025 QiMeng-GEMM: Automatically Generating High-Performance Matrix Multiplication Code by Exploiting Large Language Models
abstract
As a crucial operator in numerous scientific and engineering computing applications, the automatic optimization of General Matrix Multiplication (GEMM) with full utilization of ever-evolving hardware architectures (e.g. GPUs and RISC-V) is of paramount importance. While Large Language Models (LLMs) can generate functionally correct code for simple tasks, they have yet to produce high-performance code. The key challenge resides in deeply understanding diverse hardware architectures and crafting prompts that effectively unleash the potential of LLMs to generate high-performance code. In this paper, we propose a novel prompt mechanism called QiMeng-GEMM which enables LLMs to comprehend the architectural characteristics of different hardware platforms and automatically search for the optimization combinations for GEMM. The key of QiMeng-GEMM is a set of informative, adaptive, and iterative meta-prompts. Based on this, a searching strategy for optimal combinations of meta-prompts is used to iteratively generate high-performance code. Extensive experiments conducted on 4 leading LLMs, various paradigmatic hardware platforms, and representative matrix dimensions unequivocally demonstrate QiMeng-GEMM’s superior performance in auto-generating optimized GEMM code. Compared to vanilla prompts, our method achieves a performance enhancement of up to 113×. Even when compared to human experts, our method can reach 115% of cuBLAS on NVIDIA GPUs and 211% of OpenBLAS on RISC-V CPUs. Notably, while human experts often take months to optimize GEMM, our approach reduces the development cost by over 240×.
Qirui Zhou, Yuanbo Wen 0001, Ruizhi Chen, Ke Gao 0012, Weiqiang Xiong, Ling Li 0001, Qi Guo 0001, Yunji Chen
AAAI3
2025 QiMeng-TensorOp: One-Line Prompt is Enough for High-Performance Tensor Operator Generation with Hardware Primitives
abstract
Computation-intensive tensor operators constitute over 90% of the computations in Large Language Models (LLMs) and Deep Neural Networks. Automatically and efficiently generating high-performance tensor operators with hardware primitives is crucial for diverse and ever-evolving hardware architectures like RISC-V, ARM, and GPUs, as manually optimized implementation takes at least months and lacks portability. LLMs excel at generating high-level language codes, but they struggle to fully comprehend hardware characteristics and produce high-performance tensor operators. We introduce a tensor-operator auto-generation framework with a one-line user prompt (QiMeng-TensorOp), which enables LLMs to automatically exploit hardware characteristics to generate tensor operators with hardware primitives, and tune parameters for optimal performance across diverse hardware. Experimental results on various hardware platforms, SOTA LLMs, and typical tensor operators demonstrate that QiMeng-TensorOp effectively unleashes the computing capability of various hardware platforms, and automatically generates tensor operators of superior performance. Compared with vanilla LLMs, QiMeng-TensorOp achieves up to 1291× performance improvement. Even compared with human experts, QiMeng-TensorOp could reach 251% of OpenBLAS on RISC-V CPUs, and 124% of cuBLAS on NVIDIA GPUs. Additionally, QiMeng-TensorOp also significantly reduces development costs by 200× compared with human experts.
Xuzhi Zhang, Shaohui Peng, Qirui Zhou, Yuanbo Wen 0001, Qi Guo 0001, Ruizhi Chen, Xinguo Zhu, Weiqiang Xiong, Haixin Chen, Congying Ma, Ke Gao 0012, Yunji Chen, Ling Li 0001
IJCAI6
2025 A High-Performance and Memory-Efficient RISC-V Operating System Optimization for AIoT
abstract
The openness and flexibility of the RISC-V instruction set architecture (ISA) have driven its widespread adoption in AIoT (Artificial Intelligence of Things) devices. However, existing operating systems (OSes) for real RISC-V hardware often suffer from poor application performance and large memory footprints. To address these issues, we propose an OS optimization scheme tailored for RISC-V in AIoT devices. First, we introduce an application-transparent performance enhancement mechanism that leverages both coarse- and fine-grained process management to improve the performance of applications, particularly in AI inference. Second, we design a low-memory-footprint software stack through theoretical analysis and careful trade-offs in the adoption of software components. Lastly, we develop a lightweight OS image construction strategy algorithm tailored for RISC-V in AIoT. Using our OS optimization scheme, we build PolyOS from scratch to reduce the OS image size, thereby further lowering memory footprint. Across four real RISC-V hardware platforms, PolyOS achieves up to a 142% overall system performance improvement and up to 5.90× speedup in AI inference applications compared to baseline OSes (Armbian, Nucleisys, etc.). It also significantly reduces the runtime memory footprint of the standard C library, OpenCV, QuickJS, and AI inference applications, while shrinking the OS image size to 1/3.14–1/23.48 of its baseline OS.
Limin Cheng, Ke Gao 0012, Jiageng Yu, Ruizhi Chen, Ling Li 0001
SMC4
2025 Tracking foresters and mapping tree stem locations with decimeter-level accuracy under forest canopies using UWB
Zuoya Liu, Harri Kaartinen, Teemu Hakala, Juha Hyyppä, Antero Kukko, Ruizhi Chen
Expert Syst. Appl.6
2025 Morphology generalizable reinforcement learning via multi-level graph features
Yansong Pan, Rui Zhang 0040, Jiaming Guo, Shaohui Peng, Kaizhao Yuan, Yunkai Gao 0001, Siming Lan, Ruizhi Chen, Ling Li 0001, Xing Hu 0001, Zidong Du, Xin Zhang 0062, Wei Li 0008, Qi Guo 0001, Yunji Chen
Neurocomputing9
2025 Time-of-Arrival Estimation in Challenging Environments Using Commercial 5G NR Signals for Outdoor Positioning
abstract
The proliferation of 5G Internet of Things (IoT) devices in daily life has significantly increased the demand for 5G-based positioning services. This article presents a methodology for accurate time-of-arrival (TOA) estimation of commercial 5G new radio (NR) signals, specifically addressing the severe interference commonly encountered in challenging environments. To mitigate high noise levels, we first apply a raised cosine filter for initial denoising of channel state information (CSI), followed by frequency-domain averaging and normalization to further suppress residual noise. For precise TOA estimation, we develop an algorithm that leverages a delay lock loop (DLL) for robust signal tracking, enhanced by initialization and relocking mechanisms supported by space-alternating generalized expectation-maximization (SAGE). Additionally, we adapt several classic global positioning system (GPS) delay discriminators for compatibility with 5G NR signals, aiming to identify DLLs that perform effectively under adverse conditions. Comprehensive validation through simulations and field tests demonstrates the proposed system’s robustness and practical applicability. Furthermore, we identify delay discriminators that are particularly well-suited for deployment in challenging environments.
Wenxin Dong, Liang Chen 0007, Zhanghai Ju, Zhenhang Jiao, Ruizhi Chen
IEEE Internet Things J.5
2025 Indoor Positioning With Smartphone by Using Doppler Observations From Asynchronous Pseudolite System
abstract
Smartphones provide good outdoor positioning services through the global navigation satellite system (GNSS), playing an important role in various fields, such as the Internet of Things (IoT) and smart logistics. However, since GNSS are blocked by buildings in indoor environments, there is currently no general technology or solution for indoor positioning in outdoor environments like GNSS. As a supplement to GNSS, it has been verified that pseudolite systems can make full use of the GNSS chips embedded in smartphones to provide raw observations for indoor positioning services. Aiming at the problem of indoor positioning of smartphones, a low-cost distributed asynchronous pseudolite system and an indoor positioning method based on Doppler observations are proposed. The asynchronous pseudolite system consists of multiple dual-channel transmitters and uses Doppler raw observations to reduce the need for precise synchronization of signal transmission time. To evaluate the feasibility and positioning accuracy of the method, static and dynamic experiments were carried out in a large underground garage of an building using commercial smartphones. The field experiment shows that the indoor pseudolite positioning method proposed in this article achieves static decimeter-level and dynamic meter-level positioning accuracy for smartphones. Compared with the indoor positioning technology of smartphones based on radio frequency (RF) signals, such as Wi-Fi and Bluetooth, this study explores a new indoor positioning system for smartphones, enriches the observation information, and brings more possibilities for smartphones indoor positioning. Moreover, this study discusses an on-the-fly solution for initialization without known points using only Doppler observations and motion diversity.
Xiangchen Lu, Liang Chen 0007, Nan Shen, Ruizhi Chen
IEEE Internet Things J.6
2025 A Multigranularity Spatiotemporal Attention Model Based on Multisource Satellite Data for Monthly XCO2 Reconstruction Over China
abstract
Accurate and timely monitoring of atmospheric CO2concentrations is a crucial prerequisite for achieving carbon peaking and carbon neutrality goals. However, existing satellite observations are often limited by cloud cover, sensor constraints, and orbital characteristics, resulting in data gaps that hinder the direct analysis of CO2spatiotemporal dynamics. To address this, we develop an operational processing pipeline that leverages an integrated deep learning framework to reconstruct monthly, full-coverage XCO₂ over China at 0.05° spatial resolution for the period 2019–2024. This pipeline integrates multi-source satellite data (GOSAT, OCO-2, GF-5B) with auxiliary variables including vegetation indices, nighttime light remote sensing, NO₂ column concentrations, and ERA5 meteorological parameters. By eliminating reliance on assimilated model products, the pipeline generates monthly XCO₂ maps with a processing latency of 1–2 months after satellite observation.Validation against TCCON ground-based observations demonstrates strong reconstruction performance (R² = 0.88, RMSE = 1.50 ppm). When compared to CarbonTracker and CAMS datasets, the reconstructed dataset shows mean biases of −0.28 ppm and −0.46 ppm, and standard deviations of 0.53 ppm and 0.72 ppm, respectively, indicating that it accurately captures the spatial distribution and interannual trends of CO₂ concentrations. Comparative experiments further reveal that incorporating GF-5B data improves the overall model performance. Regional analyses show that the model effectively identifies high-emission hotspot areas. Interannual analysis indicates that the average XCO₂ in China increased from 410.30 ppm in 2019 to 422.73 ppm in 2024, corresponding to an annual growth rate between 0.4% and 0.9%.
Ruizhi Chen, Mingmin Zou, Cuihong Chen, Huizhen Xie, Chunyan Zhou, Huiqin Mao, Zhongting Wang
IEEE Trans. Geosci. Remote. Sens.1
2025 Chorus: Robust Multitasking Local Client-Server Collaborative Inference With Wi-Fi 6 for AIoT Against Stochastic Congestion Delay
abstract
The rapid growth of AIoT devices brings huge demands for DNNs deployed on resource-constrained devices. However, the intensive computation and high memory footprint of DNN inference make it difficult for the AIoT devices to execute the inference tasks efficiently. In many widely deployed AIoT use cases, multiple local AIoT devices launch DNN inference tasks randomly. Although local collaborative inference has been proposed to accelerate DNN inference on local devices with limited resources, multitasking local collaborative inference, which is common in AIoT scenarios, has not been fully studied in previous works. We consider multitasking local client-server collaborative inference (MLCCI), which achieves efficient DNN inference by offloading the inference tasks from multiple AIoT devices to a more powerful local server with parallel pipelined execution streams through Wi-Fi 6. Our optimization goal is to minimize the mean end-to-end latency of MLCCI. Based on the experiment results, we identify three key challenges: high communication costs, high model initialization latency, and congestion delay brought by task interference. We analyze congestion delay in MLCCI and its stochastic fluctuations with queuing theory and propose Chorus, a high-performance adaptive MLCCI framework for AIoT devices, to minimize the mean end-to-end latency of MLCCI against stochastic congestion delay. Chorus generates communication-efficient model partitions with heuristic search, uses a prefetch-enabled two-level LRU cache to accelerate model initialization on the server, reduces congestion delay and its short-term fluctuations with execution stream allocation based on the cross-entropy method, and finally achieves efficient computation offloading with reinforcement learning. We established a system prototype, which statistically simulated many virtual clients with limited physical client devices to conduct performance evaluations, for Chorus with real devices. The evaluation results for various workload levels show that Chorus achieved an average of$1.4\times$,$1.3\times$, and$2\times$speedup over client-only inference, and server-only inference with LRU and MLSH, respectively.
Yuzhe Luo, Ji Qi 0002, Ling Li 0001, Ruizhi Chen, Limin Cheng
IEEE Trans. Parallel Distributed Syst.4
2025 Optimizing wireless sensor network topology with node load consideration
abstract
Background With the development of the Internet, the topology optimization of wireless sensor networks has received increasing attention. However, traditional optimization methods often overlook the energy imbalance caused by node loads , which affects network performance. Methods To improve the overall performance and efficiency of wireless sensor networks , a new method for optimizing the wireless sensor network topology based on K-means clustering and firefly algorithms is proposed. The K-means clustering algorithm partitions nodes by minimizing the within-cluster variance, while the firefly algorithm is an optimization algorithm based on swarm intelligence that simulates the flashing interaction between fireflies to guide the search process. The proposed method first introduces the K-means clustering algorithm to cluster nodes and then introduces a firefly algorithm to dynamically adjust the nodes. Results The results showed that the average clustering accuracies in the Wine and Iris data sets were 86.59% and 94.55%, respectively, demonstrating good clustering performance. When calculating the node mortality rate and network load balancing standard deviation, the proposed algorithm showed dead nodes at approximately 50 iterations, with an average load balancing standard deviation of 1.7×10 4 , proving its contribution to extending the network lifespan. Conclusions This demonstrates the superiority of the proposed algorithm in significantly improving the energy efficiency and load balancing of wireless sensor networks to extend the network lifespan. The research results indicate that wireless sensor networks have theoretical and practical significance in fields such as monitoring, healthcare, and agriculture.
Ruizhi Chen
Virtual Real. Intell. Hardw.1
2024 Hypothesis, Verification, and Induction: Grounding Large Language Models with Self-Driven Skill Learning
abstract
Large language models (LLMs) show their powerful automatic reasoning and planning capability with a wealth of semantic knowledge about the human world. However, the grounding problem still hinders the applications of LLMs in the real-world environment. Existing studies try to fine-tune the LLM or utilize pre-defined behavior APIs to bridge the LLMs and the environment, which not only costs huge human efforts to customize for every single task but also weakens the generality strengths of LLMs. To autonomously ground the LLM onto the environment, we proposed the Hypothesis, Verification, and Induction (HYVIN) framework to automatically and progressively ground the LLM with self-driven skill learning. HYVIN first employs the LLM to propose the hypothesis of sub-goals to achieve tasks and then verify the feasibility of the hypothesis via interacting with the underlying environment. Once verified, HYVIN can then learn generalized skills with the guidance of these successfully grounded subgoals. These skills can be further utilized to accomplish more complex tasks that fail to pass the verification phase. Verified in the famous instruction following task set, BabyAI, HYVIN achieves comparable performance in the most challenging tasks compared with imitation learning methods that cost millions of demonstrations, proving the effectiveness of learned skills and showing the feasibility and efficiency of our framework.
Shaohui Peng, Xing Hu 0001, Qi Yi, Rui Zhang 0040, Jiaming Guo, Zikang Tian, Ruizhi Chen, Zidong Du, Qi Guo 0001, Yunji Chen, Ling Li 0001
AAAI8
2024 OCEAN-MBRL: Offline Conservative Exploration for Model-Based Offline Reinforcement Learning
abstract
Model-based offline reinforcement learning (RL) algorithms have emerged as a promising paradigm for offline RL. These algorithms usually learn a dynamics model from a static dataset of transitions, use the model to generate synthetic trajectories, and perform conservative policy optimization within these trajectories. However, our observations indicate that policy optimization methods used in these model-based offline RL algorithms are not effective at exploring the learned model and induce biased exploration, which ultimately impairs the performance of the algorithm. To address this issue, we propose Offline Conservative ExplorAtioN (OCEAN), a novel rollout approach to model-based offline RL. In our method, we incorporate additional exploration techniques and introduce three conservative constraints based on uncertainty estimation to mitigate the potential impact of significant dynamic errors resulting from exploratory transitions. Our work is a plug-in method and can be combined with classical model-based RL algorithms, such as MOPO, COMBO, and RAMBO. Experiment results of our method on the D4RL MuJoCo benchmark show that OCEAN significantly improves the performance of existing algorithms.
Rui Zhang 0040, Qi Yi, Yunkai Gao 0001, Jiaming Guo, Shaohui Peng, Siming Lan, Husheng Han, Yansong Pan, Kaizhao Yuan, Pengwei Jin, Ruizhi Chen, Yunji Chen, Ling Li 0001
AAAI12
2024 Emergent Communication for Numerical Concepts Generalization
abstract
Research on emergent communication has recently gained significant traction as a promising avenue for the linguistic community to unravel human language's origins and explore artificial intelligence's generalization capabilities. Current research has predominantly concentrated on recognizing qualitative patterns of object attributes(e.g., shape and color) and paid little attention to the quantitative relationship among object quantities which is known as the part of numerical concepts. The ability to generalize numerical concepts, i.e., counting and calculations with unseen quantities, is essential, as it mirrors humans' foundational abstract reasoning abilities. In this work, we introduce the NumGame, leveraging the referential game framework, forcing agents to communicate and generalize the numerical concepts effectively. Inspired by the human learning process of numbers, we present a two-stage training approach that sequentially fosters a rudimentary numerical sense followed by the ability of arithmetic calculation, ultimately aiding agents in generating semantically stable and unambiguous language for numerical concepts. The experimental results indicate the impressive generalization capabilities to unseen quantities and regularity of the language emergence from communication.
Enshuai Zhou, Yifan Hao 0001, Rui Zhang 0040, Zidong Du, Xishan Zhang, Xinkai Song, Chao Wang 0003, Xuehai Zhou, Jiaming Guo, Qi Yi, Shaohui Peng, Ruizhi Chen, Qi Guo 0001, Yunji Chen
AAAI14
2024 Privacy-preserving Compression for Efficient Collaborative Inference
abstract
Collaborative inference accelerates DNN inference tasks of resource-limited devices (e.g., clients) by offloading model slices to resource-rich devices (e.g., servers). During the inference procedure, outputs of model slices are transmitted among devices, causing significant intermediate data transmission overhead and posing a risk of privacy leakage of the client’s input data. Quantization has been widely used in collaborative inference to enhance communication efficiency. However, traditional quantization cannot prevent data privacy leaks. Besides, perturbation-based privacy protection methods, such as adding Laplace noise to the intermediate data of collaborative inference, do not consider communication efficiency. In this paper, we introduce Layered Laplace Random Quantization to simultaneously achieve communication efficiency and data privacy protection in collaborative inference by compressing the intermediate data with Laplace quantization noise. We also propose stability training to recover the accuracy loss caused by our method. Evaluation results show that our method achieved an average inference latency speedup of 1.2x-1.3x for different DNN models compared with the baseline methods while achieving comparable data privacy protection and recoverable accuracy loss.
Yuzhe Luo, Ji Qi 0002, Jiageng Yu, Ruizhi Chen, Ke Gao 0012, Ling Li 0001
ICPADS4
2024 CrowdLOC-S: Crowdsourced seamless localization framework based on CNN-LSTM-MLP enhanced quality indicator
Yue Yu 0003, Liang Chen 0007, Ruizhi Chen
Expert Syst. Appl.5
2024 Autonomous wireless positioning system using crowdsourced Wi-Fi fingerprinting and self-detected FTM stations
Fangli Guan, Kexin Tang, Sheng Bao, Liang Chen 0007, Ruizhi Chen, Yue Yu 0003
Expert Syst. Appl.6
2024 Dual-Step Acoustic Chirp Signals Detection Using Pervasive Smartphones in Multipath and NLOS Indoor Environments
abstract
Indoor localization techniques based on acoustic signals have been a focus of research community in the past decades due to their high accuracy and ubiquity. However, there are still some limitations that need to be overcome, such as multipath and non-line-of-sight (NLOS). To achieve robust and high precision acoustic ranging for practical applications on most smartphones, we propose a dual-step chirp signal detection algorithm consisting of coarse and fine searches. For robustness, the coarse search extracts the acoustic data segment containing the direct path by monitoring the energy changes based on the time-frequency (TF) analysis methods. For improving the accuracy and stability, adaptive slack and strict thresholds are introduced in cross-correlation function (CCF)-based fine search. Meanwhile, an extremum normalization method is proposed to alleviate the smartphones differences and near-far effects. A thresholds determination experiment and two practical applications are implemented on the proposed algorithm. Threshold determination experimental results show that in multipath and NLOS scenarios, the proposed coarse search can reach a success rate of more than 99.9% and an error rate of less than 0.4%. Furthermore, the proposed fine search offers a ranging accuracy with an average error and root-mean-square-error (RMSE) of less than 0.25 m and 0.35 m, respectively. For practical applications, ranging accuracies of 0.17 m and 0.14 m at 50%, and 0.59 m and 0.54 m at 95% are achieved in two typical indoor environments, which are superior to those achieved by two conventional CCF-based detection algorithms.
Zheng Li 0025, Ruizhi Chen, Guangyi Guo, Feng Ye 0003, Lixiong Huang, Liang Chen 0007
IEEE Internet Things J.2
2024 ChirpTracker: A Precise-Location-Aware System for Acoustic Tag Using Single Smartphone
abstract
The increasing interest in loss prevention devices using the Internet of Things, has been driven by the convenience, low cost, and low-power consumption of these devices. However, the existing technologies cannot achieve a balance between high availability over a wide area with a single smartphone and precise location awareness. In this article, a novel precise-location-aware method that integrates acoustic technology and pedestrian dead reckoning (PDR), ChirpTracker, is proposed which most smartphones support without auxiliary equipment. This system can provide wide coverage (30 m) and is suitable for many scenarios, such as searching for cars in underground parking or finding items indoors. ChirpTracker can detect the distance between the smartphone and a lost tag in real time using acoustic signals, it can monitor the relative position change of the smartphone based on deep learning-based PDR and update relative positioning of the lost tag though the observation from single base-station in motion. A technology that combines the local least squares method (LSM) and particle filter (PF) improves the convergence and the robustness of ChirpTracker through an identification strategy for a mirror position. This method was validated in experiments conducted in actual environments. The results demonstrate the effectiveness and positioning accuracy of ChirpTracker.
Xinchuang Lin, Ruizhi Chen, Lixiong Huang, Zuoya Liu, Xiaoguang Niu, Guangyi Guo, Zheng Li 0025
IEEE Internet Things J.2
2024 Submeter-Level ToF-Based Acoustic Positioning of Moving Objects With Chirp-Based Doppler Shift Compensation
abstract
Existing acoustic-based positioning solutions face difficulties achieving precise ranging and positioning, especially in dynamic situations, due to Doppler frequency shift (DFS). In this article, we present a solution that achieves precise ToF/distance measurements between the kinematic receiver and a stationary transmitter with chirp-based Doppler shift compensation (DSC). In the solution, specific chirp signals with an upchirp and downchirp branch are transmitted by the stationary transmitter. The kinematic receiver receives and detects these signals, accordingly corrects the measurements with the proposed DSC method, and estimates the real-time velocity based on a corresponding model. After obtaining the compensated ToF/distance measurements and real-time velocities of the kinematic receiver, the initial and subsequent locations of the kinematic receiver can be precisely determined with the extended Kalman filter (EKF) and Rauch-Tung-Striebel smoother (RTS). To verify the performance of our solution, experiments in ranging and positioning were conducted in an indoor open space. The results show that the developed DSC is able to achieve an average ranging accuracy of 0.1 m for the kinematic receiver with a motion velocity of larger than 1.5 m/s in line-of-sight (LOS) situations and achieves an average positioning accuracy of 0.46 m for the kinematic receiver with motion velocity up to approximately 2 m/s. Therefore, the developed approach is sufficient for realizing acoustic-based positioning in both static and dynamic situations.
Zuoya Liu, Ruizhi Chen, Changhui Jiang, Feng Ye 0003, Guangyi Guo, Liang Chen 0007, Xinchuang Lin
IEEE Internet Things J.2
2024 IALoc: Audio-Chirp-Based Indoor Tracking System - Free From IMU Sensors Dependence
abstract
The smart upgrade of large indoor venues, such as airports, exhibition centers, etc., and the rapid expansion of urban underground spaces demand modern indoor positioning technologies. However, most good positioning technologies need to fuse the inertial measurement unit (IMU) sensors to enhance the localization robustness and accuracy. In this work, an indoor audio chirp-based localization system (IALoc), which is no longer relying on the IMU sensors, is developed. We designed dedicated anchors based on embedded hardware, between which stable measurements is provided via synchronous audio networks and broadcasting strategies. The proposal distribution is improved by an empirical model of human motion in the proposed improved unscented particle filter (IUPF). The experimental results show that IALoc is able to cover the full scene with 0.6 m tracking accuracy and 1 Hz update rate in both typical indoor office and exhibition hall scenarios. As compared to UPF that carried the same number of particles, the IUPF saves 17.74% of computation time and improves the positioning accuracy by 14.29%. Since there is no need to consider the attitude of smartphones when using IUPF, it could show a considerable value of applications, such as security working and emergency rescue.
Ruizhi Chen, Guangyi Guo, Zheng Li 0025, Feng Ye 0003, Lixiong Huang, Zuoya Liu
IEEE Internet Things J.2
2024 A Benchmark of Absolute and Relative Positioning Solutions in GNSS Denied Environments
abstract
Precise positioning is fundamental to the internet of things that delivers insights into everything from large-scale business to ordinary smart life. Accurate localization and positioning in global navigation satellite system (GNSS) denied environments, such as indoor-, underground-spaces, and forests, is one of the most prosperous research fields because of the great complexity prompted by various challenging application scenarios. Different sensors, algorithms, and combinations of those have been developed in past decades, which provided a great variety of possible solutions that deliver different positioning accuracies. However, a rigorous evaluation of the positioning accuracy of different mainstream solutions is missing, mainly because of the difficulties in acquiring reliable ground truth for referencing and the lack of comparable test/application conditions. A comprehensive benchmarking was carried out in this study based on the comparisons of six solutions that consist of different combinations of five positioning technologies, i.e., 1) ultra-wideband (UWB) and inertial measurement unit (IMU); 2) UWB, IMU, and camera; 3) UWB and light detection and ranging (LIDAR); 4) UWB and radio detection and ranging (RADAR); 5) IMU, camera and LIDAR; and 6) UWB, IMU, camera and LIDAR. The five technologies, i.e., UWB, IMU, camera, RADAR, and LIDAR, were commonly regarded as those that are with high applicability, accuracy, and robustness. New anchors self-positioning algorithm and integrity monitoring algorithm were proposed to further aid the compared solutions and the benchmark. High-precision survey (millimeter) -level ground truth references were acquired at indoor and outdoor test locations and applied in the evaluations, to assist reliable quantitive benchmarks about the positioning accuracies and stabilities of the compared solutions. The strengths, limitations, and potentials of each solution were analyzed. It was revealed that all relative positioning solutions accumulate positioning errors over time. Such accumulation was of the highest significance for RADAR, followed by camera. LIDAR is presented to be the most robust solution for relative positioning. Compared to camera, LIDAR, and RADAR alone, the integration of different technologies clearly improved the performance. The tight-coupling performed slightly superior to loose-coupling, and the unscented Kalman filter with tight-coupling had a higher positioning accuracy in most cases.
Haiyun Yao, Xinlian Liang, Ruizhi Chen, Hanwen Qi, Liang Chen 0007, Yunsheng Wang 0002
IEEE Internet Things J.3
2024 Multiple Similarity Analysis-Based Deep Metric Learning for Enhancing Wi-Fi Fingerprint Indoor Localization
abstract
Wi-Fi RSSI fingerprint localization, one of the more mature solutions for indoor localization, is active in Internet of Things (IoT) applications. However, there are still some challenges to learning RSSI features. Although recent methods utilize deep metric learning techniques to discover the hidden correspondences between RSSI features and physical spatial locations, because the currently used deep metric learning methods focus more on the positive similarity (Similarity-P) of RSSI features, ignoring the constraints imposed by self-similarity (Similarity-S) and negative similarity (Similarity-N) on the model and having to ternary group RSSIs in advance. We propose a deep metric learning indoor localization method based on multiple similarities, constructed with a multiple similarity loss function in which Similarity-S pair mining of RSSI features is performed to mine positive and negative pair sets, which are then pair weighting using Similarity-P and Similarity-N. Finally, RSSI features are extracted and localized online using the WKNN method. Experiment results confirm that our proposed method can achieve excellent positioning performance with few-shot data sets.
Shuai Zhang 0016, Ruizhi Chen
IEEE Internet Things J.3
2024 IMPos: Indoor Mobile Positioning With 5G Multibeam Signals From a Single Base Station
abstract
With the widespread deployment of the fifth-generation (5G) network indoors, commercial 5G signals are highly attractive in the field of indoor positioning because of their ubiquity. Considering the user equipment (UE) requirements for user privacy protection, low computational resource consumption, and the need for location services in mobile conditions, this study developed a low-cost indoor mobile positioning system based on 5G downlink multi-beam signals, termed IMPos. In particular, this research only uses the multi-beam reference signal received power as data source, which is derived from a single commercially deployed base station (BS) and received by a single receiving antenna. Based on this data source, a machine learning method is first proposed for floor-level recognition. Thereafter, a G2Bi network based on stacked recurrent neural networks is designed to achieve UE mobile self-positioning. To evaluate the performance of IMPos, field tests are carried out in different floor scenarios. Results show that even with just one BS, IMPos achieves a floor-level recognition accuracy exceeding 95% and a mobile positioning root-mean-square error of below 1.5 m in various scenarios.
Xin Zhou 0006, Liang Chen 0007, Yanlin Ruan, Ruizhi Chen
IEEE Internet Things J.5
2024 Dynamic selection for reconstructing instance-dependent noisy labels
Jie Yang 0002, Xiaoguang Niu, Yuanzhuo Xu, Zejun Zhang 0002, Guangyi Guo, He Zhu 0002, Ruizhi Chen
Pattern Recognit.7
2024 Neural Network Aided Factor Graph Optimization for Collaborative Pedestrian Navigation
abstract
Indoor navigation and positioning services for pedestrians are challenging because of the lack of satellite signals and the unpredictability of pedestrian motion. The inertial measurement unit (IMU)-based pedestrian dead reckoning (PDR) algorithm can provide continuous position estimation for individual pedestrians. However, the accumulation of errors leads to inaccurate pedestrian position results. Radio signals such as ultra-wideband (UWB) can range between pedestrians and anchors and provide high-precision positioning information; nonetheless, radio positioning requires infrastructure deployment and maintenance in indoor environments, thus limiting the popularization and implementation of these technologies. In this paper, a neural network aided factor graph optimization (NN-FGO) method was proposed for collaborative pedestrian navigation (CPN). It integrates IMU and UWB sensors to implement PDR for individual pedestrians and CPN for the Ad-Hoc network, and it is infrastructure-free since all the sensors are wearable. For a small or sparse network, ranging constraints will be insufficient to implement an acceptable CPN. A neural network model was suggested for human activity recognition and position loopback detection, which provide virtual constraints for pedestrians. For the heterogeneous problem caused by multiple collaborative signals and constraints, FGO was employed to solve the motion states of multi-pedestrians and multi-epochs. The real experimental results revealed that NN-FGO can provide 92% accuracy in activity classification. Compared with the extended Kalman filter based CPN, the average position error decreased by 19.6% and 16.0% with triangular and parallel straight geometries, respectively.
Mingxi Wang, Jingbin Liu, Ruizhi Chen
IEEE Trans. Intell. Transp. Syst.4
2024 UltraMotion: High-Precision Ultrasonic Arm Tracking for Real-World Exercises
abstract
Home exercise and self-served gyms allow a larger population to exercise regularly without the cost of hiring private coaches. In absence of professional guidance, however, exercisers can suffer from injuries to muscles and joints. High-precision, affordable arm tracking with commercial, off-the-shelf (COTS) wearable devices has become an urgent need to prevent workout injuries and improve exercise performance. Recent studies with inertial measurement units (IMUs) or audio signals are neither computationally feasible for real-time motion tracking with satisfactory accuracy using COTS devices nor practically usable due to the interference with noisy ambient environments. In this paper, we propose UltraMotion, a real-time, high-precision ultrasonic arm motion tracking system designed for practical use. UltraMotion performs point cloud queries based on hidden Markov models (HMMs), a novel ultrasonic acoustic ranging method, and an extended Kalman filter (EKF) to predict the locations of all three arm joints, making it the first system offering shoulder locations. Experimental results with only a smartphone and a smartwatch demonstrate the effectiveness of UltraMotion in tracking shoulder, elbow, and wrist locations with impressively small median errors of 6.4 cm, 7.1 cm, and 8.5 cm in real-world environments, outperforming all previous systems, making UltraMotion an ideal choice for daily exercise.
Xiaoguang Niu, Kaiyi Zou, Da Shen, He Zhu 0002, Shaowu Wu, Guangyi Guo, Ruizhi Chen
IEEE Trans. Mob. Comput.7
2024 Indoor Localization With Multi-Beam of 5G New Radio Signals
abstract
In this work, we investigate the property of the multi-beam of 5G new radio (NR) signals for indoor localization. Specifically, the 5G NR signals are firstly sampled by a self-developed software-defined receiver, and the multi-beam is extracted via detecting the multiple synchronization signal blocks (SSBs). Secondly, with the assistance of the pilots in the multiple SSBs, the reference signal received power (RSRP) and reference signal received quality (RSRQ) of the multi-beam are calculated. Thirdly, by stacking the RSRP and RSRQ of the multi-beam as the observables, a fingerprint database is constructed. With the aim to efficiently process the fingerprint features and improve the accuracy of indoor localization, a CatBoost-based algorithm is proposed, and the parameters are further optimized by tree-structured parzen estimator (TPE). To verify the effectiveness of the proposed method, indoor field tests are carried out in an office scenario, where real 5G signals are transmitted from a commercial 5G NR base station indoors. The field tests demonstrate that, by taking the advantages of the multi-beam of 5G NR, the localization accuracy can be able to achieve the accuracy of 1.06 m in the metric of root mean squared error (RMSE), even when only one base station is heard indoors. By comparison with the single-beam, the accuracy of multi-beam has improved 48%.
Xin Zhou 0006, Liang Chen 0007, Yanlin Ruan, Ruizhi Chen
IEEE Trans. Wirel. Commun.4
2023 Conceptual Reinforcement Learning for Language-Conditioned Tasks
abstract
Despite the broad application of deep reinforcement learning (RL), transferring and adapting the policy to unseen but similar environments is still a significant challenge. Recently, the language-conditioned policy is proposed to facilitate policy transfer through learning the joint representation of observation and text that catches the compact and invariant information across various environments. Existing studies of language-conditioned RL methods often learn the joint representation as a simple latent layer for the given instances (episode-specific observation and text), which inevitably includes noisy or irrelevant information and cause spurious correlations that are dependent on instances, thus hurting generalization performance and training efficiency. To address the above issue, we propose a conceptual reinforcement learning (CRL) framework to learn the concept-like joint representation for language-conditioned policy. The key insight is that concepts are compact and invariant representations in human cognition through extracting similarities from numerous instances in real-world. In CRL, we propose a multi-level attention encoder and two mutual information constraints for learning compact and invariant concepts. Verified in two challenging environments, RTFM and Messenger, CRL significantly improves the training efficiency (up to 70%) and generalization ability (up to 30%) to the new environment dynamics.
Shaohui Peng, Xing Hu 0001, Rui Zhang 0040, Jiaming Guo, Qi Yi, Ruizhi Chen, Zidong Du, Ling Li 0001, Qi Guo 0001, Yunji Chen
AAAI6
2023 USDNL: Uncertainty-Based Single Dropout in Noisy Label Learning
abstract
Deep Neural Networks (DNNs) possess powerful prediction capability thanks to their over-parameterization design, although the large model complexity makes it suffer from noisy supervision. Recent approaches seek to eliminate impacts from noisy labels by excluding data points with large loss values and showing promising performance. However, these approaches usually associate with significant computation overhead and lack of theoretical analysis. In this paper, we adopt a perspective to connect label noise with epistemic uncertainty. We design a simple, efficient, and theoretically provable robust algorithm named USDNL for DNNs with uncertainty-based Dropout. Specifically, we estimate the epistemic uncertainty of the network prediction after early training through single Dropout. The epistemic uncertainty is then combined with cross-entropy loss to select the clean samples during training. Finally, we theoretically show the equivalence of replacing selection loss with single cross-entropy loss. Compared to existing small-loss selection methods, USDNL features its simplicity for practical scenarios by only applying Dropout to a standard network, while still achieving high model accuracy. Extensive empirical results on both synthetic and real-world datasets show that USDNL outperforms other methods. Our code is available at https://github.com/kovelxyz/USDNL.
Yuanzhuo Xu, Xiaoguang Niu, Jie Yang 0002, He Zhu 0002, Ruizhi Chen
AAAI6
2023 Online Prototype Alignment for Few-shot Policy Transfer
abstract
Domain adaptation in RL mainly deals with the changes of observation when transferring the policy to a new environment. Many traditional approaches of domain adaptation in RL manage to learn a mapping function between the source and target domain in explicit or implicit ways. However, they typically require access to abundant data from the target domain. Besides, they often rely on visual clues to learn the mapping function and may fail when the source domain looks quite different from the target domain. To address these problems, in this paper, we propose a novel framework Online Prototype Alignment (OPA) to learn the mapping function based on the functional similarity of elements and is able to achieve few-shot policy transfer within only several episodes. The key insight of OPA is to introduce an exploration mechanism that can interact with the unseen elements of the target domain in an efficient and purposeful manner, and then connect them with the seen elements in the source domain according to their functionalities (instead of visual clues). Experimental results show that when the target domain looks visually different from the source domain, OPA can achieve better transfer performance even with much fewer samples from the target domain, outperforming prior methods.
Qi Yi, Rui Zhang 0040, Shaohui Peng, Jiaming Guo, Yunkai Gao 0001, Kaizhao Yuan, Ruizhi Chen, Siming Lan, Xing Hu 0001, Zidong Du, Xishan Zhang, Qi Guo 0001, Yunji Chen
ICML7
2023 An LSTM Approach for Modelling Error of Smartphone-reported GNSS Location Under Mixed LOS/NLOS Environments
abstract
Modelling error of smartphone-reported Global Navigation Satellite System (GNSS) locations plays an important role in urban navigation under mixed LOS/NLOS environments. In the case of pedestrian navigation, the performance of GNSS error modeling significantly affects the precision of final multi-source fusion. In this work, a novel Long Short-Term Memory (LSTM) network is developed for error modeling of smartphone-reported GNSS locations combined with the detected human motion information. The LSTM network is applied to adaptively combine multi-level observations provided by GNSS and built-in sensors-based location sources under a specific time period instead of considering only adjacent timestamps. The motion features extracted from multi-level observations is then modeled as the input vector of LSTM for training and prediction purposes, and the predicted errors under two axis in the n-frame are finally modeled as the error covariance matrix and applied in the multi-sources fusion structure. The comprehensive experiments indicate the effectivity and significant improvement for integrated localization after GNSS error modeling.
Yue Yu 0003, Wenzhong Shi, Zhewei Liu, Shiyu Bai, Liang Chen 0007, Ruizhi Chen
IPIN6
2023 Context Shift Reduction for Offline Meta-Reinforcement Learning
abstract
Offline meta-reinforcement learning (OMRL) utilizes pre-collected offline datasets to enhance the agent's generalization ability on unseen tasks. However, the context shift problem arises due to the distribution discrepancy between the contexts used for training (from the behavior policy) and testing (from the exploration policy). The context shift problem leads to incorrect task inference and further deteriorates the generalization ability of the meta-policy. Existing OMRL methods either overlook this problem or attempt to mitigate it with additional information. In this paper, we propose a novel approach called Context Shift Reduction for OMRL (CSRO) to address the context shift problem with only offline datasets. The key insight of CSRO is to minimize the influence of policy in context during both the meta-training and meta-test phases. During meta-training, we design a max-min mutual information representation learning mechanism to diminish the impact of the behavior policy on task representation. In the meta-test phase, we introduce the non-prior context collection strategy to reduce the effect of the exploration policy. Experimental results demonstrate that CSRO significantly reduces the context shift and improves the generalization ability, surpassing previous methods across various challenging domains.
Yunkai Gao 0001, Rui Zhang 0040, Jiaming Guo, Qi Yi, Shaohui Peng, Siming Lan, Ruizhi Chen, Zidong Du, Xing Hu 0001, Qi Guo 0001, Ling Li 0001, Yunji Chen
NeurIPS8
2023 Efficient Symbolic Policy Learning with Differentiable Symbolic Expression
abstract
Deep reinforcement learning (DRL) has led to a wide range of advances in sequential decision-making tasks. However, the complexity of neural network policies makes it difficult to understand and deploy with limited computational resources. Currently, employing compact symbolic expressions as symbolic policies is a promising strategy to obtain simple and interpretable policies. Previous symbolic policy methods usually involve complex training processes and pre-trained neural network policies, which are inefficient and limit the application of symbolic policies. In this paper, we propose an efficient gradient-based learning method named Efficient Symbolic Policy Learning (ESPL) that learns the symbolic policy from scratch in an end-to-end way. We introduce a symbolic network as the search space and employ a path selector to find the compact symbolic policy. By doing so we represent the policy with a differentiable symbolic expression and train it in an off-policy manner which further improves the efficiency. In addition, in contrast with previous symbolic policies which only work in single-task RL because of complexity, we expand ESPL on meta-RL to generate symbolic policies for unseen tasks. Experimentally, we show that our approach generates symbolic policies with higher performance and greatly improves data efficiency for single-task RL. In meta-RL, we demonstrate that compared with neural network policies the proposed symbolic policy achieves higher performance and efficiency and shows the potential to be interpretable.
Jiaming Guo, Rui Zhang 0040, Shaohui Peng, Qi Yi, Xing Hu 0001, Ruizhi Chen, Zidong Du, Xishan Zhang, Ling Li 0001, Qi Guo 0001, Yunji Chen
NeurIPS6
2023 Emergent Communication for Rules Reasoning
abstract
Research on emergent communication between deep-learning-based agents has received extensive attention due to its inspiration for linguistics and artificial intelligence. However, previous attempts have hovered around emerging communication under perception-oriented environmental settings, that forces agents to describe low-level perceptual features intra image or symbol contexts. In this work, inspired by the classic human reasoning test (namely Raven's Progressive Matrix), we propose the Reasoning Game, a cognition-oriented environment that encourages agents to reason and communicate high-level rules, rather than perceived low-level contexts. Moreover, we propose 1) an unbiased dataset (namely rule-RAVEN) as a benchmark to avoid overfitting, 2) and a two-stage curriculum agent training method as a baseline for more stable convergence in the Reasoning Game, where contexts and semantics are bilaterally drifting. Experimental results show that, in the Reasoning Game, a semantically stable and compositional language emerges to solve reasoning problems. The emerged language helps agents apply the extracted rules to the generalization of unseen context attributes, and to the transfer between different context attributes or even tasks.
Yifan Hao 0001, Rui Zhang 0040, Enshuai Zhou, Zidong Du, Xishan Zhang, Xinkai Song, Yuanbo Wen 0001, Yongwei Zhao 0001, Xuehai Zhou, Jiaming Guo, Qi Yi, Shaohui Peng, Ruizhi Chen, Qi Guo 0001, Yunji Chen
NeurIPS15
2023 Contrastive Modules with Temporal Attention for Multi-Task Reinforcement Learning
abstract
In the field of multi-task reinforcement learning, the modular principle, which involves specializing functionalities into different modules and combining them appropriately, has been widely adopted as a promising approach to prevent the negative transfer problem that performance degradation due to conflicts between tasks. However, most of the existing multi-task RL methods only combine shared modules at the task level, ignoring that there may be conflicts within the task. In addition, these methods do not take into account that without constraints, some modules may learn similar functions, resulting in restricting the model's expressiveness and generalization capability of modular methods. In this paper, we propose the Contrastive Modules with Temporal Attention(CMTA) method to address these limitations. CMTA constrains the modules to be different from each other by contrastive learning and combining shared modules at a finer granularity than the task level with temporal attention, alleviating the negative transfer within the task and improving the generalization ability and the performance for multi-task RL. We conducted the experiment on Meta-World, a multi-task RL benchmark containing various robotics manipulation tasks. Experimental results show that CMTA outperforms learning each task individually for the first time and achieves substantial performance improvements over the baselines.
Siming Lan, Rui Zhang 0040, Qi Yi, Jiaming Guo, Shaohui Peng, Yunkai Gao 0001, Ruizhi Chen, Zidong Du, Xing Hu 0001, Xishan Zhang, Ling Li 0001, Yunji Chen
NeurIPS8
2023 Decompose a Task into Generalizable Subtasks in Multi-Agent Reinforcement Learning
abstract
In recent years, Multi-Agent Reinforcement Learning (MARL) techniques have made significant strides in achieving high asymptotic performance in single task. However, there has been limited exploration of model transferability across tasks. Training a model from scratch for each task can be time-consuming and expensive, especially for large-scale Multi-Agent Systems. Therefore, it is crucial to develop methods for generalizing the model across tasks. Considering that there exist task-independent subtasks across MARL tasks, a model that can decompose such subtasks from the source task could generalize to target tasks. However, ensuring true task-independence of subtasks poses a challenge. In this paper, we propose to \textbf{d}ecompose a \textbf{t}ask in\textbf{to} a series of \textbf{g}eneralizable \textbf{s}ubtasks (DT2GS), a novel framework that addresses this challenge by utilizing a scalable subtask encoder and an adaptive subtask semantic module. We show that these components endow subtasks with two properties critical for task-independence: avoiding overfitting to the source task and maintaining consistent yet scalable semantics across tasks. Empirical results demonstrate that DT2GS possesses sound zero-shot generalization capability across tasks, exhibits sufficient transferability, and outperforms existing methods in both multi-task and single-task problems.
Zikang Tian, Ruizhi Chen, Xing Hu 0001, Ling Li 0001, Rui Zhang 0040, Shaohui Peng, Jiaming Guo, Zidong Du, Qi Guo 0001, Yunji Chen
NeurIPS2
2023 Learning controllable elements oriented representations for reinforcement learning
Qi Yi, Rui Zhang 0040, Shaohui Peng, Jiaming Guo, Xing Hu 0001, Zidong Du, Qi Guo 0001, Ruizhi Chen, Ling Li 0001, Yunji Chen
Neurocomputing8
2023 Large-Scale Indoor Localization Solution for Pervasive Smartphones Using Corrected Acoustic Signals and Data-Driven PDR
abstract
With continuous and accelerated urbanization, a large number of location-based services (LBSs) have shifted from outdoor to indoor. The pervasive smartphone-based localization has been the subject of extensive work, including signals, algorithms, technologies, solutions, and applications. However, no single ubiquitous technology or solution exists for performing indoor positioning similar to the global navigation satellite system (GNSS) in the outdoor environment. The aim of this work is to develop a practical, precise, and economic smartphone-based localization solution. In order to address the challenges of utilizing the limited audible-band acoustic signal in pervasive smartphone localization, i.e., signal detection, correction, and evaluation, we present a low-cost anchor hardware, two-step signal detection method, data-driven pedestrian dead reckoning (PDR), and robust positioning algorithm. Moreover, we further propose acoustic measurement compensation approaches and measurement quality evaluation and control strategy (MQECS) to improve the performance of position estimation. Six phones, including Huawei Mate9, P9 Plus, OnePlus 6, Honor 8, Mi 10, and Google Pixel 3 are used to evaluate the localization performance in three typical wide-area indoor scenarios (i.e., convention center, parking lot, and dining-hall). The total testbed area is accumulated to more than 8800 square meters. The experimental results demonstrate that the proposed method achieves average positioning accuracies of 0.34 m (static) and 0.67 m (dynamic). In addition, the results show that the overall performance, repeatability, and stability are superior for different scenarios and devices.
Guangyi Guo, Ruizhi Chen, Zheng Li 0025, Xiaoguang Niu, Liang Chen 0007
IEEE Internet Things J.2
2023 Machine Learning for Time-of-Arrival Estimation With 5G Signals in Indoor Positioning
abstract
Location-based service in the indoor environment is playing a crucial role in different application scenarios. The introduction of technologies, such as ultradense network and massive multiple-input multiple-output enables fifth-generation (5G) cellular signals, as a new generation of cellular network signals, to show unique advantages in indoor positioning. This article describes 5G reference signal structures that can be used for navigation. A high-precision time-of-arrival estimation method based on 5G downlink signal is proposed that can be realized by edge computing. A software-defined receiver (SDR) based on machine learning to extract navigation observations from 5G signals is then developed. In simulation, the error sources of SDR in additive white gaussian noise channel and multipath channel were analyzed, and the possible ranging accuracy achieved by 5G signals in the developed SDR was evaluated. In field experiments, commercial 5G signals deployed by operators were collected, and the performance of SDR in practical applications was evaluated. The feasibility in practical applications of the proposed SDR is demonstrated, and high pseudorange measurement accuracy can be achieved.
Zhaoliang Liu, Liang Chen 0007, Xin Zhou 0006, Zhenhang Jiao, Guangyi Guo, Ruizhi Chen
IEEE Internet Things J.6
2023 iPos-5G: Indoor Positioning via Commercial 5G NR CSI
abstract
The fifth-generation (5G) networks have been massively deployed in commerce. The new features introduced by 5G networks are beneficial to wireless positioning. In this study, the performance of indoor positioning with commercial 5G new radio (NR) signals is investigated, and the channel state information (CSI) extracted from the downlink synchronization signal block is utilized. Considering the limited 5G NR base station (known as gNodeB) is hearable indoors, the fingerprint method is used, and an indoor positioning system termed iPos-5G is developed. The system consists of four components. First, a module of quality control is applied for CSI preprocessing. Second, an unsupervised deep-autoencoder network is utilized to reconstruct CSI features. Third, by supervised learning, a radial basis function is improved to optimize the probability model for similarity calculations. Finally, an amplitude-phase probability fusion function is proposed for positioning by weighting the coordinates of reference points. To verify the effectiveness of iPos-5G, indoor field tests are carried out in the scenarios of an office and a corridor. The test results show that iPos-5G achieves mean absolute errors of 2.14 and 2.81 m and standard deviation of the errors of 1.07 and 1.66 m, which outperforms the compared CSI fingerprint methods in terms of positioning accuracy and stability.
Yanlin Ruan, Liang Chen 0007, Xin Zhou 0006, Zhaoliang Liu, Guangyi Guo, Ruizhi Chen
IEEE Internet Things J.7
2023 SdoNet: Speed Odometry Network and Noise Adapter for Vehicle Integrated Navigation
abstract
The emerging applications of the Internet of Things (IoT), such as driverless cars, have an increasing need for precise vehicle positioning. Inertial navigation systems (INSs) became a possible component of autonomous driving systems due to low computational load, fast response, and high autonomy. However, error accumulation presents a significant challenge. Although nonholonomic constraints (NHCs) and odometry (ODO) have been demonstrated to improve INS, NHC is not always reliable, and ODO is often inaccessible in many applications. To address these issues, we propose a novel untethered pseudo-odometry, SdoNet, a convolutional neural network that estimates vehicle velocity from raw inertial measurement unit (IMU) observations to extend NHC as a 3-D velocity constraint without needing a hardware-wheeled ODO. To eliminate the influence of interference features on the accuracy of SdoNet, we improve the SdoNet by incorporating a residual module, attention mechanism, and soft threshold to guide the network to eliminate interference features. Moreover, a lightweight noise adapter network is proposed to adjust the constraint measurement noise covariance dynamically to apply the velocity constraint properly. The proposed approach is validated on the KITTI data set, demonstrating that SdoNet enhances the network’s learning ability and achieves robust and accurate velocity regression, especially in noisy IMU observations. The mean absolute speed regression error of SdoNet is lower than the two types of long short-term memory networks by 52.52% and 71.86%, respectively. Compared to the process using only NHC, the absolute translation error is reduced by approximately 44.00% after employing the pseudo-ODO velocity constraint and further reduced by around 11.31% after employing the noise adapter.
Xuan Wang 0015, Yuan Zhuang 0001, Xiaoxiang Cao, Qipeng Li, Yue Cao 0002, Ruizhi Chen
IEEE Internet Things J.7
2023 CWIWD-IPS: A Crowdsensing/Walk-Surveying Inertial/Wi-Fi Data-Driven Indoor Positioning System
abstract
Indoor positioning system plays a key role in location-based services since the widely used Global Navigation Satellite System (GNSS) is denied in indoor scenarios. Crowdsensing or walking-surveying based indoor positioning is proposed aiming at providing low-cost and high-efficient 3D location. This paper proposes a crowdsensing/walking-surveying 3D indoor positioning system by fusing the crowd-sensed inertial data and Wi-Fi fingerprinting samples using deep learning frameworks. A sine-wave-based step detector is used for pedestrian dead-reckoning (PDR) to generate original dense-trajectories. An enhanced optimization-based algorithm (Opt) and a smoothing-based algorithm (Smo) are proposed and evaluated to correct the original dense-trajectories into near-true dense-trajectories which are used to construct the inertial database and Wi-Fi radio map. A ResNet-based inertial neural-network and a BiLSTM-based Wi-Fi fingerprinting neural-network are trained on the constructed navigation database and combined by a Kalman filter to provide accurate and robust 3D localization performance. The realistic experimental results among complex indoor environments demonstrate that the proposed algorithms are proved to achieve a precise 3D indoor localization performance which is superior to several existing relative methods.
Yuan Wu 0006, Ruizhi Chen, Wenju Fu, Wei Li 0085, Haitao Zhou
IEEE Internet Things J.2
2022 Causality-driven Hierarchical Structure Discovery for Reinforcement Learning
abstract
Hierarchical reinforcement learning (HRL) has been proven to be effective for tasks with sparse rewards, for it can improve the agent's exploration efficiency by discovering high-quality hierarchical structures (e.g., subgoals or options). However, automatically discovering high-quality hierarchical structures is still a great challenge. Previous HRL methods can only find the hierarchical structures in simple environments, as they are mainly achieved through the randomness of agent's policies during exploration. In complicated environments, such a randomness-driven exploration paradigm can hardly discover high-quality hierarchical structures because of the low exploration efficiency. In this paper, we propose CDHRL, a causality-driven hierarchical reinforcement learning framework, to build high-quality hierarchical structures efficiently in complicated environments. The key insight is that the causalities among environment variables are naturally fit for modeling reachable subgoals and their dependencies; thus, the causality is suitable to be the guidance in building high-quality hierarchical structures. Roughly, we build the hierarchy of subgoals based on causality autonomously, and utilize the subgoal-based policies to unfold further causality efficiently. Therefore, CDHRL leverages a causality-driven discovery instead of a randomness-driven exploration for high-quality hierarchical structure construction. The results in two complex environments, 2D-Minecraft and Eden, show that CDHRL can discover high-quality hierarchical structures and significantly enhance exploration efficiency.
Shaohui Peng, Xing Hu 0001, Rui Zhang 0040, Ke Tang 0001, Jiaming Guo, Qi Yi, Ruizhi Chen, Xishan Zhang, Zidong Du, Ling Li 0001, Qi Guo 0001, Yunji Chen
NeurIPS7
2022 Intrinsic mixed-integer polycubes for hexahedral meshing
Manish Mandad, Ruizhi Chen, David Bommes, Marcel Campen
Comput. Aided Geom. Des.2
2022 H-WPS: Hybrid Wireless Positioning System Using an Enhanced Wi-Fi FTM/RSSI/MEMS Sensors Integration Approach
abstract
Indoor wireless localization toward the next generation Wi-Fi access point has attracted considerable attention due to the presentation of the state-of-art Wi-Fi fine time measurement (FTM) protocol. In order to improve the autonomy, accuracy, and universality of wireless positioning based on the Internet of Things (IoT) terminals, this article proposes a hybrid wireless positioning system which contains the integration of Wi-Fi FTM, crowdsourced received signal strength indicator (RSSI) fingerprinting and micro-electro-mechanical-system (MEMS) sensors (H-WPS). A light-weight pedestrian aimed inertial navigation system (PINS) is proposed, which contains multilevel constraints and a global optimization model in order to eliminate the cumulative error caused by INS update. A deep-learning-based Wi-Fi fingerprinting database generation framework is developed for crowdsourced trajectories evaluation and selection. In addition, three different multisource integration models are applied to fuse the information of PINS, Wi-Fi FTM and RSSI fingerprinting, and calibrate the Wi-Fi ranging bias in real time, which is further enhanced by a novel misclosure check and the multilayer perceptron contained signal quality evaluation strategy. The comprehensive experiments demonstrate that the proposed H-WPS achieves much more precise and universal indoor positioning performance compared with the single location source, and meter-level localization precision can be realized in the Wi-Fi FTM-covered indoor scenes.
Yue Yu 0003, Ruizhi Chen, Liang Chen 0007, Wei Li 0085, Yuan Wu 0006, Haitao Zhou
IEEE Internet Things J.2
2022 Carrier Phase Ranging for Indoor Positioning With 5G NR Signals
abstract
Indoor positioning is one of the core technologies of Internet of Things (IoT) and artificial intelligence (AI) and is expected to play a significant role in the upcoming era of AI. However, affected by the complexity of indoor environments, it is still highly challenging to achieve continuous and reliable indoor positioning. Currently, 5G cellular networks are being deployed worldwide, the new technologies of which have brought the approaches for improving the performance of wireless indoor positioning. In this article, we investigate the indoor positioning under the 5G new radio (NR), which has been standardized and being commercially operated in massive markets. Specifically, a solution is proposed and a software-defined radio (SDR) receiver is developed for indoor positioning. With our SDR indoor positioning system, the 5G NR signals are first sampled by universal software radio peripheral (USRP), and then, coarse synchronization is achieved via detecting the start of the synchronization signal block (SSB). Then, with the assistance of the pilots transmitted on the physical broadcasting channel (PBCH), multipath acquisition and delay tracking are sequentially carried out to estimate the Time of Arrival (ToA) of received signals. Furthermore, to improve the ToA ranging accuracy, the carrier phase of the first arrived path is estimated. Finally, to quantify the accuracy of our ToA estimation method, indoor field tests are carried out in an office environment, where a 5G NR base station (known as gNB) is installed for commercial use. Our test results show that, in the static test scenarios, the ToA accuracy measured by the 1-$\sigma $error interval is about 0.5 m, while in the pedestrian mobile environment, the probability of range accuracy within 0.8 m is 95%.
Liang Chen 0007, Xin Zhou 0006, Lie-Liang Yang, Ruizhi Chen
IEEE Internet Things J.5
2022 A Robust Integration Platform of Wi-Fi RTT, RSS Signal, and MEMS-IMU for Locating Commercial Smartphone Indoors
abstract
As the cornerstone of indoor location-based services (ILBSs), the smartphone-based real-time locating and tracking technologies are now becoming the key for implementing seamless indoor/outdoor location-based applications. The Wi-Fi received signal strength (RSS)-based positioning system is widely used because of the widespread deployment of Wi-Fi access points in the indoor environment. Correspondingly, the positioning performance of the RSS-based method is limited significantly by the complex and time-varying indoor environment. Contrary to the conventional RSS-based techniques, based on the introduction of a two-way ranging approach in the IEEE 802.11-REVmc2protocol, the Wi-Fi round trip time (RTT) ranging technique provides high-resolution and low-latency ranging observation on smartphones. In this work, a robust integration platform and related positioning algorithms of tightly coupled heterogeneous observables from Wi-Fi and MEMS-IMU are developed for smartphone positioning. The proposed framework optimizes the relative and absolute positioning observables in the integration process and improves the accuracy and stability as compared to the solutions, which are based on a single positioning technology. Moreover, the OQECS is established to evaluate the quality of each observation in real time before feeding the data to the adaptive filter. The experimental results demonstrate that the proposed platform achieves improvement in accuracy and robustness in both real-time tests and simulation tests performed using the polluted data. The average positioning accuracy is 0.572 m, which is 20.22% better than the results obtained from a standard EKF method.
Guangyi Guo, Ruizhi Chen, Feng Ye 0003, Zuoya Liu, Lixiong Huang, Zheng Li 0025
IEEE Internet Things J.2
2022 RSS-Based Visible Light Positioning Using Nonlinear Optimization
abstract
In recent years, indoor positioning has drawn intensive attention for both pedestrian and mobile robot applications. Among various indoor positioning technologies, visible light positioning has many advantages due to its high localization accuracy, high bandwidth, energy efficiency, long lifetime, and cost efficiency. For postprocessing or semi-real-time applications, researchers often use smoothers to improve location accuracy. However, smoothers are always local estimators and lack integrity when calculating locations. To globally optimize the positioning results and further improve the accuracy, we propose a nonlinear optimization model based on the idea of graph optimization. Innovatively, the model adds the acceleration as a constraint to become one part of the residuals and regularize the trajectory. We design a signal-to-noise ratio-based weighting strategy to suppress the outliers and better assess the errors. Moreover, we design a loop constraint to further improve the positioning accuracy. The experimental results show that our proposed model significantly improves the accuracy by 71%, which is suitable for indoor positioning.
Xiao Sun 0009, Yuan Zhuang 0001, Jianzhu Huai, Luchi Hua, Dong Chen 0041, You Li 0001, Yue Cao 0002, Ruizhi Chen
IEEE Internet Things J.8
2022 A Multimagnetometer Array and Inner IMU-Based Capsule Endoscope Positioning System
abstract
The wireless capsule endoscope (robot) has become more extensively used due to its comprehensive detection and patient-friendly experience. However, to provide better diagnostic information to medical staff, there is an urgent need for high-accuracy position information of capsule endoscopes during their working inside the human body. In this article, a capsule endoscopy positioning system using a magnetic sensor array is designed. It has two advantages. 1) Most of the existing magnetic positioning method needs to initialize the magnetic moment accurately, which is difficult to meet in practical applications. To solve this issue, this article proposes a method to determine the magnetic moment direction based on an inertial measurement unit. The proposed method can accurately estimate the direction of the magnetic moment even when the roll angle is singular. 2) This article proposes a nonlinear least-squares algorithm for capsule magnetic positioning based on the three-axis magnetometer observation. The algorithm is more robust than the Levenberg–Marquardt (LM) method that is widely used in capsule endoscopy positioning. Furthermore, its computation speed is over 100 times faster than the LM method, which successfully meets the real-time requirements. In this research, a three-axis mechanical platform and a six-axis robot arm are used to build a capsule magnetic positioning evaluation system. Preliminary results show the accuracy (RMS) of the proposed capsule endoscope positioning algorithm was better than 6 mm.
Peng Zhang 0042, Yan Xu 0025, Ruizhi Chen, Weiguo Dong, You Li 0001, Rong Yu 0002, Mingyue Dong, Zhengru Liu, Yuan Zhuang 0001, Jian Kuang 0004
IEEE Internet Things J.3
2022 Bluetooth Localization Technology: Principles, Applications, and Future Trends
abstract
The rapid development of the Bluetooth technology offers a possible solution for indoor localization scenarios. Compared with other indoor localization technologies, such as vision, light detection and ranging, ultrawide band, etc., Bluetooth has been characterized by low cost, easy deployment, low energy consumption, and potentially high localization accuracy, which enable itself to be a competitive technology in indoor location-based services, the Internet of Things, and many other fields. In this article, we first present a comprehensive survey of Bluetooth localization technology, including the measurements for localization, working principles, and method comparison. We highlight the learning-based methods and integrated localization methods. Then, we review the applications and existing commercial solutions, revealing the possible directions for the industrialization of Bluetooth localization. Finally, this article proposes several open issues of Bluetooth localization (e.g., multichannel difference, multipath, co-channel interference, and device heterogeneity) and projects several future trends.
Yuan Zhuang 0001, Jianzhu Huai, You Li 0001, Liang Chen 0007, Ruizhi Chen
IEEE Internet Things J.6
2022 TSA-SCC: Text Semantic-Aware Screen Content Coding With Ultra Low Bitrate
abstract
Due to the rapid growth of web conferences, remote screen sharing, and online games, screen content has become an important type of internet media information and over 90% of online media interactions are screen based. Meanwhile, as the main component in the screen content, textual information averagely takes up over 40% of the whole image on various commonly used screen content datasets. However, it is difficult to compress the textual information by using the traditional coding schemes as HEVC, which assumes strong spatial and temporal correlations within the image/video. State-of-the-art screen content coding (SCC) standard as HEVC-SCC still adopts a block-based coding framework and does not consider the text semantics for compression, thus inevitably blurring texts at a lower bitrate. In this paper, we propose a general text semantic-aware screen content coding scheme (TSA-SCC) for ultra low bitrate setting. This method detects the abrupt picture in a screen content video (or image), recognizes textual information (including word, position, font type, font size and font color) in the abrupt picture based on neural networks, and encodes texts with text coding tools. The other pictures as well as the background image after removing texts from the abrupt picture via inpainting, are encoded with HEVC-SCC. Compared with HEVC-SCC, the proposed method TSA-SCC reduces bitrate by up to 3× at a similar compression quality. Moreover, TSA-SCC achieves much better visual quality with less bitrate consumption when encoding the screen content video/image at ultra low bitrates.
Ling Li 0001, Ruizhi Chen, Haochen Li 0002, Guo Lu, Limin Cheng
IEEE Trans. Image Process.4
2022 Inertial Sensing Meets Machine Learning: Opportunity or Challenge?
abstract
The inertial navigation system (INS) has been widely used to provide self-contained and continuous motion estimation in intelligent transportation systems. Recently, the emergence of chip-level inertial sensors has expanded the relevant applications from positioning, navigation, and mobile mapping to location-based services, unmanned systems, and transportation big data. Meanwhile, benefit from the emergence of big data and the improvement of algorithms and computing power, machine learning (ML) has become a consensus tool that has been successfully applied in various fields. This article reviews the research on using ML technology to enhance inertial sensing from various aspects, including sensor design and selection, calibration and error modeling, navigation and motion-sensing algorithms, multi-sensor information fusion, system evaluation, and practical application. It summarizes the state of the art, advantages, and challenges on each aspect, and points out future research directions.
You Li 0001, Ruizhi Chen, Xiaoji Niu, Yuan Zhuang 0001, Zhouzheng Gao, Xin Hu 0006, Naser El-Sheimy
IEEE Trans. Intell. Transp. Syst.2
2022 Instance-Aware Semantic Segmentation of Road Furniture in Mobile Laser Scanning Data
abstract
In this paper, we present an improved framework for the instance-aware semantic segmentation of road furniture in mobile laser scanning data. In our framework, we first detect road furniture from mobile laser scanning point clouds. Then we decompose the detected pieces of road furniture into poles and their attached components, and extract the instance information of the components with different features. Most importantly, we classify the components into different categories by combining a classifier and a probabilistic graphic model named DenseCRF, which is the major contribution of this paper. For the classification of the components using DenseCRF, the unary potentials and the pairwise potentials are first obtained. The unary potentials are obtained from the classifier which takes the instance information of components as the input. The pairwise potentials are calculated considering contextual relations between components. By utilising DenseCRF, the contextual consistency of components is preserved, and the performance is significantly improved compared to our previous work. We collect three datasets to test our framework, and compare the classification performances of six different classifiers with and without DenseCRF. The combination of random forest with DenseCRF outperforms the other methods and achieves high overall accuracies of 83.7%, 96.4% and 95.3% in these three datasets. Experimental results demonstrate that our framework reliably assigns both semantic information and instance information for mobile laser scanning point clouds of road furniture.
Fashuai Li, Zhize Zhou, Ruizhi Chen, Matti Lehtomäki, Sander Oude Elberink, George Vosselman, Juha Hyyppä, Yuwei Chen 0005, Antero Kukko
IEEE Trans. Intell. Transp. Syst.4
2021 Model-Based 3D Hand Reconstruction via Self-Supervised Learning
abstract
Reconstructing a 3D hand from a single-view RGB image is challenging due to various hand configurations and depth ambiguity. To reliably reconstruct a 3D hand from a monocular image, most state-of-the-art methods heavily rely on 3D annotations at the training stage, but obtaining 3D annotations is expensive. To alleviate reliance on labeled training data, we propose S2HAND, a self-supervised 3D hand reconstruction network that can jointly estimate pose, shape, texture, and the camera viewpoint. Specifically, we obtain geometric cues from the input image through easily accessible 2D detected keypoints. To learn an accurate hand reconstruction model from these noisy geometric cues, we utilize the consistency between 2D and 3D representations and propose a set of novel losses to rationalize outputs of the neural network. For the first time, we demonstrate the feasibility of training an accurate 3D hand reconstruction network without relying on manual annotations. Our experiments show that the proposed self-supervised method achieves comparable performance with recent fully-supervised methods. The code is available at https://github.com/TerenceCYJ/S2HAND.
Yujin Chen, Zhigang Tu 0001, Linchao Bao, Ying Zhang 0021, Xuefei Zhe, Ruizhi Chen, Junsong Yuan 0001
CVPR7
2021 Interest point detection from multi-beam light detection and ranging point cloud using unsupervised convolutional neural network
abstract
Abstract Interest point detection plays an important role in many computer vision applications. This work is motivated by the light detection and ranging odometry task in autonomous driving. Existing methods are not capable of detecting enough interest points in unstructured scenarios where there are little constructions or trees around, and correspondingly light detection and ranging odometry will fail to continuous localisation. An interest point detector is proposed for detecting interest points from multi‐beam light detection and ranging point cloud using unsupervised convolutional neural network. The point cloud is projected into a two‐dimensional structured data according to the scanning geometry. Then the convolutional neural network filters trained in an unsupervised manner are used to generate a local feature map with the two‐dimensional structured data as input. Finally, interest points are obtained by extracting the grids that have significant differences with their neighbour grids. Based on an odometry benchmark, the experiments show that the proposed interest point detector can capture more local details, which contributes to more than 16% error decrease in point cloud registration in highway scenes.
Deyu Yin, Jingbin Liu, Xinlian Liang, Yunsheng Wang 0002, Shoubin Chen, Jyri Maanpää, Juha Hyyppä, Ruizhi Chen
IET Image Process.9
2021 Toward Location-Enabled IoT (LE-IoT): IoT Positioning Techniques, Error Sources, and Error Mitigation
abstract
Localization techniques are becoming key to add location context to the Internet-of-Things (IoT) data without human perception and intervention. Meanwhile, the newly emerged low-power wide-area network (LPWAN) and 5G technologies have become strong candidates for mass-market localization applications. However, various error sources have limited localization performance by using such IoT signals. This article reviews the IoT localization system through the following sequence: IoT localization system review, localization data sources, localization algorithms, localization error sources and mitigation, and localization performance evaluation. Compared to the related surveys, this article has a more comprehensive and state-of-the-art review on IoT localization methods, an original review on IoT localization error sources and mitigation, an original review on IoT localization performance evaluation, and a more comprehensive review of IoT localization applications, opportunities, and challenges. Thus, this survey provides comprehensive guidance for peers who are interested in enabling localization ability in the existing IoT systems, using IoT systems for localization, or integrating IoT signals with the existing localization sensors.
You Li 0001, Yuan Zhuang 0001, Xin Hu 0006, Zhouzheng Gao, Jia Hu 0001, Long Chen 0005, Zhe He 0002, Ling Pei, Kejie Chen, Maosong Wang, Xiaoji Niu, Ruizhi Chen, John S. Thompson, Fadhel M. Ghannouchi, Naser El-Sheimy
IEEE Internet Things J.12
2021 A Novel 3-D Indoor Localization Algorithm Based on BLE and Multiple Sensors
abstract
Indoor wireless localization using Bluetooth low energy (BLE) beacons has attracted considerable attention due to its extensive distribution and low cost properties. This article proposes a novel 3-D indoor localization algorithm which uses the combination of BLE and multiple sensors (3D-LBMS). The inertial navigation system (INS) and pedestrian dead reckoning (PDR) mechanizations are combined for accurate heading and speed estimation, which contains a multilevel constraints-based quasistatic magnetic field (QSMF) detection algorithm. In addition, dynamic-time-warping (DTW)-based BLE landmark detection algorithm is proposed to provide absolute 3-D location reference to multiple sensors-based positioning method, and the detected BLE landmark points are also used to calibrate the parameter of step-length calculation. Finally, the adaptive unscented Kalman filter (AUKF) is applied to fuse the results of INS/PDR mechanizations, QSMF and locations of detected BLE landmarks to achieve accurate and concrete multisource-based 3-D indoor localization performance. The experimental results show that the proposed 3-D-LBMS is proved to achieve meterlevel 2-D positioning accuracy and submeter level 3-D altitude estimation accuracy in typical indoor environments.
Yue Yu 0003, Ruizhi Chen, Liang Chen 0007, Xingyu Zheng, Dewen Wu, Wei Li 0085, Yuan Wu 0006
IEEE Internet Things J.2
2021 Joint Hand-Object 3D Reconstruction From a Single Image With Cross-Branch Feature Fusion
abstract
Accurate 3D reconstruction of the hand and object shape from a hand-object image is important for understanding human-object interaction as well as human daily activities. Different from bare hand pose estimation, hand-object interaction poses a strong constraint on both the hand and its manipulated object, which suggests that hand configuration may be crucial contextual information for the object, and vice versa. However, current approaches address this task by training a two-branch network to reconstruct the hand and object separately with little communication between the two branches. In this work, we propose to consider hand and object jointly in feature space and explore the reciprocity of the two branches. We extensively investigate cross-branch feature fusion architectures with MLP or LSTM units. Among the investigated architectures, a variant with LSTM units that enhances object feature with hand feature shows the best performance gain. Moreover, we employ an auxiliary depth estimation module to augment the input RGB image with the estimated depth map, which further improves the reconstruction accuracy. Experiments conducted on public datasets demonstrate that our approach significantly outperforms existing approaches in terms of the reconstruction accuracy of objects.
Yujin Chen, Zhigang Tu 0001, Ruizhi Chen, Linchao Bao, Zhengyou Zhang, Junsong Yuan 0001
IEEE Trans. Image Process.4
2020 A Novel Calibration Method between a Camera and a 3D LiDAR with Infrared Images
abstract
Fusions of LiDARs (light detection and ranging) and cameras have been effectively and widely employed in the communities of autonomous vehicles, virtual reality and mobile mapping systems (MMS) for different purposes, such as localization, high definition map or simultaneous location and mapping. However, the extrinsic calibration between a camera and a 3D LiDAR is a fundamental prerequisite to guarantee its performance. Some previous methods are inaccurate, have calibration error that is several times the beam divergence, and often require special calibration objects, thereby limiting their ubiquitous use for calibration. To overcome these shortcomings, we propose a novel and high-accuracy method for the extrinsic calibration between a camera and a 3D LiDAR. Our approach relies on the infrared images from a camera with an infrared filter, and the 2D-3D corresponding points in a scene with the corners of a wall can be extracted to calculate the six extrinsic parameters. Experiments using the Velodyne VLP-16 sensor show that the method can achieve an extrinsic accuracy at the level of the beam divergence, which is fully analyzed and validated from two different aspects. Therefore, the calibration method in this paper is highly accurate, effective and does not require special complicated calibration objects; thus, it meets the requirements of practical applications.
Shoubin Chen, Jingbin Liu, Xinlian Liang, Juha Hyyppä, Ruizhi Chen
ICRA6
2020 AtLAS: An Activity-Based Indoor Localization and Semantic Labeling Mechanism for Residences
abstract
Currently, indoor localization technology and indoor location-based services are becoming increasingly important in the area of mobile and ubiquitous computing. However, the design of an indoor location-based system confronts two challenges: 1) achieving high-precision location recognition and 2) identifying what indoor objects actually are (which is called semantic labeling). In this article, we propose AtLAS, an activity-based indoor localization and semantic labeling mechanism. The key idea is that some objects in an indoor environment, such as doors and toilets, determine predictable human behaviors in small areas, which can be reflected in unique sensor readings. AtLAS leverages this idea to determine a user's accurate location by identifying users' activities. Furthermore, we leverage the topological structure of indoor objects to mine the semantic knowledge and label the objects through gained knowledge automatically. To the best of our knowledge, AtLAS is the first attempt to build a system that leverages users' activities to conduct a high-precision indoor localization and semantic labeling system for the case of residences. The experimental results show that AtLAS can achieve a median localization accuracy of 0.57 m, and the system can localize the landmarks with a median accuracy of 0.43 m on average without 5% worst errors. AtLAS can label the objects semantically with a 5.7% false-positive rate and a 5.8% false-negative rate on average.
Xiaoguang Niu, Luyao Xie, Jiawei Wang 0020, Haiming Chen 0002, Ruizhi Chen
IEEE Internet Things J.6
2020 Precise 3-D Indoor Localization Based on Wi-Fi FTM and Built-In Sensors
abstract
More and more applications of location-based services lead to the development of indoor positioning technology. As a part of the Internet-of-Things ecosystem, most existing indoor positioning algorithms are applied to specific situations, e.g., pedestrian navigation and target detection. To meet the high-precision indoor localization requirement, IEEE 802.11 included the Wi-Fi fine-time measurement (FTM) protocol in 2016, which provides a novel approach for Wi-Fi ranging between the mobile terminal and Wi-Fi access point (AP). This article proposes a precise 3-D indoor localization algorithm based on Wi-Fi FTM and smartphone built-in sensors (3D-WFBS). The adaptive extended Kalman filter (AEKF) is used to estimate the pedestrian's real-time heading and walking speed, and the received signal strength indication and round-trip time collected from Wi-Fi APs are combined for proximity detection and providing more accurate ranging results. In addition, the unscented particle filter is applied to fuse the results of AEKF, proximity detection, and Wi-Fi ranging. The experimental results show that compared with the existing dead reckoning method and the other fusion methods, the proposed 3D-WFBS algorithm is proved to achieve meter-level indoor positioning accuracy in typical indoor scenes.
Yue Yu 0003, Ruizhi Chen, Liang Chen 0007, Wei Li 0085, Yuan Wu 0006, Haitao Zhou
IEEE Internet Things J.2
2020 Using Radar Signatures to Classify Bird Flight Modes Between Flapping and Gliding
abstract
In this work, we find that radar signatures registered by wingbeats can work as a chronophotograph in radar version to record bird flight modes. When a bird flies, its wings and body compose a corner reflector, and the incident radiation upon either face of this wingbeat corner reflector impinges onto the other face and is reflected toward the illuminator, which enhances the bird signal intensity by 0-10 dB. When a bird flaps its wings, the corner works at a right angle, and the wingbeat corner reflector strongly modulates the radar signal; therefore, the contribution is strong and stably approaches 10 dB. During gliding, the wingbeat corner reflector fades away, the modulation effect significantly decreases, and the contribution approaches 0 dB. This difference can assist in classifying bird flight modes between flapping and gliding over a long sampling period. Both the Ku-band data from a duck in an anechoic chamber and the Ku-band radar data from a pigeon in an outside environment support the ability of this signature to classify gliding and flapping modes. We present flight mode transitions between flapping and gliding extracted from the radar echoes of small and large oncoming and outgoing birds.
Jiangkun Gong, Jun Yan 0006, DeRen Li, Ruizhi Chen
IEEE Geosci. Remote. Sens. Lett.4
2020 Analyzing and Accelerating the Bottlenecks of Training Deep SNNs With Backpropagation
abstract
Spiking neural networks (SNNs) with the event-driven manner of transmitting spikes consume ultra-low power on neuromorphic chips. However, training deep SNNs is still challenging compared to convolutional neural networks (CNNs). The SNN training algorithms have not achieved the same performance as CNNs. In this letter, we aim to understand the intrinsic limitations of SNN training to design better algorithms. First, the pros and cons of typical SNN training algorithms are analyzed. Then it is found that the spatiotemporal backpropagation algorithm (STBP) has potential in training deep SNNs due to its simplicity and fast convergence. Later, the main bottlenecks of the STBP algorithm are analyzed, and three conditions for training deep SNNs with the STBP algorithm are derived. By analyzing the connection between CNNs and SNNs, we propose a weight initialization algorithm to satisfy the three conditions. Moreover, we propose an error minimization method and a modified loss function to further improve the training performance. Experimental results show that the proposed method achieves 91.53% accuracy on the CIFAR10 data set with 1% accuracy increase over the STBP algorithm and decreases the training epochs on the MNIST data set to 15 epochs (over 13 times speed-up compared to the STBP algorithm). The proposed method also decreases classification latency by over 25 times compared to the CNN-SNN conversion algorithms. In addition, the proposed method works robustly for very deep SNNs, while the STBP algorithm fails in a 19-layer SNN.
Ruizhi Chen, Ling Li 0001
Neural Comput.1
2019 SO-HandNet: Self-Organizing Network for 3D Hand Pose Estimation With Semi-Supervised Learning
abstract
3D hand pose estimation has made significant progress recently, where Convolutional Neural Networks (CNNs) play a critical role. However, most of the existing CNN-based hand pose estimation methods depend much on the training set, while labeling 3D hand pose on training data is laborious and time-consuming. Inspired by the point cloud autoencoder presented in self-organizing network (SO-Net), our proposed SO-HandNet aims at making use of the unannotated data to obtain accurate 3D hand pose estimation in a semi-supervised manner. We exploit hand feature encoder (HFE) to extract multi-level features from hand point cloud and then fuse them to regress 3D hand pose by a hand pose estimator (HPE). We design a hand feature decoder (HFD) to recover the input point cloud from the encoded feature. Since the HFE and the HFD can be trained without 3D hand pose annotation, the proposed method is able to make the best of unannotated data during the training phase. Experiments on four challenging benchmark datasets validate that our proposed SO-HandNet can achieve superior performance for 3D hand pose estimation via semi-supervised learning.
Yujin Chen, Zhigang Tu 0001, Liuhao Ge, Dejun Zhang, Ruizhi Chen, Junsong Yuan 0001
ICCV5
2018 FBNA: A Fully Binarized Neural Network Accelerator
abstract
In recent researches, binarized neural network (BNN) has been proposed to address the massive computations and large memory footprint problem of the convolutional neural network (CNN). Several works have designed specific BNN accelerators and showed very promising results. Nevertheless, only part of the neural network is binarized in their architecture and the benefits of binary operations were not fully exploited. In this work, we propose the first fully binarized convolutional neural network accelerator (FBNA) architecture, in which all convolutional operations are binarized and unified, even including the first layer and padding. The fully unified architecture provides more resource, parallelism and scalability optimization opportunities. Compared with the state-of-the-art BNN accelerator, our evaluation results show 3.1x performance, 5.4x resource efficiency and 4.9x power efficiency on CIFAR-10.
Ruizhi Chen, Pin Li, Shaolin Xie
FPL3
2018 Low Latency Spiking ConvNets with Restricted Output Training and False Spike Inhibition
abstract
Deep convolutional neural networks (ConvNets) have achieved the state-of-the-art performance on many real-world applications. However, significant computation and storage demands are required by ConvNets. Spiking neural networks (SNNs), with sparsely activated neurons and event-driven computations, show great potential to take advantage of the ultra- low power spike-based hardware architectures. Yet, training SNN with similar accuracy as ConvNets is difficult. Recent researchers have demonstrated the work of converting ConvNets to SNNs (CNN-SNN conversion) with similar accuracy. However, the energy-efficiency of the converted SNNs is impaired by the increased classification latency. In this paper, we focus on optimizing the classification latency of the converted SNNs. First, we propose a restricted output training method to normalize the converted weights dynamically in the CNN-SNN training phase. Second, false spikes are identified and the false spike inhibition theory is derived to speedup the convergence of the classification process. Third, we propose a temporal max pooling method to approximate the max pooling operation in ConvNets without accuracy loss. The evaluation shows that the converted SNNs converge in about 30 time-steps and achieve the best classification accuracy of 94% on CIFAR -10 dataset.
Ruizhi Chen, Shaolin Xie, Pin Li
IJCNN1
2018 Fast and Efficient Deep Sparse Multi-Strength Spiking Neural Networks with Dynamic Pruning
abstract
Deep convolutional neural networks (CNNs) have shown state-of-the-art accuracy for various computer vision and speech tasks. However, CNNs are computation-intensive and energy-inefficient which are difficult to be deployed in real-time systems. Event-driven Spiking Neural Networks (SNNs) are extremely power efficient, which provides an alternative for ultra-low power applications. But effective training methods for SNN are still lacking. Due to its spatio-temporal feature of SNN, conventional training method for CNN can not be employed in SNN. To address this problem, some researchers proposed to convert the corresponding weights of trained CNNs into the synapse weights of SNNs (CNNs-SNNs). Nevertheless, limited by the [0, 1] constraints on the SNN neuron outputs, the accuracy of the converted SNNs is impaired. Besides, as the SNN network becomes deeper, the convergence speed of SNN inference are unacceptably slow. In this work, we proposed an innovative deep multi-strength SNN (M-SNN) structure which relaxes the restriction of the neuron output spike strength while the event-driven feature for low-power implementations is maintained. Using this architecture, large scale SNN can be converted from CNN with comparable accuracy and fast inference speed. The evaluation results show 3.7 × convergence speedup. Moreover, with multi-strength spike, aggressive pruning strategies can be applied to reduce the computational operations by almost 85% while maintaining the same accuracy.
Ruizhi Chen, Shaolin Xie, Pin Li
IJCNN1
2018 A Localization Database Establishment Method Based on Crowdsourcing Inertial Sensor Data and Quality Assessment Criteria
abstract
Aimed at the challenge of generating indoor localization databases with daily life crowdsourcing-based inertial sensor data, this paper proposes an anchor point-based forward–backward smoothing method to obtain reliable localization solutions. More importantly, a quantitative framework is proposed to evaluate the quality of smartphone-based inertial sensor data automatically without user intervention. Through this framework, the reliability of each inertial sensor data can be evaluated and sorted. Tests with multiple people and multiple smartphones in a public office building and a shopping mall illustrate that the proposed method can provide a WiFi fingerprinting database that has similar accuracy to that generated by a supervised map-aided database-generation method. Therefore, the proposed method and framework can guide the promotion of crowdsourcing-based Internet of Things applications in the context of big data.
Peng Zhang 0042, Ruizhi Chen, You Li 0001, Xiaoji Niu, Lei Wang 0045, Ming Li 0037, Yuanjin Pan
IEEE Internet Things J.2
2013 Sound positioning using a small-scale linear microphone array
abstract
Microphone arrays, also known as acoustic antennas, have been extensively used for sound localization. Small-scale microphone arrays have especially been used in teleconferences and game consoles due to their small dimension and easy deployment. In this article, we present an approach to locating a sound source using a small linear microphone array. We describe the fundamentals of linear microphone arrays and analyze the impact of geometry in terms of positioning accuracy using the dilution of precision (DOP) concept. The generalized cross-correlation (GCC) based on the phase transform (PHAT) weighting function is used to estimate the time difference of arrivals in a microphone array. Given the time differences, we use both closed-form and iterative optimization solutions to calculate the coordinates of the sound source. In order to evaluate the performances of the solutions applied in this paper, simulations and field tests were conducted. Simulation results show that the closed-form algorithm gives a positioning error of less than 5 cm in a 10-by-10 meter room when the geometry of a microphone array is good and the signal to noise ratio (SNR) is high. Linear small microphone arrays have lower performances compared to a non-linear distributed array. When the scale of a linear array is reduced, the positioning accuracy decreases dramatically. With a small linear array, the iterative optimization algorithm gives much better performance compared to the closed-form algorithm. Field tests were conducted in an 11-by-5.6 meter room using a linear array with a length of 0.23 meters. Positioning results show an average error of 0.25 meters along the axis parallel to the linear array and 0.53 meters error along the axis which is perpendicular to the linear array.
Ling Pei, Liang Chen 0007, Robert Guinness, Jingbin Liu, Heidi Kuusniemi, Yuwei Chen 0005, Ruizhi Chen, Stefan Söderholm
IPIN7
2013 Electromyography-Based Locomotion Pattern Recognition and Personal Positioning Toward Improved Context-Awareness Applications
abstract
Personal positioning has been playing an important role in context awareness and navigation. Pedestrian dead reckoning (PDR) solution is a positioning technology used where the global positioning system (GPS) signal is not available or its signal is mightily attenuated or reflected by constructions nearby, such as inside the buildings or in GPS degraded areas such as urban city, basement. A traditional PDR solution employs a multisensor unit (integrating accelerometer, gyroscope, digital compass, barometer, etc.) to detect step occurrences, as well as to estimate the stride length. In our pilot research, we proposed a novel electromyography (EMG)-based method to fulfill that task and obtained satisfying PDR results. In this paper, a further attempt is made to investigate the feasibility of using EMG sensors in sensing muscle activities to detect the corresponding locomotion patterns, and as a result, a new approach, which recognizes different locomotion patterns using EMG signals and constructs stride length models according to the recognition results, is then proposed to improve the positioning accuracy and robustness of the EMG-based PDR solution by adapting the stride length model into different locomotion patterns. The experimental results demonstrate that EMG-based pattern recognition of four motions (walking, running, walking upstairs, walking downstairs) achieve an error rate of less than 2%. Combined with locomotion pattern recognition, the proposed EMG-based PDR solution yield a position deviation of less than 5 m within the whole distance of 404 m in a simulated indoor/outdoor field test. The proposed method is proven to be effective and practical in sensing context information, including both the user's activities and locations.
Xiang Chen 0004, Ruizhi Chen, Yuwei Chen 0005, Xu Zhang 0002
IEEE Trans. Syst. Man Cybern. Syst.3
2012 Utilizing pulsed pseudolites and high-sensitivity GNSS for ubiquitous outdoor/indoor satellite navigation
abstract
Pseudolites provide a means for bridging the gap between outdoors and indoors when GNSS (Global Navigation Satellite System) positioning is concerned. This paper presents a ubiquitous outdoor/indoor GNSS navigation platform that utilizes GPS (Global Positioning System), GLONASS, and pulsed pseudolite (PL) signals for seamless positioning. When a pseudolite signal is pulsed to efficiently transmit the GNSS-like signal only at particular time instants, interference problems between the terrestrial pseudo-satellite signals and the space-based satellite signals are significantly reduced. Pulsed pseudolites are strategically placed indoors at known locations at the ends of building corridors to assist high-sensitivity GPS and GLONASS positioning. A particle filtering solution is implemented to combine the high-sensitivity GNSS and the pseudolite proximity information in order to provide a seamless outdoor/indoor positioning platform. As demonstrated with real-life experiments, pseudolites provide a convenient navigation aid indoors for a GNSS receiver without the need for using additional hardware.
Heidi Kuusniemi, Mohammad Zahidul H. Bhuiyan, Marten Strom, Stefan Söderholm, Timo Jokitalo, Liang Chen 0007, Ruizhi Chen
IPIN7
2011 Wearable electromyography sensor based outdoor-indoor seamless pedestrian navigation using motion recognition method
abstract
Navigation and position applications are now becoming standard built-in features in a smart phone. However, locating a mobile user in GNSS unfriendly and denied environments such as urban canyons and indoor environments ubiquitously is still a challenging task. Several self-contained sensors, such as accelerometer, digital compass, gyroscope and barometer, have been adapted as assistance augmentation technologies to a GPS receiver to make a seamless outdoor-indoor pedestrian navigation system. Since the indoor environment is more complex than an open-sky environment, such GNSS signal-degraded areas are typically also contaminated with disturbance sources that affect sensor measurements, a digital compass can be disturbed significantly by e.g. an elevator that bears magnetic perturbance. And a ventilation facility may cause inconsistencies in the barometer's measurements; not to mention that the indoor surrounding attenuates or blocks the GNSS signal. In this paper, a novel outdoor-indoor seamless solution for pedestrian navigation is introduced, which is based on Electromyography (EMG) sensors. The EMG sensor measures the electrical potentials generated by muscle contractions of human body. Therefore it is immune against the environment disturbance; moreover, it has potential capability to exploit the health situation of the pedestrian, since the EMG sensor has been applied on the biomedical field for decades. In the paper, five different motions are classified to estimate the stride length, including: walking horizontally, walking up along a slope, stepping upstairs/downstairs and standing still. The stride length estimation is based on a simple empirical module where fix stride length is donated to each classified motion. In order to evaluate the EMG-based pedestrian dead reckoning (PDR) solution developed in this study, an outdoor-indoor field test had been carried out in the Finnish Geodetic Institute. The test results demonstrated that the EMG-based PDR solutions are comparable to the commercial GPS stand-alone solutions for a period of 9 minutes outdoors, which is equivalent to a walking distance of 667 meters and also demonstrates its robustness for indoor navigation for a period of 3 minutes.
Yuwei Chen 0005, Ruizhi Chen, Xiang Chen 0004
IPIN2
2011 Heading change detection for indoor navigation with a Smartphone camera
abstract
A comfortable and useful pedestrian navigation system is accurate, reasonably priced, easy to use, and light to carry. Smartphones are attractive platforms for the navigation systems due to their small size, low cost, and the fact that they are already carried routinely by many pedestrians. Pedestrian navigation is mostly needed in GNSS degraded and denied areas such as indoors and in urban canyons. Positioning in these environments is still a challenging task. Aiding from other location sensors is typically needed. Obtaining a reliable heading e.g. from a digital compass is one of the most challenging tasks indoors because of the environmental disturbances to the sensor measurements. Visual-aiding has been considered for solving this problem for the last few years. Most algorithms are however often too massive for a Smartphone with limited computing power and storage space. This paper presents a lightweight algorithm for solving the problem. It is based on vanishing points calculated from lines in consecutive images. Results from a field test with a Nokia Smartphone have demonstrated that the performance of the algorithm in calculating the heading change is much better than that of the built-in digital compass. The heading change can be detected with the proposed algorithm at about 1 Hz frequency under the PC environment. It has potential to be adopted for a Smartphone platform.
Laura Ruotsalainen, Heidi Kuusniemi, Ruizhi Chen
IPIN3