VLDB 2026 Research / reviewers in the wild / expert
Xiaoxuan Lu 0001
dblp:154/4313 · also Chris Xiaoxuan Lu
· DBLP profile ↗
67ranked-venue papers
13as first author
42since 2021 · last 2026
0000-0002-3733-4480ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 2 first-author · 23 since 2021Computer networks · 25 · 7 first-author · 13 since 2021Systems, architecture and hardware · 13 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 3 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | mmWave Radar Perception Learning Using Pervasive Visual-Inertial SupervisionabstractThis article introduces a radar perception learning framework guided by data collected from commonly equipped visual-inertial (VI) sensor suites on smart vehicles. Unlike existing approaches that rely on dense point clouds from 3D LiDARs, which are costly and not widely deployed, this method leverages the broader availability of VI data. However, visual images alone lack the ability to capture the three-dimensional motion of moving targets, which limits their effectiveness in supervising motion-related tasks. To overcome this limitation, the framework integrates multiple perception tasks such as odometry estimation, motion segmentation, and scene flow prediction into a unified learning process. The first component is an odometry estimation module that combines deterministic ego-motion models with data-driven learning results. This fusion helps accurately infer the scene flow of static background points while minimizing drift. The second component is a supervision signal extraction module that aligns optical and millimeter-wave radar measurements to guide the learning of radar scene flow and rigid transformations. This module improves the reliability of dynamic point supervision through joint constraints across sensing modalities. The third component introduces a feature-selection module designed for cross-modal learning. It enhances the accuracy of motion segmentation and enforces consistency between odometry and scene flow, resulting in more coherent radar perception outputs. Experimental evaluations show that this framework achieves superior performance in challenging conditions such as smoke-obscured environments. It surpasses state-of-the-art (SOTA) methods that depend on high-cost LiDAR systems. The implementation of VISC+ will be open-source athttps://github.com/weini-Eve/VISC Kezhong Liu, Yiwen Zhou, Mozi Chen, Jianhua He 0001, Jingao Xu, Zheng Yang 0002, Xiaoxuan Lu 0001, Shengkai Zhang |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2025 | Risk Controlled Image RetrievalabstractMost image retrieval research prioritizes improving predictive performance, often overlooking situations where the reliability of predictions is equally important. The gap between model performance and reliability requirements highlights the need for a systematic approach to analyze and address the risks associated with image retrieval. Uncertainty quantification technique can be applied to mitigate this issue by assessing uncertainty for retrieval sets, but it provides only a heuristic estimate of uncertainty rather than a guarantee. To address these limitations, we present Risk Controlled Image Retrieval (RCIR), which generates retrieval sets with coverage guarantee, i.e., retrieval sets that are guaranteed to contain the true nearest neighbors with a predefined probability. RCIR can be easily integrated with existing uncertainty-aware image retrieval systems, agnostic to data distribution and model selection. To the best of our knowledge, this is the first work that provides coverage guarantees to image retrieval. The validity and efficiency of RCIR are demonstrated on four real-world datasets: CAR-196, CUB-200, Pittsburgh, and ChestX-Det. Kaiwen Cai, Xiaoxuan Lu 0001, Xingyu Zhao 0001, Wei Huang 0035, Xiaowei Huang 0001 |
AAAI | 2 |
| 2025 | VISC: mmWave Radar Scene Flow Estimation using Pervasive Visual-Inertial SupervisionabstractThis work proposes a mmWave radar’s scene flow estimation framework supervised by data from a widespread visual-inertial (VI) sensor suite, allowing crowdsourced training data from smart vehicles. Current scene flow estimation methods for mmWave radar are typically supervised by dense point clouds from 3D LiDARs, which are expensive and not widely available in smart vehicles. While VI data are more accessible, visual images alone cannot capture the 3D motions of moving objects, making it difficult to supervise their scene flow. Moreover, the temporal drift of VI rigid transformation also degenerates the scene flow estimation of static points. To address these challenges, we propose a drift-free rigid transformation estimator that fuses kinematic model-based ego-motions with neural network-learned results. It provides strong supervision signals to radar-based rigid transformation and infers the scene flow of static points. Then, we develop an optical-mmWave supervision extraction module that extracts the supervision signals of radar rigid transformation and scene flow. It strengthens the supervision by learning the scene flow of dynamic points with the joint constraints of optical and mmWave radar measurements. Extensive experiments demonstrate that, in smoke-filled environments, our method even outperforms state-of-the-art (SOTA) approaches using costly LiDARs. Kezhong Liu, Yiwen Zhou, Mozi Chen, Jianhua He 0001, Jingao Xu, Zheng Yang 0002, Xiaoxuan Lu 0001, Shengkai Zhang |
IROS | 7 |
| 2025 | ThermoHands: A Benchmark for 3D Hand Pose Estimation from Egocentric Thermal ImagesabstractDesigning egocentric 3D hand pose estimation systems that can perform reliably in complex, real-world scenarios is crucial for downstream applications. Previous approaches using RGB or NIR imagery struggle in challenging conditions: RGB methods are susceptible to lighting variations and obstructions like handwear, while NIR techniques can be disrupted by sunlight or interference from other NIR-equipped devices. To address these limitations, we present ThermoHands, the first benchmark focused on thermal image-based egocentric 3D hand pose estimation, demonstrating the potential of thermal imaging to achieve robust performance under these conditions. The benchmark includes a multi-view and multi-spectral dataset collected from 28 subjects performing hand-object and hand-virtual interactions under diverse scenarios, accurately annotated with 3D hand poses through an automated process. We introduce a new baseline method, TherFormer, utilizing dual transformer modules for effective egocentric 3D hand pose estimation in thermal imagery. Our experimental results highlight TherFormer's leading performance and affirm thermal imaging's effectiveness in enabling robust 3D hand pose estimation in adverse conditions. Fangqiang Ding, Yunzhou Zhu 0001, Xiangyu Wen 0001, Gaowen Liu, Xiaoxuan Lu 0001 |
SenSys | 5 |
| 2025 | Learning Selective Sensor Fusion for State EstimationabstractAutonomous vehicles and mobile robotic systems are typically equipped with multiple sensors to provide redundancy. By integrating the observations from different sensors, these mobile agents are able to perceive the environment and estimate system states, e.g., locations and orientations. Although deep learning (DL) approaches for multimodal odometry estimation and localization have gained traction, they rarely focus on the issue of robust sensor fusion-a necessary consideration to deal with noisy or incomplete sensor observations in the real world. Moreover, current deep odometry models suffer from a lack of interpretability. To this extent, we propose SelectFusion, an end-to-end selective sensor fusion module that can be applied to useful pairs of sensor modalities, such as monocular images and inertial measurements, depth images, and light detection and ranging (LIDAR) point clouds. Our model is a uniform framework that is not restricted to specific modality or task. During prediction, the network is able to assess the reliability of the latent features from different sensor modalities and to estimate trajectory at both scale and global pose. In particular, we propose two fusion modules-a deterministic soft fusion and a stochastic hard fusion-and offer a comprehensive study of the new strategies compared with trivial direct fusion. We extensively evaluate all fusion strategies both on public datasets and on progressively degraded datasets that present synthetic occlusions, noisy and missing data, and time misalignment between sensors, and we investigate the effectiveness of the different fusion strategies in attending the most reliable features, which in itself provides insights into the operation of the various models. Changhao Chen, Stefano Rosa, Xiaoxuan Lu 0001, Bing Wang 0013, Agathoniki Trigoni, Andrew Markham |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Self-adapting Large Visual-Language Models to Edge Devices Across Visual Modalities
Kaiwen Cai, Zhekai Duan, Gaowen Liu, Charles Fleming, Xiaoxuan Lu 0001 |
ECCV (28) | 5 |
| 2024 | milliFlow: Scene Flow Estimation on mmWave Radar Point Cloud for Human Motion Sensing
Fangqiang Ding, Peijun Zhao, Xiaoxuan Lu 0001 |
ECCV (24) | 4 |
| 2024 | Robust 3D Object Detection from LiDAR-Radar Point Clouds via Cross-Modal Feature AugmentationabstractThis paper presents a novel framework for robust 3D object detection from point clouds via cross-modal hallucination. Our proposed approach is agnostic to either hallucination direction between LiDAR and 4D radar. We introduce multiple alignments on both spatial and feature levels to achieve simultaneous backbone refinement and hallucination generation. Specifically, spatial alignment is proposed to deal with the geometry discrepancy for better instance matching between LiDAR and radar. The feature alignment step further bridges the intrinsic attribute gap between the sensing modalities and stabilizes the training. The trained object detection models can deal with difficult detection cases better, even though only single-modal data is used as the input during the inference stage. Extensive experiments on the View-of-Delft (VoD) dataset show that our proposed method outperforms the state-of-the-art (SOTA) methods for both radar and LiDAR object detection while maintaining competitive efficiency in runtime. Jianning Deng, Gabriel Chan, Hantao Zhong, Xiaoxuan Lu 0001 |
ICRA | 4 |
| 2024 | RaTrack: Moving Object Detection and Tracking with 4D Radar Point CloudabstractMobile autonomy relies on the precise perception of dynamic environments. Robustly tracking moving objects in 3D world thus plays a pivotal role for applications like trajectory prediction, obstacle avoidance, and path planning. While most current methods utilize LiDARs or cameras for Multiple Object Tracking (MOT), the capabilities of 4D imaging radars remain largely unexplored. Recognizing the challenges posed by radar noise and point sparsity in 4D radar data, we introduce RaTrack, an innovative solution tailored for radar-based tracking. Bypassing the typical reliance on specific object types and 3D bounding boxes, our method focuses on motion segmentation and clustering, enriched by a motion estimation module. Evaluated on the View-of-Delft dataset, RaTrack showcases superior tracking precision of moving objects, largely surpassing the performance of the state of the art. We release our code and model at https://github.com/LJacksonPan/RaTrack. Zhijun Pan, Fangqiang Ding, Hantao Zhong, Xiaoxuan Lu 0001 |
ICRA | 4 |
| 2024 | Multimodal Indoor Localization Using Crowdsourced Radio MapsabstractIndoor Positioning Systems (IPS) traditionally rely on odometry and building infrastructures like WiFi, often supplemented by building floor plans for increased accuracy. However, the limitation of floor plans in terms of availability and timeliness of updates challenges their wide applicability. In contrast, the proliferation of smartphones and WiFi-enabled robots has made crowdsourced radio maps – databases pairing locations with their corresponding Received Signal Strengths (RSS) – increasingly accessible. These radio maps not only provide WiFi fingerprint-location pairs but encode movement regularities akin to the constraints imposed by floor plans. This work investigates the possibility of leveraging these radio maps as a substitute for floor plans in multimodal IPS. We introduce a new framework to address the challenges of radio map inaccuracies and sparse coverage. Our proposed system integrates an uncertainty-aware neural network model for WiFi localization and a bespoken Bayesian fusion technique for optimal fusion. Extensive evaluations on multiple real-world sites indicate a significant performance enhancement, with results showing ∼ 25% improvement over the best baseline. Zhaoguang Yi, Xiangyu Wen 0001, Qiyue Xia, Peize Li, Francisco Zampella, Firas Alsehly, Xiaoxuan Lu 0001 |
ICRA | 7 |
| 2024 | Click to Grasp: Zero-Shot Precise Manipulation via Visual Diffusion DescriptorsabstractPrecise manipulation that is generalizable across scenes and objects remains a persistent challenge in robotics. Current approaches for this task heavily depend on having a significant number of training instances to handle objects with pronounced visual and/or geometric part ambiguities. Our work explores the grounding of fine-grained part descriptors for precise manipulation in a zero-shot setting by utilizing web-trained text-to-image diffusion-based generative models. We tackle the problem by framing it as a dense semantic part correspondence task. Our model returns a gripper pose for manipulating a specific part, using as reference a user-defined click from a source image of a visually different instance of the same object. We require no manual grasping demonstrations as we leverage the intrinsic object geometry and features. Practical experiments in a real-world tabletop scenario validate the efficacy of our approach, demonstrating its potential for advancing semantic-aware robotics manipulation.Web page: https://tsagkas.github.io/click2grasp Nikolaos Tsagkas, Jack Rome, Subramanian Ramamoorthy, Oisin Mac Aodha, Xiaoxuan Lu 0001 |
IROS | 5 |
| 2024 | RadarOcc: Robust 3D Occupancy Prediction with 4D Imaging Radarabstract3D occupancy-based perception pipeline has significantly advanced autonomous driving by capturing detailed scene descriptions and demonstrating strong generalizability across various object categories and shapes. Current methods predominantly rely on LiDAR or camera inputs for 3D occupancy prediction. These methods are susceptible to adverse weather conditions, limiting the all-weather deployment of self-driving cars. To improve perception robustness, we leverage the recent advances in automotive radars and introduce a novel approach that utilizes 4D imaging radar sensors for 3D occupancy prediction. Our method, RadarOcc, circumvents the limitations of sparse radar point clouds by directly processing the 4D radar tensor, thus preserving essential scene details. RadarOcc innovatively addresses the challenges associated with the voluminous and noisy 4D radar data by employing Doppler bins descriptors, sidelobe-aware spatial sparsification, and range-wise self-attention mechanisms. To minimize the interpolation errors associated with direct coordinate transformations, we also devise a spherical-based feature encoding followed by spherical-to-Cartesian feature aggregation. We benchmark various baseline methods based on distinct modalities on the public K-Radar dataset. The results demonstrate RadarOcc's state-of-the-art performance in radar-based 3D occupancy prediction and promising results even when compared with LiDAR- or camera-based methods. Additionally, we present qualitative evidence of the superior performance of 4D radar in adverse weather conditions and explore the impact of key pipeline components through ablation studies. Fangqiang Ding, Xiangyu Wen 0001, Yunzhou Zhu 0001, Yiming Li 0003, Xiaoxuan Lu 0001 |
NeurIPS | 5 |
| 2024 | Forecasting backdraft with multimodal method: Fusion of fire image and sensor data
Tianhang Zhang, Fangqiang Ding, Zilong Wang 0029, Xiaoxuan Lu 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | Robust Metric Localization in Autonomous Driving via Doppler Compensation With Single-Chip RadarabstractMetric localization is vital to autonomous driving where it corrects cumulative errors in a long-term run. Such errors are inevitable in real scenarios where GPS signals or some other drift-free exteroceptive measurements are not available, e.g., when an automobile goes through a tunnel. Using FMCW-based mmWave radars is an attractive metric localization technique with improved robustness as RF signals can traverse small particles in harsh weather conditions like snowing, foggy, and storming, but it faces a fundamental challenge of Doppler distortion. Existing works take spatial constraints to mitigate the Doppler distortion of point clouds from mechanical radars with limited accuracy. Modern single-chip mmWave radars that provide dynamic estimates, i.e., radial velocities, bring new opportunities to develop more accurate approaches. This paper presentsDC-Loc++, a robust metric localization framework by compensating Doppler distortions using a single-chip mmWave radar. It consists of an explicit velocity-assisted Doppler compensation module for each radar sub-map, an uncertainty-aware metric registration algorithm, and a failure recovery method that validates measurement constraints to generate a more confident pose graph for optimizing vehicle poses. Extensive experiments on both nuScenes dataset and a synthetic CARLA dataset show the effectiveness ofDC-Loc++, achieving 99.2% success rate and more than 20.0%, 30.2% error reductions in terms of translation and rotation estimates, respectively, compared with existing approaches. Pengen Gao, Shengkai Zhang, Wei Wang 0050, Xiaoxuan Lu 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Deep Learning for Visual Localization and Mapping: A SurveyabstractDeep-learning-based localization and mapping approaches have recently emerged as a new research direction and receive significant attention from both industry and academia. Instead of creating hand-designed algorithms based on physical models or geometric theories, deep learning solutions provide an alternative to solve the problem in a data-driven way. Benefiting from the ever-increasing volumes of data and computational power on devices, these learning methods are fast evolving into a new area that shows potential to track self-motion and estimate environmental models accurately and robustly for mobile agents. In this work, we provide a comprehensive survey and propose a taxonomy for the localization and mapping methods using deep learning. This survey aims to discuss two basic questions: whether deep learning is promising for localization and mapping, and how deep learning should be applied to solve this problem. To this end, a series of localization and mapping topics are investigated, from the learning-based visual odometry and global relocalization to mapping, and simultaneous localization and mapping (SLAM). It is our hope that this survey organically weaves together the recent works in this vein from robotics, computer vision, and machine learning communities and serves as a guideline for future researchers to apply deep learning to tackle the problem of visual localization and mapping. Changhao Chen, Bing Wang 0013, Xiaoxuan Lu 0001, Agathoniki Trigoni, Andrew Markham |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Introduction to the Special Section on Contact-free Smart Sensing in AIoTabstractIntroduction to the Special Section on Contact-free Smart Sensing in AloTArtificial Intelligence (AI) and the Internet of Things (IoT) are two powerful forces that have been reshaping our world in recent years.When they converge, they create a new field of AIoT that enables ubiquitous intelligence through the integration of smart algorithms and connected devices.One of the key enablers of AIoT is contact-free sensing, which leverages the availability of portable and highly integrated WiFi, radar, and sonar-style devices to monitor humans and environments without physical contact.This technology has transformed the traditional computer vision-based paradigms and opened up novel possibilities for data collection and analysis.However, contactfree sensing also poses new challenges and risks for AIoT applications.The dynamic and complex wireless environments require innovative solutions for efficient data processing and interpretation.The security and privacy issues of WiFi, radar, and sonar-enabled sensing devices also demand urgent attention, as they may expose sensitive information to malicious attacks.Therefore, it is imperative to explore the potential and pitfalls of contact-free sensing in AIoT and to develop effective strategies for ensuring the robustness and reliability of AIoT applications.This special issue is dedicated to highlighting the cutting-edge methods and latest research in the field of contact-free sensing, which leverages WiFi, radar, and sonar-style devices to monitor humans and environments without physical contact.The main focus of this issue is to explore the latest machine learning analytics to extract information from the sensory data and to investigate the potential risks and countermeasures to ensure the security and privacy of sensing devices.The call for papers attracted with 44 submissions and after a rigorous review, 18 papers have been accepted for this special issue.A brief summary of some papers in this special issue is presented in the following:In "Feasibility of Remote Blood Pressure Estimation via Narrow-band Multi-wavelength Pulse Transit Time, " the authors investigate the feasibility of estimating blood pressure (BP) via pulse transit time (PTT) in a novel remote single-site manner using a modified RGB camera.A narrowband triple band-pass filter makes it possible to measure the PTT between different skin layers, harvesting information from green and near-infrared wavelengths.They design a color-channel model and a novel channel-separation method to further resolve the inter-channel influence and band overlap.The results showed a good absolute Pearson's correlation coefficient between both MW PTT and systolic BP as well as diastolic BP, pointing to the feasibility of the proposed novel remote MW BP estimation via PTT.In "LiteWiSys: A Lightweight System for WiFi-based Dual-task Action Perception, " Sheng et al. propose a lightweight system named LiteWiSys that can simultaneously detect and recognize WiFi-based human actions.This work addresses two major drawbacks of existing methods: heavy Pengfei Hu 0001, Zhe Chen 0015, Xiaoxuan Lu 0001, Xuyu Wang, Jun Luo 0001, Prasant Mohapatra |
ACM Trans. Sens. Networks | 3 |
| 2024 | End-to-End Target Liveness Detection via mmWave Radar and Vision Fusion for Autonomous VehiclesabstractThe successful operation of autonomous vehicles hinges on their ability to accurately identify objects in their vicinity, particularly living targets such as bikers and pedestrians. However, visual interference inherent in real-world environments, such as omnipresent billboards, poses substantial challenges to extant vision-based detection technologies. These visual interference exhibit similar visual attributes to living targets, leading to erroneous identification. We address this problem by harnessing the capabilities of mmWave radar, a vital sensor in autonomous vehicles, in combination with vision technology, thereby contributing a unique solution for liveness target detection. We propose a methodology that extracts features from the mmWave radar signal to achieve end-to-end liveness target detection by integrating the mmWave radar and vision technology. This proposed methodology is implemented and evaluated on the commodity mmWave radar IWR6843ISK-ODS and vision sensor Logitech camera. Our extensive evaluation reveals that the proposed method accomplishes liveness target detection with a mean average precision of 98.1%, surpassing the performance of existing studies. Shuai Wang 0008, Luoyu Mei, Zhimeng Yin 0001, Ruofeng Liu, Wenchao Jiang, Xiaoxuan Lu 0001 |
ACM Trans. Sens. Networks | 7 |
| 2023 | Hidden Gems: 4D Radar Scene Flow Learning Using Cross-Modal SupervisionabstractThis work proposes a novel approach to 4D radar-based scene flow estimation via cross-modal learning. Our approach is motivated by the co-located sensing redundancy in modern autonomous vehicles. Such redundancy implicitly provides various forms of supervision cues to the radar scene flow estimation. Specifically, we introduce a multi-task model architecture for the identified cross-modal learning problem and propose loss functions to opportunistically engage scene flow estimation using multiple cross-modal constraints for effective model training. Extensive experiments show the state-of-the-art performance of our method and demonstrate the effectiveness of cross-modal super-vised learning to infer more accurate 4D radar scene flow. We also show its usefulness to two subtasks - motion segmentation and ego-motion estimation. Our source code will be available on https://github.com/Toytiny/CMFlow. Fangqiang Ding, Andras Palffy, Dariu Gavrila, Xiaoxuan Lu 0001 |
CVPR | 4 |
| 2023 | Robust Human Detection under Visual Degradation via Thermal and mmWave Radar Fusion
Kaiwen Cai, Qiyue Xia, Peize Li, Xiaoxuan Lu 0001, John A. Stankovic |
EWSN | 4 |
| 2023 | Feature-based Visual Odometry for Bronchoscopy: A Dataset and BenchmarkabstractBronchoscopy is a medical procedure that involves the insertion of a flexible tube with a camera into the airways to survey, diagnose and treat lung diseases. Due to the complex branching anatomical structure of the bronchial tree and the similarity of the inner surfaces of the segmental airways, navigation systems are now being routinely used to guide the operator during procedures to access the lung periphery. Current navigation systems rely on sensor-integrated bronchoscopes to track the position of the bronchoscope in real-time. This approach has limitations, including increased cost and limited use in non-specialized settings. To address this issue, researchers have proposed visual odometry algorithms to track the bronchoscope camera without the need for external sensors. However, due to the lack of publicly available datasets, limited progress is made. To this end, we have developed a database of bronchoscopy videos in a phantom lung model and ex-vivo human lungs. The dataset contains 34 video sequences with over 23,000 frames with odometry ground truth data collected using electromagnetic tracking sensors. With our dataset, we empower the robotics and machine learning community to advance the field. We share our insights on challenges in endoscopic visual odometry. Furthermore, we provide benchmark results for this dataset. State-of-the-art feature extraction algorithms including SIFT, ORB, Superpoint, Shi- Tomasi, and LoFTR are tested on this dataset. The benchmark results demonstrate that the LoFTR algorithm outperforms other approaches, but still has significant errors in the presence of rapid movements and occlusions. Jianning Deng, Peize Li, Kevin Dhaliwal, Xiaoxuan Lu 0001, Mohsen Khadem |
IROS | 4 |
| 2023 | RADA: Robust Adversarial Data Augmentation for Camera Localization in Challenging ConditionsabstractCamera localization is a fundamental problem for many applications in computer vision, robotics, and autonomy. Despite recent deep learning-based approaches, the lack of robustness in challenging conditions persists due to changes in appearance caused by texture-less planes, repeating structures, reflective surfaces, motion blur, and illumination changes. Data augmentation is an attractive solution, but standard image perturbation methods fail to improve localization robustness. To address this, we propose RADA, which concentrates on perturbing the most vulnerable pixels to generate relatively less image perturbations that perplex the network. Our method outperforms previous augmentation techniques, achieving up to twice the accuracy of state-of-the-art models even under ‘unseen’ challenging weather conditions. Videos of our results can be found at https://youtu.be/niOv7-fJeCA. The source code for RADA is publicly available at https://github.com/jialuwang123321/RADA. Muhamad Risqi Utama Saputra, Xiaoxuan Lu 0001, Agathoniki Trigoni, Andrew Markham |
IROS | 3 |
| 2023 | MetaWave: Attacking mmWave Sensing with Meta-material-enhanced Tags
Zhengxiong Li, Baicheng Chen, Yi Zhu 0012, Xiaoxuan Lu 0001, Zhengyu Peng, Feng Lin 0004, Wenyao Xu, Kui Ren 0001, Chunming Qiao |
NDSS | 5 |
| 2023 | MM-Fi: Multi-Modal Non-Intrusive 4D Human Dataset for Versatile Wireless Sensingabstract4D human perception plays an essential role in a myriad of applications, such as home automation and metaverse avatar simulation. However, existing solutions which mainly rely on cameras and wearable devices are either privacy intrusive or inconvenient to use. To address these issues, wireless sensing has emerged as a promising alternative, leveraging LiDAR, mmWave radar, and WiFi signals for device-free human sensing. In this paper, we propose MM-Fi, the first multi-modal non-intrusive 4D human dataset with 27 daily or rehabilitation action categories, to bridge the gap between wireless sensing and high-level human perception tasks. MM-Fi consists of over 320k synchronized frames of five modalities from 40 human subjects. Various annotations are provided to support potential sensing tasks, e.g., human pose estimation and action recognition. Extensive experiments have been conducted to compare the sensing capacity of each or several modalities in terms of multiple tasks. We envision that MM-Fi can contribute to wireless sensing research with respect to action recognition, human pose estimation, multi-modal learning, cross-modal supervision, and interdisciplinary healthcare research. Jianfei Yang 0001, Yunjiao Zhou, Xinyan Chen 0002, Yuecong Xu, Shenghai Yuan 0001, Han Zou, Xiaoxuan Lu 0001, Lihua Xie 0001 |
NeurIPS | 8 |
| 2023 | Poster Abstract: Multimodal Indoor Localization Using Crowdsourced Radio MapsabstractTraditional Indoor Positioning Systems (IPS) use odometry, WiFi, and often building floor plans for accuracy. However, floor plan limitations have shifted attention to crowd-sourced radio maps, popularized by smartphones and WiFi-integrated robots. These maps pair locations with Received Signal Strengths (RSS) and reflect movement patterns similar to floor plans. Our research explores using radio maps as an alternative to floor plans in IPS. We've developed a new framework that combines an uncertainty-aware neural network for WiFi positioning with a Bayesian fusion method. Testing in real-world scenarios showed about a 25% performance increase compared to the leading baseline. Xiangyu Wen 0001, Zhaoguang Yi, Francisco Zampella, Firas Alsehly, Xiaoxuan Lu 0001 |
SenSys | 5 |
| 2023 | GaitFi: Robust Device-Free Human Identification via WiFi and Vision Multimodal LearningabstractAs an important biomarker for human identification, human gait can be collected at a distance by passive sensors without subject cooperation, which plays an essential role in crime prevention, security detection, and other human identification applications. Presently, most research works are based on cameras and computer vision techniques to perform gait recognition. However, vision-based methods are not reliable when confronting poor illuminations, leading to degrading performances. In this article, we propose a novel multimodal gait recognition method, namely, GaitFi, which leverages WiFi signals and videos for human identification. In GaitFi, channel state information (CSI) that reflects the multipath propagation of WiFi is collected to capture human gaits, while videos are captured by cameras. To learn robust gait information, we propose a lightweight residual convolution network (LRCN) as the backbone network and further propose the two-stream GaitFi by integrating WiFi and vision features for the gait retrieval task. The GaitFi is trained by the triplet loss and classification loss on different levels of features. Extensive experiments are conducted in the real world, which demonstrates that the GaitFi outperforms state-of-the-art gait recognition methods based on single WiFi or camera, achieving 94.2% for human identification tasks of 12 subjects. Lang Deng, Jianfei Yang 0001, Shenghai Yuan 0001, Han Zou, Xiaoxuan Lu 0001, Lihua Xie 0001 |
IEEE Internet Things J. | 5 |
| 2023 | CubeLearn: End-to-End Learning for Human Motion Recognition From Raw mmWave Radar SignalsabstractmmWave FMCW radar has attracted a huge amount of research interest for human-centered applications in recent years, such as human gesture and activity recognition. Most existing pipelines are built upon conventional discrete Fourier transform (DFT) preprocessing and deep neural network classifier hybrid methods, with a majority of previous works focusing on designing the downstream classifier to improve overall accuracy. In this work, we take a step back and look at the preprocessing module. To avoid the drawbacks of conventional DFT preprocessing, we propose a complex-weighted learnable preprocessing module, named CubeLearn, to directly extract features from raw radar signal and build an end-to-end deep neural network for mmWave FMCW radar motion recognition applications. Extensive experiments show that our CubeLearn module consistently improves the classification accuracies of different pipelines, especially, benefiting those simpler models, which are more likely to be used on edge devices due to their computational efficiency. We provide ablation studies on initialization methods and structure of the proposed module, as well as an evaluation of the running time on PC and edge devices. This work also serves as a comparison of different approaches toward data cube slicing. Through our task-agnostic design, we propose a first step toward a generic end-to-end solution for radar recognition problems. Peijun Zhao, Xiaoxuan Lu 0001, Bing Wang 0013, Agathoniki Trigoni, Andrew Markham |
IEEE Internet Things J. | 2 |
| 2022 | AutoPlace: Robust Place Recognition with Single-chip Automotive RadarabstractThis paper presents a novel place recognition approach to autonomous vehicles by using low-cost, single-chip automotive radar. Aimed at improving recognition robustness and fully exploiting the rich information provided by this emerging automotive radar, our approach follows a principled pipeline that comprises (1) dynamic points removal from instant Doppler measurement, (2) spatial-temporal feature embedding on radar point clouds, and (3) retrieved candidates refinement from Radar Cross Section measurement. Extensive experimental results on the public nuScenes dataset demonstrate that existing visual/LiDAR/spinning radar place recognition approaches are less suitable for single-chip automotive radar. In contrast, our purpose-built approach for automotive radar consistently outperforms a variety of baseline methods via a comprehensive set of metrics, providing insights into the efficacy when used in a realistic system. Kaiwen Cai, Bing Wang 0013, Xiaoxuan Lu 0001 |
ICRA | 3 |
| 2022 | DC-Loc: Accurate Automotive Radar Based Metric Localization with Explicit Doppler CompensationabstractAutomotive mmWave radar has been widely used in the automotive industry due to its small size, low cost, and complementary advantages to optical sensors (e.g., cameras, LiDAR, etc.) in adverse weathers, e.g., fog, raining, and snowing. On the other side, its large wavelength also poses fundamental challenges to perceive the environment. Recent advances have made breakthroughs on its inherent drawbacks, i.e., the multipath reflection and the sparsity of mmWave radar's point clouds. However, the frequency-modulated continuous wave modulation of radar signals makes it more sensitive to vehicles’ mobility than optical sensors. This work focuses on the problem of frequency shift, i.e., the Doppler effect distorts the radar ranging measurements and its knock-on effect on metric localization. We propose a new radar-based metric localization framework, termed DC-Loc, which can obtain more accurate location estimation by restoring the Doppler distortion. Specifically, we first design a new algorithm that explicitly compensates the Doppler distortion of radar scans and then model the measurement uncertainty of the Doppler-compensated point cloud to further optimize the metric localization. Extensive experiments using the public nuScenes dataset and CARLA simulator demonstrate that our method outperforms the state-of-the-art approach by 25.2% and 5.6% improvements in terms of translation and rotation errors, respectively. Pengen Gao, Shengkai Zhang, Wei Wang 0050, Xiaoxuan Lu 0001 |
ICRA | 4 |
| 2022 | Demo Abstract: 3D Simultaneous localization and Mapping with Power Network Electromagnetic RadiationabstractIndoor localization by leveraging the existing residential instru-ments has been widely explored. Given the properties of tempo-ral stability and spatial distinctness, the electromagnetic radiation (EMR) from the powerline network is a promising signal for location sensing. In this demo, we present a three-sensor setup to capture the powerline EMR signal from the three-dimensional (3D) space and formulate a new powerline EMR feature to implement the simultaneous localization and mapping (SLAM). Compared with the single sensor setup, our proposed approach can improve the localization accuracy to decimeter level. Zhenyu Yan 0002, Rui Tan 0001, Xiaoxuan Lu 0001 |
IPSN | 5 |
| 2022 | STUN: Self-Teaching Uncertainty Estimation for Place RecognitionabstractPlace recognition is key to Simultaneous Localization and Mapping (SLAM) and spatial perception. However, a place recognition in the wild often suffers from erroneous predictions due to image variations, e.g., changing viewpoints and street appearance. Integrating uncertainty estimation into the life cycle of place recognition is a promising method to mitigate the impact of variations on place recognition performance. However, existing uncertainty estimation approaches in this vein are either computationally inefficient (e.g., Monte Carlo dropout) or at the cost of dropped accuracy. This paper proposes STUN, a self-teaching framework that learns to simultaneously predict the place and estimate the prediction uncertainty given an input image. To this end, we first train a teacher net using a standard metric learning pipeline to produce embedding priors. Then, supervised by the pretrained teacher net, a student net with an additional variance branch is trained to finetune the embedding priors and estimate the uncertainty sample by sample. During the online inference phase, we only use the student net to generate a place prediction in conjunction with the uncertainty. When compared with place recognition systems that are ignorant of the uncertainty, our framework features the uncertainty estimation for free without sacrificing any prediction accuracy. Our experimental results on the large-scale Pittsburgh30k dataset demonstrate that STUN outperforms the state-of-the-art methods in both recognition accuracy and the quality of uncertainty estimation. Kaiwen Cai, Xiaoxuan Lu 0001, Xiaowei Huang 0001 |
IROS | 2 |
| 2022 | OdomBeyondVision: An Indoor Multi-modal Multi-platform Odometry Dataset Beyond the Visible SpectrumabstractThis paper presents a multimodal indoor odometry dataset, OdomBeyondVision, featuring multiple sensors across the different spectrum and collected with different mobile platforms. Not only does OdomBeyondVision contain the traditional navigation sensors, sensors such as IMUs, mechanical LiDAR, RGBD camera, it also includes several emerging sensors such as the single-chip mmWave radar, LWIR thermal camera and solid-state LiDAR. With the above sensors on UAV, UGV and handheld platforms, we respectively recorded the multimodal odometry data and their movement trajectories in various indoor scenes and different illumination conditions. We release the exemplar radar, radar-inertial and thermal-inertial odometry implementations to demonstrate their results for future works to compare against and improve upon. The full dataset including toolkit and documentation is publicly available at: https://github.com/MAPS-Lab/OdomBeyondVision. Peize Li, Kaiwen Cai, Muhamad Risqi Utama Saputra, Zhuangzhuang Dai, Xiaoxuan Lu 0001 |
IROS | 5 |
| 2022 | SpiralSpy: Exploring a Stealthy and Practical Covert Channel to Attack Air-gapped Computing Devices via mmWave Sensing
Zhengxiong Li, Baicheng Chen, Huining Li, Chenhan Xu, Feng Lin 0004, Xiaoxuan Lu 0001, Kui Ren 0001, Wenyao Xu |
NDSS | 7 |
| 2022 | Pedestrian Liveness Detection Based on mmWave Radar and Camera FusionabstractAutonomous driving requires vehicles to achieve fine detection of objects in the surrounding environment, especially living pedestrians. Nevertheless, in real world road environments there are living pedestrians and roadside portrait billboards. Existing vision-based object detection technologies fail to ac-curately distinguish living pedestrians from human figures. As an important sensor of autonomous driving system, mmWave radar has extra help to detect living pedestrians. In this paper, we extract the radar cross section (RCS) of the object from the low-cost mmWave radar signal as a distinguishing feature between living pedestrian and portrait billboard. Based on this observation, we propose a feature fusion network of mmWave radar and computer vision based on attention mechanism, and detect living pedestrians from fusion features. We implement the design with commodity mmWave radar IWR6843ISK-ODS and RGB camera Logitech Pro C920. The evaluation results show that our method effectively detects living pedestrians with an mAP of 97.7% and outperforms existing studies. Ruofeng Liu, Shuai Wang 0008, Wenchao Jiang, Xiaoxuan Lu 0001 |
SECON | 5 |
| 2022 | Telesonar: Robocall Alarm System by Detecting Echo Channel and Breath TimingabstractMassive fraudulent and phishing robocalls present threats to societies. The integration of artificial intelligence technologies, including dialogue and voice generation systems, renders the robocalls more deceptive. Existing countermeasures such as caller ID, call provenance, voiceprint, and fake voice detection have respective limitations and are heavyweight for end users' smartphones. This paper studies detecting the acoustic echo channel on the remote end of a call based on the received voice. The positive detection result evidencing the physical setup of an audio system is indicative of a human caller. However, the acoustic echo cancellation mechanisms of most audio systems and the use of earphone/headset diminish echoes significantly. To address these issues, the proposed Telesonar transmits short chirps during the vulnerable time of echo cancellation, detects the tiny echo remnants from the received voice, and passively analyzes the timing of caller's breath sounds to confirm a human caller. Extensive real experiments under a wide range of settings show that Telesonar correctly recognizes human callers with a rate of over 95%, while wrongly recognizing voice robots as human with a rate of 3.8%. Zhenyu Yan 0002, Rui Tan 0001, Qun Song 0001, Xiaoxuan Lu 0001 |
SenSys | 4 |
| 2022 | Graph-Based Thermal-Inertial SLAM With Probabilistic Neural NetworksabstractSimultaneous localization and mapping (SLAM) system typically employs vision-based sensors to observe the surrounding environment. However, the performance of such systems highly depends on the ambient illumination conditions. In scenarios with adverse visibility or in the presence of airborne particulates (e.g., smoke, dust, etc.), alternative modalities such as those based on thermal imaging and inertial sensors are more promising. In this article, we propose the first complete thermal–inertial SLAM system that combines neural abstraction in the SLAM front end with robust pose-graph optimization in the SLAM back end. We model the sensor abstraction in the front end by employing probabilistic deep learning parameterized by mixture density networks (MDNs). Our key strategies to successfully model this encoding from thermal imagery are the usage of normalized 14-b radiometric data, the incorporation of hallucinated visual (RGB) features, and the inclusion of feature selection to estimate the MDN parameters. To enable a full SLAM system, we also design an efficient global image descriptor that is able to detect loop closures from thermal embedding vectors. We performed extensive experiments and analysis using three datasets, namely self-collected ground robot and hand-held data taken in indoor environment, and one public dataset (SubT-tunnel) collected in underground tunnel. Finally, we demonstrate that an accurate thermal–inertial SLAM system can be realized in conditions of both benign and adverse visibility. Muhamad Risqi Utama Saputra, Xiaoxuan Lu 0001, Pedro Porto Buarque de Gusmão, Bing Wang 0013, Andrew Markham, Agathoniki Trigoni |
IEEE Trans. Robotics | 2 |
| 2021 | P2-Net: Joint Description and Detection of Local Features for Pixel and Point MatchingabstractAccurately describing and detecting 2D and 3D key-points is crucial to establishing correspondences across images and point clouds. Despite a plethora of learning-based 2D or 3D local feature descriptors and detectors having been proposed, the derivation of a shared descriptor and joint keypoint detector that directly matches pixels and points remains under-explored by the community. This work takes the initiative to establish fine-grained correspondences between 2D images and 3D point clouds. In order to directly match pixels and points, a dual fully-convolutional framework is presented that maps 2D and 3D inputs into a shared latent representation space to simultaneously describe and detect keypoints. Furthermore, an ultra-wide reception mechanism and a novel loss function are designed to mitigate the intrinsic information variations between pixel and point local regions. Extensive experimental results demonstrate that our framework shows competitive performance in fine-grained matching between images and point clouds and achieves state-of-the-art results for the task of indoor visual localization. Our source code is available at https://github.com/BingCS/P2-Net. Bing Wang 0013, Changhao Chen, Zhaopeng Cui, Jie Qin 0004, Xiaoxuan Lu 0001, Zhengdi Yu, Peijun Zhao, Zhen Dong 0005, Fan Zhu 0001, Agathoniki Trigoni, Andrew Markham |
ICCV | 5 |
| 2021 | 3D Motion Capture of an Unmodified Drone with Single-chip Millimeter Wave RadarabstractAccurate motion capture of aerial robots in 3D is a key enabler for autonomous operation in indoor environments such as warehouses or factories, as well as driving forward research in these areas. The most commonly used solutions at present are optical motion capture (e.g. VICON) and Ultrawide-band (UWB), but these are costly and cumbersome to deploy, due to their requirement of multiple cameras/anchors spaced around the tracking area. They also require the drone to be modified to carry an active or passive marker. In this work, we present an inexpensive system that can be rapidly installed, based on single-chip millimeter wave (mmWave) radar. Importantly, the drone does not need to be modified or equipped with any markers, as we exploit the Doppler signals from the rotating propellers. Furthermore, 3D tracking is possible from a single point, greatly simplifying deployment. We develop a novel deep neural network and demonstrate decimeter level 3D tracking at 10Hz, achieving better performance than classical baselines. Our hope is that this low-cost system will act to catalyse inexpensive drone research and increased autonomy. Peijun Zhao, Xiaoxuan Lu 0001, Bing Wang 0013, Agathoniki Trigoni, Andrew Markham |
ICRA | 2 |
| 2021 | Can Image Style Transfer Save Automotive Radar?abstractCompared to RGB camera and Lidar, single chip automotive radar is a promising alternative sensor with robustness to adverse weathers. But the sparseness of radar output drastically hinders its usefulness for autonomous driving tasks. Up-sampling via image style transfer could be a cure for a sparse measurement. However, it remains unknown whether style transfer can be an effective solution to automotive radar which features different and unique sparse and noisy issues. In this paper, we evaluate a variety of predominant image style transfer methods for a typical ego-vehicle pose estimation task on the public nuScenes dataset, and find that though image style transfer methods can improve the visual quality of automotive radar measurements, they can hardly contribute to the utility of radar for downstream tasks. Jianning Deng, Kaiwen Cai, Xiaoxuan Lu 0001 |
SenSys | 3 |
| 2021 | Motion Tracklet Oriented 6-DoF Inertial Tracking Using Commodity SmartphonesabstractMotion tracklets are the basic fragments of the track followed by a moving object and constitute various everyday motion behavior. An accurate estimation of motion tracklets in 3-D space can enable a wide range of applications, ranging from human computer interaction to medical rehabilitation. This paper presents a novel dataset for accurate 6-DoF motion tracklet estimation with the inertial sensors on commodity smartphones. The dataset consists of around 100 minutes of handheld motion with 3 predominant types of motion track-lets and accurate ground truth using the Vicon systems. With the presented dataset, we further benchmarked the trajectory estimation using a lightweight neural odometry model, showcasing how the dataset can be used while providing quantitative performance for downstream tasks. Our dataset, toolkit and source code available at https://github.com/MAPS-Lab/smartphone-tracking-dataset. Peize Li, Xiaoxuan Lu 0001 |
SenSys | 2 |
| 2021 | Human tracking and identification through a millimeter wave radar
Peijun Zhao, Xiaoxuan Lu 0001, Changhao Chen, Wei Wang 0226, Agathoniki Trigoni, Andrew Markham |
Ad Hoc Networks | 2 |
| 2021 | Deep Neural Network Based Inertial Odometry Using Low-Cost Inertial Measurement UnitsabstractInertial measurement units (IMUs) have emerged as an essential component in many of today's indoor navigation solutions due to their low cost and ease of use. However, despite many attempts for reducing the error growth of navigation systems based on commercial-grade inertial sensors, there is still no satisfactory solution that produces navigation estimates with long-time stability in widely differing conditions. This paper proposes to break the cycle of continuous integration used in traditional inertial algorithms, formulate it as an optimization problem, and explore the use of deep recurrent neural networks for estimating the displacement of a user over a specified time window. By training the deep neural network using inertial measurements and ground truth displacement data, it is possible to learn both motion characteristics and systematic error drift. As opposed to established context-aided inertial solutions, the proposed method is not dependent on either fixed sensor positions or periodic motion patterns. It can reconstruct accurate trajectories directly from raw inertial measurements, and predict the corresponding uncertainty to show model confidence. Extensive experimental evaluations demonstrate that the neural network produces position estimates with high accuracy for several different attachments, users, sensors, and motion types. As a particular demonstration of its flexibility, our deep inertial solutions can estimate trajectories for non-periodic motion, such as the shopping trolley tracking. Further more, it works in highly dynamic conditions, such as running, remaining extremely challenging for current techniques. Changhao Chen, Xiaoxuan Lu 0001, Johan Wahlström, Andrew Markham, Agathoniki Trigoni |
IEEE Trans. Mob. Comput. | 2 |
| 2021 | DynaNet: Neural Kalman Dynamical Model for Motion Estimation and PredictionabstractDynamical models estimate and predict the temporal evolution of physical systems. State-space models (SSMs) in particular represent the system dynamics with many desirable properties, such as being able to model uncertainty in both the model and measurements, and optimal (in the Bayesian sense) recursive formulations, e.g., the Kalman filter. However, they require significant domain knowledge to derive the parametric form and considerable hand tuning to correctly set all the parameters. Data-driven techniques, e.g., recurrent neural networks, have emerged as compelling alternatives to SSMs with wide success across a number of challenging tasks, in part due to their impressive capability to extract relevant features from rich inputs. They, however, lack interpretability and robustness to unseen conditions. Thus, data-driven models are hard to be applied in safety-critical applications, such as self-driving vehicles. In this work, we present DynaNet, a hybrid deep learning and time-varying SSM, which can be trained end-to-end. Our neural Kalman dynamical model allows us to exploit the relative merits of both SSM and deep neural networks. We demonstrate its effectiveness in the estimation and prediction on a number of physically challenging tasks, including visual odometry, sensor fusion for visual-inertial navigation, and motion prediction. In addition, we show how DynaNet can indicate failures through investigation of properties, such as the rate of innovation (Kalman gain). Changhao Chen, Xiaoxuan Lu 0001, Bing Wang 0013, Agathoniki Trigoni, Andrew Markham |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | AtLoc: Attention Guided Camera LocalizationabstractDeep learning has achieved impressive results in camera localization, but current single-image techniques typically suffer from a lack of robustness, leading to large outliers. To some extent, this has been tackled by sequential (multi-images) or geometry constraint approaches, which can learn to reject dynamic objects and illumination conditions to achieve better performance. In this work, we show that attention can be used to force the network to focus on more geometrically robust objects and features, achieving state-of-the-art performance in common benchmark, even if using only a single image as input. Extensive experimental evidence is provided through public indoor and outdoor datasets. Through visualization of the saliency maps, we demonstrate how the network learns to reject dynamic objects, yielding superior global camera pose regression performance. The source code is avaliable at https://github.com/BingCS/AtLoc. Bing Wang 0013, Changhao Chen, Xiaoxuan Lu 0001, Peijun Zhao, Agathoniki Trigoni, Andrew Markham |
AAAI | 3 |
| 2020 | Heart Rate Sensing with a Robot Mounted mmWave RadarabstractHeart rate monitoring at home is a useful metric for assessing health e.g. of the elderly or patients in post-operative recovery. Although non-contact heart rate monitoring has been widely explored, typically using a static, wall-mounted device, measurements are limited to a single room and sensitive to user orientation and position. In this work, we propose mBeats, a robot mounted millimeter wave (mmWave) radar system that provide periodic heart rate measurements under different user poses, without interfering in a users daily activities. mBeats contains a mmWave servoing module that adaptively adjusts the sensor angle to the best reflection pro le. Furthermore, mBeats features a deep neural network predictor, which can estimate heart rate from the lower leg and additionally provides estimation uncertainty. Through extensive experiments, we demonstrate accurate and robust operation of mBeats in a range of scenarios. We believe by integrating mobility and adaptability, mBeats can empower many down-stream healthcare applications at home, such as palliative care, post-operative rehabilitation and telemedicine. Peijun Zhao, Xiaoxuan Lu 0001, Bing Wang 0013, Changhao Chen, Linhai Xie, Agathoniki Trigoni, Andrew Markham |
ICRA | 2 |
| 2020 | See through smoke: robust indoor mapping with low-cost mmWave radarabstractThis paper presents the design, implementation and evaluation of milliMap, a single-chip millimetre wave (mmWave) radar based indoor mapping system targetted towards low-visibility environments to assist in emergency response. A unique feature of milliMap is that it only leverages a low-cost, off-the-shelf mmWave radar, but can reconstruct a dense grid map with accuracy comparable to lidar, as well as providing semantic annotations of objects on the map. milliMap makes two key technical contributions. First, it autonomously overcomes the sparsity and multi-path noise of mmWave signals by combining cross-modal supervision from a co-located lidar during training and the strong geometric priors of indoor spaces. Second, it takes the spectral response of mmWave reflections as features to robustly identify different types of objects e.g. doors, walls etc. Extensive experiments in different indoor environments show that milliMap can achieve a map reconstruction error less than 0.2m and classify key semantics with an accuracy of ~ 90%, whilst operating through dense smoke. Xiaoxuan Lu 0001, Stefano Rosa, Peijun Zhao, Bing Wang 0013, Changhao Chen, John A. Stankovic, Agathoniki Trigoni, Andrew Markham |
MobiSys | 1 |
| 2020 | Indoor positioning system in visually-degraded environments with millimetre-wave radar and inertial sensors: demo abstractabstractPositional estimation is of great importance in the public safety sector. Emergency responders such as fire fighters, medical rescue teams, and the police will all benefit from a resilient positioning system to deliver safe and effective emergency services. Unfortunately, satellite navigation (e.g., GPS) offers limited coverage in indoor environments. It is also not possible to rely on infrastructure based solutions. To this end, wearable sensor-aided navigation techniques, such as those based on camera and Inertial Measurement Units (IMU), have recently emerged recently as an accurate, infrastructure-free solution. Together with an increase in the computational capabilities of mobile devices, motion estimation can be performed in real-time. In this demonstration, we present a real-time indoor positioning system which fuses millimetre-wave (mmWave) radar and IMU data via deep sensor fusion. We employ mmWave radar rather than an RGB camera as it provides better robustness to visual degradation (e.g., smoke, darkness, etc.) while at the same time requiring lower computational resources to enable runtime computation. We implemented the sensor system on a handheld device and a mobile computer running at 10 FPS to track a user inside an apartment. Good accuracy and resilience were exhibited even in poorly illuminated scenes. Zhuangzhuang Dai, Muhamad Risqi Utama Saputra, Xiaoxuan Lu 0001, Agathoniki Trigoni, Andrew Markham |
SenSys | 3 |
| 2020 | milliEgo: single-chip mmWave radar aided egomotion estimation via deep sensor fusionabstractRobust and accurate trajectory estimation of mobile agents such as people and robots is a key requirement for providing spatial awareness for emerging capabilities such as augmented reality or autonomous interaction. Although currently dominated by optical techniques e.g., visual-inertial odometry these suffer from challenges with scene illumination or featureless surfaces. As an alternative, we propose milliEgo, a novel deep-learning approach to robust egomotion estimation which exploits the capabilities of low-cost mm Wave radar. Although mmWave radar has a fundamental advantage over monocular cameras of being metric i.e., providing absolute scale or depth, current single chip solutions have limited and sparse imaging resolution, making existing point-cloud registration techniques brittle. We propose a new architecture that is optimized for solving this challenging pose transformation problem. Secondly, to robustly fuse mmWave pose estimates with additional sensors, e.g. inertial or visual sensors we introduce a mixed attention approach to deep fusion. Through extensive experiments, we demonstrate our proposed system is able to achieve 1.3% 3D error drift and generalizes well to unseen environments. We also show that the neural architecture can be made highly efficient and suitable for real-time embedded applications. Xiaoxuan Lu 0001, Muhamad Risqi Utama Saputra, Peijun Zhao, Yasin Almalioglu, Pedro Porto Buarque de Gusmão, Changhao Chen, Ke Sun 0012, Agathoniki Trigoni, Andrew Markham |
SenSys | 1 |
| 2020 | Nowhere to Hide: Cross-modal Identity Leakage between Biometrics and DevicesabstractAlong with the benefits of Internet of Things (IoT) come potential privacy risks, since billions of the connected devices are granted permission to track information about their users and communicate it to other parties over the Internet. Of particular interest to the adversary is the user identity which constantly plays an important role in launching attacks. While the exposure of a certain type of physical biometrics or device identity is extensively studied, the compound effect of leakage from both sides remains unknown in multi-modal sensing environments. In this work, we explore the feasibility of the compound identity leakage across cyber-physical spaces and unveil that co-located smart device IDs (e.g., smartphone MAC addresses) and physical biometrics (e.g., facial/vocal samples) are side channels to each other. It is demonstrated that our method is robust to various observation noise in the wild and an attacker can comprehensively profile victims in multi-dimension with nearly zero analysis effort. Two real-world experiments on different biometrics and device IDs show that the presented approach can compromise more than 70% of device IDs and harvests multiple biometric clusters with purity at the same time. Xiaoxuan Lu 0001, Yang Li 0073, Yuanbo Xiangli, Zhengxiong Li |
WWW | 1 |
| 2020 | Deep-Learning-Based Pedestrian Inertial Navigation: Methods, Data Set, and On-Device InferenceabstractModern inertial measurements units (IMUs) are small, cheap, energy efficient, and widely employed in smart devices and mobile robots. Exploiting inertial data for accurate and reliable pedestrian navigation supports is a key component for emerging Internet of Things applications and services. Recently, there has been a growing interest in applying deep neural networks (DNNs) to motion sensing and location estimation. However, the lack of sufficient labelled data for training and evaluating architecture benchmarks has limited the adoption of DNNs in IMU-based tasks. In this article, we present and release the Oxford Inertial Odometry Data Set (OxIOD), a first-of-its-kind public data set for deep-learning-based inertial navigation research with fine-grained ground truth on all sequences. Furthermore, to enable more efficient inference at the edge, we propose a novel lightweight framework to learn and reconstruct pedestrian trajectories from raw IMU data. Extensive experiments show the effectiveness of our data set and methods in achieving accurate data-driven pedestrian inertial navigation on resource-constrained devices. Changhao Chen, Peijun Zhao, Xiaoxuan Lu 0001, Wei Wang 0226, Andrew Markham, Agathoniki Trigoni |
IEEE Internet Things J. | 3 |
| 2019 | MotionTransformer: Transferring Neural Inertial Tracking between DomainsabstractInertial information processing plays a pivotal role in egomotion awareness for mobile agents, as inertial measurements are entirely egocentric and not environment dependent. However, they are affected greatly by changes in sensor placement/orientation or motion dynamics, and it is infeasible to collect labelled data from every domain. To overcome the challenges of domain adaptation on long sensory sequences, we propose MotionTransformer - a novel framework that extracts domain-invariant features of raw sequences from arbitrary domains, and transforms to new domains without any paired data. Through the experiments, we demonstrate that it is able to efficiently and effectively convert the raw sequence from a new unlabelled target domain into an accurate inertial trajectory, benefiting from the motion knowledge transferred from the labelled source domain. We also conduct real-world experiments to show our framework can reconstruct physically meaningful trajectories from raw IMU measurements obtained with a standard mobile phone in various attachments. Changhao Chen, Yishu Miao, Xiaoxuan Lu 0001, Linhai Xie, Phil Blunsom, Andrew Markham, Agathoniki Trigoni |
AAAI | 3 |
| 2019 | Selective Sensor Fusion for Neural Visual-Inertial OdometryabstractDeep learning approaches for Visual-Inertial Odometry (VIO) have proven successful, but they rarely focus on incorporating robust fusion strategies for dealing with imperfect input sensory data. We propose a novel end-to-end selective sensor fusion framework for monocular VIO, which fuses monocular images and inertial measurements in order to estimate the trajectory whilst improving robustness to real-life issues, such as missing and corrupted data or bad sensor synchronization. In particular, we propose two fusion modalities based on different masking strategies: deterministic soft fusion and stochastic hard fusion, and we compare with previously proposed direct fusion baselines. During testing, the network is able to selectively process the features of the available sensor modalities and produce a trajectory at scale. We present a thorough investigation on the performances on three public autonomous driving, Micro Aerial Vehicle (MAV) and hand-held VIO datasets. The results demonstrate the effectiveness of the fusion strategies, which offer better performances compared to direct fusion, particularly in presence of corrupted data. In addition, we study the interpretability of the fusion networks by visualising the masking layers in different scenarios and with varying data corruption, revealing interesting correlations between the fusion networks and imperfect sensory input data. Changhao Chen, Stefano Rosa, Yishu Miao, Xiaoxuan Lu 0001, Andrew Markham, Agathoniki Trigoni |
CVPR | 4 |
| 2019 | mID: Tracking and Identifying People with Millimeter Wave RadarabstractThe key to offering personalised services in smart spaces is knowing where a particular person is with a high degree of accuracy. Visual tracking is one such solution, but concerns arise around the potential leakage of raw video information and many people are not comfortable accepting cameras in their homes or workplaces. We propose a human tracking and identification system (mID) based on millimeter wave radar which has a high tracking accuracy, without being visually compromising. Unlike competing techniques based on WiFi Channel State Information (CSI), it is capable of tracking and identifying multiple people simultaneously. Using a lowcost, commercial, off-the-shelf radar, we first obtain sparse point clouds and form temporally associated trajectories. With the aid of a deep recurrent network, we identify individual users. We evaluate and demonstrate our system across a variety of scenarios, showing median position errors of 0.16 m and identification accuracy of 89% for 12 people. Peijun Zhao, Xiaoxuan Lu 0001, Changhao Chen, Wei Wang 0226, Agathoniki Trigoni, Andrew Markham |
DCOSS | 2 |
| 2019 | Autonomous Learning for Face Recognition in the Wild via Ambient Wireless CuesabstractFacial recognition is a key enabling component for emerging Internet of Things (IoT) services such as smart homes or responsive offices. Through the use of deep neural networks, facial recognition has achieved excellent performance. However, this is only possibly when trained with hundreds of images of each user in different viewing and lighting conditions. Clearly, this level of effort in enrolment and labelling is impossible for wide-spread deployment and adoption. Inspired by the fact that most people carry smart wireless devices with them, e.g. smartphones, we propose to use this wireless identifier as a supervisory label. This allows us to curate a dataset of facial images that are unique to a certain domain e.g. a set of people in a particular office. This custom corpus can then be used to finetune existing pre-trained models e.g. FaceNet. However, due to the vagaries of wireless propagation in buildings, the supervisory labels are noisy and weak. We propose a novel technique, AutoTune, which learns and refines the association between a face and wireless identifier over time, by increasing the inter-cluster separation and minimizing the intra-cluster distance. Through extensive experiments with multiple users on two sites, we demonstrate the ability of AutoTune to design an environment-specific, continually evolving facial recognition system with entirely no user effort. Xiaoxuan Lu 0001, Xuan Kan, Bowen Du 0002, Changhao Chen, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni, John A. Stankovic |
WWW | 1 |
| 2019 | Autonomous Learning of Speaker Identity and WiFi Geofence From Noisy Sensor DataabstractA fundamental building block toward intelligent environments is the ability to understand who is present in a certain area. A ubiquitous way of detecting this is to exploit unique vocal characteristics as people interact with one another in common spaces. However, manually enrolling users into a biometric database is time-consuming and not robust to vocal deviations over time. Instead, consider audio features sampled during a meeting, yielding a noisy set of possible voiceprints. With a number of meetings and knowledge of participation, e.g., sniffed wireless media access control (MAC) addresses, can we learn to associate a specific identity with a particular voiceprint? To address this problem, this paper advocates an Internet of Things (IoT) solution and proposes to use co-located WiFi as supervisory weak labels to automatically bootstrap the labeling process. In particular, a novel cross-modality labeling algorithm is proposed that jointly optimizes the clustering and association process, which solves the inherent mismatching issues arising from heterogeneous sensor data. At the same time, we further propose to reuse the labeled data to iteratively update wireless geofence models and curate device specific thresholds. The extensive experimental results from two different scenarios demonstrate that our proposed method is able to achieve twofold improvement in labeling compared with conventional methods and can achieve reliable speaker recognition in the wild. Xiaoxuan Lu 0001, Yuanbo Xiangli, Peijun Zhao, Changhao Chen, Agathoniki Trigoni, Andrew Markham |
IEEE Internet Things J. | 1 |
| 2019 | Semantic Place Understanding for Human-Robot Coexistence - Toward Intelligent WorkplacesabstractRecent introductions of robots to everyday scenarios have revealed unprecedented opportunities for collaboration and social interaction between robots and people. However, to date, such interactions are hampered by a significant challenge: having a semantic understanding of their environment. Even simple requirements, such as “a robot should always be in the kitchen when a person is there,” are difficult to implement without prior training. In this paper, we advocate that robot-people coexistence can be leveraged to enhance the semantic understanding of the shared environment and improve situation awareness. We propose a probabilistic framework that combines human activity sensor data generated by smart wearables with low-level localization data generated by robots. Based on this low-level information and leveraging colocation events between a user and a robot, it can reason about the two types of semantic information: first, semantic maps, i.e., the utility of each room and, second, space usage semantics, i.e., tracking humans and robots through rooms of different utilities. The proposed system relies on two-way sharing of information between the robot and the user. In the first phase, user activities indicative of room utility are inferred from wearable devices and shared with the robot, enabling it to gradually build a semantic map of the environment. In the second phase, via colocation events, the robot teaches the user device to recognize the type of room where they are colocated. Over time, robot and user become increasingly independent and capable of semantic scene understanding. Stefano Rosa, Andrea Patanè, Xiaoxuan Lu 0001, Agathoniki Trigoni |
IEEE Trans. Hum. Mach. Syst. | 3 |
| 2019 | Efficient Indoor Positioning with Visual Experiences via Lifelong LearningabstractPositioning with visual sensors in indoor environments has many advantages: it doesn’t require infrastructure or accurate maps, and is more robust and accurate than other modalities such as WiFi. However, one of the biggest hurdles that prevents its practical application on mobile devices is the time-consuming visual processing pipeline. To overcome this problem, this paper proposes a novel lifelong learning approach to enable efficient and real-time visual positioning. We explore the fact that when following a previous visual experience for multiple times, one could gradually discover clues on how to traverse it with much less effort, e.g., which parts of the scene are more informative, and what kind of visual elements we should expect. Such second-order information is recorded as parameters, which provide key insights of the context and empower our system to dynamically optimise itself to stay localised with minimum cost. We implement the proposed approach on an array of mobile and wearable devices, and evaluate its performance in two indoor settings. Experimental results show our approach can reduce the visual processing time up to two orders of magnitude, while achieving sub-metre positioning accuracy. Hongkai Wen 0001, Ronald Clark, Sen Wang 0002, Xiaoxuan Lu 0001, Bowen Du 0002, Wen Hu 0001, Agathoniki Trigoni |
IEEE Trans. Mob. Comput. | 4 |
| 2018 | IONet: Learning to Cure the Curse of Drift in Inertial OdometryabstractInertial sensors play a pivotal role in indoor localization, which in turn lays the foundation for pervasive personal applications. However, low-cost inertial sensors, as commonly found in smartphones, are plagued by bias and noise, which leads to unbounded growth in error when accelerations are double integrated to obtain displacement. Small errors in state estimation propagate to make odometry virtually unusable in a matter of seconds. We propose to break the cycle of continuous integration, and instead segment inertial data into independent windows. The challenge becomes estimating the latent states of each window, such as velocity and orientation, as these are not directly observable from sensor data. We demonstrate how to formulate this as an optimization problem, and show how deep recurrent neural networks can yield highly accurate trajectories, outperforming state-of-the-art shallow techniques, on a wide range of tests and attachments. In particular, we demonstrate that IONet can generalize to estimate odometry for non-periodic motion, such as a shopping trolley or baby-stroller, an extremely challenging task for existing techniques. Changhao Chen, Xiaoxuan Lu 0001, Andrew Markham, Agathoniki Trigoni |
AAAI | 2 |
| 2018 | Deepauth: in-situ authentication for smartwatches via deeply learned behavioural biometricsabstractThis paper proposes DeepAuth, an in-situ authentication framework that leverages the unique motion patterns when users entering passwords as behavioural biometrics. It uses a deep recurrent neural network to capture the subtle motion signatures during password input, and employs a novel loss function to learn deep feature representations that are robust to noise, unseen passwords, and malicious imposters even with limited training data. DeepAuth is by design optimised for resource constrained platforms, and uses a novel split-RNN architecture to slim inference down to run in real-time on off-the-shelf smartwatches. Extensive experiments with real-world data show that DeepAuth outperforms the state-of-the-art significantly in both authentication performance and cost, offering real-time authentication on a variety of smartwatches. Xiaoxuan Lu 0001, Bowen Du 0002, Peijun Zhao, Hongkai Wen 0001, Yiran Shen 0001, Andrew Markham, Agathoniki Trigoni |
UbiComp | 1 |
| 2018 | Simultaneous Localization and Mapping with Power Network Electromagnetic FieldabstractVarious sensing modalities have been exploited for indoor location sensing, each of which has well understood limitations, however. This paper presents a first systematic study on using the electromagnetic field (EMF) induced by a building's electric power network for simultaneous localization and mapping (SLAM). A basis of this work is a measurement study showing that the power network EMF sensed by either a customized sensor or smartphone's microphone as a side-channel sensor is spatially distinct and temporally stable. Based on this, we design a SLAM approach that can reliably detect loop closures based on EMF sensing results. With the EMF feature map constructed by SLAM, we also design an efficient online localization scheme for resource-constrained mobiles. Evaluation in three indoor spaces shows that the power network EMF is a promising modality for location sensing on mobile devices, which is able to run in real time and achieve sub-meter accuracy. Xiaoxuan Lu 0001, Yang Li 0147, Peijun Zhao, Changhao Chen, Linhai Xie, Hongkai Wen 0001, Rui Tan 0001, Agathoniki Trigoni |
MobiCom | 1 |
| 2018 | Automatic Face Recognition Adaptation via Ambient Wireless IdentifiersabstractFace recognition is a key enabling service for smart-spaces, allowing building management agents to easily monitor 'who is where', anticipating user needs and tailoring their local environment and experiences. Although facial recognition, especially through the use of deep neural networks, has achieved stellar performance over large datasets, the majority of approaches require supervised learning, that is, to be trained with tens or hundreds of images of users in different poses and lighting conditions. In this paper, we motivate that this enrollment effort is unnecessary if the smart-space has access to a wireless identifier e.g., through a smart-phone's MAC address. By learning and refining the noisy and weak association between a user's smart-phone and facial images, AutoTune can fine-tune a deep neural network to tailor it to the environment, users and conditions of a particular camera or set of cameras. Xiaoxuan Lu 0001, Peijun Zhao, Bowen Du 0002, Hongkai Wen 0001, Andrew Markham, Stefano Rosa, Agathoniki Trigoni |
SenSys | 1 |
| 2017 | SCAN: learning speaker identity from noisy sensor dataabstractSensor data acquired from multiple sensors simultaneously is featuring increasingly in our evermore pervasive world. Buildings can be made smarter and more efficient, spaces more responsive to users. A fundamental building block towards smart spaces is the ability to understand who is present in a certain area. A ubiquitous way of detecting this is to exploit the unique vocal features as people interact with one another. As an example, consider audio features sampled during a meeting, yielding a noisy set of possible voiceprints. With a number of meetings and knowledge of participation (e.g. through a calendar or MAC address), can we learn to associate a specific identity with a particular voiceprint? Obviously enrolling users into a biometric database is time-consuming and not robust to vocal deviations over time. To address this problem, the standard approach is to perform a clustering step (e.g. of audio data) followed by a data association step, when identity-rich sensor data is available. In this paper we show that this approach is not robust to noise in either type of sensor stream; to tackle this issue we propose a novel algorithm that jointly optimises the clustering and association process yielding up to three times higher identification precision than approaches that execute these steps sequentially. We demonstrate the performance benefits of our approach in two case studies, one with acoustic and MAC datasets that we collected from meetings in a non-residential building, and another from an online dataset from recorded radio interviews. Xiaoxuan Lu 0001, Hongkai Wen 0001, Sen Wang 0002, Andrew Markham, Agathoniki Trigoni |
IPSN | 1 |
| 2017 | Towards Self-supervised Face Labeling via Cross-modality AssociationabstractFace recognition has become the de facto authentication solution in a broad spectrum of applications, from smart buildings, to industrial monitoring and security services. However, in many of those real-world scenarios, tracking or identifying people with facial recognition is extremely challenging due to the variations in the environment such as lighting conditions, camera viewing angles and subject motion. For most of the state-of-the-art face recognition systems, they need to be trained on a large dataset containing a good variety of labelled face images to work well. However, collecting and manually labelling such datasets is difficult and time consuming, probably more so than developing the algorithms. In this paper, we propose a novel framework to automatically label user identities with their face images in smart spaces, exploiting the fact that the users tend to carry their smart devices while seen by the surveillance cameras. We evaluate our method on 10 users in a smart building setting, and the experimental results show that our method can achieve > 0.9 f1 score on average. Xiaoxuan Lu 0001, Xuan Kan, Stefano Rosa, Bowen Du 0002, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni |
SenSys | 1 |
| 2016 | Standardizing location fingerprints across heterogeneous mobile devices for indoor localizationabstractThe explosive proliferation of mobile devices and the popularity of social networks have spurred extensive demands on Location Based Services (LBSs) in recent decades. The IEEE 802.11 (WiFi) based Indoor Positioning Systems (IPSs) are gaining popularity because of the wide and ubiquitous availability of WiFi infrastructures in indoor environments. Most of IPSs are adopting the fingerprinting approach to mitigate pervasive indoor multipath effects. However, the heterogeneity of mobile devices significantly degrades the localization performance of the fingerprinting approach. In this paper, we apply the Procrustes analysis method to transform the WiFi received signal strengths (RSSs) to a new type of standard location fingerprints which are tolerant of the heterogeneity of various devices. Then, a robust indoor positioning algorithm based on the standardized location fingerprints and the weighted k nearest neighbor (WKN-N) method is proposed. Extensive experiments are carried out and show that the standardized location fingerprints and the proposed positioning system address the device heterogeneity issue satisfactorily. Han Zou, Baoqi Huang, Xiaoxuan Lu 0001, Hao Jiang 0008, Lihua Xie 0001 |
WCNC | 3 |
| 2016 | Robust occupancy inference with commodity WiFiabstractAccurate occupancy information of indoor environments is one of the key prerequisites for many pervasive and context-aware services, e.g. smart building/home systems. Some of the existing occupancy inference systems can achieve impressive accuracy, but they either require labour-intensive calibration phases, or need to install bespoke hardware such as CCTV cameras, which are privacy-intrusive by default. In this paper, we present the design and implementation of a practical end-to-end occupancy inference system, which requires minimum user effort, and is able to infer room-level occupancy accurately with commodity WiFi infrastructure. Depending on the needs of different occupancy information subscribers, our system is flexible enough to switch between snapshot estimation mode and continuous inference mode, to trade estimation accuracy for delay and communication cost. We evaluate the system on a hardware testbed deployed in a 600m2workspace with 25 occupants for 6 weeks. Experimental results show that the proposed system significantly outperforms competing systems in both inference accuracy and robustness. Xiaoxuan Lu 0001, Hongkai Wen 0001, Han Zou, Hao Jiang 0008, Lihua Xie 0001, Agathoniki Trigoni |
WiMob | 1 |
| 2016 | Robust Extreme Learning Machine With its Application to Indoor PositioningabstractThe increasing demands of location-based services have spurred the rapid development of indoor positioning system and indoor localization system interchangeably (IPSs). However, the performance of IPSs suffers from noisy measurements. In this paper, two kinds of robust extreme learning machines (RELMs), corresponding to the close-to-mean constraint, and the small-residual constraint, have been proposed to address the issue of noisy measurements in IPSs. Based on whether the feature mapping in extreme learning machine is explicit, we respectively provide random-hidden-nodes and kernelized formulations of RELMs by second order cone programming. Furthermore, the computation of the covariance in feature space is discussed. Simulations and real-world indoor localization experiments are extensively carried out and the results demonstrate that the proposed algorithms can not only improve the accuracy and repeatability, but also reduce the deviation and worst case error of IPSs compared with other baseline algorithms. Xiaoxuan Lu 0001, Han Zou, Hongming Zhou, Lihua Xie 0001, Guang-Bin Huang |
IEEE Trans. Cybern. | 1 |
| 2016 | A Robust Indoor Positioning System Based on the Procrustes Analysis and Weighted Extreme Learning MachineabstractIndoor positioning system (IPS) has become one of the most attractive research fields due to the increasing demands on location-based services (LBSs) in indoor environments. Various IPSs have been developed under different circumstances, and most of them adopt the fingerprinting technique to mitigate pervasive indoor multipath effects. However, the performance of the fingerprinting technique severely suffers from device heterogeneity existing across commercial off-the-shelf mobile devices (e.g., smart phones, tablet computers, etc.) and indoor environmental changes (e.g., the number, distribution and activities of people, the placement of furniture, etc.). In this paper, we transform the received signal strength (RSS) to a standardized location fingerprint based on the Procrustes analysis, and introduce a similarity metric, termed signal tendency index (STI), for matching standardized fingerprints. An analysis of the capability of the proposed STI to handle device heterogeneity and environmental changes is presented. We further develop a robust and precise IPS by integrating the merits of both the STI and weighted extreme learning machine (WELM). Finally, extensive experiments are carried out and a performance comparison with existing solutions verifies the superiority of the proposed IPS in terms of robustness to device heterogeneity. Han Zou, Baoqi Huang, Xiaoxuan Lu 0001, Hao Jiang 0008, Lihua Xie 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2014 | Extreme learning machine with dead zone and its application to WiFi based indoor positioningabstractExtreme learning machine (ELM) as an emergent technology has shown its good performance in regression applications as well as in large dataset classification applications. It has been broadly embedded in many applications due to its fast speed of computation and accuracy. How to make good use of machine learning techniques in Indoor Positioning System (IPS) is a hot research topic in recent years. Some existing IPSs have already adopted ELM, but it suffers from signal variation and environmental dynamics in indoor settings. In this paper, extreme learning machine with dead zone (DZ-ELM) is proposed to address this problem. The consistency of this approach should be applied is studied. Simulations are also conducted to compare the performance of DZ-ELM and ELM. Lastly, real-world experimental results show that the proposed algorithm can not only provide higher accuracy but also improve the repeatability of IPSs. Xiaoxuan Lu 0001, Chengpu Yu, Han Zou, Hao Jiang 0008, Lihua Xie 0001 |
ICARCV | 1 |