EDBT 2026 Demo / reviewers in the wild / expert
Changhao Chen
dblp:215/3933
· DBLP profile ↗
34ranked-venue papers
10as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 6 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 7 since 2021Computer networks · 7 · 2 first-author · 2 since 2021Systems, architecture and hardware · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PanoNav: Mapless Zero-Shot Object Navigation with Panoramic Scene Parsing and Dynamic MemoryabstractZero-shot object navigation (ZSON) in unseen environments remains a challenging problem for household robots, requiring strong perceptual understanding and decision-making capabilities. While recent methods leverage metric maps and Large Language Models (LLMs), they often depend on depth sensors or prebuilt maps, limiting the spatial reasoning ability of Multimodal Large Language Models (MLLMs). Mapless ZSON approaches have emerged to address this, but they typically make short-sighted decisions, leading to local deadlocks due to a lack of historical context. We propose PanoNav, a fully RGB-only, mapless ZSON framework that integrates a Panoramic Scene Parsing module to unlock the spatial parsing potential of MLLMs from panoramic RGB inputs, and a Memory-guided Decision-Making mechanism enhanced by a Dynamic Bounded Memory Queue to incorporate exploration history and avoid local deadlocks. Experiments on the public navigation benchmark show that PanoNav significantly outperforms representative baselines in both SR and SPL metrics. Qunchao Jin, Changhao Chen |
AAAI | 3 |
| 2025 | M2EIT: Multi-Domain Mixture of Experts for Robust Neural Inertial Tracking
Changhao Chen, Zhongchen Shi, Wei Chen 0092, Liang Xie 0012, Erwei Yin |
ICCV | 3 |
| 2025 | ThermalLoc: A Vision Transformer-Based Approach for Robust Thermal Camera Relocalization in Large-Scale EnvironmentsabstractThermal cameras capture environmental data through heat emission, a fundamentally different mechanism compared to visible light cameras, which rely on pinhole imaging. As a result, traditional visual relocalization methods designed for visible light images are not directly applicable to thermal images. Despite significant advancements in deep learning for camera relocalization, approaches specifically tailored for thermal camera-based relocalization remain underexplored. To address this gap, we introduce ThermalLoc, a novel end-to-end deep learning method for thermal image relocalization. ThermalLoc effectively extracts both local and global features from thermal images by integrating EfficientNet with Transformers, and performs absolute pose regression using two MLP networks. We evaluated ThermalLoc on both the publicly available thermal-odometry dataset and our own dataset. The results demonstrate that ThermalLoc outperforms existing representative methods employed for thermal camera relocalization, including AtLoc, MapNet, PoseNet, and RobustLoc, achieving superior accuracy and robustness. Yangtao Meng, Xianfei Pan, Changhao Chen |
IROS | 5 |
| 2025 | DynaNav: Dynamic Feature and Layer Selection for Efficient Visual NavigationabstractVisual navigation is essential for robotics and embodied AI. However, existing foundation models, particularly those with transformer decoders, suffer from high computational overhead and lack interpretability, limiting their deployment on edge devices. To address this, we propose DynaNav, a Dynamic Visual Navigation framework that adapts feature and layer selection based on scene complexity. It employs a trainable hard feature selector for sparse operations, enhancing efficiency and interpretability. Additionally, we integrate feature selection into an early-exit mechanism, with Bayesian Optimization determining optimal exit thresholds to reduce computational cost. Extensive experiments in real-world-based datasets and simulated environments demonstrate the effectiveness of DynaNav. Compared to ViNT, DynaNav achieves a $2.6\times$ reduction in FLOPs, 42.3% lower inference time, and 32.8% lower memory usage while improving navigation performance across four public datasets. Changhao Chen |
NeurIPS | 2 |
| 2025 | LangLoc: Language-Driven Localization via Formatted Spatial Description GenerationabstractExisting localization methods commonly employ vision to perceive scene and achieve localization in GNSS-denied areas, yet they often struggle in environments with complex lighting conditions, dynamic objects or privacy-preserving areas. Humans possess the ability to describe various scenes using natural language, effectively inferring their location by leveraging the rich semantic information in these descriptions. Harnessing language presents a potential solution for robust localization. Thus, this study introduces a new task, Language-driven Localization, and proposes a novel localization framework, LangLoc, which determines the user's position and orientation through textual descriptions. Given the diversity of natural language descriptions, we first design a Spatial Description Generator (SDG), foundational to LangLoc, which extracts and combines the position and attribute information of objects within a scene to generate uniformly formatted textual descriptions. SDG eliminates the ambiguity of language, detailing the spatial layout and object relations of the scene, providing a reliable basis for localization. With generated descriptions, LangLoc effortlessly achieves language-only localization using text encoder and pose regressor. Furthermore, LangLoc can add one image to text input, achieving mutual optimization and feature adaptive fusion across modalities through two modality-specific encoders, cross-modal fusion, and multimodal joint learning strategies. This enhances the framework's capability to handle complex scenes, achieving more accurate localization. Extensive experiments on the Oxford RobotCar, 4-Seasons, and Virtual Gallery datasets demonstrate LangLoc's effectiveness in both language-only and visual-language localization across various outdoor and indoor scenarios. Notably, LangLoc achieves noticeable performance gains when using both text and image inputs in challenging conditions such as overexposure, low lighting, and occlusions, showcasing its superior robustness. Changhao Chen, Kaige Li, Yuan Xiong, Xiaochun Cao, Zhong Zhou |
IEEE Trans. Image Process. | 2 |
| 2025 | Learning Selective Sensor Fusion for State EstimationabstractAutonomous vehicles and mobile robotic systems are typically equipped with multiple sensors to provide redundancy. By integrating the observations from different sensors, these mobile agents are able to perceive the environment and estimate system states, e.g., locations and orientations. Although deep learning (DL) approaches for multimodal odometry estimation and localization have gained traction, they rarely focus on the issue of robust sensor fusion-a necessary consideration to deal with noisy or incomplete sensor observations in the real world. Moreover, current deep odometry models suffer from a lack of interpretability. To this extent, we propose SelectFusion, an end-to-end selective sensor fusion module that can be applied to useful pairs of sensor modalities, such as monocular images and inertial measurements, depth images, and light detection and ranging (LIDAR) point clouds. Our model is a uniform framework that is not restricted to specific modality or task. During prediction, the network is able to assess the reliability of the latent features from different sensor modalities and to estimate trajectory at both scale and global pose. In particular, we propose two fusion modules-a deterministic soft fusion and a stochastic hard fusion-and offer a comprehensive study of the new strategies compared with trivial direct fusion. We extensively evaluate all fusion strategies both on public datasets and on progressively degraded datasets that present synthetic occlusions, noisy and missing data, and time misalignment between sensors, and we investigate the effectiveness of the different fusion strategies in attending the most reliable features, which in itself provides insights into the operation of the various models. Changhao Chen, Stefano Rosa, Xiaoxuan Lu 0001, Bing Wang 0013, Agathoniki Trigoni, Andrew Markham |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | EffLoc: Lightweight Vision Transformer for Efficient 6-DOF Camera RelocalizationabstractCamera relocalization is pivotal in computer vision, with applications in AR, drones, robotics, and autonomous driving. It estimates 3D camera position and orientation (6-DoF) from images. Unlike traditional methods like SLAM, recent strides use deep learning for direct end-to-end pose estimation. We propose EffLoc, a novel efficient Vision Transformer for single-image camera relocalization. EffLoc’s hierarchical layout, memory-bound self-attention, and feed-forward layers boost memory efficiency and inter-channel communication. Our introduced sequential group attention (SGA) module enhances computational efficiency by diversifying input features, reducing redundancy, and expanding model capacity. EffLoc excels in efficiency and accuracy, outperforming prior methods, such as AtLoc and MapNet. It thrives on large-scale outdoor car-driving scenario, ensuring simplicity, end-to-end trainability, and eliminating handcrafted loss functions. Zhendong Xiao, Changhao Chen, Wu Wei 0001 |
ICRA | 2 |
| 2024 | HSTR: Hierarchical Scene Transformer for Multi-agent Trajectory PredictionabstractTrajectory prediction plays a pivotal role in the autonomous driving systems. Existing methods generally employ agent-centric or scene-centric approaches to represent driving scenarios. However, these methods introduce significant redundant computations or pose losses, resulting in suboptimal prediction efficiency and accuracy. To tackle these problems, a novel multi-target trajectory prediction model, named Hierarchical Scene TRansformer (HSTR), is introduced. The driving scene is decomposed into two independent components by HSTR: global and local. In the global part, global interaction information is established and shared among all predicted agents, thereby reducing redundant computations. In the local part, an individual reference frame is established for each vehicle to eliminate the impact of pose variations and extract temporal features. Moreover, an adaptive anchor point generation method is proposed to address the challenge of capturing future modalities for vehicles. This method dynamically generates corresponding anchor points based on different driving scenarios to guide the prediction of trajectories across various modalities. The model performance is verified on the argoverse1 and argoverse2 datasets, and the experimental results demonstrate that competitive performance is achieved by HSTR in terms of efficiency and precision compared to the state-of-the-art methods. Shuaiqi Fu, Xiaoyang Luo, Changhao Chen, Yanan Zhao 0004, Huachun Tan |
IV | 4 |
| 2024 | Concertorl: A reinforcement learning approach for finite-time single-life enhanced control and its application to direct-drive tandem-wing experiment platforms
Bifeng Song, Changhao Chen, Xinyu Lang |
Appl. Intell. | 3 |
| 2024 | Drone-NeRF: Efficient NeRF based 3D scene reconstruction for large-scale drone survey
Bing Wang 0013, Changhao Chen |
Image Vis. Comput. | 3 |
| 2024 | Deep Learning for Inertial Positioning: A SurveyabstractInertial sensors are widely utilized in smartphones, drones, vehicles, and wearable devices, playing a crucial role in enabling ubiquitous and reliable localization. Inertial sensor-based positioning is essential in various applications, including personal navigation, location-based security, and human-device interaction. However, low-cost MEMS inertial sensors’ measurements are inevitably corrupted by various error sources, leading to unbounded drifts when integrated doubly in traditional inertial navigation algorithms, subjecting inertial positioning to the problem of error drifts. In recent years, with the rapid increase in sensor data and computational power, deep learning techniques have been developed, sparking significant research into addressing the problem of inertial positioning. Relevant literature in this field spans across mobile computing, robotics, and machine learning. In this article, we provide a comprehensive review of deep learning-based inertial positioning and its applications in tracking pedestrians, drones, vehicles, and robots. We connect efforts from different fields and discuss how deep learning can be applied to address issues such as sensor calibration, positioning error drift reduction, and multi-sensor fusion. This article aims to attract readers from various backgrounds, including researchers and practitioners interested in the potential of deep learning-based techniques to solve inertial positioning problems. Our review demonstrates the exciting possibilities that deep learning brings to the table and provides a roadmap for future research in this field. Changhao Chen, Xianfei Pan |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Deep Learning for Visual Localization and Mapping: A SurveyabstractDeep-learning-based localization and mapping approaches have recently emerged as a new research direction and receive significant attention from both industry and academia. Instead of creating hand-designed algorithms based on physical models or geometric theories, deep learning solutions provide an alternative to solve the problem in a data-driven way. Benefiting from the ever-increasing volumes of data and computational power on devices, these learning methods are fast evolving into a new area that shows potential to track self-motion and estimate environmental models accurately and robustly for mobile agents. In this work, we provide a comprehensive survey and propose a taxonomy for the localization and mapping methods using deep learning. This survey aims to discuss two basic questions: whether deep learning is promising for localization and mapping, and how deep learning should be applied to solve this problem. To this end, a series of localization and mapping topics are investigated, from the learning-based visual odometry and global relocalization to mapping, and simultaneous localization and mapping (SLAM). It is our hope that this survey organically weaves together the recent works in this vein from robotics, computer vision, and machine learning communities and serves as a guideline for future researchers to apply deep learning to tackle the problem of visual localization and mapping. Changhao Chen, Bing Wang 0013, Xiaoxuan Lu 0001, Agathoniki Trigoni, Andrew Markham |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | DevNet: Self-supervised Monocular Depth Learning via Density Volume Construction
Kaichen Zhou, Lanqing Hong, Changhao Chen, Hang Xu 0004, Chaoqiang Ye, Qingyong Hu, Zhenguo Li |
ECCV (39) | 3 |
| 2022 | Multimodal Learning of Audio-Visual Speech Recognition with Liquid State Machine
Xuhu Yu, Changhao Chen, Junbo Tie, Shasha Guo 0001 |
ICONIP (6) | 3 |
| 2022 | An improved point feature-based sparse stereo visionabstractAbstract Since the limitation on the onboard equipment, the sparse stereo vision is becoming a suitable choice for the deployment of micro air vehicles (MAV) and small robots. However, for the point feature‐based sparse stereo, most of the current stereo algorithms ignore the similarity between feature points, so it is hard to achieve high accuracy. In addition, the problem of clustered feature distribution will still affect the performance of point feature‐based algorithms in the application. To make up for these deficiencies, the authors propose an improved features from accelerated segment test (FAST) feature detector to suppress the point detection in complex texture regions. Most importantly, the authors present a novel census transform (CT)‐based algorithm that contains two encoders ‘texture orientation’ and ‘texture gradient’ to get a more efficient census bit string for the feature point. Instead of randomly selecting pixels to calculate the bit string, we combine the texture characteristics of the census windows where feature points are located. Compared with the original CT, the processing speed of our method is improved, and the average error of our method is reduced by 18.05%. The evaluation results show the presented improved point feature‐based sparse stereo algorithm has a great value in engineering applications. Changhao Chen, Bifeng Song, Shuhui Bu |
IET Image Process. | 1 |
| 2021 | VMLoc: Variational Fusion For Learning-Based Multimodal Camera LocalizationabstractRecent learning-based approaches have achieved impressive results in the field of single-shot camera localization. However, how best to fuse multiple modalities (e.g., image and depth) and to deal with degraded or missing input are less well studied. In particular, we note that previous approaches towards deep fusion do not perform significantly better than models employing a single modality. We conjecture that this is because of the naive approaches to feature space fusion through summation or concatenation which do not take into account the different strengths of each modality. To address this, we propose an end-to-end framework, termed VMLoc, to fuse different sensor inputs into a common latent space through a variational Product-of-Experts (PoE) followed by attention-based fusion. Unlike previous multimodal variational works directly adapting the objective function of vanilla variational auto-encoder, we show how camera localization can be accurately estimated through an unbiased objective function based on importance weighting. Our model is extensively evaluated on RGB-D datasets and the results prove the efficacy of our model. The source code is available at https://github.com/Zalex97/VMLoc. Kaichen Zhou, Changhao Chen, Bing Wang 0013, Muhamad Risqi Utama Saputra, Agathoniki Trigoni, Andrew Markham |
AAAI | 2 |
| 2021 | P2-Net: Joint Description and Detection of Local Features for Pixel and Point MatchingabstractAccurately describing and detecting 2D and 3D key-points is crucial to establishing correspondences across images and point clouds. Despite a plethora of learning-based 2D or 3D local feature descriptors and detectors having been proposed, the derivation of a shared descriptor and joint keypoint detector that directly matches pixels and points remains under-explored by the community. This work takes the initiative to establish fine-grained correspondences between 2D images and 3D point clouds. In order to directly match pixels and points, a dual fully-convolutional framework is presented that maps 2D and 3D inputs into a shared latent representation space to simultaneously describe and detect keypoints. Furthermore, an ultra-wide reception mechanism and a novel loss function are designed to mitigate the intrinsic information variations between pixel and point local regions. Extensive experimental results demonstrate that our framework shows competitive performance in fine-grained matching between images and point clouds and achieves state-of-the-art results for the task of indoor visual localization. Our source code is available at https://github.com/BingCS/P2-Net. Bing Wang 0013, Changhao Chen, Zhaopeng Cui, Jie Qin 0004, Xiaoxuan Lu 0001, Zhengdi Yu, Peijun Zhao, Zhen Dong 0005, Fan Zhu 0001, Agathoniki Trigoni, Andrew Markham |
ICCV | 2 |
| 2021 | Human tracking and identification through a millimeter wave radar
Peijun Zhao, Xiaoxuan Lu 0001, Changhao Chen, Wei Wang 0226, Agathoniki Trigoni, Andrew Markham |
Ad Hoc Networks | 4 |
| 2021 | Deep Neural Network Based Inertial Odometry Using Low-Cost Inertial Measurement UnitsabstractInertial measurement units (IMUs) have emerged as an essential component in many of today's indoor navigation solutions due to their low cost and ease of use. However, despite many attempts for reducing the error growth of navigation systems based on commercial-grade inertial sensors, there is still no satisfactory solution that produces navigation estimates with long-time stability in widely differing conditions. This paper proposes to break the cycle of continuous integration used in traditional inertial algorithms, formulate it as an optimization problem, and explore the use of deep recurrent neural networks for estimating the displacement of a user over a specified time window. By training the deep neural network using inertial measurements and ground truth displacement data, it is possible to learn both motion characteristics and systematic error drift. As opposed to established context-aided inertial solutions, the proposed method is not dependent on either fixed sensor positions or periodic motion patterns. It can reconstruct accurate trajectories directly from raw inertial measurements, and predict the corresponding uncertainty to show model confidence. Extensive experimental evaluations demonstrate that the neural network produces position estimates with high accuracy for several different attachments, users, sensors, and motion types. As a particular demonstration of its flexibility, our deep inertial solutions can estimate trajectories for non-periodic motion, such as the shopping trolley tracking. Further more, it works in highly dynamic conditions, such as running, remaining extremely challenging for current techniques. Changhao Chen, Xiaoxuan Lu 0001, Johan Wahlström, Andrew Markham, Agathoniki Trigoni |
IEEE Trans. Mob. Comput. | 1 |
| 2021 | DynaNet: Neural Kalman Dynamical Model for Motion Estimation and PredictionabstractDynamical models estimate and predict the temporal evolution of physical systems. State-space models (SSMs) in particular represent the system dynamics with many desirable properties, such as being able to model uncertainty in both the model and measurements, and optimal (in the Bayesian sense) recursive formulations, e.g., the Kalman filter. However, they require significant domain knowledge to derive the parametric form and considerable hand tuning to correctly set all the parameters. Data-driven techniques, e.g., recurrent neural networks, have emerged as compelling alternatives to SSMs with wide success across a number of challenging tasks, in part due to their impressive capability to extract relevant features from rich inputs. They, however, lack interpretability and robustness to unseen conditions. Thus, data-driven models are hard to be applied in safety-critical applications, such as self-driving vehicles. In this work, we present DynaNet, a hybrid deep learning and time-varying SSM, which can be trained end-to-end. Our neural Kalman dynamical model allows us to exploit the relative merits of both SSM and deep neural networks. We demonstrate its effectiveness in the estimation and prediction on a number of physically challenging tasks, including visual odometry, sensor fusion for visual-inertial navigation, and motion prediction. In addition, we show how DynaNet can indicate failures through investigation of properties, such as the rate of innovation (Kalman gain). Changhao Chen, Xiaoxuan Lu 0001, Bing Wang 0013, Agathoniki Trigoni, Andrew Markham |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Learning With Stochastic Guidance for Robot NavigationabstractDue to the sparse rewards and high degree of environmental variation, reinforcement learning approaches, such as deep deterministic policy gradient (DDPG), are plagued by issues of high variance when applied in complex real-world environments. We present a new framework for overcoming these issues by incorporating a stochastic switch, allowing an agent to choose between high- and low-variance policies. The stochastic switch can be jointly trained with the original DDPG in the same framework. In this article, we demonstrate the power of the framework in a navigation task, where the robot can dynamically choose to learn through exploration or to use the output of a heuristic controller as guidance. Instead of starting from completely random actions, the navigation capability of a robot can be quickly bootstrapped by several simple independent controllers. The experimental results show that with the aid of stochastic guidance, we are able to effectively and efficiently train DDPG navigation policies and achieve significantly better performance than state-of-the-art baseline models. Linhai Xie, Yishu Miao, Sen Wang 0002, Phil Blunsom, Zhihua Wang 0005, Changhao Chen, Andrew Markham, Agathoniki Trigoni |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2020 | AtLoc: Attention Guided Camera LocalizationabstractDeep learning has achieved impressive results in camera localization, but current single-image techniques typically suffer from a lack of robustness, leading to large outliers. To some extent, this has been tackled by sequential (multi-images) or geometry constraint approaches, which can learn to reject dynamic objects and illumination conditions to achieve better performance. In this work, we show that attention can be used to force the network to focus on more geometrically robust objects and features, achieving state-of-the-art performance in common benchmark, even if using only a single image as input. Extensive experimental evidence is provided through public indoor and outdoor datasets. Through visualization of the saliency maps, we demonstrate how the network learns to reject dynamic objects, yielding superior global camera pose regression performance. The source code is avaliable at https://github.com/BingCS/AtLoc. Bing Wang 0013, Changhao Chen, Xiaoxuan Lu 0001, Peijun Zhao, Agathoniki Trigoni, Andrew Markham |
AAAI | 2 |
| 2020 | Heart Rate Sensing with a Robot Mounted mmWave RadarabstractHeart rate monitoring at home is a useful metric for assessing health e.g. of the elderly or patients in post-operative recovery. Although non-contact heart rate monitoring has been widely explored, typically using a static, wall-mounted device, measurements are limited to a single room and sensitive to user orientation and position. In this work, we propose mBeats, a robot mounted millimeter wave (mmWave) radar system that provide periodic heart rate measurements under different user poses, without interfering in a users daily activities. mBeats contains a mmWave servoing module that adaptively adjusts the sensor angle to the best reflection pro le. Furthermore, mBeats features a deep neural network predictor, which can estimate heart rate from the lower leg and additionally provides estimation uncertainty. Through extensive experiments, we demonstrate accurate and robust operation of mBeats in a range of scenarios. We believe by integrating mobility and adaptability, mBeats can empower many down-stream healthcare applications at home, such as palliative care, post-operative rehabilitation and telemedicine. Peijun Zhao, Xiaoxuan Lu 0001, Bing Wang 0013, Changhao Chen, Linhai Xie, Agathoniki Trigoni, Andrew Markham |
ICRA | 4 |
| 2020 | See through smoke: robust indoor mapping with low-cost mmWave radarabstractThis paper presents the design, implementation and evaluation of milliMap, a single-chip millimetre wave (mmWave) radar based indoor mapping system targetted towards low-visibility environments to assist in emergency response. A unique feature of milliMap is that it only leverages a low-cost, off-the-shelf mmWave radar, but can reconstruct a dense grid map with accuracy comparable to lidar, as well as providing semantic annotations of objects on the map. milliMap makes two key technical contributions. First, it autonomously overcomes the sparsity and multi-path noise of mmWave signals by combining cross-modal supervision from a co-located lidar during training and the strong geometric priors of indoor spaces. Second, it takes the spectral response of mmWave reflections as features to robustly identify different types of objects e.g. doors, walls etc. Extensive experiments in different indoor environments show that milliMap can achieve a map reconstruction error less than 0.2m and classify key semantics with an accuracy of ~ 90%, whilst operating through dense smoke. Xiaoxuan Lu 0001, Stefano Rosa, Peijun Zhao, Bing Wang 0013, Changhao Chen, John A. Stankovic, Agathoniki Trigoni, Andrew Markham |
MobiSys | 5 |
| 2020 | milliEgo: single-chip mmWave radar aided egomotion estimation via deep sensor fusionabstractRobust and accurate trajectory estimation of mobile agents such as people and robots is a key requirement for providing spatial awareness for emerging capabilities such as augmented reality or autonomous interaction. Although currently dominated by optical techniques e.g., visual-inertial odometry these suffer from challenges with scene illumination or featureless surfaces. As an alternative, we propose milliEgo, a novel deep-learning approach to robust egomotion estimation which exploits the capabilities of low-cost mm Wave radar. Although mmWave radar has a fundamental advantage over monocular cameras of being metric i.e., providing absolute scale or depth, current single chip solutions have limited and sparse imaging resolution, making existing point-cloud registration techniques brittle. We propose a new architecture that is optimized for solving this challenging pose transformation problem. Secondly, to robustly fuse mmWave pose estimates with additional sensors, e.g. inertial or visual sensors we introduce a mixed attention approach to deep fusion. Through extensive experiments, we demonstrate our proposed system is able to achieve 1.3% 3D error drift and generalizes well to unseen environments. We also show that the neural architecture can be made highly efficient and suitable for real-time embedded applications. Xiaoxuan Lu 0001, Muhamad Risqi Utama Saputra, Peijun Zhao, Yasin Almalioglu, Pedro Porto Buarque de Gusmão, Changhao Chen, Ke Sun 0012, Agathoniki Trigoni, Andrew Markham |
SenSys | 6 |
| 2020 | Deep-Learning-Based Pedestrian Inertial Navigation: Methods, Data Set, and On-Device InferenceabstractModern inertial measurements units (IMUs) are small, cheap, energy efficient, and widely employed in smart devices and mobile robots. Exploiting inertial data for accurate and reliable pedestrian navigation supports is a key component for emerging Internet of Things applications and services. Recently, there has been a growing interest in applying deep neural networks (DNNs) to motion sensing and location estimation. However, the lack of sufficient labelled data for training and evaluating architecture benchmarks has limited the adoption of DNNs in IMU-based tasks. In this article, we present and release the Oxford Inertial Odometry Data Set (OxIOD), a first-of-its-kind public data set for deep-learning-based inertial navigation research with fine-grained ground truth on all sequences. Furthermore, to enable more efficient inference at the edge, we propose a novel lightweight framework to learn and reconstruct pedestrian trajectories from raw IMU data. Extensive experiments show the effectiveness of our data set and methods in achieving accurate data-driven pedestrian inertial navigation on resource-constrained devices. Changhao Chen, Peijun Zhao, Xiaoxuan Lu 0001, Wei Wang 0226, Andrew Markham, Agathoniki Trigoni |
IEEE Internet Things J. | 1 |
| 2019 | MotionTransformer: Transferring Neural Inertial Tracking between DomainsabstractInertial information processing plays a pivotal role in egomotion awareness for mobile agents, as inertial measurements are entirely egocentric and not environment dependent. However, they are affected greatly by changes in sensor placement/orientation or motion dynamics, and it is infeasible to collect labelled data from every domain. To overcome the challenges of domain adaptation on long sensory sequences, we propose MotionTransformer - a novel framework that extracts domain-invariant features of raw sequences from arbitrary domains, and transforms to new domains without any paired data. Through the experiments, we demonstrate that it is able to efficiently and effectively convert the raw sequence from a new unlabelled target domain into an accurate inertial trajectory, benefiting from the motion knowledge transferred from the labelled source domain. We also conduct real-world experiments to show our framework can reconstruct physically meaningful trajectories from raw IMU measurements obtained with a standard mobile phone in various attachments. Changhao Chen, Yishu Miao, Xiaoxuan Lu 0001, Linhai Xie, Phil Blunsom, Andrew Markham, Agathoniki Trigoni |
AAAI | 1 |
| 2019 | Selective Sensor Fusion for Neural Visual-Inertial OdometryabstractDeep learning approaches for Visual-Inertial Odometry (VIO) have proven successful, but they rarely focus on incorporating robust fusion strategies for dealing with imperfect input sensory data. We propose a novel end-to-end selective sensor fusion framework for monocular VIO, which fuses monocular images and inertial measurements in order to estimate the trajectory whilst improving robustness to real-life issues, such as missing and corrupted data or bad sensor synchronization. In particular, we propose two fusion modalities based on different masking strategies: deterministic soft fusion and stochastic hard fusion, and we compare with previously proposed direct fusion baselines. During testing, the network is able to selectively process the features of the available sensor modalities and produce a trajectory at scale. We present a thorough investigation on the performances on three public autonomous driving, Micro Aerial Vehicle (MAV) and hand-held VIO datasets. The results demonstrate the effectiveness of the fusion strategies, which offer better performances compared to direct fusion, particularly in presence of corrupted data. In addition, we study the interpretability of the fusion networks by visualising the masking layers in different scenarios and with varying data corruption, revealing interesting correlations between the fusion networks and imperfect sensory input data. Changhao Chen, Stefano Rosa, Yishu Miao, Xiaoxuan Lu 0001, Andrew Markham, Agathoniki Trigoni |
CVPR | 1 |
| 2019 | mID: Tracking and Identifying People with Millimeter Wave RadarabstractThe key to offering personalised services in smart spaces is knowing where a particular person is with a high degree of accuracy. Visual tracking is one such solution, but concerns arise around the potential leakage of raw video information and many people are not comfortable accepting cameras in their homes or workplaces. We propose a human tracking and identification system (mID) based on millimeter wave radar which has a high tracking accuracy, without being visually compromising. Unlike competing techniques based on WiFi Channel State Information (CSI), it is capable of tracking and identifying multiple people simultaneously. Using a lowcost, commercial, off-the-shelf radar, we first obtain sparse point clouds and form temporally associated trajectories. With the aid of a deep recurrent network, we identify individual users. We evaluate and demonstrate our system across a variety of scenarios, showing median position errors of 0.16 m and identification accuracy of 89% for 12 people. Peijun Zhao, Xiaoxuan Lu 0001, Changhao Chen, Wei Wang 0226, Agathoniki Trigoni, Andrew Markham |
DCOSS | 4 |
| 2019 | DeepPCO: End-to-End Point Cloud Odometry through Deep Parallel Neural NetworkabstractOdometry is of key importance for localization in the absence of a map. There is considerable work in the area of visual odometry (VO), and recent advances in deep learning have brought novel approaches to VO, which directly learn salient features from raw images. These learning-based approaches have led to more accurate and robust VO systems. However, they have not been well applied to point cloud data yet. In this work, we investigate how to exploit deep learning to estimate point cloud odometry (PCO), which may serve as a critical component in point cloud-based downstream tasks or learning-based systems. Specifically, we propose a novel end-to-end deep parallel neural network called DeepPCO, which can estimate the 6-DOF poses using consecutive point clouds. It consists of two parallel sub-networks to estimate 3D translation and orientation respectively rather than a single neural network. We validate our approach on KITTI Visual Odometry/SLAM benchmark dataset with different baselines. Experiments demonstrate that the proposed approach achieves good performance in terms of pose accuracy. Wei Wang 0226, Muhamad Risqi Utama Saputra, Peijun Zhao, Pedro Porto Buarque de Gusmão, Bo Yang 0027, Changhao Chen, Andrew Markham, Agathoniki Trigoni |
IROS | 6 |
| 2019 | Autonomous Learning for Face Recognition in the Wild via Ambient Wireless CuesabstractFacial recognition is a key enabling component for emerging Internet of Things (IoT) services such as smart homes or responsive offices. Through the use of deep neural networks, facial recognition has achieved excellent performance. However, this is only possibly when trained with hundreds of images of each user in different viewing and lighting conditions. Clearly, this level of effort in enrolment and labelling is impossible for wide-spread deployment and adoption. Inspired by the fact that most people carry smart wireless devices with them, e.g. smartphones, we propose to use this wireless identifier as a supervisory label. This allows us to curate a dataset of facial images that are unique to a certain domain e.g. a set of people in a particular office. This custom corpus can then be used to finetune existing pre-trained models e.g. FaceNet. However, due to the vagaries of wireless propagation in buildings, the supervisory labels are noisy and weak. We propose a novel technique, AutoTune, which learns and refines the association between a face and wireless identifier over time, by increasing the inter-cluster separation and minimizing the intra-cluster distance. Through extensive experiments with multiple users on two sites, we demonstrate the ability of AutoTune to design an environment-specific, continually evolving facial recognition system with entirely no user effort. Xiaoxuan Lu 0001, Xuan Kan, Bowen Du 0002, Changhao Chen, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni, John A. Stankovic |
WWW | 4 |
| 2019 | Autonomous Learning of Speaker Identity and WiFi Geofence From Noisy Sensor DataabstractA fundamental building block toward intelligent environments is the ability to understand who is present in a certain area. A ubiquitous way of detecting this is to exploit unique vocal characteristics as people interact with one another in common spaces. However, manually enrolling users into a biometric database is time-consuming and not robust to vocal deviations over time. Instead, consider audio features sampled during a meeting, yielding a noisy set of possible voiceprints. With a number of meetings and knowledge of participation, e.g., sniffed wireless media access control (MAC) addresses, can we learn to associate a specific identity with a particular voiceprint? To address this problem, this paper advocates an Internet of Things (IoT) solution and proposes to use co-located WiFi as supervisory weak labels to automatically bootstrap the labeling process. In particular, a novel cross-modality labeling algorithm is proposed that jointly optimizes the clustering and association process, which solves the inherent mismatching issues arising from heterogeneous sensor data. At the same time, we further propose to reuse the labeled data to iteratively update wireless geofence models and curate device specific thresholds. The extensive experimental results from two different scenarios demonstrate that our proposed method is able to achieve twofold improvement in labeling compared with conventional methods and can achieve reliable speaker recognition in the wild. Xiaoxuan Lu 0001, Yuanbo Xiangli, Peijun Zhao, Changhao Chen, Agathoniki Trigoni, Andrew Markham |
IEEE Internet Things J. | 4 |
| 2018 | IONet: Learning to Cure the Curse of Drift in Inertial OdometryabstractInertial sensors play a pivotal role in indoor localization, which in turn lays the foundation for pervasive personal applications. However, low-cost inertial sensors, as commonly found in smartphones, are plagued by bias and noise, which leads to unbounded growth in error when accelerations are double integrated to obtain displacement. Small errors in state estimation propagate to make odometry virtually unusable in a matter of seconds. We propose to break the cycle of continuous integration, and instead segment inertial data into independent windows. The challenge becomes estimating the latent states of each window, such as velocity and orientation, as these are not directly observable from sensor data. We demonstrate how to formulate this as an optimization problem, and show how deep recurrent neural networks can yield highly accurate trajectories, outperforming state-of-the-art shallow techniques, on a wide range of tests and attachments. In particular, we demonstrate that IONet can generalize to estimate odometry for non-periodic motion, such as a shopping trolley or baby-stroller, an extremely challenging task for existing techniques. Changhao Chen, Xiaoxuan Lu 0001, Andrew Markham, Agathoniki Trigoni |
AAAI | 1 |
| 2018 | Simultaneous Localization and Mapping with Power Network Electromagnetic FieldabstractVarious sensing modalities have been exploited for indoor location sensing, each of which has well understood limitations, however. This paper presents a first systematic study on using the electromagnetic field (EMF) induced by a building's electric power network for simultaneous localization and mapping (SLAM). A basis of this work is a measurement study showing that the power network EMF sensed by either a customized sensor or smartphone's microphone as a side-channel sensor is spatially distinct and temporally stable. Based on this, we design a SLAM approach that can reliably detect loop closures based on EMF sensing results. With the EMF feature map constructed by SLAM, we also design an efficient online localization scheme for resource-constrained mobiles. Evaluation in three indoor spaces shows that the power network EMF is a promising modality for location sensing on mobile devices, which is able to run in real time and achieve sub-meter accuracy. Xiaoxuan Lu 0001, Yang Li 0147, Peijun Zhao, Changhao Chen, Linhai Xie, Hongkai Wen 0001, Rui Tan 0001, Agathoniki Trigoni |
MobiCom | 4 |