Andrew Markham

dblp:83/7169 · DBLP profile ↗
← Back
126ranked-venue papers
5as first author
63since 2021 · last 2026
0000-0001-5716-3941ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 75 · 49 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 23 since 2021Computer networks · 31 · 4 first-author · 5 since 2021Systems, architecture and hardware · 20 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 Should Robots Comply with Our Instructions or Intentions?
abstract
When people communicate, they often express their intent imperfectly, and human collaborators routinely compensate for these mistakes without issue. For example, if Alice asks for a spatula while serving soup, Bob may infer her intent and bring a ladle instead. This raises a key question for human–robot collaboration: should robots follow instructions literally or should they infer and act on human intent? We study how people expect robots to respond to ambiguous or incorrect instructions in collaborative kitchen scenarios. In this user study, participants either act directly on behalf of a robot or indirectly in observing a robot that may depart from literal instructions to pursue the inferred intent. We find that people generally prefer robots to take some action rather than refuse to comply, although people expect robots to attempt to satisfy the literal instruction (i.e., by thoroughly searching the scene) before taking an imperfect action to satisfy the intent. As large language models (LLMs) are increasingly used to model common sense, we conduct a pilot study to assess whether LLMs make the same decisions as human users about when robots should reinterpret requests.
Tiffany Horter, Andrew Markham, Agathoniki Trigoni, Serena Booth
HRI2
2025 mmDiffusion: mmWave Diffusion for Sequential 3D Human Dense Point Cloud Generation
abstract
Millimeter-wave (mmWave) point-cloud radar shows great promise in enabling responsive human-machine interfaces (e.g., through pose and gesture tracking and for emerging augmented reality approaches). However, generating dense and temporally consistent 3D human point clouds from sequential mmWave signals is challenging due to point-cloud sparsity, jitter, and noise. Existing approaches have made progress in single-frame densification, but are inaccurate over multiple frames. This work redefines the problem as a 3D point cloud denoising task, leveraging reverse diffusion processes to transform sparse mmWave data into detailed and accurate whole-body representations. Our proposed method, mmDiffusion, effectively exploits diffusion models and temporal context within mmWave sequences to learn the denoising process, resulting in denser and temporally coherent human point clouds. For the first time, we also introduce an evaluation metric tailored to measure temporal consistency for sequential 3D human point clouds. Experimental results demonstrate that mmDiffusion significantly outperforms existing methods.
Qian Xie 0001, Xinyu Hou, Qianyi Deng, Amir Patel, Agathoniki Trigoni, Andrew Markham
3DV6
2025 RiTTA: Modeling Event Relations in Text-to-Audio Generation
abstract
Existing text-to-audio (TTA) generation methods have neither systematically explored audio event relation modeling, nor proposed any new framework to enhance this capability.In this work, we systematically study audio event relation modeling in TTA generation models.We first establish a benchmark for this task by: (1) proposing a comprehensive relation corpus covering all potential relations in real-world scenarios; (2) introducing a new audio event corpus encompassing commonly heard audios; and (3) proposing new evaluation metrics to assess audio event relation modeling from various perspectives.Furthermore, we propose a gated prompt tuning strategy that improves existing TTA models' relation modeling capability with negligible extra parameters.Specifically, we introduce learnable relation and event prompt that append to the text prompt before feeding to existing TTA models 1 .
Yash Jain, Andrew Markham, Vibhav Vineet
EMNLP4
2025 SoundTRC: DNN-based Acoustic Target Region Control
abstract
We propose a deep neural network based automatic acoustic target region control framework, where the goal is to maintain all the relevant speech in the designated target region while muting all speech outside the target region in a multi-speaker conferencing room. We discuss three target regions: angle, angle-distance and distance that reflect common region-based speech control requests, and further propose a unified target region encoding strategy to encode the three different target regions into discriminative and compact target region vector. We propose a unified Cross-Attention Transformer based deep neural network, which takes the mixed speech and corresponding target region description as input and outputs ideal ratio mask that is responsible of masking out speech that is outside of the target region while suppressing the noise simultaneously. We run experiments on both simulated shoe-box like 3D room scenes and photo-realistic and complex 3D room scenes, showing the advantage of our proposed framework.
Andrew Markham, Okan Köpüklü
ICASSP2
2025 DiffRefine: Diffusion-Based Proposal Specific Point Cloud Densification for Cross-Domain Object Detection
Sang-Yun Shin, Xinyu Hou, Samuel Hodgson, Andrew Markham, Agathoniki Trigoni
ICCV5
2025 Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization
Yiyuan Yang, Shitong Xu, Agathoniki Trigoni, Andrew Markham
INTERSPEECH4
2025 COOPERA: Continual Open-Ended Human-Robot Assistance
abstract
To understand and collaborate with humans, robots must account for individual human traits, habits, and activities over time. However, most robotic assistants lack these abilities, as they primarily focus on predefined tasks in structured environments and lack a human model to learn from. This work introduces COOPERA, a novel framework for COntinual, OPen-Ended human-Robot Assistance, where simulated humans, driven by psychological traits and long-term intentions, interact with robots in complex environments. By integrating continuous human feedback, our framework, for the first time, enables the study of long-term, open-ended human-robot collaboration (HRC) in different collaborative tasks across various time-scales. Within COOPERA, we introduce a benchmark and an approach to personalize the robot's collaborative actions by learning human traits and context-dependent intents. Experiments validate the extent to which our simulated humans reflect realistic human behaviors and demonstrate the value of inferring and personalizing to human intents for open-ended and long-term HRC.
Kai Lu 0003, Ruta Desai, Xavier Puig, Andrew Markham, Agathoniki Trigoni
NeurIPS5
2025 Target Speaker Extraction through Comparing Noisy Positive and Negative Audio Enrollments
abstract
Target speaker extraction focuses on isolating a specific speaker's voice from an audio mixture containing multiple speakers. To provide information about the target speaker's identity, prior works have utilized clean audio samples as conditioning inputs. However, such clean audio examples are not always readily available. For instance, obtaining a clean recording of a stranger's voice at a cocktail party without leaving the noisy environment is generally infeasible. Limited prior research has explored extracting the target speaker's characteristics from noisy enrollments, which may contain overlapping speech from interfering speakers. In this work, we explore a novel enrollment strategy that encodes target speaker information from the noisy enrollment by comparing segments where the target speaker is talking (Positive Enrollments) with segments where the target speaker is silent (Negative Enrollments). Experiments show the effectiveness of our model architecture, which achieves over 2.1 dB higher SI-SNRi compared to prior works in extracting the monaural speech from the mixture of two speakers. Additionally, the proposed two-stage training strategy accelerates convergence, reducing the number of optimization steps required to reach 3 dB SNR by 60\%. Overall, our method achieves state-of-the-art performance in the monaural target speaker extraction conditioned on noisy enrollments. Our implementation is available at https://github.com/xu-shitong/TSE-through-Positive-Negative-Enroll .
Shitong Xu, Yiyuan Yang, Agathoniki Trigoni, Andrew Markham
NeurIPS4
2025 SoundLoc3D: Invisible 3D Sound Source Localization and Classification Using a Multimodal RGB-D Acoustic Camera
abstract
Accurately localizing 3D sound sources and estimating their semantic labels - where the sources may not be visible, but are assumed to lie on the physical surface of objects in the scene - have many real applications, including detecting gas leak and machinery malfunction. The audio-visual weak-correlation in such setting poses new challenges in deriving innovative methods to answer if or how we can use cross-modal information to solve the task. Towards this end, we propose to use an acoustic-camera rig consisting of a pinhole RGB-D camera and a coplanar four-channel microphone array (Mic-Array). By using this rig to record audio-visual signals from multiviews, we can use the cross-modal cues to estimate the sound sources 3D locations. Specifically, our framework SoundLoc3D treats the task as a set prediction problem, each element in the set corresponds to a potential sound source. Given the audio-visual weak-correlation, the set representation is initially learned from a single view microphone array signal, and then refined by actively incorporating physical surface cues revealed from multiview RGB-D images. We demonstrate the efficiency and superiority of SoundLoc3D on large-scale simulated dataset, and further show its robustness to RGB-D measurement inaccuracy and ambient noise interference.
Sang-Yun Shin, Anoop Cherian, Agathoniki Trigoni, Andrew Markham
WACV5
2025 Learning Selective Sensor Fusion for State Estimation
abstract
Autonomous vehicles and mobile robotic systems are typically equipped with multiple sensors to provide redundancy. By integrating the observations from different sensors, these mobile agents are able to perceive the environment and estimate system states, e.g., locations and orientations. Although deep learning (DL) approaches for multimodal odometry estimation and localization have gained traction, they rarely focus on the issue of robust sensor fusion-a necessary consideration to deal with noisy or incomplete sensor observations in the real world. Moreover, current deep odometry models suffer from a lack of interpretability. To this extent, we propose SelectFusion, an end-to-end selective sensor fusion module that can be applied to useful pairs of sensor modalities, such as monocular images and inertial measurements, depth images, and light detection and ranging (LIDAR) point clouds. Our model is a uniform framework that is not restricted to specific modality or task. During prediction, the network is able to assess the reliability of the latent features from different sensor modalities and to estimate trajectory at both scale and global pose. In particular, we propose two fusion modules-a deterministic soft fusion and a stochastic hard fusion-and offer a comprehensive study of the new strategies compared with trivial direct fusion. We extensively evaluate all fusion strategies both on public datasets and on progressively degraded datasets that present synthetic occlusions, noisy and missing data, and time misalignment between sensors, and we investigate the effectiveness of the different fusion strategies in attending the most reliable features, which in itself provides insights into the operation of the various models.
Changhao Chen, Stefano Rosa, Xiaoxuan Lu 0001, Bing Wang 0013, Agathoniki Trigoni, Andrew Markham
IEEE Trans. Neural Networks Learn. Syst.6
2024 SoundCount: Sound Counting from Raw Audio with Dyadic Decomposition Neural Network
abstract
In this paper, we study an underexplored, yet important and challenging problem: counting the number of distinct sounds in raw audio characterized by a high degree of polyphonicity. We do so by systematically proposing a novel end-to-end trainable neural network~(which we call DyDecNet, consisting of a dyadic decomposition front-end and backbone network), and quantifying the difficulty level of counting depending on sound polyphonicity. The dyadic decomposition front-end progressively decomposes the raw waveform dyadically along the frequency axis to obtain time-frequency representation in multi-stage, coarse-to-fine manner. Each intermediate waveform convolved by a parent filter is further processed by a pair of child filters that evenly split the parent filter's carried frequency response, with the higher-half child filter encoding the detail and lower-half child filter encoding the approximation. We further introduce an energy gain normalization to normalize sound loudness variance and spectrum overlap, and apply it to each intermediate parent waveform before feeding it to the two child filters. To better quantify sound counting difficulty level, we further design three polyphony-aware metrics: polyphony ratio, max polyphony and mean polyphony. We test DyDecNet on various datasets to show its superiority, and we further show dyadic decomposition network can be used as a general front-end to tackle other acoustic tasks.
Zhuangzhuang Dai, Agathoniki Trigoni, Long Chen 0005, Andrew Markham
AAAI5
2024 Learning Continuous 3D Words for Text-to-Image Generation
abstract
Current controls over diffusion models (e.g., through text or ControlNet) for image generation fall short in recognizing abstract, continuous attributes like illumination direction or non-rigid shape change. In this paper, we present an approach for allowing users of text-to-image models to have fine-grained control of several attributes in an image. We do this by engineering special sets of input tokens that can be transformed in a continuous manner – we call them Continuous 3D Words. These attributes can, for example, be represented as sliders and applied jointly with text prompts for fine-grained control over image generation. Given only a single mesh and a rendering engine, we show that our approach can be adopted to provide continuous user control over several 3D-aware attributes, including time-of-day illumination, bird wing orientation, dollyzoom effect, and object poses. Our method is capable of conditioning image creation with multiple Continuous 3D Words and text descriptions simultaneously while adding no overhead to the generative process. Project Page: https://ttchengab.github.io/continuous_3d_words
Ta Ying Cheng, Matheus Gadelha, Thibault Groueix, Matthew Fisher, Radomír Mech, Andrew Markham, Agathoniki Trigoni
CVPR6
2024 Spherical Mask: Coarse-to-Fine 3D Point Cloud Instance Segmentation with Spherical Representation
abstract
Coarse-to-fine 3D instance segmentation methods show weak performances compared to recent Grouping-based, Kernel-based and Transformer-based methods. We argue that this is due to two limitations: 1) Instance size over-estimation by axis-aligned bounding box(AABB) 2) False negative error accumulation from inaccurate box to the re-finement phase. In this work, we introduce Spherical Mask, a novel coarse-to-fine approach based on spherical repre-sentation, overcoming those two limitations with several benefits. Specifically, our coarse detection estimates each in-stance with a 3D polygon using a center and radial distance predictions, which avoids excessive size estimation of AABB. To cut the error propagation in the existing coarse-to-fine approaches, we virtually migrate points based on the polygon, allowing all foreground points, including false negatives, to be refined. During inference, the proposal and point mi-gration modules run in parallel and are assembled to form binary masks of instances. We also introduce two margin-based losses for the point migration to enforce corrections for the false positives/negatives and cohesion of foreground points, significantly improving the performance. Experimen-tal results from three datasets, such as ScanNetV2, S3DIS, and STPLS3D, show that our proposed method outperforms existing works, demonstrating the effectiveness of the new in-stance representation with spherical coordinates. The code is available at: https://github.com/yunshin/SphericalMask
Sang-Yun Shin, Kaichen Zhou, Madhu Vankadari, Andrew Markham, Agathoniki Trigoni
CVPR4
2024 ZeST: Zero-Shot Material Transfer from a Single Image
Ta Ying Cheng, Prafull Sharma, Andrew Markham, Agathoniki Trigoni, Varun Jampani
ECCV (1)3
2024 SSL-Net: A Synergistic Spectral and Learning-Based Network for Efficient Bird Sound Classification
abstract
Efficient and accurate bird sound classification is of important for ecology, habitat protection and scientific research, as it plays a central role in monitoring the distribution and abundance of species. However, prevailing methods typically demand extensively labeled audio datasets and have highly customized frameworks, imposing substantial computational and annotation loads. In this study, we present an efficient and general framework called SSL-Net, which combines spectral and learned features to identify different bird sounds. Encouraging empirical results gleaned from a standard field-collected bird audio dataset validate the efficacy of our method in extracting features efficiently and achieving heightened performance in bird sound classification, even when working with limited sample sizes. Furthermore, we present three feature fusion strategies, aiding engineers and researchers in their selection through quantitative analysis.
Yiyuan Yang, Kaichen Zhou, Agathoniki Trigoni, Andrew Markham
ICASSP4
2024 Deep Neural Room Acoustics Primitive
abstract
The primary objective of room acoustics is to model the intricate sound propagation dynamics from any source to receiver position within enclosed 3D spaces. These dynamics are encapsulated in the form of a 1D room impulse response (RIR). Precisely measuring RIR is difficult due to the complexity of sound propagation encompassing reflection, diffraction, and absorption. In this work, we propose to learn a continuous neural room acoustics field that implicitly encodes all essential sound propagation primitives for each enclosed 3D space, so that we can infer the RIR corresponding to arbitrary source-receiver positions unseen in the training dataset. Our framework, dubbed DeepNeRAP, is trained in a self-supervised manner without requiring direct access to RIR ground truth that is often needed in prior methods. The key idea is to design two cooperative acoustic agents to actively probe a 3D space, one emitting and the other receiving sound at various locations. Analyzing this sound helps to inversely characterize the acoustic primitives. Our framework is well-grounded in the fundamental physical principles of sound propagation, including reciprocity and globality, and thus is acoustically interpretable and meaningful. We present experiments on both synthetic and real-world datasets, demonstrating superior quality in RIR estimation against closely related methods.
Anoop Cherian, Gordon Wichern, Andrew Markham
ICML4
2024 Learning to Catch Reactive Objects with a Behavior Predictor
abstract
Tracking and catching moving objects is an important ability for robots in a dynamic world. Whilst some objects have highly predictable state evolution e.g., the ballistic trajectory of a tennis ball, reactive targets alter their behavior in response to motion of the manipulator. Reactive applications range from gently capturing living animals such as snakes or fish for biological investigations, to smoothly interacting with and assisting a person. Existing works for dynamic catching usually perform target prediction followed by planning, but seldom account for highly non-linear reactive behaviors. Alternatively, Reinforcement Learning (RL) based methods simply treat the target and its motion as part of the observation of the world-state, but perform poorly due to the weak reward signal. In this work, we blend the approach of an explicit, yet learned, target state predictor with RL. We further show how a tightly coupled predictor which ‘observes’ the state of the robot leads to significantly improved anticipatory action, especially with targets that seek to evade the robot following a simple policy. Experiments show that our method achieves an 86.4% (open plane area) and a 73.8% (room) success rate on evasive objects, outperforming monolithic reinforcement learning and other techniques. We also demonstrate the efficacy of our approach across varied targets and trajectories. All code, data, and additional videos are at this GitHub link: https://kl-research.github.io/dyncatch.
Kai Lu 0003, Jia-Xing Zhong, Bo Yang 0027, Bing Wang 0013, Andrew Markham
ICRA5
2024 Dusk Till Dawn: Self-supervised Nighttime Stereo Depth Estimation using Visual Foundation Models
abstract
Self-supervised depth estimation algorithms rely heavily on frame-warping relationships, exhibiting substantial performance degradation when applied in challenging circumstances, such as low-visibility and nighttime scenarios with varying illumination conditions. Addressing this challenge, we introduce an algorithm designed to achieve accurate selfsupervised stereo depth estimation focusing on nighttime conditions. Specifically, we use pretrained visual foundation models to extract generalised features across challenging scenes and present an efficient method for matching and integrating these features from stereo frames. Moreover, to prevent pixels violating photometric consistency assumption from negatively affecting the depth predictions, we propose a novel masking approach designed to filter out such pixels. Lastly, addressing weaknesses in the evaluation of current depth estimation algorithms, we present novel evaluation metrics. Our experiments, conducted on challenging datasets including Oxford RobotCar and MultiSpectral Stereo, demonstrate the robust improvements realized by our approach.
Madhu Vankadari, Samuel Hodgson, Sang-Yun Shin, Kaichen Zhou, Andrew Markham, Agathoniki Trigoni
ICRA5
2024 Pre-training Feature Guided Diffusion Model for Speech Enhancement
Yiyuan Yang, Agathoniki Trigoni, Andrew Markham
INTERSPEECH3
2024 Learning Generalizable Manipulation Policy with Adapter-Based Parameter Fine-Tuning
abstract
This study investigates the use of adapters in reinforcement learning for robotic skill generalization across multiple robots and tasks. Traditional methods are typically reliant on robot-specific retraining and face challenges such as efficiency and adaptability, particularly when scaling to robots with varying kinematics. We propose an alternative approach where a disembodied (virtual) hand manipulator learns a task (i.e., an abstract skill) and then transfers it to various robots with different kinematic constraints without retraining the entire model (i.e., the concrete, physical implementation of the skill). Whilst adapters are commonly used in other domains with strong supervision available, we show how weaker feedback from robotic control can be used to optimize task execution by preserving the abstract skill dynamics whilst adapting to new robotic domains. We demonstrate the effectiveness of our method with experiments conducted in the SAPIEN ManiSkill environment, showing improvements in generalization and task success rates. All code, data, and additional videos are at this GitHub link: https://kl-research.github.io/genrob.
Kai Lu 0003, Kim Tien Ly, William Hebberd, Kaichen Zhou, Ioannis Havoutis, Andrew Markham
IROS6
2024 WSCLoc: Weakly-Supervised Sparse-View Camera Relocalization via Radiance Field
abstract
Despite the advancements in deep learning for camera relocalization tasks, obtaining ground truth pose labels required for the training process remains a costly endeavor. While current weakly supervised methods excel in lightweight label generation, their performance notably declines in scenarios with sparse views. In response to this challenge, we introduce WSCLoc, a system capable of being customized to various deep learning-based relocalization models to enhance their performance under weakly-supervised and sparse view conditions. This is realized with two stages. In the initial stage, WSCLoc employs a multilayer perceptron-based structure called WFT-NeRF to co-optimize image reconstruction quality and initial pose information. To ensure a stable learning process, we incorporate temporal information as input. Furthermore, instead of optimizing SE(3), we opt for sim(3) optimization to explicitly enforce a scale constraint. In the second stage, we co-optimize the pre-trained WFT-NeRF and WFT-Pose. This optimization is enhanced by Time-Encoding based Random View Synthesis and supervised by inter-frame geometric constraints that consider pose, depth, and RGB information. We validate our approaches on two publicly available datasets, one outdoor and one indoor. Our experimental results demonstrate that our weakly-supervised relocalization solutions achieve superior pose estimation accuracy in sparse-view scenarios, comparable to state-of-the-art camera relocalization methods. We will make our code publicly available.
Kaichen Zhou, Andrew Markham, Agathoniki Trigoni
IROS3
2024 SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors
abstract
Current state-of-the-art spatial reasoning-enhanced VLMs are trained to excel at spatial visual question answering (VQA). However, we believe that higher-level 3D-aware tasks, such as articulating dynamic scene changes and motion planning, require a fundamental and explicit 3D understanding beyond current spatial VQA datasets. In this work, we present SpatialPIN, a framework designed to enhance the spatial reasoning capabilities of VLMs through prompting and interacting with priors from multiple 3D foundation models in a zero-shot, training-free manner. Extensive experiments demonstrate that our spatial reasoning-imbued VLM performs well on various forms of spatial VQA and can extend to help in various downstream robotics tasks such as pick and stack and trajectory planning.
Kai Lu 0003, Ta Ying Cheng, Agathoniki Trigoni, Andrew Markham
NeurIPS5
2024 Towards Learning Group-Equivariant Features for Domain Adaptive 3D Detection
abstract
The performance of 3D object detection in large outdoor point clouds deteriorates significantly in an unseen environment due to the inter-domain gap. To address these challenges, most existing methods for domain adaptation harness self-training schemes and attempt to bridge the gap by focusing on a single factor that causes the inter-domain gap, such as objects' sizes, shapes, and foreground density variation. However, the resulting adaptations suggest that there is still a substantial inter-domain gap left to be minimized. We argue that this is due to two limitations: 1) Biased pseudo-label collection from self-training. 2) Multiple factors jointly contributing to how the object is perceived in the unseen target domain. In this work, we propose a grouping-exploration strategy framework, Group Explorer Domain Adaptation ($\textbf{GroupEXP-DA}$), to addresses those two issues. Specifically, our grouping divides the available label sets into multiple clusters and ensures all of them have equal learning attention with the group-equivariant spatial feature, avoiding dominant types of objects causing imbalance problems. Moreover, grouping learns to divide objects by considering inherent factors in a data-driven manner, without considering each factor separately as existing works. On top of the group-equivariant spatial feature that selectively detects objects similar to the input group, we additionally introduce an explorative group update strategy that reduces the false negative detection in the target domain, further reducing the inter-domain gap. During inference, only the learned group features are necessary for making the group-equivariant spatial feature, placing our method as a simple add-on that can be applicable to most existing detectors. We show how each module contributes to substantially bridging the inter-domain gaps compared to existing works across large urban outdoor datasets such as NuScenes, Waymo, and KITTI.
Sang-Yun Shin, Madhu Vankadari, Ta Ying Cheng, Qian Xie 0001, Andrew Markham, Agathoniki Trigoni
NeurIPS6
2024 Sound3DVDet: 3D Sound Source Detection using Multiview Microphone Array and RGB Images
abstract
Spatial localization of 3D sound sources is an important problem in many real world scenarios, especially when the sources may not have any visually distinguishable characteristic; e.g., finding a gas leak, a malfunctioning motor, etc. In this paper, we cast this task in a novel audio-visual setting, by introducing an acoustic-camera rig consisting of a centered pinhole RGB camera and a uniform circular array of four coplanar microphones. Using this setup, we propose Sound3DVDet – a 3D sound source localization Transformer model that treats this task as a set prediction problem. It first learns a set of initial sound source locations (dubbed queries) from a single view of the microphone array signal, then feeds the query set to a sequence of Transformerlike layers for refinement. Each query arising from each layer repeatedly aggregates sound source cues from other views. We deeply supervise the initial sound source queries, intermediate layer queries, and the final output by measuring their respective discrepancy against ground truth queries via bipartite matching. To evaluate our method, we introduce a new dataset: Sound3DVDet Dataset, consisting of nearly 6k scenes produced using the SoundSpaces simulator. We conduct extensive experiments on our dataset and show the efficacy of our approach against closely related methods, demonstrating significant improvements in the localization accuracy. Code is available at https://github.com/yuhanghe01/Sound3DVDet.
Sang-Yun Shin, Anoop Cherian, Agathoniki Trigoni, Andrew Markham
WACV5
2024 Beyond Fusion: Modality Hallucination-based Multispectral Fusion for Pedestrian Detection
abstract
Pedestrian detection is a fundamental task for many downstream applications. Visible and thermal images, as the two most important data types, are usually used to detect pedestrians under various environmental conditions. Many state-of-the-art works have been proposed to use two-stream (i.e., two-branch) architectures to combine visible and thermal information to improve detection performance. However, conventional visible-thermal fusion-based methods have no ability to obtain useful information from the visible branch under poor visibility conditions. The visible branch could even sometimes bring noise into the combined features. In this paper, we present a novel thermal and visible fusion architecture for pedestrian detection. Instead of simply using two branches to separately extract thermal and visible features and then fusing them, we introduce a hallucination branch to learn the mapping from the thermal to the visible domain, forming a novel three-branch feature extraction module. We then adaptively fuse feature maps from all three branches (i.e., thermal, visible, and hallucination). With this new integrated hallucination branch, our network can still get relatively good visible feature maps under challenging low-visibility conditions, thus boosting the overall detection performance. Finally, we experimentally demonstrate the superiority of the proposed architecture over conventional fusion methods.
Qian Xie 0001, Ta Ying Cheng, Jia-Xing Zhong, Kaichen Zhou, Andrew Markham, Agathoniki Trigoni
WACV5
2024 EgoCap and EgoFormer: First-person image captioning with context fusion
Zhuangzhuang Dai, Andrew Markham, Agathoniki Trigoni, M. Arif Imtiazur Rahman, Lahiru N. S. Wijayasingha, John A. Stankovic, Chen Li 0009
Pattern Recognit. Lett.3
2024 Illumination-Aware Hallucination-Based Domain Adaptation for Thermal Pedestrian Detection
abstract
Thermal imagery is emerging as a viable candidate for 24-7, all-weather pedestrian detection owning to thermal sensors’ robust performance for pedestrian detection under different weather and illumination conditions. Despite the promising results obtained from combining visible (RGB) and thermal cameras in multi-spectral fusion techniques, the complex synchronization requirements, including alignment and calibration of sensors, impede their deployment in real-world scenarios. In this paper, we introduce a novel approach for domain adaptation to enhance the performance of pedestrian detection based solely on thermal images. Our proposed approach involves several stages. Firstly, we use both thermal and visible images as input during the training phase. Secondly, we leverage a thermal-to-visible hallucination network to generate feature maps that are similar to those generated by the visible branch. Finally, we design a transformer-based multi-modal fusion module to integrate the hallucinated visible and thermal information more effectively. The thermal-to-visible hallucination network acts as domain adaptation, allowing us to obtain pseudo-visual and thermal features using solely thermal input. Based on the experimental results, it is observed the mean average precision (mAP) increases by 4.72% and the miss rate decreases by 7.56% on the KAIST dataset when compared to the baseline model.
Qian Xie 0001, Ta Ying Cheng, Zhuangzhuang Dai, Vu H. Tran, Agathoniki Trigoni, Andrew Markham
IEEE Trans. Intell. Transp. Syst.6
2024 Deep Learning for Visual Localization and Mapping: A Survey
abstract
Deep-learning-based localization and mapping approaches have recently emerged as a new research direction and receive significant attention from both industry and academia. Instead of creating hand-designed algorithms based on physical models or geometric theories, deep learning solutions provide an alternative to solve the problem in a data-driven way. Benefiting from the ever-increasing volumes of data and computational power on devices, these learning methods are fast evolving into a new area that shows potential to track self-motion and estimate environmental models accurately and robustly for mobile agents. In this work, we provide a comprehensive survey and propose a taxonomy for the localization and mapping methods using deep learning. This survey aims to discuss two basic questions: whether deep learning is promising for localization and mapping, and how deep learning should be applied to solve this problem. To this end, a series of localization and mapping topics are investigated, from the learning-based visual odometry and global relocalization to mapping, and simultaneous localization and mapping (SLAM). It is our hope that this survey organically weaves together the recent works in this vein from robotics, computer vision, and machine learning communities and serves as a guideline for future researchers to apply deep learning to tackle the problem of visual localization and mapping.
Changhao Chen, Bing Wang 0013, Xiaoxuan Lu 0001, Agathoniki Trigoni, Andrew Markham
IEEE Trans. Neural Networks Learn. Syst.5
2024 Multiscale Human Activity Recognition and Anticipation Network
abstract
Deep convolutional neural networks have been leveraged to achieve huge improvements in video understanding and human activity recognition performance in the past decade. However, most existing methods focus on activities that have similar time scales, leaving the task of action recognition on multiscale human behaviors less explored. In this study, a two-stream multiscale human activity recognition and anticipation (MS-HARA) network is proposed, which is jointly optimized using a multitask learning method. The MS-HARA network fuses the two streams of the network using an efficient temporal-channel attention (TCA)-based fusion approach to improve the model's representational ability for both temporal and spatial features. We investigate the multiscale human activities from two basic categories, namely, midterm activities and long-term activities. The network is designed to function as part of a real-time processing framework to support interaction and mutual understanding between humans and intelligent machines. It achieves state-of-the-art results on several datasets for different tasks and different application domains. The midterm and long-term action recognition and anticipation performance, as well as the network fusion, are extensively tested to show the efficiency of the proposed network. The results show that the MS-HARA network can easily be extended to different application domains.
Yang Xing 0002, Stuart Golodetz, Aluna Everitt, Andrew Markham, Agathoniki Trigoni
IEEE Trans. Neural Networks Learn. Syst.4
2023 SoundSynp: Sound Source Detection from Raw Waveforms with Multi-Scale Synperiodic Filterbanks
abstract
We propose synperiodic filter banks, a novel multi-scale learnable filter bank construction strategy that all filters are synchronized by their rotating periodicity. By synchronizing in a certain periodicity, we naturally get filters whose temporal length are reduced if they carry higher frequency response, and vice versa. Such filters internally maintain a better time-frequency resolution trade-off. By further alternating the periodicity, we can easily obtain a group of synperiodic filter bank (we call synperiodic filter banks), where filters of same frequency response in different groups differ in temporal length. Convolving these filter banks with sound raw waveform achieves multi-scale perception in time domain. Moreover, applying the same filter banks to recursively process the 2x-downsampled waveform enables multi-scale perception in the frequency domain. Benefiting from the multi-scale perception in both time and frequency domains, our proposed synperiodic filter banks learn multi-scale time-frequency representation in a data-driven way. Experiments on both sound source direction of arrival (DoA) and physical location detection task show the superiority of synperiodic filter banks.
Andrew Markham
AISTATS2
2023 mmPoint: Dense Human Point Cloud Generation from mmWave
Qian Xie 0001, Qianyi Deng, Ta Ying Cheng, Peijun Zhao, Amir Patel, Agathoniki Trigoni, Andrew Markham
BMVC7
2023 3DMiner: Discovering Shapes from Large-Scale Unannotated Image Datasets
abstract
We present 3DMiner – a pipeline for mining 3D shapes from challenging large-scale unannotated image datasets. Unlike other unsupervised 3D reconstruction methods, we assume that, within a large-enough dataset, there must exist images of objects with similar shapes but varying backgrounds, textures, and viewpoints. Our approach leverages the recent advances in learning self-supervised image representations to cluster images with geometrically similar shapes and find common image correspondences between them. We then exploit these correspondences to obtain rough camera estimates as initialization for bundle-adjustment. Finally, for every image cluster, we apply a progressive bundle-adjusting reconstruction method to learn a neural occupancy field representing the underlying shape. We show that this procedure is robust to several types of errors introduced in previous steps (e.g., wrong camera poses, images containing dissimilar shapes, etc.), allowing us to obtain shape and pose annotations for images in-the-wild. When using images from Pix3D chairs, our method is capable of producing significantly better results than state-of-the-art unsupervised 3D reconstruction techniques, both quantitatively and qualitatively. Furthermore, we show how 3DMiner can be applied to in-the-wild data by reconstructing shapes present in images from the LAION-5B dataset. Project Page: https://ttchengab.github.io/3dminerOfficial.
Ta Ying Cheng, Matheus Gadelha, Sören Pirk, Thibault Groueix, Radomír Mech, Andrew Markham, Agathoniki Trigoni
ICCV6
2023 Decoupling Skill Learning from Robotic Control for Generalizable Object Manipulation
abstract
Recent works in robotic manipulation through reinforcement learning (RL) or imitation learning (IL) have shown potential for tackling a range of tasks e.g., opening a drawer or a cupboard. However, these techniques generalize poorly to unseen objects. We conjecture that this is due to the high-dimensional action space for joint control. In this paper, we take an alternative approach and separate the task of learning ‘what to do’ from ‘how to do it’ i.e., whole-body control. We pose the RL problem as one of determining the skill dynamics for a disembodied virtual manipulator interacting with articulated objects. The whole-body robotic kinematic control is optimized to execute the high-dimensional joint motion to reach the goals in the workspace. It does so by solving a quadratic programming (QP) model with robotic singularity and kinematic constraints. Our experiments on manipulating complex articulated objects show that the proposed approach is more generalizable to unseen objects with large intra-class variations, outperforming previous approaches. The evaluation results indicate that our approach generates more compliant robotic motion and outperforms the pure RL and IL baselines in task success rates. Additional information and videos are available at https://kl-research.github.io/decoupskill.
Kai Lu 0003, Bo Yang 0027, Bing Wang 0013, Andrew Markham
ICRA4
2023 Sample, Crop, Track: Self-Supervised Mobile 3D Object Detection for Urban Driving LiDAR
abstract
Deep learning has led to great progress in the detection of mobile (i.e. movement-capable) objects in urban driving scenes in recent years. Supervised approaches typically require the annotation of large training sets; there has thus been great interest in leveraging weakly, semi- or self- supervised methods to avoid this, with much success. Whilst weakly and semi-supervised methods require some annotation, self-supervised methods have used cues such as motion to relieve the need for annotation altogether. However, a complete absence of annotation typically degrades their performance, and ambiguities that arise during motion grouping can inhibit their ability to find accurate object boundaries. In this paper, we propose a new self-supervised mobile object detection approach called SCT. This uses both motion cues and expected object sizes to improve detection performance, and predicts a dense grid of 3$D$oriented bounding boxes to improve object discovery. We significantly outperform the state-of-the-art self-supervised mobile object detection method TCR on the KITTI tracking benchmark, and achieve performance that is within 30 % of the fully supervised PV-RCNN++ method for IoUs$\leq$0.5. Our source code will be made available online.
Sang-Yun Shin, Stuart Golodetz, Madhu Vankadari, Kaichen Zhou, Andrew Markham, Agathoniki Trigoni
ICRA5
2023 Fast Model Inference and Training On-Board of Satellites
abstract
Artificial intelligence onboard satellites has the potential to reduce data transmission requirements, enable real-time decision-making and collaboration within constellations. This study deploys a lightweight foundational model called RaVAEn on D-Orbit’s ION SCV004 satellite. RaVAEn is a variational auto-encoder (VAE) that generates compressed latent vectors from small image tiles, enabling several downstream tasks. In this work we demonstrate the reliable use of RaVAEn onboard a satellite, achieving an encoding time of 0.110s for tiles of a 4.8x4.8 km2area. In addition, we showcase fast few-shot training onboard a satellite using the latent representation of data. We compare the deployment of the model on the on-board CPU and on the available Myriad vision processing unit (VPU) accelerator. To our knowledge, this work shows for the first time the deployment of a multitask model onboard a CubeSat and the onboard training of a machine learning model.
Vít Ruzicka, Gonzalo Mateo-Garcia, Christopher Bridges 0001, Chris Brunskill, Cormac Purcell, Nicolas Longépé, Andrew Markham
IGARSS7
2023 RADA: Robust Adversarial Data Augmentation for Camera Localization in Challenging Conditions
abstract
Camera localization is a fundamental problem for many applications in computer vision, robotics, and autonomy. Despite recent deep learning-based approaches, the lack of robustness in challenging conditions persists due to changes in appearance caused by texture-less planes, repeating structures, reflective surfaces, motion blur, and illumination changes. Data augmentation is an attractive solution, but standard image perturbation methods fail to improve localization robustness. To address this, we propose RADA, which concentrates on perturbing the most vulnerable pixels to generate relatively less image perturbations that perplex the network. Our method outperforms previous augmentation techniques, achieving up to twice the accuracy of state-of-the-art models even under ‘unseen’ challenging weather conditions. Videos of our results can be found at https://youtu.be/niOv7-fJeCA. The source code for RADA is publicly available at https://github.com/jialuwang123321/RADA.
Muhamad Risqi Utama Saputra, Xiaoxuan Lu 0001, Agathoniki Trigoni, Andrew Markham
IROS5
2023 Multi-body SE(3) Equivariance for Unsupervised Rigid Segmentation and Motion Estimation
abstract
A truly generalizable approach to rigid segmentation and motion estimation is fundamental to 3D understanding of articulated objects and moving scenes. In view of the closely intertwined relationship between segmentation and motion estimates, we present an SE(3) equivariant architecture and a training strategy to tackle this task in an unsupervised manner. Our architecture is composed of two interconnected, lightweight heads. These heads predict segmentation masks using point-level invariant features and estimate motion from SE(3) equivariant features, all without the need for category information. Our training strategy is unified and can be implemented online, which jointly optimizes the predicted segmentation and motion by leveraging the interrelationships among scene flow, segmentation mask, and rigid transformations. We conduct experiments on four datasets to demonstrate the superiority of our method. The results show that our method excels in both model performance and computational efficiency, with only 0.25M parameters and 0.92G FLOPs. To the best of our knowledge, this is the first work designed for category-agnostic part-level SE(3) equivariance in dynamic point clouds.
Jia-Xing Zhong, Ta Ying Cheng, Kai Lu 0003, Kaichen Zhou, Andrew Markham, Agathoniki Trigoni
NeurIPS6
2023 DynPoint: Dynamic Neural Point For View Synthesis
abstract
The introduction of neural radiance fields has greatly improved the effectiveness of view synthesis for monocular videos. However, existing algorithms face difficulties when dealing with uncontrolled or lengthy scenarios, and require extensive training time specific to each new scenario. To tackle these limitations, we propose DynPoint, an algorithm designed to facilitate the rapid synthesis of novel views for unconstrained monocular videos. Rather than encoding the entirety of the scenario information into a latent representation, DynPoint concentrates on predicting the explicit 3D correspondence between neighboring frames to realize information aggregation. Specifically, this correspondence prediction is achieved through the estimation of consistent depth and scene flow information across frames. Subsequently, the acquired correspondence is utilized to aggregate information from multiple reference frames to a target frame, by constructing hierarchical neural point clouds. The resulting framework enables swift and accurate view synthesis for desired views of target frames. The experimental results obtained demonstrate the considerable acceleration of training time achieved - typically an order of magnitude - by our proposed method while yielding comparable outcomes compared to prior approaches. Furthermore, our method exhibits strong robustness in handling long-duration videos without learning a canonical representation of video content.
Kaichen Zhou, Jia-Xing Zhong, Sang-Yun Shin, Kai Lu 0003, Yiyuan Yang, Andrew Markham, Agathoniki Trigoni
NeurIPS6
2023 CubeLearn: End-to-End Learning for Human Motion Recognition From Raw mmWave Radar Signals
abstract
mmWave FMCW radar has attracted a huge amount of research interest for human-centered applications in recent years, such as human gesture and activity recognition. Most existing pipelines are built upon conventional discrete Fourier transform (DFT) preprocessing and deep neural network classifier hybrid methods, with a majority of previous works focusing on designing the downstream classifier to improve overall accuracy. In this work, we take a step back and look at the preprocessing module. To avoid the drawbacks of conventional DFT preprocessing, we propose a complex-weighted learnable preprocessing module, named CubeLearn, to directly extract features from raw radar signal and build an end-to-end deep neural network for mmWave FMCW radar motion recognition applications. Extensive experiments show that our CubeLearn module consistently improves the classification accuracies of different pipelines, especially, benefiting those simpler models, which are more likely to be used on edge devices due to their computational efficiency. We provide ablation studies on initialization methods and structure of the proposed module, as well as an evaluation of the running time on PC and edge devices. This work also serves as a comparison of different approaches toward data cube slicing. Through our task-agnostic design, we propose a first step toward a generic end-to-end solution for radar recognition problems.
Peijun Zhao, Xiaoxuan Lu 0001, Bing Wang 0013, Agathoniki Trigoni, Andrew Markham
IEEE Internet Things J.5
2023 You Only Train Once: Learning General and Distinctive 3D Local Descriptors
abstract
Extracting distinctive, robust, and general 3D local features is essential to downstream tasks such as point cloud registration. However, existing methods either rely on noise-sensitive handcrafted features, or depend on rotation-variant neural architectures. It remains challenging to learn robust and general local feature descriptors for surface matching. In this paper, we propose a new, simple yet effective neural network, termed SpinNet, to extract local surface descriptors which are rotation-invariant whilst sufficiently distinctive and general. A Spatial Point Transformer is first introduced to embed the input local surface into an elaborate cylindrical representation (SO(2) rotation-equivariant), further enabling end-to-end optimization of the entire framework. A Neural Feature Extractor, composed of point-based and 3D cylindrical convolutional layers, is then presented to learn representative and general geometric patterns. An invariant layer is finally used to generate rotation-invariant feature descriptors. Extensive experiments on both indoor and outdoor datasets demonstrate that SpinNet outperforms existing state-of-the-art techniques by a large margin. More critically, it has the best generalization ability across unseen scenarios with different sensor modalities.
Sheng Ao, Yulan Guo, Qingyong Hu, Bo Yang 0027, Andrew Markham, Zengping Chen
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 No Pain, Big Gain: Classify Dynamic Point Cloud Sequences with Static Models by Fitting Feature-level Space-time Surfaces
abstract
Scene flow is a powerful tool for capturing the motion field of 3D point clouds. However, it is difficult to directly apply flow-based models to dynamic point cloud classification since the unstructured points make it hard or even impossible to efficiently and effectively trace point-wise correspondences. To capture 3D motions without explicitly tracking correspondences, we propose a kinematics-inspired neural network (Kinet) by generalizing the kinematic concept of ST-surfaces to the feature space. By unrolling the normal solver of ST-surfaces in the feature space, Kinet implicitly encodes feature-level dynamics and gains advantages from the use of mature back-bones for static point cloud processing. With only minor changes in network structures and low computing overhead, it is painless to jointly train and deploy our framework with a given static model. Experiments on NvGesture, SHREC'17, MSRAction-3D, and NTU-RGBD demonstrate its efficacy in performance, efficiency in both the number of parameters and computational complexity, as well as its versatility to various static backbones. Noticeably, Kinet achieves the accuracy of 93.27% on MSRAction-3D with only 3.20M parameters and 10.35G FLOPS. The code is available at https://github.com/jx-zhong-for-academic-purpose/Kinet.
Jia-Xing Zhong, Kaichen Zhou, Qingyong Hu, Bing Wang 0013, Agathoniki Trigoni, Andrew Markham
CVPR6
2022 Meta-sampler: Almost-Universal yet Task-Oriented Sampling for Point Clouds
Ta Ying Cheng, Qingyong Hu, Qian Xie 0001, Agathoniki Trigoni, Andrew Markham
ECCV (2)5
2022 SQN: Weakly-Supervised Semantic Segmentation of Large-Scale 3D Point Clouds
Qingyong Hu, Bo Yang 0027, Guangchi Fang, Yulan Guo, Ales Leonardis, Agathoniki Trigoni, Andrew Markham
ECCV (27)7
2022 SoundDoA: Learn Sound Source Direction of Arrival and Semantics from Sound Raw Waveforms
Andrew Markham
INTERSPEECH2
2022 Real-Time Hybrid Mapping of Populated Indoor Scenes using a Low-Cost Monocular UAV
abstract
Unmanned aerial vehicles (UAVs) have been used for many applications in recent years, from urban search and rescue, to agricultural surveying, to autonomous underground mine exploration. However, deploying UAVs in tight, indoor spaces, especially close to humans, remains a challenge. One solution, when limited payload is required, is to use micro-UAVs, which pose less risk to humans and typically cost less to replace after a crash. However, micro-UAVs can only carry a limited sensor suite, e.g. a monocular camera instead of a stereo pair or LiDAR, complicating tasks like dense mapping and markerless multi-person 3D human pose estimation, which are needed to operate in tight environments around people. Monocular approaches to such tasks exist, and dense monocular mapping approaches have been successfully deployed for UAV applications. However, despite many recent works on both marker-based and markerless multi-UAV single-person motion capture, markerless single-camera multi-person 3D human pose estimation remains a much earlier-stage technology, and we are not aware of existing attempts to deploy it in an aerial context. In this paper, we present what is thus, to our knowledge, the first system to perform simultaneous mapping and multi-person 3D human pose estimation from a monocular camera mounted on a single UAV. In particular, we show how to loosely couple state-of-the-art monocular depth estimation and monocular 3D human pose estimation approaches to reconstruct a hybrid map of a populated indoor scene in real time. We validate our component-level design choices via extensive experiments on the large-scale ScanNet and GTA-IM datasets. To evaluate our system-level performance, we also construct a new Oxford Hybrid Mapping dataset of populated indoor scenes.
Stuart Golodetz, Madhu Vankadari, Aluna Everitt, Sang-Yun Shin, Andrew Markham, Agathoniki Trigoni
IROS5
2022 DeepCIR: Insights into CIR-based Data-driven UWB Error Mitigation
abstract
Ultra-Wide-Band (UWB) ranging sensors have been widely adopted for robotic navigation thanks to their extremely high bandwidth and hence high resolution. However, off-the-shelf devices may output ranges with significant errors in cluttered, severe non-line-of-sight (NLOS) environments. Recently, neural networks have been actively studied to improve the ranging accuracy of UWB sensors using the channel-impulse-response (CIR) as input. However, previous works have not systematically evaluated the efficacy of various packet types and their possible combinations in a two-way-ranging transaction, including poll, response and final packets. In this paper, we firstly investigate the utility of different packet types and their combinations when used as input for a neural network. Secondly, we propose two novel data-driven approaches, namely FMCIR and WMCIR, that leverage two-sided CIRs for efficient UWB error mitigation. Our approaches outperform state-of-the-art by a significant margin, further reducing range errors up to 45%. Finally, we create and release a dataset of transaction-level synchronized CIRs (each sample consists of the CIR of the poll, response and final packets), which will enable further studies in this area.
Zhuangzhuang Dai, Agathoniki Trigoni, Andrew Markham
IROS4
2022 SensatUrban: Learning Semantics from Urban-Scale Photogrammetric Point Clouds
Qingyong Hu, Bo Yang 0027, Sheikh Khalid, Agathoniki Trigoni, Andrew Markham
Int. J. Comput. Vis.6
2022 SelfVIO: Self-supervised deep monocular Visual-Inertial Odometry and depth estimation
abstract
In the last decade, numerous supervised deep learning approaches have been proposed for visual-inertial odometry (VIO) and depth map estimation, which require large amounts of labelled data. To overcome the data limitation, self-supervised learning has emerged as a promising alternative that exploits constraints such as geometric and photometric consistency in the scene. In this study, we present a novel self-supervised deep learning-based VIO and depth map recovery approach (SelfVIO) using adversarial training and self-adaptive visual-inertial sensor fusion. SelfVIO learns the joint estimation of 6 degrees-of-freedom (6-DoF) ego-motion and a depth map of the scene from unlabelled monocular RGB image sequences and inertial measurement unit (IMU) readings. The proposed approach is able to perform VIO without requiring IMU intrinsic parameters and/or extrinsic calibration between IMU and the camera. We provide comprehensive quantitative and qualitative evaluations of the proposed framework and compare its performance with state-of-the-art VIO, VO, and visual simultaneous localization and mapping (VSLAM) approaches on the KITTI, EuRoC and Cityscapes datasets. Detailed comparisons prove that SelfVIO outperforms state-of-the-art VIO approaches in terms of pose estimation and depth recovery, making it a promising approach among existing methods in the literature.
Yasin Almalioglu, Mehmet Turan, Muhamad Risqi Utama Saputra, Pedro Porto Buarque de Gusmão, Andrew Markham, Agathoniki Trigoni
Neural Networks5
2022 Learning Semantic Segmentation of Large-Scale Point Clouds With Random Sampling
abstract
We study the problem of efficient semantic segmentation of large-scale 3D point clouds. By relying on expensive sampling techniques or computationally heavy pre/post-processing steps, most existing approaches are only able to be trained and operate over small-scale point clouds. In this paper, we introduce RandLA-Net, an efficient and lightweight neural architecture to directly infer per-point semantics for large-scale point clouds. The key to our approach is to use random point sampling instead of more complex point selection approaches. Although remarkably computation and memory efficient, random sampling can discard key features by chance. To overcome this, we introduce a novel local feature aggregation module to progressively increase the receptive field for each 3D point, thereby effectively preserving geometric details. Comparative experiments show that our RandLA-Net can process 1 million points in a single pass up to 200× faster than existing approaches. Moreover, extensive experiments on five large-scale point cloud datasets, including Semantic3D, SemanticKITTI, Toronto3D, NPM3D and S3DIS, demonstrate the state-of-the-art semantic segmentation performance of our RandLA-Net.
Qingyong Hu, Bo Yang 0027, Linhai Xie, Stefano Rosa, Yulan Guo, Zhihua Wang 0005, Agathoniki Trigoni, Andrew Markham
IEEE Trans. Pattern Anal. Mach. Intell.8
2022 iMag+: An Accurate and Rapidly Deployable Inertial Magneto-Inductive SLAM System
abstract
Localisation is an important part of many applications. Our motivating scenarios are short-term construction work and emergency rescue. These scenarios also require rapid setup and robustness to environmental conditions additional to localisation accuracy. These requirements preclude the use of many traditional high-performance methods, e.g., vision-based, laser-based, Ultra-wide band (UWB) and Global Positioning System (GPS)-based localisation systems. To overcome these challenges, we introduceiMag+, an accurate and rapidly deployable inertial magneto-inductive (MI) mapping and localisation system, which only requires monitored workers to carry a single MI transmitter and an inertial measurement unit in order to localise themselves with minimal setup effort. However, one major challenge is to use distorted and ambiguous MI location estimates for localisation. To solve this challenge, we propose a novel method to use MI devices forsensing environmental distortionsfor accurate closing inertial loops. We also suggest a robust and efficient first quadrant estimator to sanitise the ambiguous MI estimates. By applying robust simultaneous localisation and mapping (SLAM), our proposed localisation method achieves excellent tracking accuracy and can improve performance significantly compared with only using a Magneto-inductive device or inertial measurement unit (IMU) for localisation.
Bo Wei 0003, Agathoniki Trigoni, Andrew Markham
IEEE Trans. Mob. Comput.3
2022 Graph-Based Thermal-Inertial SLAM With Probabilistic Neural Networks
abstract
Simultaneous localization and mapping (SLAM) system typically employs vision-based sensors to observe the surrounding environment. However, the performance of such systems highly depends on the ambient illumination conditions. In scenarios with adverse visibility or in the presence of airborne particulates (e.g., smoke, dust, etc.), alternative modalities such as those based on thermal imaging and inertial sensors are more promising. In this article, we propose the first complete thermal–inertial SLAM system that combines neural abstraction in the SLAM front end with robust pose-graph optimization in the SLAM back end. We model the sensor abstraction in the front end by employing probabilistic deep learning parameterized by mixture density networks (MDNs). Our key strategies to successfully model this encoding from thermal imagery are the usage of normalized 14-b radiometric data, the incorporation of hallucinated visual (RGB) features, and the inclusion of feature selection to estimate the MDN parameters. To enable a full SLAM system, we also design an efficient global image descriptor that is able to detect loop closures from thermal embedding vectors. We performed extensive experiments and analysis using three datasets, namely self-collected ground robot and hand-held data taken in indoor environment, and one public dataset (SubT-tunnel) collected in underground tunnel. Finally, we demonstrate that an accurate thermal–inertial SLAM system can be realized in conditions of both benign and adverse visibility.
Muhamad Risqi Utama Saputra, Xiaoxuan Lu 0001, Pedro Porto Buarque de Gusmão, Bing Wang 0013, Andrew Markham, Agathoniki Trigoni
IEEE Trans. Robotics5
2021 VMLoc: Variational Fusion For Learning-Based Multimodal Camera Localization
abstract
Recent learning-based approaches have achieved impressive results in the field of single-shot camera localization. However, how best to fuse multiple modalities (e.g., image and depth) and to deal with degraded or missing input are less well studied. In particular, we note that previous approaches towards deep fusion do not perform significantly better than models employing a single modality. We conjecture that this is because of the naive approaches to feature space fusion through summation or concatenation which do not take into account the different strengths of each modality. To address this, we propose an end-to-end framework, termed VMLoc, to fuse different sensor inputs into a common latent space through a variational Product-of-Experts (PoE) followed by attention-based fusion. Unlike previous multimodal variational works directly adapting the objective function of vanilla variational auto-encoder, we show how camera localization can be accurately estimated through an unbiased objective function based on importance weighting. Our model is extensively evaluated on RGB-D datasets and the results prove the efficacy of our model. The source code is available at https://github.com/Zalex97/VMLoc.
Kaichen Zhou, Changhao Chen, Bing Wang 0013, Muhamad Risqi Utama Saputra, Agathoniki Trigoni, Andrew Markham
AAAI6
2021 SpinNet: Learning a General Surface Descriptor for 3D Point Cloud Registration
abstract
Extracting robust and general 3D local features is key to downstream tasks such as point cloud registration and reconstruction. Existing learning-based local descriptors are either sensitive to rotation transformations, or rely on classical handcrafted features which are neither general nor representative. In this paper, we introduce a new, yet conceptually simple, neural architecture, termed SpinNet, to extract local features which are rotationally invariant whilst sufficiently informative to enable accurate registration. A Spatial Point Transformer is first introduced to map the input local surface into a carefully designed cylindrical space, enabling end-to-end optimization with SO(2) equivariant representation. A Neural Feature Extractor which leverages the powerful point-based and 3D cylindrical convolutional neural layers is then utilized to derive a compact and representative descriptor for matching. Extensive experiments on both indoor and outdoor datasets demonstrate that SpinNet outperforms existing state-of-the-art techniques by a large margin. More critically, it has the best generalization ability across unseen scenarios with different sensor modalities. The code is available at https://github.com/QingyongHu/SpinNet.
Sheng Ao, Qingyong Hu, Bo Yang 0027, Andrew Markham, Yulan Guo
CVPR4
2021 Towards Semantic Segmentation of Urban-Scale 3D Point Clouds: A Dataset, Benchmarks and Challenges
abstract
An essential prerequisite for unleashing the potential of supervised deep learning algorithms in the area of 3D scene understanding is the availability of large-scale and richly annotated datasets. However, publicly available datasets are either in relative small spatial scales or have limited semantic annotations due to the expensive cost of data acquisition and data annotation, which severely limits the development of fine-grained semantic understanding in the context of 3D point clouds. In this paper, we present an urban-scale photogrammetric point cloud dataset with nearly three billion richly annotated points, which is three times the number of labeled points than the existing largest photogrammetric point cloud dataset. Our dataset consists of large areas from three UK cities, covering about 7.6 km2of the city landscape. In the dataset, each 3D point is labeled as one of 13 semantic classes. We extensively evaluate the performance of state-of-the-art algorithms on our dataset and provide a comprehensive analysis of the results. In particular, we identify several key challenges towards urban-scale point cloud understanding. The dataset is available at https://github.com/QingyongHu/SensatUrban.
Qingyong Hu, Bo Yang 0027, Sheikh Khalid, Agathoniki Trigoni, Andrew Markham
CVPR6
2021 P2-Net: Joint Description and Detection of Local Features for Pixel and Point Matching
abstract
Accurately describing and detecting 2D and 3D key-points is crucial to establishing correspondences across images and point clouds. Despite a plethora of learning-based 2D or 3D local feature descriptors and detectors having been proposed, the derivation of a shared descriptor and joint keypoint detector that directly matches pixels and points remains under-explored by the community. This work takes the initiative to establish fine-grained correspondences between 2D images and 3D point clouds. In order to directly match pixels and points, a dual fully-convolutional framework is presented that maps 2D and 3D inputs into a shared latent representation space to simultaneously describe and detect keypoints. Furthermore, an ultra-wide reception mechanism and a novel loss function are designed to mitigate the intrinsic information variations between pixel and point local regions. Extensive experimental results demonstrate that our framework shows competitive performance in fine-grained matching between images and point clouds and achieves state-of-the-art results for the task of indoor visual localization. Our source code is available at https://github.com/BingCS/P2-Net.
Bing Wang 0013, Changhao Chen, Zhaopeng Cui, Jie Qin 0004, Xiaoxuan Lu 0001, Zhengdi Yu, Peijun Zhao, Zhen Dong 0005, Fan Zhu 0001, Agathoniki Trigoni, Andrew Markham
ICCV11
2021 SoundDet: Polyphonic Moving Sound Event Detection and Localization from Raw Waveform
abstract
We present a new framework SoundDet, which is an end-to-end trainable and light-weight framework, for polyphonic moving sound event detection and localization. Prior methods typically approach this problem by preprocessing raw waveform into time-frequency representations, which is more amenable to process with well-established image processing pipelines. Prior methods also detect in segment-wise manner, leading to incomplete and partial detections. SoundDet takes a novel approach and directly consumes the raw, multichannel waveform and treats the spatio-temporal sound event as a complete “sound-object" to be detected. Specifically, SoundDet consists of a backbone neural network and two parallel heads for temporal detection and spatial localization, respectively. Given the large sampling rate of raw waveform, the backbone network first learns a set of phase-sensitive and frequency-selective bank of filters to explicitly retain direction-of-arrival information, whilst being highly computationally and parametrically efficient than standard 1D/2D convolution. A dense sound event proposal map is then constructed to handle the challenges of predicting events with large varying temporal duration. Accompanying the dense proposal map are a temporal overlapness map and a motion smoothness map that measure a proposal’s confidence to be an event from temporal detection accuracy and movement consistency perspective. Involving the two maps guarantees SoundDet to be trained in a spatio-temporally unified manner. Experimental results on the public DCASE dataset show the advantage of SoundDet on both segment-based evaluation and our newly proposed event-based evaluation system.
Agathoniki Trigoni, Andrew Markham
ICML3
2021 RadarLoc: Learning to Relocalize in FMCW Radar
abstract
Relocalization is a fundamental task in the field of robotics and computer vision. There is considerable work in the field of deep camera relocalization, which directly estimates poses from raw images. However, learning-based methods have not yet been applied to the radar sensory data. In this work, we investigate how to exploit deep learning to predict global poses from Emerging Frequency-Modulated Continuous Wave (FMCW) radar scans. Specifically, we propose a novel end-to-end neural network with self-attention, termed RadarLoc, which is able to estimate 6-DoF global poses directly. We also propose to improve the localization performance by utilizing geometric constraints between radar scans. We validate our approach on the recently released challenging outdoor dataset Oxford Radar RobotCar. Comprehensive experiments demonstrate that the proposed method outperforms radar-based localization and deep camera relocalization methods by a significant margin.
Wei Wang 0226, Pedro Porto Buarque de Gusmão, Bo Yang 0027, Andrew Markham, Agathoniki Trigoni
ICRA4
2021 3D Motion Capture of an Unmodified Drone with Single-chip Millimeter Wave Radar
abstract
Accurate motion capture of aerial robots in 3D is a key enabler for autonomous operation in indoor environments such as warehouses or factories, as well as driving forward research in these areas. The most commonly used solutions at present are optical motion capture (e.g. VICON) and Ultrawide-band (UWB), but these are costly and cumbersome to deploy, due to their requirement of multiple cameras/anchors spaced around the tracking area. They also require the drone to be modified to carry an active or passive marker. In this work, we present an inexpensive system that can be rapidly installed, based on single-chip millimeter wave (mmWave) radar. Importantly, the drone does not need to be modified or equipped with any markers, as we exploit the Doppler signals from the rotating propellers. Furthermore, 3D tracking is possible from a single point, greatly simplifying deployment. We develop a novel deep neural network and demonstrate decimeter level 3D tracking at 10Hz, achieving better performance than classical baselines. Our hope is that this low-cost system will act to catalyse inexpensive drone research and increased autonomy.
Peijun Zhao, Xiaoxuan Lu 0001, Bing Wang 0013, Agathoniki Trigoni, Andrew Markham
ICRA5
2021 Cut, Distil and Encode (CDE): Split Cloud-Edge Deep Inference
abstract
In cloud-edge environments, running all Deep Neural Network (DNN) models on the cloud causes significant network congestion and high latency, whereas the exclusive use of the edge device for execution limits the size and structure of the DNN, impacting accuracy. This paper introduces a novel partitioning approach for DNN inference between the edge and the cloud. This is the first work to consider simultaneous optimization of both the memory usage at the edge and the size of the data to be transferred over the wireless link. The experiments were performed on two different network architectures, MobileNetV1 and VGG16. The proposed approach makes it possible to execute part of the network on very constrained devices (e.g., microcontrollers), and under poor network conditions (e.g., LoRa) whilst retaining reasonable accuracies. Moreover, the results show that the choice of the optimal layer to split the network depends on the bandwidth and memory constraints, whereas prior work suggests that the best choice is always to split the network at higher layers. We demonstrate superior performance compared to existing techniques.
Marion Sbai, Muhamad Risqi Utama Saputra, Agathoniki Trigoni, Andrew Markham
SECON4
2021 Human tracking and identification through a millimeter wave radar
Peijun Zhao, Xiaoxuan Lu 0001, Changhao Chen, Wei Wang 0226, Agathoniki Trigoni, Andrew Markham
Ad Hoc Networks7
2021 Deep Neural Network Based Inertial Odometry Using Low-Cost Inertial Measurement Units
abstract
Inertial measurement units (IMUs) have emerged as an essential component in many of today's indoor navigation solutions due to their low cost and ease of use. However, despite many attempts for reducing the error growth of navigation systems based on commercial-grade inertial sensors, there is still no satisfactory solution that produces navigation estimates with long-time stability in widely differing conditions. This paper proposes to break the cycle of continuous integration used in traditional inertial algorithms, formulate it as an optimization problem, and explore the use of deep recurrent neural networks for estimating the displacement of a user over a specified time window. By training the deep neural network using inertial measurements and ground truth displacement data, it is possible to learn both motion characteristics and systematic error drift. As opposed to established context-aided inertial solutions, the proposed method is not dependent on either fixed sensor positions or periodic motion patterns. It can reconstruct accurate trajectories directly from raw inertial measurements, and predict the corresponding uncertainty to show model confidence. Extensive experimental evaluations demonstrate that the neural network produces position estimates with high accuracy for several different attachments, users, sensors, and motion types. As a particular demonstration of its flexibility, our deep inertial solutions can estimate trajectories for non-periodic motion, such as the shopping trolley tracking. Further more, it works in highly dynamic conditions, such as running, remaining extremely challenging for current techniques.
Changhao Chen, Xiaoxuan Lu 0001, Johan Wahlström, Andrew Markham, Agathoniki Trigoni
IEEE Trans. Mob. Comput.4
2021 DynaNet: Neural Kalman Dynamical Model for Motion Estimation and Prediction
abstract
Dynamical models estimate and predict the temporal evolution of physical systems. State-space models (SSMs) in particular represent the system dynamics with many desirable properties, such as being able to model uncertainty in both the model and measurements, and optimal (in the Bayesian sense) recursive formulations, e.g., the Kalman filter. However, they require significant domain knowledge to derive the parametric form and considerable hand tuning to correctly set all the parameters. Data-driven techniques, e.g., recurrent neural networks, have emerged as compelling alternatives to SSMs with wide success across a number of challenging tasks, in part due to their impressive capability to extract relevant features from rich inputs. They, however, lack interpretability and robustness to unseen conditions. Thus, data-driven models are hard to be applied in safety-critical applications, such as self-driving vehicles. In this work, we present DynaNet, a hybrid deep learning and time-varying SSM, which can be trained end-to-end. Our neural Kalman dynamical model allows us to exploit the relative merits of both SSM and deep neural networks. We demonstrate its effectiveness in the estimation and prediction on a number of physically challenging tasks, including visual odometry, sensor fusion for visual-inertial navigation, and motion prediction. In addition, we show how DynaNet can indicate failures through investigation of properties, such as the rate of innovation (Kalman gain).
Changhao Chen, Xiaoxuan Lu 0001, Bing Wang 0013, Agathoniki Trigoni, Andrew Markham
IEEE Trans. Neural Networks Learn. Syst.5
2021 Learning With Stochastic Guidance for Robot Navigation
abstract
Due to the sparse rewards and high degree of environmental variation, reinforcement learning approaches, such as deep deterministic policy gradient (DDPG), are plagued by issues of high variance when applied in complex real-world environments. We present a new framework for overcoming these issues by incorporating a stochastic switch, allowing an agent to choose between high- and low-variance policies. The stochastic switch can be jointly trained with the original DDPG in the same framework. In this article, we demonstrate the power of the framework in a navigation task, where the robot can dynamically choose to learn through exploration or to use the output of a heuristic controller as guidance. Instead of starting from completely random actions, the navigation capability of a robot can be quickly bootstrapped by several simple independent controllers. The experimental results show that with the aid of stochastic guidance, we are able to effectively and efficiently train DDPG navigation policies and achieve significantly better performance than state-of-the-art baseline models.
Linhai Xie, Yishu Miao, Sen Wang 0002, Phil Blunsom, Zhihua Wang 0005, Changhao Chen, Andrew Markham, Agathoniki Trigoni
IEEE Trans. Neural Networks Learn. Syst.7
2020 AtLoc: Attention Guided Camera Localization
abstract
Deep learning has achieved impressive results in camera localization, but current single-image techniques typically suffer from a lack of robustness, leading to large outliers. To some extent, this has been tackled by sequential (multi-images) or geometry constraint approaches, which can learn to reject dynamic objects and illumination conditions to achieve better performance. In this work, we show that attention can be used to force the network to focus on more geometrically robust objects and features, achieving state-of-the-art performance in common benchmark, even if using only a single image as input. Extensive experimental evidence is provided through public indoor and outdoor datasets. Through visualization of the saliency maps, we demonstrate how the network learns to reject dynamic objects, yielding superior global camera pose regression performance. The source code is avaliable at https://github.com/BingCS/AtLoc.
Bing Wang 0013, Changhao Chen, Xiaoxuan Lu 0001, Peijun Zhao, Agathoniki Trigoni, Andrew Markham
AAAI6
2020 RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point Clouds
abstract
We study the problem of efficient semantic segmentation for large-scale 3D point clouds. By relying on expensive sampling techniques or computationally heavy pre/post-processing steps, most existing approaches are only able to be trained and operate over small-scale point clouds. In this paper, we introduce RandLA-Net, an efficient and lightweight neural architecture to directly infer per-point semantics for large-scale point clouds. The key to our approach is to use random point sampling instead of more complex point selection approaches. Although remarkably computation and memory efficient, random sampling can discard key features by chance. To overcome this, we introduce a novel local feature aggregation module to progressively increase the receptive field for each 3D point, thereby effectively preserving geometric details. Extensive experiments show that our RandLA-Net can process 1 million points in a single pass with up to 200x faster than existing approaches. Moreover, our RandLA-Net clearly surpasses state-of-the-art approaches for semantic segmentation on two large-scale benchmarks Semantic3D and SemanticKITTI.
Qingyong Hu, Bo Yang 0027, Linhai Xie, Stefano Rosa, Yulan Guo, Zhihua Wang 0005, Agathoniki Trigoni, Andrew Markham
CVPR8
2020 SnapNav: Learning Mapless Visual Navigation with Sparse Directional Guidance and Visual Reference
abstract
Learning-based visual navigation still remains a challenging problem in robotics, with two overarching issues: how to transfer the learnt policy to unseen scenarios, and how to deploy the system on real robots. In this paper, we propose a deep neural network based visual navigation system, SnapNav. Unlike map-based navigation or Visual-Teach-and-Repeat (VT&R), SnapNav only receives a few snapshots of the environment combined with directional guidance to allow it to execute the navigation task. Additionally, SnapNav can be easily deployed on real robots due to a two-level hierarchy: a high level commander that provides directional commands and a low level controller that provides real-time control and obstacle avoidance. This also allows us to effectively use simulated and real data to train the different layers of the hierarchy, facilitating robust control. Extensive experimental results show that SnapNav achieves a highly autonomous navigation ability compared to baseline models, enabling sparse, map-less navigation in previously unseen environments.
Linhai Xie, Andrew Markham, Agathoniki Trigoni
ICRA2
2020 Heart Rate Sensing with a Robot Mounted mmWave Radar
abstract
Heart rate monitoring at home is a useful metric for assessing health e.g. of the elderly or patients in post-operative recovery. Although non-contact heart rate monitoring has been widely explored, typically using a static, wall-mounted device, measurements are limited to a single room and sensitive to user orientation and position. In this work, we propose mBeats, a robot mounted millimeter wave (mmWave) radar system that provide periodic heart rate measurements under different user poses, without interfering in a users daily activities. mBeats contains a mmWave servoing module that adaptively adjusts the sensor angle to the best reflection pro le. Furthermore, mBeats features a deep neural network predictor, which can estimate heart rate from the lower leg and additionally provides estimation uncertainty. Through extensive experiments, we demonstrate accurate and robust operation of mBeats in a range of scenarios. We believe by integrating mobility and adaptability, mBeats can empower many down-stream healthcare applications at home, such as palliative care, post-operative rehabilitation and telemedicine.
Peijun Zhao, Xiaoxuan Lu 0001, Bing Wang 0013, Changhao Chen, Linhai Xie, Agathoniki Trigoni, Andrew Markham
ICRA8
2020 See through smoke: robust indoor mapping with low-cost mmWave radar
abstract
This paper presents the design, implementation and evaluation of milliMap, a single-chip millimetre wave (mmWave) radar based indoor mapping system targetted towards low-visibility environments to assist in emergency response. A unique feature of milliMap is that it only leverages a low-cost, off-the-shelf mmWave radar, but can reconstruct a dense grid map with accuracy comparable to lidar, as well as providing semantic annotations of objects on the map. milliMap makes two key technical contributions. First, it autonomously overcomes the sparsity and multi-path noise of mmWave signals by combining cross-modal supervision from a co-located lidar during training and the strong geometric priors of indoor spaces. Second, it takes the spectral response of mmWave reflections as features to robustly identify different types of objects e.g. doors, walls etc. Extensive experiments in different indoor environments show that milliMap can achieve a map reconstruction error less than 0.2m and classify key semantics with an accuracy of ~ 90%, whilst operating through dense smoke.
Xiaoxuan Lu 0001, Stefano Rosa, Peijun Zhao, Bing Wang 0013, Changhao Chen, John A. Stankovic, Agathoniki Trigoni, Andrew Markham
MobiSys8
2020 Indoor positioning system in visually-degraded environments with millimetre-wave radar and inertial sensors: demo abstract
abstract
Positional estimation is of great importance in the public safety sector. Emergency responders such as fire fighters, medical rescue teams, and the police will all benefit from a resilient positioning system to deliver safe and effective emergency services. Unfortunately, satellite navigation (e.g., GPS) offers limited coverage in indoor environments. It is also not possible to rely on infrastructure based solutions. To this end, wearable sensor-aided navigation techniques, such as those based on camera and Inertial Measurement Units (IMU), have recently emerged recently as an accurate, infrastructure-free solution. Together with an increase in the computational capabilities of mobile devices, motion estimation can be performed in real-time. In this demonstration, we present a real-time indoor positioning system which fuses millimetre-wave (mmWave) radar and IMU data via deep sensor fusion. We employ mmWave radar rather than an RGB camera as it provides better robustness to visual degradation (e.g., smoke, darkness, etc.) while at the same time requiring lower computational resources to enable runtime computation. We implemented the sensor system on a handheld device and a mobile computer running at 10 FPS to track a user inside an apartment. Good accuracy and resilience were exhibited even in poorly illuminated scenes.
Zhuangzhuang Dai, Muhamad Risqi Utama Saputra, Xiaoxuan Lu 0001, Agathoniki Trigoni, Andrew Markham
SenSys5
2020 milliEgo: single-chip mmWave radar aided egomotion estimation via deep sensor fusion
abstract
Robust and accurate trajectory estimation of mobile agents such as people and robots is a key requirement for providing spatial awareness for emerging capabilities such as augmented reality or autonomous interaction. Although currently dominated by optical techniques e.g., visual-inertial odometry these suffer from challenges with scene illumination or featureless surfaces. As an alternative, we propose milliEgo, a novel deep-learning approach to robust egomotion estimation which exploits the capabilities of low-cost mm Wave radar. Although mmWave radar has a fundamental advantage over monocular cameras of being metric i.e., providing absolute scale or depth, current single chip solutions have limited and sparse imaging resolution, making existing point-cloud registration techniques brittle. We propose a new architecture that is optimized for solving this challenging pose transformation problem. Secondly, to robustly fuse mmWave pose estimates with additional sensors, e.g. inertial or visual sensors we introduce a mixed attention approach to deep fusion. Through extensive experiments, we demonstrate our proposed system is able to achieve 1.3% 3D error drift and generalizes well to unseen environments. We also show that the neural architecture can be made highly efficient and suitable for real-time embedded applications.
Xiaoxuan Lu 0001, Muhamad Risqi Utama Saputra, Peijun Zhao, Yasin Almalioglu, Pedro Porto Buarque de Gusmão, Changhao Chen, Ke Sun 0012, Agathoniki Trigoni, Andrew Markham
SenSys9
2020 Deep Emergent Communication for the IoT
abstract
Learning emergent communication remains a longstanding challenge in distributed Internet of Things (IoT) settings. The need to overcome tedious, complex design of hand-engineered communication protocols coupled with superior prediction and classification capabilities, make Deep Networks attractive for distributed, cooperative IoT settings. In such settings, sensing devices must sense, communicate and provide actuation whilst executing a resource-aware operation. Reliance on the Cloud for knowledge discovery is fraught with latency, connectivity, and bandwidth issues. We continue to see the emergence of edge-centric paradigms in which sensing devices at the network edge are endowed with intelligence. In turn, these devices are equipped with self-organization capabilities, robust real-time capabilities, reduced bandwidth requirements and greater context awareness. In this paper, we propose a novel, scalable communicating Convolutional Recurrent Neural Network (C-RNN) architecture for distributed IoT settings. Our framework automatically learns emergent communication in a purely data-driven way. Extensive experimental evaluation shows that our framework can learn to solve distributed image classification tasks, optimises for communication cost, is robust to lossy-links and can scale to multiple nodes.
Prince Abudu, Andrew Markham
SMARTCOMP2
2020 Learning distributed communication and computation in the IoT
Prince Abudu, Andrew Markham
Comput. Commun.2
2020 Robust Attentional Aggregation of Deep Feature Sets for Multi-view 3D Reconstruction
abstract
We study the problem of recovering an underlying 3D shape from a set of images. Existing learning based approaches usually resort to recurrent neural nets, e.g., GRU, or intuitive pooling operations, e.g., max/mean poolings, to fuse multiple deep features encoded from input images. However, GRU based approaches are unable to consistently estimate 3D shapes given different permutations of the same set of input images as the recurrent unit is permutation variant. It is also unlikely to refine the 3D shape given more images due to the long-term memory loss of GRU. Commonly used pooling approaches are limited to capturing partial information, e.g., max/mean values, ignoring other valuable features. In this paper, we present a new feed-forward neural module, named AttSets , together with a dedicated training algorithm, named FASet , to attentively aggregate an arbitrarily sized deep feature set for multi-view 3D reconstruction. The AttSets module is permutation invariant, computationally efficient and flexible to implement, while the FASet algorithm enables the AttSets based network to be remarkably robust and generalize to an arbitrary number of input images. We thoroughly evaluate FASet and the properties of AttSets on multiple large public datasets. Extensive experiments show that AttSets together with FASet algorithm significantly outperforms existing aggregation approaches.
Bo Yang 0027, Sen Wang 0002, Andrew Markham, Agathoniki Trigoni
Int. J. Comput. Vis.3
2020 Deep-Learning-Based Pedestrian Inertial Navigation: Methods, Data Set, and On-Device Inference
abstract
Modern inertial measurements units (IMUs) are small, cheap, energy efficient, and widely employed in smart devices and mobile robots. Exploiting inertial data for accurate and reliable pedestrian navigation supports is a key component for emerging Internet of Things applications and services. Recently, there has been a growing interest in applying deep neural networks (DNNs) to motion sensing and location estimation. However, the lack of sufficient labelled data for training and evaluating architecture benchmarks has limited the adoption of DNNs in IMU-based tasks. In this article, we present and release the Oxford Inertial Odometry Data Set (OxIOD), a first-of-its-kind public data set for deep-learning-based inertial navigation research with fine-grained ground truth on all sequences. Furthermore, to enable more efficient inference at the edge, we propose a novel lightweight framework to learn and reconstruct pedestrian trajectories from raw IMU data. Extensive experiments show the effectiveness of our data set and methods in achieving accurate data-driven pedestrian inertial navigation on resource-constrained devices.
Changhao Chen, Peijun Zhao, Xiaoxuan Lu 0001, Wei Wang 0226, Andrew Markham, Agathoniki Trigoni
IEEE Internet Things J.5
2019 MotionTransformer: Transferring Neural Inertial Tracking between Domains
abstract
Inertial information processing plays a pivotal role in egomotion awareness for mobile agents, as inertial measurements are entirely egocentric and not environment dependent. However, they are affected greatly by changes in sensor placement/orientation or motion dynamics, and it is infeasible to collect labelled data from every domain. To overcome the challenges of domain adaptation on long sensory sequences, we propose MotionTransformer - a novel framework that extracts domain-invariant features of raw sequences from arbitrary domains, and transforms to new domains without any paired data. Through the experiments, we demonstrate that it is able to efficiently and effectively convert the raw sequence from a new unlabelled target domain into an accurate inertial trajectory, benefiting from the motion knowledge transferred from the labelled source domain. We also conduct real-world experiments to show our framework can reconstruct physically meaningful trajectories from raw IMU measurements obtained with a standard mobile phone in various attachments.
Changhao Chen, Yishu Miao, Xiaoxuan Lu 0001, Linhai Xie, Phil Blunsom, Andrew Markham, Agathoniki Trigoni
AAAI6
2019 Selective Sensor Fusion for Neural Visual-Inertial Odometry
abstract
Deep learning approaches for Visual-Inertial Odometry (VIO) have proven successful, but they rarely focus on incorporating robust fusion strategies for dealing with imperfect input sensory data. We propose a novel end-to-end selective sensor fusion framework for monocular VIO, which fuses monocular images and inertial measurements in order to estimate the trajectory whilst improving robustness to real-life issues, such as missing and corrupted data or bad sensor synchronization. In particular, we propose two fusion modalities based on different masking strategies: deterministic soft fusion and stochastic hard fusion, and we compare with previously proposed direct fusion baselines. During testing, the network is able to selectively process the features of the available sensor modalities and produce a trajectory at scale. We present a thorough investigation on the performances on three public autonomous driving, Micro Aerial Vehicle (MAV) and hand-held VIO datasets. The results demonstrate the effectiveness of the fusion strategies, which offer better performances compared to direct fusion, particularly in presence of corrupted data. In addition, we study the interpretability of the fusion networks by visualising the masking layers in different scenarios and with varying data corruption, revealing interesting correlations between the fusion networks and imperfect sensory input data.
Changhao Chen, Stefano Rosa, Yishu Miao, Xiaoxuan Lu 0001, Andrew Markham, Agathoniki Trigoni
CVPR6
2019 Map-aided Navigation for Emergency Searches
abstract
Real-time positioning of emergency personnel has been an active research topic for many years. However, studies on how to improve navigation accuracy by using prior information on the idiosyncratic motion characteristics of firefighters are scarce. This paper presents an algorithm for generating pseudo observations of position and orientation based on standard search patterns used by fire-fighters. The iterative closest point algorithm is used to compare walking trajectories estimated from inertial odometry with search patterns generated from digital maps. The resulting fitting errors are then used to integrate the pseudo observations into a map-aided navigation filter. Specifically, we present a sequential Monte Carlo solution where the pattern comparison is used to both update particle weights and create new particle samples. Experimental results involving professional firefighters demonstrate that the proposed pseudo observations can achieve a stable localization error of about one meter, and offer increased robustness in the presence of map errors.
Johan Wahlström, Pedro Porto Buarque de Gusmão, Andrew Markham, Agathoniki Trigoni
DCOSS3
2019 mID: Tracking and Identifying People with Millimeter Wave Radar
abstract
The key to offering personalised services in smart spaces is knowing where a particular person is with a high degree of accuracy. Visual tracking is one such solution, but concerns arise around the potential leakage of raw video information and many people are not comfortable accepting cameras in their homes or workplaces. We propose a human tracking and identification system (mID) based on millimeter wave radar which has a high tracking accuracy, without being visually compromising. Unlike competing techniques based on WiFi Channel State Information (CSI), it is capable of tracking and identifying multiple people simultaneously. Using a lowcost, commercial, off-the-shelf radar, we first obtain sparse point clouds and form temporally associated trajectories. With the aid of a deep recurrent network, we identify individual users. We evaluate and demonstrate our system across a variety of scenarios, showing median position errors of 0.16 m and identification accuracy of 89% for 12 people.
Peijun Zhao, Xiaoxuan Lu 0001, Changhao Chen, Wei Wang 0226, Agathoniki Trigoni, Andrew Markham
DCOSS7
2019 Distilling Knowledge From a Deep Pose Regressor Network
abstract
This paper presents a novel method to distill knowledge from a deep pose regressor network for efficient Visual Odometry (VO). Standard distillation relies on ''dark knowledge'' for successful knowledge transfer. As this knowledge is not available in pose regression and the teacher prediction is not always accurate, we propose to emphasize the knowledge transfer only when we trust the teacher. We achieve this by using teacher loss as a confidence score which places variable relative importance on the teacher prediction. We inject this confidence score to the main training task via Attentive Imitation Loss (AIL) and when learning the intermediate representation of the teacher through Attentive Hint Training (AHT) approach. To the best of our knowledge, this is the first work which successfully distill the knowledge from a deep pose regression network. Our evaluation on the KITTI and Malaga dataset shows that we can keep the student prediction close to the teacher with up to 92.95% parameter reduction and 2.12x faster in computation time.
Muhamad Risqi Utama Saputra, Pedro Porto Buarque de Gusmão, Yasin Almalioglu, Andrew Markham, Agathoniki Trigoni
ICCV4
2019 GANVO: Unsupervised Deep Monocular Visual Odometry and Depth Estimation with Generative Adversarial Networks
abstract
In the last decade, supervised deep learning approaches have been extensively employed in visual odometry (VO) applications, which is not feasible in environments where labelled data is not abundant. On the other hand, unsupervised deep learning approaches for localization and mapping in unknown environments from unlabelled data have received comparatively less attention in VO research. In this study, we propose a generative unsupervised learning framework that predicts 6-DoF pose camera motion and monocular depth map of the scene from unlabelled RGB image sequences, using deep convolutional Generative Adversarial Networks (GANs). We create a supervisory signal by warping view sequences and assigning the re-projection minimization to the objective loss function that is adopted in multi-view pose estimation and single-view depth generation network. Detailed quantitative and qualitative evaluations of the proposed framework on the KITTI [1] and Cityscapes [2] datasets show that the proposed method outperforms both existing traditional and unsupervised deep VO methods providing better results for both pose estimation and depth recovery.
Yasin Almalioglu, Muhamad Risqi Utama Saputra, Pedro Porto Buarque de Gusmão, Andrew Markham, Agathoniki Trigoni
ICRA4
2019 Learning Monocular Visual Odometry through Geometry-Aware Curriculum Learning
abstract
Inspired by the cognitive process of humans and animals, Curriculum Learning (CL) trains a model by gradually increasing the difficulty of the training data. In this paper, we study whether CL can be applied to complex geometry problems like estimating monocular Visual Odometry (VO). Unlike existing CL approaches, we present a novel CL strategy for learning the geometry of monocular VO by gradually making the learning objective more difficult during training. To this end, we propose a novel geometry-aware objective function by jointly optimizing relative and composite transformations over small windows via bounded pose regression loss. A cascade optical flow network followed by recurrent network with a differentiable windowed composition layer, termed CL-VO, is devised to learn the proposed objective. Evaluation on three real-world datasets shows superior performance of CL-VO over state-of-the-art feature-based and learning-based VO.
Muhamad Risqi Utama Saputra, Pedro Porto Buarque de Gusmão, Sen Wang 0002, Andrew Markham, Agathoniki Trigoni
ICRA4
2019 DeepPCO: End-to-End Point Cloud Odometry through Deep Parallel Neural Network
abstract
Odometry is of key importance for localization in the absence of a map. There is considerable work in the area of visual odometry (VO), and recent advances in deep learning have brought novel approaches to VO, which directly learn salient features from raw images. These learning-based approaches have led to more accurate and robust VO systems. However, they have not been well applied to point cloud data yet. In this work, we investigate how to exploit deep learning to estimate point cloud odometry (PCO), which may serve as a critical component in point cloud-based downstream tasks or learning-based systems. Specifically, we propose a novel end-to-end deep parallel neural network called DeepPCO, which can estimate the 6-DOF poses using consecutive point clouds. It consists of two parallel sub-networks to estimate 3D translation and orientation respectively rather than a single neural network. We validate our approach on KITTI Visual Odometry/SLAM benchmark dataset with different baselines. Experiments demonstrate that the proposed approach achieves good performance in terms of pose accuracy.
Wei Wang 0226, Muhamad Risqi Utama Saputra, Peijun Zhao, Pedro Porto Buarque de Gusmão, Bo Yang 0027, Changhao Chen, Andrew Markham, Agathoniki Trigoni
IROS7
2019 Learning Object Bounding Boxes for 3D Instance Segmentation on Point Clouds
abstract
We propose a novel, conceptually simple and general framework for instance segmentation on 3D point clouds. Our method, called 3D-BoNet, follows the simple design philosophy of per-point multilayer perceptrons (MLPs). The framework directly regresses 3D bounding boxes for all instances in a point cloud, while simultaneously predicting a point-level mask for each instance. It consists of a backbone network followed by two parallel network branches for 1) bounding box regression and 2) point mask prediction. 3D-BoNet is single-stage, anchor-free and end-to-end trainable. Moreover, it is remarkably computationally efficient as, unlike existing approaches, it does not require any post-processing steps such as non-maximum suppression, feature sampling, clustering or voting. Extensive experiments show that our approach surpasses existing work on both ScanNet and S3DIS datasets while being approximately 10x more computationally efficient. Comprehensive ablation studies demonstrate the effectiveness of our design.
Bo Yang 0027, Ronald Clark, Qingyong Hu, Sen Wang 0002, Andrew Markham, Agathoniki Trigoni
NeurIPS6
2019 Distributed Communicating Neural Network Architecture for Smart Environments
abstract
The deployment of millions of embedded sensors plagued by resource constraints in sophisticated, complex and dynamic IoT smart environments continues to inspire the need to build novel architectures and models for automated, efficient inference and communication in distributed smart settings. In such settings, practical challenges related to energy efficiency, computational power and reliability, tedious design implementation, effective communication, optimal sampling and accurate event classification, prediction and detection exist. Sensors operating in smart environments must be capable of overcoming such challenges and enable scalable monitoring of dynamic phenomena while conducting real-time operations. The development of Machine Learning (ML) continues to motivate a new wave of innovative solutions that intermarry embedded sensors, IoT, and ML to enable various applications in smart environments. We propose a distributed communicating architecture based on Recurrent Neural Networks (RNNs) that can be instantiated on smart devices observing unique data and performing automated distributed inference via hidden-state communication. Our model uses a data-driven approach to collectively solve various distributed objectives, as evidenced by a series of systematic analyses we present. Although demonstrated on a small setup (2/3) nodes, this work sets out a new direction for automatically learning to communicate to solve tasks in distributed settings.
Prince Abudu, Andrew Markham
SMARTCOMP2
2019 Autonomous Learning for Face Recognition in the Wild via Ambient Wireless Cues
abstract
Facial recognition is a key enabling component for emerging Internet of Things (IoT) services such as smart homes or responsive offices. Through the use of deep neural networks, facial recognition has achieved excellent performance. However, this is only possibly when trained with hundreds of images of each user in different viewing and lighting conditions. Clearly, this level of effort in enrolment and labelling is impossible for wide-spread deployment and adoption. Inspired by the fact that most people carry smart wireless devices with them, e.g. smartphones, we propose to use this wireless identifier as a supervisory label. This allows us to curate a dataset of facial images that are unique to a certain domain e.g. a set of people in a particular office. This custom corpus can then be used to finetune existing pre-trained models e.g. FaceNet. However, due to the vagaries of wireless propagation in buildings, the supervisory labels are noisy and weak. We propose a novel technique, AutoTune, which learns and refines the association between a face and wireless identifier over time, by increasing the inter-cluster separation and minimizing the intra-cluster distance. Through extensive experiments with multiple users on two sites, we demonstrate the ability of AutoTune to design an environment-specific, continually evolving facial recognition system with entirely no user effort.
Xiaoxuan Lu 0001, Xuan Kan, Bowen Du 0002, Changhao Chen, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni, John A. Stankovic
WWW6
2019 Autonomous Learning of Speaker Identity and WiFi Geofence From Noisy Sensor Data
abstract
A fundamental building block toward intelligent environments is the ability to understand who is present in a certain area. A ubiquitous way of detecting this is to exploit unique vocal characteristics as people interact with one another in common spaces. However, manually enrolling users into a biometric database is time-consuming and not robust to vocal deviations over time. Instead, consider audio features sampled during a meeting, yielding a noisy set of possible voiceprints. With a number of meetings and knowledge of participation, e.g., sniffed wireless media access control (MAC) addresses, can we learn to associate a specific identity with a particular voiceprint? To address this problem, this paper advocates an Internet of Things (IoT) solution and proposes to use co-located WiFi as supervisory weak labels to automatically bootstrap the labeling process. In particular, a novel cross-modality labeling algorithm is proposed that jointly optimizes the clustering and association process, which solves the inherent mismatching issues arising from heterogeneous sensor data. At the same time, we further propose to reuse the labeled data to iteratively update wireless geofence models and curate device specific thresholds. The extensive experimental results from two different scenarios demonstrate that our proposed method is able to achieve twofold improvement in labeling compared with conventional methods and can achieve reliable speaker recognition in the wild.
Xiaoxuan Lu 0001, Yuanbo Xiangli, Peijun Zhao, Changhao Chen, Agathoniki Trigoni, Andrew Markham
IEEE Internet Things J.6
2019 Dense 3D Object Reconstruction from a Single Depth View
abstract
In this paper, we propose a novel approach, 3D-RecGAN++, which reconstructs the complete 3D structure of a given object from a single arbitrary depth view using generative adversarial networks. Unlike existing work which typically requires multiple views of the same object or class labels to recover the full 3D geometry, the proposed 3D-RecGAN++ only takes the voxel grid representation of a depth view of the object as input, and is able to generate the complete 3D occupancy grid with a high resolution of 2563by recovering the occluded/missing regions. The key idea is to combine the generative capabilities of 3D encoder-decoder and the conditional adversarial networks framework, to infer accurate and fine-grained 3D structures of objects in high-dimensional voxel space. Extensive experiments on large synthetic datasets and real-world Kinect datasets show that the proposed 3D-RecGAN++ significantly outperforms the state of the art in single view 3D object reconstruction, and is able to reconstruct unseen types of objects.
Bo Yang 0027, Stefano Rosa, Andrew Markham, Agathoniki Trigoni, Hongkai Wen 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2018 IONet: Learning to Cure the Curse of Drift in Inertial Odometry
abstract
Inertial sensors play a pivotal role in indoor localization, which in turn lays the foundation for pervasive personal applications. However, low-cost inertial sensors, as commonly found in smartphones, are plagued by bias and noise, which leads to unbounded growth in error when accelerations are double integrated to obtain displacement. Small errors in state estimation propagate to make odometry virtually unusable in a matter of seconds. We propose to break the cycle of continuous integration, and instead segment inertial data into independent windows. The challenge becomes estimating the latent states of each window, such as velocity and orientation, as these are not directly observable from sensor data. We demonstrate how to formulate this as an optimization problem, and show how deep recurrent neural networks can yield highly accurate trajectories, outperforming state-of-the-art shallow techniques, on a wide range of tests and attachments. In particular, we demonstrate that IONet can generalize to estimate odometry for non-periodic motion, such as a shopping trolley or baby-stroller, an extremely challenging task for existing techniques.
Changhao Chen, Xiaoxuan Lu 0001, Andrew Markham, Agathoniki Trigoni
AAAI3
2018 Deepauth: in-situ authentication for smartwatches via deeply learned behavioural biometrics
abstract
This paper proposes DeepAuth, an in-situ authentication framework that leverages the unique motion patterns when users entering passwords as behavioural biometrics. It uses a deep recurrent neural network to capture the subtle motion signatures during password input, and employs a novel loss function to learn deep feature representations that are robust to noise, unseen passwords, and malicious imposters even with limited training data. DeepAuth is by design optimised for resource constrained platforms, and uses a novel split-RNN architecture to slim inference down to run in real-time on off-the-shelf smartwatches. Extensive experiments with real-world data show that DeepAuth outperforms the state-of-the-art significantly in both authentication performance and cost, offering real-time authentication on a variety of smartwatches.
Xiaoxuan Lu 0001, Bowen Du 0002, Peijun Zhao, Hongkai Wen 0001, Yiran Shen 0001, Andrew Markham, Agathoniki Trigoni
UbiComp6
2018 DEFO-NET: Learning Body Deformation Using Generative Adversarial Networks
abstract
Modelling the physical properties of everyday objects is a fundamental prerequisite for autonomous robots. We present a novel generative adversarial network (DEFO-NET), able to predict body deformations under external forces from a single RGB-D image. The network is based on an invertible conditional Generative Adversarial Network (IcGAN) and is trained on a collection of different objects of interest generated by a physical finite element model simulator. Defo-netinherits the generalisation properties of GANs. This means that the network is able to reconstruct the whole 3-D appearance of the object given a single depth view of the object and to generalise to unseen object configurations. Contrary to traditional finite element methods, our approach is fast enough to be used in real-time applications. We apply the network to the problem of safe and fast navigation of mobile robots carrying payloads over different obstacles and floor materials. Experimental results in real scenarios show how a robot equipped with an RGB-D camera can use the network to predict terrain deformations under different payload configurations and use this to avoid unsafe areas.
Zhihua Wang 0005, Stefano Rosa, Linhai Xie, Bo Yang 0027, Sen Wang 0002, Agathoniki Trigoni, Andrew Markham
ICRA7
2018 iMag: Accurate and Rapidly Deployable Inertial Magneto-Inductive Localisation
abstract
Localisation is of importance for many applications. Our motivating scenarios are short-term construction work and emergency rescue. Not only is accuracy necessary, these scenarios also require rapid setup and robustness to environmental conditions. These requirements preclude the use of many traditional methods e.g. vision-based, laser-based, Ultra-wide band (UWB) and Global Positioning System (GPS)-based localisation systems. To solve these challenges, we introduce iMag, an accurate and rapidly deployable inertial magneto-inductive (MI) localisation system. It localises monitored workers using a single MI transmitter and inertial measurement units with minimal setup effort. However, MI location estimates can be distorted and ambiguous. To solve this problem, we suggest a novel method to use MI devices for sensing environmental distortions, and use these to correctly close inertial loops. By applying robust simultaneous localisation and mapping (SLAM), our proposed localisation method achieves excellent tracking accuracy, and can improve performance significantly compared with only using an inertial measurement unit (IMU) and MI device for localisation.
Bo Wei 0003, Agathoniki Trigoni, Andrew Markham
ICRA3
2018 Learning with Training Wheels: Speeding up Training with a Simple Controller for Deep Reinforcement Learning
abstract
Deep Reinforcement Learning (DRL) has been applied successfully to many robotic applications. However, the large number of trials needed for training is a key issue. Most of existing techniques developed to improve training efficiency (e.g. imitation) target on general tasks rather than being tailored for robot applications, which have their specific context to benefit from. We propose a novel framework, Assisted Reinforcement Learning, where a classical controller (e.g. a PID controller) is used as an alternative, switchable policy to speed up training of DRL for local planning and navigation problems. The core idea is that the simple control law allows the robot to rapidly learn sensible primitives, like driving in a straight line, instead of random exploration. As the actor network becomes more advanced, it can then take over to perform more complex actions, like obstacle avoidance. Eventually, the simple controller can be discarded entirely. We show that not only does this technique train faster, it also is less sensitive to the structure of the DRL network and consistently outperforms a standard Deep Deterministic Policy Gradient network. We demonstrate the results in both simulation and real-world experiments.
Linhai Xie, Sen Wang 0002, Stefano Rosa, Andrew Markham, Agathoniki Trigoni
ICRA4
2018 3D-PhysNet: Learning the Intuitive Physics of Non-Rigid Object Deformations
abstract
The ability to interact and understand the environment is a fundamental prerequisite for a wide range of applications from robotics to augmented reality. In particular, predicting how deformable objects will react to applied forces in real time is a significant challenge. This is further confounded by the fact that shape information about encountered objects in the real world is often impaired by occlusions, noise and missing regions e.g. a robot manipulating an object will only be able to observe a partial view of the entire solid. In this work we present a framework, 3D-PhysNet, which is able to predict how a three-dimensional solid will deform under an applied force using intuitive physics modelling. In particular, we propose a new method to encode the physical properties of the material and the applied force, enabling generalisation over materials. The key is to combine deep variational autoencoders with adversarial training, conditioned on the applied force and the material properties.We further propose a cascaded architecture that takes a single 2.5D depth view of the object and predicts its deformation. Training data is provided by a physics simulator. The network is fast enough to be used in real-time applications from partial views. Experimental results show the viability and the generalisation properties of the proposed architecture.
Zhihua Wang 0005, Stefano Rosa, Bo Yang 0027, Sen Wang 0002, Agathoniki Trigoni, Andrew Markham
IJCAI6
2018 Identifying Sources and Sinks in the Presence of Multiple Agents with Gaussian Process Vector Calculus
abstract
In systems of multiple agents, identifying the cause of observed agent dynamics is challenging. Often, these agents operate in diverse, non-stationary environments, where models rely on hand-crafted environment-specific features to infer influential regions in the system's surroundings. To overcome the limitations of these inflexible models, we present GP-LAPLACE, a technique for locating sources and sinks from trajectories in time-varying fields. Using Gaussian processes, we jointly infer a spatio-temporal vector field, as well as canonical vector calculus operations on that field. Notably, we do this from only agent trajectories without requiring knowledge of the environment, and also obtain a metric for denoting the significance of inferred causal features in the environment by exploiting our probabilistic method. To evaluate our approach, we apply it to both synthetic and real-world GPS data, demonstrating the applicability of our technique in the presence of multiple agents, as well as its superiority over existing methods.
Adam D. Cobb, Richard Everett 0001, Andrew Markham, Stephen J. Roberts
KDD3
2018 Automatic Face Recognition Adaptation via Ambient Wireless Identifiers
abstract
Face recognition is a key enabling service for smart-spaces, allowing building management agents to easily monitor 'who is where', anticipating user needs and tailoring their local environment and experiences. Although facial recognition, especially through the use of deep neural networks, has achieved stellar performance over large datasets, the majority of approaches require supervised learning, that is, to be trained with tens or hundreds of images of users in different poses and lighting conditions. In this paper, we motivate that this enrollment effort is unnecessary if the smart-space has access to a wireless identifier e.g., through a smart-phone's MAC address. By learning and refining the noisy and weak association between a user's smart-phone and facial images, AutoTune can fine-tune a deep neural network to tailor it to the environment, users and conditions of a particular camera or set of cameras.
Xiaoxuan Lu 0001, Peijun Zhao, Bowen Du 0002, Hongkai Wen 0001, Andrew Markham, Stefano Rosa, Agathoniki Trigoni
SenSys5
2017 VINet: Visual-Inertial Odometry as a Sequence-to-Sequence Learning Problem
abstract
In this paper we present an on-manifold sequence-to-sequence learning approach to motion estimation using visual and inertial sensors. It is to the best of our knowledge the first end-to-end trainable method for visual-inertial odometry which performs fusion of the data at an intermediate feature-representation level. Our method has numerous advantages over traditional approaches. Specifically, it eliminates the need for tedious manual synchronization of the camera and IMU as well as eliminating the need for manual calibration between the IMU and camera. A further advantage is that our model naturally and elegantly incorporates domain specific information which significantly mitigates drift. We show that our approach is competitive with state-of-the-art traditional methods when accurate calibration data is available and can be trained to outperform them in the presence of calibration and synchronization errors.
Ronald Clark, Sen Wang 0002, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni
AAAI4
2017 Evolutionary Machine Learning for RTS Game StarCraft
abstract
Real-Time Strategy (RTS) games involve multiple agents acting simultaneously, and result in enormous state dimensionality. In this paper, we propose an abstracted and simplified model for the famous game StarCraft, and design a dynamic programming algorithm to solve the building order problem, which takes minimal time to achieve a specific target. In addition, Genetic Algorithms (GA) are used to find an optimal target for the opening stage.
Lianlong Wu, Andrew Markham
AAAI2
2017 VidLoc: A Deep Spatio-Temporal Model for 6-DoF Video-Clip Relocalization
abstract
Machine learning techniques, namely convolutional neural networks (CNN) and regression forests, have recently shown great promise in performing 6-DoF localization of monocular images. However, in most cases image-sequences, rather only single images, are readily available. To this extent, none of the proposed learning-based approaches exploit the valuable constraint of temporal smoothness, often leading to situations where the per-frame error is larger than the camera motion. In this paper we propose a recurrent model for performing 6-DoF localization of video-clips. We find that, even by considering only short sequences (20 frames), the pose estimates are smoothed and the localization error can be drastically reduced. Finally, we consider means of obtaining probabilistic pose estimates from our model. We evaluate our method on openly-available real-world autonomous driving and indoor localization datasets.
Ronald Clark, Sen Wang 0002, Andrew Markham, Agathoniki Trigoni, Hongkai Wen 0001
CVPR3
2017 SCAN: learning speaker identity from noisy sensor data
abstract
Sensor data acquired from multiple sensors simultaneously is featuring increasingly in our evermore pervasive world. Buildings can be made smarter and more efficient, spaces more responsive to users. A fundamental building block towards smart spaces is the ability to understand who is present in a certain area. A ubiquitous way of detecting this is to exploit the unique vocal features as people interact with one another. As an example, consider audio features sampled during a meeting, yielding a noisy set of possible voiceprints. With a number of meetings and knowledge of participation (e.g. through a calendar or MAC address), can we learn to associate a specific identity with a particular voiceprint? Obviously enrolling users into a biometric database is time-consuming and not robust to vocal deviations over time. To address this problem, the standard approach is to perform a clustering step (e.g. of audio data) followed by a data association step, when identity-rich sensor data is available. In this paper we show that this approach is not robust to noise in either type of sensor stream; to tackle this issue we propose a novel algorithm that jointly optimises the clustering and association process yielding up to three times higher identification precision than approaches that execute these steps sequentially. We demonstrate the performance benefits of our approach in two case studies, one with acoustic and MAC datasets that we collected from meetings in a non-residential building, and another from an online dataset from recorded radio interviews.
Xiaoxuan Lu 0001, Hongkai Wen 0001, Sen Wang 0002, Andrew Markham, Agathoniki Trigoni
IPSN4
2017 GraphTinker: Outlier rejection and inlier injection for pose graph SLAM
abstract
In pose graph Simultaneous Localization and Mapping (SLAM) systems, incorrect loop closures can seriously hinder optimizers from converging to correct solutions, significantly degrading both localization accuracy and map consistency. Therefore, it is crucial to enhance their robustness in the presence of numerous false-positive loop closures. Existing approaches tend to fail when working with very unreliable front-end systems, where the majority of inferred loop closures are incorrect. In this paper, we propose a novel middle layer, seamlessly embedded between front and back ends, to boost the robustness of the whole SLAM system. The main contributions of this paper are two-fold: 1) the proposed middle layer offers a new mechanism to reliably detect and remove false-positive loop closures, even if they form the overwhelming majority; 2) artificial loop closures are automatically reconstructed and injected into pose graphs in the framework of an Extended Rauch-Tung-Striebel smoother, reinforcing reliable loop closures. The proposed algorithm alters the graph generated by the front-end and can then be optimized by any back-end system. Extensive experiments are conducted to demonstrate significantly improved accuracy and robustness compared with state-of-the-art methods and various back-ends, verifying the effectiveness of the proposed algorithm.
Linhai Xie, Sen Wang 0002, Andrew Markham, Agathoniki Trigoni
IROS3
2017 Towards Self-supervised Face Labeling via Cross-modality Association
abstract
Face recognition has become the de facto authentication solution in a broad spectrum of applications, from smart buildings, to industrial monitoring and security services. However, in many of those real-world scenarios, tracking or identifying people with facial recognition is extremely challenging due to the variations in the environment such as lighting conditions, camera viewing angles and subject motion. For most of the state-of-the-art face recognition systems, they need to be trained on a large dataset containing a good variety of labelled face images to work well. However, collecting and manually labelling such datasets is difficult and time consuming, probably more so than developing the algorithms. In this paper, we propose a novel framework to automatically label user identities with their face images in smart spaces, exploiting the fact that the users tend to carry their smart devices while seen by the surveillance cameras. We evaluate our method on 10 users in a smart building setting, and the experimental results show that our method can achieve > 0.9 f1 score on average.
Xiaoxuan Lu 0001, Xuan Kan, Stefano Rosa, Bowen Du 0002, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni
SenSys6
2017 Interpreting Lion Behaviour as Probabilistic Programs
Neil Dhir, Matthijs Vákár, Matthew Wijers, Andrew Markham, Frank D. Wood
UAI4
2017 Tracking People in Highly Dynamic Industrial Environments
abstract
To date, the majority of positioning systems have been designed to operate within environments that have a long-term stable macro-structure with potential small-scale dynamics. These assumptions allow the existing positioning systems to produce and utilize stable maps. However, in highly dynamic industrial settings these assumptions are no longer valid and the task of tracking people is more challenging due to the rapid large-scale changes in structure. In this paper, we propose a novel positioning system for tracking people in highly dynamic industrial environments, such as construction sites. The proposed system leverages the existing CCTV camera infrastructure found in many industrial settings along with radio and inertial sensors within each worker’s mobile phone to accurately track multiple people. This multi-target multi-sensor tracking framework also allows our system to use cross-modality training in order to deal with the environment dynamics. In particular, we show how our system uses cross-modality training in order to automatically keep track environmental changes (i.e., new walls) by utilizing occlusion maps. In addition, we show how these maps can be used in conjunction with social forces to accurately predict human motion and increase the tracking accuracy. We have conducted extensive real-world experiments in a construction site showing significant accuracy improvement via cross-modality training and the use of social forces.
Savvas Papaioannou, Andrew Markham, Agathoniki Trigoni
IEEE Trans. Mob. Comput.2
2016 Underground Incrementally Deployed Magneto-Inductive 3-D Positioning Network
abstract
Underground mines are characterized by a network of intersecting tunnels and sharp turns, an environment which is particularly challenging for radiofrequency based positioning systems due to extreme multipath, non-line-of-sight propagation, and poor anchor geometry. Such systems typically require a dense grid of devices to enable 3-D positioning. Moreover, the precise position of each anchor node needs to be precisely surveyed, a particularly challenging task in underground environments. Magneto-inductive (MI) positioning, which provides 3-D position and orientation from a single transmitter and penetrates thick layers of soil and rock without loss, is a more promising approach, but so far has only been investigated in simple point-to-point contexts. In this paper, we develop a novel MI positioning approach to cover an extended underground 3-D space with unknown geometry using a rapidly deployable anchor network. The key to our approach is that the position of only a single anchor needs to be accurately surveyed-the positions of all secondary anchors are determined using an iterative refinement process using measurements obtained from receivers within the network. This avoids the particularly challenging and time-intensive task in an underground environment of accurately surveying the positions of all of the transmitters. We also demonstrate how measurements obtained from multiple transmitters can be fused to improve localization accuracy. We validate the proposed approach in a man-made cave and show that, with a portable system that took 5 min to deploy, we were able to provide accurate through-the-earth location capability to nodes placed along a suite of tunnels.
Traian E. Abrudan, Zhuoling Xiao, Andrew Markham, Agathoniki Trigoni
IEEE Trans. Geosci. Remote. Sens.3
2016 Magnetic Induction-Based Positioning in Distorted Environments
abstract
Ferrous and highly conductive materials distort low-frequency magnetic fields and can significantly increase magnetoinductive positioning errors. In this paper, we use the image theory in order to formulate an analytical channel model for the magnetic field of a quasi-static magnetic dipole positioned above a perfectly conducting half-space. The proposed model can be used to compensate for the distorting effects that metallic reinforcement bars (rebars) within the floor impose on the magnetic field of a magnetoinductive transmitter node in an indoor single-story environment. Good agreement is observed between the analytical solution and numerical solutions obtained from 3-D finite-element simulations. Experimental results indicate that the image theory model shows improvement over the free-space dipole model in estimating positions in the distorted environment, typically reducing positioning errors by 22% in 90% of the cases and 26% in 40% of the cases. No prior information on the geometry of the metallic distorters was available, making this essentially a “blind” technique.
Orfeas Kypris, Traian E. Abrudan, Andrew Markham
IEEE Trans. Geosci. Remote. Sens.3
2015 Robust vision-based indoor localization
abstract
Vision-based positioning has proven to be highly successful and popular in mobile robotics and computer vision applications. These methods have, however, not enjoyed the same popularity in the field of indoor localization.
Ronald Clark, Agathoniki Trigoni, Andrew Markham
IPSN3
2015 Accurate Positioning via Cross-Modality Training
abstract
In this paper we propose a novel algorithm for tracking people in highly dynamic industrial settings, such as construction sites. We observed both short term and long term changes in the environment; people were allowed to walk in different parts of the site on different days, the field of view of fixed cameras changed over time with the addition of walls, whereas radio and magnetic maps proved unstable with the movement of large structures. To make things worse, the uniforms and helmets that people wear for safety make them very hard to distinguish visually, necessitating the use of additional sensor modalities. In order to address these challenges, we designed a positioning system that uses both anonymous and id-linked sensor measurements and explores the use of cross-modality training to deal with environment dynamics. The system is evaluated in a real construction site and is shown to outperform state of the art multi-target tracking algorithms designed to operate in relatively stable environments.
Savvas Papaioannou, Hongkai Wen 0001, Zhuoling Xiao, Andrew Markham, Agathoniki Trigoni
SenSys4
2015 Distortion Rejecting Magneto-Inductive Three-Dimensional Localization (MagLoc)
abstract
Localization is a research area that, due to its overarching importance as an enabler for higher level services, has attracted a vast amount of research and commercial interest. For the most part, it can be claimed that GPS provides an unparalleled solution for outdoor tracking and navigation. However, the same cannot yet be said about positioning in GPS-denied or challenged environments, such as indoor environments, where obstructions such as floors and walls heavily attenuate or reflect high-frequency radio signals. This has led to a plethora of competing solutions targeted toward a particular application scenario, yielding a fragmented solution landscape. In this paper, we present a fresh approach to 3-D positioning based on the use of very low frequency (kHz) magneto-inductive (MI) fields. The most important property of MI positioning is that obstacles such as walls, floors, and people that heavily impact the performance of competing approaches are largely “transparent” to the quasi-static magnetic fields. MI has a number of challenges to robust operation that distort positions, including the presence of ferrous materials and sensitivity to user rotation. Through signal processing and sensor fusion across multiple system layers, we show how we can overcome these challenges. We showcase its highly accurate 3-D positioning in a number of environments, with positioning accuracy below 0.8 m even in heavily distorted areas.
Traian E. Abrudan, Zhuoling Xiao, Andrew Markham, Agathoniki Trigoni
IEEE J. Sel. Areas Commun.3
2015 Robust Indoor Positioning With Lifelong Learning
abstract
Inertial tracking and navigation systems have been playing an increasingly important role in indoor tracking and navigation. They have the competitive advantage of leveraging not requiring expensive infrastructure-only existing smart mobile devices with embedded inertial measurement units. When aided with other sources of information, such as radio data from existing WiFi/BLE infrastructure, and environment constraints from floor plans or radio maps, they often report great performance of 0.5-2 m. Given the promising results, what is it that prevents the widespread adoption of this tracking solution? We argue that pedestrian dead reckoning (PDR) techniques are often evaluated in a specific context and are not mature enough to handle variations in user motion, device type, device placement, or environment. They typically use a number of parameters that require careful context-specific tuning, which is labor intensive and requires expert knowledge. In this paper, we propose two novel approaches to address these problems. Our first contribution is a robust PDR algorithm, which is based on general physics principles that underpin human motion and is by design robust to context changes. The second contribution is a novel way of interaction between the PDR and map matching layers based on the principle of lifelong learning. Unlike traditional approaches where information flows unidirectionally from the PDR to the map matching layer, we introduce a feedback loop that can be used to automatically tune the parameters of the PDR algorithm. This is not dissimilar to the way that people improve their navigation skills when they repeatedly visit the same environment. Extensive experiments in multiple sites, with a variety of users, devices, and device placements, show that the combination of a robust PDR with a lifelong learning tracker can achieve submeter accuracy with no user effort for parameter tuning.
Zhuoling Xiao, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni
IEEE J. Sel. Areas Commun.3
2015 Accuracy Estimation for Sensor Systems
abstract
In most sensing applications, the measurements generated by sensor networks are noisy and usually annotated with some measure of uncertainty. The question that we address in this paper is how to estimate the accuracy of these uncertain sensor measurements. Existing studies on estimating the accuracy of uncertain measurements in real sensing applications are limited in three ways. First, they tend to be application-specific. Second, they typically employ learning techniques to estimate the parameters of sensor noise models, and ignore alternative state estimation approaches without learning. Third, they do not explore whether exploiting the dynamics of the monitored state can yield significant benefits. We address the above limitations as follows: we define the accuracy estimation problem in a general manner that applies to a broad spectrum of application scenarios. We present a general framework to address this problem, and show that the proposed framework can be implemented in a number of different ways. We evaluate and compare the different implementations in the context of two real sensing scenarios, and discuss how they trade accuracy for computation cost, and how this trade-off largely depends on the user's knowledge of the application scenario.
Hongkai Wen 0001, Zhuoling Xiao, Andrew Markham, Agathoniki Trigoni
IEEE Trans. Mob. Comput.3
2015 Indoor Tracking Using Undirected Graphical Models
abstract
Indoor tracking and navigation is a fundamental need for pervasive and context-aware smartphone applications. Although indoor maps are becoming increasingly available, there is no practical and reliable indoor map matching solution available at present. We present MapCraft, a novel, robust and responsive technique that is extremely computationally efficient (running in under 10 ms on an Android smartphone), does not require training in different sites, and tracks well even when presented with very noisy sensor data. Key to our approach is expressing the tracking problem as a conditional random field (CRF), a technique which has had great success in areas such as natural language processing. Unlike directed graphical models like Hidden Markov Models, CRFs capture arbitrary constraints that express how well observations support state transitions, given map constraints. In addition, we show how to further improve tracking accuracy, by tuning the parameters of the motion sensing model using an unsupervised EM-style optimization scheme. Extensive experiments in multiple sites show how MapCraft outperforms state-of-the art approaches, demonstrating excellent tracking error and accurate reconstruction of tortuous trajectories with zero training effort. As proof of its robustness, we also demonstrate how it is able to accurately track the position of a user from accelerometer and magnetometer measurements only (i.e., gyroand Wi-Fi-free). We believe that such an energy-efficient approach will enable always-on background localisation, enabling a new era of location-aware applications to be developed.
Zhuoling Xiao, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni
IEEE Trans. Mob. Comput.3
2015 Non-Line-of-Sight Identification and Mitigation Using Received Signal Strength
abstract
Indoor wireless systems often operate under non-line-of-sight (NLOS) conditions that can cause ranging errors for location-based applications. As such, these applications could benefit greatly from NLOS identification and mitigation techniques. These techniques have been primarily investigated for ultra-wide band (UWB) systems, but little attention has been paid to WiFi systems, which are far more prevalent in practice. In this study, we address the NLOS identification and mitigation problems using multiple received signal strength (RSS) measurements from WiFi signals. Key to our approach is exploiting several statistical features of the RSS time series, which are shown to be particularly effective. We develop and compare two algorithms based on machine learning and a third based on hypothesis testing to separate LOS/NLOS measurements. Extensive experiments in various indoor environments show that our techniques can distinguish between LOS/NLOS conditions with an accuracy of around 95%. Furthermore, the presented techniques improve distance estimation accuracy by 60% as compared to state-of-the-art NLOS mitigation techniques. Finally, improvements in distance estimation accuracy of 50% are achieved even without environment-specific training data, demonstrating the practicality of our approach to real world implementations.
Zhuoling Xiao, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni, Phil Blunsom, Jeff Frolik
IEEE Trans. Wirel. Commun.3
2014 Robust pedestrian dead reckoning (R-PDR) for arbitrary mobile device placement
abstract
Pedestrian dead reckoning, especially on smart-phones, is likely to play an increasingly important role in indoor tracking and navigation, due to its low cost and ability to work without any additional infrastructure. A challenge however, is that positioning, both in terms of step detection and heading estimation, must be accurate and reliable, even when the use of the device is so varied in terms of placement (e.g. handheld or in a pocket) or orientation (e.g holding the device in either portrait or landscape mode). Furthermore, the placement can vary over time as a user performs different tasks, such as making a call or carrying the device in a bag. A second challenge is to be able to distinguish between a true step and other periodic motion such as swinging an arm or tapping when the placement and orientation of the device is unknown. If this is not done correctly, then the PDR system typically overestimates the number of steps taken, leading to a significant long term error. We present a fresh approach, robust PDR (R-PDR), based on exploiting how bipedal motion impacts acquired sensor waveforms. Rather than attempting to recognize different placements through sensor data, we instead simply determine whether the motion of one or both legs impact the measurements. In addition, we formulate a set of techniques to accurately estimate the device orientation, which allows us to very accurately (typically over 99%) reject false positives. We demonstrate that regardless of device placement, we are able to detect the number of steps taken with >99.4% accuracy. R-PDR thus addresses the two main limitations facing existing PDR techniques.
Zhuoling Xiao, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni
IPIN3
2014 Lightweight map matching for indoor localisation using conditional random fields
Zhuoling Xiao, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni
IPSN3
2014 Fusion of Radio and Camera Sensor Data for Accurate Indoor Positioning
abstract
Indoor positioning systems have received a lot of attention recently due to their importance for many location-based services, e.g. indoor navigation and smart buildings. Lightweight solutions based on WiFi and inertial sensing have gained popularity, but are not fit for demanding applications, such as expert museum guides and industrial settings, which typically require sub-meter location information. In this paper, we propose a novel positioning system, RAVEL (Radio And Vision Enhanced Localization), which fuses anonymous visual detections captured by widely available camera infrastructure, with radio readings (e.g. WiFi radio data). Although visual trackers can provide excellent positioning accuracy, they are plagued by issues such as occlusions and people entering/exiting the scene, preventing their use as a robust tracking solution. By incorporating radio measurements, visually ambiguous or missing data can be resolved through multi-hypothesis tracking. We evaluate our system in a complex museum environment with dim lighting and multiple people moving around in a space cluttered with exhibit stands. Our experiments show that although the WiFi measurements are not by themselves sufficiently accurate, when they are fused with camera data, they become a catalyst for pulling together ambiguous, fragmented, and anonymous visual tracklets into accurate and continuous paths, yielding typical errors below 1 meter.
Savvas Papaioannou, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni
MASS3
2013 Comparison of Accuracy Estimation Approaches for Sensor Networks
abstract
With sensor technology gaining maturity and becoming ubiquitous, we are experiencing an unprecedented wealth of sensor data. In most sensing applications, users receive sensor measurements, which are prone to error. As a result, they are often annotated with some measure of uncertainty, such as the distribution variance or a confidence interval, and will be hereafter referred to as probabilistic measurements. The question that we address in this paper is how to estimate the accuracy of these probabilistic measurements, that is, how far they lie from the ground truth of the measured attribute. Existing studies on estimating the accuracy of probabilistic measurements in real sensing applications are limited in three ways. First, they tend to be application-specific. Second, they typically employ learning techniques to estimate the parameters of sensor noise models, and ignore alternative approaches that rely on simple state estimation without learning. Third, they do not explore whether exploiting the dynamics of the monitored state can yield significant benefits in terms of accuracy estimation. In this paper, we address the above limitations as follows: We define the problem of accuracy estimation in a general way that applies to a wide spectrum of application scenarios. We then propose a taxonomy of accuracy estimation techniques, which include both state estimation and parameter learning. These techniques are further subdivided into static and dynamic, depending on whether they exploit knowledge of system dynamics. All different approaches in the taxonomy are then applied and compared with each other in the context of two real sensing applications. We discuss how they trade accuracy for computation cost, and how this tradeoff largely depends on the user's knowledge of the application scenario.
Hongkai Wen 0001, Zhuoling Xiao, Andrew Colquhoun Symington, Andrew Markham, Agathoniki Trigoni
DCOSS4
2013 Identification and mitigation of non-line-of-sight conditions using received signal strength
abstract
Various applications, such as localisation of persons and objects could benefit greatly from non-line-of-sight (NLOS) identification and mitigation techniques. However, such techniques have been primarily investigated for ultra-wide band (UWB) signals, leaving the area of WiFi signals untouched. In this study, we propose two accurate approaches using only received signal strength (RSS) measurements from WiFi signals to identify NLOS conditions and mitigate the effects. We first explore several features from the RSS which are later demonstrated as very effective in identifying and mitigating NLOS conditions. After that, we develop and compare two major optimization problems based on a machine learning technique and hypothesis testing according to different user requirements and information available. Extensive experiments in various indoor environments have shown that our techniques can not only accurately distinguish between LOS/NLOS conditions, but also mitigate the impact of NLOS conditions as well.
Zhuoling Xiao, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni, Phil Blunsom, Jeff Frolik
WiMob3
2013 Human interactive secure key and identity exchange protocols in body sensor networks
abstract
A body sensor network (BSN) is typically a wearable wireless sensor network. Security protection is critical to BSNs, since they collect sensitive personal information. Generally speaking, security protection of BSN relies on identity (ID) and key distribution protocols. Most existing protocols are designed to run in general wireless sensor networks, and are not suitable for BSNs. After carefully examining the characteristics of BSNs, the authors propose human interactive empirical channel‐based security protocols, which include an elliptic curve Diffie–Hellman version of symmetric hash commitment before knowledge protocol and an elliptic curve Diffie–Hellman version of hash commitment before knowledge protocol. Using these protocols, dynamically distributing keys and IDs become possible. As opposite to present solutions, these protocols do not need any pre‐deployment of keys or secrets. Therefore compromised and expired keys or IDs can be easily changed. These protocols exploit human users as temporary trusted third parties. The authors, thus, show that the human interactive channels can help them to design secure BSNs.
Xin Huang 0005, Bangdao Chen, Andrew Markham, Qinghua Wang 0001, Zheng Yan 0002, A. W. Roscoe 0001
IET Inf. Secur.3
2012 Magneto-inductive networked rescue system (MINERS): taking sensor networks underground
abstract
Wireless underground networks are an emerging technology which have application in a number of scenarios. For example, in a mining disaster, flooding or a collapse can isolate portions of underground tunnels, severing wired communication links and preventing radio communication. In this paper, we explore the use of low frequency magnetic fields for communication, and present a new hardware platform that features triaxial transmitter/receiver antenna loops. We point out that the fundamental problem of the magnetic channel is the limited bitrate at long ranges, due to the extreme path loss of 60 dB/decade. To this end, we present two complementary techniques to address this limitation. Firstly, we demonstrate magnetic vector modulation, a technique which modulates the three dimensional orientation of the magnetic vector. This increases the gross bitrate by a factor of over 2.5, without an increase in transmission power or bandwidth. Secondly, we show how in a multi-hop network latencies can be dramatically reduced by receiving multiple parallel streams of frequency multiplexed data in a many-to-one configuration. These techniques are demonstrated on a working hardware platform, which for flexible operation, features a software defined magnetic transceiver. Typical communication range is approximately 30 m through rock.
Andrew Markham, Agathoniki Trigoni
IPSN1
2012 WILDSENSING: Design and deployment of a sustainable sensor network for wildlife monitoring
abstract
The increasing adoption of wireless sensor network technology in a variety of applications, from agricultural to volcanic monitoring, has demonstrated their ability to gather data with unprecedented sensing capabilities and deliver it to a remote user. However, a key issue remains how to maintain these sensor network deployments over increasingly prolonged deployments. In this article, we present the challenges that were faced in maintaining continual operation of an automated wildlife monitoring system over a one-year period. This system analyzed the social colocation patterns of European badgers ( Meles meles ) residing in a dense woodland environment using a hybrid RFID-WSN approach. We describe the stages of the evolutionary development, from implementation, deployment, and testing, to various iterations of software optimization, followed by hardware enhancements, which in turn triggered the need for further software optimization. We highlight the main lessons learned: the need to factor in the maintenance costs while designing the system; to consider carefully software and hardware interactions; the importance of rapid prototyping for initial deployment (this was key to our success); and the need for continuous interaction with domain scientists which allows for unexpected optimizations.
Vladimir Dyo, Stephen A. Ellwood, David W. Macdonald, Andrew Markham, Agathoniki Trigoni, Ricklef Wohlers, Cecilia Mascolo, Bence Pásztor, Salvatore Scellato, Kharsim Yousef
ACM Trans. Sens. Networks4
2011 The Automatic Evolution of Distributed Controllers to Configure Sensor Network Operation
abstract
Tuning the parameters that control the operation of a wireless sensor network, such as sampling rate, is not a simple task. This is partly due to the distributed nature of the problem, but is also a result of the time-varying dynamics that a network experiences. Inspired by the way in which cells alter their behaviour in response to diffused protein concentrations, an abstract representation, termed a discrete gene regulatory network (dGRN), is introduced. Each node runs an identical dGRN controller which controls node activity and interaction. The controllers are authored automatically using an evolutionary algorithm. The communication that occurs between nodes is neither specified nor designed, but emerges naturally. As a particular example, we illustrate that our approach can generate effective strategies for nodes to cooperatively track a moving target. The obtained strategies vary according to the user's accuracy requirements and the speed of the target, and are similar to those which would be expected from a network engineer. We also present results from our proof-of-concept dGRN implementation on T-Mote Sky nodes. Our approach takes high-level user application requirements and from these, automatically generates distributed parameter tuning algorithms. The dGRN framework thus greatly reduces the amount of effort involved in adjusting a sensor network's operation.
Andrew Markham, Agathoniki Trigoni
Comput. J.1
2010 Discrete Gene Regulatory Networks (dGRNs): A Novel Approach to Configuring Sensor Networks
abstract
The operation of a sensor network is determined by a large number of parameters, such as the radio duty cycle, the frequency of neighbor discovery beacons, and the rate of sampling sensors. Writing adaptive algorithms to tune these parameters in dynamic network conditions is a challenging task that requires expert knowledge, and many design-test-rewrite cycles. This paper proposes a novel nature-inspired paradigm, termed discrete Gene Regulatory Network (dGRN), for configuring sensor networks. The idea is that nodes should regulate their parameters based on their local state and state communicated from neighbor nodes, in a similar manner that cells regulate their behavior based on local levels of protein concentrations, and proteins diffused from neighbor cells. The proposed dGRN paradigm has two major strengths: 1) it is general-purpose, and can be applied to a variety of parameter tuning problems; and 2) it generates parameter tuning code automatically removing the need for a human expert. We demonstrate the feasibility of the dGRN approach in a scenario where nodes must tune their sampling rates to track a moving target with a certain accuracy. Theautomaticallygenerated code exhibits properties similar to the ones that one would expect from expert-designed code, such as aggressive sampling when the target moves fast and the sensing range is low, and relaxed sampling otherwise. Moreover, the automatically generated code causes nodes to communicate with each other to coordinate their tuning tasks, as one would expect from expert-designed code. The resulting dGRN code is evaluated both in a simulation environment, and in a real environment with eight T-Mote Sky nodes tracking a light-emitting target.
Andrew Markham, Agathoniki Trigoni
INFOCOM1
2010 Evolution and sustainability of a wildlife monitoring sensor network
abstract
As sensor network technologies become more mature, they are increasingly being applied to a wide variety of applications, ranging from agricultural sensing to cattle, oceanic and volcanic monitoring. Significant efforts have been made in deploying and testing sensor networks resulting in unprecedented sensing capabilities. A key challenge has become how to make these emerging wireless sensor networks more sustainable and easier to maintain over increasingly prolonged deployments.
Vladimir Dyo, Stephen A. Ellwood, David W. Macdonald, Andrew Markham, Cecilia Mascolo, Bence Pásztor, Salvatore Scellato, Agathoniki Trigoni, Ricklef Wohlers, Kharsim Yousef
SenSys4
2010 Revealing the hidden lives of underground animals using magneto-inductive tracking
abstract
Currently, there is no existing method for automatically tracking the location of burrowing animals when they are underground, consequently zoologists only have a partial view of their subterranean behaviour and habits. Conventional RF based methods of localization are unsuitable because electromagnetic waves are severely attenuated by soil and moisture. Here, we use an as yet unexploited method of localization, namely magneto-inductive (MI) localization. Magnetic fields are not affected by soil or water, and thus have virtually unattenuated ground penetration. In this paper, we present a method that allows the position of an animal to be determined through soil. Not only does this enable the study of behaviour, it also allows the structure of the tunnel to be automatically mapped as the animal moves through it. We describe the application for tracking wild European Badgers (Meles meles) within their burrows, providing experimental data from a two month deployment.
Andrew Markham, Agathoniki Trigoni, Stephen A. Ellwood, David W. Macdonald
SenSys1
2010 Magneto-inductive tracking of underground animals
abstract
Existing sensor network deployments for wildlife tracking (e.g. ZebraNet [1]) have concentrated on monitoring animal behaviour above-ground. However, a wide variety of animals create underground tunnels for shelter and protection whilst the animal is asleep. The extent and internal architecture of the underground structure varies considerably amongst species. For example, badgers excavate wide ranging underground tunnel systems which a number of animals inhabit in a community [6]. In addition, their tunnel structure is something which is often only determined by the destructive extreme of excavation [6]. As fossorial animals spend a large proportion of their lifetimes underground, this means that zoologists only have a partial view of their behaviour and habits. There is thus a need for a system which can localize animals whilst they are underground, in a non-invasive and automatic way.
Andrew Markham, Agathoniki Trigoni, Stephen A. Ellwood, David W. Macdonald
SenSys1
2009 Wildlife and environmental monitoring using RFID and WSN technology
abstract
Wireless Sensor Networks enable scientists to collect information about the environment with a granularity unseen before, while providing numerous challenges to software designers. Since sensor devices are often powered by small batteries, which take considerable effort to replace, it is of major importance to use energy carefully. We present two efficient ways of extending the lifetime of such systems: 1. an adaptive duty cycling protocol and 2. an adaptive data management protocol. Further, we present some details of our deployed sensor network in Wytham Woods, Oxfordshire.
Vladimir Dyo, Stephen A. Ellwood, David W. Macdonald, Andrew Markham, Cecilia Mascolo, Bence Pásztor, Agathoniki Trigoni, Ricklef Wohlers
SenSys4