VLDB 2026 Research / reviewers in the wild / expert
Sen Wang 0002
dblp:69/6403-2
· DBLP profile ↗
42ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0003-1537-8834ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 4 first-author · 13 since 2021Systems, architecture and hardware · 24 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 since 2021Computer networks · 3Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Estimating Fog Parameters From a Sequence of Stereo ImagesabstractWe propose a method which, given a sequence of stereo foggy images, estimates the parameters of a fog model and updates them dynamically. In contrast with previous approaches, which estimate the parameters sequentially and thus are prone to error propagation, our algorithm estimates all the parameters simultaneously by solving a novel optimisation problem. By assuming that fog is only locally homogeneous, our method effectively handles real-world fog, which is often globally inhomogeneous. The proposed algorithm can be easily used as an add-on module in existing visual Simultaneous Localisation and Mapping (SLAM) or odometry systems in the presence of fog. In order to assess our method, we also created a new dataset, the Stereo Driving In Real Fog (SDIRF), consisting of high-quality, consecutive stereo frames of real, foggy road scenes under a variety of visibility conditions, totalling over 40 minutes and 34 k frames. As a first-of-its-kind, SDIRF contains the camera's photometric parameters calibrated in a lab environment, which is a prerequisite for correctly applying the atmospheric scattering model to foggy images. The dataset also includes the counterpart clear data of the same routes recorded in overcast weather, which is useful for companion work in image defogging and depth reconstruction. We conducted extensive experiments using both synthetic foggy data and real foggy sequences from SDIRF to demonstrate the superiority of the proposed algorithm over prior methods. Our method not only produces the most accurate estimates on synthetic data, but also adapts better to real fog. Yining Ding, João F. C. Mota, Andrew M. Wallace, Sen Wang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Digital Beamforming Enhanced Radar OdometryabstractRadar has become an essential sensor for autonomous navigation, especially in challenging environments where camera and LiDAR sensors fail. 4D single-chip millimeter-wave radar systems, in particular, have drawn increasing attention thanks to their ability to provide spatial and Doppler information with low hardware cost and power consumption. However, most single-chip radar systems using traditional signal processing, such as Fast Fourier Transform, suffer from limited spatial resolution in radar detection, significantly limiting the performance of radar-based odometry and Simultaneous Localization and Mapping (SLAM) systems. In this paper, we develop a novel radar signal processing pipeline that integrates spatial domain beamforming techniques, and extend it to 3D Direction of Arrival estimation. Experiments using public datasets are conducted to evaluate and compare the performance of our proposed signal processing pipeline against traditional methodologies. These tests specifically focus on assessing structural precision across diverse scenes and measuring odometry accuracy in different radar odometry systems. This research demonstrates the feasibility of achieving more accurate radar odometry by simply replacing the standard FFT-based processing with the proposed pipeline. The codes are available at GitHub**https://github.com/SenseRoboticsLab/DBE-Radar. Jingqi Jiang, Shida Xu, Jiyuan Wei, Sen Wang 0002 |
ICRA | 6 |
| 2025 | AQUA-SLAM: Tightly Coupled Underwater Acoustic-Visual-Inertial SLAM With Sensor CalibrationabstractUnderwater environments pose significant challenges for visual simultaneous localization and mapping (SLAM) systems due to limited visibility, inadequate illumination, and sporadic loss of structural features in images. Addressing these challenges, this article introduces a novel, tightly coupled acoustic-visual-inertial SLAM approach, termed AQUA-SLAM, to fuse a Doppler velocity log (DVL), a stereo camera, and an inertial measurement unit (IMU) within a graph optimization framework. Moreover, we propose an efficient sensor calibration technique, encompassing the multisensor extrinsic calibration (among the DVL, camera, and IMU) and the DVL transducer misalignment calibration, with a fast linear approximation procedure for real-time online execution. The proposed methods are extensively evaluated in a tank environment with ground truth, and validated for offshore applications in the North Sea. The results demonstrate that our method surpasses current state-of-the-art underwater and visual-inertial SLAM systems in terms of localization accuracy and robustness. The proposed system will be made open-source for the community. Shida Xu, Sen Wang 0002 |
IEEE Trans. Robotics | 3 |
| 2025 | CURL-SLAM: Continuous and Compact LiDAR MappingabstractThis paper studies 3D LiDAR mapping with a focus on developing an updatable and localizable map representation that enables continuity, compactness and consistency in 3D maps. Traditional LiDAR Simultaneous Localization and Mapping (SLAM) systems often rely on 3D point cloud maps, which typically require extensive storage to preserve structural details in large-scale environments. In this paper, we propose a novel paradigm for LiDAR SLAM by leveraging the Continuous and Ultra-compact Representation of LiDAR (CURL) introduced in [1]. Our proposed LiDAR mapping approach, CURL-SLAM, produces compact 3D maps capable of continuous reconstruction at variable densities using CURL's spherical harmonics implicit encoding, and achieves global map consistency after loop closure. Unlike popular Iterative Closest Point (ICP)-based LiDAR odometry techniques, CURL-SLAM formulates LiDAR pose estimation as a unique optimization problem tailored for CURL and extends it to local Bundle Adjustment (BA), enabling simultaneous pose refinement and map correction. Experimental results demonstrate that CURL-SLAM achieves state-of-the-art 3D mapping quality and competitive LiDAR trajectory accuracy, delivering sensor-rate real-time performance (10 Hz) on a CPU. We will release the CURL-SLAM implementation to the community. Shida Xu, Yining Ding, Xianwen Kong, Sen Wang 0002 |
IEEE Trans. Robotics | 5 |
| 2024 | DISO: Direct Imaging Sonar OdometryabstractThis paper introduces a novel sonar odometry system that estimates the relative spatial transformation between two sonar image frames. Considering the unique challenges, such as low resolution and high noise, of sonar imagery for odometry and Simultaneous Localization and Mapping (SLAM), the proposed Direct Imaging Sonar Odometry (DISO) system is designed to estimate the relative transformation between two sonar frames by minimizing the aggregated sonar intensity errors of points with high intensity gradients. Moreover, DISO is implemented to incorporate a multi-sensor window optimization technique, a data association strategy and an acoustic intensity outlier rejection algorithm for reliability and accuracy. The effectiveness of DISO is evaluated using both simulated and real-world sonar datasets, showing that it outperforms the existing geometric-only method on localization accuracy and achieves state-of-the-art sonar odometry performance. We release the source codes of the DISO implementation to the community. The source code is available at https://github.com/SenseRoboticsLab/DISO. Shida Xu, Ziyang Hong 0001, Yuanchang Liu, Sen Wang 0002 |
ICRA | 5 |
| 2024 | CURL-MAP: Continuous Mapping and Positioning with CURL Representation†abstractMaps of LiDAR Simultaneous Localisation and Mapping (SLAM) are often represented as point clouds. They usually take up a huge amount of storage space for large-scale environments, otherwise much structural detail may not be kept. In this paper, a novel paradigm of LiDAR mapping and odometry is designed by leveraging the Continuous and Ultra-compact Representation of LiDAR (CURL) proposed in [1]. Termed CURL-MAP (Mapping and Positioning), the proposed approach can not only reconstruct 3D maps with a continuously varying density but also efficiently reduce map storage space by using CURL’s spherical harmonics implicit encoding. Different from the popular Iterative Closest Point (ICP) based LiDAR odometry techniques, CURL-MAP formulates LiDAR pose estimation as a unique optimisation problem tailored for CURL. Experiment evaluation shows that CURL-MAP achieves state-of-the-art 3D mapping results and competitive LiDAR odometry accuracy. We will release the CURL-MAP codes for the community. Yining Ding, Shida Xu, Ziyang Hong 0001, Xianwen Kong, Sen Wang 0002 |
ICRA | 6 |
| 2024 | Estimating Fog Parameters from an Image Sequence using Non-linear OptimisationabstractGiven a sequence of images taken in foggy weather, we seek to estimate the atmospheric light and the scattering coefficient. These are key parameters to characterise the nature of the fog, to reconstruct a clear image (defogging), and to infer scene depth. Existing methods adopt a sequential estimation strategy which is prone to error propagation. In sharp contrast, we take a more systematic approach and jointly estimate these parameters by solving a unified non-linear optimisation problem. Experimental results show that the proposed method is superior to existing ones in terms of both estimation accuracy and precision. Our method further demonstrates how image defogging and depth estimation can be linked to a visual localisation system, contributing to more comprehensive and robust perception in fog. Yining Ding, Andrew M. Wallace, Sen Wang 0002 |
WACV | 3 |
| 2024 | Deep Color Compensation for Generalized Underwater Image EnhancementabstractUnderwater images suffer from quality degradation due to the underwater light absorption and scattering. It remains challenging to enhance underwater images using deep learning-based methods since the scarcity of real-world underwater images and their enhanced counterparts. Although existing works manually select well-enhanced images as reference images to train enhancement networks in an end-to-end manner, their performance tends to be inferior in some scenarios. We argue that the manually selected reference images cannot approximate their ground truth perfectly, leading to imbalanced learning and domain shift in enhancement networks. To address this issue, we analyse widely used underwater datasets from the perspective of color spectrum distribution and surprisingly find the sound color spectrum distribution of the enhanced reference images compared to in-air datasets. Based on this perceptive observation, instead of directly learning the enhancement mapping, we propose a novel methodology to learn color compensation for general purposes. Specifically, we present a probabilistic color compensation network that estimates the probabilistic distribution of colors by multi-scale volumetric fusion of texture and color features. We further propose a novel two-stage enhancement framework that first performs color compensation and then enhancement, which is highly flexible to be integrated with an existing enhancement method without tuning. Extensive experiments on underwater image enhancement across various challenging scenarios show that our proposed approach consistently improves the results of the popular conventional and learning-based methods by a significant margin. Moreover, our enhanced images achieve superior performance on underwater salient object detection and visual 3D reconstruction, demonstrating that our method can successfully break through the generalization bottleneck of existing learning-based enhancement models. Our implementation will be made available at https://github.com/Ray2OUC/P2CNet. Yuan Rao 0001, Kunqian Li, Hao Fan 0004, Sen Wang 0002, Junyu Dong |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Large-Scale Radar Localization using Online Public MapsabstractIn this paper, we propose using online public maps, e.g., OpenStreetMap (OSM), for large-scale radar-based localization without needing a prior sensing map. This can potentially extend the localization system to anywhere worldwide without building, saving, or maintaining a sensing map, as long as an online public map covers the operating area. Existing methods using OSM only use route network or semantics information. These two sources of information are not combined in the previous works, while our proposed system fuses them to improve localization accuracy. Our experiments, on three open datasets collected from three different continents, show that the proposed system outperforms the state-of-the-art localization methods, reducing up to 50% of position errors. We release an open-source implementation for the community. Ziyang Hong 0001, Yvan R. Petillot, Shida Xu, Sen Wang 0002 |
ICRA | 5 |
| 2023 | Observability-Aware Active Extrinsic Calibration of Multiple SensorsabstractThe extrinsic parameters play a crucial role in multi-sensor fusion, such as visual-inertial Simultaneous Localization and Mapping(SLAM), as they enable the accurate alignment and integration of measurements from different sensors. However, extrinsic calibration is challenging in scenarios, such as underwater, where in-view structures are scanty and visibility is limited, causing incorrect extrinsic calibration due to insufficient motion on all degrees of freedom. In this paper, we propose an entropy-based active extrinsic calibration algorithm leverages observability analysis and information entropy to enhance the accuracy and reliability of extrinsic calibration. It determines the system observability numerically by using singular value decomposition (SVD) of the Fisher Information Matrix (FIM). Furthermore, when the extrinsic parameter is not fully observable, our method actively searches for the next best motion to recover the system's observability via entropy-based optimization. Experimental results on synthetic data, in a simulation, and using an actual underwater vehicle verify that the proposed method is able to avoid the calibration failure while improving the calibration accuracy and reliability. Shida Xu, Jonatan Scharff Willners, Ziyang Hong 0001, Yvan R. Petillot, Sen Wang 0002 |
ICRA | 6 |
| 2023 | Reliability Assessment and Safety Arguments for Machine Learning Components in System AssuranceabstractThe increasing use of Machine Learning (ML) components embedded in autonomous systems—so-called Learning-Enabled Systems (LESs)—has resulted in the pressing need to assure their functional safety. As for traditional functional safety, the emerging consensus within both, industry and academia, is to use assurance cases for this purpose. Typically assurance cases support claims of reliability in support of safety, and can be viewed as a structured way of organising arguments and evidence generated from safety analysis and reliability modelling activities. While such assurance activities are traditionally guided by consensus-based standards developed from vast engineering experience, LESs pose new challenges in safety-critical application due to the characteristics and design of ML models. In this article, we first present an overall assurance framework for LESs with an emphasis on quantitative aspects, e.g., breaking down system-level safety targets to component-level requirements and supporting claims stated in reliability metrics. We then introduce a novel model-agnostic Reliability Assessment Model (RAM) for ML classifiers that utilises the operational profile and robustness verification evidence. We discuss the model assumptions and the inherent challenges of assessing ML reliability uncovered by our RAM and propose solutions to practical use. Probabilistic safety argument templates at the lower ML component-level are also developed based on the RAM. Finally, to evaluate and demonstrate our methods, we not only conduct experiments on synthetic/benchmark datasets but also scope our methods with case studies on simulated Autonomous Underwater Vehicles and physical Unmanned Ground Vehicles. Yi Dong 0002, Wei Huang 0035, Vibhav Bharti, Victoria Cox, Alec Banks, Sen Wang 0002, Xingyu Zhao 0001, Sven Schewe, Xiaowei Huang 0001 |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2022 | Variational Simultaneous Stereo Matching and Defogging in Low Visibility
Yining Ding, Andrew M. Wallace, Sen Wang 0002 |
BMVC | 3 |
| 2022 | Autonomous Pipeline Tracking Using Bernoulli Filter for Unmanned Underwater SurveysabstractInspection of subsea pipelines is crucial for avoiding any hazards and minimizing the risks to infrastructure and the environment. These inspections are achieved using Autonomous Underwater Vehicles (AUVs) in favour of reduced operational costs. This work presents a vehicle agnostic approach for tracking subsea pipelines at close-range for autonomous guidance along the pipeline using an AUV. A multibeam echosounder is used as the primary tracking sensor augmented by fluxgate magnetometers that can track buried pipelines over short ranges until they are exposed again. A Bernoulli filter is proposed to efficiently track pipelines in presence of environmental clutter. Field experiments were carried out in a dock and at sea to evaluate the proposed system and the benefits of using a Bernoulli filter over a Kalman filter-based solution. Vibhav Bharti, Sen Wang 0002 |
IROS | 2 |
| 2022 | Monocular Depth Estimation for Equirectangular VideosabstractDepth estimation from panoramic imagery has received minimal attention in contrast to standard perspective imagery, which constitutes the majority of the literature on the key research topic. The vast - and frequently complete - field of view provided by such panoramic photographs makes them appealing for a variety of applications, including robots, autonomous vehicles, and virtual reality. Consumer-level camera systems capable of capturing such images are likewise growing more affordable, and may be desirable complements to autonomous systems' sensor packages. They do, however, introduce significant distortions and violate some assumptions regarding perspective view images. Additionally, many state-of-the-art algorithms are not designed for its projection model, and their depth estimation performance tends to degrade when being applied to panoramic imagery. This paper presents a novel technique for adapting view synthesis-based depth estimation models to omnidirectional vision. Specifically, we: 1) integrate a “virtual” spherical camera model into the training pipeline, facilitating the model training, 2) exploit spherical convolutional layers to perform convolution operations on equirectangular images, handling the severe distortion, and 3) propose an optical flow-based masking scheme to mitigate the effect of unwanted pixels during training. Our qualitative and quantitative results demonstrate that these simple yet efficient designs result in significantly improved depth estimations when compared to previous approaches. Helmi Fraser, Sen Wang 0002 |
IROS | 2 |
| 2021 | RADIATE: A Radar Dataset for Automotive Perception in Bad WeatherabstractDatasets for autonomous cars are essential for the development and benchmarking of perception systems. However, most existing datasets are captured with camera and LiDAR sensors in good weather conditions. In this paper, we present the RAdar Dataset In Adverse weaThEr (RADIATE), aiming to facilitate research on object detection, tracking and scene understanding using radar sensing for safe autonomous driving. RADIATE includes 3 hours of annotated radar images with more than 200K labelled road actors in total, on average about 4.6 instances per radar image. It covers 8 different categories of actors in a variety of weather conditions (e.g., sun, night, rain, fog and snow) and driving scenarios (e.g., parked, urban, motorway and suburban), representing different levels of challenge. To the best of our knowledge, this is the first public radar dataset which provides high-resolution radar images on public roads with a large amount of road actors labelled. The data collected in adverse weathers, e.g., fog and snowfall, is unique. Some baseline results of radar based object detection and recognition are given to show that the use of radar data is promising for automotive applications in bad weather, where vision and LiDAR fail. RADIATE also has stereo images, 32-channel LiDAR and GPS data, directed at other applications such as sensor fusion, localisation and mapping. The public dataset can be accessed at http://pro.hw.ac.uk/radiate/. Marcel Sheeny, Emanuele De Pellegrin, Saptarshi Mukherjee, Alireza Ahrabian, Sen Wang 0002, Andrew M. Wallace |
ICRA | 5 |
| 2021 | Robust Underwater Visual SLAM Fusing Acoustic SensingabstractIn this paper, we propose an approach for robust visual Simultaneous Localisation and Mapping (SLAM) in underwater environments leveraging acoustic, inertial and altimeter/depth sensors. Underwater visual SLAM is challenging due to factors including poor visibility caused by suspended particles in water, a lack of light and insufficient texture in the scene. Because of this, many state-of-the-art approaches rely on acoustic sensing instead of vision for underwater navigation.Building on the sparse visual SLAM system ORB-SLAM2, this paper proposes to improve the robustness of camera pose estimation in underwater environments by leveraging acoustic odometry, which derives a drifting estimate of the 6-DoF robot pose from fusion of a Doppler Velocity Log (DVL), a gyroscope and an altimeter or depth sensor. Acoustic odometry estimates are used as motion priors and we formulate pose residuals that are integrated within the camera pose tracking, local and global bundle adjustment procedures of ORB-SLAM2.The original design of ORB-SLAM2 supports a single map and it enters relocalisation when tracking is lost. This is a significant problem for scenarios where a robot does a continuous scanning motion without returning to a previously visited location. One of our main contributions is to enable the system to create a new map whenever it encounters a new scene where visual odometry can work. This new map is connected with its predecessor in a common graph using estimates from the proposed acoustic odometry. Experimental results on two underwater vehicles demonstrate the increased robustness of our approach compared to baseline ORB-SLAM2 in both controlled, uncontrolled and field environments. Elizabeth Vargas, Raluca Scona, Jonatan Scharff Willners, Tomasz Luczynski, Yu Cao 0007, Sen Wang 0002, Yvan R. Petillot |
ICRA | 6 |
| 2021 | Underwater Visual Acoustic SLAM with Extrinsic CalibrationabstractUnderwater scenarios are challenging for visual Simultaneous Localization and Mapping (SLAM) due to limited visibility and intermittently losing structures in image views. In this paper, we propose a visual acoustic bundle adjustment system which fuses a camera and a Doppler Velocity Log (DVL) in a graph SLAM framework for reliable underwater localization and mapping. In order to fuse the vision with the acoustic measurements, an calibration algorithm is also designed to estimate extrinsic parameters between a camera and a DVL using features detected in scenes. Experimental results in a tank and an offshore wind farm show the proposed method can achieve better robustness and localization accuracy than pure visual SLAM, especially in visually challenging scenarios, and the extrinsic calibration parameters can be accurately estimated, even when initialized with a random guess. Shida Xu, Tomasz Luczynski, Jonatan Scharff Willners, Ziyang Hong 0001, Yvan R. Petillot, Sen Wang 0002 |
IROS | 7 |
| 2021 | Learning With Stochastic Guidance for Robot NavigationabstractDue to the sparse rewards and high degree of environmental variation, reinforcement learning approaches, such as deep deterministic policy gradient (DDPG), are plagued by issues of high variance when applied in complex real-world environments. We present a new framework for overcoming these issues by incorporating a stochastic switch, allowing an agent to choose between high- and low-variance policies. The stochastic switch can be jointly trained with the original DDPG in the same framework. In this article, we demonstrate the power of the framework in a navigation task, where the robot can dynamically choose to learn through exploration or to use the output of a heuristic controller as guidance. Instead of starting from completely random actions, the navigation capability of a robot can be quickly bootstrapped by several simple independent controllers. The experimental results show that with the aid of stochastic guidance, we are able to effectively and efficiently train DDPG navigation policies and achieve significantly better performance than state-of-the-art baseline models. Linhai Xie, Yishu Miao, Sen Wang 0002, Phil Blunsom, Zhihua Wang 0005, Changhao Chen, Andrew Markham, Agathoniki Trigoni |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | DeepBEV: A Conditional Adversarial Network for Bird's Eye View GenerationabstractObtaining a meaningful, interpretable yet compact representation of the immediate surroundings of an autonomous vehicle is paramount for effective operation as well as safety. This paper proposes a solution to this by representing semantically important objects from a top-down, ego-centric bird's eye view. The novelty in this work is from formulating this problem as an adversarial learning task, tasking a generator model to produce bird's eye view representations which are plausible enough to be mistaken as a ground truth sample. This is achieved by using a Wasserstein Generative Adversarial Network based model conditioned on object detections from monocular RGB images and the corresponding bounding boxes. Extensive experiments show our model is more robust to novel data compared to strictly supervised benchmark models, while being a fraction of the size of the next best. Helmi Fraser, Sen Wang 0002 |
ICPR | 2 |
| 2020 | Interacting Vehicle Trajectory Prediction with Convolutional Recurrent Neural NetworksabstractAnticipating the future trajectories of surrounding vehicles is a crucial and challenging task in path planning for autonomy. We propose a novel Convolutional Long Short Term Memory (Conv-LSTM) based neural network architecture to predict the future positions of cars using several seconds of historical driving observations. This consists of three modules: 1) Interaction Learning to capture the effect of surrounding cars, 2) Temporal Learning to identify the dependency on past movements and 3) Motion Learning to convert the extracted features from these two modules into future positions. To continuously achieve accurate prediction, we introduce a novel feedback scheme where the current predicted positions of each car are leveraged to update future motion, encapsulating the effect of the surrounding cars. Experiments on two public datasets demonstrate that the proposed methodology can match or outperform the state-of-the-art methods for long-term trajectory prediction. Saptarshi Mukherjee, Sen Wang 0002, Andrew M. Wallace |
ICRA | 2 |
| 2020 | RadarSLAM: Radar based Large-Scale SLAM in All WeathersabstractNumerous Simultaneous Localization and Mapping (SLAM) algorithms have been presented in last decade using different sensor modalities. However, robust SLAM in extreme weather conditions is still an open research problem. In this paper, RadarSLAM, a full radar based graph SLAM system, is proposed for reliable localization and mapping in large-scale environments. It is composed of pose tracking, local mapping, loop closure detection and pose graph optimization, enhanced by novel feature matching and probabilistic point cloud generation on radar images. Extensive experiments are conducted on a public radar dataset and several self-collected radar sequences, demonstrating the state-of-the-art reliability and localization accuracy in various adverse weather conditions, such as dark night, dense fog and heavy snowfall. Ziyang Hong 0001, Yvan R. Petillot, Sen Wang 0002 |
IROS | 3 |
| 2020 | SolarSLAM: Battery-free Loop Closure for Indoor LocalisationabstractIn this paper, we propose SolarSLAM, a batteryfree loop closure method for indoor localisation. Inertial Measurement Unit (IMU) based indoor localisation method has been widely used due to its ubiquity in mobile devices, such as mobile phones, smartwatches and wearable bands. However, it suffers from the unavoidable long term drift. To mitigate the localisation error, many loop closure solutions have been proposed using sophisticated sensors, such as cameras, laser, etc. Despite achieving high-precision localisation performance, these sensors consume a huge amount of energy. Different from those solutions, the proposed SolarSLAM takes advantage of an energy harvesting solar cell as a sensor and achieves effective battery-free loop closure method. The proposed method suggests the key-point dynamic time warping for detecting loops and uses robust simultaneous localisation and mapping (SLAM) as the optimiser to remove falsely recognised loop closures. Extensive evaluations in the real environments have been conducted to demonstrate the advantageous photocurrent characteristics for indoor localisation and good localisation accuracy of the proposed method. Bo Wei 0003, Weitao Xu, Chengwen Luo 0001, Guillaume Zoppi, Dong Ma 0001, Sen Wang 0002 |
IROS | 6 |
| 2020 | Robust Attentional Aggregation of Deep Feature Sets for Multi-view 3D ReconstructionabstractWe study the problem of recovering an underlying 3D shape from a set of images. Existing learning based approaches usually resort to recurrent neural nets, e.g., GRU, or intuitive pooling operations, e.g., max/mean poolings, to fuse multiple deep features encoded from input images. However, GRU based approaches are unable to consistently estimate 3D shapes given different permutations of the same set of input images as the recurrent unit is permutation variant. It is also unlikely to refine the 3D shape given more images due to the long-term memory loss of GRU. Commonly used pooling approaches are limited to capturing partial information, e.g., max/mean values, ignoring other valuable features. In this paper, we present a new feed-forward neural module, named AttSets , together with a dedicated training algorithm, named FASet , to attentively aggregate an arbitrarily sized deep feature set for multi-view 3D reconstruction. The AttSets module is permutation invariant, computationally efficient and flexible to implement, while the FASet algorithm enables the AttSets based network to be remarkably robust and generalize to an arbitrary number of input images. We thoroughly evaluate FASet and the properties of AttSets on multiple large public datasets. Extensive experiments show that AttSets together with FASet algorithm significantly outperforms existing aggregation approaches. Bo Yang 0027, Sen Wang 0002, Andrew Markham, Agathoniki Trigoni |
Int. J. Comput. Vis. | 2 |
| 2019 | TextPlace: Visual Place Recognition and Topological Localization Through Reading Scene TextsabstractVisual place recognition is a fundamental problem for many vision based applications. Sparse feature and deep learning based methods have been successful and dominant over the decade. However, most of them do not explicitly leverage high-level semantic information to deal with challenging scenarios where they may fail. This paper proposes a novel visual place recognition algorithm, termed TextPlace, based on scene texts in the wild. Since scene texts are high-level information invariant to illumination changes and very distinct for different places when considering spatial correlation, it is beneficial for visual place recognition tasks under extreme appearance changes and perceptual aliasing. It also takes spatial-temporal dependence between scene texts into account for topological localization. Extensive experiments show that TextPlace achieves state-of-the-art performance, verifying the effectiveness of using high-level scene texts for robust visual place recognition in urban areas. Ziyang Hong 0001, Yvan R. Petillot, David Lane, Yishu Miao, Sen Wang 0002 |
ICCV | 5 |
| 2019 | Global Localization with Object-Level Semantics and TopologyabstractGlobal localization lies at the heart of autonomous navigation and Simultaneous Localization and Mapping (SLAM). The appearance-based approach has been successful, but still faces many open challenges in environments where visual conditions vary significantly over time. In this paper, we propose an integrated solution to leverage object-level dense semantics and spatial understanding of the environment for global localization. Our approach models an environment with 3D dense semantics, semantic graph and their topology. This object-level representation is then used for place recognition via semantic object association, followed by 6-DoF pose estimation by the semantic-level point alignment. Extensive experiments show that our approach can achieve robust global localization under extreme appearance changes. It is also capable of coping with other challenging scenarios, such as dynamic environments and incomplete query observations. Yvan R. Petillot, David Lane, Sen Wang 0002 |
ICRA | 4 |
| 2019 | Learning Monocular Visual Odometry through Geometry-Aware Curriculum LearningabstractInspired by the cognitive process of humans and animals, Curriculum Learning (CL) trains a model by gradually increasing the difficulty of the training data. In this paper, we study whether CL can be applied to complex geometry problems like estimating monocular Visual Odometry (VO). Unlike existing CL approaches, we present a novel CL strategy for learning the geometry of monocular VO by gradually making the learning objective more difficult during training. To this end, we propose a novel geometry-aware objective function by jointly optimizing relative and composite transformations over small windows via bounded pose regression loss. A cascade optical flow network followed by recurrent network with a differentiable windowed composition layer, termed CL-VO, is devised to learn the proposed objective. Evaluation on three real-world datasets shows superior performance of CL-VO over state-of-the-art feature-based and learning-based VO. Muhamad Risqi Utama Saputra, Pedro Porto Buarque de Gusmão, Sen Wang 0002, Andrew Markham, Agathoniki Trigoni |
ICRA | 3 |
| 2019 | Learning Object Bounding Boxes for 3D Instance Segmentation on Point CloudsabstractWe propose a novel, conceptually simple and general framework for instance segmentation on 3D point clouds. Our method, called 3D-BoNet, follows the simple design philosophy of per-point multilayer perceptrons (MLPs). The framework directly regresses 3D bounding boxes for all instances in a point cloud, while simultaneously predicting a point-level mask for each instance. It consists of a backbone network followed by two parallel network branches for 1) bounding box regression and 2) point mask prediction. 3D-BoNet is single-stage, anchor-free and end-to-end trainable. Moreover, it is remarkably computationally efficient as, unlike existing approaches, it does not require any post-processing steps such as non-maximum suppression, feature sampling, clustering or voting. Extensive experiments show that our approach surpasses existing work on both ScanNet and S3DIS datasets while being approximately 10x more computationally efficient. Comprehensive ablation studies demonstrate the effectiveness of our design. Bo Yang 0027, Ronald Clark, Qingyong Hu, Sen Wang 0002, Andrew Markham, Agathoniki Trigoni |
NeurIPS | 5 |
| 2019 | Efficient Indoor Positioning with Visual Experiences via Lifelong LearningabstractPositioning with visual sensors in indoor environments has many advantages: it doesn’t require infrastructure or accurate maps, and is more robust and accurate than other modalities such as WiFi. However, one of the biggest hurdles that prevents its practical application on mobile devices is the time-consuming visual processing pipeline. To overcome this problem, this paper proposes a novel lifelong learning approach to enable efficient and real-time visual positioning. We explore the fact that when following a previous visual experience for multiple times, one could gradually discover clues on how to traverse it with much less effort, e.g., which parts of the scene are more informative, and what kind of visual elements we should expect. Such second-order information is recorded as parameters, which provide key insights of the context and empower our system to dynamically optimise itself to stay localised with minimum cost. We implement the proposed approach on an array of mobile and wearable devices, and evaluate its performance in two indoor settings. Experimental results show our approach can reduce the visual processing time up to two orders of magnitude, while achieving sub-metre positioning accuracy. Hongkai Wen 0001, Ronald Clark, Sen Wang 0002, Xiaoxuan Lu 0001, Bowen Du 0002, Wen Hu 0001, Agathoniki Trigoni |
IEEE Trans. Mob. Comput. | 3 |
| 2018 | UnDeepVO: Monocular Visual Odometry Through Unsupervised Deep LearningabstractWe propose a novel monocular visual odometry (VO) system called UnDeepVO in this paper. UnDeepVO is able to estimate the 6-DoF pose of a monocular camera and the depth of its view by using deep neural networks. There are two salient features of the proposed UnDeepVo:one is the unsupervised deep learning scheme, and the other is the absolute scale recovery. Specifically, we train UnDeepVoby using stereo image pairs to recover the scale but test it by using consecutive monocular images. Thus, UnDeepVO is a monocular system. The loss function defined for training the networks is based on spatial and temporal dense information. A system overview is shown in Fig. 1. The experiments on KITTI dataset show our UnDeepVO achieves good performance in terms of pose accuracy. Ruihao Li 0001, Sen Wang 0002, Zhiqiang Long, Dongbing Gu |
ICRA | 2 |
| 2018 | DEFO-NET: Learning Body Deformation Using Generative Adversarial NetworksabstractModelling the physical properties of everyday objects is a fundamental prerequisite for autonomous robots. We present a novel generative adversarial network (DEFO-NET), able to predict body deformations under external forces from a single RGB-D image. The network is based on an invertible conditional Generative Adversarial Network (IcGAN) and is trained on a collection of different objects of interest generated by a physical finite element model simulator. Defo-netinherits the generalisation properties of GANs. This means that the network is able to reconstruct the whole 3-D appearance of the object given a single depth view of the object and to generalise to unseen object configurations. Contrary to traditional finite element methods, our approach is fast enough to be used in real-time applications. We apply the network to the problem of safe and fast navigation of mobile robots carrying payloads over different obstacles and floor materials. Experimental results in real scenarios show how a robot equipped with an RGB-D camera can use the network to predict terrain deformations under different payload configurations and use this to avoid unsafe areas. Zhihua Wang 0005, Stefano Rosa, Linhai Xie, Bo Yang 0027, Sen Wang 0002, Agathoniki Trigoni, Andrew Markham |
ICRA | 5 |
| 2018 | Learning with Training Wheels: Speeding up Training with a Simple Controller for Deep Reinforcement LearningabstractDeep Reinforcement Learning (DRL) has been applied successfully to many robotic applications. However, the large number of trials needed for training is a key issue. Most of existing techniques developed to improve training efficiency (e.g. imitation) target on general tasks rather than being tailored for robot applications, which have their specific context to benefit from. We propose a novel framework, Assisted Reinforcement Learning, where a classical controller (e.g. a PID controller) is used as an alternative, switchable policy to speed up training of DRL for local planning and navigation problems. The core idea is that the simple control law allows the robot to rapidly learn sensible primitives, like driving in a straight line, instead of random exploration. As the actor network becomes more advanced, it can then take over to perform more complex actions, like obstacle avoidance. Eventually, the simple controller can be discarded entirely. We show that not only does this technique train faster, it also is less sensitive to the structure of the DRL network and consistently outperforms a standard Deep Deterministic Policy Gradient network. We demonstrate the results in both simulation and real-world experiments. Linhai Xie, Sen Wang 0002, Stefano Rosa, Andrew Markham, Agathoniki Trigoni |
ICRA | 2 |
| 2018 | 3D-PhysNet: Learning the Intuitive Physics of Non-Rigid Object DeformationsabstractThe ability to interact and understand the environment is a fundamental prerequisite for a wide range of applications from robotics to augmented reality. In particular, predicting how deformable objects will react to applied forces in real time is a significant challenge. This is further confounded by the fact that shape information about encountered objects in the real world is often impaired by occlusions, noise and missing regions e.g. a robot manipulating an object will only be able to observe a partial view of the entire solid. In this work we present a framework, 3D-PhysNet, which is able to predict how a three-dimensional solid will deform under an applied force using intuitive physics modelling. In particular, we propose a new method to encode the physical properties of the material and the applied force, enabling generalisation over materials. The key is to combine deep variational autoencoders with adversarial training, conditioned on the applied force and the material properties.We further propose a cascaded architecture that takes a single 2.5D depth view of the object and predicts its deformation. Training data is provided by a physics simulator. The network is fast enough to be used in real-time applications from partial views. Experimental results show the viability and the generalisation properties of the proposed architecture. Zhihua Wang 0005, Stefano Rosa, Bo Yang 0027, Sen Wang 0002, Agathoniki Trigoni, Andrew Markham |
IJCAI | 4 |
| 2017 | VINet: Visual-Inertial Odometry as a Sequence-to-Sequence Learning ProblemabstractIn this paper we present an on-manifold sequence-to-sequence learning approach to motion estimation using visual and inertial sensors. It is to the best of our knowledge the first end-to-end trainable method for visual-inertial odometry which performs fusion of the data at an intermediate feature-representation level. Our method has numerous advantages over traditional approaches. Specifically, it eliminates the need for tedious manual synchronization of the camera and IMU as well as eliminating the need for manual calibration between the IMU and camera. A further advantage is that our model naturally and elegantly incorporates domain specific information which significantly mitigates drift. We show that our approach is competitive with state-of-the-art traditional methods when accurate calibration data is available and can be trained to outperform them in the presence of calibration and synchronization errors. Ronald Clark, Sen Wang 0002, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni |
AAAI | 2 |
| 2017 | Safety Verification of Deep Neural Networks
Xiaowei Huang 0001, Marta Z. Kwiatkowska, Sen Wang 0002, Min Wu 0011 |
CAV (1) | 3 |
| 2017 | VidLoc: A Deep Spatio-Temporal Model for 6-DoF Video-Clip RelocalizationabstractMachine learning techniques, namely convolutional neural networks (CNN) and regression forests, have recently shown great promise in performing 6-DoF localization of monocular images. However, in most cases image-sequences, rather only single images, are readily available. To this extent, none of the proposed learning-based approaches exploit the valuable constraint of temporal smoothness, often leading to situations where the per-frame error is larger than the camera motion. In this paper we propose a recurrent model for performing 6-DoF localization of video-clips. We find that, even by considering only short sequences (20 frames), the pose estimates are smoothed and the localization error can be drastically reduced. Finally, we consider means of obtaining probabilistic pose estimates from our model. We evaluate our method on openly-available real-world autonomous driving and indoor localization datasets. Ronald Clark, Sen Wang 0002, Andrew Markham, Agathoniki Trigoni, Hongkai Wen 0001 |
CVPR | 2 |
| 2017 | DeepVO: Towards end-to-end visual odometry with deep Recurrent Convolutional Neural NetworksabstractThis paper studies monocular visual odometry (VO) problem. Most of existing VO algorithms are developed under a standard pipeline including feature extraction, feature matching, motion estimation, local optimisation, etc. Although some of them have demonstrated superior performance, they usually need to be carefully designed and specifically fine-tuned to work well in different environments. Some prior knowledge is also required to recover an absolute scale for monocular VO. This paper presents a novel end-to-end framework for monocular VO by using deep Recurrent Convolutional Neural Networks (RCNNs). Since it is trained and deployed in an end-to-end manner, it infers poses directly from a sequence of raw RGB images (videos) without adopting any module in the conventional VO pipeline. Based on the RCNNs, it not only automatically learns effective feature representation for the VO problem through Convolutional Neural Networks, but also implicitly models sequential dynamics and relations using deep Recurrent Neural Networks. Extensive experiments on the KITTI VO dataset show competitive performance to state-of-the-art methods, verifying that the end-to-end Deep Learning technique can be a viable complement to the traditional VO systems. Sen Wang 0002, Ronald Clark, Hongkai Wen 0001, Agathoniki Trigoni |
ICRA | 1 |
| 2017 | SCAN: learning speaker identity from noisy sensor dataabstractSensor data acquired from multiple sensors simultaneously is featuring increasingly in our evermore pervasive world. Buildings can be made smarter and more efficient, spaces more responsive to users. A fundamental building block towards smart spaces is the ability to understand who is present in a certain area. A ubiquitous way of detecting this is to exploit the unique vocal features as people interact with one another. As an example, consider audio features sampled during a meeting, yielding a noisy set of possible voiceprints. With a number of meetings and knowledge of participation (e.g. through a calendar or MAC address), can we learn to associate a specific identity with a particular voiceprint? Obviously enrolling users into a biometric database is time-consuming and not robust to vocal deviations over time. To address this problem, the standard approach is to perform a clustering step (e.g. of audio data) followed by a data association step, when identity-rich sensor data is available. In this paper we show that this approach is not robust to noise in either type of sensor stream; to tackle this issue we propose a novel algorithm that jointly optimises the clustering and association process yielding up to three times higher identification precision than approaches that execute these steps sequentially. We demonstrate the performance benefits of our approach in two case studies, one with acoustic and MAC datasets that we collected from meetings in a non-residential building, and another from an online dataset from recorded radio interviews. Xiaoxuan Lu 0001, Hongkai Wen 0001, Sen Wang 0002, Andrew Markham, Agathoniki Trigoni |
IPSN | 3 |
| 2017 | GraphTinker: Outlier rejection and inlier injection for pose graph SLAMabstractIn pose graph Simultaneous Localization and Mapping (SLAM) systems, incorrect loop closures can seriously hinder optimizers from converging to correct solutions, significantly degrading both localization accuracy and map consistency. Therefore, it is crucial to enhance their robustness in the presence of numerous false-positive loop closures. Existing approaches tend to fail when working with very unreliable front-end systems, where the majority of inferred loop closures are incorrect. In this paper, we propose a novel middle layer, seamlessly embedded between front and back ends, to boost the robustness of the whole SLAM system. The main contributions of this paper are two-fold: 1) the proposed middle layer offers a new mechanism to reliably detect and remove false-positive loop closures, even if they form the overwhelming majority; 2) artificial loop closures are automatically reconstructed and injected into pose graphs in the framework of an Extended Rauch-Tung-Striebel smoother, reinforcing reliable loop closures. The proposed algorithm alters the graph generated by the front-end and can then be optimized by any back-end system. Extensive experiments are conducted to demonstrate significantly improved accuracy and robustness compared with state-of-the-art methods and various back-ends, verifying the effectiveness of the proposed algorithm. Linhai Xie, Sen Wang 0002, Andrew Markham, Agathoniki Trigoni |
IROS | 2 |
| 2016 | Poster Abstract: Efficient Visual Positioning with Adaptive Parameter LearningabstractPositioning with vision sensors is gaining its popularity, since it is more accurate, and requires much less bootstrapping and training effort. However, one of the major limitations of the existing solutions is the expensive visual processing pipeline: on resource-constrained mobile devices, it could take up to tens of seconds to process one frame. To address this, we propose a novel learning algorithm, which adaptively discovers the place dependent parameters for visual processing, such as which parts of the scene are more informative, and what kind of visual elements one would expect, as it is employed more and more by the users in a particular setting. With such meta- information, our positioning system dynamically adjust its behaviour, to localise the users with minimum effort. Preliminary results show that the proposed algorithm can reduce the cost on visual processing significantly, and achieve sub-metre positioning accuracy. Hongkai Wen 0001, Sen Wang 0002, Ronnie Clark, Savvas Papaioannou, Agathoniki Trigoni |
IPSN | 2 |
| 2016 | Keyframe based large-scale indoor localisation using geomagnetic field and motion patternabstractThis paper studies indoor localisation problem by using low-cost and pervasive sensors. Most of existing indoor localisation algorithms rely on camera, laser scanner, floor plan or other pre-installed infrastructure to achieve sub-meter or sub-centimetre localisation accuracy. However, in some circumstances these required devices or information may be unavailable or too expensive in terms of cost or deployment. This paper presents a novel keyframe based Pose Graph Simultaneous Localisation and Mapping (SLAM) method, which correlates ambient geomagnetic field with motion pattern and employs low-cost sensors commonly equipped in mobile devices, to provide positioning in both unknown and known environments. Extensive experiments are conducted in large-scale indoor environments to verify that the proposed method can achieve high localisation accuracy similar to state-of-the-arts, such as vision based Google Project Tango. Sen Wang 0002, Hongkai Wen 0001, Ronald Clark, Agathoniki Trigoni |
IROS | 1 |
| 2014 | Single beacon based multi-robot cooperative localization using Moving Horizon EstimationabstractThis paper studies three-dimensional multi-robot Cooperative Localization (CL) problem. Most of existing CL strategies adopt Extended Kalman Filter (EKF) or Maximum a Posteriori (MAP). In this paper, a novel approach based on Moving Horizon Estimation (MHE) is proposed. The main contribution of this paper is twofold: 1) MHE is integrated with EKF for three-dimensional CL using single mobile beacon, which can bound localization error, impose various constraints on states and noises, and make use of previous range measurements for current estimation. 2) A sufficient condition on observability of multi-robot CL is derived by using Fisher Information Matrix. Simulation is conducted to verify that the proposed MHE based CL algorithm outperforms EKF based method in terms of localization accuracy, and two scenarios where our algorithm is superior to EKF are discussed. Sen Wang 0002, Dongbing Gu, Huosheng Hu |
ICRA | 1 |
| 2013 | Single beacon based localization of AUVs using moving Horizon estimationabstractThis paper studies the underwater localization problem for a school of robotic fish, i.e., a kind of Autonomous Underwater Vehicles with limited size, power and payload. These robotic fish cannot be equipped with traditional underwater localization sensors that are big and heavy. The proposed localization system is performed by using a single surface mobile beacon which provides range measurement to bound the localization error. The main contribution of this paper lies in twofold: 1) Observability of single beacon based localization is first analyzed in the context of nonlinear discrete time system, deriving a sufficient condition on observability. 2) Moving Horizon Estimation is then integrated with Extended Kalman Filters for three-dimensional localization using single beacon, which can reduce the computational complexity, impose various constraints and make use of previous range measurements for current estimation. Extensive numerical simulations are conducted to verify the observability and high localization accuracy of the proposed underwater localization method. Sen Wang 0002, Huosheng Hu, Dongbing Gu |
IROS | 1 |