VLDB 2026 Research / reviewers in the wild / expert
Andrew Calway
dblp:c/AndrewCalway · also Andrew D. Calway
· DBLP profile ↗
64ranked-venue papers
6as first author
8since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 5 first-authorSystems, architecture and hardware · 20 · 7 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Object-based SLAM Using SuperquadricsabstractVisual SLAM uses visual information, typically point features, to localise a camera and, at the same time, map the environment. In recent years, there has been interest in using scene-understanding capabilities to enhance the mapping process and object-level SLAM systems have appeared in response. However, most of the previous work is limited to prestored object models or pre-trained networks to represent the objects, which limits working scenarios or uses representations with limited scope, such as cubes or quadrics. To address this, we propose to use superquadrics as the object representation and, in this paper, present a proof of principle SLAM system in which object-based mapping is fully integrated with camera tracking via keyframe optimisation. The system was tested on simulated and real datasets, and the results show that the system can achieve lightweight and comparatively good object representation whilst also giving good camera trajectories estimates under certain scenarios. Yifan Xing, Noe Samano, Wen Fan 0001, Andrew Calway |
IROS | 4 |
| 2024 | Geolocation on Cartographic Maps with Multi-Modal FusionabstractWe explore the geolocation problem, aiming to localize ground-view images on cartographic maps, without the need of any GPS priors. This task mimics the human wayfinding ability and offers high scalability and robustness by using the compact and semantic representations of maps. Current methods often rely on 2D maps to encode dense contextual information for ground-to-map matching. In this paper, we lift ground-to-map matching to a 2.5D space, where heights of structures (e.g. buildings) provide richer geometric information to guide the matching process. We propose a new approach to learning representative embeddings from multi-modal data. Specifically, we establish a projection relationship between 2D and 2.5D space. The projection is further used to combine multi-modal features from the 2D and 2.5D maps using an effective pixel-to-point fusion method. By encoding crucial geometric cues, our method learns discriminative location embeddings for matching panoramic images and maps. Additionally, we construct the first large-scale multi-modal geolocation dataset to validate our method and facilitate future research. Both single-image based and route based geolocation experiments are conducted to test our method. Extensive experiments demonstrate that the proposed method achieves significantly higher geolocation accuracy and faster convergence than previous 2D map-based approaches. Mengjie Zhou, Liu Liu 0009, Yiran Zhong, Andrew Calway |
IROS | 4 |
| 2022 | FD-SLAM: 3-D Reconstruction Using Features and Dense MatchingabstractIt is well known that visual SLAM systems based on dense matching are locally accurate but are also susceptible to long-term drift and map corruption. In contrast, feature matching methods can achieve greater long-term consistency but can suffer from inaccurate local pose estimation when feature information is sparse. Based on these observations, we propose an RGB-D SLAM system that leverages the advantages of both approaches: using dense frame-to-model odometry to build accurate sub-maps and on-the-fly feature-based matching across sub-maps for global map optimisation. In addition, we incorporate a learning-based loop closure component based on 3-D features which further stabilises map building. We have evaluated the approach on indoor sequences from public datasets, and the results show that it performs on par or better than state-of-the-art systems in terms of map reconstruction quality and pose estimation. The approach can also scale to large scenes where other systems often fail. Xingrui Yang 0001, Yuhang Ming 0001, Zhaopeng Cui, Andrew Calway |
ICRA | 4 |
| 2022 | CGiS-Net: Aggregating Colour, Geometry and Implicit Semantic Features for Indoor Place RecognitionabstractWe describe a novel approach to indoor place recognition from RGB point clouds based on aggregating low-level colour and geometry features with high-level implicit semantic features. It uses a 2-stage deep learning framework, in which the first stage is trained for the auxiliary task of semantic segmentation and the second stage uses features from layers in the first stage to generate discriminate descriptors for place recognition. The auxiliary task encourages the features to be semantically meaningful, hence aggregating the geometry and colour in the RGB point cloud data with implicit semantic information. We use an indoor place recognition dataset derived from the ScanNet dataset for training and evaluation, with a test set comprising 3,608 point clouds generated from 100 different rooms. Comparison with a traditional feature-based method and four state-of-the-art deep learning methods demonstrate that our approach significantly outperforms all five methods, achieving, for example, a top-3 average recall rate of 75% compared with 41% for the closest rival method. Our code is available at: https://github.com/YuhangMing/Semantic-Indoor-Place-Recognition Yuhang Ming 0001, Xingrui Yang 0001, Guofeng Zhang 0001, Andrew Calway |
IROS | 4 |
| 2021 | Global Aerial Localisation Using Image and Map EmbeddingsabstractWe present a purely vision based geolocation method for aircraft flying over urban and suburban environments. The method is based on matching aerial images with geolocated map tiles using a shared low dimensional embedded space of descriptors. The Euclidean distance between descriptors is used as a similarity measure between domains. The similarity between the observation and map locations is then integrated with visual odometry to track the aircraft’s position and yaw using a particle filter. Furthermore, we propose an efficient method to generate map descriptors in testing time based on interpolation, allowing compact representation of large areas giving the potential for high levels of scalability. We experimented in different cities with areas above 20 km2in size and preliminary results based on a database of aerial imagery demonstrate that the method gives good results. Noe Samano, Mengjie Zhou, Andrew Calway |
ICRA | 3 |
| 2021 | Object-Augmented RGB-D SLAM for Wide-Disparity RelocalisationabstractWe propose a novel object-augmented RGB-D SLAM system that is capable of constructing a consistent object map and performing relocalisation based on centroids of objects in the map. The approach aims to overcome the view dependence of appearance-based relocalisation methods using point features or images. During the map construction, we use a pre-trained neural network to detect objects and estimate 6D poses from RGB-D data. An incremental probabilistic model is used to aggregate estimates over time to create the object map. Then in relocalisation, we use the same network to extract objects-of-interest in the ‘lost’ frames. Pairwise geometric matching finds correspondences between map and frame objects, and probabilistic absolute orientation followed by application of iterative closest point to dense depth maps and object centroids gives relocalisation. Results of experiments in desktop environments demonstrate very high success rates even for frames with widely different viewpoints from those used to construct the map, significantly outperforming two appearance- based methods. Yuhang Ming 0001, Xingrui Yang 0001, Andrew Calway |
IROS | 3 |
| 2021 | Efficient Localisation Using Images and OpenStreetMapsabstractThe ability to localise is key for robot navigation. We describe an efficient method for vision-based localisation, which combines sequential Monte Carlo tracking with matching ground-level images to 2-D cartographic maps such as OpenStreetMaps. The matching is based on a learned embedded space representation linking images and map tiles, encoding the common semantic information present in both and providing potential for invariance to changing conditions. Moreover, the compactness of 2-D maps supports scalability. This contrasts with the majority of previous approaches based on matching with single-shot geo-referenced images or 3-D reconstructions. We present experiments using the StreetLearn and Oxford RobotCar datasets and demonstrate that the method is highly effective, giving high accuracy and fast convergence. Mengjie Zhou, Xieyuanli Chen, Noe Samano, Cyrill Stachniss, Andrew Calway |
IROS | 5 |
| 2021 | Evaluating Prototype Augmented and Adaptive guidance system to support Industrial Plant MaintenanceabstractWe evaluate AR for Plant maintenance by measuring how a prototype guidance system, tested under representative conditions, impacts performance. We are motivated to determine the cost-benefit of interactive guidance for hazardous, repetitive tasks and we observe an improvement of 21% efficiency, 50% accuracy and 19% reduced task load. AR has already been shown to deliver improvements in task performance, however, there is limited research exploring the integration of AR into complete task routines which presents a barrier to adoption. We apply mixed reality guidance via two within-group experiments. We measure efficiency and accuracy over a complete routine conducted under simulated conditions. Results compare AR versus Static and AR versus Adaptive. We conclude AR is best suited to demanding spatial translation and completion under pressure. We suggest AR offers potential in similar routines and propose further work to integrate in a live setting. Thomas Bale, Andrew Calway, Kirsten Cater, Chris Bevan, Robert Skilton, Thomas B. Scott 0001 |
MobileHCI | 2 |
| 2020 | You Are Here: Geolocation by Embedding Maps and Images
Noe Samano, Mengjie Zhou, Andrew Calway |
ECCV (23) | 3 |
| 2019 | Improving drone localisation around wind turbines using monocular model-based trackingabstractWe present a novel method of integrating image-based measurements into a drone navigation system for the automated inspection of wind turbines. We take a model-based tracking approach, where a 3D skeleton representation of the turbine is matched to the image data. Matching is based on comparing the projection of the representation to that inferred from images using a convolutional neural network. This enables us to find image correspondences using a generic turbine model that can be applied to a wide range of turbine shapes and sizes. To estimate 3D pose of the drone, we fuse the network output with GPS and IMU measurements using a pose graph optimiser. Results illustrate that the use of the image measurements significantly improves the accuracy of the localisation over that obtained using GPS and IMU alone. Oliver Moolan-Feroze, Konstantinos Karachalios, Dimitrios N. Nikolaidis, Andrew Calway |
ICRA | 4 |
| 2019 | Simultaneous Drone Localisation and Wind Turbine Model Fitting During Autonomous Surface InspectionabstractWe present a method for simultaneous localisation and wind turbine model fitting for a drone performing an automated surface inspection. We use a skeletal parameterisation of the turbine that can be easily integrated into a non-linear least squares optimiser, combined with a pose graph representation of the drone's 3-D trajectory, allowing us to optimise both sets of parameters simultaneously. Given images from an onboard camera, we use a CNN to infer projections of the skeletal model, enabling correspondence constraints to be established through a cost function. This is then coupled with GPS/IMU measurements taken at key frames in the graph to allow successive optimisation as the drone navigates around the turbine. We present two variants of the cost function, one based on traditional 2D point correspondences and the other on direct image interpolation within the inferred projections. Results from experiments on simulated and real-world data show that simultaneous optimisation provides improvements to localisation over only optimising the pose and that combined use of both cost functions proves most effective. Oliver Moolan-Feroze, Konstantinos Karachalios, Dimitrios N. Nikolaidis, Andrew Calway |
IROS | 4 |
| 2018 | Predicting Out-of-View Feature Points for Model-Based Camera Pose EstimationabstractIn this work we present a novel framework that uses deep learning to predict object feature points that are out-of-view in the input image. This system was developed with the application of model-based tracking in mind, particularly in the case of autonomous inspection robots, where only partial views of the object are available. Out-of-view prediction is enabled by applying scaling to the feature point labels during network training. This is combined with a recurrent neural network architecture designed to provide the final prediction layers with rich feature information from across the spatial extent of the input image. To show the versatility of these out-of-view predictions, we describe how to integrate them in both a particle filter tracker and an optimisation based tracker. To evaluate our work we compared our framework with one that predicts only points inside the image. We show that as the amount of the object in view decreases, being able to predict outside the image bounds adds robustness to the final pose estimation. Oliver Moolan-Feroze, Andrew Calway |
IROS | 2 |
| 2018 | Automated Map Reading: Image Based Localisation in 2-D Maps Using Binary Semantic DescriptorsabstractWe describe a novel approach to image based localisation in urban environments which uses semantic matching between images and a 2-D cartographic map. This contrasts with the majority of existing approaches which use image to image database matching. We use highly compact binary descriptors to represent locations, indicating the presence or not of semantic features, which significantly increases scalability and has the potential for greater invariance to variable imaging conditions. The approach is also more akin to human map reading, making it better suited to human-system interaction. In this initial study we use semantic features relating to buildings and road junctions in discrete viewing directions. CNN classifiers are used to detect the features in images and we match descriptor estimates with location tagged descriptors derived from the 2-D map to give localisation. The descriptors are not sufficiently discriminative on their own, but when concatenated sequentially along a route, their combination becomes highly distinctive and allows localisation even when using non-perfect classifiers. Performance is further improved by taking into account left or right turns over a route. Experimental results obtained using Google StreetView and OpenStreetMap data show that the approach has considerable potential, achieving localisation accuracy of around 85% using routes corresponding to approximately 200 meters. Pilailuck Panphattarasap, Andrew Calway |
IROS | 2 |
| 2016 | HDRFusion: HDR SLAM Using a Low-Cost Auto-Exposure RGB-D SensorabstractMost dense RGB/RGB-D SLAM systems require the brightness of 3-D points observed from different viewpoints to be constant. However, in reality, this assumption is difficult to meet even when the surface is Lambertian and illumination is static. One cause is that most cameras automatically tune exposure to adapt to the wide dynamic range of scene radiance, violating the brightness assumption. We describe a novel system - HDRFusion - which turns this apparent drawback into an advantage by fusing LDR frames into an HDR textured volume using a standard RGB-D sensor with auto-exposure (AE) enabled. The key contribution is the use of a normalised metric for frame alignment which is invariant to changes in exposure time. This enables robust tracking in frame-to-model mode and also compensates the exposure accurately so that HDR texture, free of artefacts, can be generated online. We demonstrate that the tracking robustness and accuracy is greatly improved by the approach and that radiance maps can be generated with far greater dynamic range of scene radiance. Shuda Li, Ankur Handa, Andrew Calway |
3DV | 4 |
| 2016 | Visual Place Recognition Using Landmark Distribution Descriptors
Pilailuck Panphattarasap, Andrew Calway |
ACCV (4) | 2 |
| 2016 | Absolute pose estimation using multiple forms of correspondences from RGB-D framesabstractWe describe a new approach to absolute pose estimation from noisy and outlier contaminated matching point sets for RGB-D sensors. We show that by integrating multiple forms of correspondence based on 2-D and 3-D points and surface normals gives more precise, accurate and robust pose estimates. This is because it gives more constraints than using one form alone and increases the available measurements, especially when dealing with sparse matching sets. We demonstrate the approach by incorporating it within a RANSAC algorithm and introduce a novel direct least-square approach to calculate pose estimates. Results from experiments on synthetic and real data demonstrate improved performance over existing methods. Shuda Li, Andrew Calway |
ICRA | 2 |
| 2015 | RGBD relocalisation using pairwise geometry and concise key point setsabstractWe describe a novel RGBD relocalisation algorithm based on key point matching. It combines two components. First, a graph matching algorithm which takes into account the pairwise 3-D geometry amongst the key points, giving robust relocalisation. Second, a point selection process which provides an even distribution of the ‘most matchable’ points across the scene based on non-maximum suppression within voxels of a volumetric grid. This ensures a bounded set of matchable key points which enables tractable and scalable graph matching at frame rate. We present evaluations using a public dataset and our own more difficult dataset containing large pose changes, fast motion and non-stationary objects. It is shown that the method significantly out performs state-of-the-art methods. Shuda Li, Andrew Calway |
ICRA | 2 |
| 2015 | Improving MAV control by predicting aerodynamic effects of obstaclesabstractBuilding on our previous work [1], in this paper we demonstrate how it is possible to improve flight control of a MAV that experiences aerodynamic disturbances caused by objects on its path. Predictions based on low resolution depth images taken at a distance are incorporated into the flight control loop on the throttle channel as this is adjusted to target undisrupted level flight. We demonstrate that a statistically significant improvement (p ≪ 0.001) is possible for some common obstacles such as boxes and steps, compared to using conventional feedback-only control. Our approach and results are encouraging toward more autonomous MAV exploration strategies. John Bartholomew, Andrew Calway, Walterio W. Mayol-Cuevas |
IROS | 2 |
| 2015 | Recognising Planes in a Single ImageabstractWe present a novel method to recognise planar structures in a single image and estimate their 3D orientation. This is done by exploiting the relationship between image appearance and 3D structure, using machine learning methods with supervised training data. As such, the method does not require specific features or use geometric cues, such as vanishing points. We employ general feature representations based on spatiograms of gradients and colour, coupled with relevance vector machines for classification and regression. We first show that using hand-labelled training data, we are able to classify pre-segmented regions as being planar or not, and estimate their 3D orientation. We then incorporate the method into a segmentation algorithm to detect multiple planar structures from a previously unseen image. Osian Haines, Andrew Calway |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | You-Do, I-Learn: Discovering Task Relevant Objects and their Modes of Interaction from Multi-User Egocentric Video
Dima Damen, Teesid Leelasawassuk, Osian Haines, Andrew Calway, Walterio W. Mayol-Cuevas |
BMVC | 4 |
| 2014 | Learning to predict obstacle aerodynamics from depth images for Micro Air VehiclesabstractMany applications of Micro Air Vehicles (MAVs) require them to operate in cluttered environments, flying in constrained spaces and close to obstacles. Such obstacles affect the airflow around the MAV and can thereby affect its flight characteristics. We describe a system for predicting these effects at a distance, using depth images obtained from an RGB-D sensor. Predictions are based on learning from prior experience gathered during training flights. We show that aerodynamic effects caused by obstacles are consistent, and demonstrate that it is practical to make predictions from experience without running a computationally expensive aerodynamic simulation. Our approach uses a Gaussian process regression, it requires minimal parameter tuning and is able to predict the acceleration that will be expected at a distance in the future. The method produces estimates within 12ms without any code optimisation and the results indicate good prediction ability with mean errors within 4–10cm/s2on a database of various obstacles. John Bartholomew, Andrew Calway, Walterio W. Mayol-Cuevas |
ICRA | 2 |
| 2013 | Place recognition from disparate viewsabstractVisual place recognition methods which use image matching techniques have shown success in recent years, however their reliance on local features restricts their use to images which are visually similar and which overlap in viewpoint. We suggest that a se-mantic approach to the problem would provide a more meaningful relationship between views of a place and so allow recognition when views are disparate and database cover-age is sparse. As initial work towards this goal we present a system which uses detected objects as the basic feature and demonstrate promising ability to recognise places from arbitrary viewpoints. We build a 2D place model of object positions and extract features which characterise a pair of models. We then use distributions learned from training ex-amples to compute the probability that the pair depict the same place and also an estimate of the relative pose of the cameras. Results on a dataset of 40 urban locations show good recognition performance and pose estimation, even for highly disparate views. 1 Rob Frampton, Andrew Calway |
BMVC | 2 |
| 2013 | Topological Map Building and Path Estimation Using Global-appearance Image DescriptorsabstractVisual-based navigation has been a source of numerous researches in the field of mobile robotics. In this paper we present a topological map building and localization algorithm using wide-angle scenes. Global-appearance descriptors are used in order to optimally represent the visual information. First, we build a topological graph that represents the navigation environment. Each node of the graph is a different position within the area, and it is composed of a collection of images that covers the complete field of view. We use the information provided by a camera that is mounted on the mobile robot when it travels along some routes between the nodes in the graph. With this aim, we estimate the relative position of each node using the visual information stored. Once the map is built, we propose a localization system that is able to estimate the location of the mobile not only in the nodes but also on intermediate positions using the visual information. The approach has been evaluated and shows good performance in real indoor scenarios under realistic illumination conditions. Francisco Amorós, Luis Payá, Óscar Reinoso, Walterio W. Mayol-Cuevas, Andrew Calway |
ICINCO (2) | 5 |
| 2013 | Context-based video codingabstractWe present a video CODEC framework which exploits extrinsic scene knowledge to condition a perspective motion model. An approximate textural-geometric model of the scene is prepared prior to coding. During coding, the locations of planar surfaces in the scene are tracked, facilitating the computation of accurate perspective motion warp parameters. These algorithms are integrated with H.264 into a hybrid CODEC framework, achieving savings of up to 48% for equivalent visual quality. Richard George Vigars, Andrew Calway, David Bull 0001 |
ICIP | 2 |
| 2013 | Visual mapping using learned structural priorsabstractWe investigate a new approach to vision based mapping, in which single image structure recognition is used to derive strong priors for initialisation of higher-level primitives in the map. This can reduce state size and speed up the building of more meaningful maps. We focus on plane mapping and use a recognition algorithm to detect and estimate the 3D orientation of planar structures in key frames, which are then used as priors for initialising planes in the map. The recognition algorithm learns the relationship between such structure and appearance from training examples offline. We demonstrate the approach in the context of an EKF based visual odometry system. Preliminary results of experiments in urban environments show that the system is able to build large maps with significant planar structure at average frames rates of around 60 fps whilst maintaining good trajectory estimation. The results suggest that the approach has considerable potential. Osian Haines, José Martínez-Carranza, Andrew Calway |
ICRA | 3 |
| 2013 | Enhancing 6D visual relocalisation with depth camerasabstractRelocalisation in 6D is relevant to a variety of Robotics applications and in particular to agile cameras exploring a 3D environment. While the use of geometry has commonly helped to validate appearance as a back-end process in several relocalisation systems before, we are interested in using 3D information to assist fast pose relocalisation computation as part of a front-end task. Our approach rapidly searches for a reduced number of visual descriptors, previously observed and stored in a database, that can be used to effectively compute the camera pose corresponding to the current view. We guide the search by means of constructing validated candidate sets using a 3D test involving the depth information obtained with an RGB-D camera (e.g. stereo of with structured light). Our experiments demonstrate that this process returns a compact quality set that works better for the pose estimation stage than when using a typical Nearest-Neighbor search over appearance only. The improvements are observed in terms of percentage of relocalised frames and speed, where the latter goes up to two orders of magnitude w.r.t. the conventional search. José Martínez-Carranza, Andrew Calway, Walterio W. Mayol-Cuevas |
IROS | 2 |
| 2012 | Real-time Learning and Detection of 3D Texture-less Objects: A Scalable ApproachabstractWe present a method for the learning and detection of multiple rigid texture-less 3D objects intended to operate at frame rate speeds for video input. The method is geared for fast and scalable learning and detection by combining tractable extraction of edgelet constellations with library lookup based on rotation- and scale-invariant descriptors. The approach learns object views in real-time, and is generative- enabling more objects to be learnt without the need for re-training. During testing, a random sample of edgelet constellations is tested for the presence of known objects. We perform testing of single and multi-object detection on a 30 objects dataset showing detections of any of them within milliseconds from the object’s visibility. The results show the scalability of the approach and its framerate performance. 1 Dima Damen, Pished Bunnun, Andrew Calway, Walterio W. Mayol-Cuevas |
BMVC | 3 |
| 2012 | Detecting planes and estimating their orientation from a single imageabstractWe propose an algorithm to detect planes in a single image of an outdoor urban scene, capable of identifying multiple distinct planes, and estimating their orientation. Using machine learning techniques, we learn the relationship between appearance and structure from a large set of labelled examples. Plane detection is achieved by classifying multiple overlapping image regions, in order to obtain an initial estimate of planarity for a set of points, which are segmented into planar and non-planar regions using a sequence of Markov random fields. This differs from previous methods in that it does not rely on line detection, and is able to predict an actual orientation for planes. We show that the method is able to reliably extract planes in a variety of scenes, and compares favourably with existing methods. Osian Haines, Andrew Calway |
BMVC | 2 |
| 2012 | Estimating Planar Structure in Single Images by Learning from Examples
Osian Haines, Andrew Calway |
ICPRAM (2) | 2 |
| 2012 | Efficient visual odometry using a structure-driven temporal mapabstractWe describe a method for visual odometry using a single camera based on an EKF framework. Previous work has shown that filtering based approaches can achieve accuracy performance comparable to that of optimisation methods providing that large numbers of features are used. However, computational requirements are significantly increased and frame rates are low. We address this by employing higher level structure - in the form of planes - to efficiently parameterise features and so reduce the filter state size and computational load. Moreover, we extend a 1-point RANSAC outlier rejection method to the case of features lying on planes. Results of experiments with both simulated and real-world data demonstrate that the method is effective, achieving comparable accuracy whilst running at significantly higher frame rates. José Martínez-Carranza, Andrew Calway |
ICRA | 2 |
| 2012 | Predicting Micro Air Vehicle landing behaviour from visual textureabstractWe introduce a framework to predict the landing behaviour of a Micro Air Vehicle (MAV) from the appearance of the landing surface. We approach this problem by learning a mapping from visual texture observed from an onboard camera to the landing behaviour on a set of sample materials. In this case we exemplify our framework by predicting the yaw angle of the MAV after landing. Our framework demonstrates the applicability of established texture classification methods usually tested on stationary camera setups for the more challenging case of textures observed from a MAV. Results for supervised training demonstrate good estimation of the landing behaviour and motivate future work to implement autonomous decision making strategies and other behaviour predictions based on imagery. John Bartholomew, Andrew Calway, Walterio W. Mayol-Cuevas |
IROS | 2 |
| 2012 | Egocentric Real-time Workspace Monitoring using an RGB-D cameraabstractWe describe an integrated system for personal workspace monitoring based around an RGB-D sensor. The approach is egocentric, facilitating full flexibility, and operates in real-time, providing object detection and recognition, and 3D trajectory estimation whilst the user undertakes tasks in the workspace. A prototype on-body system developed in the context of work-flow analysis for industrial manipulation and assembly tasks is described. The system is evaluated on two tasks with multiple users, and results indicate that the method is effective, giving good accuracy performance. Dima Damen, Andrew P. Gee, Walterio W. Mayol-Cuevas, Andrew Calway |
IROS | 4 |
| 2012 | Integrating 3D object detection, modelling and tracking on a mobile phoneabstractThis paper presents a complete system on a camera phone that integrates a texture less 3D object detector together with in-situ modelling and tracking. The result is a suite intended for AR applications on the move where objects can be captured, tracked and with the automated detection providing a bridge to either initialize tracking or resume it after measurement loss. The object detector training is online and in-situ and benefits from tight integration with the modelling and tracking processes. Pished Bunnun, Dima Damen, Andrew Calway, Walterio W. Mayol-Cuevas |
ISMAR | 3 |
| 2011 | A topometric system for wide area augmented reality
Andrew P. Gee, Matthew Webb, Ponciano Jorge Escamilla-Ambrosio, Walterio W. Mayol-Cuevas, Andrew Calway |
Comput. Graph. | 5 |
| 2010 | Unifying Planar and Point Mapping in Monocular SLAMabstractPlanar features in filter-based Visual SLAM systems require an initialisation stage that delays their use within the estimation. In this stage, surface and pose are initialised either by using an already generated map of point features [2, 3] or by using visual clues from frames [4]. This delay is unsatisfactory specially in scenarios where the camera moves rapidly such that visual features are observed for a very limited period. In this paper we present a unified approach to mapping in which points and planes are initialised alongside each other within the same framework. The best structure emerges according to what the camera observes, thus avoiding delayed initialisation for planar features. To do this we use a similar parameterisation to the one used for planar features in [3, 4]. The Inverse Depth Planar Parameterisation (IDPP), as we call it, is used to represent both planes and points. This IDPP is also combined with a point based measurement model where the planar constraint is introduced. The latter allows us to estimate and grow a planar structure if suitable, or to estimate a 3-D point if visual measurements do not support the constraint. The IDPP contains three main components: (1) A reference camera (RC); (2) the depth w.r.t. the RC of a seed 3-D point on the plane; (3) the normal of the plane. José Martínez-Carranza, Andrew Calway |
BMVC | 2 |
| 2010 | Application of multiple-wireless to a visual localisation system for emergency servicesabstractIn this paper we discuss the application of multiple-wireless technology to a practical context-enhanced service system called ViewNet. ViewNet develops technologies to support enhanced coordination and cooperation between operation teams in the emergency services and the police. Distributed localisation of users and mapping of environments implemented over a secure wireless network enables teams of operatives to search and map an incident area rapidly and in full coordination with each other and with a control centre. Sensing is based on fusing absolute positioning systems (UWB and GPS) with relative localisation and mapping from on-body or hand-held vision and inertial sensors. This paper focuses on the case for multiple-wireless capabilities in such a system and the benefits it can provide. We describe our work of developing a software API to support both WLAN and TETRA in ViewNet. It also provides a basis for incorporating future wireless technologies into ViewNet. Costas Efthymiou, Sedat Görmüs, Zhong Fan, Andrew Calway, Walterio W. Mayol-Cuevas, Angela Doufexi |
PIMRC | 4 |
| 2009 | Efficiently Increasing Map Density in Visual SLAM Using Planar Features with Adaptive MeasurementabstractThe visual simultaneous localisation and mapping (SLAM) systems now in widespread use are based on localised point features [2, 4, 5]. Although effective in many respects, the approach has limitations when considering the density and efficiency of map representation. With a dense population of features, camera tracking can be robust, able to withstand significant occlusion and large changes in camera viewpoint. But this comes at a high computational cost, typically increasing quadratically with the number of features. In this work we propose increasing map density by building in higherorder structure in the form of planar features. An important and novel aspect of the work is the manner in which the planar features are updated and used to localise the camera. We base our approach on an extended Kalman filter (EKF) monocular SLAM system developed by Chekhlov et al. [3]. This provides real-time estimates of the 3-D pose of a calibrated camera whilst simultaneously mapping the scene in terms of point based features. In order to incorporate planar structure into the real-time monocular SLAM we carry out three steps: detection of planar structure in the scene; insertion of planar features into the map; and adaptive measurement of the features. To apply the principle of adaptive measurement it is essential that planar features inserted into the map correspond to actual planar structure in the scene. For this we employ the method proposed by Martinez-Carranza and Calway [6], which uses an appearance model to detect planes defined by subsets of mapped point features (at least three points). Having detected planar features in the scene these are inserted into the map using a suitable representation within the filter state. This has two components: plane parameterisation and the reference camera. The plane is defined by yp = (θ ,φ ,ρ), where (θ ,φ) defines the unit normal of the plane in polar coordinates in the reference camera and ρ is the inverse depth of the plane centre along the ray defined by uo, with the latter being stored at initialisation of the plane, see figure 1a. Insertion of the reference camera is done by augmenting the state with a copy of the current pose, i.e. vp = v, and with initialised plane parameters yp derived from the pose and the subset of mapped points which define the plane. The reference camera serves two purposes: it references the plane in the SLAM coordinate system (with the associated uncertainties) and enables subsequent measurement of the planar feature using region based matching with respect to the current frame (key frame). To facilitate the latter the key frame image is also stored in the system. As illustrated in figure 1b, measurements for a planar feature are therefore assumed to take the following form: José Martínez-Carranza, Andrew Calway |
BMVC | 2 |
| 2008 | Appearance Based Indexing for Relocalisation in Real-Time Visual SLAMabstractPrevious work on visual SLAM has shown that indexing on space and scale facilitates the use of feature descriptors for matching in real-time systems and that this can significantly increase robustness. However, the performance gains necessarily diminish as uncertainty about camera position increases. In this paper we address this issue by introducing a further level of indexing based on appearance, using low order Haar wavelet coefficients. This enables fast look up of descriptors even when the camera is lost, hence allowing efficient relocalisation. Results of experiments on a range of real world test cases demonstrate that the method is effective, including single frame relocalisation rates up to 90\% using relatively low numbers of descriptor comparisons. Denis Chekhlov, Walterio W. Mayol-Cuevas, Andrew Calway |
BMVC | 3 |
| 2008 | Discovering Higher Level Structure in Visual SLAMabstractIn this paper, we describe a novel method for discovering and incorporating higher level map structure in a real-time visual simultaneous localization and mapping (SLAM) system. Previous approaches use sparse maps populated by isolated features such as 3-D points or edgelets. Although this facilitates efficient localization, it yields very limited scene representation and ignores the inherent redundancy among features resulting from physical structure in the scene. In this paper, higher level structure, in the form of lines and surfaces, is discovered concurrently with SLAM operation, and then, incorporated into the map in a rigorous manner, attempting to maintain important cross-covariance information and allow consistent update of the feature parameters. This is achieved by using a bottom-up process, in which subsets of low-level features are ldquofolded inrdquo to a parameterization of an associated higher level feature, thus collapsing the state space as well as building structure into the map. We demonstrate and analyze the effects of the approach for the cases of line and plane discovery, both in simulation and within a real-time system operating with a handheld camera in an office environment. Andrew P. Gee, Denis Chekhlov, Andrew Calway, Walterio W. Mayol-Cuevas |
IEEE Trans. Robotics | 3 |
| 2007 | Discovering Planes and Collapsing the State Space in Visual SLAMabstractRecent advances in real-time visual SLAM have been based primarily on mapping isolated 3-D points. This presents difficulties when seeking to extend operation to wide areas, as the system state becomes large, requiring increasing computational effort. In this paper we present a novel approach to this problem in which planar structural components are embedded within the state to represent mapped points lying on a common plane. This collapses the state size, reducing computation and improving scalability, as well as giving a higher level scene description. Critically, the plane parameters are augmented into the SLAM state in a proper fashion, maintaining inherent uncertainties via a full covariance representation. Results for simulated data and for real-time operation demonstrate that the approach is effective. 1 Andrew P. Gee, Denis Chekhlov, Walterio W. Mayol-Cuevas, Andrew Calway |
BMVC | 4 |
| 2007 | Robust Real-Time Visual SLAM Using Scale Prediction and Exemplar Based Feature DescriptionabstractTwo major limitations of real-time visual SLAM algorithms are the restricted range of views over which they can operate and their lack of robustness when faced with erratic camera motion or severe visual occlusion. In this paper we describe a visual SLAM algorithm which addresses both of these problems. The key component is a novel feature description method which is both fast and capable of repeat-able correspondence matching over a wide range of viewing angles and scales. This is achieved in real-time by using a SIFT-like spatial gradient descriptor in conjunction with efficient scale prediction and exemplar based feature representation. Results are presented illustrating robust realtime SLAM operation within an office environment. Denis Chekhlov, Mark Pupilli, Walterio W. Mayol-Cuevas, Andrew Calway |
CVPR | 4 |
| 2007 | Ninja on a Plane: Automatic Discovery of Physical Planes for Augmented Reality Using Visual SLAMabstractMost work in visual augmented reality (AR) employs predefined markers or models that simplify the algorithms needed for sensor positioning and augmentation but at the cost of imposing restrictions on the areas of operation and on interactivity. This paper presents a simple game in which an AR agent has to navigate using real planar surfaces on objects that are dynamically added to an unprepared environment. An extended Kalman filter (EKF) simultaneous localisation and mapping (SLAM) framework with automatic plane discovery is used to enable the player to interactively build a structured map of the game environment using a single, agile camera. By using SLAM, we are able to achieve real-time interactivity and maintain rigorous estimates of the system's uncertainty, which enables the effects of high quality estimates to be propagated to other features (points and planes) even if they are outside the camera's current field of view. Denis Chekhlov, Andrew P. Gee, Andrew Calway, Walterio W. Mayol-Cuevas |
ISMAR | 3 |
| 2006 | Real-Time Visual SLAM with Resilience to Erratic MotionabstractSimultaneous localisation and mapping using a single camera becomes difficult when erratic motions violate predictive motion models. This problem needs to be addressed when visual SLAM algorithms are transferred from robots or mobile vehicles onto hand-held or wearable devices. In this paper we describe a novel SLAM extension to a camera localisation algorithm based on particle filtering which provides resilience to erratic motion. The mapping component is based on auxiliary unscented Kalman filters coupled to the main particle filter via measurement covariances. This coupling allows the system to survive unpredictable motions such as camera shake, and enables a return to full SLAM operation once normal motion resumes. We present results demonstrating the effectiveness of the approach when operating within a desktop environment. Mark Pupilli, Andrew Calway |
CVPR (1) | 2 |
| 2005 | Recognizing Animals Using Motion PartsabstractWe describe a method for automatically recognizing animals in image sequences based on their distinctive locomotive movement patterns. The 2-D motion field associated with the animal is represented using a `configuration of motion parts' model, the characteristics of which are learned from training data. We adopt an unsupervised approach to learning model parameters, based on minimal a priori knowledge of the physical or locomotive characteristics of the animals concerned. Results are presented demonstrating excellent classification performance, with accuracy exceeding 98% on a test set consisting of over 100 sequences of 7 different species. Changming Kong, Andrew Calway, Majid Mirmehdi |
BMVC | 2 |
| 2005 | Real-Time Camera Tracking Using a Particle FilterabstractWe describe a particle filtering method for vision based tracking of a hand held calibrated camera in real-time. The ability of the particle filter to deal with non-linearities and non-Gaussian statistics suggests the potential to provide improved robustness over existing approaches, such as those based on the Kalman filter. In our approach, the particle filter provides recursive approximations to the posterior density for the 3-D motion parameters. The measurements are inlier/outlier counts of likely correspondence matches for a set of salient points in the scene. The algorithm is simple to implement and we present results illustrating good tracking performance using a `live' camera. We also demonstrate the potential robustness of the method, including the ability to recover from loss of track and to deal with severe occlusion. Mark Pupilli, Andrew Calway |
BMVC | 2 |
| 2005 | Recursive Estimation of 3D Motion and Surface Structure from Local Affine Flow ParametersabstractA recursive structure from motion algorithm based on optical flow measurements taken from an image sequence is described. It provides estimates of surface normals in addition to 3D motion and depth. The measurements are affine motion parameters which approximate the local flow fields associated with near-planar surface patches in the scene. These are integrated over time to give estimates of the 3D parameters using an extended Kalman filter. This also estimates the camera focal length and, so, the 3D estimates are metric. The use of parametric measurements means that the algorithm is computationally less demanding than previous optical flow approaches and the recursive filter builds in a degree of noise robustness. Results of experiments on synthetic and real image sequences demonstrate that the algorithm performs well. Andrew Calway |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2004 | Interpolating Novel Views from Image Sequences by Probabilistic Depth Carving
Annie Yao, Andrew Calway |
ECCV (2) | 2 |
| 2003 | Dense 3-D Structure from Image Sequences Using Probabilistic Depth CarvingabstractWe describe an algorithm to determine dense 3-D structure in a static scene from an image sequence captured by a moving camera. Metric camera motions are first determined using a recursive structure from motion algorithm based on tracked feature points. Dense depth information for a subset of key frames is then obtained using a novel probabilistic depth carving algorithm - analogous to space carving - in which depth probabilities obtained locally about the key frames are combined in 3-D space. An important component in this process is that opacity and occlusion relationships are modelled explicitly, enabling consistent combination of the depth probabilities. Results of experiments on a real sequence illustrate the effectiveness of the approach. Annie Yao, Andrew Calway |
BMVC | 2 |
| 2002 | Tracking Many Objects Using Subordinated CondensationabstractWe describe a novel extension to the CONDENSATION algorithm for track-ing multiple objects of the same type. Previous extensions for multiple object tracking do not scale effectively to large numbers of objects. The new ap-proach – subordinated CONDENSATION – deals effectively with arbitrary numbers of objects in an efficient manner, providing a robust means of track-ing individual objects across heavily populated and cluttered scenes. The key innovation is the introduction of bindings (subordination) amongst particles which enables multiple occlusions to be handled in a natural way within the standard CONDENSATION framework. The effectiveness of the approach is demonstrated by tracking multiple animals of the same species in cluttered wildlife footage. 1 David Tweed, Andrew Calway |
BMVC | 2 |
| 2002 | Integrated segmentation and depth ordering of motion layers in image sequences
David Tweed, Andrew Calway |
Image Vis. Comput. | 2 |
| 2001 | Moving Object Graphs and Layer Extraction from Image SequencesabstractWe describe a new approach to extracting layered representations from image sequences based on moving object graphs (MOGs). A MOG is a form of region adjacency graph which links together local motion segmentations corresponding to distinct moving regions in the scene, typically either foreground objects or the background. The local motion segmentations are obtained by fusing colour segmentations with block motion estimates and the MOGs link segmentations with consistent spatial and motion properties. Linking MOGs across frames then allows temporal consistency to be imposed and layers to be extracted. The approach provides a flexible framework within which to combine local and global constraints both spatially and temporally, enabling robust motion segmentation and layer extraction. Results of experiments on real sequences illustrate that the approach is effective. 1 David Tweed, Andrew Calway |
BMVC | 2 |
| 2001 | Multiresolution Gaussian mixture models for visual motion estimationabstractThis paper introduces a new generalisation of scale-space and pyramids, which combines statistical modelling with a spatial representation. The representation uses the familiar concept of multiple resolutions, but applied to a Gaussian mixture representation of the image - hence the title MGMM. It is shown that MGMM can approximate any probability density and can adapt to smooth motions. After a presentation of the theory, it is shown how MGMM can be applied to the estimation of visual motion. Roland Wilson, Andrew Calway |
ICIP (2) | 2 |
| 2001 | Using affine correspondence to estimate 3-D facial poseabstractWe describe a method for tracking a person's face through an image sequence and estimating the 3-D facial pose within each frame. It is based on an affine approximation to the motion of projected facial features such as eyes, mouth and nose. Tracking stability is maintained by enforcing the affine relationship amongst the motion of the features using linear regression and a Kalman filter. Facial pose is estimated using an ellipse-circle correspondence technique based on the affine transformation between the features in the current view and those in a fronto-parallel view. The method has the advantage of being simple to implement and not relying on assumed facial characteristics. Pingping Yao, Glyn Evans, Andrew Calway |
ICIP (3) | 3 |
| 2000 | Estimating the Structure of Textured Surfaces Using Local Affine FlowabstractThis paper describes a novel approach for recovering the structure and motion of a rigid textured surface from an image sequence. Camera focal length is also recovered, yielding metric estimates of the structure without the need for pre-calibration. The key innovation is the use of local affine flow parameters as the measurements within an extended Kalman filter (EKF) estimation framework, in contrast to feature correspondences or optical flow used in previous approaches. This enables surface normals to be recovered in addition to depth, unlike a feature correspondence scheme, but without the computational limitation of an optical flow approach. The method is based on equating the affine parameters to a local linearisation of the 2-D motion field and using the EKF to provide recursive estimates of the 3-D structure and motion. Experiments on both synthetic and real sequences demonstrate that the approach has considerable potential. 1 Andrew Calway |
BMVC | 1 |
| 2000 | Integrated Segmentation and Depth Ordering
David Tweed, Andrew Calway |
BMVC | 2 |
| 1998 | Image Registration using Multiresolution Frequency Domain CorrelationabstractThis paper describes a correlation based image registration method which is able to register images related by a single global affine transformation or by a transformation field which is approximately piecewise affine. The method has two key elements: an affine estimator, which derives estimates of the six affine parameters relating two image regions by aligning their Fourier spectra prior to correlating; and a multiresolution search process, which determines the global transformation field in terms of a set of local affine estimates at appropriate spatial resolutions. The method is computationally efficient and performs well for a range of different images and transformations. 1 Introduction Image registration is an important area of Computer Vision and Image Processing. It involves determining the transformation which will map pixels in one image to their corresponding or matching pixels in one or more related images, where the latter are different views and/or different ti... Stefan Krüger, Andrew Calway |
BMVC | 2 |
| 1998 | Motion Estimation using Adaptive Correlation and Local Directional Smoothing
Andrew Calway, Stefan Krüger, David Tweed |
ICIP (3) | 1 |
| 1998 | Image Sequence Analysis and Segmentation using G-blobsabstractThis paper introduces a new generalisation of the familiar scale-space and wavelet representations, designed specifically to deal with the complexities of representing motions induced in an image sequence by the movement of 3-D objects of which the scene is comprised. It does this by combining an affine representation of local image motions with a local structure model based on Gaussian functions. After a brief review of its symmetry and completeness properties, the representation is applied to the analysis of image motions in a typical video sequence. Roland Wilson, Peter R. Meulemans, Andrew Calway, Stefan Krüger |
ICIP (2) | 3 |
| 1996 | Image representation based on the affine symmetry groupabstractThe representation of 2-D signals which are symmetric under the affine group of transformations is considered. An extension of the multiresolution Fourier transform (MFT) is presented and shown to have a predictable redistribution of energy as a result of affine transformations of the input signal. An approximation based on the discrete MFT then forms the basis of a computationally efficient algorithm for estimating local affine coordinate transformations between pairs of image regions. Results of applying the technique to an image registration problem illustrate its potential. Andrew Calway |
ICIP (3) | 1 |
| 1996 | A multiresolution frequency domain method for estimating affine motion parametersabstractA novel approach to motion estimation is presented. The scheme employs a multiresolution model of 2-D motion, in which the motion within local regions at different scales is described in terms of a six parameter affine model. A frequency domain method is used to estimate the local affine motion parameters and a coarse-fine tracking algorithm is used to determine the global motion field. Results of experiments performed on synthetic and natural image sequences illustrate the effectiveness of the approach. Stefan Krüger, Andrew Calway |
ICIP (1) | 2 |
| 1992 | Multiresolution Estimation of 2-D Disparity Using a Frequency Domain Approach
Andrew Calway, Hans Knutsson, Roland Wilson |
BMVC | 1 |
| 1992 | A generalized wavelet transform for Fourier analysis: The multiresolution Fourier transform and its application to image and audio signal analysisabstractA wavelet transform specifically designed for Fourier analysis at multiple scales is described and shown to be capable of providing a local representation which is particularly well suited to segmentation problems. It is shown that, by an appropriate choice of analysis window and sampling intervals, it is possible to obtain a Fourier representation which can be computed efficiently and overcomes the limitations of using a fixed scale of window, yet by virtue of its symmetry properties allows simple estimation of such fundamental signal parameters as instantaneous frequency and onset time/position. The transform is applied to the segmentation of both image and audio signals, demonstrating its power to deal with signal events which are localized in either time/space or frequency. Feature extraction and segmentation are performed through the introduction of a class of multiresolution Markov models, whose parameters represent the signal events underlying the segmentation.> Roland Wilson, Andrew Calway, Edward R. S. Pearson |
IEEE Trans. Inf. Theory | 2 |
| 1990 | Curve extraction in images using the multiresolution Fourier transformabstractA curve extraction algorithm which is defined within a multiresolution framework is described. An image model based on local features such as lines and edges is assumed, and the parameters of this model are estimated using the multiresolution Fourier transform. The estimated features, which exist at different spatial resolutions, are then combined into curves using an appropriate curvature measure. The scheme is generally applicable and can be implemented in an efficient manner within a hierarchical data structure. Results presented demonstrate that it is capable of identifying curves in natural images.> Andrew Calway, Roland Wilson |
ICASSP | 1 |
| 1990 | Generalized quad-trees: a unified approach to multiresolution image analysis and codingabstractThis paper is an attempt to bring together a number of current ideas in multiresolution image processing into a single framework. The unification is achieved by introducing a two-level image model, comprising a quadtree containing parameters which control the evolution of the image as a sequence of successive refinements through scale-space. Examples of the method's application to image coding and analysis are used to illustrate the principles and show its usefulness. Roland Wilson, Martin Todd, Andrew Calway |
VCIP | 3 |